Video encoding and decoding
By limiting the location of the APS NAL unit in the picture unit of VVC8, the application problem of multiple versions of APS in the same picture unit is solved, and the storage and decoding process of the decoder is simplified.
Patent Information
- Application Number
- CN202180024884.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-03
- Filing Date
- 2021-03-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-03-22
AI Technical Summary
In VVC8, two versions of APS may be applied to different stripes of the same picture unit, resulting in the decoder requiring two versions of APS, increasing the complexity of storage and decoding.
Reduce the number of APS versions that need to be stored by prohibiting the prefixed APS NAL unit before the first VCL NAL of the picture unit and after the last VCL NAL.
Reduces the number of APS versions that need to be maintained in storage, simplifies the storage and decoding process of the decoder, and reduces the possibility of error states.
Smart Images

Figure CN115362684B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video encoding and decoding, and in particular to video encoding and decoding using Adaptive Parameter Sets (APS). Background Art
[0002] Recently, the Joint Video Experts Team (JVET) (a collaborative team consisting of MPEG and ITU-T Study Group 16 VCEG) began work on a new video coding standard called Versatile Video Coding (VVC). The goal of VVC is to provide significant improvements in compression performance over the existing HEVC standard (i.e., typically twice as good as before) and to be completed in 2020. Key target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) video. In total, JVET evaluated feedback from 32 organizations using formal subjective tests conducted by independent test labs. Some recommendations showed that compression efficiency was typically improved by 40% or more when compared to using HEVC. Particular results were shown on ultra-high definition (UHD) video test material. Therefore, for the final standard, we can expect compression efficiency improvements far exceeding the 50% target.
[0003] VVC provides adaptive parameter sets or APSs to transmit parameters that can be shared by one or more slices of a coded video sequence. VVC Draft 8 defines APS as a syntactic structure containing syntax elements, where the syntax elements apply to zero or more slices as determined by zero or more syntax elements found in a slice or picture header. More than one APS can be applied to slices belonging to the same coded picture. A picture unit corresponds to exactly one coded picture. A picture unit is in turn a set of network abstraction layer (NAL) units. In VVC Draft 8, any APS present in a picture unit is constrained to share the same content when it has the same APS type and the same APS identifier. In addition, when encoding picture units using several slices that can reference the APS sent before the slice NAL unit, some configurations may require additional decoding operations to maintain the APS in use in memory and / or bitstream rewrite operations when random access decoding is performed at a specific timing in a coded video sequence.
[0004] It is desirable to improve the encoding of APS and its references. Summary of the invention
[0005] According to a first aspect of the present invention, a method for encoding an image sequence in a bitstream is provided, comprising: providing a series of picture units in the bitstream, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of the VCL NAL units contains coded image data, each of the adaptation parameter set NAL units contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be contained in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be contained in the suffix APS NAL unit. NAL units; and prohibiting the inclusion of a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit.
[0006] This can solve the problem that in VVC8 two versions of the APS may apply to different slices of the same picture unit. Then, in order to decode the bitstream, the decoder must store two versions of the APS in the memory for a given value pair of the APS identifier and the APS type. In the worst case example, the decoder may have to double the memory size (to maintain two versions of each APS) to store the APS required to decode the picture unit. In addition, the decoder must maintain the order of the VCL NAL units relative to the APS NAL units to determine which VCL NAL units reference the first or second version of the APS NAL unit. By prohibiting the inclusion of a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit, some cases where two versions of the APS are needed can be eliminated.
[0007] According to a second aspect of the invention, there is provided a method for encoding an image sequence, which is the same as the method of the first aspect, except that instead of prohibiting the inclusion of a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit as in the first aspect, the second aspect involves prohibiting the inclusion of a suffix APS NAL unit in the picture unit before the last NAL unit of the picture unit.
[0008] This method is complementary to the method of the first aspect and addresses the same problem.By prohibiting the inclusion of suffix APS NAL units in a picture unit before the last NAL unit of the picture unit, some situations where two versions of APS are needed can be eliminated.
[0009] It is also possible to prohibit including a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit and prohibit including a prefix APS NAL unit in the picture unit after the first NAL unit of the picture unit.The elimination of the case where two versions of APS are needed is then further enhanced.
[0010] According to a third aspect of the present invention, a method for encoding an image sequence in a bitstream is provided, comprising: providing a series of picture units in the bitstream, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of the VCL NAL units contains coded image data, each of the adaptation parameter set NAL units contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be contained in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be contained in the suffix APS NAL unit. In the NAL unit, the APS NAL units each have an APS type and an APS identifier; and it is permitted to include a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents in the same picture unit.
[0011] Permitting the inclusion of prefix APS NAL units and suffix APS NAL units with the same APS type and the same APS identifier but different content in the same picture unit amounts to a new degree of freedom for the constraints in VVC8. In VVC8, even if APS NAL units have different (prefix and suffix) types, they cannot be present in one picture unit.
[0012] When an application performs random access to a bitstream to start decoding at a certain picture unit (random access point), the application may have to provide some APS NAL units before the VCL NAL units of the picture unit. For example, the application can insert the necessary APS NAL units at the beginning of the picture unit. However, the resulting bitstream may break some constraints of VVC8. This in turn may cause the decoding to enter an error state.
[0013] One constraint is that a suffix APS NAL unit may be inserted before the first VCL NAL unit of a PU, in contrast to the constraint that the encoder must use a prefix APS NAL unit when the APS is sent before the first VCL NAL of a PU. Furthermore, the insertion may result in a picture unit having both a suffix APS NAL unit and a prefix APS NAL unit containing APSs with the same identifier and type but with different contents, which is not allowed in VVC8.
[0014] Therefore, the application may have to rewrite the APS type (nal_unit_type) of the APS NAL unit to generate a new prefix APS NAL unit. In addition, the application may have to move and rewrite the suffix APS NAL unit as the new prefix APS NAL unit at the beginning of the next PU. If this next PU also happens to contain an APS NAL unit with the same identifier and type as the new prefix APS NAL unit, the application may also have to move and rewrite that unit.
[0015] These movement operations to make the bitstream conform to VVC8 are expensive because, in the worst case, all APS NAL units of the PU following the randomly accessed picture unit may need to be rewritten.
[0016] The method of the third aspect of the invention imposes, removes or modifies constraints on the syntactic structure to ensure that there are few or even no rewrite operations.
[0017] One embodiment also includes: prohibiting a suffix APS NAL unit associated with a particular VCL NAL unit from being used by the VCL NAL unit of a picture unit containing the suffix APS NAL unit; and allowing the VCL NAL unit of a picture unit that follows the suffix APS NAL unit in decoding order to use the suffix APS NAL unit.
[0018] Another embodiment also includes: constraints may include APS NAL units in a picture unit such that: a prefix APS NAL unit must precede any suffix APS NAL unit in the picture unit and before the last VCL NAL unit of the picture unit; and a suffix APS NAL unit must follow any prefix APS NAL unit in the picture unit and after the first VCL NAL unit of the picture unit.
[0019] Another embodiment further comprises prohibiting the inclusion of a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit.
[0020] Another embodiment also includes prohibiting the inclusion of a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit.
[0021] Another embodiment also includes prohibiting, in a picture unit, a VCL NAL unit that refers to an APS with a specific APS type and a specific APS identifier to be followed by a prefix APS NAL unit containing an APS with the same APS type and the same APS identifier. This measure can also be applied to the second aspect of the present invention without (as in the third aspect) permitting the inclusion of a prefix APS NAL unit and a suffix APS NAL unit with the same APS type and the same APS identifier but different content in the same picture unit.
[0022] Another embodiment further comprises: prohibiting, in a picture unit, a VCL NAL unit that refers to an APS with a specific APS type and a specific APS identifier to be preceded by a suffix APS NAL unit containing an APS with the same APS type and the same APS identifier. This measure can also be applied to the first aspect of the present invention without (as in the third aspect) permitting the inclusion of a prefix APS NAL unit and a suffix APS NAL unit with the same APS type and the same APS identifier but different content in the same picture unit.
[0023] Furthermore, the last two measures may be used without prohibiting (as in the first aspect) including a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit, without prohibiting (as in the second aspect) including a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit, and without permitting (as in the third aspect) including prefix APS NAL units and suffix APS NAL units having the same APS type and the same APS identifier but different contents in the same picture unit. Therefore, according to another aspect of the present invention, a method for encoding an image sequence in a bitstream is provided, comprising: providing a series of picture units in the bitstream, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of the VCL NAL units contains coded image data, each of the adaptation parameter set NAL units contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be contained in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be contained in the suffix APS NAL unit. NAL units, wherein the APS NAL units each have an APS type and an APS identifier; one or both of the following are performed: a VCL NAL unit that is prohibited from referencing an APS with a specific APS type and a specific APS identifier in a picture unit is followed by a prefix APS NAL unit that contains an APS with the same APS type and the same APS identifier; and a VCL NAL unit that is prohibited from referencing an APS with a specific APS type and a specific APS identifier in a picture unit is preceded by a suffix APS NAL unit that contains an APS with the same APS type and the same APS identifier.
[0024] According to a fourth aspect of the present invention, a method for decoding a coded image sequence is provided, comprising: receiving a bitstream having a series of picture units, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of which contains coded image data, each of which contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit. in a NAL unit; and checking the received bitstream for compliance with one or more compliance criteria, wherein one of the one or more compliance criteria is a constraint that prohibits including a prefix APS NAL unit in a picture unit after a first NAL unit of the picture unit.
[0025] According to a fifth aspect of the present invention, a method for decoding a coded image sequence is provided, wherein, instead of checking the conformity of a received bitstream using a constraint that prohibits including a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit (as in the fourth aspect), the conformity of the received bitstream is checked using a constraint that prohibits including a suffix APS NAL unit in the picture unit before the last NAL unit of the picture unit.
[0026] In one embodiment, the check involves checking compliance with both a constraint prohibiting inclusion of a suffix APSNAL unit in a picture unit before the last NAL unit of the picture unit and a constraint prohibiting inclusion of a prefix APSNAL unit in the picture unit after the first NAL unit of the picture unit.
[0027] According to a sixth aspect of the present invention, a method for decoding a coded image sequence is provided, comprising: receiving a bitstream having a series of picture units, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units comprising video coding layer NAL units, i.e., VCL NAL units and also comprising adaptation parameter set NAL units, each of the VCL NAL units comprising coded image data, each of the adaptation parameter set NAL units comprising an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, the APS NAL units that can be included in the series of picture units comprising a prefix APS NAL unit and a suffix APS NAL unit, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit, the APS The NAL units each have an APS type and an APS identifier; and a step of checking the received bitstream for compliance with one or more compliance criteria, one of the one or more compliance criteria permitting the inclusion of prefix APS NAL units and suffix APS NAL units having the same APS type and the same APS identifier but different content in the same picture unit.
[0028] In one embodiment, the conformance criteria include: prohibiting a suffix APS NAL unit associated with a particular VCL NAL unit from being used by VCL NAL units of picture units that contain the particular VCL NAL unit; and allowing the suffix APS NAL unit to be used by VCL NAL units of picture units that follow the suffix APS NAL unit in decoding order.
[0029] In another embodiment, the conformance criteria include: constraints may be included in the APS NAL units in the picture unit such that: the prefix APS NAL unit must precede any suffix APS NAL unit in the picture unit and before the last VCL NAL unit of the picture unit; and the suffix APS NAL unit must follow any prefix APS NAL unit in the picture unit and after the first VCL NAL unit of the picture unit.
[0030] In another embodiment, the conformance criteria include a constraint prohibiting inclusion of a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit.
[0031] In another embodiment, the conformance criteria include prohibiting the inclusion of a prefix APSNAL unit in a picture unit after a first NAL unit of the picture unit.
[0032] In another embodiment, the conformance criteria include: prohibiting a VCL NAL unit that refers to an APS with a specific APS type and a specific APS identifier from being followed by a prefix APS NAL unit containing an APS with the same APS type and the same APS identifier in a picture unit. This measure can also be applied to the fifth aspect of the present invention without (as in the sixth aspect) permitting the inclusion of a prefix APS NAL unit and a suffix APS NAL unit with the same APS type and the same APS identifier but different content in the same picture unit.
[0033] In another embodiment, the conformance criteria include: prohibiting a VCL NAL unit that references an APS with a specific APS type and a specific APS identifier in a picture unit from being preceded by a suffix APS NAL unit containing an APS with the same APS type and the same APS identifier. This measure can also be applied to the fourth aspect of the present invention without (as in the sixth aspect) permitting the inclusion of a prefix APS NAL unit and a suffix APS NAL unit with the same APS type and the same APS identifier but different content in the same picture unit.
[0034] Furthermore, the last two measures may be used without prohibiting (as in the fourth aspect) including a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit, without prohibiting (as in the fifth aspect) including a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit, and without permitting (as in the sixth aspect) including a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different content in the same picture unit. Therefore, according to another aspect of the present invention, a method for decoding a coded image sequence is provided, comprising: receiving a bitstream having a series of picture units, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer (NAL) units, the NAL units that may be included in the series of picture units comprising video coding layer (VCL) NAL units and also comprising adaptation parameter set NAL units, each VCL NAL unit comprising coded image data, each adaptation parameter set NAL unit comprising an adaptation parameter set (APS) having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that may be included in the series of picture units comprising a prefix APS NAL unit and a suffix APS NAL unit, wherein if the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and if the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit, each of the APS The NAL unit has an APS type and an APS identifier; and the conformance of the received bitstream is checked using one or more conformance criteria, wherein the conformance criteria include one or both of the following criteria: a VCL NAL unit that is prohibited from referencing an APS with a specific APS type and a specific APS identifier in a picture unit is followed by a prefix APS NAL unit containing an APS with the same APS type and the same APS identifier; and a VCL NAL unit that is prohibited from referencing an APS with a specific APS type and a specific APS identifier in a picture unit is preceded by a suffix APS NAL unit containing an APS with the same APS type and the same APS identifier.
[0035] In the method embodying the aforementioned first to sixth aspects and other aspects of the present invention, the NAL unit is not limited to the VCL NAL unit and the APS NAL unit. For example, the NAL unit that may be included in a series of picture units may also include a picture header NAL unit, which is neither a VCL NAL unit nor an APS NAL unit, and if present in a picture unit, it is before the first VCL NAL unit of the picture unit. In this case, the APS NAL unit referenced in the PH must not only be before the first VCL NAL unit, but also before the PH NAL unit. Instead of the PH NAL unit, a more general conception is a non-VCL NAL unit, which is neither a VCL NAL nor an APS NAL unit, which signals a reference to the APS by one or more VCL NAL units. The constraints on the ordering of APS NAL units should now be relative to these non-VCL NAL units. For example, the prefix APS NAL unit should be before the first non-VCL NAL unit and the first VCL NAL unit.
[0036] According to a seventh aspect of the present invention, there is provided an apparatus for encoding an image sequence in a bitstream, comprising: a component for providing a series of picture units in the bitstream, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of the VCL NAL units contains coded image data, each of the adaptation parameter set NAL units contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be contained in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be contained in the suffix APS NAL unit. A NAL unit; and means for prohibiting the inclusion of a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit.
[0037] According to an eighth aspect of the present invention, there is provided an apparatus for encoding an image sequence in a bitstream, comprising: a component for providing a series of picture units in the bitstream, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of the VCL NAL units contains coded image data, each of the adaptation parameter set NAL units contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be contained in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be contained in the suffix APS NAL unit. A NAL unit; and means for prohibiting the inclusion of a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit.
[0038] According to a ninth aspect of the present invention, there is provided an apparatus for encoding an image sequence in a bitstream, comprising: a component for providing a series of picture units in the bitstream, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of the VCL NAL units contains coded image data, each of the adaptation parameter set NAL units contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be contained in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be contained in the suffix APS NAL unit. NAL units, each of the APS NAL units has an APS type and an APS identifier; and means for permitting the inclusion of a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents in the same picture unit.
[0039] According to a tenth aspect of the present invention, there is provided an apparatus for decoding a coded image sequence, comprising: a component for receiving a bitstream having a series of picture units, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of which contains coded image data, each of which contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit. in a NAL unit; and means for checking the received bitstream for compliance with one or more compliance criteria, wherein one of the one or more compliance criteria is a constraint prohibiting the inclusion of a prefix APSNAL unit in a picture unit after the first NAL unit of the picture unit.
[0040] According to an eleventh aspect of the present invention, there is provided an apparatus for decoding a coded image sequence, comprising: a component for receiving a bitstream having a series of picture units, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of which contains coded image data, each of which contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit. in a NAL unit; and means for checking the received bitstream for compliance with one or more compliance criteria, wherein one of the one or more compliance criteria is a constraint prohibiting inclusion of a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit.
[0041] According to a twelfth aspect of the present invention, there is provided an apparatus for decoding a coded image sequence, comprising: a component for receiving a bitstream having a series of picture units, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of which contains coded image data, each of which contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit. NAL units, each of which has an APS type and an APS identifier; and means for checking the conformance of the received bitstream with one or more conformance criteria, one of which permits including, in the same picture unit, prefix APS NAL units and suffix APS NAL units having the same APS type and the same APS identifier but different content.
[0042] In the methods of aspects 4 to 6 and other aspects and the devices of aspects 10 to 12, in the case where the conformance check reveals a non-compliant bitstream, the decoding of the bitstream can be abandoned in whole or in part. In addition, actions such as notifying the user of the decoder of the error can be taken. The decoder can also signal the encoder that the bitstream does not conform and is not suitable for decoding. The encoder can respond by re-encoding the image sequence to produce a conforming bitstream. As described later, the conformance check is not mandatory in all decoding methods or all decoders embodying the present invention.
[0043] According to a thirteenth aspect of the present invention, there is provided a program which, when executed by a processor or a computer, causes the processor or the computer to execute the method of any one of the first to sixth aspects of the present invention.
[0044] The program may be provided separately, or may be on, carried by or in a carrier medium. The carrier medium may be non-transitory, such as a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transient, such as a signal or other transmission medium. The signal may be transmitted via any suitable network, including the Internet.
[0045] According to a fourteenth aspect of the present invention, there is provided a bitstream representing a coded image sequence and having a series of picture units, each of the picture units corresponding to a coded image and including one or more network abstraction layer (NAL) units, the NAL units that may be included in the series of picture units including video coding layer (VCL) NAL units and also including adaptation parameter set NAL units, each VCL NAL unit including coded image data, each adaptation parameter set NAL unit including an adaptation parameter set (APS) having parameters for performing one or more types of processing operations on the image data included in the one or more VCL NAL units, and the APS NAL units that may be included in the series of picture units including prefix APS NAL units and suffix APS NAL units, wherein if the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and if the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit. NAL unit; wherein none of the picture units of the series of picture units includes a prefix APS NAL unit after the first NAL unit of the picture unit.
[0046] An alternative way to represent the previous bitstream characteristic is to prohibit including a prefix APSNAL unit in a picture unit after the first NAL unit of the picture unit.
[0047] According to a fifteenth aspect of the present invention, there is provided a bitstream representing a coded image sequence and having a series of picture units, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of which contains coded image data, each of which contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit. NAL unit; wherein none of the picture units in the series of picture units includes a suffix APS NAL unit before the last NAL unit of the picture unit.
[0048] An alternative way to represent the previous bitstream characteristic is to prohibit including a suffix APSNAL unit in a picture unit before the last NAL unit of the picture unit.
[0049] Preferably, none of the picture units of the series of picture units comprises a suffix APS NAL unit before the last NAL unit of the picture unit, and none of the picture units of the series of picture units comprises a prefix APS NAL unit after the first NAL unit of the picture unit.
[0050] An alternative way of expressing the previous bitstream characteristics is to prohibit including a suffix APS NAL unit in a picture unit before the last NAL unit of the picture unit, and to prohibit including a prefix APS NAL unit in a picture unit after the first NAL unit of the picture unit.
[0051] According to a sixteenth aspect of the present invention, a bitstream is provided, which represents a coded image sequence and has a series of picture units in the bitstream, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, each of which contains coded image data, each of which contains an adaptation parameter set, i.e., APS, having parameters for performing one or more types of processing operations on the image data contained in one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and in the case where the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit. In the NAL unit, the APS NAL units each have an APS type and an APS identifier; wherein at least one picture unit in the series of picture units includes a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents.
[0052] An alternative way to represent the previous bitstream characteristics is to allow including prefix APS NAL units and suffix APS NAL units with the same APS type and the same APS identifier but different content in the same picture unit.
[0053] In one embodiment: in each picture unit of the series of picture units where there is a suffix APS NAL unit, the suffix unit is not used by the VCL NAL units of the picture units containing the specific VCL NAL unit; and for at least one picture unit having such a suffix APS NAL unit that is not used by the VCL NAL units of the picture units containing the specific VCL NAL unit, the suffix APS NAL unit is used by one or more VCL NAL units of one or more picture units that follow the suffix APS NAL unit in decoding order.
[0054] An alternative way of expressing the previous bitstream characteristic is to prohibit the use of the suffix APS NAL unit by the VCL NAL unit of the picture unit containing a particular VCL NAL unit; and to allow the suffix APS NAL unit to be used by the VCL NAL unit of the picture unit that follows the suffix APS NAL unit in decoding order.
[0055] In one embodiment: in each picture unit that includes a prefix APS NAL unit, the prefix APS NAL unit precedes any suffix APS NAL unit in the picture unit and before the last VCL NAL unit of the picture unit; and in each picture unit that includes a suffix APS NAL unit, the suffix APS NAL unit must follow any prefix APS NAL unit in the picture unit and after the first VCL NAL unit of the picture unit.
[0056] An alternative way to represent the previous bitstream characteristics is that the APS NAL units included in the series of picture units are constrained so that: the prefix APS NAL unit precedes any suffix APS NAL unit in the picture unit and before the last VCL NAL unit of the picture unit; and the suffix APS NAL unit follows any prefix APS NAL unit in the picture unit and after the first VCL NAL unit of the picture unit.
[0057] In one embodiment, none of the series of picture units includes a suffix APS NAL unit before the last NAL unit of the picture unit.
[0058] An alternative way to represent the previous bitstream characteristic is to prohibit including a suffix APS NAL unit in a picture unit after the last NAL unit of the picture unit.
[0059] In one embodiment, none of the series of picture units includes a prefix APS NAL unit after the first NAL unit of the picture unit.
[0060] An alternative way to represent the previous bitstream characteristic is to prohibit including a prefix APSNAL unit in a picture unit after the first NAL unit of the picture unit.
[0061] In one embodiment, in any picture unit including a VCL NAL unit referencing an APS having a specific APS type and a specific APS identifier, the referenced VCL NAL unit is not followed by a prefix APS NAL unit containing an APS having the same APS type and the same APS identifier. This measure can also be applied to the fifteenth aspect of the present invention without (as in the sixteenth aspect) allowing the inclusion of a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different content in the same picture unit.
[0062] An alternative way of expressing the previous bitstream characteristic is that, in a picture unit, a VCL NAL unit referencing an APS with a specific APS type and a specific APS identifier is prohibited from being followed by a prefix APS NAL unit containing an APS with the same APS type and the same APS identifier.
[0063] In one embodiment, in any picture unit including a VCL NAL unit referencing an APS having a specific APS type and a specific APS identifier, the referenced VCL NAL unit is not preceded by a suffix APS NAL unit containing an APS having the same APS type and the same APS identifier. This measure can also be applied to the fourteenth aspect of the present invention without (as in the sixteenth aspect) allowing the inclusion of a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different content in the same picture unit.
[0064] An alternative way of expressing the previous bitstream characteristic is that, within a picture unit, VCL NAL units that refer to an APS with a specific APS type and a specific APS identifier are prohibited from being preceded by a suffix APS NAL unit containing an APS with the same APS type and the same APS identifier.
[0065] Furthermore, the last two measures may be used in the absence of the bitstream characteristic (of the fourteenth aspect) (i.e., none of the picture units in the series of picture units includes a prefix APS NAL unit after the first NAL unit of the picture unit), in the absence of the bitstream characteristic (of the fifteenth aspect) (i.e., none of the picture units in the series of picture units includes a suffix APS NAL unit before the last NAL unit of the picture unit), and in the absence of the bitstream characteristic (of the sixteenth aspect) (at least one picture unit in the series of picture units includes a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents). Therefore, according to another aspect of the present invention, there is provided a bitstream representing a sequence of coded images and having a series of picture units in the bitstream, each of the picture units corresponding to a coded image and comprising one or more network abstraction layer (NAL) units, the NAL units that may be included in the series of picture units comprising video coding layer (VCL) NAL units and also comprising adaptation parameter set NAL units, each VCL NAL unit comprising coded image data, each adaptation parameter set NAL unit comprising an adaptation parameter set (APS) having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that may be included in the series of picture units comprising prefix APS NAL units and suffix APS NAL units, wherein if the APS exists in the picture unit before the first VCL NAL of the picture unit, the APS must be included in the prefix APS NAL unit, and if the APS exists in the picture unit after the last VCL NAL of the picture unit, the APS must be included in the suffix APS NAL unit, each of the APS The NAL unit has an APS type and an APS identifier; the bitstream has one or both of the following characteristics: in any picture unit including a VCL NAL unit referencing an APS having a specific APS type and a specific APS identifier, the referenced VCL NAL unit is not followed by a prefix APS NAL unit containing an APS having the same APS type and the same APS identifier; and in any picture unit including a VCL NAL unit referencing an APS having a specific APS type and a specific APS identifier, the referenced VCL NAL unit is not preceded by a suffix APS NAL unit containing an APS having the same APS type and the same APS identifier.
[0066] How to use the bitstreams of the above fourteenth to sixteenth aspects and other aspects is not particularly limited. A seventeenth aspect of the present invention provides a method of encoding an image sequence in a bitstream according to any one of the fourteenth to sixteenth aspects and other aspects.
[0067] An eighteenth aspect of the present invention provides a method for decoding an encoded image sequence, the method comprising receiving a bit stream according to any one of the fourteenth to sixteenth aspects.
[0068] In this regard, it is sufficient to receive a bit stream. No conformance check is required. For example, a decoder may simply receive a bit stream of any one of aspects 14 to 16 and decode it. For example, an embodiment further comprises: decoding a NAL unit, obtaining image data contained in a VCL NAL unit and parameters of an APS contained in an APS NAL unit, and processing the obtained image data using the obtained APS parameters.
[0069] According to a nineteenth aspect of the present invention, there is provided a bit stream generated by the encoding method of any one of the first to third aspects of the present invention.
[0070] A bitstream is typically in the form of a transient signal. However, in a non-transient form, the bitstream may be stored, for example, in a computer-readable storage device or a recording medium (such as a media storage device, etc.). DVDs, blue discs, or other optical storage media are examples of storage media for bitstreams. Therefore, according to the twentieth aspect of the present invention, there is provided a computer-readable storage medium for storing a bitstream of any one of the fourteenth to sixteenth aspects and the nineteenth aspect of the present invention.
[0071] Any features in one aspect of the invention may be applied to other aspects of the invention in any appropriate combination. In particular, method aspects may be applied to apparatus aspects, and vice versa.
[0072] Furthermore, features implemented in hardware may be implemented in software and vice versa. Any references herein to software and hardware features should be interpreted accordingly.
[0073] Any device feature as described herein may also be provided as a method feature, and vice versa. As used herein, a means plus function feature may be alternatively expressed in terms of its corresponding structure (such as a suitably programmed processor and associated memory, etc.).
[0074] It will also be appreciated that specific combinations of the various features described and defined in any aspect of the present invention may be independently implemented, provided and / or used. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Embodiments of the present invention will now be described, by way of example only, and with reference to the following drawings, in which:
[0076] Figure 1 Shows partitioning of a picture into blocks and slices according to an embodiment of the present invention;
[0077] Figure 2 An example VVC bitstream is shown;
[0078] Figure 3 is a block diagram schematically illustrating a data communication system in which one or more embodiments of the present invention may be implemented;
[0079] Figure 4 is a block diagram illustrating components of a processing device that may implement one or more embodiments of the present invention;
[0080] Figure 5 is a block diagram illustrating components of an encoder that may implement one or more embodiments of the present invention;
[0081] Figure 6 is a block diagram illustrating components of a decoder in which one or more embodiments of the present invention may be implemented;
[0082] Figure 7 is a flowchart illustrating an encoding process according to an embodiment of the present invention;
[0083] Figure 8 is a flowchart illustrating a decoding process according to an embodiment of the present invention;
[0084] Fig. 9 is shown in more detail Figure 8 A flowchart of a portion of the decoding process;
[0085] Fig.10 is shown in more detail Figure 8 Flowchart of other parts of the decoding process;
[0086] Fig.11A An example of a bitstream compliant with VVC8 is shown;
[0087] Fig. 11B shows a bit stream according to an embodiment of the present invention;
[0088] Fig. 12A Another example of a bitstream compliant with VVC8 is shown;
[0089] Fig. 12B Shows that when random access is required Fig. 12A A first variant of the bit stream;
[0090] Fig. 12C Shows that when random access is required Fig. 12A A second variation of the bit stream;
[0091] Fig.12D The corresponding embodiment of the present invention is shown Fig. 12A An example bit stream of
[0092] Fig.13 shows a bit stream according to another embodiment of the present invention;
[0093] Fig.14 is a diagram illustrating a network camera system in which one or more embodiments of the present invention may be implemented; and
[0094] Fig.15 is a diagram illustrating a smart phone in which one or more embodiments of the present invention may be implemented. DETAILED DESCRIPTION
[0095] The embodiments of the invention described below are directed to improving the encoding and decoding of images (or pictures).
[0096] In the present specification, "signaling" may refer to inserting (providing / including / encoding in) information related to one or more parameters or syntax elements into a bitstream or extracting / obtaining (decoding) the information from a bitstream, wherein the information is, for example, any one or more of information for determining an identifier of a sub-picture, the size / width / height of the sub-picture, whether only a single image portion (e.g., a slice) is included in the sub-picture, whether the slice is a rectangular slice, and / or the number of slices included in the sub-picture.
[0097] In this specification, "processing" may refer to any type of operation performed on data, for example, encoding or decoding image data of one or more images / pictures.
[0098] In this specification, the term "slice" is used as an example of an image portion (other examples of such an image portion would be an image portion comprising one or more coding tree units). It should be understood that embodiments of the present invention may also be implemented based on image portions instead of slices and appropriately modified parameters / values / syntax, such as an image portion header instead of a slice header or a slice segment header, etc. It should also be understood that various information described herein as being signaled in a slice header, a slice segment header, a sequence parameter set (SPS), or a picture parameter set (PPS) may be signaled elsewhere as long as it is able to provide the same functionality provided by signaling the information in these media. It should also be understood that any of a slice, a block group, a block, a coding tree unit (CTU) / largest coding unit (LCU), a coding tree block (CTB), a coding unit (CU), a prediction unit (PU), a transform unit (TU), or a pixel / sample block may be referred to as an image portion.
[0099] It should also be understood that when a component or tool is described as "active", the component / tool is "enabled" or "usable" or "used"; when described as "inactive", the component / tool is "disabled" or "unavailable" or "not used"; and "can be inferred" means that the relevant value or parameter can be determined / obtained from other information without explicit signaling in the bitstream. In addition, it should also be understood that when a flag is described as "active", it means that the flag indicates that the relevant component / tool is "active" (i.e., "valid").
[0100] In this specification, unless otherwise stated, relevant terms have the same definitions as in the latest VVC draft 8 (VVC8) set out below. Terms in italics have their own VVC8 definitions.
[0101] Slice: An integer number of complete blocks of a picture, or an integer number of consecutive complete CTU rows within a block, contained exclusively in a single NAL unit.
[0102] Slice header: A portion of a coded slice that contains data elements related to all blocks or CTU rows within blocks represented in the slice.
[0103] Block: A rectangular area of a CTU within a specific block column and a specific block row in a picture.
[0104] Sub-image: A rectangular area of one or more strips within an image.
[0105] Picture (or image): An array of luma samples in monochrome format or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0106] Coded picture: A coded representation of a picture that includes VCL NAL units with a specific value of nuh_layer_id within an AU and includes all CTUs of the picture.
[0107] Encoded representation: A data element represented in an encoded form.
[0108] Raster Scan: The mapping of a rectangular two-dimensional pattern to a one-dimensional pattern such that the first entry in the one-dimensional pattern is from the top row of the two-dimensional pattern scanned from left to right, followed similarly by the second, third, etc. rows of the pattern (downwards), each scanned from left to right.
[0109] Block: an M×N (M columns×N rows) array of samples, or an M×N array of transform coefficients.
[0110] Coding block: A block of M×N samples for some values of M and N such that the partitioning of CTBs into coding blocks is a partition.
[0111] Coding Tree Block (CTB): An NxN block of samples for some value of N such that the partitioning of components into CTBs is a partition.
[0112] Coding Tree Unit (CTU): A CTB of luma samples, two corresponding CTBs of chroma samples of a picture with three sample arrays, or a CTB of samples of a monochrome picture or a picture coded using three separate color planes and a syntax structure for coding the samples.
[0113] Coding Unit (CU): A coding block of luma samples of a picture with three sample arrays, two corresponding coding blocks of chroma samples, or a coding block of samples of a monochrome picture or a picture coded using three separate color planes and a syntax structure for coding the samples.
[0114] Component: An array or a single sample from one of the three arrays (luminance and two chrominance) that make up a picture in 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample from an array that makes up a picture in monochrome format.
[0115] Picture Parameter Set (PPS): A syntax structure containing syntax elements that apply to zero or more entire coded pictures as determined by syntax elements found in a picture header or slice header.
[0116] Sequence Parameter Set (SPS): A syntax structure containing syntax elements that apply to zero or more complete CLVs as determined by the contents of the syntax elements found in the PPS referenced by the syntax elements found in the picture header.
[0117] Adaptation Parameter Set (APS): A syntax structure containing syntax elements that apply to zero or more slices as determined by zero or more syntax elements found in a slice or picture header.
[0118] Network Abstraction Layer (NAL) unit: contains an indication of the type of data to follow and the syntactic structure of bytes comprising that data, in the form of RBSPs, interspersed with emulation prevention bytes as required.
[0119] Video Coding Layer (VCL) NAL unit: a collective term for coded slice NAL units and a subset of NAL units with a reserved value of nal_unit_type that are classified as VCL NAL units in this specification.
[0120] Picture Header (PH): A syntax structure containing syntax elements that apply to all slices of a coded picture.
[0121] Slice header: The portion of a coded slice that contains data elements related to all blocks or CTU rows within blocks represented in the slice.
[0122] Adaptive Loop Filter (ALF): A filtering process applied as part of the decoding process and controlled by parameters conveyed in the APS.
[0123] Luma Mapping with Chroma Scaling (LMCS): A process applied as part of the decoding process that maps luma samples to specific values and may apply scaling operations to the values of chroma samples.
[0124] Scaling list: A list that associates each frequency index with a scaling factor used for the scaling process.
[0125] Picture Unit (PU): A set of NAL units associated with each other according to a specified classification rule, which are consecutive in decoding order and contain exactly one coded picture.
[0126] Access Unit (AU): A set of PUs belonging to different layers and containing coded pictures associated with the same time for output from the DPB.
[0127] Figure 1 The partitioning of a picture into blocks and slices according to an embodiment of the present invention is shown, which is compatible with VVC8. Pictures 101 and 102 are divided into coding tree units (CTUs) represented by dotted lines. CTU is the basic unit of encoding and decoding of VVC8. For example, in VVC7, CTU can encode an area of 128×128 pixels.
[0128] A coding tree unit (CTU) may also be referred to as a block (of pixels or component samples (values)), a macroblock, or even a coding block. A coding tree unit may be used to encode / decode different image components of a picture simultaneously, or may be limited to only one image component so that different image components of a picture may be encoded / decoded separately. When the data of an image includes separate data for each component, a CTU is a group of coding tree blocks (CTBs), one CTB for each component.
[0129] like Figure 1 As shown, the picture can also be partitioned according to a block grid (i.e., divided into one or more block grids) using block boundaries represented by thin solid lines. A block is a picture portion (a portion of a picture) that is a rectangular area (of pixels / component samples) that can be defined independently of the CTU partition. For example, in VVC8, a block can also correspond to a CTU sequence to be Figure 1 In the example represented in , the partitioning technique may constrain the boundaries of blocks to be consistent / aligned with the boundaries of CTUs.
[0130] Blocks are defined so that block boundaries break spatial dependencies of the encoding / decoding process. In other words, in a given picture, blocks are defined / specified so that they can be encoded / decoded independently of other spatially "neighboring" blocks of the same picture. This means that the encoding / decoding of CTUs in a block is not based on pixels / samples or reference data from other blocks in the same picture.
[0131] Some encoding / decoding systems (e.g., embodiments of the present invention or embodiments for VVC8) provide the concept of a slice (i.e., also use a partitioning technique based on one or more slices). This mechanism enables a picture to be partitioned into one or several groups of blocks, which are collectively referred to as a slice. Each slice consists of one block or several blocks or partial blocks. As shown in pictures 101 and 102, two different types of slices are provided. The first type of slice is limited to a strip that forms a rectangular area / region in the picture as indicated by the thick solid line in picture 101. Picture 101 has a partition of the picture into six different rectangular strips (0) to (5). The second type of strip is limited to continuous blocks in raster scan order as indicated by the thick solid line in picture 102 (so that they form a sequence of blocks). Picture 102 has a partition of the picture into three different strips (0) to (2) consisting of continuous blocks in raster scan order.
[0132] Typically, rectangular strips are a structure / arrangement / configuration used to handle selection of a region of interest (ROI) in a video.
[0133] Slices can be encoded in (or decoded from) a bitstream as one or several Network Abstraction Layer (NAL) units. A NAL unit is a logical unit of data used to encapsulate data in an encoded / decoded bitstream (e.g., a packet containing an integer number of bytes, where multiple packets together form the encoded video data).
[0134] In the encoding / decoding system of VVC8, a slice is usually encoded as a single NAL unit. When a slice is encoded as several NAL units in the bitstream, each NAL unit of the slice is called a slice segment. The slice segment includes a slice segment header containing coding parameters for the slice segment. According to a variant, the header of the first slice segment NAL unit of the slice contains all coding parameters for the slice. The slice segment header of the subsequent NAL unit of the slice may contain fewer parameters than the first NAL unit. In this case, the first slice segment is an independent slice segment, and the subsequent segments are dependent slice segments (because they depend on the coding parameters of the NAL unit from the first slice segment).
[0135] Figure 2The organization (i.e., structure, configuration, or arrangement) of a bitstream according to an embodiment of the present invention that complies with the requirements of a coding system of VVC8 is shown. The bitstream 200 consists of data representing / indicating an ordered sequence of syntax elements and encoded (image) data. The syntax elements and the encoded (image) data are placed (i.e., packaged / grouped) into a series of NAL units 201 to 209. There are different NAL unit types. The network abstraction layer (NAL) provides the ability to encapsulate the bitstream into packets for different protocols, such as Real Time Protocol / Internet Protocol (RTP / IP), ISO base media file format, etc. The network abstraction layer also provides a framework for anti-packet loss.
[0136] NAL units are divided into video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain actual coded video data. Non-VCL NAL units contain additional information. The additional information can be parameters required for decoding the coded video data, or supplementary data that can enhance the usability of the decoded video data. Figure 2 The NAL units 206 in FIG. 2 correspond to slices (ie, they include the actual coded video data for the slices) and constitute Figure 2 An example bitstream of a VCL NAL unit.
[0137] All NAL units (VCL and associated non-VCL NAL units) encoding a single coded picture form one picture unit. In this example, non-VCL NAL unit 208 is associated with two VCL NAL units 206, and these three NAL units may together form one picture unit.
[0138] Different NAL units 201 to 205 and 209 correspond to different parameter sets, and these NAL units are non-VCL NAL units.
[0139] DCI stands for Decoding Capability Information. A DCI NAL unit 201 contains parameters that are constant for a given decoding process.
[0140] VPS stands for Video Parameter Set. A VPS NAL unit 202 contains parameters defined for the entire video (eg, the entire video includes one or more sequences of pictures / images), and is therefore applicable when decoding the encoded video data of the entire bitstream.
[0141] DCI NAL units may define parameters that are more static (in the sense that the parameters are stable and do not change as much during the decoding process) than parameters in VPS NAL units. In other words, the parameters of DCI NAL units change less frequently than the parameters of VPS NAL units.
[0142] SPS stands for sequence parameter set. The SPS NAL unit 203 contains parameters defined for a video sequence (i.e., a sequence of pictures or images). Specifically, the SPS NAL unit can define the sub-picture layout and associated parameters of a video sequence. The parameters associated with each sub-picture specify the coding constraints applied to the sub-picture. According to a variation, the SPS NAL unit includes a flag that is used to indicate that the temporal prediction between sub-pictures is restricted so that only data from the same sub-picture can be used during the temporal prediction process. Another flag can enable or disable a loop filter (i.e., post-filtering) across sub-picture boundaries.
[0143] PPS stands for Picture Parameter Set. A PPS NAL unit 204 contains parameters defined for a picture or group of pictures. The syntax of the PPS as specified in VVC8 includes syntax elements for specifying the size of the picture in units of luma samples, and also includes syntax elements for specifying the partitioning of each picture in units of blocks and slices. The PPS contains syntax elements that allow the slice position in a picture / frame to be determined.
[0144] APS stands for Adaptive Parameter Set. APS contains parameters for the loop filter, which is usually an adaptive loop filter (ALF) or a shaper model (or a luma mapping with chroma scaling (LMCS) model) or a scaling matrix used at the slice level.
[0145] The APS includes an aps_params_type syntax element that describes the type of parameters present in the APS. For example, aps_params_type equal to ALF_APS indicates that the APS contains ALF parameters; aps_params_type equal to LMCS_APS indicates that the APS contains LMCS parameters, and finally, when equal to SCALING_APS, indicates that scaling list parameters are present.
[0146] The second syntax element adaptation_parameter_set_id provides an identifier of the APS.
[0147] Two types of NAL units can encapsulate APS: prefix APS NAL unit 205 and suffix APS NAL unit 209. According to the VVC8 specification, when a prefix APS NAL unit is present in a PU, the prefix APS NAL unit should not follow the last VCL NAL unit of the PU. When a suffix APS NAL unit is present in a PU, the suffix APS NAL unit should not precede the first VCL NAL unit of the PU. Between the first VCL NAL unit and the last VCL NAL unit, the prefix APS NAL unit and the suffix APS NAL unit can exist in any order. For example, the first suffix APS NAL can be followed by a prefix NAL unit, followed by a VCL NAL unit and another suffix APS NAL unit.
[0148] SEI stands for Supplemental Enhancement Information. The bitstream may also contain SEI NAL units ( Figure 3 not shown).
[0149] The frequency of occurrence (or inclusion frequency) of various parameter sets (or NAL units) in the bitstream is variable. A VPS defined for the entire bitstream may appear only once in the bitstream. In contrast, an APS defined for a slice may appear once for each slice in each picture. In practice, different slices may rely on (e.g., reference) the same APS, and therefore there are typically fewer APS NAL units in the bitstream for a picture than for slices.
[0150] The AUD NAL unit 207 is an access unit delimiter NAL unit that separates two access units. An access unit is a set of NAL units that may include one or more coded pictures with the same decoding timestamp (ie, a group of NAL units related to one or more coded pictures with the same timestamp).
[0151] The PH NAL unit 208 is a picture header NAL unit that groups parameters common to a set of slices of a single coded picture.A picture may reference one or more APSs to indicate the ALF parameters, shaper models, and scaling matrices used by slices of the picture.
[0152] Each VCL NAL unit 206 contains video / image data of a slice. A slice may correspond to an entire picture or a sub-picture, a single block or multiple blocks or a small portion of a block (partial block). For example, Figure 2 A slice includes a number of blocks 220. A slice consists of a slice header 210 and a raw byte sequence payload (RBSP) 211, which contains coded pixel / component sample data encoded as coded blocks 240.
[0153] The slice header 210 (which is part of the VCL NAL unit 206) and the picture header (which is part of the PH NAL unit 208) can reference parameters in one or more APSs by signaling the identifier and type of the or each APS NAL unit containing the referenced APS. A requirement of the VVC8 specification is that the NAL unit including the APS should precede the PH or VCL NAL unit that references the APS.
[0154] Figure 3 A data communication system is shown in which one or more embodiments of the present invention may be implemented. The data communication system comprises a transmission device (in this case a server 301) operable to transmit data packets of a data stream to a receiving device (in this case a client terminal 302) via a data communication network 300. The data communication network 300 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a hybrid network consisting of several different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcasting system in which the server 301 sends the same data content to multiple clients.
[0155] The data stream 304 provided by the server 301 may consist of multimedia data representing video and audio data. In some embodiments of the present invention, the audio and video data streams may be captured by the server 301 using a microphone and a camera, respectively. In some embodiments, the data streams may be stored on the server 301 or received by the server 301 from other data providers, or generated at the server 301. The server 301 is provided with an encoder for encoding the video and audio streams, in particular to provide a compressed bit stream for transmission, which is a more compact representation of the data presented as input to the encoder.
[0156] In order to obtain a better ratio of quality of transmitted data to the amount of transmitted data, the video data may be compressed, for example, according to the HEVC format or the H.264 / AVC format or the Versatile Video Coding (VVC) format.
[0157] The client 302 receives the transmitted bitstream and decodes the reconstructed bitstream to reproduce the video image on the display device and reproduce the audio data using the speaker.
[0158] Despite Figure 2 A streaming scenario is considered in the example of , but it will be appreciated that in some embodiments of the invention, data communication between the encoder and the decoder may be carried out using, for example, a media storage device (such as an optical disc).
[0159] Figure 4A processing device 400 configured to implement at least one embodiment of the present invention is schematically illustrated. The processing device 400 may be a device such as a microcomputer, a workstation, or a light portable device. The device 400 includes a communication bus 413, which is connected to:
[0160] - a central processing unit 411 denoted as CPU, such as a microprocessor or the like;
[0161] - a read-only memory 406, denoted ROM, for storing a computer program implementing the invention;
[0162] - a random access memory 412 represented as a RAM for storing executable codes of the method of the embodiment of the present invention, and registers suitable for recording variables and parameters required for implementing the method for encoding a digital image sequence and / or the method for decoding a bit stream according to the embodiment of the present invention; and
[0163] A communication interface 402 connected to a communication network 403, via which digital data to be processed are transmitted or received.
[0164] Optionally, the device 400 may further include the following components:
[0165] - a data storage means 404, such as a hard disk, for storing a computer program for implementing the method of one or more embodiments of the present invention and data used or generated during the implementation of one or more embodiments of the present invention;
[0166] a disk drive 405 for a disk 406, which is suitable for reading data from the disk 406 or writing data to said disk;
[0167] - A screen 409 for displaying data and / or serving as a graphical interface for interaction with a user by means of a keyboard 410 or any other pointing means.
[0168] Device 400 may be connected to various peripheral devices such as digital camera 420 or microphone 408 , each of which is connected to an input / output card (not shown) to provide multimedia data to device 400 .
[0169] The communication bus provides communication and interoperability between the various elements included in or connected to the device 400. The representation of the bus is not limiting, and in particular, the central processing unit is operable to communicate instructions to any element of the device 400 directly or via other elements of the device 400.
[0170] The disk 406 may be replaced by any information medium, such as a rewritable or non-rewritable compact disk (CD-ROM), a ZIP disk or a memory card, and, in general, by an information storage component that can be read by a microcomputer or a microprocessor, the disk 406 being integrated into the device or not, possibly removable and suitable for storing one or more programs whose execution enables the implementation of the method for encoding a digital image sequence and / or the method for decoding a bit stream according to the present invention.
[0171] The executable code may be stored in a read-only memory 406, on a hard disk 404 or on a removable digital medium such as, for example, a disk 406 as previously described, etc. According to a variant, the executable code of the program may be received via the interface 402 by means of a communication network 403 to be stored in one of the storage means of the device 400, such as a hard disk 404, etc., before execution.
[0172] The central processing unit 411 is adapted to control and direct the execution of instructions or parts of software code of one or more programs according to the present invention, the execution of instructions stored in one of the above-mentioned storage means. At power-on, one or more programs stored in non-volatile memory (e.g., on hard disk 404 or in read-only memory 406) are transferred to random access memory 412 (which then contains the executable code of one or more programs) and registers for storing variables and parameters necessary for implementing the present invention.
[0173] In this embodiment, the device is a programmable device that implements the invention using software. Alternatively, however, the invention may be implemented in hardware (for example in the form of an application specific integrated circuit or ASIC).
[0174] Figure 5 A block diagram of an encoder according to at least one embodiment of the invention is shown. The encoder is represented by connected modules, each module being adapted to implement at least one corresponding step of a method for implementing at least one embodiment of encoding an image in a sequence of images according to one or more embodiments of the invention, e.g. in the form of programming instructions executed by a CPU 411 of the apparatus 400.
[0175] The encoder 500 receives as input an initial sequence 501 of digital images i0 to in. Each digital image is represented by a set of samples (sometimes also called pixels).
[0176] After implementing the encoding process, the encoder 500 outputs a bitstream 510. The bitstream 510 includes data for a plurality of coding units or image portions such as slices, each slice including a slice header for transmitting encoded values of encoding parameters used for slice encoding, and a slice body including encoded video data.
[0177] Module 502 divides the input digital images i0 to in 501 into blocks of pixels. Blocks correspond to image parts and can have variable sizes (e.g., 4×4, 8×8, 16×16, 32×32, 64×64, 128×128 pixels, and several rectangular block sizes can also be considered). A coding mode is selected for each input block. Two families of coding modes are provided: coding modes based on spatial prediction coding (intra-frame prediction) and coding modes based on temporal prediction (inter-frame coding, merge, skip). Possible coding modes are tested.
[0178] Module 503 implements an intra prediction process in which a given block to be encoded is predicted by a predictor calculated from neighboring pixels of the block to be encoded. If intra coding is selected, the selected intra predictor and an indication of the difference between the given block and its predictor are encoded to provide a residual.
[0179] Temporal prediction is implemented by a motion estimation module 504 and a motion compensation module 505. First, a reference image from a reference image set 516 is selected, and a portion of the reference image (also referred to as a reference region or image portion) is selected by the motion estimation module 504, which is the region closest (closest in terms of pixel value similarity) to a given block to be encoded. The motion compensation module 505 then uses the selected region to predict the block to be encoded. The difference between the selected reference region and the given block (also referred to as a residual block) is calculated by the motion compensation module 505. The selected reference region is indicated by motion information (e.g., a motion vector).
[0180] Thus, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the prediction from the original block. The SKIP mode is an exception. In this case any residual is ignored.
[0181] In the intra prediction implemented by module 503, the prediction direction is encoded. In the temporal prediction, at least one motion vector is encoded. In the inter prediction implemented by modules 504, 505, 516, 518, 517, at least one motion vector or information (data) for identifying such a motion vector is encoded for temporal prediction.
[0182] If inter prediction is selected, information about the motion vector and the residual block is encoded. To further reduce the bit rate, the motion is assumed to be homogeneous and the motion vector is encoded by the difference relative to the motion vector predictor. The motion vector predictor from the set of motion information predictors is obtained by the motion vector prediction and encoding module 517 from the motion vector field 518.
[0183] The encoder 500 further comprises a selection module 506 for selecting a coding mode by applying a coding cost criterion such as a rate-distortion criterion, etc. To further reduce redundancy, a transform such as a DCT is applied to the residual block by a transform module 507, and then the obtained transform data is quantized by a quantization module 508 and entropy encoded by an entropy encoding module 509. Finally, except for the SKIP mode, the encoded residual block of the current block being encoded is inserted into the bitstream 510.
[0184] The encoder 500 also performs decoding of the encoded images to generate reference images for motion estimation of subsequent images. A set 516 of reference images is stored in a memory. This enables the encoder and decoder receiving the bitstream to have the same reference frames. The inverse quantization module 511 performs inverse quantization (dequantization) of the quantized data, followed by an inverse transform by the inverse transform module 512. The inverse intra prediction module 513 uses the prediction information to determine which predictor to use for a given block, and the inverse motion compensation module 514 actually adds the residual obtained by the module 512 to the reference area obtained from the reference image set 516.
[0185] Post filtering is then applied by module 515 to filter the reconstructed pixel frame (image or image portion). The resulting filtered and reconstructed frame is added as another reference image in set 516.
[0186] Figure 6 A block diagram of a decoder 600 according to an embodiment of the present invention is shown, which can be used to receive data from an encoder. The decoder is represented by connected modules, each module being suitable for implementing the corresponding steps of the method implemented by the decoder 600, for example in the form of programming instructions to be executed by the CPU 411 of the device 400.
[0187] The decoder 600 receives a bitstream 601 including coding units (e.g., data corresponding to image parts, blocks or coding units CU), each coding unit consisting of a header containing information related to the encoded parameters and a body containing the encoded video data. Figure 2 An example structure of a bitstream in VVC is described. Figure 5 As illustrated, for a given image portion (e.g., block or CU), the coded video data is entropy encoded on a predetermined number of bits and the index of the motion vector predictor is encoded. The received coded video data is entropy decoded by module 602. The residual data is then dequantized by module 603, after which an inverse transform is applied by module 604 to obtain pixel values.
[0188] Mode data indicating the encoding mode is also entropy decoded, and based on the mode, the encoding block (unit / set / group) of the image data is intra-type decoded or inter-type decoded.
[0189] In case of intra mode, the intra inverse prediction module 605 determines the intra predictor based on the intra prediction mode specified in the bitstream.
[0190] If the mode is inter, motion prediction information is extracted from the bitstream to find (identify) the reference region used by the encoder. The motion prediction information consists of a reference frame index and a motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 610 to obtain a motion vector.
[0191] The motion vector decoding module 610 applies motion vector decoding to each image portion (e.g., current block or CU) encoded by motion prediction. Once the index of the motion vector predictor for the current block CU has been obtained, the actual value of the motion vector associated with the image portion (e.g., current block or CU) can be decoded and used to apply inverse motion compensation by module 606. The reference image portion indicated by the decoded motion vector is extracted from the reference image in the set of reference images / pictures 608 so that module 606 can perform motion compensation. The motion vector field data 611 is updated with the decoded motion vector for inverse prediction of the subsequently decoded motion vector.
[0192] Finally, a decoded block is obtained. Where appropriate, post-filtering is applied by a post-filtering module 607. The decoder 600 finally provides a decoded video signal 609.
[0193] Figure 7 A portion of an encoding method for encoding a picture of a video into a bitstream performed by an encoder 500 according to an embodiment of the present invention is shown. A processing loop 701 continuously applies steps 702 to 705 to each picture to be encoded. The encoding of the picture begins by compressing the picture samples into parts, usually slices. In step 702, the picture is divided into one or more than one slices, and the slices are compressed continuously. Compressing the slices involves splitting the slices into coding units, and each coding unit is encoded, for example, using intra-frame or inter-frame prediction. In step 703, a set of parameters for configuring a loop filter such as an adaptive loop filter (ALF) or an LMCS filter is determined. In another example, scaling parameters for quantization of the residual are determined. These parameters are typically encoded in the APS.
[0194] Each APS has an APS type (e.g., ALF_APS, LMCS_APS, or SCALING_APS) and an APS identifier. In step 704, the APS type is set according to the contents of the APS container. The encoder maintains lists of identifiers used for each APS type. These lists each contain an APS identifier for the APS for which the APS parameters of the associated APS type were determined in step 703. Prior to the first iteration of the processing loop 701, each list is initialized to an empty state.
[0195] For a given type of APS, step 704 determines an identifier value to be associated with the current APS based on previously determined APSs and their identifier values.
[0196] For example, the following applies to various types of APSs. Determine whether the APS content (APS parameters) of the current APS is the same as a previous APS having an existing identifier present in a previous identifier list of APSs of the same type. If yes, associate the existing identifier with the current APS.
[0197] Otherwise, since all APSs with existing identifiers present in the list have different content than the current APS, a new identifier must be associated with the current APS and then inserted into the list. There are a limited number of possible identifier values available for use at any given time, and if all possible identification values have been used, an existing APS identifier within the list is determined, which the current APS will replace in the list. For example, the determined APS may be the least frequently used APS or alternatively the oldest APS.
[0198] The encoder then generates NAL units containing the encoded data in step 705. In particular, the encoder generates NAL units containing the APS, slice NAL units, and optionally picture header NAL units.
[0199] The APS NAL unit signals the type and identifier of the APS. For example, the syntax elements of the APS may be as follows:
[0200]
[0201] The adaptation_parameter_set_id syntax element is the identifier value of the APS, and aps_params_type is the type of the APS. Depending on the APS type, ALF parameters alf_data(), LMCS parameters lmcs_data()) or scaling list data scaling_list_data() may be provided.
[0202] In a given picture unit, when the APS is present before the first VCL NAL unit, the encoder must use prefix APS NAL units, and when the APS follows the last VCL of the PU, the encoder must use suffix NAL units. Between the first VCL NAL unit and the last VCL NAL unit, the encoder can use prefix APS NAL units or suffix APS NAL units (unless otherwise specified in some embodiments of the present invention).
[0203] The header of a slice NAL unit or a picture header can refer to these APS NAL units by referring to the type and identifier of the APS. However, since the header of a slice NAL unit or a picture header has a syntax element with specified semantics for the APS identifier, and the semantics of the APS identifier are different for each APS type, the APS type is implicit in the semantics and can be inferred by the decoder.
[0204] The encoder signals that the picture header references a specific APS for loop filter parameters. For example, in the currently contemplated implementation in VVC8, the picture header includes the following syntax elements:
[0205]
[0206]
[0207] The picture header in this contemplated implementation includes a number of ALF APS identifiers for applying ALF filtering to the slices of the PU. These identifiers are specified, for example, by the ph_alf_aps_id_luma[i] (where i is in the range of 0 to ph_num_alf_aps_ids_luma) syntax element. The ph_num_alf_aps_ids_luma specifies the number of APS identifiers signaled in the picture header for ALF filtering of the luma component. In addition, the ph_alf_aps_id_chroma, ph_cc_alf_cb_aps_id, and ph_cc_alf_cr_aps_id syntax elements specify the ALF APS identifiers for the chroma components.
[0208] The picture header in the contemplated implementation also includes a ph_lmcs_aps_id syntax element, which indicates the identifier of the APS with LMCS_APS type (ie, aps_params_type) containing the LMCS parameters applicable to the current PU.
[0209] Similarly, the picture header includes ph_scaling_list_aps_id, which specifies the identifier of the APS with aps_params_type equal to SCALING_APS defining the scaling list data for the current PU.
[0210] It is not necessary to use all the different APS types in embodiments of the present invention, and alternative implementations with only one or two APS types can be envisioned. Furthermore, it does not matter what the specific APS type is. For example, parameters for filters other than ALF can be contemplated. The parameters are also not limited to filtering parameters.
[0211] When the APS in use is different for each slice of a PU or for two or more slices of a PU, the APS identifier may be signaled for one or more slices in the picture header. Alternatively, the APS identifier may be signaled in the slice header instead of in the picture header NAL unit (or as an overwrite value). For example, in one implementation contemplated in VVC8, the slice header may include the following syntax elements:
[0212]
[0213]
[0214] The slice header may, for example, define slice_alf_aps_id_luma[i], which is the i-th ALF APS identifier used by the slice for the luma component. For the picture header, slice_alf_aps_id_chroma, slice_cc_alf_cb_aps_id, and slice_cc_alf_cr_aps_id may indicate the identifier of the ALF APS for the chroma component.
[0215] Figure 8 A general decoding process of an encoded video sequence according to an embodiment of the present invention is shown. The decoding process of the NAL units constituting the encoded video sequence involves using a loop 801 to continuously process the NAL units of the picture units of the encoded video sequence. For each NAL unit, in step 802, the decoder determines the type of the NAL unit by parsing the NAL unit header. For example, in VVC, the NAL unit header is 2 bytes long and contains five syntactic elements in the following order:
[0216]
[0217] The first forbidden_zero_bit is a bit that should normally be equal to 0. When equal to 1, the content of the NAL unit is unspecified and should be ignored by a conforming decoder. Next, nuh_reserved_zero_bit is a bit that is equal to 0. nuh_layer_id is an integer value represented by 6 bits. It specifies the identifier of a layer in the encoded video sequence. This syntax element is followed by nal_unit_type, which is an integer encoded on 5 bits and represents the type of the NAL unit. Unique values are assigned to each different type of NAL unit. For example, for a prefix APS NAL unit, nal_unit_type can be equal to 17, and for a suffix APS NAL unit, nal_unit_type can be equal to 18. Finally, the last three bits of the 2-byte NAL unit header encode the nuh_temporal_id_plus1 syntax element. It indicates the temporal level of the NAL unit.
[0218] The decoding process then continues in step 803 where the NAL unit data is decoded according to the type of the NAL unit.
[0219] Specifically, now refer to Fig. 9 , checks in step 901 whether the NAL unit contains APS. If so, the prefix APS NAL unit and the suffix APS NAL unit (nal_unit_type equal to 17 or 18 for VVC8) are decoded as follows: First, the decoder parses the type of APS in step 902 (specified in the aps_params_type syntax element of APS) and the identifier of the APS NAL unit in step 903 (adaptation_parameter_set_id syntax element).
[0220] In step 904, the decoder then stores the APS data contained in the NAL unit in a memory. The APS data is associated with a pair of values corresponding to the type and identifier parsed in steps 902 and 903. In addition, the decoder may also associate a Boolean value with the stored APS data, the Boolean value specifying whether the current APS is provided as a suffix NAL unit or a prefix NAL unit.
[0221] In addition, the decoder may store position data indicating the position of the current APS NAL unit relative to other NAL units. For example, the position of the current APS NAL unit may be indicated by a combination of the index of the NAL unit from the beginning of the coded video sequence and the index of the PU to which it belongs. This information enables the decoder to determine the APS data to use when a slice or picture header NAL unit references an APS with a pair of APS type and APS identifier value. …
[0222] The portion of the memory storing APS data may be referred to as an APS buffer.
[0223] The decoding process of VCL (i.e., containing slice header) and picture header (PH) NAL units is as follows: Fig.10 Shown in.
[0224] In step 1001, the decoder first checks if the NAL unit type corresponds to a VCL or PH NAL unit. For VVC8, it corresponds to nal_unit_type ranging from 0 to 12, or in the case of a picture header, equal to 19. When it is verified that the NAL unit is a VCL / PH NAL unit, the decoder applies steps 1002 to 1006. In step 1002, the slice or picture header contained in the NAL unit is parsed to determine the reference to the APS. For each APS type, the decoder maintains a reference list using the APS identifier.
[0225] First, when the NAL unit contains a picture header, the reference to the APS can be applied to all slices of the PU. In step 1003, the APS identifier and APS type present in the picture header are extracted, and for each APS type, a list of references to APSs of the relevant APS type is updated.
[0226] Step 1003 involves parsing the values of the following syntax elements (when present):
[0227] - ph_lmcs_aps_id syntax element indicating the APS identifier of any APS of APS type LMCS_APS. When not present, no LMCS filtering may be applied and nothing is inserted in the list of references to APSs of this APS type. Otherwise, the parsed value is added to the list associated with the LMCS_APS type.
[0228] - ph_scaling_list_aps_id syntax element that specifies the identifier of an APS with type equal to SCALING_APS. When not present, the scaling list may use a default value and the list of references to APSs of that APS type is unchanged. Otherwise, the decoder adds the parsed value to the list associated with the SCALING_APS type.
[0229] - ph_alf_aps_id_luma[i], where i is in the range of 0 to ph_num_alf_aps_ids_luma and / or ph_cc_alf_cb_aps_id and / or ph_cc_alf_cr_aps_id and / or ph_alf_aps_id_chroma syntax elements. These syntax elements indicate the identifier of an APS with type equal to ALF_APS. When not present for a component, it may indicate that ALF is not applied to the relevant component or that a default value is used. The list of APS types remains unchanged. Otherwise, each parsed value is added to the list associated with the ALF_APS type.
[0230] When the NAL unit is a VCL NAL unit (nal_unit_type is in the range of 0 to 12 for VVC8), a slice header is included. The slice header may include a reference to the APS found by parsing the slice header in step 1002. For example, in VVC8, the slice_alf_aps_id_luma[i], slice_alf_aps_id_chroma, slice_cc_alf_cb_aps_id, and slice_cc_alf_cr_aps_id syntax elements of the slice header indicate a reference to the ALF APS. When present in the slice header, the decoder stores the parsed identifier value in the reference APS list associated with the ALF_APS type in step 1003.
[0231] Then, in step 1004, the decoder retrieves the APSs having the type and identifier present in the APS reference list determined in step 1003 from the APS buffer filled in step 904. These APSs are marked for decoding of the VCL NAL unit of the current PU. Optionally, in step 1005, the decoder checks whether the reference to the APS contained in the picture header or slice header is valid. For example, if after updating the reference list in step 1003, the list contains a reference to an APS that is not present in the APS buffer, the decoder may return an error in the sense that no APS having the same APS type and APS identifier exists in the APS buffer, and the decoder may stop decoding of the slice or PU. In practice, all APSs required for decoding the slice or picture header of one picture unit must be provided before the NAL unit referencing the APS.
[0232] In step 1006, the NAL unit is decoded. In the case of the PH NAL unit, the decoding of the picture header mainly includes parsing the parameters provided in the NAL unit. The parameters are stored in a memory for decoding the VCL NAL unit of the PU to which the PH belongs. The decoding of the VCL NAL unit involves decoding the coding unit. The decoder typically uses the parameters parsed in the picture header NAL unit (and other non-VCL NAL units) to decode the pixel values. Specifically, the decoder uses the list of references to the APS as updated in step 1003 to access the APS in the APS buffer, and then uses the APS parameters of the referenced APS to apply LMCS, scaling transforms, and ALF filtering.
[0233] return Figure 8 , in step 802, the decoder may determine other NAL unit types besides APS, PH, and VCL NAL units, such as parameter set NAL units and SEI messages, etc. In this case, the decoding of the NAL unit in step 803 involves parsing the parameters present in the NAL unit and storing them in a memory for decoding the VCL NAL unit that may refer to these parameters.
[0234] The first set of embodiments
[0235] The proposed VVC8 syntax structure may lead to some problems in practice. For example, the size of the APS buffer required to store the APS in use may be too large. In addition, the amount of processing required to manage the APS may also be too much. Fig.11A Explain these issues.
[0236] Fig.11A An example bitstream conforming to VVC8 is shown. To conform to VVC8, it is required that in a given picture unit:
[0237] (a) if an APS is present in a picture unit before the first VCL NAL of the picture unit, then the APS MUST be contained in a prefix APS NAL unit; and
[0238] (b) If an APS is present in a picture unit after the last VCL NAL of the picture unit, the APS MUST be included in the suffix APS NAL unit.
[0239] On the other hand, between the first VCL NAL unit and the last VCL NAL unit of a PU, the encoder may use a prefix APS NAL unit or a suffix APS NAL unit.
[0240] There are other constraints:
[0241] (c) A prefix APS NAL unit or a suffix APS NAL unit associated with a particular VCL NAL unit is not used by the particular VCL NAL unit, but is used by the VCL NAL units that follow the prefix APS NAL unit or the suffix APS NAL unit in decoding order.
[0242] VVC8 defines the association between VCL and non-VCL NAL units as follows:
[0243] (1) Associated non-VCL NAL unit: a non-VCL NAL unit of a VCL NAL unit (when present), where the VCL NAL unit is the associated VCL NAL unit of the non-VCL NAL unit.
[0244] (2) Associated VCL NAL unit: the preceding VCL NAL unit in decoding order of a non-VCL NAL unit with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, or in the range of UNSPES_30…UNSPES_31; or otherwise the next VCL NAL unit in decoding order.
[0245] The effect of these definitions is that a VCL NAL unit associated with a suffix NAL unit is the VCL NAL unit that precedes the associated suffix APS NAL unit in decoding order, and a VCL NAL unit associated with a prefix NAL unit is the VCL NAL unit that follows the associated suffix APS NAL unit in decoding order.
[0246] Fig.11A The compatible bitstream has NAL units of three picture units PU-01, PU-02 and PU-03.
[0247] The first picture unit PU-01 has a PH NAL unit followed by a NAL unit NAL-01 as a prefix APS NAL unit. The prefix APS NAL unit provides an APS of the first APS type (e.g., ALF type) with an identifier equal to 0. Fig.11A In PU-01, the first APS type is indicated by horizontal hatching. In PU-01, a single slice NAL unit NAL-02 follows the prefix APS NAL unit NAL-01. The slice references an APS with an APS identifier of 0 (e.g., slice_alf_aps_luma[0] is equal to 0).
[0248] In the second picture unit PU-02, the first NAL unit NAL-03 is a prefix NAL unit also having an identifier equal to 0 but of a different type (e.g., it contains LCMS parameters). This second APS type (e.g., LMCS type) is indicated by vertical hatching. The picture header (PH) NAL unit NAL-04 then refers to this APS for the LMCS parameters by indicating that ph_lmcs_aps_id is equal to 0. The subsequent slice NAL unit NAL-05 refers to the ALF APS with an identifier equal to 0 provided in the NAL unit NAL-01 of the previous picture unit PU-01. Note that the slice NAL unit NAL-05 is associated with the suffix APS NAL unit NAL-06 because the slice NAL unit NAL-05 precedes the suffix APS NAL unit NAL-06 in decoding order. This means that, under constraint (c), the VCL NAL unit NAL-05 cannot use the APS of the suffix APS NAL unit NAL-06.
[0249] Picture unit PU-02 also includes a suffix APS NAL unit NAL-06, which includes an ALF APS with an identifier equal to 0. This APS has the same type (ALF, horizontal hatching) and the same identifier (0) as the APS of NAL unit NAL-01. The encoder therefore updates the APS of ALF type and identifier 0 to the APS of NAL unit NAL-06. Slice NAL unit NAL-07 refers to the ALF APS with an identifier equal to 0, and therefore to the ALF APS of NAL-06. This is consistent with constraint (c) because NAL-07 follows NAL-06 in decoding order. Therefore, NAL-07 is not a VCL NAL unit associated with the APS NAL unit NAL-06.
[0250] In this example bitstream, slices NAL-05 and NAL-07 of PU-02 respectively reference two different ALF APSs using the same identifier value, but the order of the APS NAL units in the bitstream implies that the ALF APS parameters for the two associated slices are different (or permitted to be different; this does not preclude the encoder from making the contents of NAL-01 and NAL-06 the same). Fig.11A bitstream, the decoder must be for a given pair of values of APS identifier and APS type (in Fig. 9In step 904 of FIG. 1 , the decoder stores two versions of the APS in memory. In a worst case example, the decoder may have to double the memory size (to maintain two versions of each APS) to store the APS required to decode the picture unit. In addition, the decoder must maintain the order of the VCL NAL units relative to the APS NAL units to determine which VCL NAL units reference the first version or the second version of the APS NAL units.
[0251] To address these issues, a first set of embodiments imposes additional constraints on the syntax structure to ensure that slices of a PU reference a single version of the APS.
[0252] By the way, in VVC8, only ALF parameters (not LMCS parameters or scaling lists) are allowed to change from one slice to another in the same picture unit. However, future versions of VVC may generally allow APS parameters to change, and the following embodiments are not limited to solving the problem of two or more versions of ALF APS parameters for a slice of a PU.
[0253] First embodiment
[0254] In the first embodiment, other constraints for the bitstream encoding (in addition to the VVC8 constraints) are:
[0255] (d1) The prefix APS NAL unit must precede the VCL NAL unit of the PU (i.e., the first slice NAL unit of a picture unit).
[0256] In other words, the freedom of VVC8 is constrained so that between the first VCL NAL unit and the last VCL NAL unit of a PU, the encoder may not use a prefix APS NAL unit. As a result, the update of the APS sent in the previous picture unit after the first VCL NAL unit is prevented. The update is performed before the first VCL NAL unit, and therefore the first VCL NAL unit (or any subsequent VCL NAL unit of a picture unit) cannot refer to a previous version of the APS updated in the current picture unit.
[0257] The decoder checks whether the constraints of the bitstream are valid in step 1005. If not, the decoder may abort the decoding process.
[0258] The encoder generates NAL units such that the bitstream constraint is valid in step 705. For example, the encoder generates a prefix NAL unit just before the first VCL NAL unit in each PU.
[0259] Second embodiment
[0260] In the second embodiment, other constraints for bitstream encoding (in addition to the VVC8 constraints) are:
[0261] (d2) Suffix APS NAL unit follows the (last) VCL NAL unit.
[0262] Similar to the constraint (d1) on prefix NAL units imposed in the first embodiment, updating of the APS sent in the previous PU before the last VCL NAL unit is prevented. The APS in the suffix cannot update the APS sent in the previous PU. For example, Fig.11A The bitstream is non-compliant because the suffix APS NAL-06 is sent before the last VCL NAL unit NAL-07 of the picture unit PU-02. Therefore, the decoder may consider the bitstream to be non-compliant in step 1005 and may return a decoding warning or error to notify the problem.
[0263] Third embodiment
[0264] Of course, the other constraints (d1) and (d2) of the first embodiment and the second embodiment may be imposed in combination.
[0265] Fig. 11B is an example of a bit stream generated by an encoder according to the second embodiment or the third embodiment of the present invention. In this example, picture units PU-01, PU-02, and PU-03 are equivalent to Fig.11A The main difference is that the encoder (in step 704) constrains the order of the APS NAL units NAL-07 in PU-02: Fig.11A The equivalent of the picture unit PU-02 with the suffix APS NAL-06 is in Fig. 11B The last VCL NAL unit (now NAL-06) in picture unit PU-02 in is sent as NAL-07. The slice references in both NAL units NAL-05 and NAL-06 have an ALF APS with identifier 0 and type equal to ALF: the APS sent in prefix APS is sent in the previous PU or at the beginning of the current PU, or the APS sent in suffix APS is sent only in the previous PU.
[0266] Decoding 904 is more efficient in terms of memory consumption because a single version of the APS is needed to decode all slices of the PU.
[0267] Additionally, these APSs are provided in the previous PU or at the beginning of the current PU, which simplifies the update process of the APS buffer. Decoding the first VCL NAL unit of a PU is a confirmation that the APS buffer status is ready for decoding, which is not the case for a bitstream compliant with VVC8. Furthermore, step 1004, where the appropriate version of an APS with a given identifier and type must be selected, is simplified, since the present invention ensures that all slices of a PU will use a unique version of the APS.
[0268] although Fig. 11B An example of the second embodiment / third embodiment is presented, but it should be understood that the same or corresponding advantages are achieved in the first embodiment. When the constraints (d1) and (d2) of the first embodiment and the second embodiment are used in combination, the best advantages are achieved.
[0269] Second Group of Embodiments
[0270] Other issues arising from the VVC8 syntax structure are addressed by the second set of embodiments described below.
[0271] The APS in VVC8 enables the reuse of parameters for one or more slices of a bitstream. These one or more slices may belong to different picture units. For example, Fig. 12A The bitstream of has three picture units PU-01, PU-02, and PU-03. PU-01 contains two prefix APS NAL units NAL-02 and NAL-06 and two suffix APS NAL units NAL-04 and NAL-08. In this example, these prefix and suffix APS NAL units are interleaved with the VCL NAL units NAL-03, NAL-05, and NAL-08 of the picture unit PU-01. The APSs of NAL-02, NAL-04, and NAL-06 have different APS types (e.g., ALF, scaling list, and LMCS), respectively, but have the same identifier 0. The APS of NAL-08 has an ALF type (like the APS of NAL-02) and has an identifier 1.
[0272] Picture unit PU-02 contains two slice NAL units NAL-09 and NAL-10. The encoder determines in step 703 that the APS of PU-01 is valid for the next PU PU-02. Slice NAL-09 can, for example, refer to NAL-06, and slice NAL-10 can refer to NAL-08. When encoding PU PU-02, the encoder determines based on the content of slice NAL-10 that the APS of type LMCS with an identifier equal to 0 needs to be updated. To this end, new parameters have been generated for the APS of type LMCS with identifier 0. The suffix APS NAL unit NAL-11 includes this APS because NAL unit NAL-11 follows the last VCL NAL unit NAL-10 (according to constraint (a) above, the prefix APS NAL unit cannot follow the last VCL NAL unit of a PU).
[0273] When an application makes random access to the bitstream to start decoding at picture unit PU-02 (assuming PU-02 is a random access point), the application MUST provide NAL units NAL-06 and NAL-08 before the slice NAL units of PU PU-02. Fig. 12B As shown, NAL-06 and NAL-08 are inserted before NAL unit NAL-09 at the beginning of PU PU-02. Fig. 12B The resulting bitstream shown in breaks two constraints of VVC8. This will cause the decoding step 1005 to enter an error state.
[0274] First, NAL-08 is a suffix APS NAL unit preceding the first VCL NAL unit of a PU, and according to constraint (a) above, the encoder MUST use a prefix APS NAL unit when sending APS before the first VCL NAL of a PU. Fig. 12B In the example of , the suffix NAL unit NAL-08 is inserted before the first VCL NAL unit of PU PU-02, which is not compliant with VVC.
[0275] Second, PU-02 has a suffix APS NAL unit NAL-11 and a prefix APS NAL unit NAL-06 containing APSs with the same identifier (0) and type (LMCS) but with different contents, which is not allowed in VVC8.
[0276] Therefore, if Fig. 12C As shown, the application must rewrite the type (nal_unit_type) of the prefix APS NAL unit NAL-08 to generate a new prefix APS NAL unit NAL-23 (nal_unit_type is set equal to 18). In addition, the application must move and rewrite Fig. 12BThe application may also have to move and rewrite the APS NAL unit NAL-24 at the beginning of PU PU-03 with the suffix APS NAL unit NAL-11. If this PU-03 PU also happens to contain an APS NAL unit with the same identifier and type as NAL-24, the application may also have to move and rewrite that APS NAL unit.
[0277] These movement operations to make the bitstream conform to VVC8 are costly because, in the worst case, all APS NAL units of the PU following the randomly accessed picture unit may need to be rewritten.
[0278] To address these issues, a second set of embodiments imposes, removes, or modifies constraints on the syntactic structure to ensure that there are fewer or even no rewrite operations.
[0279] Fourth embodiment
[0280] In VVC8, the constraint (in addition to the constraints (a) and (b) above) is that any APS with the same APS type and the same identifier must have the same content. This constraint applies even if the APS is different in the sense of different APS NAL unit types (suffixes and prefixes). In other words, if new APS parameters are needed that are different from existing APS parameters, the encoder must assign a different APS type and identifier combination to the APS NAL unit carrying the new APS parameters, or if no combination is available, the existing APS must be replaced (such as the oldest existing APS, etc.).
[0281] In the fourth embodiment, the decoder allows suffix and prefix APS NAL units with the same type and identifier to have different contents. As a result, the bitstream is valid (i.e., passes the conformance check in step 1005) if the following statements are valid for a bitstream conforming to the fourth embodiment:
[0282] (e) All APS NAL units within a PU with a specific NAL unit type (nal_unit_type) and a specific value of adaptation_parameter_set_id and a specific value of aps_params_type shall have the same content
[0283] As a result, a move operation of NAL unit NAL-11 is not necessary because NAL unit NAL-06 (prefix APS NAL unit) has a different NAL unit type than NAL unit NAL-11 (suffix APS NAL unit). Fig.12D Represents a bitstream without move operations for NAL unit NAL-11.
[0284] Fifth embodiment
[0285] The fourth embodiment described above allows prefix APS NAL units and suffix APS NAL units having the same APS identifier and type to have different contents. However, one consequence of this modification is that the APS between two slices can be updated by using different types of NAL units for providing the APS. For example, now referring to Fig.13 , PU PU-01 starts with picture header NAL unit NAL-01. This PU contains two APS NAL units NAL-02 and NAL-04, which contain APSs with the same identifier and the same type but different content. NAL-02 is a prefix NAL unit, and NAL-04 is a suffix NAL unit. As a result, slice NAL-03 can refer to APS parameters in the NAL-02 APS NAL unit, while slice NAL-05 refers to parameters in the NAL-04 APS NAL unit. As a result, decoding of PU PU-01 requires additional memory to store two versions of the APS with the same combination of type and identifier values.
[0286] In a fifth embodiment, the encoder may generate suffix APS NAL units within a given PU under the constraint that the NAL unit of the current PU does not reference the APS in the suffix APS NAL unit, regardless of the position of the suffix APS NAL unit in the PU. The bitstream containing the suffix APS NAL unit shall conform to the following constraints:
[0287] (f) The suffix APS NAL unit is not used by the VCL NAL units of the PU that contains the suffix APS NAL unit, but is used by the VCL NAL units of the PU that follows the suffix APS NAL unit in the decoding order.
[0288] refer to Fig.13 For example, in the fifth embodiment, the suffix APS NAL unit NAL-04 can only be used by NAL units in subsequent PUs. Therefore, slice NAL-05 cannot refer to the parameters in the suffix APS NAL unit NAL-04. The slices NAL-09 and NAL-10 of the next PU PU-02 can refer to the APS in the suffix APS NAL unit NAL-04. However, since the APS in NAL-04 has the same identifier and type as the APS in NAL-02, these slices (NAL-09 and NAL-10) cannot refer to the initial version of the APS in NAL-02.
[0289] Sixth embodiment
[0290] In addition to the constraint of the fifth embodiment that a NAL unit of a PU should not reference an APS within a suffix APS NAL unit of that PU, the sixth embodiment prohibits certain mixes of prefix NAL units and suffix NAL units in a given PU. This means:
[0291] (g1) when prefix APS NAL units are present in a PU, these prefix APS NAL units shall not follow the last VCL NAL unit or the suffix APS NAL unit of the PU; and
[0292] (g2) When suffix APS NAL units are present in a PU, these suffix APS NAL units shall not precede the first VCL NAL unit or the prefix APS NAL unit of the PU.
[0293] In other words, the constraint (a) that the encoder must use prefix APS NAL units when APS is sent before the first VCL NAL of a PU and the constraint (b) that the encoder must use suffix NAL units when APS follows the last VCL of a PU still apply. However, the freedom of the encoder to use prefix APS NAL units or suffix APS NAL units in any mix between the first VCL NAL unit and the last VCL NAL unit of a PU is constrained. The only permitted order is a mix of prefix APS NAL units followed by suffix APS NAL units. This constraint is independent of the APS type and the APS identifier. In a variation, a constraint may apply to one APS type but not to another.
[0294] This simplifies the decoding process because once the decoder parses the first suffix APS NAL unit of the bitstream, the decoder is able to determine that the list of APSs that can be referenced in a given PU is complete.
[0295] Seventh embodiment
[0296] As in the fourth embodiment, the seventh embodiment allows suffix APS NAL units and prefix APS NAL units with the same type and identifier to have different contents. Therefore, constraint (e) applies to bitstreams conforming to the seventh embodiment:
[0297] (e) All APS NAL units within a PU with a specific NAL unit type, a specific value of adaptation_parameter_set_id, and a specific value of aps_params_type shall have the same content
[0298] Additional constraints of the second embodiment also apply:
[0299] (d2) The suffix APS NAL unit MUST follow the last VCL NAL unit.
[0300] This constraint is independent of the APS type and the APS identifier. In a variation, a constraint may apply to one APS type but not to another APS type.
[0301] Constraints (a) and (b) of VVC8 still apply. The freedom of the encoder to use prefix APS NAL units or suffix APS NAL units in any mix between the first VCL NAL unit and the last VCL NAL unit of a PU is constrained by constraint (d2). This prevents any VCL NAL unit or picture header of a PU from referencing the APS defined in the suffix APS NAL unit. In fact, in order to be referenced, the APS should be provided before the NAL unit that references it. This last constraint implies that the APS in the suffix APS NAL unit is after all NAL units that can reference the APS in a given PU. Only the VCL NAL units from the next PU in decoding order can reference these APSs.
[0302] Eighth embodiment
[0303] The eighth embodiment builds on any one of the fourth to sixth embodiments and adds additional constraints of the first embodiment:
[0304] (d1) The prefix APS NAL unit MUST precede the first VCL NAL unit.
[0305] Constraints (a) and (b) of VVC8 still apply. The freedom of the encoder to use prefix APS NAL units or suffix APS NAL units in any mix between the first and last VCL NAL units of a PU is constrained by constraint (d1).
[0306] This not only prevents complex rewrite operations, but also ensures that the decoder does not have to buffer two versions of the APS for decoding slices of a given PU, as explained with respect to the first embodiment.
[0307] Ninth embodiment
[0308] The ninth embodiment is based on the seventh embodiment and adds the above additional constraints:
[0309] (d1) The prefix APS NAL unit MUST precede the first VCL NAL unit.
[0310] Constraints (a) and (b) of VVC8 still apply. The freedom of the encoder to use prefix APS NAL units or suffix APS NAL units in any mix between the first and last VCL NAL units of a PU is constrained by constraint (d1).
[0311] This not only prevents complex rewrite operations, but also ensures that the decoder does not have to buffer two versions of the APS for decoding slices of a given PU, as explained with respect to the first embodiment.
[0312] Tenth embodiment
[0313] In the tenth embodiment, the encoder allows the suffix APS NAL units and the prefix APS NAL units to have different contents when sharing the same type and identifier APS. In addition, the conforming bitstream requires the following constraints:
[0314] (h1) Within a PU, a VCL NAL unit that references an APS with a specific identifier value and a specific type value shall not be followed by a prefix APS NAL unit containing an APS with these specific values of identifier and type.
[0315] This embodiment makes it possible to provide a prefix APS NAL unit and a suffix APS NAL unit between two VCL NAL units. If the encoder needs to generate a new APS NAL unit with an APS for the next PU, the encoder does not have to buffer the APS for several VCL NAL units.
[0316] Constraint (h1) ensures that two slices of the same PU will not reference different APSs (provided in the prefix APS NAL unit) when the same identifier and type values are used.
[0317] In a variation, the encoder may signal in the SPS whether interleaved APS is allowed using a flag in a parameter set header (such as a PPS or SPS, etc.).
[0318] Eleventh Embodiment
[0319] In an eleventh embodiment, the encoder allows suffix APS NAL units and prefix APS NAL units to have different contents when sharing the same type and identifier APS. In addition, a conforming bitstream requires the following constraints:
[0320] (h2) Within a PU, a VCL NAL unit that references an APS with a specific identifier value and a specific type value shall not be preceded by a suffix APS NAL unit containing an APS with the identifier and type having these specific values.
[0321] This embodiment makes it possible to provide a prefix APS NAL unit and a suffix APS NAL unit between two VCL NAL units. If the encoder needs to generate a new APS NAL unit with an AP for the next PU, the encoder does not have to buffer the APS for several VCL NAL units.
[0322] Constraint (h2) ensures that the suffix APS NAL units are not used for NAL units in a given PU, but only for VCL NAL units of subsequent PUs.
[0323] In a variation, the encoder may signal in the SPS whether interleaved APS is allowed using a flag in a parameter set header (such as a PPS or SPS, etc.).
[0324] Twelfth Embodiment
[0325] In the twelfth embodiment, the encoder allows suffix APS NAL units and prefix APS NAL units to have different content while sharing the same type and identifier APS. In addition, both constraints (h1) and (h2) applied in the tenth and eleventh embodiments, respectively, are necessary for a conforming bitstream.
[0326] This embodiment makes it possible to provide a prefix APS NAL unit and a suffix APS NAL unit between two VCL NAL units. If the encoder needs to generate a new APS NAL unit with an APS for the next PU, the encoder does not have to buffer the APS for several VCL NAL units.
[0327] In a variation, the encoder may signal in the SPS whether interleaved APS is allowed using a flag in a parameter set header (such as a PPS or SPS, etc.).
[0328] Other embodiments of the first group of embodiments
[0329] Some of the measures used in the embodiments of the second group of embodiments are also useful for solving the problems dealt with by the first group of embodiments. Therefore, other embodiments of the first group of embodiments are envisioned as follows. These other embodiments do not need to solve the random access problem and therefore do not involve constraint (e) of the fourth to twelfth embodiments, that is, all APS NAL units within a PU with a specific NAL unit type (nal_unit_type) and a specific value of adaptation_parameter_set_id and a specific value of aps_params_type should have the same content.
[0330] Thirteenth Embodiment
[0331] This embodiment combines the following constraints:
[0332] (h1) within a PU, a VCL NAL unit that references an APS with a specific identifier value and a specific type value shall not be followed by a prefix APS NAL unit containing an APS with those specific identifier and type values; and
[0333] (d2) The suffix APS NAL unit MUST follow the last VCL NAL unit.
[0334] Fourteenth Embodiment
[0335] This embodiment combines the following constraints:
[0336] (h2) within a PU, a VCL NAL unit that references an APS with a specific identifier value and a specific type value shall not be preceded by a suffix APS NAL unit containing an APS with the identifier and type having those specific values; and
[0337] (d1) The prefix APS NAL unit MUST precede the first VCL NAL unit.
[0338] Fifteenth Embodiment
[0339] This embodiment combines the following constraints:
[0340] (h1) within a PU, a VCL NAL unit that references an APS with a specific identifier value and a specific type value shall not be followed by a prefix APS NAL unit containing an APS with those specific identifier and type values; and
[0341] (h2) Within a PU, a VCL NAL unit that references an APS with a specific identifier value and a specific type value shall not be preceded by a suffix APS NAL unit containing an APS with the identifier and type having these specific values.
[0342] In this embodiment, neither constraint (d1) nor constraint (d2) is required.
[0343] Implementation of the embodiment of the present invention
[0344] It should also be understood that according to other embodiments of the present invention, a decoder according to the above-mentioned embodiment / variant is provided in a user terminal such as a computer, a mobile phone (cellular phone), a tablet or any other type of device (e.g., a display device) capable of providing / displaying content to a user. According to yet another embodiment, an encoder according to the above-mentioned embodiment / variant is provided in an image capture device, which also includes a camera, a video camera or a network camera (e.g., a closed-circuit television or video surveillance camera) for capturing and providing content for encoding by the encoder. See below Fig.14 and 15 Two such embodiments are provided.
[0345] Fig.14 is a diagram illustrating a network camera system 1400 including a network camera 1402 and a client device 1404 .
[0346] The network camera 1402 includes an imaging unit 1406, an encoding section 1408, a communication unit 1410, and a control unit 1412. The network camera 1402 and the client device 1404 are connected to each other via the network 300 to be able to communicate with each other.
[0347] The camera unit 1406 includes a lens and an image sensor (e.g., a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS)), and captures an image of an object and generates image data based on the image. The image may be a still image or a video image. The camera unit may also include a zoom component and / or a pan component adapted for zooming or panning (optically or digitally), respectively.
[0348] The encoding unit 1408 encodes the image data by using the encoding method described in one or more of the aforementioned embodiments / variants. The encoding unit 1408 uses at least one of the encoding methods described in the aforementioned embodiments / variants. For other examples, the encoding unit 1408 may use a combination of the encoding methods described in the aforementioned embodiments / variants.
[0349] The communication unit 1410 of the network camera 1402 transmits the encoded image data encoded by the encoding section 1408 to the client device 1404 .
[0350] In addition, the communication unit 1410 may also receive commands from the client device 1404. The commands include commands for setting parameters for encoding by the encoding section 1408.
[0351] The control unit 1412 controls other units in the network camera 1402 according to commands received by the communication unit 1410 or user input.
[0352] The client device 1404 includes a communication unit 1414 , a decoding section 1416 , and a control unit 1418 .
[0353] The communication unit 1414 of the client device 1404 may transmit a command to the web camera 1402. In addition, the communication unit 1414 of the client device 1404 receives the encoded image data from the web camera 1402.
[0354] The decoding section 1416 decodes the encoded image data by using the decoding method described in one or more than one of the aforementioned embodiments / variations. For other examples, the decoding section 1416 may use a combination of the decoding methods described in the aforementioned embodiments / variations.
[0355] The control unit 1418 of the client device 1404 controls other units in the client device 1404 according to the user operation or command received by the communication unit 1414. The control unit 1418 of the client device 1404 may also control the display device 1420 to display the image decoded by the decoding section 1416.
[0356] The control unit 1418 of the client device 1404 also controls the display device 1420 to display a GUI (Graphical User Interface) for specifying values of parameters of the network camera 1402 (e.g., parameters for encoding by the encoding unit 1408). The control unit 1418 of the client device 1404 can also control other units in the client device 1404 according to user operation input to the GUI displayed by the display device 1420.
[0357] The control unit 1418 of the client device 1404 may also control the communication unit 1414 of the client device 1404 according to the user operation input to the GUI displayed by the display device 1420 to transmit a command for specifying the value of the parameter of the network camera 1402 to the network camera 1402 .
[0358] Fig.15 is a diagram illustrating a smartphone 1500 .
[0359] The smartphone 1500 includes a communication unit 1502 , a decoding / encoding section 1504 , a control unit 1506 , and a display unit 1508 .
[0360] The communication unit 1502 receives the encoded image data via the network 9200 .
[0361] The decoding / encoding section 1504 decodes the encoded image data received by the communication unit 1502. The decoding / encoding section 1504 decodes the encoded image data by using the decoding method described in one or more than one of the aforementioned embodiments / variants. The decoding / encoding section 1504 may also use at least one of the encoding or decoding methods described in the aforementioned embodiments / variants. For other examples, the decoding / encoding section 1504 may use a combination of the decoding or encoding methods described in the aforementioned embodiments / variants.
[0362] The control unit 1506 controls other units in the smartphone 1500 according to user operations or commands received by the communication unit 1502. For example, the control unit 1506 controls the display unit 1508 to display an image decoded by the decoding / encoding section 1504.
[0363] The smartphone may also include an image recording device 1510 (eg, a digital camera and associated circuitry) for recording images or videos. Such recorded images or videos may be encoded by the decoding / encoding section 1504 under the instruction of the control unit 1506.
[0364] The smartphone may also include a sensor 1512 adapted to sense the orientation of the mobile device. Such a sensor may include an accelerometer, a gyroscope, a compass, a global positioning (GPS) unit, or a similar position sensor. Such a sensor 1512 may determine whether the smartphone changes orientation, and such information may be used when encoding the video stream.
[0365] Although the present invention has been described with reference to embodiments and variations thereof, it should be understood that the present invention is not limited to the disclosed embodiments / variations. It will be appreciated by those skilled in the art that various changes and modifications may be made without departing from the scope of the present invention as defined by the appended claims. All features disclosed in this specification (including any appended claims, abstracts and drawings), and / or all steps of any disclosed method or process, may be combined in any combination, except for at least some mutually exclusive combinations of such features and / or steps. Unless expressly stated otherwise, each feature disclosed in this specification (including any appended claims, abstracts and drawings) may be replaced by alternative features for the same, equivalent or similar purposes. Therefore, unless expressly stated otherwise, each feature disclosed is only an example of a general series of equivalent or similar features.
[0366] It should also be understood that any result of the above-mentioned comparison, determination, inference, evaluation, selection, execution, conduct or consideration (e.g., a selection made during encoding, processing or partition processing) can be indicated in data in the bitstream (e.g., a flag or information indicating the result) or can be determined / inferred from data in the bitstream, so that the indicated or determined / inferred result can be used in the processing instead of actually performing the comparison, determination, evaluation, selection, execution, conduct or consideration, such as during decoding or partition processing. It should be understood that when a "table" or "lookup table" is used, other data types such as arrays can also be used to perform the same function, as long as the data type can perform the same function (e.g., representing the relationship / mapping between different elements).
[0367] In the claims, the word "comprising" does not exclude other elements or steps and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage. Reference signs appearing in the claims are by way of illustration only and shall not have a limiting effect on the scope of the claims.
[0368] In the foregoing embodiments / variations, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or sent through a computer-readable medium as one or more instructions or codes, and may be executed by a hardware-based processing unit.
[0369] Computer-readable media may include computer-readable storage media, which corresponds to tangible media such as data storage media or communication media including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for implementing the techniques described in the present invention. A computer program product may include a computer-readable medium.
[0370] As an example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, flash memory or any other medium that can be used to store the desired program code in the form of an instruction or data structure and can be accessed by a computer. In addition, any connection can be appropriately referred to as a computer-readable medium. For example, if a coaxial cable, optical fiber cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave, etc.) is used to send instructions from a website, server or other remote source, then the coaxial cable, optical fiber cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave, etc.) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connection, carrier wave, signal or other transient media, but are directed to non-transient tangible storage media. The disk and disc used here include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and blue disc, wherein the disc usually copies data magnetically, and the disc reproduces data optically by laser. Combinations of the above should also be included within the scope of computer-readable media.
[0371] Instructions may be executed by one or more processors such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate / logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Furthermore, the techniques may be fully implemented in one or more circuits or logic elements.
[0372] Any step of the method / processing according to the present invention or the function described herein may be implemented in hardware, software, firmware or any combination thereof. If implemented in software, the step / function may be stored as one or more instructions or codes or programs or computer-readable media on one or more hardware-based processing units or sent via one or more hardware-based processing units, and executed by one or more hardware-based processing units, such as a programmable computing machine, which may be a PC ("personal computer"), a DSP ("digital signal processor"), a circuit, a circuit system, a processor and memory, a general-purpose microprocessor or central processing unit, a microcontroller, an ASIC ("application-specific integrated circuit"), a field programmable logic array (FPGA) or other equivalent integrated or discrete logic circuit system. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures or any other structures suitable for implementing the technology described herein.
[0373] Embodiments of the present invention may also be implemented by various devices or apparatuses, including wireless handsets, integrated circuits (ICs), or JC collections (e.g., chipsets). Various components, modules, or units are described herein to illustrate functional aspects of the apparatus / device configured to perform these embodiments, but do not necessarily need to be implemented by different hardware units. Instead, the various modules / units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units, which include one or more processors in conjunction with appropriate software / firmware.
[0374] Embodiments of the present invention can be implemented by a computer of a system or device that reads out and executes computer executable instructions (e.g., one or more programs) recorded on a storage medium to perform one or more modules / units / functions in the above-mentioned embodiments and / or includes one or more processing units or circuits for performing one or more functions in the above-mentioned embodiments, and can be implemented by a method performed by a computer of a system or device, for example, reading out and executing computer executable instructions from a storage medium to perform one or more functions in the above-mentioned embodiments and / or controlling one or more processing units or circuits to perform one or more functions in the above-mentioned embodiments. The computer may include a network of separate computers or separate processing units to read out and execute computer executable instructions. Computer executable instructions may be provided to a computer from a computer-readable medium such as a communication medium, for example, via a network or a tangible storage medium. The communication medium may be a signal / bit stream / carrier. A tangible storage medium is a "non-transitory computer-readable storage medium" which may include, for example, a hard disk, a random access memory (RAM), a read-only memory (ROM), a storage device of a distributed computing system, an optical disk (such as a compact disk (CD), a digital versatile disk (DVD), or a Blu-ray disk (BD) TM ), one or more of a flash memory device, a memory card, etc. At least some steps / functions may also be implemented in hardware by a machine or dedicated components such as an FPGA (“field programmable gate array”) or an ASIC (“application-specific integrated circuit”).
Claims
1. A method for encoding a sequence of images in a bit stream, include: A series of picture units is provided in the bitstream, each of the picture units corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, i.e., APS NAL units, each of the VCL NAL units contains coded image data, each of the APS NAL units contains an adaptation parameter set having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS is present in a picture unit before the first VCL NAL unit of the picture unit, the APS must be contained in a prefix APS NAL unit, and in the case where the APS is present in the picture unit after the last VCL NAL unit of the picture unit, the APS must be contained in a suffix APS NAL unit, the APS NAL units each having an APS type and an APS identifier; wherein all APS NAL units with a prefix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content, and all APS NAL units with a suffix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content; and The license includes, in the same picture unit, a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents.
2. The encoding method according to claim 1, in, In case a prefix APS NAL unit is present in a picture unit, the prefix APS NAL unit shall precede the first VCL NAL unit of the picture unit, and in case a suffix APS NAL unit is present in a picture unit, the suffix APS NAL unit shall follow the last VCL NAL unit of the picture unit.
3. A method for decoding a coded image sequence, include: receiving a bitstream having a series of picture units, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, i.e., APS NAL units, each of which contains coded image data, each of which contains an adaptation parameter set having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in case an APS is present in a picture unit before the first VCL NAL unit of the picture unit, the APS must be contained in a prefix APS NAL unit, and in case an APS is present in the picture unit after the last VCL NAL unit of the picture unit, the APS must be contained in a suffix APS NAL unit, the APS NAL units each having an APS type and an APS identifier, wherein all APS NAL units with a prefix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content, and all APS NAL units with a suffix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content, and The license includes, in the same picture unit, a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents.
4. The method for decoding according to claim 3, in, In case a prefix APS NAL unit is present in a picture unit, the prefix APS NAL unit shall precede the first VCL NAL unit of the picture unit, and in case a suffix APS NAL unit is present in a picture unit, the suffix APS NAL unit shall follow the last VCL NAL unit of the picture unit.
5. A device for encoding a sequence of images in a bit stream, include: means for providing in the bitstream a series of picture units, each of which corresponds to a coded image and comprises one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units comprise video coding layer NAL units, i.e., VCL NAL units, and also comprise adaptation parameter set NAL units, i.e., APS NAL units, each of which contains coded image data, each of which contains an adaptation parameter set having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units comprise prefix APS NAL units and suffix APS NAL units, wherein, in the case where the APS is present in a picture unit before the first VCL NAL unit of the picture unit, the APS must be contained in a prefix APS NAL unit, and in the case where the APS is present in the picture unit after the last VCL NAL unit of the picture unit, the APS must be contained in a suffix APS NAL unit, the APS NAL units each having an APS type and an APS identifier; wherein all APS NAL units with a prefix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content, and all APS NAL units with a suffix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content; and The license includes, in the same picture unit, a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents.
6. A device for decoding a coded image sequence, include: means for receiving a bitstream having a series of picture units, each of which corresponds to a coded image and includes one or more network abstraction layer units, i.e., one or more NAL units, the NAL units that can be included in the series of picture units include video coding layer NAL units, i.e., VCL NAL units, and also include adaptation parameter set NAL units, i.e., APS NAL units, each of which contains coded image data, each of which contains an adaptation parameter set having parameters for performing one or more types of processing operations on the image data contained in the one or more VCL NAL units, and the APS NAL units that can be included in the series of picture units include prefix APS NAL units and suffix APS NAL units, wherein, in case an APS is present in a picture unit before the first VCL NAL unit of the picture unit, the APS must be contained in a prefix APS NAL unit, and in case an APS is present in the picture unit after the last VCL NAL unit of the picture unit, the APS must be contained in a suffix APS NAL unit, the APS NAL units each having an APS type and an APS identifier, wherein all APS NAL units with a prefix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content, and all APS NAL units with a suffix NAL unit type and a specific APS identifier and a specific APS type within a given picture unit have the same content, and The license includes, in the same picture unit, a prefix APS NAL unit and a suffix APS NAL unit having the same APS type and the same APS identifier but different contents.
7. The apparatus for decoding according to claim 6, wherein, when a prefix APS NAL unit exists in a picture unit, the prefix APS NAL unit should precede the first VCL NAL unit of the picture unit, and when a suffix APS NAL unit exists in a picture unit, the suffix APS NAL unit should follow the last VCL NAL unit of the picture unit.
8. A computer program product comprising computer program instructions which, when executed by a processor or a computer, cause the processor or the computer to perform the method according to any one of claims 1 to 4.
9. A computer-readable storage medium storing computer program instructions which, when executed by a processor or a computer, cause the processor or the computer to perform the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Coding SEI NAL units for video coding
CN104412600A
Method and apparatus for enclosing and decoding a video bitstream for merging regions of interest
GB201904461D0