Method and apparatus for carrying text descriptions within video coding

The apparatus and method facilitate the inclusion of text descriptions within video coding by defining and signaling information messages, addressing the lack of efficient content management in existing technologies, thereby improving video decoding and identification.

WO2025219777A1PCT designated stage Publication Date: 2025-10-23NOKIA TECHNOLOGIES OY

Patent Information

Application Number
PCT/IB2025/052850
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2025-03-18
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing multimedia coding technologies lack efficient methods for carrying text descriptions within video coding, such as encoder information, which is essential for identifying and managing video content.

Method used

An apparatus and method are introduced to define and signal information messages within coded video, including text descriptions like encoder name, version, and URI, using supplemental enhancement information (SEI) messages, with persistence indicators and categorized purposes for coded or decoded pictures.

Benefits of technology

Enables effective management and identification of video content by incorporating text descriptions, ensuring persistence and categorization, enhancing the decoding process and content understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000027_0001
    Figure IMGF000027_0001
  • Figure IMGF000034_0001
    Figure IMGF000034_0001
  • Figure IMGF000035_0001
    Figure IMGF000035_0001
Patent Text Reader

Abstract

Various embodiments provide methods, apparatuses, and computer program products. An example apparatus includes at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; and signaling the information message to a decoder.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR CARRYING TEXT DESCRIPTIONS WITHIN VIDEO CODINGTECHNICAL FIELD

[0001] The examples and non-limiting embodiments relate generally to multimedia coding and, more particularly to, carrying text descriptions within video coding.BACKGROUND

[0002] It is known to provide standardized formats for encoding, signaling, or decoding of media data.SUMMARY

[0003] Example 1: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; and signaling the information message to a decoder.

[0004] Example 2: The apparatus of example 1, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0005] Example 3: The apparatus of example 1, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

[0006] Example 4 : The apparatus of example 1 , wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence.

[0007] Example 5 : The apparatus of example 4, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence untilcancelled.

[0008] Example 6: The apparatus of any of the examples 4 or 5, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

[0009] Example ?: The apparatus of example 2, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0010] Example 8: The apparatus of example 3, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

[0011] Example 9: The apparatus of any of the previous examples, wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

[0012] Example 10 : The apparatus of example 9, wherein for the one or more purposes applying to the coded pictures, and wherein the apparatus is further caused to perform: defining a persistence of the information message in terms of coded pictures in decoding order.

[0013] Example 11 : The apparatus of example 9, wherein for the one or more purposes applying to the decoded pictures, and wherein the apparatus is further caused to perform: defining a persistence of the information message in terms of decoded pictures in output order.

[0014] Example 12: The apparatus of any of the example 1 to 8, wherein the one or more purposes comprise a pre-defined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

[0015] Example 13: The apparatus of example 12, whereinthe pre-defined set of purpose comprises the tag URI and an encoder description.

[0016] Example 14: The apparatus of example 13, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the apparatus is further caused to perform: encodingthe information message with persistence flag equal to a first value and refrains from cancelling the information message with another information message with the same purpose.

[0017] Example 15: The apparatus of any of the previous examples wherein the information message comprises a supplemental enhancement information (SEI) message.

[0018] Example 16: The apparatus of example 15, wherein the SEI message is provided out-of- band in a manner that a persistence scope of the SEI message comprises a bitstream.

[0019] Example 17: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: defining an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; and signaling the information message to a decoder.

[0020] Example 18: The apparatus of example 17, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistent scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

[0021] Example 19: The apparatus of any of the examples 17 or 18, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

[0022] Example 20: The apparatus of any of the examples 17 to 19, wherein the information message further comprises text descriptions, and wherein the text descriptions comprise one or more purposes, and wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content.

[0023] Example 21: The apparatus of example 20, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0024] Example 22: The apparatus of example 20, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

[0025] Example 23 : An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: signaling at most one language tag per information message; signaling one or more text descriptions comprised in the information message; and signaling a purpose and a text string for each text description of the one or more text descriptions.

[0026] Example 24 : The apparatus of example 23 , wherein the text string comprises one or more of the following: an identifier string representing a predefined coding configuration of an encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0027] Example 25 : An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, in or from, a bitstream an information message carrying text descriptions in a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; decoding images and / or a video sequence into at least one decoded picture; and associating the information message contents with the decoded picture.

[0028] Example 26: The apparatus of example 25, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0029] Example 27: The apparatus of example 25, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

[0030] Example 28: The apparatus of example 25, wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence.

[0031] Example 29: The apparatus of example 28, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

[0032] Example 30: The apparatus of any of the examples 28 or 29, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

[0033] Example 31 : The apparatus of example 26, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0034] Example 32: The apparatus of example 27, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

[0035] Example 33 : The apparatus of any of the examples 25 to 32, wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

[0036] Example 34: The apparatus of example 33, wherein for the one or more purposes applying to the coded pictures, and wherein the apparatus is further caused to perform: receiving a persistence of the information message in terms of coded pictures in decoding order.

[0037] Example 35: The apparatus of example 33, wherein for the one or more purposes applying to the decoded pictures, and wherein the apparatus is further caused to perform: receiving a persistence of the information message in terms of decoded pictures in output order.

[0038] Example 36: The apparatus of any of the example 25 to 32, wherein the one or more purposes comprise a pre-defined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

[0039] Example 37: The apparatus of example 36, wherein the pre-defined set of purpose comprises the tag URI and an encoder description.

[0040] Example 38: The apparatus of example 37, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the apparatus is further caused to perform: receiving the information message encoded with persistence flag equal to a first value and the information message with another information message with the same purpose.

[0041] Example 39: The apparatus of any of the examples 25 to 38, wherein the information message comprises a supplemental enhancement information (SEI) message.

[0042] Example 40: The apparatus of example 39, wherein the SEI message is provided out-of- band in a manner that a persistence scope of the SEI message comprises a bitstream.

[0043] Example 41 : An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, in or from, a bitstream an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; decoding the CLVS or coded video sequence into decoded pictures; and associating the information message with decoded pictures of the CLVS or coded video sequence.

[0044] Example 42: The apparatus of example 41, wherein the association of the information with decoded pictures is performed without checking every picture in the CLVS or coded video sequence for presence of the information message of same type.

[0045] Example 43 : An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving at most one language tag per information message; receiving one or more text descriptions comprised in the information message; receiving a purpose and a text string for each text description of the one or more text descriptions; decoding images and / or a video sequence into at least one decoded picture; and associating the purpose and the text string with the at least one decoded picture.

[0046] Example 44: A method comprising: defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; and signaling the information message to a decoder.

[0047] Example 45 : The method of example 44, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0048] Example 46 : The method of example 44, wherein the tag URI comprises at least an authorityname, a date stamp, and / or a time stamp.

[0049] Example 47 : The method of example 44, wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence.

[0050] Example 48 : The method of example 47, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

[0051] Example 49: The method of any of the examples 47 or 48, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

[0052] Example 50: The method of example 45, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0053] Example 51 : The method of example 46, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

[0054] Example 52 : The method of any of the examples 44 to 51 , wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

[0055] Example 53 : The method of example 52, wherein for the one or more purposes applying to the coded pictures, and wherein the method is further caused to perform: defining a persistence of the information message in terms of coded pictures in decoding order.

[0056] Example 54: The method of example 52, wherein for the one or more purposes applying to the decoded pictures, and wherein the method is further caused to perform: defining a persistence of the information message in terms of decoded pictures in output order.

[0057] Example 55: The method of any of the example 44 to 51, wherein the one or more purposes comprise a pre-defined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

[0058] Example 56: The method of example 55, wherein the pre-defined set of purpose comprises the tag URI and an encoder description.

[0059] Example 57: The method of example 56, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the method is further caused to perform: encoding the information message with persistence flag equal to a first value and refrains from cancelling the information message with another information message with the same purpose.

[0060] Example 58: The method of any of the examples 44 to 57, wherein the information message comprises a supplemental enhancement information (SEI) message.

[0061] Example 59: The method of example 58, wherein the SEI message is provided out-of-band in a manner that a persistence scope of the SEI message comprises a bitstream.

[0062] Example 60: A method comprising: defining an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; and signaling the information message to a decoder.

[0063] Example 61 : The method of example 60, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistent scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

[0064] Example 62 : The method of any of the examples 60 or 61 , wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

[0065] Example 63 : The method of any of the examples 60 to 62, wherein the information message further comprises text descriptions, and wherein the text descriptions comprise one or more purposes, and wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content.

[0066] Example 64: The method of example 63, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0067] Example 65 : The method of example 63 , wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

[0068] Example 66: A method comprising: signaling at most one language tag per information message; signaling one or more text descriptions comprised in the information message; and signaling a purpose and a text string for each text description of the one or more text descriptions.

[0069] Example 67: The method of example 66, wherein the text string comprises one or more of the following: an identifier string representing a predefined coding configuration of an encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0070] Example 68: A method comprising: receiving, in or from, a bitstream an information message carrying text descriptions in a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; decoding images and / or a video sequence into at least one decoded picture; and associating the information message contents with the decoded picture.

[0071] Example 69: The method of example 68, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0072] Example 70 : The method of example 68, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

[0073] Example 71: The method of example 68, wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence.

[0074] Example 72 : The method of example 71 , wherein the persistence indicator uses one or moresyntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

[0075] Example 73 : The method of any of the examples 71 or 72, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

[0076] Example 74: The method of example 69, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0077] Example 75: The method of example 70, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

[0078] Example 76: The method of any of the examples 68 to 75, wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

[0079] Example 77: The method of example 76, wherein for the one or more purposes applying to the coded pictures, and wherein the method further comprises: receiving a persistence of the information message in terms of coded pictures in decoding order.

[0080] Example 78: The method of example 76, wherein for the one or more purposes applying to the decoded pictures, and wherein the method further comprises: receiving a persistence of the information message in terms of decoded pictures in output order.

[0081] Example 79: The method of any of the example 68 to 75, wherein the one or more purposes comprise a pre-defined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

[0082] Example 80: The method of example 79, wherein the pre-defined set of purpose comprises the tag URI and an encoder description.

[0083] Example 81: The method of example 80, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the method further comprises: receiving the information message encoded with persistence flag equal to a first value and the information message with another information message with the same purpose.

[0084] Example 82: The method of any of the examples 68 to 81, wherein the information message comprises a supplemental enhancement information (SEI) message.

[0085] Example 83: The method of example 82, wherein the SEI message is provided out-of-band in a manner that a persistence scope of the SEI message comprises a bitstream.

[0086] Example 84: A method comprising: receiving, in or from, a bitstream an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; decoding the CLVS or coded video sequence into decoded pictures; and associating the information message with decoded pictures of the CLVS or coded video sequence.

[0087] Example 85 : The method of example 84, wherein the association of the information with decoded pictures is performed without checking every picture in the CLVS or coded video sequence for presence of the information message of same type.

[0088] Example 86: A method comprising: receiving at most one language tag per information message; receiving one or more text descriptions comprised in the information message; receiving a purpose and a text string for each text description of the one or more text descriptions; decoding images and / or a video sequence into at least one decoded picture; and associating the purpose and the text string with the at least one decoded picture.

[0089] Example 87 : An apparatus comprising: means for defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; and means for signaling the information message to a decoder.

[0090] Example 88: The apparatus of example 87, wherein the apparatus comprises means for performing methods as described in any of the examples 45 to 59.

[0091] Example 89: An apparatus comprises: means for defining an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence(CLVS) or an entire coded video sequence; and means for signaling the information message to a decoder.

[0092] Example 90: The apparatus of example 89, wherein the apparatus further comprises means for performing methods as described in any of the examples 61 to 65.

[0093] Example 91: An apparatus comprising: means for signaling at most one language tag per information message; means for signaling one or more text descriptions comprised in the information message; and means for signaling a purpose and a text string for each text description of the one or more text descriptions.

[0094] Example 92: The apparatus of example 91, wherein the text string comprises: an identifier string representing a predefined coding configuration of an encoder; a preset string representing a predefined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0095] Example 93: An apparatus comprising: means for receiving, in or from, a bitstream an information message carrying text descriptions in a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; means for decoding images and / or a video sequence into at least one decoded picture; and means for associating the information message contents with the decoded picture.

[0096] Example 94: The apparatus of example 93, wherein the apparatus comprises means for performing methods as described in any of the examples 69 to 83.

[0097] Example 95: An apparatus comprising: means for receiving, in or from, a bitstream an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; means for decoding the CLVS or coded video sequence into decoded pictures; and means for associating the information message with decoded pictures of the CLVS or coded video sequence.

[0098] Example 96: The apparatus of example 95, wherein the association of the information with decoded pictures is performed without checking every picture in the CLVS or coded video sequence for presence of the information message of same type.

[0099] Example 97: An apparatus comprising: means for receiving at most one language tag per information message; means for receiving one or more text descriptions comprised in the informationmessage; means for receiving a purpose and a text string for each text description of the one or more text descriptions; means for decoding images and / or a video sequence into at least one decoded picture; and means for associating the purpose and the text string with the at least one decoded picture.

[0100] Example 98: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 44 to 59.

[0101] Example 99: The computer readable medium of example 98, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0102] Example 100: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 60 to 65.

[0103] Example 101: The computer readable medium of example 100, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0104] Example 102: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 66 or 67.

[0105] Example 103: The computer readable medium of example 102, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0106] Example 104: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 68 to 83.

[0107] Example 105: The computer readable medium of example 104, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0108] Example 106: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 83 or 84.

[0109] Example 107: The computer readable medium of example 106, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0110] Example 108: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as described in example 86.

[0111] Example 109: The computer readable medium of example 108, wherein the computer readable medium comprises a non-transitory computer readable medium.BRIEF DESCRIPTION OF THE DRAWINGS

[0112] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0113] FIG. 1 shows schematically an apparatus employing embodiments of the examples described herein.

[0114] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.

[0115] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.

[0116] FIG. 4 is a block diagram illustrating a system in accordance with an example.

[0117] FIG. 5 is an example apparatus, which may be implemented in hardware, and is caused to, implement examples described herein.

[0118] FIG. 6 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.

[0119] FIG. 7 is an example method performed with an encoder, based on the examples described herein.

[0120] FIG. 8 is another example method performed with an encoder, based on the examples described herein.

[0121] FIG. 9 is yet another example method performed with an encoder, based on the examples described herein.

[0122] FIG. 10 is an example method performed with a decoder, based on the examples described herein.

[0123] FIG. 11 is another example method performed with a decoder, based on the examples described herein.

[0124] FIG. 12 is yet an example method performed with a decoder, based on the examples described herein.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS

[0125] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows (the abbreviations may be appended with each other or with other characters using e.g. a hyphen or dash (-), and may be case insensitive):4CC four character code5G fifth generation cellular network technology5GC 5G core network a.k.a. also known asAVC advanced video codingCU coding unitDSP digital signal processorDU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DCE-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technologyFl orFl-C interface between CU and DU control interface gNB (or gNodeB) base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCIEC International Electrotechnical Commission loT internet of thingsISO International Organization for StandardizationISOBMFF ISO base media file formatJPEG joint photographic experts groupLTE long-term evolution mdat MediaDataBoxMIME Multipurpose Internet Mail ExtensionMME mobility management entity moov MovieBoxMP4 file format for MPEG-4 Part 14 filesMPEG moving picture experts groupMPEG-2 H.222 / H.262 as defined by the ITUMPEG-4 audio and video coding standard for ISO / IEC 14496 ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radio (5G radio)N / W or NW networkPDCP packet data convergence protocolPHY physical layerPNG portable network graphicsRAN radio access networkRFC request for commentsRLC radio link controlRRC radio resource controlRRH remote radio headRU radio unitRx receiverSDAP service data adaptation protocolSGW serving gatewaySMF session management functionSPS sequence parameter setSVC scalable video codingSI interface between eNodeBs and the EPC trak TrackBoxTx transmitterUE user equipmentUICC Universal Integrated Circuit CardUPF user plane functionURL uniform resource locatorX2 interconnecting interface between two eNodeBs in LTE networkXn interface between two NG-RAN nodes

[0126] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments may be shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments.

[0127] Described herein is a method and apparatus for carrying text descriptions within video coding.

[0128] The following describes in detail a suitable apparatus and possible method for carrying text descriptions within video coding according to embodiments. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an electronic device or apparatus 100. The apparatus 100 may be an Internet of Things (loT) apparatus configured to perform various functions, such as for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 are explained next.

[0129] The apparatus 100 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.

[0130] The apparatus 100 may comprise a housing 101 for incorporating and protecting the device. The apparatus 100 further may comprise a display 102 in the form of a liquid crystal display. In other embodiments of the examples described herein the display may be any suitable display technologysuitable to display an image or video. The apparatus 100 may further comprise a keypad 104. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.

[0131] The apparatus may comprise a microphone 106 or any suitable audio input which may be a digital or analog signal input. The apparatus 100 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 108, speaker, or an analog audio or digital audio output connection. The apparatus 100 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus 100 may further comprise a camera 109 capable of recording or capturing images and / or video. The apparatus 100 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 100 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.

[0132] The apparatus 100 may comprise a controller 110, processor or processor circuitry for controlling the apparatus 100. The controller 110 may be connected to memory 112 which in embodiments of the examples described herein may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 110. The controller 110 may further be connected to codec circuitry 114 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.

[0133] The apparatus 100 may further comprise a card reader 118 and a smart card 116, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.

[0134] The apparatus 100 may comprise radio interface circuitry 120 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 100 may further comprise an antenna 122 connected to the radio interface circuitry 120 for transmitting radio frequency signals generated at the radio interface circuitry 120 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).

[0135] The apparatus 100 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec circuitry 114 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / orstorage. The apparatus 100 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 100 described above represent examples of means for performing a corresponding function.

[0136] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 300 comprises multiple communication devices which can communicate through one or more networks. The system 300 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network, etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.

[0137] The system 300 may include both wired and wireless communication devices and / or apparatus 100 suitable for implementing embodiments of the examples described herein.

[0138] For example, the system shown in FIG. 3 shows a mobile telephone network 301 and a representation of the internet 302. Connectivity to the internet 302 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0139] The example communication devices shown in the system 300 may include, but are not limited to, an electronic device or apparatus 100, a combination of a personal digital assistant (PDA) and a mobile telephone 304, a PDA 306, an integrated messaging device (IMD) 308, a desktop computer 310, a notebook computer 312, or a head-mounted apparatus. The head-mounted apparatus may be a head-mounted display (HMD), or glasses having a device such as a camera configured to encode and / or decode images and / or video. The apparatus 100 may be stationary or mobile when carried by an individual who is moving. The apparatus 100 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.

[0140] The embodiments may also be implemented in a set-top box; e.g., a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.

[0141] Some or further apparatus may send and receive calls and messages and communicate withservice providers through a wireless connection 314 to a base station 316. The base station 316 may be connected to a network server 318 that allows communication between the mobile telephone network 301 and the internet 302. The system may include additional communication devices and communication devices of various types.

[0142] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0143] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.

[0144] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP -based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).

[0145] FIG. 4 is a block diagram illustrating a system or apparatus 400 in accordance with several examples. In an example, the encoder 402 is used to encode an image or video from the scene 404, and the encoder 402 is implemented in a transmitting apparatus 406. The encoder 402 produces a bitstream 408 comprising signaling that is received by the receiving apparatus 410, which implements a decoder 412. The encoder 402 sends the bitstream 408 that comprises the herein described signaling. Thedecoder 412 forms the image or video for the scene 404-1, and the receiving apparatus 410 would present this to the user, e.g., via a smartphone, television, or projector among many other options.

[0146] In some examples, the transmitting apparatus 406 and the receiving apparatus 410 are at least partially within a common apparatus, and for example, are located within a common housing 414. In other examples the transmitting apparatus 406 and the receiving apparatus 410 are at least partially not within a common apparatus and have at least partially different housings. Therefore in some examples, the encoder 402 and the decoder 412 are at least partially within a common apparatus, and for example are located within a common housing 414. For example, the common apparatus comprising the encoder 402 and the decoder 412 implements a codec. In other examples, the encoder 402 and the decoder 412 are at least partially not within a common apparatus and have at least partially different housings, but when together still implement a codec.

[0147] In some examples, 3D media from the capture (e.g., volumetric capture) at a viewpoint 416 of the scene 404, which includes a person 418) is converted via projection to a series of 2D representations with occupancy, geometry, attributes and / or displacements. Additional atlas information is also included in the bitstream to enable inverse reconstruction. For decoding, the received bitstream 408 is separated into its components with atlas information; occupancy, geometry, displacement, and attribute 2D representations. A 3D reconstruction is performed to reconstruct the scene 404-1 created looking at the viewpoint 416-1 with a “reconstructed” person 418-1. The “-1” are used to indicate that these are reconstructions of the original. As indicated at 420, the decoder 412 performs an operation(s) or action(s) based on the received signaling.

[0148] Encoding 422 performs encoding text description within a video bitstream based on the examples described herein. Decoding 424 performs decoding text description from the video bitstream, based on the examples described herein.

[0149] The Advanced Video Coding standard (which may be abbreviated H.264, AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, each integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) andMultiview Video Coding (MVC).

[0150] The High Efficiency Video Coding standard (which may be abbreviated H.265, HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D- HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV- HEVC, 3D-HEVC and REXT that have been made for the purpose of understanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.

[0151] Versatile Video Coding (which may be abbreviated WC, H.266, or H.266 / WC) is a video compression standard developed as the successor to HEVC. WC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3.

[0152] A specification of the AVI bitstream format and decoding process were developed by the Alliance for Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.

[0153] ITU-T Recommendation H.274, which is equivalent to ISO / IEC 23002-7, may be called "versatile supplemental enhancement information messages for coded video bitstreams" and be referred to as "versatile supplemental enhancement information" or VSEI. The VSEI standard specifies the syntax and semantics of video usability information (VUI) parameters and supplemental enhancement information (SEI) messages. The VUI parameters and SEI messages defined in the VSEI standard are designed to be conveyed within coded video bitstreams in a manner specified in a video coding specification or to be conveyed by other means determined by the specifications for systems that make use of such coded video bitstreams. The VSEI standard is intended for use with WC coded video bitstreams, although it is drafted in a manner intended to be sufficiently generic that it may also be used with other types of coded video bitstreams. VUI parameters and SEI messages may, for example, assist in processes related to decoding, display or other purposes.

[0154] Having thus introduced a suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described in detail.

[0155] Features as described herein may generally relate, for example, to the ISO base media file format (ISOBMFF).

[0156] An out-of-band transmission, signaling, or storage may refer to the capability of transmitting, signaling, or storing information in a manner that associates the information with a video bitstream. The out-of-band transmission may use a more reliable transmission mechanism compared to the protocols used for carrying coded video data, such as slices. The out-of-band transmission, signaling or storage may additionally or alternatively be used, e.g., for ease of access or session negotiation. For example, a sample entry of a track in a file conforming to the ISO Base Media File Format may comprise parameter sets, while the coded data in the bitstream is stored elsewhere in the file or in another file.

[0157] A basic building block in the ISO base media file format is called a box. Each box has a header and a payload. The box header indicates the type of the box and the size of the box in terms of bytes. Box type is typically identified by an unsigned 32-bit integer, interpreted as a four-character code (4CC). A box may enclose other boxes, and the ISO fde format specifies which box types are allowed within a box of a certain type. Furthermore, the presence of some boxes may be mandatory in each file, while the presence of other boxes may be optional. Additionally, for some box types, it may be allowable to have more than one box present in a file. Thus, the ISO base media file format may be considered to specify a hierarchical structure of boxes.

[0158] In files conforming to the ISO base media file format, the media data may be provided in one or more instances of MediaDataBox (‘mdat‘), and the MovieBox (‘moov’) may be used to enclose the metadata for timed media. In some cases, for a file to be operable, both of the ‘mdat’ and ‘moov’ boxes may be required to be present. The ‘moov’ box may include one or more tracks, and each track may reside in one corresponding TrackBox (‘trak’). Each track is associated with a handler, identified by a four-character code, specifying the track type. Video, audio, and image sequence tracks can be collectively called media tracks, and they contain an elementary media stream. Other track types comprise hint tracks and timed metadata tracks.

[0159] Tracks comprise samples, such as audio or video frames. For video tracks, a media sample may correspond to a coded picture or an access unit.

[0160] A media track refers to samples (which may also be referred to as media samples) formatted according to a media compression format (and its encapsulation to the ISO base media file format). A hint track refers to hint samples, containing cookbook instructions for constructing packets for transmission over an indicated communication protocol. A timed metadata track may refer to samples describing referred media and / or hint samples.

[0161] The 'Irak' box includes in its hierarchy of boxes the SampleDescriptionBox, which gives detailed information about the coding type used, and any initialization information needed for that coding. The SampleDescriptionBox contains an entry -count and as many sample entries as the entrycount indicates. The format of sample entries is track-type specific, but derived from generic classes (e.g. VisualSampleEntry, AudioSampleEntry). Which type of sample entry form is used for derivation of the track-type specific sample entry format is determined by the media handler of the track.

[0162] A sample entry may comprise a configuration box, which itself may comprise a configuration box, which may comprise a configuration record structure. The configuration record may comprise information that may be used to configure a decoder instance for decoding the samples mapped to the sample entry. The configuration record of some video codecs also includes parameter sets and SEI messages, which apply to the samples that reference the sample entry. The file format structures indicate for each sample which of sample entries applies.

[0163] Supplemental enhancement information (SEI) messages include metadata within a bitstream, synchronized with the coded video, to convey extra information intended to be utilized by the receiver / decoder for a variety of use cases.

[0164] Video coding standards specify the bitstream format and decoding operation, but provide significant flexibility to the video encoders that generate bitstreams conforming to the standard.

[0165] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.

[0166] A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.

[0167] In some coding formats, a coded video sequence (CVS) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.

[0168] In some coding formats, a bitstream may comprise a sequence of one or more CVSs.

[0169] A coded layer video sequence (CLVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh layer id in WC) that is decodable independently of other pictures in the same layer.

[0170] Joint video experts team (JVET) develops a reference software encoder and decoder for draft standards and existing standards which are regularly updated with improvements and released as new numbered versions. Decoder product implementations for finalized standards are expected to conform to one or more profiles defined in the standard specification and to interoperate with any future video encoder implementations that conform to a supported profile.

[0171] While a new standard is under development, new versions of the draft standard and their corresponding reference software implementations, such as the enhanced compression model (ECM) for the future H.267 standard, include significant changes to syntax and semantics which make bitstreams created by earlier versions of the encoder unable to be properly decoded by later versions of the decoder.

[0172] JVET defines common test conditions (described in JVET-AB2010 “VTM and HM common test conditions and software reference configurations for SDR 4:2:0 10 bit video” and other documents) used to configure the reference software encoder for experiments used in contributions to the committee. The configurations include Random Access (RA), All Intra (Al), Low Delay P (LDP), and Low Delay B (LDB), which represent different picture coding orders and allowable reference pictures.

[0173] Fraunhofer Versatile Video Encoder (WenC) is an open source software implementation of a versatile video coding (WC) encoder, available from [https: / / github.com / fraunhoferhhi / wenc, (last accessed on April 13, 2024)].

[0174] IETF RFC 4151 defines a mechanism for signaling a tag uniform resource indicator (URI), which includes an authority name and a time / date stamp and optionally contains additional specification.

[0175] Abstracts from the two JVET contributions (JVET-AH0052 AHG9: Text Description Information SEI Message and JVET-AH0172 AHG9: Text comment SEI message) are provided below. Both contributions propose an SEI message to carry text within the coded video bitstream.

[0176] Contribution JVET-AH0052 Text Description SEI

[0177] A Text Description Information SEI message is proposed, ft is asserted that a common SEI message which provides text information for various roles (purposes) avoids the need to define multiple SEI messages which provide specific type of text information about video. The proposed message is extensible by defining additional roles in future.

[0178] Contribution JVET-AH0172 AHG9: Text comment SEI message

[0179] This contribution is a follow-up of JVET-AG0184 proposing a new SEI message allowing to embed text comments in a video bitstream. Examples of use cases are annotations for parts of a video sequence, descriptive annotations for indexing purpose, user generated comments. In the proposed "text comment SEI message”, text comments are associated with ids to allow overlapping persistence. Compared to JVET-AG0184, use case description is augmented and a new field is proposed to specify the language of the text comment.

[0180] SEI messages in HEVC, WC, and VSEI frequently use persistence and cancel flags to identify the persistence scope of an SEI message. An example SEI message from VSEI is provided below:

[0181] sari cancel flag equal to 1 indicates that the SARI SEI message cancels the persistence of any previous SARI SEI messages in output order that applies to the current layer, sari cancel flag equal to 0 indicates that SARI follows.

[0182] sari_persistence_flag specifies the persistence of the SARI SEI message for the current layer.

[0183] sari_persistence_flag equal to 0 specifies that the SARI applies to the current decoded picture only.

[0184] sar_persistence_flag equal to 1 specifies that the SARI SEI message applies to the current decoded picture and persists for all subsequent pictures of the current layer in output order until one or more of the following conditions are true: a new coded layer video sequence (CL VS) of the current layer begins; the bitstream ends; or a picture in the current layer in an AU associated with a SARI SEI message is output that follows the current picture in output order.

[0185] The above JVET contributions don’t provide a mechanism to identify a video encoder version and / orvideo encoder configuration used to generate a bitstream in metadata carried within the bitstream.

[0186] The above JVET contributions don’t provide a mechanism to uniquely identify a bitstream in metadata carried within the bitstream.

[0187] When an SEI message included in a coded picture uses the typical persistence and cancel flags to define persistence scope, and when the persistence flag is set to 1; decoder / receiver receiving the coded picture cannot determine whether future pictures in CLVS include an SEI that cancels the persistence of the SEI message and / or includes different values of the signaled syntax elements.

[0188] Text description purposes

[0189] Various embodiments propose an information message, for example, SEI message that can signal text descriptions for a variety of purposes using a common signaling mechanism, with the purpose determined by a text description purpose indicator syntax element value.

[0190] Embodiments propose that particular values of a text description purpose indicator syntax element be defined to be used for signaling a video encoder version and a video encoder configuration used to generate the coded video in the persistence scope of the SEI message.

[0191] In an embodiment, the SEI message may be provided out-of-band in a manner that its persistence scope is effectively the bitstream. For example, the SEI message may be provided in a sample entry of an ISOBMFF track. When there is a single sample entry in the track, the SEI message applies to the entire track.

[0192] Identification within the bitstream of the encoder version and encoder configuration used to generate the bitstream is useful for implementers of both encoder and decoder products to resolve issues with interoperability. It is common that when a new encoder version refines its algorithms to further exploit flexibilities allowed by the standard, it may expose a bug in an existing decoder implementation. Having encoder version and / or encoder configuration information embedded within the bitstream will aid in debugging interoperability issues.

[0193] Strings representing the encoder version might identify, for example, “VTM10.0” or “WenC vl.ll”.

[0194] In some cases, a string of the encoder configuration may comprise an identifier string representing a predefined coding configuration of the encoder.

[0195] Strings representing the encoder configuration might identify, for example, “RA”, “Al”, “LDP”, “LDB”.

[0196] In some examples, a string of the encoder configuration may comprise a preset string representing a pre-defined computational complexity configuration of the encoder.

[0197] In some examples, a string of the encoder configuration may comprise the command line options that were used to execute the encoder.

[0198] In some examples, a string of the encoder configuration may comprise the content of a configuration file that was given as input to execute the encoder.

[0199] The embodiment is described above with reference to an encoder version and an encoder configuration. It is to be understood that embodiments may be similarly realized with any other syntax elements describing how the encoder was executed to generate the coded video data. For example, syntax elements may comprise, but may not be limited to, one or more of the following:

[0200] - Encoder vendor name

[0201] - Encoder vendor identifier, such as a tag URI comprising the vendor's domain name

[0202] - Encoder name

[0203] - Encoder version

[0204] - An identifier string that identifies a pre-defined encoder configuration

[0205] - Version of the pre-defined encoder configuration

[0206] - A computational complexity preset

[0207] - Encoder command line

[0208] - Encoder configuration file content

[0209] Embodiments are also helpful when developing a new standard, such as H.267. The bitstreams created for reference software, such as ECM, can use the embodiments to indicate the version of the encoder used to create a bitstream within an SEI message in the bitstream. Conformance test bitstreams for a new standard can also use this invention.

[0210] The embodiments described herein propose that a particular value of a text description purpose indicator syntax element be defined to identify a tag URI that identifies the coded video content (e.g., the bitstream), with the text string contents conforming to IETF RFC 4151. A tag URI includes an authority name and a time / date stamp and optionally contains additional specification. Having such information associated with a bitstream could be useful for a variety of purposes. For example, authority name can identify the service provider that encoded the video bitstream and the date / time when the bitstream was encoded. Date / timestamps may be used to distinguish among different versions of a bitstream.

[0211] Persistence scope

[0212] Embodiments described herein propose an alternate mechanism for signaling persistence scope of an SEI message. The typical persistence flag and a cancel flag can identify 3 states of the persistence scope. Use of a 2 -bit persistence indicator can represent 4 values, allowing indication of a 4th state, as follows. It should be noted that the 4 states may be mapped differently to the 4 values.)

[0213] 0 indicates cancel (similar to cancel flag equal 1)

[0214] 1 indicates does not persist (similar to persistence flag equal 0)

[0215] 2 indicates persistence until cancelled (similar to persistence flag equal 1)

[0216] 3 indicates persistence for the entire CLVS. The SEI message contents may be repeated in the CLVS for that text ID value, but not be modified.

[0217] The newly available state, represented without loss in generality by the value 3, provides useful additional information to a decoder / receiver. Without this information, in order to determine whether or not parameters signaled in an SEI message in an earlier picture apply to later pictures in the CLVS, a decoder / receiver must check every picture in the CLVS for the presence of an SEI message of the same type, and once such an SEI message is identified, the contents of the new SEI message are compared with the contents of the previous SEI message of the same type. The embodiments described herein, enable an encoder to notify the decoder / receiver that such checking is unnecessary and avoids the need to process SEI messages multiple times, simplifying usage of the SEI message by the decoder / receiver.

[0218] This persistence scope feature of the embodiments may be used for other types of SEI messages, in addition to the text description SEI message.

[0219] In case the SEI message includes an ID, the persistence scope mechanism applies to SEI messages of the same type with the same ID value.

[0220] The embodiments described above with reference to the persistence indicator value equal to 3 specifying that the persistence for the entire CLVS. It is to be understood that embodiments may be similarly realized with persistence indicator value(s) additionally or alternatively being indicative of any other pre-defined scope, such as, but not limited to, one or more of the following: a coded video sequence (including all its layers), a bitstream (including all its coded video sequences), a layer within a bitstream (including all the CLVSs of the layer within a bitstream). It is to be understood that embodiments or example semantics described with reference to a particular coded video unit, such as a bitstream, may be likewise realized with another coded video unit, such as a CLVS or a coded video sequence.

[0221] In an embodiment, the persistence scope may be additionally qualified with an indicated range of temporal sublayers. For example, when the persistence scope indicator indicates a persistence for the entire CLVS, the SEI message may include syntax element(s) indicative of a sublayer range within the CLVS that the SEI message applies to.

[0222] The embodiments described herein propose an SEI message for carrying text descriptions within a video bitstream, with the following features:

[0223] Define additional purposes for the text description, as follows:- Encoder version: identifies the encoder version used to generate the bitstream, such as VTM10.0 or WenC vl. i l- Encoder configuration: identifies the encoder configuration, such as RA, LDP, LDB, Al- Tag URL uniquely identifies the bitstreamSignaling an ID to identify the SEI message (vs. other SEI message of the same payload type)

[0224] Signaling of multiple text descriptions within the same SEI message, each with a signalled purpose indicator

[0225] Optional language signaling for the SEI messageAdditional languages are supported through use of an additional SEI message per language

[0226] Use a 2 -bit persistence indicator syntax element to represent persistence scope with 4 possible values: 0 indicates cancel (similar to cancel flag equal 1); 1 indicates does not persist (similar to persistence flag equal 0); 2 indicates persistence until cancelled (similar to persistence flag equal 1); 3 indicates persistence for the entire CLVS. The SEI message contents may be repeated in the CLVS for that text ID value, but may not be modified.

[0227] Additional text description purpose indicator values

[0228] Identification within the bitstream of the encoder version and encoder configuration used to generate the bitstream is asserted to be useful for implementers of both encoder and decoder products to resolve issues with interoperability.

[0229] Strings representing the encoder version might identify, for example, “VTM10.0” or “Wenc vl.ll”.

[0230] Strings representing the encoder configuration might identify, for example, “RA”, “Al”, “LDP”, “LDB”.

[0231] Such a feature may be useful during development of the conformance test suite for WC and earlier standards.

[0232] A tag URI provides a unique identifier that can be used to uniquely identify a bitstream. As defined in IETF RFC 4151 , a tag URI includes an authority name and a time / date stamp and optionally includes additional specification. Having such information associated with a bitstream could be useful for a variety of purposes.

[0233] In an embodiment, a text description purpose value may be defined for a tag URI that identifies the bitstream to which the coded picture(s) in the persistence scope of this SEI message belong or have belonged.

[0234] In an embodiment, an entity obtains a first CLVS from a first bitstream and a second CLVS from a second bitstream and inserts both the first CLVS and the second CLVS in a target bitstream. The target bitstream can include first and second text description information SEI messages, respectively, where the tag URI strings for the first and second CLVSs differ to indicate that the CLVSs have belonged to different original bitstreams (i.e., the first bitstream and the second bitstream, respectively).

[0235] In an additional embodiment, the entity inserts at least one third text description information SEI message with a different ID value than that in the first and second text description information SEI messages, with a purpose indicating a tag URI for the bitstream, and a tag URI value that identifies the target bitstream.

[0236] In an embodiment, a text description purpose value may be defined for a tag URI that identifies the bitstream to which the coded picture(s) in the persistence scope of this SEI message were originally encoded.

[0237] Persistence scope

[0238] A persistence indicator syntax element is signaled in the SEI message, which has following 4 example values:

[0239] 0 indicates cancel (similar to cancel flag equal 1)

[0240] 1 indicates does not persist (similar to persistence flag equal 0)

[0241] 2 indicates persistence until cancelled or CLVS ends (similar to persistence flag equal 1)

[0242] 3 indicates persistence for the entire CLVS. The SEI message contents may be repeated in the CLVS for that td id value but are not modified.

[0243] The first 3 of the 4 values correspond to states that can be accomplished using typical persistence and cancellation flags. The 4th value represents a new state that provides additional information not available using only a persistence and a cancel flag.

[0244] Decoders may find it useful to be aware that the parameters signalled in the first Text Description SEI message with a particular td id value in the CLVS applies to all pictures in the entire CLVS, without being required to check when any future Text Description SEI messages in the same CLVS cancel the persistence or contain different contents.

[0245] A semantic constraint on the allowable values of td_persistence_idc is proposed for the newly proposed values of td_purpose_idc[ i ] - Encoder version, Encoder configuration, and Tag URI - as those purposes apply to the entire CLVS rather than to individual pictures.

[0246] The proposed syntax is described below:

[0247] Text descriptions information SEI message semantics

[0248] The text descriptions information SEI message provides one or more text descriptions about one or more pictures.

[0249] td id indicates the identifier value of this text descriptions information SEI message.

[0250] td_persistence_idc equal to 0 indicates that the text descriptions information SEI message cancels the persistence of any previous text description information SEI message in output order that applies to the current layer and has the value of td id equal to that of the current text descriptions information SEI message.

[0251] td_persistence_idc equal to 1 specifies that the text descriptions information applies to the current decoded picture only.

[0252] txt_persistence_idc equal to 2 specifies that the text descriptions information SEI message applies to the current decoded picture and persists for all subsequent pictures of the current layer in output order until one or more of the following conditions are true:- A new CLVS of the current layer begins.- The bitstream ends.- A picture in the current layer in an AU associated with a text description information SEI message with the same value of td id is output that follows the current picture in output order.

[0253] td_persistence_idc equal to 3 specifies that the text descriptions information applies to the entire CLVS.

[0254] It is a requirement of bitstream conformance that when td_persistence_idc is equal to 3 and this SEI message is not the first text descriptions SEI message in a CLVS, in decoding order, a text descriptions SEI message with the same value of td id shall be present in the first PU of the CLVS in decoding order.

[0255] Text description SEI messages in a CLVS that have a particular td id value and td_persistence_idc equal to 3 shall have identical SEI payload content.

[0256] td num descriptions minusl plus 1 indicates the number of entries for td_purpose_id[ i ] and td_string[ i ] that follow.

[0257] td_language_present_flag equal to 1 indictes that the td language syntax element is present. td_language_present_flag equal to 0 indictes that the td language syntax element is not present.

[0258] td language contains a language tag as specified by IETF RFC 5646 followed by a null termination byte equal to 0x00 indicating the language of the td_string[ i ] syntax elements. The length of the td language syntax element shall be less than or equal to 63 bytes, not including the null termination byte. When not present, the language of the td_string[ i ] syntax elements is unspecified.

[0259] td_purpose_idc[ i ] indicates the purpose of the td_string[ i ] text description string, as specified in Table 1.

[0260] Values of td_purpose_idc[ i ] that are identified as reserved for future use in Rec. ITU-T H.274 | ISO / IEC 23002-7 shall not be present in bitstreams conforming to this version of this Specification.

[0261] Following table, Table 1, for definition of td_purpose_idc:Table 1

[0262] td_string[ i ] specifies i-th text description information string whose value is interpreted as specified by the td_purpose_idc[ i ]. The length of td_string[ i ] shall be in the range of 0 to 8191, inclusive.

[0263] When td_purpose_idc[ i ] is equal to 0, the interpretation of what information is conveyed in td_string[ i ] is application-defined.

[0264] When td_purpose_idc[ i ] is equal to 1, td_string[ i ] specifies copyright information that pertains to the picture(s) in the persistence scope defined by td cancel flag and td_persistence_flag.

[0265] When td_purpose_idc[ i ] is equal to 2, td_string[ i ] specifies artificial intelligence (Al) marking information that pertains to the picture(s) in the persistence scope defined by td cancel flag and td_persistence_flag.

[0266] When td_purpose_idc[ i ] is equal to 3, td_string[ i ] specifies a general text label description that pertains to the picture(s) in the persistence scope defined by td cancel flag and td_persistence_flag.

[0267] When td_purpose_idc[ i ] is equal to 4, td_string[ i ] specifies content advisory rating information conforming to US. And Canadian Rating Region Tables (RRT) [3] that pertains to the picture(s) in the persistence scope defined by td cancel flag and td persistence flag.

[0268] When td_purpose_idc[ i ] is equal to 5, td_string[ i ] specifies a description of the version of the encoding system used to generate the bitstream.

[0269] When td_purpose_idc[ i ] is equal to 6, td_string[ i ] specifies a description of the configuration of the encoding system used to generate the bitstream.

[0270] When td_purpose_idc[ i ] is equal to 7, td_string[ i ] contains a tag URI with syntax and semantics as specified in IETF RFC 4151 identifying the bitstream.

[0271] When td_purpose_idc[ i ] is equal to 5, 6, or 7, it is a requirement of bitstream conformance that the value of td_persistence_idc shall be equal to 3.

[0272] Persistence definition of the text descriptions information SEI message in relation to coded pictures or decoded pictures

[0273] Some of the text description purposes (e.g., purposes of identifying the bitstream and identifying the encoder used for encoding the bitstream) relate to coded pictures, whereas some others relate to decoded pictures, and some could relate to both coded and decoded pictures. The persistence of the text descriptions information (TDI) SEI message is currently specified in relation to decoded pictures. Hence, the specified persistence is not suitable for the purposes relating to coded pictures.

[0274] Example option 1: Categorization of text description purposes to those applying to coded pictures and those applying to decoded pictures

[0275] In an embodiment, text description purposes are categorized into two categories distinguished by whether they apply to coded or decoded pictures. For the text description purposes applying to coded pictures, define the persistence of the TDI SEI message in terms of coded pictures in decoding order. For the text description purposes applying to decoded pictures, define the persistence of the TDI SEI message in terms of decoded pictures in output order.

[0276] In an embodiment, the variable TdiCodedVideoFlag, specifying whether the text description information SEI message applies to coded or decoded video, is derived as follows: If the text description purpose is application-defined, TdiCodedVideoFlag is defined by external means (e.g., the application). Otherwise, if the text description purpose is among a pre-defined set (e.g., including purposes of identifying the bitstream and identifying the encoder used for encoding the bitstream), TdiCodedVideoFlag is set equal to 1. Otherwise, TdiCodedVideoFlag is set equal to 0.

[0277] When TdiCodedVideoFlag is equal to 1, tdi_persistence_flag equal to 1 specifies that the text description information SEI message with identifier equal to tdi descr id and purpose equal to tdi_descr_purpose applies to the current coded picture and persists for all subsequent coded pictures of the current layer in decoding order until one or more of the following conditions are true: i) A new CLVS of the current layer begins, ii) The bitstream ends, iii) A coded picture in the current layer associated with a text description information SEI message with the same values of tdi descr id and tdi_descr_purpose follows the current coded picture in decoding order.

[0278] Example option 2: Categorization of text description purposes to those applying to a CLVSand those applying to decoded pictures

[0279] In an embodiment, text description purposes are categorized into two categories distinguished by whether they apply to a CLVS or decoded pictures.

[0280] In an embodiment, a pre-defined set of text description purposes is such that applies to a CLVS. For example, the pre-defined set of text description purposes may comprise the following purposes: copyright, tag URI for the bitstream, encoder description.

[0281] In an embodiment, when a text description purpose indicated in a TDI SEI message is among the pre-defined set, the encoder encodes the TDI SEI message with persistence flag equal to 1 (may persist for multiple pictures) and refrains from cancelling the TDI SEI message with another TDI SEI message with the same text description purpose.

[0282] Example option 3: Categorization of text description purposes to those applying to coded pictures and those applying to decoded pictures and requiring some text description purposes to persist for a CLVS

[0283] In an embodiment, text description purposes are categorized into two categories distinguished by whether they apply to coded or decoded pictures. For the text description purposes applying to coded pictures, define the persistence of the TDI SEI message in terms of coded pictures in decoding order. For the text description purposes applying to decoded pictures, define the persistence of the TDI SEI message in terms of decoded pictures in output order. Moreover, a pre-defined set of text description purposes is such that applies to a CLVS. For example, the pre-defined set of text description purposes may comprise the following purposes: tag URI for the bitstream, encoder description.

[0284] In an embodiment, when a text description purpose indicated in a TDI SEI message is among the pre-defined set, the encoder encodes the TDI SEI message with persistence flag equal to 1 (may persist for multiple pictures) and refrains from cancelling the TDI SEI message with another TDI SEI message with the same text description purpose.

[0285] Example option 4: Persistence indicator indicating persistence applying to a single picture, one or more coded pictures, one or more decoded pictures, or a CLVS

[0286] In an embodiment, which may be applied to the text descriptions information SEI message or any other SEI message, a persistence indicator is included in the SEI message instead of a persistence flag. The following semantics or a subset thereof may be defined for values of the persistence indicator (for example for values 0 to 3, respectively): persistence for the current picture only, persistence fordecoded pictures in output order until cancelled, persistence for coded pictures in decoding order until cancelled, persistence for the current CLVS.

[0287] In an embodiment, a pre-defined set of text description purposes is such that applies to a CLVS. For example, the pre-defined set of text description purposes may comprise the following purposes: tag URI for the bitstream, encoder description. When a text description purpose indicated in a TDI SEI message is among the pre-defined set, the encoder encodes the TDI SEI message with persistence indicator equal to a value indicating persistence for the current CLVS.

[0288] FIG. 5 is an example apparatus 500, which may be implemented in hardware, configured to implement the examples described herein. The apparatus 500 comprises at least one processor 502 (e.g., an FPGA and / or CPU), at least one memory 504 including computer program code 505, the computer program code 505 having instructions to carry out the methods described herein, wherein the at least one memory 504 and the computer program code 505 are configured to, with the at least one processor 502, cause the apparatus 500 to implement circuitry, a process, component, module, or function (implemented with control module 506) to implement the examples described herein, including carrying text descriptions within video coding. Optionally included encoder 508 of the control module 506 implements encoding based on the examples described herein, and optionally included decoder 510 implements decoding based on the examples described herein. The at least one memory 504 may be a non-transitory memory, a transitory memory, a volatile memory (e.g. RAM), or a non-volatile memory (e.g., ROM).

[0289] The apparatus 500 includes a display and / or I / O interface 512, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, etc. The apparatus 500 includes one or more communication e.g. network (N / W) interfaces (I / F(s)) 514. The communication I / F(s) 514 may be wired and / or wireless and communicate over the Intemet / other network(s) via any communication technique including via one or more links 516. The communication I / F(s) 514 may comprise one or more transmitters or one or more receivers.

[0290] The transceiver 518 comprises one or more transmitters 520 and one or more receivers 522. The transceiver 518 and / or communication I / F(s) 514 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 524 used for communication over wireless link 526.

[0291] The control module 506 of the apparatus 500 comprises one of or both parts 506-1 and / or506-2, which may be implemented in a number of ways. The control module 506 may be implemented in hardware as control module 506-1, such as being implemented as part of the at least one processor 502. The control module 506-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 506 may be implemented as control module 506-2, which is implemented as computer program code (having corresponding instructions) 505 and is executed by the at least one processor 502. For instance, the at least one memory 504 store instructions that, when executed by the at least one processor 502, cause the apparatus 500 to perform one or more of the operations as described herein. Furthermore, the at least one processor 502, the at least one memory 504, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.

[0292] The apparatus 500 to implement the functionality of control module 506 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 500 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 500 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.

[0293] The apparatus 500 may also be distributed throughout the network including within and between apparatus 500 and any network element (such as a base station and / or terminal device and / or user equipment).

[0294] Interface 528 enables data communication and signaling between the various items of apparatus 500, as shown in FIG. 5. For example, the interface 528 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 505, including control module 506 may comprise object-oriented software configured to pass data or messages between objects within computer program code 505. The apparatus 500 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 500 may at least partially reside in a housing 530, or a subset of the various components of apparatus 500 may at least partially be located in different housings, which different housings may include housing 530.

[0295] FIG. 6 shows a schematic representation of non-volatile memory media 600a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 600b (e.g. universal serial bus (USB) memory stick) and 600c (e.g. cloud storage for downloading instructions and / or parameters 602 or receiving emailed instructions and / or parameters 602) storing instructions and / or parameters 602 whichwhen executed by a processor allows the processor to perform one or more of the operations of the methods described herein. Instructions and / or parameters 602 may represent or correspond to a non- transitory computer readable medium.

[0296] FIG. 7 is an example method 700 performed with an encoder, based on the example embodiments described herein. At 702, the method 700 includes defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes include at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content. At 704, the method 700 includes signaling the information message to a decoder.

[0297] In an example, the one or more purposes may further include one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0298] The method 700 may be performed with an encoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, for example, transmitting apparatus 406 with the encoder 402, or the apparatus 400 with the encoder 402.

[0299] FIG. 8 is another example method 800 performed with an encoder, based on the example embodiments described herein. At 802, the method 800 includes defining an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence. At 804, the method 800 includes signaling the information message to a decoder.

[0300] In an example, the persistence indicator uses one or more syntax elements comprising following values for representing a persistent scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and / or a fourth value for indicating persistence until cancelled.

[0301] In an example, the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

[0302] In an example, the information message further comprises text descriptions, and wherein thetext descriptions comprise one or more purposes, and wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content.

[0303] In an example, the one or more purposes may further include one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0304] The method 800 may be performed with an encoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, for example, the transmitting apparatus 406 with the encoder 402, or the apparatus 400 with the encoder 402.

[0305] FIG. 9 is yet another example method 900 performed with an encoder, based on the example embodiments described herein. At 902, the method 900 includes signaling at most one language tag per information message. At 904, the method 900 signaling one or more text descriptions comprised in the information message, At 906, the method 900 signaling a purpose and a text string for each text description of the one or more text descriptions.

[0306] In an example, the text string includes one or more of the following: an identifier string representing a predefined coding configuration of an encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

[0307] The method 900 may be performed with an encoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, for example, the transmitting apparatus 406 with the encoder 402, or the apparatus 400 with the encoder 402.

[0308] FIG. 10 is an example method 1000 performed with a decoder, based on the example embodiments described herein. At 1002, the method 1000 includes receiving, in or from, a bitstream an information message carrying text descriptions in a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content. At 1004, the method 1000 includes decoding images and / or a video sequence into at least one decoded picture. At 1006, the method 1000 includes associating the information message contents with the decoded picture.

[0309] In an example, the one or more purposes may further include one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

[0310] The method 1000 may be performed with a decoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, for example, the receiving apparatus 410 with the decoder 412, or the apparatus 400 with the decoder 412.

[0311] FIG. 11 is another example method 1100 performed with a decoder, based on the example embodiments described herein. At 1102, the method 1100 includes receiving, in or from, a bitstream an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence. At 1104, the method 1100 decoding the CLVS or coded video sequence into decoded pictures. At 1106, the method 1100 associating the information message with decoded pictures of the CLVS or coded video sequence.

[0312] In an example, the association of the information with decoded pictures is performed without checking every picture in the CLVS or coded video sequence for presence of the information message of same type.

[0313] The method 1100 may be performed with a decoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, for example, the receiving apparatus 410 with the decoder 412, or the apparatus 400 with the decoder 412.

[0314] FIG. 12 is yet another example method 1200 performed with a decoder, based on the example embodiments described herein. At 1202, the method 1200 includes receiving at most one language tag per information message. At 1204, the method 1200 includes receiving one or more text descriptions comprised in the information message. At 1206, the method 1200 includes receiving a purpose and a text string for each text description of the one or more text descriptions. At 1208, the method 1200 includes decoding images and / or a video sequence into at least one decoded picture. At 1210, the method 1200 includes associating the purpose and the text string with the at least one decoded picture.

[0315] The method 1200 may be performed with a decoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, for example, the receiving apparatus 410 with the decoder 412, or the apparatus 400 with the decoder 412.

[0316] As described above, FIGs. 7 to 12 include flowcharts of an apparatus (e.g. 100, 500, or anyother apparatuses described herein), method, and computer program product according to certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instmctions which embody the procedures described above may be stored by a memory (e.g. 58, 125, or 504) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g., 56, 120, or 502) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.

[0317] A computer program product is therefore defined in those instances in which the computer program instmctions, such as computer-readable program code portions, are stored by at least one non- transitory computer-readable storage medium with the computer program instructions, such as the computer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIGs. 7 to 12. In other embodiments, the computer program instmctions, such as the computer-readable program code portions, need not be stored or otherwise embodied by a non-transitory computer-readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer- readable program code portions, still being configured, upon execution, to perform the functions described above.

[0318] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions, ft will also be understood that one or more blocks of the flowcharts,and combinations of blocks in the flowcharts, may be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.

[0319] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included. Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.

[0320] In the above, some example embodiments have been described with reference to an SEI message. It needs to be understood, however, that embodiments may similarly be realized with any similar structures or data units, such as a metadata open bitstream unit (OBU), as specified in AVI or AV2, for example. Where example embodiments have been described with SEI messages included in a structure, any independently parsable structures could likewise be used in embodiments. Specific SEI SEI message syntax structures have been presented in example embodiments, but it needs to be understood that embodiments generally apply to any syntax structures with a similar intent SEI messages.

[0321] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.

[0322] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.

[0323] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodimentswithout departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0324] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.

[0325] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like.

[0326] As used herein, the term ‘circuitry’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. This description of ‘circuitry’ applies to uses of this term in this application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.

[0327] Circuitry or Circuit: As used in this application, the term ‘circuitry’ or ‘circuit’ may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware; and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0328] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example, and when applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; and signaling the information message to a decoder.

2. The apparatus of claim 1, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

3. The apparatus of claim 1, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

4. The apparatus of claim 1, wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CL VS) or an entire coded video sequence.

5. The apparatus of claim 4, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; anda fourth value for indicating persistence until cancelled.

6. The apparatus of any of the claims 4 or 5, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

7. The apparatus of claim 2, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration file given as input to execute the encoder.

8. The apparatus of claim 3, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

9. The apparatus of any of the previous claims, wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

10. The apparatus of claim 9, wherein for the one or more purposes applying to the coded pictures, and wherein the apparatus is further caused to perform: defining a persistence of the information message in terms of coded pictures in decoding order.

11. The apparatus of claim 9, wherein for the one or more purposes applying to the decoded pictures, and wherein the apparatus is further caused to perform: defining a persistence of the information message in terms of decoded pictures in output order.

12. The apparatus of any of the claim 1 to 8, wherein the one or more purposes comprise a predefined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

13. The apparatus of claim 12, wherein the pre-defined set of purpose comprises the tag URIand an encoder description.

14. The apparatus of claim 13, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the apparatus is further caused to perform: encoding the information message with persistence flag equal to a first value and refrains from cancelling the information message with another information message with the same purpose.

15. The apparatus of any of the previous claims wherein the information message comprises a supplemental enhancement information (SEI) message.

16. The apparatus of claim 15, wherein the SEI message is provided out-of-band in a manner that a persistence scope of the SEI message comprises a bitstream.

17. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: defining an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; and signaling the information message to a decoder.

18. The apparatus of claim 17, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistent scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

19. The apparatus of any of the claims 17 or 18, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

20. The apparatus of any of the claims 17 to 19, wherein the information message further comprises text descriptions, and wherein the text descriptions comprise one or more purposes,and wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content.

21. The apparatus of claim 20, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

22. The apparatus of claim 20, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.23.An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: signaling at most one language tag per information message; signaling one or more text descriptions comprised in the information message; and signaling a purpose and a text string for each text description of the one or more text descriptions.

24. The apparatus of claim 23, wherein the text string comprises one or more of the following: an identifier string representing a predefined coding configuration of an encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration fde given as input to execute the encoder.

25. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least oneprocessor, cause the apparatus at least to perform: receiving, in or from, a bitstream an information message carrying text descriptions in a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; decoding images and / or a video sequence into at least one decoded picture; and associating the information message contents with the decoded picture.

26. The apparatus of claim 25, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

27. The apparatus of claim 25, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

28. The apparatus of claim 25, wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CL VS) or an entire coded video sequence.

29. The apparatus of claim 28, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

30. The apparatus of any of the claims 28 or 29, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded videosequence is allowed and changes within the CLVS or coded video sequence are not allowed.

31. The apparatus of claim 26, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration fde given as input to execute the encoder.

32. The apparatus of claim 27, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

33. The apparatus of any of the claims 25 to 32, wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

34. The apparatus of claim 33, wherein for the one or more purposes applying to the coded pictures, and wherein the apparatus is further caused to perform: receiving a persistence of the information message in terms of coded pictures in decoding order.

35. The apparatus of claim 33, wherein for the one or more purposes applying to the decoded pictures, and wherein the apparatus is further caused to perform: receiving a persistence of the information message in terms of decoded pictures in output order.

36. The apparatus of any of the claim 25 to 32, wherein the one or more purposes comprise a pre-defined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

37. The apparatus of claim 36, wherein the pre-defined set of purpose comprises the tag URI and an encoder description.

38. The apparatus of claim 37, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the apparatus is further caused to perform: receiving the information message encoded with persistence flag equal to a first value and the informationmessage with another information message with the same purpose.

39. The apparatus of any of the claims 25 to 38, wherein the information message comprises a supplemental enhancement information (SEI) message.

40. The apparatus of claim 39, wherein the SEI message is provided out-of-band in a manner that a persistence scope of the SEI message comprises a bitstream.

41. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, in or from, a bitstream an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; decoding the CLVS or coded video sequence into decoded pictures; and associating the information message with decoded pictures of the CLVS or coded video sequence.

42. The apparatus of claim 41 , wherein the association of the information with decoded pictures is performed without checking every picture in the CLVS or coded video sequence for presence of the information message of same type.43.An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving at most one language tag per information message; receiving one or more text descriptions comprised in the information message; receiving a purpose and a text string for each text description of the one or more text descriptions; decoding images and / or a video sequence into at least one decoded picture; and associating the purpose and the text string with the at least one decoded picture.

44. A method comprising: defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; and signaling the information message to a decoder.

45. The method of claim 44, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

46. The method of claim 44, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

47. The method of claim 44, wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CL VS) or an entire coded video sequence.

48. The method of claim 47, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

49. The method of any of the claims 47 or 48, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

50. The method of claim 45, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration fde given as input to execute the encoder.

51. The method of claim 46, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

52. The method of any of the claims 44 to 51, wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

53. The method of claim 52, wherein for the one or more purposes applying to the coded pictures, and wherein the method is further caused to perform: defining a persistence of the information message in terms of coded pictures in decoding order.

54. The method of claim 52, wherein for the one or more purposes applying to the decoded pictures, and wherein the method is further caused to perform: defining a persistence of the information message in terms of decoded pictures in output order.

55. The method of any of the claim 44 to 51 , wherein the one or more purposes comprise a predefined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

56. The method of claim 55, wherein the pre-defined set of purpose comprises the tag URI and an encoder description.

57. The method of claim 56, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the method is further caused to perform: encoding the information message with persistence flag equal to a first value and refrains from cancelling the information message with another information message with the same purpose.

58. The method of any of the claims 44 to 57, wherein the information message comprises a supplemental enhancement information (SEI) message.

59. The method of claim 58, wherein the SEI message is provided out-of-band in a manner that a persistence scope of the SEI message comprises a bitstream.

60. A method comprising: defining an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CL VS) or an entire coded video sequence; and signaling the information message to a decoder.61.The method of claim 60, wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistent scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

62. The method of any of the claims 60 or 61, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

63. The method of any of the claims 60 to 62, wherein the information message further comprises text descriptions, and wherein the text descriptions comprise one or more purposes, and wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content.

64. The method of claim 63, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier;an identifier string for identifying encoder configuration; a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

65. The method of claim 63, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

66. A method comprising: signaling at most one language tag per information message; signaling one or more text descriptions comprised in the information message; and signaling a purpose and a text string for each text description of the one or more text descriptions.

67. The method of claim 66, wherein the text string comprises one or more of the following: an identifier string representing a predefined coding configuration of an encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration fde given as input to execute the encoder.

68. A method comprising: receiving, in or from, a bitstream an information message carrying text descriptions in a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; decoding images and / or a video sequence into at least one decoded picture; and associating the information message contents with the decoded picture.

69. The method of claim 68, wherein the one or more purposes further comprise one or more of the following: an encoder vendor name; an encoder vendor identifier; an identifier string for identifying encoder configuration;a version of the encoder configuration; a computational complexity preset; an encoder command line; or an encoder configuration.

70. The method of claim 68, wherein the tag URI comprises at least an authority name, a date stamp, and / or a time stamp.

71. The method of claim 68, wherein the information message further comprise a persistence indicator for indicating persistence for an entire coded layer video sequence (CL VS) or an entire coded video sequence.

72. The method of claim 71 , wherein the persistence indicator uses one or more syntax elements comprising following values for representing a persistence scope: a first value for indicating persistence for the entire CLVS or coded video sequence; a second value for indicating cancel; a third value for indicating does not persist; and a fourth value for indicating persistence until cancelled.

73. The method of any of the claims 71 or 72, wherein the information message comprises a persistence value indicating persistence for the entire CLVS or coded video sequence further indicates that retransmission of the information message within the CLVS or coded video sequence is allowed and changes within the CLVS or coded video sequence are not allowed.

74. The method of claim 69, wherein a string representing the encoder configuration comprises one or more of the following: an identifier string representing a predefined coding configuration of the encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration fde given as input to execute the encoder.

75. The method of claim 70, wherein a string representing the encoder configuration identifies the encoder configuration as a random access, an all intra, a low delay P, or a low delay B.

76. The method of any of the claims 68 to 75, wherein the one or more purposes are categorized into two categories distinguished by whether the one or more purposes apply to coded or decoded pictures.

77. The method of claim 76, wherein for the one or more purposes applying to the coded pictures, and wherein the method further comprises: receiving a persistence of the information message in terms of coded pictures in decoding order.

78. The method of claim 76, wherein for the one or more purposes applying to the decoded pictures, and wherein the method further comprises: receiving a persistence of the information message in terms of decoded pictures in output order.

79. The method of any of the claim 68 to 75, wherein the one or more purposes comprise a predefined set of purposes, and wherein the pre-defined set of purposes apply to a coded layer video sequence.

80. The method of claim 79, wherein the pre-defined set of purpose comprises the tag URI and an encoder description.

81. The method of claim 80, wherein when a purpose indicated in the SEI message is among the pre-defined set, and wherein the method further comprises: receiving the information message encoded with persistence flag equal to a first value and the information message with another information message with the same purpose.

82. The method of any of the claims 68 to 81, wherein the information message comprises a supplemental enhancement information (SEI) message.

83. The method of claim 82, wherein the SEI message is provided out-of-band in a manner that a persistence scope of the SEI message comprises a bitstream.

84. A method comprising: receiving, in or from, a bitstream an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CL VS) or an entire coded video sequence;decoding the CLVS or coded video sequence into decoded pictures; and associating the information message with decoded pictures of the CLVS or coded video sequence.

85. The method of claim 84, wherein the association of the information with decoded pictures is performed without checking every picture in the CLVS or coded video sequence for presence of the information message of same type.

86. A method comprising: receiving at most one language tag per information message; receiving one or more text descriptions comprised in the information message; receiving a purpose and a text string for each text description of the one or more text descriptions; decoding images and / or a video sequence into at least one decoded picture; and associating the purpose and the text string with the at least one decoded picture.

87. An apparatus comprising: means for defining an information message for carrying text descriptions within a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; and means for signaling the information message to a decoder.

88. The apparatus of claim 87, wherein the apparatus comprises means for performing methods as claimed in any of the claims 45 to 59.

89. An apparatus comprises: means for defining an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; and means for signaling the information message to a decoder.

90. The apparatus of claim 89, wherein the apparatus further comprises means for performing methods as claimed in any of the claims 61 to 65.

91. An apparatus comprising: means for signaling at most one language tag per information message; means for signaling one or more text descriptions comprised in the information message; and means for signaling a purpose and a text string for each text description of the one or more text descriptions.

92. The apparatus of claim 91, wherein the text string comprises: an identifier string representing a predefined coding configuration of an encoder; a preset string representing a pre-defined computational complexity configuration of the encoder; command line options used to execute the encoder; or content of a configuration fde given as input to execute the encoder.

93. An apparatus comprising: means for receiving, in or from, a bitstream an information message carrying text descriptions in a coded video, wherein the text descriptions comprise one or more purposes, wherein the one or more purposes comprise at least one of an encoder name, an encoder version, or a tag uniform resource identifier (URI) for identifying the coded video content; means for decoding images and / or a video sequence into at least one decoded picture; and means for associating the information message contents with the decoded picture.

94. The apparatus of claim 93, wherein the apparatus comprises means for performing methods as claimed in any of the claims 69 to 83.

95. An apparatus comprising: means for receiving, in or from, a bitstream an information message comprising a persistence indicator for indicating persistence for an entire coded layer video sequence (CLVS) or an entire coded video sequence; means for decoding the CLVS or coded video sequence into decoded pictures; and means for associating the information message with decoded pictures of the CLVS or coded video sequence.

96. The apparatus of claim 95, wherein the association of the information with decoded pictures is performed without checking every picture in the CLVS or coded video sequence for presence of the information message of same type.

97. An apparatus comprising: means for receiving at most one language tag per information message; means for receiving one or more text descriptions comprised in the information message; means for receiving a purpose and a text string for each text description of the one or more text descriptions; means for decoding images and / or a video sequence into at least one decoded picture; and means for associating the purpose and the text string with the at least one decoded picture.

98. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 44 to 59.

99. The computer readable medium of claim 98, wherein the computer readable medium comprises a non-transitory computer readable medium.

100. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 60 to 65.

101. The computer readable medium of claim 100, wherein the computer readable medium comprises a non-transitory computer readable medium.

102. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 66 or 67.

103. The computer readable medium of claim 102, wherein the computer readable medium comprises a non-transitory computer readable medium.

104. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 68 to 83.

105. The computer readable medium of claim 104, wherein the computer readable medium comprises a non-transitory computer readable medium.

106. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 83 or 84.

107. The computer readable medium of claim 106, wherein the computer readable medium comprises a non-transitory computer readable medium.

108. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform the methods as claimed in claim 86.

109. The computer readable medium of claim 108, wherein the computer readable medium comprises a non-transitory computer readable medium.

Citation Information

Patent Citations

  • Use of embedded signalling for backward-compatible scaling improvements and super-resolution signalling

    US20220385911A1

  • Signaling parameter value information in a parameter set to reduce the amount of data contained in an encoded video bitstream

    US20230254512A1

Cited By

  • Systems and methods for signaling text description information in video coding

    US12587667B2

  • Systems and methods for signaling text description information in video coding

    US20250317592A1

  • Systems and methods for signaling text description purpose and identifier information in video coding

    US20260025526A1