Slim mode for image file format
By implementing a reduced header mode with compact metadata in image file formats, the inefficiencies of existing formats are addressed, resulting in smaller file sizes and improved processing efficiency.
Patent Information
- Application Number
- PCT/IB2025/050418
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2025-01-14
- Publication Date
- 2025-07-24
AI Technical Summary
Existing image file formats are inefficient in terms of storage and processing due to the inclusion of redundant header information, which increases file size and processing overhead.
The implementation of a reduced header mode or compact metadata in image file formats, which omits unnecessary syntax elements and uses a slim codec brand to indicate a condensed version of codec-specific structures, allowing for efficient decoding and sharing of metadata across multiple files.
This approach reduces file size and processing overhead by minimizing redundant header information, enabling faster decoding and more efficient use of storage and bandwidth.
Smart Images

Figure IB2025050418_24072025_PF_FP_ABST
Abstract
Description
SLIM MODE FOR IMAGE FILE FORMATTECHNICAL FIELD
[0001] The examples and non-limiting embodiments relate generally to multimedia transport and, more particularly, to an image file format.BACKGROUND
[0002] It is known to provide standardized formats for signaling of media data.SUMMARY
[0003] Example 1 : An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: writing in a file: an indication of presence of a reduced header mode or compact metadata; and the reduced header comprising information for processing one or more items; and signaling the file comprising the reduced header mode. In an example, the file comprises a container file.
[0004] Example 2: The apparatus of example 1, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
[0005] Example 3: The apparatus of example 2, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoder configuration record, a network abstraction layer (NAL) unit, a NAL unit header, an open bitstream unit (OBU), an OBU header, a size or length field indicative of a NAL unit size, OBU size, payload size of a NAL unit or OBU, one or more parameter sets of one or more particular types, a sequence header, a picture header, a frame header, a slice header, or a tile group header. For example, the one more codec-specific structures may be structures applying to both the NAL unit based and the non-NAL unit-based (OBU-based) codecs; structures applying to NAL unit based codecs; or structures applying to the OBU-based codecs.
[0006] Example 4: The apparatus of any of examples 1 to 3, wherein the apparatus is further caused to perform: defining a major brand field indicating that a file structurally conforms to the file with the reduced header mode; and defining a minor version field forindicating: a slim codec brand, a condensed version of one or more codec-specific structures; or a condensed configuration item property associated with the respective one or more items.
[0007] Example 5: The apparatus of any of examples 1 to 4, wherein the apparatus is further caused to perform: including the slim codec brand in the file for indicating a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items.
[0008] Example 6: The apparatus of example 5, wherein the apparatus is caused to perform: defining or deriving an inferred syntax element, wherein the inferred syntax element is used by a reader to the reconstruct the one or more codec-specific structures.
[0009] Example 7: The apparatus of example 1, wherein the apparatus is caused to perform: using a meta-box with a version greater than zero for signaling the minimized header mode.
[0010] Example 8: The apparatus of example 1, wherein the apparatus is caused to perform: indicating the reduced header mode in a file type box, an extended type box, or an original file type box, and wherein when the reduced header mode is indicated in the file type box, the extended type box, or the original file type box, the reduced header moder imposes constraints on a first media data box.
[0011] Example 9: The apparatus of example 1, wherein the apparatus is further caused to perform: indicating a normal header mode in the file comprising the reduced header mode.
[0012] Example 10: The apparatus of example 1, wherein apparatus is further caused to perform: defining a file description box for indicating a presence of both the reduced header mode and the regular header mode in the file.
[0013] Example 11: The apparatus of example 1, wherein apparatus is further caused to perform: deriving a reduced header for the file and using the reduced header as input to a hash generation algorithm; deriving the reduced header excluding pre-defined syntax elements and using the reduced header excluding the pre-defined syntax elements as input to the hash generation algorithm; deriving the reduced header excluding width and height fields for the file and using the reduced header excluding the width and height as an input to the hash generationalgorithm; or using a meta-box excluding pre-defined syntax structures or the pre-defined syntax elements as an input to the hash generation algorithm.
[0014] Example 12: The apparatus of example 1, wherein apparatus is further caused to perform: including items sharing same properties in the file.
[0015] Example 13: The apparatus of example 1, wherein the apparatus is further caused to perform: coding and enumerating the reduced header mode in a web page html source, and wherein multiple files reuse the codes, and wherein the apparatus is further caused to perform: signaling byte-ranges for item data associated with the multiple files.
[0016] Example 14: The apparatus of example 1, wherein the file is referenced and / or declared in a hyper text markup latigsage (HTML) formatted document, and wherein when the file is referenced and / or declared in the hyper text maxkup language (HTML) formatted document plurality of files share the reduced header mode metadata.
[0017] Example 15: The apparatus of example 1, wherein the apparatus is further caused to perform: defining an indicator in the HTML context to enable download one of the files, of the plurality of files, with the reduced header mode and apply processed reduced header mode metadata to the plurality of files without re-downloading respective reduced header mode metadata.
[0018] Example 16: An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving and parsing a file comprising: an indication of presence of a reduced header mode or compact metadata; and a reduced header comprising data for processing one or more items in the file; and processing the reduced header data to extract information needed for decoding of the one or more items in the file. In an example, the file comprises a container file.
[0019] Example 17: The apparatus of example 16, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
[0020] Example 18: The apparatus of example 16, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoderconfiguration record, one or more parameter sets of one or more particular types, a sequence header, a picture header, or a slice header.
[0021] Example 19: The apparatus of any of the examples 16 to 18, wherien the apparatus is further caused to perform: parsing a slim codec brand from the file, wherein the slim codec brand in indicates a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items; and wherein the apparatus is further caused to perform: parsing a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items from the slim coded brand.
[0022] Example 20: The apparatus of any of examples 18 or 19, wherein the apparatus is further caused to perform: parsing an inferred syntax element; and using the inferred syntax element to reconstruct the one or more codec-specific structures.
[0023] Example 21: The apparatus of example 20, wherein the apparatus is further caused to perform: reconstructing one or more codec-specific structures from the respective parsed condensed structures by adding the inferred syntax elements; and including a value of the inferred syntax element into the reconstructed one or more codec-specific structures.
[0024] Example 22: The apparatus of example 16, wherein items sharing same properties are comprised in the same file.
[0025] Example 23: The apparatus of example 22, wherein the file is downloaded once and item properties are shared between the one or more items.
[0026] Example 24: The apparatus of example 16, wherein the reduced hear mode is coded and enumerated in a web page html source, and wherein multiple files reuse the codes, and wherein the apparatus is caused to perform: receiving the web page; downloading the reduced header mode; and reusing the coded reduced header moder for different files, without re-downloading coded reduced header mode.
[0027] Example 25: A method comprising: writing in a file: an indication of presence of a reduced header mode or compact metadata; and the reduced header comprising information for processing one or more items; and signaling the file comprising the reduced header mode. In an example, the file comprises a container file.
[0028] Example 26: The method of example 25, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
[0029] Example 27: The method of example 26, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoder configuration record, a network abstraction layer (NAL) unit, a NAL unit header, an open bitstream unit (OBU), an OBU header, a size or length field indicative of a NAL unit size, OBU size, payload size of a NAL unit or OBU, one or more parameter sets of one or more particular types, a sequence header, a picture header, a frame header, a slice header, or a tile group header.
[0030] Example 28: The method of any of examples 25 to 27 further comprising: defining a major brand field indicating that a file structurally conforms to the file with the reduced header mode; and defining a minor version field for indicating: a slim codec brand, a condensed version of one or more codec- specific structures; or a condensed configuration item property associated with the respective one or more items.
[0031] Example 29: The method of any of examples 25 to 28 further comprising: including the slim codec brand in the file for indicating a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items.
[0032] Example 30: The method of example 29 further comprising defining or deriving an inferred syntax element, wherein the inferred syntax element is used by a reader to the reconstruct the one or more codec-specific structures.
[0033] Example 31: The method of example 25 further comprising: using a meta-box with a version greater than zero for signaling the minimized header mode.
[0034] Example 32: The method of example 25 further comprising: indicating the reduced header mode in a file type box, an extended type box, or an original file type box, and wherein when the reduced header mode is indicated in the file type box, the extended type box, or the original file type box, the reduced header moder imposes constraints on a first media data box.
[0035] Example 33: The method of example 25 further comprising: indicating a normal header mode in the file comprising the reduced header mode.
[0036] Example 34: The method of example 25 further comprising: defining a file description box for indicating a presence of both the reduced header mode and the regular header mode in the file.
[0037] Example 35: The method of example 25 further comprising: deriving a reduced header for the file and using the reduced header as input to a hash generation algorithm; deriving the reduced header excluding pre-defined syntax elements and using the reduced header excluding the pre-defined syntax elements as input to the hash generation algorithm; deriving the reduced header excluding width and height fields for the file and using the reduced header excluding the width and height as an input to the hash generation algorithm; or using a metabox excluding pre-defined syntax structures or the pre-defined syntax elements as an input to the hash generation algorithm.
[0038] Example 36: The method of example 25 further comprising: including items sharing same properties in the file.
[0039] Example 37: The method of example 25 further comprising: coding and enumerating the reduced header mode in a web page html source, and wherein multiple files reuse the codes, and wherein the method further comprises: signaling byte -ranges for item data associated with the multiple files.
[0040] Example 38: The method of example 25, wherein the file is referenced and / or declared in a hyper text markup language (HTML) formatted document, and wherein when the file is referenced and / or declared in the hyper text markup language (HTML) formatted document plurality of files share the reduced header mode metadata.
[0041] Example 39: The method of example 25 further comprising defining an indicator in the HTML context to enable download one of the files, of the plurality of files, with the reduced header mode and apply processed reduced header mode metadata to the plurality of files without re-downloading respective reduced header mode metadata.
[0042] Example 40: A method comprising: receiving and parsing a file comprising: an indication of presence of a reduced header mode or compact metadata; and a reduced headercomprising data for processing one or more items in the file; and processing the reduced header data to extract information needed for decoding of the one or more items in the file. In an example, the file comprises a container file.
[0043] Example 41: The method of example 40, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
[0044] Example 42: The method of example 40, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoder configuration record, one or more parameter sets of one or more particular types, a sequence header, a picture header, or a slice header.
[0045] Example 43: The method of any of the examples 40 to 42 further comprising: parsing a slim codec brand from the file, wherein the slim codec brand in indicates a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items; and wherein the method further comprises: parsing a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items from the slim coded brand.
[0046] Example 44: The method of any of examples 42 or 43 further comprising: parsing an inferred syntax element; and using the inferred syntax element to reconstruct the one or more codec-specific structures.
[0047] Example 45: The method of example 44 further comprising: reconstructing one or more codec-specific structures from the respective parsed condensed structures by adding the inferred syntax elements; and including a value of the inferred syntax element into the reconstructed one or more codec-specific structures.
[0048] Example 46: The method of example 40, wherein items sharing same properties are comprised in the same file.
[0049] Example 47 : The method of example 46, wherein the file is downloaded once and item properties are shared between the one or more items.
[0050] Example 48: The method of example 40, wherein the reduced hear mode is coded and enumerated in a web page html source, and wherein multiple files reuse the codes, and wherein the method further comprises: receiving the web page; downloading the reduced header mode; and reusing the coded reduced header moder for different files, without re- downloading coded reduced header mode.
[0051] Example 49: An apparatus comprising means for performing the methods as described in any of the examples 25 to 39.
[0052] Example 50: An apparatus comprising means for performing the methods as described in any of the examples 40 to 48.
[0053] Example 51: A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 25 to 39.
[0054] Example 52: The computer readable medium of example 51, wherein the computer readable medium comprises a non-transitory computer readable medium.
[0055] Example 53: A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 40 to 48.
[0056] Example 54: The computer readable medium of example 53, wherein the computer readable medium comprises a non-transitory computer readable medium.
[0057] Example 55: An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: writing, in a file, a slim codec brand indicative of a condensed configuration item property associated with an image item; and writing, in the file, the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition.
[0058] Example 56: The apparatus of example 55, wherein the file further comprises an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
[0059] Example 57: The apparatus of any of examples 55 or 56, wherein for writing the condensed configuration item property, the apparatus is caused to perform: writing a condensed video coding layer (VCL) network abstraction layer (NAL) unit by excluding the NAL unit length and / or the NAL unit header from an item data stored in the file.
[0060] Example 58: The apparatus of the example 57, wherein for writing the condensed configuration item property, the apparatus is further caused to perform: including other NAL units of a bitstream into the condensed configuration item property.
[0061] Example 59: An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving and parsing a file comprising: a slim codec brand indicative of a condensed configuration item property associated with an image item; and the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition; and processing the slim codec brand to extract information needed for decoding of one or more items in the file.
[0062] Example 60: The apparatus of example 59, wherein the one or more items comprise an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
[0063] Example 61: The apparatus of any of the examples 59 or 60, wherein the condensed configuration item property comprises a condensed video coding layer (VCL) network abstraction layer (NAL) unit excluding the NAL unit length and / or the NAL unit header from an item data stored in the file, and wherein the apparatus is further caused to perform: reconstructing a VCL NAL unit from the condensed VCL NAL unit by including the NAL unit length set equal to an item size and / or the NAL unit header set equal to pre-defined values defined in the slim codec brand.
[0064] Example 62: A method comprising: writing, in a file, a slim codec brand indicative of a condensed configuration item property associated with an image item; and writing, in the file, the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition.
[0065] Example 63: The method of example 62, wherein the file further comprises an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
[0066] Example 64: The method of any of examples 62 or 63, wherein writing the condensed configuration item property comprises: writing a condensed video coding layer (VCL) network abstraction layer (NAL) unit by excluding the NAL unit length and / or the NAL unit header from an item data stored in the file.
[0067] Example 65: The method of the example 64, wherein writing the condensed configuration item property further comprises: including other NAL units of a bitstream into the condensed configuration item property.
[0068] Example 66: A method comprising; receiving and parsing a file comprising: a slim codec brand indicative of a condensed configuration item property associated with an image item; and the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition; and processing the slim codec brand to extract information needed for decoding of one or more items in the file.
[0069] Example 67 : The method of example 66, wherein the one or more items comprise an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
[0070] Example 68: The method of any of the examples 66 or 67, wherein the condensed configuration item property comprises a condensed video coding layer (VCL) network abstraction layer (NAL) unit excluding the NAL unit length and / or the NAL unit header from an item data stored in the file, and wherein the method further comprises: reconstructing a VCL NAL unit from the condensed VCL NAL unit by including the NAL unit length set equal to an item size and / or the NAL unit header set equal to pre-defined values defined in the slim codec brand.BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0072] FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.
[0073] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.
[0074] FIG. 3 shows a block diagram of a general structure of a video encoder.
[0075] FIG. 4 illustrates a high efficiency image file format image file.
[0076] FIG. 5 is an example apparatus, which may be implemented in hardware, and is caused to, implement examples described herein.
[0077] FIG. 6 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.
[0078] FIG. 7 is an example method to implement the embodiments described herein, in accordance with an embodiment.
[0079] FIG. 8 is another example method to implement the embodiments described herein, in accordance with another embodiment.
[0080] FIG. 9 is yet another example method to implement the embodiments described herein, in accordance with an embodiment.
[0081] FIG. 10 is still another example method to implement the embodiments described herein, in accordance with another embodiment.
[0082] FIG. 11 is a block diagram of one possible and non-limiting system in which the example embodiments may be practiced.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS
[0083] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows:4CC four character code5G fifth generation cellular network technology5GC 5G core network a.k.a. also known asAVC advanced video codingCU central unitDSP digital signal processorDU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DCE-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technologyFl or Fl-C interface between CU and DU control interface gNB (or gNodeB) base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCIEC International Electrotechnical Commission loT internet of thingsISO International Organization for StandardizationISOBMFF ISO base media file formatJPEG joint photographic experts groupLTE long-term evolution mdat MediaDataBoxMIME Multipurpose Internet Mail ExtensionMME mobility management entity moov MovieBoxMP4 file format for MPEG-4 Part 14 filesMPEG moving picture experts groupMPEG-2 H.222 / H.262 as defined by the ITUMPEG-4 audio and video coding standard for ISO / IEC 14496 ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radio (5G radio)N / W or NW networkPDCP packet data convergence protocolPHY physical layerPNG portable network graphicsRAN radio access networkRFC request for commentsRLC radio link controlRRC radio resource controlRRH remote radio headRU radio unitRx receiverSDAP service data adaptation protocolSGW serving gatewaySMF session management functionSPS sequence parameter setSVC scalable video codingSI interface between eNodeBs and the EPC trak TrackBoxTx transmitterUE user equipmentUICC Universal Integrated Circuit CardUPF user plane functionURL uniform resource locatorX2 interconnecting interface between two eNodeBs in LTE networkXn interface between two NG-RAN nodes
[0084] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments are shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments.
[0085] Additionally, as used herein, the term ‘circuitry’ refers to (a) hardware-only circuit implementations (e.g., implementations in analog circuitry and / or digital circuitry); (b) combinations of circuits and computer program product(s) comprising software and / or firmware instructions stored on one or more computer readable memories that work together to cause an apparatus to perform one or more functions described herein; and (c) circuits, such as, for example, a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation even when the software or firmware is not physically present. This definition of ‘circuitry’ applies to all uses of this term herein, including in any claims. As a further example, as used herein, the term ‘circuitry’ also includes an implementation comprising one or more processors and / or portion(s) thereof and accompanying software and / or firmware. As another example, the term ‘circuitry’ as used herein also includes, for example, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, other network device, and / or other computing device.
[0086] As defined herein, a ‘computer-readable storage medium,’ which refers to a non- transitory physical storage medium (e.g., volatile or non-volatile memory device), may be differentiated from a ‘computer-readable transmission medium,’ which refers to an electromagnetic signal.
[0087] A method, apparatus and computer program product are provided in accordance with example embodiments for slim mode for image file format.
[0088] In an example, the following describes in detail suitable apparatus and possible mechanisms for implementing slim mode for image file format. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an apparatus 50. The apparatus may be an internet of things (loT) apparatus configured to perform various functions, for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 will be explained next.
[0089] The apparatus 50, may for example be, a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or a lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.
[0090] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 may further comprise a display 32, for example, in the form of a liquid crystal display, light emitting diode display, organic light emitting diode display, and the like. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display media or multimedia content, for example, an image or a video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0091] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 38, speaker, or an analogue audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera 42 capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.
[0092] The apparatus 50 may comprise a controller 56, a processor or a processor circuitry for controlling the apparatus 50. The controller 56 may be connected to a memory 58 which in embodiments of the examples described herein may store both data in the form of an image, audio data and video data, and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio, image and / or video data or assisting in coding and / or decoding carried out by the controller.
[0093] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example, a universal integrated circuit card (UICC) and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
[0094] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals, for example, for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).
[0095] The apparatus 50 may comprise a camera 42 capable of recording or detecting individual frames which are then passed to the codec circuitry 54 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
[0096] FIG. 3 shows a block diagram of a general structure of a video encoder. FIG. 3 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers. FIG. 3 illustrates a video encoder comprising a first encoder section 501 for a base layer and a second encoder section 502 for an enhancement layer. Each of the first encoder section 501 and the second encoder section 502may comprise similar elements for encoding incoming pictures. The encoder sections 501, 502 may comprise a pixel predictor 302, 402, prediction error encoder 303, 403 and prediction error decoder 304, 404. FIG. 3 also shows an embodiment of the pixel predictor 302, 402 as comprising an inter-predictor 306, 406, an intra-predictor 308, 408, a mode selector 310, 410, a filter 316, 416, and a reference frame memory 318, 418. The pixel predictor 302 of the first encoder section 501 receives base layer picture(s) / image(s) 300 of a video stream to be encoded at both the inter -predictor 306 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 308 (which determines a prediction for an image block based only on the already processed parts of current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 310. The intra-predictor 308 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer image(s) 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives enhancement layer picture(s) / images(s) 400 of a video stream to be encoded at both the interpredictor 406 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 410. The intra-predictor 408 may have more than one intra-prediction modes. Hence, each mode may perform the intraprediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer pictures 400.
[0097] Depending on which encoding mode is selected to encode the current block, the output of the inter-predictor 306, 406 or the output of one of the optional intra-predictor modes or the output of a surface encoder within the mode selector is passed to the output of the mode selector 310, 410. The output of the mode selector 310, 410 is passed to a first summing device 321, 421. The first summing device may subtract the output of the pixel predictor 302, 402 from the base layer image(s) 300 / enhancement layer image(s) 400 to produce a first prediction error signal 320, 420 which is input to the prediction error encoder 303, 403.
[0098] The pixel predictor 302, 402 further receives from a preliminary reconstructor 339, 439 the combination of the prediction representation of the image block 312, 412 and the output 338, 438 of the prediction error decoder 304, 404. The preliminary reconstructed image 314, 414 may be passed to the intra-predictor 308, 408 and to the filter 316, 416. The filter 316, 416 receiving the preliminary representation may filter the preliminary representation andoutput a final reconstructed image 340, 440 which may be saved in the reference frame memory 318, 418. The reference frame memory 318 may be connected to the inter-predictor 306 to be used as the reference image against which a future base layer image 300 is compared in interprediction operations. Subject to the base layer being selected and indicated to be source for inter-layer sample prediction and / or inter-layer motion information prediction of the enhancement layer according to some embodiments, the reference frame memory 318 may also be connected to the inter-predictor 406 to be used as the reference image against which a future enhancement layer image(s) 400 is compared in inter-prediction operations. Moreover, the reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which the future enhancement layer image(s) 400 is compared in interprediction operations.
[0099] Filtering parameters from the filter 316 of the first encoder section 501 may be provided to the second encoder section 502 subject to the base layer being selected and indicated to be source for predicting the filtering parameters of the enhancement layer according to some embodiments.
[0100] The prediction error encoder 303, 403 comprises a transform unit 342, 442 and a quantizer 344, 444. The transform unit 342, 442 transforms the first prediction error signal 320, 420 to a transform domain. The transform is, for example, the DCT transform. The quantizer 344, 444 quantizes the transform domain signal, for example, the DCT coefficients, to form quantized coefficients.
[0101] The prediction error decoder 304, 404 receives the output from the prediction error encoder 303, 403 and performs the opposite processes of the prediction error encoder 303, 403 to produce a decoded prediction error signal 338, 438 which, when combined with the prediction representation of the image block 312, 412 at the second summing device 339, 439, produces the preliminary reconstructed image 314, 414. The prediction error decoder may be considered to comprise a dequantizer 346, 446, which dequantizes the quantized coefficient values, for example, DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 348, 448, which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 348, 448 includes reconstructed block(s). The prediction error decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.
[0102] The entropy encoder 330, 430 receives the output of the prediction error encoder 303, 403 and may perform a suitable entropy encoding / variable length encoding on the signal to provide a compressed signal. The outputs of the entropy encoders 330, 430 may be inserted into a bitstream, for example, by a multiplexer 508.
[0103] The one or more apparatuses described in FIGs 1 to 3 may be caused to implement slim mode for image file format, for example, for implement slim mode for high efficiency image file format.
[0104] ISO / IEC 23008-12 specifies high efficiency image file format or HEIF, which is designed to enable the interchange of images and image sequences, as well as their associated metadata. Images may be stored as items using the support for untimed data storage in ISOBMFF, utilizing the MetaBox. A file may include any number of image items.
[0105] ISO base media file format
[0106] Available media file format standards include International Standards Organization (ISO) base media file format (ISO / IEC 14496-12, which may be abbreviated ISOBMFF), Moving Picture Experts Group (MPEG)-4 file format (ISO / IEC 14496-14, also known as the MP4 format), the file format for Network Abstraction Layer (NAL) unit structured video (ISO / IEC 14496-15), and and High Efficiency Video Coding standard (HEVC or H.265 / HEVC).
[0107] Some concepts, structures, and specifications of ISOBMFF are described below as an example of a container file format, based on which some embodiments may be implemented. The features of the disclosure are not limited to ISOBMFF, but rather the description is given for one possible basis on top of which at least some embodiments may be partly or fully realized.
[0108] A basic building block in the ISO base media file format is called a box. Each box has a header and a payload. The box header indicates the type of the box and the size of the box in terms of bytes. Box type is typically identified by an unsigned 32-bit integer, interpreted as a four character code (4CC). A box may enclose other boxes, and the ISO file format specifies which box types are allowed within a box of a certain type. Furthermore, the presence of some boxes may be mandatory in each file, while the presence of other boxes may be optional. Additionally, for some box types, it may be allowable to have more than one box present in afile. Thus, the ISO base media file format may be considered to specify a hierarchical structure of boxes.
[0109] According to ISOBMFF, a file includes metadata encapsulated into boxes and may also include media data encapsulated into boxes. Media data may alternatively or additionally be present in other file(s) that a referenced by a file conforming to ISOBMFF.
[0110] The syntax of boxes may be specified using the syntax description language (SDL) defined as part of ISO / IEC 14496-1. Standardization is currently ongoing to extract the SDL into a new standard, ISO / IEC 14496-34, with some updates.
[0111] The syntax of a Box is as follows: aligned(8) class Box (unsigned int(32) boxtype, optional unsigned int(8)
[0016] extended_type) { unsigned int(32) size; unsigned int(32) type = boxtype; if (size==l) { unsigned int(64) largesize;} else if (size==0) { / / box extends to end of file} if (boxtype=='uuid') { unsigned int(8)
[0016] usertype = extended_type;}}
[0112] In the above syntax, size is an integer that specifies the number of bytes in this box, including all its fields and included boxes; when size is 1 then the actual size is in the field largesize; when size is 0, then this box must be in a top-level container, and be the last box in that container (typically, a file or data object delivered over a protocol), and its contents extend to the end of that container (normally only used for a MediaDataBox). type identifies the box type; user extensions use an extended type, and in this case, the type field is set to 'uuid'.
[0113] A FullBox extends the Box syntax by adding version and flags fields into the box header. The version field is an integer that specifies the version of this format of the box. The flags field is a map or a bit field of flags. Parsers may be required that to ignore and skip boxes that have an unrecognized version value. The syntax of a FullBox may be specified as follows: aligned(8) class FullBox(unsigned int(32) boxtype, unsigned int(8) v, bit(24) f) extends Box(boxtype) { unsigned int(8) version = v; bit(24) flags = f;}
[0114] The prefix Ox in a value may indicate that the value is a hexadecimal or base-16 value.
[0115] ISOBMFF defines the FileTypeBox as below:Box Type: 'ftyp'Container: File, or OriginalFileTypeBoxMandatory: YesQuantity: Exactly one (but see below)
[0116] Files with no FileTypeBox should be read as when they included a FileTypeBox with Major_brand='mp41', minor_version=0, and the single compatible brand 'mp4T.
[0117] A media-file structure may be compatible with more than one detailed specification, and it is therefore not always possible to speak of a single ‘type’ or ‘brand’ for the file. This means that the utility of the file name extension and multipurpose internet mail extension (MIME) type are somewhat reduced.
[0118] This box shall be placed as early as possible in the file (e.g. after any obligatory signature, but before any significant variable-size boxes such as a MovieBox, MediaDataBox, or FreeSpaceBox). It identifies which specification is the ‘best use’ of the file, and a minor version of that specification; and also a set of other specifications to which the file complies. Readers implementing this format should attempt to read files that are marked as compatible with any of the specifications that the reader implements. Any incompatible change in a specification should therefore register a new ‘brand’ identifier to identify files conformant to the new specification.
[0119] The minor version is informative only. It does not appear for compatible-brands, and is not used to determine the conformance of a file to a standard. It may allow more precise identification of the major specification, for inspection, debugging, or improved decoding.
[0120] Files would normally be externally identified (e.g. with a file extension or MIME type) that identifies the ‘best use’ (major brand), or the brand that the author believes will provide the greatest compatibility.
[0121] Syntax aligned(8) class GeneralTypeBox(code) extends Box(code) { unsigned int(32) major_brand; unsigned int(32) minor_version; unsigned int(32) compatible_brands[]; / / to end of the boxaligned(8) class FileTypeBox extends GeneralTypeBox fftyp'){}
[0122] SemanticsThis box identifies the specifications to which this file complies.Each brand is a four character code, registered with ISO, that identifies a precise specification. major_brand - is a brand identifier minor_version - is an informative integer for the minor version of the major brand compatible_brands - is a list, to the end of the box, of brands
[0123] A file-level FileTypeBox identifies which specification is the ‘best use’ of the file (in its major_brand syntax element), a minor version of that specification, and also a set of other specifications to which the file complies (the compatible_brands syntax element array). The major_brand value may be repeated in the compatible_brands list. Readers should attempt to read files that are marked as compatible with any of the specifications that the reader implements. A file-level FileTypeBox is usually placed early in the file.
[0124] In ISOBMFF, the ExtendedTypeBox may be placed after the FileTypeBox, any SegmentTypeBox, or any TrackTypeBox, or used as an item property to indicate that a reader should only process the file, the segment, the track, or the item, respectively, when it supports the processing requirements of all the brands in at least one of the included TypeCombinationBoxes, or at least one brand in the preceding FileTypeBox, SegmentTypeBox, or TrackTypeBox, or in the BrandProperty associated with the same item, respectively. The TypeCombinationBox expresses that the associated file, segment, track, or item may include any boxes or other code points required to be supported in any of the brands listed in the TypeCombinationBox and that the associated file, segment, track, or item complies with the intersection of the constraints of the brands listed in the TypeCombinationBox.
[0125] An ISO base media file may have further transformation of its box structure, for example, it may use compressed top-level boxes. In this case, the resulting file may no longer be compliant with the brand promises of the original file, as it requires the support for new tools (such as compressed top-level boxes). For example, a file using a compressed MovieBox is no longer compliant to any brand defined prior to the introduction of compressed boxes.
[0126] The OriginalFileTypeBox is used to encapsulate brand information applying to the original file before transformation but not valid in the transformed domain. There is at most one OriginalFileTypeBox after each FileTypeBox or SegmentTypeBox. An OriginalFileTypeBox includes the FileTypeBox and when present, the ExtendedTypeBox, of the file before transformation. In cases where a file uses multiple transformations, nested OriginalFileTypeBox is used, and there is at most one OriginalFileTypeBox in each Origin alFileT ypeB ox .
[0127] The processing model for a file reader is equivalent to reversing the transformation (e.g., decompressing a file with compressed boxes), removing both FileTypeBox / SegmentTypeBox and OriginalFileTypeBox, and inserting in their place the child boxes of the OriginalFileTypeBox. In the case of compressed top-level boxes, the resulting replacement and decompression of the file is a compliant uncompressed ISOBMFF file.
[0128] In files conforming to the ISO base media file format, the media data may be provided in one or more instances of MediaDataBox (‘mdat‘) and the MovieBox (‘moov’) may be used to enclose the metadata for timed media. In some cases, for a file to be operable, both of the ‘mdat’ and ‘moov’ boxes may be required to be present. The ‘moov’ box may include one or more tracks, and each track may reside in one corresponding TrackBox (‘trak’). Each track is associated with a handler, identified by a four-character code, specifying the track type. Video, audio, and image sequence tracks may be collectively called media tracks, and they include an elementary media stream. Other track types comprise hint tracks and timed metadata tracks.
[0129] Tracks comprise samples, such as audio or video frames. For video tracks, a media sample may correspond to a coded picture or an access unit.
[0130] A media track refers to samples (which may also be referred to as media samples) formatted according to a media compression format (and its encapsulation to the ISO base media file format). A hint track refers to hint samples, including cookbook instructions for constructing packets for transmission over an indicated communication protocol. A timed metadata track may refer to samples describing referred media and / or hint samples.
[0131] The 'trak' box includes in its hierarchy of boxes the SampleDescriptionBox, which gives detailed information about the coding type used, and any initialization information needed for that coding. The SampleDescriptionBox includes an entry-count and as many sample entriesas the entry-count indicates. The format of sample entries is track-type specific but derived from generic classes (e.g., VisualSampleEntry, AudioSampleEntry). Which type of sample entry form is used for derivation of the track-type specific sample entry format is determined by the media handler of the track.
[0132] A sample entry may comprise a decoder configuration box. A decoder configuration box may comprise a decoder configuration record, which may be defined as a data structure that may contain information that may assist in initializing decoder instance(s) for decoding the samples mapped to the sample entry.
[0133] The track reference mechanism may be used to associate tracks with each other. The TrackReferenceBox includes box(es), each of which provides a reference from the including track to a set of other tracks. These references are labeled through the box type (e.g., the four-character code of the box) of the included box(es).
[0134] The ISO Base Media File Format includes three mechanisms for timed metadata that may be associated with particular samples: sample groups, timed metadata tracks, and sample auxiliary information. A derived specification may provide similar functionality with one or more of these three mechanisms.
[0135] A sample grouping in the ISO base media file format and its derivatives, such as ISO / IEC 14496-15, may be defined as an assignment of each sample in a track to be a member of one sample group, based on a grouping criterion. A sample group in a sample grouping is not limited to being contiguous samples and may include non-adjacent samples. As there may be more than one sample grouping for the samples in a track, each sample grouping may have a type field to indicate the type of grouping. Sample groupings may be represented by two linked data structures: (1) a SampleToGroupBox (sbgp box) represents the assignment of samples to sample groups; and (2) a SampleGroupDescriptionBox (sgpd box) includes a sample group entry for each sample group describing the properties of the group. There may be multiple instances of the SampleToGroupBox and SampleGroupDescriptionBox based on different grouping criteria. These may be distinguished by a type field used to indicate the type of grouping. SampleToGroupBox may comprise a grouping_type_parameter field that may be used, e.g., to indicate a sub-type of the grouping.
[0136] In ISOMBFF, an edit list provides a mapping between the presentation timeline and the media timeline. Among other things, an edit list provides for the linear offset of thepresentation of samples in a track, provides for the indication of empty times and provides for a particular sample to be dwelled on for a certain period of time. The presentation timeline may be accordingly modified to provide for looping, such as for the looping videos of the various regions of the scene. One example of the box that includes the edit list, the EditListBox, is provided below: aligned(8) class EditListBox extends FullBox(‘elst’, version, flags) { unsigned int(32) entry_count; for (i=l; i <= entry_count; i++) { if (version==l) { unsigned int(64) segment_duration; int(64) media lime;} else { / / version==0 unsigned int(32) segment_duration; int(32) media lime;} int(16) media_rate_integer; int(16) media_rate_fraction = 0;}}
[0137] In ISOBMFF, an EditListBox may be included in EditBox, which is included in a TrackBox ('trak').
[0138] In this example of the edit list box, flags specifies the repetition of the edit list. By way of example, setting a specific bit within the box flags (the least significant bit, e.g., flags & 1 in ANSI-C notation, where & indicates a bit- wise AND operation) equal to 0 specifies that the edit list is not repeated, while setting the specific bit (e.g., flags & 1 in ANSI-C notation) equal to 1 specifies that the edit list is repeated. The values of box flags greater than 1 may be defined to be reserved for future extensions. As such, when the edit list box indicates the playback of zero or one samples, (flags & 1) shall be equal to zero. When the edit list is repeated, the media at time 0 resulting from the edit list follows immediately the media having the largest time resulting from the edit list such that the edit list is repeated seamlessly.
[0139] In ISOBMFF, a Track group enables grouping of tracks based on certain characteristics or the tracks within a group have a particular relationship. Track grouping, however, does not allow any image items in the group.
[0140] The syntax of TrackGroupBox in ISOBMFF is as follows: aligned(8) class TrackGroupBox extends Box('trgr') {} aligned(8) class TrackGroupTypeBox(unsigned int(32) track_group_type) extends FullBox(track_group_type, version = 0, flags = 0){ unsigned int(32) track_group_id; / / the remaining data may be specified for a particular track_group_type}
[0141] track_group_type indicates the grouping type and may be set to one of the following values, or a value registered, or a value from a derived specification or registration:- 'msrc' indicates that this track belongs to a multi-source presentation. The tracks that have the same value of track_group_id within a TrackGroupTypeBox oftrack_group_type 'msrc' are mapped as being originated from the same source. For example, a recording of a video telephony call may have both audio and video for both participants, and the value of track_group_id associated with the audio track and the video track of one participant differs from value of track_group_id associated with the tracks of the other participant.- The pair of track_group_id and track_group_type identifies a track group within the file. The tracks that include a particular TrackGroupTypeBox having the same value of track_group_id and track_group_type belong to the same track group.
[0142] The Entity grouping is similar to track grouping but enables grouping of both tracks and image items in the same group.
[0143] The syntax of Entity ToGroupB ox in ISOBMFF is as follows: aligned(8) class EntityToGroupBox(grouping_type, version, flags) extends FullBox(grouping_type, version, flags) { unsigned int(32) group_id; unsigned int(32) num_entities_in_group; for(i=0; i<num_entities_in_group; i++) unsigned int(32) entity _id;}
[0144] group_id is a non-negative integer assigned to the particular grouping that shall not be equal to any group_id value of any other EntityToGroupBox, any item_ID value of the hierarchy level (file, movie, or track) that includes the GroupsListBox, or any track_ID value (when the GroupsListBox is included in the file level).
[0145] num_entities_in_group specifies the number of entity_id values mapped to this entity group.
[0146] entity _id is resolved to an item, when an item with item_ID equal to entity _id is present in the hierarchy level (file, movie or track) that includes the GroupsListBox, or to a track, when a track with track_ID equal to entity _id is present and the GroupsListBox is included in the file level.
[0147] Files conforming to the ISOBMFF may include any non-timed objects, referred to as items, meta items, or metadata items, in a meta box (four-character code: ‘meta’). While the name of the meta box refers to metadata, items may generally include metadata or media data. The meta box may reside at the top level of the file, within a movie box (four-character code: ‘moov’), and within a track box (four-character code: ‘trak’), but at most one meta box may occur at each of the file level, movie level, or track level. The meta box may be required to include a ' hd I r ’ box indicating the structure or format of the ‘meta’ box contents. The meta box may list and characterize any number of items that may be referred and each one of them may be associated with a file name and are uniquely identified with the file by item identifier (item_id) which is an integer value. The metadata items may be for example stored in the ’idat’ box of the meta box or in an ’mdat’ box or reside in a separate file. When the metadata is located external to the file then its location may be declared by the DatalnformationBox (four-character code: ‘dinf ). In the specific case that the metadata is formatted using extensible Markup Language (XML) syntax and is required to be stored directly in the MetaBox, the metadata may be encapsulated into either the XMLBox (four-character code: ‘xml ‘) or the Binary XMLBox (four-character code: ‘bxml’). An item may be stored as a contiguous byte range, or it may be stored in several extents, each being a contiguous byte range. In other words, items may be stored fragmented into extents, e.g. to enable interleaving. An extent is a contiguous subset of the bytes of the resource. The resource may be formed by concatenating the extents.
[0148] A common base structure is used to include general untimed metadata. This structure is called the MetaBox as it was originally designed to carry metadata, i.e. data that is annotating other data. However, it is now used for a variety of purposes including the carriage of data that is not annotating other data, especially when present at ‘file level’.
[0149] In some versions of ISOBMFF, the MetaBox is required to include a HandlerBox indicating the structure or format of the MetaBox contents..
[0150] All other included boxes are specific to the format specified by the HandlerBox.
[0151] The other boxes defined here may be defined as optional or mandatory for a given format. When they are used, then they shall take the form specified here. These optional boxes include a DatalnformationBox, which documents other files in which metadata values (e.g. pictures) are placed, and an ItemLocationBox, which documents where in those files each item is located (e.g. in the common case of multiple pictures stored in the same file).
[0152] At most one MetaBox may occur at each of the file level, segment, movie level, or track level.
[0153] When an ItemProtectionBox occurs, then some or all of the metadata, including possibly the primary resource, may have been protected and be un-readable unless the protection system is taken into account.
[0154] In ISOBMFF, the MetaBox is unusual in that it is a container box yet extends FullBox, not Box. In QuickTime file format, which uses the same registration space for four- character codes as ISOBMFF, the MetaBox ('meta') is structurally equivalent to a Box, i.e., excludes the version and flags fields, yet uses the same 4CC as in ISOBMFF. Parsers may need to conclude whether ISOBMFF definition of MetaBox (based on FullBox) or the QuickTime file format definition of MetaBox has been used in a file.
[0155] Metadata items are identified by item_ID. Within a given MetaBox, a given item_ID shall uniquely refer to a single item. When an item is updated in movie fragments, the item_ID refers to the latest received version.
[0156] Derived specifications may further restrict the criteria for uniqueness: unique among the item_IDs in both file and movie-level boxes, or unique within that set extended with the track_ID of the tracks in a movie box. The item_ID value of 0 should not be used, and shall not be used when the set is extended to include track_IDs.
[0157] There are three scopes for item_IDs: file and segments; MovieBox and MovieFragmentBox; and TrackBox and TrackFragmentBox. In other words, there shall be only one item with a given item_ID within a given scope (e.g. in the TrackBox and all TrackFragmentBox with the same track_ID). aligned(8) class MetaBox (handler_type)extends FullBox('meta', version = 0, 0) {HandlerBox(handler_type) theHandler;Primary ItemB ox primary _resource; / / optionalDatalnformationBox file_locations; / / optionalItemLocationBox item_locations; / / optionalItemProtectionBox protections; / / optionalItemlnfoBox item_infos; / / optionalIPMPControlBox IPMP_control; / / optionalItemReferenceBox item_refs; / / optionalItemDataBox item_data; / / optionalBox other_boxes[]; / / optional
[0158] The structure or format of the metadata is declared by the handler. In the case that the primary data is identified by a primary item, and that primary item has an item information entry with an item_type, the handler type may be the same as the item_type.
[0159] The ItemPropertiesBox enables the association of any item with an ordered set of item properties. Item properties may be regarded as small data records. The ItemPropertiesBox includes two parts: ItemPropertyContainerBox that includes an implicitly indexed list of item properties, and one or more ItemProperty Associations ox(es) that associate items with item properties.
[0160] High Efficiency Image File Format (HEIF)
[0161] High Efficiency Image File Format (HEIF) is a standard developed by the Moving Picture Experts Group (MPEG) for storage of images and image sequences. Among other things, the standard facilitates file encapsulation of data coded according to the High Efficiency Video Coding (HEVC) standard. HEIF includes features building on top of the used ISO Base Media File Format (ISOBMFF).
[0162] The ISOBMFF structures and features are used to a large extent in the design of HEIF. The basic design for HEIF comprises still images that are stored as items and image sequences that are stored as tracks. An item in HEIF is defined as the data that does not require timed processing, as opposed to sample data, and is described by the boxes included in a MetaBox
[0163] In the context of HEIF, the following boxes may be included within the root-level 'meta' box and may be used as described in the following. In HEIF, the handler value of the Handler box of the 'meta' box is 'picf. The resource (whether within the same file, or in an external file identified by a uniform resource identifier) including the coded media data is resolved through the Data Information (’dinf) box, whereas the Item Location filoc') box stores the position and sizes of every item within the referenced file. The Item Reference firef) box documents relationships between items using typed referencing. When there is an item among a collection of items that is in some way to be considered the most important compared to others then this item is signaled by the Primary Item fpitm') box. Apart from the boxes mentioned here, the 'meta' box is also flexible to include other boxes that may be necessary to describe items.
[0164] Any number of image items may be included in the same file. Given a collection of images stored by using the 'meta' box approach, it sometimes is essential to qualify certain relationships between images. Examples of such relationships include indicating a cover image for a collection, providing thumbnail images for some or all of the images in the collection, and associating some or all of the images in a collection with an auxiliary image such as an alpha plane. A cover image among the collection of images is indicated using the 'pitm' box. A thumbnail image or an auxiliary image is linked to the primary image item using an item reference of type 'thmb' or 'auxl', respectively.
[0165] A decoder configuration item property (a.k.a. a configuration item property) may comprise information that may assist in initializing the decoder(s) that may be used to decoderthe image item(s) associated with the decoder configuration item property. The decoder configuration item property may comprise a decoder configuration record.
[0166] Predict! vely coded image items have a decoding dependency to one or more other coded image items. An example for such an image item could be a P frame stored as an image item in a burst entity group that has IPPP. . . structure, with the P frames dependent only on the preceding I frames.
[0167] Capability to have predict! vely coded image items has certain benefits especially in content re-editing and cover image selection:- Image sequences may be converted to image items with no transcoding.- Any sample of an image sequence track may be selected as a cover image. The cover image does not need to be intra-coded.- Devices that do not have a video or image encoder are capable of updating the cover image of a file including an image sequence track.- Storage efficiency is further achieved by re-using the predictively coded picture rather than re-encoding it as I frame and storing as an additional image item. Moreover, image quality degradation is also avoided.- Re-encoding might not be allowed or preferred by the copyright owner. Predictively coded image items avoid the need of re-encoding of any image from an image sequence track.
[0168] Predictively coded image items are linked to the coded image items they directly and indirectly depend on by item references of type 'pred'. The list of referenced items in item references of type 'pred' shall indicate the decoding order. When concatenated, the encoded media data of items with item_ID equal to to_item_ID for all values of j from 0 to reference_count - 1, inclusive, in increasing order of j, followed by the item with item_ID equal to from_item_ID shall form a bitstream that conforms to the decoder configuration item property of the predictively coded image item.
[0169] In order to decode the predictively coded image item, there shall be no other decoding dependencies other than the image items referenced by item references of type 'pred'.
[0170] The predictively coded image item shall be associated with exactly one RequiredReferenceTypesProperty including one reference type with the value 'pred'.
[0171] A file having the 'pred' brand in the compatible_brands of a TypeCombinationBox associated with the FileTypeBox may include predictively coded image items and shall conform to following constraints:- when 'mifl' brand is among the compatible brands array of the FileTypeBox then the primary item shall be independently coded. Additionally, an alternate group including this primary item and possibly predictively coded image items may exist.- when the 'pred' brand is in the compatible_brands of a TypeCombinationBox associated with the FileTypeBox, and the 'mifl' brand is not among the compatible brands array of the FileTypeBox then the primary item and possibly all items from the alternate group including the primary item may be predictively coded image items- For each predictively coded image item present in the file, the file shall also include all items that the predictively coded image item depends on (by item references of type 'pred').
[0172] The RequiredReferenceTypesProperty descriptive item property lists the item reference types that a reader shall understand and process to decode the associated image item. The respective essential flag shall be equal to 1 in ItemPropertyAssociationBox.
[0173] In the absence of this property, required reference types are not explicitly listed, but may still exist.
[0174] Syntax aligned(8) class RequiredReferenceTypesProperty extends ItemFullPropertyfrref, version = 0, flags = 0){ unsigned int(8) reference_type_count; for (i=0; i< reference_type_count; i++) { unsigned int(32) reference_type[i];}}
[0175] Semantics
[0176] reference_type_count indicates the number of reference types that are required to understand and process to decode the associated image item.
[0177] reference_type[i] indicates a reference type that is required to understand and process to decode.
[0178] Internet media types
[0179] Internet media types, also known as Multipurpose Internet Mail Extension (MIME) types, are used by various applications to identify the type of a resource or a file. MIME types consist of a media type, a subtype, and zero or more optional parameters.
[0180] As described, MIME is an extension to an email protocol which makes it possible to transmit and receive different kinds of data files on the Internet, for example, video, audio, images, software, and the like. An internet media type is an identifier used on the Internet to indicate the type of data that a file contains. Such internet media types may also be called as content types. Several MIME type / subtype combinations exist that may indicate different media formats. Content type information may be included by a transmitting entity in a MIME header at the beginning of a media transmission. A receiving entity thus may need to examine the details of such media content to determine when the specific elements may be rendered given an available set of codecs. Especially, when the end system has limited resources, or the connection to the end systems has limited bandwidth, it may be helpful to know from the content type alone when the content may be rendered.
[0181] Two optional parameters, ‘codecs’ and ‘profiles’, are specified to be used with various MIME types or type / subtype combinations to allow for unambiguous specification of the codecs employed by the media formats included within, or the profile(s) of the overall container format.
[0182] By labelling content with the specific codecs indicated to render the included media, receiving systems may determine when the codecs are supported by the end system, and when not, may take appropriate action (such as rejecting the content, sending notification of thesituation, transcoding the content to a supported type, fetching and installing the required codecs, further inspection to determine when it will be sufficient to support a subset of the indicated codecs, and the like). For file formats derived from the ISOBMFF, the codecs parameter may be considered to comprise a comma-separated list of one or more list items. When a list item of the codecs parameter represents a track of an ISOBMFF compliant file, the list item may comprise a four-character code of the sample entry of the track.
[0183] The profiles MIME parameter may provide an overall indication, to the receiver, of the specifications with which the content complies. This is an indication of the compatibility of the container format and its contents to some specification. The receiver may be able to work out the extent to which it may handle and render the content by examining to see which of the declared profiles it supports, and what they mean. The profiles parameter for an ISOBMFF file may be specified to comprise a list of the compatible brands included in the file.
[0184] Uniform Resource Identifier (URI)
[0185] A uniform resource identifier (URI) may be defined as a string of characters used to identify a name of a resource. Such identification enables interaction with representations of the resource over a network, using specific protocols. A URI is defined through a scheme specifying a concrete syntax and associated protocol for the URI. A URI comprises a scheme part (identifying e.g. the protocol for the URI) and a hierarchical part identifying the resource, and these two parts are separated by a colon character. A URI may optionally comprise a query part (separated by the character '?') and / or a fragment part (separated by the character '#'). The uniform resource locator (URL) and the uniform resource name (URN) are forms of URI. A URL may be defined as a URI that identifies a web resource and specifies the means of acting upon or obtaining the representation of the resource, specifying both its primary access mechanism and network location. A URN may be defined as a URI that identifies a resource by name in a particular namespace. A URN may be used for identifying a resource without implying its location or how to access it.
[0186] URL fragment identifiers (which may also be referred to as URL forms) may be specified for a particular content type to access a part of the resource, such as a file, indicated by the base part of the URL (without the fragment identifier). URL fragment identifiers may be identified for example by a hash ('#') character within the URL. For the ISOBMFF, it may be specified that URL fragments "#X" refer to a track with track_ID equal to X, "#item_ID=" and "#item_name=" refer to file level meta box(es), "# / item_ID=" and "# / item_name=" refer to themeta box(es) in the Movie box, and "#track_ID=X / item_ID=" and "#track_ID=X / item_name=" refer to meta boxes in the track with track_ID equal to X, including the meta boxes potentially found in movie fragments.
[0187] Fundamentals of video / image coding
[0188] Video codec includes an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that may decompress the compressed video representation back into a viewable form. Typically encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).
[0189] Hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, e.g., the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder may control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).
[0190] In temporal prediction, the sources of prediction are previously decoded pictures (a.k.a. reference pictures). In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction is applied similarly to temporal prediction but the reference picture is the current picture and only previously decoded samples may be referred in the prediction process. Interlayer or inter- view prediction may be applied similarly to temporal prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal prediction only, while in other cases inter prediction may refer collectively to temporal prediction and any of intra block copy, inter-layer prediction, and inter-view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0191] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction may be performed in spatial or transform domain, i.e., either sample values or transform coefficients may be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0192] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters may be entropy-coded more efficiently when they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0193] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organisation for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multi view Video Coding (MVC).
[0194] The High Efficiency Video Coding standard (which may be abbreviated HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV-HEVC, 3D-HEVC and REXT thathave been made for the purpose of understanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.
[0195] SHVC, MV-HEVC, and 3D-HEVC use a common basis specification, specified in Annex F of the version 2 of the HEVC standard. This common basis comprises for example high-level syntax and semantics, e.g., specifying some of the characteristics of the layers of the bitstream, such as inter-layer dependencies, as well as decoding processes, such as reference picture list construction including inter-layer reference pictures and picture order count derivation for multi-layer bitstream. Annex F may also be used in potential subsequent multilayer extensions of HEVC. It is to be understood that even though a video encoder, a video decoder, encoding methods, decoding methods, bitstream structures, and / or embodiments may be described in the following with reference to specific extensions, such as SHVC and / or MV- HEVC, they are generally applicable to any multi-layer extensions of HEVC, and even more generally to any multi-layer video coding scheme.
[0196] The Versatile Video Coding standard (which may be abbreviated VVC, H.266, or H.266 / VVC) was developed by the Joint Video Experts Team (JVET), which is a collaboration between the ISO / IEC MPEG and ITU-T VCEG. VVC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3. Extensions to VVC are presently under development.
[0197] A specification of the AVI bitstream format and decoding process were developed by the Alliance of Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.
[0198] An elementary unit for the input to an encoder and the output of a decoder, respectively, in most cases is a picture. A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.
[0199] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:- Luma (Y) only (monochrome).- Luma and two chroma (YCbCr or YCgCo).- Green, Blue and Red (GBR, also known as RGB).- Arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).
[0200] In the following, these arrays may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr or Cg and Co; regardless of the actual color representation method in use. The actual color representation method in use may be indicated, e.g., in a coded bitstream e.g., using the Video Usability Information (VUI) syntax. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) or the array or a single sample of the array that compose a picture in monochrome format.
[0201] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.
[0202] Some chroma formats may be summarized as follows:- In monochrome sampling there is only one sample array, which may be nominally considered the luma array.- In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.- In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.- In 4:4:4 sampling when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.
[0203] Coding formats or standards may allow to code sample arrays as separate color planes into the bitstream and respectively decode separately coded color planes from the bitstream. When separate color planes are in use, each one of them is separately processed (by the encoder and / or the decoder) as a picture with monochrome sampling.
[0204] Video codec includes an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically encoder discards some information in the original video sequence in order to represent the video in a morecompact form (that is, at lower bitrate).
[0205] Typical hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).
[0206] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures (a.k.a. reference pictures).
[0207] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter-view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0208] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0209] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0210] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0211] Image and video codecs may use a set of filters, which may enhance the visual quality of the predicted visual content. Filters may be applied either in-loop or out-of-loop, or both. In-loop filters (which may be also called loop filters) are used in reconstructing prediction reference that may be used for predicting forthcoming video signal. In other words, in the case of in-loop filters, the filter applied on one block in the currently encoded frame may affect the encoding of another block in the same frame and / or in another frame which is predicted from the current frame. An in-loop filter may affect the bitrate and / or the visual quality. In fact, an enhanced block may cause a smaller residual (difference between original block and predicted- and-filtered block), thus requiring less bits to be encoded. An out-of-the loop filter (which may also be called a post-processing filter or a post-filter) may be applied on a frame or part of a frame after it has been reconstructed, the filtered visual content may not be used as a source for prediction, and thus it may only impact the visual quality of the frames that are output by the decoder.
[0212] In typical video codecs the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decodedpictures.
[0213] In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks.
[0214] Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in temporal reference picture.
[0215] Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signalled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.
[0216] In typical video codecs the prediction residual after motion compensation is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.
[0217] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor X to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + Rwhere C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).
[0218] An out-of-band transmission, signaling, or storage may refer to the capability of transmitting, signaling, or storing information in a manner that associates the information with a video bitstream. The out-of-band transmission may use a more reliable transmission mechanism compared to the protocols used for carrying coded video data, such as slices. The out-of-band transmission, signaling or storage may additionally or alternatively be used, e.g., for ease of access or session negotiation. For example, a sample entry of a track in a file conforming to the ISO Base Media File Format may comprise parameter sets, while the coded data in the bitstream is stored elsewhere in the file or in another file. Another example of out- of-band transmission, signaling, or storage comprises including information, such as NN and / or NN updates in a file format track that is separate from track(s) including coded video data.
[0219] The phrase along the bitstream (e.g., indicating along the bitstream) or along a coded unit of a bitstream (e.g., indicating along a coded tile) may be used in claims and described embodiments to refer to transmission, signaling, or storage in a manner that the ‘out- of-band’ data is associated with, but not included within, the bitstream or the coded unit, respectively. The phrase decoding along the bitstream or along a coded unit of a bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of- band transmission, signaling, or storage) that is associated with the bitstream or the coded unit, respectively. For example, the phrase along the bitstream may be used when the bitstream is included in a container file, such as a file conforming to the ISO Base Media File Format, and certain file metadata is stored in the file in a manner that associates the metadata to the bitstream, such as boxes in the sample entry for a track including the bitstream, a sample group for the track including the bitstream, or a timed metadata track associated with the track including the bitstream. In another example, the phrase along the bitstream may be used when the bitstream is made available as a stream over a communication protocol and a media description, such as a streaming manifest, is provided to describe the stream.
[0220] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.
[0221] A syntax element may be defined as an element of data represented in a bitstream.A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.
[0222] Syntax structures may be specified, for example, using arithmetic, logical, relational, bit- wise, and assignment operators similar to those available in many programming languages. For example, & may indicate a bit-wise ‘AND’ operation. Furthermore, syntax structures may be specified with reference to mathematical functions.
[0223] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper case letter and without any underscore characters. Variables starting with an upper case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning the originating syntax structure of the variable. Variables starting with a lower case letter may only be used in relation to the syntax structure or function they have been defined for.
[0224] An elementary unit for the output of a video encoder and the input of a video decoder, respectively, may be a network abstraction layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format encapsulating NAL units may be used for transmission or storage environments that do not provide framing structures. The bytestream format may separate NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders may run a byte-oriented start code emulation prevention algorithm, which may add an emulation prevention byte to the NAL unit payload, when a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet and stream-oriented systems, start code emulation prevention may be performed regardless of whether the bytestream format is in use or not. A NAL unit may be defined as a syntax structure including an indication of the type of data to follow and bytes including that data in the form of a raw byte sequence payload interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure including an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits including syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0225] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.
[0226] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.
[0227] In some formats or standards, a first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams.
[0228] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, the header may comprise a size of the payload in bytes.
[0229] In some coding standards, NAL units include a header and payload. The NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g., called nuh_layer_id in H.265 / HEVC and H.266 / VVC), which may be used, e.g., for indicating spatial or quality layers, views of a multiview video, or auxiliary layers (such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30-frames-per-second subset of a 60- frames-per-second bitstream.
[0230] Bitstreams or coded video sequences may be encoded to be temporally scalable as follows. Each picture may be assigned to a particular temporal sub-layer. A temporal sublayer may be equivalently called a sub-layer, temporal sublayer, sublayer, or temporal level. Temporal sub-layers may be enumerated, e.g., from 0 upwards. The lowest temporal sub-layer, sub-layer 0, may be decoded independently. Pictures at temporal sub-layer 1 may be predicted from reconstructed pictures at temporal sub-layers 0 and 1. Pictures at temporal sub-layer 2 may be predicted from reconstructed pictures at temporal sub-layers 0, 1, and 2, and so on. Inother words, a picture at temporal sub-layer N does not use any picture at temporal sub-layer greater than N as a reference for inter prediction. The bitstream created by excluding all pictures greater than or equal to a selected sub-layer value and including pictures remains conforming.
[0231] Each picture of a temporally scalable bitstream may be assigned with a temporal identifier (also known as temporal layer identifier, temporal sublayer identifier, or temporal layer ID), which may be, for example, assigned to a variable Temporalld. The temporal identifier may, for example, be indicated in a NAL unit header or in an OBU extension header. Temporalld equal to 0 corresponds to the lowest temporal level. The bitstream created by excluding all coded pictures having a Temporalld greater than or equal to a selected value and including all other coded pictures remains conforming. Consequently, a picture having Temporalld equal to tid_value does not use any picture having a Temporalld greater than tid_value as a prediction reference.
[0232] NAL units may be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units.
[0233] A non-VCL NAL unit may be, for example, one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0234] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or alike and the latter type can end a picture unit or alike. An SEI NAL unit includes one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may include the syntax and semantics for the specified SEI messages but a process for handling the messages in therecipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.
[0235] Some video coding specifications enable metadata OBUs. A metadata OBU comprises a type field, which specifies the type of metadata.
[0236] A coded video sequence (CVS) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.
[0237] A coded layer video sequence (CLVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh_layer_id in VVC) that is decodable independently of other pictures in the same layer.
[0238] Some codecs use a concept of picture order count (POC). A value of POC is derived for each picture and is non-decreasing with increasing picture position in output order. POC therefore indicates the output order of pictures. POC may be used in the decoding process for example for implicit scaling of motion vectors and for reference picture list initialization. Furthermore, POC may be used in the verification of output order conformance. The variable including a POC value of a picture may be referred to as PicOrderCntVal.
[0239] A Decoded Picture Buffer (DPB) may be used in the encoder and / or in the decoder. There may be two reasons to buffer decoded pictures, for references in inter prediction and for reordering decoded pictures into output order. Some coding formats, such as HEVC, provide a great deal of flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Hence, the DPB may include a unified decoded picture buffering process for reference pictures and output reordering. A decoded picture may be removed from the DPB when it is no longer used as a reference and is not needed for output.
[0240] Output order may be defined as the order in which the decoded pictures are output from the decoded picture buffer (for the decoded pictures that are to be output from the decoded picture buffer).
[0241] Output time may be defined as a time when a decoded picture is to be output from a decoder or from the DPB of a decoder (for the decoded pictures that are to be output from the DPB), for example as specified by a hypothetical reference decoder specification according to the output timing DPB operation.
[0242] Pictures having the same output order may be defined to mean the same as pictures having the same output time.
[0243] Decoding order may be defined as the order in which syntax elements are processed by the decoding process. It may be required that syntax elements are ordered in a bitstream in their decoding order.
[0244] An identifier may be defined as a syntax element that identifies a syntax structure. A value of the identifier may for example differ in different instances of the same syntax structure, such as a parameter set. A particular instance of the syntax structure may be referenced through its identifier value. For example, a parameter set that is referenced by the (de)coding of a coded video slice may be identified by providing the identifier value of the parameter set in a header of the coded video slice.
[0245] An indicator (ide) may be defined as a syntax element whose value indicates a selection among more than two values (for which semantics have been specified). An indicator syntax element may have _idc postfix in its name.
[0246] A uniform resource identifier (URI) may be defined as a string of characters used to identify a name of a resource. Such identification enables interaction with representations of the resource over a network, using specific protocols. A URI is defined through a scheme specifying a concrete syntax and associated protocol for the URI. The uniform resource locator (URL) and the uniform resource name (URN) are forms of URI. A URL may be defined as a URI that identifies a web resource and specifies the means of acting upon or obtaining the representation of the resource, specifying both its primary access mechanism and network location. A URN may be defined as a URI that identifies a resource by name in a particular namespace. A URN may be used for identifying a resource without implying its location orhow to access it.Scalable video coding may refer to coding structure where one bitstream may include multiple representations of the content, for example, at different bitrates, resolutions or frame rates. In these cases, the receiver may extract the desired representation depending on its characteristics (e.g., resolution that matches best the display device). Alternatively, a server or a network element may extract the portions of the bitstream to be transmitted to the receiver depending on e.g., the network characteristics or processing capabilities of the receiver. A meaningful decoded representation may be produced by decoding only certain parts of a scalable bit stream. A scalable bitstream may consist of a “base layer” providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer typically depends on the lower layers. E.g., the motion and mode information of the enhancement layer may be predicted from lower layers. Similarly, the pixel data of the lower layers may be used to create prediction for the enhancement layer.
[0247] In some scalable video coding schemes, a video signal may be encoded into a base layer and one or more enhancement layers. An enhancement layer may enhance, for example, the temporal resolution (i.e., the frame rate), the spatial resolution, or simply the quality of the video content represented by another layer or part thereof. Each layer together with all its dependent layers is one representation of the video signal, for example, at a certain spatial resolution, temporal resolution, and quality level. In this document, a scalable layer together with all of its dependent layers is referred to as a “scalable layer representation”. The portion of a scalable bitstream corresponding to a scalable layer representation may be extracted and decoded to produce a representation of the original signal at certain fidelity.
[0248] Scalability modes or scalability dimensions may include but are not limited to the following:Quality scalability: Base layer pictures are coded at a lower quality than enhancement layer pictures, which may be achieved for example using a greater quantization parameter value (i.e., a greater quantization step size for transform coefficient quantization) in the base layer than in the enhancement layer. Quality scalability may be further categorized into fine-grain or fine-granularity scalability (FGS), medium-grain or medium-granularity scalability (MGS), and / or coarse-grain or coarse-granularity scalability (CGS), as described below.Spatial scalability: Base layer pictures are coded at a lower resolution (e.g., have fewer samples) than enhancement layer pictures. Spatial scalability and quality scalability, particularly its coarse-grain scalability type, may sometimes be considered the same type of scalability.View scalability, which may also be referred to as multiview coding. The base layer represents a first view, whereas an enhancement layer represents a second view. A view may be defined as a sequence of pictures representing one camera or viewpoint. It may be considered that in stereoscopic or two-view video, one video sequence or view is presented for the left eye while a parallel view is presented for the right eye.Depth scalability, which may also be referred to as depth-enhanced coding. A layer or some layers of a bitstream may represent texture view(s), while other layer or layers may represent depth view(s).
[0249] It should be understood that many of the scalability types may be combined and applied together.
[0250] The term “layer” may be used in context of any type of scalability, including view scalability and depth enhancements. An enhancement layer may refer to any type of an enhancement, such as SNR, spatial, multiview, and / or depth enhancement. A base layer may refer to any type of a base video sequence, such as a base view, a base layer for SNR / spatial scalability, or a texture base view for depth-enhanced video coding.
[0251] A sender, a gateway, a client, or another entity may select the transmitted layers and / or sub-layers of a scalable video bitstream. Terms layer extraction, extraction of layers, or layer down-switching may refer to transmitting fewer layers than what is available in the bitstream received by the sender, the gateway, the client, or another entity. Layer up-switching may refer to transmitting additional layer(s) compared to those transmitted prior to the layer up-switching by the sender, the gateway, the client, or another entity, e.g., restarting the transmission of one or more layers whose transmission was ceased earlier in layer downswitching. Similarly, to layer down-switching and / or up-switching, the sender, the gateway, the client, or another entity may perform down- and / or up-switching of temporal sub-layers. The sender, the gateway, the client, or another entity may also perform both layer and sub-layer down-switching and / or up-switching. Layer and sub-layer down-switching and / or up-switching may be carried out in the same access unit or alike (e.g., virtually simultaneously) or may be carried out in different access units or alike (e.g., virtually at distinct times).
[0252] A scalable video encoder for quality scalability (also known as Signal-to-Noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non-scalable video encoder and decoder may be used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer and / or reference picture lists for an enhancement layer. In case of spatial scalability, the reconstructed / decoded baselayer picture may be upsampled prior to its insertion into the reference picture lists for an enhancement-layer picture. The base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picture similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a baselayer reference picture as an inter prediction reference and indicate its use with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as an inter prediction reference for the enhancement layer. When a decoded base-layer picture is used as the prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.
[0253] While the previous paragraph described a scalable video codec with two scalability layers with an enhancement layer and a base layer, it needs to be understood that the description may be generalized to any two layers in a scalability hierarchy with more than two layers. In this case, a second enhancement layer may depend on a first enhancement layer in encoding and / or decoding processes, and the first enhancement layer may therefore be regarded as the base layer for the encoding and / or decoding of the second enhancement layer. Furthermore, it needs to be understood that there may be inter-layer reference pictures from more than one layer in a reference picture buffer or reference picture lists of an enhancement layer, and each of these inter-layer reference pictures may be considered to reside in a base layer or a reference layer for the enhancement layer being encoded and / or decoded. Furthermore, it needs to be understood that other types of inter-layer processing than referencelayer picture upsampling may take place instead or additionally. For example, the bit-depth of the samples of the reference- layer picture may be converted to the bit-depth of the enhancement layer and / or the sample values may undergo a mapping from the color space of the reference layer to the color space of the enhancement layer.
[0254] A scalable video coding and / or decoding scheme may use multi-loop coding and / or decoding, which may be characterized as follows. In the encoding / d ecoding, a base layer picture may be reconstructed / decoded to be used as a motion-compensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as a reference forinter-layer (or inter-view or inter-component) prediction. The reconstructed / decoded base layer picture may be stored in the DPB. An enhancement layer picture may likewise be reconstructed / decoded to be used as a motion-compensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as reference for inter-layer (or interview or inter-component) prediction for higher enhancement layers, when any. In addition to reconstructed / decoded sample values, syntax element values of the base / reference layer or variables derived from the syntax element values of the base / reference layer may be used in the inter-layer / inter-component / inter-view prediction.
[0255] Inter-layer prediction may be defined as prediction in a manner that is dependent on data elements (e.g., sample values or motion vectors) of reference pictures from a different layer than the layer of the current picture (being encoded or decoded). Many types of inter-layer prediction exist and may be applied in a scalable video encoder / decoder. The available types of inter-layer prediction may for example depend on the coding profile according to which the bitstream or a particular layer within the bitstream is being encoded or, when decoding, the coding profile that the bitstream or a particular layer within the bitstream is indicated to conform to. Alternatively or additionally, the available types of inter-layer prediction may depend on the types of scalability or the type of a scalable codec or video coding standard amendment (e.g. SHVC, MV-HEVC, or 3D-HEVC) being used.
[0256] A direct reference layer may be defined as a layer that may be used for interlayer prediction of another layer for which the layer is the direct reference layer. A direct predicted layer may be defined as a layer for which another layer is a direct reference layer. An indirect reference layer may be defined as a layer that is not a direct reference layer of a second layer but is a direct reference layer of a third layer that is a direct reference layer or indirect reference layer of a direct reference layer of the second layer for which the layer is the indirect reference layer. An indirect predicted layer may be defined as a layer for which another layer is an indirect reference layer. An independent layer may be defined as a layer that does not have direct reference layers. In other words, an independent layer is not predicted using inter-layer prediction. A non-base layer may be defined as any other layer than the base layer, and the base layer may be defined as the lowest layer in the bitstream. An independent non-base layer may be defined as a layer that is both an independent layer and a non-base layer.
[0257] Similarly, to MVC, in MV-HEVC, inter-view reference pictures may be included in the reference picture list(s) of the current picture being coded or decoded. SHVC uses multiloop decoding operation (unlike the SVC extension of H.264 / AVC). SHVC may be consideredto use a reference index based approach, e.g., an inter-layer reference picture may be included in a one or more reference picture lists of the current picture being coded or decoded (as described above).
[0258] For the enhancement layer coding, the concepts and coding tools of HEVC base layer may be used in SHVC, MV-HEVC, and / or alike. However, the additional inter-layer prediction tools, which employ already coded data (including reconstructed picture samples and motion parameters a.k.a motion information) in reference layer for efficiently coding an enhancement layer, may be integrated to SHVC, MV-HEVC, and / or alike codec.
[0259] Video coding specifications may include a set of constraints for associating data units (e.g., NAL units in H.264 / AVC or HEVC) into access units. These constraints may be used to conclude access unit boundaries from a sequence of NAL units. For example, the following is specified in the HEVC standard:An access unit includes one coded picture with nuh_layer_id equal to 0, zero or more VCL NAL units with nuh_layer_id greater than 0 and zero or more non-VCL NAL units.Let firstBIPicNalUnit be the first VCL NAL unit of a coded picture with nuh_layer_id equal to 0. The first of any of the following NAL units preceding firstBIPicNalUnit and succeeding the last VCL NAL unit preceding firstBIPicNalUnit, when any, specifies the start of a new access unit: o access unit delimiter NAL unit with nuh_layer_id equal to 0 (when present), o VPS NAL unit with nuh_layer_id equal to 0 (when present), o SPS NAL unit with nuh_layer_id equal to 0 (when present), o PPS NAL unit with nuh_layer_id equal to 0 (when present), o Prefix SEI NAL unit with nuh_layer_id equal to 0 (when present), o NAL units with nal_unit_type in the range of RSV_NVCL41..RSV_NVCL44 with nuh_layer_id equal to 0 (when present), o NAL units with nal_unit_type in the range of UNSPEC48..UNSPEC55 with nuh_layer_id equal to 0 (when present).
[0260] The first NAL unit preceding firstBIPicNalUnit and succeeding the last VCL NAL unit preceding firstBIPicNalUnit, when any, may only be one of the above-listed NAL units.When there is none of the above NAL units preceding firstBIPicNalUnit and succeeding the last VCL NAL preceding firstBIPicNalUnit, when any, firstBIPicNalUnit starts a new access unit.
[0261] Access unit boundary detection may be based on but may not be limited to one or more of the following:Detecting that a VCL NAL unit of a base-layer picture is the first VCL NAL unit of an access unit, e.g., on the basis that: o the VCL NAL unit includes a block address or alike that is the first block of the picture in decoding order; and / or o the picture order count, picture number, or similar decoding or output order or timing indicator differs from that of the previous VCL NAL unit(s).Having detected the first VCL NAL unit of an access unit, concluding based on pre-defined rules, e.g., based on nal_unit_type which non-VCL NAL units that precede the first VCL NAL unit of an access unit and succeed the last VCL NAL unit of the previous access unit in decoding order belong to the access unit.
[0262] The Versatile Video Coding (VVC) includes new coding tools compared to HEVC or H.264 / AVC. These coding tools are related to, for example, intra prediction; interpicture prediction; transform, quantization and coefficients coding; entropy coding; in-loop filter; screen content coding; 360-degree video coding; high-level syntax and parallel processing. Some of these tools are briefly described in the following:Intra prediction o 67 intra mode with wide angles mode extension o Block size and mode dependent 4 tap interpolation filter o Position dependent intra prediction combination (PDPC) o Cross component linear model intra prediction (CCLM) o Multi-reference line intra prediction o Intra sub-partitions o Weighted intra prediction with matrix multiplicationInter-picture predictiono Block motion copy with spatial, temporal, history-based, and pairwise average merging candidates o Affine motion inter prediction o sub-block based temporal motion vector prediction o Adaptive motion vector resolution o 8x8 block-based motion compression for temporal motion prediction o High precision (1 / 16 pel) motion vector storage and motion compensation with 8-tap interpolation filter for luma component and 4-tap interpolation filter for chroma component o Triangular partitions o Combined intra and inter prediction o Merge with motion vector difference (MVD) (MMVD) o Symmetrical MVD coding o Bi-directional optical flow o Decoder side motion vector refinement o Bi-prediction with CU-level weightTransform, quantization and coefficients coding o Multiple primary transform selection with DCT2, DST7 and DCT8 o Secondary transform for low frequency zone o Sub-block transform for inter predicted residual o Dependent quantization with max QP increased from 51 to 63 o Transform coefficient coding with sign data hiding o Transform skip residual codingEntropy Coding o Arithmetic coding engine with adaptive double windows probability updateIn loop filter o In-loop reshaping o Deblocking filter with strong longer filter o Sample adaptive offset o Adaptive Loop FilterScreen content coding: o Current picture referencing with reference region restriction 360-degree video coding o Horizontal wrap-around motion compensationHigh-level syntax and parallel processingo Reference picture management with direct reference picture list signalling o Tile groups with rectangular shape tile groups
[0263] In VVC, each picture may be partitioned into coding tree units (CTUs) similar to HEVC. A CTU may be split into smaller Cus using quaternary tree structure. Each CU may be partitioned using quad-tree and nested multi-type tree including ternary and binary split. There are specific rules to infer partitioning in picture boundaries. The redundant split patterns are disallowed in nested multi-type partitioning.
[0264] In some video coding schemes, such as HEVC and VVC, a picture is divided into one or more tile rows and one or more tile columns. The partitioning of a picture to tiles forms a tile grid that may be characterized by a list of tile column widths and a list of tile row heights. A tile may be required to include an integer number of elementary coding blocks, such as CTUs in HEVC and VVC. Consequently, tile column widths and tile row heights may be expressed in the units of elementary coding blocks, such as CTUs in HEVC and VVC.
[0265] A tile may be defined as a sequence of elementary coding blocks, such as CTUs in HEVC and VVC, that covers one "cell" in the tile grid, i.e., a rectangular region of a picture. Elementary coding blocks, such as CTUs, may be ordered in the bitstream in raster scan order within a tile.
[0266] Some video coding schemes may allow further subdivision of a tile into one or more bricks, each of which includes a number of CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick. However, a brick that is a true subset of a tile is not referred to as a tile.
[0267] In some video coding schemes, such as H.264 / AVC, HEVC and VVC, a coded picture may be partitioned into one or more slices. A slice may be decodable independently of other slices of a picture and hence a slice may be considered as a preferred unit for transmission. In some video coding schemes, such as H.264 / AVC, HEVC, and VVC, a video coding layer (VCL) NAL unit includes exactly one slice.
[0268] A slice may comprise an integer number of elementary coding blocks, such asCTUs in HEVC or VVC.
[0269] In some video coding schemes, such as VVC, a slice includes an integer number of tiles of a picture or an integer number of CTU rows of a tile.
[0270] In some video coding schemes, two modes of slices may be supported, namely the raster-scan slice mode and the rectangular slice mode. In the raster-scan slice mode, a slice includes a sequence of tiles in a tile raster scan of a picture. In the rectangular slice mode, a slice includes an integer number of tiles of a picture or an integer number of CTU rows of a tile that collectively form a rectangular region of the picture.
[0271] A non-VCL NAL unit may be for example one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, a picture header (PH) NAL unit, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Some non-VCL NAL units, such as parameter sets and picture headers, may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units might not be necessary for the reconstruction of decoded sample values.
[0272] Some coding formats specify parameter sets that may carry parameter values needed for the decoding or reconstruction of decoded pictures. Some examples of different types of parameter sets are briefly described in this paragraph. A video parameter set (VPS) may include parameters that are common across multiple layers in a coded video sequence or describe relations between layers. Parameters that remain unchanged through a coded video sequence (in a single-layer bitstream) or in a coded layer video sequence may be included in a sequence parameter set (SPS). In addition to the parameters that may be needed by the decoding process, the sequence parameter set may optionally include video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. A picture parameter set (PPS) includes such parameters that are likely to be unchanged in several coded pictures. A picture parameter set may include parameters that may be referred to by the coded image segments of one or more coded pictures. A header parameter set (HPS) has been proposed to include such parameters that may change on picture basis. In VVC, an Adaptation Parameter Set (APS) may comprise parameters for decoding processes of different types, such as adaptive loop filtering or luma mapping with chroma scaling.
[0273] A parameter set may be activated when it is referenced e.g., through its identifier.For example, a header of an image segment, such as a slice header, may include an identifierof the PPS that is activated for decoding the coded picture including the image segment. A PPS may include an identifier of the SPS that is activated, when the PPS is activated. An activation of a parameter set of a particular type may cause the deactivation of the previously active parameter set of the same type.
[0274] Instead of or in addition to parameter sets at different hierarchy levels (e.g., sequence and picture), video coding formats may include header syntax structures, such as a sequence header or a picture header. A sequence header may precede any other data of the coded video sequence in the bitstream order. A picture header may precede any coded video data for the picture in the bitstream order.
[0275] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units. A prefix SEI NAL unit may start a picture unit or alike; and a suffix SEI NAL unit may end a picture unit or alike. Hereafter, an SEI NAL unit may equivalently refer to a prefix SEI NAL unit or a suffix SEI NAL unit. An SEI NAL unit includes one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation.
[0276] Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for specific use. The standards may include the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications may require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient may be specified.
[0277] A coded picture is a coded representation of a picture.
[0278] A bitstream may be defined as a sequence of bits, which may in some coding formats or standards be in the form of a NAL unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.
[0279] A first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams. In some coding formats or standards, the end of the first bitstream may be indicated by a specific NAL unit, which may be referred to as the end of bitstream (EOB) NAL unit and which is the last NAL unit of the bitstream.
[0280] In some coding formats, a coded video sequence (CVS) may be defined as such a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream. In some coding formats, an end of sequence (EOS) NAL unit may be used to indicate the end of a CVS.
[0281] In some coding formats, such as AVI, a coded video sequence comprises one or more temporal units. A temporal unit includes a series of OBUs starting from a temporal delimiter, optional sequence headers, optional metadata OBUs, a sequence of one or more frame headers, each followed by zero or more tile group OBUs as well as optional padding OBUs. A temporal unit may be defined to comprise all the OBUs that are associated with a specific, distinct time instant. A temporal unit may comprise a temporal delimiter OBU, and all the OBUs that follow, up to but not including the next temporal delimiter. A temporal delimiter OBU may be defined as an indication that the following OBUs will have a different presentation / decoding time stamp from the one of the last frame prior to the temporal delimiter.
[0282] In some coding formats, a bitstream may comprise one or more coded video sequences.
[0283] The AVI codec supports input video signals in the 4:0:0 (monochrome), 4:2:0, 4:2:2, and 4:4:4 formats. The allowed pixel representations are 8, 10, and 12 bit. The AVI codec operates on pixel blocks. Each pixel block is processed in a predictive -trans form coding scheme, where the prediction comes from either intraframe reference pixels, interframe motion compensation, or some combinations of the two. The residuals undergo a 2-D unitary transform to further remove the spatial correlations, and the transform coefficients are quantized. Boththe prediction syntax elements and the quantized transform coefficient indexes are then entropy coded using arithmetic coding. There are three optional in-loop postprocessing filter stages to enhance the quality of the reconstructed frame for reference by subsequent coded frames. A normative film grain synthesis unit is also available to improve the perceptual quality of the displayed frames.
[0284] The AVI bitstream is packetized into open bitstream units (OBUs). An ordered sequence of OBUs is fed into the AVI decoding process.
[0285] An OBU may comprise a variable length string of bytes. An OBU includes a header and a payload and may include the OBU size, which may follow the OBU header and precede the OBU payload within the OBU. The header identifies the OBU type (in AVI, obu_type of 4 bits). In AVI, OBU header includes obu_has_size_field (1 bit), which indicates when the OBU includes the OBU size or when the OBU size is obtained otherwise (e.g., see the length delimited bitstream format described below).
[0286] A bitstream of OBUs may be formatted as a so-called low-overhead bitstream format, which comprises a sequence of OBUs, wherein the OBU header indicates the OBU size (without the size field) in bytes. Alternatively, a bitstream of OBUs may be formatted as a so- called length delimited bitstream format, in which each temporal unit starts with a size field (temporal_unit_size) indicating the size of the temporal unit payload in bytes, and each frame unit starts with a size field (frame_unit_size) indicating the size of the frame unit payload in bytes, and each OBU is preceded by a size field indicating the size of the OBU in bytes.
[0287] The OBU types may include the following.Sequence Header includes information that applies to the entire sequence, e.g., sequence profile (see Section VIII) and whether to enable certain coding tools.Temporal Delimiter indicates the frame presentation time stamp. All displayable frames following a temporal delimiter OBU will use this time stamp, until the next temporal delimiter OBU arrives. A temporal delimiter and its subsequent OBUs of the same time stamp are referred to as a temporal unit. In the context of scalable coding, the compression data associated with all representations of a frame at various spatial and fidelity resolutions will be in the same temporal unit.Frame Header sets up the coding information for a given frame, including signaling inter or intraframe type, indicating the reference frames and signaling probability model update method.Tile Group includes the tile data associated with a frame. Each tile may be independently decoded. The collective reconstructions form the reconstructed frame after potential loop filtering.Frame includes the frame header and tile data. The frame OBU is largely equivalent to a frame header OBU and a tile group OBU but allows less overhead cost.Metadata carries information, such as high dynamic range, scalability, and timecode.Tile List includes tile data similar to a tile group OBU. However, each tile here has an additional header that indicates its reference frame index and position in the current frame. This allows the decoder to process a subset of tiles and display the corresponding part of the frame, without the need to fully decode all the tiles in the frame.
[0288] MPEG-5 Part 2 Low Complexity Enhancement Video Coding (LCEVC) is published as ISO / IEC 23094-2. LCEVC works by encoding a lower resolution (and potentially also lower bit depth) version of a source video using any existing codec (the “base codec”) and then coding the differences between the lower resolution video and the full resolution source, up to mathematically lossless coding when needed, using a different compression method (the “enhancement”).
[0289] This enhancement is achieved by a combination of processing an input video at a lower resolution with an existing single-layer codec and using a simple and small set of highly specialized tools to correct impairments, upscale and add details to the processed video.
[0290] Encoder
[0291] The encoding process to create an LCEVC conformant bitstream and may be depicted in three major steps.
[0292] Base codec
[0293] Firstly, the input sequence is fed into two consecutive non-normative downscalers and is processed according to the chosen scaling modes. Any combination of the three available options (2-dimensional scaling, 1 -dimensional scaling in the horizontal direction only or no scaling) may be used. The output then invokes the base codec which produces a base bitstream according to its own specification. This encoded base may be included as part of the LCEVC bitstream.
[0294] Enhancement sub-layer 1
[0295] The reconstructed base picture may be upscaled to undo the downscaling process and is then subtracted from the first-order downscaled input sequence in order to generate the sub-layer 1 (L-l) residuals. These residuals form the starting point of the encoding process of the first enhancement sub-layer. A number of coding tools, which will be described further in the following section, process the input, and generate entropy encoded quantized transform coefficients.
[0296] Enhancement sub-layer 2
[0297] As a last step of the encoding process, the enhancement data for sub-layer 2 (L- 2) needs to be generated. In order to create the residuals, the coefficients from sub-layer 1 are processed by an in-loop LCEVC decoder to achieve the corresponding reconstructed picture. Depending on the chosen scaling mode, the reconstructed picture is processed by an upscaler. Finally, the residuals are calculated by a subtraction of the input sequence and the upscaled reconstruction. Similar to sub-layer 1, the samples are processed by a few coding tools. In addition, a temporal prediction may be applied on the transform coefficients in order to achieve a better removal of redundant information. The entropy encoded quantized transform coefficients of sub-layer 2, as well as a temporal layer specifying the use of the temporal prediction on a block basis, are included in the LCEVC bitstream.
[0298] Decoder
[0299] For the creation of the output sequence, the decoder analyses the LCEVC conformant bitstream. The process may again be divided into three parts:
[0300] Base codec
[0301] In order to generate the Decoded Base Picture (Layer 0) the base decoder is fed with the extracted base bitstream. According to the chosen scaling mode in the configuration, this reconstructed picture might be upscaled and is afterwards called Preliminary Intermediate Picture.
[0302] Enhancement sub-layer 1
[0303] Following the base layer, the enhancement part needs to be decoded. Firstly, the coefficients belonging to enhancement sub-layer 1 are decoded using the inverse tools of the encoding process. Additionally, an L- 1 filter might be applied in order to smooth the boundaries of a transform block. The output is then referred to as Enhancement Sub-Layer 1 and is added to the preliminary intermediate picture which results in the Combined Intermediate Picture. Again, depending on the scaling mode, an upscaler might be applied and the resulting Preliminary Output Picture has then the same dimensions as the overall output picture.
[0304] Enhancement sub-layer 2
[0305] As a final step, the second enhancement sub-layer is decoded. According to the temporal layer, a temporal prediction might be applied to the dequantized transform coefficients. This Enhancement Sub-Layer 2 is then added to the Preliminary Output Picture to form the Combined Output Picture as a final output of the decoding process.
[0306] Bitstream structure
[0307] The LCEVC bitstream includes a base layer, which may be at a lower resolution, and an enhancement layer including up to two sub-layers. This subsection briefly explains the structure of this bitstream and how the information may be extracted. While the base layer may be created using any video encoder and is not specified further in the LCEVC specification, the enhancement layer is expected to follow the structure as specified. Similar to other MPEG codecs, the syntax elements are encapsulated in network abstraction layer (NAL) units which also help synchronize the enhancement layer information with the base layer decoded information. Depending on the position of the frame within a group of pictures (GOP), additional data specifying the global configuration and controlling the decoder may be present. The data of one enhancement picture is encoded into several chunks. These data chunks are hierarchically organized. For each processed plane (nPlanes), up to two enhancement sublayers (nLevels) are extracted. Each of them again unfolds into numerous coefficient groups ofentropy encoded transform coefficients. The amount depends on the chosen type of transform (nLayers). Additionally, when the temporal prediction is used, for each processed plane an additional chunk with temporal data for Enhancement sub-layer 2 is present.
[0308] MPEG-5 LCEVC has some similarities with a scalable codec (e.g., spatial scalability thanks to the upsampler) but it is also substantially different for the following reasons:Generally, in scalable codec, the base layer is encoded with the same standard of the enhancement layer. As specified in the description LCEVC is codec agnostic. The base layer used in LCEVC may be any codec. This particular feature allows LCEVC to be used with any standard, such as H.264 / AVC or H.266 / VVC, and also with any other video codec (e.g., AVI, VP8, VP9 etc.). The MPEG-5 LCEVC structure is using simple tools specifically designed for the sparse nature of the residual data which allow to keep the complexity low and limit the overhead associated with the enhancement layers, a common problem of scalable codecs. This makes possible to have a software version of the MPEG-5 LCEVC that may run on existing hardware and on top of existing base codec with no need to develop a specific hardware for it. As a consequence, the base codec may work more efficiently and faster given the ability of LCEVC to work with a base codec running at a quarter of the resolution.Differently to most of scalable codecs, MPEG-5 LCEVC provides two levels of enhancement that may be applied at different stages or resolutions. Each level has its own independent quantization module and sublayer of the bitstream that may easily decoupled from the other. This also allows bitrate allocation flexibility to cope with different type of content. It may be noted that MPEG-5 LCEVC offers up to two cascade scaling processes in order to further improve the efficiency of the base layer. Each scaler may be user defined, along the following degrees of freedom: kernel size, type of upscaling (e.g., which sub-layer, L-l or L-2) and kernel values. MPEG-5 LCEVC offers 4 normative upsamplers and one 4 taps user defined kernel. Scalable codecs are generally offering only one fixed scaling engine and it is not programmable.MPEG-5 LCEVC may handle different bit depths up to 14 bits per pixel in the main profile. The standard allows the base layer to work on a different bit depth compared the input signal one. This operation may effectively enhancea base layer working at a lower bit depth to a higher one contributing to maintain the fidelity of the input signal. An example of this application is delivering high dynamic range (HDR) with technologies that cannot deliver more than 8-bit per color component, like AVC High Profile.
[0309] A uniform resource identifier (URI) may be defined as a string of characters used to identify a name of a resource. Such identification enables interaction with representations of the resource over a network, using specific protocols. A URI is defined through a scheme specifying a concrete syntax and associated protocol for the URI. The uniform resource locator (URL) and the uniform resource name (URN) are forms of URI. A URL may be defined as a URI that identifies a web resource and specifies the means of acting upon or obtaining the representation of the resource, specifying both its primary access mechanism and network location. A URN may be defined as a URI that identifies a resource by name in a particular namespace. A URN may be used for identifying a resource without implying its location or how to access it.
[0310] In many video communication or transmission systems, transport mechanisms, and multimedia container file formats, there are mechanisms to transmit or store a scalability layer separately from another scalability layer of the same bitstream, e.g. to transmit or store the base layer separately from the enhancement layer(s). It may be considered that layers are stored in or transmitted through separate logical channels. For example in ISOBMFF, the base layer may be stored as a track and each enhancement layer may be stored in another track, which may be linked to the base-layer track using so-called track references.
[0311] WebP image format
[0312] WebP is a modern image format that provides superior lossless and lossy compression for images on the web. Lossless WebP supports transparency (also known as alpha channel). For cases when lossy RGB compression is acceptable, lossy WebP also supports transparency. Lossy, lossless and transparency are all supported in animated WebP images.
[0313] Lossy WebP compression uses predictive coding to encode an image, the same method used by the VP8 video codec to compress keyframes in videos. Predictive coding uses the values in neighboring blocks of pixels to predict the values in a block, and then encodesonly the difference. Lossless WebP compression uses already seen image fragments in order to exactly reconstruct new pixels. It may also use a local palette when no interesting match is found.
[0314] A WebP file includes VP8 or VP8L image data, and a container based on RIFF.
[0315] More information on WebP is available from [htEll^ cYeloper^google oro / sp^d / we^ (last accessed on January 11, 2024)]
[0316] Resource Interchange File Format (RIFF)
[0317] WebP is an image format that uses either (i) the VP8 key frame encoding to compress image data in a lossy way or (ii) the WebP lossless encoding. These encoding schemes should make images more efficient than older formats, such as JPEG, GIF, and PNG. WebP is optimized for fast image transfer over the network (for example, for websites). The WebP format has feature parity (color profile, metadata, animation, and the like.) with other formats as well.
[0318] The WebP container (that is, the RIFF container for WebP) allows feature support over and above the basic use case of WebP (e.g.„ a file including a single image encoded as a VP8 key frame). The WebP container provides additional support for the following:Lossless Compression: An image may be losslessly compressed, using the WebP Lossless Format.Metadata: An image may have metadata stored in Exchangeable Image File Format (Exif) or Extensible Metadata Platform (XMP) format.Transparency: An image may have transparency, that is, an alpha channel.Color Profile: An image may have an embedded ICC profile as described by the International Color Consortium.Animation: An image may have multiple frames with pauses between them, making it an animation.
[0319] Further information on RIFF with WebP is available from ner and(last accessed on January 11, 2024)]
[0320] Current considerations on Slim HEIF
[0321] Minimized Image Box
[0323] The minimized image box provides a more compact way to represent carriage of image items in a file. Its main use case is for very small images where the usage of traditional carriage using the MetaBox would result in considerable overhead compared to the image data payload.
[0324] When MinimizedlmageBox is present, a file-level MetaBox shall not be present in the file. However, some parts of the body of a MetaBox may be embedded in the MinimizedlmageBox when the has_extended_meta flag is set to one.
[0325] The major_brand of the FileTypeBox may be specified in derived specifications to signal pre-defined values for inl'e lype and codec_config_type. However, when no such codec specific brand exists, the ’mini’ brand may be used, in which case has_explicit_codec_types shall be set to 1.
[0326] A file that includes a MinimizedlmageBox shall have major_brand of the FileTypeBox set to ’mini’, or to some other brand that explicitly states that the file conforms to the ’mini’ brand.
[0327] The MinimizedlmageBox may be followed by a MovieBox.
[0328] The MinimizedlmageBox may be followed by a MediaDataBox.
[0329] When processing the MinimizedlmageBox it is expanded to a full file-level MetaBox, which shall then be treated the same way as when the file included this MetaBox from the start.
[0330] Syntax aligned(8) class MinimizedlmageBox extends Box('mini') { bit(2) version; sqlite_varint width_minus_one; sqlite_varint height_minus_one; / / Color and bit-depth bit(l) is_float; if (is_float) { bit(2) float_precision;} else { bit(4) bit_depth_minus_one;} bit(l) is_monochrome; if (is_monochrome == 0) { bit(l) is_subsampled;} bit(l) full_range; bit(2) colour_type; if (colour_type == 0) { / / sRGB colour space colour_primaries = 1 ; transfer_characteristics = 13; matrix_coefficients = 6;} else if (colour_type == 1) { bit(5) colour_primaries; bit(5) transfer_characteristics; bit(5) matrix_coefficients;}else if (colour_type == 2) { bit(8) colour_primaries; bit(8) transfer_characteristics; bit(8) matrix_coefficients;} else { colour_primaries = 2; transfer_characteristics = 2; bit(8) matrix_coefficients; sqlite_varint icc_data_size_minus_one;} / / Item metadata bit(l) has_explicit_codec_types; if (has_explicit_codec_types) { unsigned int(32) infe_type; unsigned int(32) codec_config_type;} sqlite_varint main_item_codec_config_size; sqlite_varint main_item_data_size_minus_one; / / Other items bit(l) has_alpha; if (has_alpha) { / / Alpha has the following requirements: / / 1. Same dimensions, bit depth, and codec as main / / 2. Monochrome / / If hasAlpha is 1 and alpha size is 0, it means that / / the main image codec / / supports interleaved alpha bit(l) alpha_is_premultiplied; sqlite_varint alpha_item_codec_config_size; sqlite_ varint alph a_item_d ata_size ;}Boolean has_separate_alpha_item = has_alpha && alpha_item_data_size >; bit(l) has_extended_meta;if (has_extended_meta) { sqlite_varint extended_meta_size_minus_one;} bit(l) has_exif; if (has_exif) { sqlite_varint exif_data_size_minus_one;} bit(l) has_xmp; if (has_xmp) { sqlite_varint xmp_data_size_minus_one;} / / Pad bits until byte-aligned trailing_bits(); / / Payload data / / Codec config body data for alpha and main if (has_alpha && alpha_item_codec_config_size > 0) { unsigned int(8) alpha_item_codec_config [alpha_item_codec_config_size] ;} if (main_item_codec_config_size > 0) { unsigned int(8) main_item_codec_config [main_item_codec_config_size] ;} / / Extended 'meta' box if (has_extended_meta) { unsigned int(8) extended_meta[extended_meta_size_minus_one + 1];} / / ICC profile data if (colour_type == 3) { unsigned int(8) icc_data[icc_data_size_minus_one + 1];} / / Alpha and main elementary stream payloads if (has_separate_alpha_item) { unsigned int(8) alpha_data[alpha_item_data_size];} unsigned int(8) main_data[main_item_data_size_minus_one + 1]; / / Metadata payloads if (has_exif) { unsigned int(8) exif_data[exif_data_size_minus_one + 1];} if (has_xmp) { unsigned int(8) xmp_data[xmp_data_size_minus_one + 1];}}
[0331] Semantics
[0332] version: version of the MinimizedlmageBox. The current version shall be set to 0.
[0333] width_minus_one: specifies the width minus one of the reconstructed image in pixels, as specified in ImageSpatialExtentsProperty in clause 6.5.3
[0334] height_minus_one: specifies the height minus one of the reconstructed image in pixels, as specified in ImageSpatialExtentsProperty in clause 6.5.3
[0335] is_float: specifies whether float_precision or bit_depth_minus_one are signalled. When is_float is set to 1, it indicates that the float_precision is signalled, otherwise bit_depth_minus_one is signalled.
[0336] float_precision: specifies the format of floating-point numbers used for the pixel values as defined by IEEE 754-2008. The values 0, 1, and 2 correspond to half-precision float (binaryl6), single-precision float (binary32), and double-precision float (binary64) formats, respectively. Other values are reserved for a future specification. When is_float is set to 0, the value is undefined.
[0337] bit_depth_minus_one: indicates the maximum number of bits, minus one, per channel for the pixels of the reconstructed image of every associated image item.
[0338] is_monochrome: when set to 1 indicates that there is exactly one channel of coded colour samples, otherwise there are exactly three channels of coded colour samples.
[0339] is_subsampled: 0 indicates that there is the same number of samples in each colour channel. 1 indicates that the chroma planes are subsampled both vertically and horizontally compared to the luma plane (4:2:0). Set to 0 when is_monochrome is 1. The meaning of the value of is_subsampled shall match the contents of the main_item_codec_config and the main_data.
[0340] full_range: carries a VideoFullRangeFlag value as defined in ISO / IEC 23091-2
[0341] colour lype: specifies the colour encoding type. When set to 0 indicates sRGB. When set to 1 or 2 it implies the on-screen colours as signalled in Colour Informations ox with colour_type='nclx'. When set to 3 it indicates that an ICC Profile is present.
[0342] colour_primaries: carries a ColourPrimaries value as defined in ISO / IEC 23091- 2
[0343] transfer_characteristics: carries a Transfercharacteristics value as defined in ISO / IEC 23091-2
[0344] matrix_coefficients: carries a MatrixCoefficients value as defined in ISO / IEC 23091-2
[0345] icc_data_size_minus_one: specifies the size of ICC profile data minus one when the colour_type field indicates it is present in bytes. Undefined when the value of colour_type is not equal to 3.
[0346] has_explicit_codec_types: when set to 1 indicates that both infe lype and codec_config_type are explicitly signalled, otherwise their types are implied from the major_brand of the FileTypeBox. Shall be set to 1 when major_brand does not explicitly specify their default values.
[0347] infe_type: corresponds to the item_type field of the version 2 of the ItemlnfoEntry box. Defined by the major brand when has_explicit_codec_types is set to 0.
[0348] codec_config_type: corresponds to the codec configuration box type. Defined by the major brand when has_exxplicit_codec_types is set to 0.
[0349] main_item_codec_config_size: specifies the size of the configuration for the main image item.
[0350] main_item_data_size_minus_one: specifies the size minus one of the data for the main image item in bytes.
[0351] has_alpha: when set to 0 indicates that the image is opaque, otherwise the image has an alpha layer, whether the codec has native translucency support or an auxiliary image item is used.
[0352] alpha_is_premultiplied: when set to 1 indicates that main values are premultiplied by alpha, otherwise main values are not pre-multiplied.
[0353] alpha_item_codec_config_size: specifies the size of the configuration for the alpha image item in bytes. When set to 0 indicates that the codec does not need any configuration data for alpha or may reuse the one from the main image. The value is set to 0 when hasAlpha is 0. Undefined when has_alpha is not set to 1.
[0354] alpha_item_data_size: specifies the size of the data for the alpha image item in bytes. When has_alpha is set to 1, the value 0 indicates that the codec has native translucency support and that the alpha samples are coded alongside the colour samples in the main_data chunk. Zero when has_alpha is not set to 1.
[0355] has_extended_meta: when set to 1 indicates the presence of an extended MetaBox within the MinimizedlmageBox, otherwise it indicates the absence of it.
[0356] extended_meta_size_minus_one: specifies the size minus one of the extended metadata in bytes. Undefined when has_extended_meta is not set to 1.
[0357] has_exif: when set to 1 indicates the presence of an Exif metadata chunk, otherwise it indicates the absence of it.
[0358] exif_data_size_minus_one: specifies the size minus one of the Exif metadata in bytes. Undefined when has_exif is not set to 1.
[0359] has_xmp: when set to 1 indicates the presence of an XMP metadata chunk, otherwise it indicates the absence of it.
[0360] xmp_data_size_minus_one: specifies the size minus one of the XMP metadata in bytes. Undefined when has_xmp is not set to 1.
[0361] trailing_bits: padding bits to ensure payloads are 8-bit aligned. Shall be 0.
[0362] alpha_item_codec_config: specifies the optional alpha image codec configuration data. When has_alpha is set to 0 or alpha_item_codec_config_size is 0, alpha_item_codec_config is not present.
[0363] main_item_codec_config: specifies the main image item codec configuration data.
[0364] extended_meta: specifies the optional extended metadata. When has_extended_meta is set to 0 extended_meta is not present.
[0365] icc_data: specifies the optional ICC profile data. When colour_type is not set to 3 icc_data is not present.
[0366] alpha_data: specifies the optional alpha image data. When has_alpha is set to 0 or alpha_item_data_size is 0, alpha_data is not present.
[0367] main_data: specifies the main image data.
[0368] exif_data: specifies the optional Exif metadata. When has_exif is set to 0 exif_data is not present.
[0369] xmp_data: specifies the optional XMP metadata. When has_xmp is set to 0 xmp_data is not present.
[0370] sqlite_varint() function reads a varint coded according available from (last accessed on December 12,
[0371] For encapsulating images, the metadata required for describing the content is quite significant in HEIF. When we consider small and very small images in some cases the metadata size could be appreciable against the size of the image itself. For example, web content may include images that are small (on the order of 32x32 pixels). An example of HEIF image file 450 is shown in FIG. 4. The image file 450 includes the mandatory FileTypeBox 452 followed by the file-level MetaBox 454 which carries the data related to image items, for example, an item location 456, an item protection 458, an item references 460, and several other boxes related to the type of item. For example, when it includes an image encoded with HE VC then the image related decoder configuration and other properties are also stored in the file. The decoder configuration record, the sequence parameter set and the picture parameter set for an HEVC image may collectively have a size of several tens bytes; yet the coded slice data of a small HEVC image, e.g., of 32x32 pixels, may be just tens or hundreds of bytes. As it may be seen the size of the metadata related to an image item may be very significant especially when the size of the image data itself is very small.
[0372] There is a need to reduce the overhead of metadata needed for encapsulating smaller images to make the processing of such images lightweight.
[0373] In the following paragraphs the terms reduced header mode, compact metadata, compressed metadata, condensed image, slim HEIF, and slim AVIF are used interchangeably and are to be considered as alternatives of each other.
[0374] Reduced Header mode
[0375] According to embodiments of the invention, a file includes a reduced header or reduced metadata that is used for processing the media data of items included in the file.
[0376] An example embodiment describes a method comprising: writing in a container file: o the indication of presence of a reduced header mode or compact metadata o the reduced header including information for processing one or more item(s) in the file.
[0377] Another example embodiment describes a client-side method comprising: receiving and parsing a file comprising: o the indication of presence of a reduced header mode or compact metadata; and o the reduced header comprising information for processing one or more item(s) in the file; and processing the reduced header data to extract information needed for the decoding of the one or more item(s) in the file
[0378] Processing HTML with Multiple-Images! items)
[0379] When such files is referenced and / or declared in an HTML formatted document, several files may share the same reduced header data. In such a case, there may be an indicator in the HTML context (e.g., an HTML attribute, flag, variable, etc.) which may enable the client to download one of the files with reduced header mode and then apply the processed reduced header information to multiple other files without re-downloading their reduced header information.
[0380] In an embodiment, byte-size indicated partial download of one or more such files may result in fast access to media data and speed-up the retrieval and decoding process for the files.
[0381] In another embodiment, after partially downloading such files, the client may insert the reduced header to the appropriate part of the file so that it may be parsed and decoded.
[0382] Codec specific optimization
[0383] In embodiments described in this section, the image may be encoded with an NAL -unit-based-codec (for example, AVC, HEVC, VVC, EVC, LCEVC and their multi-layer extensions) or with a non-NAL -unit-based codec (for example, OBU-based AVI and AV2).
[0384] In embodiments described in this section, one or more codec-specific structures may comprise, but may not be limited to, one or more of the following: configuration item property, decoder configuration record, one or more parameter sets of one or more particulartypes (such as VPS, SPS, PPS, and / or APS), a sequence header, a picture header, or a slice header.
[0385] In embodiments described in this section, a slim codec brand is included in a slim HEIF file or parsed from a slim HEIF file. In an embodiment, the major_brand field of FileTypeBox is indicative that the file is structurally conforming to the slim HEIF, and minor_version field of FileTypeBox is indicative of the slim codec brand. In another embodiment, any brand among major_brand or compatible_brands in FileTypeBox is indicative that the file is structurally conforming to the slim HEIF, and the condensed image box comprises the slim codec brand. The slim codec brand may be indicated with a pre-defined number of bits, such as a 16-bit unsigned integer or a 32-bit unsigned integer, such as a four- character code.
[0386] In an embodiment, the indicated slim codec brand is indicative of a condensed version of one or more codec-specific structures. A condensed version is otherwise like the respective codec-specific structure but excludes certain syntax elements pre-defined in the slim codec brand definition. These syntax elements may be referred to as inferred syntax elements. In an embodiment, a file writer indicates a slim codec brand in a file and authors one or more condensed codec-specific structures defined for the slim codec brand into the file. In an embodiment, a file reader parses a slim codec brand from a file and parses one or more condensed codec-specific structures defined for the slim codec brand from the file.
[0387] In an example embodiment, the indicated slim codec brand is indicative of a condensed configuration item property associated with the respective image item (for example, HEVC configuration item property associated with hvcl image item) which carries the corresponding condensed decoder configuration record. A condensed decoder configuration record is otherwise like the respective decoder configuration record but excludes certain syntax elements pre-defined in the slim codec brand definition. In an embodiment, a file writer indicates a slim codec brand in a file and authors a condensed decoder configuration record defined for the slim codec brand into the file. In an embodiment, a file reader parses a slim codec brand from a file and parses a condensed decoder configuration record defined for the slim codec brand from the file.
[0388] In an embodiment, an inferred syntax element has a value that is pre-defined in the slim codec brand.
[0389] In an embodiment, an inferred syntax element has a value that is derived from one or more other syntax elements in the file.
[0390] In an embodiment, a file reader reconstructs one or more codec-specific structures from the respective parsed condensed structures by adding inferred syntax elements. The file reader includes a pre-defined or derived value for an inferred syntax element into the reconstructed one or more codec-specific structures.
[0391] In an example embodiment, condensed SPS and PPS NAL units of VVC are defined to exclude sps_pic_width_max_in_luma_samples, sps_pic_height_max_in_luma_samples, pps_pic_width_in_luma_samples, and pps_pic_height_in_luma_samples under a VVC slim codec brand. In an embodiment, a file writer writes condensed SPS and PPS NAL units in a decoder configuration record of a configuration item property of a file. A file writer writes the image width and image height into metadata of the file, e.g., in the ImageSpatialExtentsProperty that defines the width and height of the associated image item or in the width and height fields of the condensed image box or alike (e.g., width_minus_one and height_minus_one in the MinimizedlmageBox). In an embodiment, a VVC slim codec brand is parsed from a file. In response to parsing the VVC slim codec brand, a file reader reconstructs sps_pic_width_max_in_luma_samples and the sps_pic_height_max_in_luma_samples in the SPS NAL unit and the pps_pic_width_in_luma_samples and the pps_pic_height_in_luma_samples in the PPS NAL unit from the values from the metadata of the file, e.g., from the ImageSpatialExtentsProperty or from the width and height fields of the condensed image box or alike.
[0392] In an embodiment, a file writer does not include start code emulation prevention bytes in the units (e.g., NAL units or OBUs) stored in a condensed decoder configuration record. Likewise, in an embodiment, a file reader omits removal of start code emulation prevention bytes when parsing the units (e.g., NAL units or OBUs) stored in a condensed decoder configuration record.
[0393] In an example embodiment, the image is encoded with a pre-defined set of configuration rules; for example, the image is coded as an IDR picture with a single slice per picture resulting in an encoded picture with only one VCL NAL unit.
[0394] In an embodiment, a file writer authors condensed VCL NAL unit by excluding the NAL unit length and / or the NAL unit header from the item data stored in the file. The filewriter includes other NAL units of the bitstream, such as SPS and PPS NAL units, when any, into the configuration item property. In an embodiment, a file reader reconstructs a VCL NAL unit from the condensed VCL NAL unit by including the NAL unit length set equal to the item size and / or the NAL unit header set equal to pre-defined values defined in the slim codec brand.
[0395] In another embodiment, according to a pre-defined set of configuration rules, a bitstream comprises a sequence of OBUs, wherein the number of OBUs in the sequence is predefined, and OBU types of each OBU in the sequence of OBUs are pre-defined. For example, it may be pre-defined that the bitstream comprises a temporal delimiter OBU, a sequence header OBU, a frame header OBU, and a tile group OBU. In addition to OBU types, also the other OBU header fields for the sequence of OBUs are pre-defined. A file writer according to the embodiment modifies the sequence of OBUs of the bitstream to form a slim bitstream stored in the file. A file reader according to the embodiment reconstructs the sequence of OBUs of the bitstream from the slim bitstream stored in the file.
[0396] In an embodiment, the file writer according to the embodiment modifies the sequence of OBUs of the bitstream to form a slim bitstream with one or more of the following: 1) excludes the OBU header of the OBUs of the bitstream; 2) excludes the temporal delimiter OBU (when included in the sequence of OBUs); 3) excludes the sequence header OBU (when included in the sequence of OBUs) or pre-defined syntax elements of the sequence header OBU; 4) excludes the frame header OBU or pre-defined syntax elements of the frame header OBU; 5) excludes pre-defined syntax elements of tile group header. Pre-defined syntax elements of the sequence header OBU and / or the frame header OBU and / or tile group header may have pre-defined values. For example, in the AV 1 sequence header OBU, still_picture may be pre-defined to be equal to 1 and reduced_still_picture_header may be pre-defined to be equal to 1. In another example, the first tile group header may be pre-defined to start at the top-left tile of the frame, e.g., tg_start of AV 1 tile group header may be pre-defined to be equal to 0. Such pre-defined syntax elements may be excluded from the sequence header OBU and / or the frame header OBU and / or tile group header. Pre-defined syntax elements may have values that may be concluded from the file metadata. For example, in the AV 1 sequence header OBU, the value of max_frame_width_minus_l and max_frame_height_minus_l may be concluded from the file metadata.
[0397] In an embodiment, the file reader according to the embodiment the sequence of OBUs of the bitstream from the slim bitstream with one or more of the following: 1) writes and includes the OBU header based on the knowledge of the pre-defined values for the OBU headerfor each OBU in the sequence of OBUs of a bitstream; 2) writes and includes a temporal delimiter OBU (when included in the sequence of OBUs); 3) writes and includes a sequence header OBU (when included in the sequence of OBUs) or pre-defined syntax elements of the sequence header OBU; 4) writes and includes the frame header OBU or pre-defined syntax elements of the frame header OBU; 5) writes and includes pre-defined tile group header syntax elements. For items 3 and / or 4 and / or 5, the file reader writes and includes pre-defined syntax elements of the sequence header OBU and / or the frame header OBU and / or tile group header that have pre-defined values to the reconstructed sequence header OBU and / or the frame header OBU and / or tile group header. For items 3 and / or 4 and / or 5, the file reader derives values from the file metadata and writes and includes the respective syntax elements of the sequence header OBU and / or the frame header OBU and / or tile group header. For example, in the AV 1 sequence header OBU, the values of max_frame_width_minus_l and max_frame_height_minus_l are derived from the file metadata (such as image spatial extents item property or the width and height fields included in the MimizedlmageBox).
[0398] In an embodiment, the file writer according to the embodiment excludes the OBU size fields of the OBUs of the slim bitstream. A player according to the embodiment, decodes an OBU of the slim bitstream until the end of data of the OBU and then concludes that the next byte represents the next OBU.
[0399] In an alternate embodiment, elements of the pre-defined set of configuration rules may be signalled separately.
[0400] In an example embodiment, an indication of the item data of a picture comprising a single NAL unit (which may be referred to as single_nalu_flag) and an indication of VCL NAL unit(s) being of type IDR (which may be referred to as idr_nalu_header_flag) may be used to indicate elements of the pre-defined set of configuration rules, e.g., in the metabox with version 1 or in the condensed image header box. In an example, the semantics of single_nalu_flag and idr_nalu_header_flag may be defined as follows. single_nalu_flag equal to 0 indicates that the image item may comprise more than one NAL unit in the item data and NAL unit length is present in the item data. single_nalu_flag equal to 1 indicates that the image item has only one NAL unit in the item data and NAL unit length is absent in the item data. idr_nalu_header_flag equal to 0 indicates that the NAL unit header is present in NAL unit(s) of the item data. idr_nalu_header_flag equal to 1 indicates that the NAL unit header is absent in the NAL unit(s) of the item data and has pre-defined values, such as indicating an IDR NALunit type. For example, this embodiment would save around 2 to 4 bytes of NAL unit length and 2 bytes of NAL unit header in HE VC and VVC slim codec brands.
[0401] In another example embodiment, one or more of the following may be used to indicate elements of the pre-defined set of configuration rules used for the slim bitstream stored in the file: 0) pre_defined_obu_seq_flag, when equal to 1, indicates that the bitstream comprises a sequence of OBUs, wherein the number of OBUs in the sequence is pre-defined, and OBU types of each OBU in the sequence of OBUs are pre-defined; 1) obu_header_excluded_flag, when equal to 1, indicates that the slim bitstream excludes the OBU header of the OBUs of the bitstream; 2) temporal_delimiter_excluded_flag, when equal to 1, indicates that the slim bitstream excludes the temporal delimiter OBU (if included in the sequence of OBUs);3) sequence_header_trimmed_flag, when equal tol, indicates that the slim bitstream excludes the sequence header OBU (when included in the sequence of OBUs) or the slim bitstream excludes pre-defined syntax elements of the sequence header OBU;4) frame_header_trimmed_flag, when equal tol, indicates that the slim bitstream excludes the frame header OBU or the slim bitstream excludes pre-defined syntax elements of the frame header OBU; 5) tile_group_header_trimmed_flag, when equal to 1, indicates the slim bitstream excludes pre-defined syntax elements of the tile group header. The flags described above may, for example, be present in in the MetaBox with version 1 or in the condensed image header box or alike, or may be pre-defined to certain values based on a slim codec brand.
[0402] Codec-independent optimization
[0403] In an embodiment, the MetaBox with a version higher than zero is used for signaling the minimized header mode. For example, version equal to 1 may be used.
[0404] In an embodiment, the MetaBox with a version higher than zero and one or more bits in the flags field set equal to 1 are used for signaling the minimized header mode. For example, version equal to 1 and (flags & 1) equal to 1 may be used to indicate the minimized header mode.
[0405] In an embodiment, when the MetaBox with a version higher than zero, such as version equal to 1, it may be specified that the semantics of the flags field of the MetaBox header is determined by the first child box within the MetaBox. When the first child box within the MetaBox is the MimizedlmageBox or alike, the flags field of the MetaBox may be specified for the reduced header more.
[0406] In an embodiment, when the MetaBox with a version higher than zero, such as version equal to 1, it may be specified that the semantics of the flags field of the MetaBox header are specified for the reduced header mode.
[0407] When flags field of the MetaBox are specified for the reduced header mode, the flags field may comprise, but may not be limited to, one or more of the following types of information: o Syntax elements related to the metadata and / or items described by the MinimizedlmageBox or alike, such as bit depth, chroma format, full range flag. The respective syntax elements need not be present in the MinimizedlmageBox or alike. o An indication of a codec, such as a slim codec brand or alike. o Hash value, as described in some embodiments below. o For the indicated codec, a condensed version of one or more codec-specific structures and / or a condensed configuration item property associated with the respective one or more items.
[0408] In an embodiment, when version of the MetaBox is equal a value greater than 0, such as equal to 1, the following applies: It may be specified that the HandlerBox shall not or may not be present in the MetaBox. When the HandlerBox is absent, it may be specified that a child box in the MetaBox, such as the first child box in the MetaBox, shall specify how the handler lype value is inferred. Furthermore, it may be specified that this child box causes certain inference of complete child boxes of the MetaBox or syntax elements within child boxes of the MetaBox. Specifications apply as when the inferred child boxes and syntax elements values were present in the MetaBox. It may be specified that when a file reader does not support parsing of the first child box of the MetaBox, the file reader shall omit parsing of the entire MetaBox.
[0409] In an embodiment, when the minimized header mode is in use as signalled by means external to the MetaBox, such as a brand in the FileTypeBox, the MetaBox with a version higher than zero includes a MinimizedlmageBox or alike that specifies the condensed headers for images. For example, MetaBox version equal to 1 may be used. It may be required that the MinimizedlmageBox or alike is the first child box within the MetaBox.
[0410] In an embodiment, a MetaBox that has a version higher than zero and the MinimizedlmageBox or alike as a child box may be used to indicate the minimized headermode. It may be required that the MinimizedlmageBox or alike is the first child box within the MetaBox.
[0411] A MetaBox in ISOBMFF is a FullBox, whereas in QuickTime file format, the MetaBox is structured as a Box. Thus, to enable the reader to conclude whether the ISOBMFF MetaBox or the QuickTime MetaBox is in use, the 32 bits of version and flags in the ISOBMFF MetaBox must be distinguishable from the size field of the first child box within the MetaBox. In an embodiment, a file writer writes a MetaBox with version equal to 1 only when the size of the MetaBox is less than 0x01000000 + 8. When this constraint is followed, a reader is enabled to conclude that the next 32-bit unsigned integer after the box header type equal to 'meta' is not valid size field for the first child box within a MetaBox derived from a Box and thus the MetaBox is derived from a FullBox. In an embodiment, when the version of the ISOBMFF MetaBox is greater than 0, a reader interprets the size of the first child box in a QuickTime MetaBox being greater than or equal to 0x01000000 (e.g., 2 to the power of 24 or 224). When the size of the MetaBox itself is less than 0x01000000 + 8, this would be an illegal situation, and the reader concludes that the file uses the ISOBMFF MetaBox syntax.
[0412] In an embodiment, when a reader gets a QuickTime file with MetaBox derived from Box, the size field of the first child box of the MetaBox would be interpreted as the version and flags fields of the ISOBMFF-compliant MetaBox (i.e., derived from FullBox). When the size of that first child box is less than or equal to OxOOFFFFFF, the version and flags field in the ISOBMFF-compliant MetaBox would necessarily be interpreted as zero and non-zero, respectively, which is not allowed in the ISOBMFF specification of the MetaBox, which requires the flags field to be equal to 0 when version is equal to 0. Hence, the reader concludes that the file has MetaBox derived from Box. When the size of the MetaBox is less than OxOOFFFFF, the reader may conclude that the first child box cannot have a size that is more than that and the above-described method to conclude when the MetaBox is derived from a Box is applicable.
[0413] It is to be understood that the embodiments above may be similarly realized with thresholds for the size other than presented above (e.g. 0x01000000 + 8 or OxOOFFFFFF).
[0414] In an embodiment, when the MinimizedlmageBox or alike is present in a MetaBox, the file is further authored as follows: The syntax of the MinimizedlmageBox is such that it does not include item data, EXIF data or XMP data. For example, the syntax of MinimizedlmagBox may exclude alpha_data, main_data, exif_data and xmp_data as presentedin the MinimizedlmageBox syntax earlier. An ItemDataBox is present in the MetaBox and starts with the data defined by the MinimizedlmageBox, such as the item data for the alpha map (when present), the item data for the main image, the EXIF data (when present), and the XMP data (when present). For example, with reference to the syntax of MinimizedlmageBox presented earlier, the ItemDataBox shall start with the following data without any gaps: alpha_data (when has_alpha is equal to 1), main_data, exif_data (when has_exif is equal to 1) and xmp_data (when has_xmp is equal to 1) concatenated in that order. Any child boxes of MetaBox may be optionally present to describe additional items beyond those described with the MinimizedlmageBox or alike or to include additional non-essential metadata to items described in MinimizedlmageBox or alike.
[0415] In an embodiment, when the MinimizedlmageBox or alike is present in a MetaBox, the file is further authored as follows: The syntax of the MinimizedlmageBox is such that it may include item data, EXIF data or XMP data. For example, the syntax of MinimizedlmagBox may include alpha_data, main_data, exif_data and xmp_data as presented in the MinimizedlmageBox syntax earlier. Any child boxes of MetaBox may be optionally present to describe additional items beyond those described with the MinimizedlmageBox or alike or to include additional non-essential metadata to items described in MinimizedlmageBox or alike.
[0416] In an embodiment, has_extended_meta and extended_meta of MinimizedlmageBox as presented earlier are excluded from the syntax of the MinimizedlmageBox or alike included in a MetaBox. The MetaBox may describe, in its child boxes, any additional metadata beyond that described in the MinimizedlmageBox.
[0417] In an embodiment, a reduced header mode or alike being indicated in a FileTypeBox, ExtendedTypeBox, or OriginalFileTypeBox imposes that the first MediaDataBox of the file is structured as when it were a condensed image box or MinimizedlmageBox or alike. Such a MediaDataBox may be referred to as a self-contained MediaDataBox.
[0418] In an embodiment, a file may be authored in a way such that it allows the presence of a complete HEIF header with a MetaBox having version 0 and also as a file including the condensed / minimized headers called the multi-purpose HEIF file. For example, the MediaDataBox may be structured as when it were a condensed image box or MinimizedlmageBox or alike, and slim HEIF parsers may use the information of theMediaDataBox to parse condensed headers and respective image items, and additionally the file may include a MetaBox that legacy HEIF parsers may use. The same approach may be applied to AVIF or similar ISOBMFF based image-containing files.
[0419] In an embodiment, the box order of the file is subject to specific constraints, such as the MediaDataBox may be required to precede the file-level MetaBox (if any).
[0420] In an embodiment, the file is authored such that a byte range request allows for the processing of the condensed / minimized headers independently from the complete HEIF header with a MetaBox having version 0. For example, when MediaDataBox of the file is a self-contained MediaDataBox, and the box order of the file is such that the MediaDataBox is required to precede the file-level MetaBox, a slim HEIF parser may parse the resource that has byte range covering the MediaDataBox but excluding the MetaBox. The same approach may be applied to AVIF or similar ISOBMFF based image-containing files.
[0421] Embodiments described in this section enable multi-purpose file processing, which allows the creation of the file only once both for slim HEIF and conventional HEIF, and repurpose it later. The same approach may be applied to AVIF or similar ISOBMFF based image-containing files.
[0422] In an embodiment, an optional MIME parameter is used to signal the byte range to obtain a slim HEIF file or condensed / minimized headers.
[0423] In an embodiment, the profiles MIME parameter is extended to include the minor_version for the first listed brand, which is the major_brand in the file. For example, when the first four-character code of the profiles MIME parameter is followed by a dot (.), the minor_version value follows, which may, for example, be represented by an eight-character base-16 (e.g., hexadecimal) string.
[0424] In an embodiment, an optional MIME parameter is used to signal the byte range to obtain a file conforming to one or more brands indicated in the optional MIME parameter and / or in another MIME parameter within the same MIME type. For example, a nestedfile MIME parameter may have the following structure: ByteRange(,ProfileIndex)+, using a regular expression syntax, where keywords in italics are considered variables that are resolved by replacing them with their values, ( ) indicates a string of one or more characters, * indicates the inclusion of the string enclosed within the preceding parentheses by 0 or more times, + indicatesthe inclusion of the string enclosed within the preceding parentheses by 1 or more times, and ? indicates the inclusion of the string enclosed within the preceding parentheses by 0 times or 1 time, and alphanumeric characters are included as such. A ByteRange may have a structure like the value of bytes field in HTTP GET request (e.g., start-end, where start is the first byte index and end is the last byte index, and start may be empty to indicate the start of the file), e.g., or may just have the byte count starting from the beginning of the file. Each Profilelndex value points to the list index of the values listed in the profiles MIME parameter. The list indices start from a pre-defined value, such as 0 or 1. In an example, the MIME type comprises profiles="shif,mifl,heic" and nestedfile="-213,0", which indicate that the file complies with 'shif (which here indicate a slim HEIF brand), 'mifl' (the structural brand of HEIF) and 'heic' (the HEVC image brand) and when the first 213 bytes of the file are extracted into a separate file, that separate file complies with the listed brand having index 0 in the profiles parameter (i.e., 'shif').
[0425] In an embodiment, for EXIF and XMP metadata in the MinimizedlmageBox (or alike) or the self-contained MediaDataBox may be compressed using a generic compression method such as DEFLATE. The compression method may be pre-defined, or the compression method may be signalled either in the MinimizedlmageBox (or alike) or in self-contained MediaDataBox.
[0426] Hash and file type definition
[0427] A container file may include only the reduced header mode.
[0428] The indication of presence of a reduced header mode should be indicated either in the FileTypeBox or OriginalFileTypeBox or at a level prior to the reduced header mode (boxes).
[0429] A container file may include both the reduced header mode and the normal header mode (mode that is currently supported in HEIF).
[0430] To indicate the presence of both reduced header mode and normal mode in a file, a new box is defined called the file description box.
[0431] In an embodiment, a file writer derives a hash value using a pre-defined hash generation algorithm, such as a particular 32-bit cyclic redundancy check (CRC) (e.g., asspecified in ISO 3309 or ITU-T V.42). The file writer uses particular syntax elements and / or bytes of the file as input to the hash generation algorithm.
[0432] In an embodiment, a file writer derives a reduced header for the file (e.g., the payload of a MinimizedlmageBox excluding the item data) and uses the reduced header as input to the hash generation algorithm.
[0433] In an embodiment, a file writer derives a reduced header excluding pre-defined syntax elements (and excluding the item data) and uses the reduced header excluding the predefined syntax elements as input to the hash generation algorithm.
[0434] In an embodiment, a file writer derives a reduced header excluding the width and height fields (e.g., width_minus_one and height_minus_one of MinimizedlmageBox) for the file and uses the reduced header excluding the width and height as input to the hash generation algorithm.
[0435] In an embodiment, a file writer uses the MetaBox excluding pre-defined syntax structures or syntax elements as input to the hash generation algorithm. For example, the ItemLocationBox may be excluded from the derivation of the hash value.
[0436] In an embodiment, the value of the minor_version field present in the FileTypeBox may be filled with a hash value that may be registered and allows for a mapping into a MetaBox.
[0437] In an alternate embodiment, the hash value may be present in the MediaDataBox to allow for a longer hash value.
[0438] In an alternate embodiment, a hash value may be split into two parts where the initial part is stored in the minor_version field of the FileTypeBox and the extended hash value stored in the MediaDatabox.
[0439] In an embodiment, the hash value may be derived based on a pre-defined set for example the width and height of the image item.
[0440] In an embodiment, the hash value is included in a MIME parameter, which may be, e.g., referred to as item-metadata-hash.
[0441] In an embodiment, the byte range comprising the MediaDataBox(es) including the item data is included in a MIME parameter, which may be, e.g., referred to as media-bytes.
[0442] In an embodiment, a receiver or alike obtains MIME parameters including item- metadata-hash and media-bytes for several files. The receiver or alike obtains a first file among those several files. When a second file has the same value of item-metadata-hash as the first file, the receiver or alike obtains the byte range indicated by media-bytes of the second file and reconstructs the second file from the obtained byte range and the input data used for the hash generation algorithm.
[0443] Optimization for web page use case
[0444] In an embodiment, images sharing the same properties are included in the same HEIF file.
[0445] In an embodiment, a web page or alike is authored to comprise images specified through URL fragments of a HEIF file, for example #item_id = XXX, for individual images / image items.
[0446] In an embodiment, since URL fragments are resolved in the client side, the complete HEIF file is downloaded only once. Item properties, and the like may hence be shared between images.
[0447] In an embodiment, metadata of HEIF items may be coded (e.g., using base64 coding) and enumerated in the web page html source. Multiple slim HEIF files may reuse such codes and only byte -ranges for media data may be signalled. Relationship between the coded metadata and the slim HEIF file may be signalled via MIME type parameters or HTML attributes. A client which receives such a web page may first download the coded metadata and then reuse the coded metadata for different slim HEIF files, without re-downloading the metadata. There may be multiple such coded metadata listed or referenced in the html page.
[0448] Changes to MinimizedlmageBox
[0449] In an example embodiment, the minimizedlmageBox may be renamed to MinimizedMetadataBox which allows image items together with other items to be present inside the MinimizedMetadataBox.
[0450] In an embodiment, the MinimizedlmageBox is encapsulated inside a MetaBox with a new version for example MetaBox with version 1. For example, the MetaBox with a new version is the container for the MinimizedlmageBox or MinimizedMetadataBox.
[0451] In an embodiment, a HEIF file includes more than one item, for example two image items. The HEIF file includes two MinimizedlmageBox at the file level. Where the image item in the first MinimizedlmageBox is the first image item and the image item in the second minimizedlmageBox is the second image item.
[0452] In an embodiment, when the information about the item_ids in the minimizedlmageBox is not present the file reader implicitly assumes the item_id of the first image item to be item_id=l and the item_id of the second image item to be item_id=2, alternatively, the file reader may assume the item_id values of the two image items to be any value in ascending order bounded by the range of the item_id parameter represented by unsigned int(32) or unsigned int( 16).
[0453] In an embodiment, the item_id of the image item included in the minimizedlmageBox may be present in the extended_meta[] data of the minimizedlmageBox, for example, as part of the PrimaryitemBox or as part ItemlnfoBox included in the extended_meta[] data.
[0454] In an alternate embodiment, the minimizedlmageBox may include more than one item, for example, two image item. In an example implementation embodiment, the minimizedlmageBox structure is updated to carry more than one image item. aligned(8) class MinimizedlmageBox extends Box('mini') { bit(2) version; sqlite_varint item_count; for (i=0; i<item_count; i++) { sqlite_varint width_minus_one; sqlite_varint height_minus_one;unsigned int(8) main_data[main_item_data_size_minus_one + 1]; / / Metadata payloads if (has_exif) { unsigned int(8) exif_data[exif_data_size_minus_one + 1];} if (has_xmp) { unsigned int(8) xmp_data[xmp_data_size_minus_one + 1];}}}
[0455] Where item_count is the number of items included in the minimizedlmagebox.
[0456] The MinimizedlmageBox carries the following parameters shown below, in addition to others.{ bit(2) version; bit(l) is_float; bit(l) is_monochrome; bit(l) full_range; bit(2) colour lype; bit(l) has_explicit_codec_types; bit(l) has_alpha;Boolean has_separate_alpha_item = has_alpha && alpha_item_data_size > 0; bit(l) has_extended_meta; bit(l) has_exif; bit(l) has_xmp;}
[0457] In an example implementation embodiment, when the MinimizedlmageBox is encapsulated under the Metabox with a new version, the above parameters may be signalled using the flags field of the MetaBox.
[0458] One implementation example may be as below.
[0459] The following flags are allowed to be set in the tr_flags:0x000001 is-float.0x000002 is-monochrome.0x000004 full-range0x000008 has-explicit-codec-type0x000030 colour- type0x000040 alpha-present0x000080 extended-data-present0x000100 exif -present0x000200 xmp-present
[0460] In an alternate implementation embodiment, the MinimizedlmageBox carries the information needed to represent an image item with a specific output width and height while all other information is present optionally, an example of this case is shown below.{ bit(2) version; bit(l) is_float; bit(l) is_mo nochrome; bit(l) full_range; bit(l) has_explicit_codec_types; bit(l) has_extended_metadata; bit(l) reserved = 0; if (has_extended_metadata) { / / other information related to colour lype, xmp, exif, alpha channel are optionally present }}
[0461] Animation information
[0462] In an embodiment, the MinimizedlmageBox includes information related to animation. Where animation includes multiple frames with pauses between them.
[0463] In an example embodiment, the MinimizedlmageBox structure may be extended to include animation information as below.{ bit(l) has_animation_data; if (has_animation_data) { sqlite_varint anim_data_size_minus_one; unsigned int(32) background_color; unsigned int(16) loop_count; unsigned int(8) anim_data[anim_data_size_minus_one + 1];}}
[0464] When has_animation_data is set to zero, it indicates that no animation related data is present in the MinimizedlmageBox. when has_animation_data is set to one, it indicates that animation related data is present in the MinimizedlmageBox. anim_data_size_minus_one specifies the size of the animation data. background_color specifies the background color to be used for animation. The default background color of the canvas in [Blue, Green, Red, Alpha] byte order. This color may be used to fill the unused space on the canvas around the frames, as well as the transparent pixels of the first frame. loop_count specifies the number of times to loop the animation. When loop_count is 0, this means infinitely. anim_data[] carries the animation related configuration data together with the encoded payload.
[0465] In an embodiment, the animation data may include information about the configuration data needed to decode the animation payload such as the decoder configuration record of codec used for encoding the animated data. For example, when HE VC is used for encoding the animation data then it carries the HEVCDecoderConfigurationRecord as specified in ISO / IEC 14996-15.
[0466] In an embodiment, the animation data may include information about a reference space on which the animation data is to be displayed / Rendered. Where the reference space may be defined as a 2D coordinate system with the origin (0,0) located at the top-left corner and a maximum size defined by reference_width and reference_height; the x-axis is oriented from left to right and the y-axis from top to bottom. The animation is obtained after applying theimplicit resampling caused by the difference between the size of the reference space and the size of the animated frame. When the animation frame has transformative item properties, then the implicit resampling shall be performed on the animation frame before the first of its transformative item properties is applied.
[0467] In an embodiment, the animation data may include the time to wait before displaying the next frame, for example in 1 -millisecond units. The interpretation of the Frame Duration of 0 (and often <= 10) may defined by the implementation.
[0468] In an embodiment, the animation data may include information about the blending method to be used for each frame of the animation. For example, the blending method indicates how transparent pixels of the current frame are to be blended with corresponding pixels of the previous reference space. The blending methods may further indicate to use alphablending. After disposing of the previous frame, render the current frame on the reference space using alpha-blending. When the current frame does not have an alpha channel, assume the alpha value is 255. The blending method may indicate to not use blending. After disposing of the previous frame, render the current frame on the reference space by overwriting the rectangle covered by the current frame.
[0469] In an embodiment, the animation data may include information about disposal method. The disposal method indicates how the current frame is to be treated after it has been displayed (before rendering the next frame) on the reference space. The disposal method may indicate not to dispose and leave the reference space as is. The disposal method may indicate to dispose to the background color. Fill the rectangle on the reference space covered by the current frame with the background color.
[0470] In an embodiment, the animation data may include the encoded bitstream of the frames used for animation. The reader / client determines the frame boundaries together with frame duration, frame count, frame duration and frame type after parsing the encoded bitstream based on the codec used for encoding, for example, HEVC encoded bitstream.
[0471] Sample-to-reducedHeaderltem sample group
[0472] In an embodiment, samples of a track may be linked to items in reduced header mode / minimizedimage box using the sample-to-reducedheaderitem sample grouping (alternatively, any other suitable name may be used).
[0473] In an embodiment, the minimizedimagebox including the referred items are resolved as specified in semantics below.
[0474] In an embodiment, the sample-to-reducedheaderitem sample grouping is allowed for any types of tracks, and its syntax and semantics are unchanged regardless of the track handler type.
[0475] In an embodiment, in the absence of this sample group, the entire track-level minimixedimagebox, when any, is applicable to every sample.
[0476] class SampleToReducedHeaderltemEntryO extends SampleGroupDescriptionEntry('stmi'){ unsigned int(32) reduced_header_handler_type; unsigned int(32) num_items; for(i = 0; i < num_items; i++) { unsigned int(32) item_id[i];}}
[0477] In an embodiment the reduced_header_handler_type informs about the type of metadata schema used by the minimixedimagebox which is referenced by the items in this sample group (e.g., when the minimixedimagebox explicitly carries the handler type information within the box).
[0478] In an embodiment the reduced_header_handler_type is implied to be same as the type used by the minimixedimagebox which is referenced by the items in this sample group (i.e., when the minimixedimagebox does not carry the handler type information within the box and it implied to be equal to type ‘pict’).
[0479] In an embodiment, when there are multiple minimixedimageboxes with the same handler types (e.g., when the minimixedimagebox explicitly carries the handler type information within the box), the minimixedimagebox referred to in this sample group description entry is the first minimixedimagebox fulfilling one of the following ordered constraints:A minimixedimagebox included in the current track, with handler lype equal to reduced_header_handler_type.A minimixedimagebox included in MovieBox, with handler lype equal to reduced_header_handler_type.A minimixedimagebox included in the root level of the file, with handler_type equal to reduced_header_handler_type.In an embodiment, when there are multiple minimixedimageboxes (i.e., when the minimixedimagebox does not carry the handler type information within the box and it implied to be equal to type ‘pict’), the minimixedimagebox referred to in this sample group description entry is the first minimixedimagebox fulfilling one of the following ordered constraints:The first minimixedimagebox when included in the current track; orThe first minimixedimagebox when included in MovieBox; orThe first minimixedimagebox when included in the root level of the file.
[0480] In an embodiment, num_items counts the number of items referenced by this sample group.
[0481] In an alternate embodiment, num_items may not be present in which case it is implied to be equal to 1.
[0482] In an embodiment, item_id[i] specifies the item_ID value of an item that applies to or is valid for the sample mapped to this sample group description entry.
[0483] In an alternate embodiment, item_id[i] may not be present, in which case implicitly assumes the item_id of the first image item to be item_id=l and the item_id of the second image item to be item_id=2, alternatively, the file reader may assume the item_id values of the one or more image items to be any value in ascending order bounded by the range of the item_id parameter represented by unsigned int(32) or unsigned int( 16).
[0484] In an embodiment, the minimizedlmageBox may be present within the Metabox with a new version in such a case the SampleToMetadataltemEntry may be extended with a new version to map the samples of a track to the items in the MinimizedlmageBox.
[0485] In an embodiment, to map the samples of a track to the items in the minimizedlmagebox the minimizedlmagebox in the metabox would explicitly carry the handler lype information.
[0486] EntityToGroupBox updates
[0487] In an embodiment, a Entity ToGroupB ox with a specific grouping type may include items in the Minimizedlmagebox.
[0488] In an embodiment, when Entity groups are specified in GroupsListBox of a filelevel MetaBox they refer to items in the MinimizedlmageBox at the file-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple MinimizedlmageBoxes are at the file-level, the Entity group refers to the first MinimizedlmageBox at the file-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple items are present in the MinimizedlmageBox, the Entity group refers to the first item in the MinimizedlmageBox at the file-level.
[0489] In an embodiment, when Entity groups are specified in GroupsListBox of a filelevel MetaBox they refer to items in the MinimizedlmageBox included in the Metabox at the file-level. When a MinimizedlmageBox is within the MetaBox at the file-level, the Metabox may not include any other items outside the MinimizedlmageBox. When the MinimizedlmageBox does not include the item_ID parameter and multiple MinimizedlmageBoxes are included within the MetaBox at the file-level, the Entity group refers to the first MinimizedlmageBox included in the Metabox at the file-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple items are present in the MinimizedlmageBox within the MetaBox at the file-level, the Entity group refers to the first item in the MinimizedlmageBox included in the Metabox at the file-level.
[0490] In an embodiment, when Entity groups are specified in GroupsListBox of a movie-level MetaBox they refer to items in the MinimizedlmageBox at the movie-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple MinimizedlmageBoxes are at the movie-level, the Entity group refers to the first MinimizedlmageBox at the movie-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple items are present in the MinimizedlmageBox, the Entity group refers to the first item in the MinimizedlmageBox at the movie-level.
[0491] In an embodiment, when Entity groups are specified in GroupsListBox of a movie-level MetaBox they refer to items in the MinimizedlmageBox included in the Metabox at the movie-level. When a MinimizedlmageBox is within the MetaBox at the movie-level, theMetabox may not include any other items outside the MinimizedlmageBox. When the MinimizedlmageBox does not include the item_ID parameter and multiple MinimizedlmageBoxes are included within the MetaBox at the movie-level, the Entity group refers to the first MinimizedlmageBox included in the Metabox at the movie-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple items are present in the MinimizedlmageBox within the MetaBox at the movie-level, the Entity group refers to the first item in the MinimizedlmageBox included in the Metabox at the movie-level.
[0492] In an embodiment, when Entity groups are specified in GroupsListBox of a tracklevel MetaBox they refer to items in the MinimizedlmageBox at the track-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple MinimizedlmageBoxes are at the track-level, the Entity group refers to the first MinimizedlmageBox at the track-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple items are present in the MinimizedlmageBox, the Entity group refers to the first item in the MinimizedlmageBox at the track-level.
[0493] In an embodiment, when Entity groups are specified in GroupsListBox of a tracklevel MetaBox they refer to items in the MinimizedlmageBox included in the Metabox at the track-level. When a MinimizedlmageBox is within the MetaBox at the track-level, the Metabox may not include any other items outside the MinimizedlmageBox. When the MinimizedlmageBox does not include the item_ID parameter and multiple MinimizedlmageBoxes are included within the MetaBox at the track-level, the Entity group refers to the first MinimizedlmageBox included in the Metabox at the track-level. When the MinimizedlmageBox does not include the item_ID parameter and multiple items are present in the MinimizedlmageBox within the MetaBox at the track-level, the Entity group refers to the first item in the MinimizedlmageBox included in the Metabox at the track-level.
[0494] Implementation of SlimHEIF files and / or Reduced Header Mode HEIF image files with web resources.
[0495] The term web resources used below includes (the data itself and / or the link to the data) image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for example, JPEG’s, PNG’s and so on) unless stated otherwise.
[0496] In an embodiment, an entity (for example., a web client or a web browser or a user agent): detects the support (for example, in a client device) of images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats; based on the detection. indicates the support (for example, to a server) of content format(s) (or willingness to support) including images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats. optionally request a web resource which is formatted based on the supported format
[0497] In an embodiment, the said entity (for example., a web client or a web browser or a user agent) may indicate support of only images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format even when it supports images / items encapsulated in different other formats so as to prioritize images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format over others.
[0498] In an embodiment, when the said entity (for example, a web client or a web browser or a user agent) supports images / items encapsulated in multiple formats including images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format may indicate the following: list the images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format as the first in the list of multiple formats as an indication that it prioritizes the images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format over others or list the images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format with a specific priority value.
[0499] In an embodiment, an entity (for example, a server): receives the indication (from example, a web client or a web browser or a user agent) for the support of content format(s) (or willingness to support) including images / itemsencapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats; parses the received indication and selects the format(s) supported both by the indicated entity and format(s) supported by itself; where the selected format(s) may include images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats; based on the selected format(s), the entity may further create a web resource or update an existing web resource to include (the data itself and / or the link to the data) image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for example, JPEG’s, PNG’s and so on); and send the said web resource to the requestor entity (from example., a web client or a web browser or a user agent) and optionally together with the list of selected formats(s) which includes images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats.
[0500] In an embodiment, the said entity (for example, a server) may select only images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format even when it supports images / items encapsulated in different other formats so as to prioritize images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format over others.
[0501] In an embodiment, when the list of selected format(s) include multiple formats, the said entity (for example, a server) may choose only images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format due to the following: the list of selected format(s) includes the images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format as the first in the list of multiple formats and the order in this list is used to prioritizes the images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format over others or the list of selected format(s) includes the images / items encapsulated in LegacyHEIF format and / or images / items encapsulated in Slim HEIF mode format and / or ReducedHeader Mode HEIF format with a specific priority value which is higher than the priority value for other formats images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format provide the least format overhead over other formats images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format provide the least CPU cycle or least processing complexity or least battery usage over other formats.
[0502] In an embodiment, when the list of selected formats does not include any image format: the said entity (for example, a server) supports only other formats (for example JPEG’ s) and the other entity (from example., a web client or a web browser or a user agent) supports only images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format; the said entity (for example, a server) may choose to re-encode / transcode the images / entities in other formats to images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format; the said entity (for example, an edge server) may redirect the request to another entity (for example cache CDN or origin server or a proxy server) which hosts the web resource with images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format; the said entity (for example, a server) supports only images / items encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and the other entity (from example., a web client or a web browser or a user agent) supports only other formats (for example, JPEG’s); the said entity (for example, a server) may choose to re-encode / transcode the images / entities encapsulated in Legacy HEIF format and / or images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format to other formats the said entity (for example, an edge server) may redirect the request to another entity (for example cache CDN or origin server or a proxy server) which hosts the web resource with images / items in other formats.
[0503] In an embodiment, an entity, for example, a web server may host webpages or web resources which includes (the data itself and / or the link to the data) image files encodedwith SlimHEIF and / or Reduced Header Mode HEIF format and / or HEIF format together with image files encoded with other formats (for example, JPEG’s, PNG’s and so on).
[0504] In an embodiment, the said entity, for example, a web server may distribute / cache the said web resource to other entities on the network for example origin servers, proxy servers, a CDN’s and edge servers.
[0505] In an embodiment, the said entity, for example, a web server or the said other entities, for example origin servers, proxy servers, a CDN’ s and edge servers on the network may send as a response to a request or through PUSH mechanism or through other mechanisms, the said web resources to one or more entities (for example web browsers).
[0506] In an embodiment, an entity (for example web browsers or an app in mobile device) may receive as a response to a request or through PUSH mechanism or through other mechanisms, the said web resources.
[0507] In an embodiment, an entity (for example web browsers) requests a web page hosted on a server using HTTP Request mechanism.
[0508] In an example embodiment, an entity (for example., a web client or a web browser or a user agent) sends the said HTTP Request with a Accept request header indicating the support of content format(s) (or willingness to support) including images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats.
[0509] In an example embodiment, the support for images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats may be indicated using mime types registered for the specific formats.
[0510] In an example embodiment, the mime type for images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format may be indicated as images / heif+reduced or as images / reducedheif or any other suitable mime type which clearly distinguishes the file which includes the following: file includes only images / items in Slim HEIF mode format and / or Reduced Header Mode HEIF format;file includes only images / items in Slim HEIF mode format and / or Reduced Header Mode HEIF format with a specific codec type; file includes only images / items in Slim HEIF mode format and / or Reduced Header Mode HEIF format with different codec types; files include both images / items in Legacy HEIF format and Slim HEIF mode format and / or Reduced Header Mode HEIF format; files include both images / items in Legacy HEIF format and Slim HEIF mode format and / or Reduced Header Mode HEIF format with a specific codec type; files include both images / items in Legacy HEIF format and Slim HEIF mode format and / or Reduced Header Mode HEIF format with different codec types; file includes only images / items in HEIF format; file includes only images / items in HEIF format with a specific codec type; or file includes only images / items in HEIF format with different codec types.
[0511] In an example embodiment, an entity (for example server) receives: the HTTP request with Accept header including the indication for the support of images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats; parses the HTTP Accept header and selects the images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats; responds with the requested web resource including (the data itself and / or the link to the data) image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for example, JPEG’s, PNG’s and so on); and furthermore, it responds with the selected image / item formats through Content- Type response header.
[0512] In an example embodiment, the entity (for example server) may respond with requested web resource(s) without changing the extension on the image file(s) in the URL of the web resource. The same URL for a specific image file extension is reused to respond with a different image format. For example, the image file extension could be xyz.jpeg, however the response image file may be of type SlimHEIF and / or Reduced Header Mode HEIF format.
[0513] In an example embodiment, the entity (for example a web browser):receives the response web resource together with the said Content-Type response header including the indication about the image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for -^example, JPEG’s, PNG’s and so on); parses the Content-Type response header and concludes that the received web resource may include URLs to images / items with a specific extension are to be interpreted to be of certain format including image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for -example, JPEG’s, PNG’s and so on); and processes the received web resource based on the said conclusion.
[0514] However, the fact that the same URL may have multiple representations means that all the intermediate entities (for example caches) must have a way to differentiate the various Content-Type responses and Accept headers.
[0515] In an example embodiment, the said entity (for example a server) uses a Vary header together with Content-Type header to indicate other entities (for example, upstream clients) that the image format should be varied based on requesting entities (for example, web browser) HTTP Request Accept header values.
[0516] In an example below, a web browser indicates in its Accept header the support of images / items encapsulated with HEIF and SlimHEIF and / or Reduced Header Mode HEIF format.(client) > Accept: image / jpeg, image / png, image / heif, image / slimheif(server) > Content-Type: image / slimheif> Vary: Accept> (object)
[0517] In the example above, the value image / slimheif indicates the support of images / items encapsulated with SlimHEIF and / or Reduced Header Mode HEIF format. Any other suitable value which indicates the support of images / items encapsulated with SlimHEIF and / or Reduced Header Mode HEIF format may be used.
[0518] The server selected the image format to be image / slimheif given the choice of 4 different formats and responds in the Content-Type header its choice. Any upstream cache maysafely cache and serve the slimheif object to any entity which provides the same Accept header. When another entity sends a different header value, for example without image / slimheif, then a different format will be served and cached.
[0519] Furthermore, it is well known that the HTTP request with Accept Header extension could be expensive in terms of bitrate.
[0520] In an embodiment, an entity (for example a web browser) compresses the indication to support content format(s) (or willingness to support) including images / items encapsulated in Legacy HEIF format and images / items encapsulated in Slim HEIF mode format and / or Reduced Header Mode HEIF format and / or other image formats. The indication is part of the compressed payload.
[0521] For example, the web browser may use HTTP 2.0 header compression methods to compress the Accept header and the server may compress the Content-Type and Vary headers.
[0522] In an example embodiment, an entity (for example edge server / CDN) may not support the image / items encapsulated with SlimHEIF and / or Reduced Header Mode HEIF format as indicated in the HTTP Request Accept header in such a case the entity may redirect the request to a new entity (for example a proxy server which supports such formats or the origin server with dynamic image format optimization support) or may itself perform image format re-encoding / transcoding of the images to the image / items encapsulated with SlimHEIF and / or Reduced Header Mode HEIF format.
[0523] It is also well known that certain entities (for example, web browsers) do not provide a list or provide a partial list of image formats supported by them through the HTTP Request Accept header.
[0524] In an embodiment, entities (for example, servers) which do not receive or receive only a partial list of image formats supported by the requesting entity through the HTTP Request Accept headers may use any of the techniques defined below to provide web resources including image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for -example, JPEG’s, PNG’s and so on).
[0525] In an embodiment, the said entity (for example, servers) uses the HTML5 <picture> element to list multiple, alternative image targets in priority order where the list includes image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for -^example, JPEG’s, PNG’s and so on); the web resource is formatted as HTML5 document with the <picture> element as defined above and sent to the requesting entity (for example, web browser)
[0526] In an embodiment, the requesting entity (for example, web browser) receives the web resource formatted as HTML5 document with the <picture> element which lists multiple, alternative image targets in priority order where the list includes image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for -^example, JPEG’s, PNG’s and so on); the requesting entity parses the received web resource formatted as HTML5 document with the <picture> element and selects the image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format together with image files encoded with other formats (for -example, JPEG’s, PNG’s and so on).
[0527] In an embodiment, the said entity (for example, servers) uses feature detection methods (for example libraries such as modernizr) to detect the features supported by requesting entities (for example, web browsers); based on the feature detected when the requesting entity supports image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format the web resource may include image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format when the requesting entity does not support image files encoded with SlimHEIF and / or Reduced Header Mode HEIF format and / or Legacy HEIF format the web resource may include image files encoded with other formats (for -^example, JPEG’s, PNG’s and so on).
[0528] FIG. 5 is an example apparatus, which may be implemented in hardware, caused to implement slim mode for image file format. The apparatus 550 comprises at least one processor 552 (e.g., an FPGA and / or CPU), one or more memories 554 including computer program code 555, the computer program code 555 having instructions to carry out the methods described herein, wherein the one or more memories 554 and the computer program code 555 are configured to, with the at least one processor 552, cause the apparatus 550 to implement circuitry, a process, component, module, or function (implemented with control module 556) to implement the examples described herein, including implementing slim mode for image fileformat. Optionally included encoder 580 of the control module 556 performs encoding, and optionally included decoder 590 implements decoding. The memory 554 may be a non- transitory memory, a transitory memory, a volatile memory (e.g., RAM), or a non-volatile memory (e.g., ROM).
[0529] The apparatus 550 includes a display and / or I / O interface 558, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, and the like. The apparatus 550 includes one or more communication, e.g., network (N / W) interfaces (I / F(s)) 560. The communication I / F(s) 560 may be wired and / or wireless and communicate over the Internet / other network(s) via any communication technique including via one or more links 574. The communication I / F(s) 560 may comprise one or more transmitters or one or more receivers.
[0530] The transceiver 566 comprises one or more transmitters 568 and one or more receivers 570. The transceiver 566 and / or communication I / F(s) 560 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 564 used for communication over wireless link 572.
[0531] The control module 556 of the apparatus 550 comprises one of or both parts 556- 1 and / or 556-2, which may be implemented in a number of ways. The control module 556 may be implemented in hardware as control module 556-1, such as being implemented as part of the at least one processor 552. The control module 556-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 556 may be implemented as control module 556-2, which is implemented as computer program code (having corresponding instructions) 555 and is executed by the at least one processor 552. For instance, the one or more memories 554 store instructions that, when executed by the at least one processors 552, cause the apparatus 550 to perform one or more of the operations as described herein. Furthermore, the at least one processor 552, one or more memories 554, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
[0532] The apparatus 550 to implement the functionality of control 556 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 550 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 550 may be part of a self-organizing / optimizing network (SON) node or other node, such as a node in a cloud.
[0533] The apparatus 550 may also be distributed throughout the network (e.g., internet) including within and between apparatus 550 and any network element.
[0534] Interface 562 enables data communication and signaling between the various items of apparatus 550, as shown in FIG. 5. For example, the interface 562 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g., instructions) 555, including control 556 may comprise object-oriented software configured to pass data or messages between objects within computer program code 555. The apparatus 550 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 550 may at least partially reside in a common housing 578, or a subset of the various components of apparatus 550 may at least partially be located in different housings, which different housings may include housing 578.
[0535] FIG. 6 shows a schematic representation of non-volatile memory media 600a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 600b (e.g. universal serial bus (USB) memory stick) and 600c (e.g. cloud storage for downloading instructions and / or parameters 602 or receiving emailed instructions and / or parameters 602) storing instructions and / or parameters 602 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein.
[0536] FIG.7 is an example method 700 to implement the embodiments described herein, in accordance with an embodiment. At 702, the method 700 includes writing in a file: an indication of presence of a reduced header mode or compact metadata and the reduced header comprising information for processing one or more items. At 704, the method 700 incudes signaling the file comprising the reduced header mode.
[0537] The method 700 may be performed with an apparatus described herein, for example, any apparatus of FIG. 1 to FIG. 3, an apparatus of FIG. 5, or any apparatus of FIG. 11, or any other apparatus described herein.
[0538] FIG. 8 is another example method 800 to implement the embodiments described herein, in accordance with an embodiment. At 802, the method 800 includes, receiving and parsing a file comprising: an indication of presence of a reduced header mode or compact metadata and a reduced header comprising data for processing one or more items in the file. At 804, the method 800 includes, processing the reduced header data to extract information needed for decoding of the one or more items in the file.
[0539] The method 800 may be performed with an apparatus described herein, for example, any apparatus of FIG. 1 to FIG. 3, an apparatus of FIG. 5, any apparatus of FIG. 11, or any other apparatus described herein.
[0540] FIG. 9 is yet another example method 900 to implement the embodiments described herein, in accordance with an embodiment. At 902, the method 900 includes writing, in a file, a slim codec brand indicative of a condensed configuration item property associated with an image item. At 904, the method 900 includes and writing, in the file, the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition.
[0541] In an embodiment, the file further comprises an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
[0542] In an embodiment, writing the condensed configuration item property comprises: writing a condensed video coding layer (VCL) network abstraction layer (NAL) unit by excluding the NAL unit length and / or the NAL unit header from an item data stored in the file.
[0543] In an embodiment, writing the condensed configuration item property further comprises: including other NAL units of a bitstream into the condensed configuration item property.
[0544] The method 900 may be performed with an apparatus described herein, for example, any apparatus of FIG. 1 to FIG. 3, an apparatus of FIG. 5, or any apparatus of FIG. 11, or any other apparatus described herein.
[0545] FIG. 10 is still another example method 1000 to implement the embodiments described herein, in accordance with another embodiment. At 1002, the method 1000 includes receiving and parsing a file comprising: a slim codec brand indicative of a condensed configuration item property associated with an image item; and the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition. At 1004, the method 1000 includes and processing the slim codec brand to extract information needed for decoding of one or more items in the file.
[0546] In an embodiment, the one or more items comprise an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
[0547] In an embodiment, the condensed configuration item property comprises a condensed video coding layer (VCL) network abstraction layer (NAL) unit excluding the NAL unit length and / or the NAL unit header from an item data stored in the file, and the method 1000 further comprises: reconstructing a VCL NAL unit from the condensed VCL NAL unit by including the NAL unit length set equal to an item size and / or the NAL unit header set equal to pre-defined values defined in the slim codec brand.
[0548] The method 1000 may be performed with an apparatus described herein, for example, any apparatus of FIG. 1 to FIG. 3, an apparatus of FIG. 5, or any apparatus of FIG. 11, or any other apparatus described herein.
[0549] Referring to FIG. 11, this figure shows a block diagram of one possible and nonlimiting example in which the examples may be practiced. A user equipment (UE) 110, radio access network (RAN) node 170, and network element(s) 190 are illustrated. In the example of FIG. 11, the user equipment (UE) 110 is in wireless communication with a wireless network 100. A UE is a wireless device that may access the wireless network 100. The UE 110 includes one or more processors 120, one or more memories 125, and one or more transceivers 130 interconnected through one or more buses 127. Each of the one or more transceivers 130 includes a receiver, Rx, 132 and a transmitter, Tx, 133. The one or more buses 127 may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers 130 are connected to one or more antennas 128. The one or more memories 125 include computer program code 123.The UE 110 includes a module 140, comprising one of or both parts 140-1 and / or 140-2, which may be implemented in a number of ways. The module 140 may be implemented in hardware as module 140-1, such as being implemented as part of the one or more processors 120. The module 140-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the module 140 may be implemented as module 140-2, which is implemented as computer program code 123 and is executed by the one or more processors 120. For instance, the one or more memories 125 and the computer program code 123 may be configured to, with the one or more processors 120, cause the user equipment 110 to perform one or more of the operations as described herein. The UE 110 communicates with a radio access network (RAN) node 170 via a wireless link 111.
[0550] The RAN node 170 in this example is a base station that provides access by wireless devices such as the UE 110 to the wireless network 100. The RAN node 170 may be, for example, a base station for fifth generation cellular network technology (5G), also called New Radio (NR). In 5G, the RAN node 170 may be a NG-RAN node, which is defined as either a gNB (e.g., base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC) or an ng (new generation) -eNB. A gNB is a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to a 5G core network (5GC) (such as, for example, the network element(s) 190). The ng-eNB, is a node providing evolved universal terrestrial radio access (E-UTRA), for example, the LTE radio access technology, user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC. The NG-RAN node may include multiple gNBs, which may also include a central unit (CU) (gNB-CU) 196 and distributed unit(s) (DUs) (gNB-DUs), of which DU 195 is shown. Note that the DU may include or be coupled to and control a radio unit (RU). The gNB-CU is a logical node hosting radio resource control (RRC), service data adaptation protocol (SDAP) and PDCP protocols of the gNB or RRC and packet data convergence protocol (PDCP) protocols of the en-gNB (e.g., node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in E-UTRA- NR dual connectivity (EN-DC)) that controls the operation of one or more gNB-DUs. The gNB- CU terminates the interface between CU and DU control interface (Fl or Fl-C) interface connected with the gNB -DU. The Fl interface is illustrated as reference 198, although reference 198 also illustrates a link between remote elements of the RAN node 170 and centralized elements of the RAN node 170, such as between the gNB-CU 196 and the gNB-DU 195. The gNB-DU is a logical node hosting radio link control (RLC), MAC and physical layer (PHY) layers of the gNB or en-gNB, and its operation is partly controlled by gNB-CU. One gNB-CUIllsupports one or multiple cells. One cell is supported by only one gNB-DU. The gNB-DU terminates the Fl interface 198 connected with the gNB-CU. Note that the DU 195 is considered to include the transceiver 160, for example, as part of a RU, but some examples of this may have the transceiver 160 as part of a separate RU, for example, under control of and connected to the DU 195. The RAN node 170 may also be an eNB (evolved NodeB) base station, for example, long term evolution (LTE), or any other suitable base station or node.
[0551] The RAN node 170 includes one or more processors 152, one or more memories 155, one or more network interfaces (N / W I / F(s)) 161, and one or more transceivers 160 interconnected through one or more buses 157. Each of the one or more transceivers 160 includes a receiver, Rx, 162 and a transmitter, Tx, 163. The one or more transceivers 160 are connected to one or more antennas 158. The one or more memories 155 include computer program code 153. The CU 196 may include the processor(s) 152, memories 155, and network interfaces 161. Note that the DU 195 may also include its own memory / memories and processor(s), and / or other hardware, but these are not shown.
[0552] The RAN node 170 includes a module 150, comprising one of or both parts 150- 1 and / or 150-2, which may be implemented in a number of ways. The module 150 may be implemented in hardware as module 150-1, such as being implemented as part of the one or more processors 152. The module 150-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the module 150 may be implemented as module 150-2, which is implemented as computer program code 153 and is executed by the one or more processors 152. For instance, the one or more memories 155 and the computer program code 153 are configured to, with the one or more processors 152, cause the RAN node 170 to perform one or more of the operations as described herein. Note that the functionality of the module 150 may be distributed, such as being distributed between the DU 195 and the CU 196, or be implemented solely in the DU 195.
[0553] The one or more network interfaces 161 communicate over a network such as via the links 176 and 131. Two or more gNBs 170 may communicate using, for example, link 176. The link 176 may be wired or wireless or both and may implement, for example, an Xn interface for 5G, an X2 interface for LTE, or other suitable interface for other standards.
[0554] The one or more buses 157 may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, wireless channels, and the like.For example, the one or more transceivers 160 may be implemented as a remote radio head (RRH) 195 for LTE or a distributed unit (DU) 195 for gNB implementation for 5G, with the other elements of the RAN node 170 possibly being physically in a different location from the RRH / DU, and the one or more buses 157 could be implemented in part as, for example, fiber optic cable or other suitable network connection to connect the other elements (for example, a central unit (CU), gNB-CU) of the RAN node 170 to the RRH / DU 195. Reference 198 also indicates those suitable network link(s).
[0555] It is noted that description herein indicates that ‘cells’ perform functions, but it should be clear that equipment which forms the cell may perform the functions. The cell makes up part of a base station. That is, there may be multiple cells per base station. For example, there could be three cells for a single carrier frequency and associated bandwidth, each cell covering one- third of a 360 degree area so that the single base station’s coverage area covers an approximate oval or circle. Furthermore, each cell may correspond to a single carrier and a base station may use multiple carriers. So when there are three 120 degree cells per carrier and two carriers, then the base station has a total of 6 cells.
[0556] The wireless network 100 may include a network element or elements 190 that may include core network functionality, and which provides connectivity via a link or links 181 with a further network, such as a telephone network and / or a data communications network (for example, the Internet). Such core network functionality for 5G may include access and mobility management function(s) (AMF(S)) and / or user plane functions (UPF(s)) and / or session management function(s) (SMF(s)). Such core network functionality for LTE may include MME (Mobility Management Entity) / SGW (Serving Gateway) functionality. These are merely example functions that may be supported by the network element(s) 190, and note that both 5G and LTE functions might be supported. The RAN node 170 is coupled via a link 131 to the network element 190. The link 131 may be implemented as, for example, an NG interface for 5G, or an SI interface for LTE, or other suitable interface for other standards. The network element 190 includes one or more processors 175, one or more memories 171, and one or more network interfaces (N / W I / F(s)) 180, interconnected through one or more buses 185. The one or more memories 171 include computer program code 173. The one or more memories 171 and the computer program code 173 are configured to, with the one or more processors 175, cause the network element 190 to perform one or more operations.
[0557] The wireless network 100 may implement network virtualization, which is the process of combining hardware and software network resources and network functionality intoa single, software-based administrative entity, a virtual network. Network virtualization involves platform virtualization, often combined with resource virtualization. Network virtualization is categorized as either external, combining many networks, or parts of networks, into a virtual unit, or internal, providing network-like functionality to software containers on a single system. Note that the virtualized entities that result from the network virtualization are still implemented, at some level, using hardware such as processors 152 or 175 and memories 155 and 171, and also such virtualized entities create technical effects.
[0558] The computer readable memories 125, 155, and 171 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The computer readable memories 125, 155, and 171 may be means for performing storage functions. The processors 120, 152, and 175 may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multi-core processor architecture, as non-limiting examples. The processors 120, 152, and 175 may be means for performing functions, such as controlling the UE 110, RAN node 170, network element(s) 190, and other functions as described herein.
[0559] In general, the various embodiments of the user equipment 110 may include, but are not limited to, cellular telephones such as smart phones, tablets, personal digital assistants (PDAs) having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and playback appliances having wireless communication capabilities, Internet appliances permitting wireless Internet access and browsing, tablets with wireless communication capabilities, as well as portable units or terminals that incorporate combinations of such functions.
[0560] One or more of modules 140-1, 140-2, 150-1, and 150-2 may be caused to implement slim mode for image file format. Computer program code 173 may also be caused to implement for implement slim mode for image file format.
[0561] As described above, FIGs. 7 to 10 include flowcharts of an apparatus (e.g. 50, 550, or any other apparatuses described herein), method, and computer program product accordingto certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by a memory (e.g. 58, 125, or 554) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g. 56, 120, or 552) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.
[0562] A computer program product is therefore defined in those instances in which the computer program instructions, such as computer-readable program code portions, are stored by at least one non-transitory computer -readable storage medium with the computer program instructions, such as the computer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIGs. 7 to 10. In other embodiments, the computer program instructions, such as the computer -readable program code portions, need not be stored or otherwise embodied by a non-transitory computer -readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer-readable program code portions, still being configured, upon execution, to perform the functions described above.
[0563] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or moreblocks of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.
[0564] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included. Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.
[0565] In the above, syntax and semantics for MinimizedlmageBox has been presented. It is to be understood that embodiments are not limited to the presented syntax and semantics but apply generally to any modifications of the syntax and semantics. For example, instead of sqlite_varint, any variable- or fixed-length representation format for unsigned integers may be used.
[0566] Some examples and embodiments have been described with reference to the MinimizedlmageBox. It is to be understood that examples and embodiments are not limited to a box of this name but apply to any similar box or boxes, or more generally to any similar syntax structure(s).
[0567] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.
[0568] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0569] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associateddrawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0570] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.
[0571] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field- programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like.
[0572] As used herein, the term ‘circuitry’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. This description of ‘circuitry’ applies to uses of this term inthis application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.
[0573] Circuitry or Circuit: As used in this application, the term ‘circuitry’ or ‘circuit’ may refer to one or more or all of the following:
[0574] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware; and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(c) hardware circuit(s) and or processor(s), such as a microprocessor s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0575] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example, and when applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
Claims
CLAIMSWhat is claimed is:
1. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: writing in a file: an indication of presence of a reduced header mode or compact metadata; and the reduced header comprising information for processing one or more items; and signaling the file comprising the reduced header mode.
2. The apparatus of claim 1, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
3. The apparatus of claim 2, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoder configuration record, a network abstraction layer (NAL) unit, a NAL unit header, an open bitstream unit (OBU), an OBU header, a size or length field indicative of a NAL unit size, OBU size, payload size of a NAL unit or OBU, one or more parameter sets of one or more particular types, a sequence header, a picture header, a frame header, a slice header, or a tile group header.
4. The apparatus of any of claims 1 to 3, wherein the apparatus is further caused to perform: defining a major brand field indicating that a file structurally conforms to the file with the reduced header mode; and defining a minor version field for indicating: a slim codec brand, a condensed version of one or more codec-specific structures; or a condensed configuration item property associated with the respective one or more items.
5. The apparatus of any of claims 1 to 4, wherein the apparatus is further caused to perform:including the slim codec brand in the file for indicating a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items.
6. The apparatus of claim 5, wherein the apparatus is caused to perform: defining or deriving an inferred syntax element, wherein the inferred syntax element is used by a reader to the reconstruct the one or more codec-specific structures.
7. The apparatus of claim 1, wherein the apparatus is caused to perform: using a meta-box with a version greater than zero for signaling the minimized header mode.
8. The apparatus of claim 1, wherein the apparatus is caused to perform: indicating the reduced header mode in a file type box, an extended type box, or an original file type box, and wherein when the reduced header mode is indicated in the file type box, the extended type box, or the original file type box, the reduced header moder imposes constraints on a first media data box.
9. The apparatus of claim 1, wherein the apparatus is further caused to perform: indicating a normal header mode in the file comprising the reduced header mode.
10. The apparatus of claim 1, wherein apparatus is further caused to perform: defining a file description box for indicating a presence of both the reduced header mode and the regular header mode in the file.
11. The apparatus of claim 1, wherein apparatus is further caused to perform: deriving a reduced header for the file and using the reduced header as input to a hash generation algorithm; deriving the reduced header excluding pre-defined syntax elements and using the reduced header excluding the pre-defined syntax elements as input to the hash generation algorithm; deriving the reduced header excluding width and height fields for the file and using the reduced header excluding the width and height as an input to the hash generation algorithm; or using a meta-box excluding pre-defined syntax structures or the pre-defined syntax elements as an input to the hash generation algorithm.
12. The apparatus of claim 1, wherein apparatus is further caused to perform: including items sharing same properties in the file.
13. The apparatus of claim 1, wherein the apparatus is further caused to perform: coding and enumerating the reduced header mode in a web page html source, and wherein multiple files reuse the codes, and wherein the apparatus is further caused to perform: signaling byteranges for item data associated with the multiple files.
14. The apparatus of claim 1, wherein the file is referenced and / or declared in a hyper text markup language (HTML) formatted document, and wherein when the file is referenced and / or declared in the hyper text markup language (HTML) formatted document plurality of files share the reduced header mode metadata.
15. The apparatus of claim 1, wherein the apparatus is further caused to perform: defining an indicator in the HTML context to enable download one of the files, of the plurality of files, with the reduced header mode and apply processed reduced header mode metadata to the plurality of files without re-downloading respective reduced header mode metadata.
16. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving and parsing a file comprising: an indication of presence of a reduced header mode or compact metadata; and a reduced header comprising data for processing one or more items in the file; and processing the reduced header data to extract information needed for decoding of the one or more items in the file.
17. The apparatus of claim 16, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
18. The apparatus of claim 16, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoder configuration record, one or more parameter sets of one or more particular types, a sequence header, a picture header, or a slice header.
19. The apparatus of any of the claims 16 to 18, wherien the apparatus is further caused to perform: parsing a slim codec brand from the file, wherein the slim codec brand in indicates a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items; and wherein the apparatus is further caused to perform: parsing a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items from the slim coded brand.
20. The apparatus of any of claims 18 or 19, wherein the apparatus is further caused to perform: parsing an inferred syntax element; and using the inferred syntax element to reconstruct the one or more codec-specific structures.
21. The apparatus of claim 20, wherein the apparatus is further caused to perform: reconstructing one or more codec-specific structures from the respective parsed condensed structures by adding the inferred syntax elements; and including a value of the inferred syntax element into the reconstructed one or more codec-specific structures.
22. The apparatus of claim 16, wherein items sharing same properties are comprised in the same file.
23. The apparatus of claim 22, wherein the file is downloaded once and item properties are shared between the one or more items.
24. The apparatus of claim 16, wherein the reduced hear mode is coded and enumerated in a web page html source, and wherein multiple files reuse the codes, and wherein the apparatus is caused to perform: receiving the web page; downloading the reduced header mode; and reusing the coded reduced header moder for different files, without re-downloading coded reduced header mode.
25. A method comprising:writing in a file: an indication of presence of a reduced header mode or compact metadata; and the reduced header comprising information for processing one or more items; and signaling the file comprising the reduced header mode.
26. The method of claim 25, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
27. The method of claim 26, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoder configuration record, a network abstraction layer (NAL) unit, a NAL unit header, an open bitstream unit (OBU), an OBU header, a size or length field indicative of a NAL unit size, OBU size, payload size of a NAL unit or OBU, one or more parameter sets of one or more particular types, a sequence header, a picture header, a frame header, a slice header, or a tile group header.
28. The method of any of claims 25 to 27 further comprising: defining a major brand field indicating that a file structurally conforms to the file with the reduced header mode; and defining a minor version field for indicating: a slim codec brand, a condensed version of one or more codec-specific structures; or a condensed configuration item property associated with the respective one or more items.
29. The method of any of claims 25 to 28 further comprising: including the slim codec brand in the file for indicating a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items.
30. The method of claim 29 further comprising defining or deriving an inferred syntax element, wherein the inferred syntax element is used by a reader to the reconstruct the one or more codec-specific structures.
31. The method of claim 25 further comprising: using a meta-box with a version greater than zero for signaling the minimized header mode.
32. The method of claim 25 further comprising: indicating the reduced header mode in a file type box, an extended type box, or an original file type box, and wherein when the reduced header mode is indicated in the file type box, the extended type box, or the original file type box, the reduced header moder imposes constraints on a first media data box.
33. The method of claim 25 further comprising: indicating a normal header mode in the file comprising the reduced header mode.
34. The method of claim 25 further comprising: defining a file description box for indicating a presence of both the reduced header mode and the regular header mode in the file.
35. The method of claim 25 further comprising: deriving a reduced header for the file and using the reduced header as input to a hash generation algorithm; deriving the reduced header excluding pre-defined syntax elements and using the reduced header excluding the pre-defined syntax elements as input to the hash generation algorithm; deriving the reduced header excluding width and height fields for the file and using the reduced header excluding the width and height as an input to the hash generation algorithm; or using a meta-box excluding pre-defined syntax structures or the pre-defined syntax elements as an input to the hash generation algorithm.
36. The method of claim 25 further comprising: including items sharing same properties in the file.
37. The method of claim 25 further comprising: coding and enumerating the reduced header mode in a web page html source, and wherein multiple files reuse the codes, and wherein the method further comprises: signaling byte-ranges for item data associated with the multiple files.
38. The method of claim 25, wherein the file is referenced and / or declared in a hyper text markup language (HTML) formatted document, and wherein when the file is referenced and / or declared in the hyper text markup language (HTML) formatted document plurality of files share the reduced header mode metadata.
39. The method of claim 25 further comprising defining an indicator in the HTML context to enable download one of the files, of the plurality of files, with the reduced header mode and apply processed reduced header mode metadata to the plurality of files without re-downloading respective reduced header mode metadata.
40. A method comprising: receiving and parsing a file comprising: an indication of presence of a reduced header mode or compact metadata; and a reduced header comprising data for processing one or more items in the file; and processing the reduced header data to extract information needed for decoding of the one or more items in the file.
41. The method of claim 40, wherein the one or more items are encoded with a network abstract layer (NAL) unit based codec or an non-NAL unit based codec.
42. The method of claim 40, wherein one or more codec-specific structures comprise one or more of the following: a configuration item property, a decoder configuration record, one or more parameter sets of one or more particular types, a sequence header, a picture header, or a slice header.
43. The method of any of the claims 40 to 42 further comprising: parsing a slim codec brand from the file, wherein the slim codec brand in indicates a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items; and wherein the method further comprises: parsing a condensed version of one or more codec-specific structures or a condensed configuration item property associated with the respective one or more items from the slim coded brand.
44. The method of any of claims 42 or 43 further comprising: parsing an inferred syntax element; and using the inferred syntax element to reconstruct the one or more codec-specific structures.
45. The method of claim 44 further comprising:reconstructing one or more codec-specific structures from the respective parsed condensed structures by adding the inferred syntax elements; and including a value of the inferred syntax element into the reconstructed one or more codec-specific structures.
46. The method of claim 40, wherein items sharing same properties are comprised in the same file.
47. The method of claim 46, wherein the file is downloaded once and item properties are shared between the one or more items.
48. The method of claim 40, wherein the reduced hear mode is coded and enumerated in a web page html source, and wherein multiple files reuse the codes, and wherein the method further comprises: receiving the web page; downloading the reduced header mode; and reusing the coded reduced header moder for different files, without re-downloading coded reduced header mode.
49. An apparatus comprising means for performing the methods as claimed in any of the claims 25 to 39.
50. An apparatus comprising means for performing the methods as claimed in any of the claims 40 to 48.
51. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 25 to 39.
52. The computer readable medium of claim 51, wherein the computer readable medium comprises a non-transitory computer readable medium.
53. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 40 to 48.
54. The computer readable medium of claim 53, wherein the computer readable medium comprises a non-transitory computer readable medium.
55. An apparatus comprising at least one processor; and at least one non -transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: writing, in a file, a slim codec brand indicative of a condensed configuration item property associated with an image item; and writing, in the file, the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition.
56. The apparatus of claim 55, wherein the file further comprises an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
57. The apparatus of any of claims 55 or 56, wherein for writing the condensed configuration item property, the apparatus is caused to perform: writing a condensed video coding layer (VCL) network abstraction layer (NAL) unit by excluding the NAL unit length and / or the NAL unit header from an item data stored in the file.
58. The apparatus of the claim 57, wherein for writing the condensed configuration item property, the apparatus is further caused to perform: including other NAL units of a bitstream into the condensed configuration item property.
59. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving and parsing a file comprising: a slim codec brand indicative of a condensed configuration item property associated with an image item; and the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition; andprocessing the slim codec brand to extract information needed for decoding of one or more items in the file.
60. The apparatus of claim 59, wherein the one or more items comprise an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
61. The apparatus of any of the claims 59 or 60, wherein the condensed configuration item property comprises a condensed video coding layer (VCL) network abstraction layer (NAL) unit excluding the NAL unit length and / or the NAL unit header from an item data stored in the file, and wherein the apparatus is further caused to perform: reconstructing a VCL NAL unit from the condensed VCL NAL unit by including the NAL unit length set equal to an item size and / or the NAL unit header set equal to pre-defined values defined in the slim codec brand.
62. A method comprising: writing, in a file, a slim codec brand indicative of a condensed configuration item property associated with an image item; and writing, in the file, the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition.
63. The method of claim 62, wherein the file further comprises an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
64. The method of any of claims 62 or 63, wherein writing the condensed configuration item property comprises: writing a condensed video coding layer (VCL) network abstraction layer (NAL) unit by excluding the NAL unit length and / or the NAL unit header from an item data stored in the file.
65. The method of the claim 64, wherein writing the condensed configuration item property further comprises: including other NAL units of a bitstream into the condensed configuration item property.
66. A method comprising; receiving and parsing a file comprising: a slim codec brand indicative of a condensed configuration item property associated with an image item; and the condensed configuration item property for carrying corresponding condensed decoder configuration record, wherein the condensed configuration item property is similar or substantially similar to a respective decoder configuration record but excludes syntax elements pre-defined in the slim coded brand definition; and processing the slim codec brand to extract information needed for decoding of one or more items in the file.
67. The method of claim 66, wherein the one or more items comprise an image coded as an instantaneous decoder refresh (IDR) picture with a single slice per picture resulting in an encoded picture with one video coding layer (VLC) network abstraction layer (NAL) unit.
68. The method of any of the claims 66 or 67, wherein the condensed configuration item property comprises a condensed video coding layer (VCL) network abstraction layer (NAL) unit excluding the NAL unit length and / or the NAL unit header from an item data stored in the file, and wherein the method further comprises: reconstructing a VCL NAL unit from the condensed VCL NAL unit by including the NAL unit length set equal to an item size and / or the NAL unit header set equal to pre-defined values defined in the slim codec brand.
Citation Information
Patent Citations
Systems and methods for accessing digital data
US20150334413A1
High level syntax for video coding and decoding
US20230145618A1