A method for signalling large images in heif
The HEIF standard addresses the challenges of encoding and decoding large images by leveraging its features on ISOBMFF, enabling efficient storage and decoding of high-resolution images through tile-based access, surpassing the limitations of geoTIFF.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2026-01-09
- Publication Date
- 2026-07-23
AI Technical Summary
Existing technologies face challenges in efficiently encoding and decoding large images, particularly those with resolutions exceeding tera or peta pixels, such as geospatial images, due to limitations in current file formats like geoTIFF, which do not leverage the capabilities of the High Efficiency Image File Format (HEIF) for such high-resolution imagery.
Implementing the HEIF standard for encoding and decoding large images, utilizing its features built on the ISO Base Media File Format (ISOBMFF) to efficiently store and manage high-resolution images, including mechanisms for tile-based access and decoding of image tiles independently.
The HEIF standard enables efficient storage and decoding of large images, allowing for seamless tile-based access and improved handling of high-resolution imagery, overcoming limitations of geoTIFF and enhancing the functionality for geospatial applications.
Smart Images

Figure EP2026050446_23072026_PF_FP_ABST
Abstract
Description
A METHOD FOR SIGNALLING LARGE IMAGES IN HEIFCROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 745,545, filed January 15, 2025. The entire content of the above-referenced application is hereby incorporated by reference.TECHNICAL FIELD
[0002] The example and non-limiting embodiments relate generally to coding of images and, more particularly, to the use of HEIF to code large images.BACKGROUND
[0003] It is known, in image processing, to provide large images using the geoTIFF format.SUMMARY
[0004] The following summary is merely intended to be illustrative. The summary is not intended to limit the scope of the claims.
[0005] According to some aspects, there is provided the subject matter of the independent claims. Some further aspects are defined in the dependent claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The foregoing aspects and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0007] FIG. 1 is a block diagram of one possible and non-limiting example system in which the example embodiments may be practiced;
[0008] FIG. 2 is a block diagram of one possible and non-limiting exemplary system in which the example embodiments may be practiced;
[0009] FIGs. 3-24 are diagrams illustrating features as described herein; and
[0010] FIGs. 25-26 are flowcharts illustrating steps as described herein.DETAILED DESCRIPTION OF EMBODIMENTS
[0011] The following describes suitable apparatus and possible mechanisms for practicing example embodiments of the present disclosure. Accordingly, reference is first made to FIG. 1 , which shows an exampleblock diagram of an apparatus 50. The apparatus may be configured to perform various functions such as, for example, gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus, or the like. A device configured to encode a video scene may (optionally) comprise one or more microphones for capturing the scene and / or one or more sensors, such as cameras, for capturing information about the physical environment in which the scene is captured. Alternatively, a device configured to encode a video scene may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. A device configured to decode and / or render the video scene may be configured to receive a Moving Picture Experts Group immersive codec family (MPEG-I) bitstream comprising the encoded video scene. A device configured to decode and / or render the video scene may comprise one or more speakers / audio transducers and / or displays, and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers / audio transducers and / or displays. A device configured to decode and / or render the video scene may comprise a user equipment, a head / mounted display, or another device capable of rendering to a user an augmented reality (AR), virtual reality (VR) and / or mixed reality (MR) experience. Additionally, or alternatively, a device may be configured to encode and / or decode image information and / or untimed images.
[0012] The electronic device 50 may for example be a mobile terminal or user equipment of a wireless communication system. Alternatively, the electronic device may be a computer or part of a computer that is not mobile. It should be appreciated that example embodiments of the present disclosure may be implemented within any electronic device or apparatus which may process data. The electronic device 50 may comprise a device that can access a network and / or cloud through a wired or wireless connection. The electronic device 50 may comprise one or more processors 56, one or more memories 58, and one or more transceivers 52 interconnected through one or more buses. The one or more processors 56 may comprise a central processing unit (CPU) and / ora graphical processing unit (GPU). Each of the one or more transceivers 52 includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. A “circuit” may include dedicated hardware or hardware in association with software executable thereon. The one or more transceivers may be connected to one or more antennas 44. The one or more memories 58 may include computer readable code. The one or more memories 58 and the computer readable code may be configured to, with the one or more processors 56, cause the electronic device 50 to perform one or more of the operations as described herein.
[0013] The electronic device 50 may connect to a node of a network. The network node may comprise one or more processors, one or more memories, and one or more transceivers interconnected through one ormore buses. Each of the one or more transceivers includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers may be connected to one or more antennas. The one or more memories may include computer readable code. The one or more memories and the computer readable code may be configured to, with the one or more processors, cause the network node to perform one or more of the operations as described herein.
[0014] The electronic device 50 may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The electronic device 50 may further comprise an audio output device 38 which in example embodiments of the present disclosure may be any one of: an earpiece, speaker, or an analogue audio or digital audio output connection. The electronic device 50 may also comprise a battery (or in other example embodiments of the present disclosure the device may be powered by any suitable mobile energy device such as solar cell, fuel cell, or clockwork generator). The electronic device 50 may further comprise a camera 42 or other sensor capable of recording or capturing images and / or video. Additionally, or alternatively, the electronic device 50 may further comprise a depth sensor. The electronic device 50 may further comprise a display 32. The electronic device 50 may further comprise an infrared port for short range line of sight communication to other devices. In other example embodiments of the present disclosure the apparatus 50 may further comprise any suitable short-range communication solution such as for example a BLUETOOTH™ wireless connection or a USB / firewire wired connection.
[0015] It should be understood that an electronic device 50 configured to perform example embodiments of the present disclosure may have fewer and / or additional components, which may correspond to what processes the electronic device 50 is configured to perform. For example, an apparatus configured to encode a video might not comprise a speaker or audio transducer and may comprise a microphone, while an apparatus configured to render the decoded video might not comprise a microphone and may comprise a speaker or audio transducer. For example, an apparatus configured to encode and / or decode images may not comprise a speaker, audio transducer, or microphone; an apparatus configured to encode images may comprise an image sensor but not an image display, and an apparatus configured to decode images may comprise an image display but not an image sensor.
[0016] Referring now to FIG. 1, the electronic device 50 may comprise a controller 56, processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to memory 58 which in example embodiments of the present disclosure may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connectedto codec circuitry 54 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.
[0017] The electronic device 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader, for providing user information and being suitable for providing authentication information for authentication and authorization of the user / electronic device 50 at a network. The electronic device 50 may further comprise an input device 34, such as a keypad, one or more input buttons, or a touch screen input device, for providing information to the controller 56.
[0018] The electronic device 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).
[0019] The electronic device 50 may comprise a microphone 38, camera 42, and / or other sensors capable of recording or detecting audio signals, image / video signals, and / or other information about the local / virtual environment, which are then passed to the codec 54 or the controller 56 for processing. The electronic device 50 may receive the audio / image / video signals and / or information about the local / virtual environment for processing from another device prior to transmission and / or storage. The electronic device 50 may also receive either wirelessly or by a wired connection the audio / image / video signals and / or information about the local / virtual environment for encoding / decoding. The structural elements of electronic device 50 described above represent examples of means for performing a corresponding function.
[0020] The memory 58 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The memory 58 may be a non-transitory memory. The memory 58 may be means for performing storage functions. The controller 56 may be or comprise one or more processors, which may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multicore processor architecture, as non-limiting examples. The controller 56 may be means for performing functions.
[0021] The electronic device 50 may be configured to perform capture of a volumetric scene according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a camera 42 or other sensor capable of recording or capturing images and / or video. The electronic device 50 mayalso comprise one or more transceivers 52 to enable transmission of captured content for processing at another device. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.
[0022] With respect to FIG. 2, an example of a system within which example embodiments of the present disclosure can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, E-UTRA, LTE, CDMA, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a BLUETOOTH™ personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and / or the Internet. A wireless network may implement network virtualization, which is the process of combining hardware and software network resources and network functionality into a single, software-based administrative entity, a virtual network. Network virtualization involves platform virtualization, often combined with resource virtualization. Network virtualization is categorized as either external, combining many networks, or parts of networks, into a virtual unit, or internal, providing network-like functionality to software containers on a single system. For example, a network may be deployed in a tele cloud, with virtualized network functions (VNF) running on, for example, data center servers. For example, network core functions and / or radio access network(s) (e.g. CloudRAN, O-RAN, edge cloud) may be virtualized. Note that the virtualized entities that result from the network virtualization are still implemented, at some level, using hardware such as processors and memories, and also such virtualized entities create technical effects.
[0023] It may also be noted that operations of example embodiments of the present disclosure may be carried out by a plurality of cooperating devices (e.g. cRAN).
[0024] The system 10 may include both wired and wireless communication devices and / or electronic devices suitable for implementing example embodiments of the present disclosure.
[0025] For example, the system shown in FIG. 2 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0026] The example communication devices shown in the system 10 may include, but are not limited to, an apparatus 15, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, and a head-mounted display (HMD) 17. The electronic device 50 may comprise any of those example communication devices. In an example embodiment of the present disclosure, more than one of these devices, or a plurality of one or more ofthese devices, may perform the disclosed process(es). These devices may connect to the internet 28 through a wireless connection 2.
[0027] The example embodiments of the present disclosure may also be implemented in a set-top box; i.e. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding. The example embodiments of the present disclosure may also be implemented in cellular telephones such as smart phones, tablets, personal digital assistants (PDAs) having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and playback appliances having wireless communication capabilities, Internet appliances permitting wireless Internet access and browsing, tablets with wireless communication capabilities, as well as portable units or terminals that incorporate combinations of such functions. Additionally, or alternatively, the example embodiments of the present disclosure may also be implemented in a client, a server, a file reader, a file writer, a player, a file parser, an image caching server, etc.
[0028] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24, which may be, for example, an eNB, gNB, access point, access node, other node, etc. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.
[0029] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), BLUETOOTH™, IEEE 802.11, 3GPP Narrowband Internet of Things (loT) and any similar wireless communication technology. A communications device involved in implementing various example embodiments of the present disclosure may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0030] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may referto a logical connection over a multiplexed medium, capable of conveying several logicalchannels. A channel may be used for conveying an information signal, for example a bitstream, which may be a MPEG-I bitstream, from one or several senders (or transmitters) to one or several receivers.
[0031] Having thus introduced one suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described with greater specificity.
[0032] Features as described herein may generally relate to coding and decoding of images and / or video (i.e. a timed sequence of images). A codec consists of an encoder that transforms input into a compressed representation suited for storage / transmission, and a decoder that can decompress the compressed representation back into a viewable form. Typically, the encoder discards some information in the original input to be able to represent the input in a more compact form (that is, at a lower bitrate). A codec may be configured to code an image and / or video.
[0033] Typical hybrid video codecs, such as H.264 / AVC, H.265 / HEVC and H.266 / WC, encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) (310) may be predicted (320) for example by motion compensation means (finding and indicating an area in one of the previously coded pictures that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, may be coded (330). This may typically be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it) (340), quantizing the resulting transform coefficients (350), and entropy coding the quantized coefficients (360). By varying the fidelity of the quantization process, the encoder may control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate). An example of the encoding process is illustrated in FIG. 3. While a video encoder is described, a similar process may be performed for encoding of an image.
[0034] In some video codecs, such as H.265 / HEVC and H.266 / WC, the video pictures may be divided into coding units (CU) covering the area of the picture. A CU may consist of one or more prediction units (PU) defining the prediction process for the samples within the CU and one or more transform units (TU) defining the prediction error coding process for the samples in the said CU. Typically, a CU may consist of a rectangular block of samples with a size selectable from a predefined set of possible CU sizes. A CU with the maximum allowed size may typically be named as LCU (largest coding unit) or CTU (coding tree unit), and the video picture may be divided into non-overlapping CTUs. A CTU may be further split into a combination of smaller CUs, e.g. by recursively splitting the CTU and resultant CUs. Each resulting CU typically may have at least one PU and at least one TU associated with it. Each PU and TU may be further split into smaller PUs and TUs to increasegranularity of the prediction and prediction error coding processes, respectively. Each PU may have prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g. motion vector information for inter predicted PUs and intra prediction directionality information for intra predicted PUs). Similarly, each TU may be associated with information describing the prediction error decoding process for the samples within the TU (including e.g. DCT coefficient information). It is typically signaled at CU level whether prediction error coding may be applied or not for each CU. In the case there is no prediction error residual associated with the CU, it may be considered that there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and Tus, may typically be signaled in the bitstream, and may allow the decoder to reproduce the intended structure of these units.
[0035] The decoder may reconstruct the output video by applying prediction means (410) similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain) (420). After applying prediction and prediction error decoding means, the decoder may sum up the prediction and prediction error signals (pixel values) to form the output video frame (440). The decoder (and encoder) may also apply additional filtering means (430) to improve the quality of the output video before passing it for display and / or storing it as a prediction reference for the forthcoming frames in the video sequence. An example of the decoding process is illustrated in FIG. 4. While a video decoder is described, a similar process may be performed for decoding of an image.
[0036] FIG. 5 illustrates diagrams of an example encoder (502) and an example decoder (540), which may be used to code images, for example JPEG or PNG images. In the encoder (502), input pictures (504) may be divided into a CU or CTU (506), a prediction block may be subtracted (508) to form a residual (510), which may be transformed (512) and quantized (514) before coding (516) as compressed bits (518) into a bitstream. The quantized transformed coefficients may also be dequantized / inverse quantized (520) and inverse transformed (522), and combined with the output of a prediction block (524). The result may then be used in intra prediction (526) and may be (e.g. in parallel) in-loop filtered (528), included in a decoded picture buffer (530), and used in inter prediction (532). In the decoder (540), compressed bits (542) may be decoded (544), dequantized (546), inverse transformed (548), and combined with the output of a prediction block (550). The result may then be used in intra prediction (552) and may be (e.g. in parallel) in-looped filtered (554), included in a decoded picture buffer (556), and used in inter prediction (558). The contents of the decoded picture buffer (556) may be output (560).
[0037] As shown in FIG. 5, the decoder (540) may actually be part of the coding loop of the encoder (502) in a reverse way (e.g. 520-532). The quantized transformed coefficients in the encoder (e.g. after quantization block (514), or the output of context-adaptive binary arithmetic coding (CABAC) (344) in the decoder) may be dequantized (520, 546) and inverse transformed (522, 548), generating the coded residual block (e.g. 524, 550). The (intra or inter) prediction block (526, 532, 552, 558) may then be added to the coded residual block (524, 550), generating the reconstructed block. In-loop filtering may be performed over the reconstructed block (528, 554), forming the final reconstructed block. The final reconstructed blocks may be stored in a decoded picture buffer (530, 556) for output (560), as well as for possible use of future coding.
[0038] The example codecs illustrated in FIGs. 3-5 are merely examples; example embodiments of the present disclosure may be practiced in conjunction with other codecs, or codecs including fewer, more, or different steps.
[0039] Features as described herein may generally relate to encoding and decoding of large images, or images whose resolution are in the order of tera pixels, peta pixels, or beyond. In geospatial applications, the image resolution commonly exceeds 300,000 by 300,000 pixels, and the image sizes are growing. Geospatial images are images of geographical locations, which are typically captured with satellites. The images may be captured using different sensors, and not necessarily with traditional cameras (e.g. infrared sensors, heat sensors, electromagnetic sensors, light detection and ranging sensors, etc.). Geospatial images are typically accessed by a viewer or application over a network using "cloud optimization" techniques.
[0040] Image storage techniques, such as tiles / grids and image pyramids / overviews, are used for simplified access to parts of, and lower resolution versions of, the image. Tiles are usually set based on end device capabilities. 512x512 or 1 Kx1 K tile configurations are common to match handheld and desktop devices. As a user navigates the large image space, individual tiles are pulled / requested depending on, for example, pan and zoom commands from the user (e.g. typically through a browser interface). Hypertext transfer protocol (HTTP) byte range requests are typically used to achieve efficiency and to enable functionality using simple built-in browser and web server capabilities.
[0041] Currently, cloud optimized geoTIFF is used for large images or geospatial images. However, the high efficiency image file format (HEIF) provides significant feature / capabi lity benefits over geoTIFF, so there is significant interest in the standardization community for using HEIF for such large geographic images.
[0042] The HEIF standard builds on top of the International Standards Organization (ISO) base media file format (ISO / IEC 14496-12, which may be abbreviated ISOBMFF). Other available media file format standards include Moving Picture Experts Group (MPEG)-4 file format (ISO / IEC 14496-14, also known as theMP4 format), file format for NAL (Network Abstraction Layer) unit structured video (ISO / IEC 14496-15), and High Efficiency Video Coding standard (HEVC or H.265 / HEVC).
[0043] Some concepts, structures, and specifications of ISOBMFF are described below as an example of a container file format, based on which some example embodiments of the present disclosure may be implemented. The aspects of the disclosure are not limited to ISOBMFF, but rather the description is given as one possible basis on top of which at least some example embodiments may be partly or fully realized.
[0044] A basic building block in the ISO base media file format is called a box. Each box has a header and a payload. The box header indicates the type of the box and the size of the box in terms of bytes. Box type is typically identified by an unsigned 32-bit integer, interpreted as a four-character code (4CC). A box may enclose other boxes, and the ISO file format specifies which box types are allowed within a box of a certain type. Furthermore, the presence of some boxes may be mandatory in each file, while the presence of other boxes may be optional. Additionally, for some box types, it may be allowable to have more than one box present in a file. Thus, the ISO base media file format may be considered to specify a hierarchical structure of boxes.
[0045] In files conforming to the ISO base media file format, the media data may be provided in one or more instances of MediaDataBox (‘mdat’), and the MovieBox (‘moov’) may be used to enclose the metadata for timed media. In some cases, for a file to be operable, both of the ‘mdat’ and ‘moov’ boxes may be required to be present. The ‘moov’ box may include one or more tracks, and each track may reside in one corresponding TrackBox (‘trak’). Each track is associated with a handler, identified by a four-character code, specifying the track type. Video, audio, and image sequence tracks can be collectively called media tracks, and they contain an elementary media stream. Other track types comprise hint tracks and timed metadata tracks.
[0046] Tracks comprise samples, such as audio or video frames. For video tracks, a media sample may correspond to a coded picture or an access unit.
[0047] A media track refers to samples (which may also be referred to as media samples) formatted according to a media compression format (and its encapsulation to the ISO base media file format). A hint track refers to hint samples, containing cookbook instructions for constructing packets for transmission over an indicated communication protocol. A timed metadata track may refer to samples describing referred media and / or hint samples.
[0048] The 'trak' box includes in its hierarchy of boxes the SampleDescriptionBox, which gives detailed information about the coding type used, and any initialization information needed for that coding. The SampleDescriptionBox contains an entry-count and as many sample entries as the entry-count indicates. The format of sample entries is track-type specific but derived from generic classes (e.g. VisualSampleEntry,AudioSampleEntry). Which type of sample entry form is used for derivation of the track-type specific sample entry format is determined by the media handler of the track.
[0049] The track reference mechanism may be used to associate tracks with each other. The TrackReferenceBox includes box(es), each of which provides a reference from the containing track to a set of other tracks. These references are labeled through the box type (e.g., the four-character code of the box) of the contained box(es).
[0050] In ISOMBFF, an edit list provides a mapping between the presentation timeline and the media timeline. Among other things, an edit list provides for the linear offset of the presentation of samples in a track, provides for the indication of empty times, and provides for a particular sample to be dwelled on for a certain period of time. The presentation timeline may be accordingly modified to provide for looping, such as for the looping of videos of the various regions of the scene. One example of the box that includes the edit list, the EditListBox, is provided in FIG. 6. In ISOBMFF, an EditListBox may be contained in EditBox, which is contained in TrackBox ('trak').
[0051] In the example of FIG. 6, flags (610) may specify the repetition of the edit list. By way of example, setting a specific bit within the box flags (the least significant bit, i.e., flags & 1 in ANSI-C notation, where & indicates a bit-wise AND operation) equal to 0 specifies that the edit list is not repeated, while setting the specific bit (i.e., flags & 1 in ANSI-C notation) equal to 1 specifies that the edit list is repeated. The values of box flags (610) greater than 1 may be defined to be reserved for future extensions. As such, when the edit list box indicates the playback of zero or one samples, (flags & 1) may be equal to zero. When the edit list is repeated, the media at time 0 resulting from the edit list follows immediately the media having the largest time resulting from the edit list, such that the edit list is repeated seamlessly.
[0052] In ISOBMFF, a Track group enables grouping of tracks based on certain characteristics, or the tracks within a group having a particular relationship. Track grouping, however, does not allow any image items in the group. An example of the syntax of TrackGroupBox in ISOBMFF is illustrated in FIG. 7. track_group_type (710) indicates the grouping_type and may be set to a group ID, or a registered value, ora value from a derived specification or registration.
[0053] ' msrc' indicates that this track belongs to a multi-source presentation. The tracks that have the same value of track_group_id within a TrackGroupTypeBox of track_group_type 'msrc' are mapped as being originated from the same source. For example, a recording of a video telephony call may have both audio and video for both participants, and the value of track_group_id associated with the audio track and the video track of one participant differs from value of track_group_id associated with the tracks of the other participant.
[0054] The pair of track_group_id and track_group_type identifies a track group within the file. The tracks that contain a particular TrackGroupTypeBox having the same value of track_group_id and track_group_type belong to the same track group.
[0055] The Entity grouping is similar to track grouping, but enables grouping of both tracks and image items in the same group.
[0056] An example of the syntax of EntityToGroupBox in ISOBMFF is illustrated in FIG. 8. groupjd (810) is a non-negative integer assigned to the particular grouping that may not be equal to any groupjd value of any other EntityToGroupBox, any itemJD value of the hierarchy level (file, movie, or track) that contains the GroupsListBox, or any trackJD value (i.e. when the GroupsListBox is contained in the file level). num_entitiesjn_group (820) specifies the number of entityjd values mapped to this entity group, entityjd (830) is resolved to an item, when an item with itemJD equal to entityjd is present in the hierarchy level (file, movie or track) that contains the GroupsListBox, or to a track, when a track with track_ID equal to entityjd is present and the GroupsListBox is contained in the file level.
[0057] Files conforming to the ISOBMFF may contain any non-timed objects, referred to as items, meta items, or metadata items, in a meta box (four-character code: ‘meta’). While the name of the meta box refers to metadata, items can generally contain metadata or media data. The meta box may reside at the top level of the file, within a movie box (four-character code: ‘moov’), and / or within a track box (four-character code: rak’), but at most one meta box may occur at each of the file level, movie level, or track level. The meta box may be required to contain a ‘hdlr’ box indicating the structure or format of the ‘meta’ box contents. The meta box may list and characterize any number of items that can be referred to, and each one of them can be associated with a file name, and are uniquely identified with the file by item identifier (itemjd), which is an integer value. The metadata items may be, for example, stored in the 'idat' box of the meta box, or in an 'mdat' box, or reside in a separate file. If the metadata is located external to the file, then its location may be declared by the DatalnformationBox (four-character code: ‘dinf’).
[0058] In the specific case that the metadata is formatted using extensible Markup Language (XML) syntax and is required to be stored directly in the MetaBox, the metadata may be encapsulated into either the XMLBox (four-character code: ‘xml ’) or the BinaryXMLBox (four-character code: ‘bxml’). An item may be stored as a contiguous byte range, or it may be stored in several extents, each being a contiguous byte range. In other words, items may be stored fragmented into extents, e.g. to enable interleaving. An extent is a contiguous subset of the bytes of the resource. The resource can be formed by concatenating the extents.
[0059] A common base structure is used to contain general untimed metadata. This structure is called the MetaBox, as it was originally designed to carry metadata, i.e. data that is annotating other data. However, itis now used for a variety of purposes, including the carriage of data that is not annotating other data, especially when present at ‘file level’. Referring now to FIG. 9, illustrated is an example of syntax for MetaBox, including some optional information. The MetaBox is required to contain a HandlerBox (910) indicating the structure or format of the MetaBox contents. All other contained boxes are specific to the format specified by the HandlerBox. The structure or format of the metadata is declared by the handler. In the case that the primary data is identified by a primary item (960), and that primary item has an item information entry with an item_type, the handler type may be the same as the item_type.
[0060] At most one MetaBox may occur at each of the file level, segment level, movie level, or track level.
[0061] It may be noted that the MetaBox is a container box that extends FullBox (920), not Box.
[0062] The other boxes defined here may be defined as optional or mandatory for a given format. If they are used, then they may take the form specified here. These optional boxes include a DatalnformationBox (930), which documents other files in which metadata values (e.g. pictures) are placed, and an ItemLocationBox (940), which documents where in those files each item is located (e.g. in the common case of multiple pictures stored in the same file). If the DatalnformationBox carries uniform resource locators (URL) to external files, then the ItemLocationBox may be empty. The external files may be self-contained HEIF files, with no need of any offsets. A DataEntryTiledltemURLBox may be present in the DatalnformationBox.
[0063] If an ItemProtectionBox (950) occurs, then some or all of the metadata, including possibly the primary resource, may have been protected and be un-readable unless the protection system is taken into account.
[0064] Metadata items are identified by itemJD. Within a given MetaBox, a given itemJD may uniquely refer to a single item. When an item is updated in movie fragments, the item_ID refers to the latest received version. Derived specifications may further restrict the criteria for uniqueness of the item_ID: item_ID may be unique among the itemJDs in both file and movie-level boxes, or unique within that set extended with the track_ID of the tracks in a movie box. The item_l D value of 0 should not be used, and may not be used when the set is extended to include trackJDs.
[0065] There are three scopes for itemJDs: file and segments; MovieBox and MovieFragmentBox; and TrackBox and TrackFragmentBox. In other words, there may be only one item with a given itemJD within a given scope (e.g. in the TrackBox and all TrackFragmentBox with the same trackJD).
[0066] The ItemPropertiesBox, which is not illustrated in FIG. 8, enables the association of any item with an ordered set of item properties. Item properties may be regarded as small data records. The ItemPropertiesBox consists of two parts: ItemPropertyContainerBox that contains an implicitly indexed list of item properties, and one or more ItemProperty Association Box(es) that associate items with item properties.
[0067] Features as described herein may generally relate to HEIF, which is a standard developed by the Moving Picture Experts Group (MPEG) for storage of images and image sequences. Among other things, the standard facilitates file encapsulation of data coded according to the High Efficiency Video Coding (HEVC) standard. HEIF includes features building on top of the used ISO Base Media File Format (ISOBMFF), as noted above.
[0068] The ISOBMFF structures and features are used to a large extent in the design of HEIF. The basic design for HEIF comprises still images that are stored as items, and image sequences that are stored as tracks. An item in HEIF is defined as the data that does not require timed processing, as opposed to sample data, and is described by the boxes contained in a MetaBox.
[0069] In the context of HEIF, the following boxes may be contained within the root-level 'meta' box and may be used as described in the following. In HEIF, the handler value of the Handler box of the 'meta' box is 'pict'. The resource (whether within the same file, or in an external file identified by a uniform resource identifier) containing the coded media data is resolved through the Data Information ('d inf) box, whereas the Item Location ('Hoc') box stores the position and sizes of every item within the referenced file. The Item Reference ('iref') box documents relationships between items using typed referencing. If there is an item among a collection of items that is in some way to be considered the most important compared to others, then this item is signaled by the Primary Item ('pitm') box. Apart from the boxes mentioned here, the 'meta' box is also flexible to include other boxes that may be necessary to describe items.
[0070] Any number of image items can be included in the same file. Given a collection of images stored by using the 'meta' box approach, it sometimes is essential to qualify certain relationships between images. Examples of such relationships include indicating a cover image for a collection, providing thumbnail images for some or all of the images in the collection, and associating some or all of the images in a collection with an auxiliary image such as an alpha plane. A cover image among the collection of images is indicated using the 'pitm' box. A thumbnail image or an auxiliary image is linked to the primary image item using an item reference of type 'thmb' or 'auxl', respectively.
[0071] An overview image is described by a grid derived image item or a tiled pre-derived coded image item or a ti le / ti led image item whose reconstructed image is formed from generating a lower resolution, ‘binned’ version of the reconstructed image of a base image item. The base image item is also a tiled image item. The tiling may be implemented using a feature of a specific video codec such as H.265 / HEVC, versatile video coding (WC), or any upcoming video codec, or by using a grid derived image item, or by using a tiled image item. When a grid derived image item is used, the items input to the grid define the tiles. Derived image items may not be used as inputs to the image grid, due to the need for in-place byte range accessing of content (e.g. for on-demand access of tiles of an image). Individual tiles may be written contiguously in memory, thereby allowing access with a single read or write action.
[0072] A pre-derived coded image item representing an overview image or an image item representing the base image that are tiled using a feature of a specific codec may be stored in such a way that each extent identifies the data range corresponding to a tile, and may be associated with a ConstrainedExtentsGrid Property indicating the constraint on the extents and describing the tiling grid.
[0073] In cases where the binned resolution results in a fractional, or incomplete, tile at the end of a row (or column), the last tile in a row (or column) of tiles may be padded with the value zero at the end of the row (or column) to complete the last tile in the row (or column). If necessary, the clean aperture transformative property ('clap') may be applied to crop padded rows and / or columns. The number of tiles in a row (or column) of tiles is determined by dividing the width (or height) of the overview image by the tile size in X (or tile size in Y) and rounding up.
[0074] The image format of the overview images is the same as the base image, i.e. number of bands, bit depth, color format, etc. Overview images can be stacked together with the base image as a series of progressively binned images in an image pyramid entity group.
[0075] The constrained extents grid property may be defined as follows: Box type: 'cexg'; Property type: Descriptive item property; Container: ItemPropertyContainerBox; Mandatory (per item): No; Quantity (per item): At most one. The ConstrainedExtentsGridProperty descriptive item property indicates that each extent of the associated image item in the ItemLocationBox is constrained to enclose data units of the item that are extractable as a contiguous byte range and are independently decodable and renderable as image tiles.
[0076] The configuration data needed to decode each extent independently may be present in the ExtentDecoderConfigurationRecord within the ConstrainedExtentsGridProperty. If the configuration data is not present in the ConstrainedExtentsGridProperty, all data units or properties required to configure the decoder and decode an image tile may be declared in the decoder configuration, as well as initialization properties associated with the image item.
[0077] The reconstructed image of the associated image item is formed from one or more image tiles in a given grid order within a larger canvas.
[0078] The image tiles corresponding to the extents are inserted in row-major order, top-row first, left to right, in the order of the extents for the associated image item within the ItemLocationBox. The value of extent_count within the ItemLocationBox may be equal to (1+rows_minus_one)*(1-K;olumns_minus_one). All image tiles may have exactly the same width and height, image_tile_width and image_tile_height. The reconstructed image is formed by tiling the image tiles into a grid with a column width equal to image_tile_widthand a row height equal to image_tile_height, without gap or overlap. The grid of image tiles may completely “cover” the reconstructed image of the associated image item, where image_tile_width*columns is greater than or equal to image_width and image_tile_height*rows is greater than or equal to image_height, where image_width and image_height are signalled in the ImageSpatialExtentsProperty associated with the image item.
[0079] The flags field is used to signal image tiles related information. The following flags values are defined:
[0080] - 0x000001 field Jength_flag: when set to 1 specifies that the length of the fields image_tile_width and image_tile_height is 32 bits. When fieldjength_flag is set to 0 specifies that the length of the fields image_tile_width and image_tile_height is 16 bits.
[0081] - 0x000002 ti le_config_i nfo_present_f lag : when set to 1 , specifies that the decoder configuration data needed to decode each extent independently is present in the ConstrainedExtentsGridProperty. When ti le_config_i nfo_present_f lag is set to 0, all data units or properties required to configure the decoder and decode the extent independently may be declared in the decoder configuration and initialization properties associated with the image item.
[0082] Referring now to FIG. 10, illustrated is an example of syntax for the ConstrainedExtentsGridProperty. image_tile_width, image_tile_height (1010) specify respectively the width and height in pixels of the image tiles. rows_minus_one, columns_minus_one (1020) specify the number of rows of image tiles, and the number of image tiles per row. The value is one less than the number of rows or columns, respectively. Image tiles enclosed in extents populate the top row first, followed by the second row and following rows, in the order of extents. ExtentDecoderConfigurationRecord (1030) is the decoder configuration record needed to decode the corresponding extents. The decoder configuration record is specific and defined by the image coding format used for encoding the extent data.
[0083] The image pyramid entity group may be defined as follows: Box Type: 'pymd'; Container: GroupsListBox in a MetaBox at file level; Mandatory: No; Quantity: Zero or more. The ImagePyramidEntityGroup indicates a set of image items, formed as a base image item and a series of progressively binned overview image items, which together form the layers of an image pyramid. The ImagePyramidEntityGroup provides overall information for the individual tiles inside the overview image items and base image item of the image pyramid. The image format of the overview images may be the same as the base image (i.e. number of bands, bit depth, color format, etc.). This entity group may contain entityjd values that point to a base image item and a set of overview image items and may contain no entityjd values that point to tracks. The entities may be listedin the order of lowest resolution overview image item to the highest resolution overview image item, followed finally by the base image item of the image pyramid.
[0084] The flags field is used to signal image tiles related information. The following flags values are defined as follows:
[0085] - 0x000001 tile_info_present_flag: when set to 1, specifies that the tile information is present. When set to 0, the tile information is not present and is derived as described below.
[0086] - 0x000002 ti le_info_constant_f lag : when set to 1 , specifies that the tile size is constant across all the layers of the image pyramid. When set to 0, the tile size may not be constant across all the layers of the image pyramid.
[0087] When the flag ti le_info_p resent_f lag is set, the tile information of a layer of the image pyramid may be defined as follows:- tileWidth = tile_size_x- tileHeight = tile_size_y- tileColumns = ceil(ispe.image_width / tileWidth)- tileRows = ceil(ispe.image_height / tileHeight)where ispe is the ImageSpatialExtentsProperty associated with the image item.
[0088] When the flag tile_info_present_flag is not set, the tile information of a layer of the image pyramid is derived depending on the image item as described in TABLES 1-4 below.
[0089] TABLE 1 describes an example of tile information based on ConstrainedExtentsGridProperty associated with the image item:TABLE 1
[0090] TABLE 2 describes an example of tile information based on Tiled image item 'tili':TABLE 2
[0091] TABLE 3 describes an example of tile information based on a grid derived image item with ImageGrid pay load:TABLE 3
[0092] TABLE 4 describes an example of tile information based on 'uncC item property as defined in ISO / IEC 23001-17 associated with the image item:TABLE 4
[0093] There may be multiple ImagePyramidEntityGroups in the same file with different groupjd values.
[0094] All the entities of a same ImagePyramidEntityGroup, or only some of them, may also be members of a same entity group of type 'prgr' if they are stored in the file for allowing a progressive refinement. They may also be members of a same entity group of type 'altr' if they are proposed by the content creator as alternatives to be displayed for players not supporting the ImagePyramidEntityGroup.
[0095] When using region partition groups jointly with an image pyramid, the area covered by a region partition group should correspond to the area of a tile of the image pyramid.
[0096] A region item may be associated with an image item within an ImagePyramidEntityGroup either: through an item reference of type 'cdsc' from the region item to the image item; or through a RegionPartitionGroupBox associated with the image item via an item reference of type 'rpds' and referencing the itemJD of the region item.
[0097] A region item associated with a base image or an overview image within a same ImagePyramidEntityGroup may be applied to the output image of any image item within this ImagePyramidEntityGroup by applying the implicit resampling caused by the difference between the reference space of the region item and the size of the image. It may be noted that a player can use the item reference of type 'base' of a merge region item to filter the region items that are inherited from other image items in the ImagePyramidEntityGroup.
[0098] Referring now to FIG. 11, illustrated is an example of syntax for an ImagePyramidEntityGroup. num_entities_in_group (1110) may be as defined for EntityToGroupBox (i.e. 820 of FIG. 8). In addition, it may also specify the number of layers of the image pyramid. tile_size_x, tile_size_y (1120) may indicate the size in pixels of a tile in the width and height dimension, respectively, for all layers of the image pyramid ifnu m_of_ti le_i nfo (1130) is equal to 1 , or for each layer of the image pyramid in the order of the lowest resolution overview image item to the highest resolution overview image item if num_of_tile_info (1130) is greater than 1.
[0099] The PixellnformationProperty descriptive item property indicates the number and bit depth of colour and alpha / depth components, if present, in the reconstructed image of the associated image item. Referring now to FIG. 12, illustrated is an example of syntax for PixellnformationProperty. In the following, properties of PixellnformationProperty are described.
[0100] bits_per_channel (1205): This field indicates the bits per channel for the pixels of the reconstructed image of the associated image item. The value of this field shall not be 0.
[0101] px_flags&1 (1210): If not 0, indicates that the channeljdc (1215), component_format (1220), and channel _label_f lag (1225) fields are present.
[0102] px_flags&2 (1230): If not 0, indicates that subsampling information is present. Only applicable when px_flags&1 (1210) is not 0.
[0103] channeljdc (1215): This field indicates the contents of the channel. A value of 0 indicates colour / grayscale. A value of 1 indicates alpha. A value of 2 indicates depth. Values 3-7 are reserved for future use. At most one channel shall have a channeljdc (1215) of 1.
[0104] component_format (1220): This field indicates the data type of the channel as defined by the component_format values in ISO / IEC 23001-17 where component J)it_depth is considered to be equal to bits_per_channel (1205).
[0105] channeljabel_flag (1245): This flag indicates the presence of the channel Jabel (1250).
[0106] subsampling_type (1235) : This field indicates the subsampling type as specified by GenericSubsamplingType in Rec. ITU-T H.273 | ISO / IEC 23091-2.
[0107] subsampling Jocation (1240): This field indicates the subsampling sample location as specified by GenericSubsamplingSampleLocType in Rec. ITU-T H.273 ISO / IEC 23091-2.
[0108] channeljabel (1250): The human readable description of the channel.
[0109] Features as described herein may generally relate to tiled image items. A tiled image item is constructed of uniform, independently coded tiles arranged in rows, columns, and optionally extra dimensions, to form a rectangular image or n-dimensional hyperrectangle. The tiles are identical in size, format, coding, and makeup and may be compressed or uncompressed. An image item of type ‘tili’ is a tiled image, with each tile coded independently from other tiles. Input tiles may be stored either in separate external files or in a contiguous range of addressing space to support byte range addressing and retrieval of individual tiles with a single read. It may be noted that the image coding method may be defined by a writer using a valid image codec 4CC. It may be noted that, as opposed to a ‘grid’ image item, where the declaration and addressing of tiles occurs in the file-scoped MetaBox, a ‘tili’ image item with contiguous range of addressing space may have a single declaration parameter in the file-scoped MetaBox and an associated addressing table, with offsets and extents for each tile, stored with the image tiles, typically in a media data box. This has the advantage that the required file ranges of the addressing table, which can be large for terapixel images, may also be loaded on-demand.
[0110] The tiled image item (‘tili’) may be associated with Tiled ImageConfigurationBox (‘tilC’) and ImageSpatialExtentsProperty, which carries the width and height of the overall tiled image. All necessary configuration properties for the given image item type may be defined in the ‘tilC’ property and stored in the tile_image_property[] array.
[0111] When the overall image dimensions are not an even multiple of the image tile size, the rows may be padded on the right to complete the last tile in each row of tiles, and the columns may be padded on the bottom to complete the last tile in each column of tiles. The width and height parameters in the ImageSpatialExtentsProperty may be set to the size of the image containing valid image content, effectively achieving a crop of the padded boundary area.
[0112] The tiled image configuration may be defined as follows: Box type: 'tilC; Property type: Descriptive item property; Container: ItemPropertyContainerBox; Mandatory (per item): Yes, for image items of type 'tili'; Quantity (per item): One. The TiledlmageConfigurationBox may specify parameters associated with a tiled image item (‘tili’). These parameters may include the tile resolution, and the image item type used to code and store individual tile content. Configuration information may also include the number and size of additional dimensions when coding n-dimensional hyperrectangles. This may include support for the coding of multi and hyperspectral imagery where each band in a ‘tili’ tile region is separately retrievable.
[0113] Referring now to FIG. 13, illustrated is an example of syntax for TiledlmageConfigurationBox. tile_width, tile_height (1305) may be set to the size of a single tile width and height. All tiles may have the same size. Tiles at the right or bottom border of the overall image may include padding when the tile width and / or height are not integer multiples of the overall ‘tili’ item width or height. In this case, the ImageSpatialExtentsProperty may be set to the boundary of the true image width and height to achieve a crop of the padded area.
[0114] external_tiles_urls (1310), when set to 0, may specify that the input tile images are each stored using the Tiled ImageOffsetTable and the tile related configuration information is present in TiledlmageConfigurationBox. external_tiles_urls (1320), when set to 1, may specify that the input tile images are each stored in external files as indicated by the URLs in DataEntryTiledltemURLBox and the tile related configuration information is not present in TiledlmageConfigurationBox. In this case, the tile image item data may be empty.
[0115] tile_item_type (1355) may specify the image item type used for all the individual tile images. In a ‘tili’ item, each tile may be coded separately so they can be extracted and decoded independently. tile_item_type (1355) may be set to a valid four-character code fora coded image item (e.g., ‘hvd ’ for h265 compression, ‘j2k1 ’ for JPEG2000, or ‘unci’ for uncompressed). When required by the image item type, all necessary image properties may be stored in tile_image_property[] (1340). Certain codecs (jpg, etc.) may not require any configuration properties.
[0116] number_of_extra_dimensions (1320) may specify the number of extra dimensions if the image resembles a (number_of_extra_dimensions+2)-dimensional hyperrectangle. For a 2D image, number_of_extra_dimensions (1320) may be 0.
[0117] dimension_size[i] (1325) may specify the size of dimension i+2 of the n-dimensional hyperrectangle. Note that the size of the first two dimensions may be the image_width and image_height specified in the ImageSpatialExtentsProperty of the ‘tili’ item.
[0118] sequential_order (1330), when true, may indicate that the compressed image tile data is stored consecutively in sequential order.
[0119] number_of_tile_properties (1335) may specify the number of tile properties stored in the Tiled ImageConfiguration Box.
[0120] tile_image_property[] (1340) may be the image item properties used when decoding a tile image. This may include at least all mandatory item properties for an image item of type tile_item_type (1355), with the exception of ImageSpatialExtentsProperty. If tile_image_property[] (1340) does not contain an ImageSpatialExtentsProperty, the decoder may synthesize an ImageSpatialExtentsProperty with tile_width and tile_height (1305) as the size.
[0121] offset_field_length (1345) may define the number of bits used to store the offset to the image data of a specific tile in the TiledlmageOffsetTable.
[0122] size_field Jength (1350) may define the number of bits used to store the length of the image data of a specific tile in the TiledlmageOffsetTable.
[0123] Tiled image item data may comprise the payload of a tiled image item (‘tili’), which may consist of a TiledlmageOffsetTable and the tiles of the item when the external_tiles_urls (1310) is set to 0.
[0124] The TiledlmageOffsetTable may contain offset pointers and size information for each tile in the item. This table may then be followed by the set of coded image tile data. Referring now to FIG. 14, illustrated is an example of syntax for TiledlmageOffsetTable. tile_start_offset[i] (1410) may point to the start of the coded data of a tile. The position may be given relative to the start of the TiledlmageOffsetTable. If a specific tile is empty and does not contain image content, the tile is not coded and the tile_start_offset[i] (1410) entry may beset to 0. This situation may occur when an image is generated on a canvas and certain portions of the overall image only contain canvas with no image pixels. Readers may interpret a tile_start_offset[i] (1410) value equal to 0 as an empty tile with no media content. Note, a tile_start_offset[i] (1410) value is not a file offset, but an offset into the item's data that may potentially span several iloc extents (i.e. item location extents). tile_size[i] (1420), if present, may indicate the number of bytes of the coded tile bitstream.
[0125] The number of tile offsets stored in the table (NumTiles) may be computed as follows:TileColumns = ceil(lmageSpatialExtentsProperty.image_width / tile_width);TileRows = ceil(lmageSpatialExtentsProperty.image_height / tile_height);NumTiles = TileColumns * TileRowsfor (i=0; i<number_of_extra_dimensions; i++) {NumTiles = NumTiles * dimension_size[i];}
[0126] TileRows and TileColumns may be the number of tiles in a row within the overall image, and the number of tiles in a column within the overall image, respectively. image_width and image_height may be the dimensions of the entire image as specified in the mandatory ImageSpatialExtentsProperty item property. number_of_extra_dimensions and dimension_size[] may be defined in the TiledlmageConfigurationBox property associated with the ‘tili’ item. NumTiles may represent the number of tiles in the entire tiled image item.
[0127] The entries in the offset table may be ordered in row-major sequence. Fora 2D image with a single coded layer, they may be indexed as [y][x], where: x = tile column; y = tile row. For a 3D tiled image item, they may be indexed as [z][y][x], where: x = tile column; y = tile row; z= depth coordinate. Fora general n-dimensional hyperrectangle, the tiles may be indexed as [zn-1] [zn-2] ,..[z3] [z2][y][x], where zi may be the n-2 extra dimensions, x may be the inner most looping variable, followed by y, and then z2 to zn-1.
[0128] The coded tile data may be stored in the file in an arbitrary order, resulting in the tile_start_offset entries not necessarily being in increasing address order.
[0129] When size_field_length==0, the tile_size[i] variables may not be present, and the decoder may infer them from the difference between the tile_start_offset entries. For the case where tiles are stored in sequential order (flags & 0x10 == 0x10), the tile_size[i] may be computed as tile_start_offset[i+1] -ti le_start_offset[i], except for the last tile, which may extend until the end of the data. If the tiles are not stored in sequential order, the decoder may first sort the tile start offsets before computing the size from the offset differences. In this case, the decoder may not be able to read the offset table on-demand. For on-demand applications, the tile sizes may be included. When multiple tiles contain the same content, the ti le_start_offset entries for these tiles may point to the same data block. In this case, sequential ordering may not be used.I
[0130] The data entry tiled item URL box may be defined as follows: Box Type: 'deti'; Container: DataReferenceBox; Mandatory: No; Quantity: Zero or more. The DataEntry Tiled ItemURLBox may identify the location of the external files which carry the input image items of a tiled image item ('tili'). Each location identified by the DataEntryTiled ItemURLBox may correspond to an input image item with a specific itemJD.
[0131] Referring now to FIG. 15, illustrated is an example of the syntax of DataEntryTiledltemURLBox. input_items_size_index (1510) may specify the size of the parameters no_of_input_items (1520) in bytes, with value 0 indicating that the size is of 1 byte, up to the value 7 indicating that the size is to be 8 bytes.
[0132] The parameter no_of_input_items in DataEntryTiledltemURLBox may be equal to:TileColumns = (lmageSpatialExtentsProperty.image_width +TiledlmageConfigurationBox.tile_width- 1) / TiledlmageConfigurationBox.tile_width;TileRows = (lmageSpatialExtentsProperty.image_height + Ti led I mageConfig uration Box.ti le_height- 1) / TiledlmageConfigurationBox.tile_height;no_of_input_items = TileColumns * TileRowsfor (i=0; i<number_of_extra_dimensions; i++) {no_of_input_items = no_of_input_items * dimension_size[i];}where location (1530), in FIG. 15, indicates the location of the referred file as a URL. The URL may be an absolute ora relative URL, and the located resource may be a compliant HEIF file. Relative URLs may be relative to the file that contains the location.
[0133] It may be noted that, when the number of tiles in a tiled image item is high, for example, in geospatial images, the overhead of storing URLs in DataEntryTiledltemURLBox of each tile may be high.
[0134] It may be noted that the item properties related to each tile within a tiled image item is currently defined in the TilelmageConfigurationBox. This may impact the location of presence of Itemproperty child boxes within a HEIF file, and changes to the derivation process of an output image of an image item in a HEIF file may be needed. In other words, currently proposed use of HEIF files to contain large images include all metadata information into a single configuration property which breaks the HEIF file reader processing structure.
[0135] It may be noted that, currently, there is no standardized definition on how a file is created for a tiled image item, nor a definition of a process followed by a file reader when using a file with a tile / tiled image item.
[0136] In an example embodiment, when a file contains an offset, it may enable tiles to be downloaded on demand.
[0137] In an example embodiment, a server, image caching server, encoder, file writer, or file creator may generate or output an initial data segment comprising a reduced URL structure. The file structure of the initial data segment may be according to example embodiments of the present disclosure. The server may host large images (e.g. large geospatial images) which may be divided into spatial regions or tiles, along with metadata to access each tile individually. The initial data segment may comprise a MetaBox, and may comprise an indication of either URLs or an offset table configured to enable a client, decoder, file parser, or file reader to understand and / or request one or more tiles of a large image. The file reader behavior may be according to an example embodiment of the present disclosure. The file reader may download / receive the initial data segment providing metadata related to all tiles of the large image (e.g. URLs for individual tile access). Based on the initial data segment, the file reader may request one or more tiles of an image, for example based on user input, user movement, a current viewing direction / orientation or coordinates of the user, and / or viewport tracking. The file reader, or a UE comprising the file reader, may have enough memory to store the initial data segment as well as some number of tiles, and may be capable of processing the initial data segment and extracting URLs to be used in constructing HTTP requests for tiles, which may be processed for output.
[0138] Optionally, intermediate caching servers) may cache all tiles or a limited set of tiles for immediate access by the file reader. Accordingly, the reader may transmit requests for tiles to the intermediate caching server(s) rather than the server.
[0139] In an example embodiment, an offsettable, for example TiledlmageOffsetTable, may be indicated in a tiled image item with 4cc ‘tili’.
[0140] In an example embodiment, when a tiled image item with 4cc ‘tili’, having a contiguous range of addressing space to support byte range addressing is used, the ItemLocationBox associated with the tiled image item may contain only one item extent, wherein the item extent may include both the offsettable needed for byte range addressing of individual or a set of tiles, and all the item data of the image item.
[0141] In an alternative example embodiment, when a tiled image item with 4cc ‘tili’, having a contiguous range of addressing space to support byte range addressing is used, the ItemLocationBox associated with the tiled image item may contain only one item extent, wherein the item extent may include only the offset table needed for byte range addressing of individual or a set of tiles, and may not include the item data of the image item.
[0142] In an alternative example embodiment, when a tiled image item with 4cc ‘tili’, having a contiguous range of addressing space to support byte range addressing is used, the ItemLocationBox associated with the tiled image item contains only one item extent, wherein the item extent may not include the offsettable needed for byte range addressing of individual or a set of tiles, and may only include the item data of the image item.
[0143] In another example embodiment, when a tiled image item with 4cc ‘tili’, having a contiguous range of addressing space to support byte range addressing is used, the ItemLocationBox associated with the tiled image item may contain two or more item extents, which may be contiguous or non-contiguous. In an example embodiment, the first item extent may include / indicate the byte range of the offset table needed for byte range addressing of individual or a set of tiles, and the remaining item extent(s) may include the item data of the image item.
[0144] In an example embodiment, the offset table needed for byte range addressing of individual or a set of tiles of a tiled image item with 4cc ‘tili’ may be indicated within the DataEntryTiledltemURLBox of the DataReferenceBox.
[0145] In an example embodiment, the DataEntryTiledltemllRLBox of the DataReferenceBox may be extended to include the offset table needed for byte range addressing of an individual tile or a set of tiles of a tiled image item with 4cc ‘tili’.
[0146] Referring now to FIG. 16, illustrated is an example of the syntax of Tiled ImageOffsetTable according to an example embodiment of the present disclosure, ti le_start_offset[i] (1630) may point to the start of the coded data of the ith tile. The position may be given relative to the location identified for the corresponding item by the ItemLocationBox. ti le_size[i] (1640), if present, may indicate the number of bytes of the ith coded tile bitstream. If a specific tile is empty and does not contain image content, the tile is not coded and / or the tile_start_offset[i] (1630) entry and the tile_size[i] (1640) entry is set to 0. offset_field_index (1610) specifies the size of tile_start_offset[i] (1630) in bits., with value 0 indicating size is of 32 bits up to the value 3 indicating the size to be 64 bits. size_field_index (1620) specifies the size of tile_size[i] (1640) in bits, with value 0 indicating size is of 0 bits up to the value 3 indicating the size to be 64 bits. The parameter no_of_input_items (1650) may be the same as the value in DataEntryTiledltemllRLBox, an example of the syntax of which is illustrated in FIG.17, at 1710.
[0147] Referring now to FIG. 18, illustrated is an example of the syntax of DataEntryTiledltemllRLBox according to an example embodiment of the present disclosure, ti le_url_present_flag (1810) when set to 1 may indicate that the URLs of the tiles are present in the DataEntryTiledltemURLBox. tile_url_present_flag (1810) when set to 0 may indicate that TiledlmageOffsetTable (as defined in FIG. 16) (1820) is present in the DataEntryTiledltemURLBox.
[0148] In an example embodiment, alternate to tile_url_present_flag (1810), a new parameter may be used called tile_offset_table_present_flag, the semantics of which may be defined as follows. tile_offset_table_present_flag when set to 0 may indicate that the URLs of the tiles are present in theDataEntryTiledltemURLBox. tile_offset_table_present_flag when set to 1 may indicate that TiledlmageOffsetTable (as defined in FIG. 16) is present in the DataEntryTiledltemllRLBox.
[0149] In an example embodiment, when either the TiledlmageOffsetTable or the URLs of the tiles are present in the DataEntryTiledltemURLBox as defined above, the DataEntryTiledltemURLBox may be renamed to DataEntry Tiled ItemBox or DataEntryTiledltemAccessBox or to any other suitable name.
[0150] In an example embodiment, the offset table needed for byte range addressing of individual or a set of tiles of a tiled image item with 4cc ‘tili’ may be indicated within a new data entry box called DataEntry TiledltemOffsetTableBox of the DataReferenceBox.
[0151] Referring now to FIG. 19, illustrated is an example of the syntax of DataEntryTiledltemOffsetTableBox according to an example embodiment of the present disclosure, ti le_start_offset[i] (1910) may point to the start of the coded data of the ith tile. The position may be given relative to the location identified for the corresponding item by the ItemLocationBox. tile_size[i] (1920), if present, may indicate the number of bytes of the ith coded tile bitstream. If a specific tile is empty and does not contain image content, the tile is not coded and the ti le_start_offset[i] (1910) entry and / or ti le_size[i] (1920) entry may be set to 0.
[0152] In an alternate example embodiment, the syntax of DataEntry TiledltemOffsetTableBox may be defined as illustrated in FIG. 20, where the DataEntryTiledltemOffsetTableBox may contain offset (2010) and size (2020) of the tile offset table, ti le_offset_table_start_offset (2010) may specify the offset to the start of the tile offset table. The location may be relative to the referenced data in the itemlocation associated to the image item. Alternatively, the location may be relative to the reference indicated in the DataEntry TiledltemOffsetTableBox, wherein the DataEntry TiledltemOffsetTableBox may additionally include the following parameters:{unsigned int(4) reserved = 0;unsigned int(4) construction_method;}where construction_method may be taken from the set 0 (file); all other values may be reserved, ti le_offset_tab le_size (2020), if present, may indicate the size in number of bytes of the tile offset table.
[0153] In an example embodiment, when a tiled image item with 4cc ‘tili’, having a contiguous range of addressing space to support byte range addressing is used, and if the offset table needed for byte range addressing of individual or a set of tiles is present in one of the DataEntryBox of the DataReferenceBox, as defined above, then the item extent(s) in the ItemLocationBox associated with the tiled image item may notinclude the offset table needed for byte range addressing of individual or a set of tiles, and may only include all the item data of the image item.
[0154] In an example embodiment, when a tiled image item with 4cc ‘tili’ is used, and if the tiled image item contains external URLs of each tile, then the data_reference_index parameter in the ItemLocationBox associated with the tiled image item may point to the DataEntryBox of the DataReferenceBox containing only the URLs of each tile. Otherwise, if the tiled image item contains offsettable needed for byte range addressing of individual or a set of tiles, then the data_reference_index parameter in the ItemLocationBox associated with the tiled image item may point to the DataEntryBox of the DataReferenceBox containing only the tile image offset table.
[0155] In an example embodiment, when individual or a set of tiles of a tiled image item with 4cc ‘tili’ is mapped using the URLs using any of the Data Entry Boxes defined above within the DataReferenceBox, the URLs within those data entry boxes may be signalled using one or more of the following methods according to example embodiments of the present disclosure.
[0156] The following set of parameters may be present in any of the Data Entry Boxes defined above within the DataReferenceBox which contains URLs:{unsigned int(64) tilelDstart;utf8string baseurl;utf8string urlextension;utf8string tileitemrequesttemplate;}Where the baseurl contains the base URL for the tiles. The urlextension contains URL extensions which is used in URL construction as defined below in TABLE 5. The tileitemrequesttemplate contains the template which is used in URL construction as defined below. The tilelDstart indicates the tile ID of the first tile of the tile image item.TABLE 5
[0157] In the example above, assuming that the first tile is selected, the URL constructed results in http: / / cdn.example.com / movies / 134532 / image / Representation1 / 1000.heif by concatenating the baseurl and the urlextension and the 7’ character followed by the value of the tileitemrequesttemplate, where the tilelDstart isreplaced by the actual value of the tile starting from the value given in tilelDstart up to the value tilelDstart+no_of_input_items incrementing by one with the tiles starting from top left to right bottom.
[0158] In an example embodiment, URLs and / or the offsettable may be compressed or encoded. In an example embodiment, the server / writer may return either a valid HEIF file with the tiles mentioned above, or a single file with a single tile, or even just the media data of multiple tiles concatenated and decodable by a video decoder when parameter sets are embedded to the beginning of it (e.g. coming from the decoder configuration record of the HEIF file).
[0159] In an example embodiment, a file reader may indicate, to a file writer, the format according to which the initial data segment comprising the reduced URL structure should be generated and provided to the file reader. The format may be preconfigured to the file reader, or the UE comprising the file reader.
[0160] In an example embodiment, when a tiled image item with 4cc ‘tili’ is present it may be associated with PixellnformationProperty, an example of the syntax of which is illustrated in FIG. 21. The PixellnformationProperty item property may be extended with the version==0 and by using the box flags or by using a new version==1 with the following additional parameters.
[0161] channel_dimension_flag (2110), when set to 0, indicates that no channel size information is present. channel_dimension_flag (2110), when set to 1, indicates that channel size information is present and is equal to channel_size[i] (2120).
[0162] ch an ne l_size[i] (2120) indicates the size of the channel. Where the width of the channel is obtained as follows:
[0163] Channel_width[i] = ceil(lmageSpatialExtentsProperty.image_width) * channel_size[i])
[0164] Channel_height[i] = ceil(lmageSpatialExtentsProperty.image_height) * channel_size[i])
[0165] In an alternate example embodiment, the channel size parameter may be split into two parameters called channel_width_size[i] and the channel_height_size[i]. The channel_width_size[i] and the chan nel_height_size[i] may specify the absolute size of the channel in width and height dimensions, respectively:Channel_width[i] = channel_width_size[i]Channel_height[i] = channel_width_size[i]
[0166] Alternatively, channel_width_size[i] and channel_height_size[i] may specify the multiple of the size of the channel in width and height dimensions respectively.Channel_width[i] = ceil(lmageSpatialExtentsProperty.image_width) * channel_width_size[i]) Channel_height[i] = ceil(lmageSpatialExtentsProperty.image_height) * channel_height_size[i])
[0167] In an example embodiment, the Tiled ImageConfigurationBoxftilC’). may specify parameters associated with a tiled image item (‘tili’). These parameters may include the tile resolution, and / or the image itemtype (2210) used to code and store individual tile content. Additionally it may contain a TileltemPropertyAssociationBox (2220). Examples of syntax for the TiledlmageConfigurationBox and the TileltemPropertyAssociationBox are illustrated in FIG. 22.
[0168] Referring now to FIG. 23, illustrated is an example file according to an example embodiment of the present disclosure. The file may be, for example a HEIF compliant file (2310). The HEIF file (2310) may comprise a MetaBox (2320). The MetaBox (2320) may comprise a URL box (2330), or a data structure configured to include an offsettable (2340). The MetaBox (2320) may be comprised at the file level, movie level, or track level.
[0169] Referring now to FIG. 24, illustrated is an example file according to an example embodiment of the present disclosure. The file may be, for example a HEIF compliant file (2410). The HEIF file (2410) may comprise a MetaBox (2420). The MetaBox may be comprised at the file level, movie level, or track level. The MetaBox (2420) may comprise a URL box (2430), which may comprise URLs (2440) of external files, for example separate HEIF file(s) (2450), which may comprise metadata (2460) for the image.
[0170] A technical effect of example embodiments of the present disclosure may be to reduce signaling overhead with respect to large images.
[0171] FIG. 25 illustrates the potential steps of an example method 2500. The example method 2500 may include: obtaining an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image, 2510; and providing the initial data segment to at least one file reader, 2520. The example method 2500 may be performed, for example, with a server, a file writer, a UE, an encoder, a codec, a capture device, a satellite, a (image) caching server, etc.
[0172] In an example embodiment, the TilelmageConfigurationBoxftilC’) or the PixellnformationPropertyBox (‘pixi’) may additionally carry information about how the channels / dimensions of the tile image are stored in the file.
[0173] In an example embodiment, the tile image may be stored starting from the top left tile to the bottom right tile in a row-wise manner:- with each channel / dimension separately from each other one after another- or with all the channels / dimensions together (e.g. each pixel has RGBA channels)- or with only a subset of channels together followed by the remaining channels (e.g. first with RGB together and then the alpha plane separately)
[0174] In an example embodiment, a file reader / parser may download / retrieve tiles from all the channels / dimensions, by either using the corresponding URL or the tile offset table, and may concatenate the data from all the channels / dimensions before displaying the image.
[0175] In an example embodiment, the DataEntryTiledltemURLBox, for example as in FIG. 17, and / or FIG. 18, need not contain the parameter no_of_input_items. The parameter no_of_input_items, for example as in FIG. 17 and / or FIG. 18, and NumTiles in DataEntryTiledltemOffsetTableBox of for example FIG. 19 may be calculated as below:no_of_input_items =ceil(lmageSpatialExtentsProperty.image_width / TiledlmageConfigurationBox.tile_width)*ceil(lmageSpatialExtentsProperty.image_height / TiledlmageConfigurationBox.tile_height) if the channels / dimensions of the image are stored separately from each other then:no_of_input_items = no_of_input_items(calculated above) * no_of_channelsif only a subset of the channels / dimensions of the image are stored together and the remaining channels separately then:no_of_input_items = no_of_input_items(calculated above) * no_of_additional_dimensionswhere no_of_additional_dimensions may indicate channels / dimensions stored separately.
[0176] In an example embodiment, when creating a file with a tiled image item with 4cc ‘tili’, having support for either a contiguous range of addressing space to support byte range addressing or URLs for individual tile access, the following ordering of boxes may be suggested:- The file-level MetaBox may precede the MediaDataBox(es), if any (i.e., if the tiled image item(s) are contained in the MediaDataBox). When tiled image items(s) are contained in the ItemDataBox, the ItemDataBox may be arranged to be the last box within the containing MetaBox.
[0177] In an example embodiment, when creating a file with a tiled image item with 4cc ‘tili’, having support for either a contiguous range of addressing space to support byte range addressing or URLs for individual tile access the item data for tiled image item(s) may be suggested to have an order from top left tile to bottom right tile with row-wise arrangement which the player can then retrieve, decode and display.
[0178] In an example embodiment, when creating a file with a tiled image item with 4cc ‘tili’, having support for either a contiguous range of addressing space to support byte range addressing or URLs for individual tile access, a new brand may be defined which may specify a set of constraints for files with a tiled image item with 4cc ‘tili’. In this case the specified brand may be included in the FileTypeBox.
[0179] In an example embodiment, a new the brand called tiled image specific brand or an overview image specific brand with 4cc 'tibr' (any other suitable name and 4cc may be used) may be specified as follows. A coded image item may be specified to conform to the 'tibr' brand when all the following constraints are true:— The item is an overview image item— If the overview image item is a pre-derived coded image item or an image item that are tiled using a feature of a specific codec it is associated with ConstrainedExtentsGridProperty item property — The image item is part of the ImagePyramidEntityGroup
[0180] In an example embodiment, a tiled image item may be specified to conform to the 'tibr' brand when all the following constraints are true:— The item has type 'tili'— The item is associated with TilelmageConfigurationBox item property
[0181] In an example embodiment, files may include 'tibr' among the compatible brands. The files conforming to the 'tibr' brand may additionally be constrained as follows: Each file including 'tibr' as a compatible brand may contain a base item that is present in the file, may be either the primary item or any item from the alternate group containing the base item or any overview item from the ImagePyramidEntityGroup in which the base item is part of or the lowest resolution overview item or any item from the alternate group containing the lowest resolution overview item from the ImagePyramidEntityGroup in which the base item is part of or the highest resolution overview item or any item from the alternate group containing the highest resolution overview item from the ImagePyramidEntityGroup of which the base item is part.
[0182] In an example embodiment, readers conforming to the 'tibr' brand may support displaying a base item that is either the primary item or any item from the alternate group containing the base item or any overview item from the ImagePyramidEntityGroup in which the base item is part of or the lowest resolution overview item or any item from the alternate group containing the lowest resolution overview item from the ImagePyramidEntityGroup in which the base item is part of or the highest resolution overview item or any item from the alternate group containing the highest resolution overview item from the ImagePyramidEntityGroup of which the base item is part.
[0183] FIG. 26 illustrates the potential steps of an example method 2600. The example method 2600 may include: receiving an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image, 2610; determining to request at least one of the one or more tiles based, at least partially, on the at least one data structure, 2620; and transmitting, to at least one server, a request for the at least one tile, 2630. The example method 2600 may be performed, for example, with a client, a file reader, a file parser, a UE, a decoder, a codec, a viewer device, an application, etc.
[0184] According to a first aspect, there is provided the subject matter of: obtain an initial data segment, wherein the initial data segment may comprise, at least, at least one data structure, wherein the at least one data structure may comprise information for processing one or more tiles of an image; and provide the initial data segment to at least one file reader. The first aspect may be implemented in algorithm(s), encoded in a distribution medium or computer program product(s), and as method(s) performed by apparatus(es). An apparatus may comprise means for causing the apparatus at least to perform the method(s). The means may comprise at least one processor; and at least one memory storing algorithm(s) as instructions that, when executed by the at least one processor, cause the apparatus at least to perform the method(s).
[0185] The first aspect may include an single feature or combination of features from:Obtaining the initial data segment may comprise: generate the initial data segment based, at least partially, on one or more format preferences for the initial data segment received from the at least one file reader. Receive, from at least one file reader, a request for at least one tile of the image; and transmit, to the at least one file reader, the at least one requested tile. The initial data segment may comprise a meta box. The at least one data structure may comprise one or more uniform resource locators respectively associated with the one or more tiles. The at least one data structure may comprise an offsettable. The offsettable may comprise: one or more offset pointers associated with respective starting points of the one or more tiles, and size information associated with the one or more tiles. The at least one data structure may comprise a data entry tiled item uniform resource locator box. A data reference box may comprise the data entry tiled item uniform resource locator box. The data reference box may further comprise a size of the at least one data structure. The data entry tiled item uniform resource locator box may comprise one or more locations of the one or more tiles. The initial data segment may further comprise at least one of: item data associated with the image, or an item location box. The one or more tiles may be respectively associated with a contiguous range of addressing space configured for supporting byte range addressing. The initial data segment may comprise a flag, wherein the flag may be configured to indicate that the at least one data structure may comprise one of: an offsettable, or one or more locations associated with the one or more tiles. The initial data segment may comprise an indication of a location of an external file comprising the image. The external file may comprise a high efficiency image file format compliant file.
[0186] According to a second aspect, there is provided the subject matter of: receive an initial data segment, wherein the initial data segment may comprise, at least, at least one data structure, wherein the at least one data structure may comprise information for processing one or more tiles of an image; determine to request at least one of the one or more tiles based, at least partially, on the at least one data structure; andtransmit, to at least one server, a request for the at least one tile. The second aspect may be implemented in algorithm(s), encoded in a distribution medium or computer program product(s), and as method(s) performed by apparatus(es). An apparatus may comprise means for causing the apparatus at least to perform the method(s). The means may comprise at least one processor; and at least one memory storing algorithm(s) as instructions that, when executed by the at least one processor, cause the apparatus at least to perform the method(s).
[0187] The second aspect may include any single feature or any combination of features from:Transmit, to the at least one server, one or more format preferences for the initial data segment. Receive the at least one requested tile. The initial data segment may comprise a meta box. The at least one data structure may comprise one or more uniform resource locators respectively associated with the one or more tiles. The at least one data structure may comprise an offsettable. The offset table may comprise: one or more offset pointers associated with respective starting points of the one or more tiles, and size information associated with the one or more tiles. The at least one data structure may comprise a data entry tiled item uniform resource locator box. A data reference box may comprise the data entry tiled item uniform resource locator box. The data reference box may further comprise a size of the at least one data structure. The data entry tiled item uniform resource locator box may comprise one or more locations of the one or more tiles. The initial data segment may further comprise at least one of: item data associated with the image, or an item location box. The one or more tiles may be respectively associated with a contiguous range of addressing space configured for supporting byte range addressing. The initial data segment may comprise a flag, wherein the flag may be configured to indicate that the at least one data structure may comprise one of: an offset table, or one or more locations associated with the one or more tiles. The initial data segment may comprise an indication of a location of an external file comprising the image. The external file may comprise a high efficiency image file format compliant file. Determining to request the at least one of the one or more tiles may comprise: obtain location information of a user, wherein the location information may comprise at least one of: an orientation of the user, a position of the user, coordinates associated with the user, a viewport associated with the user, user movement, or user input; and determine the at least one tile to be requested based, at least partially, on the obtained location information. Generate the request for the at least one tile based, at least partially, on: a base uniform resource locator, a uniform resource locator extension, a request template, and an identifier of a first tile of the at least one tile.
[0188] As used in this application, the term “circuitry” or “means” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in analog, digital and / or quantum circuitry) and (b) combinations of hardware circuit(s) and software, such as (as applicable): (i) a combination ofanalog, digital and / or quantum hardware circuit(s) with software / firmware and (ii) any or all portions of hardware processor(s) (including digital and / or quantum processor(s)) with software, and memory(ies) that work together to cause an apparatus, such as a mobile device, computing device, or server, to perform various functions) and (c) any or all portions of hardware circuit(s), such as a microprocessor(s), processor(s) and / or quantum processor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0189] A processor, memory, and / or example algorithms (which may be encoded as instructions, program, or code) may be provided as example means for providing or causing performance of operation.
[0190] The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e. tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
[0191] As used herein, the terms “the at least one” and “the one or more” mean “any one of the at least one” and “any one of the one or more”, respectively.
[0192] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications can be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modification and variances which fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:
1. An apparatus comprising:at least one processor; andat least one memory storing instructions that, when executed with the at least one processor, cause the apparatus at least to:obtain an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image; andprovide the initial data segment to at least one file reader.
2. The apparatus of claim 1, wherein the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:receive, from at least one file reader, a request for at least one tile of the image; and transmit, to the at least one file reader, the at least one requested tile.
3. The apparatus of any of claims 1 or 2, wherein the initial data segment further comprises at least one of:item data associated with the image,an item location box, ora meta box.
4. The apparatus of any of claims 1 to 3, wherein the at least one data structure comprises one or more uniform resource locators respectively associated with the one or more tiles.
5. The apparatus of any of claims 1 to 3, wherein the at least one data structure comprises an offset table, wherein the offset table comprises:one or more offset pointers associated with respective starting points of the one or more tiles, andsize information associated with the one or more tiles.
6. The apparatus of any of claims 1 to 5, wherein the at least one data structure comprises a data entry tiled item uniform resource locator box, wherein the data entry tiled item uniform resource locator box comprises one or more locations of the one or more tiles.
7. The apparatus of claim 6, wherein a data reference box comprises the data entry tiled item uniform resource locator box, wherein the data reference box further comprises a size of the at least one data structure.
8. The apparatus of any of claims 1 to 7, wherein the one or more tiles are respectively associated with a contiguous range of addressing space configured for supporting byte range addressing.
9. The apparatus of any of claims 1 to 8, wherein the initial data segment comprises a flag, wherein the flag is configured to indicate that the at least one data structure comprises one of:an offset table, orone or more locations associated with the one or more tiles.
10. The apparatus of any of claims 1 to 9, wherein the initial data segment comprises an indication of a location of an external file comprising the image, wherein the external file comprises a high efficiency image file format compliant file.
11. A method comprising:obtaining an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image; andproviding, with a server, the initial data segment to at least one file reader.
12. An apparatus comprising means for:obtaining an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image; andproviding the initial data segment to at least one file reader.3613. A computer-readable medium comprising instructions stored thereon for performing at least the following:causing obtaining of an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image; andcausing providing of the initial data segment to at least one file reader.
14. An apparatus comprising:at least one processor; andat least one memory storing instructions that, when executed with the at least one processor, cause the apparatus at least to:receive an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image;determine to request at least one of the one or more tiles based, at least partially, on the at least one data structure; andtransmit, to at least one server, a request for the at least one tile.
15. The apparatus of claim 14, wherein the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:transmit, to the at least one server, one or more format preferences for the initial data segment.
16. The apparatus of claims 14 or 15, wherein the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:receive the at least one requested tile.
17. The apparatus of any of claims 14 to 16, wherein the initial data segment further comprises at least one of:item data associated with the image,an item location box, ora meta box.
18. The apparatus of any of claims 14 to 17, wherein the at least one data structure comprises one or more uniform resource locators respectively associated with the one or more tiles.
19. The apparatus of any of claims 14 to 17, wherein the at least one data structure comprises an offset table, wherein the offset table comprises:one or more offset pointers associated with respective starting points of the one or more tiles, andsize information associated with the one or more tiles.
20. The apparatus of any of claims 14 to 19, wherein the at least one data structure comprises a data entry tiled item uniform resource locator box, wherein the data entry tiled item uniform resource locator box comprises one or more locations of the one or more tiles.
21. The apparatus of claim 20, wherein a data reference box comprises the data entry tiled item uniform resource locator box, wherein the data reference box further comprises a size of the at least one data structure.
22. The apparatus of any of claims 14 to 21 , wherein the one or more tiles are respectively associated with a contiguous range of addressing space configured for supporting byte range addressing.
23. The apparatus of any of claims 14 to 22, wherein the initial data segment comprises a flag, wherein the flag is configured to indicate that the at least one data structure comprises one of:an offset table, orone or more locations associated with the one or more tiles.
24. The apparatus of any of claims 14 to 23, wherein the initial data segment comprises an indication of a location of an external file comprising the image, wherein the external file comprises a high efficiency image file format compliant file.
25. The apparatus of any of claims 14 to 24, wherein determining to request the at least one of the one or more tiles comprises the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:obtain location information of a user, wherein the location information comprises at least one of:an orientation of the user,a position of the user,coordinates associated with the user,a viewport associated with the user,user movement, oruser input; anddetermine the at least one tile to be requested based, at least partially, on the obtained location information.
26. The apparatus of any of claims 14 to 25, wherein the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:generate the request for the at least one tile based, at least partially, on:a base uniform resource locator,a uniform resource locator extension,a request template, andan identifier of a first tile of the at least one tile.
27. A method comprising:receiving, with a file reader, an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image;determining to request at least one of the one or more tiles based, at least partially, on the at least one data structure; andtransmitting, to at least one server, a request for the at least one tile.
28. An apparatus comprising means for:39receiving an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image;determining to request at least one of the one or more tiles based, at least partially, on the at least one data structure; andtransmitting, to at least one server, a request for the at least one tile.
29. A computer-readable medium comprising instructions stored thereon for performing at least the following:causing receiving of an initial data segment, wherein the initial data segment comprises, at least, at least one data structure, wherein the at least one data structure comprises information for processing one or more tiles of an image;determining to request at least one of the one or more tiles based, at least partially, on the at least one data structure; andcausing transmitting, to at least one server, of a request for the at least one tile.40