Signaling a large number of images with grids in high efficiency image file format
By employing 64-bit representations and IMDA mapping, the apparatus efficiently handles and navigates large images in HEIF format, addressing the challenges of high-resolution image access and interaction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing technologies face challenges in efficiently handling and accessing large images, particularly those with resolutions exceeding 300,000 pixels by 300,000 pixels, especially when navigating and zooming interactions are involved, and often rely on cloud optimization techniques like geoTIFF which may not fully optimize performance.
The implementation of an apparatus and method that supports 64-bit representations, configures ItemID parameters beyond 32-bit integer ranges, and uses IdentifiedMediaDataBox (IMDA) mapping to order image boxes, allowing for grid-derived image items and new item properties to enhance image retrieval and rendering, particularly in HEIF format.
This approach enables efficient display and navigation of large images by supporting higher resolution representations and interactions, improving performance and functionality in handling extremely large images.
Smart Images

Figure EP2025081107_07052026_PF_FP_ABST
Abstract
Description
SIGNALING A LARGE NUMBER OF IMAGES WITH GRIDS IN HIGH EFFICIENCYIMAGE FILE FORMATTECHNICAL FIELD
[0001] Various example embodiments relate generally to signaling images, such as large quantities of images, with grids of the high efficiency image file (HEIF) format.BACKGROUND
[0002] In some examples, such as in geospatial applications, image resolution may exceed 300,000 pixels by 300,000 pixels, and image sizes continue to grow. In at least some technologies, images are accessed over a network using "cloud optimization" techniques. Image storage techniques such as tiles, grids, image pyramids and image overviews may be used for simplified access to at least portions of and / or lower resolution versions of an image. Tiles may be set based on end device capabilities. For example, to accommodate some handheld computer devices (e.g., mobile phones, smartphones, tablets, wearable devices, laptop computers, etc.) and / or non-mobile computer devices (desktop computers, etc.), tile resolutions such as 512 pixels by 512 pixels and / or 1000 pixels by 1000 pixels may be used. As a user navigates the large image space, individual tiles may be pulled, for example, depending on pan and / or zoom commands from the user (e.g., via a browser interface). Hypertext Transfer Protocol (HTTP) byte range requests may be used to achieve efficiency and / or to enable functionality using some built-in browser and web server capabilities. In at least some technologies, cloud optimized geospatial Tagged Image File Format (geoTIFF) may be used.BRIEF DESCRIPTION
[0003] According to an aspect of the invention, there is provided an apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: (a) display an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; (b) configure one or more parameters of the image item to support 64-bit representations;(c) configure at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and / or (d) order one or more boxes of the image item by rewriting theat least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping. In some examples, the features of the client comprise at least one of: (a) a current viewing position; (b) a pan interaction; (c) a zoom interaction; and / or (d) a different interaction. In some examples, the instructions, when executed by the at least one processor, further cause the apparatus at least to perform storing one or more images as grid derived image items supported by at least one codec. In some examples, the instructions, when executed by the at least one processor, further cause the apparatus at least to perform grouping one or more images to form at least one of an overview or a reduced-resolution subfile representing a reduced-resolution version of a higher- resolution image item. In some examples, the specifying further comprises defining a new item property for the grid derived image item. In some examples, the instructions, when executed by the at least one processor, further cause the apparatus at least to perform defining a new version of MiniBox, wherein the new version is configured to carry at least a subset of items or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0004] According to an aspect of the invention, there is provided a method, comprising: (a) displaying an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; (b) configuring one or more parameters of the image item to support 64-bit representations; (c) configuring at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and / or (d) ordering one or more boxes of the image item by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping. In some examples, the features of the client comprise at least one of: (a) a current viewing position; (b) a pan interaction; (c) a zoom interaction; and / or a different interaction. In some examples, the method further comprises storing one or more images as grid derived image items supported by at least one codec. In some examples, the method further comprises grouping one or more images to form at least one of an overview or a reduced-resolution subfile representing a reduced-resolution version of a higher- resolution image item. In some examples, the method further comprises specifying a new brand for a grid derived image item. In some examples, the specifying further comprises defining a new item property for the grid derived image item. In some examples, the method further comprises defining a new version of MiniBox, wherein the new version is configured to carry at least a subset of items or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0005] According to an aspect of the invention, there is provided a non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to: (a) display an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; (b) configure one or more parameters of the image item to support 64-bit representations; (c) configure at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and / or (d) order one or more boxes of the image item by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping. In some examples, the features of the client comprise at least one of: (a) a current viewing position; (b) a pan interaction; (d) a zoom interaction; and / or (d) a different interaction. In some examples, the program instructions, when executed by the apparatus, further cause the apparatus at least to perform storing one or more images as grid derived image items supported by at least one codec. In some examples, the program instructions, when executed by the apparatus, further cause the apparatus at least to perform grouping one or more images to form at least one of an overview or a reduced-resolution subfile representing a reduced-resolution version of a higher-resolution image item. In some examples, the program instructions, when executed by the apparatus, further cause the apparatus at least to perform specifying a new brand for a grid derived image item. In some examples, the specifying further comprises defining a new item property for the grid derived image item. In some examples, the program instructions, when executed by the apparatus, further cause the apparatus at least to perform defining a new version of MiniBox, wherein the new version is configured to carry at least a subset of items or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0006] According to an aspect of the invention, there is provided an apparatus, comprising: (a) means for displaying an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; (b) means for configuring one or more parameters of the image item to support 64-bit representations; (c) means for configuring at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and / or (d) means for ordering one or more boxes of the image item by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping. In some examples, the features of the client comprise at least one of: (a) a current viewing position; (b) a pan interaction; (c) a zoom interaction; and / or (d) a different interaction.In some examples, the apparatus further comprises means for storing one or more images as grid derived image items supported by at least one codec. In some examples, the apparatus further comprises means for grouping one or more images to form at least one of an overview or a reduced-resolution subfile representing a reduced-resolution version of a higher-resolution image item. In some examples, the apparatus further comprises means for specifying a new brand for a grid derived image item. In some examples, the specifying further comprises defining a new item property for the grid derived image item. In some examples, the apparatus further comprises means for defining a new version of MiniBox, wherein the new version is configured to carry at least a subset of items or item properties used for large-scale tiled hierarchical image retrieval rendering.LIST OF THE DRAWINGS
[0007] In the following, the invention will be described in greater detail with reference to the embodiments and the accompanying drawings, in which:
[0008] Fig. 1 shows an example of an apparatus which may implement one or more examples disclosed herein;
[0009] Fig. 2 shows an example of a communication network to which one or more examples disclosed herein may be applied;
[0010] Fig. 3 shows a network diagram in which example apparatuses may be implemented in accordance with some embodiments disclosed herein;
[0011] Fig. 4 shows exemplary syntax;
[0012] Fig. 5 shows exemplary syntax;
[0013] Fig. 6 shows exemplary syntax;
[0014] Fig. 7 shows exemplary syntax;
[0015] Fig. 8 shows exemplary syntax;
[0016] Fig. 9 shows exemplary syntax;
[0017] Fig. 10 shows exemplary syntax;
[0018] Figs. 11A, 11B, 11C, 11D, HE, 11F, 11G, 11H, 111, 11 J, and 1 IK show exemplary syntax;
[0019] Fig. 12 shows an exemplary architecture;
[0020] Fig. 13 shows exemplary syntax;
[0021] Figs. 14A, 14B, and 14C show exemplary syntax;
[0022] Figs. 15 A, 15B, and 15C show exemplary syntax;
[0023] Fig. 16 shows exemplary syntax;
[0024] Fig. 17 shows exemplary syntax;
[0025] Fig. 18 shows exemplary syntax;
[0026] Fig. 19 shows exemplary syntax;
[0027] Fig. 20 shows exemplary syntax;
[0028] Figs. 21 A and 21B show exemplary syntax;
[0029] Figs. 22A and 22B show exemplary syntax;
[0030] Fig. 23 shows exemplary syntax;
[0031] Fig. 24 shows exemplary syntax;
[0032] Fig. 25 shows exemplary syntax;
[0033] Fig. 26 shows exemplary syntax;
[0034] Figs. 27A and 27B show exemplary storage schemes;
[0035] Fig. 28 shows exemplary syntax;
[0036] Fig. 29 shows exemplary syntax;
[0037] Figs. 30A, 30B, and 30C show exemplary syntax;
[0038] Fig. 31 shows exemplary syntax;
[0039] Fig. 32 shows an exemplary architecture;
[0040] Fig. 33 shows an exemplary architecture;
[0041] Fig. 34 shows an exemplary architecture;
[0042] Figs. 35A and 35B show exemplary syntax;
[0043] Fig. 36 shows exemplary syntax; and
[0044] Fig. 37 shows an example of a method.DESCRIPTION OF EMBODIMENTS
[0045] The following embodiments are exemplary. Although the specification may refer to “an”, “one”, or “some” embodiment(s) in several locations of the text, this does not necessarily mean that each reference is made to the same embodiment(s), or that a particular feature onlyapplies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments. Further, when a particular feature, structure, or characteristic is described in connection of an embodiment, it is within the knowledge of one skilled in the art to apply such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described. It shall be understood that although the terms “first”, “second”, and / or the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0046] For the purposes of the present disclosure, the phrases “at least one of A or B”, “at least one of A and B”, and “A and / or B” mean (A), (B), or (A and B). For the purposes of the present disclosure, the phrase “A, B, and / or C” means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
[0047] Embodiments described herein may be implemented by at least one apparatus. In some embodiments, the apparatus comprises one or more servers configured to communicate with one or more terminal devices via a network. For example, as shown in Fig. 3 and further described herein, an apparatus may comprise or be embodied as an application server 301 configured to communicate with a terminal device 303 via a mobile telephony network 202, Internet 204, and / or the like. Additionally, or alternatively, in some embodiments, the apparatus comprises a terminal device. For example, as shown in Fig. 3 and described herein, an apparatus may comprise or be embodied as a client application 305 that is installed or accessed by a terminal device 303. It will be understood and appreciated that the functionality and operations described herein may be implemented by different apparatuses (e.g., application servers, terminal devices, and / or the like) without departing from the scope and spirit of the disclosure.
[0048] Embodiments described herein may be implemented by at least one apparatus. Referring now to Fig. 1, an example apparatus 100 configured to implement one or more examples described herein is provided. The apparatus 100 may be an electronic device. The apparatus 100 may be configured to perform various functions, for example, such as gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus, and / or the like. In some examples, an apparatus configured to encode a video scene (e.g., such as the apparatus 100) may comprise (e.g., optionally) one or more microphones for capturing the scene and / or one or more sensors, such as cameras, for capturing information about the physicalenvironment in which the scene is captured. Additionally or alternatively, in some examples, an apparatus configured to encode a video scene (e.g., such as the apparatus 100) may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. Additionally or alternatively, in some examples, an apparatus configured to decode and / or render a video scene (e.g., such as the apparatus 100) may be configured to receive a Moving Picture Experts Group immersive codec family (MPEG-I) bitstream comprising an encoded video scene. Additionally or alternatively, in some examples, an apparatus configured to decode and / or render a video scene (e.g., such as the apparatus 100) may comprise one or more speakers, audio transducers, and / or displays and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers, audio transducers, and / or displays. Additionally or alternatively, in some examples, an apparatus configured to decode and / or render a video scene (e.g., such as the apparatus 100) may comprise a user equipment (UE), a head and / or mounted display, and / or a device capable of rendering to a user an augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) experience.
[0049] The apparatus 100 may, for example, be a mobile terminal and / or UE of a wireless communication system. Additionally or alternatively, the apparatus 100 may be a computer and / or a part of a computer which is not mobile. It should be appreciated that embodiments of the present disclosure may be implemented within any electronic device or apparatus which may process data. The apparats 100 may comprise a device that can access a network and / or cloud through a wired and / or wireless connection.
[0050] The apparatus 100 may comprise a controller 102. which may comprise one or more processors and / or processing circuitry for controlling the apparatus 100. The controller 102 may be connected to a memory 104 which may be configured to store data such as: image data and / or audio data, and / or instructions for implementation on the controller 102. The controller 102 may be coupled (e.g., connected) to codec circuitry 106. The codec circuitry 106 may be configured to code, encode, and / or decode audio and / or video data. Additionally or alternatively, the codec circuitry may be configured to assist in coding, encoding, and / or decoding performed by the controller 102.
[0051] The apparatus 100 may comprise one or more processors, one or more memories, and / or one or more transceivers which may be interconnected via one or more buses. The one or more processors may comprise a central processing unit (CPU) and / or a graphical processingunit (GPU). At least one of the one or more transceivers (e.g., a respective one of the one or more transceivers) may include a receiver and / or a transmitter. The one or more buses may be address, data, and / or control buses. The one or more buses may include any interconnection mechanism, for example, such as a series of lines on a motherboard and / or integrated circuit, fiber optics, other optical communication equipment, and / or the like. The one or more transceivers may be connected to one or more antennas. The one or more memories may include program code. The one or more memories and the program code may be configured to, with the one or more processors, cause the apparatus 100 to perform one or more operations as described herein.
[0052] The apparatus 100 may comprise a card reader 110 and / or a smart card 108, for example, such as a universal integrated circuit card (UICC) and / or a UICC reader. The UICC and / or UICC reader may be configured to provide user information and / or provide authentication information for authentication and / or authorization of the apparatus 100 at a network.
[0053] The apparatus 100 may couple (e.g., connect) to a node of a network. The network node may comprise one or more processors, one or more memories, and / or one or more transceivers which may be interconnected via one or more buses. At least one of the one or more transceivers (e.g., a respective one of the one or more transceivers) may include a receiver and / or a transmitter. The one or more buses may be address, data, and / or control buses. The one or more buses may include any interconnection mechanism, for example, such as a series of lines on a motherboard and / or integrated circuit, fiber optics, other optical communication equipment, and / or the like. The one or more transceivers may be connected to one or more antennas. The one or more memories may include program code. The one or more memories and the program code may be configured to, with the one or more processors, cause the network node to perform one or more operations as described herein.
[0054] The apparatus 100 may comprise an input device 112, for example, such as a keypad, one or more input buttons, a touch screen input device, and / or the like configured to provide information to the controller 102. The apparatus 100 may comprise radio interface circuitry 114 connected to the controller 102. The radio interface circuitry 114 may be configured to generate wireless communication signals, for example, for communication with a cellular communications network, a wireless communications system, a wireless local area network (WLAN), and / or the like. The apparatus 100 may comprise one or more antennae 116 connectedto the radio interface circuitry 114. The one or more antennae 116 may be configured to transmit radiofrequency (RF) signals generated at the radio interface circuitry 114 to one or more other apparatuses and / or configured to receive RF signals from one or more other apparatuses.
[0055] The apparatus 100 may comprise a microphone 118, an audio output device 120, a camera 122, and / or other sensors configured to record and / or detect audio signals, image signals, video signals, and / or other information about a local and / or virtual environment. The information about the local and / or virtual environment recorded and / or detected by the microphone 118, the audio output device 120, the camera 122, and / or other sensors may transmit (e.g., pass) such information to the codec circuitry 106 and / or the controller 102, for example, for processing. The apparatus 100 may receive such information for processing from one or more other devices for processing from one or more other devices, for example, prior to transmission and / or storage. The apparatus 100 may receive, via wired and / or wireless connection, the information. One or more structural elements of the apparatus 100 described herein may represent examples of means for performing a function, for example, such as a corresponding function.
[0056] The memory 104 may be of any type suitable to a local technical environment and / or may be implemented using any suitable data storage technology, for example, such as semiconductor-based memory devices, flash memory, magnetic memory devices and / or systems, optical memory devices and / or systems, fixed memory, removable memory, and / or other types of memory or systems. The memory 104 may be a non-transitory memory. The controller 102 may be or comprise one or more processors, which may be of any type suitable to the local technical environment, and / or may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), processors based on multi-core processor architectures, and / or other types of processors, to provide non-limiting examples, The controller 102 may be a means for performing one or more functions.
[0057] The apparatus 100 may comprise a microphone and / or other audio input, which may be a digital and / or analog signal input. The apparatus 100 may comprise an audio output device which, in one or more embodiments of the present disclosure, may be any one of: an earpiece, a speaker, and / or an analog audio and / or digital audio output connection. The apparatus 100 may comprise one or more batteries. In some examples, the apparatus 100 may be powered by one or more mobile energy devices, for example, such as solar cells, fuel cells, clockwork generators, and / or the like. The apparatus 100 may comprise a camera and / or other sensor capable ofrecording and / or capturing images and / or video. Additionally, or alternatively, the apparatus 100 may comprise a depth sensor. The apparatus 100 may comprise a display 124. The apparatus 100 may comprise an infrared port for short range line-of-sight communication to other devices. In some embodiments, the apparatus 100 may comprise other types of short-range communication technologies, for example, such as Bluetooth, wireless connections, universal serial bus (USB) connections, firewire connections, wired connections, and / or other types of connections.
[0058] It should be understood that an apparatus, such as the apparatus 100, configured to perform one or more example embodiments of the present disclosure may have fewer and / or additional components, which may correspond to one or more processes the apparatus is configured to perform. For example, an apparatus configured to encode a video ay not comprise a speaker or audio transducer and may comprise a microphone, while an apparatus configured to render a decoded vide may not comprise a microphone and may comprise a speaker or audio transducer.
[0059] An apparatus, such as the apparatus 100, may be configured to perform capture of a volumetric scene according to example embodiments of the present disclosure. For example, the apparatus 100 may comprise the camera 122 and / or other sensors capable or recording and / or capturing images and / or video. The apparatus may comprise one or more transceivers configured to enable transmission of captured content for processing at another device. The apparatus may comprise one or more transceivers configured to enable reception of captured content for processing at the apparatus. Such an apparatus may or may not include all elements shown in the example of Fig. 1.
[0060] An apparatus, such as the apparatus 100, may be configured to perform processing of volumetric video content according to one or more example embodiments of the present disclosure. For example, the apparatus may comprise: a controller (e.g., the controller 102) for processing images to produce volumetric video content; a controller (e.g., the controller 102) for processing volumetric video content to project three-dimensional (3D) information into two- dimensional (2D) information, patches, and / or auxiliary information; a codec (e.g., the codec circuitry 106) for encoding 2D information, patches, and / or auxiliary information into a bitstream for transmission to another device via a radio interface (e.g., the radio interface 114); and / or other elements. Such an apparatus may or may not include all elements shown in the example of Fig. 1.
[0061] An apparatus, such as the apparatus 100, may be configured to perform encoding and / or decoding of 2D information representative of volumetric video content according to one or more example embodiments of the present disclosure. For example, the apparatus may comprise a codec (e.g., the codec circuitry 106) for encoding and / or decoding 2D information representative of volumetric video content. Such an apparatus may or may not include all elements shown in the example of Fig. 1.
[0062] An apparatus, such as the apparatus 100, may be configured to perform rendering of decoded 3D volumetric video according to one or more example embodiments of the present disclosure. For example, the apparatus may comprise a controller (e.g., the controller 102) for projecting 2D information to reconstruct 3D volumetric video and / or a display (e.g., the display 124) for rendering decoded 3D volumetric video. Such an apparatus may or may not include all elements shown in the example of Fig. 1.
[0063] Embodiments described may be implemented in a communication network, such as any of the following radio access technologies (RATs): Worldwide Interoperability for Microwave Access (WiMAX), Global System for Mobile communications (GSM, 2G), GSM EDGE radio access Network (GERAN), General Packet Radio Service (GRPS), Universal Mobile Telecommunication System (UMTS, 3G) based on basic wideband-code division multiple access (W-CDMA), high-speed packet access (HSPA), Long Term Evolution (LTE), LTE-Advanced, and enhanced LTE (eLTE), 5G (also called NR), or any future RAT such as 6G. Moreover, communications within the communication network may utilize any proper wireless communication technology, comprising but not limited to: Code Division Multiple Access (CDMA), Frequency Division Multiple Access (FDMA), Time Division Multiple Access (TDMA), Frequency Division Duplex (FDD), Time Division Duplex (TDD), Multiple-Input Multiple-Output (MIMO), Orthogonal Frequency Division Multiple (OFDM), and / or Discrete Fourier Transform spread OFDM (DFT-s-OFDM).
[0064] As used herein, the terms “network device” and / or “network node” refer to a node in a communication network via which user equipment may access the network and / or which is capable of controlling radio communication and managing radio resources within a cell. The network node or network device may be referred to as a base station (BS), an access point (AP), or an access node. The network device may be, depending on the applied technology, for example, a node B (NodeB or NB), an evolved NodeB (eNodeB or eNB), an NR NB (alsoreferred to as a gNB), a Remote Radio Unit (RRU), a radio head (RH), a remote radio head (RRH), a relay, an Integrated Access and Backhaul (IAB) node, a low power node, a nonterrestrial network (NTN) or non-ground network device such as a satellite network device, a low earth orbit (LEO) satellite and a geosynchronous earth orbit (GEO) satellite, or an aircraft network device.
[0065] Referring now to Fig. 2, an example of a system 200, within which one or more embodiments of the present disclosure may be utilized and / or implemented, is shown. The system 200 comprises multiple communication devices which may communicate via one or more networks. The system 200 may comprise any combination of wired and / or wireless networks including, but not limited to: a wireless cellular network (e.g., GSM, UMTS, Evolved UMTS Terrestrial Radio Access (E-UTRA), Long Term Evolution (LTE), CDMA, 4G, 5G, 6G, etc.), a WLAN such as defined by any one or more standards (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.x standards), a short-range personal area network (e.g., a Bluetooth personal area network), an Ethernet local area network, a token ring local area network, a wide area network, the Internet, and / or other networks. The system 200 may include wired and / or wireless communication devices and / or electronic devices configured to implement one or more embodiments of the present disclosure.
[0066] The example of Fig. 2 shows an exemplary mobile telephone network 202 and a representation of the Internet 204. Connectivity to the internet 204 may include, but is not limited to, long range wireless connections, short range wireless connections, telephone lines, cable lines, power lines, and / or other wireless and / or wired connections.
[0067] Example communication devices included in the system 200 may include, but are not limited to, a first apparatus 206 (e.g., a mobile and / or non-mobile telephone), a second apparats 208 (e.g., a combination of a personal digital assistant (PDA) and a mobile telephone), a third device 210 (e.g., a PDA), a fourth device 212 (e.g., an integrated messaging device (IMD), a desktop computer 214, a notebook and / or laptop computer 216, a head-mounted display 218, and / or other devices. The apparatus 100 may comprise any such communication devices. In an example embodiment of the present disclosure, more than one of these devices, or a plurality of one or more of these devices, may perform one or more disclosed processes. Any one or more of the devices may connect to the Internet 204 via a wireless connection 220.
[0068] Embodiments of the present disclosure may be implemented in other types of devices, for example, such as: a set-top box (e.g., a digital television (TV) receiver), which may or may not have a display and / or wireless capabilities; tablet and / or laptop personal computers (PCs), which may have hardware and / or software to process neural network data; various operating systems; chipsets, processors, DSPs, and / or embedded systems offering hardware- and / or software-based coding. Embodiments of the present disclosure may be implemented in cellular telephones such as smart phones having wireless communication capabilities, tablets having wireless communication capabilities, PDAs having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and / or playback appliances having wireless communication capabilities, Internet-of-Things (loT) devices having wireless communication capabilities, Internet appliances permitting wireless Internet access and / or browsing, portable units and / or terminals incorporating one or more combinations of such functions, and / or other devices.
[0069] At least one of the example communication devices included in the system 200 may send and / or receive calls and / or messages, and / or communicate with service providers via a wireless connection 222 to a base station 224. The base station 224 may be an eNB, a gNB, and / or the like. The base station 224 may be connected to a network server 226 which may allow communication between the mobile telephone network 202 and the Internet 204. The system 200 may include additional communication devices and / or communication devices of various types.
[0070] The example communication devices included in the system 200 may communicate using various transmission technologies including, but not limited to, CDMA, GSM, UMTS, TDMA, frequency division multiple access (FDMA), transmission control protocol internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), electronic mail (e-mail), instant messaging service (IMS), rich communication service (RCS), Bluetooth, IEEE 802.11, 3rdGeneration Partnership Project (3GPP) Narrowband loT, and / or any other wireless communication technologies. A communications device involved in implementing various embodiments of the present disclosure may communicate using various media including, but not limited to, radio, infrared (IR), laser, cable connections, and / or any other suitable connection.
[0071] As used herein, the term “channel” may refer to a physical channel and / or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire. A logical channel may refer to a logical connection over a multiplexed medium, for example, capable of conveying an information signal (e.g., a bitstream such as an MPEG-I bitstream from one or more senders and / or transmitters to one or more receivers).
[0072] Referring now to Fig. 3, an example network diagram 300 is provided. In various embodiments, the processes and functionality described herein are implemented by an application server 301, a terminal device 303, and / or the like. For example, an application server 301 may comprise, or be embodied as, one or more apparatuses 100. As another example, a terminal device 303 may comprise one or more apparatuses 100’, which may be embodied as one or more client applications 305.
[0073] In some embodiments, the application server 301 embodies one or more computing environments comprising computing resources configured to communicate and provide data, services, and / or the like to clients. For example, the application server 301 may comprise a computing environment configured to provide image viewing services, image data, audio data and / or the like to one or more terminal devices 303. In some embodiments, the application server 301 is associated with a client application 305. For example, in a geospatial context, an application server 301 may include an image-based mapping service by which a client application 305 (e.g., a mapping application, navigational application, and / or the like) may access or view geospatial images. In various embodiments, the application server 301 includes non-transitory memory, volatile memory, and / or the like that is configured to store information, data, content, applications, instructions, or the like for enabling the apparatus 100 to carry out various functions in accordance with an example embodiment disclosed herein. For example, the application server 301 may include memory configured to store instructions for encoding or decoding hierarchical images, or segments thereof, in accordance with a file format data structure.
[0074] In some embodiments, the terminal device 303 is a user equipment (UE) configured to render media on a display, such as images, videos, and / or the like. For example, the terminal device 303 may be a mobile telephone (e.g., smartphone, and / or the like), tablet, phablet, personal digital assistant, wearable device, game console, Internet of Things (loT) device, infotainment system, streaming device, navigation device, and / or the like. In variousembodiments, the client application 305 is configured to encode or decode hierarchical images, or segments thereof, in accordance with a file format data structure described herein. In some embodiments, the terminal device 303 includes memory that is accessible to the client application 305. The memory of the terminal device 303 may include information, data, content, applications, instructions, or the like for enabling the apparatus 100’ to carry out various functions in accordance with an example embodiment disclosed herein.
[0075] In some embodiments, the application server 301 is configured to provision and receive data to and from the terminal device 303 via the client application 305 and the mobile telephony network 202 or Internet 204. For example, the application server 301 may provision image data to the client application 305, the image data including a plurality of encoded segments (e.g., gridded portions of the image, tiles that embody various resolutions of the image, and / or the like). Additionally, or alternatively, the application server 301 may receive such image data from the client application 305. In some embodiments, the application server 301 is configured to receive features from the terminal device 301 via the client application 305. For example, the terminal device 301 may include one on or more input devices configured to receive user input for defining features of image or other media interactions, such as a zooming interaction, panning interaction, a current viewing position, and / or the like.
[0076] In some examples, such as in geospatial applications, image resolution may exceed 300,000 pixels by 300,000 pixels, and image sizes continue to grow. In at least some technologies, images are accessed over a network using "cloud optimization" techniques. Image storage techniques such as tiles and / or grids, and / or image pyramids and / or image overviews may be used for simplified access to at least portions of and / or lower resolution versions of an image. Tiles may be set based on end device capabilities. For example, to accommodate some handheld computer devices (e.g., mobile phones, smartphones, tablets, wearable devices, laptop computers, etc.) and / or non-mobile computer devices (desktop computers, etc.), tile resolutions such as 512 pixels by 512 pixels and / or 1000 pixels by 1000 pixels may be used. As a user navigates the large image space, individual tiles may be pulled, for example, depending on pan and / or zoom commands from the user (e.g., via a browser interface). Hypertext Transfer Protocol (HTTP) byte range requests may bey used to achieve efficiency and / or to enable functionality using some built-in browser and web server capabilities. In at least some technologies, cloud optimized geospatial Tagged Image File Format (geoTIFF) may be used. Embodiments of thepresent disclosure provide for HEIF technologies, for example, for such extremely large images, providing feature benefits, capability benefits, and / or other advantages compared with other technologies.
[0077] Some media file format standards which may be used to implement one or more embodiments described herein include International Standards Organization (ISO) base media file format (ISO / IEC 14496-12, which may be abbreviated ISOBMFF), Moving Picture Experts Group (MPEG)-4 file format (ISO / IEC 14496-14, also known as the MP4 format), file format for NAL (Network Abstraction Layer) unit structured video (ISO / IEC 14496-15), High Efficiency Video Coding standard (HEVC or H.265 / HEVC), and / or other format standards (e.g., developing file formats and / or file format standards).
[0078] In files conforming to the ISO base media file format, the media data may be provided in one or more instances of MediaDataBox (‘mdat’) and the MovieBox (‘moov’) may be used to enclose the metadata for timed media. In some cases, for a file to be operable, both of the ‘mdat’ and ‘moov’ boxes may be required to be present. The ‘moov’ box may include one or more tracks, and a respective track may reside in one corresponding TrackBox (‘trak’). Each track is associated with a handler, identified by a four-character code, specifying the track type. Video, audio, and image sequence tracks can be collectively called media tracks, and they contain an elementary media stream. Other track types comprise hint tracks and timed metadata tracks.
[0079] Tracks comprise samples, such as audio or video frames. For video tracks, a media sample may correspond to a coded picture or an access unit.
[0080] A media track refers to samples (which may also be referred to as media samples) formatted according to a media compression format (and its encapsulation to the ISO base media file format). A hint track refers to hint samples, containing cookbook instructions for constructing packets for transmission over an indicated communication protocol. A timed metadata track may refer to samples describing referred media and / or hint samples.
[0081] The 'trak' box includes in its hierarchy of boxes the SampleDescriptionBox, which gives detailed information about the coding type used, and any initialization information needed for that coding. The SampleDescriptionBox contains an entry-count and as many sample entries as the entry-count indicates. The format of sample entries is track-type specific but derived from generic classes (e.g. VisualSampleEntry, AudioSampleEntry). Which type of sample entry formis used for derivation of the track-type specific sample entry format is determined by the media handler of the track.
[0082] The track reference mechanism can be used to associate tracks with one another. The TrackReferenceBox includes box(es), wherein the box(es) provide(s) a reference from the containing track to a set of other tracks. These references are labeled through the box type (e.g., the four-character code of the box) of the contained box(es).
[0083] In ISOMBFF, an edit list provides a mapping between the presentation timeline and the media timeline. Among other things, an edit list provides for the linear offset of the presentation of samples in a track, provides for the indication of empty times and provides for a particular sample to be dwelled on for a certain period of time. The presentation timeline may be accordingly modified to provide for looping, such as for the looping videos of the various regions of the scene.
[0084] Referring now to Fig. 4, an example of a box including an edit list (the EditListBox) is provided. In ISOBMFF, an EditListBox may be contained in EditBox, which is contained in TrackBox ('trak'). In this example of the edit list box, flags specify the repetition of the edit list. By way of example, setting a specific bit within the box flags (the least significant bit, i.e., flags & 1 in ANSI-C notation, where & indicates a bit- wise AND operation) equal to 0 specifies that the edit list is not repeated, while setting the specific bit (i.e., flags & 1 in ANSI-C notation) equal to 1 specifies that the edit list is repeated. The values of box flags greater than 1 may be defined to be reserved for future extensions. As such, when the edit list box indicates the playback of zero or one samples, (flags & 1) shall be equal to zero. When the edit list is repeated, the media at time 0 resulting from the edit list follows immediately the media having the largest time resulting from the edit list such that the edit list is repeated seamlessly.
[0085] In ISOBMFF, a Track group enables grouping of tracks based on certain characteristics or the tracks within a group have a particular relationship. Track grouping, however, does not allow any image items in the group.
[0086] Referring now to Fig. 5, exemplary syntax of TrackGroupBox in ISOBMFF is provided. In the example of Fig. 5, track group type indicates the grouping type and shall be set to one of the following values, or a value registered, or a value from a derived specification or registration: 'msrc' indicates that this track belongs to a multi-source presentation. The tracks that have the same value of track group id within a TrackGroupTypeBox of track group type 'msrc'are mapped as being originated from the same source. For example, a recording of a video telephony call may have both audio and video for both participants, and the value of track group id associated with the audio track and the video track of one participant differs from value of track group id associated with the tracks of the other participant. The pair of track group id and track group type identifies a track group within the file. The tracks that contain a particular TrackGroupTypeBox having the same value of track group id and track group type belong to the same track group.
[0087] Entity grouping is similar to track grouping but enables grouping of both tracks and image items in that same group. Referring now to Fig. 6, exemplary syntax of EntityToGroupBox in ISOBMFF is provided. In the example of Fig. 6, group id is a nonnegative integer assigned to the particular grouping that shall not be equal to any group id value of any other EntityToGroupBox, any item ID value of the hierarchy level (file, movie, or track) that contains the GroupsListBox, or any track ID value (when the GroupsListBox is contained in the file level). In the example of Fig. 6, num entities in group specifies the number of entity id values mapped to this entity group. In the example of Fig. 6, entity id is resolved to an item, when an item with item ID equal to entity id is present in the hierarchy level (e.g., file, movie, and / or track) that contains the GroupsListBox, or to a track, when a track with track ID equal to entity id is present and the GroupsListBox is contained in the file level.
[0088] Files conforming to the ISOBMFF may contain any non-timed objects, referred to as items, meta items, and / or metadata items, in a meta box (four-character code: ‘meta’). While the name of the meta box refers to metadata, items can generally contain metadata or media data. The meta box may reside at the top level of the file, within a movie box (four-character code: ‘moov’), and within a track box (four-character code: ‘trak’), but at most one meta box may occur at the file level, movie level, and / or track level. The meta box may be required to contain a ‘hdlr’ box indicating the structure or format of the ‘meta’ box contents. The meta box may list and characterize any number of items that can be referred and / or respective items may be associated with a file name and are uniquely identified with the file by item IDentifier (item id) which is an integer value. The metadata items may be for example stored in the 'idat' box of the meta box or in an 'mdaf box or reside in a separate file. If the metadata is located external to the file then its location may be declared by the DatalnformationBox (four-character code: ‘dinf ). In the specific case that the metadata is formatted using extensible Markup Language (XML)syntax and is required to be stored directly in the MetaBox, the metadata may be encapsulated into either the XMLBox (four-character code: ‘xml ‘) or the Binary XMLBox (four-character code: ‘bxml’). An item may be stored as a contiguous byte range, or it may be stored in several extents, wherein respective extents are contiguous byte ranges. In other words, items may be stored fragmented into extents, e.g. to enable interleaving. An extent is a contiguous subset of the bytes of the resource. The resource can be formed by concatenating the extents.
[0089] A common base structure is used to contain general untimed metadata. This structure is called the MetaBox, as it was originally designed to carry metadata — data that is annotating other data. However, it may be used for a variety of purposes including the carriage of data that is not annotating other data, for example, when present at ‘file level’. The MetaBox is required to contain a HandlerBox indicating the structure or format of the MetaBox contents. Other contained boxes (e.g., all other contained boxes) are specific to the format specified by the HandlerBox. The other boxes defined hereing may be defined as optional or mandatory for a given format. If they are used, then they shall take the form specified herein. These optional boxes include a DatalnformationBox, which documents other files in which metadata values (e.g., pictures) are placed, and / or an ItemLocationBox, which documents where in those files respective items are located (e.g. in the common case of multiple pictures stored in the same file). At most one MetaBox may occur at the file level, segment, movie level, and / or track level. If an ItemProtectionBox occurs, then some or all of the metadata, including possibly the primary resource, may have been protected and be un-readable unless the protection system is taken into account. The MetaBox is a container box extending FullBox.
[0090] Metadata items are identified by item ID. Within a given MetaBox, a given item ID shall uniquely refer to a single item. When an item is updated in movie fragments, the item ID refers to the latest received version. Derived specifications may further restrict the criteria for uniqueness: unique among the item IDs in both file and movie-level boxes, and / or unique within that set extended with the track ID of the tracks in a movie box.
[0091] In some examples, there are three scopes for item IDs: file and segments; MovieBox and MovieFragmentBox; and TrackBox and TrackFragmentBox. In other words, there shall be only one item with a given item ID within a given scope (e.g. in the TrackBox and all TrackFragmentBox with the same track ID).
[0092] Referring now to Fig. 7, an exemplary metadata format is provided. The structure or format of the metadata is declared by the handler. In the case that the primary data is identified by a primary item, and that primary item has an item information entry with an item type, the handler type may be the same as the item type. The ItemPropertiesBox enables the association of any item with an ordered set of item properties. Item properties may be regarded as small data records. The ItemPropertiesBox consists of two parts: ItemPropertyContainerBox that contains an implicitly indexed list of item properties, and one or more ItemProperty AssociationBox(es) that associate items with item properties.
[0093] High Efficiency Image File Format (HEIF) is a standard developed by the Moving Picture Experts Group (MPEG) for storage of images and image sequences. Among other things, the standard facilitates file encapsulation of data coded according to the High Efficiency Video Coding (HEVC) standard. HEIF includes features building on top of the used ISO Base Media File Format (ISOBMFF).
[0094] The ISOBMFF structures and features are used to a large extent in the design of HEIF. The basic design for HEIF comprises still images that are stored as items and image sequences that are stored as tracks. An item in HEIF is defined as the data that does not require timed processing, as opposed to sample data, and is described by the boxes contained in a MetaBox.
[0095] In the context of HEIF, the following boxes may be contained within the root-level 'meta' box and may be used as described in the following. In HEIF, the handler value of the Handler box of the 'meta' box is 'pict'. The resource (e.g., within the same file, or in an external file identified by a uniform resource identifier) containing the coded media data is resolved through the Data Information ('dinf ) box, whereas the Item Location ('iloc') box stores the position and sizes of every item within the referenced file. The Item Reference ('iref ) box documents relationships between items using typed referencing. If there is an item among a collection of items that is in some way to be considered the most important compared to others, then this item is signaled by the Primary Item ('pitm') box. Apart from the boxes mentioned here, the 'meta' box is also flexible to include other boxes that may be necessary to describe items.
[0096] Any number of image items can be included in the same file. Given a collection of images stored by using the 'meta' box approach, it sometimes is essential to qualify certain relationships between images. Examples of such relationships include indicating a cover imagefor a collection, providing thumbnail images for some or all of the images in the collection, and associating some or all of the images in a collection with an auxiliary image such as an alpha plane. A cover image among the collection of images is indicated using the 'pitm' box. A thumbnail image or an auxiliary image is linked to the primary image item using an item reference of type 'thmb' or 'auxf, respectively.
[0097] HEIF defines defines a derived image item (an item with an item type value of 'grid') whose reconstructed image is formed from one or more input images in a given grid order within a larger canvas.
[0098] The input images are inserted in row-major order, top-row first, left to right, in the order of SingleltemTypeReferenceBox of type 'dimg' for this derived image item within the ItemReferenceBox. In the SingleltemTypeReferenceBox of type 'dimg', the value of from item lD identifies the derived image item of type 'grid', the value of reference count shall be equal to rows*columns, and the values of to item ID identify the input images. All input images shall have exactly the same width and height; call those tile width and tile height. The tiled input images shall completely “cover” the reconstructed image grid canvas, where tile width* columns is greater than or equal to output width and tile_height*rows is greater than or equal to output height.
[0099] The reconstructed image is formed by tiling the input images into a grid with a column width equal to tile width and a row height equal to tile height, without gap or overlap, and then trimming on the right and the bottom to the indicated output width and output height.
[0100] If the desired input images are not of a consistent size, then derived image items that scale or crop them, as needed to make them consistent, can be used; other specifications can, however, restrict whether derived image items are permissible as input to the image grid derived image item. When removing an item that is marked as an input image of an image grid item, the content of the image grid item might need to be rewritten.
[0101] Referring now to Fig. 8, exemplary syntax of a gird-derived image item is provided. Semantics of the parameters in the grid derived image item may be as follows: a. version shall be equal to 0. Readers shall not process an ImageGrid with an unrecognized version number; b. (flags & 1) equal to 0 specifies that the length of the fields output width, output height, is 16 bits, (flags & 1) equal to 1 specifies that the length of thefields output width, output height, is 32 bits. The values of flags greater than 1 are reserved; c. output width, output height: specifies the width and height, respectively, of the reconstructed image on which the input images are placed. The image area of the reconstructed image is referred to as the canvas; and / or d. rows minus one, columns minus one: specifies the number of rows of input images, and the number of input images per row. The value is one less than the number of rows or columns respectively. Input images populate the top row first, followed by the second and following, in the order of item references.
[0102] In some examples, grid derived image items are limited to 256 tiles by 256 tiles because the rows minus one and columns minus one are stored as 8 bit integers.
[0103] The Draft International Standard amendment 1 of HEIF (ISO / IEC 230008- 12:2024 / AMD l:2024(E) WG03N1297_MDS24143) specifies the Constrained Extents Grid Property, which may be defined as follows: a. Box type: 'cexg' b. Property type: Descriptive item property c. Container: ItemPropertyContainerBox d. Mandatory (per item): No e. Quantity (per item): At most one
[0104] The ConstrainedExtentsGridProperty descriptive item property indicates that respective extents of the associated image item in the itemLocationBox are constrained to enclose data units of the item that are extractable as a contiguous byte range and are independently decodable and renderable as image tiles.
[0105] Some (e.g., all) data units or properties required to configure the decoder and decode an image tile are declared in the decoder configuration and initialization properties associated with the image item. The reconstructed image of the associated image item is formed from one or more image tiles in a given grid order within a larger canvas.
[0106] The image tiles corresponding to the extents are inserted in row-major order, top-row first, left to right, in the order of the extents for the associated image item within the ItemLocationBox. The value of extent count within the ItemLocationBox shall be equal to (l+rows_minus_one)*(l+columns_minus_one). Some (e.g., all) image tiles shall have exactlythe same width and height, image tile width and image tile height. The reconstructed image is formed by tiling the image tiles into a grid with a column width equal to image tile width and a row height equal to image tile height, without gap or overlap. The grid of image tiles shall completely “cover” the reconstructed image of the associated image item, where image tile width* columns is greater than or equal to image width and image_tile_height*rows is greater than or equal to image height, where image width and image height are signalled in the ImageSpatialExtentsProperty associated with the image item.
[0107] Referring now to Fig. 9, exemplary syntax of the ConstrainedExtentsGridProperty is provided. Semantics of the parameters of the constrained extents grid property may be as follows: a. (flags & 1) equals to 0 specifies that the length of the fields image tile width and image tile height is 16 bits, (flags & 1) equals to 1 specifies that the length of the fields image_tile_width and image_tile_height is 32 bits. The values of flags greater than 1 are reserved; b. image_tile_width, image_tile_height: specify respectively the width and height in pixels of the image tiles; and / or c. rows minus one, columns minus one: specify the number of rows of image tiles, and the number of image tiles per row. The value is one less than the number of rows or columns respectively. Image tiles enclosed in extents populate the top row first, followed by the second row and following rows, in the order of extents.
[0108] An overview image is described by a grid derived image item or a tiled pre-derived coded image item whose reconstructed image is formed from generating a lower resolution, ‘binned’ version of the reconstructed image of a base image item. The base image item is also a tiled image item. The tiling may be implemented using a feature of a specific codec, or by using a grid derived image item. When a grid derived image item is used, the input items to the grid define the tiles. Derived image items shall not be used as inputs to the image grid, due to the need for in place byte range accessing of content. Individual tiles shall be written contiguously in memory, thereby allowing access with a single read or write action.
[0109] A pre-defined coded image item representing an overview image or an image item representing the base image that are tiled using a feature of a specific codec shall be stored in such a way that respective extents identify that data range corresponding to a tile, and shall beassociated with a ConstrainedExtentsGridProperty indicating the constraint on the extents and describing the tiling grid.
[0110] An overview image shall be tiled using the same tiling scheme as the base image, for example, if tiles in the base image are X by Y pixels, they are X by Y pixels in the overview image. In cases where the binned resolution results in a fractional, or incomplete tile at the end of a row (column), the last tile in a row (column) of tiles shall be padded with the value zero at the end of the row (column) to complete the last tile in the row (column). The clean aperture transformative property ('clap') may be applied to crop padded rows and / or columns. The number of tiles in a row (column) of tiles is determined by dividing the width (height) of the overview image by the tile size in X (tile size in Y) and rounding up.
[0111] The image format of the overview images is the same as the base image, for example, the overview images may have the same number of bands, bit depth, color format, etc. as the base image.
[0112] Overview images are associated with the original full resolution base image, using a reference of type 'base' and can be stacked together with the base image as a series of progressively binned images in an Image Pyramid Entity Group, which may be defined as follows: a. Box Type: 'pymd' b. Container: GroupsListBox in a MetaBox at file level c. Mandatory: No d. Quantity: Zero or more
[0113] The ImagePyramidEntityGroup indicates a set of image items, formed as a base image item and a series of progressively binned overview image items, which together form an image pyramid. At least one overview image item (e.g., a respective overview image item) has a reference to the original full resolution base image item, using a reference of type 'base'. The ImagePyramidEntityGroup also provides overall information for the individual tiles inside the overview image items and base image item of the image pyramid.
[0114] The image format of the overview images shall be the same as the base image (e.g., same number of bands, bit depth, color format, etc.). This entity group shall contain entity id values that point to a base image item and a set of overview image items and shall contain no entity id values that point to tracks. The entities shall be listed in the order of lowest resolutionoverview image item to the highest resolution overview image item, followed finally by the base image item of the image pyramid. There may be multiple ImagePyramidEntityGroups in the same file with different group id values.
[0115] All the entities of a same ImagePyramidEntityGroup, or only some of them, can also be members of a same entity group of type 'prgr' if they are stored in the file for allowing a progressive refinement. They can also be members of a same entity group of type 'altr' if they are proposed by the content creator as alternatives to be displayed for players not supporting the ImagePyramidEntityGroup. When using region partition groups jointly with an image pyramid, the area covered by a region partition group should correspond to the area of a tile of the image pyramid.
[0116] A region item may be associated with an image item within an ImagePyramidEntityGroup, for example, via at least one of: (a) an item reference of type 'cdsc' from the region item to the image item; and / or (b) a RegionPartitionGroupBox associated with the image item via an item reference of type 'rpds' and referencing the item ID of the region item. A region item associated with a base image or an overview image within a same ImagePyramidEntityGroup may be applied to the output image of any image item within this ImagePyramidEntityGroup by applying the implicit resampling caused by the difference between the reference space of the region item and the size of the image.
[0117] A player can use the item reference of type 'base' of a merge region item to filter the region items that are inherited from other image items in the ImagePyramidEntityGroup. Referring now to Fig. 10, exemplary syntax of the ImagePyramidEntityGroup is provided. Semantics of the parameters of the image pyramid entity group may be as follows: a. num entities in group is as defined for EntityToGroupBox. In addition, it also specifies the number of layers of the image pyramid; b. tile size x, tile_size_y indicate the size in pixels of a tile in the width and height dimension, respectively, for all layers of the image pyramid; c. layer binning indicates for respective layers of the pyramid the level of binning between the base image and the overview image. A 2x2 binning is defined to be a layer binning of 2, a 4x4 binning is defined to be 4, etc. The width and height for an overview image with layer_binning of 2 is half the width and half the height of the base image, etc. A base image has a layer_binning of 1; and / ord. tiles in layer row minusl, tiles in layer column minusl indicate the number of tiles minus one in a row and a column, respectively, of a specific layer. If the layer is represented by a grid derived image item, tiles_in_layer_row_minusl is equal to rows minus one and tiles in layer column minusl is equal to columns minus one. If the layer is represented by a tiled pre-derived coded image item with a ConstrainedExtentsGridProperty, then tiles in layer row minusl is equal to rows minus one and tiles in layer column minusl is equal to columns minus one.
[0118] The ISO / IEC 23008-12:2024 / CDAM 2:2024(E) Amendment 2: WG03N1298_2414 specifies the Low-overhead image file format.
[0119] The low-overhead image file format provides a more compact representation of the image file format for at least some use cases. This format is designed for small and simple files where the traditional use of the MetaBox would result in significant overhead relative to the size of the image and / or metadata payloads. In this format, the top-level MetaBox is replaced by a MinimizedlmageBox, which logically maintains the presence of the MetaBox by representing its contents.
[0120] For example, if a parser encounters a MinimizedlmageBox, it may expand it to a MetaBox. The minimized image box format provides a more compact representation of the MetaBox for a subset of use cases. It is meant to be used for small and simple files where the full MetaBox would result in considerable overhead compared to the image data payload.
[0121] Referring now to Figs. 11 A-l IK, exemplary syntax of the MinimizedlmageBox is provided. Semantics of the parameters of the MinimizedlmageBox may be as follows: a. version: specifies the version of the MinimizedlmageBox. The version shall be set to 0 in this version of this document; b. small dimensions flag: if set to 0, the length of the fields signaled among width minusl, height minusl, gainmap width minusl and gainmap_height_minusl is 7 bits; otherwise, it is 15 bits; c. width minusl : plus 1 specifies the width of the reconstructed image in pixels; d. height minusl : plus 1 specifies the height of the reconstructed image in pixels; e. orientation minusl : plus 1 specifies the Exif orientation value as defined in JEITA CP-3451E section 4.6.4.A "Orientation";f. icc flag: equal to 1 indicates that the main image is associated with an ICC profile as defined in ISO 15076-1 or ICC.1
[0013] ; g. exif flag: equal to 1 indicates the presence of Exif metadata; h. xmp flag: equal to 1 indicates the presence of XMP metadata; i. full range flag: is a binary value representing the VideoFullRangeFlag as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2. This signaling applies exclusively to the main image; j. chroma subsampling: 0 specifies that there is exactly one channel of coded color samples (monochrome), otherwise there are exactly three channels of coded color samples (one luma channel and two chroma channels). A value of 1 specifies that these chroma channels are subsampled both horizontally and vertically by a factor 2 (i.e. 4:2:0). A value of 2 specifies that these chroma channels are subsampled by a factor 2 horizontally (i.e. 4:2:2). A value of 3 specifies that there is no subsampling of these chroma channels (i.e. 4:4:4); k. chroma is horizontally centered: 0 specifies that the chroma samples of the main image are co-located horizontally with the luma samples of the main image, otherwise they are horizontally centered between the luma samples of the main image. 0 unless chroma subsampling is 1 or 2; l. chroma is vertically centered: 0 specifies that the chroma samples of the main image are co-located vertically with the luma samples of the main image, otherwise they are vertically centered between the luma samples of the main image. 0 unless chroma subsampling is 1 ; m. float flag: specifies the format of the pixel values of the reconstructed main and alpha image items as the channel format values, as specified in PixellnformationProperty with version 1 in clause 6.5.6; n. bit_depth_log2_minus4: specifies the format of floating-point numbers used for the pixel values of the reconstructed main and alpha image items. The values 0, 1, and 2 respectively correspond to the bits_per_channel values 16, 32 and 64, as specified in PixellnformationProperty with version 1 in clause 6.5.6. Other values are reserved. When float_flag is set to 0, the value is undefined;o. high bit depth flag: 0 specifies that the number of bits per channel for the pixel values of the reconstructed main and alpha image items, as specified in PixellnformationProperty with version 1 in clause 6.5.6, is 8. Otherwise bit_depth_minus9 is signaled. When float flag is set to 1, the value is undefined. p. bit_depth_minus9: specifies the number of bits, minus nine, per channel for the pixel values of the reconstructed main and alpha image items, as specified in PixellnformationProperty with version 1 in clause 6.5.6. When high bit depth flag is set to 0 or float flag is set to 1, the value is undefined; q. alpha flag: 0 specifies that the image is opaque. Otherwise the image has an alpha layer, whether the codec has native translucency support or an alpha auxiliary image item is used; r. alpha_is_premultiplied: when set to 1 specifies that the color channels are premultiplied by the alpha channel, otherwise the color channels are not premultiplied. Ignored if alpha flag is 0; s. explicit cicp flag: equal to 0 indicates the sRGB on-screen colors as the values of ColourPrimaries and Transfercharacteristics, as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2, respectively set to 1 and 13 if icc_flag is 0, and to 2 and 2 otherwise. 0 specifies sRGB on-screen colors as the value of MatrixCoefficients, as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2, set to 2 if chroma subsampling is 0, and to 6 otherwise. When the value is equal to 1 it indicates that these values are signaled explicitly; t. colour_primaries: carries a ColourPrimaries value as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the main image; u. transfer characteristics: carries a Transfercharacteristics value as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the mam image; v. matrix coefficients: carries a MatrixCoefficients value as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the mam image; w. explicit codec types flag: 0 specifies that the minor version of the FileTypeBox carries a brand defining a single coded image item type and a single codec configuration property box type. Shall be set to 1 otherwise;x. infe type: carries the coded image item type. Corresponds to the item type field of the version 2 of the ItemlnfoEntry box. Defined by the brand carried by the minor version of the FileTypeBox if explicit codec types flag is 0; y. codec config type: carries the codec configuration property box type. Defined by the brand carried by the minor version of the FileTypeBox if explicit_codec_types_flag is 0; z. hdr flag: 0 specifies that the image is SDR and has no associated HDR-related signaling. Otherwise the image is either SDR with a SDR-to-HDR gain map, or HDR with an optional HDR-to-SDR gain map; aa. gainmap flag: 0 specifies that the file has no tone-mapped image and no associated HDR-related ISO 21496-1 gain map. Otherwise the file contains a tone-mapped image and is associated with a gain map, whether the codec has native gain map support or a separate gain map image item is used. 0 if hdr flag is 0; bb. gainmap width minusl : carries the width minus one of the gain map image in pixels; cc. gainmap height minusl: carries the height minus one of the gain map image in pixels; dd. gainmap matrix coefficients: carries a MatrixCoefficients value as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the gam map image; ee. gainmap full range flag: carries a VideoFullRangeFlag as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the gain map image; ff. gainmap chroma subsampling: 0 specifies that there is exactly one channel of coded gain map samples (monochrome), otherwise there are exactly three channels of coded gain map samples (one luma channel and two chroma channels). A value of 1 specifies that these chroma channels are subsampled both horizontally and vertically by a factor 2 (i.e. 4:2:0). A value of 2 specifies that these chroma channels are subsampled by a factor 2 horizontally (i.e. 4:2:2). A value of 3 specifies that there is no subsampling of these chroma channels (i.e. 4:4:4);gg. gainmap chroma is horizontally centered: 0 specifies that the chroma samples of the gain map image are co-located horizontally with the luma samples of the gain map image, otherwise they are horizontally centered between the luma samples of the gain map image. Ignored unless gainmap chroma subsampling is 1 or 2; hh. gainmap chroma is vertically centered: 0 specifies that the chroma samples of the gain map image are co-located vertically with the luma samples of the gain map image, otherwise they are vertically centered between the luma samples of the gain map image. Ignored unless gainmap chroma subsampling is 1; ii. gainmap float flag: specifies the format of the pixel values of the reconstructed gain map image item as the channel format values, as specified in PixellnformationProperty with version 1 in clause 6.5.6; jj. gainmap_bit_depth_log2_minus4: specifies the format of floating-point numbers used for the pixel values of the reconstructed gain map image item. The values 0, 1, and 2 respectively correspond to the bits_per_channel values 16, 32 and 64, as specified in PixellnformationProperty with version 1 in clause 6.5.6. Other values are reserved. When float_flag is set to 0, the value is undefined; kk. gainmap high bit depth flag: 0 specifies that the number of bits per channel for the pixel values of the reconstructed gain map image item, as specified in PixellnformationProperty with version 1 in clause 6.5.6, is 8. Otherwise gainmap_bit_depth_minus9 is signaled. When gainmap float flag is set to 1 , the value is undefined;11. gainmap_bit_depth_minus9: specifies the number of bits, minus nine, per channel for the pixel values of the reconstructed gain map image item, as specified in PixellnformationProperty with version 1 in clause 6.5.6. When high bit depth flag is set to 0 or float flag is set to 1 , the value is undefined; mm. tmap icc flag: if 1, specifies that the tone-mapped image is associated with an ICC profile as defined in ISO 15076-1 or ICC.1
[0023] , 0 if gainmap flag is 0; nn. tmap explicit cicp flag: 0 specifies sRGB on-screen colors as the values of ColourPrimaries, Transfercharacteristics and MatrixCoefficients, as defined inRec. ITU-T H.273 | ISO / IEC 23091-2, associated with the tone-mapped image, set to 1, 13 and 6, respectively. Otherwise these values are signaled explicitly; oo. tmap_colour_primaries: carries a ColourPrimaries value as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the tone-mapped image; pp. tmap transfer characteristics: carries a Transfercharacteristics value as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the tone-mapped image; qq. tmap matrix coefficients: carries a MatrixCoefficients value as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the tone-mapped image; rr. tmap full range flag: carries a VideoFullRangeFlag as defined in Rec. ITU-T H.273 | ISO / IEC 23091-2 for the tone-mapped image. Set to 1 if tmap_explicit_cicp_flag is 0; ss. clli flag: 1 specifies that there is signaling for ContentLightLevel attached to the main image. Otherwise no such signaling is present. 0 if hdr_flag is 0; tt. mdcv flag: 1 specifies that there is signaling for MasteringDisplayColourVolume attached to the main image. Otherwise no such signaling is present. 0 if hdr flag is 0; uu. cclv flag: 1 specifies that there is signaling for ContentColourVolume attached to the main image. Otherwise no such signaling is present. 0 if hdr_flag is 0; vv. amve_flag: 1 specifies that there is signaling for AmbientViewingEnvironment attached to the main image. Otherwise no such signaling is present. 0 if hdr flag is 0; ww. reve_flag: 1 specifies that there is signaling for ReferenceViewingEnvironment attached to the main image. Otherwise no such signaling is present. 0 if hdr_flag is 0; xx. ndwt_flag: 1 specifies that there is signaling for NominalDiffuseWhite attached to the main image. Otherwise no such signaling is present. 0 if hdr_flag is 0; yy. tmap clli flag: 1 specifies that there is signaling for ContentLightLevel attached to the tone-mapped image. Otherwise no such signaling is present. 0 if gainmap_flag is 0;zz. tmap mdcv flag: 1 specifies that there is signaling for MasteringDisplayColourVolume attached to the tone-mapped image. Otherwise no such signaling is present. 0 if gainmap_flag is 0; aaa. tmap cclv flag: 1 specifies that there is signaling for ContentColourVolume attached to the tone-mapped image. Otherwise no such signaling is present. 0 if gainmap_flag is 0; bbb. tmap amve flag: 1 specifies that there is signaling for AmbientViewingEnvironment attached to the tone-mapped image. Otherwise no such signaling is present. 0 if gainmap_flag is 0; ccc. tmap_reve_flag: 1 specifies that there is signaling for ReferenceViewingEnvironment attached to the tone-mapped image. Otherwise no such signaling is present. 0 if gainmap_flag is 0; ddd. tmap ndwt flag: 1 specifies that there is signaling for NominalDiffuseWhite attached to the tone-mapped image. Otherwise no such signaling is present. 0 if gainmap_flag is 0; eee. clli: The box body of the ContentLightLevelBox as defined in ISO / LEC 14496-12 attached to the main image. Only present if clli flag is 1; fff mdcv: The box body of the MasteringDisplayColourVolumeBox as defined in ISO / IEC 14496-12 attached to the main image. Only present if clli flag mdcv is 1; ggg. cclv: The box body of the ContentColourVolumeBox as defined in ISO / IEC 14496-12 attached to the main image. Only present if clli flag cclv is 1; hhh. amve: The box body of the AmbientViewingEnvironmentBox as defined in ISO / IEC 14496-12 attached to the main image. Only present if clli flag amve is 1; iii. reve: The box body of the ReferenceViewingEnvironmentBox attached to the main image. Only present if clli_flag_reve is 1 ; jjj. ndwt: The box body of the NominalDiffuseWhiteBox attached to the main image. Only present if clli_flag_ndwt is 1 ;kkk. tmap clli: The box body of the ContentLightLevelBox as defined in ISO / IEC 14496-12 attached to the tone-mapped image. Only present gainmap flagif tmap clli flag is 1 ;111. tmap mdcv: The box body of the MasteringDisplayColourVolumeBox as defined in ISO / IEC 14496-12 attached to the tone-mapped image. Only present gainmap_flagif tmap_clli_flag_mdcv is 1 ; mmm. tmap cclv: The box body of the ContentColourVolumeBox as defined in ISO / IEC 14496-12 attached to the tone-mapped image. Only present gainmap_flagif tmap_clli_flag_cclv is 1 ; nnn. tmap amve: The box body of the AmbientViewingEnvironmentBox as defined in ISO / IEC 14496-12 attached to the tone-mapped image. Only present gainmap_flagif tmap_clli_flag_amve is 1 ; ooo. tmap reve: The box body of the ReferenceViewingEnvironmentBox attached to the tone-mapped image. Only present gainmap flagif tmap_clli_flag_reve is 1 ; ppp. tmap ndwt: The box body of the NominalDiffuseWhiteBox attached to the tone-mapped image. Only present gainmap flagif tmap clli flag ndwt is 1; qqq. few metadata bytes flag: 0 specifies that the length of the signaled fields among icc data size minusl, tmap icc data size minusl, gainmap metadata size, exif data size minusl and xmp data size minusl is 10 bits, otherwise 20 bits. Undefined unless one of these fields is signaled; rrr. few_codec_config_bytes_flag: 0 specifies that the length of the signaled fields among gainmap item codec config size, main item codec config size and alpha_item_codec_config_size is 3 bits, otherwise 12 bits; sss. few_item_data_bytes_flag: 0 specifies that the length of the signaled fields among gainmap item data size, main item data size minusl and alpha item data size is 15 bits, otherwise 28 bits; ttt. icc data size minusl: carries the size minus one in bytes of the ICC profile as defined in ISO 15076-1 or ICC.1
[0023] , associated with the main image. Undefined if icc_flag is 0uuu. tmap icc data size minusl : carries the size minus one in bytes of the ICC profile as defined in ISO 15076-1 or ICC.1
[0023] , associated with the tonemapped image. Undefined if tmap icc flag is 0; vw. gainmap metadata size: carries the size of the gain map metadata. 0 if gainmap_flag is 0; www. gainmap item data size: carries the size of the coded sample data for the HDR-related gain map image item in bytes. If gainmap flag is set to 1, a size of 0 is reserved for future use. 0 if gainmap_flag is 0; xxx. gainmap item codec config size: carries the size of the codec configuration for the gain map auxiliary image item in bytes. The value 0 specifies that the codec does not need any configuration data for the gain map. 0 if gainmap item data size is 0; yyy. main item codec config size: carries the size of the codec configuration for the main image item in bytes; zzz. main item data size minusl : carries the size minus one of the coded sample data for the main image item in bytes; aaaa. alpha item data size: carries the size of the coded sample data for the alpha auxiliary image item in bytes. If alpha flag is set to 1, the value 0 specifies that the codec has native translucency support and that the alpha samples are coded alongside the color samples in the main item data chunk. 0 if alpha flag is 0; bbbb. alpha item codec config size: carries the size of the codec configuration for the alpha auxiliary image item in bytes. The value 0 specifies that the codec does not need any configuration data for alpha. 0 if alpha item data size is 0; cccc. exif data size minusl : specifies the size minus one of the Exif metadata in bytes. -1 if exif_flag is 0; dddd. xmp data size minusl : specifies the size minus one of the XMP metadata in bytes. -1 if xmp_flag is 0; eeee. trailing bits: padding bits to ensure payloads are 8-bit aligned shall be 0;ffff. alpha item codec config: carries the optional alpha image codec configuration data. When alpha item codec config size is 0, alpha item codec config is not present; gggg. gainmap item codec config: carries the HDR-related gain map image item codec configuration data. When gainmap item codec config size is 0, gainmap item codec config is not present; hhhh. main item codec config: carries the main image item codec configuration data. When main item codec config size is 0, main item codec config is not present; iiii. icc data: carries the ICC profile data of the main image as defined in ISO 15076-1 or ICC.1
[0023] , When icc flag is 0, icc data is not present; jjjj. tmap icc data: carries the ICC profile data of the optional HDR-related tonemapped image as defined in ISO 15076-1 or ICC.1
[0023] , When tmap icc flag is 0, tmap icc data is not present; kkkk. gainmap metadata: Gain map metadata as defined by the GainMapMetadata struct in ISO 21496-1. Not present if gainmap metadata size is 0;1111. alpha item data: carries the coded sample data of the optional alpha image. When alpha item data size is 0, alpha item data is not present; mmmm. gainmap item data: carries the coded sample data of the optional gain map image. When gainmap item data size is 0, gainmap item data is not present; nnnn. main item data: carries the coded sample data of the main image; oooo. exif data: specifies the optional Exif metadata. When exif flag is set to 0, exif data is not present; and / or pppp. xmp data: specifies the optional XMP metadata. When xmp flag is set to 0, xmp data is not present.
[0122] A MinimizedlmageBox has a one-to-one mapping to a MetaBox. Persons having skill in the art shall treat a MinimizedlmageBox as if it were the equivalent MetaBox that is transformed from MinimizedlmageBox as specified herein. When a person having skill in the art encounters the MinimizedlmageBox, it shall be understood that the MinimizedlmageBox willcreate the equivalent MetaBox in memory and populate its contents based on the parsed contents of the MinimizedlmageBox.
[0123] File writers can choose between two formats: they can either write a traditional image file based on the MetaBox or opt for a low-overhead image file format based on MinimizedlmageBox.
[0124] The Open Geospatial Consortium Cloud Optimized GeoTIFF (COG) Standard may rely on: (1) two characteristics of the TIFF v6 format (tiles and reduced resolution subfiles); (2) GeoTIFF keys for georeferenced; and / or (3) the HTTP range, which allows for efficient downloading of parts of imagery and grid coverage data on the web and to make fast data visualization of TIFF or BigTIFF files and fast geospatial processing workflows possible.
[0125] COG-aware applications can download only the information they need to visualize or process the data on the web. The COG standard formalizes the requirements for a TIFF file to become a COG file and for the HTTP server to make COG files available in a fast fashion on the web.
[0126] TIFF is a flexible, adaptable file format for handling images and data within a single file, by including the header tags (e.g., size, definition, image-data arrangement, applied image compression, etc.) that provide metadata about the images. The ability to store image data in a lossless format makes a TIFF file a useful image archive. TIFF can be used to store grey scale, color, or RGB images as well as integer of floating point data, making it ideal as a support for storing the rangeset of a 2D grid coverage data.
[0127] To improve TIFF performance over the web, COG may rely on two characteristics of the TIFF v6 format, the georeference GeoTIFF keys and a relatively unused HTTP property.This way, COG allows for efficient downloading of parts of imagery and grid coverage data on the web, enables fast data visualization, and facilitates faster geospatial processing workflows. This particular type of TIFF has been recently used to set up a large series of remote sensing images on cloud providers repositories (e.g., Amazon Web Services), enabling cloud processing at lower traffic. COG-aware software may be configured to request just the portions of data that it needs, improving access time and bandwidth.
[0128] COG is based, at least in part, on the Geo TIFF standard. In some examples, legacy software may be able to read COG files with no additional modifications. The amount of data available for geospatial analytics has increased considerably in recent years. Therefore,downloading the data into a single computer is often not feasible. Data producers that provide data in the COG format can help decrease how much data is downloaded and copied. This is because online software systems do not need to keep their own copy of the data for efficient access. New online software can access the content efficiently, while old versions can still download the data completely. This avoids the need to have two copies of the file: one for fast access and another for download purposes.
[0129] COG may rely on two complementary approaches: (1) the ability of Geo TIFF to store the raw pixels of the image organized in an efficient way using tiles and overviews; and (2) HTTP GET Range request, which allow web clients to request only portions of a file that they need. Using the first approach, COG organizes the Geo TIFF so the latter requests can easily select and get the parts of the file that are useful for processing.
[0130] The Tiling and Reduced-Resolution Subfiles (sometimes called overviews) in the GeoTIFF format support structure for COG files so that the HTTP GET Range queries can request just the part of the file that is relevant.
[0131] Reduced-Resolution Subfiles come into play when the client wants to render a quick image of the whole or a big part of the area represented in the file. Instead of downloading every pixel, the software can just request a smaller, already created, lower resolution version. The structure of the COG file on an HTTP Range supporting web server enables client software to easily find and download just the part of the whole file that is needed.
[0132] Tiles come into play when some small area of the overall extent of the COG file needs to be processed or visualized. This could be part of a reduced-resolution subfile, or it could be at full resolution. Tile organization makes all the relevant bytes of an area (a tile) to be in the same part of the file, so the software can use HTTP GET Range request to get only the tiles it needs.
[0133] In the context for a TIFF file, Tiling is a strategy for dividing the content in the TIFF file differently than using the classical Strips. In the Strips approach the data are organized into sequences of lines (rows) while tiling creates a number of internal rectangular tiles stored in the actual image. Strips divide the content of an image vertically (rows) but not horizontally (columns). With Tiling, a much quicker access to a certain area or two dimensional bounding box is possible as the relevant data is closer in the file and the portion of bytes that needs to be read is smaller than in the strips approach.
[0134] Reduced-Resolution Subfiles (e.g., overviews) are down-sampled versions of the same image included in the same TIFF file. This means that an overview is a zoomed out version from the original image. It has less detail but is also smaller. For visualization purposes or for analytical processes that do not require full resolution, a COG can provide Reduced-Resolution Subfiles that match different scale denominators or cell sizes required by clients. Reduced- Resolution Subfiles increase the size of the file but also increase performance.
[0135] HTTP Version 1.1 introduced a range header in the GET requests that supports requesting only a fragment of a resource. If the server advertises "Accept-Ranges: bytes" in its response headers of a HEAD or GET request, the server is telling the client that bytes of data can be requested in parts, in separated requests. The client can request just the bytes that it needs from the server at any time. In a web environment, this is very useful for serving files such as video. By using range requests, clients do not need to download the entire file to begin playing it. In the case of COG, HTTP range is useful to get only the tiles needed to be processed or shown. This is done by getting the headers and IFDs of the TIFF file first and using this information to determine the conversion between tile indices to byte ranges containing the needed tiles. A client trying to show a COG file on the screen can request the resolutions needed and only the tiles needed to cover the screen. Once the user moves or pans, other GET range requests will get the new needed resolutions and tiles.
[0136] Some problems with the above-mentioned methods include: limitations in range of various parameters of boxes used to define image items; lack of definitions for large image retrieval, decoding, and / or rendering in HEIF; and / or other problems. The HEIF format uses many boxes to define image items; however, the parameters used in such boxes may be limited in range, for example, due to their underlying bit representation. For example, the Item ID field in the Item Location Box may be limited to unsigned 32 bits, which may not be sufficient for representing large geospatial images with a grid / tile division. The HEIF format defines grid derived image items; however, retrieval, decoding, and / or rendering of large images (e.g., large geospatial images) using tiles and / or grid storage are not defined in HEIF.
[0137] To at least partially tackle these and other problems, there is proposed a solution for signaling images, such as large quantities of images, with tiles and / or grids of the high efficiency image file (HEIF) format.
[0138] In some embodiments, “cloud optimized rendering” (and / or any one or more of: “on- demand rendering”, “tile-based rendering”, “grid-based rendering”, “large-scale tiled hierarchical image retrieval rendering”, and / or any other suitable name) may refer to displaying image content in successive steps, wherein respective steps display certain portions (e.g., small spatial regions) of the image content based on a current viewing position, pan interaction, zoom interaction, and / or other feature of a user and / or client.
[0139] Referring now to Fig. 12, shown is an example system architecture which may be used to realize the said methods, apparatuses, and computer program products defined in this invention. The figure shows multiple instances of the image called the overview images, where each image is of different resolution and having different grid / tiling scheme. In the example there are three overview images with 4x8, 3x6 and 2x4 grid / tiling scheme. At a given instance of time, a client / user may be viewing only a small portion of the whole image at a given resolution. Based on the user interaction grid / tiles from a different region and different resolution may be retrieved by the client and displayed to the user.
[0140] In an embodiment, in large-scale tiled hierarchical image retrieval rendering, the portion of the image displayed may vary from an entire image (e.g., at the lowest resolution) to a very small area of the image (e.g., a portion of an intermediate resolution image or highest resolution image).
[0141] In an embodiment, large-scale tiled hierarchical image retrieval rendering displays image content in successive steps, where a respective step may improve the perceived image quality over that of the previous step (e.g., if the user zooms into a specific region of the image, then regions from higher quality and / or resolution are rendered) and is superimposed over the image content of the previous step in the same displaying window; and / or the image content of the previous step is flushed out from the displaying window, and the new content is rendered on the displaying window.
[0142] In an embodiment, large-scale tiled hierarchical image retrieval rendering displays image content in successive steps where a respective step may decrease the perceived image quality over that of the previous step (e.g., if the user zooms out of a specific region of the image, then regions from lower quality / resolution is rendered) and is superimposed over the image content of the previous step in the same displaying window; and / or the image content ofthe previous step is flushed out from the displaying window, and the new content is rendered on the displaying window.
[0143] In an embodiment, large-scale tiled hierarchical image retrieval rendering displays image content in successive steps where a respective step may render a different part of the same image over that of the previous step (e.g., if the user pans to a different region of the image, then regions from same quality / resolution is rendered) and is superimposed over the image content of the previous step in the same displaying window; and / or the image content of the previous step is flushed out from the displaying window, and the new content is rendered on the displaying window.
[0144] In an embodiment, large-scale tiled hierarchical image retrieval rendering displays image content in successive steps where a respective step may render a different part of the different image over that of the previous step (e.g., if the user pans and together zooms in / out to a different region of the image, then regions from different quality and / or resolution is rendered) and is superimposed over the image content of the previous step in the same displaying window; and / or or the image content of the previous step is flushed out from the displaying window, and the new content is rendered on the displaying window.
[0145] In an embodiment, the first rendering step in large-scale tiled hierarchical image retrieval rendering may result in a base quality and / or resolution image and / or may result in a portion and / or region of a full quality and / or resolution image, or a portion and / or region of an intermediate quality and / or resolution image.
[0146] In an embodiment, when the first rendering step in large-scale tiled hierarchical image retrieval rendering is of a portion and / or region of a full quality and / or resolution image or a portion and / or region of an intermediate quality and / or resolution image: the rendering may start at a position defined by the content provider called “the initialization position” (e.g., wherein the initialization position may be part of the bitstream carried together with image content); and / or the rendering may start at the center of the image and / or at any position chosen (e.g., randomly) by the player and / or Tenderer.
[0147] In an embodiment, given a navigation command (e.g., pan and / or zoom) during large- scale tiled hierarchical image retrieval rendering, a first homography, inpainting, affine transformation, and / or any other suitable transformation of the previously rendered image is provided before loading the required optimal quality. This enables reducing the latency in case anew tile or quality portion is required. In an embodiment, missing areas may be rendered with uniform color (e.g., black and / or gray) in this pre-rendering step.
[0148] In an embodiment, for example, to support large-scale tiled hierarchical image retrieval rendering, the HEIF format may support at least one of the following features: a. images may be stored as grid derived image items and / or images encoded with some form of grids and / or tiles inherently supported by the codecs (e.g., motion constrained tile set in HEVC encoded images, subpictures in WC coded images, and / or the like). Grid and / or tile support may be inherent with the storage, for example, if image data is uncompressed (without any encoding); b. images may be grouped together to form overviews and / or reduced resolution sub-files which represent the same content but are of different resolutions (e.g., from a very low resolution to a very high resolution) and have the grid- and / or tile-based storage support; c. images may be additionally and / or optionally mapped to geospatial metadata, for example, if they contain geospatial data and / or are used for geospatial applications; and / or d. storage constraints for large-scale tiled hierarchical image retrieval rendering may be defined by specifying a new brand and / or by defining a new item property for images which are grid derived image items and / or images encoded with some form of grids and / or tiles inherently supported by the codecs.
[0149] In an embodiment, a Metabox with version = 0, may be used to carry at least some of one or more items and / or item properties (e.g., all items and / or item properties) used for large- scale tiled hierarchical image retrieval rendering. Additionally or alternatively, a new version (e.g., version = 1) of the Metabox is defined to carry at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0150] In an embodiment, a payload of the Metabox containing at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering is compressed, for example, using a deflate algorithm.
[0151] In an embodiment, a compressed version of the MetaBox is defined to carry large- scale tiled hierarchical image retrieval image items. The compressed MetaBox has a new 4cc value, for example, ‘cldo’ (and / or any other suitable 4cc value) indicating that the compressedMetaBox contains at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering. The MetaBox may be compressed, for example, via the deflate algorithm and / or any other suitable compression method.
[0152] In an embodiment, a new box may be defined with a new 4cc value, for example, ‘cldo’ (and / or any other suitable 4cc value). The new box replaces the MetaBox at the file-level and carries at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0153] In an embodiment, an HEIF file may be not allowed to carry both the ‘cldo’ box and the MetaBox. Additionally or alternatively, an HEIF file may contain both the ‘cldo’ box and the MetaBox, wherein the ‘cldo’ box contains items and / or item properties used for large-scale tiled hierarchical image retrieval rendering and the MetaBox contains other items used by the application.
[0154] In an embodiment, a new version of mini box is defined. The new version of mini box may carry at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0155] In an embodiment, the container box (e.g., MetaBox with version = 1) which carries at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering may contain: (a) only other child boxes; (b) only the parameters used for defining at least some of (e.g., all) the items and / or item properties used for cloud optimized rendering; and / or (c) a combination of other child boxes and / or parameters used for defining at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0156] In an embodiment, the container box (e.g., MetaBox with version = 1) which carries at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering may contain: (a) the compressed data of only other child boxes; (b) the compressed data of only the parameters used for defining at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering; and / or (c) the compressed data of a combination of other child boxes and parameters used for defining at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0157] In an embodiment, the container box (e.g., MetaBox with version = 1) which carries at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering may contain a HandlerBox to indicate the format of the container box.
[0158] In an embodiment, the container box (e.g., MetaBox with version = 1) which carries at least some of (e.g., all) the items and / or item properties used for large-scale tiled hierarchical image retrieval rendering may not contain the HandlerBox but rather contain only a parameter, for example, such as handler type, which indicates the format of the container box.
[0159] In an embodiment, the container box (e.g., MetaBox with version = 1) which carries at least some of (e.g., all) the items and item properties used for large-scale tiled hierarchical image retrieval may not contain both the HandlerBox and the handler type parameter, but rather that the format of the container box may be inferred by a new brand definition and / or due to the container box being constrained to be used for a specific handler type (e.g., ‘pict’).
[0160] In an embodiment, a new handler_type may be defined with 4cc value ‘geos’ (and / or any other suitable 4cc value) indicating that the format of the structure and / or format of the container box is meant to handle items and / or item properties used for cloud large-scale tiled hierarchical image retrieval.
[0161] In an embodiment, some or all the items, including the primary item (if any), used for large-scale tiled hierarchical image retrieval may be associated with a HandlerProperty item property. The handler type in the handlerProperty is a new 4cc value ‘geos’ (and / or any other suitable 4cc value) indicating that the associated item is meant to be used for large-scale tiled hierarchical image retrieval.
[0162] In an embodiment, item data of the one or more items used for large-scale tiled hierarchical image retrieval may be stored within a single MediaDataBox(‘mdat’) of the HEIF file.
[0163] In an embodiment, the item data of the one or more items used for large-scale tiled hierarchical image retrieval may be stored in one or more MediaDataBox(‘mdat’) of the HEIF file. For example, the item data of a respective item may be stored in a separate MediaDataBox, and / or the item data of a subset of items may be stored in a separate MediaDataBox.
[0164] In an embodiment, the item data of the one or more input items may be stored in one or more MediaDataBox(‘mdat’) of the HEIF file, for example, if the one or more items used for large-scale tiled hierarchical image retrieval may be formed by a collection of other input items(e.g., grid derive image item). For example, the item data of a respective input item to a grid derived image item may be stored in a separate MediaDataBox, and / or the item data of a subset of input item to a grid derived image item may be stored in a separate MediaDataBox.
[0165] In an embodiment, the item data of one or more item extents (e.g., if respective subpictures are represented in an item extent) may be stored in one or more MediaDataBox(‘mdat’) of the HEIF file, for example, if the one or more items used for large- scale tiled hierarchical image retrieval have an inherent grid property (e.g., such as images encoded with VVC subpicture, HEVC motion-constrained tiles, uncompressed images, and / or the like).
[0166] In an embodiment, the item data of a subset of items (e.g., the subset may include zero items, one or more items, etc.) used for large-scale tiled hierarchical image retrieval may be stored in one or more MediaDataBox (‘mdat’) of the HEIF file. The item data of one or more remaining (e.g., all) items may be stored in other external files. The external files may be identified by either DataEntryUrlBox or a DataEntryUrnBox.
[0167] In an embodiment, the item data of the one or more items used for large-scale tiled hierarchical image retrieval may be stored in one or more IdentifiedMediaDataBox (‘imda’) of the HEIF file. For example, the item data of a respective item may be stored in a separate IdentifiedMediaDataBox, and / or the item data of a subset of items may be stored in a separate IdentifiedMediaDataBox.
[0168] In an embodiment, the item data of the one or more input items may be stored in one or more IdentifiedMediaDataBox of the HEIF file, for example, if the one or more items used for large-scale tiled hierarchical image retrieval may be formed by a collection of other input items (e.g., grid derived image item). For example, the item data of a respective input item to a grid derived image item in a separate IdentifiedMediaDataBox, or the item data of a subset of input item to a grid derived image item in a separate IdentifiedMediaDataBox.
[0169] In an embodiment, the item data of the one or more item extents (e.g., if respective subpictures are represented in an item extent) may be stored in IdentifiedMediaDataBox of the HEIF file, for example, if the one or more items used for large-scale tiled hierarchical image retrieval may have an inherent grid property (e.g., images encoded with WC subpicture, HEVC motion-constrained tiles, and / or uncompressed images).
[0170] In an embodiment, a DataEntrylmdaBox may be used to identify the IdentifiedMediaDataBox containing the media data accessed through data reference index corresponding to this DataEntrylmdaBox, for example, if the one or more items (and / or the one or more item extents) used for large-scale tiled hierarchical image retrieval are stored in one or more IdentifiedMediaDataBox of the HEIF file. The DataEntrylmdaBox contains the value of imda identifier of the referred IdentifiedMediaDataBox. The value of imda identifier of the referred IdentifiedMediaDataBox may be equal to the item ID value of the item whose data is present in the corresponding IdentifiedMediaDataBox. Alternatively, the value of imda identifier of the referred IdentifiedMediaDataBox may be equal to a unique identifier value with the unique identifier value associated with the corresponding item in one of the associated boxes for the item and the item data present in the corresponding IdentifiedMediaDataBox.
[0171] In an embodiment, a DataEntrySeqNumlmdaBox may be used to identify the IdentifiedMediaDataBox containing the media data accessed through the data reference index corresponding to this DataEntrySeqNumlmdaBox, for example, if the one or more items (and / or the one or more item extents) used for large-scale tiled hierarchical image retrieval are stored in one or more IdentifiedMediaDataBox of the HEIF file. The value of imda identifier of the referred IdentifiedMediaDataBox may be equal to the item ID value of the item whose data is present in the corresponding IdentifiedMediaDataBox. Additionally or alternatively, the value of imda identifier of the referred IdentifiedMediaDataBox may be equal to a unique identifier value with the unique identifier value associated with the corresponding item in one of the associated boxes for the item and the item data present in the corresponding IdentifiedMediaDataBox.
[0172] In an embodiment, a new data entry box may be defined — called the DataEntryltemIDImdaBox — and / or may be used to identify the IdentifiedMediaDataBox containing the media data accessed through the data reference index corresponding to this DataEntryltemIDImdaBox, for example, if the one or more items (and / or the one or more item extents) used for large-scale tiled hierarchical image retrieval are stored in one or more IdentifiedMediaDataBox of the HEIF file. The value of imda identifier of the referred IdentifiedMediaDataBox may be equal to the item ID value of the item whose data is present in the corresponding IdentifiedMediaDataBox.
[0173] Referring now to Fig. 13, exemplary syntax of DataEntryltemUniquelDImdaBox with a 4cc value equal to 'iuim' (and / or any other suitable 4cc value) is provided.
[0174] In an embodiment, the item data of the one or more items used for large-scale tiled hierarchical image retrieval may not be stored in the ItemDataBox (‘idat’) box.
[0175] In an embodiment, the one or more items used for large-scale tiled hierarchical image retrieval may be formed by a collection of other input items, which may be very large in number. For example, a grid derived image item may be formed by other input image items; a WC base item may be formed by a collection of WC subpicture items.
[0176] In an embodiment, the container box (for example MetaBox with version = 1) which is used for large-scale tiled hierarchical image retrieval may contain a very large number of items, wherein respective items are identified by an item id. The boxes and parameters defined in ISOBMFF and HEIF may be limited by the underlying bit representation to carry a large number of items. For example, the ItemLocationBox defined in ISOBMFF with version = 2 has a 32-bit representation of item count, whereas large-scale tiled hierarchical image retrieval may support an item id parameter having a 64-bit representation. To support large-scale tiled hierarchical image retrieval, at least some existing boxes and / or parameters used for representing such items may need to be modified to accommodate at least some (e.g., all) of the items (e.g., with a 64-bit representation) in the container box.
[0177] In an embodiment, a new version of ItemLocationBox may be defined to support a large number (e.g., greater than a 32-bit range) of item count and / or item IDs. Referring now to Figs. 14A, 14B, and 14C, exemplary syntax of the ItemLocationBox, as defined in ISOBMFF, is provided with modifications.
[0178] In an embodiment, a new version (e.g., version = 3) of ItemLocationBox may enable indicating a count of consecutive item ID values with one loop entry in the ItemLocationBox. Referring now to Figs. 15A, 15B, and 15C, exemplary syntax of the ItemLocationBox in ISOBMFF is provided with modifications. For example, if item batch flag is equal to 1, data reference index may refer to an IdentifiedMediaDataBox that is identified by an item ID value. Additionally or alternatively, if item batch flag is equal to 1, the syntax may exclude data reference index and imply reference to an IdentifiedMediaDataBox that is identified by an item ID value. If item batch flag is equal to 1, one loop entry specifies (item_count_per_entry_minus2 + 2) items with consecutive item_ID values starting from the item ID value signaled for the respective loop entry. The item data for respective items specifiedthis way nay ve present in the IdentifiedMediaDataBox identified by the respective inferred item ID value.
[0179] In an embodiment, a new box may be defined called the ItemLocationLargeBox with 4cc value ‘ ilol’ (and / or any other suitable 4cc value) which replaces the ItemLocationBox, for example, to allow a large number of items to be represented. Referring now to Fig. 16, exemplary syntax of ItemLocationLargeBox is provided. Semantics of the parameters of the Item Location Large Box may be as follows: a. offset_size is taken from the set {0-255} and may indicate the length in bytes of the extent offset field; b. length_size is taken from the set {0-255} and may indicate the length in bytes of the extent length field; c. base_offset_size is taken from the set {0-255} and may indicate the length in bytes of the base offset field; d. index_size is taken from the set {0-255} and may indicate the length in bytes of the item reference index field; e. item count size is taken from the set {0-255} and may indicate the length in bytes of the item count and the item ID fields; f. data_reference_size is taken from the set {0-255} and may indicate the length in bytes of the data reference index fields; and / or g. extent_count_size is taken from the set {0-255} and may indicate the length in bytes of the extent count fields.
[0180] In an embodiment, a new version of Primary ItemBox may be defined, for example, to support larger item ID values (e.g., greater than a 32-bit unsigned integer range). Referring now to Fig. 16, exemplary syntax of the PrimaryltemBox in ISOBMFF with modifications is provided. The modifications may comprise including version = 2, for example, to support a 64- bit unsigned integer range.
[0181] In an embodiment, a new box may be defined called the PrimaryltemLargeBox with 4cc value ‘piml’ (and / or any other suitable 4cc value) which may replace the PrimaryltemBox, for example, to allow a large value of item IDs to be represented as a primary item. Referring now to Fig. 18, exemplary syntax of PrimaryltemLargeBox is provided. Semantics of the parameters of the Primary Item Large Box may be as follows:a. item_id_size is taken from the set {0-255} and may indicate the length in bytes of the item id field.
[0182] As described herein, one or more versions (e.g., including alternate versions) of images having different resolutions and / or tiling grids may be used for large-scale tiled hierarchical image retrieval.
[0183] In an embodiment, an image at the lowest resolution among a group of images used for large-scale tiled hierarchical image retrieval may be marked as the Primary item.
[0184] In an embodiment, an image with a highest resolution among the group of images used for large-scale tiled hierarchical image retrieval may be marked as the Primary item.
[0185] In an embodiment, an arbitrary image among the group of images used for large-scale tiled hierarchical image retrieval may be marked as the Primary item.
[0186] In an embodiment, a thumbnail image of the group of images used for large-scale tiled hierarchical image retrieval may be marked as the Primary item.
[0187] In an embodiment, the container box which is used for large-scale tiled hierarchical image retrieval may not contain the PrimaryltemBox or the PrimaryltemLargeBox or any indication of which image is the primary image. In such a case, any one or more of: the first image; the image with the lowest item ID; the last image; the image with highest item ID; the image with lowest resolution; and / or the image with the highest resolution among the group of images used for cloud optimized rendering may be interpreted as the primary item.
[0188] One or more items used for large-scale tiled hierarchical image retrieval may be encrypted with an item protection scheme. In some examples, not all the items may be encrypted such that a subset of the items used for large-scale tiled hierarchical image retrieval may be encrypted. For example, if the number of protection schemes relied upon by the items used for large-scale tiled hierarchical image retrieval is greater than the range allowed in ItemProtectionBox, the ItemProtectionBox may be modified. In an embodiment, a new version of ItemProtectionBox may be defined, for example, to support larger protection count values (e.g., greater than a 16-bit unsigned integer range). Referring now to Fig. 19, exemplary syntax of the ItemProtectionBox in ISOBMFF is provided with modifications. The modifications may include version = 1, for example, to support 32- and / or 64-bit unsigned integer range.
[0189] In an embodiment, a new box may be defined called the ItemProtectionLargeBox with 4cc value ‘ iprl’ (and / or any other suitable 4cc value) which may replace theItemProtectionBox, for example, to allow a large number of protection count to be used. Referring now to Fig. 20, exemplary syntax of ItemProtectionLargeBox is provided. Semantics of the parameters of the Item Protection Large Box may be as follows: a. protection count size is taken from the set {0-255} and may indicate the length in bytes of the protection count field.
[0190] In some examples, if one or more items used for large-scale tiled hierarchical image retrieval is encrypted and / or if content encoding may have changed the format of the data in the item, ItemlnfoBox mnay be present. For example, if the number of items used for cloud optimized rendering is greater than the range allowed in ItemlnfoBox, the ItemlnfoBox may be be modified. In an embodiment, a new version of ItemlnfoBox may be defined, for example, to support larger entry count and / or Item ID values (e.g., greater than a 32-bit unsigned integer range). Referring now to Figs. 21 A and 21B, exemplary syntax of the ItemlnfoBox in ISOBMFF is provided with modifications. The modifications may include version = 4, for example, to support a 32- and / or / 64-bit unsigned integer range.
[0191] In an embodiment, a new box may be defined called the ItemlnfoLargeBox with 4cc value ‘iinl’ (and / or any other suitable 4cc value) which may replace the ItemlnfoBox, for example, to allow a large number of entry count to be used. In an embodiment, a new box may be defined called the ItemlnfoEntryLargeBox with 4cc value ‘infl’ (and / or any other suitable 4cc value) which may the ItemlnfoEntry to allow a large number of item IDs to be used. Referring now to Figs. 22A and 22B, exemplary syntax of ItemlnfoLargeBox and ItemlnfoEntryLargeBox is provided. Semantics of the parameters of the Item Info Large Box and Item Info Entry Large Box may be as follows: a. entry_count_size is taken from the set {0-255} and indicates the length in bytes of the entry count field; b. item ID size is taken from the set {0-255} and indicates the length in bytes of the item ID field; c. item_protection_size is taken from the set {0-255} and indicates the length in bytes of the item_protection field; d. content encoding flag when set to 1 indicates that a content encoding is applied to the item. Whem set 0 indicates that no content encoding has been applied to the item. When item type is mime then content encoding flag is set to 1 ; and / ore. content extension flag when set to 1 indicates that the item information has content extension information. When set 0 indicates that the item information does not have any content extension.
[0192] In some examples, if one or more items used for large-scale tiled hierarchical image retrieval may be a derived image item (e.g., a grid derived image item) and / or the derived image item may be formed by the number of input items, ItemReferenceBox may be be present. For example, if the number of input items used for a derived image item in cloud optimized rendering is greater than the range allowed in ItemReferenceBox, the ItemReferenceBox may be modified. In an embodiment, a new version of ItemReferenceBox may be defined, for example, to support larger reference count and Item ID values (e.g., greater than a 32-bit unsigned integer range). Referring now to Fig. 23, exemplary syntax of the ItemReferenceBox in ISOBMFF is provided with modifications. The modifications may include version = 2, for example, to support a 32- and / or 64-bit unsigned integer range. Semantics of the parameters of the Item Reference Box may be as follows: a. item_id_size is taken from the set {0-255} and may indicate the length in bytes of the from item id and to item id fields; and / or b. reference count size is taken from the set {0-255} and may indicate the length in bytes of the reference count field.
[0193] In some examples, if one or more items are used for large-scale tiled hierarchical image retrieval, the one or more items may be associated with one or more item properties, for example, such as a codec configuration property (e.g., if the one or more items are encoded), ItemPropertiesBox may be present. For example, if the number of input items used for cloud optimized rendering is greater than the range allowed in ItemPropertiesBox child boxes (ItemProperty AssociationBox), the ItemPropertiesBox child boxes may be modified. In an embodiment, a new version of ItemProperty AssociationBox may be defined, for example, to support larger Item ID values (e.g., greater than a 32-bit unsigned integer range). Referring now to Fig. 24, exemplary syntax of the ItemProperty AssociationBox in ISOBMFF is provided with modifications. The modifications may include version = 2, for example, to support a 32- and / or 64-bit unsigned integer range.
[0194] In an embodiment, a new box may be defined called the ItemProperty AssociationLargeBox with 4cc value ‘ipal’ (and / or any other suitable 4cc value)which may replace the ItemProperty AssociationBox, for example, to allow a large number of entry count and item IDs to be used. In an embodiment, one or more items are used for large- scale tiled hierarchical image retrieval may have common item properties. Associating respective individual items to the same item property may increase the metadata, for example, if the number of items is very large (e.g., such as in large-scale tiled hierarchical image retrieval). In an embodiment, the ItemProperty AssociationLargeBox may allow for associating an array of item IDs to a specific Item Property. Referring now to Fig. 25, exemplary syntax ofItemProperty AssociationLargeBox is provided. Semantics of the parameters of the Item Property Association Large Box may be as follows: a. entry_count_size is taken from the set {0-255} and indicates the length in bytes of the entry count field; b. item id size is taken from the set {0-255} and indicates the length in bytes of the item id field; c. item id array flag when set to 1 indicates that item association is for a range of item IDs as indicated by the value of num of item ids. The num of item ids is expected to be greater than 1 when item id array flag is set to 1; and / or d. item id array flag when set to 0 indicates that item association is for a single item ID and the num of item ids value is set to 1.
[0195] In an embodiment, the item ID parameter may be used in multiple boxes, including ItemlocationBox, PrimaryltemBox, ItemlnfoEntry, from item ID and / or to item ID in SingleltemTypeReferenceBox and SingleltemTypeReferenceBoxLarge,ItemProperty AssociationBox, and / or in the newly defined boxes above. Herein the item ID parameter may be optimized to reduce the number of bits used for its representation.
[0196] In an embodiment, the structures containing the item ID parameter may be further extended or modified. Referring now to Fig. 26, exemplary syntax of extensions and / or modifications for optimize the item ID representation is provided. Semantics of the parameters of the extensions and / or modifications may be as follows: a. the parameter item id size index is represented by a 3 -bit unsigned integer to cover the range of item ID values which can be represented up to a 64-bit representation.
[0197] In an embodiment, for example, if higher bit representation of the item ID values is desired, the bit representation of item id size index may be increased. The trailing bitsQ function may be used to make sure that the structures are byte-aligned.
[0198] In an embodiment, for example, if the image used for large-scale tiled hierarchical image retrieval rendering is a grid derived image item, the input image items to the grid derived image item may be inserted in any one or more of the following ways: a. (by default) row-major order, top-row first, left to right, in the order of SingleltemTypeReferenceBox / SingleltemTypeReferenceLargeBox of type 'dimg' for this derived image item within the ItemReferenceBox; b. column-major order, left-column first, top to bottom, in the order of SingleltemTypeReferenceBox / SingleltemTypeReferenceLargeBox of type 'dimg' for this derived image item within the ItemReferenceBox; c. a first zig-zag order, starting from the top row, as shown in Fig. 27A, in the order of SingleltemTypeReferenceBox / SingleltemTypeReferenceLargeBox of type 'dimg' for this derived image item within the ItemReferenceBox; and / or d. a second zig-zag order, starting from the top row, as shown in Fig. 27B, in the order of SingleltemTypeReferenceBox / SingleltemTypeReferenceLargeBox of type 'dimg' for this derived image item within the ItemReferenceBox.
[0199] In an embodiment, if the image used for large-scale tiled hierarchical image retrieval rendering is a grid derived image item, and if the input image items to the grid derived image item may be inserted in any of the ways defined above, then the ImageGrid structure may be extended to include a parameter which indicates the arrangement order. Referring now to Fig. 28, exemplary syntax of the extensions and / or modifications to the ImageGrid structure is provided. Semantics of the parameters of the extensions and / or modifications may be as follows: a. arrangement direction identifies the direction to apply for input image item arrangement; arrangement direction takes one of the following values: i. input image items are inserted in row-major order, top-row first, left to right; ii. input image items are inserted in column-major order, left-column first, top to bottom;iii. input image items are inserted in a first zig-zag order, starting from the top row; iv. input image items are inserted in a second zig-zag order, starting from the left column; and / or v. other values.
[0200] In an embodiment, for example, if the image used for large-scale tiled hierarchical image retrieval rendering is a grid derived image item and / or if the input image items to the grid derived image item may be inserted in any of the ways defined herein, a new item property may be defined called the InputArrangementltemProperty with the 4cc value (‘iaip’) (and / or any other suitable 4cc value).
[0201] In an embodiment, for example, if the InputArrangementltemProperty is associated with the grid derived image item, the InputArrangementltemProperty indicates that the input image items to the grid derived image item are arranged in different order than the row-major order, top-row first, left to right. Referring now to Fig. 29, exemplary syntax of InputArrangementltemProperty is provided. Semantics of the parameters in InputArrangementltemProperty may be as follows: a. arrangement direction identifies the direction to apply for input image item arrangement; arrangement direction takes one of the following values: i. input image items are inserted in row-major order, top-row first, left to right; ii. input image items are inserted in column-major order, left-column first, top to bottom; iii. input image items are inserted in a first zig-zag order, starting from the top row; iv. input image items are inserted in a second zig-zag order, starting from the left column; and / or v. other values.
[0202] In an embodiment, for example, if the image used for large-scale tiled hierarchical image retrieval rendering is a tiled pre-derived image item, the data in the image extents present in the ItemLocationBox of the associated tiled pre-derived image item may be stored in any one or more of the following ways:a. (by default) row-major order, top-row first, left to right, in the order of extents within the ItemLocationBox; b. column-major order, left-column first, top to bottom, in the order of extents within the ItemLocationBox; c. a first zig-zag order, starting from the top row as shown in Fig. 27A, in the order of extents within the ItemLocationBox; and / or d. a second zig-zag order, starting from the top row as shown in Fig. 27B, in the order of extents within the ItemLocationBox.
[0203] In an embodiment, for example, if the image used for large-scale tiled hierarchical image retrieval rendering is a tiled pre-derived image item and / or the data in the image extents present in the ItemLocationBox of the associated tiled pre-derived image item is arranged in any of the ways defined herein, the ItemLocationBox may be modified or extended to include a parameter which indicates the arrangement order. Referring now to Figs. 30A-30C, exemplary syntax of the extensions and / or modifications to ItemLocationBox is provided. Semantics of the parameters of the extensions and / or modifications may be as follows: a. extent arrangement direction identifies the direction of the extent arrangement; extent arangement direction takes one of the following values: i. the extents are arranged in row-major order, top-row first, left to right; ii. the extents are arranged in column-major order, left-column first, top to bottom; iii. the extents are arranged in a first zig-zag order, starting from the top row; iv. the extents are arranged in a second zig-zag order, starting from the left column; and / or v. other values.
[0204] In an embodiment, for example, if the image used for large-scale tiled hierarchical image retrieval rendering is a tiled pre-derived image item and / or the data in the image extents is present in the ItemLocationBox of the associated tiled pre-derived image item, a new item property may be defined called the ExtentArrangementltemProperty with the 4cc value (‘eaip’) (and / or any other suitable 4cc value).
[0205] In an embodiment, for example, if the ExtentArrangementltemProperty is associated with the tiled pre-derived image item, it indicates that the extent in the ItemLocationBoxassociated to the tiled pre-derived image item are arranged in an order as defined in the ExtentArrangementltemProperty. Referring now to Fig. 31, exemplary syntax of the ExtentArrangementltemProperty is provided. Semantics of the parameters in ExtentArrangementltemProperty may be as follows: a. extent arrangement direction identifies the direction of the extent arrangement; extent arangement direction takes one of the following values: i. the extents are arranged in row-major order, top-row first, left to right; ii. the extents are arranged in column-major order, left-column first, top to bottom; iii. the extents are arranged in a first zig-zag order, starting from the top row; iv. the extents are arranged in a second zig-zag order, starting from the left column; and / or v. other values.
[0206] In an embodiment, for example, if an image is a grid derived image item used for large-scale tiled hierarchical image retrieval rendering with a specific item ID, the input image items to the grid derived image item may or may not be present in the same file (e.g., the same HEIF file) as the grid derived image item.
[0207] In an embodiment, for example, if an image is a grid derived image item used for large-scale tiled hierarchical image retrieval rendering, any one or more of the following may be true: a. the grid derived image item may contain only the ItemReference of type ‘dimg’ from the grid derived image item to the input image item, with the input image items arranged in the order of SingleltemTypeReferenceBox as defined in the above invention embodiments; b. the item related metadata and the media data of the input image items to the grid derived image item may be present in the MediaDataBox (‘mdat’) and / or in the IdentifiedmediaDataBox (‘imda’) of the file; c. alternatively, the item related metadata and the media data of the input image items to the grid derived image item may be present in the external file; d. the grid derived image item with a specific item ID may be associated with an ItemLocationBox entry with a respective item ID, which may comprise theextent count equal to the number of input image items to the grid derived image item; e. the data in each extent corresponds to an input image item and is a conformant image file format. For example, the data is a conformant HEIF file containing a single image image item which is used as input to the grid derived image item. The HEIF file has a structure with FileTypeBox followed by a MetaBox which contains the item related information of the input image item. Alternateively, the HEIF file has a structure with FileTypeBox followed by a MinimizedlmageBox which contains the item related information of the single input image item; and / or f. a client which retrieves each extent should parse the data of the extent as a regular image file format to extract the item related metadata and the media data.
[0208] In an embodiment, one or more new boxes may be defined called: DataEntryMiniBox with 4cc value ‘demb’ (and / or any other suitable 4cc value); DataEntryHEIFBox with 4cc value ‘dehb’ (and / or any other suitable value); DataEntrylmageBox with 4cc value ‘deib’ (and / or any other suitable 4cc value); DataEntrylmageltemBox with 4cc value ‘deii’ (and / or any other suitable 4cc value); DataEntryltemBox with 4cc value ‘deit’ (and / or any other suitable 4cc value); and / or any other suitable name with a suitable 4cc value. In some examples, if the new data entry box identifies the container box having the input image items to the grid derived image item and / or the input image item is contained in an HEIF file with the MiniBox (minimized image box design of the low-overhead HEIF) design, the data entry box may be called DataEntryMiniBox. In some examples, if the new data entry box identifies the container box having the input image items to the grid derived image item and / or the input image item is contained in an HEIF file with the MetaBox design, the data entry box may be called DataEntryHEIFBox. In some examples, if the new data entry box identifies the container box having any image item and / or the image item is contained in an HEIF file with the MetaBox design and / or in an HEIF file with MiniBox design, the data entry box may be called DataEntrylmageBox and / or DataEntrylmageltemBox. In some examples, if the new data entry box identifies the container box having any item and / or the item is contained in an HEIF file with the MetaBox design and / or in an HEIF file with MiniBox design, the data entry box may be called DataEntryltemBox. In some examples, if the new data entry box identifies the box, the DataEntrylmdaBox identifies the IdentifiedMediaDataBox containing the media data accessedthrough the data reference index corresponding to the DataEntrylmdaBox. The DataEntrylmdaBox may contain the value of imda identifier of the referred IdentifiedMediaDataBox. The media data offsets may be relative to the first byte of the payload of the referred IdentifiedMediaDataBox such that media data is offset 0 points to the first byte of the payload of the referred IdentifiedMediaDataBox. In some examples, the grid derived image item with a specific item ID may have an ItemReferenceType ‘dimg’ with the from item lD equal to the item ID of the grid derived image item and / or the to item ID listing at least some (e.g., all) of the item IDs of the input images which form the grid derived image item, wherein the reference count is equal to the extent count.
[0209] Referring now to Fig. 32, an example architecture for implementing one or more embodiments described herein is provided. The architecture comprises a grid derived image item with item ID = 1. The grid derived image item has an item reference of type ‘dimg’ from the grid derived image item to four input image items with item IDs 11 to 14. The grid derived image item has four extents in the ItemLocationBox corresponding to the four input image items. The item related metadata and the media data of the input image items are not present together with the grid derived image item as part of the metabox in the image file. The item related metadata and the media data of the input image items are present in the Metadatabox of the image file. A respective extent of the grid derived image item points to the corresponding input image item and the data in the extent is a conformant HEIF file contain a single input image item in a MetaBox.
[0210] In an embodiment, a new item may be defined called the hierarchical item with 4cc value equal to ‘ hiit’ (and / or any other suitable 4cc value).
[0211] In an embodiment, a hierarchical item may not contain any media data (e.g., no extents) and / or may be a result of processing other input items, where the item related metadata and the media data of the input items to the hierarchical item may be present in the MediaDataBox (‘mdat’) and / or in the IdentifiedmediaDataBox (‘imda’) of the file. Additionally or alternatively, the item related metadata and the media data of the input items to the hierarchical item may be present in the external file.
[0212] In an embodiment, a hierarchical item may refer to other items in external files using the DataReferenceBox and / or are used as inputs to the hierarchical item.
[0213] In an embodiment, the number of input items used by the hierarchical item is defined by the specific hierarchical item structure and / or the corresponding number of entries equal to the number of input items may be present in the DataReferenceBox.
[0214] In an embodiment, the hierarchical item may have no item body (e.g., no extents) such that at least some of the data relied upon (e.g., all the data relied upon) to process the input items may be present in the external file referred to by the DataReferenceBox.
[0215] In an embodiment, for example, if the hierarchical item contains any item body (e.g., one or more extents), the specific hierarchical item may define the process relied upon to parse the data in the extent.
[0216] In an embodiment, a hierarchical item may be associated with one or more item properties. The item properties are applied after processing all the input items of the said hierarchical item.
[0217] In an embodiment, for example, if the hierarchical item is an image item, it may be defined as the hierarchical image item with the 4cc value ‘himg’.
[0218] In an embodiment, for example, if the reconstruction process of a hierarchical image item results in a grid of images, it may be defined as a grid hierarchical image item defined with the 4cc value ‘hgri’.
[0219] In an embodiment, at least some (e.g., all) input images items to a grid hierarchical image item may have the same widths and heights, which may be referred to as tile_width and tile_height, respectively. The tiled input images may completely “cover” the reconstructed image grid canvas, where tile width* columns is greater than or equal to output width and tile_height*rows is greater than or equal to output height.
[0220] In an embodiment, the reconstructed image may be formed by tiling the input images into a grid with a column width equal to tile width and a row height equal to tile height, without gap or overlap, and then trimming on the right and the bottom to the indicated output width and output height.
[0221] In an embodiment, for example, if the desired input images are not of a consistent size, then input image items that scale or crop them, as relied upon to make them consistent, can be used.
[0222] Referring now to Fig. 33, an example architecture for implementing one or more embodiments described herein is provided. The architecture comprises a hierarchical item withitem lD = 1. The hierarchical item lists four input items with item IDs 11 to 14. The hierarchical item has four extents in the ItemLocationBox corresponding to the four input items. The item related metadata and the media data of the input items are not present together with the hierarchical item as part of the metabox in the file. The item related metadata and the media data of the input items are present in the Metadatabox of the file. Respective extents of the hierarchical item point to the corresponding input item, and / or the data in the extent is a conformant HEIF file comprising a single input item in a MetaBox.
[0223] Referring now to Fig 34, an example architecture for implementing one or more embodiments described herein is provided. The architecture comprises a hierarchical item with item ID = 1. The hierarchical item lists four input items with item IDs 11 to 14. The item related metadata and the media data of the input items are not present in the same file as the hierarchical item. The item related metadata and the media data of the input items are present in the external file. Respective external files are conformant HEIF files comprising a single input item in a MetaBox.
[0224] Geospatial applications use large geospatial images using the Cloud Optimized GeoTiff (COG) standard. One of the main features of cloud optimized GeoTiff is the support of HTTP byte range requests, which allows spatial random-access of at least one or more portions of a large geospatial image with tiles, based on user interaction.
[0225] At least the following features are configured to support large geospatial images: a. The HEIF standard specifies two different ways for representing images with tiles: i. the Grid derived image items; and ii. image items with codec-specific inherent tiling (e.g., high efficiency video coding (HEVC) motion-constrained tile sets (MCTS), WC subpictures, and / or the like). b. The DAM1 text on ISO / IEC 23008-12 (MDS24143_WG03N01297) specifies at least the following definitions to support the signalling of large geospatial images: i. the ConstrainedExtentsGridProperty descriptive item property; ii. the overview images; and / or iii. the ImagePyramidEntityGroup.
[0226] In an embodiment, an item type equal to ‘thim’ is called a tiled hierarchical image item, which uses input image items present in external files.
[0227] In an embodiment, the TiledHi erarchi callmag eitem may indicate default information about the input image items which are located in external file.
[0228] In an embodiment, the external files which contain the input image items are HEIF Image file format conformant.
[0229] In an embodiment, the external files which contain the input image items may be formatted with any image file format, for example, including but not limited to AVIF, JPEG, PNG, JPEG2000, and / or any other format which supports carriage of images.
[0230] In an embodiment, the input image items may be contained within the MetaBox of the external HEIF file.
[0231] In an alternate embodiment, the input image items may be contained within the MinimizedlmageBox (‘mini’) of the external HEIF file.
[0232] In an embodiment, the location of input items to the TiledHierarchicallmageltem is identified by the corresponding DataEntryHierarchicalltemBox in the DataReferenceBox, which is mapped to the TiledHierarchicallmageltem through data reference index in the ItemLocationB ox.
[0233] In an embodiment, a respective external file contains one input image item (e.g., only one input image item) and / or does not contain entity grouping (e.g., any entity grouping).
[0234] In an embodiment, a respective external file contains one or more input image items and / or may or may not contain entity grouping (e.g., any entity grouping).
[0235] In an embodiment, the input image item in an external file may contain item properties associated with the input image item.
[0236] In an embodiment, for example, if the input image item in an external file contains item properties associated with the input image item, the input image used for TiledHierarchicallmageltem may be obtained after decoding the input image in the external file and / or applying the associated transformations (e.g., all the associated transformations).
[0237] In an embodiment, the tiled hierarchical image item may contain item properties associated with the tiled hierarchical image item.
[0238] In an embodiment, the handler type of the tiled hierarchical image item is equal to the handler_ type of the input image item(s) in the external file.
[0239] In an embodiment, the reconstructed image of the TiledHierarchicallmageltem is formed from one or more input image items from external files in a given grid order within a larger canvas.
[0240] In an embodiment, the input image items in the external file are inserted in row-major order, top-row first, left to right, in the order as it is listed in the TiledHierarchicallmageltem.
[0241] In an embodiment, the value of the parameter no of input items is equal to ( 1 +ro ws minus one) * ( 1 +columns_minus_one) .
[0242] In an embodiment, at least some (e.g., all) input image items may have width and height equal to, image_tile_width and image_tile_height, respectively.
[0243] In an embodiment, for example, if any of the input image item to the tiled hierarchical image item have different width and height, not equal to image tile width and image tile height, respectively, then the input image item may be resized to the image tile width and image tile height, respectively, before it is used for tiled hierarchical image item.
[0244] In an embodiment, the reconstructed tiled hierarchical image is formed by tiling the input images into a grid with a column width equal to image tile width and a row height equal to image tile height, without gap or overlap.
[0245] In an embodiment, the grid of input images at least partially “cover” (e.g., completely cover) the reconstructed tiled hierarchical image, where image tile width* columns is greater than or equal to image width and image_tile_height*rows is greater than or equal to image height, where image width and image height are signaled in the ImageSpatialExtentsProperty associated with the TiledHierarchicallmageltem.
[0246] Referring now to Figs. 35A-35B, example syntax of TiledHierarchicallmageltem is provided. Semantics of the parameters of TiledHierarchicallmageltem may be as follows: a. version may be equal to 0; b. input items size index specifies the size of the parameters no of input items and item ID in bytes. With value 1 indicating the size is of 1 byte up to the value 7 indicating the size to be 8 bytes; c. default_item_protection_flag when set to 0 specifies that the input image items in the external file are not encrypted with any protection schemes. When default_item_protection_flag is set to 1 , specifies that the input image items in the external file are encrypted with a protection scheme;d. default item type specifies the item type of the input image item. All the input image items in the external files shall be of same item type; e. no of input items is an integer which specifies the number of input image items to the TiledHi erarchi callmageltem; f. item ID specifies the item ID of the input image items in the external file; g. image_tile_width, image_tile_height: specify respectively the width and height in pixels of the input image items; and / or h. rows minus one, columns minus one: specify the number of rows of image tiles, and the number of image tiles per row. The value is one less than the number of rows or columns respectively.
[0247] In an embodiment, at least some (e.g., any) information in the TileHierachi callmageltem struct may be overridden by the information in the external file containing the input image items. For example, the default_item_protection_flag in the TileHierachicallmageltem struct may be set to 0, indicating that the input image items in the external file are not encrypted. Some of the input image items to the Tiled hierarchical image item may be encrypted with a known protection scheme. The information about the encryption of input image item may be known (e.g., may only be known) after the external file containing the input image item is retrieved and parsed by the file reader / player.
[0248] In an embodiment, the tiled hierarchical image item may be an overview image and / or a base image.
[0249] In an embodiment, for example, if the tiled hierarchical image item is an entity within the ImagePyramidEntityGroup, then the parameters image tile width, image tile height and rows minus one, columns minus one may be present in the TiledHierarchi callmageltem struct and / or may be gated using a flag. For example, a new flag called tile_info_present_flag (and / or any other suitable name) may be present in the TiledHierarchicallmageltem struct.
[0250] In an embodiment, for example, if the tile_info_present_flag present in the TiledHierarchicallmageltem struct is set to 0, the parameters image tile width, image tile height and rows minus one, columns minus one, are not present in the TiledHierarchicallmageltem struct.
[0251] In an embodiment, when the tile_info_present_flag present in the TiledHierarchicallmageltem struct is set to 1, the parameters image tile width, image tile height and rows minus one, columns minus one, are present in the TiledHierarchicallmageltem struct.
[0252] In an embodiment, for example, if the parameters image tile width, image tile height and rows minus one, columns minus one, are not present in the TiledHierarchicallmageltem struct, and / or if the tiled hierarchical image item is an entity within the ImagePyramidEntityGroup, the parameter values may be infered as follows: a. image tile width is set equal to the tile size x parameter of theImagePyramidEntityGroup; b. image tile height is set equal to the tile_size_y parameter of theImagePyramidEntityGroup; c. rows minus one, is set equal to the tiles in layer row minusl parameter of the ImagePyramidEntityGroup; and / or d. columns minus one is set equal to the tiles in layer column minusl parameter of the ImagePyramidEntityGroup.
[0253] In an embodiment, the TiledHierarchicallmageltem struct may contain the handler type of respective input image items.
[0254] In an embodiment, a new DataEntryHierarchicalltemBox may be defined in the DataReferenceBox of the DatalnformationBox as follows: a. Box Type: 'dehi' b. Container: DataReferenceBox c. Mandatory: No. d. Quantity: Zero or more.
[0255] Referring now to Fig 36, example syntax of DataEntryHierarchicalltemBox is provided. Semantics of the parameters of DataEntryHierarchicalltemBox may be as follows: a. input items size index specifies the size of the parameters no of input items and referenced item ID in bytes. With value 1 indicating size is of 1 byte up to the value 7 indicating the size to be 8 bytes; b. location_URN_flag if this flag is set, it indicates that the location field is a URN string, otherwise (not set) the location string is a URL;c. referenced item ID indicates the identifier (item ID) of the input image item in the referred external file; and / or d. location indicates the location of the referred file as a URN or URL, depending on the flag location URN flag. If the indicated location is a URL, it can be an absolute or a relative URL, and the located resource shall be a compliant HEIF file. Relative URLs are relative to the file that contains this location. In an embodiment, the location URN flag may not be present and / or the semantic of the location parameter may represent (e.g., may always represent) a URL.
[0256] In an embodiment, the DataEntryHierarchicalltemBox identifies the location of the external files which carry the input image items of a TiledHierar chicallmag eitem. The parameter no of input items in DataEntryHierarchicalltemBox may be equal to the parameter no of input items in TiledHi erarchi callmageltem.
[0257] In an embodiment, respective locations identified by the DataEntryHierarchicalltemBox correspond to input image items with specific item IDs.
[0258] In an embodiment, large-scale tiled hierarchical image retrieval rendering-based decoding may be defined. ‘Large-scale tiled hierarchical image retrieval rendering-based decoding’ may refer to decoding a bitstream with a single decoder instance in successive steps, wherein respective steps result in an output image of the same size as the previous step. Effects of this approach may include utilizing the same decoder instance without resetting at respective steps of rendering. In an embodiment, large-scale tiled hierarchical image retrieval renderingbased decoding, if enabled in a file, may be used with large-scale tiled hierarchical image retrieval rendering.
[0259] In an embodiment, large-scale tiled hierarchical image retrieval rendering-based refinement may be defined. ‘Large-scale tiled hierarchical image retrieval rendering-based refinement’ may refer to large-scale tiled hierarchical image retrieval rendering of image content in a file while downloading the file.
[0260] Referring now to Fig. 37, a method 3700 is provided. As shown in block 3702, an image item may be displayed in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client, and wherein the features of the client include at least one of: a current viewing position; a pan interaction; a zoom interaction; and / or a different interaction. As shown in block 3704, one or more parameters of the image item may beconfigured to support 64-bit representations. As shown in block 3706, at least one ItemID parameter of the image item may be configured to support representations greater than a 32-bit integer range. As shown in block 3708, one or more boxes of the image item may be ordered by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping. As shown in block 3710, one or more images may be stored as grid derived image items supported by at least one codec. As shown in block 3712, one or more images may be grouped to form at least one of an overview or a reduced-resolution subfile representing a reduced- resolution version of a higher-resolution image item. As shown in block 3714, a new brand may be specified for a grid derived image item, wherein the specifying further comprises defining a new item property for the grid derived image item. As shown in block 3716, a new version of MiniBox may be defined, wherein the new version is configured to carry at least a subset of items or item properties used for large-scale tiled hierarchical image retrieval rendering.
[0261] In an embodiment, at least some of the processes described herein may be carried out by an apparatus comprising means for carrying out at least some of the described processes. Means for performing method steps as disclosed herein may include software and / or hardware components of the apparatus 100. For example, the at least one controller 102, the memory 104, and the instructions comprised by the memory 104 (e.g., computer program code) form means for carrying out the method or methods as disclosed herein, and any of the embodiments thereof. As used herein, the term “means” is to construed in singular form, i.e., referring to a single element, or in plural form, i.e., referring to a combination of single elements. Therefore, terminology “means for [performing A, B, C]” is to be interpreted to cover an apparatus in which there is only one means for performing A, B, and C, or where there are separate means for performing A, B, and C, or partially or fully overlapping means for performing A, B, C. Further, terminology “means for performing A, means for performing B, means for performing C” is to be interpreted to cover an apparatus in which there is only one means for performing A, B, and C, or where there are separate means for performing A, B, and C, or partially or fully overlapping means for performing A, B, C.
[0262] Even though the invention has been described above with reference to an example according to the accompanying drawings, it is clear that the invention is not restricted thereto but can be modified in several ways within the scope of the appended claims. Therefore, all words and expressions should be interpreted broadly, and they are intended to illustrate, not to restrict,the embodiment. It will be obvious to a person skilled in the art that, as technology advances, the inventive concept can be implemented in various ways. Further, it is clear to a person skilled in the art that the described embodiments may, but are not required to, be combined with other embodiments in various ways.
Claims
CLAIMS:
1. An apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: display an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; configure one or more parameters of the image item to support 64-bit representations; configure at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and order one or more boxes of the image item by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping.
2. The apparatus of claim 1, wherein the features of the client comprise at least one of: a current viewing position; a pan interaction; a zoom interaction; or a different interaction.
3. The apparatus of any of claims 1 or 2, wherein the instructions, when executed by the at least one processor, further cause the apparatus at least to perform: storing one or more images as grid derived image items supported by at least one codec.
4. The apparatus of any of claims 1 to 3, wherein the instructions, when executed by the at least one processor, further cause the apparatus at least to perform: grouping one or more images to form at least one of an overview or a reduced-resolution subfile representing a reduced-resolution version of a higher-resolution image item.
5. The apparatus of any of claims 1 to 4, wherein the instructions, when executed by the at- 67 -least one processor, further cause the apparatus at least to perform: specifying a new brand for a grid derived image item.
6. The apparatus of any of claims 1 to 5, wherein the specifying further comprises defining a new item property for the grid derived image item.
7. The apparatus of any of claims 1 to 6, wherein the instructions, when executed by the at least one processor, further cause the apparatus at least to perform: defining a new version of MiniBox, wherein the new version is configured to carry at least a subset of items or item properties used for large-scale tiled hierarchical image retrieval rendering.
8. A method comprising: displaying an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; configuring one or more parameters of the image item to support 64-bit representations; configuring at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and ordering one or more boxes of the image item by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping.
9. The method of claim 8, wherein the features of the client comprise at least one of: a current viewing position; a pan interaction; a zoom interaction; or a different interaction.
10. The method of any of claims 8 or 9, further comprising: storing one or more images as grid derived image items supported by at least one codec.
11. The method of any of claims 8 to 10, further comprising:- 68 -grouping one or more images to form at least one of an overview or a reduced-resolution subfile representing a reduced-resolution version of a higher-resolution image item.
12. The method of any of claims 8 to 11, further comprising: specifying a new brand for a grid derived image item.
13. The method of any of claims 8 to 12, wherein the specifying further comprises defining a new item property for the grid derived image item.
14. The method of any of claims 8 to 13, further comprising: defining a new version of MiniBox, wherein the new version is configured to carry at least a subset of items or item properties used for large-scale tiled hierarchical image retrieval rendering.
15. A non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus at least to: display an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; configure one or more parameters of the image item to support 64-bit representations; configure at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and order one or more boxes of the image item by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping.
16. The non-transitory computer readable medium of claim 15, wherein the features of the client comprise at least one of: a current viewing position; a pan interaction; a zoom interaction; or a different interaction.- 69 -17. The non-transitory computer readable medium of any of claims 15 or 16, wherein the program instructions, when executed by the apparatus, further cause the apparatus to: store one or more images as grid derived image items supported by at least one codec.
18. The non-transitory computer readable medium of any of claims 15 to 17, wherein the program instructions, when executed by the apparatus, further cause the apparatus to: group one or more images to form at least one of an overview or a reduced-resolution subfile representing a reduced-resolution version of a higher-resolution image item.
19. The non-transitory computer readable medium of any of claims 15 to 18, wherein the program instructions, when executed by the apparatus, further cause the apparatus to: specify a new brand for a grid derived image item.
20. An apparatus comprising: means for displaying an image item in successive steps, wherein respective steps display predetermined portions of the image item based on features of a client; means for configuring one or more parameters of the image item to support 64-bit representations; means for configuring at least one ItemID parameter of the image item to support representations greater than a 32-bit integer range; and means for ordering one or more boxes of the image item by rewriting the at least one ItemID parameter via an IdentifiedMediaDataBox (IMDA) mapping.- 70 -
Citation Information
Patent Citations
File format with identified media data box mapping with track fragment box
EP4068781A1
Method and apparatus for late binding in media content
US20220167042A1