Photo coding operations for different image displays

By carrying metadata of primary and non-primary image formats in the image (container) file, the problem of consumer devices supporting multiple image formats is solved, enabling efficient and accurate image rendering on different display devices and improving the consistency and quality of image display.

CN121816747APending Publication Date: 2026-04-07DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In the prior art, consumer devices typically have a limited number of image codecs installed, which cannot effectively support multiple image formats, leading to incorrect interpretation of image content and the generation of artifacts, especially inconsistent image rendering between different display devices.

Method used

By carrying metadata of primary and secondary image formats in the image (container) file, compositor metadata and display management metadata are used to optimize image reconstruction and rendering, generating image content adapted to different display devices.

Benefits of technology

It enables efficient and accurate rendering of image content on different display devices, making full use of the device's display capabilities and reducing image artifacts and inconsistencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121816747A_ABST
    Figure CN121816747A_ABST
Patent Text Reader

Abstract

A primary image of a first image format is encoded into an image file specified for the first image format. A non-primary image in a second image format is encoded into one or more companion segments of the image file. The second image format is different from the first image format. A display image derived from the reconstructed image is rendered with a recipient device of the image file. The reconstructed image is generated from one of the primary image or the non-primary image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This patent application claims the priority of U.S. Provisional Application No. 63 / 528,608, filed July 24, 2023, and U.S. Provisional Application No. 63 / 606,424, filed December 5, 2023, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure generally relates to images. More specifically, embodiments of this disclosure relate to photo decoding operations for different image displays. Background Technology

[0004] Display technology has evolved to support the transmission and rendering of image content (e.g., photographs) based on specific image formats. For example, JPEG image encoders and decoders can support image content decoded in JPEG image format. Other image encoders and decoders can support image content decoded in different image formats.

[0005] Consumer or end-user devices, such as handheld devices, typically have a limited set of image codecs installed or configured, each of which can support a specific image format from a limited set of image formats. Therefore, if an image is not encoded and delivered in an image-video format, the device may not be able to find a suitable image decoder to decode and help render the image content. Even if rendered, the rendered image content may include incorrect interpretations or representations of the received image content, producing visible artifacts in color and brightness values.

[0006] As the inventors of this paper recognize, there is a desire to improve the technology for composing image content data that can be used to support the display capabilities of a wide variety of display devices.

[0007] The methods described in this section are methods that can be pursued, but are not necessarily methods that were previously conceived or pursued. Therefore, unless otherwise indicated, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise indicated, problems identified with respect to one or more methods should not be assumed, based on this section, to have been identified in any prior art. Attached Figure Description

[0008] Embodiments of the invention are illustrated in the accompanying drawings by way of example and not by way of limitation, and like reference numerals in the drawings indicate similar elements, wherein:

[0009] Figure 1 An example of the delivery pipeline processing is depicted;

[0010] Figure 2A and Figure 2B The diagram illustrates an example image codec architecture;

[0011] Figures 3A to 3C The illustration shows an example decoding syntax or structure for an image (container) file, an APP11 tag fragment, or one or more data boxes;

[0012] Figure 4A and Figure 4B The example processing flow is illustrated;

[0013] Figure 5 A simplified block diagram of an example hardware platform on which a computer or computing device as described herein can be implemented is illustrated;

[0014] Figure 6A and Figure 6B An example (image / photo) capture device is described;

[0015] Figure 6C An example image processing device is described;

[0016] Figure 7 An example image / photo receiving device is depicted;

[0017] Figure 8 A sample image / photo packaging operation is described; and

[0018] Figure 9 The illustration shows an example image metadata compression operation. Detailed Implementation

[0019] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent, however, that this disclosure may be practiced even without these specific details. In other instances, well-known structures and devices have not been described exhaustively to avoid unnecessarily obscuring, obscuring, or confusing this disclosure.

[0020] Overview

[0021] The techniques described herein can be used to package or encode photos or still images in primary and / or non-primary image formats along with image metadata in an image (container) file, enabling downstream receiving devices to reconstruct photos or still images of different formats, dynamic ranges, color spaces, bit depths, etc.

[0022] Image (container) files (e.g., JPEG image files, etc.) can be specified as carrying photos / images (e.g., JPEG images, etc.) in a primary image format (e.g., JPEG, etc.). Accompanying (attendant) (data) fragments (such as APP11 fragments, etc.) of image (container) files can be used to carry photos / images (e.g., non-JPEG images, etc.) in a non-primary image format (e.g., non-JPEG, etc.) and / or image metadata.

[0023] Image metadata can carry operational parameters or their values ​​that have been optimized by upstream devices. These optimized operational parameters or values ​​can be used in image reconstruction operations, such as forward and / or backward reshaping operations.

[0024] Reconstructed photos or still images generated from photos / images in an image (container) file can be optimized for rendering on the image display of the downstream receiving device by using image metadata carried in the same image (container) file.

[0025] The example embodiments described herein relate to packaging and encoding photographs in an image (container) file. A primary image of a first image format is encoded into an image file specified for the first image format. A secondary image of a second image format is encoded into one or more accompanying segments of the image file. The second image format differs from the first image format. A receiving device uses the image file to render a display image derived from the reconstructed image. The reconstructed image is generated from either the primary or secondary image.

[0026] The example embodiments described herein relate to decoding and rendering photographs from an image (container) file for image reconstruction and rendering. An image file specifying a first image format is received. This image file is encoded with a primary image in the first image format. A secondary image in a second image format is decoded from one or more accompanying fragments of the image file. The second image format differs from the first image format. A display image derived from the reconstructed image is rendered on an image display. The reconstructed image is generated from the secondary image decoded from the image file.

[0027] Example video delivery processing pipeline

[0028] Figure 1 An example processing of the image delivery pipeline (100) is depicted, illustrating the various stages from image capture / generation to image displays of different types or capabilities. The example image displays may include, but are not limited to, high dynamic range (HDR) image displays, standard dynamic range (SDR) image displays, image displays operating in conjunction with an end-user or personal computer, mobile devices, home theaters, televisions, head-mounted displays, wearable displays, etc.

[0029] Image frames (102) are captured or generated using an image generation block (105). The image frames (102) may be captured digitally (e.g., by a digital camera or an image signal processor (ISP) operating in a specific mode or camera setting, etc.) or generated by a computer (e.g., using computer animation, etc.) to provide image data (107). In some embodiments, the image data (107) may be edited or transformed into a post-ISP or post-ISP processed image by post-ISP image processing operations (e.g., automatically, manually, automatically, etc., with or without human input) before being passed to the next processing stage / stage in the image delivery pipeline (100).

[0030] The image data (107) is then provided to the processor for post-production image processing (115). Post-production image processing (115) may include (e.g., partially or fully automatically, partially or fully manually, using image enhancement or processing applications running on a computing device, image cropping, visual effects, global or local hue and / or color adjustments, etc.) adjusting or modifying the colors or brightness in the image to enhance image quality or achieve a specific look for the image according to the creative intent of the image content creator. Therefore, post-production image processing (115) operates on the image data (107) to produce a release version of one or more images to be decoded into an image signal (such as an image (container) file).

[0031] In some operational scenarios, one or more primary images and / or secondary images (117) — for example, a single JPEG image as a primary image, a primary image plus zero, one or more secondary images, etc. — can be decoded into an image signal or an image (container) file.

[0032] In some operational scenarios, the primary and / or secondary images (117) may have been (e.g., in post-production image processing (115)) forward or backward remodeled to generate a relatively efficient, relatively high-quality image—compared to the input or source image from the ISP or post-ISP image processing operation—for decoding into an image signal or image (container) file. Therefore, the decoding block (120) may receive the primary and / or secondary images (117) as the remodeled image. It should be noted that in other operational scenarios, the primary or secondary images (117) encoded in the image (container) file may not represent the remodeled image.

[0033] In some embodiments, the decoding block (120) and / or post-processing image processing (115) can implement, for example, Figure 2AThe codec framework is shown. The primary and / or secondary images (117) can be compressed / encoded by the decoding block (120) into an image signal or image (container) file (122) (e.g., a decoded bitstream, binary file, JPEG image file, etc.). In some embodiments, the decoding block (120) may include one or more image encoders or codecs, such as those associated with industry-standard or proprietary specification delivery formats, to generate the image (container) file (122).

[0034] In some embodiments, the image (117) retains the content creator’s intent (also referred to as “artist’s intent”) and generates the image in post-production image processing (115) based on that content creator’s intent.

[0035] In some embodiments, the image container file (122) is an image signal that conforms to one or more (image) decoding or decoding syntax specifications.

[0036] The image signal or image (container) file (122) may also include or be decoded with image metadata (117-1), including but not limited to compositor metadata. The image metadata (117-1) may be generated by the decoding block (120) and / or the post-production block (115). The compositor metadata—e.g., forward and / or backward reshaping maps, lookup tables, etc.—may be used by the downstream decoder to perform forward / backward reshaping (e.g., tone mapping, inverse tone mapping, etc.) on the primary image and / or non-primary image (117) to generate one or more other images, including but not limited to the display image, which may also be rendered relatively accurately on one or more other image displays, in addition to the primary and / or non-primary image (117) being optimized for use on one or more reference image displays on which it is rendered.

[0037] As used herein, reshaping (e.g., forward reshaping, backward reshaping, tone mapping, inverse tone mapping, etc.) can refer to image processing operations that convert between different EOTFs, different color spaces, different dynamic ranges, etc. Additionally, optionally, or alternatively, backward or inverse reshaping refers to image processing operations that convert a requantized image back to the original EOTF domain (e.g., gamma or PQ, etc.) or to a different EOTF domain for further downstream processing, such as display management.

[0038] The image container file (122) further encodes one or more portions of image metadata (117-1), including but not limited to one or more specific display management (DM) metadata portions, which can be used by a downstream decoder to perform specific display management operations on the decoded or backward-reconstructed image for a specific image display to generate a display image optimized for rendering on the specific image display. Examples of display management operations and corresponding DM metadata portions in the (image) metadata are described in U.S. Patent Application Publication No. 2022 / 0164931, “Display management for high dynamic range images,” by Robin Atkins, Jaclyn Anne Pytlarz, and Elizabeth G. Pieri, the entire contents of which are incorporated herein by reference.

[0039] The image (container) file (122) is then delivered to a downstream receiver or receiving device, such as a decoding and playback device, a media source device, a media streaming client device, a television (e.g., a smart TV), a set-top box, a cinema, etc. In the receiver (or downstream device), the image (container) file (122) is decoded by a decoding block (130) to generate a decoded image (182), which may be identical to one of the primary image and / or the secondary image (117) affected by quantization errors generated during compression performed by the decoding block (120) and decompression performed by the decoding block (130).

[0040] Some or all of the image metadata (117-1) transmitted to the receiving device along with the primary and / or secondary images (117) in the image (container) file—including, but not limited to, synthesizer metadata—can be generated automatically, in real-time, in offline processing, etc., by the decoding block (120) and / or post-processing image processing (115). In some embodiments, the image data (117-1) is provided to the decoding block (120) and / or post-processing image processing (115) for synthesizer metadata generation. Synthesizer metadata generation can be performed automatically with little or no human interaction.

[0041] Composer metadata can be used to provide or generate image content optimized for a wide variety of display devices or image displays. Composer metadata can also be used to generate additional images that are not available or sent in the image (container) file (122). Therefore, as long as the primary and / or secondary images (117) in the image (container) file (122) and the composer metadata are available, the techniques described herein can be used to generate or synthesize image content specifically optimized for non-reference image displays. Image content for these non-reference image displays can be optimized to explore the full or relatively wide range of display capabilities of these non-reference displays.

[0042] Additionally, optionally, or alternatively, the DM metadata in the image metadata can be used by a downstream decoder to perform display management operations on the backward-reconstructed image, generating device-specific display images for rendering on a wide variety of display devices or image displays.

[0043] In an operational scenario where the receiver operates together with (or is attached to) an image display 140, the image display 140 is supported by or is the target of the decoded image (182) (which is the same as or substantially similar to one of the primary and / or secondary images (117) affected by encoding errors), and the receiver can render the decoded image on the image display (140).

[0044] In an operational scenario where the receiver operates (or is attached to) an image display 140-1, which is not supported by or targeted by any of the decoded images (182), the receiver can extract compositor metadata (e.g., lookup table or LUT-based compositor metadata, polynomial-based compositor metadata, multichannel multi-regression (MMR) compositor data, tensor product B-spline (TPB) compositor metadata, non-TPB compositor metadata, etc.) from the image (container) file (122) and use the compositor metadata to synthesize an inverse / backward reconstructed image (132) (also referred to as a constructed or reconstructed image) based at least in part on one of the compositor metadata and / or the decoded image (182). Furthermore, the receiver can extract DM metadata from the image (container) file (122) and apply DM operations (135) to the reconstructed image (132) based on the DM metadata to generate a corresponding display image (137) for rendering on a display device (140-1) (e.g., a non-referenced image).

[0045] Image displays that can generate optimized display images for rendering according to the techniques described herein may include image displays with various dynamic ranges.

[0046] As used herein, the term “dynamic range” (DR) can refer to the ability of the human visual system (HVS) to perceive the range of intensity (e.g., luminance, brightness) in an image (e.g., from the darkest black (dark) to the brightest white (bright)).

[0047] As used in this article, the term "high dynamic range (HDR)" refers to a DR width that spans approximately 14-15 or more orders of magnitude across the human visual system (HVS). In practice, a DR with a wide range of intensity that humans can simultaneously perceive can be slightly truncated relative to HDR.

[0048] In practice, an image comprises one or more color components in a color space (e.g., lightness Y and chromaticity Cb and Cr), where each color component is represented by a precision of n bits per pixel (e.g., n=8, etc.). Using non-linear luminance decoding (e.g., gamma decoding), images where n≤8 (e.g., color 24-bit JPEG images, etc.) are considered to have standard dynamic range, while images where n>8 can be considered to have enhanced dynamic range.

[0049] As used herein, the term "PQ" refers to Perceptual Luminance Amplitude Quantization. The human visual system responds to increasing levels of illumination in a highly nonlinear manner. A person's ability to see a stimulus is influenced by the luminance of the stimulus, the size of the stimulus, the spatial frequency constituting the stimulus, and the luminance level to which the eye adapts at a particular moment of viewing the stimulus. In some embodiments, the perceptual quantizer function maps linear input gray levels to output gray levels that better match the contrast sensitivity threshold in the human visual system. An example PQ mapping function is described in SMPTE ST 2084:2014 "High Dynamic Range EOTF of Mastering Reference Displays" (hereinafter referred to as "SMPTE"), which, given a fixed stimulus size, selects the minimum visible contrast step size for each luminance level (e.g., stimulus level, etc.) based on the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model).

[0050] A reference electro-optical transfer function (EOTF) for a given display characterizes the relationship between the color values ​​(e.g., luminance) of the input image signal and the color values ​​(e.g., screen luminance) of the output screen produced by the display. For example, a reference EOTF for flat panel displays is defined by referencing ITU Rec. ITU-R BT. 1886, “Reference electro-optical transfer function for flat panel displays used in HDTV studio production” (March 2011), which is incorporated herein by reference in its entirety. It supports 200 to 1,000 cd / m² relative to HDR. 2 A display with a brightness of nits represents low dynamic range (LDR), also known as standard dynamic range (SDR). Further examples of EOTF are defined or described in SMPTE 2084 and Rec. ITU-R BT.2100, “Image parameter values ​​for high dynamic range television for use in production and international programme exchange” (06 / 2017), which are incorporated herein by reference in their entirety.

[0051] Codec framework

[0052] Figure 2A and Figure 2B The diagram illustrates an example image codec architecture. More specifically, Figure 2A The diagram illustrates an example encoder-side codec architecture, which can be implemented using one or more computational processors in the upstream image encoder. Figure 2B The diagram illustrates an example codec architecture on the decoder side, which can also be implemented using one or more computational processors in a downstream image decoder (e.g., a receiver, etc.).

[0053] In such Figure 2A The encoder-side codec architecture shown receives raw images or photographs, such as those generated by the camera's ISP or post-ISP image processing tools, as input. These raw images or photographs can be used to export or generate image (container) files to be included or contained within. Figure 1 The main image and / or non-main image (117) in 122). For ease of reference, the main image and / or non-main image (117) may be referred to herein as the main image (in the image (container) file).

[0054] By way of illustration and not limitation, image generator 162—which may represent or include one or more image transformation or mapping tools, etc.—is used to generate one or more primary and / or secondary images (117) derived from or corresponding to the source image. In some embodiments, image generator (162) may perform forward and / or inverse tone mapping operations.

[0055] In the encoder-side codec architecture, the image metadata generator 150 (e.g., Figure 1 The decoding block (120) and / or part of the post-production image processing (115) receive some or all of the main images and / or non-main images (117) as input to generate image metadata (117-1), such as synthesizer metadata, DM metadata, etc.

[0056] In the encoder-side architecture, compression block 142 (e.g., Figure 1 A portion of the decoding block (120), etc., compresses / encodes the primary image and / or non-primary images (117) in image data 144, which is carried or included in the image data. Figure 1 The image signal or image (container) file (122) may be included. Image metadata (117-1) (denoted as "rpu") generated by the image metadata generator (150) may also be included or encoded (e.g., by...). Figure 1 The decoding block (120) etc. is inserted into the image signal or image (container) file (122).

[0057] In the encoder-side architecture, image metadata (117-1) can be carried separately in designated decoded segments of an image signal or image (container) file (122). These designated decoded segments can be separate from one or more specific decoded segments in the image signal or image (container) file (122) used to carry or include the primary image and / or non-primary images (117). For example, image metadata (117-1) can be encoded in designated image metadata segments or syntax elements in the image (container) file (122), while the primary image and / or non-primary images (117) are encoded in designated image data segments (or one or more corresponding syntax elements) in the same image signal or image (container) file (122).

[0058] In the encoder-side architecture, synthesizer metadata in the image metadata (117-1) of the image signal or image (container) file (122) can be used to enable downstream receivers to reshape or map the primary image and / or non-primary images (117) (e.g., forward, backward, reverse, etc.) to one or more reconstructed images (e.g., approximate or identical to one or more non-primary images (148)) for one or more other image displays besides the reference image display supported by the primary image and / or non-primary images (117). Example image displays may include, but are not limited to, any of the following: image displays with display capabilities similar to the reference display, image displays with display capabilities different from the reference display, image displays with additional DM operations (which are used to map reconstructed image content to display image content for the image display), etc.

[0059] In such Figure 2B In the decoder-side architecture shown, the received image signal, encoded with the main image and / or non-main image (117) and image metadata (117-1), is used as input.

[0060] Decompress block 154 (e.g., Figure 1 A portion of the decoding block (130, etc.) decompresses / decodes the compressed image data in the image signal or image (container) file (122) into a decoded image (182). The decoded image (182) may be identical to one of the primary image and / or non-primary image (117) affected by quantization errors in the compression block (142) and decompression block (154). The decoded image (182) may be output to the (reference) image display in the output image signal 156 (e.g., via an HDMI interface, via a video link, etc.) and rendered on the (reference) image display.

[0061] In addition, the image reshaping block 158 extracts image metadata (117-1), such as synthesizer metadata (or backward reshaping metadata), from the input image signal or image (container) file (122), constructs a reshaping function (e.g., backward, reverse, forward, etc.) based on the synthesizer metadata in the extracted image metadata, and performs a reshaping operation on the decoded image (182) based on the reshaping function to generate one or more reshaped images (132) (or reconstructed images), which are used for one or more other image displays in addition to one or more reference image displays.

[0062] In some operational scenarios, the receiver may not perform DM operations to simplify device operation. In some operational scenarios, DM metadata may be transmitted to the receiver along with compositor metadata in the image signal or image (container) file (122) and the primary and / or secondary images (117). Display management operations specific to an image display having different display capabilities than the reference display may be performed on the reshaped or reconstructed image (132) based at least in part on the DM metadata in the image metadata (117-1), for example, to generate a corresponding device-specific display image to be rendered on the actual image display.

[0063] In some operational scenarios, including but not limited to forward-reconstructed SDR images, SDR images may be included or packaged as primary or secondary images in image (container) files as described herein. Other images (including but not limited to HDR images) may be included or packaged together with SDR images, or reconstructed from SDR images (using compositor metadata).

[0064] In other operating scenarios, an image with a dynamic range other than SDR can be included or packaged as a primary or secondary image in an image (container) file as described herein. Additionally, optionally, or alternatively, other image metadata and / or other compositor metadata can be included or packaged in the same image (container) file to construct other images of various image formats, and / or various dynamic ranges, and / or various color spaces, and / or various precisions and / or bit depths.

[0065] Image container file

[0066] An image encoder generates photo / image data and accompanying image metadata, and encodes or compresses the photo / image data and image metadata (e.g., specifically identified, with specific parameter values ​​such as rpu_type = 8) into an image (container) file using decoding syntax conforming to an image container specification (e.g., a specific version, etc.). The image file encoded by the image encoder can be signaled, transmitted, or otherwise delivered directly or indirectly to a downstream receiving device such as an image decoder. The image decoder can use the same decoding syntax to decode the (encoded or compressed) photo / image (payload) data and image metadata from the image file.

[0067] This specification can provide standards-based or proprietary specifications for constructing some or all of the syntax elements (or decoded fragments) of the decoding syntax associated with including or packaging photo / image (payload) data and image metadata in image files.

[0068] The example image (container) file specifications described herein may include, but are not limited to, any of the following: ISO / IEC 10918-1:1994, Information Technology – Digital compression and coding of continuous-tone still images: Requirements and guidelines (similarly, and later defined in ISO / IEC 18477-1:2020-05, Information Technology – Digital compression and coding of continuous-tone still images – Part 1: Core coding system specification); ISO / IEC 10918-4:1999, Information Technology – Digital compression and coding of continuous-tone still images: Registration of JPEG profiles, SPIFF profiles, SPIFF tags, SPIFF color spaces, APPn markers, SPIFF compression Types and Registration Authorities (REGAUT); ISO / IEC 10918-5:2013, Information Technology – Digital compression and coding of continuous-tone still images: JPEG File Interchange Format (JFIF); CIPA DC-008-2019 / JEITACP-3451E, Standard of the Camera & Imaging Products Association, Exchangeableimage file format for digital still cameras: EXIF ​​Version 2.32; ISO / IEC 14496-12:2015, Information Technology – Coding of Audio-Visual Objects, Part 12: ISOBase Media File Format, available from http: / / www.iso.org; ISO / IEC 14496-15:2017, Information Technology – Coding of Audio-Visual Objects, Part 15: Carriage of Network Abstraction Layer (NAL) Unit Structured Video in ISO Base Media File Format; ISO / IEC 14496-15:2017, Amendment 2, 2019-01, Information Technology—Coding of Audio-Visual Objects, Part 15: Carriage of Network Abstraction Layer (NAL) Unit Structured Video in ISO Base Media; Recommendation ITU-T H.265 / ISO / IEC 23008-2:2017 Information technology – High efficiency coding and media delivery in heterogeneous environments – Part 2: High efficiency video coding; ISO / IEC 23008-2:2017 / AMD2:2018 23008-12:2017, Information Technology – High efficiency coding and media delivery in heterogeneous environments – Part 12: Image File Format; ISO / IEC 18477-1:2020-05, Information Technology – Scalable compression and coding of continuous-tone still images – Part 1: Core coding system specification (also known as JPEG-XT specification); ISO / IEC 18477-3:2015-12-15, Information Technology – Scalable compression and coding of continuous-tone still images – Part 3: Box file format; and AV1 Bitstream & Decoding Process Specification, v1.0.0 and Errata 1, 2019-01-08, from the Alliance for Open Media, available at https: / / github.com / AOMediaCodec / av1-spec, all of which are incorporated herein by reference in their entirety.

[0069] Figure 3A The diagram illustrates an example of a relatively high-level decoding syntax / fragment structure for an image (container) file. The decoding syntax / fragment structure can represent, but is not necessarily limited to, sequential and progressive decoding syntaxes used to support image codec operations based on sequential DCT, progressive DCT, and lossless modes.

[0070] Figure 3A The decoding syntax / fragment structure of an image file includes several syntax elements or decoding fragments to support including / packaging JPEG images along with image metadata into an image (container) file. For example, these syntax elements or fragments may include APP11 tag syntax elements or fragments in the image (container) file. Some or all of these APP11 tag syntax elements or fragments can be used to carry photographic or image (payload) data of one or more images (such as one or more HEVC or AV1 images).

[0071] like Figure 3A As shown, conditional tagging syntax elements or fragments, such as restart tag number m (RSTm), can be obtained from... Figure 3A The decoding structure of the image (container) file is omitted or excluded; for example, in some operating scenarios, the (scan) restart may not be enabled during the decoding operation.

[0072] Such as Figure 3A The first subset of syntactic elements or fragments, such as the Start of Image (SOI) syntactic element or fragment and the End of Image (EOI) syntactic element or fragment, can represent the first or highest level (syntactic element or fragment) of the decoding syntax of an image (container) file.

[0073] Such as Figure 3A A second subset of syntax elements or fragments, such as APP1 or EXIF ​​markers, optional APP0 or JFIF markers, APP11 syntax elements or fragments, one or more DQT or quantization table syntax elements or fragments, and start-of-frame (SOF) syntax elements or fragments, can represent the next or second level (syntax elements or fragments) of the decoding syntax of an image (container) file. This second subset of syntax elements or fragments at the second level of the decoding syntax can be used to include or encompass one or more scans (e.g., in decoding operations), which may be preceded by one or more specific tables, such as one or more DQT tables. Figure 3A As shown, the image (container) file may exclude or have no defined line number tag (DNL) syntax elements or fragments (e.g., after defining or specifying the syntax elements or fragments for the first scan).

[0074] Such as Figure 3A A third subset of syntactic elements or fragments, such as one or more Huffman table (DHT) syntactic elements or fragments and one or more SOS (scan start, last-1) syntactic elements or fragments, can represent the third level (syntactic elements or fragments) of the decoding syntax of an image (container) file. For example... Figure 3A As shown, one or more SOS syntactic elements or fragments may be preceded by a specific table, such as one or more Huffman Table (DHT) syntactic elements or fragments.

[0075] Such as Figure 3A The fourth subset of syntactic elements or fragments such as (one or more) ECS (one or more) entropy decoded fragments, last-1) in an image (container) file can represent the fourth level (syntactic element or fragment) of the decoded syntax.

[0076] In some operational scenarios, some or all of the DAC (Define Arithmetic Regulation), DRI (Define Reboot Interval), or COM (Comment) syntax elements or fragments may be excluded or absent from the image (container) file.

[0077] like Figure 3A As shown, the image (container) file includes application fragments 0, 1, and 11 (APP0, APP1, and APP11). As previously mentioned, ISO / IEC 10918-4:1999 / Amd.1:2013(E) describes or specifies an example list of application-specific markers (APPn).

[0078] It should be pointed out that in some operational scenarios, such as Figure 3A The specific ordering shown can be used to pack or include some or all of one or more specific tables (e.g., APP1, APP0, APP11, DQT, etc.) at the beginning (or before) of a frame. This specific order can be (e.g., explicitly, by default, precisely, etc.) implemented, enforced, or followed by the image encoder and / or image decoder as the order in which these decoding syntaxes or fragments are written into / read from image (container) files, with the aim of supporting accelerated discovery of image (container) files decoded using the decoding syntax described herein, or some or all of the specific content packed or included in these files.

[0079] However, it should be noted that in some operational scenarios, Figure 3AThis specific ordering can be optional and therefore not necessarily implemented, enforced, or followed precisely by the image codec. Therefore, the image decoder or unpacker used to discover or decode the contents of an image (container) file must support or tolerate other orderings of the image (payload) data and / or image metadata carried in the image (container) file, including but not limited to other orderings of one or more specific table application-tagged fragments. Additionally, optionally, or alternatively, in some image applications, as described herein, the image (container) file may carry, include, or package only a single APP11 tagged fragment.

[0080] The decoded structure of an image (container) file, or its syntactic elements or fragments, can be signaled, identified, or indicated (from the encoder to the decoder) through specific extended tags in the image (container) file.

[0081] Support for multiple image types

[0082] As described herein, image (container) files can be used to support carrying and / or constructing multiple types of images that depict the same visual semantic content (e.g., characters, objects, motion, visual background, etc.) as the primary (or master) image specified in the image (container) file. Some or all of these carried or constructible images from the image (container) file may be derived directly or indirectly from or originate from the same source image, such as images or photographs generated by an image signal processor of a camera with specific settings.

[0083] For illustrative purposes, an image (container) file can be a JPEG image file, which includes a main (or master) image file component, such as Figure 3A One or more ECS fragments may be used to carry, store, or contain (decoded) JPEG images. Various types of images, including but not limited to JPEG images, may be carried or constructed based on image data and / or image metadata included in the JPEG image file, according to the techniques described herein.

[0084] In the first example, in addition to the primary (or master) image file component carrying the JPEG image, the image (container) file also includes one or more accompanying data components, such as one or more APP11s (e.g., such as...) used to carry, store, or contain one or more HEVC images, and / or one or more AV1 / HDR still images, and / or image metadata (“rpu”), etc. Figure 3ASyntactic elements or fragments (as shown in the image). Image metadata can be applied to some or all of the HEVC or AV1 / HDR still images to generate corresponding (e.g., SDR, HDR, etc.) display images. These corresponding display images generated from images included in an image (container) file can be specifically optimized for a particular image display using specific image processing operations and / or specific operation parameters defined or included in the image metadata of the image (container) file to make relatively full use of the corresponding capabilities of those image displays.

[0085] Image (container) files may include things like APP1 (e.g., ... Figure 3A The syntactic elements or fragments such as EXIF ​​shown in the image (container) file describe specific characteristics of the JPEG image. The location depth, chromaticity format, and color space of the images already included in the image (container) file and / or additional images that can be synthesized or generated at least in part based on the included images and the included image metadata can be defined or specified in the image metadata carried or stored in the image (container) file.

[0086] In the second example, in addition to the main (or master) image file component carrying the JPEG images, the image (container) file also includes one or more accompanying data components, such as a first image metadata portion used to carry, store, or contain (overall) image metadata (“rpu”), zero or one or more HEVC images, and / or zero or one or more AV1 / HDR still images, and / or a second image metadata portion, etc., of which there are one or more APP11 (e.g., such as...). Figure 3A The first image metadata portion can be applied to a JPEG image to generate a first corresponding (e.g., SDR, HDR, etc.) display image. However, if HEVC or AV1 / HDR still images are present in the image (container) file, the second image metadata portion can be applied to each of some or all of the HEVC or AV1 / HDR still images to generate a second (e.g., SDR, HDR, etc.) display image. These display images can be specifically optimized for a particular image display using specific corresponding image processing operations and / or specific corresponding operation parameters defined or included in the corresponding image metadata portion of the image data in the image (container) file to make relatively full use of the corresponding capabilities of those image displays.

[0087] Similar to the first example, in the second example, the image (container) file may include things like APP1 (e.g., ...). Figure 3AThe syntactic elements or fragments such as EXIF ​​shown in the image (container) file are used to describe specific characteristics of a JPEG image. The location depth, chroma format, and color space of the images already included in the image (container) file and / or additional images that can be synthesized or generated at least in part based on the included images and the included image metadata can be defined or specified in the image metadata carried or stored in the image (container) file.

[0088] As used herein, the term "HEVC image" can refer to an HEVC Main10 still image as defined in Recommendation ITU-TH.265 / ISO / IEC 23008-2 cited herein. In some operational scenarios, HEVC images include or carry HEVC VUI parameters. HEVC images may conform to or be consistent with their HEVC VUIs (e.g., settings, values, etc.), which include, but are not limited to, some or all of the following: video range, primary colors, transfer characteristics, color matrix, and chroma sampling location.

[0089] The terms "AV1 image" or "AV1 / HDR image" can refer to an AV1 Main or High still image as defined in the AV1 Bitstream & Decoding Process Specification cited herein. In some operational scenarios, AV1 / HDR images include or carry AV1 color description parameters. AV1 / HDR images may conform to or be consistent with their AV1 parameters (e.g., settings, values, etc.), which include, but are not limited to, some or all of the following: video range, primary colors, transfer characteristics, color matrix, and chroma sampling position.

[0090] For illustrative purposes only, it has been described that image (container) files such as JPEG image files can carry one or more HEVC and / or AV1 / HDR images, as well as JPEG images. It should be noted that in other operational scenarios, other types of images may be carried along with JPEG images as a supplement to or alternative to one or more HEVC and / or AV1 / HDR images. For example, in some operational scenarios, zero or more HEIC images may be carried along with JPEG images as a supplement to or alternative to one or more HEVC and / or AV1 / HDR images. The term "HEIC image" can refer to an efficient image file format image that conforms to a specific specification related to HEVC (such as the specification defined in ISO / IEC 23008-12:2017, the contents of which are incorporated herein by reference in their entirety). In some operational scenarios, HEIC images include or carry HEVC VUI parameters. HEIC images may conform to or be consistent with their HEVC VUI, which includes, but is not limited to, some or all of the following: video range, primary colors, transfer characteristics, color matrix, and chroma sampling position.

[0091] Encoding operations

[0092] As described herein, an image (container) file may be specifically encoded by an image encoder to include image (payload) data and / or image metadata, with the aim of distributing the visual semantic content of an original image or photograph through multiple types of images, which may be decoded or constructed by an image decoder at the receiving end of the image (container) file from the image (payload) data and / or image metadata encoded in the image (container) file.

[0093] Image encoders can encode image (container) files with image (payload) data and / or image metadata according to one or more standards-based or proprietary specifications, including but not limited to some or all of the following: ISO / IEC 10918-1:1994; ISO / IEC 18477-1:2020-05; ISO / IEC 10918-4:1999; ISO / IEC 10918-5:2013; ISO / IEC 14496-12:2015; JEITA CP-345E / CIPA DC-008-2019 (EXIF), etc.

[0094] Not all parameters and fields (or their values) defined in the (applicable) specifications are actually used to encode the image (container) file. A subset of the parameters / fields (or values) defined in these specifications may be excluded or restricted from the image (container) file. Some or all of such exclusions or restrictions on this subset of parameters / fields in the image (container) file may be explicitly signaled or indicated by the image encoder to the receiving image decoder in the image (container) file.

[0095] In some operational scenarios, the image (container) file includes or is encoded with at least one APP11 tag, which has a specific identifier, such as an identifier string with a specific string value such as 'DI', to indicate that the image (container) file contains image (payload) data and / or image metadata to be used to construct or derive other images and / or different display images.

[0096] Image (container) files encoded by an image encoder may correspond to a specific version of an applicable specification (e.g., format, packaging, "0", "1", etc.) that specifies or defines what image (payload) data and / or image metadata are encoded in the image (container) file. For example, an applicable specification for encoding an image (container) file may specify or define that, in addition to a JPEG image serving as the primary (or master) image of the image (container) file, the image (container) file (such as a JPEG image file) also carries or encodes zero or one or more HEVC Main10 still images and / or one or more AV1 Main or High encoded images. The applicable specification may be specifically enhanced from one or more other standards-based specifications or proprietary specifications to include, specify, or define specific encoding operations and syntax for distributing images / photographs of multiple image types using a single image file or a single image container file.

[0097] For each of the HEVC or AV1 / HDR (compressed) still images carried or included in, or constructed from, an image (container) file, HDR characteristics (e.g., bit depth, chroma format, color space, etc.) may be defined or specified in the image metadata of the image (container) file according to applicable (decoding or decoding syntax) specifications.

[0098] In this example, the JPEG image will be encoded or carried in the image (container) file as the primary (or master) image. The image encoder may include or invoke a primary image codec (such as the JPEG image codec) to perform the JPEG image encoding operation.

[0099] An image encoder or JPEG image codec can perform image compression operations to generate JPEG compressed image data from a received image (e.g., input, source, raw, uncompressed, less compressed, etc.). JPEG compressed image data represents a JPEG image and can be encoded as (payload) data or included in the main (or master) image data component of an image (container) file. A JPEG image can include multiple pixels or pixel values ​​at multiple pixel locations (e.g., within an image frame, in a spatial array such as a two-dimensional pixel array, etc.).

[0100] In some operational scenarios, JPEG images can be represented in the YCbCr color space, which has three color components: Y, Cb, and Cr. Each pixel value can contain the lightness (Y) and chrominance (Cb / Cr) component pixel values ​​of the corresponding pixel or pixel location. Each (Y, Cb, or Cr) component pixel value can have a specific positioning depth, such as 8-bit depth. JPEG images can be sampled or subsampled to chrominance sampling (or subsampling) format 4:2:0 based on a centered chrominance position. This chrominance position is based on the TIFF default value used in many personal computer applications.

[0101] It should be noted that in other operating scenarios, chroma component values ​​can be sampled or subsampled using a chroma sampling (or subsampling) format other than 4:2:0 and / or a chroma position sampling other than centered chroma position sampling. Additionally, optionally, or alternatively, a color space other than YcbCr with different color components can be used to represent the main image in the image (container) file as described herein. Additionally, optionally, or alternatively, component pixel values ​​can be a different bit depth than 8 bits.

[0102] In this example, the full (codeword) value range corresponding to all possible values ​​of the bit depth (e.g., all possible 8-bit values ​​in this example, etc.) can be used to encode or represent component (Y, Cb, or Cr) pixel values ​​in the primary (or master) image as described herein. For example, the luminance (Y) component value can be represented in the full range [0, 255] from a reference black (e.g., 0, etc.) to a reference white (e.g., 255, etc.). The chrominance (Cb or Cr) component value can be represented in [0, 255], where 128 is the reference minimum color (e.g., gray value, etc.) and 0 / 255 is the reference maximum color (one or more).

[0103] In some operational scenarios, a JPEG image encoded in an image (container) file is derived from a raw (e.g., received, input, source, etc.) JPEG image, which includes a raw or input APP1 (EXIF) application mark fragment 1. The JPEG image encoded in the image (container) file includes an APP1 (EXIF) application mark fragment 1 that corresponds to (e.g., an equivalent, consistent, appropriate, etc.) of the raw or input APP1 (EXIF) application mark fragment 1 of the raw JPEG image, or is contained within the same image (container) file along with that APP1 (EXIF) application mark fragment 1.

[0104] The JPEG image encoded in the image (container) file may also include (to be) appropriately used and consistent with the JPEG image encoded in the image (container) file, or be included in the same image (container) file together with the APP0 (JFIF) application mark fragment 1.

[0105] As noted herein, an image (container) file may include image (payload) data representing additional images besides the primary (or master) image, and image metadata that may be used with some or all of the image (payload) data to construct or derive other images, including images of a different type from the primary (or master) image and / or different display images optimized for rendering on different types of image displays.

[0106] An image encoder may include or invoke one or more non-primary image codecs, such as (one or more) HEVC and / or AV1 / HDR (and / or HEIC) image codecs, to perform HEVC and / or AV1 / HDR image encoding operations.

[0107] An image encoder or non-primary image codec may perform image compression operations to generate HEVC and / or AV1 / HDR (and / or HEIC) compressed image data corresponding to a JPEG image in the primary data component of an image (container) file and / or corresponding to a received (e.g., input, source, raw, uncompressed, less compressed, etc.) image used to generate the JPEG file. The HEVC and / or AV1 / HDR (and / or HEIC) compressed image data represents one or more HEVC and / or AV1 / HDR (and / or HEIC) images and may also be encoded as payload data or included in other image data components (such as one or more APP11 data fragments) of the image (container) file, in addition to or separately from the primary (or master) image data component. Each of the one or more HEVC and / or AV1 / HDR (and / or HEIC) images may include multiple pixels or multiple pixel values ​​at pixel locations (e.g., in an image frame, in a spatial array such as a two-dimensional pixel array, etc.).

[0108] Each of the one or more HEVC and / or AV1 / HDR (and / or HEIC) images (still images) represented or included in one or more APP11 fragments of an image (container) file may be encoded or decoded using the same type of image codec.

[0109] In some operational scenarios, applicable specifications may specify or define that some or all of the still images (as extended image data) included in one or more APP11 segments of an image (container) file are decoded using the same type of image codec. In these operational scenarios, the image (payload) data includes non-primary (or extended) images, one or more HEVC images, or one or more AV1 images, within one or more APP11 segments of the image (container) file. In other words, these one or more APP11 segments may not carry a mixture of extended image data (one or more HEVC images and one or more AV1 images).

[0110] Image data (e.g., compressed image data, image payload, etc.) can be encoded in a YCbCr 4:2:0 image, with each color component or primary color (Y, Cb, or Cr) being 10 bits.

[0111] The color components or primary colors representing the color space of the image data (e.g., YCbCr, etc.) can be defined or specified in one or more applicable image data decoding specifications such as BT.2100-2, and can be notified in the image container file by signaling, for example as follows: HEVC VUI or AV1 parameter = '9'.

[0112] Similarly, the transfer characteristics of image data can be specified or defined as PQ or HLG according to the applicable image data decoding specification, and can be notified by signaling in the image container file, for example: for PQ, HEVC VUI or AV1 parameter = '16', or for HLG, HEVC VUI or AV1 parameter = '18'. In some operational scenarios, all non-primary images (such as HEVC or AV1 still images) included in the image container file use one and only one type of transfer characteristic: PQ or HLG.

[0113] The color matrix associated with the color space (e.g., YCbCr, etc.) can be defined or specified in one or more applicable image data decoding specifications such as BT.2020-2 (NCL) and BT.2100-1, and can be notified in the image container file by signaling, for example as follows: HEVC VUI or AV1 parameter = '9'.

[0114] Top-left chroma sampling (which can be signaled in the image container file, for example: HEVC VUI or AV1 parameter = '2') can be used for the 4:2:0 chroma sampling format.

[0115] A finite range of codewords (which can be signaled in the image container file, for example: HEVC VUI or AV1 parameter = '0') can be used to decode codewords in the color components of a color space (e.g., YCbCr, etc.). For example, a first finite codeword range of 64–940 can be used to represent Y codewords from reference black to reference white. A second finite codeword range of 64–940 can be used to represent each of the Cb or Cr codewords from the reference minimum color to the reference maximum color, where 512 is a gray value.

[0116] Encoded / compressed image data (e.g., HEVC compressed image data, etc.) in an image container file can be carried as Network Abstraction Layer (NAL) unit data and written / stored in the image container file (or image signal) as an H.265 Annex B byte stream format.

[0117] Encoding extended image data

[0118] In addition to the primary (or main) image, image data of one or more non-primary (or extended) images can also be cascaded within an image (container) file. Figure 3AThe APP11 fragments shown are formed from image (payload) data carried within the decoding syntax. According to the applicable image decoding (syntactic) specification, these APP11 fragments may be specified or include specific fragment markers, such as the string value 'DI'. The applicable image decoding (syntactic) specification may specify or define application-marked fragments, such as APP11 for 'tables / other items'. APP11 fragments or syntactic elements may be located at the beginning of a frame or at the beginning of a scan within a frame.

[0119] Figure 3B The diagram illustrates an example syntax element or decoding structure of an APP11 tag fragment. As shown, an APP11 syntax element includes multiple parameters, which are decoded separately in multiple component syntax elements as part of the overall APP11 syntax element. These parameters or component syntax elements (representing different data fields in the APP11 tag fragment) can be ordered and sequentially (unless otherwise explicitly indicated, there are no empty bytes (one or more) and / or padding bytes (one or more) between successive parameters or syntax elements), such as... Figure 3B As shown in the image.

[0120] In some operational scenarios, such as Figure 3B As shown, multiple parameters in an APP 11 tag fragment or syntax element may include: an APP11 tag parameter (2 bytes); a length parameter (2 bytes); an ID string parameter (2 bytes); followed by a null byte (1 byte); then a photo / image (payload) data syntax element; and so on. The photo / image (payload) data syntax element includes the very first byte of the entire photo / image (payload) data syntax element (e.g., the very first (lower level, etc.) syntax element), carrying a format / packaging version number parameter (1 byte).

[0121] If the format / packaging version number (version #) parameter in the photo / image (payload) data syntax element is set to a specific value, such as one (1), then the format / packaging version number parameter is followed by a syntax element for the specific payload version (1 in this example). Depending on the applicable specification, all syntax elements constituting the entire photo / image (payload) data syntax element (e.g., components, lower levels, etc.) can be consistent with or identified by the indicated format / packaging version number parameter.

[0122] For example, rather than a limitation, such as Figure 3B As shown, the two-byte (or 16-bit) APP11 tag parameter can be used to carry or specify a specific tag value, such as 0xFFEB, to identify that this fragment represents or carries (e.g., Dolby image, canonical definition, etc.) image (payload) data tag fragment or syntax element.

[0123] The two-byte (or 16-bit) APP11 length parameter following the APP11 tag parameter (e.g., immediately after, as defined by the applicable specification, etc.) can be used to carry a value indicating (e.g., the entire, minus two bytes, current, etc.) the length of the APP11 tag fragment. The length of the APP11 tag fragment can include the size of multiple parameters (including any (one or more) intermediate null bytes), as well as the size of the photo / image (payload) data syntax element that is separately included in the current APP11 tag fragment, and can exclude the two bytes of the APP11 tag itself (0xFFEB in this example).

[0124] A two-byte (or 16-bit) ID string parameter can be used to carry a special or specifically designated value x4449 (corresponding to ASCII: 'D' 'I') to distinguish the current APP11 tag fragment from any other APP11 tag fragment(s) used for purposes other than carrying photo / image (payload) data specified or defined by the applicable specification. In response to determining that the (decoded) value of the two bytes in the APP11 tag fragment corresponding to this parameter is different from or does not match the special or specifically designated value, the receiving decoding device may ignore or avoid using that APP11 tag fragment to retrieve photo or image (payload) data.

[0125] The first byte of the photo / image (payload) data syntax element, namely the format / packaging version number parameter, can be used to define or specify a particular payload version (e.g., "1", etc.) syntax element that follows the format / packaging version number parameter in the photo / image (payload) data syntax element. The specific payload version syntax element can be used to store or carry the actual photo / image (payload) data of non-primary or extended images (such as HEVC or AV1 images).

[0126] In some operational scenarios, all APP11 tagged fragments encoded in the same image (container) file as any non-primary or extended photo / image (payload) data should carry the same value for the specified format / packaging version number parameter for each of these APP11 tagged fragments.

[0127] In this example, the format / packaging version number parameter in the APP11 tag fragment or syntax element has a specific value '1'. Accordingly, a specific payload version syntax element (such as a payload version 1 syntax element) is used to encode or store the actual non-primary or extended photo / image (payload) data.

[0128] In some operational scenarios, payload version 1 syntax elements logically include one or more (data) boxes decoded within the corresponding syntax elements. These boxes, or the corresponding syntax elements, describe (e.g., non-primary or extended, HEVC, AV1, etc.) the size and location of image data and / or image metadata (“rpu”) used for primary and / or non-primary image data. Image metadata may include specific image metadata items or portions associated with or corresponding to specific image items or portions in non-primary or extended image data (such as HEVC or AV1 still image data) and primary or master image data (such as JPEG image data).

[0129] As used herein, an image metadata item or portion associated with or corresponding to an image data item or portion refers to a part of the image metadata that is specifically designated to carry or include specific operational parameters for specific image processing operations that can be performed on the (associated or corresponding) image data item or portion.

[0130] Figure 3C The illustration shows an example decoding syntax / structure of a (data) box or corresponding syntax element within one or more boxes included in a version 1 syntax element of the payload for an APP11 token fragment. As shown, the box includes multiple parameters to be decoded separately in multiple component syntax elements as part of the whole box. These parameters or component syntax elements—representing different data fields within the box—may be ordered sequentially and consecutively (unless otherwise explicitly indicated, there are no empty bytes (one or more) and / or padding bytes (one or more) between successive parameters or syntax elements), such as... Figure 3C As shown in the image.

[0131] In some operational scenarios, such as Figure 3C As shown, multiple parameters in the boxes of the APP11 tagged fragment payload version 1 syntax elements may include: a two-byte (or 16-bit) box instance number parameter, denoted as "EN"; a four-byte (or 32-bit) group sequence number parameter, denoted as "Z"; a four-byte (or 32-bit) box length parameter, denoted as "LBox"; a four-byte (or 32-bit) box type parameter, denoted as "TBox"; an optional eight-byte (or 64-bit) box length extension, denoted as "XLBox"; version 1 payload data syntax elements; and so on.

[0132] A two-byte (or 16-bit) box instance number (En) parameter can be used to allow or support (e.g., APP11, etc.) tagged fragments carrying (data) boxes of the same or identical (box) type but with different data or content portions. This parameter can be used to distinguish these boxes of the same box type. Data or content portions belonging to or residing in logically different boxes of the same box type differ in the corresponding value of the box instance number (En) parameter. The receiving decoding device can concatenate the data or content portions—or payload data—of the boxes in the tagged fragments, where these boxes have the same box type (BType) parameter value, but different (e.g., consecutive, sequential, etc.) values, such as the box instance number (En) parameter values ​​being in ascending order.

[0133] In some operational scenarios, in response to determining that the box type (BType) parameter has a specific value (such as 'RPJP') to indicate its association with primary or master image data (such as JPEG image data), the box instance number is set (by the image encoder) to equal 0x0001. This setting in the box within the image (container) file indicates or signals to the receiving image decoder (e.g., in a single box used for JPEG) that the image metadata portion of the box is associated with or corresponds to the primary or master image data (such as JPEG image data).

[0134] It should be noted that in some operational scenarios, any APP11-tagged fragment in an image (container) file may not contain any boxes to carry or include the image metadata portion of a box associated with or corresponding to the primary or master image data (e.g., in a single box used for JPEG, etc.).

[0135] A box can be used to carry or include non-primary or extended image data items or portions, such as HEVC or AV1 image data items or portions. This box can carry a specific value (e.g., 4 bytes, a string, etc.) for the box type (BType) parameter within the box, such as 'HEVC' or 'AV01'.

[0136] Another box can be used to carry or include image metadata items or portions for non-primary or extended image data items or portions included or carried in the previous box. This other box can carry a specific (e.g., 4 bytes, string, etc.) corresponding value for the box type (BType) parameter in this other box, such as 'RPHE' or 'RPAV'—corresponding to a specific value of 'HEVC' or 'AV01' in the (previous) box.

[0137] The two boxes (the former carrying an image data item or portion, and the latter carrying the corresponding image metadata item or portion to be used for processing that image data item or portion) should each carry the same or identical value for the box instance number (En) parameter (e.g., "01", etc.).

[0138] In the first example, the first box can be an 'HEVC' box that carries or includes HEVC image data items / parts. The second box can be an 'RPHE' box that carries or includes RPHE image metadata items / parts for the HEVC image data items or parts carried or included in the first box. Both the 'HEVC' box and the 'RPHE' box can carry the same or identical value for the box instance number (En) parameter (e.g., "01", etc.) in each box.

[0139] In the second example, the first box can be the 'AV01' box, which carries or includes AV1 image data items / parts. The second box can be the 'RPAV' box, which carries or includes RPAV image metadata items / parts for the AV1 image data items or parts carried or included in the first box. Both the 'AV01' box and the 'RPAV' box can carry the same or identical value for the box instance number (En) parameter (e.g., "01", etc.) in each box.

[0140] In some operational scenarios, an image (container) file contains multiple boxes of the same box type (BType). The receiving decoding device of the image (container) file uses the box instance number parameter (or data field) as an instruction or reference provided by the encoder to the receiving decoding device to sort and merge the image data or metadata items / parts carried or included in these multiple boxes of the same box type with the payload version 1 data syntax elements into a single whole box or a single whole image data or metadata item / part.

[0141] The four-byte (or 32-bit) group sequence number (Z) parameter can be set by the image encoder in each of the groups (e.g., all groups, etc.) used to transmit or transfer the box. This parameter can be used to specify a specific order—e.g., ascending order of the group sequence number parameter values ​​within a group—in which the payload data of the group of boxes (e.g., all payload data, etc.) will be merged into the overall payload data of the box. The concatenation of payload data can be performed in a specific order. The value of the group sequence number parameter for the first group among all groups of a box of a specific box type (e.g., a given instance, etc.) can be set to 0x0001 or 0x00000001.

[0142] The 4-byte (or 32-bit) box length (LBox) parameter can be used to specify the length of a box. The value of the box length (LBox) parameter can be measured or set to the sum of: (a) the combined size of all payload data carried or encoded by payload version 1 data syntax elements across all boxes of the same box type (which can be set to a specific value in the enumerator); (b) the size of a single copy / instance of the box type (BType) parameter (4 bytes in this example); (c) the size of a single copy / instance of the box length (LBox) parameter (4 bytes in this example); and (d) the length of a single copy / instance of the box length extension (XLBox; optional) parameter, if present (8 bytes in this example). The value of the box length (LBox) parameter can exclude the group sequence number (Z) parameter, the box instance number (En) parameter, the format / packaging version number parameter, null bytes, the ID string parameter, and the (APP11 marker) length parameter (e.g., ...). Figure 3B (as shown in the image) or the size of the APP11 marker parameter.

[0143] In the first example, the box has 32 bytes of payload version 1 data without using the box length extension parameter. This box has a value of 32 (payload version 1 data) + 4 (BType) + 4 (LBox) = 40 bytes for the box length (LBox) parameter. If the box is evenly split across two APP11 tag fragments, then each tag fragment has a value of 2 (APP11 tag) + 2 ((APP11 tag) length) + 2 (ID string) + 1 (null byte) + 1 (format / packaging version number) + 2 (En) + 4 (Z) + (4 (LBox) + 4 (TBox) + 16 (version 1 payload data; half of the 32 bytes of pre-split payload version 1 data in the box)) = 38 bytes for the APP11 tag fragment length.

[0144] If the length of the box is greater than, for example, (2) 32 - 1) - If the box length threshold is 7 bytes, then the box length extension (XLBox) parameter can be specified to indicate that the box is extended accordingly. Otherwise, the box length extension (XLBox) parameter can be omitted from the Box syntax element; therefore, the XLBox size is zero (0). Example box extensions can be found in the previously mentioned ISO / IEC 18477-3:2015-12-15.

[0145] The four-byte (or 32-bit) box type (TBox) parameter can be used to specify the specific type and associated context of the payload data carried in the box. Table 1 below illustrates example box types, their respective values, ASCII encodings, and constraints.

[0146] Table 1

[0147]

[0148] In addition to or as an alternative to the box types shown in Table 1, additional box types may be used, for example, to specify or define additional image metadata about the image carried in or to be constructed from the image (container) file. The receiving image decoding device may ignore box types that the receiving decoding device does not understand, or perform a no-op on them.

[0149] Example descriptions of box type "LCHK" can be found in ISO / IEC 18477-3:2015-12-15. In response to detecting or determining that the checksum calculated on the received data of a box undergoing this checksum operation or constraint differs from the checksum recorded in the box, the image decoding device may abort the decoding operation while providing the user with information such as an error message. Additionally, optionally, or alternatively, the image decoding device may decode only the primary JPEG image and reject non-primary or extended images, such as HEVC or AV1 still images. Additionally, optionally, or alternatively, the image decoding device may decide to attempt full decoding even if one or more checksums fail.

[0150] The version 1 payload data syntax element in the box carried as part of the photo / image (payload) data of the APP11 tagged fragment can be used to carry specific content data of the box, such as image data items / parts such as HEVC or AV1 still image data (items / parts), or image metadata items / parts.

[0151] Table 2 below provides information such as Figure 3B and Figure 3C The example sizes, allowed values, and meanings of the syntactic elements or parameters shown are provided.

[0152] Table 2

[0153]

[0154] HEVC still images

[0155] HEVC still images may be included or encoded as an H.265 bitstream in an image (container) file as described herein. The H.265 bitstream may conform to one or more applicable decoding specifications, such as the aforementioned Recommendation ITU-TH.265 / ISO / IEC 23008-2:2017.

[0156] In some operational scenarios, the parameters carried or decoded in the H.265 bitstream can be set as follows.

[0157] The “nuh_layer_id” parameter can be set to “0”. An H.265 bitstream should contain only one image—for example, an H.265 bitstream containing only one HEVC image derived from a single source image, from which the main image in the image (container) file is derived.

[0158] For decoding operations used to decode H.265 bitstreams, the "general_profile_idc" parameter can be set to "2" in the HEVC Main10 still image profile.

[0159] The "general_one_picture_only_constraint_flag" can be set to "1" in the HEVC Main 10 still image profile.

[0160] The “general_level_idc” parameter can be set to less than or equal to “183”.

[0161] In addition to the parameters described above in Recommendation ITU-T H.265 / ISO / IEC 23008-2:2017, constraints can be signaled or set in the data fields of the sequence parameter set and / or picture parameter set of the H.265 bitstream (e.g., HEVC, etc.) to inform the receiving device and enable it to perform relatively efficient decoding operations. Additionally, optionally, or alternatively, constraints can be signaled (by the receiving device) or set in the data fields of the AV1 OBU.

[0162] In some operational scenarios, one or more constraints—or corresponding parameters (e.g., video availability information or VUI) carried or decoded in the H.265 bitstream—can be set as follows.

[0163] For HEVC, the "bit_depth_luma_minus8" parameter can be set to "2". For HEVC, the "bit_depth_chroma_minus8" parameter can be set to "bit_depth_luma_minus8".

[0164] For HEVC, the "chroma_format_idc" parameter can be set to "1". For AV1, the "subsampling_x" parameter can be set to "1", and the "subsampling_y" parameter can be set to "1".

[0165] For HEVC, the "vui_parameters_present_flag" parameter can be set to "1".

[0166] For HEVC, the "video_signal_type_present_flag" parameter can be set to "1". The "video_format" parameter can be set to "0".

[0167] For HEVC, the "color_description_present_flag" parameter can be set to "1".

[0168] For HEVC, the "chroma_loc_info_present_flag" parameter can be set to "1". For HEVC, the "chroma_sample_loc_type_top_field" parameter can be set to "2" (or the top left position). For HEVC, the "chroma_sample_loc_type_bottom_field" parameter can be set to "2" (or the top left position).

[0169] For AV1, the "chroma_sampling_position" parameter can be set to "2".

[0170] For AV1, the "seq_profile" parameter can be set to "0" and the "high_bitdepth" parameter can be set to "1".

[0171] Image decoding operation

[0172] Image decoding devices can decode and / or process received image (container) files as described in accordance with one or more applicable encoding specifications. Image decoding devices can support both isotopic and center-level chroma sampling or subsampling of the chroma image data carried or encoded in the image (container) file.

[0173] Table 3 below illustrates example types of image decoding devices and their corresponding capabilities that can receive and process received image (container) files.

[0174] Table 3

[0175]

[0176] Example device configuration

[0177] Figure 6AAn example (image / photograph) capture device (600) that can be used to implement the techniques described herein is depicted. The example capture device described herein may include, but is not limited to: a camera that supports HDR and / or SDR image capture, a mobile device with one or more cameras, a head-mounted user device with a camera, a wearable user device with a camera, etc.

[0178] One or more (image) sensors (602) are used to capture or generate a raw image, which may have a bit depth of 12-16 bits. The raw image can be captured with specific camera settings for aperture, shutter speed, exposure, focal length, etc. The raw image can be processed by an ISP (604) (e.g., error correction, local and global image adjustments, etc.) to generate an ISP-post image or input image for use by a photo processing core (606), e.g., encoder side, etc. In some operational scenarios, the bit depth of the ISP-post image or input image may be 10-12 bits and includes High Dynamic Range (HDR) Perceptual Quantization (PQ) codewords represented in the BT.2100 color space. The ISP-post image or input image generated by the ISP (604) may optionally be used as a preview image (612).

[0179] A decoding block such as (e.g., encoder side, etc.) a photo processing core (606) receives the ISP post-processed or input image from the ISP (604) and interacts with the image / video codec (e.g., Figure 6A 608 and 610 (located outside the photo processing core (606), etc.) operate in combination to generate and encode images / photos in different image formats from the same input image, and package the generated / encoded images / photos into an image (container) file (614).

[0180] For example, the photo processing core (606) may generate a first (e.g., 10-bit, intermediate, reshaped, original, etc.) image (which is derived from an input image received by the photo processing core (606)) or provide it to an HEVC codec, such as HEVC encoding (608). HEVC encoding (608) receives, processes, converts, compresses the first image and / or encodes it into a corresponding encoded HEVC image.

[0181] Additionally, optionally, or alternatively, the photo processing core (606) may generate a second (e.g., 8-bit, intermediate, reshaped, etc.) image (derived from the input image received by the photo processing core (606)) or provide it to a JPG codec, such as JPG encoding (610). JPG encoding (610) receives, processes, converts, compresses the second image, and / or encodes it into a corresponding encoded JPG image.

[0182] Encoded HEVC images and encoded JPG images derived from the same input image can be sent as outputs by HEVC encoding (608) and JPG (610), and received as inputs by the photo processing core (606) and packaged (e.g., along with image metadata, etc.) into an image (container) file (614).

[0183] In some operational scenarios, such as Figure 6A Some or all of the subsystems or processing modules / blocks shown in the diagram can be implemented in the same capture device (600).

[0184] Figure 6B Alternative example (image / photograph) capture devices (600-1) that can be used to implement the techniques described herein are depicted. The capture device (600-1) may be one or more cameras that support HDR and / or SDR image capture, a mobile device with one or more cameras, a head-mounted user device with one or more cameras, a wearable user device with one or more cameras, etc.

[0185] One or more (image) sensors (602) are used to capture or generate a raw image, which may have a bit depth of 12-16 bits. The raw image can be captured using specific camera settings for aperture, shutter speed, exposure, focal length, etc. The raw image can be processed by an ISP (604) (e.g., error correction, local and global image adjustments, etc.) to generate an ISP-post image for HEVC codec or HEVC encoding (606). The ISP-post image generated by the ISP (604) can be compressed or encoded by HEVC encoding (606) to generate an encoded version of the ISP-post image (e.g., HEVC, HEIC, etc.).

[0186] In some operational scenarios, additionally, image processing operations such as image analysis (616) and / or metadata generation (e.g., information generated from image analysis (616) etc.) can be performed on the post-ISP image and / or the encoded version of the post-ISP image, and then the (final) encoded version of the post-ISP image is packaged in a relatively efficient capture device image (container) file (such as a HEIC image file).

[0187] Image metadata generated from image analysis (616) can be included in a capture device image (container) file and can be used by a downstream device receiving the capture device image (container) file along with an encoded version of the ISP-backed image to generate an optimized display image for rendering on an image display.

[0188] Figure 6C Depicts what can be used with capture devices (e.g., Figure 6BAn example image processing device (650) operates together with 600-1, etc., to implement the technology described herein. The image processing device (650) may be a mobile or non-mobile computing device, a cloud-based image processing system, an image / photo processing and / or storage service, a cloud-based photo library, etc.

[0189] A decoding block such as (e.g., encoder side, etc.) a photo processing core (606) receives a capture device image (container) file containing an encoded version of the ISP-post image, such as a HEIC image file. The capture device image (container) file may also include image metadata. The photo processing core (606) can extract the image metadata and the encoded (e.g., HEVC, HEIC, etc.) version of the ISP-post image from the capture device image (container) file.

[0190] The photo processing core (606) can work with image / video codecs (e.g., Figure 6C 618 and 610; located outside the photo processing core (606), etc.) operate in combination to combine the same ISP to generate and decode / encode images / photos of different image formats and package the generated / encoded images / photos into an image (container) file (614).

[0191] For example, the photo processing core (606) can operate in conjunction with or invoke an HEVC codec (such as HEVC decoder (618)) to generate a decoded version of the ISP-post image. HEVC decoder (618) receives, processes, converts, decompresses, and / or decodes the encoded version of the ISP-post image in a captured device image (container) file into a corresponding decoded version of the ISP-post image.

[0192] The photo processing core (606) can generate (e.g., 8-bit, intermediate, reshaped, etc.) an image—which is derived from a decoded version of the image after ISP, or provided to a JPG codec, such as JPG Encoder (610). JPG Encoder (610) receives, processes, converts, compresses, and / or encodes the image into a corresponding encoded JPG image.

[0193] Encoded versions of ISP-encoded images (e.g., HEVC, HEIC, etc.) received in a capture device image (container) file and encoded JPG images derived from JPG encoding (610) can be packaged (e.g., along with image metadata) into an image (container) file (614) by the photo processing core (606). The image metadata in the image (container) file (614) may include portions of image metadata generated or received based on image metadata extracted from the capture device image (container) file—e.g., portions related to the HEVC or HEIC encoding version. The image metadata in the image (container) file (614) may also include other portions of image metadata generated by the photo processing core (606)—e.g., portions related to the JPG image.

[0194] Figure 7 An example (image / photograph) receiving and decoding device (700) that can be used to implement the technology described herein is depicted. The decoding device (700) may be an image display that supports HDR and / or SDR image rendering, a mobile or non-mobile computing device with one or more image displays, a television, etc.

[0195] The photo processing core (606) in the decoding device (700) (e.g., the decoder side, etc.) can receive an image (container) file containing two or more images of different image formats (such as HEVC images and JPG images, etc.). For example, and not limitingly, the photo processing core (606) operates in conjunction with image / video codecs (such as JPG codecs, HEVC codecs, AV1 codecs, etc.). By way of illustration and not limitation, image / video codec refers to the HEVC codec, such as HEVC decoder (618). The photo processing core (606) invokes HEVC decoder (618) to decode an HEVC image, which may be an HEVC-encoded version of an ISP-encoded image captured by a camera device.

[0196] The decoding device (700) may include or operate with an image display (such as a display panel). Panel configuration (704)—which may specify display capabilities such as dynamic range, color gamut, and image refresh rate—may be generated or provided to the photo processing core (606).

[0197] Additionally, optionally, or alternatively, image metadata may be extracted by the photo processing core (606) from the received image (container) file (614).

[0198] Based at least in part on panel configuration (704) and / or image metadata, the photo processing core (606) generates a display image from an HEVC image. For illustrative purposes only, the display image represents an RGB image. The display image may be transmitted or provided to an image display or display panel for rendering via an RGB image buffer (702).

[0199] Figure 6D An example (e.g., encoder side, etc.) photo processing core (or subsystem) 606 is depicted, which may be included in or implemented in an image processing device to implement the techniques described herein. The image processing device including or implementing the photo processing core (606) may be a mobile or non-mobile computing device, a cloud-based image processing system, an image / photo processing and / or storage service, a cloud-based photo library, a capture device, a separate device operating with the capture device, a decoding device, a rendering device, etc. In various operational scenarios, the photo processing core (606) may include or implement more or fewer components and / or processing blocks.

[0200] For example, and not limited to, the photo processing core (620) receives an input image (620), which may (but is not necessarily limited to) come from an ISP, a photo image file generated by a camera, etc. In some operational scenarios, the photo processing core (620) may convert or reshape the first-depth (e.g., 10-12 bits, etc.) input image (620) into an intermediate image of second-depth (e.g., 10 bits, etc.).

[0201] The photo processing core (606) operates together with the HEVC encoder (608) and invokes the HEVC encoder to generate an encoded HEVC image (or an HEVC encoded version of the input image or a converted / reconstructed image, etc.).

[0202] The photo processing core (606) includes a color volume mapper (624), which can be used to generate (e.g., color volume mapped, gamut mapped, color space mapped, reshaped, etc.) mapped images from an input image (620).

[0203] The photo processing core (606) operates together with the JPG encoding (610) and invokes the JPG encoding to generate an encoded JPG image (or a JPG encoded version of a mapped image, etc.).

[0204] Additionally, optionally, or alternatively, the photo processing core (606) includes a metadata analysis processing block (622), which can be used to analyze some or all of the input image, intermediate images generated from the input image, reshaped, transformed, or mapped, encoded images, etc., to generate image metadata. This image metadata may include operating parameters that the receiving decoding device can use to generate or reconstruct images optimized for rendering on various devices or image displays with different system and / or display capabilities.

[0205] Some or all of the encoded images (e.g., HEVC, JPG, etc.) and / or image metadata can be combined or decoded into an image (container) file (614) by a packing processing block (626).

[0206] like Figure 8 As shown, the packing processing block (626) (also known as the "photo packer") can take EXIF ​​(Exchangeable Image File Format) metadata, color volume mapping metadata, HEVC-encoded images, JPG-encoded images, etc., as input. The EXIF ​​metadata and JPG-encoded images can be used by the packing processing block (626) to populate or generate syntax elements corresponding to the JPG header and EXIF ​​in the JPG image (container) file. Additionally, optionally, or alternatively, some or all of the (input) EXIF ​​metadata, color volume mapping metadata, HEVC-encoded images, etc., can be included by the packing processing block (626) as one or more photo payloads to populate or generate syntax elements corresponding to one or more APP11 tag fragments in the same JPG image (container) file.

[0207] The photo processing (or processing) core described herein can be implemented using GPUs, SOCs, DSPs, ASICs, FPGAs, CPUs, or other computing resources for full-quality and / or degraded images. Degraded images (such as gallery images or thumbnails) can be generated in a "reduced computation" mode with relatively low computational resource usage. For example, for gallery or thumbnail views, the primary image (e.g., the JPEG base layer, etc.) can be rendered or displayed without additional processing.

[0208] like Figures 6A to 6C and Figure 7As shown, the photo processing core can operate in conjunction with or rely on an external codec to manipulate the input / output. Photo / image processing operations associated with the photo processing core can be broken down or divided into multiple processing blocks. These processing blocks can implement methods for generating the "most efficient" package (e.g., an image container file with relatively few different image formats and / or little or no image metadata, etc.) and / or the "most compatible" package (e.g., an image container file with more different image formats and / or a larger amount of image metadata, etc.) for a given input image.

[0209] like Figure 6D (and Figure 8 As shown in the diagram, the photo packing operation can be reduced to a single "packing" processing block or performed by a single "packing" processing block.

[0210] Metadata compression

[0211] In some operational scenarios, the image metadata carried or included in the image (container) file 122, such as that related to image reshaping or prediction mapping, can be compressed by the upstream encoding device to reduce the size of the image metadata transmitted from the encoding device to the downstream receiver (decoding) device.

[0212] Image (container) files 122 or bitstreams may carry configuration, profile, and / or (one or more) level parameters whose values ​​may indicate different schemes or types of metadata compression to be supported by downstream devices (such as playback devices). For example, configuration, profile, and / or (one or more) level parameters may be set to (one or more) first specific values ​​to indicate limited metadata compression—or a first metadata compression scheme or type—using a relatively small buffer (e.g., metadata, etc.) in the downstream device. Additionally, optionally, or alternatively, configuration, profile, and / or (one or more) level parameters may be set to (one or more) second specific values ​​to indicate extended metadata compression—or a second metadata compression scheme or type—using a relatively large buffer (e.g., metadata, etc.) in the downstream device. Additionally, optionally, or alternatively, configuration, profile, and / or (one or more) level parameters may be set to (one or more) third specific values ​​to indicate that upstream and / or downstream devices do not perform or support metadata compression.

[0213] In some operational scenarios, finite metadata compression schemes or types allow for at most one (1) reference buffer for the full synthesizer (or image remodeling) coefficient payload and at most one (1) reference buffer for the DM coefficients. In contrast, extended metadata compression schemes or types allow image remodeling and / or DM operations to use more buffers.

[0214] For example, rather than limiting it, image metadata compression can be performed on the portion of the image metadata containing parameters for prediction or reshaping operations.

[0215] In some operational scenarios, the first flag decoded in the image (container) file (122) (e.g., denoted as "use_prev_di_rpu_flag", a specific syntax element, etc.) can be set by the upstream device to a specific value (e.g., use_prev_di_rpu_flag = 1, etc.) to signal to the downstream device or indicate that the image metadata portion (e.g., a specific image metadata type, a type identified by another syntax element (such as di_rpu_type = 8, etc.) of the previously sent / decoded image (e.g., in the same image (container) file 122, etc.) can be reused by the downstream device for the current image. Therefore, some or all of the image processing operations (such as image reshaping or prediction operations) that the downstream device wants to perform on or for the current image can share or utilize the same previously sent image metadata portion, thereby reducing or omitting some or all of the separate (e.g., reshaping, prediction, synthesizer, etc.) image metadata portions that the upstream device explicitly or specifically decodes or transmits to the downstream device in the image (container) file 122 for the current image.

[0216] On the other hand, the first flag (“use_prev_di_rpu_flag”) decoded in the image (container) file (122) can be set by the upstream device to a second specific value (e.g., use_prev_di_rpu_flag = 0, etc.) to signal to the downstream device that the image metadata portion for the current image (e.g., a specific image metadata type, a type identified by another syntactic element (such as di_rpu_type = 8, etc.) is explicitly or specifically sent, transmitted, included, or present in the image (container) file 122. Therefore, the downstream device can use the image metadata portion explicitly or specifically decoded / transmitted by the upstream device for the current image in the image (container) file 122 and decoded or retrieved by the downstream device from the image (container) file 122 for the current image to perform some or all image processing operations, such as image reshaping or prediction operations, on the current image or for the current image. In some operational scenarios, an image or picture can be indicated or carried as a keyframe in the image (container) file 122, for which the value of the first flag ("use_prev_di_rpu_flag") is set to zero (0).

[0217] The second flag decoded in the image (container) file (122) (e.g., denoted as "prev_di_rpu_id", second specific syntax element, etc.) can be set by the upstream device to a specific value (e.g., a specific value in the range of 0 to 15, etc.) to signal to the downstream device or indicate which specific part of the previously sent image metadata (e.g., for a previous image, etc.) will be used to perform image processing, reshaping, or prediction operations on the current image or for the current image.

[0218] Image metadata previously sent by an upstream device and received by a downstream device may include multiple (previously sent) image metadata portions held or stored in the memory or cache at the downstream device. Each of the multiple image metadata portions at the downstream device may be labeled or identified using a corresponding specific metadata portion identifier among multiple metadata portion identifiers used for the multiple (previously sent) image metadata portions. The corresponding specific metadata portion identifier may be within a specific value range such as 0 to 15.

[0219] A second flag (“prev_di_rpu_id”) with a specific value within a specific range can be used by an upstream device to signal to a downstream device or to indicate that a specific (previously sent) portion of image metadata from among multiple image metadata portions held or cached at the downstream device will be used by the downstream device to perform image processing, reshaping, or prediction on the current image or for the current image. This specific image metadata portion may have been previously sent for a previous image, where a third flag set for that previous image (represented as “di_rpu_id”, a specific syntax element, etc.) has the same metadata portion identifier value as the second flag set for the current image (“prev_di_rpu_id”, etc.).

[0220] Additionally, optionally, or alternatively, in response to determining that the second flag (“prev_di_rpu_id”) is absent in the current image in image (container) file 122, the downstream device may continue to internally set the second flag to a value outside a specific range, such as -1 (e.g., an invalid rpu_id). This helps to avoid or prevent the downstream device from subsequently performing image processing, reshaping, or prediction on or for the current image using incorrect (previously sent) image metadata portions.

[0221] The second flag (“prev_di_rpu_id”) decoded in the image (container) file (122) can be set by the upstream device to a second specific value (e.g., prev_di_rpu_id = 0, etc.) to signal or instruct the downstream device that a specific portion of image metadata, such as image reshaping or prediction for the current image (e.g., inter-layer, etc.), is explicitly or specifically sent from the upstream device to the downstream device, included in or present in the image (container) file 122. Therefore, some or all image processing operations—which may include, but are not necessarily limited to, image reshaping or prediction operations—can be performed on or for the current image using the (currently sent) specific portion of image metadata. Some or all of the operation parameters in the (currently sent) specific portion of image metadata can be specified using (currently sent) data structures (e.g., denoted as “di_rpu_data_mapping()”, a specific set of syntax elements, etc.) explicitly or specifically decoded for the current image in the image (container) file 122.

[0222] Therefore, according to the technology described herein, a downstream (receiving or decoding) device can store, retain, or cache portions of received image metadata, including optimized operational parameter values ​​generated by the upstream device and included in the image (container) file 122 (such as in the “di_rpu_data_mapping()” data structure), and if signaled or instructed by the upstream device through the image (container) file 122, such values ​​are used in subsequent operations performed by the downstream device or referenced as reference data for current and subsequent images.

[0223] As noted, each of the multiple image metadata portions (e.g., operating parameters, such as prediction coefficients) that are stored or cached as reference data at the downstream device can be clearly labeled, indexed, or identified with a corresponding value within a specific range (which may be the same value as the third flag “di_rpu_id” set by the upstream device and received by the downstream device).

[0224] In response to the determination that the reference data held or cached by the receiving device already includes a previously received (or otherwise held / cached) portion of image metadata with the same tag / index / identifier value as the currently transmitted portion of the image metadata of the current image, the downstream device can update the reference data by overwriting the previously received portion of image metadata with the currently received portion of image metadata (including any prediction coefficients explicitly transmitted for the current image).

[0225] For example, rather than restricting, the upstream device can determine that, for the current image, the value of the first flag (“use_prev_di_rpu_flag”) should be set to one (1). In response, the upstream device does not explicitly transmit or include inter-layer prediction coefficients for the current image in the image (container) file 122. Accordingly, the downstream device reads or reuses the operation parameters in the previously received or stored data structure (“di_rpu_data_mapping()”), which is tagged with a specific label / index / identifier value equal to the value of the second flag (“prev_di_rpu_id”) of the currently received image metadata unit (“rpu” with metadata unit flag or header field “di_rpu_type” = 8) for the current image, and processes the current image or picture using the previously received / stored operation parameters (including inter-layer prediction coefficients). Previously received or stored data structures can be identified by the value of a second flag ("prev_di_rpu_id") carried in or signaled in the currently received image metadata unit ("rpu" with "di_rpu_type" = 8).

[0226] However, if the value of the second flag ("prev_di_rpu_id") in the current image metadata unit (with "di_rpu_type" = 8) does not match any label, index, or identifier of the stored data structure ("di_rpu_data_mapping()") or any stored inter-layer prediction coefficients in the reference data held at the downstream device, then the downstream device or its image metadata parser may infer that the stored data structure containing the inter-layer prediction coefficients—with a label / index / identifier value equal to zero (0)—should be read and reused for the current image (e.g., as a fallback).

[0227] Furthermore, if there are no stored data structures or inter-layer prediction coefficients with a label / index / identifier value of 0 in the reference data, then downstream devices or image metadata parsers can fall back to using inter-layer prediction coefficients that represent (e.g., default, trivial, supported, etc.) a 1:1 linear mapping to process the current image or picture.

[0228] Different images—e.g., those with different image formats, characteristics, or qualities (such as different dynamic ranges, bit depths, color precisions, etc.)—can be logically represented as different image layers, generated from image data and accompanying image metadata in image (container) file 122. In some operational scenarios, image (container) file 122 may or may not explicitly include all image data and / or all image metadata required to generate a specific image or image layer within these different images or image layers. Instead, inter-layer image processing operations (such as inter-layer prediction, reshaping, mapping, etc.) can be used to generate some or all of the image data constituting a specific image.

[0229] Image processing operations described herein—which may include, but are not limited to, any of (e.g., interlayer, etc.) image prediction, image reshaping, image mapping, etc.—can be performed by signaling from an upstream device in image metadata or by transmitting polynomial pivot points, polynomial coefficients, etc., to a downstream device. Example image prediction, reshaping, and / or mapping methods described herein may include, but are not limited to, any of linear interpolation, second-order polynomial interpolation, multi-color channel, multi-regression prediction (MMR), etc., for example, as described in U.S. Patent No. 10,021,390, the entire contents of which are incorporated herein by reference. Additionally, optionally, or alternatively, image prediction, reshaping, and / or mapping methods described herein may include tensor product, B-spline prediction (TPB), for example, as described in U.S. Patent Application Publication No. 2022 / 0408081, the entire contents of which are incorporated herein by reference.

[0230] In some operational scenarios, image metadata compression can be performed on the DM metadata portion that contains display management (DM) operation parameters.

[0231] For example, a fourth flag decoded in the image (container) file (122) (e.g., indicated as "use_prev_level_md_flag", a specific syntax element, etc.) can be set by the upstream device to a specific value (e.g., use_prev_level_md_flag = 1, etc.) to signal to the downstream device or indicate that a previously sent portion of the DM metadata (e.g., a specific image metadata type, a type identified by another syntax element (such as di_rpu_type = 8), etc.) used for a previously sent / decoded image (e.g., in the same image (container) file 122, etc.) can be reused by the downstream device to perform DM operations for the current image. Therefore, the DM operations performed by the downstream device for the current image can share or utilize the same previously sent portion of the DM metadata, thereby reducing or omitting some or all of the separate (e.g., DM, etc.) DM metadata portions that the upstream device explicitly or specifically decodes or transmits to the downstream device in the image (container) file 122 for the current image.

[0232] On the other hand, the fourth flag (“use_prev_level_md_flag”) decoded in the image (container) file (122) can be set by the upstream device to a second specific value (e.g., use_prev_level_md_flag = 0, etc.) to signal to the downstream device that the DM metadata portion (e.g., a specific image metadata type, a type identified by another syntax element (such as di_rpu_type = 8), etc.) related to the DM operation to be performed for the current image is explicitly or specifically sent, transmitted, included, or present in the image (container) file 122. Therefore, the downstream device can perform a DM operation for the current image using the DM metadata portion explicitly or specifically encoded / transmitted by the upstream device for the current image in the image (container) file 122 and decoded or retrieved by the downstream device from the image (container) file 122 for the current image.

[0233] The fifth flag decoded in the image (container) file (122) (e.g., denoted as "prev_level_md_id", second specific syntax element, etc.) can be set by the upstream device to a specific value (e.g., a specific value in the range of 0 to 15, etc.) to signal to the downstream device or indicate which specific part of the previously sent image metadata (e.g. for previous images, etc.) will be used to perform DM operations for the current image.

[0234] DM metadata previously sent by an upstream device and received by a downstream device may include multiple (previously sent) DM metadata portions held or stored in memory or cache at the downstream device. Each of the multiple DM metadata portions at the downstream device may be marked or identified by a corresponding specific DM metadata portion identifier among multiple DM metadata portion identifiers respectively used for the multiple (previously sent) DM metadata portions. The corresponding specific DM metadata portion identifier may be in a specific value range such as 0 to 15.

[0235] A fifth flag (“prev_level_md_id”) with a specific value within a specific range can be used by an upstream device to signal to a downstream device that a specific (previously sent) DM metadata portion among multiple DM metadata portions held or cached at the downstream device will be used by the downstream device to perform DM operations for the current image. This specific DM metadata portion may have been previously sent for a previous image, where a third flag (“di_rpu_id”) set for that previous image has the same DM metadata portion identifier value as the fifth flag (“prev_level_md_id”, etc.) set for the current image.

[0236] Additionally, optionally, or alternatively, in response to determining that a fifth flag (“prev_level_md_id”) does not exist in the image (container) file 122 for the current image, the downstream device may continue to internally set the fifth flag to a value outside a specific range, such as -1 (e.g., an invalid rpu_id). This helps to avoid or prevent the downstream device from subsequently performing DM operations on the current image using an incorrect (previously sent) DM metadata portion.

[0237] The fifth flag (“prev_level_md_id”) decoded in the image (container) file (122) can be set by the upstream device to a second specific value (e.g., prev_level_md_id = 0, etc.) to signal or indicate to the downstream device that a specific DM metadata portion for the current image is explicitly or specifically sent, transmitted, included, or present in the image (container) file 122 from the upstream device to the downstream device. Therefore, DM operations can be performed on the current image using the (currently sent) specific DM metadata portion. Some or all of the operation parameters in the (currently sent) specific DM metadata portion can be specified using the (currently sent) data structure (e.g., denoted as “di_dm_data_payload()”, a specific set of syntax elements, etc.) explicitly or specifically decoded for the current image in the image (container) file 122.

[0238] Therefore, according to the technology described herein, a downstream (receiving or decoding) device can store, retain, or cache portions of the received DM metadata, including optimized operational parameter values ​​generated by the upstream device and included in the image (container) file 122 (such as in the “di_dm_data_payload()” data structure), and if so signaled or indicated by the upstream device through the image (container) file 122, then it is used or referenced in subsequent operations performed by the downstream device as reference data for the current and subsequent images.

[0239] As noted, each of the multiple DM metadata portions stored as reference data or cached at the downstream device—for example, operational parameters used for DM operations—can be clearly tagged, indexed, or identified with a corresponding value within a specific range (which may have the same value as the third flag “di_rpu_id” set by the upstream device and received by the downstream device).

[0240] In response to determining that the reference data held or cached by the receiving device already includes a previously received (or otherwise held / cached) DM metadata portion having the same tag / index / identifier value as the third flag (“di_rpu_id”) of the currently transmitted DM metadata portion of the current image, the downstream device can update the reference data by overwriting the previously received DM metadata portion with the currently received DM metadata portion explicitly transmitted for the current image.

[0241] For example, rather than restricting, the upstream device can determine that, for the current image, the value of the fourth flag (“use_prev_level_md_flag”) should be set to one (1). In response, the upstream device does not explicitly transmit or include DM operation parameters for the current image in the image (container) file 122. Accordingly, the downstream device reads or reuses the operation parameters in a previously received or stored data structure (“di_dm_data_payload()”) marked with a specific tag / index / identifier value equal to the value of the fifth flag (“prev_level_md_id”) of the currently received image metadata unit (“rpu” with metadata unit flag or header field “di_rpu_type” = 8) for the current image, and performs DM operation for the current image or picture using the previously received / stored DM operation parameters. Previously received or stored data structures can be identified by downstream devices using the value of the fifth flag ("prev_level_md_id") carried or signaled in the currently received image metadata unit ("rpu" with "di_rpu_type" = 8). Downstream devices can (for previous images) read previously stored data structures ("di_dm_data_payload()") from the last image metadata unit ("rpu" with "di_rpu_type" = 8) where the fourth flag ("use_prev_level_md_flag") is set to 0 and copy them to the DM metadata buffer of the current image or picture. If the level(s) present or tagged in the stored DM metadata corresponds to the DM metadata portion identifier of the current metadata unit ("rpu" with "di_rpu_type" = 8) of the current image or picture, then this copy can overwrite any existing DM metadata portions of the same level(s) in the current DM metadata buffer.

[0242] However, if the value of the fifth flag ("prev_level_md_id") in the current image metadata unit (with "di_rpu_type" = 8) does not match any tag, index, or identifier of the stored data structure ("di_dm_data_payload()") or any stored DM operation parameter in the reference data held at the downstream device, then the downstream device or the image metadata parser therein can infer that the stored data structure containing DM operation parameters with a tag / index / identifier value equal to zero (0) should be read and reused in the DM operation of the current image (e.g., as a fallback).

[0243] The steps for copying or reusing DM metadata mentioned above can be performed in the order described. Subsequently, the portion of DM metadata stored in the current DM metadata buffer for the current image is used to perform DM operations on the current image or picture.

[0244] A metadata compression subsystem or module may be implemented by an upstream device to support or perform metadata compression operations on image metadata related to or associated with image prediction, image reshaping, or display management. In some operational scenarios, some or all of the metadata compression operations (such as limited metadata compression) may be performed in the (image or picture) encoding order. Example encoding orders may be, but are not limited to, one of the following: encoding a key image frame first, followed by encoding one or more non-key image frames(s) that reference the key image frame; encoding a referenced image frame first, followed by encoding one or more image frames(s) that reference the referenced image, etc. Additionally, optionally, or alternatively, in some operational scenarios, an image metadata stream or substream with metadata compression may be generated by the upstream device in display order.

[0245] Figure 9 The illustration shows an example image metadata compression operation that can be implemented or performed by an upstream device.

[0246] Box 902 includes receiving a request to encode an image metadata portion (“RPU”) related to image prediction, image reshaping, etc., of an image frame. The upstream device then proceeds to determine whether the image frame is a keyframe.

[0247] Box 904 includes, in response to determining that an image frame is a keyframe, the image metadata portion (RPU) of this (reference) keyframe is... ref ) is set as the input image metadata section (RPU) in ), or is included based on it. Box 910 includes setting the first flag (“use_prev_di_rpu_flag”) to zero (0).

[0248] Box 906 includes, in response to determining that an image frame is not a keyframe, the upstream device continues to determine the input image metadata portion (RPU) of this (non-reference) keyframe. in Is it equal to the image metadata portion (RPU) that has already been set or included for the keyframe? ref ).

[0249] Box 908 includes the input image metadata portion (RPU) in response to determining this (non-reference) keyframe. in This is equivalent to the image metadata portion (RPU) that has already been set or included for keyframes. refThe upstream device further continues to set the first flag ("use_prev_di_rpu_flag") to one (1) and the second flag ("prev_di_rpu_id") to the same value as the third flag ("di_rpu_id") that has been set for the keyframe or included in the image metadata part (RPUref).

[0250] Figure 9 A’s processing flow can be repeated or iteratively executed on all images or pictures (or image frames) in the order of encoding, thereby generating compressed image metadata units (RPUs with compressed synthesizer metadata) for image processing operations such as (e.g., inter-layer) image prediction or reshaping operations.

[0251] Downstream devices or their metadata parsers can extract or receive image metadata (“RPU” portions or payloads from image (container) files 122 or bitstreams generated by upstream devices in encoded order.

[0252] The metadata parser may maintain a (e.g., a single etc.) first metadata buffer to store or cache the first (e.g., the entire etc.) data structure (“di_rpu_data_payload()”) of the current image metadata portion, with the first flag (“use_prev_di_rpu_flag”) of the current image metadata portion set to zero (0) and the third flag (“di_rpu_id”) also set to zero (0).

[0253] Additionally, optionally, or alternatively, the metadata parser may maintain another (e.g., a single, etc.) second metadata buffer to store or cache a second (e.g., the entire, etc.) DM metadata structure (“di_dm_data_payload()” or “di_dm_data_payload2()”) for which one or more flag indications or signaling notifications are explicitly carried or included in the image (container) file 122 or bitstream.

[0254] If the first data structure (“di_rpu_data_payload()”) does not exist, for example, as indicated by the first flag (“use_prev_vdr_rpu_flag” = 1), then the metadata parser can recover or reuse the operating parameters (e.g., synthesizer coefficients, etc., related to image prediction / mapping / reshaping) from the first metadata buffer.

[0255] If the second data structure (“di_dm_data_payload()” or “di_dm_data_payload2()”) does not exist, for example, as indicated by one or more corresponding flags used for DM metadata compression, then the metadata parser can recover or reuse the operating parameters (e.g., parameters or coefficients related to display management) from the second metadata buffer.

[0256] Example processing flow

[0257] Figure 4A An example processing flow according to an embodiment of the present invention is illustrated. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, inverse mapping generation and application systems, etc.) may perform this processing flow. In block 402, the image processing system encodes a primary image of a first image format into an image file specified for the first image format.

[0258] In box 404, the image processing system encodes a non-primary image in a second image format into one or more accompanying segments of the image file. The second image format is different from the first image format.

[0259] In box 406, the image processing system causes the display image derived from the reconstructed image to be rendered using the receiving device of the image file. The reconstructed image is generated from either the primary image or a secondary image.

[0260] In the embodiments, the primary image represents a JPEG image; the image file represents a JPEG image file; and the non-primary image represents a non-JPEG image.

[0261] In this embodiment, both the primary image and the secondary image are derived from the same source image.

[0262] In an embodiment, a non-primary image refers to one of one or more non-primary images in a second image format that are encoded together with the primary image in an image file.

[0263] In one embodiment, one or more segments are encoded in an image file as application 11 (APP11) tagged segments.

[0264] In one embodiment, one or more image metadata portions are encoded in one or more second segments of the image file.

[0265] In one embodiment, one or more image metadata portions include a specific image metadata portion that carries specific operational parameters for a specific image processing operation that the receiving device wants to perform on one of the primary or non-primary images.

[0266] In the embodiments, a particular image processing operation includes one or more of the following: forward image reshaping, backward image reshaping, inverse image mapping, image mapping, color space conversion, codeword linear mapping, codeword nonlinear mapping, display management operation, perceptual quantization-based mapping, mapping based on one or more transfer functions, and other image processing operations performed by the receiving device.

[0267] In an embodiment, a particular image metadata portion is generated by concatenating one or more boxes carried in one or more application 11 (APP11) tagged fragments included in one or more second fragments.

[0268] In an embodiment, a non-primary image of the second image format represents one of the following: an HEVC image, an AV1 image, or another non-JPEG image.

[0269] Figure 4B An example processing flow according to an embodiment of the present invention is illustrated. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, prediction models and feature selection systems, inverse mapping generation and application systems, etc.) may perform this processing flow. In block 452, the image decoding system receives an image file specified for a first image format. The image file encodes a primary image in the first image format.

[0270] In box 454, the image decoding system decodes a non-primary image in a second image format from one or more accompanying segments of the image file. The second image format differs from the first image format.

[0271] In box 456, the image decoding system enables the display image derived from the reconstructed image to be rendered on the image display. The reconstructed image is generated from one of the primary or secondary images using the image metadata carried in the image file.

[0272] In an embodiment, the image decoding system further performs the following: receiving a second image file designated for a first image format, wherein the second image file encodes a second primary image of the first image format and the second image file does not encode another image other than the second primary image; decoding the second primary image of the first image format from one or more second accompanying segments of the second image file; and causing a second display image derived from the second primary image to be reconstructed and rendered on an image display.

[0273] In an embodiment, the display image is generated through one or more image processing operations; one or more operation parameters for the one or more image processing operations are decoded from the image metadata portion carried in the image file.

[0274] In embodiments, computing devices (such as display devices, mobile devices, set-top boxes, multimedia devices, etc.) are configured to perform any of the aforementioned methods. In embodiments, an apparatus includes a processor and is configured to perform any of the aforementioned methods. In embodiments, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause the performance of any of the aforementioned methods.

[0275] In one embodiment, a computing device includes one or more processors and one or more storage media storing an instruction set, which, when executed by the one or more processors, causes any of the aforementioned methods to be performed.

[0276] It should be noted that although individual embodiments are discussed herein, any combination of the embodiments and / or some of the embodiments discussed herein may be combined to form other embodiments.

[0277] Example computer system implementation

[0278] Embodiments of the present invention can be implemented using computer systems, systems configured with electronic circuit systems and components, integrated circuit (IC) devices such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs), and / or devices comprising one or more such systems, devices, or components. The computer and / or IC can execute, control, or perform instructions related to adaptive perceptual quantization of images with enhanced dynamic range, such as those described herein. The computer and / or IC can calculate any of the various parameters or values ​​associated with the adaptive perceptual quantization described herein. Image and video embodiments can be implemented in hardware, software, firmware, and various combinations thereof.

[0279] Some embodiments of the present invention include a computer processor that executes software instructions that cause a processor to perform the methods of the present disclosure. For example, one or more processors, such as displays, encoders, set-top boxes, code converters, etc., can implement the methods related to adaptive perceptual quantization of HDR images described above by executing software instructions in a processor-accessible program memory. Embodiments of the present invention can also be provided in the form of a program product. A program product may include any non-transitory medium carrying a set of computer-readable signals including instructions that, when executed by a data processor, cause the data processor to perform the methods of embodiments of the present invention. A program product according to embodiments of the present invention can be any of a wide variety of forms. A program product may include, for example, physical media (such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc.). The computer-readable signals on the program product may optionally be compressed or encrypted.

[0280] When referring to components (e.g., software modules, processors, components, devices, circuits, etc.) above, unless otherwise stated, references to that component (including references to “device”) should be interpreted to include equivalents of that component (any component that performs the function of said component (e.g., functionally equivalent), including components that are not structurally equivalent to the disclosed structure but perform the function in the illustrative example embodiments of the invention).

[0281] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to execute these techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) persistently programmed to execute these techniques, or may include one or more general-purpose hardware processors programmed to execute these techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hardwired and / or program logic to implement these techniques.

[0282] For example, Figure 5 This is a block diagram illustrating a computer system 500 on which embodiments of the present invention can be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor.

[0283] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502, for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions executed by processor 504. When stored in non-transitory storage media accessible to processor 504, these instructions make computer system 500 a dedicated machine customized to perform the operations specified in the instructions.

[0284] Computer system 500 also includes a read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. Storage device 510 (such as a disk or optical disk) is provided and coupled to bus 502 for storing information and instructions.

[0285] Computer system 500 can be coupled to display 512 (such as an LCD) via bus 502 for displaying information to the computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516 (such as a mouse, trackball, or arrow keys) for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Such input devices typically have two degrees of freedom on two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane.

[0286] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware and / or program logic (which, in conjunction with the computer system, enable or program the computer system 500 as a special-purpose machine). According to one embodiment, computer system 500 performs the techniques described herein in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. These instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequence of instructions contained in main memory 506 causes processor 504 to perform the processing steps described herein. In alternative embodiments, hardwired circuitry may be used instead of or in combination with software instructions.

[0287] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips or cassette tapes.

[0288] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0289] Various forms of media can be used to transfer one or more sequences of one or more instructions to processor 504 for execution. For example, instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 transfers the data to main memory 506, from which processor 504 retrieves and executes the instructions. Instructions received by main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.

[0290] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which is connected to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem providing data communication connectivity with a corresponding type of telephone line. As another example, communication interface 518 may be a Local Area Network (LAN) card to provide data communication connectivity with a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying streams of digital data representing various types of information.

[0291] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection via local network 522 to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526. ISP 526 then provides data communication services via a global packet data communication network (now commonly referred to as the "Internet" 528). Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals carrying digital data streams. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.

[0292] Computer system 500 can send messages and receive data, including program code, through one or more networks, network links 520, and communication interfaces 518. In the Internet example, server 530 can send requested code to the application through the Internet 528, ISP 526, local network 522, and communication interface 518.

[0293] The received code can be executed by processor 504 upon receipt and / or stored in storage device 510 or other non-volatile memory for later execution.

[0294] Equivalents, extensions, alternatives and miscellaneous

[0295] In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary from implementation to implementation. Therefore, the sole and exclusive indication of the claimed embodiments of the invention and the content of the applicant's intent as a claimed embodiment of the invention is contained in the set of claims issued from this application, in the specific form of such claims, including any subsequent amendments. Any definitions of terms included in these claims that are expressly set forth herein will determine the meaning of the terms used in the claims. Therefore, any limitations, elements, characteristics, features, advantages, or attributes not expressly stated in the claims should not in any way limit the scope of such claims. Consequently, the description and drawings are considered illustrative rather than restrictive.

[0296] Exemplary examples of enumeration

[0297] This invention may be practiced in any of the forms described herein, including but not limited to the exemplary embodiments (EEEs) enumerated below, which describe the structure, features, and functionality of some portions of embodiments of the invention.

[0298] EEE1. A method comprising:

[0299] Encode the primary image of the first image format into the image file specified for the first image format;

[0300] Encode a non-primary image in a second image format into one or more accompanying segments of the image file, wherein the second image format is different from the first image format;

[0301] The display image derived from the reconstructed image is rendered using the receiving device of the image file, wherein the reconstructed image is generated from one of the primary image or the non-primary image.

[0302] EEE2. The method of EEE1, wherein the non-primary image is represented by non-primary image data divided into one or more payload data portions respectively included in one or more accompanying segments of the image file; wherein each of the one or more accompanying segments of the image file includes a corresponding box designated to contain the corresponding payload data portion of the one or more payload data portions.

[0303] EEE3. The method of EEE2, wherein each of the one or more accompanying segments of the image file includes a first data length field and a second data length field; wherein, in response to determining that the first length field is not set to a specific reserve value among a plurality of reserve values, the total number of bytes of all the one or more payload data portions is indicated in the first length field; wherein, in response to determining that the first length field is set to the specific reserve value, the total number of bytes of all the one or more payload data portions is indicated in the second length field.

[0304] The method of EEE4. EEE2, wherein each of the one or more accompanying segments of the image file includes a data type field having a specific data type value indicating the second image format.

[0305] The method of any one of EEE1-EEE4, wherein the primary image represents a JPEG image; wherein the image file represents a JPEG image file; and wherein the non-primary image represents a non-JPEG image.

[0306] The method of any one of EEE1-EEE5, wherein both the primary image and the secondary image are derived from the same source image.

[0307] The method of any one of EEE7, EEE1-EEE6, wherein the non-primary image represents one of one or more non-primary images of the second image format encoded together with the primary image in the image file.

[0308] The method of any one of EEE8, EEE1-EEE7, wherein one or more segments are encoded as application 11 (APP11) marked segments in the image file.

[0309] The method of any of EEE9, EEE1-EEE8, wherein one or more image metadata portions are encoded in one or more second segments of the image file.

[0310] A method of any one of EEE10, EEE1-EEE9, wherein the one or more image metadata portions include an image metadata portion containing an explicitly specified first operation parameter value for generating a first image from the image file; wherein the image file does not have an explicitly specified second operation parameter value for generating a second different image from the image file; wherein the image file includes one or more flags in place of the explicitly specified second operation parameter value to indicate the reuse of the explicitly specified first operation parameter value to generate the second different image from the image file.

[0311] The method of any one of EEE1-EEE10, wherein the explicitly specified first operation parameter value relates to one or more of the following: image prediction, image mapping, image reshaping, display management, or other image processing operations.

[0312] The method of EEE12. EEE9, wherein the one or more image metadata portions include a specific image metadata portion, the specific image metadata portion carrying specific operational parameters for a specific image processing operation to be performed by the receiving device on one of the primary image or the non-primary image.

[0313] EEE13. The method of EEE12, wherein the specific image processing operation includes one or more of the following: image forward reshaping, image backward reshaping, image inverse mapping, image mapping, color space conversion, codeword linear mapping, codeword nonlinear mapping, display management operation, perceptual quantization-based mapping, mapping based on one or more transfer functions, or other image processing operation performed by the receiving device.

[0314] EEE14. The method of EEE12, wherein a particular image metadata portion is generated by concatenating one or more boxes carried in one or more application 11 (APP11) tagged segments included in one or more second segments.

[0315] The method of any one of EEE1-EEE14, wherein the non-primary image of the second image format represents one of the following: an HEVC image, an AV1 image, or another non-JPEG image.

[0316] EEE16. A method comprising:

[0317] Receive an image file specified for a first image format, wherein the image file encodes a main image of the first image format;

[0318] Decode a non-primary image in a second image format from one or more accompanying segments of the image file, wherein the second image format is different from the first image format, and wherein both the primary image and the non-primary image are derived from the same source image;

[0319] This causes a display image derived from the reconstructed image to be rendered on an image display, wherein the reconstructed image is generated from one of the primary image or the non-primary image using image metadata carried in the image file.

[0320] The methods of EEE17 and EEE16 also include:

[0321] Receive a second image file designated for the first image format, wherein the second image file encodes a second primary image of the first image format, and wherein the second image file does not encode another image other than the second primary image;

[0322] Decode the second primary image of the first image format from one or more second accompanying segments of the second image file;

[0323] This causes the second display image derived from the second primary image to be reconstructed and rendered on the image display.

[0324] The method of EEE18, EEE16, or EEE17, wherein the displayed image is generated by one or more image processing operations; wherein one or more operation parameters for the one or more image processing operations are decoded from image metadata portions carried by the image file.

[0325] The method of any of EEE19, EEE16-18 further includes: maintaining at least one image metadata buffer to store at least one image metadata portion associated with the first image received with the image file; and applying one or more image processing operations associated with a second different image using a first operation parameter value specified in the at least one image metadata portion in the at least one image metadata buffer.

[0326] EEE20. A method comprising:

[0327] Capture raw images using a capture device;

[0328] The raw image is processed into an ISP-processed image using the image signal processor (ISP) of the capture device;

[0329] The ISP-generated image is converted into two images of different image formats using two codecs of different image formats in the capture device;

[0330] The photo processing subsystem of the capture device packages the two images of different image formats into a single image file;

[0331] The display image is generated by the receiving device of the single image file from one of the two images of different image formats and rendered on an image display that operates with the receiving device.

[0332] EEE21. A method comprising:

[0333] The photo processing device receives an input image file containing a first image in a first image format;

[0334] The first codec of the first image format in the photo processing device is invoked to decode the first image of the first image format into a decoded image;

[0335] The decoded image is converted into a second image of a second image format using a second codec of a second different image format in the photo processing device;

[0336] The photo processing subsystem of the capture device packages the first image in the first image format and the second image in the second image format into a single image file;

[0337] The display image is generated by the receiving device of the single image file from one of the two images of different image formats and rendered on an image display that operates with the receiving device.

[0338] EEE22. A method comprising:

[0339] The receiving device receives an image file containing two or more images in different image formats;

[0340] The codec of one of the different image formats in the photo receiving device is invoked to decode one of the two or more images of different image formats into a decoded image;

[0341] The photo receiving device generates a display image from the decoded image, at least in part based on the image display configuration of the photo receiving device;

[0342] The photo receiving device causes the displayed image to be rendered on the image display of the photo receiving device in the image display configuration of the photo receiving device.

[0343] EEE23. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause to perform any one of the methods described in EEE1-EEE22.

[0344] EEE24. An apparatus including a processor and configured to perform the method described in any one of EEE1-EEE22.

Claims

1. A method comprising: Encode the primary image of the first image format into the image file specified for the first image format; Encode a non-primary image in a second image format into one or more accompanying segments of the image file, wherein the second image format is different from the first image format; The display image derived from the reconstructed image is rendered using the receiving device of the image file, wherein the reconstructed image is generated from one of the primary image or the non-primary image.

2. The method of claim 1, wherein the non-primary image is represented by non-primary image data divided into one or more payload data portions respectively included in the one or more accompanying segments of the image file; wherein each of the one or more accompanying segments of the image file includes a corresponding frame designated to contain a corresponding payload data portion of the one or more payload data portions.

3. The method of claim 2, wherein each of the one or more accompanying segments of the image file includes a first data length field and a second data length field; wherein, In response to determining that the first length field is not set to a specific reserved value among a plurality of reserved values, the total number of bytes of all one or more payload data portions is indicated in the first length field; wherein, in response to determining that the first length field is set to the specific reserved value, the total number of bytes of all one or more payload data portions is indicated in the second length field.

4. The method according to any one of claims 1-3, wherein the primary image represents a JPEG image; wherein the image file represents a JPEG image file; and wherein the non-primary image represents a non-JPEG image.

5. The method of any one of claims 1-4, wherein both the primary image and the secondary image are derived from the same source image.

6. The method of any one of claims 1-5, wherein the one or more segments are encoded as application 11 (APP11) tagged segments in the image file.

7. The method of any one of claims 1-6, wherein one or more image metadata portions are encoded in one or more second segments of the image file; wherein the one or more image metadata portions include an image metadata portion containing an explicitly specified first operation parameter value for generating a first image from the image file; wherein the image file does not have an explicitly specified second operation parameter value for generating a second different image from the image file; wherein the image file includes one or more flags in place of the explicitly specified second operation parameter value to indicate the reuse of the explicitly specified first operation parameter value for generating the second different image from the image file.

8. The method of any one of claims 1-7, wherein the one or more image metadata portions include a specific image metadata portion carrying specific operational parameters for a specific image processing operation that the receiving device is to perform on one of the primary image or the non-primary image.

9. The method of any one of claims 1-8, wherein the non-primary image of the second image format represents one of the following: an HEVC image, an AV1 image, or another non-JPEG image.

10. A method comprising: Receive an image file specified for a first image format, wherein the image file encodes a main image of the first image format; Decode a non-primary image in a second image format from one or more accompanying segments of the image file, wherein the second image format is different from the first image format, and wherein both the primary image and the non-primary image are derived from the same source image; This causes a display image derived from the reconstructed image to be rendered on an image display, wherein the reconstructed image is generated from one of the primary image or the non-primary image using image metadata carried in the image file.

11. The method of claim 10, further comprising: Receive a second image file designated for the first image format, wherein the second image file encodes a second primary image of the first image format, and wherein the second image file does not encode another image other than the second primary image; Decode the second primary image of the first image format from one or more second accompanying segments of the second image file; This causes the second display image derived from the second primary image to be reconstructed and rendered on the image display.

12. The method of claim 10 or 11, wherein the displayed image is generated by one or more image processing operations; wherein one or more operation parameters for the one or more image processing operations are decoded from image metadata portions carried in the image file.

13. The method of any one of claims 10-12, further comprising: At least one image metadata buffer is maintained to store at least one portion of image metadata associated with the first image received with the image file; One or more image processing operations related to a second different image are applied using a first operation parameter value specified in at least one image metadata portion of the at least one image metadata buffer.

14. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause to perform the method as described in any one of claims 1-13.

15. An apparatus comprising a processor and configured to perform the method as claimed in any one of claims 1-13.

Citation Information

Patent Citations

  • Multiple color channel multiple regression predictor

    US10021390B2

  • Display management for high dynamic range images

    US20220164931A1

  • Tensor-product b-spline predictor

    US20220408081A1