Signaling and storage of jpeg ai images and image collections in media files

By using the HEIF format to store JPEG AI encoded and decoded images, the shortcomings of JPEG AI encoded and decoded image signaling and storage design are solved, enabling more efficient video data processing and decoding.

CN122641867APending Publication Date: 2026-08-25DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580010488.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-04
Filing Date
2025-01-16
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

The lack of effective signaling and storage design for JPEG AI encoded and decoded images in existing technologies leads to low efficiency in video data processing.

Method used

The High Efficiency Image File Format (HEIF) specified by ISO/IEC 23008-12 is used to store JPEG AI encoded or decoded images or image sets, and the JPEG AI image codec defined by ISO/IEC 6048-1 is used for encoding and decoding, with explicit signaling format to support parsing by the JPEG AI decoder.

Benefits of technology

It improves the efficiency and parsability of video data processing, ensures the effective storage and decoding of JPEG AI encoded images, and enhances the storage and transmission quality of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122641867A_ABST
    Figure CN122641867A_ABST
Patent Text Reader

Abstract

A mechanism of processing video data is disclosed. The mechanism includes determining that a media file format is based on a High Efficiency Image File Format (HEIF) specification for storage of an image or a set of images coded using a Joint Photographic Experts Group Artificial Intelligence (JPEG AI) image codec. A conversion between visual media data and a bitstream is performed based on the JPEG AI image codec.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority and claim to U.S. Provisional Patent Application No. 63 / 655,881, filed June 4, 2024, and U.S. Provisional Patent Application No. 63 / 621,820, filed January 17, 2024. All of the foregoing patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] This patent document relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology

[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention

[0005] The first aspect relates to a method for processing video data, comprising: for the storage of an image or collection of images encoded and decoded using the Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, determining that the media file format is based on the High Efficiency Image File Format (HEIF) specification; and performing a conversion between visual media data and bitstream based on the JPEG AI image codec.

[0006] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the aspects described above.

[0007] The third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec apparatus performs the methods of any of the preceding aspects.

[0008] The fourth aspect relates to a non-transitory computer-readable recording medium for storing a bitstream of video, the bitstream being generated by a method performed by a video processing apparatus, wherein the method includes: for the storage of an image or set of images encoded or decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, determining that the media file format is based on the High Efficiency Image File Format (HEIF) specification; and generating the bitstream based on the determination.

[0009] The fifth aspect relates to a method for storing a bitstream of video, comprising: for storing an image or set of images encoded or decoded using the Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, determining that the media file format is based on the High Efficiency Image File Format (HEIF) specification; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] For clarity, any of the embodiments described above may be combined with any one or more other embodiments described above to create new embodiments within the scope of this disclosure.

[0011] These and other features will become clearer through the following detailed description of the embodiments with reference to the accompanying drawings and claims. Attached Figure Description

[0012] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.

[0013] Figure 1 This is a schematic diagram of the example bitstream structure.

[0014] Figure 2 This is a block diagram illustrating an example video processing system.

[0015] Figure 3 This is a block diagram of an example video processing device.

[0016] Figure 4 This is a flowchart of an example method for video processing.

[0017] Figure 5 This is a block diagram illustrating an example video codec system.

[0018] Figure 6 This is a block diagram showing an example encoder.

[0019] Figure 7 This is a block diagram showing an example decoder.

[0020] Figure 8 This is a schematic diagram of an example encoder. Detailed Implementation

[0021] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the scope of the appended claims and the full scope of their equivalents.

[0022] The use of chapter headings in this document is for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, edited text changes are shown in bold italics (indicating deleted text) and bold (indicating added text) relative to the Multi-Functional Video Codec (VVC) specification and / or the Supplemental Enhancement Information (SEI) Message (VSEI) standard for encoding and decoding video bitstreams.

[0023] 1. Preliminary Discussion

[0024] This document relates to the Joint Image Experts Group Artificial Intelligence (JPEG AI) image file format. Specifically, this disclosure relates to the signaling and storage of JPEG AI-encoded images and image sets in media files based on the High Efficiency Image File Format (HEIF), which in turn is based on the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF). This idea can be applied alone or in various combinations to images encoded and decoded by any neural network (NN)-based codec (e.g., JPEG AI (i.e., ISO / IEC 60481, Information Technology – Learning-Based Image Codec Systems (JPEG AI) – Part 1: Core Codec Systems)) and any image file format (e.g., the JPEG AI image file format).

[0025] 2. Further discussion

[0026] 2.1 File Format Standards

[0027] Media streaming applications can be based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transmission methods and can rely on file formats such as ISO Basic Media File Format (ISOBMFF) [1]. One such streaming system is HTTP-based Dynamic Adaptive Streaming (DASH) [2]. In order to use video formats with ISOBMFF and DASH, video format-specific file format specifications such as Advanced Video Codec (AVC) and High Efficiency Video Codec (HEVC) file formats [3] will be needed to encapsulate the video content in ISOBMFF tracks as well as DASH representations and segments. Information about the video bitstream, such as grade, layer, and level, can be exposed as file format-level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, such as for selecting appropriate media segments, both for initialization at the start of a streaming session and for stream adaptation during a streaming session.

[0028] Similarly, in order to use an image format with ISOBMFF, you can use image format-specific file format specifications, such as the AVC image file format and HEVC image file format in [4].

[0029] 2.2. Image and Video Encoding and Decoding Based on Neural Networks (NNs)

[0030] Deep learning is rapidly developing across various fields, particularly in computer vision and image processing. Inspired by the tremendous success of deep learning in computer vision, many researchers have shifted their focus from image / video compression techniques to neural image / video compression. Neural networks are designed through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Significant progress has been made in neural network-based image / video compression techniques. It has been reported that image compression algorithms based on example neural networks have achieved rate-distortion (RD) performance comparable to Multifunctional Video Coding (VVC), a video coding standard developed by the Joint Video Experts Group (JVET), comprised of experts from the Moving Picture Experts Group (MPEG) and the Video Codec Experts Group (VCEG). With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, due to the inherent difficulty of the problem, neural network-based video coding is still in its early stages.

[0031] 2.2.1 Image / Video Compression

[0032] Image / video compression (also known as image / video encoding / decoding) generally refers to the computational techniques used to compress images or videos into binary code for easier storage and transmission. The binary code may or may not support lossless reconstruction of the original image or video; this is called lossless compression and lossy compression. Most efforts focus on lossy compression because lossless reconstruction is unnecessary in most cases. The performance of image or video compression algorithms is typically evaluated from two aspects: compression ratio and reconstruction quality. Compression ratio is directly related to the amount of binary code; less is better. Reconstruction quality is measured by comparing the reconstructed image or video with the original image or video; higher is better.

[0033] Image / video compression techniques can be divided into two branches: classical video encoding / decoding methods and neural network-based video compression methods. Classical video encoding / decoding schemes employ transform-based solutions, where researchers utilize statistical dependencies in latent variables (e.g., discrete cosine transform (DCT) or wavelet coefficients) by carefully hand-designing entropy encoding / decoding to model dependencies in the quantization domain. Neural network-based video compression takes two forms: neural network-based encoding / decoding tools and end-to-end neural network-based video compression. The former is embedded as an encoding / decoding tool within a classical video codec, serving only as part of the framework; the latter is a separate framework developed based on neural networks, independent of classical video codecs.

[0034] A series of classic video codec standards have been developed to accommodate the ever-increasing amount of visual content. The International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) has two expert groups, the Joint Group of Picture Experts (JPEG) and the Moving Picture Experts Group (MPEG), and the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) also has its own Video Codec Experts Group (VCEG), which is used for the standardization of image or video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, began working on the video codec standard Multifunction Video Codec (VVC). The first version of VVC was released in July 2020. Compared to HEVC, VVC reduces the bit rate by an average of 50% while maintaining the same visual quality.

[0035] Many researchers are working on neural network-based image encoding and decoding for use in neural network-based image / video compression. However, the network architectures used in example designs are relatively shallow, resulting in unsatisfactory performance. Thanks to abundant data and powerful computing resources, neural network-based methods have been better utilized in a variety of applications. Currently, neural network-based image / video compression has shown promising improvements and proven its feasibility. However, this technology is still far from mature, and many challenges need to be addressed.

[0036] 2.2.2 Neural Networks

[0037] Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. They typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform it into different kinds of representations. Note that these representations are not manually designed; instead, they are learned from massive amounts of data using general machine learning procedures. Deep learning eliminates the need for handcrafted representations and is therefore considered particularly suitable for processing natural, unstructured data, such as acoustic and visual signals, which has been a long-standing challenge in the field of artificial intelligence.

[0038] 2.2.3. Neural Networks for Image and Video Compression

[0039] Example neural networks used in image compression methods can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding / decoding strategies, while the latter is a transform-based solution. Sometimes, these two methods are combined.

[0040] Similar to classic video encoding and decoding techniques, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology lagged behind that of neural network-based image compression, but due to its complexity, more effort is needed to address the challenges. Some researchers are working on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to eliminate inter-frame redundancy. Inter-frame prediction is a key step in these works. Motion estimation and compensation are used, but this has only recently been achieved through trained neural networks.

[0041] Depending on the target scenario, research on neural network-based video compression can be divided into two categories: random access and low latency. In the random access case, decoding can begin at any point in the sequence, dividing the entire sequence into multiple separate segments, each of which can be decoded independently. The low latency case aims to reduce decoding time, allowing earlier frames to be used as reference frames for decoding subsequent frames.

[0042] 2.2.4. JPEG AI Image Encoding and Decoding Standard

[0043] 2.2.4.1. (9.2) Stream Layout ...

[0045] The overall grammatical structure of the image is as follows:

[0046]

[0047] Each bitstream begins with a 16-bit marker. All markers used in this specification are as follows:

[0048] ...

[0050] 2.2.4.2. (9.3) Image header

[0051] This substream contains information about the image height. ,width Potential space piece location and size, control flags for each tool, scaling factors for primary and secondary components, rate control parameters (for primary components) and secondary components )of - Information on the learnable model index and displacement.

[0052] 2.2.4.2.1 (9.3.1) Syntax Table

[0053] ...

[0055] 2.2.4.2.2. (9.3.1.2) Grading and Level Syntax

[0056] ...

[0058] 2.2.4.2.3. (9.3.2) Image header semantics

[0059] The following service information is transmitted via signaling:

[0060] picture_header_size is the number of bytes in the image header excluding the first two bytes of the marker;

[0061] Adding 64 to img_width specifies the width of the input image (from 64 to 65600).

[0062] img_height plus 64 specifies the height of the input image (from 64 to 65600);

[0063] picture_format is the data format of the output image (YUV420 = 0, YUV444 = 1, sRGB = 2, YUV422 = 3).

[0064] bit_depth is the bit depth of the output image ("0" corresponds to 8, "1" corresponds to 10); ...

[0066] stream_profile_idc indicates the stream profile that the bitstream conforms to.

[0067] The increment of 1 in num_decoder_profiles_minus1 specifies the number of decoder profiles supported by the bitstream.

[0068] decoder_profile_idc[i] indicates the i-th supported decoder profile provided by the bitstream.

[0069] level_idc indicates the level that the bitstream conforms to. ...

[0071] 2.2.5 The ongoing draft of the JPEG AI image encoding and decoding standard

[0072] The draft specification for the JPEG AI image codec standard is included in the JPEG output document WG1N100864, ​​which is the output of the 103rd JPEG conference.

[0073] The latest JPEG AI draft specification utilizes some of the neural network-based image encoding and decoding methods mentioned above. The following describes some features of the latest JPEG AI specification, including grade and level signaling. ...

[0075] 1. Bitstream Structure and Entropy Encoder

[0076] 1.1 Overview

[0077] This appendix specifies the bitstream structure (subclause 11.2) and the entropy decoder operation. The encoder must produce a stream that conforms to the specified order of elements.

[0078] 1.2 Stream Layout

[0079] A bitstream consists of a start and end point marked by stream markers, and multiple marker segments. A marker segment is the portion of the bitstream that begins with a marker. Each marker segment has byte-aligned boundaries. The markers and marker segments are as follows:

[0080] - SOC - Start of stream marker;

[0081] - PIH - The beginning of the image header marker segment, followed by the image header substream (sub-clause 11.4).

[0082] - SOZ - The start of the z-stream marker segment, followed by a sub-stream of the super-prior information tensor z, including [0] ( )and 1[1] ( )

[0083] - SORp - The start of the residual stream of the primary component marker segment, followed by sub-streams of the primary component residuals, including... [0] Data ( );

[0084] - SORs - The beginning of the residual stream of the minor component marker segment, followed by the sub-stream of the minor component residuals, including [1] Data ( );

[0085] - TOH - The beginning of the tool header marker segment, followed by the tool information substream (sub-clause 11.4.2).

[0086] - RDI - The beginning of the rendering information marker segment, followed by the rendering information substream (sub-clause 11.8).

[0087] - SOQ - The beginning of the quality graph marker segment, followed by subflows (sub-clause 11.5.2).

[0088] - UDI - The beginning of a user-defined information marker segment, followed by a sub-stream (sub-clause 11.7).

[0089] - EOC - End of stream marker.

[0090] Figure 1 This is a schematic diagram of an example layout of the bitstream structure. Markers are shown as boxes. Each mark (except for the Start of Stream (SOC) and End of Stream (EOC)) is followed by a bitstream segment. Each bitstream segment begins with a syntax element specifying the segment size. The start and end of each bitstream segment are identified using the marks and the segment size. Parsing of a bitstream segment is terminated after resolving the number of bytes specified in its size. The order of the bitstream segments is as follows: Figure 1 The data is displayed as vertical lines. The order is restricted as follows: the bitstream begins with SOC and ends with the EOC marker. Figure 1 The encoded segments displayed as boxes on either side of the vertical line can be placed in the bitstream in any order, but must follow the PIH (Picture Header) bitstream segment. Some bitstream segments are mandatory. They are in... Figure 1 The boxes shown to the left of the vertical lines are connected by solid lines. Optional stream segments are... Figure 1 The boxes shown are located to the right of the vertical lines and are connected by dashed lines.

[0091] The bitstream segments that begin with the marker start of z stream (SOZ), the minor start of residual (SORs), the major start of residual (SORp), and the start of quality map (SOQ) are parsed using a motion estimation table asymmetric numerical system (me-tANS) entropy encoder (sub-clause 11.5.4).

[0092] Each bitstream begins with a 16-bit marker. All markers used in this specification and their interpretations are listed in Table 1.

[0093] Table 1 - Stream Markers

[0094]

[0095] The flags with code assignments 0xff85 to 0xff87 are reserved for future use in this version of the standard. If any of the flag values ​​0xff85 to 0xff87 appear in the bitstream, subsequent bitstream segments must be ignored. The flags with code assignments 0xff8d to 0xff8f are reserved for future use in future versions of the standard. If any of the flag values ​​0xff8d to 0xff8f appear in the bitstream, decoding must be aborted; that is, the decoder must ignore the entire bitstream.

[0096] 1.3 Substream Syntax

[0097] 1.3.1 Substream Syntax Structure

[0098] The bitstream contains substreams. Each substream, which includes the substream associated with the reserved tag, begins with the tag ID, the substream size, and the payload data.

[0099] Apart from the image header substream, all other substreams follow the syntax of the substream() syntax structure, where the byte_alignment() syntax structure follows the substream_size syntax element and is also present at the end of the substream.

[0100] The size of the substream allows for the discarding of substreams associated with reserved tags in future versions of the standard, even when decoded by older decoders. The decoder can read the tag ID, and if it does not support the tag, the decoder must discard the remaining bytes of the substream indicated by the size field.

[0101]

[0102]

[0103]

[0104] 1.3.2 Semantics of Substreams

[0105] marker_id is the marker identifier for the substream. Possible values ​​are listed in sub-clause 11.2.

[0106] substream_size is the size of the substream (in bytes), excluding the two bytes used for the marker and the size of substream_size itself (in bytes).

[0107] The alignment_zero_bit must be equal to 0.

[0108] 1.4 Image Header

[0109] 1.4.1 Image Header Syntax

[0110] This marker segment contains information about the image height. ,width Potential domain slice location and size, control flags for each tool, scaling factors for primary and secondary components, rate control parameters (for primary components) and secondary components The learnable model index and displacement information.

[0111]

[0112]

[0113] 1.4.2 Image Header Semantics

[0114] The following service information is transmitted via signaling:

[0115] The picture_header_size is the number of bytes in the picture header substream excluding the first two bytes of the marker, the picture_header_size, and the byte_alignment() syntax structure.

[0116] stream_profile_idc indicates the stream profile that the bitstream conforms to.

[0117] The increment of 1 in num_decoder_profiles_minus1 specifies the number of decoder profiles supported by the bitstream.

[0118] decoder_profile_idc[i] indicates the i-th supported decoder profile provided by the bitstream. The variable decoderID is equal to decoder_profile_idc[i], which is the identifier of the synthesized transform network (sub-item 10.3).

[0119] level_idc indicates the level that the bitstream conforms to.

[0120] The value of `img_width_minus64` plus 64 specifies the width of the input image (from 64 to 65599). The variable `W` = `img_width` = `img_width_minus64` + 64.

[0121] The value of `img_height_minus64` plus 64 specifies the height of the input image (from 64 to 65599). The variable `H` = `img_height` = `img_height_minus64` + 64.

[0122] diff_display_img_width is a value ranging from 0 to 63;

[0123] display_image_width = img_width - diff_display_img_width

[0124] The display_image_width specifies that the pixels in columns 0 to 1 of the decoded image are used for display, and the pixels in the remaining columns are not used for display.

[0125] diff_display_img_height is a value ranging from 0 to 63;

[0126] display_image_height = img_height - diff_display_img_height

[0127] The display_image_height specifies that the pixels in rows 0 to 1 of the decoded image are used for display, and the pixels in the remaining rows are not used for display.

[0128] bit_depth_idc is the bit depth of the output image ("0" corresponds to 8, "1" corresponds to 10; other values ​​are reserved).

[0129] s_ver_minus1 is the definition of s ver = 1 + the value of one bit of s_ver_minus1; s ver This is the ratio of the height of the primary component to the height of the secondary component in the output image. ver The permissible values ​​are specified in Table 3. ver The usage of this is specified in sub-clauses 5.3, 8.6, 8.7, and 8.8.

[0130] s_hor_minus1 is the definition of s hor =1 + the value of one bit of s_hor_minus1, s hor This is the ratio of the width of the primary component to the width of the secondary component in the output image. ver The permissible values ​​are specified in Table 3. ver The usage of this is specified in sub-clauses 5.3, 8.6, 8.7, and 8.8.

[0131] c_ver_minus1 is the definition of c ver = 1 + the value of one bit of c_ver_minus1; c ver It is the ratio of the height of the primary component to the height of the secondary component in the encoded / decoded image. ver The permissible values ​​are specified in Table 3: 1 s ver c ver 2. c ver The usage of this is specified in sub-clauses 5.3, 8.6 and 10.3.

[0132] c_hor_minus1 is the definition of c hor =1 + the value of one bit of c_hor_minus1, where c hor It is the ratio of the width of the primary component to the width of the secondary component in the encoded / decoded image. hor The permissible values ​​are specified in Table 3: 1 s hor c hor 2. c hor The usage of this is specified in sub-clauses 5.3, 8.6 and 10.3.

[0133] `independent_beta_uv` is a flag (false / true). `independent_beta_uv` being true indicates that the beta displacement parameters of the principal and minor components are different. If `independent_beta_uv` being false, then the beta displacement parameters of the principal and minor components are the same.

[0134] `beta_displacement_log_plus_2048[comp]` minus 2048 is a parameter indicating the displacement between the rate control parameter `beta` selected by the encoder for the `comp` component and the reference rate control parameter `beta` associated with the index of the model used (the `model_id` syntax element). This displacement is on a logarithmic scale. `betaDisplacementLog[comp]` must be in the range of -1069 to 702 (inclusive) to maintain the performance of variable rate encoding / decoding.

[0135] betaDisplacementLog[comp] = clip(-1069,702, beta_displacement_log_plus_2048[comp] - 211)

[0136] Note - The reference rate control parameter beta (β) mentioned here is a parameter used during model training to control the ratio between bit rate and distortion. "Model" here refers to the model associated with the index (model_id) of the model used.

[0137] When the syntax element beta_displacement_log_plus_2048[1] does not exist (independent_beta_uv equals 0), betaDisplacementLog[1] = betaDisplacementLog[0].

[0138] `model_id` is an identifier for a pre-stored checkpoint with model weights; `model_id` = 0, 1, 2, or 3. Other values ​​(i.e., 4 to 15) are reserved for future expansions of the standard.

[0139] synthesis_tile_enable[comp] is an enable flag for slice processing of the primary component (comp==0) and secondary component (comp==1) in a synthesis transform.

[0140] `synthesis_tile_size[comp]` is the size of the tile used for the major components (comp == 0) and minor components (comp == 1) in the synthesis transform, counted by the elements of the latent tensor. The variable `SynthesisTransformTileSize[0] = synthesis_tile_size[0]` 16. VariableSynthesisTransfromTileSize[1] =synthesis_tile_size[1] 16

[0141] `synthesis_tile_overlap[comp]` is the size of the tile used for the primary component (comp == 0) and secondary component (comp == 1) in the synthesis transform overlap region. If it does not exist in the bitstream, `synthesis_tile_overlap[0] = 0` and `synthesis_tile_overlap[1] = 0`. The variable `SynthesisTransfromTileOverlap[0] = synthesis_tile_overlap[0]` 16. VariableSynthesisTransfromTileOverlap[1] =synthesis_tile_overlap[1] 16

[0142] A region_partitioning_flag of 1 indicates that the residual tensor, latent tensor prediction, and input and output tensors in the reconstruction process are partitioned into rectangular sub-tensor sets. A region_partitioning_flag of 0 indicates that the residual tensor, input tensor, and output tensor in latent tensor prediction are not partitioned.

[0143] The increment of 1 in num_ver_splits_minus1 specifies the number of vertical partitions of the residual tensor and the tensors in the potential prediction and reconstruction processes. When num_ver_splits_minus1 does not exist, its value is presumed to be 0.

[0144] The derived variable NumVerSubTensorSplits is equal to num_ver_splits_minus1 + 1.

[0145] The maximum and minimum values ​​of NumVerSubTensorSplits are constrained by the grade and level of the bitstream.

[0146] Incrementing 1 to num_hor_splits_minus1 specifies the number of horizontal splits. When num_hor_splits_minus1 does not exist, its value is presumed to be 0.

[0147] The derived variable NumHorSubTensorSplits is equal to num_hor_splits_minus1 + 1.

[0148] The maximum and minimum values ​​of NumHorSubTensorSplits are constrained by the grade and level of the bitstream.

[0149] `mcm_overlap_in_latent_samples` specifies the amount of overlap used in multi-stage context modeling, expressed in units of the number of latent samples. The variable `McmOverlap` is set to equal `mcm_overlap_in_latent_samples`. 2. If it does not exist, set it to 0.

[0150] `hyper_decode_overlap_in_latent_samples` specifies the amount of overlap used during hyper-decoding, in units of the number of latent samples. The variable `HyperDecoderOverlap` is set to equal `hyper_decode_overlap_in_latent_samples`. 2. If it does not exist, set it to 0.

[0151] A region_residual_in_its_own_substream_flag value of 1 indicates that the primary or secondary residual data for each region is in the substream. Residual substreams from multiple regions can exist in the bitstream. A region_residual_in_its_own_substream_flag value of 0 indicates that there is only one substream for the primary residual data and only one substream for the secondary residual data.

[0152] The variable VerSubTensorSize is set to equal floor(floor((img_height + 127) / 128) / NumHorResSplits). 128.

[0153] The variable HorSubTensorSize is set to equal floor(floor((img_width + 127) / 128) / NumVerResSplits). 128.

[0154] Tensors RCVer[6][NumVerSubTensorSplits][3] and RCHor[6][NumHorSubTensorSplits][3] are set as follows:

[0155] For d = 0..5;

[0156] For i = 0..NumHorSplits-1;

[0157] - RCVer [d][i][0] = i VerSubTensorSize / 2 d

[0158] - RCVer [d][i][1]=(i==( NumVerSubTensorSplits -1))? h d : (i+1) VerSubTensorSize / 2 d

[0159] - RCVer [d][i][2] = RCVer [d][i][1] - RCVer [d][i][0]

[0160] For j = 0..NumVerSplits-1;

[0161] - RCHor [d][j][0] = j HorSubTensorSize / 2 d

[0162] - RCHor [d][j][1] = (j==(NumHorSubTensorSplits - 1)) ? w d :(j+1) HorSubTensorSize / 2 d

[0163] - RCHor [d][j][2] = RCHor [d][j][1] - RCHor [d][j][0]

[0164] `cube_group_flag[comp]` is a 1-bit unsigned integer. `cube_group_flag=0` indicates that no `cube_flag` is signaled in a group, and all cube flags in the group are set to 1. `cube_group_flag=1` indicates that the cube flag of a group is signaled.

[0165] cube_flag[comp] is a variable with dimensions (h4+7). 3) ((w4+7) 3) A 1D array containing cube flags for the principal components (comp = 0). A value of 1 indicates the value for the residual tensor. A cube applies skip mode. A value of 0 indicates the value applied to the residual tensor. A cube is used to disable skip mode. If the comp value is 0, this flag controls the primary component residual signaling; otherwise, it controls the secondary component residual signaling. Dimensions h4 and w4 are defined in Table 2.

[0166] colour_transform_idx specifies the transformation between the internal color format and the output color format. colour_transform_idx equal to 0 indicates no color transformation, colour_transform_idx equal to 1 indicates a color transformation with the same parameters as specified in sub-clause 8.8, and colour_transform_idx equal to 2 indicates a color transformation with user-defined parameters.

[0167] colour_transform_matrix[i][j] is the color transformation matrix, which is transmitted via signal only when colour_transform_idx equals 2.

[0168] colour_transform_offset[i] is the offset for color transformation, transmitted only when colour_transform_idx equals 2.

[0169] `rvs_enable_flag[comp]` is a flag used in RVS for the `comp` component. A flag of 0 indicates that RVS is disabled, and a flag of 1 indicates that RVS is enabled.

[0170] grfs_enable_flag[comp] is an enable flag for the per-channel gain unit refinement scaling tool. If grfs_enable_flag[comp] is equal to 1, the residual tensor elements of the comp component in the channel indicated by grfs_channel_flag[comp] are scaled.

[0171] `grfs_channel_flag[comp]` is an array with 1-bit flags. Its size depends on the value of `comp`. For `comp` equal to 0, the size is C. p And for comp = 1, the size is C. sFor channels where grfs_channel_flag[comp] equals 0, the absolute value of the residual tensor of the comp component decreases; for channels where grfs_channel_flag[comp] equals 1, the absolute value of the residual tensor of the comp component increases. Here, comp=0 indicates the primary component, comp=1 indicates the secondary component, and the primary component is C. p and secondary C s The number of channels in the component potential tensor is defined in Table 2.

[0172] gain_3D_enable_flag is an enable flag for local quality control. If gain_3D_enable_flag is equal to 1, the residual tensor elements are scaled according to the rules specified in sub-clause 14.2.

[0173] `quality_map_entropy_index` is an indicator of the sigma index defined for decoding quality map information in me-tANS. The variable `QMapEntropyIndex` is set to `quality_map_entropy_index`.

[0174] `multi_threading_z` is a flag indicating that multiple threads are used for CodeStreamZ decoding.

[0175] log2_num_threads_z_minus1 is an indicator of the number of threads used for CodeStreamZ decoding.

[0176] The variable NumberOfThreadsZ is set as follows:

[0177] If multi_threading_z = 0, then NumberOfThreadsZ = 1.

[0178] If multi_threading_z = 1, then NumberOfThreadsZ = 2 << (log2_num_threads_z_minus1 + 1).

[0179] multi_threading_r[comp] is a flag that indicates that multiple threads are used for decoding CodeStreamR[comp].

[0180] log2_num_threads_r_minus1 [comp] is an indicator of the number of threads decoded by CodeStreamR[comp] (comp=0..1).

[0181] For comp=0.1

[0182] The variable NumberOfThreadsR[comp] is set as follows.

[0183] - If multi_threading_r[comp] = 0, then NumberOfThreadsR[comp] = 1.

[0184] - If multi_threading_r[comp] = 1, then NumberOfThreadsR[comp] = 2 << (log2_num_threads_r_minus1[comp] + 1).

[0185] `multi_threading_q` is a flag indicating that multiple threads are used for CodeStreamQ decoding.

[0186] log2_num_threads_q_minus1 is an indicator of the number of threads used for CodeStreamQ decoding.

[0187] The variable NumberOfThreadsQ is set as follows:

[0188] If multi_threading_q = 0, then NumberOfThreadsQ = 1.

[0189] If multi_threading_q = 1, then NumberOfThreadsQ = 2 << (log2_num_threads_q_minus1 + 1).

[0190] When present, additional_picture_header_bit[i] can have any value.

[0191] Suppose that the variable NumPhBitsAtThisPoint is the total number of bits of all syntax elements from the stream_profile_idc syntax element in the picture_header() syntax structure up to (but not including) the additional_picture_header_bit[0] syntax element (if present).

[0192] The value of NumPhBitsAtThisPoint must be less than or equal to picture_header_size 8. In a bitstream conforming to this version and standard, the picture_header_size 8 - The value of NumPhBitsAtThisPoint must be less than 8. The decoder must allow picture_header_size. The value of 8 - NumPhBitsAtThisPoint is greater than or equal to 8, and the values ​​of all instances of additional_picture_header_bit[i] must be ignored.

[0193] Note - An instance of additional_picture_header_bit[i] can be an extension bit only, a byte alignment bit only, or an extension bit followed by a byte alignment bit. ...

[0195] 3. Technical problems solved by publicly available technical solutions

[0196] There is a lack of design for signaling and storage of JPEG AI encoded and decoded images in media files.

[0197] 4. List of solutions and implementation examples

[0198] To address the aforementioned issues, the following summarized methods are disclosed. These aspects should be considered as examples for explaining general concepts, and not interpreted in a narrow sense. Furthermore, these examples can be applied individually or in any combination.

[0199] 1) In one example, for the storage of images or collections of images encoded or decoded using the JPEG AI image codec specified in ISO / IEC 6048-1, the media file format is based on the High Efficiency Image File Format (HEIF) specified in ISO / IEC 23008-12.

[0200] 2) In one example, a codec image item (also known as a JPEGAI image item) of a specific type (e.g., 'jai0') is a JPEG AI bitstream that carries a codec image as defined in ISO / IEC 6048-1.

[0201] 3) In one example, it is specified that each bitstream presented to the JPEG AI decoder is the content of a codec image item of a specific type (e.g., 'jai0').

[0202] 4) In one example, for basic information carrying JPEG AI image items, the JPEG AI header item attribute is specified, where the box type of the item attribute is a specific type, such as 'jaih'.

[0203] a. In one example, the JPEG AI header attribute carries at least one or more of the following information:

[0204] i. Information about the stream profile that the JPEG AI bitstream carried in the JPEG AI image item conforms to.

[0205] ii. Information regarding the decoder profile provided by the JPEG AI bitstream carried in the JPEG AI image item.

[0206] iii. Information regarding the level of JPEG AI bitstream conformance carried in JPEG AI image items.

[0207] iv. Information regarding the color format of the output image decoded from the JPEG AI bitstream carried in the JPEG AI image item.

[0208] v. Information about the bit depth of the output image decoded from the JPEG AI bitstream carried in the JPEG AI image item.

[0209] vi. Information about color sampling mode, scaling factor, encoding / decoding image color format, and output image color format.

[0210] vii. Information about dividing the image into regions, such as the number of region rows and the number of region columns.

[0211] i. Information about whether each codec region is in its own substream.

[0212] ii. Information regarding whether the segmented regions are independent of each other.

[0213] 5) In one example, the JPEG AI codec auxiliary image is specified to use a specific item_type value, such as 'jai0'.

[0214] 6) In one example, specify a specific file brand for JPEG AI encoded images and image sets, such as 'jaii'.

[0215] a. In one example, it is specified that the encoded / decoded image item conforms to the 'jaii' brand when all of the following constraints are true:

[0216] i. This item has a type 'jai0' associated with the JPEG AI header attribute, wherein the type 'jai0' and the JPEG AI header attribute are specified by items 2) and 4) above and their sub-items.

[0217] ii. This item is not associated with any other type of basic item attribute except for 'ispe', 'pixi', 'jxpl', 'colr', 'rloc', 'irot', 'clap' and 'imir' as specified in ISO / IEC 23008-12:2017.

[0218] 7) In one example, specify the media subtype of media type 'image' for JPEG AI codec images and image collections, for example, named 'jaii'.

[0219] a. In one example, the presence of an image item of type 'jaii' is specified by including an item description in the itemtypes parameter, wherein the item type string of the item description begins with 'jai0', followed by a period ('.'), and further followed by a series of values ​​separated by periods ('.'), which include a subset of the information carried in the JPEG AI header item attributes specified by item 4) and its sub-items above, each value being encoded as a hexadecimal number.

[0220] 5. Examples

[0221] The following are some example embodiments of the aspects outlined in Section 4 of the previous article, which can be applied to potential standard specifications for the JPEG AI image file format.

[0222] 5.1 First Embodiment

[0223] This example applies to all items outlined in Section 4 of the previous article.

[0224] 1. Scope

[0225] This document specifies the container file format for JPEG AI streams as defined in ISO / IEC 6048-1. It defines a file format for processing still images and moving image sequence files on computer platforms, allowing for internet-based communication and other communications.

[0226] This document uses an existing file format specification and extends it for embedding JPEG AI streams.

[0227] 2. Standardized Citations

[0228] The following documents are cited in the text such that some or all of them constitute the requirements of this document. For dated references, only the cited version applies. For undated references, the latest version of the cited document (including any revisions) applies.

[0229] ISO / IEC 14496-12, Encoding and decoding of audiovisual objects - Part 12: ISO basic media file format

[0230] ISO / IEC 23008-12:2017, Information technology – Efficient encoding / decoding and media transmission in heterogeneous environments – Part 12: Image file formats

[0231] ISO / IEC 6048-1, Information technology – Learning-based image encoding and decoding systems (JPEG AI) – Part 1: Core encoding and decoding systems

[0232] ISO / IEC 6048-2, Information technology – Learning-based image encoding and decoding systems (JPEG AI) – Part 2: Classification

[0233] Recommendation ITU-T H.273 | ISO / IEC 23091-2, Code-code independent code points - Part 2: Video

[0234] 3. Terms and Definitions

[0235] For the purposes of this document, the terms and definitions given in ISO / IEC 14496-12, ISO / IEC 6048-1, ISO / IEC 6048-2, and ISO / IEC 23008-12, as well as the following terms and definitions, shall apply.

[0236] ISO and IEC maintain a database of terms used in their standardization at the following addresses:

[0237] - ISO online browsing platform: available at https: / / www.iso.org / obp

[0238] - IEC Electronic Encyclopedia: Available at http: / / www.electropedia.org / 3.1

[0240] box

[0241] A structured collection of data describing an image or the image decoding process. 3.2

[0243] Box type

[0244] The types of information stored together with box (3.1) 3.3

[0246] byte

[0247] 8-bit group 3.4

[0249] Code points that are independent of encoding and decoding

[0250] The definition of color space is based on code points with enumerated values.

[0251] Note 1: Code points defined in Recommendation ITU-T H.273 | ISO / IEC 23091-2. 3.5

[0253] High-efficiency image file format

[0254] Image file formats that can embed still images and motion sequences (3.7)

[0255] Note 1: Based on ISO / IEC 23008-12. 3.6

[0257] Image collection

[0258] An unordered collection of images without implicit or signal-transmitted presentation order or presentation timestamps. 3.7

[0260] motion sequence

[0261] Movie

[0262] Timing sequence of images (3.9) 3.8

[0264] sample

[0265] <isobmff>All data associated with a single time

[0266] Note 1: This definition is used in Appendices B and C as data associated with a codec image in a sequence. 3.9

[0268] Timing sequence

[0269] A linearly ordered sequence of media entities (such as images), where each entity is presented at a well-defined timestamp.

[0270] 4. Abbreviations

[0271] For the purposes of this document, the abbreviations given in ISO / IEC 14496-12, ISO / IEC 6048-1, ISO / IEC 6048-2, ISO / IEC 23008-12, and the following abbreviations shall apply.

[0272]

[0273] 5. Naming conventions for numerical values

[0274] Integers are represented as bit patterns, hexadecimal values, or decimal numbers. Both bit patterns and hexadecimal values ​​have a numerical value and a specific associated length in bits.

[0275] Hexadecimal representation, indicated by the prefix "0x", can be used instead of binary representation to represent bit patterns that are multiples of 4 in length. For example, 0x41 represents an octet where only its second most significant bit and its least significant bit are equal to 1. In a table called the "encode / decode table," the values ​​specified under the "encode / decode" heading are bit pattern values ​​(strings of numbers defined as equal to 0 or 1, where the leftmost bit is considered the most significant bit). Other values ​​that do not begin with the prefix "0x" are decimal values. When used in expressions, hexadecimal values ​​are interpreted as having a value equal to the value of the corresponding bit pattern evaluated as an unsigned integer (i.e., the value of a number formed by a bit pattern prefixed with a sign bit equal to 0, and the result interpreted as a two's complement representation of the integer value). For example, the hexadecimal value 0xF is equivalent to the 4-bit pattern '1111' and is interpreted in expressions as equal to the decimal number 15.

[0276] 6. Consistency

[0277] This document shares a common definition of file structure (a sequence of objects, referred to here as boxes, or atoms in other similar file formats) and a common definition of the general structure of objects (size and type).

[0278] The file format representing an image or image sequence must be as specified in Appendix B. All of these specifications require the reader to ignore any unrecognizable objects.

[0279] In any case where there are differences or conflicts, this document takes precedence over the document on which it is based; however, no such conflict is known to exist.

[0280] For better readability and understanding, the syntax descriptions of different file formats are presented in the same way as the basic format.

[0281] 7. Color Specifications

[0282] JPEG AI (as defined by ISO / IEC 6048-1) describes only the encoded bitstream of an image. For the correct display or interpretation of an image, it is crucial to correctly represent the color space of the image data. Therefore, the corresponding container file format must transmit the correct color space via signaling. The format defined for JPEG AI in this document transmits the color space as specified in Recommendation ITU-T H.273 | ISO / IEC 23091-2 via signaling.

[0283] Appendix A (Standardized) Use of JPEG AI Bitstream in HEIF Image File Format

[0284] A.1 Overview

[0285] This appendix specifies a format for encapsulating JPEG AI images, image sets, and image sequences in the HEIF image file format as defined in ISO / IEC 23008-12. The branding of individual images, image sets, and image sequences is specified in section A.4.

[0286] Note: In the context of ISOBMFF (ISO / IEC 14496-12), the sample refers to "all data associated with a single time." In this appendix, it refers to data associated with a single encoded image, not "pixels."

[0287] A.2 JPEG AI Images and Image Collections

[0288] A.2.1 Overview

[0289] Clause A.2 specifies the requirements for ISO / IEC 23008-12 documents that include JPEG AI encoded image items. When the brand specified in sub-clause A.3.1 is one of the compatible brands for the document, the requirements specified in clause A.2 must be complied with.

[0290] The provisions of Clause 6 of ISO / IEC 23008-12:2017 apply.

[0291] A.2.2. Image Encoding / Decoding Items

[0292] The encoded image item of type 'jai0' is a JPEG AI stream, which carries an encoded image as defined in ISO / IEC 6048-1.

[0293] Each bitstream presented to the JPEG AI decoder is the content of the encoded and decoded image item.

[0294] A.2.3 JPEG AI Header Attributes

[0295] A.2.3.1 Definition

[0296]

[0297] Each encoded image item representing a JPEG AI image, as defined in ISO / IEC 23008-12, may have an associated property as JPEGAIHeaderProperty.

[0298] For the 'jaiH' item property associated with an image item of type 'jai0', the basic bit of the ItemPropertyAssociationBox must be equal to 1.

[0299] A.2.3.2 Syntax

[0300] class JPEGAIHeaderProperty()

[0301] extends ItemFullProperty ('jaiH', 0, version=0){

[0302] unsigned int(8) stream_profile_idc;

[0303] unsigned int(8) num_decoder_profiles_minus1;

[0304] for (i=0; i <= num_decoder_profiles_minus1; i++)

[0305] unsigned int(8) decoder_profile_idc[i];

[0306] unsigned int(8) level_idc

[0307] bit(3) reserved = '111'b;

[0308] unsigned int(2) colour_format_idc;

[0309] unsigned int(3) bit_depth_idc;

[0310] }

[0311] Or, the syntax is as follows:

[0312] class JPEGAIHeaderProperty()

[0313] extends ItemFullProperty ('jaiH', 0, version=0){

[0314] unsigned int(4) stream_profile_idc;

[0315] unsigned int(4) num_decoder_profiles_minus1;

[0316] for (i=0; i <= num_decoder_profiles_minus1; i++)

[0317] unsigned int(4) decoder_profile_idc[i];

[0318] [[ID=—>32]]unsigned int(—>4) level_idc

[0319] if(num_decoder_profiles_minus1 % 2 == 1)

[0320] unsigned int(4) byte_alignment_4bits = '1111'b;

[0321] unsigned int(3) bit_depth_idc;

[0322] unsigned int(1) s_ver_minus1;

[0323] unsigned int(1) s_hor_minus1;

[0324] unsigned int(1) c_ver_minus1; It should be noted that there seems to be an error in the original text where "unsigned int(—>4) level_idc" and "unsigned int(—>3) bit_depth_idc" have incorrect notations. I've translated them as they are but this might need to be corrected in the source.

[0325] unsigned int(1) c_hor_minus1;

[0326] bit(1) reserved = '1'b;

[0327] unsigned int(1) regions_independent_flag;

[0328] unsigned int(7) num_region_rows_minus1;

[0329] unsigned int(7) num_region_cols_minus1;

[0330] }

[0331] A.2.3.3 Semantics

[0332] - For the bitstream (hereinafter referred to as "stream") carried in an image item associated with the attributes of a sample entry containing a configuration box in the application, stream_profile_idc, num_decoder_profiles_minus1, decoder_profile_idc[i], level_idc, colour_format_idc, and bit_depth_idc contain matching values ​​of the fields stream_profile_idc, num_decoder_profiles_minus1, decoder_profile_idc[i], level_idc, picture_format, and bit_depth_idc as defined in ISO / IEC 6048-1.

[0333] Alternatively, the semantics are as follows:

[0334] - For the bitstream (hereinafter referred to as "stream") carried in an image item associated with an attribute of a sample entry containing a configuration box in the application, stream_profile_idc, num_decoder_profiles_minus1, decoder_profile_idc[i], level_idc, bit_depth_idc, s_ver_minus1, s_hor_minus1, c_ver_minus1, c_hor_minus1, regions_independent_flag, num_region_rows_minus1, and num_region_cols_minus1 contain the fields stream_profile_idc, num_decoder_profiles_minus1, decoder_profile_idc[i] as defined in ISO / IEC 6048-1. The matching values ​​of ], level_idc, bit_depth_idc, s_ver_minus1, s_hor_minus1, c_ver_minus1, c_hor_minus1, region_residual_in_its_own_substream_flag, num_ver_splits_minus1, and num_hor_splits_minus1.

[0335] A.2.4 JPEG AI-assisted image

[0336] The URN specified for HEVC in B.2.4 of ISO / IEC 23008-12:2017 can also be used with JPEG AI.

[0337] The auxiliary image for JPEG AI encoding and decoding uses the item_type value 'jai0'.

[0338] Currently, there is no defined value for the aux_subtype byte array in AuxiliaryTypeProperty for JPEG AI-assisted images.

[0339] Note: The auxiliary image is considered the same as the main image. Among other things, this means that the initialization data for the auxiliary image in JPEG AI encoding and decoding is provided by attributes. The same image item can be associated with AuxiliaryTypeProperty.

[0340] A.3 JPEG AI Specific Brand

[0341] A.3.1 JPEG AI Images and Image Collection Brands

[0342] A.3.1.1 Overview

[0343] The brand 'jaii' is specified in the following sub-terms.

[0344] The encoding / decoding image item conforms to the 'jaii' brand when all of the following constraints are true:

[0345] - This item has the type 'jai0' and conforms to the provisions in A.2.

[0346] - This item is not associated with any other type of basic item property except for 'ispe', 'pixi', 'jxpl', 'colr', 'rloc', 'irot', 'clap', and 'imir'.

[0347] For the following definition, a crop-rotate-mirror derived image item is defined as a derived image item of type 'iden', and it is not associated with any other type of base item property except 'irot', 'clap', and 'imir'.

[0348] A.3.1.2 Requirements for HEIF files

[0349] The document must include 'mif1' in the compatible brand name, thus complying with the requirements of A.2.1.1 of ISO / IEC 23008-12:2017. Furthermore, the document must comply with the requirements of A.2.

[0350] Documents conforming to the 'jaii' brand are subject to the following additional constraints:

[0351] Documents that include 'jaii' as a compatible brand must contain an item that exists in the document, is a primary item, or is an alternative group containing a primary item, and satisfy one of the following constraints:

[0352] - This item is a codec image item conforming to the 'jaii' brand as specified in A.3.1.1.

[0353] - This item is a crop-rotate-mirror derived image item, and each source image item of this item is a crop-rotate-mirror derived image item or a codec image item conforming to the 'jaii' brand as specified in A.3.1.1.

[0354] A.3.1.3 Requirements for HEIF Readers

[0355] The reader must meet the requirements specified in A.2.1.2 of ISO / IEC 23008-12:2017.

[0356] Readers compliant with the 'jaii' brand must support displaying the main item or any item in an alternative group containing the main item and must meet one of the following constraints:

[0357] - This item is a codec image item conforming to the 'jaii' brand as specified in A.3.1.1.

[0358] - This item is a crop-rotate-mirror derived image item, and each source image item of this item is a crop-rotate-mirror derived image item or a codec image item conforming to the 'jaii' brand as specified in A.3.1.1.

[0359] The document reader should support displaying images with opacity information specified by the associated auxiliary image, either directly in the Picture() bitstream or as a separate bitstream with aux_type equal to urn:mpeg:hevc:2015:auxid:1 (as specified in ISO / IEC 23008-12).

[0360] A.4 JPEG AI Encoding / Decoding Images in ISO / IEC 23008-12 Image File Media Type Registration

[0361] A.4.1 Overview

[0362] The file extensions and media types of files within this family typically reflect the primary brand in the FileTypeBox. When the primary brand indicates a brand associated with A.3.1 (Single Image and Collection of Images), the media type defined here should be used. The media type may also be used when such a brand is a compatible brand. Sub-clause A.4.2 provides for media type registration, in accordance with the Internet Engineering Task Force (IETF) Request for Comment (RFC) 6838.

[0363] A.4.2 Registration

[0364]

[0365] 7. References

[0366] [1] IETF RFC 6838, Media Type Specification and Registration Procedure

[0367] [2] ISO / IEC 14496-12: "Information technology - Encoding and decoding of audiovisual objects - Part 12: ISO basic media file format".

[0368] [3] ISO / IEC 23009-1: "Information technology - Dynamic adaptive streaming based on HTTP (DASH) - Part 1: Media presentation description and segment format".

[0369] [4] ISO / IEC 14496-15: "Information technology - Encoding and decoding of audiovisual objects - Part 15: Carrying of structured video in the Network Abstraction Layer (NAL) unit of the ISO Basic Media File Format".

[0370] [5] ISO / IEC 23008-12: "Information technology - efficient encoding and decoding and media transmission in heterogeneous environments - Part 12: image file formats".

[0371] Figure 2 This is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0372] System 4000 may include an encoding / decoding component 4004 capable of implementing the various encoding / decoding or coding methods described in this document. Encoding / decoding component 4004 can reduce the average bit rate from the video input 4002 to the output of encoding / decoding component 4004 to produce an encoded / decoded representation of the video. Encoding / decoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 4004 may be stored or transmitted via a communication connection such as that represented by component 4006. The bitstream (or encoded / decoded) representation of the video received at input 4002, whether stored or communicated, may be used by component 4008 to generate pixel values ​​or transmit displayable video to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "encoding / decoding" operations or tools, it should be understood that encoding / decoding tools or operations used at the encoder will be followed by corresponding decoding tools or operations that inversely reproduce the encoding / decoding results by the decoder.

[0373] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), etc. The technologies described in this document can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0374] Figure 3 This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 may be configured to implement one or more methods described herein. The memories(multiple) 4104 may be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.

[0375] Figure 4 This is a flowchart of an example method 4200 for video processing. In step 4202, method 4200 determines that the media file format for a storage of an image or set of images encoded or decoded using the Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec is based on the High Efficiency Image File Format (HEIF) specification. In step 4204, a conversion between visual media data and a bitstream is performed based on the JPEG AI image codec. This conversion may include encoding at the encoder, decoding at the decoder, or a combination thereof.

[0376] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions cause the processor to execute method 4200 when executed by the processor. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video encoding / decoding device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video encoding / decoding device to execute method 4200 when executed by a processor.

[0377] Figure 5 This is a block diagram illustrating an example video encoding / decoding system 4300 that can utilize the techniques disclosed herein. The video encoding / decoding system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and this destination device 4320 may be referred to as a video decoding device.

[0378] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of such sources. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or transmitter. Encoded video data may be transmitted directly to destination device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by destination device 4320.

[0379] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may acquire encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320, wherein the destination device 4320 may be configured to interface with an external display device.

[0380] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVM) standard, and other existing and / or further standards.

[0381] Figure 6 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 5 The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0382] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406), a residual generation unit 4407, a transform unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414.

[0383] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture containing the current video block.

[0384] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.

[0385] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.

[0386] The mode selection unit 4403 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding), for example, based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).

[0387] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.

[0388] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0389] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 4404 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0390] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0391] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0392] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.

[0393] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0394] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0395] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0396] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components in the current video block.

[0397] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.

[0398] Transform unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0399] After the transform unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0400] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the buffer 4413.

[0401] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0402] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0403] Figure 7 This is a block diagram illustrating an example of a video decoder 4500, which can be... Figure 5 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0404] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.

[0405] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.

[0406] The motion compensation unit 4502 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.

[0407] The motion compensation unit 4502 can use interpolation filters, such as those used by the video encoder 4400 during the encoding of a video block, to calculate sub-integer pixels for a reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate a prediction block.

[0408] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0409] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.

[0410] The reconstruction unit 4506 can add the residual block to the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0411] Figure 8 This is a schematic diagram of the example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding offsets and applying finite impulse response (FIR) filters respectively, and utilizing the encoded / decoded side information through signal transmission offsets and filter coefficients. ALF 4606 is located in the last processing stage of each image and can be considered as a tool to attempt to capture and repair artifacts caused by previous stages.

[0412] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy codec component 4618. The entropy codec component 4618 entropy codes and decodes the prediction results and the quantized transform coefficients and transmits them toward a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.

[0413] The following provides a list of preferred solutions as examples.

[0414] The following solutions illustrate examples of the techniques discussed in this article.

[0415] 1. A method for processing media data, comprising: for storage of an image or collection of images encoded or decoded using the Joint Image Experts Group Artificial Intelligence (JPEGAI) image codec, determining that the media file format is based on the High Efficiency Image File Format (HEIF) specification; and performing a conversion between visual media data and a bitstream based on the JPEGAI image codec.

[0416] 2. According to the method of Solution 1, wherein the codec image item of a specific type, represented as a JPEG AI image item, is Picture().

[0417] 3. The method according to any one of solutions 1-2, wherein each bitstream presented to the JPEG AI decoder is the content of a specific type of encoded image item.

[0418] 4. The method according to any one of solutions 1-3, wherein the specific type is 'jai0'.

[0419] 5. The method according to any one of solutions 1-4, wherein, for carrying basic information of JPEG AI image items, JPEG AI header item attributes are specified, wherein the box type of the item attribute is a second specific type.

[0420] 6. The method according to any one of solutions 1-5, wherein the second specific type is 'jaih'.

[0421] 7. The method according to any one of solutions 1-6, wherein the JPEG AI header attribute carries information about the stream profile to which the JPEG AI bitstream carried in the JPEG AI image item conforms.

[0422] 8. The method according to any one of solutions 1-7, wherein the JPEG AI header attribute carries information about the decoder profile provided by the JPEG AI bitstream carried in the JPEG AI image item.

[0423] 9. The method according to any one of solutions 1-8, wherein the JPEG AI header attribute carries information about the level of conformance of the JPEG AI bitstream carried in the JPEG AI image item.

[0424] 10. The method according to any one of solutions 1-9, wherein the JPEG AI header attribute carries information about the color format of the output image generated by decoding the JPEG AI bitstream carried in the JPEG AI image item.

[0425] 11. The method according to any one of solutions 1-10, wherein the JPEG AI header attribute carries information about the bit depth of the output image generated by decoding the JPEG AI bitstream carried in the JPEG AI image item.

[0426] 12. The method according to any one of solutions 1-11, wherein the auxiliary image of the JPEG AI encoding / decoding uses a specific item_type value, wherein the specific item_type value is 'jai0'.

[0427] 13. The method according to any one of solutions 1-12, wherein a specific filename 'jaii' is specified for the images and image sets encoded and decoded by JPEG AI.

[0428] 14. The method according to any one of solutions 1-13, wherein the codec image item is defined as 'jaii' when all of the following constraints are satisfied:

[0429] This item has the type 'jai0' associated with the JPEG AI header attribute; and

[0430] This item is not associated with any other type of basic item property except for 'ispe', 'pixi', 'jxpl', 'colr', 'rloc', 'irot', 'clap', and 'imir'.

[0431] 15. The method according to any one of solutions 1-14, wherein, for images and image sets encoded and decoded by JPEG AI, a media subtype of media type 'image' is specified, and wherein the media subtype is 'jaii'.

[0432] 16. The method according to any one of solutions 1-15, wherein the presence of an image item of type 'jaii' is transmitted by including an item description in the itemtypes parameter, wherein the item type string of the item description begins with 'jai0', followed by a period ('.'), and further followed by a series of values ​​separated by periods ('.'), wherein the values ​​separated by periods ('.') include a subset of information carried in the JPEG AI header item attributes, each value being encoded as a hexadecimal number.

[0433] 17. The method according to any one of solutions 1-16, wherein the conversion includes encoding visual media data into a bitstream.

[0434] 18. The method according to any one of solutions 1-16, wherein the conversion includes decoding visual media data from a bitstream.

[0435] 19. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-18.

[0436] 20. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform the method according to any one of solutions 1-18 when executed by a processor.

[0437] 21. A non-transitory computer-readable recording medium for storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method comprises: for the storage of an image or set of images encoded or decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, determining a media file format based on the High Efficiency Image File Format (HEIF) specification; and generating a bitstream based on the determination.

[0438] 22. A method for storing a bitstream of video, comprising: for the storage of an image or set of images encoded or decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, determining a media file format based on a High Efficiency Image File Format (HEIF) specification; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0439] 23. A method, apparatus, or system described in this document.

[0440] The following solutions illustrate further examples of the techniques discussed in this article.

[0441] 1. A method for processing media data, comprising: for storage of an image or collection of images encoded or decoded using the Joint Image Experts Group Artificial Intelligence (JPEGAI) image codec, determining that the media file format is based on the High Efficiency Image File Format (HEIF) specification; and performing a conversion between visual media data and a bitstream based on the JPEGAI image codec.

[0442] 2. The method according to Solution 1, wherein the specific type of encoded image item represented as a JPEG AI image item is a JPEG AI bitstream and carries an encoded image.

[0443] 3. The method according to any one of solutions 1-2, wherein each bitstream presented to the JPEG AI decoder is the content of a specific type of encoded image item.

[0444] 4. The method according to any one of solutions 1-3, wherein the specific type is 'jai0'.

[0445] 5. The method according to any one of solutions 1-4, wherein the auxiliary image of the JPEG AI encoding / decoding uses a specific item_type value.

[0446] 6. The method according to any one of solutions 1-5, wherein the specific item_type value is 'jai0'.

[0447] 7. The method according to any one of solutions 1-6, wherein a specific file brand is specified for JPEG AI encoded images and image sets.

[0448] 8. The method according to any one of solutions 1-7, wherein the specific file brand is 'jaii'.

[0449] 9. The method according to any one of solutions 1-8, wherein the encoded image item conforms to 'jaii' when the encoded image item has a type 'jai0' associated with the JPEG AI header item attribute.

[0450] 10. The method according to any one of solutions 1-9, wherein, in addition to 'ispe', 'pixi', 'jxpl', 'colr', 'rloc', 'irot', 'clap' and 'imir', the codec image item conforms to 'jaii' when it is not associated with any other type of basic item attribute.

[0451] 11. The method according to any one of solutions 1-10, wherein a media subtype of media type 'image' is specified for JPEG AI encoded and decoded images and image sets.

[0452] 12. The method according to any one of solutions 1-12, wherein the media subtype is 'jaii'.

[0453] 13. The method according to any one of solutions 1-12, wherein the presence of an image item of type 'jaii' is transmitted by including an item description in the itemtypes parameter, the item type string of which begins with 'jai0', followed by a period, and further followed by a series of period-separated values.

[0454] 14. The method according to any one of solutions 1-13, wherein the dot-separated values ​​comprise a subset of the information carried in the JPEG AI header attribute.

[0455] 15. The method according to any one of solutions 1-14, wherein each value separated by a dot is encoded or decoded into a hexadecimal number.

[0456] 16. The method according to any one of solutions 1-15, wherein the conversion includes encoding visual media data into a bitstream.

[0457] 17. The method according to any one of solutions 1-15, wherein the conversion includes decoding visual media data from a bitstream.

[0458] 18. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method described in any one of solutions 1-17.

[0459] 19. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec device performs the method described in any one of solutions 1-17.

[0460] 20. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method comprises: for the storage of an image or set of images encoded or decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, determining a media file format based on the High Efficiency Image File Format (HEIF) specification; and generating a bitstream based on the determination.

[0461] 21. A method for storing a bitstream of video, comprising: for the storage of an image or set of images encoded or decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, determining a media file format based on the High Efficiency Image File Format (HEIF) specification; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0462] In the described solution, the encoder conforms to the format rules by generating a codec representation based on those rules. In the described solution, the decoder parses the syntax elements in the codec representation using known information about their presence or absence, based on the format rules, to produce the decoded video.

[0463] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits at the same position in the bitstream defined by the syntax or bits propagated at different positions. For example, a macroblock can be encoded based on the error residual value after transformation and encoding / decoding, and can also use bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on this determination, knowing whether some fields may be present or absent, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0464] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0465] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0466] The processing and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the devices can be implemented as special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0467] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0468] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in the claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.

[0469] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the specific order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the division of various system components in the embodiments described in this patent document should not be construed as requiring such division in all embodiments.

[0470] Only a few implementations and examples are described, and other implementations, improvements and variations can be made based on what is described and shown in this patent document.

[0471] When there is no intermediary component other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there is an intermediary component other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.

[0472] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be considered illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0473] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as coupled can be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Those skilled in the art can identify examples of other changes, substitutions, and modifications, which can be made without departing from the spirit and scope of this disclosure.< / isobmff>

Claims

1. A method for processing media data, comprising: For the storage of images or collections of images encoded or decoded using the Joint Picture Experts Group Artificial Intelligence (JPEG AI) image codec, the media file format is determined based on the High Efficiency Image File Format (HEIF) specification; and The conversion between visual media data and bitstream is performed based on the JPEG AI image codec.

2. The method as described in claim 1, wherein, A specific type of encoded image item, represented as a JPEG AI image item, is a JPEG AI bitstream and carries an encoded image.

3. The method according to any one of claims 1-2, wherein, Each bitstream presented to the JPEG AI decoder is the content of a specific type of encoded image item.

4. The method according to any one of claims 1-3, wherein, The specific type is 'jai0'.

5. The method according to any one of claims 1-4, wherein, The auxiliary image for JPEG AI encoding and decoding uses a specific item_type value.

6. The method according to any one of claims 1-5, wherein, The specific item_type value is 'jai0'.

7. The method according to any one of claims 1-6, wherein, Specify a file brand for JPEG AI encoded images and image collections.

8. The method according to any one of claims 1-7, wherein, The specific file brand is 'jaii'.

9. The method according to any one of claims 1-8, wherein, When a codec image item has a type 'jai0' associated with a JPEG AI header item attribute, the codec image item conforms to 'jaii'.

10. The method according to any one of claims 1-9, wherein, In addition to 'ispe', 'pixi', 'jxpl', 'colr', 'rloc', 'irot', 'clap', and 'imir', a codec image item conforms to 'jaii' when it is not associated with any other type of basic item attribute.

11. The method according to any one of claims 1-10, wherein, Specifies the media subtype of media type 'image' for JPEG AI encoded and decoded images and image collections.

12. The method according to any one of claims 1-12, wherein, The media subtype is 'jaii'.

13. The method according to any one of claims 1-12, wherein, The presence of an image item of type 'jaii' is signaled by including an item description in the itemtypes parameter, wherein the item type string of the item description begins with 'jai0', followed by a period, and further followed by a series of period-separated values.

14. The method according to any one of claims 1-13, wherein, The dot-separated values ​​include a subset of the information carried in the JPEG AI header attributes.

15. The method according to any one of claims 1-14, wherein, Each value separated by a period is encoded into a hexadecimal number.

16. The method according to any one of claims 1-15, wherein, The conversion includes encoding the visual media data into the bitstream.

17. The method according to any one of claims 1-15, wherein, The conversion includes decoding the visual media data from the bitstream.

18. An apparatus for processing video data, comprising: processor; and a non-transitory memory thereon having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method as described in any one of claims 1-17.

19. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec apparatus performs the method as described in any one of claims 1-17.

20. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: For the storage of images or collections of images encoded or decoded using the Joint Picture Experts Group Artificial Intelligence (JPEG AI) image codec, the media file format is determined based on the High Efficiency Image File Format (HEIF) specification; and The bit stream is generated based on the determination.

21. A method for storing a video bitstream, comprising: For the storage of images or collections of images encoded or decoded using the Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, the media file format is determined based on the High Efficiency Image File Format (HEIF) specification; Based on the determination, a bit stream is generated; as well as The bit stream is stored in a non-transitory computer-readable recording medium.