Signaling and storage of timing image sequences encoded and decoded using JPEG AI in media files

CN122556068APending Publication Date: 2026-08-11DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

随着能够接收和显示视频的连接用户设备的数量增加,对数字视频使用的带宽需求可能继续增长

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122556068A_ABST
    Figure CN122556068A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. This mechanism includes: defining a media file format for storing a sequence of timed images encoded and decoded using a Joint Image Experts Group (JPEG AI) codec; and performing a conversion between visual media data and a bitstream based on this media file format.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This patent application claims the benefit of U.S. Patent Application No. 63 / 621,828, filed January 17, 2024, which is incorporated herein by reference. Technical Field

[0003] This disclosure relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology

[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention

[0005] The first aspect relates to a method for processing media data, comprising: defining a media file format for storing a sequence of timed images encoded and decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) codec; and performing a conversion between visual media data and a bitstream based on the media file format.

[0006] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the JPEG AI codec is specified in International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) 6048-1.

[0007] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that a video track carrying a sequence of timed images encoded and decoded using a JPEG AI image codec is referred to as a JPEG AI video track.

[0008] Alternatively, in any of the above aspects, another implementation of that aspect provides that the media file format is based on the International Organization for Standardization Basic Media File Format (ISOBMFF).

[0009] Alternatively, in any of the above aspects, another implementation of that aspect provides that the media file format is based on the High Efficiency Image File Format (HEIF), wherein HEIF is based on the International Organization for Standardization Basic Media File Format (ISOBMFF).

[0010] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that each sample in the JPEG AI video track carries a JPEG AI codec image.

[0011] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that, for a standards-compliant file containing at least one JPEG AI video track, the file type box uses a specific brand as the major brand, and the specific brand is referred to as 'jai0' or 'jais'.

[0012] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that, for a standards-compliant file containing at least one JPEG AI video track, the file type box includes a specific brand from a list of compatible brands.

[0013] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that, when the particular brand is a compatible brand, the sample entry type of the image sequence track is 'jaim'.

[0014] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the syntax element track_enabled is equal to 1 when the particular brand is a compatible brand.

[0015] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the syntax element track_in_movie is equal to 1 when the particular brand is a compatible brand.

[0016] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that, when a particular brand is a compatible brand, the DataEntryBox to which the value of the syntax element data_reference_index of each sample entry is mapped satisfies (entry_flags & 1) equals 1.

[0017] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that a brand-specific reader is capable of displaying image sequence tracks with sample entry type 'jaim'.

[0018] Alternatively, in any of the above aspects, another implementation of that aspect provides that a brand-specific reader is capable of displaying image sequence tracks with the syntax element track_enabled equal to 1.

[0019] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that a brand-specific reader is capable of displaying image sequence tracks where the syntax element track_in_movie is equal to 1.

[0020] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the value of the Compressorname field includes "\016Motion JPEG AI", where \016 is 14 and the length of the string is one byte.

[0021] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a media subtype that specifies the media type "image" for a JPEG AI codec image sequence carried in an ISOBMFF file or HEIF file.

[0022] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the media subtype of media type 'image' is referred to as 'jais'.

[0023] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the presence of a sample entry of type 'jaim' is transmitted via a signal by including a value in the codec parameters, the first element of which is 'jaim', followed by a period ('.') and then a series of values ​​separated by periods ('.'), wherein the values ​​separated by periods ('.') comprise a subset of the information carried in the JPEG AI header, and each of the values ​​separated by periods ('.') is encoded as a hexadecimal number.

[0024] Alternatively, in any of the above aspects, another implementation of that aspect provides that the conversion includes encoding visual media data into a bitstream.

[0025] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the conversion includes decoding visual media data from a bitstream.

[0026] The second aspect relates to an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of the disclosed embodiments.

[0027] The third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video codec apparatus performs the method of any of the disclosed embodiments.

[0028] The fourth aspect relates to a non-transitory computer-readable recording medium for storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining a media file format for storing a sequence of timed images encoded and decoded using a Joint Picture Experts Group Artificial Intelligence (JPEG AI) codec; and generating a bitstream based on the media file format.

[0029] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining a media file format for storing a sequence of timed images encoded and decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) codec; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0030] The sixth aspect relates to the methods, apparatus, or systems described in this disclosure.

[0031] For clarity, any of the embodiments described above may be combined with any one or more of the other embodiments described above to create new embodiments within the scope of this disclosure.

[0032] These and other features will become clearer from the following detailed description by referring to the accompanying drawings and claims. Attached Figure Description

[0033] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.

[0034] Figure 1 This is a block diagram illustrating an example video processing system.

[0035] Figure 2 This is a block diagram of an example video processing device.

[0036] Figure 3 This is a flowchart of an example method for video processing.

[0037] Figure 4 This is a block diagram illustrating an example video codec system.

[0038] Figure 5 This is a block diagram showing an example encoder.

[0039] Figure 6 This is a block diagram showing an example decoder.

[0040] Figure 7 This is a schematic diagram of an example encoder. Detailed Implementation

[0041] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but modifications can be made within the scope of the appended claims and their equivalents.

[0042] The use of chapter headings in this disclosure is for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this disclosure, edited changes to the text are shown in bold italics (indicating deleted text) and bold (indicating added text) relative to the Multi-Functional Video Codec (VVC) specification and / or the SEI Message (VSEI) standard for encoding and decoding video bitstreams.

[0043] 1. Preliminary Discussion

[0044] This disclosure relates to image file formats encoded and decoded using Joint Image Experts Group Artificial Intelligence (JPEG AI). Specifically, this disclosure relates to the signaling and storage of timing image sequences encoded and decoded using JPEG AI in media files, which may be based on the High Efficiency Image File Format (HEIF) (which in turn is based on the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF)) or directly on ISOBMFF. These ideas can be applied individually or in various combinations for images encoded and decoded by any neural network (NN)-based codec (e.g., JPEG AI (i.e., ISO / IEC 6048-1, Information Technology – Learning-based Image Codec Systems (JPEGAI) – Part 1: Core Codec Systems)) and any image file format (e.g., JPEG AI image sequence file format or motion JPEG AI file format).

[0045] 2. Further discussion

[0046] 2.1 File Format Standards

[0047] Media streaming applications can be based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transmission methods and can rely on file formats such as ISO Basic Media File Format (ISOBMFF) [1]. One such streaming system is HTTP-based Dynamic Adaptive Streaming (DASH) [2]. In order to use video formats with ISOBMFF and DASH, video format-specific file format specifications such as the Advanced Video Codec (AVC) file format and the High Efficiency Video Codec (HEVC) file format [3] will be needed to encapsulate the video content in ISOBMFF tracks as well as DASH representations and segments. Information about the video bitstream (e.g., grade, layer, and level) can be exposed as file format-level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes (e.g., for selecting appropriate media segments, both for initialization at the start of a streaming session and for streaming adaptation during a streaming session).

[0048] Similarly, in order to use the image format with ISOBMFF, image format-specific file format specifications can be used, such as the AVC image file format and HEVC image file format in [4].

[0049] 2.2. Image and Video Encoding and Decoding Based on Neural Networks (NN)

[0050] Deep learning has made rapid progress in various fields, especially in computer vision and image processing. Inspired by the tremendous success of deep learning in computer vision, many researchers have shifted their attention from image / video compression techniques to neural image / video compression. Neural networks are designed through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in nonlinear transformations and classification. Significant progress has been made in neural network-based image / video compression techniques. Examples of neural network-based image compression algorithms have reportedly achieved rate-distortion (RD) performance comparable to Multifunctional Video Coding (VVC), a video codec standard developed by the Joint Video Experts Group (JVET), comprised of experts from the Moving Picture Experts Group (MPEG) and the Video Codec Experts Group (VCEG). With the continuous improvement of neural image compression performance, neural network-based video compression has become an actively developing research area. However, due to the inherent difficulty of the problem, neural network-based video coding and decoding is still in its early stages.

[0051] 2.2.1 Image / Video Compression

[0052] Image / video compression (also known as image / video encoding / decoding) generally refers to the computational technique of compressing images or videos into binary code for convenient storage and transmission. Binary code may or may not support lossless reconstruction of the original image or video; this is called lossless compression and lossy compression. Since lossless reconstruction is not necessary in most cases, most efforts focus on lossy compression. The performance of image or video compression algorithms is typically evaluated from two aspects: compression ratio and reconstruction quality. The compression ratio is directly related to the amount of binary code; less is better. Reconstruction quality is measured by comparing the reconstructed image or video with the original image or video; higher is better.

[0053] Image / video compression techniques can be divided into two branches: classical video encoding / decoding methods and neural network-based video compression methods. Classical video encoding / decoding schemes employ transform-based solutions, where researchers model dependencies in the quantization domain through carefully hand-designed entropy encoding / decoding, thereby leveraging statistical dependencies in latent variables such as Discrete Cosine Transform (DCT) or wavelet coefficients. Neural network-based video compression has two approaches: neural network-based encoding / decoding tools and end-to-end neural network-based video compression. The former is embedded as an encoding / decoding tool within a classical video codec, serving only as part of the framework; while the latter is a standalone framework developed based on neural networks without relying on a classical video codec.

[0054] A series of classic video codec standards have been developed to accommodate the ever-increasing amount of visual content. The International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) has two expert groups, the Joint Group of Picture Experts (JPEG) and the Moving Picture Experts Group (MPEG), and the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) also has its own Video Codec Experts Group (VCEG) for standardizing image or video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, began working on the Multi-Functional Video Codec (VVC) standard. The first version of VVC was released in July 2020. It has been reported that, compared to HEVC, VVC reduces the bit rate by an average of 50% while maintaining the same visual quality.

[0055] Many researchers have focused on neural network-based image encoding and decoding for use in neural network-based image / video compression. However, the network architectures used in example designs are relatively shallow, resulting in unsatisfactory performance. Thanks to abundant data and powerful computing resources, neural network-based methods have been better utilized in a variety of applications. Currently, neural network-based image / video compression has shown promising improvements and demonstrated its feasibility. However, this technology is still far from mature, and many challenges remain to be solved.

[0056] 2.2.2 Neural Networks

[0057] Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. They typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform it into different kinds of representations. Note that these representations are not manually designed; instead, they are learned from massive amounts of data using general machine learning procedures. Deep learning eliminates the need for manually designed representations and is therefore considered particularly suitable for processing raw, unstructured data, such as acoustic and visual signals, which has been a long-standing challenge in the field of artificial intelligence.

[0058] 2.2.3. Neural Networks for Image and Video Compression

[0059] Example neural networks used in image compression methods can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding / decoding strategies, while the latter is a transform-based solution. Sometimes, these two methods are combined.

[0060] Similar to classic video encoding and decoding techniques, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology lagged behind that of neural network-based image compression, but due to its complexity, it requires more effort to overcome its challenges. Some researchers have focused on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to eliminate inter-frame redundancy. Inter-frame prediction is a key step in these works. Motion estimation and compensation have been employed, but only recently have they been implemented using trained neural networks.

[0061] Video compression research based on neural networks can be divided into two categories according to the target scenario: random access and low latency. In the case of random access, decoding can start from any point in the sequence, the entire sequence is divided into multiple individual segments, and each segment can be decoded independently. The low latency case aims to reduce decoding time, so that earlier frames can be used as reference frames to decode subsequent frames.

[0062] 2.2.4. JPEG AI Image Encoding and Decoding Standard

[0063] At the time of writing (January 17, 2024), the JPEG AI image codec standard is an image codec standard standardized by the JPEG Working Group (WG), which is WG 1 of ISO / IEC JTC 1 SC 29. The ISO / IEC number for the JPEG AI standard is ISO / IEC 6048. The latest JPEG AI draft specification is contained in the JPEG output document WG1N100660.

[0064] The latest JPEG AI draft specification utilizes some of the neural network-based image encoding and decoding methods mentioned above. The following describes or summarizes some features of the latest JPEG AI specification and possible methods for grade and level signaling. The section numbers in parentheses are the same as those in document WG1N100660.

[0065] 2.2.4.1. (9.2) Stream Layout ...

[0067] The overall grammatical structure of the image is as follows:

[0068] Each bitstream begins with a 16-bit marker. All markers used in this specification are as follows: ...

[0070] 2.2.4.2. (9.3) Image header

[0071] This substream contains information about the image height. ,width Potential space piece location and size, control flags for each tool, scaling factors for primary and secondary components, - Learnable model index and displacement (primary components) for rate control parameters Secondary components are (information).

[0072] 2.2.4.2.1 (9.3.1) Syntax Table ...

[0074] 2.2.4.2.2. (9.3.1.2) Grading and Level Syntax ...

[0076] 2.2.4.2.3. (9.3.2) Image header semantics

[0077] The following service information is transmitted via signal: picture_header_size is the number of bytes in the image header excluding the first two bytes of the marker; img_width plus 64 specifies the width of the input image (from 64 to 65600). img_height plus 64 specifies the height of the input image (from 64 to 65600); picture_format is the data format of the output image (YUV420 = 0, YUV444 = 1, sRGB = 2, YUV422 = 3). bit_depth is the bit depth of the output image ("0" corresponds to 8, "1" corresponds to 10); ... stream_profile_idc indicates the stream profile that the bitstream conforms to; The increment of 1 in num_decoder_profiles_minus1 specifies the number of decoder profiles supported by the bitstream; decoder_profile_idc[i] indicates the i-th supported decoder profile provided by the bitstream; level_idc indicates the level that the bitstream conforms to. ...

[0079] 3. The technical problem solved by the disclosed technical solution

[0080] There is a lack of design for signaling and storage of timed image sequences using JPEG AI encoding and decoding in media files.

[0081] 4. List of solutions and implementation examples

[0082] To address the aforementioned issues, the methods outlined below are disclosed. These aspects should be considered as examples for interpreting general concepts, and not interpreted in a narrow sense. Furthermore, these examples can be applied individually or in any combination.

[0083] 1) In one example, a media file format is specified for storing a sequence of timing images encoded and decoded using the JPEG AI image codec specified in ISO / IEC 6048-1, and a video track carrying the sequence of timing images encoded and decoded using the JPEG AI image codec is called a JPEG AI video track.

[0084] a. In one example, the media file format is specified based on the ISO Basic Media File Format (ISOBMFF).

[0085] b. In one example, the media file format is defined based on the High Efficiency Image File Format (HEIF), which in turn is based on the ISO Basic Media File Format (ISOBMFF).

[0086] 2) In one example, it is specified that each sample in the JPEG AI video track carries a JPEG AI codec image.

[0087] 3) In one example, it is specified that for a standard-compliant file containing at least one JPEG AI video track, the file type box must use a specific brand (e.g., 'jai0' or 'jais') as the major brand, or include the specific brand in the list of compatible brands.

[0088] a. In one example, it is specified that when a specific brand is in a compatible brand, there must be an image sequence track with sample entry type 'jaim', track_enabled equal to 1, track_in_movie equal to 1, and each sample entry has a data_reference_index value, so that it is mapped to a DataEntryBox with (entry_flags & 1) equal to 1.

[0089] b. In one example, it is specified that a particular brand of reader must be able to display image sequence tracks with sample entry type 'jaim', track_enabled equal to 1, and track_in_movie equal to 1.

[0090] 4) In one example, a specific sample entry type (also known as a sample entry name), such as 'jaim', is specified for use by motion JPEG AI video tracks.

[0091] a. In one example, a sample entry of a specific sample entry type includes a configuration box that contains at least one or more of the following information: i. Information regarding the stream profile of the JPEG AI codec image carried in the sample associated with the sample entry.

[0092] ii. Information regarding the decoder profile provided by the JPEG AI-encoded images carried in the samples associated with the sample entries.

[0093] iii. Information regarding the level of conformity of the JPEG AI codec image carried in the sample associated with the sample entry.

[0094] iv. Information regarding the color format of the output image generated from the JPEG AI-encoded image carried in the sample associated with the decoded sample entry.

[0095] v. Information about the bit depth of the output image generated from the JPEG AI-encoded image carried in the sample associated with the sample entry.

[0096] vi. Information about the frame rate of the sequence of JPEG AI-encoded images carried in the samples associated with the sample entry.

[0097] vii. Information about the bit rate of the sequence of JPEG AI-encoded images carried in the samples associated with the sample entry.

[0098] 5) In one example, the recommended value for the Compressorname field is specified as "\016Motion JPEGAI", where \016 is 14, which is the string length in bytes.

[0099] 6) In one example, a media subtype of media type 'image' is specified for a JPEG AI codec image sequence carried in an ISOBMFF file or HEIF file, for example named 'jais'.

[0100] a. In one example, the existence of a sample entry of type 'jaim' is specified by signaling in the following manner: a value is included in the codec parameters, the first element of which is 'jaim', followed by a period ('.'), and then a series of values ​​separated by periods ('.'), wherein the values ​​separated by periods ('.') include a subset of the information carried in the JPEG AI header item attributes as specified by item 4)a and its sub-items above, and each value is encoded as a hexadecimal number.

[0101] 5. Examples

[0102] The following are some example implementations of the aspects summarized in Section 4, which can be applied to potential standard specifications for the JPEG AI image sequence file format and / or the motion JPEG AI file format.

[0103] 5.1 First Embodiment

[0104] This example applies to all the projects summarized in Section 4.

[0105] 1. Scope

[0106] This disclosure specifies a container file format for JPEG AI streams as defined in ISO / IEC 6048-1. It defines a file format for processing timed image sequence (also known as motion picture sequence) files on a computer platform, allowing for internet-based communication and other communications.

[0107] This disclosure uses existing specifications for the file format and extends them for embedding JPEG AI streams.

[0108] 2. Standardized Citations

[0109] The following documents are cited in such a way that some or all of them constitute a requirement of this disclosure. For dated citations, only the cited version applies. For undated citations, the latest version of the cited document (including any revisions) applies.

[0110] ISO / IEC 14496-12 , Coding of audio-visual objects - Part 12: ISO base media file format

[0111] ISO / IEC 23008-12:2017 ,Information technology-High efficiency coding and media delivery in heterogeneous environments - Part 12: Image File Format

[0112] ISO / IEC 6048-1, Information technology - Learning-based image coding system (JPEG AI) - Part 1: Core coding system

[0113] ISO / IEC 6048-2, Information technology - Learning-based image coding system (JPEG AI) - Part 2: Profiling

[0114] Rec. ITU-T H.273 | ISO / IEC 23091-2, Coding-independent code points -Part 2: Video

[0115] 3. Terms and Definitions

[0116] For the purposes of this disclosure, the terms and definitions given in ISO / IEC 14496-12, ISO / IEC 6048-1, ISO / IEC 6048-2, ISO / IEC 23008-12 and below shall apply.

[0117] ISO and IEC maintain a terminology database for standardization, available at the following URL: - ISO online browsing platform: available at https: / / www.iso.org / obp - IEC Electronics Encyclopedia: Available at http: / / www.electropedia.org / 3.1

[0119] box

[0120] A structured collection of data describing an image or the image decoding process. 3.2

[0122] Box type

[0123] The types of information stored together with box (3.1) 3.3

[0125] byte

[0126] 8-bit group 3.4

[0128] Code points that are independent of encoding and decoding

[0129] The code points of the color space are defined based on enumeration values.

[0130] Note 1 to item: Code points defined in Recommendation ITU-T H.273 | ISO / IEC 23091-2. 3.5

[0132] High-efficiency image file format

[0133] Image file formats that can embed still images and motion sequences (3.7)

[0134] Note 1 to item: Based on ISO / IEC 23008-12. 3.6

[0136] Image collection

[0137] An unordered collection of images without implicit or signal-transmitted presentation order or presentation timestamps. 3.7

[0139] motion sequence

[0140] Film

[0141] Timing sequence of images (3.9) 3.8

[0143] sample

[0144] <isobmff>All data associated with a single time

[0145] Note 1 to entry: This definition is used in Appendices B and C as data associated with a codec image in the sequence. 3.9

[0147] Timing sequence

[0148] A linearly ordered sequence of media entities (such as images), where each entity is presented at a well-defined timestamp.

[0149] 4. Abbreviations

[0150] For the purposes of this disclosure, the abbreviations given in ISO / IEC 14496-12, ISO / IEC 6048-1, ISO / IEC 6048-2, ISO / IEC 23008-12, and the following abbreviations shall apply.

[0151] CICP code points that are independent of encoding and decoding

[0152] HEIF High-Efficiency Image File Format

[0153] ISOBMFF ISO Basic Media File Format

[0154] 5. Naming conventions for numerical values

[0155] Integers are represented as bit patterns, hexadecimal values, or decimal numbers. Bit patterns and hexadecimal values ​​have both numerical values ​​and an associated specific length in bits.

[0156] Hexadecimal representation (signed by prefixing a hexadecimal number with "0x") can be used in place of binary representation to represent bit patterns that are multiples of 4 in length. For example, 0x41 represents an octet pattern where only its second most significant bit and least significant bit are equal to 1. The values ​​specified under the "Encoding / Decoding" heading in a table called the "encoding / decoding table" are bit pattern values ​​(defined as strings of digits equal to 0 or 1, where the leftmost bit is considered the most significant bit). Other values ​​without the "0x" prefix are decimal values. When used in expressions, hexadecimal values ​​are interpreted as having the value of the corresponding bit pattern, evaluated as the binary representation of an unsigned integer (i.e., as the value of the number formed by prefixing the bit pattern with a sign bit equal to 0 and interpreting the result as the two's complement representation of the integer value). For example, the hexadecimal value 0xF is equivalent to the 4-bit pattern '1111' and is interpreted in expressions as equal to the decimal number 15.

[0157] 6. Consistency

[0158] This public disclosure shares a general definition of file structure (a sequence of objects, referred to herein as a box, or an atom in other similar file formats) and a general definition of the general structure (size and type) of objects.

[0159] File formats representing images or image sequences must conform to the specifications in Appendices A and B. All of these specifications require readers to ignore objects they cannot recognize.

[0160] In any event of discrepancies or conflicts, this disclosure takes precedence over those documents on which it is based; however, no such conflicts are currently known.

[0161] For better readability and understanding, the syntax descriptions of different file formats are presented in the same way as in the basic format.

[0162] 7. Color Standards

[0163] JPEG AI (as defined in ISO / IEC 6048-1) describes only the encoded bitstream of an image. For the image to be displayed or interpreted correctly, it is crucial that the color space of the image data be correctly represented. For this purpose, the corresponding container file format must transmit the correct color space via signaling. The format defined for JPEG AI in this disclosure transmits the color space as specified in Recommendation ITU-T H.273 | ISO / IEC 23091-2 via signaling.

[0164] 8. The organization disclosed herein

[0165] Appendix A specifies the integration of JPEG AI streams into ISOBMFF (as defined in ISO / IEC 14496-12) to use image sequences as movies in the file format.

[0166] Appendix B specifies the integration of JPEG AI bitstreams into HEIF file formats (as defined in ISO / IEC 23008-12), thereby allowing the integration of JPEG AI encoded and decoded image sequences.

[0167] Appendix A (Standardized) Use of JPEG AI Bitstream in ISOBMFF - Motion JPEG AI

[0168] A.1 Overview

[0169] This appendix specifies the use of JPEG AI encoding and decoding for timing sequences of images within files based on the ISO Basic Media File Format (defined in ISO / IEC 14496-12), denoted as Motion JPEG AI. The Motion JPEG AI file format is designed to contain one or more motion sequences of JPEG AI compressed images and their timing. It is intended to serve as a 'building block,' specifying only the video format. Applications will be expected to incorporate Motion JPEG AI with appropriate audio, metadata, etc., for a complete application specification; this specification typically selects the Motion JPEG AI profile and level, and may also specify the application profile and level suitable for integration.

[0170] Motion JPEG AI is expected to be used in a variety of applications, especially where JPEG AI encoding / decoding technology is already available for other reasons, or where a high-quality frame-based approach without inter-frame encoding / decoding is suitable. These application areas include: - Digital still camera, - Error-prone environments such as wireless and internet, - Video capture, - High-quality digital video recording for professional broadcasting and film production (from film systems to digital systems). - as well as high-resolution medical and satellite imaging.

[0171] Motion JPEG AI is a flexible format that allows for a wide range of uses, such as editing, displaying, exchanging, and streaming.

[0172] Note that in the context of ISOBMFF (ISO / IEC 14496-12), a sample is "all data associated with a single time." In this appendix, it refers to data associated with a single encoded image, not "pixels."

[0173] A.2 Compatibility and Technical Derivation

[0174] A.2.1 Series Members

[0175] This is a 'building block' specification; it defines how to store motion JPEG AI sequences in a file format based on the ISO Basic Media File Format. It is one of a series of specifications with a common format.

[0176] Since this is a building block specification, if audio is required, appropriate audio support should be selected from other specifications that use the ISO Basic Media File Format (ISO / IEC 14496-12).

[0177] These specifications share a common definition of file structure (a sequence of objects, referred to here as a box, or an atom in other similar file formats) and a common definition of the general structure of objects (size and type).

[0178] All of these specifications require readers to ignore objects they cannot recognize.

[0179] In any case where differences or conflicts exist, this specification takes precedence over those specifications on which it is based; however, no such conflicts are currently known.

[0180] A.2.2. Consistency

[0181] The implementation of the Motion JPEG AI decoder must support the decoding of video tracks using JPEG AI encoding and decoding technology. Files conforming to this specification must contain at least one Motion JPEG AI video track, and the file type box must list 'jai0' as the major brand, or include 'jai0' as a brand in the compatibility list. Additional brands can be defined based on derivations and applications of this specification.

[0182] A.3 Sample Entries and Sample Format of Motion Sequences

[0183] A.3.1 Overview

[0184] The sample entries and sample formats of the JPEG AI stream in ISOBMFF are derived from the syntax in ISO / IEC 14496-12 and are defined in A.3.2 to A.3.5.

[0185] A.3.2 Definition

[0186] Sample entry type: 'jaim'

[0187] Container: Sample description box ('stsd')

[0188] Mandatory: Yes

[0189] Quantity: There can be one or more sample entries.

[0190] Box type: 'jaiC'

[0191] Container: Motion JPEG AI sample entry ('jaim')

[0192] Mandatory: Yes

[0193] Quantity: One

[0194] When the sample entry name is 'jaim', the sample (3.8) is in the format of a JPEG AI stream, which carries an encoded and decoded image as defined in ISO / IEC 6048-1.

[0195] Each image presented to the JPEG AI decoder is the content of a sample.

[0196] The VisualSampleEntry, its constituent boxes, and the values ​​present in the bitstream described by these boxes must be consistent within the allowed range of field format and precision. This consistency includes, but is not limited to, width and height information and resolution declarations (within the allowed precision of different representations). Conflicting files are non-compliant with the standard, and the reader may attempt to determine which values ​​are correct or reject the file.

[0197] The width and height fields in Visual Sample Entry indicate the highest resolution component of the image (which is usually, but not required, the brightness component in an image where not all components have the same spatial sampling density).

[0198] If the encoded / decoded image contains an alpha plane, then an appropriate value for 'depth' must be used, as indicated by the Visual Sample Entry.

[0199] Color information can be provided in one or more ColourInformationBoxes. These should be placed sequentially in the sample entries, starting with the most accurate (and likely the most expensive to process) and ending with the least accurate. These are advisory and pertain to rendering and color conversion, and there is no standardized behavior associated with them; the reader can choose to use the most appropriate one. ColourInformationBoxes with unknown color types can be ignored. The values ​​of the colour_type field, other than those documented here, are reserved.

[0200] ColourInformationBox is specific to VideoSampleEntry as defined in ISO / IEC 14496-12.

[0201] A.3.3 Grammar

[0202] / / Visual sequence

[0203] class JAIMSampleEntry() extends VisualSampleEntry ('jaim'){

[0204] JAIMConfigurationBox();

[0205] ColourInformationBox(); / / As defined in ISO / IEC 14496-12

[0206] }

[0207] class JAIMConfigurationBox extends Box('jaiC') {

[0208] unsigned int(8) configurationVersion = 1;

[0209] unsigned int(8) stream_profile_idc;

[0210] unsigned int(8) num_decoder_profiles_minus1;

[0211] for (i=0; i <= num_decoder_profiles_minus1; i++)

[0212] [[ID=2۳]]unsigned int(8) decoder_profile_idc[i];

[0213] unsigned int(8) level_idc

[0214] bit(2) reserved = '11'b;

[0215] unsigned int(2) colour_format_idc;

[0216] unsigned int(3) bit_depth_idc;

[0217] unsigned int(1) constantFrameRate;

[0218] unsigned int(16) avgFrameRate;

[0219] BitRateBox(); / / Optional

[0220] }

[0221] A.3.4 Semantics

[0222] In JAIMConfigurationBox(): - stream_profile_idc, num_decoder_profiles_minus1, decoder_profile_idc[i], level_idc, colour_format_idc, and bit_depth_idc contain values ​​that match the fields stream_profile_idc, num_decoder_profiles_minus1, decoder_profile_idc[i], level_idc, picture_format, and bit_depth_idc as defined in ISO / IEC 6048-1, for each JPEG AI bitstream (hereinafter referred to as "stream") carried in the sample to which the sample entry of this configuration box applies.

[0223] A constantFrameRate value of 1 indicates that the stream has a constant frame rate. A value of 0 indicates that the stream may or may not have a constant frame rate.

[0224] - avgFrameRate gives the average frame rate of the stream, in frames per (256 seconds). A value of 0 indicates an unspecified average frame rate.

[0225] In the Visual Sample Entry: - It is recommended, but not mandatory, to set the value of Compressorname to "\016Motion JPEG AI" (\016 is 14, which is the string length in bytes). - depth takes one of the following values; other values ​​are reserved, and if found, the composition behavior is undefined. 0x18 - The image is a color image without alpha. 0x28 - The image is a color image with alpha.

[0226] Appendix B (Standardized) Use of JPEG AI Encoded Image Sequences in HEIF Image File Format

[0227] B.1 Overview

[0228] This appendix specifies a format for encapsulating JPEG AI-encoded image sequences in the HEIF image file format as defined in ISO / IEC 23008-12. The brand of the image sequence is specified in B.3.

[0229] Note that in the context of ISOBMFF (ISO / IEC 14496-12), a sample is "all data associated with a single time." In this appendix, it refers to data associated with a single encoded image, not "pixels."

[0230] B.2 JPEG AI image sequence

[0231] B.2.1. Overview

[0232] Item B.2 specifies the requirements for all files containing one or more JPEG AI codec image sequence tracks. The requirements specified in B.2 must be followed when the branding specified in sub-item B.3.2 is among the compatible brandings of the file.

[0233] Item 7 is subject to the standard of ISO / IEC 23008-12:2017.

[0234] B.2.2 Derived from ISO / IEC 14496-12

[0235] As defined in A.3, sample entries of type 'jaim' must be used for image sequence tracks encoded in JPEG AI, using JAIMSampleEntry() and the sample format specified in A.3.

[0236] For a track containing JPEG AI image sequences, all samples (3.8) are synchronization samples.

[0237] B.3 JPEG AI Specific Brand

[0238] B.3.1 JPEG AI Image Sequence Brand

[0239] B.3.1.1 Overview

[0240] The brand 'jais' is specified in the following sub-entries.

[0241] B.3.1.2 Requirements for HEIF files

[0242] The file must include 'msf1' in the compatible branding to comply with the specifications in ISO / IEC 23008-12:2017, A.3.1.1. Additionally, the file must comply with the specifications in B.2. For at least one image sequence track conforming to the specifications in B.2, the value of track_enabled must be equal to 1, and the value of track_in_movie must be equal to 1.

[0243] When the 'jais' brand is in a compatible brand, there must be image sequence tracks with sample entries of type 'jaim', track_enabled equal to 1, and track_in_movie equal to 1, and each sample entry must have a data_reference_index value so that it is mapped to a DataEntryBox with (entry_flags & 1) equal to 1.

[0244] B.3.1.3 Requirements for HEIF Readers

[0245] The requirements for the reader specified in ISO / IEC 23008-12:2017, A.3.1.2 must be met.

[0246] Readers with the 'jais' brand must be able to display image sequence tracks with sample entry type 'jaim', track_enabled equal to 1, and track_in_movie equal to 1.

[0247] The reader must support all values ​​allowed for the matrix syntax elements of TrackHeaderBox according to ISO / IEC 23008-12:2017, 7.2.1, and must comply with the CleanApertureBox of visual sample entries when displaying image sequence tracks of sample entry type 'jaim'.

[0248] In other words, the reader is required to support rotations of 0, 90, 180, and 270 degrees controlled by matrix syntax elements, as well as clipping controlled by CleanApertureBox.

[0249] It should support displaying image sequence tracks with opacity information, which can be specified as part of the JPEGAI bitstream or via an associated auxiliary track with aux_track_type equal to urn:mpeg:hevc:2015:auxid:1.

[0250] B.4 JPEG AI Encoded Image Sequences in ISO / IEC 23008-12 Image File Media Type Registration

[0251] B.4.1 Overview

[0252] File extensions and media types derived from the ISO Basic Media File Format typically reflect the major brand in the FileTypeBox. When the major brand indicates a brand relevant to sub-entry B.4.2 (Image Sequences), the media type defined here should be used. The media type may also be used when such a brand is a compatible brand. Sub-entry B.4.2 follows the Internet Engineering Task Force (IETF) Request for Comment (RFC) 6838 to provide media type registration.

[0253] B.4.2 Registration

[0254] Media type name: Image

[0255] Media subtype name: jais

[0256] Required parameters: None

[0257] Optional parameters: Same as the media type image / heif. The presence of a sample entry of type 'jaim' is transmitted via signaling in the following way: a value is included in the codec parameters, the first element of which is 'jaim', followed by a period ('.'), and then a series of period ('.') separated values ​​from JAIMConfigurationBox() (as specified in entry A.3 of ISO / IEC 6048-5, starting from stream_profile_idc up to and including level_idc, each value being encoded as a hexadecimal number).

[0258] Encoding Notes: Binary

[0259] Note: None

[0260] Safety Precautions: See the media type image / heif. Additionally, sample entries of type 'jaim' contain variable-length structures and have extensible syntax. Both of these aspects introduce potential security risks to the implementation. In particular, the variable-length structures are susceptible to buffer overflows, and the extensible syntax could lead to malicious operations.

[0261] Interoperability Considerations: Similar to the media type image / heif. Additionally, sample entries of type 'jaim' may conform to one of several gradations and / or require one of several capabilities (e.g., as specified in ISO / IEC 6048-2), and not all gradations and / or capabilities are necessarily supported by the receiving decoder. Therefore, the decoder may attempt to process this content but ultimately find that it cannot be rendered partially or completely.

[0262] Published standard: ISO / IEC 6048-5 Information technologies-Learning-based Image Coding - Part 5: JPEG AI file formats

[0263] Applications: Multimedia, Imaging, Images, Science

[0264] Note on fragment identifiers: Same as media type image / heif

[0265] Usage restrictions: None

[0266] Additional Information: Deprecated alias: N / A (Multiple) Magic Numbers: None (Multiple) file extensions: jais (Multiple) Macintosh file type codes: N / A Object identifier: N / A Intended use: General Note: None

[0267] 7. References

[0268] [1] IETF RFC 6838, Media Type Specifications and RegistrationProcedures.

[0269] [2] ISO / IEC 14496-12: "Information technology - Coding of audio-visual objects - Part 12: ISO base media file format".

[0270] [3] ISO / IEC 23009-1: "Information technology - Dynamic adaptive streaming over HTTP (DASH) - Part 1: Media presentation description and segment formats".

[0271] [4] ISO / IEC 14496-15: "Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unitstructured video in the ISO base media file format".

[0272] [5] ISO / IEC 23008-12: "Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image FileFormat".

[0273] Figure 1 This is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0274] System 4000 may include an encoding component 4004 capable of implementing the various encoding / decoding or encoding methods described in this disclosure. Encoding component 4004 may reduce the average bit rate from the video input 4002 to the output of encoding component 4004 to produce an encoded representation of the video. Encoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding component 4004 may be stored or transmitted via a communication connection as indicated by component 4006. The stored or communicatively transmitted bitstream (or encoded) representation of the video received at input 4002 may be used by component 4008 to generate pixel values ​​or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "encoding" operations or tools, it is understood that encoding tools or operations are used by encoders, and corresponding decoding tools or operations that reverse the encoded result will be performed by decoders.

[0275] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), etc. The technologies described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0276] Figure 2 This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 can be configured to implement one or more methods described herein. The memories(multiple) 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.

[0277] Figure 3 This is a flowchart of an example method 4200 for video processing. In step 4202, method 4200 determines and specifies a media file format for storing a sequence of timed images encoded and decoded using a Joint Picture Experts Group (JPEG AI) codec. In step 4204, a conversion between visual media data and a bitstream is performed based on the media file format. The conversion may include encoding at the encoder, decoding at the decoder, or a combination thereof.

[0278] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions cause the processor to execute method 4200 when executed by the processor. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video codec device executes method 4200.

[0279] Figure 4 This is a block diagram illustrating an example video encoding / decoding system 4300 from which the techniques of this disclosure can be utilized. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310, and this target device 4320 may be referred to as a video decoding device.

[0280] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.

[0281] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320 or may be external to target device 4320, wherein target device 4320 may be configured to interface with an external display device.

[0282] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.

[0283] Figure 5 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 4 The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0284] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414.

[0285] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0286] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.

[0287] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.

[0288] The mode selection unit 4403 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding), for example, based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).

[0289] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.

[0290] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0291] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0292] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0293] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0294] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.

[0295] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0296] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0297] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0298] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0299] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.

[0300] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0301] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0302] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the buffer 4413.

[0303] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0304] Entropy encoding unit 4414 can receive data from other functional components of video encoder 4400. When entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0305] Figure 6 This is a block diagram illustrating an example of a video decoder 4500. The video decoder 4500 can be... Figure 4 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0306] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is the overall inversion of the encoding process described with respect to the video encoder 4400.

[0307] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.

[0308] The motion compensation unit 4502 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.

[0309] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of a video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate the prediction block.

[0310] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0311] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 performs inverse quantization (i.e., dequantization) on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.

[0312] The reconstruction unit 4506 can add the residual block to the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0313] Figure 7 This is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image and reduce the mean square error between the original and reconstructed samples by adding an offset and by applying a finite impulse response (FIR) filter, respectively. The side information of the encoding and decoding is transmitted via the offset and filter coefficients. ALF 4606 is located in the last processing stage of each image and can be thought of as a tool to attempt to capture and repair artifacts caused by previous stages.

[0314] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference images obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy coding component 4618. The entropy coding component 4618 entropy-codes the prediction results and the quantized transform coefficients and transmits them toward a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.

[0315] The following is a list of some preferred solutions.

[0316] The following solutions illustrate examples of the techniques discussed in this article.

[0317] 1. A method for processing media data, comprising: determining that a media file format is specified for storing a sequence of timed images encoded and decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, and that a video track carrying the sequence of timed images encoded and decoded using the JPEG AI image codec is referred to as a JPEG AI video track; and performing a conversion between visual media data and a bitstream based on the JPEG AI image codec.

[0318] 2. The method according to Solution 1, wherein the media file format is defined based on the ISO Basic Media File Format (ISOBMFF).

[0319] 3. The method according to Solution 1, wherein the media file format is defined based on the High Efficiency Image File Format (HEIF), wherein the HEIF is based on the ISO Basic Media File Format (ISOBMFF).

[0320] 4. The method according to any one of solutions 1-3, wherein each sample in the JPEG AI video track carries a JPEG AI encoded image.

[0321] 5. The method according to any one of solutions 1-4, wherein for a standard-compliant file containing at least one JPEG AI video track, the file type box must designate a specific brand as the major brand, or include said specific brand in the list of compatible brands.

[0322] 6. The method according to any one of solutions 1-5, wherein the specific brand is 'jai0' or 'jais'.

[0323] 7. The method according to any one of solutions 1-6, wherein when the specific brand is in the compatible brand, there must be an image sequence track with sample entry type 'jaim', track_enabled equal to 1, track_in_movie equal to 1, and each sample entry has a data_reference_index value such that it is mapped to a DataEntryBox with (entry_flags& 1) equal to 1.

[0324] 8. The method according to any one of solutions 1-7, wherein the reader of the particular brand shall be able to display image sequence tracks with sample entry type 'jaim', track_enabled equal to 1, and track_in_movie equal to 1.

[0325] 9. The method according to any one of solutions 1-8, wherein a specific sample entry type (also referred to as a sample entry name) is specified for use by motion JPEG AI video tracks.

[0326] 10. The method according to any one of solutions 1-9, wherein the specific sample entry type is 'jaim'.

[0327] 11. The method according to any one of solutions 1-10, wherein the sample entry of the specific sample entry type includes a configuration box, the configuration box including one or more of the following information: Information regarding the stream grade that the JPEG AI codec image carried in the sample associated with the sample entry conforms to; Information regarding the decoder quality provided by the JPEG AI codec image carried in the sample associated with the sample entry; Information regarding the level of conformity of the JPEG AI codec image carried in the sample associated with the sample entry; Information regarding the color format of the output image generated from the JPEG AI codec image carried in the sample associated with the sample entry; Information regarding the bit depth of the output image generated from the JPEG AI codec image carried in the sample associated with the sample entry; Information regarding the frame rate of the sequence of JPEG AI encoded / decoded images carried in samples associated with the sample entry; and Information regarding the bit rate of the sequence of JPEG AI encoded / decoded images carried in the samples associated with the sample entry.

[0328] 12. The method according to any one of solutions 1-11, wherein the recommended value of the Compressorname field is "\016Motion JPEG AI", where \016 is 14, i.e., the string length in bytes.

[0329] 13. The method according to any one of solutions 1-12, wherein a media subtype of media type 'image' is specified for the JPEG AI codec image sequence carried in an ISOBMFF file or HEIF file.

[0330] 14. The method according to any one of solutions 1-13, wherein the media subtype is 'jais'.

[0331] 15. The method according to any one of solutions 1-14, wherein the presence of a sample entry of type 'jaim' is transmitted via signaling by including a value in the codec parameters, the first element of the value being 'jaim', followed by a period ('.') after 'jaim', and further followed by a series of values ​​separated by periods ('.'), wherein the values ​​separated by periods ('.') comprise a subset of the information carried in the JPEG AI header item attributes, each value being encoded as a hexadecimal number.

[0332] 16. The method according to any one of solutions 1-15, wherein the conversion includes encoding the visual media data into the bitstream.

[0333] 17. The method according to any one of solutions 1-15, wherein the conversion includes decoding the visual media data from the bitstream.

[0334] 18. An apparatus for processing video data, comprising: a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-17.

[0335] 19. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec apparatus performs the method according to any one of solutions 1-17.

[0336] 20. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method comprises: determining that a media file format is specified for storing a sequence of timed images encoded and decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, and that a video track carrying the sequence of timed images encoded and decoded using the JPEG AI image codec is referred to as a JPEG AI video track; and generating a bitstream based on the determination.

[0337] 21. A method for storing a bitstream of video, comprising: determining that a media file format is specified for storing a sequence of timed images encoded and decoded using a Joint Image Experts Group Artificial Intelligence (JPEG AI) image codec, and that a video track carrying the sequence of timed images encoded and decoded using the JPEG AI image codec is referred to as a JPEG AI video track; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0338] 22. A method, apparatus or system described in this disclosure.

[0339] In the described solution, the encoder conforms to the format rules by generating an encoded representation based on those rules. In the described solution, the decoder parses the syntax elements in the encoded representation using known information about their presence or absence, based on the format rules, to generate the decoded video.

[0340] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits at co-positions or propagated at different positions in the bitstream defined by the syntax. For example, a macroblock can be encoded based on the error residual value after transformation and encoding / decoding, and can also use bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on this determination, knowing whether some fields may be present or absent, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate the encoded / decoded representation accordingly by including or excluding syntax fields from the encoded / decoded representation.

[0341] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information to be transmitted to a suitable receiver device.

[0342] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the related program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0343] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the apparatus can be implemented as special-purpose logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).

[0344] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0345] While this disclosure contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in the context of individual embodiments in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in the claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.

[0346] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the particular order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be construed as requiring such partitioning in all embodiments.

[0347] Only a few implementations and examples are described, and other implementations, improvements and variations may be made based on what is described and shown in this disclosure.

[0348] When there is no intermediary component other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there is an intermediary component other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.

[0349] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are intended to be illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0350] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings can be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of changes, substitutions, and modifications will be apparent to those skilled in the art upon reference to this disclosure and can be made without departing from the spirit and scope of this disclosure.< / isobmff>

Claims

1. A method for processing media data, comprising: A media file format is defined for storing timed image sequences encoded and decoded using the Joint Image Experts Group Artificial Intelligence (JPEG AI) codec; as well as Perform the conversion between visual media data and bitstream based on the media file format.

2. The method as described in claim 1, wherein, The JPEG AI codec is specified in International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) 6048-1.

3. The method as described in claim 1 or 2, wherein, A video track carrying a sequence of timed images encoded and decoded using the JPEG AI image codec is called a JPEG AI video track.

4. The method according to any one of claims 1-3, wherein, The media file format is based on the International Organization for Standardization Basic Media File Format (ISOBMFF).

5. The method according to any one of claims 1-4, wherein, The media file format is based on the High Efficiency Image File Format (HEIF), which is based on the International Organization for Standardization Basic Media File Format (ISOBMFF).

6. The method according to any one of claims 1-5, wherein, Each sample in the JPEG AI video track carries a JPEG AI encoded image.

7. The method according to any one of claims 1-6, wherein, For a standard-compliant file containing at least one JPEG AI video track, the file type box uses a specific brand as the majorbrand, and the specific brand is referred to as 'jai0' or 'jais'.

8. The method according to any one of claims 1-6, wherein, For standards-compliant files containing at least one JPEG AI video track, the file type box includes a specific brand from the list of compatible brands.

9. The method of claim 7 or 8, wherein, When the specific brand is a compatible brand, the sample entry type for the image sequence track is 'jaim'.

10. The method according to any one of claims 7-9, wherein, When the specific brand is a compatible brand, the syntax element track_enabled equals 1.

11. The method according to any one of claims 7-10, wherein, When the specific brand is a compatible brand, the syntax element track_in_movie equals 1.

12. The method according to any one of claims 7-11, wherein, When the specific brand is a compatible brand, the DataEntryBox mapped to the value of the syntax element data_reference_index for each sample entry satisfies (entry_flags & 1) equals 1.

13. The method according to any one of claims 7-12, wherein, The reader of that specific brand is able to display image sequence tracks with sample entry type 'jaim'.

14. The method according to any one of claims 7-13, wherein, The specific brand of reader is able to display image sequence tracks with the syntax element track_enabled equal to 1.

15. The method according to any one of claims 7-14, wherein, The specific brand of reader is able to display the image sequence track when the syntax element track_in_movie is equal to 1.

16. The method according to any one of claims 1-15, wherein, The value of the Compressorname field includes "\016Motion JPEG AI", where \016 is 14 and the string length is one byte.

17. The method according to any one of claims 1-16, wherein, Specifies the media subtype of 'image' for JPEG AI codec image sequences carried in ISOBMFF or HEIF files.

18. The method of claim 17, wherein, The media subtype of the media type 'image' is called 'jais'.

19. The method according to any one of claims 1 to 18, wherein, The presence of a sample entry of type 'jaim' is transmitted via signaling by including a value in the codec parameters, the first element of which is 'jaim', followed by a period ('.'), and then a series of period ('.') separated values, wherein the period ('.') separated values ​​comprise a subset of the information carried in the JPEG AI header, and each of the period ('.') separated values ​​is encoded as a hexadecimal number.

20. The method according to any one of claims 1-19, wherein, The conversion includes encoding the visual media data into the bitstream.

21. The method according to any one of claims 1-19, wherein, The conversion includes decoding the visual media data from the bitstream.

22. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When executed by the processor, the instructions cause the processor to perform the method as described in any one of claims 1-21.

23. A non-transitory computer-readable medium comprising a computer program product for use with a video encoding / decoding device, wherein, The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video codec device performs the method as described in any one of claims 1-21.

24. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: A media file format is defined for storing timed image sequences encoded and decoded using the Joint Image Experts Group Artificial Intelligence (JPEG AI) codec; and A bitstream is generated based on the media file format.

25. A method for storing a video bitstream, comprising: A media file format is defined for storing timed image sequences encoded and decoded using the Joint Image Experts Group Artificial Intelligence (JPEG AI) codec; Based on the determination, a bit stream is generated; as well as The bit stream is stored in a non-transitory computer-readable recording medium.

26. A method, apparatus or system described in this disclosure.