Icc profile metadata for video streams

CN122700510APending Publication Date: 2026-09-04TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580010819.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-02-24
Filing Date
2025-02-25
Publication Date
2026-09-04

Smart Images

  • Figure CN122700510A_ABST
    Figure CN122700510A_ABST
Patent Text Reader

Abstract

An example method of video decoding includes receiving a video bitstream including a set of pictures, the video bitstream corresponding to a source device. The method also includes identifying color profile metadata for the source device based on a supplemental enhancement information (SEI) message for the video bitstream, and reconstructing the set of pictures using information from the video bitstream. The method also includes causing the set of pictures to be presented at an output device using color characteristics from the color profile metadata.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] priority

[0002] This application is a continuation-to-file of U.S. Patent Application No. 19 / 061,900, filed February 24, 2025, which claims priority to U.S. Provisional Patent Application No. 63 / 558,070, filed February 26, 2024, entitled “ICC Profile Metadata for VideoStreams,” which is incorporated herein by reference in its entirety. Technical Field

[0003] The disclosed implementations generally relate to video coding, including but not limited to systems and methods for incorporating color profile metadata into a video bitstream and using the color profile metadata for video reconstruction. Background Technology

[0004] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital camera devices, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data across communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video encoding can be used to compress video data according to one or more video encoding standards before transmission or storage. Video encoding can be performed by hardware and / or software on electronic / client devices or servers providing cloud services.

[0005] Video coding typically uses prediction methods that leverage the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form using a lower bitrate while avoiding or minimizing video quality degradation. Several video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC released the HEVC / H.265 standard in 2013 (Revision 1), 2014 (Revision 2), 2015 (Revision 3), and 2016 (Revision 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed as a successor to HEVC. ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (Revision 1) and 2022 (Revision 2). AOMedia Video 1 is an open video coding format designed as an alternative to HEVC. On January 8, 2019, a validation version 1.0.0 with errata table 1 was released. Enhanced Compression Model (ECM) is a video coding standard currently under development. ECM aims to significantly improve compression efficiency, surpassing existing standards such as HEVC / H.265 and VVC, essentially enabling higher quality video at lower bitrates.

[0006] Another technique used in video coding standards is Supplemental Enhancement Information (SEI) messages, which enable the carrying of supplementary information to the encoded video within the encoded bitstream. Such SEI information may or may not be directly related to the video coding process, i.e., as specified by the video standard. In some cases, the information in the SEI message is related to application processing that is performed synchronously with or immediately following the video decoding process. Such applications may include rendering processes that use certain SEI messages to adjust the brightness or color space of decoded video frames before they are rendered by a display device. Another such application process arranges portions of the decoded video into a specific pattern defined by SEI messages for, for example, displaying 360-degree video on a head-mounted device. Typically, a wide range of applications can be supported by the information provided in the SEI messages.

[0007] For the H.264, AVC, H.265, and HEVC standards, SEI messages are specified in the main coding specifications, namely "Advanced Video Coding" and "High-Efficiency Video Coding," respectively. For H.266 and VVC, as examples, application-only SEI messages can be specified in a separate specification titled "Versatile supplemental enhancement information messages for coded video bitstreams" (VSEI), while SEI messages that can affect decoding processing can be specified in the main coding specification "Versatile Video Coding." Additionally, the International Color Consortium has developed configuration files to facilitate the management of color between different input and output devices. Summary of the Invention

[0008] Among other things, this disclosure describes providing color profile metadata (e.g., International Color Consortium (ICC) profile metadata) in SEI messages. ICC profile specifications may be published by ISO, for example, as ISO 15076. As an example, an input device (e.g., a digital camera) may use specific settings that describe the color characteristics used to capture raw red, green, and blue image samples in an image captured by the input device. These settings may be described in an ICC color profile used by the input device. The input device may store its ICC profile as metadata in the image captured by the camera. Similarly, an output device (e.g., a color printer) may have its own color characteristics, which are also described in its ICC profile settings. Through the availability of the input device's ICC color profile stored in the metadata of the captured image, the output device can use both the input device's and the output device's ICC profiles to reconcile color values ​​used between the two devices (e.g., to maintain color fidelity in the image when printing the captured image on a color printer). In a similar manner, a display device (another output device) may have its own ICC profile that describes the display's color characteristics. With the presence of an ICC color profile for the camera used to capture images, the display can manage RGB colors more effectively when displaying images, thereby maintaining the fidelity of the image appearance.

[0009] According to some implementations, a video decoding method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) comprising a set of images, wherein the video bitstream corresponds to a source device; (ii) identifying color profile metadata for the source device based on a Supplemental Enhancement Information (SEI) message for the video bitstream; (iii) reconstructing the set of images using information from the video bitstream; and (iv) using color characteristics from the color profile metadata to render the set of images at an output device.

[0010] According to some implementations, a video encoding method includes: (i) receiving video data (e.g., a source video sequence) comprising a set of pictures corresponding to a source device; (ii) identifying color profile metadata for the source device; (iii) encoding the set of pictures and signaling it in a video bitstream; and (iv) signaling the color profile metadata in a supplementary enhancement information (SEI) message for the video bitstream.

[0011] According to some implementations, a method for processing visual media data includes: (i) obtaining a source video sequence including a set of pictures corresponding to a source device; (ii) identifying color profile metadata for the source device; and (iii) encoding the set of pictures in a video bitstream, wherein the video bitstream includes the encoded set of pictures and supplementary enhancement information (SEI) messages indicating the color profile metadata.

[0012] According to some implementations, a video processing method includes: (i) setting an image format metadata type identifier to indicate that ICC profile metadata is included in an SEI message associated with the current image; (ii) representing the image format metadata type identifier in a bitstream using a signal; and (iii) encoding the current image in the bitstream based on the image format metadata type identifier.

[0013] According to some embodiments, a computing system, such as a streaming system, server system, personal computer system, or other electronic device, is provided. The computing system includes a control circuitry system and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes encoder components and decoder components (e.g., a transcoder). According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions executable by the computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0014] Therefore, methods, apparatus, and systems for encoding and decoding video are disclosed. Such methods, apparatus, and systems may supplement or replace conventional methods, apparatus, and systems for encoding / decoding video. Not all features and advantages described in the specification are necessarily included, and in particular, some additional features and advantages will be apparent to those skilled in the art in light of the accompanying drawings, specification, and claims provided in this disclosure. Furthermore, it should be noted that the language used in the specification has been chosen primarily for readability and instruction purposes and is not necessarily intended to depict or limit the subject matter described herein. Attached Figure Description

[0015] To provide a more detailed understanding of this disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings only show relevant features of this disclosure and are therefore not necessarily intended to be limiting, as those skilled in the art will understand upon reading this disclosure that other valid features may be permissible.

[0016] Figure 1 This is a block diagram illustrating an example communication system according to some implementations.

[0017] Figure 2A This is a block diagram illustrating example elements of an encoder component according to some embodiments.

[0018] Figure 2B This is a block diagram illustrating example elements of a decoder component according to some embodiments.

[0019] Figure 3 This is a block diagram illustrating an example server system according to some implementation methods.

[0020] Figure 4A Example NAL units and SEI headers according to some implementations are shown.

[0021] Figure 4B An example of generative AI post-filtering processing according to some implementations is shown.

[0022] Figure 4C An example capture system for embedding ICC profile metadata into an image, according to some implementations, is shown.

[0023] Figure 4D An example of carrying ICC profile metadata within the payload of an SEI message, according to some implementations, is shown.

[0024] Figure 4EAn example is shown of a Uniform Resource Identifier (URI) referencing ICC profile metadata within the payload of an SEI message, according to some implementations.

[0025] Figure 4F Example syntax for ICC profile metadata SEI messages is shown according to some implementations.

[0026] Figure 5A An example video decoding process according to some implementation methods is shown.

[0027] Figure 5B An example video encoding process according to some implementation methods is shown.

[0028] By convention, the various features shown in the accompanying drawings are not necessarily drawn to scale, and similar reference numerals may be used to indicate similar features throughout the specification and the drawings. Detailed Implementation

[0029] This disclosure describes video / image compression techniques that include signaling color profile metadata (e.g., ICC profile metadata) associated with a source device in an SEI message used for the video bitstream. Including color profile metadata for the source device allows the color characteristics of the source device to be taken into account at the output device, which can improve the fidelity / appearance of the video data.

[0030] Example systems and devices

[0031] Figure 1 This is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0032] Source device 102 includes a video source 104 (e.g., a camera device component or media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from video source 104 can have a higher data volume compared to the encoded video bitstream 108 generated by encoder component 106. Because the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from video source 104, it requires less bandwidth to transmit and less storage space to store compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video to network 110).

[0033] One or more networks 110 represent any number of networks that transmit information between source device 102, server system 112, and / or electronic device 120, including, for example, wired (connected) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet.

[0034] One or more networks 110 include server systems 112 (e.g., distributed / cloud computing systems). In some embodiments, server system 112 is or includes streaming servers (e.g., configured to store and / or distribute video content such as encoded video streams from source device 102). Server system 112 includes codec components 114 (e.g., configured to encode and / or decode video data). In some embodiments, codec components 114 include encoder components and / or decoder components. In various embodiments, codec components 114 are instantiated as hardware, software, or a combination thereof. In some embodiments, codec components 114 are configured to decode encoded video bitstream 108 and re-encode video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings based on encoded video bitstream 108. In some embodiments, server system 112 serves as a Media-Aware Network Element (MANE). For example, server system 112 can be configured to trim the encoded video bitstream 108 to tailor potentially different bitstreams for one or more of the electronic devices 120. In some implementations, MANE is provided separately from server system 112.

[0035] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode encoded video data 116 to generate an outgoing video stream that can be displayed on a display or other type of presentation device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or includes media storage). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0036] The source device and / or multiple electronic devices 120 are sometimes referred to as “terminal devices” or “user devices”. In some implementations, one or more of the electronic devices 120 and / or the source device 102 are instances of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing equipment, and / or other types of electronic devices.

[0037] In an example operation of communication system 100, source device 102 transmits an encoded video bitstream 108 to server system 112. For example, source device 102 may encode a stream of images captured by the source device. Server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using codec component 114. For example, server system 112 may apply encoding to video data that is better suited for network transmission and / or storage. Server system 112 may transmit encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video images.

[0038] Figure 2AThis is a block diagram illustrating example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0 or YCrCb 4:4:4). In some embodiments, the video source 104 is a storage device storing previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. Video data can be provided as multiple individual images that are given motion when viewed sequentially. Each image can be organized as a spatial array of pixels, where, depending on the sampling structure, color space, etc., each pixel may include one or more samples. The relationship between pixels and samples will be readily understood by those skilled in the art.

[0039] Encoder component 106 is configured to encode and / or compress images of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, encoder component 106 is configured to perform a conversion between the source video sequence and a visual media data bitstream (e.g., a video bitstream). Implementing an appropriate encoding rate is a function of controller 204. In some embodiments, controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by controller 204 may include rate control-related parameters (e.g., image skipping, quantizer, and / or the λ value of rate-distortion optimization techniques), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of controller 204 can be readily identified by those skilled in the art, as these functions may belong to encoder component 106 optimized for a particular system design.

[0040] In some implementations, encoder component 106 is configured to operate within an encoding / decoding loop. In a simplified example, the encoding / decoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder 210. Decoder 210 reconstructs the symbols in a manner similar to that of the (remote) decoder to create sample data (when compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference image memory 208. Because decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory 208 are also bit-accurate between the local and remote encoders. In this way, the encoder's prediction portion interprets the same sample values ​​as the sample values ​​that the decoder interprets during decoding using the prediction as reference image samples.

[0041] The operation of decoder 210 can be combined with a remote decoder, for example, as shown below. Figure 2B The operation of the decoder component 122 is the same as described in the detailed description. However, a brief reference is provided. Figure 2B Since symbols are available and the encoding / decoding of symbols into encoded video sequences by entropy encoder 214 and parser 254 can be lossless, the entropy decoding portion of decoder component 122, including buffer memory 252 and parser 254, does not need to be fully implemented in local decoder 210.

[0042] Besides parsing / entropy decoding, the decoder techniques described in this paper can exist in the corresponding encoders in essentially the same form. For this reason, the subject matter focuses on decoder operations. Furthermore, the description of the encoder techniques can be simplified, as the encoder techniques are inverses of the decoder techniques.

[0043] As part of the operation of the source encoder 202, the source encoder 202 can perform motion-compensated predictive coding, which predictively encodes the input frame with reference to one or more previously encoded frames from the video sequence designated as reference frames. In this manner, the encoding engine 212 encodes the differences between pixel blocks of the input frame and pixel blocks of the reference frame, which can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.

[0044] Decoder 210 decodes encoded video data based on symbols created by source encoder 202, from frames that can be designated as reference frames. The operation of encoding engine 212 can advantageously handle lossy processing. When encoded video data is processed by video decoder (… Figure 2AWhen decoded at (not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. Decoder 210 can replicate the decoding process performed on the reference frame by a remote video decoder, and the reconstructed reference frame can be stored in reference image memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame that shares common content (no transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.

[0045] Predictor 206 can perform a prediction search against encoding engine 212. That is, for a new frame to be encoded, predictor 206 can search the reference image memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. Predictor 206 can operate pixel-by-pixel based on the sample blocks to find appropriate prediction references. As determined by the search results obtained by predictor 206, the input image can have prediction references obtained from multiple reference images stored in reference image memory 208.

[0046] The outputs of all the aforementioned functional units can undergo entropy encoding in entropy encoder 214. Entropy encoder 214 converts these symbols into a coded video sequence by lossless compression of the symbols generated by the various functional units according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, and / or arithmetic coding).

[0047] In some implementations, the output of entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence created by entropy encoder 214 in preparation for transmission via communication channel 218, which can be a hardware / software link to a storage device storing the encoded video data. The transmitter can be configured to combine the encoded video data from source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some implementations, the transmitter can transmit additional data along with the encoded video. Source encoder 202 can include such data as part of the encoded video sequence. Additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, Supplementary Enhancement Information (SEI) messages, fragments of Visual Usability Information (VUI) parameter sets, etc.

[0048] Controller 204 can manage the operation of encoder component 106. During encoding, controller 204 can assign a specific encoding picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, pictures can be assigned as intra-frame pictures (I-pictures), prediction pictures (P-pictures), or bidirectional prediction pictures (B-pictures). Intra-frame pictures can be encoded and decoded without using any other frames in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are familiar with those variations of I-pictures and their corresponding applications and characteristics, and therefore they will not be repeated here. Predictive pictures can be encoded and decoded using inter-frame prediction or intra-frame prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block. Bidirectional prediction pictures can be encoded and decoded using inter-frame prediction or intra-frame prediction, which uses at most two motion vectors and reference indexes to predict sample values ​​for each block. Similarly, multiple prediction images can use more than two reference images and associated metadata to reconstruct a single block.

[0049] The source image can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, determined by the encoding assignment of the corresponding images applied to the blocks. For example, blocks of image I can be unpredictably encoded, or blocks of image I can be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same image. Pixel blocks of image P can be unpredictably encoded with reference to a previously encoded reference image via spatial prediction or temporal prediction. Blocks of image B can be unpredictably encoded with reference to one or two previously encoded reference images via spatial prediction or temporal prediction.

[0050] Video can be captured as multiple source images (video images) in a time-series manner. Intra-frame image prediction (often simply called intra-prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In the example, a specific image in the encoding / decoding process (referred to as the current image) is segmented into blocks. Where a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference images.

[0051] Encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard described herein. In operation, encoder component 106 can perform various compression operations, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0052] Figure 2B This is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 is coupled to channel 218 and display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to loop filter 256 and configured to transmit data to display 124 (e.g., via a wired or wireless connection).

[0053] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data along with other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective user entities (not depicted). The receiver may separate the encoded video sequences from other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. Decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0054] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-frame image prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference image memory 266, and a current image memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuit systems. The decoder component 122 can be implemented at least partially in software.

[0055] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to combat network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, a separate buffer memory is provided outside decoder component 122 (e.g., to combat network jitter) in addition to buffer memory 252 inside decoder component 122 (e.g., buffer memory 252 is configured to handle playback timing). Buffer memory 252 may not be necessary or may be small when receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network. Buffer memory 252 may be required to make the best use of packet networks such as the Internet; buffer memory 252 may be relatively large and / or have an adaptive size and may be implemented at least partially in an operating system or similar component outside decoder component 122.

[0056] Parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. Symbols may include, for example, information for managing the operation of decoder component 122, and / or information for controlling a presentation device such as display 124. Control information for the presentation device may be in the form of, for example, Supplementary Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). Parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow principles known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 may extract a subgroup parameter set from the encoded video sequence for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. Subgroups may include group of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser 254 can also extract information such as transform coefficients, quantizer parameter values, and motion vectors from the encoded video sequence.

[0057] Depending on the type of encoded video picture or a subset of encoded video pictures (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and how these units are involved can be controlled by subgroup control information parsed from the encoded video sequence by parser 254. For simplicity, such subgroup control information flow between parser 254 and the following multiple units is not depicted.

[0058] The decoder component 122 can be conceptually subdivided into multiple functional units, and in some implementations, these units interact closely with each other and can be at least partially integrated with each other. However, for the sake of brevity, this paper retains the conceptual subdivision of the functional units.

[0059] The scaler / inverse transform unit 258 receives quantized transform coefficients and control information (such as which transform to use, block size, quantization factor, and quantization scaling matrix) as symbols 270 from the parser 254. The scaler / inverse transform unit 258 can output blocks including sample values, which can be input to the aggregator 268. In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use prediction information from previously reconstructed images, but can use prediction information from previously reconstructed portions of the current image. Such prediction information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 can use surrounding reconstructed information obtained from the current (partially reconstructed) image from the current image memory 264 to generate blocks of the same size and shape as the blocks in the reconstruction. The aggregator 268 can add the prediction information already generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.

[0060] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-frame coded blocks and potentially to motion-compensated blocks. In such cases, the motion-compensated prediction unit 260 can access the reference image memory 266 to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols 270 belonging to the blocks, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory 266 from which the motion-compensated prediction unit 260 obtains the predicted samples can be controlled by motion vectors. Motion vectors can be used by the motion-compensated prediction unit 260 in the form of symbols 270, which can have, for example, X components, Y components, and reference image components. Motion compensation can also include, for example, interpolation of sample values ​​obtained from the reference image memory 266 when using subsampled precise motion vectors, and motion vector prediction mechanisms.

[0061] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video bitstream and available to loop filter unit 256 as symbols 270 from parser 254. However, video compression techniques may also respond to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to sample values ​​obtained from previous reconstruction and loop filtering. The output of loop filter unit 256 may be a sample stream, which can be output to a presentation device such as display 124 and stored in reference picture memory 266 for future inter-frame picture prediction.

[0062] Once reconstructed, certain coded images can be used as reference images for future predictions. Once a coded image is reconstructed and identified as a reference image (e.g., by parser 254), the current reference image can become part of the reference image memory 266, and a new current image memory can be reallocated before reconstructing subsequent coded images begins.

[0063] Decoder component 122 can perform decoding operations according to a predetermined video compression technique that may be recorded in a standard such as any standard described herein. As specified in a video compression technique document or standard, and particularly in a configuration file therein, the encoded video may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard. Furthermore, in order to conform to some video compression techniques or standards, the complexity of the encoded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstruction sample rate (measured, for example, megasamples per second), maximum reference image size, etc. In some cases, the limitations set by the level can be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata represented in the encoded video sequence for HRD buffer management.

[0064] Figure 3 This is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes a control circuitry system 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry system 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuitry system includes a field-programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application-specific integrated circuit).

[0065] Network interface 304 can be configured to interface with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). These communication networks can be local, wide area, metropolitan area, vehicle-mounted, industrial, real-time, latency-tolerant, etc. Examples of communication networks include: local area networks such as Ethernet; wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicle and industrial networks including CANbus, etc. Such communication can be one-way receiving (e.g., broadcast TV), one-way transmitting (e.g., to a CANbus device), or bidirectional (e.g., to other computer systems using local or wide area digital networks). Such communication can include communication to one or more cloud computing networks.

[0066] User interface 306 includes one or more output devices 308 and / or one or more input devices 310. Input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. Output devices 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.

[0067] Memory 314 may include high-speed random access memory (e.g., DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state memory devices). Memory 314 may optionally include one or more storage devices remote from the control circuitry system 302. Memory 314, or alternatively, the non-volatile solid-state memory devices within memory 314 include non-transitory computer-readable storage media. In some embodiments, memory 314 or the non-transitory computer-readable storage media of memory 314 stores programs, modules, instructions, and data structures, or subsets or supersets thereof: Operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks; A network communication module 318 is used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); Codec module 320 is used to perform various functions related to encoding and / or decoding data such as video data. In some embodiments, codec module 320 is an instance of codec component 114. Codec module 320 includes, but is not limited to, one or more of the following: Decoding module 322, which performs various functions related to decoding encoded data, such as those previously described with respect to decoder component 122; and Encoding module 340, which performs various functions related to encoding data, such as those previously described with respect to encoder component 106; and Image memory 352 is used to store images and image data, for example, for use by encoding / decoding module 320. In some embodiments, image memory 352 includes one or more of the following: reference image memory 208, buffer memory 252, current image memory 264, and reference image memory 266.

[0068] In some implementations, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described with respect to the scaler / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra-frame picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter 256).

[0069] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described with respect to source encoder 202 and / or encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include... Figure 3 A subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0070] Each of the modules identified above and stored in memory 314 corresponds to a set of instructions for performing the functions described herein. The modules identified above (e.g., instruction sets) do not need to be implemented as separate software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various implementations. For example, the codec module 320 may optionally not include separate decoding and encoding modules, but instead use the same set of modules to perform both sets of functions. In some implementations, memory 314 stores a subset of the modules and data structures identified above. In some implementations, memory 314 stores additional modules and data structures not described above.

[0071] although Figure 3 A server system 112 according to some embodiments is shown, but Figure 3 This is intended more as a functional description of various features that can exist in one or more server systems than as a structural diagram of the implementation described herein. In practice, items shown individually may be combined and some items may be separated. For example, Figure 3Some items shown individually can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement server system 112 and how features are distributed among them will vary depending on the implementation and, optionally, in part, depend on the amount of data traffic processed by the server system during peak usage periods and during average usage periods.

[0072] Example encoding techniques

[0073] The encoding processes and techniques described below can be performed at the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). As previously discussed, video codecs typically include several aspects, including segmentation, intra / inter-frame prediction, transform coding, quantization, entropy coding, and intra-loop filtering. Additionally, the video encoder may include, signal, and / or parse supplementary data for the video bitstream. This supplementary data may include temporal enhancement layers, spatial enhancement layers, and / or SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, fragments of visual usability information (VUI) parameter sets, etc. The following discussion relates to such supplementary data information provided along with the video bitstream.

[0074] As previously described, SEI messages can be used to signal color profile metadata within a video stream. The metadata can be stored directly in the SEI message payload, depending on the capabilities of the video coding standard. Alternatively, the SEI message can be created using information identifying the ICC profile metadata resource to be obtained from a source outside the video bitstream (e.g., a Uniform Resource Identifier (URI)).

[0075] Many applications use SEI messages along with the associated encoded video stream to access metadata that is important / necessary for the application to operate the video stream. One example is a stereoscopic video application where the frame packing SEI message describes the organization of the left-eye and right-eye views of individual video frames, packing them into a single frame. Another example application that relies on metadata carried within the SEI message is Dolby Vision®, which uses metadata carried in the SEI message to describe the transfer function to be applied to the input video signal by a display system capable of emitting high dynamic range output.

[0076] Emerging applications of generative artificial intelligence (AI) are also an area where SEI messages can facilitate the deployment of new services. In this application space, images, along with corresponding supplemental information, can be provided to one or more neural network models that create or “generate” new or similar image-based results. A popular example of a non-image-based generative AI tool is ChatGPT. Adobe Firefly is an example of an image-based generative AI tool. However, for generative AI applications that rely on images within encoded video streams, limiting individual SEI messages to address every nuanced input that the underlying model might require is impractical. An example of metadata that could be useful for generative AI (and other applications) is information stored in ICC profile specifications, which can also be published by ISO as ISO 15076.

[0077] As previously discussed, compressed video can be enhanced, for example, by providing supplementary enhancement information in the form of SEI messages or VUIs, within the corresponding video bitstream. Video coding standards may include specification sections for SEI and VUI. SEI and VUI information may also be specified in separate specifications that can be referenced by the video coding specification.

[0078] Figure 4A An example layout of encoded video sequences (CVS) is shown. Figure 4A The encoded video sequence is subdivided into Network Abstraction Layer Units (NAL Units). Example NAL Unit 501 includes a NAL Unit header 502, which may include one or more reserved bits. In some implementations, the NAL Unit header 502 includes 16 bits, such as forbidden_zero_bit 503 and nuh_reserved_zero_bit 504, which may be unused and can be zero in the NAL Unit. The NAL Unit header 502 may also include nuh_layer_id 505. Several bits of nuh_layer_id 505 (e.g., 3 bits) may indicate the layer to which the NAL Unit belongs (e.g., spatial, SNR, or multi-view enhancement). The NAL Unit header 502 may also include nuh_unit_type 505. Several bits of nuh_nal_unit_type 505 (e.g., 5 bits) indicate the type of NAL Unit 501. In some implementations, 22 NAL unit type values ​​are defined; for example, 6 NAL unit types are reserved, and 4 NAL unit type values ​​are unspecified and can be used by other specifications (e.g., outside of video codecs). Several bits (e.g., 3) of the NAL unit header can indicate the temporal layer to which the NAL unit belongs, e.g., nuh_temporal_id_plus1506.

[0079] A encoded picture may contain one or more VCL NAL units and / or one or more non-VCL NAL units. VCL NAL units may contain encoded data conceptually belonging to the video coding layer as previously discussed. Non-VCL NAL units may contain data conceptually not belonging to the video coding layer. Using H.266 as an example, NAL units can be categorized as parameter sets, picture headers, NAL tags, prefix / suffix SE NAL units, padding data, and reserved / unspecified NAL unit types.

[0080] The parameter set includes information that may be necessary for decoding and can be applied to more than one encoded picture. The parameter set and conceptually similar NAL units can be NAL unit types such as DCI_NUT (e.g., corresponding Decoding Capability Information (DCI)), VPS_NUT (e.g., corresponding to the Video Parameter Set (VPS), establishing layer relationships, etc.), SPS_NUT (e.g., corresponding to the Sequence Parameter Set (SPS), establishing parameters used and kept constant throughout the CVS, etc.), PPS_NUT (e.g., corresponding to the Picture Parameter Set (PPS), establishing parameters used and kept constant within the encoded picture, etc.), and prefix and suffix adaptation parameter sets (e.g., PREFIX_APS_NUT and SUFFIX_APS_NUT). The parameter set may include information required by the decoder to decode the VCL NAL unit and can therefore be referred to here as a "canonical" NAL unit. The Picture Header (PH_NUT) can also be considered a "canonical" NAL unit.

[0081] NAL units that mark certain positions in a NAL unit stream may include NAL units with NAL unit types AUD_NUT (e.g., corresponding to an access unit separator), EOS_NUT (e.g., corresponding to the end of a sequence), and EOB_NUT (e.g., corresponding to the end of a bitstream). These NAL units can be considered non-canonical, and are also referred to as informative in the sense that a standards-compliant decoder does not need them for its decoding processing, although the decoder needs to be able to receive these NAL units in the NAL unit stream.

[0082] Prefix and suffix SEI NAL unit types (e.g., PREFIX_SEI_NUT and SUFFIX_SEI_NUT) indicate NAL units that contain prefix and suffix supplementary enhancement information. For example, in H.266, those NAL units are considered informational because they are not required for the decoding process.

[0083] The padding data NAL cell type (e.g., FD_NUT) indicates the padding data, which can be random and can be used to "waste" bits in the NAL cell stream or bit stream. Such padding data may be necessary for transmissions in certain synchronous transmission environments.

[0084] Still refer to Figure 4A The diagram illustrates the layout of a NAL unit stream containing an encoded image 511 in decoding order 510, where the encoded image 511 contains NAL units of some of the types discussed previously. For example, early in the NAL unit stream, DCI 512, VPS 513, and SPS 514 can establish parameters that the decoder can use to decode the encoded image of the CVS, which includes the encoded image 511 of the NAL unit stream.

[0085] The encoded picture 511 may (e.g., in the order depicted or in any other order that conforms to the video coding technology or standard in use) include the prefix -APS 516, picture header 517, prefix -SEI 518, one or more VCL NAL units 519 and suffix -SEI 520.

[0086] For some SEI messages, prefix and suffix SEI NAL units are used such that the content of the message is known before the encoding of a given image begins, while other content is only known after the image is encoded. Allowing certain SEI messages to appear earlier or later in the NAL unit stream of the encoded image via prefix and suffix SEI enables the reduction / avoidance of buffering. As an example, in an encoder, the sampling time of the image to be encoded is known before the image is encoded, and therefore the image timing SEI message can be a prefix SEI message (e.g., prefix-APS 516). On the other hand, a decoded image hash SEI message containing the hash of the sample values ​​of the decoded image and which can be used, for example, to debug the encoder implementation is a suffix SEI message (e.g., suffix-SEI 518) because the encoder cannot compute the hash of the reconstructed samples before the image has been encoded. The position of the prefix and suffix SEI NAL units is not limited to their position in the NAL unit stream. The terms “prefix” and “suffix” imply which coded pictures or NAL units a prefix / suffix SEI message can belong to, and the details of that applicability can be specified, for example, in the semantic description of a given SEI message.

[0087] Figure 4A A simplified syntax diagram of a NAL unit containing a prefixed or suffixed SEI message 520 is also shown. This syntax is a container format for multiple SEI messages that can be carried in a single NAL unit. Like other NAL units, the SEI NAL unit 520 begins with a NAL unit header 521. The header 521 is followed by one or more SEI messages. Figure 4A The diagram depicts two SEI messages, 530-1 and 530-2. Each SEI message 530 within an SEI NAL unit 520 includes a payload_type_byte 522 (e.g., 8 bits), a payload_size_byte 523 (e.g., 8 bits), and a corresponding payload 524. The payload_type_byte 522 can specify which of a plurality of SEI types (e.g., one of 256 different SEI types) applies to the SEI message, and the payload_size_byte 523 specifies the number of bytes in the SEI payload. This structure can be repeated until the payload_type_byte is observed to be equal to 0xff, indicating the end of the NAL unit. The syntax of the payload 524 can depend on the SEI message (e.g., it can have a length between 0 and 255 bytes).

[0088] Figure 4B This is a functional block diagram illustrating an example encoding 106 and decoding 122 system employing an example generative AI post-filtering process 609. The system may include a video source 104, such as a digital camera device, which creates, for example, a source video sequence input to encoder 106. Figure 4B A separate source 601 for supplementary metadata 602 for source 104 is also shown. In some embodiments, supplementary metadata 602 may come from the video source 104 itself (e.g., many digital camera devices create supplementary metadata while capturing source images). Figure 4B An encoder 106 is shown that receives both supplementary metadata and the source video sequence. The supplementary metadata 602 can be obtained by the encoder 106 from a separate source 601, or it can be obtained directly as output from the video source 104.

[0089] exist Figure 4B In this process, the output from encoder 106 is an encoded video stream 603, which includes one or more sequences of encoded image data 604 and an SEI message 605. The SEI message 605 may reference supplementary metadata 602 or carry supplementary metadata 602 in its payload. The encoded video stream 603 is input to decoder 122, which outputs a decoded video stream 606, which includes a sequence of reconstructed image data 607 and a payload 608 of supplementary metadata SEI messages. The decoded video stream 606 may be input to generative AI post-filtering processing 609. Output 610 may come from generative AI post-filtering processing 609.

[0090] Figure 4C A capture system that embeds ICC profile metadata within a JPEG image is shown. Although Figure 4CAn example with a JPEG image is shown, but similar techniques can be used for other types of images. In this figure, a digital camera device 701 captures a scene and generates a corresponding JPEG image 705. The first part of the JPEG image 705 is represented by the hexadecimal number sequence shown in the figure (and its corresponding ASCII interpretation), where "0xFFD8" represents the "start of image" JPEG marker 703 as defined in the JPEG image encoding standard. The second part of the JPEG image 705 is represented as a continuation 704 of the JPEG image 705. In this continuation 704, the APP2 marker 706 (defined by the sequence "0xFFE2" in the JPEG standard) marks the beginning of the ICC profile metadata 708 (or other types of metadata in other embodiments). The end of the ICC profile metadata 707 indicates the end of the profile metadata, the address of which can be calculated by appending 0228 hexadecimal bytes to the address of the APP2 marker.

[0091] Figure 4D The diagram shows ICC profile metadata 708, beginning with APP2 marker 706, encapsulated within SEI message 802. SEI message 802 can be specified by video standards for carrying an ICC profile metadata payload (and / or other types of payload) in an encoded video stream created by an encoder. In some implementations, the presence of SEI message 802 is indicated by a signal via SEI NAL unit 801.

[0092] Figure 4E The diagram illustrates an SEI message used to carry ICC profile metadata, where ICC profile metadata 708 is referenced by URI 901 within the payload of SEI message 901. Figure 4E In this context, the ICC configuration file metadata resides at location 902, separate from the SEI message, or in location 902.

[0093] Figure 4F A system is illustrated that uses ICC profile metadata embedded within a video sequence 1001 created by source 104 in a generative AI post-filtering process 609. The sequence 1001 from source 104 is input to encoder 106. The output from encoder 106 is an encoded video stream 603, which includes, for example, encoded video data 604 and an ICC profile metadata SEI message 1002. The encoded video stream 603 is reconstructed by decoder 122, which outputs decoded video stream 606. Decoded stream 606 includes reconstructed image data 605 and an ICC profile metadata payload 1003. Decoded stream 606 is input to generative AI process 609, which creates a generative AI processing output 610.

[0094] Table 1 below shows example syntax for SEI messages in the ICC configuration file.

[0095]

[0096] Table 1 - Example ICC Configuration File Metadata SEI Message Syntax

[0097] As shown above, the cancellation flag (icc_profile_cancel_flag) can be used to disable the persistence of previously processed ICC profile SEI messages. For example, if icc_profile_cancel_flag is set to a value indicating "false", the ICC profile mode ID (icc_profile_mode_id) can indicate whether the payload of the SEI message is the ICC profile metadata itself or a URI for a location used for the ICC profile metadata. For example, if the mode ID equals 0, the ICC profile data payload byte (icc_profile_data_payload_byte) receives data bytes from the SEI payload. If the mode ID equals 1, the ICC profile data URI (icc_profile_data_URI) receives a data string from the SEI payload.

[0098] In some implementations, different types of image format metadata can be represented by signals via SEI messages. For example, image format types may include Exchangeable Image File Format (EXIF) types, JPEG File Exchange Format (JFIF) types, Extensible Metadata Platform (XMP) types, ICC types, and / or Tagged Image File Format (TIFF) types. Table 2 shows example syntax for SEI messages involving more than one image format type.

[0099]

[0100] Table 2 - Example Image Profile Metadata SEI Message Syntax

[0101] As previously discussed, a cancellation flag equal to 1 (cancel_flag) can instruct the SEI message to cancel the persistence of any previous image format metadata SEI messages in output order, and a cancellation flag equal to 0 can instruct subsequent image format information to follow. The persistence flag (persistence_flag) specifies the persistence of the image format metadata SEI message for the current layer. For example, a persistence flag equal to 0 can instruct that the image format metadata SEI message applies only to the current image, and a persistence flag equal to 1 can instruct that the image format metadata SEI message applies to the current image and persists for subsequent images in the current layer in output order (e.g., until a new CLVS begins or the bitstream ends). The metadata payload quantity variable (num_metadata_payloads) indicates the number of metadata payloads following the SEI message. The bit_equal_to_zero flag should always be zero and used for SEI message validation. The URI present_flag indicates whether the metadata is included within the payload or obtained via a URI specified in the payload. For example, if the URI presence flag is equal to 0, the image format information can be obtained directly from the SEI message payload, and if the URI presence flag is equal to 1, the image format information is obtained by using the URI. The payload length variable (payload_len_minus1) indicates the length of the image format information payload. In some implementations, the payload length (e.g., payload size) is limited to a threshold (or optionally equal to). For example, the payload length may be limited to less than 65535 bytes. The data payload byte (data_payload_byte) contains the syntax and semantics of the image format type. For example, in the case of indicating the ICC profile type, the data payload byte can be used to indicate the ICC version (e.g., major and / or minor version). The data URI variable (data_uri) contains the URI (e.g., with the syntax and semantics specified in IETF Internet Standard 66).

[0102] The image format type (type_id) indicates the type of metadata payload of the SEI message. For example, a value of 0 can indicate the EXIF ​​type, values ​​1 to 3 can indicate the JFIF type, a value of 4 can indicate the XMP type, and a value of 5 can indicate the ICC profile type.

[0103] In some implementations, the image format (e.g., ICC profile) SEI message is a prefixed SEI message. For example, the SEI payload can be specified as shown in Table 3 below.

[0104]

[0105] Table 3 - Example Syntax of SEI Payload

[0106] Figure 5A This is a flowchart illustrating a method 500 for decoding video according to some embodiments. Method 500 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 500 is executed by executing instructions stored in the computing system's memory (e.g., memory 314).

[0107] The system receives (502) a video bitstream (e.g., an encoded video sequence) comprising a set of images, wherein the video bitstream corresponds to a source device (e.g., source device 102). The system identifies (504) the source device's color profile metadata (e.g., image format metadata) based on the SEI message of the video bitstream. The system uses the information from the video bitstream to reconstruct (506) the set of images. The system uses color characteristics from the color profile metadata to render (508) the set of images at the output device. In some implementations, the color profile metadata includes ICC metadata (e.g., an ICC profile). In this way, the ICC metadata can be extracted from or referenced by the SEI message.

[0108] Figure 5B This is a flowchart illustrating a method 550 for encoding video according to some embodiments. Method 550 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 550 is executed by executing instructions stored in the computing system's memory (e.g., memory 314).

[0109] The system receives (552) video data (e.g., a source video sequence) including a set of images corresponding to a source device (e.g., source device 102). The system identifies (554) color profile metadata for the source device. The system encodes the set of images and represents it in a signal (556) in the video bitstream. The system represents the color profile metadata in a signal (558) in an SEI message for the video bitstream. As previously described, the encoding process may reflect the decoding process described herein (e.g., representing and parsing additional data for the video bitstream in a signal). For the sake of brevity, these details are not repeated here.

[0110] although Figure 5A and Figure 5BMultiple logical stages are shown in a specific order, but stages that are not dependent on the order can be reordered, and other stages can be combined or split. Some reorderings or other groupings not specifically mentioned will be obvious to those skilled in the art, and therefore the orderings and groupings presented herein are not exhaustive. Furthermore, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.

[0111] Now let's turn to some example implementations: (A1) In one aspect, some implementations include a method for video decoding (e.g., method 500). In some implementations, the method is performed at a computing system having memory and one or more processors (e.g., server system 112). In some implementations, the method is performed at an encoding / decoding module (e.g., encoding / decoding module 320). The method includes: (i) receiving a video bitstream comprising a set of images (e.g., an encoded video sequence), wherein the video bitstream corresponds to a source device; (ii) identifying color profile metadata (and / or other image format metadata) for the source device based on an SEI message for the video bitstream; (iii) reconstructing the image set using information from the video bitstream; and (iv) using color characteristics from the color profile metadata to render the image set at an output device. For example, the SEI message may contain ICC profile metadata. As an example, the SEI message may be included as part of a video bitstream obtained from an image source.

[0112] The advantage of using color metadata (e.g., ICC profiles) is that it provides a mechanism for unambiguously mapping a video stream to the color space of a reference output, where the color information in the video stream is insufficient to perform such a mapping. Mapping to the output color space can be particularly important when the video is displayed on a computer monitor where fidelity might otherwise be reduced.

[0113] (A2) In some embodiments of A1, the method further includes parsing an image format metadata (IFM) type identifier that indicates the type of metadata included in the SEI message, wherein, if the IFM type identifier indicates that the SEI message contains color profile metadata, the color profile metadata is identified. For example, the IFM type may include Exif, JFIF, XMP, ICC, TIFF, and / or other types of metadata.

[0114] (A3) In some implementations of A1 or A2, the color profile metadata includes the International Color Consortium (ICC) profile. For example, information stored in the International Color Consortium (ICC) profile specification may also be published by ISO as ISO 15076.

[0115] (A4) In some implementations of A3, the color profile metadata indicates the ICC major version number and the ICC minor version number. For example, a first variable (ICCmajorVer) can be set to equal data_payload_byte[i][8]>>4, and a second variable (ICCminorVer) can be set to equal data_payload_byte[i][9]>>4, and the first and second variables are interpreted as the major and minor versions of the ICC profile.

[0116] (A5) In some embodiments of any one of A1 to A4, color profile metadata is obtained from the payload of the SEI message. For example, the metadata may be stored directly in the payload of the SEI message according to the capabilities of the video coding standard.

[0117] (A6) In some embodiments of any one of A1 to A4, the SEI message includes a Uniform Resource Identifier (URI) string for color profile metadata. For example, an SEI message may be created using a URI that identifies the exact ICC profile metadata resource to be obtained from a source outside the video bitstream. As an example, the ICC profile metadata may reside in a location separate from or within the SEI message.

[0118] (A7) In some embodiments of any one of A1 to A6, the method further includes applying filtering to the image set using color profile metadata. In some embodiments, the filtering is generative artificial intelligence (AI) filtering, for example, as... Figure 4F As shown.

[0119] (A8) In some embodiments of any one of A1 to A7, the video bitstream corresponds to a source video sequence that includes image data and color profile metadata. For example, a first portion of an image may be represented by a sequence of hexadecimal numbers (e.g., as specified in the JPEG image coding standard), and a second portion of the same image may include profile metadata for the image.

[0120] (A9) In some embodiments of any one of A1 to A8, the method further includes determining the presence of an SEI message based on syntax elements in the Network Abstraction Layer (NAL). For example, the presence of an SEI message can be indicated by a signal from an SEI NAL unit. In some embodiments, the SEI message is parsed only if the NAL indicates that the SEI message exists. In some embodiments, based on the determination that the NAL indicates that the SEI message does not exist, the image set is rendered without recognizing color profile metadata.

[0121] (A9) In some embodiments of any one of A1 to A9, the SEI message includes a first indicator indicating whether a previously processed SEI message of the same type should not be persisted. For example, a cancellation flag can be used to disable the persistence of previously processed ICC profile SEI messages. The first indicator may be referred to as the cancellation flag (e.g., icc_profile_cancel_flag or cancel_flag).

[0122] (A10) In some embodiments of A9, the method further includes: if a first indicator indicates that a previously processed SEI message should be persisted, determining, based on a second indicator, whether the color profile metadata is obtained from the payload of the SEI message or from a URI. For example, an ICC profile mode ID may indicate whether the payload of the SEI message is the ICC profile metadata itself or a URI for the location of the ICC profile metadata. The second indicator may be referred to as a URI flag (e.g., uri_present_flag or icc_profile_mode_id).

[0123] (B1) In another aspect, some implementations include a video encoding method (e.g., method 550). In some implementations, the method is performed at a computing system (e.g., server system 112) having memory and one or more processors. In some implementations, the method is performed at an encoding / decoding module (e.g., encoding / decoding module 320). The method includes: (i) receiving video data (e.g., a source video sequence) comprising a set of pictures corresponding to a source device; (ii) identifying color profile metadata for the source device; (iii) encoding the set of pictures and signaling it in a video bitstream; and (iv) signaling the color profile metadata in an SEI message for the video bitstream.

[0124] (B2) In some implementations of B1, the color profile metadata includes the International Color Consortium (ICC) profile.

[0125] (B3) In some implementations of B1 or B2, the color profile metadata is represented by a signal in the payload of the SEI message.

[0126] (B4) In some embodiments of any one of B1 to B3, the SEI message includes a Uniform Resource Identifier (URI) string for color profile metadata. For example, the SEI message may include a payload with some metadata and a URI with additional metadata.

[0127] (B5) In some embodiments of any one of B1 to B4, the video data includes image data and color profile metadata.

[0128] (B6) In some embodiments of any one of B1 to B5, the method further includes using a signal to indicate whether the SEI message exists in the network abstraction layer (NAL) of the video bitstream.

[0129] (B7) In some embodiments of any one of B1 to B6, the method further includes signaling via a first indicator in the SEI message whether a previously processed SEI message of the same type should not be persisted.

[0130] (C1) In another aspect, some implementations include a method for processing visual media data. In some implementations, the method is performed at a computing system (e.g., server system 112) having memory and one or more processors. In some implementations, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) obtaining a source video sequence comprising multiple frames; and (ii) performing a conversion between the source video sequence and a video bitstream of visual media data according to format rules, wherein the video bitstream comprises multiple encoded frames, and the format rules specify: (a) obtaining color profile metadata for a set of images using SEI messages of the video bitstream; and (b) reconstructing the multiple encoded frames using color characteristics from the color profile metadata.

[0131] (D1) In another aspect, some implementations include a method for video processing. In some implementations, the method is performed at a computing system (e.g., server system 112) having memory and one or more processors. In some implementations, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) setting an image format metadata type identifier to indicate that ICC profile metadata is included in an SEI message associated with the current image; (ii) representing the image format metadata type identifier in a bitstream using a signal; and (iii) encoding the current image in the bitstream based on the image format metadata type identifier.

[0132] In another aspect, some embodiments include a computing system (e.g., server system 112) comprising a control circuitry system (e.g., control circuitry system 302) and a memory (e.g., memory 314) coupled to the control circuitry system, the memory storing one or more sets of instructions configured to be executed by the control circuitry system, the set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A10, B1 to B7, C1, and D1 above). In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by the control circuitry system of the computing system, the set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A10, B1 to B7, C1, and D1 above).

[0133] Unless otherwise stated, any syntactic element described herein can be a High-Level Syntax (HLS). As used herein, HLS is represented by signals at a level higher than the block level. For example, an HLS can correspond to a sequence level, frame level, slice level, or tile level. As another example, HLS elements can be represented by signals in a Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), slice header, picture header, tile header, and / or CTU header.

[0134] It should be understood that although the terms “first,” “second,” etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of embodiments and appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or,” as used herein, refers to and covers any and all possible combinations of one or more associated listed items. It will also be understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, N This refers to the number of variables. Unless otherwise specified, NDifferent instances of can refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.

[0135] As used herein, the term "if" may be interpreted, depending on the context, as "when the prerequisite is true," or "after the prerequisite is true," or "in response to determining that the prerequisite is true," or "based on determining that the prerequisite is true," or "in response to detecting that the prerequisite is true." Similarly, the phrases "if it is determined [the prerequisite is true]," "if [the prerequisite is true]," or "when [the prerequisite is true]" may be interpreted, depending on the context, as meaning "after determining that the prerequisite is true," or "in response to determining that the prerequisite is true," or "based on determining that the prerequisite is true," or "after detecting that the prerequisite is true," or "in response to detecting that the prerequisite is true."

[0136] For illustrative purposes, the foregoing description has been described with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the claims to the precise forms disclosed. In view of the above teachings, many modifications and variations are possible. These embodiments were chosen and described in order to best illustrate the operating principles and practical applications, thereby enabling others skilled in the art to implement them.

Claims

1. A method for video decoding, the method being performed at a computing system having a memory and one or more processors, the method comprising: Receive a video bitstream comprising a set of images, wherein the video bitstream corresponds to a source device; Metadata for the color profile of the source device is identified based on the Supplemental Enhancement Information (SEI) message for the video bitstream; The image set is reconstructed using information from the video bitstream; and The color characteristics derived from the color profile metadata are used to render the image set at the output device.

2. The method of claim 1, further comprising parsing an image format metadata (IFM) type identifier indicating the type of metadata included in the SEI message, wherein, If the IFM type identifier indicates that the SEI message contains the color profile metadata, then the color profile metadata is identified.

3. The method according to claim 1, wherein, The color profile metadata includes the International Color Consortium (ICC) profile.

4. The method according to claim 3, wherein, The color profile metadata indicates the ICC major version number and the ICC minor version number.

5. The method according to claim 1, wherein, The color profile metadata is obtained from the payload of the SEI message.

6. The method according to claim 1, wherein, The SEI message includes a Uniform Resource Identifier (URI) string used for the color profile metadata.

7. The method of claim 1 further includes applying filtering processing to the image set using the color profile metadata.

8. The method according to claim 1, wherein, The video bitstream corresponds to a source video sequence that includes image data and the color profile metadata.

9. The method of claim 1, further comprising determining the presence of the SEI message based on syntax elements in the Network Abstraction Layer (NAL).

10. The method according to claim 1, wherein, The SEI message includes a first indicator indicating whether a previously processed SEI message of the same type should not be persisted.

11. The method of claim 10, further comprising: If the first indicator indicates that the previously processed SEI message should be persisted, the color profile metadata is determined from the payload of the SEI message or from the URI according to the second indicator.

12. A method for video encoding, the method being performed at a computing system having a memory and one or more processors, the method comprising: Receive video data including a set of images corresponding to the source device; Identify color profile metadata for the source device; The image set is encoded and represented as a signal in the video bitstream; as well as The color profile metadata is represented by a signal in the Supplemental Enhancement Information (SEI) message used for the video bitstream.

13. The method according to claim 12, wherein, The color profile metadata includes the International Color Consortium (ICC) profile.

14. The method according to claim 12, wherein, The color profile metadata is represented by a signal in the payload of the SEI message.

15. The method according to claim 12, wherein, The SEI message includes a Uniform Resource Identifier (URI) string used for the color profile metadata.

16. The method according to claim 12, wherein, The video data includes image data and the color profile metadata.

17. The method of claim 12, further comprising using a signal to indicate whether the SEI message exists in the Network Abstraction Layer (NAL) of the video bitstream.

18. The method of claim 12, further comprising signaling via a first indicator in the SEI message whether a previously processed SEI message of the same type should not be persisted.

19. A non-transitory computer-readable storage medium for storing a video bitstream, the video bitstream being generated by a video coding method comprising: Receive video data including a set of images corresponding to the source device; Identify color profile metadata for the source device; as well as The image set is encoded; and The video bitstream includes an encoded set of images and a Supplemental Enhancement Information (SEI) message indicating the metadata of the color profile.

20. The non-transitory computer-readable storage medium according to claim 19, wherein, The SEI message contains the color profile metadata or indicates the storage location of the color profile metadata.