Sei message for carrying text data for generative artificial intelligence application in video stream

CN122556079APending Publication Date: 2026-08-11TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2026-08-11

Smart Images

  • Figure CN122556079A_ABST
    Figure CN122556079A_ABST
Patent Text Reader

Abstract

A method for processing a video bitstream includes receiving a video bitstream comprising: (i) one of an image and a video, and (ii) a Supplemental Enhancement Information (SEI) message associated with one of the image and the video. The SEI message includes text data intended for use in generative artificial intelligence (AI) processing. The method for processing the video bitstream includes extracting the text data from the SEI message. The SEI message does not indicate whether one of the images or the video has been modified by another generative AI process.
Need to check novelty before this filing date? Find Prior Art

Description

Related applications

[0001] This application is a continuation of U.S. Application No. 19 / 012,648, filed January 7, 2025, which claims priority to U.S. Provisional Application No. 63 / 618,856, filed January 8, 2024. The entire disclosure of the prior application is incorporated herein by reference. Technical Field

[0002] This disclosure describes aspects generally related to video encoding and decoding, including Supplemental Enhancement Information (SEI) messages. Background Technology

[0003] The background description provided herein is for the purpose of presenting the overall context of this disclosure. To the extent that the work described in this background section is intended, neither the work of the currently named inventors nor any aspect of the description that would not otherwise be considered prior art at the time of filing is expressly or implicitly acknowledged as prior art to this disclosure.

[0004] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. For instance, a video codec can use a technique called intra-frame prediction, which can compress images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress images based on temporal redundancy. For example, inter-frame prediction can utilize motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] Various aspects of this disclosure include methods and apparatus for video processing, which includes, for example, encoding, decoding, and post-filtering such as neural network post-filtering. In some examples, the apparatus for video processing includes a processing circuit system.

[0006] A method for processing a video bitstream may include: receiving a video bitstream comprising: (i) one of an image and a video, and (ii) a Supplemental Enhancement Information (SEI) message associated with one of the image and the video; and extracting text data from the SEI message. The SEI message includes text data intended for use in generative artificial intelligence (AI) processing. The SEI message does not indicate whether one of the images or the video has been modified by another generative AI process.

[0007] A method for generating Supplemental Enhancement Information (SEI) messages includes: obtaining text data intended for use in generative artificial intelligence (AI) processing; and encoding a video bitstream comprising: (i) one of an image and a video, and (ii) an SEI message associated with one of the image and the video. The SEI message comprises text data. The SEI message does not indicate whether one of the image and the video has been modified by either generative AI processing.

[0008] Some aspects of this disclosure provide a method for processing visual media data. The method includes processing a bitstream of visual media data according to format rules. The bitstream includes: (i) one of an image and a video, and (ii) a Supplemental Enhancement Information (SEI) message associated with one of the images and videos. The SEI message includes text data intended for use in generative artificial intelligence (AI) processing and does not indicate whether one of the images and videos has been modified by another generative AI process. The format rules specify the extraction of text data from the SEI message.

[0009] According to another aspect of this disclosure, an apparatus is provided. The apparatus includes a processing circuitry system. The processing circuitry system can be configured to perform any of the described methods for video processing.

[0010] This disclosure also provides a non-transitory computer-readable medium for storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video processing. Attached Figure Description

[0011] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0012] Figure 1 Block diagrams of communication systems in some examples are shown.

[0013] Figure 2 Block diagrams of some example video processing systems are shown.

[0014] Figure 3Exemplary block diagrams of video decoders are shown in some examples.

[0015] Figure 4 Block diagrams of video encoders are shown in some examples.

[0016] Figure 5 The layout of some examples of Coded Video Sequences (CVS) is shown.

[0017] Figure 6 A functional block diagram of an encoding and decoding system according to one aspect of this disclosure is shown.

[0018] Figure 7 An AI (Artificial Intelligence, AI) image or video generator application is shown according to one aspect of this disclosure.

[0019] Figure 8 A functional block diagram of an encoding and decoding system employing generative AI post-filtering processing according to one aspect of this disclosure is shown.

[0020] Figure 9 An example of a functional block diagram illustrating a use case for AI text data SEI messages, according to one aspect of this disclosure, is shown.

[0021] Figure 10 An example of AI text data packaged into an SEI message according to one aspect of this disclosure is shown.

[0022] Figure 11 An example of the syntax for SEI messages according to one aspect of this disclosure is shown.

[0023] Figure 12 A flowchart outlining some aspects of video processing methods according to this disclosure is shown.

[0024] Figure 13 A flowchart outlining some aspects of video processing methods according to this disclosure is shown.

[0025] Figure 14 It is a schematic diagram of a computer system based on one aspect. Detailed Implementation

[0026] Some aspects of this disclosure provide techniques for video encoding and decoding (e.g., encoding and decoding). In some examples, this technique is used for SEI messages, which instruct the use of text data in a video stream for generative artificial intelligence applications.

[0027] Video encoding and decoding technologies can compress video data. In some examples, uncompressed digital video can comprise a series of images, each with a spatial dimension of, for example, 1920×1080 luminance samples and associated chrominance samples. This series of images can have a fixed or variable image rate (also informally referred to as the frame rate), such as 60 images per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920×1080 luminance sample resolution at a 60 Hz frame rate) requires close to 1.5 Gbit / s of bandwidth. One hour of such video would require over 600 gigabytes (GByte) of storage space.

[0028] Video encoding and decoding techniques (e.g., encoding and decoding techniques) can reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth or storage space requirements, in some cases by two or more orders of magnitude. Lossless compression, lossy compression, and combinations thereof can be used. Lossless compression refers to the technique of reconstructing an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, lossy compression is widely used. The amount of distortion tolerated depends on the application; for example, users of some consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can be reflected in the fact that higher allowable / tolerable distortion can result in a higher compression ratio.

[0029] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform, quantization, and entropy coding, some of which will be discussed below.

[0030] This document discloses systems and methods for performing video encoding and decoding in at least one processor. For example, the video decoding method may include: receiving at least one picture including an SEI message, the SEI message including at least one syntax element indicating at least one of a forward truncation function or a backward truncation function; and reconstructing an encoded picture including samples, wherein the numbering range of the samples is determined by at least one of the forward truncation function or the backward truncation function indicated by the SEI message.

[0031] In certain environments, such as video encoding and decoding for machine consumption (as opposed to human consumption), it may not be necessary for the bitstream to exceed a quality threshold based on human perception. Alternatively, this quality may not be sufficient for human consumption, but it may be adequate for machine consumption.

[0032] In some examples, techniques for reducing the range of sample values ​​are used, which may employ preprocessing performed before encoding and corresponding postprocessing after decoding. According to one aspect of this disclosure, doing so with natural sequences may result in artifacts such as banding, which is unsatisfactory to human receivers but may be acceptable for machine consumption (e.g., image analysis for high-contrast content such as barcodes). In some examples, the bit depth of a VVC image can be reduced from 10 bits to 5 bits—corresponding to 32 grayscale and color component levels. For example, in the original image, 10 bits are used to represent grayscale levels, and after the bit depth reduction, 5 bits (e.g., the most significant 5 bits out of 10) are used to represent grayscale levels. In some examples, when such preprocessing is performed before encoding, the decoder and associated processor need to know how the preprocessing was performed. Some aspects of this disclosure provide techniques for providing such information for video encoding and decoding.

[0033] Figure 1 A block diagram of a communication system (100) is shown in some examples. The system (100) includes at least two terminals interconnected via a network (150), for example... Figure 1 The diagram shows a first terminal (110) and a second terminal (120). In some examples, one-way data transmission is performed in the communication system (100). In one example, for one-way data transmission, the first terminal (110) can encode video data at its local location for transmission to the second terminal (120) via the network (150). The second terminal (120) can receive encoded video data from other terminals, such as the first terminal (110), from the network (150), decode the encoded data, and display the recovered video data. Note that one-way data transmission is typically used in media service applications, etc.

[0034] Figure 1 A second pair of terminals, such as a third terminal (130) and a fourth terminal (140), is also shown. This second pair of terminals is configured to support bidirectional transmission of encoded video, such as bidirectional transmission of encoded video that may occur during a video conference. For bidirectional transmission of data, each of the third terminal (130) and the fourth terminal (140) can encode video data captured at a local location for transmission to other terminals via a network (150). Each of the third terminal (130) and the fourth terminal (140) can also receive encoded video data transmitted by other terminals, decode the encoded data, and display the recovered video data on a local display device.

[0035] Note that, although in Figure 1In this disclosure, terminals (110), (120), (130), and (140) are shown as servers, personal computers, and smartphones, but this disclosure is not limited to such terminal examples. Aspects of this disclosure may include applications in the case of laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (150) refers to any number of networks that transmit encoded video data between terminals (110), (120), (130), and (140), including, for example, wired and / or wireless communication networks. Network (150) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network (150) may be irrelevant to the operation of this disclosure unless explained below. In some examples, the network (150) includes a Media Aware Network Element (MANE, 160), which may be included in the transmission path between, for example, a third terminal (130) and a fourth terminal (140). In some examples, the MANE (160) may selectively forward portions of media data to respond to network congestion, media switching, media mixing, archiving, and similar tasks typically performed by service providers rather than end users. Such a MANE may be able to parse and respond to limited portions of the media transmitted over the network, such as syntactic elements associated with video codec technologies or standard network abstraction layers.

[0036] Figure 2 Block diagrams of some example video processing systems (200) are shown. The video processing system (200) is an example of the application of the disclosed subject matter—video encoders and video decoders—in a streaming environment. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs (Compact Discs), DVDs (Digital Versatile Discs), memory sticks, etc.

[0037] The video processing system (200) includes a capture subsystem (213) that may include a video source (201), such as a digital camera device, which creates, for example, an uncompressed video picture stream (202). In the example, the video picture stream (202) includes samples captured by the digital camera device. The video picture stream (202) is depicted as a thick line to emphasize the high data volume when compared with encoded video data (204) (or encoded video bitstream), which may be processed by an electronic device (220) including a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (204) (or encoded video bitstream) is depicted as a thin line to emphasize its lower data volume when compared to the video picture stream (202). This encoded video data (204) (or encoded video bitstream) can be stored on a streaming server (205) for future use. One or more streaming client subsystems, for example... Figure 2 Client subsystems (206) and (208) can access a streaming server (205) to retrieve copies (207) and (209) of encoded video data (204). Client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and creates an outgoing stream (211) of video images that can be rendered on a display (212) (e.g., a screen) or other rendering device (not depicted). In some streaming systems, the encoded video data (204), (207), and (209) (e.g., video bitstreams) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T (International Telecommunication Union-Telecommunication Standardization Sector, ITU-T) Recommendation H.265. In this example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The topics that are exposed can be used in the context of VVC.

[0038] Note that electronic devices (220) and (230) may include other components (not shown). For example, electronic device (220) may include a video decoder (not shown), and electronic device (230) may also include a video encoder (not shown).

[0039] Figure 3A block diagram of a video decoder (310) is shown. The video decoder (310) may be included in an electronic device (330). The electronic device (330) may include a receiver (331) (e.g., a receiving circuitry system). The video decoder (310) may be used in place of... Figure 2 The video decoder (210) in the example.

[0040] The receiver (331) can receive, for example, one or more encoded video sequences included in a bitstream, to be decoded by the video decoder (310). In one aspect, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Encoded video sequences can be received from a channel (301), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (331) can receive encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective user entities (not depicted). The receiver (331) can separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (315) can be coupled between the receiver (331) and the entropy decoder / parser (320) (hereinafter referred to as the "parser (320)"). In some applications, the buffer memory (315) is part of the video decoder (310). In other applications, the buffer memory can be external to the video decoder (310) (not depicted). In other applications, a buffer memory (not depicted) may exist outside the video decoder (310) to prevent network jitter, for example, and another buffer memory (315) may exist inside the video decoder (310) to handle broadcast timing, for example. The buffer memory (315) may not be necessary, or it may be small, when the receiver (331) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isochronous synchronization network. For the purpose of utilizing packet networks such as the Internet, a buffer memory (315) may be required. This buffer memory (315) may be relatively large and advantageously have an adaptive size, and may be implemented at least partially in the operating system or in a similar element (not depicted) outside the video decoder (310).

[0041] The video decoder (310) may include a parser (320) to reconstruct symbols (321) from the encoded video sequence. These symbols include: information for managing the operation of the video decoder (310), and potential information for controlling rendering devices such as rendering devices (312) (e.g., a display screen), which are not part of the electronic device (330) but may be coupled to it, such as... Figure 3As shown. Control information for the rendering device can be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (320) can parse / decode the received encoded video sequence. Encoding and decoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (320) can extract a subgroup parameter set of at least one subgroup of pixels in the subgroups of the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. Subgroups can include Group of Picture (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser (320) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0042] The parser (320) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (315) to create symbols (321).

[0043] Depending on the type of the encoded video picture or a portion thereof (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (321) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (320). For clarity, such subgroup control information flow between the parser (320) and the multiple units below is not depicted.

[0044] In addition to the functional blocks already mentioned, the video decoder (310) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0045] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives quantization transform coefficients as symbols (321) from the parser (320) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) can output a block containing sample values ​​that can be input into the aggregator (355).

[0046] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to intra-coded blocks. An intra-coded block is a block that does not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses surrounding reconstructed information extracted from the current picture buffer (358) to generate blocks of the same size and shape as the blocks in the reconstruction. For example, the current picture buffer (358) buffers partially reconstructed and / or fully reconstructed current images. In some cases, the aggregator (355) adds the predictive information already generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.

[0047] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to inter-frame encoded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (353) can access the reference image memory (357) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (321) belonging to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory (357) from which the motion compensation prediction unit (353) extracts its prediction samples can be controlled by motion vectors, which are provided to the motion compensation prediction unit (353) in the form of symbols (321), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (357), motion vector prediction mechanisms, etc., when using subsample precise motion vectors.

[0048] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). Video compression techniques may include in-loop filtering techniques controlled by parameters available to the loop filter unit (356) as symbols (321) included in the encoded video sequence (also referred to as the encoded video bitstream) from the parser (320). Video compression may also respond to metadata acquired during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, as well as to previously reconstructed and loop-filtered sample values.

[0049] The output of the loop filter unit (356) can be a sample stream, which can be output to the rendering device (312) and stored in the reference image memory (357) for use in future inter-frame image prediction.

[0050] Once fully reconstructed, certain encoded images can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and that encoded image (by, for example, the parser (320)) is identified as the reference image, the current image buffer (358) can become part of the reference image memory (357), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0051] The video decoder (310) can perform decoding operations according to standards or predetermined video compression technologies such as those specified in ITU-T Recommendation H.265. An encoded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and in the sense that a profile is documented in the video compression technology or standard. Specifically, the profile may select certain tools from all available tools in the video compression technology or standard as tools available only under that profile. For compliance, the complexity of the encoded video sequence may also be required to be within limits defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata managed by the HRD buffer used to signal in the encoded video sequence.

[0052] On one hand, the receiver (331) can receive the encoded video along with additional (redundant) data. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (310) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be, for example, in the form of temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0053] Figure 4 An exemplary block diagram of a video encoder (403) is shown. The video encoder (403) is included in an electronic device (420). The electronic device (420) includes a transmitter (440) (e.g., a transmission circuit system). The video encoder (403) can be used in place of Figure 2 The video encoder (203) in the example.

[0054] The video encoder (403) can obtain data from the video source (401) (which is not...). Figure 4 In one example, an electronic device (420) receives a video sample, and the video source (401) can capture a video image to be encoded by a video encoder (403). In another example, the video source (401) is part of the electronic device (420).

[0055] A video source (401) can provide a sequence of source video samples in the form of a digital video sample stream to be encoded by a video encoder (403). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (401) can be a storage device storing previously prepared video. In a video conferencing system, the video source (401) can be a camera device capturing local image information as a video sequence. Video data can be provided as multiple individual pictures that are given motion when viewed sequentially. The pictures themselves can be organized as a spatial array of pixels, where each pixel can include one or more samples, depending on the sampling structure, color space, etc., used. The following description focuses on samples.

[0056] According to one aspect, the video encoder (403) can encode and compress the images of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints as required. Implementing an appropriate encoding rate is a function of the controller (450). In some aspects, the controller (450) controls and is functionally coupled to other functional units as described below. For clarity, the coupling is not depicted. Parameters set by the controller (450) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. The controller (450) can be configured to have other suitable functions belonging to the video encoder (403) optimized for a specific system design.

[0057] In some respects, the video encoder (403) is configured to operate within an encoding / decoding loop. As an oversimplification, in this example, the encoding / decoding loop may include a source encoder (430) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and decoded and a reference image) and a (local) decoder (433) embedded within the video encoder (403). The decoder (433) reconstructs the symbols, creating sample data in a manner similar to how the (remote) decoder would also create them. The reconstructed sample stream (sample data) is input to a reference image memory (434). Since decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (434) are also bit-accurate between the local and remote encoders. In other words, the encoder's prediction portion "sees" the exact same sample values ​​that the decoder "sees" during prediction as the reference image sample. This basic principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related techniques.

[0058] The operation of the "local" decoder (433) can be combined with that of the "remote" decoder, as already mentioned above. Figure 3 The operation of the video decoder (310) described in detail is the same. However, a brief reference is also provided. Figure 3 Since symbols are available and the encoding of symbols into an encoded video sequence by the entropy encoder (445) and the decoding of symbols by the parser (320) can be lossless, the entropy decoding portion of the video decoder (310), which includes the buffer (315) and the parser (320), may not be fully implemented in the local decoder (433).

[0059] On the one hand, decoder techniques other than parsing / entropy decoding present in the decoder exist in the corresponding encoder with the same or substantially the same functional form. Therefore, the subject matter disclosed focuses on decoder operation. Because encoder techniques are inverses of the fully described decoder techniques, the description of encoder techniques can be simplified. A more detailed description is provided below in certain sections.

[0060] In some examples, during operation, the source encoder (430) may perform motion-compensated predictive coding, which predictively codes the input image with reference to one or more previously encoded images from the video sequence designated as "reference images". In this way, the encoding engine (432) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which can be selected as a predictive reference for the input image.

[0061] The local video decoder (433) can decode encoded video data of a picture that can be designated as a reference picture based on symbols created by the source encoder (430). The operation of the encoding engine (432) can be advantageously for lossy processing. When the encoded video data can be decoded by the video decoder (433), Figure 4 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process performed on the reference image by the video decoder and can store the reconstructed reference image in the reference image memory (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0062] The predictor (435) can perform a prediction search against the encoding engine (432). That is, for a new image to be encoded, the predictor (435) can search in the reference image memory (434) for sample data (as candidate reference pixel blocks) or certain metadata such as reference image motion vectors, block shapes, etc. that can be used as appropriate prediction references for the new image. The predictor (435) can operate pixel-by-pixel based on the sample blocks to find appropriate prediction references. In some cases, as determined by the search results obtained by the predictor (435), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (434).

[0063] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, the setting of parameters and subgroup parameters for encoding video data.

[0064] The outputs of all the functional units mentioned above can be entropy encoded in the entropy encoder (445). The entropy encoder (445) converts the symbols generated by the various functional units into an encoded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0065] The transmitter (440) can buffer an encoded video sequence, such as that created by the entropy encoder (445), in preparation for transmission via a communication channel (460), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (440) can combine the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0066] The controller (450) can manage the operation of the video encoder (403). During encoding, the controller (450) can assign a certain encoded image type to each encoded image, which may affect the encoding techniques that can be applied to the corresponding image. For example, images can typically be assigned to one of the following image types:

[0067] Intra-frame pictures (I-pictures) can be encoded and decoded without using any other pictures in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0068] Predictive images (P-images) can be encoded and decoded using intra-frame or inter-frame prediction that uses motion vectors and reference indices to predict sample values ​​for each block.

[0069] Bidirectional predictive images (B-images) can be encoded and decoded using intra-frame or inter-frame predictions that predict sample values ​​for each block using two motion vectors and a reference index. Similarly, multiple predictive images can be used for the reconstruction of a single block using more than two reference images and associated metadata.

[0070] Source images are typically spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignment of the corresponding images applied to the blocks. For example, blocks of image I can be non-predictively coded, or blocks of image I can be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of image P can be predictively coded with reference to a previously coded reference image via spatial prediction or via temporal prediction. Blocks of image B can be predictively coded with reference to one or two previously coded reference images via spatial prediction or via temporal prediction.

[0071] The video encoder (403) can perform encoding operations according to standards or predetermined video coding technologies such as ITU-T H.266 Recommendation. In the operation of the video encoder (403), various compression operations can be performed, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technology or standard being used.

[0072] On one hand, the transmitter (440) can transmit additional data along with the encoded video. The source encoder (430) can include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, video availability information (VUI) parameter set fragments, etc.

[0073] Video can be captured as multiple source images (video images) in a time-series manner. Intra-frame image prediction (often simply called intra-prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In the example, a specific image during encoding / decoding—referred to as the current image—is segmented into blocks. Where a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference images.

[0074] In some aspects, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted using a combination of the first and second reference blocks.

[0075] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0076] According to some aspects of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the HEVC (High-Efficiency Video Coding) standard, pictures in a video picture sequence are segmented into Coding Tree Units (CTUs) for compression. The CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU comprises three Coding Tree Blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively divided into one or more Coding Units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type used for that CU, such as inter-frame prediction or intra-frame prediction. Based on temporal and / or spatial predictability, the CU is divided into one or more Prediction Units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, prediction operations in encoding / decoding are performed on a block-by-block basis. Using a luma prediction block as an example, a prediction block comprises a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0077] Note that the video encoders (203) and (403) and the video decoders (210) and (310) can be implemented using any suitable technology. On one hand, the video encoders (203) and (403) and the video decoders (210) and (310) can be implemented using one or more integrated circuits. On the other hand, the video encoders (203) and (403) and the video decoders (210) and (310) can be implemented using one or more processors that execute software instructions.

[0078] This disclosure includes video encoding and decoding, such as carrying text for AI applications (also known as AI text) within an encoded video stream for video-based applications.

[0079] Video encoders and video decoders can utilize techniques from a wide range of categories, including, for example, motion compensation, transform, quantization, entropy coding, carrying supplementary information (e.g., metadata that can describe the images in the encoded bitstream), etc.

[0080] On the one hand, compressed video and / or images in a video bitstream can be enhanced by supplemental enhancement information (e.g., in the form of Supplemental Enhancement Information (SEI) messages or VUIs). In some examples, video coding standards may include specifications for SEI and VUI. In some examples, SEI and VUI information may also be specified in a separate specification that can be referenced by the video coding specification.

[0081] The technology used for video coding standards may include one or more SEI messages that enable, for example, the carrying of supplementary information to the encoded video within the encoded bitstream. Such SEI information may or may not be directly related to the video coding process specified by, for example, a video standard (e.g., H.264|AVC, H.265|HEVC, H.266|VVC, etc.). In many cases, the information in the SEI messages may be related to application processing performed in conjunction with or immediately following the video decoding process. For example, such applications may include rendering processes that use certain SEI messages to adjust the brightness or color space of decoded video frames (also called pictures) before being rendered by a display device.

[0082] On the one hand, within current standards that utilize SEI messages (e.g., H.264|AVC, H.265|HEVC, and H.266|VVC), SEI messages can be divided into two categories: a first category that may affect video decoding processing, and a second category that does not affect video decoding processing (e.g., for external applications). SEI messages that do not affect decoding processing can be specified in a separate specification—for example, a specification entitled "General Supplemental Enhancement Information Messages for Encoded Video Bitstreams" (VSEI). SEI messages that may affect decoding processing can be specified in a master coding specification such as "General Video Coding".

[0083] On one hand, encoded video bitstreams can be used in applications leveraging artificial intelligence and machine learning techniques, which relates to a recent standardization area within the ITU-T / ISO / IEC Joint Video Experts Group (JVET). For example, for this standardization work, a set of SEI messages can be specified, for instance, in version 3.0 of the VSEI specification for use in such applications. In some examples, SEI messages enable the carrying (or referencing via a Uniform Resource Identifier (URI)) of one or more neural networks that will be applied to decoded images within the video stream. In some examples, not all AI-based applications can opt to utilize SEI messages (e.g., newly specified SEI messages), as SEI messages can be specified to reference or carry neural network models. In some examples, there are AI applications where the neural network does not need to be carried in the encoded video stream or does not need to be referenced from the encoded video stream, such as emerging applications for generative AI. In this application space, images, along with corresponding supplementary information, can be provided to one or more neural network models that can create or “generate” new or similar image-based results. The neural network models used for these applications can be directly referenced from the applications and may not need to be sent in the encoded video stream. Examples of these generative AI image applications can include AI art generators, Canva's AI image generator, and more.

[0084] Some aspects of this disclosure include video processing such as encoding and decoding, which includes, for example, carrying text for AI applications (also referred to as AI text) within an encoded video stream for video-based applications. For example, an SEI message for carrying text data for generative artificial intelligence applications is described.

[0085] Figure 5The layout of an encoded video sequence (CVS) is shown in some examples (e.g., according to H.266). The encoded video sequence is subdivided into Network Abstraction Layer (NAL) units (NAL units), for example... Figure 5 The NAL unit (501) in the text. Figure 5 In some examples, the NAL unit (501) may include a NAL unit header (502). In some examples, the NAL unit header (502) includes 16 bits. Figure 5 In the example, the NAL unit header (502) includes a first bit (e.g., forbidden_zero_bit) (503) and a second bit (e.g., nuh_reserved_zero_bit) (504). In the example, the first and second bits are not used by H.266 and can be set to zero in an H.266 compliant NAL unit.

[0086] exist Figure 5 In the example, the NAL unit header (502) includes a syntax element nuh_layer_id (505) with multiple bits (e.g., six bits). In the example, the three bits of nuh_layer_id (505) can indicate the layer (spatial, SNR, or multi-view enhancement) to which the NAL unit (501) belongs. The NAL unit header (502) includes a five-bit nuh_nal_unit_type (506) that defines the type of the NAL unit (501). In some examples (e.g., H.266), of the 32 values ​​represented by five bits, 22 NAL unit type values ​​are defined for the NAL unit type, six NAL unit types are reserved, and four NAL unit type values ​​are unspecified and can be used by specifications other than H.266. The NAL unit header (502) includes a three-bit nuh_temporal_id_plus1 (506) to indicate the temporal layer to which the NAL unit (501) belongs.

[0087] NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units can include data representing the values ​​of samples in a video frame, while non-VCL NAL units can contain associated additional information, such as one or more parameter sets, SEI, etc.

[0088] In some examples, an encoded image may include one or more VCL NAL units and zero or more non-VCL NAL units. VCL NAL units may contain encoded data that conceptually belongs to the video coding layer as described above. Non-VCL NAL units may contain data that conceptually does not belong to a video coding layer. Using H.266 as an example, non-VCL NAL units can be categorized as follows:

[0089] (1) A parameter set, which includes information that can be used for decoding processing and can be applied to more than one encoded picture. A parameter set and a conceptually similar NAL unit can be, for example, the following NAL unit types (NUTs): DCI_NUT (Decoding Capability Information, DCI), VPS_NUT (Video Parameter Set, VPS, establishing layer relationships, etc.), SPS_NUT (Sequence Parameter Set, SPS, establishing parameters used and kept constant throughout the encoded video sequence CVS, etc.), PPS_NUT (Picture Parameter Set, PPS, establishing parameters used and kept constant within an encoded picture, etc.), and PREFIX_APS_NUT and SUFFIX_APS_NUT (prefix and suffix adaptation parameter sets). The parameter set may include information required by the decoder to decode the VCL NAL unit, and is therefore referred to herein as a "canonical" NAL unit.

[0090] (2) Picture header (PH_NUT), which is also a “normal” NAL unit.

[0091] (3) Mark NAL units at specific positions in the NAL unit stream. The third category includes NAL units with NAL unit types AUD_NUT (Access Unit Separator), EOS_NUT (End of Sequence), and EOB_NUT (End of Bitstream). These units are non-canonical, also known as informative, in the sense that a compatible decoder does not require them for its decoding processing. In the example, a compatible decoder might need to be able to receive them in the NAL unit stream.

[0092] (4) Prefix and suffix SEI NAL unit types (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), which indicate NAL units containing prefix and suffix supplementary enhancement information. In some examples (e.g., in H.266), the fourth category of NAL units is informative because they are not required for the decoding process.

[0093] (5) Padding data NAL cell type FD_NUT indicates padding data, which can be random and can be used to “waste” bits in the NAL cell stream or bit stream, which may be necessary for transmission in some synchronous transmission environments.

[0094] (6) Reserved and unspecified NAL cell types.

[0095] Figure 5 The layout of the NAL unit stream (510) in decoding order (590) in some examples is also shown. The NAL unit stream (510) includes encoded images (511). The NAL unit stream (510) includes DCI (512), VPS (513), and SPS (514) at positions earlier than the encoded images (511). The DCI (512), VPS (513), and SPS (514) can be combined to establish parameters that the decoder can use to decode the encoded images of the encoded video sequence (CVS), which includes the encoded images (511) in the NAL unit stream (510).

[0096] exist Figure 5 In the example, the encoded images (511) may be in the order depicted or conform to the video coding technology or standard in use (e.g., Figure 5 Any other order of H.266 shown includes: prefix APS (516), picture header (PH) (617), prefix SEI (518), one or more VCL NAL units (519) and suffix SEI (520).

[0097] In some examples, the prefix SEI NAL unit (518) and suffix SEI NAL unit (520) were configured during standardization such that, for some SEI messages, the content of the message could be known before the encoding of a given image begins, while in other examples, the content could only be known after the image had been encoded. Enabling certain SEI messages to appear earlier or later in the NAL unit stream of the encoded image via prefix SEI and suffix SEI avoids buffering. For example, in an encoder, the sampling time of the image to be encoded is known before the image is encoded, and therefore, the image timing SEI message could be a prefix SEI message (518). On the other hand, the decoded image hash SEI message (which contains the hash of the sample values ​​of the decoded image and can be used, for example, to debug the encoder implementation) is a suffix SEI message (520) because the encoder cannot compute the hash of the reconstructed sample before the image is encoded. The positions of the prefix and suffix SEI NAL units are not limited to their location within the NAL unit stream. The phrases “prefix” and “suffix” can imply what encoded picture or NAL unit a prefix / suffix SEI message can belong to, and details of that applicability can be specified, for example, in the semantic description of a given SEI message.

[0098] Figure 5 A diagram is also shown illustrating the syntax of a NAL unit (551) containing prefixed or suffixed SEI messages. This syntax is a container format for multiple SEI messages that can be carried within a single NAL unit (also known as an SEI NAL unit). For clarity, details of the simulation prevention syntax specified in H.266 are omitted here. As other NAL units, an SEI NAL unit may begin with a NAL unit header (521). Following the NAL unit header (521) are one or more SEI messages, such as... Figure 5 The first SEI message (530) and the second SEI message (540) within the NAL unit (551). Each SEI message within the NAL unit (551) may include an 8-bit payload_type_byte specifying one of 256 different SEI types, for example, by Figure 5 The payload_type_byte (532) and payload_type_byte (542) are shown in the NAL unit (551). Each SEI message within the NAL unit (551) may include an 8-bit payload_size_byte specifying several bytes of the SEI payload, for example, by Figure 5The payload_size_byte (533) and payload_size_byte (543) are shown in the NAL unit (551). Each SEI message within the NAL unit (551) may include an SEI payload of several bytes specified by payload_size_byte, for example... Figure 5 The payload (534) and payload (544) are specified in the structure. In some examples, this structure may be repeated until the payload_type_byte is observed to be equal to 0xff, which indicates the end of the NAL unit. The syntax of the payload may depend on the SEI message and may have any suitable length, such as between 0 bytes and 255 bytes.

[0099] Figure 6A functional block diagram of an encoding and decoding system (600) according to one aspect of this disclosure is shown. The encoding and decoding system (600) may employ post-filtering processing (e.g., neural network post-filtering) (613), wherein one or more neural network models may be carried in the payload of an SEI message, or one or more neural network models may be referenced in the SEI message, for example, via a URI to a source outside the encoded video stream. The encoding and decoding system (600) may include a video source, such as a video source (201) (e.g., a digital camera device), which creates a source video sequence as input to an encoder such as an encoder (203). In the example, in addition to the source video sequence, the encoder (203) may also receive input from, for example, a separate source (601) containing one or more neural network models that may be used in the post-filtering processing (613). The output from the encoder (203) may include an encoded video stream (604). The encoded video stream (604) may include one or more sequences of encoded image data (602) and an SEI message (603), which may reference or carry neural network model information in the payload of the SEI message (603). The encoded video stream (604) may be input to a decoder—for example, a decoder (210) that may output a decoded video stream (607). The decoded video stream (607) may include one or more sequences of reconstructed image data (605) and a payload (606) of a neural network SEI message. The decoded video stream (607) may be input to a neural network post-filtering process (613), in which a neural network filter controller (608) may perform a series of steps. The series of steps may include selecting image data (609) from the data in the decoded video stream (607) and building a neural network pipeline (610) based on the SEI payload (606). In the example, the neural network pipeline (610) includes a sequence of one or more neural network filters (611) based on the SEI payload (606). The output (612) from the neural network pipeline (610) may also be the output from a neural network post-filtering process (613).

[0100] Figure 7 An AI image or video generator application according to one aspect of this disclosure is illustrated. The AI ​​image and / or video generator (703) can receive an image and / or video sequence (701) and a text prompt (702). The text prompt (702) can guide the AI ​​image or video generator processing performed by the AI ​​image and / or video generator (703). An output image or video sequence (704) can be output from the AI ​​image or video generator (703). See also... Figure 7In the example shown, the image and / or video sequence (701) is an image of a person riding a bicycle in a city. The text prompt (702) is the message “A panda is riding a bicycle through the city.” When the AI ​​image or video generator (703) receives the image and text prompt, the AI ​​image or video generator (703) can generate an output based on the text prompt (702) and the image and / or video sequence (701), which is an image of a panda riding a bicycle in a city.

[0101] Figure 8 A functional block diagram of an encoding and decoding system (800) employing generative AI post-filtering (803) according to one aspect of this disclosure is shown. The encoding and decoding system (800) may include a video source, such as a video source (201) (e.g., a digital camera device), which creates, for example, a source video sequence input to an encoder such as an encoder (203). (Refer to...) Figure 8 The supplementary data (also known as supplementary metadata) (805) can be further described by a separate source (801) of the video source (201). The encoder (203) can receive both the supplementary metadata (805) and the source video sequence.

[0102] The output from the encoder (203) may include an encoded video stream (804). The encoded video stream (804) may include one or more sequences of encoded image data (802) and an SEI message (also known as a supplementary metadata SEI message) (883), which may reference or carry supplementary metadata (e.g., supplementary metadata (805)) in the payload of the SEI message (883). The encoded video stream (804) may be input to a decoder such as a decoder (210). The decoder (210) may output a decoded video stream (807). The decoded video stream (807) may include one or more sequences of reconstructed image data (885) and a payload (806) of a supplementary data SEI message. The decoded video stream (807) may be input to a generative AI process (also known as generative AI post-filtering or AI image / video generative processing) (803). The output (also known as generative AI processing output) (804) may come from the generative AI post-filtering process (803).

[0103] Figure 9 An example of a functional block diagram illustrating a use case for AI text data SEI messages, according to one aspect of this disclosure, is shown. Input (e.g., an input image or input video sequence) (901) can be provided to an encoder, such as an encoder (203). Text prompts (902) can be provided to guide generative AI processing (or AI image / video generative processing), such as generative AI processing (803), to generate output based on the text prompts (902). Figure 9 In the example shown, the text prompt is the string "A panda rides a bicycle through the city," and the output can include an image of a panda riding a bicycle in the city. In the example, the text prompt (902) is provided by the user (993). The text prompt (902) and the input (901) can be provided at appropriate times. In the example, the text prompt (902) and the input (901) are provided approximately simultaneously.

[0104] An encoded video stream with AI text data SEI messages can be transmitted to a decoder such as a decoder (210). In the example, the encoded video stream with AI text data SEI messages is transmitted over a network to the cloud (904). The decoder (210) can generate a decoded video stream (908). The decoded video stream (908) can be input to a generative AI process (803). The generative AI process (803) can generate an output that can be input to a second encoder (903). In this scenario, the optional second encoder (903) can generate an encoded video stream (905). The encoded video stream (905) can include, for example, a video sequence or image sequence of a panda riding a bicycle in a city guided by a text prompt (902).

[0105] Figure 10 An example of AI text data (1002) packaged into an SEI message (e.g., an AI text cue SEI message) (1000) is shown, which may be specified by video standards to be used in video outputs generated by encoders (e.g., such as...). Figure 8 or Figure 9 The encoder (203) shown in the image generates an encoded video stream that carries AI text data. An example of the AI ​​text data (1002) is as follows: Figure 9 The text prompt shown (902) or as... Figure 7 The text prompt shown is (702). In some examples, such as in the specifications of video standards, the presence of the SEI message payload (1002) is notified by a signal from the SEI NAL unit (1001).

[0106] Figure 11 An example of the syntax for an SEI message (e.g., an AI data SEI message) (1100) according to one aspect of this disclosure is shown. The SEI message (1100) may include a cancellation flag (e.g., ait_data_cancel_flag) (1101). In one aspect, if the cancellation flag (e.g., ait_data_cancel_flag) (1101) is set to a first value (e.g., a zero value, such as...), ... Figure 11As shown), the remaining portion of the SEI payload of the SEI message (e.g., AI data SEI message) (1100) can then be processed. If the cancellation flag (e.g., ait_data_cancel_flag) (1101) is set to a second value (e.g., a non-zero value, such as...), then... Figure 11 As shown), there is no remaining SEI payload to be processed for the SEI message (e.g., AI data SEI message) (1100).

[0107] If the cancellation flag (e.g., ait_data_cancel_flag) (1101) is set to a first value (e.g., zero), a persistence flag (e.g., ait_data_persistence_flag) (1102) can be stored. In the example, if the persistence flag (e.g., ait_data_persistence_flag) (1102) is equal to zero, the AI ​​data SEI message (1100) can be applied only to the current image. If the persistence flag (e.g., ait_data_persistence_flag) (1102) is equal to a value of 1, the AI ​​data SEI message (1100) can be applied to the current image (e.g., the currently decoded image) and can continue to be applied to all subsequent images of the current layer in output order until one or more of the following conditions are true. One or more of the following conditions may include: 1) the start of a new CLVS for the current layer; 2) the end of the bitstream; and 3) in the AU associated with the SEI message used to implement text data for AI applications, the images in the current layer are outputs after the current image in output order.

[0108] When the cancellation flag (e.g., ait_data_cancel_flag) (1101) is set to a first value (e.g., zero), the data string (e.g., ait_data_string) (1103) can receive the string of the payload from the AI ​​data SEI message (1100).

[0109] According to one aspect of this disclosure, a video stream (also referred to as a video bitstream) may include text data intended for use in generative AI processing or generative AI applications. The text data intended for use in generative AI processing or generative AI applications may be encoded in an SEI message (e.g., an AI data SEI message) (1100). In another aspect, the text data may be encoded as the payload of the AI ​​data SEI message (1100). For example, the text data may be encoded in a data string (e.g., an ait_data_string) (1103).

[0110] On one hand, textual data intended for use in generative AI processing can be obtained. A video bitstream can be encoded, comprising: (i) one of an image and a video, and (ii) an SEI message associated with one of the image and the video. The SEI message may include textual data. According to one aspect of this disclosure, the SEI message does not indicate whether one of the image and the video has been modified by any generative AI processing. An example of textual data is... Figure 7 The text prompt shown (702) or Figure 9 The text prompt shown is (902). In the example, for example, in... Figures 10 to 11 As shown, text data is carried in the payload of the SEI message. In the example, the payload of the SEI message includes instructions for the SEI message (e.g., Figure 11 The remaining portion of the payload shown in (1100) has markings to be processed (e.g., Figure 11 The cancellation flag (1101) shown is illustrated. In the example, generative AI processing uses, for example, the cancellation flag (1101). Figure 6 This is achieved through post-neural filtering of one or more neural networks, as shown. In the example, the SEI message indicates the post-neural filter information for the post-neural filtering process.

[0111] According to one aspect of this disclosure, the video bitstream, including the SEI message, can be received, for example, by a receiving system. Text data intended for use in generative AI processing or generative AI applications can be extracted from the SEI message.

[0112] On one hand, the video bitstream includes (i) one of the images and the video, and (ii) an SEI message associated with one of the images and the video, and the SEI message includes text data intended for generative AI processing. The text data can be extracted from the SEI message, which does not indicate whether one of the images and the video has been modified by another generative AI process.

[0113] In the example, after decoding one of the images and videos, when generative AI processing is to be applied to one of the images and videos, the generative AI processing can be used to further modify the decoded image or video based on text data.

[0114] In the example, after decoding one of the images and videos, when generative AI processing is not applied to one of the images and videos, the decoded images and videos are not modified based on text data through generative AI processing.

[0115] In the example, the text data includes instructions for enabling generative AI to modify one of the images and videos.

[0116] In the example, generative AI processing is implemented using a neural network post-filtering process that includes one or more neural networks.

[0117] In the example, the SEI message does not indicate that either the image or the video has been modified by other generative AI processing, nor does it indicate that either the image or the video has not been modified by other generative AI processing.

[0118] Figure 12 A flowchart outlining a process (1200) according to one aspect of this disclosure is shown. The process (1200) can be used in a video encoder. In various aspects, the process (1200) is executed by a processing circuit system—for example, a processing circuit system that performs the functions of a video encoder (203), a processing circuit system that performs the functions of a video encoder (403), etc. In some aspects, the process (1200) is implemented as software instructions, so that the processing circuit system executes the process (1200) when the processing circuit system executes the software instructions. The process begins at (S1201) and proceeds to (S1210).

[0119] At (S1210), text data intended for use in generative artificial intelligence (AI) processing can be obtained.

[0120] At (S1220), a video bitstream can be encoded, comprising: (i) one of an image and a video, and (ii) an SEI message associated with one of the image and the video. The SEI message may include text data. The SEI message does not indicate whether one of the image and the video has been modified by either generative AI processing.

[0121] In the example, generative AI processing is implemented using a neural network post-filtering process that includes one or more neural networks.

[0122] In the example, the SEI message indicates post-neural network filter information for post-filtering processing. In the example, text data is carried in the payload of the SEI message. In the example, the SEI message payload includes flags indicating that the remainder of the SEI message payload is yet to be processed.

[0123] Then, the process proceeds to (S1299) and terminates.

[0124] Process (1200) can be adjusted as appropriate. Steps in process (1200) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0125] Figure 13A flowchart outlining one aspect of the processing (1300) according to this disclosure is shown. Process (1300) can be used in a receiving system such as a video decoder. In various aspects, processing (1300) is performed by processing circuitry—for example, a processing circuitry that performs the functions of a video decoder (210), a processing circuitry that performs the functions of a video decoder (310), etc. In some aspects, processing (1300) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes processing (1300). The processing begins at (S1301) and proceeds to (S1310).

[0126] At (S1310), a video bitstream can be received, comprising: (i) one of an image and a video, and (ii) a Supplemental Enhancement Information (SEI) message associated with one of the image and the video. The SEI message may include text data intended for generative artificial intelligence (AI) processing. In the example, the text data includes instructions for causing generative AI processing to modify one of the image and the video. In the example, the text data is carried in the payload of the SEI message. In the example, the payload of the SEI message includes a flag indicating that the remainder of the SEI message payload is pending processing.

[0127] At (S1320), text data can be extracted from the SEI message. On the one hand, the SEI message does not indicate whether one of the images and videos has been modified by another generative AI. In the example, the SEI message does not indicate whether one of the images and videos has been modified by another generative AI, nor does it indicate whether one of the images and videos has not been modified by another generative AI.

[0128] In this example, generative AI processing is applied to one of the images and videos. After decoding one of the images and videos, generative AI processing is used to modify the decoded image or video based on text data.

[0129] In this example, no generative AI processing is applied to either the image or the video. After decoding either the image or the video, no generative AI processing is used to modify the decoded image or video based on text data.

[0130] In the example, generative AI processing is implemented using post-neural filtering, which includes one or more neural networks. In the example, the SEI message indicates the post-neural filter information for the post-neural filtering process.

[0131] Then, the process proceeds to (S1399) and terminates.

[0132] Process (1300) can be adjusted as appropriate. Steps in process (1300) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0133] A method for processing visual media data is disclosed. This method may include processing a bitstream of visual media data according to format rules. The bitstream may include (i) one of an image and a video, and (ii) a Supplemental Enhancement Information (SEI) message associated with one of the images and videos. The SEI message includes text data intended for use in generative artificial intelligence (AI) processing and does not indicate whether one of the images and videos has been modified by another generative AI process. The format rules specify the extraction of text data from the SEI message.

[0134] The techniques described above, such as those involving SEI messages, can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 14 A computer system (1400) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0135] Computer software can be coded using any suitable machine code or computer language. Machine code or computer language can be subjected to mechanisms such as assembly, compilation, and linking to create code that includes instructions. These instructions can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, etc.

[0136] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0137] Figure 14 The components shown for the computer system (1400) are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing the aspects of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of the components shown in the exemplary aspects of the computer system (1400).

[0138] The computer system (1400) may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input made by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device can also be used to capture certain media that are not necessarily directly related to conscious input made by humans, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images acquired from still image capturing devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0139] Human-machine interface input devices may include one or more of the following (only one of each is depicted): keyboard (1401), mouse (1402), touchpad (1403), touch screen (1410), data glove (not shown), joystick (1405), microphone (1406), scanner (1407), and camera device (1408).

[0140] The computer system (1400) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback via a touchscreen (1410), data gloves (not shown), or joystick (1405), but tactile feedback devices that are not used as input devices may also exist); audio output devices (e.g., speakers (1409), headphones (not depicted)); visual output devices (e.g., screens (1410), including CRT (Cathode Ray Tube, CRT) screens, LCD (Liquid Crystal Display, LCD) screens, plasma screens, OLED (Organic Light Emitting Diode, OLED) screens, each screen may or may not have touchscreen input capability, each screen may or may not have tactile feedback capability—some of the screens may be able to output two-dimensional visual output or more than three-dimensional output in a manner such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and ashtrays (not depicted)); and printers (not depicted).

[0141] The computer system (1400) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read-Only Memory, ROM) / RW (1420) having media such as CD / DVD (1421), thumb drives (1422), removable hard disk drives or solid-state drives (1423), conventional magnetic media such as magnetic tape and floppy disks (not depicted), devices based on dedicated ROM / ASIC (Application Specific Integrated Circuit, ASIC) / PLD (Programable Logic Device, PLD) such as security dongles (not depicted), etc.

[0142] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the subject matter disclosed herein does not cover transmission media, carrier waves, or other transient signals.

[0143] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The network may be, for example, wireless, wired, or optical. The network may also be local area, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include: local area networks such as Ethernet and wireless LANs; cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; cable or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and industrial networks including CANBus (Controller Area Network Bus), etc. Some networks typically require external network interface adapters that attach to certain general-purpose data ports or peripheral buses (1449) (such as, for example, the USB (Universal Serial Bus, USB) port of the computer system (1400); other networks are typically integrated into the core of the computer system (1400) by attaching to system buses as described below (e.g., to an Ethernet interface in a PC (Personal Computer, PC) computer system or to a cellular network interface in a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be one-way receiving (e.g., broadcasting TV), one-way transmitting (e.g., to a CANBus device), or bidirectional, such as to other computer systems using local area digital networks or wide area digital networks. Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0144] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1440) of the computer system (1400).

[0145] The core (1440) may include one or more central processing units (CPU) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1443), hardware accelerators (1444) for certain tasks, graphics adapters (1450), etc. These devices, along with read-only memory (ROM) (1445), random access memory (1446), and internal mass storage devices (1447) such as internal non-user-accessible hard disk drives, SSDs (Solid-State Drives), etc., can be connected via the system bus (1448). In some computer systems, the system bus (1448) may be accessed in the form of one or more physical plugs to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1448) or may be attached to the core's system bus (1448) via a peripheral bus (1449). In the example, a screen (1410) may be connected to the graphics adapter (1450). Peripheral bus architectures include PCI (Peripheral Component Interconnect), USB, etc.

[0146] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (1445) or RAM (Random Access Memory) (1446). Transient data can also be stored in RAM (1446), while permanent data can be stored, for example, in an internal mass storage device (1447). Fast storage and retrieval of any memory device in the memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage devices (1447), ROMs (1445), RAMs (1446), etc.

[0147] Computer-readable media may have computer code thereon for performing operations of various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0148] By way of example and not limitation, a computer system (1400) having an architecture, and in particular a core (1440), may be functionalized by a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium may be a medium associated with a user-accessible mass storage device as described above, and certain storage devices of the core (1440) having non-transitory characteristics, such as a mass storage device (1447) or ROM (1445) within the core. Software implementing various aspects of this disclosure may be stored in such a device and executed by the core (1440). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software may cause the core (1440), and in particular the processor therein (including a CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to the processes defined by the software. Alternatively or as an alternative, the computer system may provide functionality by means of hard-wired logic or otherwise embodied in circuitry (e.g., an accelerator (1444)), which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may cover logic, and references to logic may also include software. Where appropriate, references to computer-readable media may cover circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure covers any suitable combination of hardware and software.

[0149] The use of “at least one of…” or “one of…” in this disclosure is intended to include any one or a combination of the described elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include either A or B or (A and B). The use of “one of…” does not exclude any combination of the described elements where applicable, such as when the elements are not mutually exclusive.

[0150] While several exemplary aspects have been described in this disclosure, modifications, substitutions, and various alternative equivalents fall within the scope of this disclosure. It will therefore be understood that those skilled in the art will be able to conceive of numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.

[0151] The above disclosure also covers the features described below. Features can be combined in various ways, and are not limited to the combinations described below.

[0152] (1) A method for processing a video bitstream, the method comprising: receiving a video bitstream comprising: (i) one of an image and a video, and (ii) a supplementary enhancement information (SEI) message associated with the image and the video, the SEI message comprising text data intended for use in generative artificial intelligence (AI) processing; and extracting the text data from the SEI message, wherein the SEI message does not indicate whether one of the image and the video has been modified by another generative AI processing.

[0153] (2) According to the method of feature (1), wherein, after decoding one of the images and videos, when generative AI processing is to be applied to one of the images and videos, the generative AI processing is used to modify one of the decoded images and videos based on text data.

[0154] (3) According to the method of feature (1), wherein, after decoding one of the images and videos, when generative AI processing is not applied to one of the images and videos, the decoded images and videos are not modified based on text data using generative AI processing.

[0155] (4) The method according to any one of features (1) to (3), wherein the text data includes instructions for causing generative AI processing to modify one of the images and videos.

[0156] (5) The method according to any one of features (1) to (4), wherein the generative AI processing is implemented using a neural network post-filtering process comprising one or more neural networks.

[0157] (6) According to the method of feature (5), wherein the SEI message indicates the neural network post-filter information of the neural network post-filtering process.

[0158] (7) The method according to any one of features (1) to (6), wherein the text data is carried in the payload of the SEI message.

[0159] (8) According to the method of feature (7), wherein the payload of the SEI message includes a flag indicating that the remainder of the payload of the SEI message is pending processing.

[0160] (9) The method according to any one of features (1) to (8), wherein the SEI message does not indicate that one of the images and videos has been modified by other generative AI processing, and the SEI message does not indicate that one of the images and videos has not been modified by other generative AI processing.

[0161] (10) A method for generating a supplementary enhancement information (SEI) message, the method comprising: obtaining text data intended for use in generative artificial intelligence (AI) processing; and encoding a video bitstream comprising: (i) one of an image and a video, and (ii) an SEI message associated with one of the image and the video, the SEI message comprising text data, wherein the SEI message does not indicate whether one of the image and the video has been modified by either generative AI processing.

[0162] (11) According to the method of feature (10), the generative AI processing is implemented using a neural network post-filtering process that includes one or more neural networks.

[0163] (12) According to the method of feature (11), wherein the SEI message indicates the neural network post-filter information of the neural network post-filtering process.

[0164] (13) The method according to any one of features (10) to (12), wherein the text data is carried in the payload of the SEI message.

[0165] (14) According to the method of feature (13), wherein the payload of the SEI message includes a flag indicating that the remainder of the payload of the SEI message is pending processing.

[0166] (15) A method for processing visual media data, the method comprising: processing a bitstream of visual media data according to a format rule, wherein the bitstream includes: (i) one of an image and a video, and (ii) a supplementary enhancement information (SEI) message associated with one of the image and the video, the SEI message including text data intended for use in generative artificial intelligence (AI) processing and not indicating whether one of the image and the video has been modified by another generative AI processing; and the format rule specifying the extraction of text data from the SEI message.

[0167] (16) The method according to feature (15), wherein the format rules specify that after decoding one of the images and videos, generative AI processing is used to further modify one of the decoded images and videos based on text data.

[0168] (17) The method according to feature (15), wherein the format rule specifies that after decoding one of the images and videos, the decoded images and videos are not modified using generative AI processing based on text data.

[0169] (18) The method according to any one of features (15) to (17), wherein the text data includes instructions for causing the generative AI to modify one of the images and videos.

[0170] (19) The method according to any one of features (15) to (18), wherein the generative AI processing is implemented using a neural network post-filtering process comprising one or more neural networks.

[0171] (20) The method according to any one of features (15) to (19), wherein the text data is carried in the payload of the SEI message.

[0172] (21) An apparatus for processing video bitstreams, the apparatus comprising a processing circuit system configured to perform the method according to any one of features (1) to (9).

[0173] (22) An apparatus for generating supplemental enhancement information (SEI) messages, the apparatus comprising a processing circuit system configured to perform the method according to any one of features (10) to (14).

[0174] (23) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method according to any one of features (1) to (20).

Claims

1. A method for processing a video bitstream, the method comprising: The video bitstream is received, the video bitstream comprising: (i) one of an image and a video, and (ii) a Supplemental Enhancement Information (SEI) message associated with the image and the video, the SEI message comprising text data intended for use in generative artificial intelligence (AI) processing; and Extract the text data from the SEI message, wherein, The SEI message does not indicate whether one of the images and videos has been modified by another generative AI.

2. The method of claim 1, wherein, After decoding one of the images and videos, when the generative AI processing is to be applied to one of the images and videos, the generative AI processing is used to modify one of the decoded images and videos based on the text data.

3. The method of claim 1, wherein, After decoding one of the images and videos, when the generative AI processing is not applied to one of the images and videos, the generative AI processing is not used to modify one of the decoded images and videos based on the text data.

4. The method of any one of claims 1 to 3, wherein, The text data includes instructions for causing the generative AI process to modify one of the images and videos.

5. The method of any one of claims 1 to 4, wherein, The generative AI processing is implemented using a neural network post-filtering process that includes one or more neural networks.

6. The method of claim 5, wherein, The SEI message indicates the neural network post-filter information for the neural network post-filtering process.

7. The method of any one of claims 1 to 6, wherein, The text data is carried in the payload of the SEI message.

8. The method of claim 7, wherein, The payload of the SEI message includes a flag indicating that the remainder of the SEI message's payload is pending processing.

9. The method of any one of claims 1 to 8, wherein, The SEI message does not indicate that either the image or the video has been modified by other generative AI processing, nor does it indicate that either the image or the video has not been modified by the other generative AI processing.

10. A method for generating Supplemental Enhancement Information (SEI) messages, the method comprising: Obtain text data intended for use in generative artificial intelligence (AI) processing; as well as The video bitstream is encoded, the video bitstream comprising: (i) one of an image and a video, and (ii) the SEI message associated with the one of the image and the video, the SEI message comprising the text data, wherein, The SEI message does not indicate whether either the image or the video has been modified by any generative AI.

11. The method of claim 10, wherein, The generative AI processing is implemented using a neural network post-filtering process that includes one or more neural networks.

12. The method of claim 11, wherein, The SEI message indicates the neural network post-filter information for the neural network post-filtering process.

13. The method of any one of claims 10 to 12, wherein, The text data is carried in the payload of the SEI message.

14. The method of claim 13, wherein, The payload of the SEI message includes a flag indicating that the remainder of the SEI message's payload is pending processing.

15. A method for processing visual media data, the method comprising: The bitstream of the visual media data is processed according to format rules, wherein... The bitstream includes: (i) one of an image and a video, and (ii) a Supplemental Enhancement Information (SEI) message associated with one of the images and videos, the SEI message comprising textual data intended for generative artificial intelligence (AI) processing and not indicating whether one of the images and videos has been modified by another generative AI process; and The formatting rules specify how to extract the text data from the SEI message.