Tagging generative artificial intelligence generated or modified content
By introducing content created or modified by SEI message marker generative AI in video encoding, the problem of identifying AI-generated content is solved, effective content tagging and supervision is achieved, and the reliability of content authenticity verification is improved.
Patent Information
- Application Number
- CN202480005783.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-01
- Filing Date
- 2024-11-20
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is difficult to effectively mark and identify video content created or modified by generative artificial intelligence, making it difficult to verify the authenticity of the content, and there is a social risk of abuse of AI technology.
By introducing auxiliary enhancement information (SEI) messages during the video encoding process, marking the contents of generative AI creation or modification, including tool identification, timestamps and other information, for easy identification and supervision.
It realizes effective marking of content generated or modified by AI, improves the reliability of content authenticity verification, reduces consumer troubles, and complies with regulatory requirements in multiple countries.
Smart Images

Figure CN120418833A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 542,739, filed on October 5, 2023, U.S. Provisional Patent Application No. 63 / 605,391, filed on December 1, 2023, and U.S. Patent Application No. 18 / 903,509, filed on October 1, 2024, the disclosures of which are incorporated herein by reference in their entireties. Technical Field
[0003] Embodiments of the present disclosure relate to video encoding and decoding. Specifically, embodiments of the present disclosure relate to marking the involvement of generative artificial intelligence in the creation or modification of an image or video in the form of Supplemental Enhancement Information (SEI) messages. Background Art
[0004] Uncompressed digital video can consist of a series of pictures, each picture having certain spatial dimensions. For example, it has 1920×1080 luma samples and associated chroma samples. The series of pictures can have a fixed or variable picture rate (also commonly referred to as frame rate), e.g., 60 pictures per second or 60 Hertz (Hz). Uncompressed video has significant requirements for bitrate. For example, a 1080p60 4:2:0 video (1920×1080 luma sample resolution at 60Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. Such a video requires more than 600 GB of storage space for one hour.
[0005] Accordingly, video encoding and decoding can reduce redundancy in the input video signal through compression. Compression can help reduce the above - mentioned bandwidth or storage space requirements, and in some cases, can reduce them by two or more orders of magnitude. Lossless compression, lossy compression, and combinations thereof can all be used for video encoding and decoding.
[0006] Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be used for the intended application. Lossy compression is widely used in video. The amount of distortion tolerated by lossy compression depends on the application; for example, consumer users of certain streaming applications can tolerate higher distortion compared to users of television distribution applications. The achievable compression ratio can reflect that the higher the allowable / tolerable distortion, the higher the compression ratio that can be produced.
[0007] Generative artificial intelligence (hereinafter referred to as "generative AI" or "AI") can be used to create or modify images or image sequences (e.g., videos). For example, in the related art, there are some applications and web pages on the World Wide Web that allow images (pixel maps) to be generated based on a string input into the application. The image can be downloaded and processed by an image compression tool, which includes, for example, an encoder that conforms to one of the still image profiles of the H.266 standard. Thus, image manipulation or video manipulation is possible. For example, so-called deepfakes are known, which obtain a source image or video and manipulate it in a way that the original content creator never intended. A common example is an audiovisual sequence of a politician, where the audio stream has been modified so that the politician makes a statement that he / she would not make, while the video is manipulated to synchronize with the modified audio (lip movements, gestures, etc.).
[0008] The quality of image and video generation / manipulation is such that a trained observer or an image analysis tool is required to identify that the image or video has been created or manipulated by generative AI. Thus, there is a need for annotation to identify that an image or video has been created or manipulated by generative AI, and also for associated standards. Summary of the Invention
[0009] According to an embodiment, there is provided a method and apparatus including computer code for video processing. The method includes setting a first value of a text description usage parameter in a bitstream, the first value of the text description usage parameter indicating the type of information included in an information string of text description information in the bitstream; when the first value indicates that the type of information included in the information string of text description information includes marker information associated with one or more artificial intelligence (AI) processing procedures used, setting artificial intelligence marker information associated with one or more pictures in the bitstream as the information string of text description information; signaling the text description usage parameter in the bitstream; and signaling the information string of text description information.
[0010] According to an embodiment, a device for video processing is provided. The device may include at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code. The program code may include receiving code configured to cause the at least one processor to receive an encoded video bitstream including pictures; first obtaining code configured to cause the at least one processor to obtain a first value of a text description usage parameter in the encoded video bitstream, the first value of the text description usage parameter indicating the type of information included in the text description information string in the encoded video bitstream; second obtaining code configured to cause the at least one processor to obtain artificial intelligence marking information associated with one or more pictures in the encoded video bitstream as the text description information string in the encoded video bitstream when the first value indicates that the type of information included in the text description information string includes marking information, and the marking information is associated with one or more artificial intelligence (AI) processing processes used; and reconstruction code configured to cause the at least one processor to reconstruct pictures from the encoded video bitstream based on the first value and the AI marking information.
[0011] According to an embodiment, a non-volatile computer-readable medium storing instructions is provided. The instructions may include: one or more instructions that, when executed by one or more processors of a device, cause the one or more processors to perform a conversion between a visual media file and a bitstream of the visual media file, where the bitstream includes at least one of the following: a text description usage parameter indicating the type of information included in the text description information string; and artificial intelligence (AI) marking information associated with one or more pictures in the visual media file, and when the text description usage parameter indicates that the type of information included in the text description information string includes marking information, and the marking information is associated with one or more artificial intelligence processing processes used, the AI marking information is the text description information string; where the format rule specifies that the text description usage parameter and the text description information string are syntax parameters in a supplementary enhancement information (SEI) message. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0013] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment of the present disclosure.
[0014] Figure 2 is a schematic diagram of a simplified block diagram of a streaming system according to an embodiment of the present disclosure.
[0015] Figure 3Schematic diagram of a simplified block diagram of a video decoder according to an embodiment of the present disclosure.
[0016] Figure 4 Schematic diagram of a simplified block diagram of a video encoder according to an embodiment of the present disclosure.
[0017] Figure 5 Schematic diagram of a NAL unit and SEI header according to an embodiment of the present disclosure.
[0018] Figure 6 Illustration of SEI message syntax according to an embodiment of the present disclosure.
[0019] Figure 7 Schematic diagram of a system according to an embodiment of the present disclosure.
[0020] Figure 8 Exemplary diagram of a computer system suitable for implementing various embodiments. Detailed Description of Specific Embodiments
[0021] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0022] Figure 1 Illustrates a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional transmission of data, a first terminal 103 may encode video data at a local location for transmission via the network 105 to another terminal 102. The second terminal 102 may receive the encoded video data of another terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media service applications and the like.
[0023] Figure 1 Illustrates a second pair of terminals 101 and 104, which are provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional transmission of data, each terminal 101 and 104 may encode video data captured at a local location for transmission via the network 105 to another terminal. Each terminal also may receive the encoded video data transmitted by another terminal, may decode the encoded data, and may display the recovered video data on a local display device.
[0024] In Figure 1Among them, the terminals 101, 102, 103, and 104 can be illustrated as servers, personal computers, and smart phones, but the principles disclosed in this application are not limited thereto. The embodiments disclosed in this application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 105 represents any number of networks for transmitting encoded video data among the terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. The communication network 105 can exchange data in circuit-switched and / or packet-switched channels. The network can include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of the network 105 may be irrelevant to the operations disclosed in this application.
[0025] As an example, Figure 2 illustrates the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-supported applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0026] The streaming system can include an acquisition subsystem 203, and the acquisition subsystem can include a video source 201 such as a digital camera, and the video source, for example, creates an uncompressed video sample stream 213. Compared with the encoded video bitstream, the video sample stream 213 is emphasized as a high-data-volume video sample stream, which can be processed by an encoder 202 coupled to the camera 201. The encoder 202 can include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the sample stream, the encoded video bitstream 204 is emphasized as a lower-data-volume encoded video bitstream, which can be stored on the streaming server 205 for future use. One or more streaming clients 212 and 207 can access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 can include a video decoder 211, and the video decoder 211 decodes the incoming copy 208 of the encoded video bitstream and generates an output video sample stream 210 that can be presented on a display 209 or another presentation device. In some streaming systems, the video bitstreams 204, 206, and 208 can be encoded according to certain video coding / compression standards. Examples of these standards are mentioned above and are further described herein.
[0027] Figure 3 is a functional block diagram of a video decoder 300 according to an embodiment of the present application.
[0028] The receiver 302 may receive one or more codec video sequences to be decoded by the decoder 300; in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from the channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data and other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities. The receiver 302 may separate the encoded video sequences from the other data. To prevent network jitter, the buffer memory 303 may be coupled between the receiver 302 and the entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer 303, or the buffer may be made smaller. For use on a service packet network such as the Internet, the buffer 303 may also be required, which may be relatively large and may have an adaptive size.
[0029] The video decoder 300 may include a parser 304 to reconstruct symbols 313 from the entropy-encoded video sequences. The categories of these symbols include information for managing the operation of the decoder 300 and potential information for controlling a display device such as the display 312, which is not a part of the decoder but may be coupled to the decoder. The control information for the display device may be a parameter set segment of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser 304 may parse / entropy-decode the received encoded video sequences. The encoding of the encoded video sequences may be performed according to video coding techniques or standards and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequences based on at least one parameter corresponding to a group. The subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Unit (CU), blocks, Transform Unit (TU), Prediction Unit (PU), etc. The entropy decoder / parser 304 may also extract information from the encoded video sequences, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0030] The parser 304 can perform entropy decoding / parsing operations on the video sequence received from the buffer 303, thereby creating symbols 313. The parser 304 can receive the encoded data and selectively decode the special symbols 313. In addition, the parser 304 can determine whether to provide the special symbols 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0031] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter picture and intra picture, inter block and intra block) and other factors, the reconstruction of the symbol 313 may involve multiple different units. Which units are involved and the way they are involved can be controlled by the subgroup control information parsed by the parser 304 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 304 and multiple units below are not described.
[0032] In addition to the functional blocks already mentioned, the decoder 300 can be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.
[0033] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives the quantized transform coefficients as symbols 313 and control information from the parser 304, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit 305 can output a block including sample values, and the sample values can be input into the aggregator 310.
[0034] In some cases, the output samples of the scaler / inverse transform unit 305 may belong to an intra-coded block; that is, a block that does not use the predictive information from the previously reconstructed picture but may use the predictive information from the previously reconstructed part of the current picture. Such predictive information can be provided by the intra picture prediction unit 307. In some cases, the intra picture prediction unit 307 generates a surrounding block with the same size and shape as the block being reconstructed using the reconstructed information extracted from the partially reconstructed current picture 309. In some cases, the aggregator 310 adds the prediction information generated by the intra prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 based on each sample.
[0035] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit 306 may access the reference picture memory 308 to extract samples for prediction. After motion-compensating the extracted samples according to the symbol 313, these samples may be added by the aggregator 310 to the output of the scaler / inverse transform unit (referred to as residual samples or a residual signal in this case) to generate output sample information. The motion compensation unit obtaining the prediction samples from an address within the reference picture memory may be controlled by a motion vector, and the motion vector is in the form of the symbol 313 for use by the motion compensation unit, where the symbol 313 may include, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of the sample values extracted from the reference picture memory, a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.
[0036] The output samples of the aggregator 310 may be employed by various loop filtering techniques in the loop filter unit 311. The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters may be used for the loop filter unit 311 as the symbol 313 from the parser 304. However, in other embodiments, the video compression technique may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0037] The output of the loop filter unit 311 may be a sample stream that may be output to the display device 312 and stored in the reference picture memory 557 for subsequent inter-picture prediction.
[0038] Once fully reconstructed, certain encoded pictures may be used as reference pictures for future prediction. Once an encoded picture has been fully reconstructed and the encoded picture is identified (by, for example, the parser 304) as a reference picture, the current reference picture 309 may become part of the reference picture buffer 308, and a new current picture buffer may be reallocated before starting to reconstruct subsequent encoded pictures.
[0039] Video decoder 300 may perform decoding operations according to a predetermined video compression technique documented in, for example, the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard being used, i.e., conform to the video compression technique or standard specified in the video compression technique document or standard, particularly the syntax of the video compression technique or standard specified in its profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.
[0040] In an embodiment, receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by video decoder 300 to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0041] Figure 4 May be a functional block diagram of video encoder 400 according to an embodiment disclosed in the present application.
[0042] Encoder 400 may receive video samples from video source 401 (not part of the encoder), and the video source may capture video images to be encoded by encoder 400.
[0043] Video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by encoder 303. The digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCb, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, video source 401 may be a storage device storing previously prepared video. In a video conferencing system, video source 401 may be a camera that captures local image information as a video sequence. Video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial arrays of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0044] According to an embodiment, encoder 400 may encode and compress pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by an application. Enforcing an appropriate encoding speed is a function of controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For simplicity, the couplings are not labeled in the figure. Parameters set by the controller may include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will easily recognize other functions of controller 402 as these functions relate to video encoder 400 optimized for a particular system design.
[0045] The working principle of some video encoders is the "encoding loop" that is easily recognizable to those skilled in the art. As a simple description, the encoding loop can consist of the encoding part of encoder 402 (hereinafter referred to as the source encoder), which is responsible for creating symbols based on the input picture to be encoded and reference pictures, and the (local) decoder 406 embedded in encoder 400. The decoder 406 reconstructs the symbols to create sample data that the (remote) decoder will also create (because in the video compression technology considered in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to the reference picture memory 405. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content of the reference picture buffer is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well-known to those skilled in the art.
[0046] The operation of the "local" decoder 406 can be the same as that of the "remote" decoder 300 described in detail above in combination with Figure 3 However, briefly referring additionally to Figure 4 , when the symbols are available and the entropy encoder 408 and the parser 304 can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the decoder 300, including the channel 301, the receiver 302, the buffer 303, and the parser 304, may not be fully implemented in the local decoder 406.
[0047] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions are only required in certain areas and are provided below.
[0048] As part of its operation, the source encoder 403 can perform motion-compensated predictive coding. Referring to one or more previously encoded frames designated as "reference frames" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 407 encodes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, and the reference frame can be selected as the prediction reference for the input frame.
[0049] The local video decoder 406 can decode the encoded video data of the frames that can be specified as reference frames based on the symbols created by the source encoder 403. The operation of the encoding engine 407 can be a lossy process. When the encoded video data can be decoded at the video decoder, the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that can be performed by the video decoder on the reference frames and can cause the reconstructed reference frames to be stored in the reference picture cache 405. In this way, the encoder 400 can locally store a copy of the reconstructed reference frames, which has the same content (without transmission errors) as the reconstructed reference pictures that will be obtained by the remote video decoder.
[0050] The predictor 404 can perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 can search in the reference picture memory 405 for sample data (as candidate reference pixel blocks) or some metadata that can be used as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor 404 can operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor 404, it can be determined that the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 405.
[0051] The controller 402 can manage the encoding operations of the video encoder 403, including, for example, setting parameters and subgroup parameters for encoding the video data.
[0052] The outputs of all the above functional units can be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., so as to convert the symbols into an encoded video sequence.
[0053] The transmitter 409 can buffer the encoded video sequence created by the entropy encoder 408 to prepare for transmission through the communication channel 411, which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter 409 can merge the encoded video data from the video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams.
[0054] The controller 402 can manage the operation of the encoder 400. During encoding, the controller 405 can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can generally be assigned to any of the following frame types:
[0055] An intra picture (I picture) is a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.
[0056] A predictive picture (P picture) is a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0057] A bi-predictive picture (B picture) is a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0058] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined according to the coding assignment of the corresponding pictures applied to the blocks. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0059] Video encoder 400 can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, video encoder 400 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0060] In an embodiment, transmitter 409 can transmit additional data when transmitting the encoded video. Source encoder 403 can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, SEI messages, VUI parameter set fragments, etc.
[0061] Compressed video can be enhanced in the video bitstream by auxiliary enhancement information (e.g., in the form of Supplemental Enhancement Information (SEI) messages, or in the form of Video Usability Information (VUI)). Video coding standards can include specification sections for SEI and VUI. SEI information and VUI information can also be specified in a separate specification that can be referenced by the video coding specification.
[0062] This application discloses an SEI message for marking content created and / or modified by a generative AI engine to improve the technical implementation of potential regulatory requirements and to help improve the identification of AI-generated content. The disclosed subject matter further relates to video encoding and decoding, and more particularly, to including a mark in an SEI message to mark the content of a video segment generated using generative artificial intelligence.
[0063] Figure 5 An exemplary layout of an encoded video sequence (CVS) according to the H.266 standard is illustrated. The encoded video sequence is subdivided into Network Abstraction Layer units (NAL units). An exemplary NAL unit 501 can include a NAL unit header 502, which in turn includes 16 bits as follows: forbidden_zero_bit 503 and nuh_reserved_zero_bit 504, which may not be used by the H.266 standard and may be zero in a NAL unit compliant with the H.266 standard. Three bits of nuh_layer_id 505, which can indicate the layer (spatial, SNR, or multi-view enhancement) to which the NAL unit belongs. Five bits of nuh_nal_unit_type, which defines the type of the NAL unit. In the H.266 standard, 22 NAL unit type values are defined for the NAL unit types defined in the H.266 standard, six NAL unit types are reserved, and four NAL unit type values are unspecified and can be used by specifications other than the H.266 standard. Finally, three bits of the NAL unit header indicate the temporal layer nuh_temporal_id_plusl506 to which the NAL unit belongs.
[0064] An encoded picture can contain one or more Video Coding Layer (VCL) NAL units, and zero or more non-VCL NAL units. VCL NAL units can contain encoded data that conceptually belongs to the video coding layer introduced above. Non-VCL NAL units can contain data that conceptually belongs to the video coding layer and can be characterized in the following:
[0065] (1) Parameter sets, which include information that is necessary for the decoding process and can be applied to more than one coded picture. Parameter sets and conceptually similar NAL units can be NAL unit types such as DCI_NUT (Decoding Capability Information (DCI)), VPS_NUT (Video Parameter Set (VPS), establishing layer relationships, etc.), SPS_NUT (Sequence Parameter Set (SPS), establishing parameters used and remaining constant in the entire coded video sequence (CVS), etc.), PPS_NUT (Picture Parameter Set (PPS), establishing parameters used and remaining constant within a coded picture, etc.), and PREFIX_APS_NUT and SUFFIX_APS_NUT (prefix and suffix adaptation parameter sets). Parameter sets can include information required by the decoder to decode VCL NAL units and are thus referred to as "normative" NAL units.
[0066] (2) Picture Header (PH_NUT), which is also a "normative" NAL unit.
[0067] (3) NAL units that mark certain positions in the NAL unit stream. These NAL units include NAL units with the following NAL unit types: AUD_NUT (Access Unit Delimiter), EOS_NUT (End of Sequence), and EOB_NUT (End of Bitstream). These NAL units are non-normative and are also referred to as informative in the sense that a conforming decoder does not require them in its decoding process, although the decoder needs to be able to receive them in the NAL unit stream.
[0068] (4) Prefix and suffix SEI NAL unit types (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), which indicate NAL units containing prefix auxiliary enhancement information and suffix auxiliary enhancement information. In the H.266 standard, these NAL units are informative as they are not required for the decoding process.
[0069] (5) Filler data NAL unit type FD_NUT, which indicates filler data; the data can be random and can be used to "consume" bits in the NAL unit stream or bitstream, which may be necessary for transmission in certain synchronous transmission environments.
[0070] (6) Reserved and undefined NAL unit types.
[0071] Reference Figure 5, which is a layout of a NAL unit stream containing encoded pictures 511 in decoding order 510, including some of the types of NAL units introduced previously. Somewhere early in the NAL unit stream, the DCI 512, VPS 513, and SPS 514 can be combined to establish parameters that the decoder can use to decode the encoded pictures (including the encoded pictures 511 of the NAL unit stream) of an encoded video sequence (CVS).
[0072] The encoded picture 511 can contain: a prefix APS 516, a picture header 517, a prefix SEI 518, one or more VCL NAL units 519, and a suffix SEI 520 (in the order depicted, or any other order consistent with the video coding technology or standard in use).
[0073] During standard development, prefix SEI NAL units and suffix SEI NAL units (e.g., 518 and 520) were created for the following reasons: For some SEI messages, the content of the message is known before the encoding and decoding of a given picture begins, while other content can only be known after the picture has been encoded and decoded. By having prefix SEI and suffix SEI, allowing certain SEI messages to appear earlier or later in the NAL unit stream of an encoded picture can avoid buffering. As an example, in the encoder, the sampling time of the picture to be encoded is known before the picture is encoded, so the picture timing SEI message can be a prefix SEI message 516.
[0074] On the other hand, the decoded picture hash SEI message (which contains the hash of the sample values of the decoded picture and can be used, for example, to debug the implementation of the encoder) is a suffix SEI message 518 because the encoder cannot compute the hash of the reconstructed samples before the picture has been encoded. The positions of the prefix SEI NAL unit and the suffix SEI NAL unit are not limited to their positions in the NAL unit stream. The phrases "prefix" and "suffix" can imply which encoded pictures or NAL units the prefix / suffix SEI messages may pertain to, and the details of this applicability can be described, for example, in the semantic description of a given SEI message.
[0075] Refer again to Figure 5, which is a simplified syntax diagram of a NAL unit containing a prefix SEI message or a suffix SEI message 520. This syntax is a container format for multiple SEI messages that can be carried in a single NAL unit. For clarity, the anti-collision syntax details specified in the H.266 standard are omitted here. As with other NAL units, the SEI NAL unit starts with a NAL unit header 521. After the NAL unit header is one or more SEI messages; two SEI messages (e.g., 530, 531) are depicted and described below. Each SEI message within the SEI NAL unit includes an 8-bit payload_type_byte 522, which specifies one of 256 different SEI types; an 8-bit payload_size_byte 523, which specifies the number of bytes of the SEI payload, and a payload of payload_size_byte bytes 524. This structure can be repeated until a payload_type_byte equal to 0xff is observed, which indicates the end of the NAL unit. The syntax of the payload 524 depends on the SEI message, and the payload 524 can be of any length between 0 and 255 bytes.
[0076] From a social perspective, generative AI has become a challenging new technology. Until a few years ago, modifying or "tampering" with media such as audio, images, and videos required specialized, highly skilled personnel and often a significant amount of time and cost, but now, it can be achieved for free or at a relatively low cost by individuals with access to a computer and the internet. Content creation is increasingly relying on AI-based tools. Many people view using artificial intelligence to create new original content as just another tool, similar to a paintbrush or a violin.
[0077] However, using AI to create content can be for malicious reasons (forging the speeches of politicians or other media stars, creating false documentary evidence that may be mistaken for original, etc.). The quality of such modifications is good enough that trained observers are needed to distinguish between false and real (microphone / camera-captured) content without relying on semantics. It is widely believed that the risk of abusing this technology is significant enough to endanger society as a whole. Therefore, AI marking technology is needed.
[0078] In this disclosure, the terms creation, modification, and generation are used. Creation can refer to the situation where AI creates media relying only on its own models / databases and content generally available on the internet without specific input media guidance. In contrast, AI modifying content means that AI modifies the media supplied by the user based on the user's instructions. Generation is used to refer to either creation or modification.
[0079] Explicit marker / warning segments in audio or visual form (such as screen content showing letters / words / sentences indicating the use of AI) take up the consumer's time and are therefore generally considered annoying and are easily deleted. In addition, if the content mixes AI creation and / or AI modification with original content, it may be difficult and even more annoying to mark only the AI-based segments without marking the remaining content, without further increasing the consumer's annoyance.
[0080] Watermarking is a technique that can be employed in audio and video. This technique marks AI-generated content, for example, by embedding invisible / ininaudible markers, or markers that can be visible / audible but are not as annoying as explicit marker segments. Watermarks come in various forms, and some of these forms can persist through possible multiple encoding / decoding steps. Invisible watermarks can be made to attract the user's attention through a dedicated watermark-finding application. Watermarking can conceptually be applied in the compressed domain or the source / reconstruction domain, and techniques exist for both. Watermarks can be designed such that they are difficult to remove without damaging the audio / video content. However, the computational cost of encoding and decoding watermarks is high, and for some media types, there is currently no watermark standard that allows for interoperable implementations. Watermarks also consume a large amount of bitrate overhead; especially for more covert and more robust versions of watermarks. This is even more so if the information conveyed within the watermark goes beyond a simple Boolean on / off signal.
[0081] AI marking can also be done by inserting metadata into the audiovisual stream. Metadata can be very efficient and can be placed at the exact location of a picture / audio frame, enabling precise marking of AI-generated content. Metadata can convey as much or as little information as needed and has a low computational cost in many cases. However, in many cases, metadata can be easily modified or removed from the content.
[0082] According to an embodiment, all three of the above-mentioned marking options (and possibly other marking options) can be used. The marker segments can be short and can be in a form that enables traditional consumer devices to present an AI warning. Watermarks can serve as a reasonable anti-interference marking mechanism for individual or a group of audio frames, pictures, or AI-based video segments. Finally, metadata may be the best technique for identifying what modifications have been applied, which AI engine has been used, the timestamp of creation / modification, etc.
[0083] According to an embodiment, metadata in the form of Supplementary Enhancement Information (SEI) messages can be used for marking purposes.
[0084] A NAL unit may carry one or more SEI messages. The concept of SEI messages can also be applied to other NAL unit-based coding techniques, even if the term "SEI message" may not exist in those specifications. Additionally, the disclosed subject matter can also be applicable outside of NAL unit-based codecs, as long as the bitstream format of the codec or bitstream formatter being discussed allows for the insertion of a bitstream into the encoded media bits and its format can be associated with metadata. Such codecs or bitstream formatters can include audio or system layer multiplexing standards. For clarity, the following description takes the aforementioned SEI message syntax or a syntax closely related thereto, but those skilled in the art will be able to adapt or modify the subject matter of the present invention to meet the application requirements outside of SEI messages. Additionally, the following description focuses on a video stream consisting of one or more pictures encoded in a format that supports SEI messages similar to the H.266 standard, but the present disclosure is not limited thereto.
[0085] According to an embodiment, the AI mark may at least include the following:
[0086] · A content identifier that identifies the content as AI-generated content / AI-modified content;
[0087] · An identifier of the tool used to create / modify the content; and
[0088] · A timestamp of the content creation / content modification.
[0089] Other possible information that can be considered is the identifier of the content used before the modification (if any), and the instructions received by the AI for creating / modifying the content.
[0090] In an embodiment, the metadata suitable for marking one or more encoded pictures in an encoded video sequence as being generated or modified by AI may have a syntax as Figure 6 shown. Two syntaxes are disclosed, but the present disclosure is not limited thereto.
[0091] As Figure 6As shown, the ai_mark_cancel_flag 610, when it is true, can cancel the previously received ai_mark 601 SEI message and reset the status of any variables related to the AI mark to undefined - however, also see the ai_mark_persistence_flag 612 below. When the flag ai_mark_cancel_flag 610 is false 611, the current picture can be at least marked as being generated by AI. Additionally, this syntax can enable the existence of additional syntax elements, the combination of which can mark the current picture or the next picture (depending on whether the SEI message is a prefix SEI message or a suffix SEI message) as being generated by AI or modified by AI.
[0092] When the ai_mark_persistence_flag 612 is 0, it can indicate that the remaining information within the ai_mark SEI 601 is only related to the current decoded picture. When the ai_mark_persistence_flag 612 is 1, it can indicate that the SEI message for marking AI can be applied to the current decoded picture and can persist to all subsequent pictures in output order until one or more of the following conditions become true:
[0093] · A new encoded video sequence starts;
[0094] · The bitstream ends; and
[0095] · A picture with an ai_mark SEI message is output, which follows the current picture in output order.
[0096] The form of the ai_mark SEI message 601 can be designed to include the marking information in a string 620. In the same embodiment or another embodiment, the format of the string 620 can be free form. In some embodiments, there can be a text description usage parameter that indicates the purpose of the text description SEI message (e.g., the string 620). The value of the text description usage parameter can be from 0 to 255, including 0 and 255. As an example, a value of 2 for the text description usage parameter can indicate that the string 620 contains AI marking information related to one or more pictures. It should be noted that Figure 6 is an exemplary, rather than a restrictive example.
[0097] In the same embodiment or another embodiment, when the string 620 is not an empty string, the string 620 may contain AI tagging information related to one or more pictures within the duration of the SEI message. Additionally, the string may contain information about machine learning-based processing, the intended use of the decoded pictures, or other aspects related to the associated pictures.
[0098] In the same embodiment or another embodiment, semantics may require that the string 620 be in a certain structured format. This format may be defined in the semantics. As an example, it may be required that the format of the string conform to the c-style format string "Tool: %s Timestamp: %ld". Those familiar with the C programming language will readily understand that %s refers to a null-terminated character sequence (where the null is ignored when printing the character sequence using printf() or sprintf()), and %ld refers to a long (e.g., 64-bit length) signed integer. Semantics may require that the character sequence replacing %s can be a suitable identifier of the AI tool used to generate the content, such as a URI, and the 64-bit integer replacing %ld can be the time when the AI-generated content was created, measured in seconds starting from January 1, 1970, which is called Unix time or epoch time.
[0099] In the same embodiment or another embodiment, the format of the string 620 can be in json format, where the json codepoints include but are not necessarily limited to the tool identifier and the date and time. In this case, the string 620 can be as follows:
[0100] {
[0101] "tool": "https: / / www.starryai.com"
[0102] "time": "Nov 27,2023,14:33Z" [[ID=z16]]
[0103] }
[0104] where "https: / / www.starryai.com" is the URI of the AI tool in use, and "Nov 27,2023,14:33Z" represents the timestamp when the AI engine was running. The above time code is self-evident to those skilled in the art. The "Z" at the end of the time code represents the "Zulu" time zone, commonly known as Coordinated Universal Time.
[0105] In the same or another embodiment, the string 620 may follow other structured content representation languages such as, for example, XML or ASN.1. In an embodiment, the syntax restrictions may be detailed in a free-form language in the semantics, or may be enforced by mandating the use of a schema or equivalent.
[0106] Further referring Figure 6 , in an embodiment, the information that the standards setting committee deems necessary may also be enforced by using the H.266 standard (or equivalent) syntax. The advantage of doing so is high encoding and decoding efficiency, because the video coding syntax is designed with minimal overhead in mind, such as using SEI with the marker AI, which is much more efficient than character-based syntax mechanisms such as JSON or XML.
[0107] In addition to the ai_mark_cancel_flag 610 and the ai_mark_persistence_flag 612, the SEI message 630 may also include the ai_tool_id_string 631 and the ai_timestamp_string 632 that are unconditionally present. The term "unconditionally" here means that they are not gated by a single presence flag, which is contrary to the syntax elements introduced later.
[0108] The semantics may require that the ai_tool_id_string 631 contain appropriate identification information for the AI tool used to generate the video content, such as, for example, a URI. Similar concepts have been introduced above in the context of the JSON syntax.
[0109] To encode the timestamp 632, a reasonable option may be a timestamp string conforming to RFC 3339. RFC 3339 specifies a subset of some of the options provided in ISO IS 8601 and mandates certain fields, including year, month, day, hour, and minute (where only the year is mandated in ISO 8601). In this regard, RFC 3339 is sometimes considered a practical variant of ISO8601 for specifying time on the Internet. For example, a string in the RFC 3339 format may be in a format such as "2023-11-27T14:33Z". However, many other date / time representations are also known in the art and may equally be used to encode the timestamp 632. Advantageously, the semantic description of the SEI message may specify the format to be used to enable automatic parsing.
[0110] If the regulatory language introduced is the same globally and in the foreseeable future, SEI messages with only the content presented so far may be sufficient for AI tagging purposes. However, SEI messages cannot be extended to remain backward compatible. In other words, a format needs to be found that ideally includes all the data elements that regulatory bodies or legislatures worldwide may require in the foreseeable future. Creating such an exhaustive list may be difficult to achieve. Additionally, regulatory bodies in different countries may set requirements that are not only different but also conflicting, such as when it comes to the presence of information related to individuals and different priorities regarding privacy rights and tagging requirements.
[0111] Embodiments of the present disclosure include and provide methods to overcome the above difficulties.
[0112] In an embodiment, a field can be gated by a presence flag. In this example, as an illustration, ai_person_id634 is gated by ai_person_id_presence_flag 633. If the presence flag 633 is 0, the syntax enforces that the ai_person_id string 634 does not exist; if the presence flag 633 is 1, the string ai_person_id634 exists. For example, this mechanism can be used as follows: When generating AI in a country where the regulatory body requires the inclusion of a personal identifier, then the AI needs to ask the individual it represents for such information and make it available to the bitstream encoder to include in the SEI message, where the flag 633 is set to 1 and the personal id is encoded into the string 634. On the other hand, in a country where personal privacy is prioritized over tagging, the AI will not ask the user, and the encoder will set the flag 633 to 0, thus ignoring the string 634 with the personal identifier. The advantage of this mechanism is that for conditionally mandatory information (conditioned on the regulatory body mandating its presence), only a single bit is used, unless detailed information is required.
[0113] In the same or another embodiment, without using a gating flag, semantics can also indicate that, from a standard perspective, a string of length 0 is sufficient. For integers, values that are unlikely or impossible to occur can be used. For dates, a common date earlier than the generative AI application can be used, such as January 1, 1970.
[0114] Regarding the lack of scalability of SEI messages, in the same or another embodiment, one solution could be that, in addition to the various syntax elements described above (some of which are gated by presence flags, or allow strings of length zero, or equivalents for other date types as described above), it could also include a free-form string "ai_other_regulatory" 635, into which AI / encoders compliant with regulatory requirements can place relevant tagging information that cannot be represented by other fields.
[0115] In the same or another embodiment, the SEI message can contain an instruction 637 gated by a presence flag 636, which is used by the AI to generate content. Such an instruction can be in the form of a URI, a plain text field, a JSON-encoded string, etc.
[0116] In the same or another embodiment, it may be advantageous to supply this information when the AI is modifying user-supplied content rather than creating content based on, for example, its internal model. As a possible implementation, there could be a string ai_base_content 639 gated by a presence flag 638. For example, the string could include a comma-separated list of URIs pointing to the content being modified. For example, in ai_bas_content 639, an empty string could indicate that specific content has been provided to the AI, but that content cannot be obtained using a URI (e.g., a live stream that has not been stored, private content not available on the Internet, etc.). In this case, another role of the presence flag 638 is to indicate that existing content has been modified without identifying the content itself.
[0117] Figure 7 A system involving generative AI is illustrated. A producer user 701 may have created an instruction 702 for a generative AI engine 704. These instructions can be free-form text of a few sentences (e.g., "Create a 10-second video showing a panda riding a bicycle through the city"), or more complex. For example, the instruction itself could be a video stream providing the instruction in the form of sign language. In the context of AI creation, as introduced above, these instructions can be the information supplied by the user to which the AI responds uniquely. In the context of AI modification, the AI can also receive user-specified content 703 to be modified, shown here as a film reel. Hybrid forms of these two techniques are also possible; for example, the content 703 to be modified can include metadata that contains instructions for the AI engine 704 on how to modify the content.
[0118] Details of the mechanism for supplying user-provided instructions 702 to AI engine 704 and for supplying only modified content 703 when content is modified may be omitted here. In some cases, the physical location of the AI engine may be in a data center operated by a third-party provider, in which case it is likely that instructions 702 and content 703 are supplied to it via the Internet, for example, using an interactive web page. However, other forms of transmitting instructions 702 and content 703 from user 701 to AI engine 704 are also conceivable.
[0119] AI engine 704 may adopt instructions 702 and content 703 (in the modification use case) to create an output medium that may be uncompressed. In this example, the medium is a series of uncompressed images (and timing information) that together form an uncompressed video stream 705. The AI engine may also generate certain metadata 706, for example, including data representing various syntax elements described above in conjunction with Figure 6 the various syntax elements described.
[0120] Both the uncompressed video stream 705 and the metadata 706 may be input into a video encoder 707. The video encoder 707 may compress the uncompressed video stream 705 into a compressed video stream 708. The compressed video stream 708 may include SEI messages 709 associated with certain encoded pictures in the encoded video bitstream 708. Three encoded pictures (e.g., 710, 711, 712) are shown in this example, and only the first encoded picture 710 has an associated SEI message 709.
[0121] The encoded video bitstream 708 may be distributed to consuming users in a manner known to those skilled in the art. Consuming users may include the producing user 701. As depicted here, the encoded video stream 709 is stored in a file 713 and then streamed by a streaming server 714 to the terminal of the consuming user via a network 715. The terminal may include a video decoder 719 for reconstructing the streamed encoded video stream 709 into a series of decoded pictures 720. The decoder 719 may also extract AI marker information and other metadata from the SEI messages 709 included in the encoded video stream 708. Thus, the AI marker information is available at the terminal of the consuming user at this moment.
[0122] The received marker information may be used in accordance with the regulations of the regulatory environment in which the decoder and the display are located, which may be different from the regulatory environment that specifies what content the AI needs to include.
[0123] Different scenarios that can be envisioned include:
[0124] (1) Regulations may mandate that consumers be informed that content is AI-generated. In this case, at the beginning of a video segment marked as containing AI-generated content, the decoder 719, renderer 720, or display 721 may insert a visible warning on the screen 722, visible to the consumer 723. Such a warning may be in the form of text, a logo, a sound, or any other suitable form.
[0125] (2) The decoder 719, renderer 720 or display 721 may be required to provide the consuming user with the option of being warned by the aforementioned device. Alternatively or in addition, the user may be able to set his decoder 719, renderer 720, display 721 so that the AI-generated content may not be displayed, which is similar to the MPAA ratings that are widely used to mark whether content is suitable for certain age groups to consume with or without adult supervision. A combination of MPAA enforcement and AI-based labeling warnings is also possible. For example, a terminal may be set so that it treats AI-based content as "R" rated, in which case, assuming the terminal is configured correctly, only mature audiences will be exposed to the AI-based content.
[0126] In an even more permissive setup, the decoder / renderer / display may ignore the marker SEI message.
[0127] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Figure 8 A computer system 800 suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0128] The computer software may be encoded using any suitable machine code or computer language that may be subjected to assembly, compilation, linking, or similar mechanisms to create code comprising instructions that may be executed directly or through interpretation, microcode execution, or the like by one or more computer central processing units (CPUs), graphics processing units (GPUs), or the like.
[0129] The instructions may be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, and the like.
[0130] Figure 8 The components shown for computer system 800 are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present application. Nor should the configuration of components be interpreted as having any dependency or requirement on any one or combination of components shown in the exemplary embodiment of computer system 800.
[0131] The computer system 800 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs by one or more human users through, for example, tactile inputs (such as: key presses, swipes, data glove movements), audio inputs (such as: speech, taps), visual inputs (such as: gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that may not be directly related to conscious human input, such as audio (such as: conversations, music, ambient sounds), images (such as: scanned images, photographic images obtained from a still image camera), and video (such as, two-dimensional video, three-dimensional video including stereoscopic video).
[0132] The input human-machine interface devices may include one or more of the following (each depicted only one): keyboard 801, mouse 802, trackpad 803, touch screen 810, joystick 805, microphone 806, scanner 808, camera 807.
[0133] The computer system 800 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as, the tactile feedback of the touch screen 810 or the joystick 805, but there may also be tactile feedback devices that do not act as input devices), audio output devices (such as: speakers 809, headsets (not depicted)), visual output devices (such as, screen 810, including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or output greater than three-dimensional through, for example, stereoscopic flat painting output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted).
[0134] The computer system 800 may also include human-accessible storage devices and associated media of the storage devices, such as, optical media, including CD / DVD ROM / RW 820 with media such as CD / DVD (821), thumb drives 822, removable hard disk drives or solid state drives 823, legacy magnetic media such as tapes and floppy disks (not depicted), dedicated devices based on ROM / application specific integrated circuit (ASIC) / programmable logic device (PLD), such as, security protection devices (not depicted), and so on.
[0135] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0136] The computer system 800 may also include an interface 899 to one or more communication networks 898. The network 898 may be, for example, wireless, wired, or optical. The network 898 may also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of the network 898 include, for example, Ethernet, local area networks of wireless LANs, cellular networks including Global System for Mobile Communications (GSM), Third Generation (3G), Fourth Generation (4G), Fifth Generation (5G), Long Term Evolution (LTE), etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular networks and industrial networks including Controller Area Network Bus (CANBus), etc. Some networks 898 typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (850 and 851) (e.g., Universal Serial Bus (USB) ports of the computer system 800); other networks are typically integrated into the core of the computer system 800 by attaching to the system bus as described below (e.g., integrated into a PC computer system through an Ethernet interface, or integrated into a smart phone computer system through a cellular network interface). By using any of these networks 898, the computer system 800 can communicate with other entities. Such communication can be one-way reception only (e.g., broadcast TV), one-way transmission only (e.g., CANBus connected to certain CANBus devices), or two-way, for example, using a local digital network or a wide area digital network to connect to other computer systems. Certain protocols and protocol stacks can be used on each of the networks and network interfaces as described above.
[0137] The above-described human-machine interface device, human-accessible storage device, and network interface can be attached to the core 840 of the computer system 800.
[0138] The core 840 may include one or more central processing units (CPUs) 841, a graphics processing unit (GPU) 842, a graphics adapter 817, a dedicated programmable processing unit 843 in the form of a Field Programmable Gate Area (FPGA), a hardware accelerator 844 for certain tasks, etc. These devices, together with a read-only memory (ROM) 845, a random access memory 846, and internal mass storage devices 847 such as internal non-user-accessible hard disk drives, solid state drives (SSDs), etc., can be connected through a system bus 848. In some computer systems, the system bus 848 can be accessed in the form of one or more physical plugs to enable expansion through additional CPUs, GPUs, etc. Peripheral devices can be attached directly or through a peripheral bus 849 to the system bus 848 of the core. Architectures for peripheral buses include Peripheral Component Interconnect (PCI), USB, etc.
[0139] The CPU 841, GPU 842, FPGA 843, and accelerator 844 can execute certain instructions that, when combined, can constitute the aforementioned computer code. The computer code can be stored in the ROM 845 or the RAM 846. Transitional data can also be stored in the RAM 846, while permanent data can be stored, for example, in the internal mass storage device 847. Fast storage and retrieval of any memory device can be achieved by using a cache memory that can be closely associated with one or more of the CPU 841, GPU 842, mass storage device 847, ROM 845, RAM 846, etc.
[0140] Computer code for performing various computer-implemented operations can be present on a computer-readable medium. The medium and the computer code can be those designed and constructed specifically for the purposes of this application, or can be of the kind well-known and available to those skilled in the art of computer software.
[0141] By way of example and not limitation, a computer system having the architecture 800 and particularly the core 840 can provide functions resulting from software executed by processors (including CPU, GPU, FPGA, accelerator, etc.) embodied on one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage device introduced above and certain non-transitory storage devices of the core 840 (e.g., the core internal mass storage device 847 or the ROM 845). The software implementing various embodiments of this application can be stored in such devices and executed by the core 840. Depending on specific requirements, the computer-readable media can include one or more memory devices or chips. The software can cause the core 840 and specifically the processors therein (including CPU, GPU, FPGA, etc.) to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in the RAM 846 and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functions resulting from logic hard-wired or otherwise embodied in circuitry (e.g., accelerator 844) that can operate in place of or in conjunction with the software to execute the specific processes or specific parts of the specific processes described herein. Where appropriate, references to software can encompass logic, and vice versa. Where appropriate, references to computer-readable media can encompass circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both such circuits. This application encompasses any suitable combination of hardware and software.
[0142] Although this application describes several exemplary embodiments, within the scope of this application, there can be various modifications, permutations, and various alternative equivalents. Therefore, it should be understood that within the spirit and scope of the application, those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, can embody the principles of this application.
[0143] The above - disclosed content also includes the features mentioned below. These features can be combined in various ways and are not limited to the combinations mentioned below.
[0144] (1) A video encoding / decoding method, the method comprising: setting a first value of a text description usage parameter in a bitstream, the first value of the text description usage parameter indicating the type of information included in a text description information string in the bitstream; when the first value indicates that the type of information included in the text description information string includes marker information, the marker information being associated with one or more artificial intelligence (AI) processing processes used, setting artificial intelligence marker information associated with one or more pictures in the bitstream as the text description information string; signaling the text description usage parameter in the bitstream; and signaling the text description information string.
[0145] (2) The method according to feature (1), wherein the text description usage parameter and the text description information string are signaled in a supplementary enhancement information (SEI) message.
[0146] (3) The method according to any one of features (1) to (2), wherein the AI marker information associated with one or more pictures includes an indication that one or more pictures are generated using a machine - learning - based processing process.
[0147] (4) The method according to any one of features (1) to (3), wherein the AI marker information associated with one or more pictures further includes: a tool identifier indicating a tool used to generate one or more pictures by using a machine - learning - based processing process; a timestamp indicating a time at which the tool was used to generate one or more pictures by using a machine - learning - based processing process; or one or more instructions used by the tool to generate one or more pictures.
[0148] (5) The method according to any one of features (1) to (4), wherein when the text description information string is not an empty string, the text description information string includes AI marker information associated with pictures within the scope of the SEI message.
[0149] (6) The method according to any one of features (1) to (5), wherein the value of the text description usage parameter is between 0 and 255.
[0150] (7) A method according to any one of features (1) to (6), wherein the text description information string is in string format.
[0151] (8) A method according to any one of features (1) to (7), further comprising: receiving an encoded video bitstream including pictures; obtaining a first value of a text description usage parameter in the encoded video bitstream, the first value of the text description usage parameter indicating the type of information included in the text description information string in the encoded video bitstream; in response to the first value indicating that the type of information included in the text description information string includes markup information associated with one or more artificial intelligence (AI) processing processes used, obtaining artificial intelligence markup information associated with one or more pictures in the encoded video bitstream as the text description information string in the encoded video bitstream; and reconstructing pictures from the encoded video bitstream based on the first value and the AI markup information.
[0152] (9) A method according to any one of features (1) to (8), wherein the text description usage parameter and the text description information string are signaled in a supplementary enhancement information (SEI) message.
[0153] (10) A method according to any one of features (1) to (9), wherein the AI markup information associated with one or more pictures includes an indication that one or more pictures are generated using a machine learning-based processing process.
[0154] (11) A method according to any one of features (1) to (10), wherein the AI markup information associated with one or more pictures further includes: a tool identifier indicating a tool used to generate one or more pictures by using a machine learning-based processing process; a timestamp indicating a time at which the tool was used to generate one or more pictures by using a machine learning-based processing process; or one or more instructions
[0155] used by the tool to generate one or more pictures.
[0156] (12) A method according to any one of features (1) to (11), wherein when the text description information string is not an empty string, the text description information string includes AI markup information associated with at least one picture within the scope of the SEI message.
[0157] (13)A method according to any one of features (1) to (12), further comprising: performing a conversion between a visual media file and a bitstream of the visual media file according to formatting rules, wherein the bitstream includes: a text description usage parameter indicating the type of information included in the text description information string; and artificial intelligence (AI) marker information associated with one or more pictures in the visual media file, and when the text description usage parameter indicates that the type of information included in the text description information string includes marker information that is associated with one or more artificial intelligence (AI) processing processes used, the AI marker is the text description information string; and wherein the formatting rules specify that the text description usage parameter and the text description information string are syntax parameters in a supplementary enhancement information (SEI) message.
[0158] (14)A method according to any one of features (1) to (13), wherein the text description information string is in string format.
[0159] (15)An apparatus for video decoding, comprising a processing circuit configured to perform the method according to any one of features (1) to (14).
[0160] (16)An apparatus for video encoding, comprising a processing circuit configured to perform the method according to any one of features (1) to (14).
[0161] (17)A non-volatile computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of features (1) to (14).
Claims
1. A video processing method, characterized in that, The method is executed by at least one processor, and the method includes: Setting a first value of a text description usage parameter in a bitstream, the first value of the text description usage parameter indicating the type of information included in a text description information string in the bitstream; When the first value indicates that the type of information included in the text description information string includes markup information that is associated with one or more artificial intelligence (AI) processing procedures used, setting artificial intelligence markup information associated with one or more pictures in the bitstream as the text description information string; Signaling the text description usage parameter in the bitstream; and Signaling the text description information string.
2. The method according to claim 1, characterized in that, The text description usage parameter and the text description information string are signaled in a supplementary enhancement information (SEI) message.
3. The method according to claim 2, wherein The AI markup information associated with the one or more pictures includes an indication that the one or more pictures are generated using a machine learning-based processing procedure.
4. The method according to claim 3, characterized in that, The AI markup information associated with the one or more pictures further includes: A tool identifier indicating a tool that is used to generate the one or more pictures by using the machine learning-based processing procedure; A timestamp indicating a time at which the tool was used to generate the one or more pictures by using the machine learning-based processing procedure; or One or more instructions used by the tool to generate the one or more pictures.
5. The method according to claim 2, wherein When the text description information string is not an empty string, the text description information string includes AI markup information associated with pictures within the scope of the SEI message.
6. The method according to claim 1, characterized in that, The value of the text description usage parameter is between 0 and 255.
7. The method according to claim 1, characterized in that, The text description information string is in string format.
8. A video processing device, characterized in that, The apparatus includes: At least one memory storing program code; and At least one processor configured to access the at least one memory and operate according to the instructions of the program code, the program code including: A receiving code configured to cause the at least one processor to receive an encoded video bitstream including pictures; A first obtaining code configured to cause the at least one processor to obtain a first value of a text description usage parameter in the encoded video bitstream, the first value of the text description usage parameter indicating the type of information included in a text description information string in the encoded video bitstream; A second obtaining code configured to cause the at least one processor to obtain, in response to the first value indicating that the type of information included in the text description information string includes markup information that is associated with one or more artificial intelligence (AI) processing procedures used, artificial intelligence markup information associated with one or more pictures in the encoded video bitstream as the text description information string in the encoded video bitstream; and A reconstruction code configured to cause the at least one processor to reconstruct the pictures from the encoded video bitstream based on the first value and the AI markup information.
9. The device according to claim 8, characterized in that, The text description usage parameter and the text description information string are signaled in a Supplemental Enhancement Information (SEI) message.
10. The device according to claim 9, characterized in that, The AI marking information associated with the one or more pictures includes an indication that the one or more pictures are generated using a machine learning-based processing procedure.
11. The device according to claim 10, characterized in that, The AI marking information associated with the one or more pictures further includes: A tool identifier indicating a tool used to generate the one or more pictures by using the machine learning-based processing procedure; A timestamp indicating a time at which the tool was used to generate the one or more pictures by using the machine learning-based processing procedure; or One or more instructions used by the tool to generate the one or more pictures.
12. The device according to claim 9, characterized in that, When the text description information string is not an empty string, the text description information string includes AI marking information associated with at least one picture within the scope of the SEI message.
13. The device according to claim 8, characterized in that, The value of the text description usage parameter is between 0 and 255.
14. The device according to claim 8, characterized in that, The text description information string is in string format.
15. A non - volatile computer - readable medium for video processing, characterized in that, The non-volatile computer-readable medium stores one or more instructions that are configured to cause at least one processor to: Perform a conversion between a visual media file and a bitstream of the visual media file according to format rules, wherein the bitstream includes: A text description usage parameter indicating the type of information included in the text description information string; and Artificial Intelligence (AI) marking information associated with one or more pictures in the visual media file, and when the text description usage parameter indicates that the type of information included in the text description information string includes marking information associated with one or more artificial intelligence (AI) processing procedures used, the AI marking is the text description information string; and wherein the format rules specify that the text description usage parameter and the text description information string are syntax parameters in a Supplemental Enhancement Information (SEI) message.
16. The non-volatile computer-readable medium according to claim 15, wherein The AI marking information associated with the one or more pictures includes an indication that the one or more pictures are generated using a machine learning-based processing procedure.
17. The non-volatile computer-readable medium according to claim 16, wherein The AI marking information associated with the one or more pictures further includes: A tool identifier indicating a tool used to generate the one or more pictures by using the machine learning-based processing procedure; A timestamp indicating a time at which the tool was used to generate the one or more pictures by using the machine learning-based processing procedure; or One or more AI generation instructions used by the tool to generate the one or more pictures.
18. The non-volatile computer-readable medium according to claim 15, wherein When the text description information string is not an empty string, the text description information string includes AI marking information associated with pictures within the scope of the SEI message.
19. The non-volatile computer-readable medium according to claim 15, wherein The value of the text description usage parameter is between 0 and 255. The non-volatile computer-readable medium according to claim 15, wherein The text description information string is in string format.