Image encoding / decoding method and device, and recording medium with bitstream stored thereon
The method and device allow for efficient management and signaling of AI usage restrictions in video bitstreams, addressing the challenge of enforcing AI usage policies by differentiating between AI inference, training, and generative AI usage, thereby facilitating compliance with specified constraints.
Patent Information
- Application Number
- PCT/KR2025/009450
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-07-02
- Publication Date
- 2026-01-08
AI Technical Summary
Existing video encoding and decoding technologies lack the ability to efficiently manage and signal Artificial Intelligence (AI) usage restrictions, making it difficult to enforce and check AI usage constraints and recommendations within video bitstreams.
A method and device for configuring and signaling AI usage restrictions within video bitstreams, allowing for the definition of AI usage information in Network Abstraction Layer (NAL) units, enabling the differentiation between AI inference, training, and generative AI usage, and providing a mechanism to encode and decode these restrictions.
Enables easy checking and enforcement of AI usage constraints and recommendations within video bitstreams, ensuring compliance with specified AI usage policies.
Smart Images

Figure KR2025009450_08012026_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and recording medium storing bitstream
[0001] The present invention relates to a video encoding / decoding method and device, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed.
[0003] There are various technologies for image compression, such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and entropy encoding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, and these technologies can be used to effectively compress and transmit or store image data.
[0004] The present disclosure provides a method and device for configuring content usage information.
[0005] The present disclosure provides a method and device for signaling content usage information.
[0006] The present disclosure provides a method and device for configuring AI usage restrictions.
[0007] The present disclosure provides a method and device for signaling AI usage restrictions.
[0008] The present disclosure provides a method and apparatus for composing a text description.
[0009] The present disclosure provides a method and apparatus for signaling a text description.
[0010] A video decoding method and device according to the present disclosure can receive a bitstream including an encoded video picture and restore the encoded video picture included in the bitstream. Here, the bitstream can include AI (Artificial Intelligence) usage restriction information indicating restrictions on the use of AI. The AI usage restriction information can be obtained from a network abstraction layer (NAL) unit of the bitstream.
[0011] In the image decoding method and device according to the present disclosure, the AI usage restriction information of the first value may indicate that it cannot be used for AI inference, and the AI usage restriction information of the second value may indicate that it cannot be used for AI training.
[0012] In the image decoding method and device according to the present disclosure, the AI usage restriction information of the first value may indicate that it cannot be used for AI training, the AI usage restriction information of the second value may indicate that it cannot be used for generative AI, and the AI usage restriction information of the third value may indicate that it cannot be used in any AI application.
[0013] In the image decoding method and device according to the present disclosure, a smaller value may be assigned to the AI usage restriction information indicating that it is unusable for AI training than to the AI usage restriction information indicating that it is unusable for the generative AI.
[0014] In the video decoding method and device according to the present disclosure, the AI usage restriction information may be related to four restrictions on the use of the AI.
[0015] In the image decoding method and device according to the present disclosure, the AI usage restriction information of the first value may indicate that it cannot be used for AI training, the AI usage restriction information of the second value may indicate that it cannot be used for AI inference, the AI usage restriction information of the third value may indicate that it cannot be used for generative AI, and the AI usage restriction information of the fourth value may indicate that it cannot be used in any AI application.
[0016] In the image decoding method and device according to the present disclosure, a smaller value may be assigned to the AI usage restriction information indicating that the AI cannot be used for AI training than to the AI usage restriction information indicating that the AI cannot be used for AI inference.
[0017] In the video decoding method and device according to the present disclosure, the bitstream may further include restriction number information indicating the number of restriction entries signaled.
[0018] In the video decoding method and device according to the present disclosure, the AI usage constraint information can be signaled from the bitstream based on the constraint number information.
[0019] A video encoding method and device according to the present disclosure can receive a video picture to be encoded, encode the received video picture to generate video information about the encoded video picture, generate AI (Artificial Intelligence) usage restriction information indicating restrictions on AI usage, and generate a bitstream including the video information and the AI usage restriction information. Here, the AI usage restriction information can be encoded in a network abstraction layer (NAL) unit of the bitstream.
[0020] A computer-readable digital storage medium is provided, which stores encoded video / image information that causes a decoding device according to the present disclosure to perform a video decoding method.
[0021] A computer-readable digital storage medium storing video / image information generated by a video encoding method according to the present disclosure is provided.
[0022] A method and device for transmitting video / image information generated by a video encoding method according to the present disclosure are provided.
[0023] According to the present disclosure, by defining content usage information, it is possible to easily check restrictions and recommendations for content usage applied to a bitstream.
[0024] According to the present disclosure, by defining AI usage constraints, it is possible to easily check the constraints on AI usage applied to a bitstream and the context for the constraints.
[0025] According to the present disclosure, by defining a text description, it is possible to easily check the purpose, constraints, recommendations, etc. of the text description applied to a bitstream.
[0026] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0027] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0028] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0029] FIG. 4 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.
[0030] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.
[0031] FIG. 6 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.
[0032] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.
[0033] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0034] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Similar reference numerals have been used to designate similar components throughout the description of each drawing.
[0035] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0036] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0037] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0038] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268).
[0039] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0040] In this specification, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.
[0041] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.
[0042] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0043] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0044] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0045] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0046] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0047] Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction."
[0048] Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously.
[0049] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0050] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device).
[0051] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0052] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0053] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0054] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0055] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0056] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0057] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0058] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0059] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0060] For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0061] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0062] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0063] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0064] The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0065] The intra prediction unit (222) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is only an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0066] The inter prediction unit (221) can derive a prediction block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0067] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal.
[0068] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size.
[0069] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0070] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from quantized transform coefficients.
[0071] Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).
[0072] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0073] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0074] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0075] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks within the current picture and transfer them to the intra prediction unit (222).
[0076] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0077] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0078] The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0079] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit of decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.
[0080] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310).
[0081] Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).
[0082] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0083] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0084] The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode.
[0085] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be included and signaled in the video / image information.
[0086] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0087] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.
[0088] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block.
[0089] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0090] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0091] The (corrected) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (332) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (331).
[0092] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.
[0093] FIG. 4 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.
[0094] A bitstream including an encoded video picture can be received (S400).
[0095] The encoded video picture of the bitstream can be restored (S410).
[0096] Video information about an encoded video picture can be extracted from a bitstream. The encoded video picture can be restored based on the extracted video information.
[0097] Example 1
[0098] A bitstream may contain content usage information (CUI). Content usage information relates to how a picture is used, and may include restrictions and recommendations regarding the use of the picture, including the user and intended use. Below, we will examine content usage information in more detail.
[0099] Content usage information may include a CUI cancellation flag (cui_cancel_flag). If cui_cancel_flag is 1, this may indicate that the persistence of previous content usage information is canceled, and if cui_cancel_flag is 0, this may indicate that the content usage information is persisted.
[0100] Content usage information may include a CUI persistence flag (cui_persistence_flag). cui_persistence_flag may indicate the persistence of the content usage information. If cui_persistence_flag is 0, this may indicate that the content usage information can only be applied to the current picture. If cui_persistence_flag is 1, this may indicate that the content usage information can be applied to the current picture and all subsequent pictures.
[0101] Content usage information may include entity information (cui_applicable_entities). cui_applicable_entities may indicate the user and purpose of the picture. If (cui_applicable_entities & bitMask) is not 0, it may indicate that the entity corresponding to the bitMask in Table 1 is included as the user and purpose. The value of cui_applicable_entities may be constrained to be in the range of 0 to 7. In this case, if the value of cui_applicable_entities is in the range of 8 to 65,535, it may be reserved for future use, and decoders with certain specifications must ignore the value of cui_applicable_entities.
[0102] bitMaskInterpretation0x01AI0x02General users0x04Content creator
[0103] For example, if cuiForAIFlag is 1, this indicates that the picture's user and purpose include AI (Artificial Intelligence), and if cuiForAIFlag is 0, this indicates that the picture's user and purpose do not include AI. cuiForAIFlag can be derived as follows.
[0104] cuiForAIFlag = ( ( cui_applicable_entities & 0x01 ) > 0 ) ? 1:0
[0105] If cuiForGeneralUserFlag is 1, this indicates that the picture's user and purpose include general users. If cuiForGeneralUserFlag is 0, this may indicate that the picture's user and purpose do not include general users. cuiForGeneralUserFlag can be derived as follows.
[0106] cuiForGeneralUserFlag = ( ( cui_applicable_entities & 0x02 ) > 0 ) ? 1:0
[0107] If cuiForProfessionalUserFlag is 1, this indicates that the picture is used by a professional user or content creator. If cuiForProfessionalUserFlag is 0, this indicates that the picture is not used by a professional user or content creator. cuiForProfessionalUserFlag can be derived as follows.
[0108] cuiForProfessionalUserFlag= ( ( cui_applicable_entities & 0x04 ) > 0 ) ? 1:0
[0109] In Table 1, AI indicates that the content is used by AI or is intended for AI use. General user refers to a typical user. This refers to a user whose primary purpose is consumption, such as viewing. Content creator refers to a user who creates content, and refers to a user whose primary purpose is to produce and distribute content.
[0110] If the content's intended users and purposes are determined based on cui_applicable_entities, usage methods (e.g., restrictions, recommendations, etc.) can be defined for those intended users and purposes. For example, usage methods for specific intended users and purposes can be defined as in Table 2 or Table 3.
[0111] content_usage_information( payloadSize ) { Descriptorcui_cancel_flagu(1)if( !cui_cancel_flag ) { cui_persistence_flagu(1)cui_applicable_entitiesue(v)if(cuiForAIFlag){ allowed_for_training_flagu(1)allowed_for_inferencing_flagu(1)allowed_for_editing_flagu(1)}}}
[0112] In Table 2, if cuiForAIFlag is 1 (or if AI is included in the user entity and purpose), the usage method for AI can be defined. The usage method for AI can be defined by at least one of the allowed_for_training_flag, allowed_for_inferencing_flag, or allowed_for_editing_flag. However, this is only an example, and other restrictions, recommendations, usage methods, etc. may be defined.
[0113] allowed_for_training_flag, allowed_for_inferencing_flag, and allowed_for_editing_flag can indicate whether they are allowed to be used for a specific purpose. Alternatively, allowed_for_training_flag, allowed_for_inferencing_flag, and allowed_for_editing_flag can indicate whether they are allowed to be used for a specific purpose. Depending on the clarity of permission, they can be defined as "if allowed" and "if may be allowed." If allowed_for_training_flag, allowed_for_inferencing_flag, and allowed_for_editing_flag are 0, this can be defined as a case where it is not known whether they are allowed to be used or whether they are allowed to be used.
[0114] Specifically, if allowed_for_training_flag is 1, this may indicate that the picture associated with the content usage information may be used for AI training purposes. Otherwise, this may indicate that the picture associated with the content usage information may not be used for AI training purposes. Alternatively, if allowed_for_training_flag is 1, this may indicate that the picture associated with the content usage information is permitted (or may be permitted) to be used for AI training purposes. Otherwise, this may indicate that the picture associated with the content usage information is not permitted (or may not be permitted) to be used for AI training purposes.
[0115] If allowed_for_inferencing_flag is 1, this may indicate that the picture associated with the content usage information may be used for AI inferencing purposes. Otherwise, this may indicate that the picture associated with the content usage information may not be used for AI inferencing purposes. Alternatively, if allowed_for_inferencing_flag is 1, this may indicate that the picture associated with the content usage information is permitted (or may be permitted) to be used for AI inferencing purposes. Otherwise, this may indicate that the picture associated with the content usage information is not permitted (or may not be permitted) to be used for AI inferencing purposes. Here, inferencing means that AI obtains a result using the picture, and may include at least one of discrimination and generation.
[0116] If allowed_for_editing_flag is 1, this may indicate that the picture associated with the content usage information may be edited. Otherwise, this may indicate that the picture associated with the content usage information may not be edited. Alternatively, if allowed_for_editing_flag is 1, this may indicate that editing of the picture associated with the content usage information is permitted (or may be permitted). Otherwise, this may indicate that editing of the picture associated with the content usage information is not permitted (or may not be permitted). Here, editing may include cropping, resizing, quality enhancement, or color adjustment.
[0117] content_usage_information( payloadSize ) { Descriptorcui_cancel_flagu(1)if( !cui_cancel_flag ) { cui_persistence_flagu(1)cui_applicable_entitiesue(v)if(cuiForAIFlag){ not_allowed_for_training_flagu(1)not_allowed_for_inferencing_flagu(1)not_allowed_for_editing_flagu(1)}}}
[0118] Table 3 defines not_allowed_for_training_flag, not_allowed_for_inferencing_flag, and not_allowed_for_editing_flag, unlike Table 2. While allowed_for_training_flag, allowed_for_inferencing_flag, and allowed_for_editing_flag in Table 2 indicate "usable" or "allowed" when their value is 1, not_allowed_for_training_flag, not_allowed_for_inferencing_flag, and not_allowed_for_editing_flag in Table 3 indicate "unusable" or "not allowed" when their value is 1.
[0119] Specifically, if not_allowed_for_training_flag is 1, this may indicate that the picture associated with the content usage information cannot be used for AI training purposes. Otherwise, this may indicate that the picture associated with the content usage information may be used for AI training purposes. Alternatively, if not_allowed_for_training_flag is 1, this may indicate that the picture associated with the content usage information is not (or may not be) permitted to be used for AI training purposes. Otherwise, this may indicate that the picture associated with the content usage information is (or may be) permitted to be used for AI training purposes.
[0120] If not_allowed_for_inferencing_flag is 1, this may indicate that the picture associated with the content usage information cannot be used for AI inferencing purposes. Otherwise, this may indicate that the picture associated with the content usage information may be used for AI inferencing purposes. Alternatively, if not_allowed_for_inferencing_flag is 1, this may indicate that the picture associated with the content usage information is not (or may not be) permitted to be used for AI inferencing purposes. Otherwise, this may indicate that the picture associated with the content usage information is (or may be) permitted to be used for AI inferencing purposes. Here, inferencing means that the AI obtains a result using the picture, and may include at least one of discrimination and generation.
[0121] If not_allowed_for_editing_flag is 1, this may indicate that the picture associated with the content usage information cannot be edited. Otherwise, this may indicate that the picture associated with the content usage information may be edited. Alternatively, if not_allowed_for_editing_flag is 1, this may indicate that editing of the picture associated with the content usage information is not permitted (or may not be permitted). Otherwise, this may indicate that editing of the picture associated with the content usage information is permitted (or may be permitted). Here, editing may include cropping, resizing, quality enhancement, or color adjustment.
[0122] Content usage information may include category information (cui_usage_category). cui_usage_category may indicate the usage category of the picture. If (cui_usage_category & bitMask) is not 0, it may indicate that the usage category corresponding to the bitMask in Table 4 is included. The value of cui_usage_category may be constrained to be in the range of 0 to 7. In this case, if the value of cui_usage_category is in the range of 8 to 65,535, it may be reserved for future use, and decoders of certain specifications must ignore cui_usage_category. The usage category may indicate the usage type of the content or picture (e.g., how, by whom, in what environment the picture is used, etc.), and this may be the target or criterion to which the usage method defined by the content usage information is applied.
[0123] bitMaskInterpretation0x01AI0x02Region0x04Editing
[0124] For example, if cuiForAIFlag is 1, this may indicate that the picture's usage category includes AI, and if cuiForAIFlag is 0, this may indicate that the picture's usage category does not include AI. CuiForAIFlag may be derived as follows.
[0125] cuiForAIFlag = ( (cui_usage_category & 0x01 ) > 0 ) ? 1:0
[0126] If cuiForRegionFlag is 1, it indicates that the picture's usage category includes the region, and if cuiForRegionFlag is 0, it may indicate that the picture's usage category does not include the region. cuiForRegionFlag can be derived as follows.
[0127] cuiForRegionFlag = ( ( cui_usage_category & 0x02 ) > 0 ) ? 1:0
[0128] If cuiForEditingFlag is 1, it indicates that the picture's usage category includes editing, and if cuiForEditingFlag is 0, it indicates that the picture's usage category does not include editing. cuiForEditingFlag can be derived as follows.
[0129] cuiForEditingFlag= ( ( cui_usage_category & 0x04 ) > 0 ) ? 1:0
[0130] If a usage category is determined based on cui_usage_category, constraints and / or recommendations for that usage category may be defined. Constraints and / or recommendations for a usage category may be defined using any of the methods listed in Tables 5 to 7.
[0131] content_usage_information( payloadSize ) {Descriptorcui_cancel_flagu(1)if( !cui_cancel_flag ) { cui_persistence_flagu(1)cui_usage_categoryue(v)if(cuiForAIFlag){ allowed_for_training_flagu(1)allowed_for_inferencing_flagu(1)if(cuiForRegionFlag){allowed_in_region_a_flagu(1)allowed_in_region_b_flagu(1)allowed_in_region_c_flagu(1)}if(cuiForEditingFlag){allowed_for_cropping_flagu(1)allowed_for_colorization_flagu(1)allowed_for_msaking_flagu(1)}}}}
[0132] content_usage_information( payloadSize ) {Descriptorcui_cancel_flagu(1)if( !cui_cancel_flag ) { cui_persistence_flagu(1)cui_usage_categoryue(v)if(cuiForAIFlag){ not_allowed_for_training_flagu(1)not_allowed_for_inferencing_flagu(1)if(cuiForRegionFlag){not_allowed_in_region_a_flagu(1)not_allowed_in_region_b_flagu(1)not_allowed_in_region_c_flagu(1)}if(cuiForEditingFlag){not_allowed_for_cropping_flagu(1)not_allowed_for_colorization_flagu(1)not_allowed_for_msaking_flagu(1)}}}}
[0133] content_usage_information( payloadSize ) { Descriptorcui_cancel_flagu(1)if( !cui_cancel_flag ) { cui_persistence_flagu(1)cui_usage_categoryue(v)if(cuiForAIFlag){ allowed_for_training_flagu(1)allowed_for_inferencing_flagu(1)if(cuiForRegionFlag){cui_usage_region_codest(v)}if(cuiForEditingFlag){allowed_for_cropping_flagu(1)allowed_for_colorization_flagu(1)allowed_for_msaking_flagu(1)}}}}
[0134] If cuiForAIFlag is 1 (or AI is included in the usage category), the content usage information may include restrictions and recommendations regarding the types of use of the content or picture due to AI. For example, at least one of the allowed_for_training_flag, allowed_for_inferencing_flag, not_allowed_for_training_flag, or not_allowed_for_inferencing_flag may be defined.
[0135] If allowed_for_training_flag is 1, this may indicate that pictures associated with content usage information may be used for AI training purposes, and otherwise, this may indicate that pictures associated with content usage information may not be used for AI training purposes. Alternatively, if allowed_for_training_flag is 1, this may indicate that pictures associated with content usage information are permitted (or may be permitted) to be used for AI training purposes, and otherwise, this may indicate that pictures associated with content usage information are not permitted (or may not be permitted) to be used for AI training purposes.
[0136] If allowed_for_inferencing_flag is 1, this may indicate that the picture associated with the content usage information may be used for AI inferencing purposes, and otherwise, this may indicate that the picture associated with the content usage information may not be used for AI inferencing purposes. Alternatively, if allowed_for_inferencing_flag is 1, this may indicate that the picture associated with the content usage information is permitted (or may be permitted) to be used for AI inferencing purposes, and otherwise, this may indicate that the picture associated with the content usage information is not permitted (or may not be permitted) to be used for AI inferencing purposes. Here, inferencing means that the AI obtains a result by using the picture, and may include at least one of discrimination and generation.
[0137] If not_allowed_for_training_flag is 1, this may indicate that the picture associated with the content usage information cannot be used for AI training purposes, otherwise, this may indicate that the picture associated with the content usage information may be used for AI training purposes. Alternatively, if not_allowed_for_training_flag is 1, this may indicate that the picture associated with the content usage information is not (or may not be) permitted to be used for AI training purposes, otherwise, this may indicate that the picture associated with the content usage information is (or may be permitted to be) permitted to be used for AI training purposes.
[0138] If not_allowed_for_inferencing_flag is 1, this indicates that the picture associated with the content usage information cannot be used for AI inferencing purposes, otherwise, this may indicate that the picture associated with the content usage information may be used for AI inferencing purposes. Alternatively, if not_allowed_for_inferencing_flag is 1, this may indicate that the picture associated with the content usage information is not (or may not be) permitted to be used for AI inferencing purposes, otherwise, this may indicate that the picture associated with the content usage information is (or may be permitted to be) permitted to be used for AI inferencing purposes. Here, inferencing means that the AI obtains a result by using the picture, and may include at least one of discrimination or generation.
[0139] If CuiForRegionFlag is 1, at least one of allowed_in_region_a_flag, allowed_in_region_b_flag, allowed_in_region_c_flag, not_allowed_in_region_a_flag, not_allowed_in_region_b_flag, or not_allowed_in_region_c_flag may be defined.
[0140] Including a region in a usage category may mean that constraints and recommendations for content and pictures based on the region are included in the content usage information.
[0141] If allowed_in_region_a_flag, allowed_in_region_b_flag, and allowed_in_region_c_flag are 1, this may indicate that the use (e.g., viewing) of the content or picture is permitted (or may be permitted, or is possible) in Region A, Region B, and Region C, respectively. Otherwise, this may indicate that the use of the content or picture is not permitted (or may not be permitted, or is not possible) in Region A, Region B, and Region C, respectively. Alternatively, if allowed_in_region_a_flag, allowed_in_region_b_flag, and allowed_in_region_c_flag are 0, this may indicate that it is not known whether the use is permitted or is not possible, respectively.
[0142] If not_allowed_in_region_a_flag, not allowed_in_region_b_flag, not allowed_in_region_c_flag are 1, this may indicate that the use (e.g., viewing) of the content or picture is not (or may not be) permitted in Region A, Region B, and Region C, respectively. Otherwise, this may indicate that the use of the content or picture is not (or may not be) permitted in Region A, Region B, and Region C, respectively. Alternatively, if not_allowed_in_region_a_flag, not allowed_in_region_b_flag, not allowed_in_region_c_flag are 0, this may indicate that it is not known whether the use is permitted or permitted, respectively.
[0143] Examples for each Region are as follows:
[0144] Region A: North, Central, South America, Japan, Korea, and Southeast Asia.
[0145] Region B: Europe (EU), Africa, Middle East, Australia, and New Zealand.
[0146] Region C: Russia, India, China, and the rest of the world.
[0147] If cuiForRegionFlag is 1, a string representing the constraint, cui_usage_region_code, may be defined instead of a flag indicating a region-specific constraint. cui_usage_region_code may be defined in Table 6.
[0148] cui_usage_region_code specifies the region code in the format of Blu-ray region codes:
[0149] A: The Americas and their dependencies, Taiwan, Japan, Hong Kong, Macau, Korea, and Southeast Asia.
[0150] B: Africa, West Asia, most of Europe, Australia, New Zealand, and their dependencies.
[0151] C: Central Asia, China, Mongolia, South Asia, Belarus, Ukraine, Russia, Kazakhstan, Moldova, and their dependencies.
[0152] When using region codes instead of bitmasks, the information can be simplified and standardized by referring to the Blu-ray region code standard. However, this is only an example, and DVD region codes, or codes or phrases indicating separate regions, may be used instead of Blu-ray region codes.
[0153] If cuiForEditingFlag is 1, at least one of allowed_for_cropping_flag, allowed_for_colorization_flag, allowed_for_msaking_flag, not_allowed_for_cropping_flag, not_allowed_for_colorization_flag, or not_allowed_for_msaking_flag may be defined.
[0154] Including Editing in the Usage Category may mean that the Content Usage Information includes constraints and recommendations for editing and reconfiguring content and pictures.
[0155] If allowed_for_cropping_flag, allowed_for_colorization_flag, and allowed_for_msaking_flag are 1, this may indicate that editing such as cropping, colorization, and masking are permitted (or may be permitted, or may be possible) for the corresponding picture, respectively. Otherwise, this may indicate that editing such as cropping, colorization, and masking are not permitted (or may not be permitted, or may not be possible) for the corresponding picture, respectively. If allowed_for_cropping_flag, allowed_for_colorization_flag, and allowed_for_msaking_flag are 0, this may indicate that it is not known whether they are permitted or not, respectively.
[0156] If not_allowed_for_cropping_flag, not_allowed_for_colorization_flag, and not_allowed_for_msaking_flag are 1, this may indicate that editing such as cropping, colorization, and masking are not allowed (or may not be allowed, or are not allowed) for the corresponding picture, respectively. Otherwise, this may indicate that editing such as cropping, colorization, and masking are not allowed (or may not be allowed, or are not allowed) for the corresponding picture, respectively. If not_allowed_for_cropping_flag, not_allowed_for_colorization_flag, and not_allowed_for_msaking_flag are 0, this may indicate that it is not known whether they are allowed or not, respectively.
[0157] content_usage_information(payloadSize) {Descriptorcui_cancel_flagu(1)if(!cui_cancel_flag) {cui_persistence_flagu(1)cui_usage_categoryue(v)if(cuiForAIFlag)cui_usage_for_aiue(v)if(cuiForRegionFlag)cui_usage_for_regionue(v)if(cuiForEditingFlag)cui_usage_for_editingue(v)}}
[0158] The cui_usage_category in Table 8 can indicate the usage category of the picture. If (cui_usage_category & bitMask) is not 0, it can indicate that the usage category corresponding to the bitMask in Table 9 is included. If cui_usage_category is 0, the usage category may be defined by the application. The value of cui_usage_category may be constrained to be in the range of 0 to 7. In this case, if the value of cui_usage_category is in the range of 8 to 65,535, it may be reserved for future use, and decoders with certain specifications should ignore cui_usage_category.
[0159] bitMaskInterpretation0x01AI process usage. This SEI includes information on how associated pictures may be used with AI processes0x02Region-based usage. This SEI includes information on how associated pictures may be used based on region.0x04Editing usage. This SEI includes information on how associated pictures may be used for editing or modification purposes
[0160] cuiForAIFlag can specify whether cui_usage_category represents a usage category that includes AI process usage. CuiForAIFlag can be derived as follows:
[0161] cuiForAIFlag = ( (cui_usage_category & 0x01 ) > 0 ) ? 1:0
[0162] cuiForRegionFlag can specify whether cui_usage_category represents a usage category that includes region-based usage. cuiForRegionFlag can be derived as follows:
[0163] cuiForRegionFlag = ( ( cui_usage_category & 0x02 ) > 0 ) ? 1:0
[0164] cuiForEditingFlag can specify whether cui_usage_category represents a usage category that includes editing usage. cuiForEditingFlag can be derived as follows:
[0165] cuiForEditingFlag= ( ( cui_usage_category & 0x04 ) > 0 ) ? 1:0
[0166] AI usage constraint information (cui_usage_for_ai) can indicate constraints on AI usage. For example, cui_usage_for_ai can indicate AI-related usage information. If (cui_usage_for_ai & bitMask) is not 0, this can indicate that the usage information corresponding to the bitMask in Table 10 applies. If cui_usage_for_ai is 0, the usage information can be defined by the application. The value of cui_usage_for_ai can be constrained to be in the range of 0 to 7. In this case, if the value of cui_usage_for_ai is in the range of 8 to 65,535, it can be reserved for future use, and decoders with certain specifications must ignore cui_usage_for_ai.
[0167] bitMaskInterpretation0x01Usage information related to AI training; Indicates that pictures related to content usage information can be used for AI training0x02Usage information related to AI inference; Indicates that pictures related to content usage information can be used for AI inference0x04Usage information related to editing: Indicates that pictures related to content usage information may be edited for AI utilization purposes
[0168] cuiAllowedForAITrainingFlag can specify whether the picture associated with the SEI message is usable for AI training purposes. cuiAllowedForAIInferencingFlag can specify whether the picture associated with the SEI message is usable for AI inference purposes. cuiAllowedForAIEditingFlag can specify whether the picture associated with the SEI message is usable for AI editing purposes. CuiAllowedForAITrainingFlag, cuiAllowedForAIInferencingFlag, and cuiAllowedForAIEditingFlag can be derived as follows.
[0169] cuiAllowedForAITrainingFlag = ( (cui_usage_for_ai & 0x01) > 0 ) ? 1:0
[0170] cuiAllowedForAIInferenceFlag = ( ( cui_usage_for_ai & 0x02) > 0 ) ? 1:0
[0171] cuiAllowedForAIEditingFlag = ( ( cui_usage_for_ai & 0x04) > 0 ) ? 1:0
[0172] That is, if cui_usage_for_ai is the first value (e.g., (cui_usage_for_ai & 0x01) is greater than 0), this may indicate that the picture associated with the SEI message is available for AI training purposes. In other words, if cui_usage_for_ai is the first value, this may indicate that the picture associated with the SEI message is not available for AI inference or AI editing purposes.
[0173] If cui_usage_for_ai is a secondary value (e.g., (cui_usage_for_ai & 0x02) is greater than 0), this may indicate that the picture associated with the SEI message is available for AI inference purposes. In other words, if cui_usage_for_ai is a secondary value, this may indicate that the picture associated with the SEI message is not available for AI training or AI editing purposes.
[0174] If cui_usage_for_ai is a third value (e.g., (cui_usage_for_ai & 0x04) is greater than 0), this may indicate that the picture associated with the SEI message is available for AI editing purposes. In other words, if cui_usage_for_ai is a second value, this may indicate that the picture associated with the SEI message is not available for AI training or AI inference purposes.
[0175] The region usage constraint information (cui_usage_for_region) can indicate region-related usage information. If (cui_usage_for_region & bitMask) is not 0, it can indicate that the usage information corresponding to the bitMask in Table 11 applies. If cui_usage_for_region is 0, the usage information can be defined by the application. The value of cui_usage_for_region can be constrained to be in the range of 0 to 7. In this case, if the value of cui_usage_for_region is in the range of 8 to 65,535, it can be reserved for future use, and decoders with certain specifications should ignore cui_usage_for_region.
[0176] bitMaskInterpretation0x01Usage information related to Region A(The Americas and their dependencies, Taiwan, Japan, Hong Kong, Macau, Korea, and Southeast Asia): Indicates that pictures related to content usage information can be used in Region A.0x02Usage information related to Region B(Africa, West Asia, most of Europe, Australia, New Zealand, and their dependencies): Indicates that pictures related to content usage information can be used in Region B.0x04Usage information related to Region C( Central Asia, China, Mongolia, South Asia, Belarus, Ukraine, Russia, Kazakhstan, Moldova, and their dependencies): Indicates that pictures related to content usage information can be used in Region C.
[0177] cuiAllowedInRegionAFlag can specify whether cui_usage_for_region represents usage information including usage in region A. cuiAllowedInRegionBFlag can specify whether cui_usage_for_region represents usage information including usage in region B. cuiAllowedInRegionCFlag can specify whether cui_usage_for_region represents usage information including usage in region C. cuiAllowedInRegionAFlag, cuiAllowedInRegionBFlag, and cuiAllowedInRegionCFlag can be derived as follows, respectively.
[0178] cuiAllowedInRegionAFlag = ( (cui_usage_for_region & 0x01 ) > 0 ) ? 1:0
[0179] cuiAllowedInRegionBFlag = ( ( cui_usage_for_region & 0x02 ) > 0 ) ? 1:0
[0180] cuiAllowedInRegionCFlag = ( ( cui_usage_for_region & 0x04 ) > 0 ) ? 1:0
[0181] cui_usage_for_region can be defined as ue(v). Usage information can be defined and distinguished using a bitmask. Alternatively, cui_usage_for_region can be defined as a string st(v). In this case, cui_usage_for_region can specify a region code in a standardized region code format, such as a Blu-ray region code or DVD region code. When using a region code instead of a bitmask, the information can be simplified and standardized by referring to the Blu-ray region code standard.
[0182] The editing usage constraint information (cui_usage_for_editing) can indicate editing-related usage information. If (cui_usage_for_editing & bitMask) is not 0, it can indicate that the usage information corresponding to the bitMask in Table 12 applies. If cui_usage_for_editing is 0, the usage information can be defined by the application. The value of cui_usage_for_editing can be constrained to be in the range of 0 to 7. In this case, if the value of cui_usage_for_editing is in the range of 8 to 65,535, it can be reserved for future use, and decoders with certain specifications should ignore cui_usage_for_editing.
[0183] bitMaskInterpretation0x01Usage information related to cropping: Identifies whether the picture can be cropped0x02Usage information related to colorization: Identifies whether the picture can be colorized0x04Usage information related to masking: Identifies whether the picture can be made
[0184] cuiAllowedForCroppingFlag can specify whether cui_usage_for_editing indicates usage information including cropping. cuiAllowedForColorizationFlag can specify whether cui_usage_for_editing indicates usage information including colorization. cuiAllowedForMaskingFlag can specify whether cui_usage_for_editing indicates usage information including masking. cuiAllowedForCroppingFlag, cuiAllowedForColorizationFlag, and cuiAllowedForMaskingFlag can be derived as follows.
[0185] cuiAllowedForCroppingflag = ( (cui_usage_for_editing & 0x01) > 0 ) ? 1:0
[0186] cuiAllowedForColorizationFlag = ( (cui_usage_for_editing & 0x02) > 0 ) ? 1:0
[0187] cuiAllowedForMaskingFlag = ( (cui_usage_for_editing & 0x04) > 0 ) ? 1:0
[0188] content_usage_information(payloadSize) {Descriptor cui_cancel_flag u(1) if (!cui_cancel_flag) {cui_persistence_flag u(1)cui_num_context_minus1 u(8) for (i = 0; i <= cui_num_context_minus1; i++) { cui_context_present_flag[i] u(1)if(cui_context_present_flag[i]) cui_context[i] ue(v) cui_usage_category[i] ue(v) if (cuiForAIFlag[i]) { cui_usage_for_ai[i] ue(v)} if (cuiForEditingFlag[i]) { cui_usage_for_editing[i] ue(v)}}}}
[0189] The context count information (cui_num_context_minus1) can specify the number of context entries. For example, the value obtained by adding 1 to cui_num_context_minus1 can be set as the number of context entries.
[0190] The context presence flag (cui_context_present_flag[i]) can indicate whether the ith context information (cui_context[i]) exists. For example, when cui_context_present_flag[i] is 1, this indicates that the ith cui_context exists, and when cui_context_present_flag[i] is 0, this indicates that the ith cui_context does not exist. When cui_context_present_flag[i] is 0, at least one of the ith category information (cui_usage_category[i]), the ith AI usage restriction information (cui_usage_for_ai[i]), or the ith editing usage restriction information (cui_usage_for_editing[i]) can be applied to all cases regardless of the context.
[0191] cui_context[i] may represent the ith context for at least one of cui_usage_for_ai[i] or cui_usage_for_editing[i]. The value of cui_context may be constrained to be in the range 0 to 3, inclusive. If the value of cui_context is in the range 4 to 65,535, it may be reserved for future use, and decoders of certain specifications should ignore cui_usage_for_editing.
[0192] cui_contextInterpretation0commercial use1non-commercial use2official government use3research and academic use
[0193] cui_usage_category[i] may indicate the ith usage category of the picture as defined in Table 15. If (cui_usage_category[i] & bitMask) is not 0, it may indicate that the usage category corresponding to the bitMask value in Table 15 is included. If cui_usage_category[i] is 0, the usage category may be defined by the application. The value of cui_usage_category[i] may be constrained to be in the range of 0 to 3, inclusive. If the value of cui_usage_category[i] is in the range of 4 to 65,535, it is reserved for future use, and decoders with specific specifications should ignore cui_usage_category[i].
[0194] bitMaskInterpretation0x01AI process usage. This SEI includes information on how associated pictures may be used with AI processes0x02Editing usage. This SEI includes information on how associated pictures may be used for editing or modification purposes
[0195] cuiForAIFlag[i] can specify whether cui_usage_category[i] represents a usage category that includes AI process usage. cuiForEditingFlag[i] can specify whether cui_usage_category[i] represents a usage category that includes AI editing usage. cuiForAIFlag[i] and cuiForEditingFlag[i] can be derived as follows.
[0196] cuiForAIFlag[i] = ( (cui_usage_category[i] & 0x01 ) > 0 ) ? 1 : 0 cuiForEditingFlag[i]= ( ( cui_usage_category[i] & 0x02 ) > 0 ) ? 1:0
[0197] cui_usage_for_ai[i] may represent the ith usage information related to AI as defined in Table 16, if any. If (cui_usage_for_ai[i] & bitMask) is not 0, it may indicate that the usage information corresponding to the bitMask value in Table 16 applies. If cui_usage_for_ai[i] is 0, the usage information may be application-defined. The value of cui_usage_for_ai[i] may be constrained to be in the range of 0 to 7. If the value of cui_usage_for_ai[i] is in the range of 8 to 65,535, it is reserved for future use, and decoders with specific specifications should ignore cui_usage_for_ai[i].
[0198] bitMaskInterpretation0x01Usage information related to AI training; Indicates that pictures related to content usage information may be used for AI training0x02Usage information related to AI inference; Indicates that pictures related to content usage information may be used for AI inference0x04Usage information related to AI synthesis: Indicates that pictures related to content usage information may be used for AI synthesis
[0199] cuiAllowedForAITrainingFlag[i] can specify whether the picture associated with the SEI message is usable for AI training purposes. cuiAllowedForAIInferencingFlag[i] can specify whether the picture associated with the SEI message is usable for AI inference purposes. cuiAllowedForAISynthesisFlag[i] can specify whether the picture associated with the SEI message is usable for AI synthesis purposes. cuiAllowedForAITrainingFlag[i], cuiAllowedForAIInferencingFlag[i], and cuiAllowedForAISynthesisFlag[i] can be derived as follows.
[0200] cuiAllowedForAITrainingFlag[i] = ( (cui_usage_for_ai[i] & 0x01 ) > 0 ) ? 1:0
[0201] cuiAllowedForAIInferenceFlag[i] = ( ( cui_usage_for_ai[i] & 0x02 ) > 0 ) ? 1:0
[0202] cuiAllowedForAISynthesisFlag[i] = ( ( cui_usage_for_ai[i] & 0x04 ) > 0 ) ? 1:0
[0203] cui_usage_for_editing[i] may represent the ith usage information related to editing as defined in Table 17. If (cui_usage_for_editing[i] & bitMask) is not 0, it may indicate that the usage information corresponding to the bitMask value in Table 17 applies. If cui_usage_for_editing[i] is 0, the usage information may be application-defined. The value of cui_usage_for_editing[i] may be constrained to be in the range of 0 to 7. If the value of cui_usage_for_editing[i] is in the range of 8 to 65,535, it is reserved for future use, and decoders with specific specifications should ignore cui_usage_for_editing[i].
[0204] bitMaskInterpretation0x01Usage information related to cropping: Identifies that the picture may be cropped0x02Usage information related to colorization: Identifies that the picture may be colorized0x04Usage information related to masking: Identifies that the picture may be masked
[0205] cuiAllowedForCroppingflag[i] can specify whether the picture associated with the SEI message is available for cropping. cuiAllowedForColorizationFlag[i] can specify whether the picture associated with the SEI message is available for colorization. cuiAllowedForMaskingFlag[i] can specify whether the picture associated with the SEI message is available for masking. cuiAllowedForCroppingflag[i], cuiAllowedForColorizationFlag[i], and cuiAllowedForMaskingFlag[i] can be derived as follows.
[0206] cuiAllowedForCroppingflag[i] = ( (cui_usage_for_editing[i] & 0x01) > 0 ) ? 1:0
[0207] cuiAllowedForColorizationFlag[i] = ( (cui_usage_for_editing[i] & 0x02) > 0 ) ? 1:0
[0208] cuiAllowedForMaskingFlag[i] = ( (cui_usage_for_editing[i] & 0x04) > 0 ) ? 1:0
[0209] content_usage_information(payloadSize) {Descriptor cui_cancel_flag u(1) if (!cui_cancel_flag) {cui_persistence_flag u(1)cui_num_context u(8) for (i = 0; i <= cui_num_context; i++) {If(cui_num_context ) cui_context[i] ue(v) cui_usage_category[i] ue(v) if (cuiForAIFlag[i]) { cui_usage_for_ai[i] ue(v)} if (cuiForEditingFlag[i]) { cui_usage_for_editing[i] ue(v)}}}}
[0210] The context count information (cui_num_context) can specify the number of context entries. For example, the value of cui_num_context can be set to the number of context entries. If the value of cui_num_context is 0, this can indicate that cui_context does not exist. If the value of cui_num_context is 0, the information defined in at least one of cui_usage_category, cui_usage_for_ai, or cui_usage_for_editing can be applied to all cases regardless of context.
[0211] cui_context[i], cui_usage_category[i], cui_usage_for_ai[i], and cui_usage_for_editing[i] are as shown in Table 13.
[0212] content_usage_information(payloadSize) {Descriptor cui_cancel_flag u(1) if (!cui_cancel_flag) {cui_persistence_flag u(1)cui_num_context_minus1 u(8) for (i = 0; i <= cui_num_context_minus1; i++) { cui_context[i] ue(v) cui_usage_category[i] ue(v) if (cuiForAIFlag[i]) { cui_usage_for_ai[i] ue(v)} if (cuiForEditingFlag[i]) { cui_usage_for_editing[i] ue(v)}}}}
[0213] cui_context[i] may represent the ith context for at least one of cui_usage_for_ai[i] or cui_usage_for_editing[i], as defined in Table 20. The value of cui_context may be constrained to be in the range 0 to 3, inclusive. If the value of cui_context is in the range 4 to 65,535, it is reserved for future use, and decoders of certain specifications MUST ignore cui_context.
[0214] cui_contextInterpretation0All contexts1commercial use2non-commercial use3official government use4research and academic use
[0215] cui_num_context_minus1, cui_usage_category[i], cui_usage_for_ai[i], and cui_usage_for_editing[i] in Table 19 are as seen with reference to Table 13.
[0216] content_usage_information(payloadSize) {Descriptor cui_cancel_flag u(1) if (!cui_cancel_flag) {cui_persistence_flag u(1)cui_num_context_minus1 u(8) for (i = 0; i <= cui_num_context_minus1; i++) { cui_context_present_flag[i] u(1)if(cui_context_present_flag[i]) cui_context[i] ue(v) cui_usage[i] ue(v)}}}
[0217] cui_usage[i] can represent the ith usage information defined in Table 22. If (cui_usage[i] & bitMask) is not 0, it can indicate that the usage information corresponding to the bitMask value in Table 22 applies. If cui_usage[i] is 0, the usage information can be defined by the application. The value of cui_usage[i] can be constrained to be in the range of 0 to 63. If the value of cui_usage[i] is in the range of 64 to 65,535, it is reserved for future use, and decoders with specific specifications should ignore cui_usage[i].
[0218] bitMaskInterpretation0x01Usage information related to AI training; Indicates that pictures related to content usage information may be used for AI training0x02Usage information related to AI inference; Indicates that pictures related to content usage information may be used for AI inference0x04Usage information related to AI synthesis: Indicates that pictures related to content usage information may be used for AI synthesis0x08Usage information related to cropping: Identifies that the picture may be cropped0x10Usage information related to colorization: Identifies that the picture may be colorized0x20Usage information related to masking: Identifies that the picture may be masked
[0219] cui_num_context_minus1, cui_context_present_flag[i], and cui_context[i] in Table 21 are as described with reference to Table 13.
[0220] Content usage information can be structured as shown in Table 23 below. cui_num_context, cui_context[i], and cui_usage[i] in Table 23 are the same as previously discussed, and any duplicate explanation will be omitted here.
[0221] content_usage_information(payloadSize) {Descriptor cui_cancel_flag u(1) if (!cui_cancel_flag) {cui_persistence_flag u(1)cui_num_context u(8) for (i = 0; i <= cui_num_context; i++) {if(cui_num_context) cui_context[i] ue(v) cui_usage[i] ue(v)}}}
[0222] Content usage information can be structured as shown in Table 24 below. cui_num_context_minus1, cui_context[i], and cui_usage[i] in Table 24 are as discussed above, and any duplicate explanations will be omitted here.
[0223] content_usage_information(payloadSize) {Descriptor cui_cancel_flag u(1) if (!cui_cancel_flag) {cui_persistence_flag u(1)cui_num_context_minus1 u(8) for (i = 0; i <=cui_num_context_minus1; i++) { cui_context[i] ue(v) cui_usage[i] ue(v)}}}
[0224] Content usage information may be structured as shown in Table 25 below.
[0225] content_usage_information(payloadSize) {Descriptorcui_cancel_flagu(1)if(!cui_cancel_flag) {cui_persistence_flagu(1)cui_usage_categoryue(v)if(cuiForAIFlag){cui_usage_for_aiue(v)cui_context_for_ai_present_flagu(1)if(cui_context_for_ai_present_flag)cui_context_f or_aiue(v)}if(cuiForEditingFlag){cui_usage_for_editingue(v)cui_context_for_editing_present_flagu(1)if(cui_context_for_editing_present_flag)cui_context_for_editingue(v)}}}
[0226] cui_usage_category, cui_usage_for_ai, and cui_usage_for_editing in Table 25 are the same as previously discussed, and duplicate explanations will be omitted here.
[0227] cui_context_for_ai_present_flag can indicate whether context information (cui_context_for_ai) exists for the value of cui_usage_for_ai. For example, if cui_context_for_ai_present_flag is 1, it can indicate that cui_context_for_ai exists, and if cui_context_for_ai_present_flag is 0, it can indicate that cui_context_for_ai does not exist. If cui_context_for_ai_present_flag is 0, the usage information defined in cui_usage_for_ai can be applied to all cases regardless of context.
[0228] cui_context_for_ai can represent the context of cui_usage_for_ai defined in Table 26. If (cui_context_for_ai & bitMask) is not 0, it can indicate that the context corresponding to the bitMask value in Table 26 is applied. If cui_context_for_ai is greater than 0 and (cui_context_for_ai & bitMask) is 0, the context corresponding to the bitMask value may not be applied. cui_context_for_ai can be restricted not to have a value of 0.
[0229] BitmaskInterpretation0x0001commercial use0x0002non-commercial use0x0004official government use0x0004research and academic use0x0008 to 0xFFFFReserved for future use of ITU and ISO
[0230] cui_context_for_editing_present_flag can indicate whether context information (cui_context_for_editing) exists for the value of cui_usage_for_editing. For example, if cui_context_for_editing_present_flag is 1, it can indicate that cui_context_for_editing exists, and if cui_context_for_editing_present_flag is 0, it can indicate that cui_context_for_editing does not exist. If cui_context_for_editing_present_flag is 0, the usage information defined in cui_usage_for_editing can be applied to all cases regardless of context.
[0231] cui_context_for_editing can represent the context of cui_usage_for_editing defined in Table 27. If (cui_context_for_editing & bitMask) is not 0, it can indicate that the context corresponding to the bitMask value in Table 27 is applied. If cui_context_for_editing is greater than 0 and (cui_context_for_editing & bitMask) is 0, the context corresponding to the bitMask value may not be applied. cui_context_for_editing can be restricted to not have a value of 0.
[0232] BitmaskInterpretation0x0001commercial use0x0002non-commercial use0x0004official government use0x0004research and academic use0x0008 to 0xFFFFReserved for future use of ITU and ISO
[0233] Content usage information may be structured as shown in Table 28 below.
[0234] content_usage_information( payloadSize ) {Descriptorcui_cancel_flagu(1)if( !cui_cancel_flag ) {cui_persistence_flagu(1)cui_usage_categoryue(v)if(cuiForAIFlag)cui_usage_for_aiue(v)if(cuiForEditingFlag)cui_usage_for_editingue(v)}}
[0235] The cui_usage_category in Table 28 is as discussed above.
[0236] cui_usage_for_ai can represent AI-related usage information defined in Table 29. If (cui_usage_for_ai & bitMask) is not 0, it can indicate that the usage information corresponding to the bitMask value in Table 29 applies. If cui_usage_for_ai is 0, the usage information can be defined by the application. The value of cui_usage_for_ai can be constrained to be in the range of 0 to 7. If the value of cui_usage_for_ai is in the range of 8 to 65,535, it is reserved for future use, and decoders with certain specifications should ignore cui_usage_for_ai.
[0237] bitMaskInterpretation0x01Usage information related to AI training; Indicates that pictures related to content usage information should not be used for AI training0x02Usage information related to AI inference; Indicates that pictures related to content usage information should not be used for AI inference0x04Usage information related to AI synthesis: Indicates that pictures related to content usage information should not be used for AI synthesis
[0238] cuiDisallowedForAITrainingFlag can specify whether the picture associated with the SEI message is unavailable for AI training purposes. cuiDisallowedForAIInferencingFlag can specify whether the picture associated with the SEI message is unavailable for AI inference purposes. cuiDisallowedForAISynthesisFlag can specify whether the picture associated with the SEI message is unavailable for AI synthesis purposes. cuiDisallowedForAITrainingFlag, cuiDisallowedForAIInferencingFlag, and cuiDisallowedForAISynthesisFlag can be derived as follows.
[0239] cuiDisallowedForAITrainingFlag = ( (cui_usage_for_ai & 0x01 ) > 0 ) ? 1:0
[0240] cuiDisallowedForAIInferenceFlag = ( ( cui_usage_for_ai & 0x02 ) > 0 ) ? 1:0
[0241] cuiDisallowedForAISynthesisFlag = ( ( cui_usage_for_ai & 0x04 ) > 0 ) ? 1:0
[0242] cui_usage_for_editing can represent editing-related usage information as defined in Table 30. If (cui_usage_for_editing & bitMask) is not 0, it can indicate that the usage information corresponding to the bitMask value in Table 30 applies. If cui_usage_for_editing is 0, the usage information can be defined by the application. The value of cui_usage_for_editing can be constrained to be in the range of 0 to 7. If the value of cui_usage_for_editing is in the range of 8 to 65,535, it is reserved for future use, and decoders with certain specifications should ignore cui_usage_for_editing.
[0243] bitMaskInterpretation0x01Usage information related to cropping: Identifies that the picture should not be cropped0x02Usage information related to colorization: Identifies that the picture should not be colorized0x04Usage information related to masking: Identifies that the picture should not be masked
[0244] cuiDisallowedForCroppingFlag can specify whether the picture associated with the SEI message is unavailable for cropping. cuiDisallowedForColorizationFlag can specify whether the picture associated with the SEI message is unavailable for colorization. cuiDisallowedForMaskingFlag can specify whether the picture associated with the SEI message is unavailable for masking. cuiDisallowedForCroppingFlag, cuiDisallowedForColorizationFlag, and cuiDisallowedForMaskingFlag can be derived as follows.
[0245] cuiDisallowedForCroppingflag = ( (cui_usage_for_editing & 0x01 ) > 0 ) ? 1:0
[0246] cuiDisallowedForColorizationFlag = ( ( cui_usage_for_editing & 0x02 ) > 0 ) ? 1:0
[0247] cuiDisallowedForMaskingFlag = ( ( cui_usage_for_editing & 0x04 ) > 0 ) ? 1:0
[0248] The content usage information may be configured in a supplemental enhancement information (SEI) message of the bitstream. The SEI message may be included in a network abstraction layer (NAL) unit of the bitstream. All or part of the content usage information may be configured in an SEI message in which the AI usage restriction information described below is defined. Alternatively, the content usage information according to the present disclosure may be configured in a high level syntax of the bitstream. Here, the high level syntax may be at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Alternatively, the content usage information according to the present disclosure may be defined as a separate NAL unit type in the bitstream.
[0249] Example 2
[0250] The bitstream may contain AI usage restrictions (AUR). AI usage restrictions may indicate at least one of the following: usage restrictions for AI applications, optional context, or contact information.
[0251] AI usage constraints may include the AUR cancellation flag (aur_cancel_flag). If aur_cancel_flag is 1, this may indicate that the SEI message cancels the persistence of a previous AI usage constraint in the output order. If aur_cancel_flag is 0, this may indicate that an AI usage constraint follows.
[0252] AI usage constraints can include the AUR persistence flag (aur_persistence_flag). aur_persistence_flag can indicate the persistence of AI usage constraints for the current layer. For example, if aur_persistence_flag is 0, this can indicate that the AI usage constraint applies only to the currently decoded picture. If aur_persistence_flag is 1, this can indicate that the AI usage constraint applies to the currently decoded picture and persists for all subsequent pictures belonging to the current layer.
[0253] AI usage restrictions can include information about the number of restrictions (aur_num_restrictions_minus1). aur_num_restrictions_minus1 can indicate the number of signaled restriction entries. For example, the number of signaled restriction entries can be aur_num_restrictions_minus1 plus 1.
[0254] AI usage restrictions may include AI usage restriction information (aur_restriction). aur_restriction may indicate restrictions or no restrictions on usage for a specific purpose. aur_restriction may be signaled based on aur_num_restrictions_minus1. aur_restriction may be signaled as many times as the number of restriction entries specified by aur_num_restrictions_minus1.
[0255] For example, aur_restriction can be defined as shown in Table 31 below.
[0256] aur_restrictionInterpretation0Do not use for AI training1Do not use for generative (modification or creation) AI2Do not use in any AI application
[0257] If the value of aur_restriction is the first value (for example, if the value of aur_restriction is 0), this may indicate that it is not used (or cannot be used) for AI training. If the value of aur_restriction is the second value (for example, if the value of aur_restriction is 1), this may indicate that it is not used (or cannot be used) for generative AI. If the value of aur_restriction is the third value (for example, if the value of aur_restriction is 2), this may indicate that it is not used (or cannot be used) in any AI application.
[0258] The value of aur_restriction can be constrained to be in the range 0 to 2. If the value of aur_restriction is in the range 3 to 65,535, it is reserved for future use, and decoders of certain specifications should ignore aur_restriction.
[0259] Alternatively, aur_restriction can be defined as shown in Table 32 below.
[0260] aur_restrictionInterpretation0No restriction1Do not use for AI training2Do not use for AI inferencing3Do not use for generative (modification or creation) AI4Do not use in any AI application
[0261] If the value of aur_restriction is the first value (for example, if the value of aur_restriction is 1), this may indicate that it is not used (or cannot be used) for AI training. If the value of aur_restriction is the second value (for example, if the value of aur_restriction is 2), this may indicate that it is not used (or cannot be used) for AI inferencing. If the value of aur_restriction is the third value (for example, if the value of aur_restriction is 3), this may indicate that it is not used (or cannot be used) for generative AI. If the value of aur_restriction is the fourth value (for example, if the value of aur_restriction is 4), this may indicate that it is not used (or cannot be used) in any AI application.
[0262] The value of aur_restriction can be constrained to be in the range 0 to 4. If the value of aur_restriction is in the range 5 to 65,535, it is reserved for future use, and decoders of certain specifications should ignore aur_restriction.
[0263] AI usage constraints may include a context presence flag (aur_context_present_flag). aur_context_present_flag may indicate whether context information (aur_context) exists for the value of aur_restriction. For example, if aur_context_present_flag is 1, it indicates that aur_context exists, and if aur_context_present_flag is 0, it indicates that aur_context does not exist. If aur_context_present_flag is 0, the constraints defined in aur_restriction can be applied in all cases regardless of context. aur_context_present_flag may be signaled based on aur_num_restrictions_minus1. aur_context_present_flag may be signaled as many times as the number of constraint entries according to aur_num_restrictions_minus1.
[0264] AI usage restrictions can include contextual information (aur_context). aur_context can represent the context for aur_restriction. For example, aur_context can be defined as shown in Table 33.
[0265] BitmaskInterpretation0x0001commercial use0x0002non-commercial use0x0004official government use0x0008research and academic use0x0010 to 0xFFFFReserved for future use of ITU and ISO
[0266] If (aur_context & bitMask) is not 0, this may indicate that the context corresponding to the bitMask value in Table 33 applies. If aur_context is greater than 0 and (aur_context & bitMask) is 0, the context corresponding to the bitMask value may not apply. aur_context may be restricted to not have a value of 0. aur_context may be signaled based on at least one of aur_num_restrictions_minus1 or aur_context_present_flag. aur_context may be signaled as many times as the number of restriction entries according to aur_num_restrictions_minus1. aur_context may be signaled based on aur_context_present_flag having a value of 1, and may not be signaled based on aur_context_present_flag having a value of 0.
[0267] AI usage constraints may include a contact information presence flag (aur_contact_info_present_flag). aur_contact_info_present_flag may indicate whether contact information (aur_contact_info) for the AI usage constraints exists. aur_payload_bit_equal_to_zero must be equal to 0. aur_contact_info may include a string indicating that the contact information for the entity for which additional information can be obtained.
[0268] For example, AI usage constraints can be structured as shown in Table 34 below.
[0269] AI_usage_restrictions ( payloadSize ) {Descriptoraur_cancel_flagu(1)if (!aur_cancel_flag) {aur_persistence_flagu(1)aur_num_restrictions_minus1u(8)for( i = 0; i <= aur_num_restrictions_minus1; i++ ){aur_restrictionue(v)aur_context_present_flagu(1)if (aur_context_present_flag)aur_contextue(v)}aur_contact_info_present_flagu(1)if (aur_contact_info_present_flag) {while (!byte_aligned())aur_payload_bit_equal_to_zerof(1)aur_contact_inpost(v)}}}
[0270] The AI usage constraint information may be configured in a supplemental enhancement information (SEI) message of the bitstream. The SEI message may be included in a network abstraction layer (NAL) unit of the bitstream. Alternatively, the AI usage constraint information according to the present disclosure may be configured in a high-level syntax of the bitstream. Here, the high-level syntax may be at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Alternatively, the AI usage constraint information according to the present disclosure may be defined as a separate NAL unit type within the bitstream.
[0271] Example 3
[0272] A bitstream may include a text description. The text description may represent a text description of one or more pictures. For example, the text description may be structured as shown in Table 35 below.
[0273] text_description( payloadSize ) {Descriptortxt_descr_idu(14)txt_cancel_flagu(1)if( !txt_cancel_flag ) {txt_persistence_flagu(1)txt_descr_purposeu(8)txt_num_strings_minus1u(8)for( i = 0; i <= txt_num_strings_minus1; i++ ) {txt_descr_string_lang[ i ]st(v)txt_descr_string[ i ]st(v)}}}
[0274] txt_descr_id can represent the identifier value of the corresponding text description. The value of txt_descr_id can be constrained to be in the range of 1 to 16383. The value 0 can be reserved.
[0275] If txt_cancel_flag is 1, this may indicate that the text description applied to the current layer will cancel the persistence of the previous text description with the same txt_descr_id in the output order. If txt_cancel_flag is 0, this may indicate that a text description will follow.
[0276] The txt_persistence_flag can indicate the persistence of text information for the current layer. For example, if txt_persistence_flag is 0, this can indicate that the text description applies only to the currently decoded picture. If txt_persistence_flag is 1, this can indicate that the text description applies to the currently decoded picture and is maintained for all subsequent pictures belonging to the current layer in output order.
[0277] txt_descr_purpose can indicate the purpose of the text description, as defined in Table 36. The value of text_descr_purpose can be constrained to be in the range 0 to 6. Values in the range 7 to 255 for text_descr_purpose are reserved for future use, and decoders of a particular specification should allow values of text_descr_purpose in the range 0 to 255.
[0278] ValueInterpretation0Application defined1Copyright information2AI marking information3General comment information4Content advisory rating information conforming to US. And Canadian Rating Region Tables (RRT)5Tag URI for identifying the bitstream6Content usage information7..255Reserved
[0279] The value of txt_num_strings_minus1 plus 1 can represent the number of entries for txt_descr_string_lang[i] and txt_descr_string[i].
[0280] txt_descr_string_lang[i] can represent the language of txt_descr_string[i]. The language of txt_descr_string[i] can be restricted to be specified by a language tag defined in IETF RFC 5646. The length of txt_descr_string_lang[i] can be restricted to be in the range 0 to 49.
[0281] txt_descr_string[i] can represent the ith text description information string interpreted as the value specified in txt_descr_purpose.
[0282] If txt_descr_purpose is 0, the interpretation of the information contained in txt_descr_string can be defined by the application.
[0283] If txt_descr_purpose is 1, txt_descr_string[i] may represent copyright information related to pictures within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[0284] When txt_descr_purpose is 2, txt_descr_string[i], if not a null string, may represent AI marking information related to pictures within the persistence scope of this SEI message. When txt_descr_purpose is 2, the string may contain information about machine learning-based processing, the intended use of the decoded picture, or other aspects of the related picture.
[0285] If txt_descr_purpose is 3, txt_descr_string[i] may represent a plain text label description associated with a picture within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[0286] If txt_descr_purpose is 4, txt_descr_string[i] may represent content recommendation rating information compliant with the U.S. and Canadian Rating Territory Tables (RRT) for pictures within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[0287] If txt_descr_purpose is 5, txt_descr_string[i] contains a tag URI with the syntax and semantics defined in IETF RFC 4151, which can identify CLVS.
[0288] If txt_descr_purpose is 6, txt_descr_string[i], when not a null string, may represent content usage information related to a picture within the persistence scope of this SEI message. If txt_descr_purpose is 6, the string may contain usage information for the picture. The usage information may include constraints and recommendations that may be related to the AI process, intended usage area, editorial rights, or other relevant usage aspects.
[0289] cuiAllowedForAITrainingFlag, cuiAllowedForAIInferenceFlag, cuiAllowedForAIEditingFlag, cuiAllowedInRegionAFlag, cuiAllowedInRegionBFlag, cuiAllowedInRegionCFlag, cuiAllowedForCroppingFlag, cuiAllowedForColorizationFlag, and cuiAllowedForMaskingFlag may indicate that the usage information includes AI training, AI inference, AI editing, usage in Region A, Region B, Region C, cropping, colorization, and masking, respectively.
[0290] Table 37 below is an example of a free-form string for content usage information.
[0291] cuiAllowedForAITrainingFlag:1cuiAllowedForAIInferenceFlag:1cuiAllowedForAIEditingFlag:1cuiAllowedInRegionAFlag:1cuiAllowedInRegionBFlag:1cuiAllowedInRegionCFlag:1cuiAllowedForCroppingFlag: 0cuiAllowedForColorizationFlag:0cuiAllowedForMaskingFlag: 0
[0292] The text description information may be configured in a supplemental enhancement information (SEI) message of the bitstream. The SEI message may be included in a network abstraction layer (NAL) unit of the bitstream. Alternatively, the text description information according to the present disclosure may be configured in a high-level syntax of the bitstream. Here, the high-level syntax may be at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Alternatively, the text description information according to the present disclosure may be defined as a separate NAL unit type within the bitstream.
[0293] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.
[0294] Referring to FIG. 5, the decoding device (300) may include a receiving unit (500), a video information extraction unit (510), and a video restoration unit (520).
[0295] The receiving unit (500) can receive a bitstream including an encoded video picture.
[0296] The video information extraction unit (510) can extract video information about an encoded video picture from a bitstream. In addition, the video information extraction unit (710) can extract at least one of content usage information, AI usage restrictions, or text description from the bitstream, as described with reference to FIG. 4.
[0297] The video restoration unit (520) can restore an encoded video picture based on the extracted video information.
[0298] FIG. 6 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.
[0299] A video picture to be encoded can be received (S600).
[0300] The received video picture can be encoded to generate video information about the video picture (S610).
[0301] A bitstream including video information about a video picture can be generated (S420).
[0302] Additionally, at least one of content usage information, AI usage restrictions, or text descriptions applicable to the bitstream may be generated, as described with reference to FIG. 4. At least one of the generated content usage information, AI usage restrictions, or text descriptions may be included in the bitstream.
[0303] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.
[0304] Referring to FIG. 7, the encoding device (200) may include a receiving unit (700), a video compression unit (710), and a bitstream generation unit (720).
[0305] The receiving unit (700) can receive one or more video pictures to be encoded.
[0306] The video compression unit (710) may encode one or more received video pictures to generate video information about the video pictures. The video compression unit (710) may generate at least one of content usage information, AI usage constraints, or text descriptions applied to the bitstream.
[0307] The bitstream generation unit (720) can generate a bitstream including the video information. The bitstream generation unit (720) can further generate a bitstream including at least one of the generated content usage information, AI usage restrictions, or text description.
[0308] In the embodiments described above, the methods are described based on a flowchart as a series of steps or blocks. However, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0309] The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0310] When the embodiments in this document are implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above-described function. The module can be stored in memory and executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chipset, logic circuit, and / or data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored on a digital storage medium.
[0311] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0312] In addition, the processing method to which the embodiment(s) of the present specification are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0313] Additionally, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0314] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0315] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0316] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.
[0317] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0318] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.
[0319] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0320] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0321] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0322] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
Claims
A step of receiving a bitstream including an encoded video picture; and A step of restoring an encoded video picture included in the above bitstream, The above bitstream includes AI usage restriction information indicating restrictions on the use of AI (Artificial Intelligence), A method in which the above AI usage restriction information is obtained from a NAL (network abstraction layer) unit of the bitstream. In the first paragraph, The above AI usage restriction information of the first value indicates that it cannot be used for AI inference, A method in which the above AI usage restriction information of the second value indicates that it cannot be used for AI training. In the first paragraph, The above AI usage restriction information of the first value indicates that it cannot be used for AI training, The above AI usage restriction information of the second value indicates that it cannot be used for generative AI, A method for indicating that the AI usage restriction information of the third value cannot be used in any AI application. In the third paragraph, A method wherein a smaller value is assigned to the AI usage restriction information indicating that the AI cannot be used for AI training than to the AI usage restriction information indicating that the AI cannot be used for the generative AI. In the first paragraph, The above AI usage restriction information is related to four restrictions on the use of the above AI. In paragraph 5, The above AI usage restriction information of the first value indicates that it cannot be used for AI training, The above AI usage restriction information of the second value indicates that it cannot be used for AI inference, The above AI usage restriction information of the third value indicates that it cannot be used for generative AI, A method for indicating that the above AI usage restriction information of the fourth value cannot be used in any AI application. In paragraph 6, A method wherein a smaller value is assigned to the AI usage restriction information indicating that the AI is unusable for AI training than to the AI usage restriction information indicating that the AI is unusable for AI inference. In the first paragraph, A method wherein the bitstream further includes restriction count information indicating the number of restriction entries being signaled. In paragraph 8, A method in which the above AI usage restriction information is signaled from the bitstream based on the number of restriction information. A step of receiving a video picture to be encoded; A step of encoding the received video picture to generate video information about the encoded video picture; A step of generating AI usage restriction information indicating restrictions on the use of AI (Artificial Intelligence); and Including a step of generating a bitstream including the above video information and the above AI usage restriction information, A method wherein the above AI usage restriction information is encoded in a NAL (network abstraction layer) unit of the bitstream. A computer-readable storage medium storing a bitstream generated by the method according to Article 10. A step of generating a bitstream; wherein the bitstream is generated based on the steps of: receiving a video picture to be encoded; encoding the received video picture to generate video information about the encoded video picture; and generating AI usage constraint information indicating constraints on the use of AI (Artificial Intelligence). Including a step of transmitting data including the above bitstream, A method wherein the above AI usage restriction information is encoded in a NAL (network abstraction layer) unit of the bitstream.
Citation Information
Patent Citations
KR20240053550A