Method for decoding image information, method for encoding image, method for storing bitstream of image information, and method for transmitting bitstream of image information
By employing feature-based optimization techniques for encoding and decoding, the method addresses the challenge of high data volume in high-resolution images, reducing costs through efficient compression.
Patent Information
- Application Number
- PCT/KR2025/000450
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-17
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to a significant increase in the amount of information transmitted, resulting in higher costs for storage and transmission, necessitating a more efficient image compression technology.
A method for encoding and decoding image information based on features, utilizing feature-based optimization techniques such as rate-distortion optimization, preprocessing, and quantization parameter adaptation to optimize image encoding and decoding processes.
This approach reduces the amount of data required for high-resolution images, thereby lowering storage and transmission costs while maintaining image quality.
Smart Images

Figure KR2025000450_17072025_PF_FP_ABST
Abstract
Description
A method for decoding image information, a method for encoding an image, a method for storing a bitstream of image information, and a method for transmitting a bitstream of image information
[0001] The present disclosure relates to a method for decoding image information, a method for encoding image information, and a method for transmitting a bitstream of image information.
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has been increasing across various fields. As image data becomes higher resolution and higher quality, the amount of information transmitted, or bits, increases relative to conventional image data. This increase in information or bits transmitted leads to increased transmission and storage costs.
[0003] Accordingly, a highly efficient image compression technology is required to effectively transmit, store, and play high-resolution, high-quality image information.
[0004] The present disclosure provides a method for decoding image information based on a feature, a method for encoding image information, and a method and device for transmitting a bitstream of image information.
[0005] According to one aspect of the present disclosure, a method for decoding image information may include acquiring image information including encoder optimization information, determining a type of encoder optimization based on the encoder optimization information, and processing the image based on the type of encoder optimization. The encoder optimization information may include optimization type information indicating a type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes the image based on features of the image.
[0006] According to one aspect of the present disclosure, a device for decoding image information may include a memory and a processor connected to the memory. The processor may include obtaining image information including encoder optimization information, determining a type of encoder optimization based on the encoder optimization information, and processing the image based on the type of encoder optimization. The encoder optimization information may include optimization type information indicating a type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes the image based on features of the image.
[0007] Whether or not the above feature-based optimization is performed can be determined based on the logical product of the above optimization type information and a predetermined first bitmask.
[0008] The above encoder optimization information may further include feature-based optimization type information indicating the type of the feature-based optimization.
[0009] The above feature-based optimization may include at least one of feature-based rate-distortion optimization, feature-based preprocessing, or feature-based quantization parameter (QP) adaptation / rate control based on the value of the feature-based optimization type information.
[0010] The above feature-based optimization may include at least one of feature-based rate-distortion optimization, feature-based preprocessing, or feature-based quality adaptation / rater control based on a logical product of the feature-based optimization type information and a predetermined second bitmask.
[0011] The above encoder optimization information may further include some feature usage information indicating whether some or all features are used for optimization.
[0012] The encoder optimization information may further include optimization suitability information indicating whether the encoder optimization is suitable for user viewing or machine analysis.
[0013] The above optimization suitability information may include a 3-bit or 4-bit indicator indicating suitability for user viewing and suitability for machine analysis.
[0014] According to one aspect of the present disclosure, a method for encoding an image may include determining a type of encoder optimization, processing an image based on the type of encoder optimization, generating encoder optimization information based on the type of encoder optimization, and encoding image information including the encoder optimization information. The encoder optimization information may include optimization type information indicating the type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes an image based on a feature of the image.
[0015] According to one aspect of the present disclosure, a device for encoding a video may include a memory and a processor connected to the memory. The processor may include determining a type of encoder optimization, processing the video based on the type of encoder optimization, generating encoder optimization information based on the type of encoder optimization, and encoding video information including the encoder optimization information. The encoder optimization information may include optimization type information indicating the type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes the video based on features of the video.
[0016] According to one aspect of the present disclosure, a method for storing a bitstream of image information in a computer-readable storage medium may include obtaining a bitstream of the image information and storing data including the bitstream in the storage medium. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined according to the image. The encoder optimization information may include optimization type information indicating a type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes the image based on a feature of the image.
[0017] According to one aspect of the present disclosure, a computer-readable storage medium may store data including a bitstream of image information. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined according to the image. The encoder optimization information may include optimization type information indicating the type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes the image based on features of the image.
[0018] According to one aspect of the present disclosure, a method for transmitting a bitstream of image information may include obtaining a bitstream of the image information and transmitting data including the bitstream. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined according to the image. The encoder optimization information may include optimization type information indicating a type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes the image based on features of the image.
[0019] According to one aspect of the present disclosure, a device for transmitting a bitstream of image information may include a processor for obtaining a bitstream of the image information, and a transmission unit for transmitting data including the bitstream. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined according to the image. The encoder optimization information may include optimization type information indicating a type of encoder optimization. The encoder optimization may include feature-based optimization that optimizes the image based on a feature of the image.
[0020] According to the present disclosure, a method for decoding image information based on a feature, a method for encoding image information, and a method and device for transmitting a bitstream of image information can be provided.
[0021] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.
[0022] FIG. 1 is a schematic diagram of a VCM system to which embodiments of the present disclosure can be applied.
[0023] FIG. 2 is a schematic diagram illustrating a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0024] FIG. 3 is a schematic diagram of an image / video encoder to which embodiments of the present disclosure can be applied.
[0025] FIG. 4 is a schematic diagram of an image / video decoder to which embodiments of the present disclosure can be applied.
[0026] FIG. 5 is a flowchart schematically illustrating a feature / feature map encoding procedure to which embodiments of the present disclosure may be applied.
[0027] FIG. 6 is a flowchart schematically illustrating a feature / feature map decoding procedure to which embodiments of the present disclosure may be applied.
[0028] Figure 7 is a diagram showing an example of a feature extraction method using a feature extraction network.
[0029] Figure 8a is a diagram showing data distribution characteristics of a video source.
[0030] Figure 8b is a diagram showing the data distribution characteristics of the feature set.
[0031] FIG. 9 is a diagram for explaining a transformation and inverse transformation process to which embodiments of the present disclosure can be applied.
[0032] FIG. 10 is a diagram illustrating an example of a quantization group according to one embodiment of the present disclosure.
[0033] Figure 11 shows a block diagram of CABAC for encoding one syntax element.
[0034] Figures 12 and 13 are drawings for explaining the entropy encoding procedure.
[0035] Figures 14 and 15 are drawings for explaining the entropy decoding procedure.
[0036] Figure 16 illustrates an example of a VCM hierarchy.
[0037] Figure 17 is a diagram showing an example of a bitstream composed of encoded abstract features and NNAL information.
[0038] Figure 18 illustrates an example of encoding optimization and use of an encoded bitstream that includes privacy processing.
[0039] Figure 19 shows an example of data attribute and privacy information definitions.
[0040] Figure 20 shows an example where an area required for image analysis and an area containing personal information exist together.
[0041] FIG. 21 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.
[0042] FIG. 22 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.
[0043] FIG. 23 is a diagram illustrating an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0044] FIG. 24 is a diagram showing another example of a content streaming system to which embodiments of the present disclosure can be applied.
[0045] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.
[0046] In describing embodiments of the present disclosure, detailed descriptions of known configurations or functions will be omitted if they are deemed to obscure the gist of the present disclosure. Furthermore, portions unrelated to the description of the present disclosure in the drawings have been omitted, and similar portions have been designated with similar reference numerals.
[0047] In the present disclosure, when a component is said to be "connected," "coupled," or "connected" to another component, this may include not only a direct connection, but also an indirect connection in which another component exists in between. Furthermore, when a component is said to "include" or "have" another component, unless otherwise specifically stated, this does not exclude the other component, but rather implies that the other component may be included.
[0048] In this disclosure, terms such as first, second, etc. are used solely to distinguish one component from another, and do not limit the order or importance of components unless specifically stated otherwise. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0049] In this disclosure, distinct components are used to clearly illustrate their respective characteristics, and do not necessarily imply that the components are separated. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not specifically mentioned, such integrated or distributed embodiments are also included within the scope of this disclosure.
[0050] In the present disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments comprising a subset of the components described in one embodiment are also within the scope of the present disclosure. Furthermore, embodiments including other components in addition to the components described in various embodiments are also within the scope of the present disclosure.
[0051] The present disclosure relates to encoding and decoding of images, and terms used in the present disclosure may have their usual meanings commonly used in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure.
[0052] The present disclosure may be applied to a method disclosed in the Versatile Video Coding (VVC) standard and / or the Video Coding for Machines (VCM) standard. In addition, the present disclosure may be applied to a method disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or a next-generation video / image coding standard (e.g., H.267 or H.268, etc.).
[0053] The present disclosure presents various embodiments related to video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other. In the present disclosure, "video" may mean a set of a series of images over time. An "image" may be information generated by artificial intelligence (AI). Input information used in the process of AI performing a series of tasks, information generated during the information processing process, and output information may be used as images. A "picture" generally means a unit representing one image at a specific time point, and a slice / tile is a coding unit that constitutes a part of a picture in encoding. A single picture may be composed of one or more slices / tiles. In addition, a slice / tile may include one or more CTUs (coding tree units). The CTUs may be divided into one or more CUs. A tile is a rectangular area within a specific tile row and a specific tile column within a picture, and may be composed of multiple CTUs. A tile row may be defined as a rectangular area of CTUs, and may have a height equal to the height of the picture, and a width specified by a syntax element signaled from a bitstream portion such as a picture parameter set. A tile row may be defined as a rectangular area of CTUs, and may have a width equal to the width of the picture, and a height specified by a syntax element signaled from a bitstream portion such as a picture parameter set. A tile scan is a predetermined sequential ordering method of CTUs that divide a picture. Here, CTUs may be sequentially ordered within a tile according to a CTU raster scan, and tiles within a picture may be sequentially ordered according to a raster scan order of tiles in the picture.A slice may contain an integer number of complete tiles, or a contiguous integer number of complete CTU rows within a tile of a picture. A slice may be exclusively contained in a single NAL unit. A picture may consist of one or more tile groups. A tile group may contain one or more tiles. A brick may represent a rectangular region of CTU rows within a tile of a picture. A tile may contain one or more bricks. A brick may represent a rectangular region of CTU rows within a tile. A tile may be divided into multiple bricks, and each brick may contain one or more CTU rows belonging to the tile. A tile that is not divided into multiple bricks may also be treated as a brick.
[0054] In the present disclosure, "pixel" or "pel" may refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component.
[0055] In one embodiment, particularly when applied to VCM, a pixel / pixel value may represent independent information of each component or a pixel / pixel value of a component generated through combination, synthesis, or analysis when there is a picture composed of a set of components with different characteristics and meanings. For example, in an RGB input, only the pixel / pixel value of R may be represented, only the pixel / pixel value of G may be represented, or only the pixel / pixel value of B may be represented. For example, only the pixel / pixel value of the Luma component synthesized using the R, G, and B components may be represented. For example, only the pixel / pixel value of the image information extracted through analysis of the R, G, and B components may be represented.
[0056] In the present disclosure, a "unit" may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., Cb, Cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows. In one embodiment, particularly when applied to VCM, a unit may represent a basic unit containing information for performing a specific task.
[0057] In the present disclosure, the "current block" may mean one of the following: a "current coding block," a "current coding unit," a "block to be encoded," a "block to be decoded," or a "block to be processed." When prediction is performed, the "current block" may mean a "current prediction block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, the "current block" may mean a "current transformation block" or a "block to be transformed." When filtering is performed, the "current block" may mean a "block to be filtered."
[0058] Additionally, in the present disclosure, "current block" may mean "luma block of the current block" unless explicitly described as a chroma block. "Chroma block of the current block" may be expressed by explicitly including explicit description of chroma block, such as "chroma block" or "current chroma block."
[0059] In this disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Additionally, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."
[0060] In this disclosure, "or" may be interpreted as "and / or." For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B." Alternatively, "or" in this disclosure may mean "additionally or alternatively."
[0061] The present disclosure relates to VCM (Video / image coding for machines).
[0062] VCM (Virtual Compression Modeling) is a compression technology that encodes / decodes a portion of a source image / video or information obtained from the source image / video for machine vision purposes. In VCM, the target of encoding / decoding can be referred to as a feature. A feature can refer to information extracted from the source image / video based on the task purpose, requirements, surrounding environment, etc. Features may have a different information format than the source image / video, and accordingly, the compression method and representation format of the feature may also differ from those of the video source.
[0063] VCMs can be applied to a variety of applications. For example, in surveillance systems that recognize and track objects or people, VCMs can be used to store or transmit object recognition information. Furthermore, in intelligent transportation or smart traffic systems, VCMs can be used to transmit vehicle location information collected from GPS, sensing information collected from LIDAR and radar, and various vehicle control information to other vehicles or infrastructure. Furthermore, in smart city applications, VCMs can be used to perform individual tasks for interconnected sensor nodes or devices.
[0064] The present disclosure provides various embodiments related to feature / feature map coding. Unless otherwise specified, the embodiments of the present disclosure may be implemented individually or in combination of two or more.
[0065] VCM System Overview
[0066] FIG. 1 is a schematic diagram of a VCM system to which embodiments of the present disclosure can be applied.
[0067] Referring to FIG. 1, the VCM system may include an encoding device (10) and a decoding device (20).
[0068] The encoding device (10) can compress / encode features / feature maps extracted from source images / videos to generate a bitstream, and transmit the generated bitstream to a decoding device (20) via a storage medium or a network. The encoding device (10) may also be referred to as a feature encoding device. In a VCM system, features / feature maps can be generated in each hidden layer of a neural network. The size and number of channels of the generated feature maps can vary depending on the type of neural network or the location of the hidden layer. In the present disclosure, the feature maps can be referred to as a feature set.
[0069] The encoding device (10) may include a feature acquisition unit (11), an encoding unit (12), and a transmission unit (13).
[0070] The feature acquisition unit (11) can acquire features / feature maps for source images / videos. In some embodiments, the feature acquisition unit (11) can acquire features / feature maps from an external device, for example, a feature extraction network. In this case, the feature acquisition unit (11) performs a feature reception interface function. Alternatively, the feature acquisition unit (11) can acquire features / feature maps by executing a neural network (e.g., CNN, DNN, etc.) using the source images / videos as input. In this case, the feature acquisition unit (11) performs a feature extraction network function.
[0071] According to an embodiment, the encoding device (10) may further include a source image generation unit (not shown) for obtaining a source image / video. The source image generation unit may be implemented as an image sensor, a camera module, or the like, and may obtain the source image / video through a process of capturing, synthesizing, or generating an image / video. In this case, the generated source image / video may be transmitted to a feature extraction network and used as input data for extracting a feature / feature map.
[0072] The encoding unit (12) can encode the feature / feature map acquired by the feature acquisition unit (11). The encoding unit (12) can perform a series of procedures such as prediction, transformation, and quantization to increase encoding efficiency. The encoded data (encoded feature / feature map information) can be output in the form of a bitstream. A bitstream including the encoded feature / feature map information can be referred to as a VCM bitstream.
[0073] The transmission unit (13) can transmit feature / feature map information or data output in the form of a bitstream to a decoding device (20) via a digital storage medium or network in the form of a file or streaming. Here, the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) can include elements for generating a media file having a predetermined file format or elements for transmitting data via a broadcasting / communication network.
[0074] The decoding device (20) can obtain feature / feature map information from the encoding device (10) and restore the feature / feature map based on the obtained information.
[0075] The decryption device (20) may include a receiving unit (21) and a decryption unit (22).
[0076] The receiving unit (21) can receive a bitstream from the encoding device (10), obtain feature / feature map information from the received bitstream, and transmit it to the decoding unit (22).
[0077] The decoding unit (22) can decode the feature / feature map based on the acquired feature / feature map information. The decoding unit (22) can perform a series of procedures, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of the encoding unit (14), to increase decoding efficiency.
[0078] According to an embodiment, the decryption device (20) may further include a task analysis / rendering unit (23).
[0079] The task analysis / rendering unit (23) can perform task analysis based on the decrypted feature / feature map. In addition, the task analysis / rendering unit (23) can render the decrypted feature / feature map into a form suitable for task execution. Various machine-oriented tasks can be performed based on the task analysis results and the rendered feature / feature map.
[0080] Above, the VCM system can encode / decode features extracted from source images / videos according to user and / or machine requests, task objectives, and surrounding environments, and perform various machine-directed tasks based on the decoded features. The VCM system can also be implemented by extending / redesigning a video / image coding system, and can perform various encoding / decoding methods defined in the VCM standard.
[0081] VCM pipeline
[0082] FIG. 2 is a schematic diagram illustrating a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0083] Referring to FIG. 2, the VCM pipeline (200) may include a first pipeline (210) for encoding / decoding images / videos and a second pipeline (220) for encoding / decoding features / feature maps. In the present disclosure, the first pipeline (210) may be referred to as a video codec pipeline, and the second pipeline (220) may be referred to as a feature codec pipeline.
[0084] The first pipeline (210) may include a first stage (image / video encoder) (211) that encodes an input image / video and a second stage (image / video decoder) (212) that decodes the encoded image / video to generate a restored image / video. The restored image / video may be used for human viewing, i.e., for human vision.
[0085] The second pipeline (220) may include a third stage (feature extraction network) (221) for extracting features / feature maps from input images / videos, a fourth stage (VCM encoder) (222) for encoding the extracted features / feature maps, and a fifth stage (VCM decoder) (223) for decoding the encoded features / feature maps to generate restored features / feature maps. The restored features / feature maps may be used for machine (vision) tasks. Here, the machine (vision) task may refer to a task in which images / videos are consumed by a machine. The machine (vision) task may be applied to service scenarios such as, for example, surveillance, intelligent transportation, smart city, intelligent industry, and intelligent content. In some embodiments, the restored features / feature maps may also be used for human vision.
[0086] In some embodiments, the encoded feature / feature map in the fourth stage (222) may be transferred to the first stage (221) and used to encode an image / video. In this case, an additional bitstream may be generated based on the encoded feature / feature map, and the generated additional bitstream may be transferred to the second stage (222) and used to decode the image / video.
[0087] In some embodiments, the decrypted feature / feature map in the fifth stage (223) may be transferred to the second stage (222) and used to decrypt an image / video.
[0088] Although FIG. 2 illustrates a case where the VCM pipeline (200) includes a first pipeline (210) and a second pipeline (220), this is merely exemplary and embodiments of the present disclosure are not limited thereto. For example, the VCM pipeline (200) may include only the second pipeline (220), or the second pipeline (220) may be extended into multiple feature codec pipelines.
[0089] Meanwhile, in the first pipeline (210), the first stage (211) may be performed by an image / video encoder, and the second stage (212) may be performed by an image / video decoder. In addition, in the second pipeline (220), the third stage (221) may be performed by a VCM encoder (or a feature / feature map encoder), and the fourth stage (222) may be performed by a VCM decoder (or a feature / feature map decoder). The encoder / decoder structure will be described in detail below.
[0090] Encoder
[0091] FIG. 3 is a schematic diagram of an image / video encoder to which embodiments of the present disclosure can be applied.
[0092] Referring to FIG. 3, the image / video encoder (300) may include an image partitioner (310), a prediction unit (predictor) 320, a residual processor (residual processor) 330, an entropy encoder (entropy encoder) 340, an adder (adder) 350, a filter (filter) 360, and a memory (memory) 370. The prediction unit (320) may include an inter prediction unit (321) and an intra prediction unit (322). The residual processor (330) may include a transformer (transformer) 332, a quantizer (quantizer) 333, a dequantizer (dequantizer) 334, and an inverse transformer (inverse transformer) 335. The residual processing unit (330) may further include a subtractor (331). The addition unit (350) may be referred to as a reconstructor or a recontructed block generator. The image segmentation unit (310), the prediction unit (320), the residual processing unit (330), the entropy encoding unit (340), the addition unit (350), and the filtering unit (360) described above may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. In addition, the memory (370) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component described above may further include the memory (370) as an internal / external component.
[0093] The image segmentation unit (310) can segment an input image (or picture, frame) input to the image / video encoder (300) into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). The coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a Quad-tree binary-tree ternary-tree (QTBTTT) structure. For example, one coding unit may be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The image / video coding procedure according to the present disclosure may be performed based on a final coding unit that is no longer divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, etc., or, if necessary, the coding unit may be recursively divided into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0094] The term "unit" may be used interchangeably with terms such as "block" or "area" in some cases. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel.
[0095] The image / video encoder (300) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter-prediction unit (321) or an intra-prediction unit (322) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (332). In this case, as illustrated, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input image signal (original block, original sample array) within the image / video encoder (300) may be referred to as a subtraction unit (331). The prediction unit can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or CU unit. The prediction unit can generate various information regarding prediction, such as prediction mode information, and transmit it to the entropy encoding unit (340). The information regarding prediction can be encoded in the entropy encoding unit (340) and output in the form of a bitstream.
[0096] The intra prediction unit (322) can predict the current block by referring to samples in the current picture. At this time, the referenced samples may be located in the neighborhood of the current block or may be located away from it depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail in the prediction direction. However, this is only an example, and a greater or lesser number of directional prediction modes may be used depending on the settings. The intra prediction unit (322) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0097] The inter prediction unit (321) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. Temporal neighboring blocks may be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures including temporal neighboring blocks may be referred to as collocated pictures (colPic). For example, the inter prediction unit (321) may construct a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (321) may use the motion information of neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0098] The prediction unit (320) can generate a prediction signal based on various prediction methods. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index.
[0099] The prediction signal generated by the prediction unit (320) can be used to generate a reconstructed signal or a residual signal. The transformation unit (332) can apply a transformation technique to the residual signal to generate transform coefficients. For example, the transformation technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transformation obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transformation process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size.
[0100] The quantization unit (333) quantizes the transform coefficients and transmits them to the entropy encoding unit (340), and the entropy encoding unit (340) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be called residual information. The quantization unit (333) can rearrange the quantized transform coefficients in the form of a block into the form of a one-dimensional vector based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of a one-dimensional vector. The entropy encoding unit (340) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (340) may encode, together or separately, information necessary for image / video restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded image / video information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The image / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the image / video information may further include general constraint information. In addition, the image / video information may further include a method of generating and using the encoded information, a purpose, etc. In the present disclosure, information and / or syntax elements transmitted / signaled from an image / video encoder to an image / video decoder may be included in the image / video information.Image / video information can be encoded through the encoding process described above and included in a bitstream. The bitstream can be transmitted through a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits a signal output from an entropy encoding unit (340) and / or a storage unit (not shown) that stores the signal can be configured as an internal / external element of the image / video encoder (300), or the transmission unit can be included in the entropy encoding unit (340).
[0101] The quantized transform coefficients output from the quantization unit (333) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (334) and the inverse transform unit (335), a residual signal (residual block or residual samples) can be restored. The addition unit (350) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (321) or the intra prediction unit (322). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (350) may be called a restoration unit or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0102] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0103] The filtering unit (360) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (360) can apply various filtering methods to the restoration picture to generate a modified restoration picture and store the modified restoration picture in the memory (370), specifically, in the DPB of the memory (370). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (360) can generate various information regarding filtering and transmit the information to the entropy encoding unit (340). The information regarding filtering can be encoded by the entropy encoding unit (340) and output in the form of a bitstream.
[0104] The modified restored picture transmitted to the memory (370) can be used as a reference picture in the inter prediction unit (321). Through this, prediction mismatch in the encoder and decoder stages can be avoided, and encoding efficiency can also be improved.
[0105] The DPB of the memory (370) can store the modified restored picture to be used as a reference picture in the inter prediction unit (321). The memory (370) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (321) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (370) can store restored samples of restored blocks in the current picture and can transfer the stored restored samples to the intra prediction unit (322).
[0106] Meanwhile, the VCM encoder (or feature / feature map encoder) may have a structure basically identical / similar to the image / video encoder (300) described with reference to FIG. 3 in that it performs a series of procedures such as prediction, transformation, and quantization to encode a feature / feature map. However, the VCM encoder differs from the image / video encoder (300) in that it targets a feature / feature map for encoding, and accordingly, the name of each unit (or component) (e.g., image segmentation unit (310), etc.) and its specific operation contents may be different from the image / video encoder (300). The specific operation contents of the VCM encoder will be described in detail later.
[0107] Decoder
[0108] FIG. 4 is a schematic diagram of an image / video decoder to which embodiments of the present disclosure can be applied.
[0109] Referring to FIG. 4, the image / video decoder (400) may include an entropy decoder (410), a residual processor (420), a predictor (430), an adder (440), a filter (450), and a memory (460). The predictor (430) may include an inter-prediction unit (431) and an intra-prediction unit (432). The residual processor (420) may include a dequantizer (421) and an inverse transformer (422). The entropy decoding unit (410), residual processing unit (420), prediction unit (430), addition unit (440), and filtering unit (450) described above may be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. In addition, the memory (460) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (460) as an internal / external component.
[0110] When a bitstream including video / image information is input, the video / video decoder (400) can restore the video / image in accordance with the process in which the video / image information is processed in the video / video encoder (300) of FIG. 3. For example, the video / video decoder (400) can derive units / blocks based on block division-related information obtained from the bitstream. The video / video decoder (400) can perform decoding using a processing unit applied in the video / video encoder. Accordingly, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored video signal decoded and output by the video / video decoder (400) can be played back through a playback device.
[0111] The image / video decoder (400) can receive a signal output from the encoder (300) of FIG. 3 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (410). For example, the entropy decoding unit (410) can parse the bitstream to derive information (e.g., image / video information) necessary for image restoration (or picture restoration). The image / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the image / video information may further include general constraint information. In addition, the image / video information may include a method of generating and using decoded information, a purpose, etc. The image / video decoder (400) can decode a picture further based on information on the parameter set and / or the general constraint information. Signaling / received information and / or syntax elements can be decoded through a decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (410) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image restoration and the quantized values of transform coefficients for the residual. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (410) is provided to the prediction unit (inter prediction unit (432) and intra prediction unit (431)), and residual values on which entropy decoding is performed by the entropy decoding unit (410), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (420). The residual processing unit (420) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (410) can be provided to the filtering unit (450). Meanwhile, a receiving unit (not shown) that receives a signal output from an image / video encoder may be further configured as an internal / external element of the image / video decoder (400), or the receiving unit may be a component of an entropy decoding unit (410). Meanwhile, the image / video decoder according to the present disclosure may be called an image / video decoding device, and the image / video decoder may be divided into an information decoder (image / video information decoder) and a sample decoder (image / video sample decoder). In this case, the information decoder may include an entropy decoding unit (410), and the sample decoder may include at least one of an inverse quantization unit (321), an inverse transformation unit (322), an addition unit (440), a filtering unit (450), a memory (460), an inter prediction unit (432), and an intra prediction unit (431).
[0112] The inverse quantization unit (421) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (421) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the image / video encoder. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0113] In the inverse transform unit (422), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0114] The prediction unit (430) can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit (410), the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.
[0115] The prediction unit (420) can generate a prediction signal based on various prediction methods. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index can be signaled by including it in the image / video information.
[0116] The intra prediction unit (431) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit (431) can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0117] The inter prediction unit (432) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (432) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the mode of inter prediction for the current block.
[0118] The addition unit (440) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (including the inter prediction unit (432) and / or intra prediction unit (431)). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.
[0119] The addition unit (440) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture.
[0120] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0121] The filtering unit (450) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (450) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (460), specifically, the DPB of the memory (460). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0122] The (corrected) reconstructed picture stored in the DPB of the memory (460) can be used as a reference picture in the inter prediction unit (432). The memory (460) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (432) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (460) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (431).
[0123] Meanwhile, the VCM decoder (or feature / feature map decoder) may have a structure basically identical to / similar to the image / video decoder (400) described above with reference to FIG. 4 in that it performs a series of procedures such as prediction, inverse transformation, and inverse quantization to decode a feature / feature map. However, the VCM decoder differs from the image / video decoder (400) in that it targets a feature / feature map for decoding, and accordingly, the name of each unit (or component) (e.g., DPB, etc.) and its specific operation contents may be different from the image / video decoder (400). The operation of the VCM decoder may correspond to the operation of the VCM encoder, and the specific operation contents will be described in detail later.
[0124] Feature / feature map encoding procedure
[0125] FIG. 5 is a flowchart schematically illustrating a feature / feature map encoding procedure to which embodiments of the present disclosure may be applied.
[0126] Referring to FIG. 5, the feature / feature map encoding procedure may include a prediction procedure (S510), a residual processing procedure (S520), and an information encoding procedure (S530).
[0127] The prediction procedure (S510) can be performed by the prediction unit (320) described above with reference to FIG. 3.
[0128] Specifically, the intra prediction unit (322) can predict the current block (i.e., the set of feature elements to be currently encoded) by referring to feature elements within the current feature / feature map. Intra prediction can be performed based on the spatial similarity of feature elements constituting the feature / feature map. For example, feature elements included in the same region of interest (RoI) within an image / video can be estimated to have similar data distribution characteristics. Accordingly, the intra prediction unit (322) can predict the current block by referring to the restored feature elements within the region of interest including the current block. At this time, the referenced feature elements may be positioned adjacent to the current block or spaced apart from the current block depending on the prediction mode. Intra prediction modes for feature / feature map encoding can include a plurality of non-directional prediction modes and a plurality of directional prediction modes. Non-directional prediction modes may include prediction modes corresponding to, for example, the DC mode and the planar mode of an image / video encoding procedure. Furthermore, directional modes may include prediction modes corresponding to, for example, the 33 directional modes or the 65 directional modes of an image / video encoding procedure. However, this is merely an example, and the types and number of intra-prediction modes may be set / changed in various ways depending on the embodiment.
[0129] The inter prediction unit (321) can predict the current block based on a reference block (i.e., a set of referenced feature elements) specified by motion information on a reference feature / feature map. Inter prediction can be performed based on the temporal similarity of feature elements constituting the feature / feature map. For example, temporally consecutive features can have similar data distribution characteristics. Therefore, the inter prediction unit (321) can predict the current block by referring to the restored feature elements of features that are temporally adjacent to the current feature. At this time, the motion information for specifying the referenced feature elements can include a motion vector and a reference feature / feature map index. The motion information can further include information regarding the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing in the current feature / feature map and temporal neighboring blocks existing in the reference feature / feature map. The reference feature / feature map including the reference block and the reference feature / feature map including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be referred to as collocated reference blocks, and the reference feature / feature map including the temporal neighboring blocks may be referred to as collocated feature / feature maps. The inter prediction unit (321) may construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference feature / feature map index of the current block.Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit (321) can use the motion information of the surrounding blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the surrounding blocks can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference. In addition to the intra prediction and inter prediction described above, the prediction unit (320) can generate a prediction signal based on various prediction methods.
[0130] The prediction signal generated by the prediction unit (320) can be used to generate a residual signal (residual block, residual feature elements) (S520). The residual processing procedure (S520) can be performed by the residual processing unit (330) described above with reference to FIG. 3. In addition, (quantized) transform coefficients can be generated through a transformation and / or quantization procedure for the residual signal, and the entropy encoding unit (340) can encode information about the (quantized) transform coefficients as residual information in the bitstream (S530). In addition to the residual information, the entropy encoding unit (340) can encode information necessary for feature / feature map restoration, such as prediction information (e.g., prediction mode information, motion information, etc.), in the bitstream.
[0131] Meanwhile, the feature / feature map encoding procedure may further include a procedure (S530) for encoding information for feature / feature map restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, as well as a procedure for generating a restored feature / feature map for the current feature / feature map and a procedure (optional) for applying in-loop filtering to the restored feature / feature map.
[0132] The VCM encoder can derive (modified) residual feature(s) from the quantized transform coefficient(s) through inverse quantization and inverse transformation, and can generate a restored feature / feature map based on the predicted feature(s) and the (modified) residual feature(s) which are the outputs of step S510. The restored feature / feature map generated in this way can be identical to the restored feature / feature map generated by the VCM decoder. If an in-loop filtering procedure is performed on the restored feature / feature map, a modified restored feature / feature map can be generated through the in-loop filtering procedure on the restored feature / feature map. The modified restored feature / feature map can be stored in a decoded feature buffer (DFB) or memory, and can be used as a reference feature / feature map in a subsequent feature / feature map prediction procedure. In addition, (in-loop) filtering-related information (parameters) can be encoded and output in the form of a bitstream. Through the in-loop filtering procedure, noise that may occur during feature / feature map coding can be removed, and the performance of feature / feature map-based tasks can be improved. Furthermore, by performing the in-loop filtering procedure in both the encoder and decoder stages, the consistency of prediction results can be guaranteed, the reliability of feature / feature map coding can be improved, and the amount of data transmission for feature / feature map coding can be reduced.
[0133] Feature / feature map decoding procedure
[0134] FIG. 6 is a flowchart schematically illustrating a feature / feature map decoding procedure to which embodiments of the present disclosure may be applied.
[0135] Referring to FIG. 6, the feature / feature map decoding procedure may include an image / video information acquisition procedure (S610), a feature / feature map restoration procedure (S620 to S640), and an in-loop filtering procedure (S650) for the restored feature / feature map. The feature / feature map restoration procedure may be performed based on a prediction signal and a residual signal obtained through the inter / intra prediction (S620) and residual processing (S630), inverse quantization for quantized transform coefficients, and inverse transformation) processes described in the present disclosure. A modified restored feature / feature map may be generated through the in-loop filtering procedure for the restored feature / feature map, and the modified restored feature / feature map may be output as a decoded feature / feature map. The decoded feature / feature map is stored in the decoded feature buffer (DFB) or memory, and can be used as a reference feature / feature map in the inter prediction procedure during subsequent decoding of the feature / feature map. In some cases, the in-loop filtering procedure described above may be omitted. In this case, the reconstructed feature / feature map can be output as a decoded feature / feature map, and can be stored in the decoded feature buffer (DFB) or memory, and can be used as a reference feature / feature map in the inter prediction procedure during subsequent decoding of the feature / feature map.
[0136] Feature extraction methods and data distribution characteristics
[0137] Figure 7 is a diagram showing an example of a feature extraction method using a feature extraction network.
[0138] Referring to FIG. 7, a feature extraction network (700) can output a feature set (720) of a video source (710) by receiving a video source (710) as input and performing a feature extraction operation. The feature set (720) includes a plurality of features (C0, C1, ..., C) extracted from the video source (710). n ) can be included and can be expressed as a feature map. Each feature (C0, C1, ..., C n) contains multiple feature elements and may have different data distribution characteristics.
[0139] In FIG. 7, W, H, and C may represent the width, height, and number of channels of the video source (710), respectively. Here, the number of channels (C) of the video source (710) may be determined based on the image format of the video source (710). For example, if the video source (710) has an RGB image format, the number of channels (C) of the video source (710) may be 3.
[0140] In addition, W', H' and C' may represent the width, height and number of channels of the feature set (720), respectively. The number of channels (C') of the feature set (720) is the number of features (C0, C1, ..., C) extracted from the video source (710). n ) may be equal to the total number (n+1). In one example, the number of channels (C') of the feature set (720) may be greater than the number of channels (C) of the video source (710).
[0141] The properties (W', H', C') of the feature set (720) may vary depending on the properties (W, H, C) of the video source (710). For example, as the number of channels (C) of the video source (710) increases, the number of channels (C') of the feature set (720) may also increase. In addition, the properties (W', H', C') of the feature set (720) may vary depending on the type and properties of the feature extraction network (700). For example, when the feature extraction network (700) is implemented as an artificial neural network (e.g., CNN, DNN, etc.), each feature (C0, C1, ..., C n ) may also vary depending on the location of the layer outputting the feature set (720) properties (W', H', C').
[0142] The video source (710) and the feature set (720) may have different data distribution characteristics. For example, the video source (710) may generally consist of one (grayscale image) channel or three (RGB image) channels. The pixels included in the video source (710) may have the same integer value range for all channels and may have non-negative values. In addition, each pixel value may be evenly distributed within a predetermined integer value range. In contrast, the feature set (720) may consist of a variety of channels (e.g., 32, 64, 128, 256, 512, etc.) depending on the type of the feature extraction network (700) (e.g., CNN, DNN, etc.) and the layer location. The feature elements included in the feature set (720) may have different real value ranges for each channel and may also have negative values. Additionally, each feature element value can be concentrated in a specific area within a predetermined real number value range.
[0143] Fig. 8a is a diagram showing data distribution characteristics of a video source, and Fig. 8b is a diagram showing data distribution characteristics of a feature set.
[0144] First, referring to Fig. 8a, the video source is composed of three channels: R, G, and B, and each pixel value can have an integer value range from 0 to 255. In this case, the data type of the video source can be expressed as an 8-bit integer type.
[0145] In contrast, referring to Fig. 8b, the feature set consists of 64 channels (features), and each feature element value can have a real value range from -∞ to +∞. In this case, the data type of the feature set can be expressed as a 32-bit floating point type.
[0146] A feature set can have floating-point feature element values, and each channel (or feature) can have different data distribution characteristics. An example of the data distribution characteristics for each channel of a feature set is shown in Table 1.
[0147] [Table 1]
[0148]
[0149] Referring to Table 1, the feature set has a total of n+1 channels (C0, C1, ..., C n ) can be composed of. The mean (μ), standard deviation (σ), maximum (Max) and minimum (Min) values of the feature elements are provided for each channel (C0, C1, ..., C n ) may be different for each channel. For example, the mean (μ) of the feature elements included in channel 0 (C0) may be 10, the standard deviation (σ) may be 20, the maximum (Max) may be 90, and the minimum (Min) may be 60. In addition, the mean (μ) of the feature elements included in channel 1 (C1) may be 30, the standard deviation (σ) may be 10, the maximum (Max) may be 70.5, and the minimum (Min) may be -70.2. In addition, the mean (μ) of the feature elements included in channel n (Cn) may be 100, the standard deviation (σ) may be 5, the maximum (Max) may be 115.8, and the minimum (Min) may be 80.2.
[0150] Quantization of features / feature maps can be performed, for example, based on the different data distribution characteristics for each channel described above. Floating-point type feature / feature map data can be converted to integer type through quantization.
[0151] Meanwhile, due to spatiotemporal similarities between consecutive frames, feature sets and / or channels sequentially extracted from a video source may have identical / similar data distribution characteristics. An example of the data distribution characteristics of consecutive feature sets is shown in Table 2.
[0152] [Table 2]
[0153]
[0154] In Table 2, f F0 refers to the first feature set extracted from frame 0 (F0), and f F1 refers to the second feature set extracted from frame 1 (F1), and f F2 refers to the third feature set extracted from frame 2 (F2).
[0155] Referring to Table 2, the first to third consecutive feature sets (f F0 , f F1 , f F2 ) can have the same / similar mean (μ), standard deviation (σ), maximum (Max) and minimum (Min) values.
[0156] Additionally, due to spatiotemporal similarities between consecutive frames, corresponding channels within feature sets sequentially extracted from a video source may have identical or similar data distribution characteristics. An example of the data distribution characteristics of corresponding channels within consecutive feature sets is shown in Table 3.
[0157] [Table 3]
[0158]
[0159] In Table 3, f F0C0 refers to the first channel in the first feature set extracted from frame 0 (F0), and f F1C0 refers to the first channel in the second feature set extracted from frame 1 (F1).
[0160] Referring to Table 3, the first channel (f) within the first feature set F0C0 ) and the first channel (f) in the second feature set corresponding to it F1C0 ) can have the same / similar mean (μ), standard deviation (σ), maximum value (Max), and minimum value (Min).
[0161] Prediction for features / feature maps can be performed based on, for example, the similarity of data distribution characteristics between the feature sets or channels described above.
[0162] Transformation / Inverse transformation
[0163] As described above, the image / video encoder can derive a residual block based on a predicted block, and can apply transformation and quantization to the derived residual block to derive quantized transform coefficients. Similarly, the VCM encoder can derive a residual block based on a predicted block (or feature elements), and can apply transformation and quantization to the derived residual block to derive quantized transform coefficients for a feature / feature map. Information about the quantized transform coefficients (residual information) can be included in the residual coding syntax, encoded, and then output in the form of a bitstream.
[0164] A decoder (video decoder and VCM decoder) can obtain information about quantized transform coefficients (residual information) from a bitstream and decode the information to derive quantized transform coefficients. In addition, the decoder can derive a residual block through inverse quantization / inverse transformation based on the quantized transform coefficients. As described above, at least one of quantization / inverse quantization and / or transform / inverse transformation can be omitted. If transform / inverse transformation is omitted, the transform coefficient can be referred to as a coefficient or a residual coefficient, or can still be referred to as a transform coefficient for consistency of expression. Whether transform / inverse transformation is omitted can be signaled based on a predetermined flag, for example, transform_skip_flag. Meanwhile, in VCM, whether transform / inverse transformation is omitted can be determined on a feature / feature map basis or on a channel basis.
[0165] Transformation / inverse transformation can be performed based on transformation kernel(s). For example, according to embodiments, a multiple transform selection (MTS) scheme can be applied. In this case, some of a plurality of transformation kernel sets can be selected and applied to the current block. The transformation kernels can be referred to by various terms such as transformation matrix, transformation type, etc. For example, a transformation kernel set can represent a combination of vertical transformation kernels (vertical transformation kernels) and horizontal transformation kernels (horizontal transformation kernels).
[0166] MTS index information, e.g., mts_idx, may be signaled to indicate a set of transformation kernels for the current block or feature / feature map. An example of MTS index information is shown in Table 4.
[0167] [Table 4]
[0168]
[0169] In Table 4, trTypeHor can represent a horizontal transform kernel, and trTypeVer can represent a vertical transform kernel. A trTypeHor / trTypeVer value of 0 can represent DCT2, a trTypeHor / trTypeVer value of 1 can represent DST7, and a trTypeHor / trTypeVer value of 2 can represent DCT8. However, this is just an example, and other values may be mapped to other DCT / DST or transform kernels by convention.
[0170] In the present disclosure, the MTS-based transform / inverse transform is applied as a primary transform / inverse-transform, and a secondary transform / inverse-transform may be further applied. In this case, the transform at the encoder stage may proceed in the order of the primary transform and the secondary transform, and the transform at the decoder stage may proceed in the order of the secondary inverse transform and the primary inverse transform. In some embodiments, the secondary transform may be applied only to coefficients in the upper left low-frequency region of the coefficient block to which the primary transform is applied. Such a secondary transform may be referred to as a low frequency non-separable transform (LFNST). In addition, in some embodiments, the number of coefficients may be reduced as a result of performing the secondary transform. Such a secondary transform may be referred to as a reduced secondary transform (RST).
[0171] FIG. 9 is a diagram for explaining a transformation and inverse transformation process to which embodiments of the present disclosure may be applied. In FIG. 9, the transformation unit (910) may correspond to the transformation unit (332) of FIG. 3 , and the inverse transformation unit (920) may correspond to the inverse transformation unit (335) of FIG. 3 or the inverse transformation unit (422) of FIG. 4 .
[0172] Referring to FIG. 9, the conversion unit (910) may include a first conversion unit (911) and a second conversion unit (912).
[0173] The primary transform unit (911) can apply a primary transform to the residual data (A) (or residual feature elements) to generate (primary) transform coefficients (B). In the present disclosure, the primary transform may be referred to as a core transform. The primary transform may be performed based on an MTS scheme. When the existing MTS is applied, a transformation from the spatial domain to the frequency domain may be applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, etc. to generate transform coefficients (or primary transform coefficients). Here, DCT type 2, DST type 7, DCT type 8, etc. may be referred to as a transform type, a transform kernel, or a transform core. However, this is merely an example, and the primary transform may also be applied to MTS kernels having different configurations from the above-described MTS kernels.
[0174] The secondary transform unit (912) can apply a secondary transform to the (primary) transform coefficients (B) to generate (secondary) transform coefficients (C). The secondary transform may be a non-separable transform, such as the aforementioned LFNST or RST. For example, when a 4x4 block is used, the non-separable secondary transform may be performed as follows.
[0175] A 4x4 input block X can be expressed as in Equation 1 below.
[0176] [Formula 1]
[0177]
[0178] When the input block X is represented in vector form, the vector X can be expressed as in the following equation 2.
[0179] [Formula 2]
[0180]
[0181] In this case, the non-separable second-order transform can be calculated as shown in Equation 3 below.
[0182] [Formula 3]
[0183]
[0184] Here, vector F represents a transformation coefficient vector, T represents a 16x16 non-separable transformation matrix, and · represents the multiplication of a matrix and a vector.
[0185] A 16x1 transform coefficient vector F can be derived through the above mathematical formula, and the vector F can be reconstructed into 4x4 blocks according to the scan order (e.g., horizontal, vertical, diagonal, or a predetermined / stored scan order).
[0186] Next, the inverse transformation unit (820) may include a (inverse) secondary transformation unit (821) and a (inverse) primary transformation unit (822).
[0187] The (inverse) secondary transform unit (921) can apply an (inverse) secondary transform to the inverse quantized (secondary) transform coefficients (C') to generate (first) (inverse) transform coefficients (B'). Here, the (inverse) secondary transform may correspond to the reverse process of the secondary transform performed by the transform unit (910). The (inverse) secondary transform may be applied to the upper left corner of the coefficient block according to an embodiment.
[0188] The (inverse) first transformation unit (922) can generate residual samples (A') by applying the (inverse) first transformation to the (inverse) transformation coefficients (B').
[0189] Quantization / dequantization
[0190] As described above, the quantization unit (333) of the image / video encoder (300) can apply quantization to transform coefficients to derive quantized transform coefficients, and the inverse quantization unit (334) of the image / video encoder (300) or the inverse quantization unit (421) of the image / video decoder (400) can apply inverse quantization to the quantized transform coefficients to derive transform coefficients. Similarly, the VCM encoder can apply quantization to transform coefficients to derive quantized transform coefficients, and the inverse quantization unit of the VCM encoder or the inverse quantization unit of the VCM decoder can apply inverse quantization to the quantized transform coefficients to derive transform coefficients. In general, in feature / feature map coding, the quantization rate can be changed, and the compression rate can be adjusted using the changed quantization rate. From an implementation perspective, a quantization parameter (QP) can be used instead of directly using the quantization rate in consideration of complexity. For example, quantization parameters with integer values from 0 to 63 can be used, and each quantization parameter value can correspond to an actual quantization rate. The quantization parameter (QPY) for the luma component (luma sample) and the quantization parameter (QPC) for the chroma component (chroma sample) can be set differently.
[0191] The quantization process takes a transform coefficient (C) as input, divides it by a quantization rate (Qstep), and obtains a quantized transform coefficient (C`) based on this. In this case, considering the computational complexity, the quantization rate can be multiplied by a scale to make it into an integer, and a shift operation can be performed by a value corresponding to the scale value. The quantization scale can be derived based on the product of the quantization rate and the scale value. In other words, the quantization scale can be derived according to the QP. The quantized transform coefficient (C`) can also be derived based on this by applying the quantization scale to the transform coefficient (C).
[0192] The inverse quantization process is the reverse process of the quantization process. By multiplying the quantized transform coefficient (C`) by the quantization rate (Qstep), a restored transform coefficient (C``) can be obtained based on this. In this case, a level scale can be derived according to the quantization parameter, and by applying the level scale to the quantized transform coefficient (C`), a restored transform coefficient (C``) can be derived based on this. The restored transform coefficient (C``) may be somewhat different from the original transform coefficient (C) due to loss in the transformation and / or quantization process. Therefore, inverse quantization is performed in the encoder in the same way as in the decoder.
[0193] Meanwhile, an adaptive frequency-based weighted quantization technique that adjusts the quantization strength according to frequency may be applied. The adaptive frequency-based weighted quantization technique is a method of applying different quantization strengths to different frequencies. The adaptive frequency-based weighted quantization can apply different quantization strengths to each frequency using a predefined quantization scaling matrix. That is, the quantization / dequantization process described above may be performed further based on the quantization scaling matrix. For example, a different quantization scaling matrix may be used depending on the size of the current block and / or whether the prediction mode applied to the current block to generate the residual signal of the current block is inter-prediction or intra-prediction. The quantization scaling matrix may be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix may be predefined. In addition, for frequency-adaptive scaling, frequency-based quantization scale information for the quantization scaling matrix may be configured / encoded in the encoder and signaled to the decoder. The frequency-specific quantization scale information may be referred to as quantization scaling information. The frequency-specific quantization scale information may include scaling list data (scaling_list_data). A (modified) quantization scaling matrix may be derived based on the scaling list data. In addition, the frequency-specific quantization scale information may include a presence flag (present flag) information indicating whether the scaling list data exists. Alternatively, if the scaling list data is signaled at a higher level (e.g., sequence level, feature set group level, etc.), information indicating whether the scaling list data is modified at a lower level (e.g., feature set level, channel level, etc.) may be further included.
[0194] Feature quantization / dequantization can be performed based on a predetermined quantization group. Specifically, multiple quantization intervals can be set based on the data distribution characteristics of the feature set. The set quantization intervals have different data distribution ranges and can be defined as a single quantization group. In addition, feature quantization / dequantization operations can be performed by converting the data distribution of each channel within the feature set into the data distribution range of one of the quantization intervals.
[0195] FIG. 10 is a diagram illustrating an example of a quantization group according to one embodiment of the present disclosure.
[0196] Referring to FIG. 10, a quantization group may include four quantization intervals (A, B, C, D) having different data distribution ranges. Each quantization interval (A, B, C, D) within a quantization group is defined using a minimum value and a maximum value, and may be set based on the data distribution characteristics of the current feature / feature map. For example, quantization interval A may be set to [-1, 3], quantization interval B may be set to [0, 2], quantization interval C may be set to [-2, 4], and quantization interval D may be set to [-2, 1].
[0197] The VCM encoder can quantize each channel (or feature) in a feature set based on preset quantization intervals (A, B, C, D) without having to separately calculate the maximum and minimum values of feature elements. For example, channel 1 can be quantized based on quantization interval A having the most similar data distribution range. In addition, channel 2 can be quantized based on quantization interval B having the most similar data distribution range. In addition, channel 3 can be quantized based on quantization interval C having the most similar data distribution range. In addition, channel 4 can be quantized based on quantization interval D having the most similar data distribution range. In this case, the VCM encoder can encode / signal the number, minimum value, and maximum value of quantization intervals as feature quantization-related information. Additionally, the VCM encoder can signal quantization interval index information, which represents the quantization interval used for encoding the current channel (or feature), as feature quantization-related information. Accordingly, there is no need to signal the number of quantization bits for each channel and the maximum and minimum values of feature elements, which can reduce the amount of transmission bits and further improve encoding / signaling efficiency.
[0198] The VCM decoder can form a quantization group identical to the quantization group formed by the VCM encoder based on the number of quantization intervals received from the VCM encoder and the minimum and maximum values of each interval. Alternatively, the VCM decoder can form a quantization group based on the data distribution characteristics of the restored feature sets. In addition, the VCM decoder can dequantize the current feature based on the quantization interval identified by the quantization interval index information received from the VCM encoder.
[0199] An example of a feature quantization operation based on a quantization group is as shown in Equations 4 and 5.
[0200] [Formula 4]
[0201]
[0202] [Formula 5]
[0203]
[0204] In Equation 4, Fn is the feature set (Fset RxC ) can mean the nth channel (n is an integer greater than or equal to 1). Here, R can mean the width of the feature set, and C can mean the height of the feature set. In addition, max(Fn) can mean the maximum value of the feature elements in the nth channel (Fn), and min(Fn) can mean the minimum value of the feature elements in the nth channel (Fn).
[0205] Referring to Equation 4, the nth channel (Fn) can be normalized to have feature element values between 0 and 1 based on the maximum and minimum values of the feature elements within the channel.
[0206] In formula 5, Fn norm is a feature set (Fset) RxC ) can mean the normalized nth channel of the feature set. Here, R can mean the width of the feature set, and C can mean the height of the feature set. In addition, max(F Gm ) is the quantization interval (F) applied to the nth channel (Fn). Gm ) means the maximum value of, and min(F Gm ) is the quantization interval (F) applied to the nth channel (Fn). Gm ) can mean the minimum value.
[0207] Referring to Equation 5, the normalized nth channel (Fn norm ) is the quantization interval (F) applied to the nth channel (Fn). Gm ) of the maximum value (max(F Gm )) and minimum (min(F Gm )) can be quantized based on.
[0208] Meanwhile, quantization-related information may be encoded in the bitstream for inverse quantization. Depending on the embodiment, the quantization-related information may include global quantization information, activation function information, quantization bit count, and quantization interval index information. The global quantization information may indicate whether the quantization bit count is set for each quantization interval or is set to be the same for all quantization intervals. The activation function information may indicate the type of activation function applied to the current feature. For example, the activation function information may indicate whether the activation function applied to the current feature is a first activation function that must signal both the maximum and minimum values of each quantization interval or a second activation function that only needs to signal the maximum value of each quantization interval. The quantization bit count may be set for each quantization interval based on the global quantization information described above, or may be set to be the same for all quantization intervals. If the quantization bit count is set to be the same for all quantization intervals, the quantization bit count may be signaled only once for all quantization intervals.
[0209] In addition, the quantization-related information may further include the number of quantization intervals within the quantization group and the minimum and maximum values of each quantization interval. In this case, the minimum and maximum values of each quantization interval may be adaptively signaled based on the type of the activation function. For example, if the activation function applied to the current feature is the first activation function (e.g., Leaky ReLU), both the minimum and maximum values of each quantization interval may be signaled. In contrast, if the activation function applied to the current feature is the second activation function (e.g., ReLU), the minimum value of each quantization interval is estimated to be 0, and only the maximum value of each quantization interval may be signaled.
[0210] Entropy coding
[0211] As described above, part or all of the feature / feature map information may be encoded by the entropy encoding unit (340), and part or all of the feature / feature map information may be decoded by the entropy decoding unit (410). In this case, the feature / feature map information may be encoded / decoded in units of syntax elements, similar to image / video information. In the present disclosure, encoding / decoding information may include encoding / decoding by the method described in this paragraph.
[0212] Figure 11 shows a block diagram of CABAC for encoding one syntax element.
[0213] Referring to Fig. 11, the encoding process of CABAC can first convert the input signal into a binary value through binarization if the input signal is a syntax element rather than a binary value. If the input signal is already a binary value, it can be bypassed without going through binarization. Here, each binary 0 or 1 that constitutes the binary value can be called a bin. For example, if the binary string (bin string) after binarization is 110, 1, 1, and 0 can each be called a bin. The bin(s) for one syntax element can represent the value of the corresponding syntax element.
[0214] Binarized bins can be input to a regular coding engine or a bypass coding engine. The regular coding engine can assign a context model that reflects the probability value for each bin and encode the bin based on the assigned context model. The regular coding engine can update the probability model for each bin after coding. Bins coded in this way are called context-coded bins. The bypass coding engine can omit the process of estimating the probability of each input bin and updating the probability model applied to the bin after coding. Instead of assigning a context, the engine can apply a uniform probability distribution (e.g., 50:50) to the input bins, thereby improving coding speed. Bins coded in this way are called bypass bins. The context model can be assigned and updated for each bin subject to context coding (regular coding), and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, the context index (ctxidx) that points to the context model for each of the regular coded bins can be derived as the sum of ctxInc (context index increment) and ctxIdxOffset (context index offset). Here, ctxInc can be derived differently for each bin. ctxIdxOffset can be represented as the lowest value of ctxIdx. ctxIdxOffset can generally be determined according to the slice type, and the context model for one syntax element in the slice can be distinguished / derived based on ctxInc.
[0215] In the entropy encoding process, it is possible to determine whether encoding will be performed through the regular coding engine or the bypass coding engine, and to switch the coding path. Entropy decoding can be performed in reverse order, similar to the entropy encoding process.
[0216] The entropy coding described above can be performed, for example, as shown in FIGS. 12 and 13.
[0217] Figures 12 and 13 are drawings for explaining the entropy encoding procedure.
[0218] Referring to FIGS. 12 and 13, an encoder (entropy encoding unit) may perform an entropy coding procedure on feature / feature map information. The feature / feature map information may include prediction-related information (e.g., inter / intra prediction distinction information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Such entropy coding may be performed on a syntax element basis. The entropy encoding unit may be the entropy encoding unit (340) of the image / video encoder (300) of FIG. 3 described above.
[0219] The encoder can perform binarization on the target syntax element (S1210). Here, binarization can be based on various binarization methods, such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element can be predefined. The binarization procedure can be performed by the binarization unit (1301) within the entropy encoding unit (1300).
[0220] The encoder can perform entropy encoding on the target syntax element (S1220). The encoder can encode the empty string of the target syntax element based on a regular coding-based (context-based) or bypass coding-based entropy coding technique such as CABAC (context-adaptive arithmetic coding) or CAVLC (context-adaptive variable length coding), and the output thereof can be included in the bitstream. The entropy encoding procedure can be performed by the entropy encoding processing unit (1302) within the entropy encoding unit (1300). As described above, the output bitstream can be transmitted to the decoder via a (digital) storage medium or a network.
[0221] Figures 14 and 15 are drawings for explaining the entropy decoding procedure.
[0222] Referring to FIGS. 14 and 15, a decoder (entropy decoding unit) can decode encoded feature / feature map information. The feature / feature map information can include prediction-related information (e.g., inter / intra prediction distinction information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or can include various syntax elements related thereto. Entropy coding can be performed in units of syntax elements. The entropy decoding unit can be the entropy decoding unit (410) of the image / video decoder (400) of FIG. 4 described above.
[0223] The decoder can perform binarization on the target syntax element (S1410). Here, the binarization can be based on various binarization methods such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element can be predefined. The decoder can derive available empty strings (empty string candidates) for available values of the target syntax element through the binarization process. The binarization process can be performed by the binarization unit (1501) within the entropy decoding unit (1500).
[0224] The decoder can perform entropy decoding on the target syntax element (S1420). The decoder can sequentially decode and parse each bin for the target syntax element from the input bit(s) in the bitstream, and compare the derived bin string with the available bin strings for the corresponding syntax element. If the derived bin string is equal to one of the available bin strings, the value corresponding to the bin string can be derived as the value of the corresponding syntax element. If not, the next bit in the bitstream can be further parsed and the above-described procedure can be performed again. Through this process, it is possible to signal specific information (specific syntax element) using variable-length bits without using start bits or end bits in the bitstream. This allows relatively fewer bits to be allocated to low values, thereby improving overall coding efficiency.
[0225] The decoder can decode each bin within a bin string from a bitstream based on a context-based or bypass-based entropy coding technique such as CABAC or CAVLC. The entropy decoding procedure can be performed by an entropy decoding processing unit (1502) within an entropy decoding unit (1500). As described above, the bitstream can include various information for feature / feature map decoding. As described above, the bitstream can be transmitted to a decoding device via a (digital) storage medium or a network.
[0226] In the present disclosure, a table including syntax elements (a syntax table) may be used to indicate signaling of information from an encoder to a decoder. The order of the syntax elements in the table including the syntax elements used in the present disclosure may indicate a parsing order of the syntax elements from a bitstream. An encoder may construct and encode a syntax table such that the syntax elements can be parsed by a decoder in the corresponding parsing order, and the decoder may parse and decode the syntax elements of the corresponding syntax table from a bitstream in the corresponding parsing order to obtain the values of the syntax elements.
[0227] Coding hierarchy and structure
[0228] Figure 16 illustrates an example of a VCM hierarchy.
[0229] Referring to FIG. 16, the VCM hierarchy may be composed of a feature extraction layer (1610), a neural network (feature) abstraction layer (1620), and a feature coding layer (1630).
[0230] The feature extraction layer (1610) refers to a layer for extracting features from an input source, and may also include the results of the extraction.
[0231] The feature coding layer (1630) refers to a layer for compressing extracted features, and may also include the results of compression.
[0232] The neural network (feature) abstraction layer (1620) can abstract information generated by the feature extraction layer (1610) (e.g., information about extracted features / feature maps) and pass it to the feature coding layer (1630). The neural network abstraction layer (1720) can hide the internals of the feature extraction layer (1610) through information abstraction and provide a consistent feature interface function. Accordingly, even if the compression target changes due to a change in the tool (e.g., CNN, DNN, etc.), the feature coding layer (1630) can perform a consistent feature coding procedure. In the present disclosure, the neural network abstraction layer (NNAL) may also be referred to as a feature abstraction layer.
[0233] The interface between the feature extraction layer (1610) and the neural network abstraction layer (1620) and the interface between the feature coding layer (1630) and the neural network abstraction layer (1620) can be defined in advance, and the operation in the neural network abstraction layer (1620) can be set to be changed later.
[0234] Figure 17 is a diagram showing an example of a bitstream composed of encoded abstract features and NNAL information.
[0235] A bitstream configured as illustrated in Fig. 17 may be referred to as an NNAL (Neural Network Abstraction Layer) unit. An NNAL unit may be an independent feature restoration unit. The input features for an NNAL unit may be extracted from the same layer within a neural network. Accordingly, the input features for an NNAL unit may be forced to have the same characteristics. For example, the same feature extraction method may be applied to the input features for an NNAL unit.
[0236] An NNAL unit may include an NNAL unit header and an NNAL unit payload. The NNAL unit header may include all information necessary to utilize encoded features for a task. The NNAL unit payload may include abstracted feature information. The NNAL unit payload may include a group header and group data. The group header may include configuration information of feature group data, such as information about the temporal order, number, and common properties of feature channels constituting the feature group. A feature channel may mean an encoded feature unit. The group data may include a plurality of feature channels and a coding indicator, and each feature channel may include type information, prediction information, side information, and residual information (data). In this case, the type information may indicate an encoding method code, and the prediction information may indicate a prediction method. Additionally, side information may represent additional information required for decoding (e.g., entropy coding, quantization-related information, etc.), and residual information may include information about encoded feature elements (i.e., a set of feature value information).
[0237] Example
[0238] Embodiments described in the present disclosure relate to techniques related to coding layers and structures, particularly to methods for optimizing encoding applied according to the purpose of a user or a machine and to methods for expressing suitable uses and properties of an encoded bitstream, including content related to privacy protection properties for users and machines (e.g., de-identification).
[0239] The need for privacy protection processing of images (e.g., de-identification) is as follows.
[0240] 1. Privacy: Images used for AI learning and inferencing may contain sensitive or personally identifiable information.
[0241] 2. Legal Compliance and Ethical Use: Many countries and regions have strict regulations regarding the protection of personal information. De-identification can help you comply with these laws.
[0242] 3. Data Sharing: Anonymization allows you to share data safely without exposing sensitive information.
[0243] Figure 18 illustrates an example of encoding optimization and the use of encoded bitstreams that include privacy processing. In this example, optimization can remove unnecessary information or enhance necessary information during image acquisition. In certain cases, personal information can be unnecessary. In these cases, optimization can protect and remove personal information. This ensures privacy and legal compliance, as sensitive personal information is removed before the image is distributed and used, enabling ethical use and distribution. While the need for privacy processing is clear, the encoded bitstream presents a problem: information related to privacy processing is not defined. The following issues can arise when this information is not defined:
[0244] 1) Incorrect usage possible
[0245] To protect your privacy, personally identifiable information may have been removed or altered. Using these images for personal identification purposes may result in misidentification.
[0246] 2) Determining the legality of using the received video
[0247] If you do not know whether the data you receive or collect contains personal information and whether privacy protection applies, it is difficult to determine the legality of its use, storage, and distribution.
[0248] 3) Management difficulties
[0249] If the bitstream itself does not contain information about the properties of the information, and a problem occurs in the management system, it may be difficult to identify and manage the properties of the bitstream.
[0250] Example 1
[0251] Since the introduction of the General Data Protection Regulation (GDPR) and other similar laws worldwide to protect personal information, storing and using data containing unauthorized personal information, or personally identifiable information (PII), such as faces or license plates, has become legally problematic. To minimize legal issues, the nature of the data being used—whether it contains PII (e.g., video or images)—must be defined. Furthermore, if privacy protection measures are applied to obscure PII, the methods used must be clearly defined.
[0252] Figure 19 illustrates an example of defining data attributes and privacy information. Data attributes and privacy can indicate whether personal information is included. If personal information is included, information indicating the type and nature of the personal information can be defined. Furthermore, information indicating whether privacy protection is applied can be defined. If privacy protection is applied, information regarding the application method can be defined. This can be information regarding the application method, but it can also be information regarding the properties of data that have been changed due to privacy protection. For example, applying privacy protection can change the properties of data from data containing personally identifiable information to pseudonymized data, and information regarding the changed properties is defined.
[0253] Tables 5 to 9 show examples of how to express data attributes and privacy information (Fig. 19) and define related information in privacy_protection_info.
[0254] [Table 5]
[0255]
[0256] [Table 6]
[0257]
[0258] [Table 7]
[0259]
[0260] [Table 8]
[0261]
[0262] [Table 9]
[0263]
[0264] privacy_protection_info may be defined to include only a portion of the information defined in Fig. 19. privacy_protection_info may be defined within an encoded video / image bitstream, and the definition location may be multiple locations within the bitstream as needed. For example, it may be defined in the NAL (Network Abstraction Layer) Header, Sequence parameter set, Picture parameter set, SEI (Supplemental enhancement information), etc. The definition location is not limited to the locations exemplified here, and may be any high-level parameter set.
[0265] Table 5 shows an example that only defines whether privacy protection is applied and how it is protected in privacy_protection_info.
[0266] privacy_protection_cancel_flag can indicate that the persistence of a previously applied privacy method is canceled if its value is 1. If its value is 0, it can indicate that privacy_protection_persistence_flag and privacy_protection_type are defined subsequently, and that the privacy protection method and scope identified by the privacy_protection_persistence_flag and privacy_protection_type defined subsequently can be applied.
[0267] The privacy_protection_persistence_flag can indicate the persistence of the privacy protection method indicated by privacy_protection_type. A value of 0 indicates that the privacy protection method identified by privacy_protection_type may only be applied to the current picture. A value of 1 indicates that the privacy protection method identified by privacy_protection_type may be applied to the current picture and all subsequent pictures.
[0268] privacy_protection_type is information to express the properties of the applied privacy protection.
[0269] [Table 10]
[0270]
[0271] [Table 11]
[0272]
[0273] The example in Table 10 illustrates the applied method in a lookup table format. Using the value of privacy_protection_type in the example in Table 10, you can identify whether the applied method is blurring, masking, or replacing.
[0274] Table 11 shows how the method information defined in Table 10 is expressed as a result of privacy_protection_type and bitmask. Table 12 shows an example, and the value of privacy_protection_type may be defined differently from the value in Table 10.
[0275] [Table 12]
[0276]
[0277] In this case, it is possible to express multiple properties. blurring_Flag, replacing_Flag, masking_Flag, adversarial_attack_Flag, and anonymization_Flag are information indicating whether the privacy protection methods, which are blurring, replacing, masking, adversarial attack, and anonymization properties, are applied, respectively. If the value is 1, it means that the method was applied, and if the value is 0, it means that the method was not applied.
[0278] The example in Table 13 illustrates the attributes of data that have changed due to the applied privacy protection. In some cases, information about the attributes of data that have changed due to the application of privacy protection may be more valuable than information about the applied privacy protection method.
[0279] [Table 13]
[0280]
[0281] [Table 14]
[0282]
[0283] Table 13 shows examples of cases where personal information attributes of data have been pseudonymized, de-identified, or removed due to privacy protection.
[0284] Table 14 shows how the method information defined in Table 13 is expressed as a result of privacy_protection_type and bitmask. Table 15 shows an example, and the value of privacy_protection_type can be defined differently from the value in Table 13.
[0285] [Table 15]
[0286]
[0287] Pseudonymized_data_Flag, anonymized_data_Flag, Removed_PII_flag, and Identified_data_flag are information indicating whether the personal information attributes of data have been changed to pseudonymized data, anonymized data, data with PII removed, or identified data due to the application of a personal information protection method, respectively. If the value is 1, it means that the attribute has been changed, and if the value is 0, it means that the attribute has not been changed.
[0288] The privacy protection methods defined in Tables 10, 11, 13, 14, 16, and 17 are examples, and other methods may also be defined.
[0289] Protected_info_idc represents information for identifying protected information.
[0290] [Table 16]
[0291]
[0292] In Table 16, if the Protected_info_idc value is 00, it means that information that can identify a person, such as a face, is protected. If the Protected_info_idc value is 01, it means that information that can identify a means of transportation, such as a license plate, is protected. If the Protected_info_idc value is 10, it means that information that can infer a location, such as a business name or letters or pictures on a sign, is protected. If the Protected_info_idc value is 11, it means that information that can infer an object, such as letters or pictures of a product or object, is protected. It is clear that additional protected information can be defined.
[0293] [Table 17]
[0294]
[0295] Table 17 provides an example of how to identify multiple protected pieces of information using a bitmask. In this case, the Protected_info_idc value may be defined differently than in Table 16.
[0296] In Table 18, Protected_personal_identifiable_data_Flag, Protected_vehicle_identifiable_data_Flag, Protected_location_identifiable_data_Flag, and Protected_object_identifiable_data_Flag represent the protection of information that can identify a person, such as a face; the protection of information that can identify a vehicle, such as a license plate; the protection of information that can identify a location, such as a business name or a sign, or the protection of information that can identify a product or object, such as a product or object. If the value is 1, it means that protection has been applied to the information, and if the value is 0, it means that protection has not been applied to the information.
[0297] [Table 18]
[0298]
[0299] Table 6 above shows an example that defines the purpose and target of personal information protection as well as whether and how personal information protection is applied as defined in Table 5 in privacy_protection_info.
[0300] The prevent_human_perception_flag flag, if set to 1, indicates that the purpose of the applied privacy protection method is to prevent human identification of personal information. If set to 0, it indicates that the purpose of the applied privacy protection method is not to prevent human identification of personal information.
[0301] The prevent_machine_perception_flag, if set to 1, indicates that the purpose of the applied privacy protection method is to prevent machine or AI-based identification of personal information. A value of 0 indicates that the purpose of the applied privacy protection method is not to prevent machine or AI-based identification of personal information.
[0302] Tables 7 and 8 above show examples in which compliant_regulation_info, which is information indicating what regulations are to be met through personal information protection, is defined in Tables 5 and 6, respectively.
[0303] compliant_regulation_info represents information indicating the regulations to be met through personal information protection. Tables 7 and 8 above show examples of compliant_regulation_info, which represents the regulations to be met through personal information protection, defined in Tables 5 and 6, respectively.
[0304] In the example in Table 19, the regulations to be met are defined in a lookup table format based on the value of compliant_regulation_info.
[0305] [Table 19]
[0306]
[0307] In this case, the value of compliant_regulation_info defines information about one targeted regulatory entity.
[0308] Table 20 represents the target regulation defined in Table 19 as a result of compliant_regulation_info and bitmask.
[0309] [Table 20]
[0310]
[0311] In this case, compliant_regulation_info can express multiple target regulations. Table 21 shows an example. compliant_GDPR_Flag, compliant_CCPA_Flag, compliant_PIPEDA_Flag, and compliant_PIPL_Flag are information indicating whether the regulations that are met (or intended to be met) by applying the personal information protection method are GDPR, CCPA, PIPEDA, and PIPL, respectively. A value of 1 means that the regulation is met, and a value of 0 means that the regulation is not met.
[0312] [Table 21]
[0313]
[0314] It is clear that the regulations defined in Tables 19 and 20 are examples and that other regulations can also be defined.
[0315] Second Example
[0316] The first embodiment presented an example that solely defined information about the application and purpose of personal information protection.
[0317] The second embodiment defines the application and purpose of personal information protection as one of the optimization methods, as shown in Table 29 below. That is, a personal information protection-related optimization type is defined in optimization_type.
[0318] Tables 22 to 24 show examples of how to express optimization objectives, properties, and scope of application.
[0319] [Table 22]
[0320]
[0321] [Table 23]
[0322]
[0323] [Table 24]
[0324]
[0325] optimization_cancel_flag can indicate that the persistence of a previously applied optimization is canceled if its value is 1. If its value is 0, it can indicate that optimization_persistence_flag, optimization_for_machine_analysis_flag, and optimization_type are defined subsequently, and that the optimizations and scopes identified by optimization_persistence_flag, optimization_for_machine_analysis_flag, and optimization_type defined subsequently can be applied.
[0326] The optimization_persistence_flag can indicate the persistence of the optimization indicated by optimization_type. A value of 0 indicates that the optimization identified by optimization_type may only be applied to the current picture. A value of 1 indicates that the optimization identified by optimization_type may be applied to the current picture and all subsequent pictures.
[0327] optimization_for_machine_analysis_flag can indicate that the optimization objective and the encoded bitstream are for machine analysis (suitable for performing machine tasks) if its value is 1. If its value is 0, the optimization objective and the encoded bitstream may or may not be for machine analysis (if the value is '1', it indicates that it is for machine analysis. If the value is '0', it may / may not be for machine analysis).
[0328] optimization_for_human_viewing_flag can indicate that the optimization purpose and the encoded bitstream are intended for human viewing (suitable for human viewing) if its value is 1. If its value is 0, the optimization purpose and the encoded bitstream may or may not be intended for human viewing (if the value is '1', it indicates that it is for human viewing. If the value is '0', it may / may not be for human viewing).
[0329] Optimization can impact encoding quality in ways other than intended. For example, optimization for machine analysis might result in improved quality for human viewing. In this case, the optimization_for_machine_analysis_flag could be set to 1 and optimization_for_human_viewing_flag could be set to 0 to clearly define the intended purpose. Another example is optimization for human viewing might result in improved quality for machine analysis. In this case, optimization_for_machine_analysis_flag could be set to 0 and optimization_for_human_viewing_flag could be set to 1.
[0330] There may be restrictions on the definition of the values of optimization_for_human_viewing_flag and optimization_for_machine_analysis_flag. For example, if the optimization goal is limited to human viewing and machine analysis, optimization_for_human_viewing_flag and optimization_for_machine_analysis_flag may not both have the value '0'.
[0331] In this case, if the value of the preceding flag among optimization_for_human_viewing_flag and optimization_for_machine_analysis_flag is '0', the value of the succeeding flag becomes '1', so there is no need to encode the corresponding value in the bitstream, and the decoder or receiver can process the value as '1'. Tables 23 and 24 above show examples.
[0332] optimization_type indicates the properties of the optimization method.
[0333] [Table 25]
[0334]
[0335] [Table 26]
[0336]
[0337] The optimization properties defined in Tables 25 and 26 can be identified. The optimization properties defined in Tables 25 and 26 are only examples, and new optimization properties can be defined using this structure. It is of course possible to identify optimization properties and methods by including only some of the optimization properties defined in Tables 25 and 26, or to include other optimization properties and methods not defined in Tables 25 and 26. Tables 25 and 26 represent examples that can perform the same function but are expressed in different ways.
[0338] Table 26 is another example of expressing multiple optimization methods via a bitmask for optimization_type, and Tables 27 and 28 show examples.
[0339] [Table 27]
[0340]
[0341] [Table 28]
[0342]
[0343] optimizationForPIIProtectionFlag indicates whether personal information protection optimization is included in the applied optimization method. A value of 1 indicates that personal information protection optimization is included in the applied optimization method, while a value of 0 indicates that personal information protection optimization is not included in the applied optimization method.
[0344] Table 29 shows examples of additionally expressing required additional information depending on the optimization properties.
[0345] [Table 29]
[0346]
[0347] compliant_regulation_info represents information indicating the regulations to be met through privacy protection. The example in Table 30 defines the regulations to be met in a lookup table format, based on the value of compliant_regulation_info.
[0348] [Table 30]
[0349]
[0350] In this case, the value of compliant_regulation_info defines information about one target regulation. Table 31 represents the target regulation defined in Table 30 as the result of compliant_regulation_info and a bitmask.
[0351] [Table 31]
[0352]
[0353] In this case, compliant_regulation_info can express multiple target regulations. Table 32 shows an example. compliant_GDPR_Flag, compliant_CCPA_Flag, compliant_PIPEDA_Flag, and compliant_PIPL_Flag indicate whether the regulations being met (or intended to be met) by applying the personal information protection method are GDPR, CCPA, PIPEDA, and PIPL, respectively. A value of 1 indicates that the regulation is met, while a value of 0 indicates that the regulation is not met.
[0354] [Table 32]
[0355]
[0356] It is clear that the regulations defined in Tables 30 and 31 are examples and that other regulations can also be defined.
[0357] privacy_protection_type is information used to express the properties of the applied privacy protection. The example in Table 33 is an example of the applied method expressed in lookup table format.
[0358] [Table 33]
[0359]
[0360] In the example in Table 33, the value of privacy_protection_type can be used to identify whether the applied method is blurring, masking, or replacement.
[0361] [Table 34]
[0362]
[0363] Table 34 represents the method information defined in Table 33 as a result of privacy_protection_type and bitmask, and Table 35 below shows an example.
[0364] [Table 35]
[0365]
[0366] In this case, it is possible to express multiple properties. blurring_Flag, replacing_Flag, masking_Flag, adversarial_attack_Flag, and anonymization_Flag are information indicating whether privacy protection methods such as blurring, replacing, masking, adversarial attack, and anonymization properties are applied, respectively. If the value is 1, it means that the method was applied, and if the value is 0, it means that the method was not applied.
[0367] The example in Table 36 shows the attribute information of data that has changed due to applied privacy protection.
[0368] [Table 36]
[0369]
[0370] In some cases, information about changes in the personal information attributes of data due to the application of privacy protection may be more important than information about the privacy protection methods applied. Table 36 presents examples of cases where the personal information attributes of data have been pseudonymized, de-identified, or removed due to the application of privacy protection.
[0371] [Table 37]
[0372]
[0373] Table 37 represents the method information defined in Table 36 as a result of privacy_protection_type and bitmask, and Table 38 below shows an example.
[0374] [Table 38]
[0375]
[0376] Pseudonymized_data_Flag, anonymized_data_Flag, Removed_PII_flag, and Identified_data_flag are information indicating whether the personal information attributes of Data have been changed to pseudonymized data, anonymized data, data with PII removed, or identified data due to the application of personal information protection methods, respectively. If the value is 1, it means that the attribute has been changed, and if the value is 0, it means that the attribute has not been changed.
[0377] The privacy protection methods defined in Tables 33, 34, 36, and 37 are examples and other methods may also be defined.
[0378] Protected_info_idc represents information for identifying protected information.
[0379] [Table 39]
[0380]
[0381] In Table 39, if the Protected_info_idc value is 00, it means that information that can identify a person, such as a face, is protected. If the Protected_info_idc value is 01, it means that information that can identify a means of transportation, such as a license plate, is protected. If the Protected_info_idc value is 10, it means that information that can infer a location, such as a business name or letters or pictures on a sign, is protected. If the Protected_info_idc value is 11, it means that information that can infer an object, such as letters or pictures of a product or object, is protected. It is clear that additional protected information can be defined.
[0382] [Table 40]
[0383]
[0384] Table 40 provides an example of how to identify multiple protected pieces of information using a bitmask. In this case, the Protected_info_idc value may be defined differently than in Table 39.
[0385] In Table 41, Protected_personal_identifiable_data_Flag, Protected_vehicle_identifiable_data_Flag, Protected_location_identifiable_data_Flag, and Protected_object_identifiable_data_Flag represent protection of information that can identify a person, such as a face; protection of information that can identify a vehicle, such as a license plate; protection of information that can identify a location, such as a business name or a sign's letters or pictures; and protection of information that can identify a product or object, such as letters or pictures. If the value is 1, it means that protection has been applied to the corresponding information, and if the value is 0, it means that protection has not been applied to the corresponding information.
[0386] [Table 41]
[0387]
[0388] Third Example
[0389] Figure 20 shows an example where an area required for image analysis and an area containing personal information exist together.
[0390] In the example of Figure 20, region A is required to detect a person. There is an SEI that defines information about the region required for a specific purpose. Similarly, information about regions unnecessary for specific purposes or regions requiring caution when used for specific purposes may also be required. For example, personally identifiable information may be present in the defined required region or other regions, and these regions may be subject to modification or removal optimization for privacy protection. An SEI that defines information about these regions may also be required to prevent misuse. For example, region B in Figure 20 corresponds to a human face and contains personally identifiable information. In this case, the personally identifiable information in region B can be removed or modified, and information about the corresponding region can be defined.
[0391] Tables 42, 43, 44, and 45 provide examples of how to represent areas for image analysis and privacy protection areas.
[0392] [Table 42]
[0393]
[0394]
[0395] [Table 43]
[0396]
[0397]
[0398] Tables 42 and 43 independently define information regarding whether each defined area is a privacy-protected area. For example, information (e.g., size, location, etc.) is defined for Area A and Area B, respectively.
[0399] [Table 44]
[0400]
[0401]
[0402]
[0403] Table 44 shows an example of a dependent representation where a privacy protection region or a privacy region is partially contained within a defined annotated region.
[0404] [Table 45]
[0405]
[0406]
[0407]
[0408] In Table 45, if num_protected_region_minus1 is 0, the areas defined by ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] may all be privacy protection processing areas or areas containing personal information.
[0409] The information defined in Tables 42, 43, 44 and 45 may be reconstructed through some or all of them.
[0410] The annotated regions SEI message conveys parameters for identifying annotated regions using bounding boxes that represent the size and location of identified objects. Using this SEI message requires defining the following variables:
[0411] - Cropped picture width and height in luma samples, represented as CroppedWidth and CroppedHeight, respectively.
[0412] - Chroma subsampling width and height, SubWidthC and SubHeightC, respectively
[0413] - Conformance cropping window left offset, ConfWinLeftOffset
[0414] - Conformance cropping window top offset, ConfWinTopOffset
[0415] Below are examples of semantics for Table 42, Table 43, Table 44, and Table 45.
[0416] If ar_cancel_flag is 1, this indicates that the annotation region SEI message cancels the persistence of a previous annotation region SEI message, where the previous annotation region SEI relates to one or more layers to which the annotation region SEI message applies. If ar_cancel_flag is 0, this indicates that annotation region information follows.
[0417] If ar_cancel_flag is 1 or a new CVS for the current layer is started, the variables LabelAssigned[i], ObjectTracked[i], and ObjectBoundingBoxAvail are set to 0 for i in the range 0 to 255.
[0418] If ar_not_optimized_for_viewing_flag is 1, it indicates that the decoded picture to which the annotation region SEI message applies is not optimized for user viewing, but is optimized for some other purpose, such as algorithmic object classification performance. If ar_not_optimized_for_viewing_flag is 0, it indicates that the decoded picture to which the annotation region SEI message applies may or may not be optimized for user viewing.
[0419] If ar_not_optimized_for_machine_analysis_flag is 1, it indicates that the decoded pictures to which the annotation region SEI message applies are not optimized for machine analysis. If ar_not_optimized_for_machine_analysis_flag is 0, it indicates that the decoded pictures to which the annotation region SEI message applies may or may not be optimized for machine analysis.
[0420] If ar_true_motion_flag is 1, it indicates that motion information in the coded picture to which the annotation region SEI message applies has been selected for the purpose of accurately representing object motion for objects in the annotation region. If ar_true_motion_flag is 0, it indicates that motion information in the coded picture to which the annotation region SEI message applies may or may not be selected for the purpose of accurately representing object motion for objects in the annotation region.
[0421] If ar_occluded_object_flag is 1, each of the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicates the size and position of an object, or a portion of an object, that is not visible or only partially visible within the cropped decoded picture. If ar_occluded_object_flag is 0, the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicate the size and position of the object that is fully visible within the cropped decoded picture. It is a bitstream conformance requirement that the value of ar_occluded_object_flag be the same for all annotated_regions( ) syntax structures in the CVS.
[0422] If ar_partial_object_flag_present_flag is 1, it indicates that the ar_partial_object_flag[ar_object_idx[i]] syntax element is present. If ar_partial_object_flag_present_flag is 0, it indicates that the ar_partial_object_flag[ar_object_idx[i]] syntax element is not present. It is a bitstream conformance requirement that the value of ar_partial_object_flag_present_flag be the same for all annotated_regions( ) syntax structures in the CVS.
[0423] If ar_object_label_present_flag is 1, it indicates that label information corresponding to objects in the annotation areas exists. If ar_object_label_present_flag is 0, it indicates that label information corresponding to objects in the annotation areas does not exist.
[0424] If ar_privacy_protection_region_present_flag is 1, it indicates that the ar_privacy_protection_region_flag[ar_object_idx[i]] syntax element is present. If ar_privacy_protection_region_present_flag is 0, it indicates that the ar_privacy_protection_region_flag[ar_object_idx[i]] syntax element is not present. It is a bitstream conformance requirement that the value of ar_privacy_protection_region_present_flag be the same for all annotated_regions( ) syntax structures in CVS.
[0425] If ar_rivacy_protection_region_flag[ar_object_idx[i]] is 1, it specifies that the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]] and ar_bounding_box_height[ar_object_idx[i]] indicate the size and position of the privacy protection region within the decoded picture. If ar_privacy_protection_region_flag[ar_object_idx[i]] is 0, it specifies that the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicate the size and location of a privacy region that may or may not be present within the decoded picture. If not present, the value of ar_partial_object_flag[ar_object_idx[i]] is inferred from the previous annotation region SEI message (if present) in the output order of CVS.
[0426] If ar_include_privacy_protection_region_flag[ar_object_idx[i]] is 1, it indicates that the defined annotated region partially contains or includes personal identification information, and privacy protection is applied to that region. If ar_include_privacy_protection_region_flag[ar_object_idx[i]] is 0, it indicates that the defined annotated region may not partially contain or include personal identification information, and privacy protection is applied to that region.
[0427] num_protected_region_minus1 represents the number of regions containing protected personally identifiable information or the total number of regions containing personally identifiable information minus 1. The value of "num_protected_region" must be in the range of 0 to 255.
[0428] pii_region_top[j], pii_region_left[j], pii_region_width[j], and pii_region_height[j] specify the coordinates of the upper left corner and the width and height, respectively, of the privacy zone or the area containing personally identifiable information within the area defined by ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]].
[0429] If ar_object_confidence_info_present_flag is 1, it indicates that the ar_object_confidence[ar_object_idx[i]] syntax elements are present.
[0430] If ar_object_confidence_info_present_flag is 0, it indicates that the ar_object_confidence[ar_object_idx[i]] syntax element is not present. It is a bitstream conformance requirement that the value of ar_object_confidence_present_flag be the same for all annotated_regions( ) syntax structures in CVS.
[0431] ar_object_confidence_length_minus1 + 1 specifies the length in bits of the ar_object_confidence[ar_object_idx[i]] syntax elements. It is a bitstream conformance requirement that the value of ar_object_confidence_length_minus1 be the same for all annotated_regions( ) syntax structures in CVS.
[0432] If ar_object_label_language_present_flag is 1, it indicates that the ar_object_label_language syntax element is present. If ar_object_label_language_present_flag is 0, it indicates that the ar_object_label_language syntax element is not present.
[0433] ar_bit_equal_to_zero must be 0.
[0434] ar_object_label_language contains language tags as specified in IETF RFC 5646, with a null-terminating byte of 0x00. The length of the ar_object_label_language syntax element must be less than or equal to 255 bytes, excluding the null-terminating byte. If it is not present, the label's language is not specified.
[0435] ar_num_label_updates represents the total number of labels associated with the annotation regions to be signaled. The value of ar_num_label_updates must be in the range of 0 to 255.
[0436] ar_label_idx[i] represents the index of the signaled label. The value of ar_label_idx[i] must be in the range of 0 to 255.
[0437] If ar_label_cancel_flag is 1, the persistence range of the ar_label_idx[i]th label is canceled. If ar_label_cancel_flag is 0, it indicates that the signaled value is assigned to the ar_label_idx[i]th label.
[0438] ar_label[ar_label_idx[i]] specifies the contents of the ar_label_idx[i]th label. The length of the ar_label[ar_label_idx[i]] syntax element must be less than or equal to 255 bytes, excluding the terminating null byte.
[0439] ar_num_object_updates represents the number of object updates being signaled. ar_num_object_updates must be in the range 0 to 255.
[0440] ar_object_idx[i] is the index of the object parameters being signaled. ar_object_idx[i] must be in the range 0 to 255.
[0441] If ar_object_cancel_flag is 1, the persistence scope of the ar_object_idx[i]th object is canceled. If ar_object_cancel_flag is 0, it indicates that the parameters related to the tracked ar_object_idx[i]th object will be signaled.
[0442] If ar_object_label_update_flag is 1, it indicates that the object label will be signaled. If ar_object_label_update_flag is 0, it indicates that the object label will not be signaled.
[0443] ar_object_label_idx[ar_object_idx[i]] represents the index of the label corresponding to the ar_object_idx[i]th object. If ar_object_label_idx[ar_object_idx[i]] does not exist, its value is inferred from the previous annotation area SEI message (if any) in the output order of the same CVS.
[0444] If ar_bounding_box_update_flag is 1, it indicates that the object bounding box parameters will be signaled. If ar_bounding_box_update_flag is 0, it indicates that the object bounding box parameters will not be signaled.
[0445] If ar_bounding_box_cancel_flag is 1, ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], ar_bounding_box_height[ar_object_idx[i]]. The persistence scopes of ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]] are canceled. If ar_bounding_box_cancel_flag is 0, it indicates that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], ar_bounding_box_height[ar_object_idx[i]], ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]] syntax elements will be signaled.
[0446] ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] specify the coordinates of the top left corner and the width and height, respectively, of the bounding box of the ar_object_idx[i]-th object in the cropped decoded picture relative to the conformance cropping window specified by the active SPS.
[0447] The value of ar_bounding_box_left[ar_object_idx[i]] must be in the range 0 to CroppedWidth / SubWidthC - 1.
[0448] The value of ar_bounding_box_top[ar_object_idx[i]] must be in the range 0 to CroppedHeight / SubHeightC - 1.
[0449] The value of ar_bounding_box_width[ar_object_idx[i]] must be in the range 0 to CroppedWidth / SubWidthC - ar_bounding_box_left[ar_object_idx[i]].
[0450] The value of ar_bounding_box_height[ar_object_idx[i]] must be in the range of 0 to CroppedHeight / SubHeightC - ar_bounding_box_top[ar_object_idx[i]].
[0451] The identified object rectangle contains luma samples with horizontal picture coordinates from SubWidthC * (ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]]) to SubWidthC * (ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]] + ar_bounding_box_width[ar_object_idx[i]]) - 1, and vertical picture coordinates from SubHeightC * (ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]]) to SubHeightC * (ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]] + ar_bounding_box_height[ar_object_idx[i]]) - 1.
[0452] The values of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] are maintained in the output order in CVS for each value of ar_object_idx[i]. If not present, the values of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], or ar_bounding_box_height[ar_object_idx[i]] are inferred from the previous annotation area SEI message (if present) in the output order in CVS.
[0453] If ar_partial_object_flag[ar_object_idx[i]] is 1
[0454] Specifies that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] syntax elements indicate the size and position of objects that may be only partially visible within a cropped decoded picture.
[0455] If ar_partial_object_flag[ar_object_idx[i]] is 0, it specifies that the ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] syntax elements represent the size and position of objects that are partially visible or not visible within the cropped decoded picture. If not present, the value of ar_partial_object_flag[ar_object_idx[i]] is inferred from the previous annotation region SEI message (if present) in the output order of CVS.
[0456] ar_object_confidence[ar_object_idx[i]] represents the confidence associated with the ar_object_idx[i]-th object in units of 2-(ar_object_confidence_length_minus1 + 1), such that a higher value of ar_object_confidence[ar_object_idx[i]] indicates a higher confidence. The length of the ar_object_confidence[ar_object_idx[i]] syntax element can be ar_object_confidence_length_minus1 + 1 bits. If not present, the value of ar_object_confidence[ar_object_idx[i]] is inferred from the previous annotation area SEI message (if present) in the output order of CVS.
[0457] Example 4
[0458] Generative face video (GFV) is proposed to encode / decode human faces by combining neural network-based technology and traditional video encoding / decoding technology.
[0459] Table 46 shows an example of the syntax of the proposed GFV.
[0460] [Table 46]
[0461]
[0462]
[0463]
[0464] In generative videos like this, personally identifiable information may be altered. Furthermore, facial information may be intentionally altered depending on the base picture and face generation network used to generate the face. Therefore, caution is required when using the video for personal identification purposes. To prevent misuse, it is necessary to define whether personally identifiable information is altered or whether it is appropriate for personal identification purposes.
[0465] Tables 47, 48, and 49 show examples of generative face video SEI message syntax that define whether personally identifiable information is converted or whether it is suitable for personal identification purposes.
[0466] [Table 47]
[0467]
[0468]
[0469]
[0470] [Table 48]
[0471]
[0472]
[0473]
[0474] [Table 49]
[0475]
[0476]
[0477]
[0478] The not_optimized_for_machine_analysis_flag flag, if its value is 1, indicates that the generated / decoded result may not be suitable (optimized) for machine analysis. Being unsuitable means that the results of machine analysis may differ from the intended results or may be inaccurate. For example, if the shape or color of an object is changed, or some information is removed, the results of object recognition may be inaccurate. A value of 0 indicates that the generated / decoded result may not be suitable for machine analysis. This syntax can be applied not only to GFVs but also to other generated / decoded results.
[0479] gfv_not_optimized_for_personal_identification_flag, if its value is 1, means that the result generated / decrypted via GFV may not be suitable (optimized) for personal identification purposes. If its value is 0, it means that the result generated / decrypted via GFV may not be suitable for personal identification purposes.
[0480] gfv_matrix_transformed_flag, if its value is 1, indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., base picture). If its value is 0, it indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) are the same as the attributes of the input data (e.g., base picture).
[0481] gfv_matrix_property is information that represents the information properties of each matrix.
[0482] Table 47 shows an example of defining gfv_matrix_property when gfv_matrix_transformed_flag is 1. In this case, gfv_matrix_property can be defined as follows.
[0483] The properties defined by gfv_matrix_property may include information transformation, removal, enhancement, etc. to achieve the purpose of the matrix. If the value is 0, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed differently from the properties of the input data (e.g., base picture). If the value is 1, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed or partially removed from the properties of the input data (e.g., base picture) for the purpose of protecting personally identifiable information. The above are examples, and additional information representing the properties of the matrix may be defined. gfv_matrix_property[i] represents the properties of the matrix represented by gfv_matrix_type_idx[i].
[0484] Table 48 shows an example of defining gfv_matrix_transformed_flag alone.
[0485] Table 49 shows an example of defining gfv_matrix_property alone, in which case the definition of gfv_matrix_property could be as follows.
[0486] gfv_matrix_property is information that represents the information properties of each matrix. The properties defined by gfv_matrix_property may include transformation, removal, enhancement, etc. of information to achieve the purpose of the matrix. For example, if the value is 0, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) are the same as those of the input data (e.g., base picture). If the value is 1, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed differently from the properties of the input data (e.g., base picture). If the value is 2, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed or partially removed from the properties of the input data (e.g., base picture) for the purpose of protecting personally identifiable information. The above are examples, and additional information representing the properties of the matrix may be defined. gfv_matrix_property[i] represents the properties of the matrix represented by gfv_matrix_type_idx[i].
[0487] gfv_matrix_transformed_flag, if its value is 1, indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., base picture). If its value is 0, it indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) are the same as the attributes of the input data (e.g., base picture).
[0488] The above example illustrates the usage of the defined gfv_matrix_transformed_flag and gfv_matrix_property properties. Other uses are also possible. For example, both properties can be defined together without any separate conditions.
[0489] The GFV (generative face video) SEI message specifies a face parameter transformation network, denoted as TranslatorNN( ), and a face picture generator neural network, denoted as GenerativeNN( ), representing face parameters. Here, TranslatorNN( ) can be used to transform face parameters of various formats signaled in the SEI message into parameters of a fixed format. GenerativeNN( ) can be used to generate output pictures using the format of the face parameters and previously decoded output pictures.
[0490] Note 1 - Facial parameters can be determined from source pictures prior to encoding. These source pictures can be referred to as driving pictures.
[0491] NOTE 2 - The previously decoded output picture input to GenerativeNN( ) can be a base picture (a decoded output picture that provides a reference texture from which a face picture can be generated) and optionally a picture that can be fused by GenerativeNN( ) to enhance background texture and face details. If the current picture is not a base picture, the GFV SEI message can be used to generate a face picture based on the previously decoded base picture, the face parameters passed by the GFV SEI message, and optionally the current decoded picture for fusion purposes.
[0492] The use of these SEI messages requires the definition of the following variables:
[0493] - Input picture width and height in luma samples, represented as CroppedWidth and CroppedHeight, respectively.
[0494] - Luma sample array baseCroppedYPic for the decoded output picture, chroma sample arrays baseCroppedCbPic and baseCroppedCrPic, and BasePicture corresponding to the source base picture.
[0495] - The luma sample array driveCroppedYPic for the decoded output picture and the chroma sample arrays driveCroppedCbPic and driveCroppedCrPic, which are represented by DrivePicture corresponding to the source driving picture.
[0496] - Bit depth for the luma sample array of the input picture BitDepth Y
[0497] - BitDepth, the bit depth of the chroma sample arrays (if any) of the input picture. C
[0498] - Chroma format indicator, indicated by ChromaFormatIdc
[0499] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.
[0500] gfv_id contains an identifier that can be used to identify facial feature information and specify a neural network that can be used with GenerativeNN( ). The value of gfv_id is 0 to 2. 32_ -2 must be in the range of 256 to 511 and 231 to 2 32_ -gfv_id values in the range -2 are reserved for future use in ITU-T | ISO / IEC. The decoder can use values in the range 256 to 511 or 231 to 2. 32_ GFV SEI messages containing gfv_id in the range -2 to 1000 must be ignored.
[0501] NOTE - For example, if there is more than one face in the output photo, different values of gfv_id in different GFV SEI messages can be used to identify different faces.
[0502] If gfv_base_pic_flag is 1, it indicates that the currently decoded output picture corresponds to the base picture. If gfv_base_pic_flag is 0, it indicates that the currently decoded output picture does not correspond to the base picture.
[0503] The following constraints apply to the gfv_base_pic_flag value:
[0504] - If the GFV SEI message is the first GFV SEI message in decoding order with a specific gfv_id value within the current CLVS, the gfv_base_pic_flag value must be 1.
[0505] - If the gfv_base_pic_flag of a GFV SEI message with a particular gfv_id value is 0, then this SEI message is associated with the current decoded picture and all subsequent decoded pictures of the current layer in output order, up to and including the end of the current CLVS or the decoded picture that follows the current decoded picture in output order within the current CLVS, and this SEI message is associated with the earlier of subsequent GFV SEI messages in output order that have the gfv_base_pic_flag 0 and have the particular gfv_id value within the current CLVS. (When a GFV SEI message that has a particular gfv_id value has gfv_base_pic_flag being equal to 0, this SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer, in output order, until the end of the current CLVS or up to but excluding the decoded picture that follows the current decoded picture in output order within the current CLVS and is associated with a subsequent GFV SEI message, in decoding order, having gfv_base_pic_flag equal to 0 and that particular gfv_id value within the current CLVS, whichever is earlier.)
[0506] gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_payload_byte[i] specify a neural network that can be used with TranslatorNN( ). gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_payload_byte[i] have the same syntax and semantics as nnpfc_base_flag, nnpfc_mode_idc, nnpfc_reserved_zero_bit_a, nnpfc_tag_uri, nnpfc_uri, nnpfc_payload_byte[i], respectively.
[0507] If present, gfv_drive_pic_fusion_flag, if 1, indicates that the currently decoded picture can be input to GenerativeNN( ). The currently decoded picture corresponds to a driving picture that can be used for fusion. If gfv_drive_pic_fusion_flag, if 0, indicates that the currently decoded picture should not be input to GenerativeNN( ).
[0508] Note 3 - For example, a gfv_drive_pic_fusion_flag value of 1 could be used to indicate that the currently decoded picture can enhance facial details or handle background changes.
[0509] Note 4 - Fusion outputs a picture using three inputs: a base picture, keypoints and / or features from a matrix passed in a GFV SEI message, and the currently decoded picture.
[0510] Note 5 - If the currently decoded picture corresponds to a driving picture, it must be marked as not intended for output.
[0511] If not_optimized_for_machine_analysis_flag is 1, it indicates that the generated / decoded result may not be optimized (unsuitable) for machine analysis purposes. Unsuitability means that the machine analysis result may differ from the intended result or may not be accurate. If not_optimized_for_machine_analysis_flag is 0, it indicates that the generated / decoded result may be suitable (or not suitable) for machine analysis purposes. This syntax can be applied not only to GFVs but also to other generated / decoded results.
[0512] If gfv_not_optimized_for_personal_identification_flag is 1, it indicates that the result generated / decoded via GFV may not be optimized (suitable) for personal identification purposes. If gfv_not_optimized_for_personal_identification_flag is 0, it indicates that the result generated / decoded via GFV may be suitable (suitable) for personal identification purposes.
[0513] gfv_matrix_property represents information about the properties of each matrix. The properties defined by gfv_matrix_property include transformation, removal, and enhancement. For example, if gfv_matrix_property is 0, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head) are the same as the properties of the input data (e.g., base picture). If gfv_matrix_property is 1, it indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. If gfv_matrix_property is 2, it indicates that the information represented by the matrix (e.g., mouth, eyes, head) have been transformed or partially removed for the purpose of protecting personally identifiable information. Information defining the properties of the matrix can be additionally specified.
[0514] When gfv_matrix_transformed_flag is set to 1, the semantics of gfv_matrix_property defined under this condition can be specified as follows.
[0515] gfv_matrix_property represents information about the properties of each matrix. The properties defined by gfv_matrix_property include transformations, removals, and enhancements. For example, a value of 0 indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. A value of 1 indicates that the information represented by the matrix, such as the properties of the mouth, eyes, and hair, have been transformed or partially removed for the purpose of protecting personally identifiable information. The information defining the properties of the matrix can be additionally specified.
[0516] If gfv_matrix_transformed_flag is 1, it indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head) have been transformed, for example, enhanced, removed, or replaced, to be different from the attributes of the input data (e.g., base picture). If gfv_matrix_transformed_flag is 0, it indicates that the attributes of the information represented by the matrix are the same as the attributes of the input data.
[0517] If gfv_coordinate_present_flag is 1, it indicates that coordinate information of keypoints exists. If gfv_coordinate_present_flag is 0, it indicates that coordinate information of keypoints does not exist.
[0518] It is a bitstream conformance requirement that the value of gfv_coordinate_present_flag must be 1 if gfv_matrix_type_idx[i] is 0 or 1 for all i from 0 to gfv_num_matrix_types_minus1.
[0519] gfv_coordinate_precision_factor_minus1 plus 1 represents the length in bits of gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i].
[0520] gfv_num_kps_minus1 plus 1 represents the number of keypoints. The value of gfv_num_kp_minus1 is from 0 to 2. 10 - Must be in the range of 1.
[0521] If gfv_kp_pred_flag is 1, it indicates that the syntax elements gfv_coordinate_dx_abs[i], gfv_coordinate_dy_abs[i], and gfv_coordinate_dz_abs[i] are present, and the syntax elements gfv_coordinate_dx_sign_flag[i], gfv_coordinate_dy_sign_flag[i], and gfv_coordinate_dz_sign_flag[i] may be present. If gfv_kp_pred_flag is 0, it indicates that gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i] are present, and that syntax elements gfv_coordinate_x_sign_flag[i], gfv_coordinate_y_sign_flag[i], and gfv_coordinate_z_sign_flag[i] may be present.
[0522] If gfv_coordinate_z_present_flag is 1, it indicates that the z-axis coordinate information of the keypoints exists. If gfv_coordinate_z_present_flag is 0, it indicates that the z-axis coordinate information of the keypoints does not exist.
[0523] gfv_coordinate_z_max_value_minus1 plus 1 represents the maximum absolute value of the z-axis coordinates of keypoints.
[0524] gfv_coordinate_x_abs[i] represents the normalized absolute value of the x-axis coordinate of the i-th keypoint.
[0525] gfv_coordinate_x_sign_flag[i] specifies the sign of the x-coordinate of the i-th keypoint. If gfv_coordinate_x_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0526] gfv_coordinate_y_abs[i] specifies the normalized absolute value of the y-coordinate of the i-th keypoint.
[0527] gfv_coordinate_y_sign_flag[i] specifies the sign of the y-coordinate of the i-th keypoint. If gfv_coordinate_y_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0528] gfv_coordinate_z_abs[i] represents the normalized absolute value of the z-axis coordinate of the i-th keypoint.
[0529] gfv_coordinate_z_sign_flag[i] specifies the sign of the z-coordinate of the i-th keypoint. If gfv_coordinate_z_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0530] gfv_coordinate_dx_abs[i] represents the absolute difference value of the normalized value of the x-axis coordinate of the i-th keypoint.
[0531] gfv_coordinate_dx_sign_flag[i] specifies the sign of the difference value of the x-coordinate of the i-th keypoint. If gfv_coordinate_dx_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0532] gfv_coordinate_dy_abs[i] specifies the absolute difference value of the normalized y-coordinate of the i-th keypoint.
[0533] gfv_coordinate_dy_sign_flag[i] specifies the sign of the difference value of the y-coordinate of the i-th keypoint. If gfv_coordinate_yd_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0534] gfv_coordinate_dz_abs[i] specifies the absolute difference value of the normalized z-axis coordinate of the i-th keypoint.
[0535] gfv_coordinate_dz_sign_flag[i] specifies the sign of the difference value of the z-coordinate of the i-th keypoint. If gfv_coordinate_dz_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0536] The variables coordinateDeltaX[i], coordinateDeltaY[i], and coordinateDeltaZ[i], which represent the delta x-axis coordinate, delta y-axis coordinate, and delta z-axis coordinate of the i-th keypoint, respectively, are derived as shown in Table 50 below.
[0537] [Table 50]
[0538]
[0539] The variables coordinateX[i], coordinateY[i], and coordinateZ[i], which represent the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the i-th keypoint, respectively, are derived as in Table 51 below when gfv_kp_pred_flag is 0, and are derived as in Table 52 below when gfv_kp_pred_flag is 1.
[0540] [Table 51]
[0541]
[0542] [Table 52]
[0543]
[0544] Here, BaseKpCoordinateX[i], BaseKpCoordinateY[i], and BaseKpCoordinateZ[i], which represent the x-axis, y-axis, and z-axis coordinates of the i-th key point for the base picture, respectively, are derived as shown in Table 53 below.
[0545] [Table 53]
[0546]
[0547] If gfv_matrix_present_flag is 1, it indicates that matrix parameters are present. If gfv_matrix_present_flag is 0, it indicates that matrix parameters are not present.
[0548] gfv_matrix_element_precision_factor_minus1 plus 1 represents the length of gfv_matrix_element_dec[i][j][k][m] in bits.
[0549] gfv_num_matrix_types_minus1 plus 1 indicates the number of matrix types signaled in the SEI message. The value of gfv_matrix_type_num_minus1 is between 0 and 2. 6 - Must be in the range of 1.
[0550] gfv_matrix_type_idx[i] represents the index of the i-th matrix type as specified in Table 54.
[0551] [Table 54]
[0552]
[0553] Note 2. The undefined matrix type is used to represent matrix types other than affine transformation matrices, covariance matrices, rotation matrices, translation matrices, and compact feature matrices. This can be used by users to extend the matrix type.
[0554] If gfv_num_matrices_equal_to_num_kps_flag[i] is 1, it indicates that the number of matrices of the i-th matrix type is equal to gfv_num_kps_minus1 + 1. If gfv_num_matrices_equal_to_num_kps_flag[i] is 0, it indicates that the number of matrices of the i-th matrix type is not equal to gfv_num_kps_minus1 + 1.
[0555] gfv_num_matrices_info[i] provides information for deriving the number of matrices of the i-th matrix type.
[0556] gfv_matrix_width_minus1[i] plus 1 represents the width of the matrix of the i-th matrix type.
[0557] gfv_matrix_height_minus1[i] plus 1 represents the height of the matrix of the i-th matrix type.
[0558] If gfv_matrix_for_3D_space_flag[i] is 1, it indicates that the matrix of the i-th matrix type is defined in 3D space. If gfv_matrix_for_3D_space_flag[i] is 0, it indicates that the matrix of the i-th matrix type is defined in 2D space.
[0559] If gfv_matrix_width_minus1[i] does not exist, the following is inferred:
[0560] - If gfv_matrix_type_idx[i] is 0, 1, or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is present and equal to 1, then gfv_matrix_width_minus1[i] is inferred to be equal to 2.
[0561] - Otherwise, if matrix_type_idx[i] is 0, 1, or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is present and equal to 0, then gfv_matrix_width_minus1[i] is inferred to be equal to 1.
[0562] - Otherwise (if matrix_type_idx[i] is 5 or 6), gfv_matrix_width_minus1[i] is inferred to be equal to 0.
[0563] If gfv_matrix_height_minus1[i] does not exist, the following is inferred:
[0564] - If matrix_type_idx is 0, 1, 4, 5, or 6, and either gfv_coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is present and equal to 1, then gfv_matrix_height_minus1[i] is inferred to be equal to 2.
[0565] - Otherwise (gfv_matrix_type_idx is 0, 1, 4, 5, or 6, and either gfv_coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is 0), gfv_matrix_height _minus1[i] is inferred to be equal to 1.
[0566] The variables matrixWidth[i] and matrixHeight[i], which represent the width and height of the matrix of the i-th matrix type, are derived as shown in Table 55 below.
[0567] [Table 55]
[0568]
[0569] gfv_num_matrices_minus1[i] plus 1 represents the number of matrices of the i-th matrix type. The variable numMatrices[i], which represents the number of matrices of the i-th matrix type, is derived as shown in Table 56 below.
[0570] [Table 56]
[0571]
[0572] gfv_matrix_element_int[i][j][k][m] represents the integer part of the matrix element value at position (k, m) of the jth matrix of the ith matrix type.
[0573] gfv_matrix_element_dec[i][j][k][m] represents the decimal part of the matrix element value at position (k, m) of the jth matrix of the ith matrix type.
[0574] gfv_matrix_element_sign_flag[i][j][k][m] indicates the sign of the matrix element at position (k, m) of the jth matrix of the ith matrix type. If gfv_matrix_element_sign_flag[i][j][k][m] does not exist, it is inferred to be equal to 0.
[0575] matrixElementVal[i][j][k][m], which represents the matrix element value at position (k, m) of the jth matrix of the ith matrix type, is derived as shown in Table 57 below.
[0576] [Table 57]
[0577]
[0578] GenerativeNN ( ) is a process for generating sample values of the output picture corresponding to the driving picture. It is called only when gfc_base_pic_flag is 0.
[0579] The input to TranslatorNN() is:
[0580] - sigKeyPoint and sigMatrix
[0581] The output of TranslatorNN() is:
[0582] - convKeyPoint and convMatrix
[0583] The input to GenerativeNN() is:
[0584] - gfv_base_pic_flag가 0이고 gfv_drive_pic_fusion_flag가 0이며 ChromaFormatIdc가 0이면: inputBaseY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0585] - gfv_base_pic_flag가 0이고 gfv_drive_pic_fusion_flag가 0이며 ChromaFormatIdc가 0이 아니면: inputBaseY, inputBaseCb, inputBaseCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0586] - gfv_base_pic_flag가 0이고 gfv_drive_pic_fusion_flag가 1이며 ChromaFormatIdc가 0이면: inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0587] - gfv_base_pic_flag가 0이고 gfv_drive_pic_fusion_flag가 1이며 ChromaFormatIdc가 0이 아니면: inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0588] GenerativeNN( )의 출력은:
[0589] - 루나 샘플 어레이 genY
[0590] - If ChromaFormatIdc is not 0, there are two chroma sample arrays genCb and genCr.
[0591] The following process is used to generate a video picture, as shown in Table 58:
[0592] [Table 58]
[0593]
[0594] To derive the input of TranslatorNN(), the process DeriveSigParam() is specified as follows:
[0595] The keypoint coordinate array sigKeyPoint and matrix sigMatrix are derived as shown in Table 59 below.
[0596] [Table 59]
[0597]
[0598] The input values for GenerativeNN( ) are real numbers, and the functions InpY( ) and InpC( ) are specified as shown in Table 60 below.
[0599] [Table 60]
[0600]
[0601] The output values in GenerativeNN( ) are real numbers, and the functions OutY( ) and OutC( ) are specified as shown in Table 61 below.
[0602] [Table 61]
[0603]
[0604] To derive the input of GenerativeNN ( ), the process DeriveInputTensors( ) is specified as follows:
[0605] If gfv_base_pic_flag is 1, the BasePicture input tensors inputBaseY, inputBaseCb, and inputBaseCr are derived as shown in Table 62 below.
[0606] [Table 62]
[0607]
[0608] If gfv_drive_pic_fusion_flag is 1, the DrivePicture luma sample arrays inputDriveY, inputDriveCb, and inputDriveCr are derived as shown in Table 63 below.
[0609] [Table 63]
[0610]
[0611] If gfv_base_pic_flag is 0, the keypoint coordinate array inputDriveKeyPoint and matrix inputDriveMatrix for the current picture are derived as shown in Table 64 below.
[0612] [Table 64]
[0613]
[0614] If gfv_base_pic_flag is 1, the keypoint coordinate array inputBaseKeyPoint and matrix inputBaseMatrix for the base picture are derived as shown in Table 65 below.
[0615] [Table 65]
[0616]
[0617] To derive the output, the process StoreOutputTensors( ) is specified as follows:
[0618] If gfv_base_pic_flag is 0, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as shown in Table 66 below.
[0619] [Table 66]
[0620]
[0621] If gfv_base_pic_flag is 1, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as shown in Table 67.
[0622] [Table 67]
[0623]
[0624] Example 5
[0625] Tables 68, 69, and 70 show examples of defining whether personally identifiable information is converted or suitable for personal identification purposes in the generated face video SEI message syntax.
[0626] [Table 68]
[0627]
[0628]
[0629]
[0630] [Table 69]
[0631]
[0632]
[0633]
[0634] [Table 70]
[0635]
[0636]
[0637]
[0638] optimized_for_machine_analysis_flag, if its value is 0, means that the generated / decoded result may not be suitable (optimized) for machine analysis. Being unsuitable means that the result of machine analysis may be different from the intended result or may not be accurate. For example, if the shape or color of the object is changed or some information is removed, the result of object recognition may be inaccurate. If its value is 1, it means that the generated / decoded result may be suitable for machine analysis. This syntax can be applied not only to GFV but also to other generated / decoded results.
[0639] gfv_optimized_for_personal_identification_flag, if its value is 0, means that the result generated / decrypted via GFV may not be suitable (optimized) for personal identification purposes. If its value is 1, it means that the result generated / decrypted via GFV may be suitable for personal identification purposes.
[0640] gfv_matrix_transformed_flag, if its value is 1, indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., base picture). If its value is 0, it indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) are the same as the attributes of the input data (e.g., base picture).
[0641] gfv_matrix_property is information that represents the information properties of each matrix.
[0642] Table 68 shows an example of defining gfv_matrix_property when gfv_matrix_transformed_flag is 1. In this case, gfv_matrix_property can be defined as follows.
[0643] The properties defined by gfv_matrix_property may include information transformation, removal, enhancement, etc. to achieve the purpose of the matrix. If the value is 0, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed differently from the properties of the input data (e.g., base picture). If the value is 1, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed or partially removed from the properties of the input data (e.g., base picture) for the purpose of protecting personally identifiable information. The above are examples, and additional information representing the properties of the matrix may be defined. gfv_matrix_property[i] represents the properties of the matrix represented by gfv_matrix_type_idx[i].
[0644] Table 69 shows an example of defining gfv_matrix_transformed_flag alone.
[0645] Table 70 shows an example of defining gfv_matrix_property alone, in which case the definition of gfv_matrix_property could be as follows.
[0646] gfv_matrix_property is information that represents the information properties of each matrix. The properties defined by gfv_matrix_property may include transformation, removal, enhancement, etc. of information to achieve the purpose of the matrix. For example, if the value is 0, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) are the same as those of the input data (e.g., base picture). If the value is 1, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed differently from the properties of the input data (e.g., base picture). If the value is 2, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed or partially removed from the properties of the input data (e.g., base picture) for the purpose of protecting personally identifiable information. The above are examples, and additional information representing the properties of the matrix may be defined. gfv_matrix_property[i] represents the properties of the matrix represented by gfv_matrix_type_idx[i].
[0647] gfv_matrix_transformed_flag, if its value is 1, indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., base picture). If its value is 0, it indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head, etc.) are the same as the attributes of the input data (e.g., base picture).
[0648] The above example illustrates the usage of the defined gfv_matrix_transformed_flag and gfv_matrix_property properties. Other uses are also possible. For example, both properties can be defined together without any separate conditions.
[0649] The GFV SEI message specifies a face parameter transformation network, denoted as TranslatorNN( ), and a face picture generator neural network, denoted as GenerativeNN( ), representing face parameters. Here, TranslatorNN( ) can be used to transform face parameters of various formats signaled in the SEI message into parameters of a fixed format. GenerativeNN( ) can be used to generate output pictures using the format of the face parameters and previously decoded output pictures.
[0650] Note 1 - Facial parameters can be determined from source pictures prior to encoding. These source pictures can be referred to as driving pictures.
[0651] NOTE 2 - The previously decoded output picture input to GenerativeNN( ) can be a base picture (a decoded output picture that provides a reference texture from which a face picture can be generated) and optionally a picture that can be fused by GenerativeNN( ) to enhance background texture and face details. If the current picture is not a base picture, the GFV SEI message can be used to generate a face picture based on the previously decoded base picture, the face parameters passed by the GFV SEI message, and optionally the current decoded picture for fusion purposes.
[0652] The use of these SEI messages requires the definition of the following variables:
[0653] - Input picture width and height in luma samples, represented as CroppedWidth and CroppedHeight, respectively.
[0654] - Luma sample array baseCroppedYPic for the decoded output picture, chroma sample arrays baseCroppedCbPic and baseCroppedCrPic, and BasePicture corresponding to the source base picture.
[0655] - The luma sample array driveCroppedYPic for the decoded output picture and the chroma sample arrays driveCroppedCbPic and driveCroppedCrPic, which are represented by DrivePicture corresponding to the source driving picture.
[0656] - Bit depth for the luma sample array of the input picture BitDepthY
[0657] - BitDepthC, the bit depth for the chroma sample arrays of the input picture (if any).
[0658] - Chroma format indicator, indicated by ChromaFormatIdc
[0659] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.
[0660] gfv_id contains an identifier that can be used to identify facial feature information and specify a neural network that can be used with GenerativeNN( ). The value of gfv_id is 0 to 2. 32_ -2 must be in the range of 256 to 511 and 231 to 2 32_ -gfv_id values in the range -2 are reserved for future use in ITU-T | ISO / IEC. The decoder can use values in the range 256 to 511 or 231 to 2. 32_ GFV SEI messages containing gfv_id in the range -2 to 1000 must be ignored.
[0661] NOTE - For example, if there is more than one face in the output photo, different values of gfv_id in different GFV SEI messages can be used to identify different faces.
[0662] If gfv_base_pic_flag is 1, it indicates that the currently decoded output picture corresponds to the base picture. If gfv_base_pic_flag is 0, it indicates that the currently decoded output picture does not correspond to the base picture.
[0663] The following constraints apply to the gfv_base_pic_flag value:
[0664] - If the GFV SEI message is the first GFV SEI message in decoding order with a specific gfv_id value within the current CLVS, the gfv_base_pic_flag value must be 1.
[0665] - If the gfv_base_pic_flag of a GFV SEI message with a particular gfv_id value is 0, then this SEI message is associated with the current decoded picture and all subsequent decoded pictures of the current layer in output order, up to and including the end of the current CLVS or the decoded picture that follows the current decoded picture in output order within the current CLVS, and this SEI message is associated with the earlier of subsequent GFV SEI messages in output order that have the gfv_base_pic_flag 0 and have the particular gfv_id value within the current CLVS. (When a GFV SEI message that has a particular gfv_id value has gfv_base_pic_flag being equal to 0, this SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer, in output order, until the end of the current CLVS or up to but excluding the decoded picture that follows the current decoded picture in output order within the current CLVS and is associated with a subsequent GFV SEI message, in decoding order, having gfv_base_pic_flag equal to 0 and that particular gfv_id value within the current CLVS, whichever is earlier.)
[0666] gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_payload_byte[i] specify a neural network that can be used with TranslatorNN( ). gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_payload_byte[i] have the same syntax and semantics as nnpfc_base_flag, nnpfc_mode_idc, nnpfc_reserved_zero_bit_a, nnpfc_tag_uri, nnpfc_uri, nnpfc_payload_byte[i], respectively.
[0667] If present, gfv_drive_pic_fusion_flag, if 1, indicates that the currently decoded picture can be input to GenerativeNN( ). The currently decoded picture corresponds to a driving picture that can be used for fusion. If gfv_drive_pic_fusion_flag, if 0, indicates that the currently decoded picture should not be input to GenerativeNN( ).
[0668] Note 3 - For example, a gfv_drive_pic_fusion_flag value of 1 could be used to indicate that the currently decoded picture can enhance facial details or handle background changes.
[0669] Note 4 - Fusion outputs a picture using three inputs: a base picture, keypoints and / or features from a matrix passed in a GFV SEI message, and the currently decoded picture.
[0670] Note 5 - If the currently decoded picture corresponds to a driving picture, it must be marked as not intended for output.
[0671] If optimized_for_machine_analysis_flag is 0, it indicates that the generated / decoded results may not be optimized (suitable) for machine analysis purposes. Being unsuitable means that the results of the machine analysis may differ from the intended results or may not be accurate. If optimized_for_machine_analysis_flag is 1, it indicates that the generated / decoded results may be suitable for machine analysis purposes. This syntax can be applied not only to GFVs but also to other generated / decoded results.
[0672] If gfv_optimized_for_personal_identification_flag is 0, it indicates that the result generated / decoded via GFV may not be optimized (suitable) for personal identification purposes. If gfv_not_optimized_for_personal_identification_flag is 0, it indicates that the result generated / decoded via GFV may be suitable for personal identification purposes.
[0673] gfv_matrix_property represents information about the properties of each matrix. The properties defined by gfv_matrix_property include transformation, removal, and enhancement. For example, if gfv_matrix_property is 0, it indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, head) are the same as the properties of the input data (e.g., base picture). If gfv_matrix_property is 1, it indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. If gfv_matrix_property is 2, it indicates that the information represented by the matrix (e.g., mouth, eyes, head) have been transformed or partially removed for the purpose of protecting personally identifiable information. Information defining the properties of the matrix can be additionally specified.
[0674] When gfv_matrix_transformed_flag is set to 1, the semantics of gfv_matrix_property defined under this condition can be specified as follows.
[0675] gfv_matrix_property represents information about the properties of each matrix. The properties defined by gfv_matrix_property may include transformations, removals, and enhancements. For example, a value of 0 indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. A value of 1 indicates that the information represented by the matrix, such as the properties of the mouth, eyes, and hair, have been transformed or partially removed for the purpose of protecting personally identifiable information. The information defining the properties of the matrix may be additionally specified.
[0676] If gfv_matrix_transformed_flag is 1, it indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, head) have been transformed, for example, enhanced, removed, or replaced, to be different from the attributes of the input data (e.g., base picture). If gfv_matrix_transformed_flag is 0, it indicates that the attributes of the information represented by the matrix are the same as the attributes of the input data.
[0677] If gfv_coordinate_present_flag is 1, it indicates that coordinate information of the keypoint exists. If gfv_coordinate_present_flag is 0, it indicates that coordinate information of the keypoint does not exist.
[0678] It is a bitstream conformance requirement that gfv_coordinate_present_flag must be 1 if gfv_matrix_type_idx[i] is 0 or 1 for all i from 0 to gfv_num_matrix_types_minus1.
[0679] gfv_coordinate_precision_factor_minus1 plus 1 represents the length in bits of gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i].
[0680] gfv_num_kps_minus1 plus 1 represents the number of keypoints. The value of gfv_num_kp_minus1 is from 0 to 2. 10 - Must be in the range of 1.
[0681] If gfv_kp_pred_flag is 1, it indicates that the syntax elements gfv_coordinate_dx_abs[i], gfv_coordinate_dy_abs[i], and gfv_coordinate_dz_abs[i] are present, and the syntax elements gfv_coordinate_dx_sign_flag[i], gfv_coordinate_dy_sign_flag[i], and gfv_coordinate_dz_sign_flag[i] may be present. If gfv_kp_pred_flag is 0, it indicates that gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i] are present, and that syntax elements gfv_coordinate_x_sign_flag[i], gfv_coordinate_y_sign_flag[i], and gfv_coordinate_z_sign_flag[i] may be present.
[0682] If gfv_coordinate_z_present_flag is 1, it indicates that the z-axis coordinate information of the keypoints exists. If gfv_coordinate_z_present_flag is 0, it indicates that the z-axis coordinate information of the keypoints does not exist.
[0683] gfv_coordinate_z_max_value_minus1 plus 1 represents the maximum absolute value of the z-axis coordinates of keypoints.
[0684] gfv_coordinate_x_abs[i] represents the normalized absolute value of the x-axis coordinate of the i-th keypoint.
[0685] gfv_coordinate_x_sign_flag[i] specifies the sign of the x-coordinate of the i-th keypoint. If gfv_coordinate_x_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0686] gfv_coordinate_y_abs[i] specifies the normalized absolute value of the y-coordinate of the i-th keypoint.
[0687] gfv_coordinate_y_sign_flag[i] specifies the sign of the y-coordinate of the i-th keypoint. If gfv_coordinate_y_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0688] gfv_coordinate_z_abs[i] represents the normalized absolute value of the z-axis coordinate of the i-th keypoint.
[0689] gfv_coordinate_z_sign_flag[i] specifies the sign of the z-coordinate of the i-th keypoint. If gfv_coordinate_z_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0690] gfv_coordinate_dx_abs[i] represents the absolute difference value of the normalized value of the x-axis coordinate of the i-th keypoint.
[0691] gfv_coordinate_dx_sign_flag[i] specifies the sign of the difference value of the x-coordinate of the i-th keypoint. If gfv_coordinate_dx_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0692] gfv_coordinate_dy_abs[i] specifies the absolute difference value of the normalized y-coordinate of the i-th keypoint.
[0693] gfv_coordinate_dy_sign_flag[i] specifies the sign of the difference value of the y-coordinate of the i-th keypoint. If gfv_coordinate_yd_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0694] gfv_coordinate_dz_abs[i] specifies the absolute difference value of the normalized z-axis coordinate of the i-th keypoint.
[0695] gfv_coordinate_dz_sign_flag[i] specifies the sign of the difference value of the z-coordinate of the i-th keypoint. If gfv_coordinate_dz_sign_flag[i] does not exist, it is inferred to be equal to 0.
[0696] The variables coordinateDeltaX[i], coordinateDeltaY[i], and coordinateDeltaZ[i], which represent the delta x-axis coordinate, delta y-axis coordinate, and delta z-axis coordinate of the i-th keypoint, respectively, are derived as shown in Table 71 below.
[0697] [Table 71]
[0698]
[0699] The variables coordinateX[i], coordinateY[i], and coordinateZ[i], which represent the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the i-th keypoint, respectively, are derived as in Table 72 below when gfv_kp_pred_flag is 0, and are derived as in Table 73 below when gfv_kp_pred_flag is 1.
[0700] [Table 72]
[0701]
[0702] [Table 73]
[0703]
[0704] Here, BaseKpCoordinateX[i], BaseKpCoordinateY[i], and BaseKpCoordinateZ[i], which represent the x-axis, y-axis, and z-axis coordinates of the i-th key point of the base picture, respectively, are derived as shown in Table 74 below.
[0705] [Table 74]
[0706]
[0707] If gfv_matrix_present_flag is 1, it indicates that matrix parameters are present. If gfv_matrix_present_flag is 0, it indicates that matrix parameters are not present.
[0708] gfv_matrix_element_precision_factor_minus1 plus 1 represents the length of gfv_matrix_element_dec[i][j][k][m] in bits.
[0709] gfv_num_matrix_types_minus1 plus 1 indicates the number of matrix types signaled in the SEI message. The value of gfv_matrix_type_num_minus1 is between 0 and 2. 6 - Must be in the range of 1.
[0710] gfv_matrix_type_idx[i] represents the index of the i-th matrix type as specified in Table 75.
[0711] [Table 75]
[0712]
[0713] Note 2. The undefined matrix type is used to indicate a matrix type other than an affine transformation matrix, a covariance matrix, a rotation matrix, a translation matrix, or a compact feature matrix. This can be used by the user to specify the matrix type.
[0714] If gfv_num_matrices_equal_to_num_kps_flag[i] is 1, it indicates that the number of matrices of the i-th matrix type is equal to gfv_num_kps_minus1 + 1. If gfv_num_matrices_equal_to_num_kps_flag[i] is 0, it indicates that the number of matrices of the i-th matrix type is not equal to gfv_num_kps_minus1 + 1.
[0715] gfv_num_matrices_info[i] provides information for deriving the number of matrices of the i-th matrix type.
[0716] gfv_matrix_width_minus1[i] plus 1 represents the width of the matrix of the i-th matrix type.
[0717] gfv_matrix_height_minus1[i] plus 1 represents the height of the matrix of the i-th matrix type.
[0718] If gfv_matrix_for_3D_space_flag[i] is 1, it indicates that the matrix of the i-th matrix type is defined in 3D space. If gfv_matrix_for_3D_space_flag[i] is 0, it indicates that the matrix of the i-th matrix type is defined in 2D space.
[0719] If gfv_matrix_width_minus1[i] does not exist, the following is inferred:
[0720] - If gfv_matrix_type_idx[i] is 0, 1, or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is present and equal to 1, then gfv_matrix_width_minus1[i] is inferred to be equal to 2.
[0721] - Otherwise, if matrix_type_idx[i] is 0, 1, or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is present and equal to 0, then gfv_matrix_width_minus1[i] is inferred to be equal to 1.
[0722] - Otherwise (if matrix_type_idx[i] is 5 or 6), gfv_matrix_width_minus1[i] is inferred to be equal to 0.
[0723] If gfv_matrix_height_minus1[i] does not exist, the following is inferred:
[0724] - If matrix_type_idx is 0, 1, 4, 5, or 6, and either gfv_coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is present and equal to 1, then gfv_matrix_height_minus1[i] is inferred to be equal to 2.
[0725] - Otherwise (gfv_matrix_type_idx is 0, 1, 4, 5, or 6, and either gfv_coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is 0), gfv_matrix_height _minus1[i] is inferred to be equal to 1.
[0726] The variables matrixWidth[i] and matrixHeight[i], which represent the width and height of the matrix of the i-th matrix type, are derived as shown in Table 76 below.
[0727] [Table 76]
[0728]
[0729] gfv_num_matrices_minus1[i] plus 1 represents the number of matrices of the i-th matrix type. The variable numMatrices[i], which represents the number of matrices of the i-th matrix type, is derived as shown in Table 77 below.
[0730] [Table 77]
[0731]
[0732] gfv_matrix_element_int[i][j][k][m] represents the integer part of the matrix element value at position (k, m) of the jth matrix of the ith matrix type.
[0733] gfv_matrix_element_dec[i][j][k][m] represents the decimal part of the matrix element value at position (k, m) of the jth matrix of the ith matrix type.
[0734] gfv_matrix_element_sign_flag[i][j][k][m] indicates the sign of the matrix element at position (k, m) of the jth matrix of the ith matrix type. If gfv_matrix_element_sign_flag[i][j][k][m] does not exist, it is inferred to be equal to 0.
[0735] matrixElementVal[i][j][k][m], which represents the matrix element value at position (k, m) of the jth matrix of the ith matrix type, is derived as shown in Table 78 below.
[0736] [Table 78]
[0737]
[0738] GenerativeNN ( ) is a process for generating sample values of the output picture corresponding to the driving picture. It is called only when gfc_base_pic_flag is 0.
[0739] The input to TranslatorNN() is:
[0740] - sigKeyPoint and sigMatrix
[0741] The output of TranslatorNN() is:
[0742] - convKeyPoint and convMatrix
[0743] The input to GenerativeNN() is:
[0744] - If gfv_base_pic_flag is 0, gfv_drive_pic_fusion_flag is 0, and ChromaFormatIdc is 0: inputBaseY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0745] - If gfv_base_pic_flag is 0, gfv_drive_pic_fusion_flag is 0, and ChromaFormatIdc is non-zero: inputBaseY, inputBaseCb, inputBaseCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0746] - If gfv_base_pic_flag is 0, gfv_drive_pic_fusion_flag is 1, and ChromaFormatIdc is 0: inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0747] - If gfv_base_pic_flag is 0, gfv_drive_pic_fusion_flag is 1, and ChromaFormatIdc is non-zero: inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix
[0748] The output of GenerativeNN( ) is:
[0749] - Luna sample array genY
[0750] - If ChromaFormatIdc is not 0, there are two chroma sample arrays genCb and genCr.
[0751] The following process, shown in Table 79, is used to generate a video picture.
[0752] [Table 79]
[0753]
[0754] To derive the input of TranslatorNN(), the process DeriveSigParam() is specified as follows:
[0755] The keypoint coordinate array sigKeyPoint and matrix sigMatrix are derived as shown in Table 80 below.
[0756] [Table 80]
[0757]
[0758] The input values for GenerativeNN( ) are real numbers, and the functions InpY( ) and InpC( ) are specified as shown in Table 81 below.
[0759] [Table 81]
[0760]
[0761] The output values in GenerativeNN( ) are real numbers, and the functions OutY( ) and OutC( ) are specified as shown in Table 82 below.
[0762] [Table 82]
[0763]
[0764] To derive the input of GenerativeNN ( ), the process DeriveInputTensors( ) is specified as follows:
[0765] If gfv_base_pic_flag is 1, the BasePicture input tensors inputBaseY, inputBaseCb, and inputBaseCr are derived as shown in Table 83 below.
[0766] [Table 83]
[0767]
[0768] If gfv_drive_pic_fusion_flag is 1, the DrivePicture luma sample arrays inputDriveY, inputDriveCb, and inputDriveCr are derived as shown in Table 84 below.
[0769] [Table 84]
[0770]
[0771] If gfv_base_pic_flag is 0, the keypoint coordinate array inputDriveKeyPoint and matrix inputDriveMatrix for the current picture are derived as shown in Table 85 below.
[0772] [Table 85]
[0773]
[0774] If gfv_base_pic_flag is 1, the keypoint coordinate array inputBaseKeyPoint and matrix inputBaseMatrix for the base picture are derived as shown in Table 86 below.
[0775] [Table 86]
[0776]
[0777] To derive the output, the process StoreOutputTensors( ) is specified as follows:
[0778] If gfv_base_pic_flag is 0, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as shown in Table 87 below.
[0779] [Table 87]
[0780]
[0781] If gfv_base_pic_flag is 1, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as shown in Table 88.
[0782] [Table 88]
[0783]
[0784] Example 6
[0785] Traditional video codecs aim to minimize pixel-level distortion to optimize for human viewing. Minimizing pixel-level distortion may not be optimal for machine learning. This is because machine learning relies on features rather than pixel-level details. Features represent various aspects of input data, such as edges, textures, shapes, or high-level semantic information. They can also be the output of neural networks representing specific features of the input picture. Research on feature-based optimization for machine learning is reported as follows.
[0786] 1) Feature-based RDO. This method optimizes rate-distortion by calculating distortion using extracted features. To extract features, a pretrained model can be used, or a machine-perception-aware metric model can be defined through training.
[0787] 2) Feature-based preprocessing. This method minimizes the loss of important features and reduces unnecessary information for machine analysis. This makes compression more suitable and reduces bitrate without negatively impacting machine analysis.
[0788] The Encoder Optimization Information SEI message is used to indicate whether a video is optimized for user viewing or machine analysis, or what type of optimization was applied during preprocessing or encoding. The encoder can apply feature-based optimization to improve the quality of regions of the decoded picture with significant features compared to regions with few or no significant features. The presence of significant features may vary depending on the feature extraction method. Feature extraction methods can be defined based on the type and objectives of the task, which affect both user and machine perception. Therefore, providing information about feature-based optimization in the bitstream of the coded video can be useful.
[0789] If eoi_cancel_flag is 1, it indicates that the persistence of the encoder optimization information SEI message contained in the previous PU in the output order is canceled. If eoi_cancel_flag is 0, it indicates that information about optimizations provided during preprocessing or encoding follows.
[0790] eoi_persistence_flag specifies the persistence of the optimization information applied to this SEI message. If eoi_persistence_flag is 0, it specifies that the optimization information applies only to the current picture. If eoi_persistence_flag is 1, it specifies that the optimization information applies to the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions are true:
[0791] - A new CLVS of the current layer begins.
[0792] - The bitstream ends.
[0793] - A picture in the current layer in an AU associated with a content protection information SEI message is output that follows the current picture in output order.
[0794] If eoi_for_human_viewing_flag is 1, it specifies that the purpose for the applied optimization includes human viewing. If eoi_for_human_viewing_flag is 0, it specifies that the purpose for the applied optimization does not include human viewing.
[0795] If eoi_for_machine_analysis_flag is 1, it specifies that the objective for the applied optimization includes machine analysis. If eoi_for_machine_analysis_flag is 0, it specifies that the objective for the applied optimization may or may not include machine analysis.
[0796] eoi_type represents the types of optimization methods specified in Table 89 below.
[0797] [Table 89]
[0798]
[0799] Here, if (eoi_type & bitMask) is not 0, it indicates that the optimization type containing the bitmask value in Table 89 was applied. If eoi_type is greater than 0 and (eoi_type & bitMask) is 0, the optimization type containing the bitmask value was not applied. If eoi_type is 0, the optimization determined by the application was used.
[0800] The variables EoiObjectBasedFlag, EoiTemporalResamplingFlag, EoiSpatialResamplingFlag, EoiTemporalQualityFlag, EoiSpatialQualityFlag, and EoiFeatureBasedFlag, which specify whether eoi_type represents an optimization type including object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and feature-based optimization, respectively, can be derived as shown in Table 90 below.
[0801] [Table 90]
[0802]
[0803] NOTE - For example, when certain top-level temporal sub-layers are encoded with coarse quantization where quality variations are annoying to human viewers but do not affect machine task performance, eoi_for_human_viewing_flag and eoi_for_machine_analaysis_flag can be set to 0 and 1 respectively, and eoi_type can be set to a value such that EoiTemporalQualityFlag is 1.
[0804] If eoi_persistence_flag is 0, it is a bitstream conformance requirement that EoiTemporalResamplingFlag must be 0 and EoiTemporalQualityFlag must be 0.
[0805] eoi_object_based_idc, if present, indicates the type of object-based optimization specified in Table 91.
[0806] [Table 91]
[0807]
[0808] Here, a non-zero value of (eoi_object_based_idc & bitMask) indicates that the object-based optimization type associated with the bitmask value in Table 91 has been applied. If eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) is 0, the object-based optimization type associated with the bitmask value has not been applied. If eoi_object_based_idc is 0, the application-defined type of object-based optimization has been applied. In a bitstream conforming to this specification, the value of eoi_object_based_idc shall be in the range 0 to 7, inclusive. The values 8 to 65,535 for eoi_object_based_idc are reserved for future use in ITU-T | ISO / IEC and shall not be present in a bitstream conforming to this specification. Decoders conforming to this specification must ignore eoi_object_based_idc when the eoi_object_based_idc value is in the range 8 to 65,535.
[0809] If eoi_temporal_resampling_type_flag is 0, it specifies that the temporal resampling optimization is a subsampling operation. If eoi_temporal_resampling_type_flag is 1, it specifies that the temporal resampling optimization is an upsampling operation.
[0810] If eoi_num_int_pics is greater than 0, it indicates that the encoding system has a constant count of pictures excluded (if eoi_temporal_resampling_type_flag is 0) or added (if eoi_temporal_resampling_type_flag is 1) between each pair of coded pictures in output order within the persistence of this message. If eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the count of pictures excluded by the encoding system between each pair of coded pictures in output order. If eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the count of pictures added by the encoding system between each pair of source pictures in encoding.
[0811] If eoi_num_int_pics is 0, it indicates that the count of pictures that the encoding system excludes between each pair of coded pictures in output order (if eoi_temporal_resampling_type_flag is 0) or adds between each pair of source pictures for encoding (if eoi_temporal_resampling_type_flag is 1) within this SEI message is unknown or variable.
[0812] The eoi_num_int_pics value must be in the range 0 to 63.
[0813] eoi_object_based_idc, if present, indicates the type of object-based optimization specified in Table 91. Here, if (eoi_object_based_idc & bitMask) is non-zero, it indicates that the object-based optimization type associated with the bitmask value in Table 91 is applied.
[0814] eoi_feature_optimization_type_idc can be defined to identify a single optimization method as shown in Table 92, or it can be defined to identify multiple optimization methods via bitmasks as shown in Table 93.
[0815] [Table 92]
[0816]
[0817] [Table 93]
[0818]
[0819] eoi_feature_optimization_type_idc, if present, indicates the type of feature-based optimization specified in Table 92.
[0820] eoi_feature_optimization_type_idc, if present, indicates the type of feature-based optimization specified in Table 93. Here, if (eoi_feature_optimization_type_idc & bitMask) is not 0, it indicates that feature-based optimization related to the bitmask value in Table 93 was applied.
[0821] If eoi_partial_feature_use_flag is 0, it indicates that all features were used for optimization. If eoi_partial_feature_use_flag is 1, it indicates that only some features were used for optimization.
[0822] Table 94 shows an example of the encoder optimization information SEI message syntax.
[0823] [Table 94]
[0824]
[0825] Example 7
[0826] Tables 95 and 96 illustrate examples of how to express optimization objectives, properties, and scopes of application. This embodiment includes a method of representing optimization objectives and status with a single identifier instead of a Flag (e.g., optimization_for_machine_analysis_flag, optimization_human_viewing_flag in the second embodiment) defined for each optimization objective.
[0827] [Table 95]
[0828]
[0829] [Table 96]
[0830]
[0831] optimization_cancel_flag can indicate that the persistence of a previously applied optimization is canceled if its value is 1. If its value is 0, it can indicate that optimization_persistence_flag, optimization_for_machine_analysis_flag, and optimization_type are defined subsequently, and that the optimizations and scopes identified by optimization_persistence_flag, optimization_for_machine_analysis_flag, and optimization_type defined subsequently can be applied.
[0832] The optimization_persistence_flag can indicate the persistence of the optimization indicated by optimization_type. A value of 0 indicates that the optimization identified by optimization_type may only be applied to the current Picture. A value of 1 indicates that the optimization identified by optimization_type may be applied to the current Picture and all subsequent Pictures.
[0833] optimization_purpose_idc identifies the optimization purpose and optimization status. Optimization purposes can include human viewing and machine analysis. Optimization status can be further categorized into: 1) Unknown suitability for the purpose, 2) Unsuitable, 3) Suitable but not optimized for the purpose, and 4) Suitable and optimized for the purpose.
[0834] Tables 97, 98, 99, 100, and 101 provide examples of optimization_purpose_idc definitions. The defined optimization objectives and states are examples, and new optimization properties can be defined using this structure.
[0835] [Table 97]
[0836]
[0837]
[0838] Table 97 shows an example of applying the four optimization states defined above to two optimization objectives: human viewing and machine analysis. In this case, 16 pieces of information are identified, and optimization_purpose_idc is defined as 4 bits, as shown in Table 95.
[0839] [Table 98]
[0840]
[0841] Table 98 is an example of defining identification information based on the requirement that at least one of the defined optimization objectives must have a clear optimization status. This is because the SEI represents optimization information, and its definition can imply that optimization is being applied for a specific purpose. In this case, seven pieces of information are identified, and optimization_purpose_idc is defined as 3 bits, as shown in Table 96.
[0842] Table 98 is an example and the identification information may vary depending on the level of prerequisites and conditions.
[0843] [Table 99]
[0844]
[0845] Table 99 adds the case where the suitability of optimization is unknown (Index 111) to Table 98. This can be defined when the impact and suitability of the applied optimization method on human viewing and machine analysis are not clearly known. This can be defined when the viewing conditions of the receiver or the type of machine analysis to be performed are unclear. Alternatively, even when the optimization method is defined by the application, the encoder or transmitter may not be able to determine the impact, and thus define the corresponding identifier.
[0846] [Table 100]
[0847]
[0848] Table 100 excludes cases in Table 99 where optimization was performed for both human viewing and machine analysis purposes (index 000 in Table 99).
[0849] [Table 101]
[0850]
[0851] Table 101 excludes cases where optimizations are optimized for both human viewing and machine analysis purposes, and cases where the suitability of an optimization is unknown.
[0852] Tables 99, 100, and 101 exclude or include specific identifiers based on the identifier definition conditions in Table 98. This can be applied equally to Table 97, where specific identifiers can be excluded or included individually or multiple identifiers simultaneously. Furthermore, it is self-evident that the optimization objectives and status assigned to each identifier index can be changed. That is, the order of the indexes defined in Tables 97, 98, 99, 100, and 101 are examples and can be changed. This can also be defined based on the importance of the optimization objective, utilization ratio, etc.
[0853] optimization_type indicates the properties of the optimization method.
[0854] The optimization properties defined in Table 102 and Table 103 can be identified.
[0855] [Table 102]
[0856]
[0857] [Table 103]
[0858]
[0859] The optimization properties defined in Tables 102 and 103 are examples, and new optimization properties can be defined using this structure. It is of course possible to identify optimization properties and methods by including only some of the optimization properties defined in Tables 102 and 103, or by including other optimization properties and methods not defined in Tables 102 and 103. Tables 102 and 103 represent examples that perform the same function but are expressed differently.
[0860] Table 103 is another example of expressing multiple optimization methods via a bit mask for optimization_type, and Tables 104 and 105 show examples.
[0861] [Table 104]
[0862]
[0863] [Table 105]
[0864]
[0865] optimizationForPIIProtectionFlag indicates whether personal information protection optimization is included in the applied optimization method. A value of 1 indicates that personal information protection optimization is included in the applied optimization method, while a value of 0 indicates that it is not included in the applied optimization method.
[0866] FIG. 21 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.
[0867] The operations illustrated in FIG. 21 do not correspond to essential components of the decoding method according to an embodiment, and at least some of the operations illustrated in FIG. 21 may be omitted or other operations not illustrated in FIG. 21 may be added.
[0868] The terms or names described in FIG. 21 (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms described in FIG. 21. For example, the image information described in FIG. 21 may include various information according to the embodiments described in the present disclosure, and may include information described in at least one of Tables 1 to 105.
[0869] The operations illustrated in FIG. 21 may be executed by a decoding device including a memory and a processor electrically connected to the memory, and may be executed by, for example, a processor.
[0870] The decoding device can obtain encoder optimization information (S2110).
[0871] For example, a processor of a decoding device can obtain image information and obtain encoder optimization information from the image information.
[0872] Encoder optimization information can take on various forms or be given various names. Furthermore, encoder optimization information can take on various forms or be given various names.
[0873] For example, encoder optimization information may be a syntax element or a syntax structure containing one or more syntax elements. For example, encoder optimization information may be expressed in various ways, such as encoder_optimization_info( ), but is not limited thereto. Hereinafter, encoder optimization information is described as encoder_optimization_info, but is not limited thereto.
[0874] Encoder optimization information (encoder_optimization_info) may include syntax elements such as optimization persistence cancellation information (eg eoi_cancel_flag), optimization persistence information (eg eoi_persistence_flag), optimization information for human viewing (eg eoi_for_human_viewing_flag), optimization information for machine analysis (eg eoi_for_machine_analysis_flag), optimization type information (eg eoi_type), feature-based optimization type information (eg eoi_feature_optimization_type), and / or partial feature use information (eg eoi_partial_feature_use_flag).
[0875] Syntax elements or information included in encoder optimization information (e.g. encoder_optimization_info) may be in various forms, such as a 1-bit flag, an indicator (idc) of 2 or more bits, or a string of characters / numbers.
[0876] The decoding device can determine the type of encoder optimization (S2120).
[0877] For example, a processor of a decoding device may determine the type of encoder optimization based on information about the encoder optimization, particularly optimization type information. The type of encoder optimization may include object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization.
[0878] Encoder optimization information may include optimization type information indicating the type of encoder optimization. The optimization type information may include information regarding whether object-based optimization is performed, information regarding whether temporal resampling is performed, information regarding whether spatial resampling is performed, information regarding whether temporal quality optimization is performed, information regarding whether spatial quality optimization is performed, and / or information regarding whether feature-based optimization is performed.
[0879] Optimization type information can be expressed in various ways, such as optimization_type or eoi_type, but is not limited thereto. Hereinafter, optimization type information is expressed as eoi_type, but is not limited thereto. Furthermore, optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of letters / numbers.
[0880] The processor can identify the type of encoder optimization based on the optimization type information. For example, the processor can identify whether the optimization is object-based, temporal resampling, spatial resampling, temporal quality, spatial quality, and / or feature-based based on a logical AND between the optimization type information and a predetermined bitmask.
[0881] Here, the type of encoder optimization may include feature-based optimization. Features represent information about various aspects of the input data, such as edges, textures, shapes, or high-level semantic information. Features may also be the output of a neural network representing specific features of the input picture.
[0882] The processor can identify whether the encoder optimization includes feature-based optimization based on the optimization type information. As shown in Table 89 described above, the processor can identify whether the encoder optimization includes feature-based optimization based on a logical AND between the optimization type information and the predetermined bitmask 0x20. In addition, as shown in Table 90 described above, the processor can obtain the value of the variable EoiFeatureBasedFlag based on the logical AND between the optimization type information and the predetermined bitmask 0x20, and identify whether the encoder optimization includes feature-based optimization based on the value of the variable EoiFeatureBasedFlag.
[0883] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization and some feature usage information indicating whether some or all features are used for optimization.
[0884] Feature-based optimization may include at least one of feature-based rate-distortion optimization (RDO) for machine learning, feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control.
[0885] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization. The feature-based optimization type information may include information regarding whether feature-based rate-distortion optimization (RDO) is performed, whether feature-based preprocessing is performed, and / or whether feature-based quantization parameter (QP) adaptation / rate control is performed.
[0886] Feature-based optimization type information can be expressed in various ways, such as eoi_feature_optimization_type, but is not limited thereto. Furthermore, feature-based optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of letters / numbers.
[0887] The processor can identify the type of feature-based optimization based on the feature-based optimization type information. For example, the processor can identify whether feature-based rate-distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control is performed based on the value of the feature-based optimization type information. As another example, the processor can identify whether feature-based rate-distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control is performed based on a logical product between the feature-based optimization type information and a predetermined bitmask.
[0888] Feature-based optimization can use some or all features for optimization.
[0889] Encoder optimization information may include feature usage information indicating whether some or all features are used for optimization. For example, encoder optimization information may include some or all feature usage information.
[0890] Some feature usage information may indicate in feature-based optimization whether some features or all features are used for optimization.
[0891] Some feature usage information can be expressed in various ways, such as eoi_partial_feature_use_flag, but is not limited thereto. Furthermore, some feature usage information, such as feature-based optimization type information, can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of characters / numbers.
[0892] The processor can identify whether some features or all features are used for optimization based on some feature usage information. For example, the processor can identify that all features are used for optimization based on the value of some feature usage information being 0. Alternatively, the processor can identify that some features are used for optimization based on the value of some feature usage information being 1.
[0893] Unlike the example described above, if the encoder optimization information includes all feature usage information, the processor can identify that all features are used for optimization based on the value of all feature usage information being 1. Additionally, the processor can identify that some features are used for optimization based on the value of all feature usage information being 0.
[0894] The decoding device can process the image (S2130).
[0895] For example, a processor of a decoding device may process an image based on the type of encoder optimization. For example, the processor of the decoding device may perform optimization for user viewing and / or optimization for machine analysis. The processor may perform object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization. The processor may perform feature-based rate-distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control. Additionally, the processor may utilize some or all features for feature-based optimization.
[0896] As described above, the decoding device can determine whether feature-based optimization is performed based on encoding optimization information, and can process the image based on the type of feature-based optimization and / or the use of optimization for some or all features. Thus, the VCM system can provide feature-based optimization for machine learning.
[0897] FIG. 22 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.
[0898] The operations illustrated in FIG. 22 do not correspond to essential components of the encoding method according to an embodiment, and at least some of the operations illustrated in FIG. 22 may be omitted or other operations not illustrated in FIG. 22 may be added.
[0899] The terms or names described in FIG. 22 (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms described in FIG. 22. For example, the image information described in FIG. 22 may include various information according to the embodiments described in the present disclosure, and may include information described in at least one of Tables 1 to 105.
[0900] The operations illustrated in FIG. 22 may be executed by a decoding device including a memory and a processor electrically connected to the memory, and may be executed by, for example, a processor.
[0901] The encoding device can determine the type of encoder optimization (S2210).
[0902] For example, a processor included in an encoding device may determine the type of encoder optimization based on the characteristics of the video and / or the use of the video. For example, the processor may determine whether to optimize for user viewing and / or for machine analysis based on the use of the video. The processor may determine whether to perform object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization based on the characteristics of the video. The processor may determine whether to perform feature-based rate-distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control. Furthermore, the processor may determine whether to utilize some or all features for feature-based optimization.
[0903] The encoding device can process images (S2220).
[0904] For example, a processor of an encoding device may perform image processing for optimization based on the type of encoder optimization. For example, encoder optimization may include optimization for user viewing and / or optimization for machine analysis. Encoder optimization may include object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization. Feature-based optimization may include feature-based rate-distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control. Additionally, feature-based optimization may use some or all features.
[0905] The processor of the encoding device can perform optimization for user viewing and / or optimization for machine analysis. The processor can perform object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization. The processor can perform feature-based rate-distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control. Additionally, the processor can utilize some or all features for feature-based optimization.
[0906] The encoding device can generate encoder optimization information (S2230).
[0907] For example, a processor of an encoding device can generate encoder optimization information based on a type of encoder optimization.
[0908] Encoder optimization information may include optimization type information indicating the type of encoder optimization. The optimization type information may include information regarding whether object-based optimization is performed, information regarding whether temporal resampling is performed, information regarding whether spatial resampling is performed, information regarding whether temporal quality optimization is performed, information regarding whether spatial quality optimization is performed, and / or information regarding whether feature-based optimization is performed.
[0909] Optimization type information can be expressed in various ways, such as optimization_type or eoi_type, but is not limited thereto. Hereinafter, optimization type information is expressed as eoi_type, but is not limited thereto. Furthermore, optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of letters / numbers.
[0910] The processor may generate optimization type information based on the type of encoder optimization. For example, the processor may generate optimization type information based on whether the optimization is object-based, temporal resampling, spatial resampling, temporal quality, spatial quality, and / or feature-based.
[0911] Here, the type of encoder optimization may include feature-based optimization. Features represent information about various aspects of the input data, such as edges, textures, shapes, or high-level semantic information. Features may also be the output of a neural network representing specific features of the input picture.
[0912] The processor can generate optimization type information based on whether the encoder optimization includes feature-based optimization.
[0913] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization and some feature usage information indicating whether some or all features are used for optimization.
[0914] Feature-based optimization may include at least one of feature-based rate-distortion optimization (RDO) for machine learning, feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptation / rate control.
[0915] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization. The feature-based optimization type information may include information regarding whether feature-based rate-distortion optimization (RDO) is performed, whether feature-based preprocessing is performed, and / or whether feature-based quantization parameter (QP) adaptation / rate control is performed.
[0916] Feature-based optimization type information can be expressed in various ways, such as eoi_feature_optimization_type, but is not limited thereto. Furthermore, feature-based optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of letters / numbers.
[0917] The processor may generate feature-based optimization type information based on the type of feature-based optimization. For example, the processor may generate feature-based optimization type information based on whether feature-based rate-distortion optimization (RDO) is performed, whether feature-based preprocessing is performed, and / or whether feature-based quantization parameter (QP) adaptation / rate control is performed.
[0918] Feature-based optimization can use some or all features for optimization.
[0919] Encoder optimization information may include feature usage information indicating whether some or all features are used for optimization. For example, encoder optimization information may include some or all feature usage information.
[0920] Some feature usage information may indicate in feature-based optimization whether some features or all features are used for optimization.
[0921] Some feature usage information can be expressed in various ways, such as eoi_partial_feature_use_flag, but is not limited thereto. Furthermore, some feature usage information, such as feature-based optimization type information, can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of characters / numbers.
[0922] The processor may generate some feature usage information based on whether some features or all features are used for optimization. For example, the processor may set the value of some feature usage information to 0 based on whether all features are used for optimization. Alternatively, the processor may set the value of some feature usage information to 1 based on whether some features are used for optimization.
[0923] Unlike the example described above, if the encoder optimization information includes all feature usage information, the processor may set the values of all feature usage information to 1 based on whether all features are used for optimization. Alternatively, the processor may set the values of all feature usage information to 0 based on whether some features are used for optimization.
[0924] The encoding device can encode image information (S2240).
[0925] For example, a processor of an encoding device can encode image information that includes encoder optimization information.
[0926] Encoded video information is converted into a bitstream, and the bitstream can be stored in a storage medium or transmitted through a transmission device.
[0927] As described above, the encoding device can optimize an image based on the type of feature-based optimization and / or the optimization use of some or all features, and can generate encoding optimization information based on whether feature-based optimization is used. Furthermore, the encoding device can encode the encoding optimization information. Thus, the VCM system can provide feature-based optimization for machine analysis.
[0928] FIG. 23 is a diagram illustrating an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0929] Referring to FIG. 23, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0930] The encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data, generates a bitstream, and transmits it to a streaming server. Alternatively, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.
[0931] A bitstream can be generated by a video encoding method and / or a video encoding device to which an embodiment of the present disclosure is applied, and a streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0932] A streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary, informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. The content streaming system may include a separate control server, in which case the control server may control commands and responses between each device within the content streaming system.
[0933] A streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0934] Examples of user devices include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.
[0935] Each server within a content streaming system can be operated as a distributed server, in which case the data received from each server can be processed in a distributed manner.
[0936] FIG. 24 is a diagram showing another example of a content streaming system to which embodiments of the present disclosure can be applied.
[0937] Referring to FIG. 24, in an embodiment such as VCM, depending on the performance of the device, the user's request, the characteristics of the task to be performed, etc., the task may be performed on the user terminal or on an external device (e.g., a streaming server, an analysis server, etc.). In this way, in order to transmit the information necessary for performing the task to the external device, the user terminal may directly or through an encoding server generate a bitstream containing the information necessary for performing the task (e.g., information such as the task, neural network, and / or purpose).
[0938] The analysis server can decrypt encoded information received from the user terminal (or encoding server) and then perform a task requested by the user terminal. The analysis server can transmit the results obtained through the task performance back to the user terminal or to another linked service server (e.g., a web server). For example, the analysis server can perform a task to determine fire and transmit the results obtained to a fire-related server. The analysis server may include a separate control server, in which case the control server can control commands / responses between each device associated with the analysis server and the server. In addition, the analysis server can request desired information from the web server based on information about the tasks the user device wishes to perform and can perform. When the analysis server requests a desired service from the web server, the web server can transmit this request to the analysis server, and the analysis server can transmit data corresponding to the request to the user terminal. In this case, the control server of the content streaming system can control commands / responses between each device within the streaming system.
[0939] In this disclosure, the names of the syntax elements described above are arbitrarily assigned for clarity of explanation and are not intended to limit the names of the syntax elements. Furthermore, each syntax element may also be referred to as information. Furthermore, while syntax elements may be obtained from a bitstream, they may also be derived from other syntax elements, which may also be included in embodiments of the present disclosure.
[0940] Additionally, the bitstream generated by the image encoding method may be stored in a non-transitory computer-readable recording medium.
[0941] Additionally, as another example, a bitstream generated by a video encoding method may be transmitted to another device (e.g., a video decoding device, etc.). In this case, the method for transmitting the bitstream may include a process for transmitting the bitstream.
[0942] While the exemplary methods of this disclosure are presented as a series of operations for clarity of description, this is not intended to limit the order in which the steps are performed, and individual steps may be performed simultaneously or in different orders, if desired. To implement a method according to this disclosure, additional steps may be included in addition to the steps illustrated, some steps may be excluded and the remaining steps included, or some steps may be excluded and additional steps included.
[0943] In the present disclosure, a video encoding device or video decoding device performing a predetermined operation (step) may perform an operation (step) of checking the conditions or circumstances under which the operation (step) is performed. For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the video encoding device or video decoding device may perform an operation of checking whether the predetermined condition is satisfied and then perform the predetermined operation.
[0944] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0945] The embodiments described in this disclosure may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units depicted in each drawing may be implemented and performed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.
[0946] In addition, the decoder (decoding device) and encoder (encoding device) to which the embodiment(s) of the present disclosure are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argument reality) device, a video phone video device, a transportation terminal (e.g., a vehicle (including an autonomous vehicle) terminal, a robot terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, OTT video (Over the top video) devices may include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, and DVRs (Digital Video Recorders).
[0947] In addition, the processing method to which the embodiment(s) of the present disclosure are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0948] Additionally, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0949] Embodiments according to the present disclosure can be used to encode / decode video.
Claims
1. As a method of decoding image information, Obtaining the image information including encoder optimization information; Determining the type of encoder optimization based on the above encoder optimization information; Processing the image based on the type of the above encoder optimization, The above encoder optimization information includes optimization type information indicating the type of the encoder optimization, A method wherein the above encoder optimization includes feature-based optimization that optimizes an image based on features of the image.
2. In paragraph 1, A method in which whether or not the above feature-based optimization is performed is determined based on the logical product of the above optimization type information and a predetermined first bitmask.
3. In paragraph 1, A method wherein the above encoder optimization information further includes feature-based optimization type information indicating a type of the feature-based optimization.
4. In paragraph 3, A method wherein the feature-based optimization comprises at least one of feature-based rate-distortion optimization, feature-based preprocessing or feature-based quantization parameter (QP) adaptation / rate control based on the value of the feature-based optimization type information.
5. In paragraph 3, A method wherein the feature-based optimization comprises at least one of feature-based rate-distortion optimization, feature-based preprocessing or feature-based quality adaptation / rater control based on a logical product of the feature-based optimization type information and a predetermined second bitmask.
6. In paragraph 1, A method wherein the encoder optimization information further includes some feature usage information indicating whether some or all features are used for optimization.
7. In paragraph 1, A method wherein the encoder optimization information further includes optimization suitability information indicating whether the encoder optimization is suitable for user viewing or machine analysis.
8. In paragraph 7, A method wherein the above optimization suitability information includes a 3-bit or 4-bit indicator indicating suitability of user viewing and suitability of machine analysis.
9. As a method of encoding video, Determine the type of encoder optimization; Processing an image based on the type of encoder optimization described above; Generate encoder optimization information based on the type of the above encoder optimization; Encoding image information including the above encoder optimization information, The above encoder optimization information includes optimization type information indicating the type of the encoder optimization, A method wherein the above encoder optimization includes feature-based optimization that optimizes an image based on features of the image.
10. A method for storing a bitstream of video information in a computer-readable storage medium, Obtaining a bitstream of the above image information; Comprising storing data including the bitstream in the storage medium, The above image information includes encoder optimization information, The above encoder optimization information is generated based on the type of encoder optimization determined according to the image. The above encoder optimization information includes optimization type information indicating the type of the encoder optimization, A method wherein the above encoder optimization includes feature-based optimization that optimizes an image based on features of the image.
11. A method for transmitting a bitstream of video information, Obtaining a bitstream of the above image information; Including transmitting data including the above bitstream, The above image information includes encoder optimization information, The above encoder optimization information is generated based on the type of encoder optimization determined according to the image. The above encoder optimization information includes optimization type information indicating the type of the encoder optimization, A method wherein the above encoder optimization includes feature-based optimization that optimizes an image based on features of the image.
Citation Information
Patent Citations
Photocatalyst composition, method for manufacturing thereof and method for coating using sam
KR1020250056489A
Video coding based on feature extraction and picture synthesis
US20230343099A1
Systems and methods for joint optimization training and encoder side downsampling
WO2023023229A1
A method, an apparatus and a computer program product for video encoding and video decoding
WO2023194651A1
Cited By
Systems and methods for signaling encoder optimization privacy information in video coding
US12641294B2
Systems and methods for signaling encoder optimization privacy information in video coding
US20260107019A1