Methods for decoding image information, methods for encoding images, methods for storing image information bitstreams, and methods for transmitting image information bitstreams.

By using feature-based optimization methods to decode and encode image information, the problem of increased costs in high-resolution image transmission and storage is solved, achieving efficient image compression and decoding.

CN122498144APending Publication Date: 2026-07-31LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-01-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

With the increasing demand for high-resolution, high-quality images, existing technologies face the problem of rising costs due to the increased amount of information in image data transmission and storage, necessitating efficient image compression technologies.

Method used

Image information is decoded, encoded, and transmitted as a bitstream using feature-based optimization methods. Image processing is performed using encoder optimization information and feature information, including feature-based rate distortion optimization, preprocessing, and adaptive rate control of quantization parameters.

Benefits of technology

It achieves efficient encoding and decoding of image information, reduces transmission and storage costs, and is suitable for the needs of human viewing and machine analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122498144A_ABST
    Figure CN122498144A_ABST
Patent Text Reader

Abstract

A method for decoding image information may include the following steps: obtaining image information including encoder optimization information; determining the type of encoder optimization based on the encoder optimization information; and processing the image based on the type of encoder optimization, wherein the encoder optimization information includes optimization type information indicating the type of encoder optimization, and the encoder optimization includes feature-based optimization that optimizes the image based on its features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to methods for decoding image information, methods for encoding image information, and methods for transmitting bit streams of image information. Background Technology

[0002] In recent years, the demand for high-resolution, high-quality images (such as HD (High Definition) and UHD (Ultra High Definition) images has been increasing across various fields. As the resolution and quality of image data become higher, the amount of information or bits transmitted increases relative to conventional image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0003] Therefore, efficient image compression technology is needed to effectively send, store, and reproduce high-resolution, high-quality image information. Summary of the Invention

[0004] Technical issues

[0005] This disclosure aims to provide a method for decoding image information based on features, a method for encoding image information, and a method and apparatus for transmitting bit streams of image information.

[0006] Technical solution

[0007] According to one aspect of this disclosure, a method for decoding image information may include the following steps: obtaining image information including encoder optimization information, determining the type of encoder optimization based on the encoder optimization information, and processing the image based on the type of encoder optimization. The encoder optimization information may include optimization type information indicating the type of encoder optimization. The encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0008] According to one aspect of this disclosure, an apparatus for decoding image information may include a memory and a processor connected to the memory. The processor may include: acquiring image information including encoder optimization information, determining a type of encoder optimization based on the encoder optimization information, and processing the image based on the type of encoder optimization. The encoder optimization information may include optimization type information indicating the type of encoder optimization. The encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0009] Whether to apply feature-based optimization can be determined based on the logical product (AND) of the optimization type information and the predefined first-order mask.

[0010] Encoder optimization information may also include feature-based optimization type information that indicates the type of feature-based optimization.

[0011] Based on the value of the feature-based optimization type information, feature-based optimization may include at least one of feature-based rate distortion optimization, feature-based preprocessing, or feature-based quantization parameter (QP) adaptation / rate control.

[0012] Based on the logical product (AND) of the feature-based optimization type information and the predefined second bit mask, feature-based optimization can include at least one of feature-based rate distortion optimization, feature-based preprocessing, or feature-based quality adaptation / rate control.

[0013] Encoder optimization information may also include partial feature usage information indicating whether some features or all features were used for optimization.

[0014] Encoder optimization information may also include optimization suitability information that indicates whether the encoder optimization is suitable for human viewing or machine analysis.

[0015] Optimization suitability information may include 3-bit or 4-bit indicators that indicate suitability for human viewing and suitability for machine analysis.

[0016] According to one aspect of this disclosure, a method for encoding an image may include the following steps: determining a type of encoder optimization, processing the image based on the type of encoder optimization, generating encoder optimization information based on the type of encoder optimization, and encoding image information including the encoder optimization information. The encoder optimization information may include optimization type information indicating the type of encoder optimization. Encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0017] According to one aspect of this disclosure, an apparatus for encoding an image may include a memory and a processor connected to the memory. The processor may include: determining a type of encoder optimization; processing the image based on the type of encoder optimization; generating encoder optimization information based on the type of encoder optimization; and encoding image information including the encoder optimization information. The encoder optimization information may include optimization type information indicating the type of encoder optimization. Encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0018] According to one aspect of this disclosure, a method for storing a bitstream of image information in a computer-readable storage medium may include the steps of: obtaining a bitstream of image information, and storing data including the bitstream in the storage medium. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined according to the image. The encoder optimization information may include optimization type information indicating the type of encoder optimization. Encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0019] According to one aspect of this disclosure, a computer-readable storage medium can store data including a bitstream of image information. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined from the image. The encoder optimization information may include optimization type information indicating the type of encoder optimization. The encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0020] According to one aspect of this disclosure, a method for transmitting a bitstream of image information may include the steps of: obtaining a bitstream of image information, and transmitting data including the bitstream. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined according to the image. The encoder optimization information may include optimization type information indicating the type of encoder optimization. Encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0021] According to one aspect of this disclosure, an apparatus for transmitting a bitstream of image information may include: a processor for acquiring the bitstream of image information; and a transmitter for transmitting data including the bitstream. The image information may include encoder optimization information, and the encoder optimization information may be generated based on a type of encoder optimization determined from the image. The encoder optimization information may include optimization type information indicating the type of encoder optimization. Encoder optimization may include feature-based optimization for optimizing the image based on its features.

[0022] Beneficial effects

[0023] According to this disclosure, a method for decoding image information based on features, a method for encoding image information, and a method and apparatus for transmitting bit streams of image information can be provided.

[0024] The effects that can be obtained from this disclosure are not limited to those described above, and other effects not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains based on the following description. Attached Figure Description

[0025] Figure 1 This is a schematic diagram illustrating a VCM system to which embodiments of the present disclosure can be applied.

[0026] Figure 2 This is a schematic diagram illustrating a VCM pipeline structure to which embodiments of the present disclosure can be applied.

[0027] Figure 3 This diagram schematically illustrates an image / video encoder to which embodiments of the present disclosure can be applied.

[0028] Figure 4 This diagram schematically illustrates an image / video decoder to which embodiments of the present disclosure can be applied.

[0029] Figure 5 This is a flowchart illustrating a feature / feature map encoding process that can be applied to embodiments of the present disclosure.

[0030] Figure 6 This is a flowchart illustrating a feature / feature map decoding process that can be applied to embodiments of the present disclosure.

[0031] Figure 7 This is a diagram illustrating a feature extraction network to which embodiments of the present disclosure can be applied.

[0032] Figure 8a This is a diagram illustrating an example of the data distribution characteristics of a video source.

[0033] Figure 8b This is a diagram illustrating an example of the data distribution characteristics of a feature set.

[0034] Figure 9 The diagram illustrates the transformation and inverse transformation processes applicable to embodiments of this disclosure.

[0035] Figure 10 This is a diagram illustrating a quantization group according to an embodiment of the present disclosure.

[0036] Figure 11 This is a block diagram illustrating Context Adaptive Binary Arithmetic Encoding (CABAC) for encoding a syntax element.

[0037] Figure 12 and Figure 13 This is a diagram illustrating the entropy coding process.

[0038] Figure 14 and Figure 15 This is a diagram illustrating the entropy decoding process.

[0039] Figure 16 This is a diagram illustrating an example of a machine video coding (VCM) layer architecture.

[0040] Figure 17 This is a diagram illustrating an example of a bitstream consisting of encoded abstract features and neural network abstraction layer (NNAL) information.

[0041] Figure 18 Examples of encoding optimizations and the use of encoded bitstreams involving privacy-preserving processing are illustrated.

[0042] Figure 19 Examples of data attributes and privacy protection information definitions are provided.

[0043] Figure 20 An example is shown where the areas required for image analysis coexist with areas containing personal information.

[0044] Figure 21 This is a diagram illustrating a method for decoding image information according to an embodiment of the present disclosure.

[0045] Figure 22 This is a diagram illustrating a method for encoding image information according to an embodiment of the present disclosure.

[0046] Figure 23 This is a diagram illustrating an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0047] Figure 24 This is a diagram illustrating another example of a content streaming system to which embodiments of the present disclosure can be applied. Detailed Implementation

[0048] In the following description, embodiments of the present disclosure will be detailed with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.

[0049] In describing embodiments of this disclosure, detailed explanations of well-known configurations or functions have been omitted where they are deemed to obscure the essential points of this disclosure. Furthermore, portions irrelevant to the description of this disclosure have been omitted from the accompanying drawings, and similar reference numerals have been assigned to similar portions.

[0050] In this disclosure, when certain components are described as “connected,” “coupled,” or “linked” to another component, this may include not only direct connections but also indirect connections in which another component may be present in between. Additionally, when a component is described as “including” or “having” another component, this means that, unless otherwise expressly stated, it does not exclude other components but may further include additional components.

[0051] In this disclosure, unless otherwise expressly stated, the terms first, second, etc., are used only to distinguish one component from another and do not limit the order or importance of the components. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0052] In this disclosure, distinguishable components are described to clearly explain their respective characteristics, and this does not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed across multiple hardware or software units. Therefore, such integrated or distributed implementations are also included within the scope of this disclosure unless they are explicitly described.

[0053] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments comprising a subset of the components described in one embodiment are also included within the scope of this disclosure. Furthermore, embodiments that include additional components beyond those described in the various embodiments are also included within the scope of this disclosure.

[0054] This disclosure relates to the encoding and decoding of images, and the terms used herein may have their common meanings as commonly used in the art to which this disclosure pertains, unless these terms are redefined in this disclosure.

[0055] This disclosure can be applied to methods disclosed in Versatile Video Coding (VVC) standards and / or Video Coding for Machines (VCM) standards. Additionally, this disclosure can be applied to methods disclosed in Basic Video Coding (EVC) standards, AOMedia Video 1 (AV1) standards, second-generation Audio Video Coding (AVS2) standards, or next-generation video / image coding standards (e.g., H.267 or H.268).

[0056] This disclosure presents various implementations related to video / image encoding and decoding, and unless otherwise stated, these implementations can be combined with each other. In this disclosure, "video" can refer to a collection of images arranged in chronological order. "Image" can be information generated by artificial intelligence (AI). Input information used by AI in performing a series of tasks, information generated during information processing, and output information can be used as images. In this disclosure, "picture" generally refers to a unit indicating a single image at a specific point in time, and a slice / tile is a coding unit that constitutes a part of a picture. A picture can consist of at least one slice / tile. Additionally, a slice / tile can include at least one codec tree unit (CTU). A CTU can be partitioned into at least one CU. A tile is a rectangular region existing within a specific tile row and a specific tile column within a picture, and can consist of multiple CTUs. A tile column can be defined as a rectangular region of a CTU and can have the same height as the picture and a width specified by syntax elements signaled from a bitstream portion (such as a picture parameter set). A tile row can be defined as a rectangular region of a CTU and can have the same width as the image and a height specified by a syntax element signaled from a bitstream portion (such as an image parameter set). Tile scanning is a method of predefined ordering of CTUs within a tile partition. Here, CTUs can be ordered according to the raster scan order of CTUs within a tile, and tiles within an image can be ordered sequentially according to the raster scan order of tiles in the image. A slice can include an integer number of complete tiles or an integer number of sequential complete CTU rows within a tile of an image. Slices can be specifically included in a single NAL unit. An image can consist of at least one tile group. A tile group can include at least one tile. A brick can indicate a rectangular region of a CTU row within a tile of an image. A tile can include at least one brick. A brick can indicate a rectangular region of a CTU row within a tile. A tile can be partitioned into multiple bricks, and each brick can include at least one CTU row belonging to the tile. A tile that is not partitioned into multiple bricks can also be considered a brick.

[0057] In this disclosure, a "pixel" or "cell" can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as the corresponding term for a pixel. A sample can generally indicate a pixel or a pixel value, and can indicate only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0058] In implementations, particularly when applied to VCM, when an image exists consisting of a set of components with different characteristics and meanings, the pixel / pixel value can indicate the pixel / pixel value of the component generated by combining, synthesizing, and analyzing the independent information of each component. For example, in RGB input, it can indicate only the pixel / pixel value of R, only the pixel / pixel value of G, or only the pixel / pixel value of B. For example, it can indicate only the pixel / pixel value of the luminance component synthesized using R, G, and B components. For example, it can indicate only the pixel / pixel value of the information or image extracted by analyzing the R, G, and B components.

[0059] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include at least one of a specific region of an image or information associated with that region. A unit may include a luminance block and two chrominance (e.g., Cb, Cr) blocks. Depending on the context, the term "unit" may be used interchangeably with "sample array," "block," "region," etc. Typically, an M×N block may include a set (or array) of samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows. In one embodiment, particularly when applied to VCM, a unit may refer to a basic unit that includes information for performing a specific task.

[0060] In this disclosure, the term "current block" may refer to one of "current codec block," "current codec unit," "encoding target block," "decoding target block," or "processing target block." When performing prediction, "current block" may refer to "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" may refer to "current transform block" or "transform target block." When performing filtering, "current block" may refer to "filter target block."

[0061] Additionally, in this disclosure, unless explicitly stated as a chroma block, "current block" may refer to "the luminance block of the current block". "The chroma block of the current block" can be expressed by explicitly including an explicit description of a chroma block such as "chroma block" or "current chroma block".

[0062] In this disclosure, " / " and "," can mean "and / or". For example, "A / B" and "A, B" can mean "A and / or B". In addition, "A / B / C" and "A, B, C" can mean "at least one of A, B and / or C".

[0063] In this disclosure, "or" can mean "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in this disclosure, "or" can also mean "additionally or alternatively".

[0064] This disclosure relates to video / image encoding and decoding (VCM) for machines.

[0065] VCM refers to compression techniques that encode / decode a portion of a source image / video or information for use in machine vision. In VCM, the encoding / decoding target can be called a feature. Features can refer to information extracted from the source image / video based on task objectives, requirements, and the surrounding environment. Features can have different information formats than the source image / video, and therefore, feature compression methods and representation formats can also differ from the video source.

[0066] VCMs can be applied to a wide range of applications. For example, in surveillance systems that identify and track objects or people, VCMs can be used to store or transmit object identification information. Additionally, in intelligent transportation or smart mobility systems, VCMs can be used to transmit vehicle location information collected from GPS, sensor information collected from LiDAR, radar, etc., and various vehicle control information to other vehicles or infrastructure. Furthermore, in the field of smart cities, VCMs can be used to perform individual tasks for interconnected sensor nodes or devices.

[0067] This disclosure provides various implementations of feature / feature map encoding and decoding. Unless otherwise specifically stated, the implementations of this disclosure can be implemented individually or in combination of at least two.

[0068] Overview of VCM Systems

[0069] Figure 1 This is a schematic diagram illustrating a VCM system to which embodiments of the present disclosure can be applied.

[0070] refer to Figure 1 The VCM system may include an encoding device 10 and a decoding device 20.

[0071] Encoding device 10 can compress / encode features / feature maps extracted from source images / videos to generate a bitstream, and send the generated bitstream to decoding device 20 via a storage medium or network. Encoding device 10 can also be referred to as a feature encoding device. In a VCM system, features / feature maps can be generated in each hidden layer of a neural network. The size and number of channels in the generated feature map can vary depending on the type of neural network or the location of the hidden layers. In this disclosure, a feature map can be referred to as a feature set, and a feature or feature map can be referred to as "feature information".

[0072] The encoding device 10 may include a feature acquisition unit 11, an encoder 12, and a transmitter 13.

[0073] Feature acquirer 11 can obtain features / feature maps from the source image / video. According to one implementation, feature acquirer 11 can obtain features / feature maps from an external device (e.g., a feature extraction network). In this case, feature acquirer 11 performs a feature receiving interface function. Alternatively, feature acquirer 11 can obtain features / feature maps by executing a neural network (e.g., CNN, DNN, etc.) using the source image / video as input. In this case, feature acquirer 11 performs a feature extraction network function.

[0074] According to one embodiment, the encoding apparatus 10 may further include a source image generator (not shown) for obtaining source images / videos. The source image generator can be implemented using an image sensor, camera module, etc., and the source images / videos can be obtained through a process of capturing, synthesizing, or generating images / videos. In this case, the generated source images / videos can be sent to a feature extraction network and used as input data for extracting features / feature maps.

[0075] Encoder 12 can encode the features / feature maps obtained by feature acquirer 11. Encoder 12 can perform a series of processes, such as prediction, transformation, and quantization, to increase encoding efficiency. The encoded data (encoded feature / feature map information) can be output in the form of a bitstream. The bitstream including the encoded feature / feature map information can be referred to as the VCM bitstream.

[0076] Transmitter 13 can acquire feature / feature map information or data output in bitstream form, and can transmit the acquired information or data to decoding device 20 or another external object in the form of file or streaming via digital storage medium or network. Here, digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files with a predetermined file format, or elements for transmitting data via broadcast / communication network. Transmitter 13 can be provided as a transmission device separate from encoder 12, and in this case, the transmission device can include at least one processor for acquiring feature / feature map information or data output in bitstream form and a transmitter for transmitting it in the form of file or streaming.

[0077] The decoding device 20 can obtain feature / feature map information from the encoding device 10 and reconstruct the feature / feature map based on the obtained information.

[0078] The decoding device 20 may include a receiver 21 and a decoder 22.

[0079] Receiver 21 can receive bit streams from encoding device 10 and obtain feature / feature map information from the received bit streams to send to decoder 22.

[0080] Decoder 22 can decode features / feature maps based on the obtained feature / feature map information. Decoder 22 can perform a series of processes corresponding to the operations of encoder 14, such as dequantization, inverse transform, prediction, etc., to increase decoding efficiency.

[0081] According to one embodiment, the decoding device 20 may further include a task analysis / rendering unit 23.

[0082] The task analysis / rendering unit 23 can perform task analysis based on decoded features / feature maps. Furthermore, the task analysis / rendering unit 23 can render decoded features / feature maps into a form suitable for task execution. Based on the task analysis results and the rendered features / feature maps, various (machine-oriented) tasks can be executed.

[0083] Therefore, a VCM system can encode / decode features extracted from source images / videos based on user and / or machine requests, task objectives, and the surrounding environment, and perform various (machine-oriented) tasks based on the decoded features. A VCM system can also be implemented by extending / redesigning the video / image encoding / decoding system and can execute various encoding / decoding methods defined in the VCM standard.

[0084] VCM production line

[0085] Figure 2 This is a schematic diagram illustrating a VCM pipeline structure to which embodiments of the present disclosure can be applied.

[0086] refer to Figure 2 The VCM pipeline 200 may include a first pipeline 210 for encoding / decoding images / videos and a second pipeline 220 for encoding / decoding features / feature maps. In this disclosure, the first pipeline 210 may be referred to as a video codec pipeline, and the second pipeline 220 may be referred to as a feature codec pipeline.

[0087] The first pipeline 210 may include a first stage 211 for encoding the input image / video and a second stage 212 for decoding the encoded image / video to generate a reconstructed image / video. The reconstructed image / video can be used for human viewing, i.e., human vision.

[0088] The second pipeline 220 may include a third stage 221 for extracting features / feature maps from the input image / video, a fourth stage 222 for encoding the extracted features / feature maps, and a fifth stage 223 for decoding the encoded features / feature maps to generate reconstructed features / feature maps. The reconstructed features / feature maps can be used for machine (vision) tasks. Here, machine (vision) tasks can refer to tasks where machines consume images / videos. Machine vision tasks can be applied to service scenarios such as, for example, surveillance, intelligent transportation, smart cities, smart industries, and smart content. According to one implementation, the reconstructed features / feature maps can also be used for human vision.

[0089] According to the implementation, the features / feature maps encoded in the fourth stage 222 can be sent to the first stage 221 and used to encode the image / video. In this case, an additional bitstream can be generated based on the encoded features / feature maps, and the generated additional bitstream can be sent to the second stage 222 and used to decode the image / video.

[0090] According to the implementation, the features / feature maps decoded in the fifth stage 223 can be sent to the second stage 222 and used to decode the images / videos.

[0091] although Figure 2 The illustration shows a VCM pipeline 200 including a first pipeline 210 and a second pipeline 220, but this is merely exemplary, and embodiments of this disclosure are not limited thereto. For example, the VCM pipeline 200 may include only the second pipeline 220, or the second pipeline 220 may be extended to multiple feature codec pipelines.

[0092] Meanwhile, in the first pipeline 210, the first stage 211 can be executed by an image / video encoder, and the second stage 212 can be executed by an image / video decoder. Additionally, in the second pipeline 220, the third stage 221 can be executed by a VCM encoder (or a feature / feature map encoder), and the fourth stage 222 can be executed by a VCM decoder (or a feature / feature map decoder). The encoder / decoder structure is described in detail below.

[0093] encoder

[0094] Figure 3 This diagram schematically illustrates an image / video encoder to which embodiments of the present disclosure can be applied.

[0095] refer to Figure 3The image / video encoder 300 may include an image partitioner 310, a predictor 320, a residual processor 330, an entropy encoder 340, an adder 350, a filter 360, and a memory 370. The predictor 320 may include an inter-frame predictor 321 and an intra-frame predictor 322. The residual processor 330 may include a transformer 332, a quantizer 333, a dequantizer 334, and an inverse transformer 335. The residual processor 330 may also include a subtractor 331. The adder 350 may be referred to as a reconstructor or a reconstruction block generator. According to one embodiment, the image partitioner 310, predictor 320, residual processor 330, entropy encoder 340, adder 350, and filter 360 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 370 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may further include the memory 370 as an internal / external component.

[0096] Image partitioner 310 can partition an input image (or picture, frame) input to image / video encoder 300 into at least one processing unit. As an example, a processing unit may be referred to as a codec unit (CU). Codec units can be recursively partitioned from codec tree units (CTUs) or maximum codec units (LCUs) according to a quadtree-binary-trinary tree (QTBTTT) structure. For example, a codec unit can be partitioned into multiple codec units with greater depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary structure. Alternatively, a binary tree structure can be applied first. The image / video encoding / decoding process according to this disclosure can be performed based on the final codec unit that is no longer partitioned. In this case, the maximum codec unit can be used as the final codec unit based on factors such as encoding / decoding efficiency according to image features, or, if necessary, the codec unit can be recursively partitioned into deeper codec units, and the optimally sized codec unit can be used as the final codec unit. Here, the encoding / decoding process may include processes such as prediction, transformation, and reconstruction as described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be partitioned or divided from the final encoding / decoding unit described above, respectively. The prediction unit may be a unit for predicting samples, and the transformation unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.

[0097] In some cases, a unit can be used interchangeably with terms such as block, region, etc. Typically, an MxN block can refer to a set of transform coefficients or a set of samples consisting of M columns and N rows. Samples can typically indicate pixels or pixel values, or they can indicate only pixel / pixel values ​​of the luminance component, or only pixel / pixel values ​​of the chrominance component. Samples can be used as items corresponding to pixels or cells.

[0098] The image / video encoder 300 generates a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 321 or intra-frame predictor 322 from the input image signal (original block, original sample array), and the generated residual signal is sent to the converter 332. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the image / video encoder 300 can be called the subtractor 331. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes the prediction samples of the current block. The predictor can determine whether intra-frame prediction or inter-frame prediction is applied to the current block or the unit of the CU. The predictor can generate various prediction-related information, such as prediction mode information, and send it to the entropy encoder 340. The prediction-related information can be encoded by the entropy encoder 340 and output in the form of a bitstream.

[0099] Intra-predictor 322 can predict the current block by referencing samples within the current image. In this case, the reference samples can be located in a neighboring region of the current block, or they can be located further away depending on the prediction mode. In intra-prediction, the prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes and planar modes. Directional modes can include, for example, 33 or 65 directional prediction modes depending on the granularity of the prediction orientation. However, this is just an example, and more or fewer directional prediction modes can be used depending on the configuration. Intra-predictor 322 can also determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0100] Inter-frame predictor 321 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing within the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a co-located reference block or co-located CU (colCU), and the reference image including the temporally neighboring block may be referred to as a co-located image (colPic). For example, inter-frame predictor 321 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the inter-frame predictor 321 can use motion information of neighboring blocks as motion information of the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0101] Predictor 320 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction for the prediction of a block, and can also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for the prediction of a block. The IBC prediction mode or palette mode can be used for content image / video codecs, such as Screen Content Codec (SCC). IBC essentially performs prediction within the current frame, but because it derives a reference block within the current frame, it can be similar to inter-frame prediction operation. In other words, IBC can use at least one of the inter-frame prediction methods described in this disclosure. The palette mode can be considered as an example of intra-frame codec or intra-frame prediction. When a palette mode is applied, sample values ​​within the frame can be signaled based on information associated with the palette table and palette index.

[0102] The predicted signal generated by predictor 320 can be used to generate a reconstructed signal or a residual signal. Transformer 332 can generate transform coefficients by applying a transform method to the residual signal. For example, the transform method can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to the transform obtained from a graph when the relationship information between pixels is indicated as a graph. CNT refers to the transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform process can be applied to pixel blocks of the same square size or non-square, variable-size blocks.

[0103] The quantizer 333 quantizes the transform coefficients and sends them to the entropy encoder 340, which encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. This information about the quantized transform coefficients can be referred to as residual information. The quantizer 333 can reorder the block-like quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can generate information about the quantized transform coefficients based on this one-dimensional vector form. The entropy encoder 340 can perform various encoding methods, such as Exponential Columbus, Context Adaptive Variable Length Codec (CAVLC), Context Adaptive Binary Arithmetic Codec (CABAC), etc. The entropy encoder 340 can encode the quantized transform coefficients together with or separately from information necessary for video / image reconstruction (e.g., values ​​of syntax elements). The encoded information (e.g., encoded video / image information) can be sent or stored as a bitstream in a Network Abstraction Layer (NAL) unit. The image / video information may further include information about various parameter sets, such as Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. Furthermore, the image / video information may further include methods for generating and using encoded information, their purpose, etc. In this disclosure, information and / or syntax elements sent / signed from the image / video encoder to the image / video decoder may be included in the image / video information. The image / video information may be encoded using the encoding process described above and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmission and / or a storage unit (not shown) for storing the signal output from the entropy encoder 340 may be configured as an internal / external element of the image / video encoder 300, or the transmitter may be included in the entropy encoder 340.

[0104] The quantization transform coefficients output from quantizer 333 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients through dequantizer 334 and inverse transform 335. Adder 350 can generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from inter-frame predictor 321 or intra-frame predictor 322. When there is no residual for the processing target block, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 350 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block within the current image, and, as described later, can also be used for inter-frame prediction of the next image by filtering.

[0105] Simultaneously, luminance mapping with chroma scaling can be applied during image encoding and / or reconstruction.

[0106] Filter 360 can apply filtering to the reconstructed signal to enhance subjective / objective image quality. For example, filter 360 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be stored in memory 370, specifically in the DPB of memory 370. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 360 can generate various filtering-related information and send it to entropy encoder 340. The filtering-related information can be encoded by entropy encoder 340 and output as a bitstream.

[0107] The modified reconstructed image sent to memory 370 can be used as a reference image in inter-frame predictor 321. This avoids prediction mismatch between the encoder and decoder sides and improves coding efficiency.

[0108] The DPB of memory 370 can store modified reconstructed images for use as reference images in inter-frame predictor 321. Memory 370 can store motion information of blocks in which motion information within the current image and / or motion information of blocks within reconstructed images is derived (or encoded). The stored motion information can be sent to inter-frame predictor 321 as motion information for spatially or temporally neighboring blocks. Memory 370 can store reconstructed samples of reconstructed blocks in the current image and send the stored reconstructed samples to intra-frame predictor 322.

[0109] At the same time, the VCM encoder (or feature / feature map encoder) can have essentially the same reference. Figure 3The image / video encoder 300 described has the same / similar structure as the image / video encoder 300 because it performs a series of processes such as prediction, transformation, quantization, etc., to encode features / feature maps. However, the VCM encoder differs from the image / video encoder 300 because it targets the features / feature maps to be encoded, and therefore, its name in each unit (or component) (e.g., image partitioner 310, etc.) and its specific operational details derived from the image / video encoder 300 may differ. The specific operational details of the VCM encoder will be described in detail later.

[0110] decoder

[0111] Figure 4 This diagram schematically illustrates an image / video decoder to which embodiments of the present disclosure can be applied.

[0112] refer to Figure 4 The image / video decoder 400 may include an entropy decoder 410, a residual processor 420, a predictor 430, an adder 440, a filter 450, and a memory 460. The predictor 430 may include an inter-frame predictor 431 and an intra-frame predictor 432. The residual processor 420 may include a dequantizer 421 and an inverse transformer 422. According to one embodiment, the entropy decoder 410, residual processor 420, predictor 430, adder 440, and filter 450 may be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 460 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may also include the memory 460 as an internal / external component.

[0113] When the input includes a bitstream containing video / image information, the image / video decoder 400 can respond to... Figure 3 The image / video encoder 300 reconstructs the image / video through a process of processing image / video information. For example, the image / video decoder 400 can derive units / blocks based on block partitioning information obtained from the bitstream. The image / video decoder 400 can perform decoding using processing units applied in the image / video encoder. Therefore, the decoding processing unit can be, for example, an encoding / decoding unit, and the encoding / decoding unit can be partitioned according to a quadtree structure, binary tree structure, and / or ternary tree structure from the encoding / decoding tree unit or the maximum encoding / decoding unit. At least one transform unit can be derived from the encoding / decoding unit. Furthermore, the reconstructed image signal decoded and output by the image / video decoder 400 can be played by a playback device.

[0114] The image / video decoder 400 can be used in bitstream form from... Figure 3The encoder in the process receives the signal output and can decode the received signal using the entropy decoder 410. For example, the entropy decoder 410 can parse the bitstream to derive the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. Additionally, the video / image information may further include general constraint information. Furthermore, the image / video information may include the method of generating the decoding information, its usage, purpose, etc. The image / video decoder 400 can further decode the picture based on information about the parameter sets and / or general constraint information. The signal notification / received information and / or syntax elements can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 410 can decode the information in the bitstream based on encoding and decoding methods such as exponential Golomb coding, CAVLC, or CABAC, and can output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the transform coefficients associated with the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model by using information about the target syntax element, the decoding information of neighboring and target blocks, or information about symbols / bins decoded in previous steps, predict the probability of bin occurrence based on the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining it by using information about decoded symbols / bins of the context model for the next symbol / bin. In the information decoded by the entropy decoder 410, prediction-related information can be provided to the predictors (inter-frame predictor 432 and intra-frame predictor 431), and the residual values ​​(i.e., quantization transform coefficients and related parameter information) entropied by the entropy decoder 410 can be input to the residual processor 420. The residual processor 420 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, in the information decoded by the entropy decoder 410, filtering-related information can be provided to the filter 450. Meanwhile, the receiver (not shown) that receives the signal output from the image / video encoder can be additionally configured as an internal / external component of the image / video decoder 400, or the receiver can be a component of the entropy decoder 410. Furthermore, the image / video decoder according to this disclosure can also be referred to as an image / video decoding device, and the image / video decoder can be divided into an information decoder (image / video information decoder) and / or a sample decoder (image / video sample decoder).In this case, the information decoder may include an entropy decoder 410, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 440, a filter 450, a memory 460, an inter-frame predictor 432, and an intra-frame predictor 431.

[0115] Dequantizer 421 can dequantize the quantized transform coefficients and the output transform coefficients. Dequantizer 421 can reorder the quantized transform coefficients in the form of two-dimensional blocks. In this case, the reordering can be performed based on the coefficient scan order performed in the image / video encoder. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (i.e., quantization step size information) and obtain the transform coefficients.

[0116] The inverse transformer 422 can perform an inverse transformation on the transform coefficients to obtain the residual signal (residual block, residual sample array).

[0117] Predictor 430 can perform prediction for the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether intra-frame prediction or inter-frame prediction is applied to the current block based on prediction-related information output from entropy decoder 410, and can determine a specific intra-frame / inter-frame prediction mode (prediction method).

[0118] Predictor 420 can generate prediction signals based on various prediction methods. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame and inter-frame prediction simultaneously for the prediction of a block. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for block prediction. The IBC prediction mode or palette mode can be used for content image / video encoding / decoding in games, such as Screen Content Codec (SCC). IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered an example of intra-frame encoding / decoding or intra-frame prediction. When a palette mode is applied, information related to the palette table and palette index can be included in the image / video information and signaled.

[0119] The intra-predictor 431 can predict the current block by referencing samples within the current image. The reference samples can be located in the neighborhood of the current block or positioned far from the current block based on the prediction pattern. In intra-prediction, the prediction pattern can include multiple non-directional patterns and multiple directional patterns. The intra-predictor 431 can determine the prediction pattern applied to the current block by using prediction patterns applied to neighboring blocks.

[0120] Inter-frame predictor 432 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include information about the inter-frame prediction orientation (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks within the current image and temporally neighboring blocks in the reference image. For example, inter-frame predictor 432 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction-related information may include information indicating the inter-frame prediction mode of the current block.

[0121] Adder 440 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictors (including inter-frame predictor 432 and / or intra-frame predictor 431). When there is no residual for the processing target block, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.

[0122] Adder 440 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current image, or, as described later, can be filtered out, or used for inter-frame prediction of the next image.

[0123] Simultaneously, luminance mapping with chroma scaling can be applied during image decoding.

[0124] Filter 450 can apply filtering to the reconstructed signal to enhance subjective / objective image quality. For example, filter 450 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and can send the modified reconstructed image to memory 460, specifically to the DPB in memory 460. Various filtering methods may include, for example, unblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.

[0125] The (modified) reconstructed image stored in the DPB of memory 460 can be used as a reference image in inter-frame predictor 432. Memory 460 can store motion information of blocks in which motion information within the current image is derived (or decoded) and / or the motion information of blocks in the reconstructed image is sent to inter-frame predictor 432 as motion information for spatially or temporally neighboring blocks. Memory 460 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 431.

[0126] Meanwhile, the VCM decoder (or feature / feature map decoder) can have the same features as described above (see reference). Figure 4 The image / video decoder 400 described below has the same / similar structure as the image / video decoder 400 because it performs a series of processes, such as prediction, inverse transform, dequantization, etc., to decode features / feature maps. However, the VCM decoder differs from the image / video decoder 400 because it targets the features / feature maps used for decoding, and therefore, its name in each unit (or component) (e.g., DPB, etc.) and its specific operational details from those of the image / video decoder 400 may differ. The operations of the VCM decoder can correspond to the operations of the VCM encoder, and their specific operational details will be described in detail later.

[0127] Feature / feature map encoding process

[0128] Figure 5 This is a flowchart illustrating a feature / feature map encoding process that can be applied to embodiments of the present disclosure.

[0129] refer to Figure 5 The feature / feature map encoding process may include a prediction process S510, a residual processing process S520, and an information encoding process S530.

[0130] The prediction process S510 can be referenced above. Figure 3 The predictor 320 is described and executed.

[0131] Specifically, the intra predictor 322 can predict the current block (i.e., the currently encoded set of feature elements) by referencing feature elements in the current feature / feature map. Intra-prediction can be performed based on the spatial similarity of feature elements in the configured feature / feature map. For example, it can be estimated that feature elements included in the same region of interest (RoI) within an image / video have similar data distribution characteristics. Therefore, the intra predictor 322 can predict the current block by referencing pre-reconstructed feature elements in the RoI that includes the current block. In this case, the referenced feature elements can be located near the current block or can be located separately from the current block depending on the prediction mode. The intra-prediction modes used for feature / feature map encoding can include multiple non-directional prediction modes and multiple directional prediction modes. Non-directional prediction modes can include, for example, prediction modes corresponding to DC mode and planar mode in the image / video encoding process. In addition, directional modes can include, for example, prediction modes corresponding to 33 directional modes or 65 directional modes in the image / video encoding process. However, this is only an example, and according to one implementation, the type and number of intra-prediction modes can be configured / changed in various ways.

[0132] Inter-frame predictor 321 can predict the current block based on a reference block (i.e., a set of reference feature elements) specified by motion information on a reference feature / feature map. Inter-frame prediction can be performed based on the temporal similarity of feature elements in the configuration feature / feature map. For example, temporally consecutive features may have similar data distribution characteristics. Therefore, inter-frame predictor 321 can predict the current block by referencing pre-reconstructed feature elements of the current feature and temporally adjacent features. In this case, the motion information used to specify the reference feature elements may include motion vectors and reference feature / feature map indices. The motion information may further include information related to inter-frame prediction orientation (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current feature / feature map and temporally neighboring blocks existing in the reference feature / feature map. The reference feature / feature map including the reference block and the reference feature / feature map including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a co-located reference block, etc., and the reference feature / feature map including the temporally neighboring block may be referred to as a co-located feature / feature map. Inter-frame predictor 321 can configure a candidate list of motion information based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference feature / feature map index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip mode and merge mode, inter-frame predictor 321 can use the motion information of neighboring blocks as the motion information of the current block. For skip mode, unlike merge mode, residual signals may not be sent. For motion vector prediction (MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference. In addition to the intra-frame prediction and inter-frame prediction described above, predictor 320 can also generate prediction signals based on various prediction methods.

[0133] The prediction signal generated by predictor 320 can be used to generate the residual signal (residual block, residual eigenvalue) S520. This can be seen from the above reference... Figure 3 The described residual processor 330 executes the residual processing procedure S520. Furthermore, the (quantization) transform coefficients can be generated through a transform and / or quantization process for the residual signal, and the entropy encoder 340 can encode the information associated with the (quantization) transform coefficients into residual information in the bitstream S530. In addition to the residual information in the bitstream, the entropy encoder 340 can also encode information necessary for feature / feature map reconstruction, such as prediction information (e.g., prediction pattern information, motion information, etc.).

[0134] Meanwhile, the feature / feature map encoding process may further include a process for generating a reconstructed feature / feature map for the current feature / feature map, and a process for applying in-loop filtering to the reconstructed feature / feature map (optional), as well as a process S530 for encoding information for feature / feature map reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bit stream.

[0135] The VCM encoder can derive (modified) residual features from the quantization transform coefficients through dequantization and inverse transform, and can generate reconstructed features / feature maps based on the predicted features as the output of S510 and the (modified) residual features. The reconstructed features / feature maps generated in this way can be identical to those generated by the VCM decoder. When an in-loop filtering process is performed on the reconstructed features / feature maps, a modified reconstructed features / feature maps can be generated through this in-loop filtering process. The modified reconstructed features / feature maps can be stored in the decoded feature buffer (DFB) or memory, and then used as reference features / feature maps during feature / feature map prediction. Furthermore, (in-loop) filtering-related information (parameters) can be encoded and output as a bitstream. The in-loop filtering process removes noise that may occur during feature / feature map encoding and decoding, and improves the performance of feature / feature map-based tasks. Additionally, the in-loop filtering process can be performed on both the encoder and decoder sides to ensure the recognition of prediction results, improve the reliability of feature / feature map encoding and decoding, and reduce the amount of data transmitted for feature / feature map encoding and decoding.

[0136] Feature / Feature Map Decoding Process

[0137] Figure 6 This is a flowchart illustrating a feature / feature map decoding process that can be applied to embodiments of the present disclosure.

[0138] refer to Figure 6The feature / feature map decoding process may include an image / video information acquisition process S610, feature / feature map reconstruction processes S620 to S640, and an intra-loop filtering process S650 for reconstructing the feature / feature map. The feature / feature map reconstruction process can be performed based on the prediction signal and residual signal obtained through inter-frame / intra-frame prediction S620, residual processing S630, and dequantization and inverse transform processes for the quantization transform coefficients described in this disclosure. A modified reconstructed feature / feature map can be generated through the intra-loop filtering process for reconstructing the feature / feature map, and the modified reconstructed feature / feature map can be output as a decoded feature / feature map. The decoded feature / feature map can be stored in a decoded feature buffer (DFB) or memory, and then used as a reference feature / feature map in the inter-frame prediction process when decoding the feature / feature map. In some cases, the above-described intra-loop filtering process may be omitted. In this case, the reconstructed features / feature maps can be output as decoded features / feature maps and stored in the decoded feature buffer (DFB) or memory, and then used as reference features / feature maps in the inter-frame prediction process when decoding the features / feature maps.

[0139] Feature extraction and data distribution characteristics

[0140] Figure 7 This is a diagram illustrating an example of a feature extraction method using Feature Extraction Network 700.

[0141] refer to Figure 7 The feature extraction network 700 can receive a video source (image / video, 710) and perform feature extraction operations to output a feature set 720 of the video source 710. The feature set 720 may include multiple features (C0, C1, ..., C...) extracted from the video source 710. n ), and can be represented as a feature map. Features (C0, C1, ..., C n Each of the features can include multiple feature elements and can have different data distribution characteristics.

[0142] exist Figure 7 In the diagram, W, H, and C represent the width, height, and number of channels of the video source 710, respectively. Here, the number of channels C of the video source 710 can be determined based on the image format of the video source 710. For example, when the video source 710 has an RGB image format, the number of channels C of the video source 710 can be 3.

[0143] Additionally, W', H', and C' represent the width, height, and number of channels of feature set 720, respectively. The number of channels C' of feature set 720 can be equal to the features (C0, C1, ..., C...) extracted from video source 710. nThe total number of channels (n+1). In one example, the number of channels C' of feature set 720 can be greater than the number of channels C of video source 710.

[0144] The attributes (W', H', C') of feature set 720 can vary depending on the attributes (W, H, C) of video source 710. For example, as the number of channels C of video source 710 increases, the number of channels C' of feature set 720 can also increase. Furthermore, the attributes (W', H', C') of feature set 720 can vary depending on the type and attributes of feature extraction network 700. For example, when feature extraction network 700 is implemented as an artificial neural network (e.g., CNN, DNN, etc.), the attributes (W', H', C') of feature set 720 can depend on the output of each feature (C0, C1, ..., C...). n The position of the layer varies.

[0145] Video source 710 and feature set 720 can have different data distribution characteristics. For example, video source 710 typically consists of one channel (grayscale image) or three channels (RGB image). Pixels included in video source 710 can have the same range of integer values ​​across all channels and can have non-negative values. Furthermore, each pixel value can be uniformly distributed within a predetermined range of integer values. In contrast, feature set 720 can consist of different numbers of channels (e.g., 32, 64, 128, 256, 512, etc.) depending on the type of feature extraction network 700 (e.g., CNN, DNN, etc.) and layer location. Feature elements included in feature set 720 can have different ranges of real values ​​for each channel and can also have negative values. Additionally, each feature element value can be densely distributed within a specific region of a predetermined range of real values.

[0146] Figure 8a It is a graph illustrating the data distribution characteristics of the video source, and Figure 8b It is a graph illustrating the data distribution characteristics of the feature set.

[0147] refer to Figure 8a A video source can consist of three channels in total—R, G, and B channels (R channel, G channel, B channel), and each pixel value can have an integer value ranging from 0 to 255. In this case, the data type of the video source can be represented as an 8-bit integer type.

[0148] Conversely, refer to Figure 8b The feature set can consist of 64 channels (features), and each feature element value can have a range from... arrive The real value range. In this case, the data type of the feature set can be represented as a 32-bit floating-point type.

[0149] The feature set can have feature element values ​​of type floating point, and can have different data distribution characteristics for each channel (or feature). Table 1 shows examples of data distribution characteristics for each channel of the feature set.

[0150] [Table 1]

[0151] Referring to Table 1, the feature set can consist of a total of n+1 channels (C0, C1, ..., C...). n It consists of ) channels. For each channel (C0, C1, ..., C... n The mean (μ), standard deviation (σ), maximum (Max), and minimum (Min) of the feature elements may differ. For example, the mean (μ) of the feature elements included in channel 0 (C0) can be 10, the standard deviation (σ) can be 20, the maximum (Max) can be 90, and the minimum (Min) can be 60. Conversely, the mean (μ) of the feature elements included in channel 1 (C1) can be 30, the standard deviation (σ) can be 10, the maximum (Max) can be 70.5, and the minimum (Min) can be -70.2. Furthermore, channel n (C... n The mean (μ) of the feature elements included in the ) can be 100, the standard deviation (σ) can be 5, the maximum value (Max) can be 115.8, and the minimum value (Min) can be 80.2.

[0152] As mentioned above, feature / feature map quantization can be performed based on the different data distribution characteristics of each channel. Floating-point feature / feature map data can be converted to integer type through quantization.

[0153] Furthermore, due to the spatiotemporal similarity between consecutive frames, the feature sets and / or channels extracted continuously from the video source may have the same / similar data distribution characteristics. Table 2 shows examples of the data distribution characteristics of consecutive feature sets.

[0154] [Table 2]

[0155] In Table 2, f F0 f represents the first feature set extracted from frame 0 (F0). F1 Let f represent the second feature set extracted from frame 1 (F1), and f F2 This represents the third feature set extracted from frame 2 (F2).

[0156] Referring to Table 2, the first to third consecutive feature sets (f) F0 f F1 f F2They may have the same / similar mean (μ), standard deviation (σ), maximum (Max), and minimum (Min).

[0157] Furthermore, due to the spatiotemporal similarity between consecutive frames, the corresponding channels within a continuously extracted feature set from a video source may have the same / similar data distribution characteristics. Table 3 shows examples of the data distribution characteristics of the corresponding channels for consecutive feature sets.

[0158] [Table 3]

[0159] In Table 3, f F0C0 Let f represent the first channel within the first feature set extracted from frame 0 (F0), and f F1C0 This represents the first channel within the second feature set extracted from frame 1 (F1).

[0160] Referring to Table 3, the first channel (f) of the first feature set F0C0 ) and the second channel (f) of the corresponding second feature set F0C1 ) can have the same / similar mean (μ), standard deviation (σ), maximum (Max), and minimum (Min).

[0161] Prediction of features / feature maps can be performed based on, for example, the similarity of data distribution features between feature sets or channels.

[0162] Transform / Inverse Transform

[0163] As described above, an image / video encoder can derive a residual block based on a prediction block, and derive quantized transform coefficients by applying transform and quantization to the derived residual block. Similarly, a VCM encoder can derive a residual block based on a prediction block (or feature elements), and derive quantized transform coefficients for the feature / feature map by applying transform and quantization to the derived residual block. Information about the quantized transform coefficients (residual information) can be included in the residual coding syntax, encoded, and then output as a bitstream.

[0164] Decoders (image / video decoders and VCM decoders) can obtain information about the quantization transform coefficients (residual information) from the bitstream and derive the quantization transform coefficients by decoding this information. Furthermore, the decoder can derive residual blocks based on the quantization transform coefficients through dequantization / inverse transform. As mentioned above, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When transform / inverse transform is omitted, for consistency of expression, transform coefficients can be referred to as "coefficients" or "residual coefficients," or they can still be called "transform coefficients." Whether transform / inverse transform is omitted can be signaled based on a flag (e.g., transform_skip_flag). Furthermore, in VCM, whether transform / inverse transform is omitted can be determined by feature / feature map or by channel.

[0165] Transformation / inverse transformation can be performed based on a transform kernel. For example, depending on the implementation, a multiple transform selection (MTS) scheme can be applied. In this case, some from a set of multiple transform kernels can be selected and applied to the current block. Transform kernels can be referred to by various terms such as transform matrix, transform type, etc. For example, a set of transform kernels can represent a combination of vertical transform kernels (vertical transform kernels) and horizontal transform kernels (horizontal transform kernels).

[0166] To indicate the transform kernel set or feature / feature map for the current block, MTS index information (e.g., mts_idx) can be signaled. Examples of MTS index information are given in Table 4.

[0167] [Table 4]

[0168] In Table 4, `trTypeHor` can represent a horizontal transform kernel, and `trTypeVer` can represent a vertical transform kernel. A `trTypeHor / trTypeVer` value of 0 can represent DCT2, a `trTypeHor / trTypeVer` value of 1 can represent DCT7, and a `trTypeHor / trTypeVer` value of 2 can represent DCT8. However, these are exemplary, and other values ​​can be mapped to other DCT / DST or transform kernels based on predefined rules.

[0169] In this disclosure, an MTP-based transform / inverse transform can be applied as the primary transform / inverse transform, and a secondary transform / inverse transform can be further applied. In this case, the encoder-side transform can be performed in the order of primary transform and secondary transform, and the decoder-side transform can be performed in the order of secondary inverse transform and primary inverse transform. According to an embodiment, the secondary transform can be applied only to the coefficients in the upper left low-frequency region of the coefficient block to which the primary transform is applied. This secondary transform can be called a "low-frequency non-separable transform (LFNST)". Furthermore, according to an embodiment, as a result of performing the secondary transform, the number of coefficients can be reduced. This secondary transform can be called a "reduced secondary transform (RST)".

[0170] Figure 9 This is a diagram illustrating the transformation and inverse transformation processes applicable to embodiments of this disclosure. Figure 9 In the middle, converter 910 can correspond to Figure 3 The converter 332, and the inverse converter 920 can correspond to Figure 3 Inverter 335 or Figure 4 The inverse converter 422.

[0171] Reference Figure 9 The converter 910 may include a main converter 911 and a secondary converter 912.

[0172] The master converter 911 can generate (master) transform coefficients B by applying the master transform to the residual data A (or residual feature elements). In this disclosure, the master transform can be referred to as the "core transform". The master transform can be performed based on the MTS scheme. When applying an existing MTS, a transform from the spatial domain to the frequency domain based on DCT type 2, DCT type 7, DCT type 8, etc., is applied to the residual signal (or residual block) so that transform coefficients (or master transform coefficients) can be generated. Here, DCT type 2, DCT type 7, DCT type 8, etc., can be referred to as "transform type", "transform kernel", or "transform core". However, these are exemplary, and the master transform can also be applied to MTS kernels other than the aforementioned MTS kernels.

[0173] The quadratic transformer 912 can generate (quadratic) transform coefficients C by applying a quadratic transform to the (primary) transform coefficients B. The quadratic transform is an inseparable transform, such as the aforementioned LFNST or RST. For example, when using a 4×4 block, an inseparable quadratic transform can be performed as follows.

[0174] The 4×4 input block X can be represented as shown in Equation 1 below.

[0175] [Formula 1]

[0176] When the input block X is represented in vector form, the vector X can be shown in Equation 2 below.

[0177] [Equation 2]

[0178] In this case, Equation 3 can be used to calculate the inseparable quadratic transformation.

[0179] [Formula 3]

[0180] Here, vector F is the transformation coefficient vector, T is a 16×16 inseparable transformation matrix, and It is the multiplication of matrices and vectors.

[0181] Based on the above formula, a 16×1 transformation coefficient vector F can be derived, and vector F can be reconfigured into 4×4 blocks according to the scan order (e.g., horizontal, vertical, diagonal, or predetermined / pre-stored scan order).

[0182] Subsequently, the inverse converter 820 may include an (inverse) secondary converter 821 and an (inverse) primary converter 822.

[0183] The inverse quadratic transformer 921 can generate the main (inverse) transformation coefficients B' by applying the inverse quadratic transformation to the dequantized quadratic transformation coefficients C'. Here, the inverse quadratic transformation can correspond to the inverse process of the quadratic transformation performed by the transformer 910. According to the embodiment, the inverse quadratic transformation can be applied to the upper left corner of the coefficient block.

[0184] The (inverse) master converter 922 can generate residual samples A' by applying the (inverse) master transformation to the (master) (inverse) transformation coefficients B'.

[0185] Quantization / Dequantization

[0186] As described above, the quantizer 333 of encoder 300 can derive quantized transform coefficients by applying quantization to transform coefficients, and the dequantizer 334 of image / video encoder 300 or the dequantizer 421 of image / video encoder 400 can derive transform coefficients by applying dequantization to quantized transform coefficients. Similarly, a VCM encoder can derive quantized transform coefficients by applying quantization to transform coefficients, and the dequantizer of a VCM encoder or the dequantizer of a VCM decoder can derive transform coefficients by applying dequantization to quantized transform coefficients. Typically, the quantization rate can be changed during feature / feature map encoding, and the changed quantization rate can be used to adjust the compression ratio. From an implementation perspective, complexity can be considered by using quantization parameters (QPs) instead of directly using the quantization rate. For example, QPs with integer values ​​from 0 to 63 can be used, and each QP value can correspond to an actual quantization rate. The QP for the luma component (luma sample) can be QPY, and the QP for the chroma component (chroma sample) can be QPC.

[0187] During the quantization process, the transform coefficients C can be input and divided by the quantization rate Qstep, and the quantization transform coefficients C' can be obtained based on the result. In this case, considering computational complexity, the quantization rate can be multiplied by a scaling factor (scale) to convert it to an integer, and a shift operation corresponding to the scaling value can be performed. The quantization scaling factor can be derived from the product of the quantization rate and the scaling value. In other words, the quantization scaling factor can be derived from QP. The transform coefficients C and the quantization scaling factor can be applied, and the quantization transform coefficients C' can be derived based on the result.

[0188] During the dequantization process, which is the inverse of the quantization process, the quantization transform coefficients C' can be multiplied by the quantization rate Qstep, and the recovered transform coefficients C'' can be obtained based on the result. In this case, a level scaling factor can be derived from Qstep, and the level scaling factor can be applied to the quantization transform coefficients C'. Based on this result, the recovered transform coefficients C'' can be derived. Due to losses in the transform and / or quantization processes, the recovered transform coefficients C'' may differ slightly from the initial transform coefficients. Therefore, the encoder performs dequantization in the same manner as the decoder.

[0189] Furthermore, adaptive frequency-weighted quantization (IFQ) techniques can be applied to adjust the quantization intensity according to the frequency. IFQ is a method of applying different quantization intensities to a frequency. According to IFQ, a predefined quantization scaling matrix can be used to apply different frequency-specific quantization intensities. In other words, the aforementioned quantization / dequantization process can be performed on top of the quantization scaling matrix. For example, depending on the size of the current block and / or whether the prediction mode applied to the current block to generate the residual signal of the current block is inter-frame prediction or intra-frame prediction, different quantization scaling matrices can be used. The quantization scaling matrix can be called a "quantization matrix" or a "scaling matrix." The quantization scaling matrix can be predefined. Furthermore, for frequency-adaptive scaling, the frequency-specific quantization scaling information of the quantization scaling matrix can be configured / encoded by the encoder and signaled to the decoder. This frequency-specific quantization scaling information can be called "quantization scaling information." The frequency-specific quantization scaling information can include scaling list data (scaling_list_data). Based on the scaling list data, a (modified) quantization scaling matrix can be derived. Additionally, the frequency-specific quantization scaling information can include presence flags indicating the existence of scaling list data. On the other hand, when the scaling list data is signaled at a higher level (e.g., sequence level, feature set group level, etc.), information indicating whether the scaling list data is modifiable at a lower level (e.g., feature set level, channel level, etc.) may be further included.

[0190] Feature quantization / dequantization can be performed based on a specific quantization group. Specifically, multiple quantization segments can be set based on the data distribution characteristics of the feature set. These quantization segments can have different data distribution ranges and can be defined as a quantization group. Furthermore, feature quantization / dequantization calculations can be performed by transforming the data distribution of each channel in the feature set to any given data distribution range.

[0191] Figure 10 This is a diagram illustrating a quantization group according to an embodiment of the present disclosure.

[0192] Reference Figure 10 A quantization group can include four quantization segments A, B, C, and D with different data distribution ranges. Each of the quantization segments A, B, C, and D in a quantization group can be defined using minimum and maximum values, and is set based on the data distribution characteristics of the current feature / feature map. For example, quantization segments A, B, C, and D can be set to [-1, 3], [0, 2], [-2, 4], and [-2, 1], respectively.

[0193] The VCM encoder can quantize each channel (or feature) in the feature set based on preset quantization segments A, B, C, and D without separately calculating the maximum and minimum values ​​of the feature elements. For example, the first channel (channel 1) can be quantized based on quantization segment A, which has the most similar data distribution range. The second channel (channel 2) can be quantized based on quantization segment B, which has the most similar data distribution range. The third channel (channel 3) can be quantized based on quantization segment C, which has the most similar data distribution range. The fourth channel (channel 4) can be quantized based on quantization segment D, which has the most similar data distribution range. In this case, the VCM encoder can encode / signal the number, minimum, and maximum values ​​of the quantization segments as feature quantization-related information. Furthermore, the VCM encoder can signal the quantization segment index information used to encode the current channel (or feature) as feature quantization-related information. Therefore, it is unnecessary to signal the number of quantization bits for each channel and the maximum and minimum values ​​of the feature elements, thereby reducing the number of transmitted bits and improving encoding / signaling efficiency.

[0194] The VCM decoder can generate the same quantization set as the one generated by the VCM encoder, based on the number of quantization segments received from the VCM encoder and the minimum and maximum values ​​of each segment. Alternatively, the VCM decoder can generate the quantization set based on the data distribution characteristics of the previously reconstructed feature set. The VCM decoder can dequantize the current feature based on the quantization segments identified from the quantization segment index information received from the VCM encoder.

[0195] Examples of feature quantization operations based on quantization groups are shown in Equations 4 and 5.

[0196] [Formula 4]

[0197] [Formula 5]

[0198] In Equation 4, F n It can be a feature set The nth channel (n is an integer greater than or equal to 1). Here, R can be the width of the feature set, and C can be the height of the feature set. Furthermore, max(F n ) can be the maximum value of the feature elements in the nth channel, and min(F) n ) can be the minimum value of the feature element in the nth channel.

[0199] Referring to Equation 4, the nth channel F can be modified based on the maximum and minimum values ​​of the characteristic elements in the channel. n Normalize to have characteristic element values ​​between 0 and 1.

[0200] In Equation 5, It can be a feature set The normalized nth channel. Here, R can be the width of the feature set, and C can be the height of the feature set. Furthermore, max( ) can be applied to the nth channel F n Quantization segment The maximum value, and min( ) can be applied to the nth channel F n Quantization segment The minimum value.

[0201] Referring to Equation 5, it can be based on the application to the nth channel F n Quantization segment The maximum value of max( ) and minimum value min( For the normalized nth channel Quantify it.

[0202] Furthermore, for dequantization, quantization-related information can be encoded in the bitstream. According to the implementation, quantization-related information may include global quantization information, activation function information, the number of quantization bits, and quantization segment index information. Global quantization information may indicate whether the number of quantization bits is set individually for each quantization segment or uniformly for all quantization segments. Activation function information may indicate the type of activation function applied to the current feature. For example, activation function information may indicate whether the activation function applied to the current feature is a first activation function that involves signaling both the maximum and minimum values ​​of each quantization segment or a second activation function that involves signaling only the maximum value of each quantization segment. The number of quantization bits can be set individually for each quantization segment based on the aforementioned global quantization information, or the number of quantization bits can be set uniformly for all quantization segments. When the number of quantization bits is set uniformly for all quantization segments, the number of quantization bits can be signaled only once for all quantization segments.

[0203] Additionally, quantization-related information may include the number of quantization segments in the quantization group and the minimum and maximum values ​​of each quantization segment. Here, the minimum and maximum values ​​of each quantization segment can be adaptively signaled based on the type of activation function. For example, when the activation function applied to the current feature is a first activation function (e.g., a leaky rectified linear unit (ReLU)), both the minimum and maximum values ​​of each quantization segment can be signaled. On the other hand, when the activation function applied to the current feature is a second activation function (e.g., ReLU), the minimum value of each quantization segment can be estimated as 0, and only the maximum value of each quantization segment can be signaled.

[0204] Entropy coding

[0205] As described above, some or all of the feature / feature map information can be encoded by the entropy encoder 340, and some or all of the feature / feature map information can be decoded by the entropy decoder 410. In this case, the feature / feature map information can be encoded / decoded according to syntax elements such as image / video information. In this disclosure, "information is encoded / decoded" can include "information is encoded / decoded in the manner described in this paragraph".

[0206] Figure 11 This is a block diagram illustrating Context Adaptive Binary Arithmetic Encoding (CABAC) for encoding a syntax element.

[0207] Reference Figure 11 When the input signal is a syntax element rather than a binary value, the CABAC encoding process first involves converting the input signal into a binary value through binarization. When the input signal is already a binary value, binarization can be bypassed. Here, each binary digit 0 or 1 that constitutes the binary value can be called a "bin". For example, when the binarized binary string (bin string) is 110, each of 1, 1, and 0 can be a single bin. A bin of a syntax element can represent the value of the syntax element.

[0208] Binarized bins can be input into either a regular encoding engine or a bypass encoding engine. A regular encoding engine can assign a context model reflecting probability values ​​to the corresponding bin and encode the bin based on the assigned context model. The regular encoding engine can encode each bin and then update the probability model for that bin. These encoded bins can be called "context-encoded bins." A bypass encoding engine can omit the process of estimating the probabilities of the input bins and the process of updating the probability model applied to the bins after encoding. Encoding speed can be improved by using a uniform probability distribution (e.g., 50:50) instead of assigning context to encode the input bins. These encoded bins can be called "bypass bins." A context model can be assigned and updated for each context-encoded (regularly encoded) bin, and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived from ctxInc. As a specific example, the context index ctxidx indicating the context model for each of the regular encoded bins can be derived as the sum of the context index increment ctxInc and the context index offset ctxIdxOffset. Here, ctxInc can be exported differently for each bin. ctxIdxOffset can be the lowest value of ctxIdx. Typically, ctxIdxOffset can be determined based on the type of slice, and the context model for a single syntax element within the slice can be determined / derived from ctxInc.

[0209] It can be determined whether encoding is performed using a regular encoding engine or a bypass encoding engine during the entropy encoding process, and the encoding path can be switched. As entropy decoding, the same process can be performed in reverse order.

[0210] For example, it can be like Figure 12 and Figure 13 The aforementioned entropy encoding is performed as shown.

[0211] Figure 12 and Figure 13 This is a diagram illustrating the entropy coding process.

[0212] Refer to together Figure 12 and Figure 13 An encoder (entropy encoder) can perform entropy encoding on feature / feature map information. Feature / feature map information may include prediction-related information (e.g., inter-frame / intra-frame prediction determination information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, intra-loop filtering-related information, etc., or may include various related syntax elements. Entropy encoding can be performed on a syntax element-by-syntax basis. An entropy encoder can be one of the above. Figure 3 Image / video encoder 300 and entropy encoder 340.

[0213] The encoder can binarize the target syntax element (S1210). Here, the binarization can be based on various binarization methods (such as truncated Rice binarization, fixed-length binarization, etc.), and a binarization method for the target syntax element can be predefined. The binarization process can be performed by the binarizer 1301 in the entropy encoder 1300.

[0214] The encoder can perform entropy encoding on the target syntax element (S1220). The encoder can perform regular encoding (context-based) or bypass encoding on the bin string of the target syntax element based on CABAC, context-adaptive variable-length encoding (CAVLC), etc., and the output can be included in the bitstream. The entropy encoding process can be executed by the entropy encoding processor 1302 in the entropy encoder 1300. As described above, the output bitstream can be sent to the decoder via a (digital) storage medium or network.

[0215] Figure 14 and Figure 15 This is a diagram illustrating the entropy decoding process.

[0216] Reference Figure 14 and Figure 15 The decoder (entropy decoder) can decode the encoded feature / feature map information. This feature / feature map information may include prediction-related information (e.g., inter-frame / intra-frame prediction determination information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, intra-loop filtering-related information, etc., or may include various related syntax elements. Entropy encoding can be performed on a syntax element-by-syntax basis. The entropy decoder can be one of the above. Figure 4 Image / video decoder 400 and entropy decoder 410.

[0217] The decoder can binarize the target syntax element (S1410). Here, binarization can be based on various binarization methods (such as truncated Rice binarization, fixed-length binarization, etc.), and a binarization method for the target syntax element can be predefined. The decoder can derive available bin strings (bin string candidates) for the available values ​​of the target syntax element through the binarization process. The binarization process can be performed by the binarizer 1501 in the entropy decoder 1500.

[0218] The decoder can perform entropy decoding (S1420) on the target syntax element. The decoder can sequentially decode and parse the bins of the target syntax element from the input bits in the bitstream, thereby comparing the derived bin strings with the available bin strings for the syntax element. When the derived bin string matches one of the available bin strings, the value corresponding to that bin string can be derived as the value of the syntax element. Otherwise, the next bit in the bitstream can be further parsed, and the above process can be repeated. This process allows specific information (specific syntax elements) to be signaled in the bitstream using variable-length bits instead of start and stop bits. In this way, a relatively small number of bits can be allocated to low values, and overall encoding efficiency can be improved.

[0219] The decoder can perform context-based decoding or bypass-based decoding on bins in a bin string from a bitstream, based on entropy coding techniques (such as CABAC, CAVLC, etc.). The entropy decoding process can be executed by the entropy decoding processor 1502 in the entropy decoder 1500. As mentioned above, the bitstream can include various information for feature / feature map decoding. As mentioned above, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.

[0220] In this disclosure, a table including syntax elements (syntax table) can be used to indicate information signaled from the encoder to the decoder. The order of the syntax elements in the table including the syntax elements used in this disclosure can represent the parsing order of the syntax elements from the bitstream. The encoder can generate and encode the syntax table such that the syntax elements can be parsed by the decoder in the parsing order, and the decoder can obtain the values ​​of the syntax elements by parsing and decoding the syntax elements of the syntax table from the bitstream in the parsing order.

[0221] Coding layer and architecture

[0222] Figure 16 This is a diagram illustrating an example of a VCM layer architecture.

[0223] Reference Figure 16 The VCM layer architecture can consist of a feature extraction layer 1610, a neural network (feature) abstraction layer 1620, and a feature encoding layer 1630.

[0224] The feature extraction layer 1610 is a layer used to extract features from the input source, and may also include the extraction results.

[0225] The feature coding layer 1630 is a layer used to compress the extracted features, and may also include the compressed result.

[0226] The neural network (feature) abstraction layer 1620 can abstract the information generated by the feature extraction layer 1610 (e.g., information about the extracted features / feature maps) and send the abstracted information to the feature encoding layer 1630. By abstracting the information, the neural network abstraction layer (NNAL) 1720 can hide the internals of the feature extraction layer 1610 and provide a consistent feature interface function. Therefore, even when the compression target changes due to changes in tools (e.g., CNN, DNN, etc.), the feature encoding layer 1630 can perform a consistent feature encoding process. In this disclosure, the NNAL may also be referred to as the "feature abstraction layer".

[0227] The interface between the feature extraction layer 1610 and NNAL 1620 and the interface between the feature encoding layer 1630 and NNAL 1620 can be predefined, and the operation of NNAL 1620 can be set to be changeable later.

[0228] Figure 17 This is a diagram illustrating an example of a bitstream consisting of encoded abstract features and NNAL information.

[0229] like Figure 17 The bitstream configured as shown can be called an "NNAL unit". An NNAL unit can be a single, independent feature reconstruction unit. Input features for a single NNAL unit can be extracted from the same layer in the neural network. Therefore, the input features for a single NNAL unit can be forced to have the same properties. For example, the same feature extraction method can be applied to the input features for a single NNAL unit.

[0230] An NNAL unit may include an NNAL unit header and an NNAL unit payload. The NNAL unit header may include all the information needed to utilize the encoded features according to the task. The NNAL unit payload may include abstract feature information. The NNAL unit payload may include a group header and group data. The group header may include information about the composition of the feature group data, such as the temporal order, number, and common properties of the feature channels that make up the feature group. Feature channels may be encoded feature units. Group data may include multiple feature channels and encoding indicators, and each of the feature channels may include type information, prediction information, side information, and residual information (data). In this case, type information may indicate the encoding method, and prediction information may indicate the prediction method. Furthermore, side information may indicate additional information required for decoding (e.g., entropy coding, quantization-related information, etc.), and residual information may include information about the encoded feature elements (i.e., a set of feature value information).

[0231] Implementation

[0232] The embodiments described in this disclosure relate to technologies associated with coding layers and architectures, and more specifically, to methods for representing coding optimization methods applied according to the purpose of a user or machine, and methods for representing the appropriate use and attributes of an encoded bitstream (including privacy-preserving attributes for users and machines, such as de-identification).

[0233] The need for privacy-preserving processing of images (e.g., de-identification) is as follows.

[0234] 1. Privacy protection: Images used for AI training and inference may contain sensitive or personally identifiable information.

[0235] 2. Legal Compliance and Ethical Use: Many countries and regions have strict regulations on privacy protection. De-identification enables compliance with such laws.

[0236] 3. Data sharing: De-identification enables secure data sharing without disclosing sensitive information.

[0237] Figure 18 Examples of encoding optimization involving privacy-preserving processing and the use of encoded bitstreams are illustrated. In this example, unnecessary information can be removed or necessary information can be enhanced during image acquisition through optimization. In some cases, personal information can constitute unnecessary information. In such cases, personal information can be protected and removed through optimization. Because sensitive personal information is removed before the image is distributed and used, privacy protection and legal compliance are achievable, and ethical use and distribution are realized. Although the necessity of privacy-preserving processing is clear, there is a problem of not defining the information associated with privacy-preserving processing in the encoded bitstream. When such information is not defined, the following problems may arise.

[0238] 1) Risks of improper use

[0239] For privacy protection purposes, personally identifiable information may be removed or modified. When such images are used for personal identification purposes, problems may arise due to misidentification.

[0240] 2) Determining the legality of the received video.

[0241] When it is unknown whether the received or collected data includes personal information and whether privacy protections have been applied, it is difficult to determine the legality of using, storing, and distributing such data.

[0242] 3) Management difficulties

[0243] When the bitstream itself does not contain information about data attributes, it may be difficult to identify and manage the attributes of the corresponding bitstream if there are problems in the management system.

[0244] First Implementation Method

[0245] Because of the global introduction of the General Data Protection Regulation (GDPR) and other similar laws for privacy protection, storing and using data that includes unauthorized personal information or personally identifiable information (PII) such as faces and vehicle license plates is legally problematic. To minimize legal issues, the attributes of the data used must be defined (such as whether the data (e.g., videos, images) includes personally identifiable information), and when privacy protections have been applied to obscure personally identifiable information, information about the methods applied must also be defined.

[0246] Figure 19 Examples of data attribute and privacy protection information definitions are provided. Data attributes and privacy protections can indicate whether personal information is included. When personal information is included, information indicating the type and attributes of that personal information can also be defined. Additionally, information indicating whether privacy protections have been applied can be defined. When privacy protections have been applied, information about the applied methods can be defined. This can be information about the applied methods, or it can be attribute information of the data that has changed as a result of applying privacy protections. For example, as a result of applying privacy protections, the attributes of the data can be changed from data containing personally identifiable information to pseudonymous data, and the information defining the changed attributes is the content of the definition.

[0247] Tables 5 to 9 illustrate examples of information used to represent data attributes and privacy protection information. Figure 19 Examples of the methods are provided and related information is defined in privacy_protection_info.

[0248] [Table 5]

[0249] [Table 6]

[0250] [Table 7]

[0251] [Table 8]

[0252] [Table 9]

[0253] privacy_protection_info can also be defined to include only Figure 19 This refers to some information defined within the `privacy_protection_info` parameter. `privacy_protection_info` can be defined within the encoded video / image bitstream, and the definition location can be multiple locations within the bitstream as needed. For example, `privacy_protection_info` can also be defined in the Network Abstraction Layer (NAL) header, sequence parameter set, image parameter set, Supplemental Enhancement Information (SEI), etc. The definition location is not limited to the locations illustrated here and can be any high-level parameter set.

[0254] Table 5 illustrates examples where privacy protection is defined in privacy_protection_info only, specifying whether privacy protection has been applied and the method of protection.

[0255] When equal to 1, `privacy_protection_cancel_flag` indicates that the persistence of previously applied privacy protection methods is canceled. When equal to 0, `privacy_protection_cancel_flag` indicates that `privacy_protection_persistence_flag` and `privacy_protection_type` must be followed, and that the privacy protection methods and scope identified by `privacy_protection_persistence_flag` and `privacy_protection_type` can be applied.

[0256] `privacy_protection_persistence_flag` indicates the persistence of the privacy protection method specified by `privacy_protection_type`. When equal to 0, `privacy_protection_persistence_flag` indicates that the privacy protection method identified by `privacy_protection_type` can be applied only to the current image. When equal to 1, `privacy_protection_persistence_flag` indicates that the privacy protection method identified by `privacy_protection_type` can be applied to the current image and all subsequent images.

[0257] privacy_protection_type is information used to represent the attributes of the privacy protections applied.

[0258] [Table 10]

[0259] [Table 11]

[0260] The examples in Table 10 illustrate how lookup tables are applied. In the examples in Table 10, the value of privacy_protection_type can be used to identify whether the applied method is obfuscation, masking, replacement, etc.

[0261] Table 11 shows the method information defined in Table 10 as the result of a bitwise AND operation between privacy_protection_type and a bitmask. Table 12 shows an example of this, and the value of privacy_protection_type can be defined differently from the values ​​in Table 10.

[0262] [Table 12]

[0263] In this case, multiple attributes can be represented. blurring_Flag, replacing_Flag, masking_Flag, adversarial_attack_Flag, and anonymization_Flag are information indicating whether privacy protection methods such as blurring, replacement, masking, adversarial perturbation, and anonymization have been applied. A value of 1 means that the corresponding method has been applied, and a value of 0 means that the corresponding method has not been applied.

[0264] The examples in Table 13 illustrate attribute information of data that has changed due to the applied privacy protections. In some cases, attribute information of data that has changed due to the application of privacy protections may be required, without information about the privacy protection method itself.

[0265] [Table 13]

[0266] [Table 14]

[0267] Table 13 illustrates examples of situations where, as a result of application privacy protection, the privacy attributes of the data have been pseudonymous, de-identified, or removed.

[0268] Table 14 shows the method information defined in Table 13 as the result of a bitwise AND operation between privacy_protection_type and a bitmask. Table 15 illustrates an example of this, and the value of privacy_protection_type can be defined differently from the values ​​in Table 13.

[0269] [Table 15]

[0270] Pseudonymized_data_Flag, anonymized_data_Flag, Removed_PII_flag, and Identified_data_flag indicate whether, as a result of the applied privacy protection method, the privacy attributes of the data have been changed to pseudonymized data, anonymized data, data with PII removed, or identified data, respectively. A value of 1 means the data has been changed to the corresponding attribute, and a value of 0 means the data has not been changed to the corresponding attribute.

[0271] The privacy protection methods defined in Tables 10, 11, 13, 14, 16, and 17 are examples, and other methods may also be defined.

[0272] Protected_info_idc represents information used to identify the type of information being protected.

[0273] [Table 16]

[0274] In Table 16, a Protected_info_idc value of 00 means that information that can identify a person (such as a face) is protected. A Protected_info_idc value of 01 means that information that can identify a vehicle or other means of transport (such as a license plate) is protected. A Protected_info_idc value of 10 means that information that can infer a location (such as text or images on signs or storefronts) is protected. A Protected_info_idc value of 11 means that information that can infer a specific object (such as text or images on a product or object) is protected. It is obvious that additional types of information to be protected can be defined separately.

[0275] [Table 17]

[0276] Table 17 illustrates examples of methods for identifying various types of protected information via bitmasks. In this case, the value of Protected_info_idc can be defined differently from the values ​​in Table 16.

[0277] In Table 18, Protected_personal_identifiable_data_Flag, Protected_vehicle_identifiable_data_Flag, Protected_location_identifiable_data_Flag, and Protected_object_identifiable_data_Flag indicate that information that can identify a person (such as a face) has been protected; information that can identify a vehicle or other means of transport (such as a license plate) has been protected; information that can infer a location (such as text or images on signs or storefronts) has been protected; and information that can infer a specific object (such as text or images on products or objects) has been protected. A value of 1 means that protection has been applied to the corresponding information, and a value of 0 means that protection has not yet been applied to the corresponding information.

[0278] [Table 18]

[0279] Table 6 above illustrates the following example: privacy_protection_info not only defines whether privacy protection has been applied and the protection methods defined in Table 5, but also defines the purpose and objectives of privacy protection.

[0280] When equal to 1, `prevent_human_perception_flag` indicates that the purpose of the applied privacy protection method is to prevent human identification of personal information. When equal to 0, `prevent_human_perception_flag` indicates that the purpose of the applied privacy protection method is not to prevent human identification of personal information.

[0281] When equal to 1, prevent_machine_perception_flag indicates that the purpose of the applied privacy protection method is to prevent machines or AI from recognizing personal information. When equal to 0, prevent_machine_perception_flag indicates that the purpose of the applied privacy protection method is not to prevent machines or AI from recognizing personal information.

[0282] Tables 7 and 8 above illustrate examples of how compliant_regulation_info is defined in Tables 5 and 6, respectively. compliant_regulation_info is information indicating what regulations are intended to be met through privacy protection.

[0283] `compliant_regulation_info` indicates information indicating what regulations are intended to be met through privacy protection. Tables 7 and 8 above illustrate examples of `compliant_regulation_info` defined in Tables 5 and 6 respectively, where `compliant_regulation_info` is information indicating what regulations are intended to be met through privacy protection.

[0284] In the example in Table 19, the regulations to be met are defined in a lookup table format based on the value of compliant_regulation_info.

[0285] [Table 19]

[0286] In this case, the value of compliant_regulation_info defines information for a target regulation.

[0287] Table 20 shows the target regulations defined in Table 19 as the result of bitwise AND of compliant_regulation_info and bitmask.

[0288] [Table 20]

[0289] In this context, `compliant_regulation_info` can represent multiple target regulations. Table 21 illustrates examples, where `compliant_GDPR_Flag`, `compliant_CCPA_Flag`, `compliant_PIPEDA_Flag`, and `compliant_PIPL_Flag` indicate whether the regulation being met (or intended to be met) by applying privacy-preserving methods is GDPR, CCPA, PIPEDA, or PIPL. A value of 1 means the corresponding regulation is met, and a value of 0 means the corresponding regulation is not met.

[0290] [Table 21]

[0291] The regulations defined in Tables 19 and 20 are examples, and it is obvious that other regulations may also be defined.

[0292] Second Implementation Method

[0293] The first embodiment illustrates an example of independently defining information regarding the application and purpose of privacy protection.

[0294] The second implementation defines the application and purpose of privacy protection as one of the optimization methods, as shown in the example in Table 29 below. That is, the privacy-related optimization type is defined in optimization_type.

[0295] Tables 22 to 24 illustrate examples of methods used to represent optimization objectives, attributes, and application scope.

[0296] [Table 22]

[0297] [Table 23]

[0298] [Table 24]

[0299] When equal to 1, `optimization_cancel_flag` indicates that the persistence of previously applied optimizations is canceled. When equal to 0, `optimization_cancel_flag` indicates that optimizations and their scope can be applied according to the `optimization_persistence_flag`, `optimization_for_machine_analysis_flag`, and `optimization_type`.

[0300] `optimization_persistence_flag` indicates the persistence of the optimization indicated by `optimization_type`. When equal to 0, `optimization_persistence_flag` indicates that the optimization identified by `optimization_type` can be applied only to the current image. When equal to 1, `optimization_persistence_flag` indicates that the optimization identified by `optimization_type` can be applied to the current image and all subsequent images.

[0301] When equal to 1, `optimization_for_machine_analysis_flag` indicates the purpose of the optimization and the encoded bitstream is intended for machine analysis (suitable for performing machine tasks). When equal to 0, the purpose of the optimization and the encoded bitstream may or may not be intended for machine analysis (if the value is "1", this indicates that the purpose of the optimization and the encoded bitstream are for machine analysis; if the value is "0", the purpose of the optimization and the encoded bitstream may or may not be for machine analysis).

[0302] When equal to 1, `optimization_for_human_viewing_flag` indicates the purpose of optimization and the encoded bitstream is intended for human viewing (suitable for human viewing). When equal to 0, the purpose of optimization and the encoded bitstream may or may not be intended for human viewing (if the value is "1", this indicates that the purpose of optimization and the encoded bitstream is for human viewing; if the value is "0", the purpose of optimization and the encoded bitstream may or may not be for human viewing).

[0303] Optimization can unintentionally affect code quality. For example, optimization might be performed for machine analysis, but it could also result in improved quality for human viewing. In this case, `optimization_for_machine_analysis_flag` could be set to 1 and `optimization_for_human_viewing_flag` to 0 to clearly define the intended purpose. As another example, optimization might be performed for human viewing, but it could also result in improved quality for machine analysis. In this case, `optimization_for_machine_analysis_flag` could be set to 0 and `optimization_for_human_viewing_flag` to 1.

[0304] There can be restrictions on the definitions of the values ​​of `optimization_for_human_viewing_flag` and `optimization_for_machine_analysis_flag`. For example, when the purpose of optimization is limited to human viewing and machine analysis, `optimization_for_human_viewing_flag` and `optimization_for_machine_analysis_flag` may not both have the value "0".

[0305] In this case, if the previous flag in `optimization_for_human_viewing_flag` and `optimization_for_machine_analysis_flag` is "0", then the subsequent flag is "1", and the value does not need to be encoded in the bitstream; the decoder or receiver can process the value as "1". Tables 23 and 24 above illustrate examples of this.

[0306] The `optimization_type` attribute indicates the optimization method.

[0307] [Table 25]

[0308] [Table 26]

[0309] The optimization attributes defined in Tables 25 and 26 can be identified. The optimization attributes defined in Tables 25 and 26 are examples, and new optimization attributes can be defined using the corresponding structures. Needless to say, optimization attributes and methods can be identified by including only any subset (not all) of the optimization attributes defined in Tables 25 and 26, or can be configured to include other optimization attributes and methods not defined in Tables 25 and 26. Tables 25 and 26 perform the same function but represent examples in different ways.

[0310] Table 26 represents another example of multiple optimization methods using a bitmask for optimization_type, and Tables 27 and 28 illustrate examples of them.

[0311] [Table 27]

[0312] [Table 28]

[0313] The `optimizationForPIIProtectionFlag` indicates whether the applied optimization method includes personal information protection optimization. A value of 1 means that the applied optimization method includes personal information protection optimization, and a value of 0 means that personal information protection optimization is not included.

[0314] Table 29 illustrates examples of additional information required based on the optimized attributes.

[0315] [Table 29]

[0316] `compliant_regulation_info` indicates information indicating which regulations are intended to be met through privacy protection. In the example in Table 30, the regulations to be met are defined in a lookup table format based on the value of `compliant_regulation_info`.

[0317] [Table 30]

[0318] In this case, the value of compliant_regulation_info defines information for a target regulation. Table 31 shows the target regulations defined in Table 30 as the result of a bitwise AND operation between compliant_regulation_info and a bitmask.

[0319] [Table 31]

[0320] In this context, `compliant_regulation_info` can represent multiple target regulations. Table 32 illustrates examples, where `compliant_GDPR_Flag`, `compliant_CCPA_Flag`, `compliant_PIPEDA_Flag`, and `compliant_PIPL_Flag` indicate whether the regulation being met (or intended to be met) by applying privacy-preserving methods is GDPR, CCPA, PIPEDA, or PIPL. A value of 1 means the corresponding regulation is met, and a value of 0 means the corresponding regulation is not met.

[0321] [Table 32]

[0322] The regulations defined in Tables 30 and 31 are examples, and it is obvious that other regulations may also be defined.

[0323] `privacy_protection_type` is information used to represent the attributes of the privacy protections applied. The examples in Table 33 represent the applied methods in a lookup table format.

[0324] [Table 33]

[0325] In the example in Table 33, the value of privacy_protection_type can be used to identify whether the applied method is obfuscation, masking, replacement, etc.

[0326] [Table 34]

[0327] Table 34 shows the method information defined in Table 33 as the result of bitwise AND of privacy_protection_type and bitmask, and Table 35 below illustrates examples of it.

[0328] [Table 35]

[0329] In this case, multiple attributes can be represented. blurring_Flag, replacing_Flag, masking_Flag, adversarial_attack_Flag, and anonymization_Flag are information indicating whether privacy protection methods such as blurring, replacement, masking, adversarial perturbation, and anonymization have been applied. A value of 1 means that the corresponding method has been applied, and a value of 0 means that the corresponding method has not been applied.

[0330] The examples in Table 36 represent attribute information of data that has changed due to the privacy protections applied.

[0331] [Table 36]

[0332] In some cases, information about changes to privacy attributes that occur as a result of applied privacy protection may be needed, rather than information about the privacy protection method itself. Table 36 illustrates examples of situations where, as a result of applied privacy protection, the privacy attributes of data have been pseudonymous, de-identified, or removed.

[0333] [Table 37]

[0334] Table 37 shows the method information defined in Table 36 as the result of bitwise AND of privacy_protection_type and bitmask, and Table 38 below illustrates examples of it.

[0335] [Table 38]

[0336] Pseudonymized_data_Flag, anonymized_data_Flag, Removed_PII_flag, and Identified_data_flag indicate whether, as a result of the applied privacy protection method, the privacy attributes of the data have been changed to pseudonymized data, anonymized data, data with PII removed, or identified data, respectively. A value of 1 means the data has been changed to the corresponding attribute, and a value of 0 means the data has not been changed to the corresponding attribute.

[0337] The privacy protection methods defined in Tables 33, 34, 36 and 37 are examples, and other methods may also be defined.

[0338] Protected_info_idc represents information used to identify the type of information being protected.

[0339] [Table 39]

[0340] In Table 39, a Protected_info_idc value of 00 means that information that can identify a person (such as a face) is protected. A Protected_info_idc value of 01 means that information that can identify a vehicle or other means of transport (such as a license plate) is protected. A Protected_info_idc value of 10 means that information that can infer a location (such as text or images on signs or storefronts) is protected. A Protected_info_idc value of 11 means that information that can infer a specific object (such as text or images on a product or object) is protected. It is obvious that additional types of information to be protected can be defined separately.

[0341] [Table 40]

[0342] Table 40 illustrates examples of methods for identifying various types of protected information via bitmasks. In this case, the value of Protected_info_idc can be defined differently from the values ​​in Table 39.

[0343] In Table 41, Protected_personal_identifiable_data_Flag, Protected_vehicle_identifiable_data_Flag, Protected_location_identifiable_data_Flag, and Protected_object_identifiable_data_Flag indicate that information that can identify a person (such as a face) has been protected; information that can identify a vehicle or other means of transport (such as a license plate) has been protected; information that can infer a location (such as text or images on signs or storefronts) has been protected; and information that can infer a specific object (such as text or images on products or objects) has been protected. A value of 1 means that protection has been applied to the corresponding information, and a value of 0 means that protection has not yet been applied to the corresponding information.

[0344] [Table 41]

[0345] Third Implementation Method

[0346] Figure 20 An example is shown where the areas required for image analysis coexist with areas containing personal information.

[0347] exist Figure 20 In the example, region A is needed to detect people. There exists an SEI (Search Engine Information) defining information about regions needed for a specific purpose. Similarly, information about regions unnecessary for a specific purpose or requiring attention when used for that purpose may also be needed. For example, personally identifiable information may reside in the defined required region or other regions, and modifications or removal optimizations may be applied to such regions for privacy protection purposes. It may also be necessary to define an SEI (Search Engine Information) regarding corresponding regions to prevent misuse. For example, in... Figure 20 In the case of region B, which corresponds to a face and contains personally identifiable information, the personally identifiable information in region B can be removed or modified, and information about the corresponding region can be defined.

[0348] Tables 42, 43, 44, and 45 illustrate examples of methods for representing regions and privacy-preserving regions used for image analysis.

[0349] [Table 42]

[0350] [Table 43]

[0351] Tables 42 and 43 independently define information about whether privacy protections have been applied to each defined region. For example, information about region A and region B (e.g., size, location, etc.) is defined separately.

[0352] [Table 44]

[0353] Table 44 illustrates examples of situations where the privacy protection area or personal information area is partially included within the defined annotation area.

[0354] [Table 45]

[0355] In Table 45, when num_protected_region_minus1 is 0, the region defined by ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] can be a completely privacy-protected processing region or a region containing personal information.

[0356] The information defined in Tables 42, 43, 44 and 45 can be reorganized by subsets or combinations thereof.

[0357] The SEI message for annotation regions conveys parameters used to identify annotation regions using bounding boxes that indicate the size and position of the identified objects. Using this SEI message requires defining the following variables.

[0358] - The cropped image width and height, expressed in units of brightness samples, are denoted as CroppedWidth and CroppedHeight, respectively.

[0359] - Chroma subsampling width and height, denoted as SubWidthC and SubHeightC respectively.

[0360] - Consistent clipping window left offset, ConfWinLeftOffset

[0361] - Consistent clipping window top offset, ConfWinTopOffset

[0362] The following describes examples of the semantics for Tables 42, 43, 44, and 45.

[0363] `ar_cancel_flag` equal to 1 indicates that the annotation region SEI message cancels the persistence of the previous annotation region SEI message. Here, the previous annotation region SEI is associated with one or more layers to which the annotation region SEI message was applied. `ar_cancel_flag` equal to 0 indicates that the annotation region information follows.

[0364] When ar_cancel_flag equals 1 or a new CVS for the current layer begins, the variables LabelAssigned[i], ObjectTracked[i], and ObjectBoundingBoxAvail are set to 0 for the range of 0 to 255 (inclusive).

[0365] `ar_not_optimized_for_viewing_flag` equal to 1 indicates that the SEI message applied to the annotated region was not optimized for human viewing, but rather for some other purpose, such as the performance of the algorithm's object classification. `ar_not_optimized_for_viewing_flag` equal to 0 indicates that the SEI message applied to the annotated region may or may not be optimized for human viewing.

[0366] `ar_not_optimized_for_machine_analysis_flag` equal to 1 indicates that the SEI message applied to the annotated region was not optimized for machine analysis. `ar_not_optimized_for_machine_analysis_flag` equal to 0 indicates that the SEI message applied to the annotated region may or may not be optimized for machine analysis.

[0367] An `ar_true_motion_flag` value of 1 indicates that the motion information in the encoded image to which the annotation region SEI message has been applied has been selected for the purpose of accurately representing the motion of objects within the annotation region. An `ar_true_motion_flag` value of 0 indicates whether or not the motion information in the encoded image to which the annotation region SEI message has been applied has been selected for the purpose of accurately representing the motion of objects within the annotation region.

[0368] `ar_occluded_object_flag` equal to 1 specifies that each of the syntax elements `ar_bounding_box_top[ar_object_idx[i]]`, `ar_bounding_box_left[ar_object_idx[i]]`, `ar_bounding_box_width[ar_object_idx[i]]`, and `ar_bounding_box_height[ar_object_idx[i]]` indicates the size and position of an object or a portion of an object that is not visible or is only partially visible within the cropped, decoded image. `ar_occluded_object_flag` equal to 0 specifies that each of the syntax elements `ar_bounding_box_top[ar_object_idx[i]]`, `ar_bounding_box_left[ar_object_idx[i]]`, `ar_bounding_box_width[ar_object_idx[i]]`, and `ar_bounding_box_height[ar_object_idx[i]]` indicates the size and position of an object that is fully visible within the cropped, decoded image. The requirement for bitstream consistency is that the value of ar_occluded_object_flag should be the same for all annotated_regions() syntax structures in CVS.

[0369] A value of 1 for `ar_partial_object_flag_present_flag` indicates the existence of the syntax element `ar_partial_object_flag[ar_object_idx[i]]`. A value of 0 for `ar_partial_object_flag_present_flag` indicates the absence of the syntax element `ar_partial_object_flag[ar_object_idx[i]]`. The requirement for bitstream consistency is that the value of `ar_partial_object_flag_present_flag` should be identical for all `annotated_regions()` syntax structures in CVS.

[0370] An `ar_object_label_present_flag` value of 1 indicates that there is label information corresponding to the object in the annotation region. An `ar_object_label_present_flag` value of 0 indicates that there is no label information corresponding to the object in the annotation region.

[0371] A value of 1 for `ar_privacy_protection_region_present_flag` indicates the existence of the syntax element `ar_privacy_protection_region_flag[ar_object_idx[i]]`. A value of 0 for `ar_privacy_protection_region_present_flag` indicates the absence of the syntax element `ar_privacy_protection_region_flag[ar_object_idx[i]]`. The requirement for bitstream consistency is that the value of `ar_privacy_protection_region_present_flag` should be identical for all `annotated_regions()` syntax structures in CVS.

[0372] The value of ar_privacy_protection_region_flag[ar_object_idx[i]] equals 1. The syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicate the size and position of the privacy-protected region within the decoded image. The value of ar_privacy_protection_region_flag[ar_object_idx[i]] being equal to 0 specifies that the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicate the size and location of a privacy-protected region that may or may not exist within the decoded image. When it does not exist, the value of ar_partial_object_flag[ar_object_idx[i]] is inferred from the previously annotated region SEI message (if it exists) in the order of CVS output.

[0373] `ar_include_privacy_protection_region_flag[ar_object_idx[i]]` equal to 1 indicates that the defined annotation region partially includes a region where privacy protection has been applied, or partially includes personally identifiable information. `ar_include_privacy_protection_region_flag[ar_object_idx[i]]` equal to 0 indicates that the defined annotation region may or may not partially contain personally identifiable information, and privacy protection is applied to that region.

[0374] `num_protected_region_minus1` indicates the number of regions containing protected personally identifiable information or the total number of regions containing personally identifiable information minus 1. The value of `num_protected_region` should be in the range of 0 to 255 (inclusive).

[0375] pii_region_top[j], pii_region_left[j], pii_region_width[j], and pii_region_height[j] specify the coordinates of the top-left corner, as well as the width and height, of the privacy-protected region or the region containing personally identifiable information within the area defined by ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]], respectively.

[0376] The ar_object_confidence_info_present_flag being equal to 1 indicates the existence of the ar_object_confidence[ar_object_idx[i]] syntax element.

[0377] A value of 0 for ar_object_confidence_info_present_flag indicates that the ar_object_confidence[ar_object_idx[i]] syntax element does not exist. The requirement for bitstream consistency is that the value of ar_object_confidence_present_flag should be identical for all annotated_regions() syntax structures in CVS.

[0378] `ar_object_confidence_length_minus1 + 1` specifies the length in bits of the `ar_object_confidence[ar_object_idx[i]]` syntax element. The requirement for bitstream consistency is that the value of `ar_object_confidence_length_minus1` should be identical for all `annotated_regions()` syntax structures in CVS.

[0379] An ar_object_label_language_present_flag value of 1 indicates the existence of the ar_object_label_language syntax element. An ar_object_label_language_present_flag value of 0 indicates the non-existence of the ar_object_label_language syntax element.

[0380] ar_bit_equal_to_zero should be equal to 0.

[0381] `ar_object_label_language` includes the language label with a null-terminated byte of 0x00 as specified in IETF RFC 5646. The length of the `ar_object_label_language` syntax element should be less than or equal to 255 bytes, excluding the null-terminated byte. If it does not exist, the language of the label is not specified.

[0382] `ar_num_label_updates` indicates the total number of labels associated with the annotation area to be signaled. The value of `ar_num_label_updates` should be in the range of 0 to 255 (inclusive).

[0383] ar_label_idx[i] indicates the index of the label that is signaled. The value of ar_label_idx[i] should be in the range of 0 to 255 (inclusive).

[0384] An ar_label_cancel_flag of 1 cancels the persistent range of the ar_label_idx[i]. An ar_label_cancel_flag of 0 indicates that the value signaled by the signal has been assigned to the ar_label_idx[i].

[0385] ar_label[ar_label_idx[i]] specifies the content of the ar_label_idx[i]. The length of the ar_label[ar_label_idx[i]] syntax element should be less than or equal to 255 bytes, excluding the null terminator.

[0386] `ar_num_object_updates` indicates the number of object updates to be signaled. `ar_num_object_updates` should be in the range of 0 to 255 (inclusive).

[0387] ar_object_idx[i] is the index of the object parameter to be signaled. ar_object_idx[i] should be in the range of 0 to 255 (inclusive).

[0388] An ar_object_cancel_flag of 1 cancels the persistent scope of the ar_object_idx[i]. An ar_object_cancel_flag of 0 indicates that the parameters associated with the tracked ar_object_idx[i] object will be signaled.

[0389] An ar_object_label_update_flag value of 1 indicates that the object label will be notified by signaling. An ar_object_label_update_flag value of 0 indicates that the object label will not be notified by signaling.

[0390] `ar_object_label_idx[ar_object_idx[i]]` indicates the index of the label corresponding to the `ar_object_idx[i]`-th object. If `ar_object_label_idx[ar_object_idx[i]]` does not exist, its value is inferred from the previous annotation region SEI message (if present) in the same CVS output order.

[0391] An ar_bounding_box_update_flag value of 1 indicates that the object bounding box parameters will be signaled. An ar_bounding_box_update_flag value of 0 indicates that the object bounding box parameters will not be signaled.

[0392] ar_bounding_box_cancel_flag equals 1 to cancel ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_obje The persistence scope of ct_idx[i]], ar_bounding_box_height[ar_object_idx[i]], ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]]. ar_bounding_box_cancel_flag equal to 0 indicates that the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width will be signaled [ar_object_idx[i]], ar_bounding_box_height[ar_object_idx[i]], ar_partial_object_flag[ar_object_idx[i]], and ar_object_confidence[ar_object_idx[i]].

[0393] ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] specify the coordinates of the top-left corner, width, and height of the bounding box of the ar_object_idx[i]-th object in the cropped, decoded image associated with the consistent cropping window specified by the active SPS, respectively.

[0394] The value of ar_bounding_box_left[ar_object_idx[i]] should be in the range of 0 to CroppedWidth / SubWidthC-1 (inclusive).

[0395] The value of ar_bounding_box_top[ar_object_idx[i]] should be in the range of 0 to CroppedHeight / SubHeightC-1 (inclusive).

[0396] The value of ar_bounding_box_width[ar_object_idx[i]] should be in the range of 0 to CroppedWidth / SubWidthC -ar_bounding_box_left[ar_object_idx[i]] (inclusive of 0 and CroppedWidth / SubWidthC -ar_bounding_box_left[ar_object_idx[i]]).

[0397] The value of ar_bounding_box_height[ar_object_idx[i]] should be in the range of 0 to CroppedHeight / SubHeightC - ar_bounding_box_top[ar_object_idx[i]] (inclusive of 0 and CroppedHeight / SubHeightC - ar_bounding_box_top[ar_object_idx[i]]).

[0398] The identified object rectangle contains a brightness sample, which has a value from SubWidthC (ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]]) to SubWidthC (ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]] + ar_bounding_box_width[ar_object_idx[i]]) - 1 (including SubWidthC (ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]]) and SubWidthC (ConfWinLeftOffset + ar_bounding_box_left[ar_object_idx[i]] + ar_bounding_box_width[ar_object_idx[i]]) - 1) are the horizontal image coordinates, and have a value from SubHeightC (ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]]) to SubHeightC (ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]] + ar_bounding_box_height[ar_object_idx[i]]) - 1 (including SubHeightC (ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]]) and SubHeightC (ConfWinTopOffset + ar_bounding_box_top[ar_object_idx[i]] + ar_bounding_box_height[ar_object_idx[i]]) - 1) is the vertical image coordinate.

[0399] For each value of ar_object_idx[i], maintain the values ​​of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] in CVS in the output order. If they do not exist, infer the values ​​of ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], or ar_bounding_box_height[ar_object_idx[i]] from the previously annotated region SEI message (if it exists) in the output order of CVS.

[0400] ar_partial_object_flag[ar_object_idx[i]] equal to 1 specifies

[0401] The syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicate the size and position of objects that are only partially visible within the cropped, decoded image.

[0402] The value of ar_partial_object_flag[ar_object_idx[i]] equal to 0 specifies that the syntax elements ar_bounding_box_top[ar_object_idx[i]], ar_bounding_box_left[ar_object_idx[i]], ar_bounding_box_width[ar_object_idx[i]], and ar_bounding_box_height[ar_object_idx[i]] indicate the size and position of objects that may be partially visible or invisible within the cropped, decoded image. When not present, the value of ar_partial_object_flag[ar_object_idx[i]] is inferred from the previously annotated region SEI messages (if present) in the order of CVS output.

[0403] `ar_object_confidence[ar_object_idx[i]]` indicates the confidence associated with the `ar_object_idx[i]`-th object in units of 2 - (ar_object_confidence_length_minus1 + 1), such that a higher value of `ar_object_confidence[ar_object_idx[i]]` indicates a higher confidence. The length of the syntax element `ar_object_confidence[ar_object_idx[i]]` can be `ar_object_confidence_length_minus1 + 1` bits. When it does not exist, the value of `ar_object_confidence[ar_object_idx[i]]` is inferred from the previously annotated SEI messages (if present) in the CVS output order.

[0404] Fourth Implementation Method

[0405] Generative Face Video (GFV) has been proposed, which encodes / decodes faces through a combination of neural network-based techniques and conventional video encoding / decoding techniques.

[0406] Table 46 illustrates examples of the proposed GFV syntax.

[0407] [Table 46]

[0408] In the context of generative videos, modifications to personally identifiable information are permissible. Furthermore, facial information can be intentionally modified based on the base image and facial generation network used as the basis for facial generation. Therefore, in some cases, when used for personal identification purposes, care must be taken to determine whether personally identifiable information has been altered or whether the data is suitable for such purposes to prevent misuse.

[0409] Tables 47, 48, and 49 illustrate examples of defining whether personally identifiable information has been altered or whether the data is suitable for personal identification purposes in the generative facial video SEI message syntax.

[0410] [Table 47]

[0411] [Table 48]

[0412] [Table 49]

[0413] A value of 1 for `not_optimized_for_machine_analysis_flag` indicates that the generated / decoded result may not be optimized (may not be suitable) for machine analysis purposes. "Unsuitable" means the machine analysis result may differ from the expected result or may be inaccurate. For example, when the shape or color of an object changes or some information is removed, the object recognition result may be inaccurate. When equal to 0, `not_optimized_for_machine_analysis_flag` indicates that the generated / decoded result may not be unsuitable for machine analysis purposes. This syntax can be applied not only to GFV but also to other generated / decoded results.

[0414] A value of 1 for `gfv_not_optimized_for_personal_identification_flag` indicates that the result generated / decoded by GFV may not be optimized (may not be suitable) for personal identification purposes. A value of 0 indicates that `gfv_not_optimized_for_personal_identification_flag` may not be unsuitable for personal identification purposes.

[0415] A value of 1 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., the base image). When equal to 0, `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) are the same as the attributes of the input data (e.g., the base image).

[0416] gfv_matrix_property is information about the attributes that indicate the information of each matrix.

[0417] Table 47 illustrates an example of defining gfv_matrix_property when gfv_matrix_transformed_flag equals 1. In this case, gfv_matrix_property can be defined as follows.

[0418] The properties defined by `gfv_matrix_property` can include transformations, removals, enhancements, etc., of information used to achieve matrix purposes. When equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed differently from the properties of the input data (e.g., the base image). When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information in the input data (e.g., the base image). The items listed above are examples, and additional information indicating the properties of the matrix can be defined separately. `gfv_matrix_property[i]` indicates the properties of the matrix represented by `gfv_matrix_type_idx[i]`.

[0419] Table 48 illustrates an example of defining gfv_matrix_transformed_flag independently.

[0420] Table 49 illustrates an example of defining gfv_matrix_property independently. In this case, the definition of gfv_matrix_property can be as follows.

[0421] `gfv_matrix_property` is information indicating the properties of each matrix. Properties defined by `gfv_matrix_property` can include transformations, removals, enhancements, etc. As an example, when equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) are the same as the properties of the input data (e.g., the base image). When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed differently from the properties of the input data (e.g., the base image). When equal to 2, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information in the input data (e.g., the base image). The items listed above are examples, and additional information indicating properties of the matrix can be defined separately. gfv_matrix_property[i] indicates the property of the matrix represented by gfv_matrix_type_idx[i].

[0422] A value of 1 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., the base image). When equal to 0, `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix are the same as the attributes of the input data.

[0423] The examples above illustrate the usage of `gfv_matrix_transformed_flag` and `gfv_matrix_property` as defined, and other usage forms are also possible. For example, two pieces of information can be defined together without separate conditions.

[0424] Generative Face Video (GFV) SEI messages specify facial parameters and a facial parameter transformer neural network (denoted as TranslatorNN()) and a facial image generator neural network (denoted as GenerativeNN()). Here, TranslatorNN() can be used to transform facial parameters of various formats signaled by the SEI message into parameters of a fixed format. GenerativeNN() can be used to generate an output image using the format of the facial parameters and a previously decoded output image.

[0425] Note 1 - Facial parameters can be determined from the source image before encoding. This source image can be called the driving image.

[0426] Note 2 - The previously decoded output image input to GenerativeNN() can be a base image (a decoded output image that provides a reference texture from which a face image can be generated) and optionally, an image that can be fused by GenerativeNN() to enhance background texture and facial details. When the current image is not a base image, GFV SEI messages can be used to generate a face image for fusion purposes based on the previously decoded base image, facial parameters conveyed by the GFV SEI message, and optionally the current decoded image.

[0427] To use this SEI message, you need to define the following variables.

[0428] - The width and height of the input image, expressed in units of brightness samples, are denoted as CroppedWidth and CroppedHeight, respectively.

[0429] - A luminance sample array baseCroppedYPic and chrominance sample arrays baseCroppedCbPic and baseCroppedCrPic for the decoded output image corresponding to the source base image (denoted as BasePicture).

[0430] - The luminance sample array driveCroppedYPic and the chrominance sample arrays driveCroppedCbPic and driveCroppedCrPic for the decoded output image corresponding to the source driving image (denoted as DrivePicture).

[0431] - BitDepth of the luminance sample array for the input image Y

[0432] - BitDepth of the chroma sample array for the input image (if it exists) C

[0433] - Chroma format indicator, denoted as ChromaFormatIdc.

[0434] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.

[0435] `gfv_id` includes an identifier that can be used to identify facial feature information and specify the neural network that can be used as a GenerativeNN(). The value of `gfv_id` should be between 0 and 2. 32 - The range of 2 (inclusive of 0 and 2) 32 - 2) Within. gfv_id in the range of 256 to 511 and in 2 31 Up to 2 32 Values ​​in the range of -2 are reserved for future use by ITU-T|ISO / IEC. The decoder should ignore values ​​included in the range of 256 to 511 or in 2... 31 Up to 2 32 - GFV SEI messages with gfv_id in the range of 2.

[0436] Note - For example, when there is more than one face in the output image, different values ​​of gfv_id in different GFV SEI messages can be used to identify different faces.

[0437] A value of 1 for `gfv_base_pic_flag` indicates that the currently decoded output image corresponds to the base image. A value of 0 for `gfv_base_pic_flag` indicates that the currently decoded output image does not correspond to the base image.

[0438] The following constraints apply to the value of gfv_base_pic_flag.

[0439] - When the GFV SEI message is the first GFV SEI message in the current CLVS with a specific gfv_id value in decoding order, the value of gfv_base_pic_flag should be equal to 1.

[0440] - When a GFV SEI message with a specific gfv_id value has a gfv_base_pic_flag equal to 0, the SEI message is associated with the currently decoded picture and all subsequent decoded pictures of the current layer in output order, until the end of the current CLVS, or until, but not including, the decoded picture in the current CLVS that follows the currently decoded picture in output order and is associated with subsequent GFV SEI messages in the current CLVS in decoding order with a gfv_base_pic_flag equal to 0 and the specific gfv_id value, whichever comes first.

[0441] `gfv_nn_base_flag`, `gfv_nn_mode_idc`, `gfv_nn_reserved_zero_bit_a`, `gfv_nn_tag_uri`, `gfv_nn_uri`, and `gfv_nn_payload_byte[i]` specify neural networks that can be used as TranslatorNN(). `gfv_nn_base_flag`, `gfv_nn_mode_idc`, `gfv_nn_reserved_zero_bit_a`, `gfv_nn_tag_uri`, `gfv_nn_uri`, and `gfv_nn_payload_byte[i]` have the same syntax and semantics as `nnpfc_base_flag`, `nnpfc_mode_idc`, `nnpfc_reserved_zero_bit_a`, `nnpfc_tag_uri`, `nnpfc_uri`, and `nnpfc_payload_byte[i]`, respectively.

[0442] When present, a `gfv_drive_pic_fusion_flag` of 1 indicates that the currently decoded image can be input into `GenerativeNN()`. The currently decoded image corresponds to the driving image that can be used for fusion. A `gfv_drive_pic_fusion_flag` of 0 indicates that the currently decoded image should not be input into `GenerativeNN()`.

[0443] Note 3 - For example, a value of gfv_drive_pic_fusion_flag equal to 1 can be used to indicate whether the currently decoded image can enhance facial details or handle background changes.

[0444] Note 4 - Fusion uses three inputs to output the image: the base image, features from keypoints and / or matrices conveyed in the GFV SEI message, and the currently decoded image.

[0445] Note 5 - When the currently decoded image corresponds to the driving image, it should be marked as not intended for output.

[0446] A `not_optimized_for_machine_analysis_flag` value of 1 indicates that the generated / decoded result may not be optimized (may not be suitable) for machine analysis purposes. "Unsuitable" means the machine analysis result may differ from the expected result or may be inaccurate. A `not_optimized_for_machine_analysis_flag` value of 0 indicates that the generated / decoded result is suitable (not unsuitable) for machine analysis purposes. This syntax can be applied not only to GFV but also to other generated / decoded results.

[0447] A value of 1 for `gfv_not_optimized_for_personal_identification_flag` indicates that the result generated / decoded by GFV may not be optimized (may not be suitable) for personal identification purposes. A value of 0 for `gfv_not_optimized_for_personal_identification_flag` indicates that the result generated / decoded by GFV may be suitable (not unsuitable) for personal identification purposes.

[0448] `gfv_matrix_property` indicates information about the properties of each matrix. Properties defined by `gfv_matrix_property` include transformations, removals, enhancements, etc. For example, when equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) are the same as the properties of the input data (e.g., a base image). When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. When equal to 2, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information. Additional information defining the properties of the matrix can be specified.

[0449] When gfv_matrix_transformed_flag is set to 1, the semantics of gfv_matrix_property defined under this condition can be specified as follows.

[0450] `gfv_matrix_property` indicates information about the properties of each matrix. Properties defined by `gfv_matrix_property` include transformations, removals, enhancements, etc. For example, when equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information. Additional information defining the properties of the matrix can be specified.

[0451] A value of 1 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., the base image). A value of 0 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix are the same as the attributes of the input data.

[0452] A value of 1 for `gfv_coordinate_present_flag` indicates the existence of keypoint coordinate information. A value of 0 for `gfv_coordinate_present_flag` indicates the absence of keypoint coordinate information.

[0453] The requirement for bitstream consistency is that when gfv_matrix_type_idx[i] is equal to 0 or 1 for all i in the range from 0 to gfv_num_matrix_types_minus1 (inclusive), the value of gfv_coordinate_present_flag should be equal to 1.

[0454] Incrementing gfv_coordinate_precision_factor_minus1 by 1 indicates the bit length of gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i].

[0455] Incrementing `gfv_num_kps_minus1` by 1 indicates the number of keypoints. The value of `gfv_num_kp_minus1` should be between 0 and 2. 10 -1 range (including 0 and 2) 10 -1) inside.

[0456] A value of 1 for `gfv_kp_pred_flag` indicates the existence of the syntax elements `gfv_coordinate_dx_abs[i]`, `gfv_coordinate_dy_abs[i]`, and `gfv_coordinate_dz_abs[i]`, and may also contain the syntax elements `gfv_coordinate_dx_sign_flag[i]`, `gfv_coordinate_dy_sign_flag[i]`, and `gfv_coordinate_dz_sign_flag[i]`. A value of 0 for `gfv_kp_pred_flag` indicates the existence of `gfv_coordinate_x_abs[i]`, `gfv_coordinate_y_abs[i]`, and `gfv_coordinate_z_abs[i]`, and may also contain the syntax elements `gfv_coordinate_x_sign_flag[i]`, `gfv_coordinate_y_sign_flag[i]`, and `gfv_coordinate_z_sign_flag[i]`.

[0457] A value of 1 for `gfv_coordinate_z_present_flag` indicates the presence of z-axis coordinate information for keypoints. A value of 0 for `gfv_coordinate_z_present_flag` indicates the absence of z-axis coordinate information for keypoints.

[0458] The value of gfv_coordinate_z_max_value_minus1 plus 1 indicates the maximum absolute value of the z-axis coordinate of the key point.

[0459] gfv_coordinate_x_abs[i] indicates the normalized absolute value of the x-axis coordinate of the i-th key point.

[0460] `gfv_coordinate_x_sign_flag[i]` specifies the sign of the x-axis coordinate of the i-th keypoint. When `gfv_coordinate_x_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0461] gfv_coordinate_y_abs[i] indicates the normalized absolute value of the y-axis coordinate of the i-th key point.

[0462] `gfv_coordinate_y_sign_flag[i]` specifies the sign of the y-axis coordinate of the i-th keypoint. When `gfv_coordinate_y_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0463] gfv_coordinate_z_abs[i] indicates the normalized absolute value of the z-axis coordinate of the i-th key point.

[0464] `gfv_coordinate_z_sign_flag[i]` specifies the sign of the z-axis coordinate of the i-th keypoint. When `gfv_coordinate_z_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0465] gfv_coordinate_dx_abs[i] indicates the absolute difference of the normalized values ​​of the x-axis coordinates of the i-th key point.

[0466] `gfv_coordinate_dx_sign_flag[i]` specifies the sign of the difference in the x-axis coordinates of the i-th keypoint. When `gfv_coordinate_dx_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0467] gfv_coordinate_dy_abs[i] specifies the absolute difference of the normalized values ​​of the y-axis coordinates of the i-th keypoint.

[0468] `gfv_coordinate_dy_sign_flag[i]` specifies the sign of the difference in the y-axis coordinates of the i-th keypoint. When `gfv_coordinate_dy_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0469] gfv_coordinate_dz_abs[i] specifies the absolute difference of the normalized values ​​of the z-axis coordinates of the i-th keypoint.

[0470] `gfv_coordinate_dz_sign_flag[i]` specifies the sign of the difference in the z-axis coordinates of the i-th keypoint. When `gfv_coordinate_dz_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0471] As shown in Table 50 below, derive the variables coordinateDeltaX[i], coordinateDeltaY[i], and coordinateDeltaZ[i], which represent the incremental x-axis coordinate, incremental y-axis coordinate, and incremental z-axis coordinate of the i-th key point, respectively.

[0472] [Table 50]

[0473] When gfv_kp_pred_flag equals 0, variables coordinateX[i], coordinateY[i], and coordinateZ[i] representing the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the i-th keypoint are derived as shown in Table 51 below. When gfv_kp_pred_flag equals 1, the variables are derived as shown in Table 52 below.

[0474] [Table 51]

[0475] [Table 52]

[0476] Here, as shown in Table 53 below, we derive BaseKpCoordinateX[i], BaseKpCoordinateY[i], and BaseKpCoordinateZ[i], which represent the x-axis coordinates, y-axis coordinates, and z-axis coordinates of the i-th keypoint of the base image, respectively.

[0477] [Table 53]

[0478] A value of 1 for `gfv_matrix_present_flag` indicates the presence of a matrix parameter. A value of 0 for `gfv_matrix_present_flag` indicates the absence of a matrix parameter.

[0479] Incrementing gfv_matrix_element_precision_factor_minus1 by 1 indicates the length of gfv_matrix_element_dec[i][j][k][m] in bits.

[0480] Incrementing `gfv_num_matrix_types_minus1` by 1 indicates the number of matrix types signaled in the SEI message. The value of `gfv_matrix_type_num_minus1` should be between 0 and 2. 6 -1 range (including 0 and 2) 6 -1) inside.

[0481] gfv_matrix_type_idx[i] indicates the index of the i-th matrix type as specified in Table 54.

[0482] [Table 54]

[0483] Note 2. An undefined matrix type is used to indicate a matrix type that is not an affine transformation matrix, covariance matrix, rotation matrix, translation matrix, or compact eigenma matrix. It can be used by the user to extend the matrix type.

[0484] A value of 1 for `gfv_num_matrices_equal_to_num_kps_flag[i]` indicates that the number of matrices of the i-th matrix type is equal to `gfv_num_kps_minus1 + 1`. A value of 0 for `gfv_num_matrices_equal_to_num_kps_flag[i]` indicates that the number of matrices of the i-th matrix type is not equal to `gfv_num_kps_minus1 + 1`.

[0485] gfv_num_matrices_info[i] provides information about the number of matrices of type i used to derive the matrix.

[0486] gfv_matrix_width_minus1[i] increments by 1 to indicate the width of the matrix of type i.

[0487] The increment of gfv_matrix_height_minus1[i] indicates the height of the matrix of type i.

[0488] A value of 1 for `gfv_matrix_for_3D_space_flag[i]` indicates that the matrix of type `i` is defined in three-dimensional space. A value of 0 for `gfv_matrix_for_3D_space_flag[i]` indicates that the matrix of type `i` is defined in two-dimensional space.

[0489] When gfv_matrix_width_minus1[i] does not exist, it is inferred as follows.

[0490] - When gfv_matrix_type_idx[i] is equal to 0, 1 or 4, and one of coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] exists and is equal to 1, gfv_matrix_width_minus1[i] is inferred to be equal to 2.

[0491] Otherwise, gfv_matrix_width_minus1[i] is inferred to be equal to 1 when matrix_type_idx[i] is equal to 0, 1 or 4, and one of coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] exists and is equal to 0.

[0492] Otherwise (when matrix_type_idx[i] equals 5 or 6), gfv_matrix_width_minus1[i] is inferred to be equal to 0.

[0493] When gfv_matrix_height_minus1[i] does not exist, it is inferred as follows.

[0494] - When matrix_type_idx is equal to 0, 1, 4, 5 or 6, and one of gfv_coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] exists and is equal to 1, gfv_matrix_height_minus1[i] is inferred to be equal to 2.

[0495] Otherwise (when gfv_matrix_type_idx is equal to 0, 1, 4, 5 or 6, and one of gfv_coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] is equal to 0), gfv_matrix_height_minus1[i] is inferred to be equal to 1.

[0496] As shown in Table 55 below, derive the variables matrixWidth[i] and matrixHeight[i], which represent the width and height of the matrix of type i, respectively.

[0497] [Table 55]

[0498] `gfv_num_matrices_minus1[i]` increments by 1 to indicate the number of matrices of type `i`. The variable `numMatrices[i]` representing the number of matrices of type `i` is derived as shown in Table 56 below.

[0499] [Table 56]

[0500] gfv_matrix_element_int[i][j][k][m] indicates the integer part of the matrix element value at position (k, m) of the j-th matrix of type i.

[0501] gfv_matrix_element_dec[i][j][k][m] indicates the fractional part of the matrix element value at position (k, m) of the j-th matrix of type i.

[0502] `gfv_matrix_element_sign_flag[i][j][k][m]` indicates the sign of the matrix element at position (k, m) of the j-th matrix of type i. When `gfv_matrix_element_sign_flag[i][j][k][m]` does not exist, it is inferred to be equal to 0.

[0503] As shown in Table 57 below, the matrixElementVal[i][j][k][m] represents the value of the matrix element at position (k, m) of the i-th matrix of type j.

[0504] [Table 57]

[0505] GenerativeNN() is the process used to generate sample values ​​for the output image corresponding to the driving image. GenerativeNN() is only called when gfc_base_pic_flag equals 0.

[0506] The input to TranslatorNN() is: - sigKeyPoint and sigMatrix The output of TranslatorNN() is: - convKeyPoint and convMatrix The input to GenerativeNN() is: - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 0, and ChromaFormatIdc equals 0: inputBaseY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 0, and ChromaFormatIdc is not equal to 0: inputBaseY, inputBaseCb, inputBaseCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 1, and ChromaFormatIdc equals 0: inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 1, and ChromaFormatIdc is not equal to 0: inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix The output of GenerativeNN() is: - Luminance sample array genY - When ChromaFormatIdc is not equal to 0, the two chromaticity sample arrays genCb and genCr Use the following procedure, as shown in Table 58 below, to generate video images: [Table 58]

[0507] The procedure DeriveSigParam() is specified as follows to derive the input to TranslatorNN().

[0508] As shown in Table 59 below, export the keypoint coordinate array sigKeyPoint and the matrix sigMatrix.

[0509] [Table 59]

[0510] The input values ​​to GenerativeNN() are real numbers, and the functions InpY() and InpC() are specified as shown in Table 60 below.

[0511] [Table 60]

[0512] The output values ​​from GenerativeNN() are real numbers, and the functions OutY() and OutC() are specified as shown in Table 61 below.

[0513] [Table 61]

[0514] The procedure DeriveInputTensors() is specified as follows to derive the input to GenerativeNN().

[0515] When gfv_base_pic_flag equals 1, the BasePicture input tensors inputBaseY, inputBaseCb, and inputBaseCr are exported as shown in Table 62 below.

[0516] [Table 62]

[0517] When gfv_drive_pic_fusion_flag equals 1, the DrivePicture luminance sample arrays inputDriveY, inputDriveCb, and inputDriveCr are derived as shown in Table 63 below.

[0518] [Table 63]

[0519] When gfv_base_pic_flag equals 0, the keypoint coordinate array inputDriveKeyPoint and matrix inputDriveMatrix for the current image are exported as shown in Table 64 below.

[0520] [Table 64]

[0521] When gfv_base_pic_flag equals 1, the keypoint coordinate array inputBaseKeyPoint and matrix inputBaseMatrix for the base image are exported as shown in Table 65 below.

[0522] [Table 65]

[0523] The procedure StoreOutputTensors() is specified as follows to export output.

[0524] When gfv_base_pic_flag equals 0, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are exported as shown in Table 66 below.

[0525] [Table 66]

[0526] When gfv_base_pic_flag equals 1, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as shown in Table 67 below.

[0527] [Table 67]

[0528] Fifth Implementation Method

[0529] Tables 68, 69, and 70 illustrate examples of defining whether personally identifiable information has been altered or whether the data is suitable for personal identification purposes in the generative facial video SEI message syntax.

[0530] [Table 68]

[0531] [Table 69]

[0532] [Table 70]

[0533] A `optimized_for_machine_analysis_flag` value of 0 indicates that the generated / decoded results may not be optimized (may not be suitable) for machine analysis purposes. "Unsuitable" means the machine analysis results may differ from the expected results or may be inaccurate. For example, object recognition results may be inaccurate when the shape or color of an object changes or some information is removed. When equal to 1, `optimized_for_machine_analysis_flag` indicates that the generated / decoded results are suitable for machine analysis purposes. This syntax can be applied not only to GFV but also to other generated / decoded results.

[0534] A value of 0 for `gfv_optimized_for_personal_identification_flag` indicates that the result generated / decoded by GFV can be unoptimized (unsuitable) for personal identification purposes. A value of 1 for `gfv_optimized_for_personal_identification_flag` indicates that the result generated / decoded by GFV can be suitable for personal identification purposes.

[0535] A value of 1 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., the base image). When equal to 0, `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) are the same as the attributes of the input data (e.g., the base image).

[0536] gfv_matrix_property is information about the attributes that indicate the information of each matrix.

[0537] Table 68 illustrates an example of defining gfv_matrix_property when gfv_matrix_transformed_flag equals 1. In this case, gfv_matrix_property can be defined as follows.

[0538] The properties defined by `gfv_matrix_property` can include transformations, removals, enhancements, etc., of information used to achieve matrix purposes. When equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed differently from the properties of the input data (e.g., the base image). When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information in the input data (e.g., the base image). The items listed above are examples, and additional information indicating the properties of the matrix can be defined separately. `gfv_matrix_property[i]` indicates the properties of the matrix represented by `gfv_matrix_type_idx[i]`.

[0539] Table 69 illustrates an example of defining gfv_matrix_transformed_flag independently.

[0540] Table 70 illustrates an example of defining gfv_matrix_property independently. In this case, the definition of gfv_matrix_property can be as follows.

[0541] `gfv_matrix_property` is information indicating the properties of each matrix. Properties defined by `gfv_matrix_property` can include transformations, removals, enhancements, etc. As an example, when equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) are the same as the properties of the input data (e.g., the base image). When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed differently from the properties of the input data (e.g., the base image). When equal to 2, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information in the input data (e.g., the base image). The items listed above are examples, and additional information indicating properties of the matrix can be defined separately. gfv_matrix_property[i] indicates the property of the matrix represented by gfv_matrix_type_idx[i].

[0542] A value of 1 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., the base image). When equal to 0, `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) are the same as the attributes of the input data (e.g., the base image).

[0543] The examples above illustrate the usage of `gfv_matrix_transformed_flag` and `gfv_matrix_property` as defined, and other usage forms are also possible. For example, two pieces of information can be defined together without separate conditions.

[0544] The GFV SEI message specifies the facial parameters and the facial parameter transformer neural network (denoted as TranslatorNN()) and the facial image generator neural network (denoted as GenerativeNN()). Here, TranslatorNN() can be used to transform the facial parameters in various formats signaled by the SEI message into parameters of a fixed format. GenerativeNN() can be used to generate an output image using the format of the facial parameters and a previously decoded output image.

[0545] Note 1 - Facial parameters can be determined from the source image before encoding. This source image can be called the driving image.

[0546] Note 2 - The previously decoded output image input to GenerativeNN() can be a base image (a decoded output image that provides a reference texture from which a face image can be generated) and optionally, an image that can be fused by GenerativeNN() to enhance background texture and facial details. When the current image is not a base image, GFV SEI messages can be used to generate a face image for fusion purposes based on the previously decoded base image, facial parameters conveyed by the GFV SEI message, and optionally the current decoded image.

[0547] To use this SEI message, you need to define the following variables.

[0548] - The width and height of the input image, expressed in units of brightness samples, are denoted as CroppedWidth and CroppedHeight, respectively.

[0549] - A luminance sample array baseCroppedYPic and chrominance sample arrays baseCroppedCbPic and baseCroppedCrPic for the decoded output image corresponding to the source base image (denoted as BasePicture).

[0550] - The luminance sample array driveCroppedYPic and the chrominance sample arrays driveCroppedCbPic and driveCroppedCrPic for the decoded output image corresponding to the source driving image (denoted as DrivePicture).

[0551] - BitDepth of the luminance sample array for the input image Y

[0552] - BitDepth of the chroma sample array for the input image (if it exists) C

[0553] - Chroma format indicator, denoted as ChromaFormatIdc

[0554] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.

[0555] `gfv_id` includes an identifier that can be used to identify facial feature information and specify the neural network that can be used as a GenerativeNN(). The value of `gfv_id` should be between 0 and 2. 32 - The range of 2 (inclusive of 0 and 2) 32 - 2) Within. gfv_id in the range of 256 to 511 and in 2 31 Up to 2 32 Values ​​in the range of -2 are reserved for future use by ITU-T|ISO / IEC. The decoder should ignore values ​​included in the range of 256 to 511 or in 2... 31 Up to 2 32 - GFV SEI messages with gfv_id in the range of 2.

[0556] Note - For example, when there is more than one face in the output image, different values ​​of gfv_id in different GFV SEI messages can be used to identify different faces.

[0557] A value of 1 for `gfv_base_pic_flag` indicates that the currently decoded output image corresponds to the base image. A value of 0 for `gfv_base_pic_flag` indicates that the currently decoded output image does not correspond to the base image.

[0558] The following constraints apply to the value of gfv_base_pic_flag.

[0559] - When the GFV SEI message is the first GFV SEI message in the current CLVS with a specific gfv_id value in decoding order, the value of gfv_base_pic_flag should be equal to 1.

[0560] - When a GFV SEI message with a specific gfv_id value has a gfv_base_pic_flag equal to 0, the SEI message is associated with the currently decoded picture and all subsequent decoded pictures of the current layer in output order, until the end of the current CLVS, or until, but not including, the decoded picture in the current CLVS that follows the currently decoded picture in output order and is associated with subsequent GFV SEI messages in the current CLVS in decoding order with a gfv_base_pic_flag equal to 0 and the specific gfv_id value, whichever comes first.

[0561] `gfv_nn_base_flag`, `gfv_nn_mode_idc`, `gfv_nn_reserved_zero_bit_a`, `gfv_nn_tag_uri`, `gfv_nn_uri`, and `gfv_nn_payload_byte[i]` specify neural networks that can be used as TranslatorNN(). `gfv_nn_base_flag`, `gfv_nn_mode_idc`, `gfv_nn_reserved_zero_bit_a`, `gfv_nn_tag_uri`, `gfv_nn_uri`, and `gfv_nn_payload_byte[i]` have the same syntax and semantics as `nnpfc_base_flag`, `nnpfc_mode_idc`, `nnpfc_reserved_zero_bit_a`, `nnpfc_tag_uri`, `nnpfc_uri`, and `nnpfc_payload_byte[i]`, respectively.

[0562] When present, a `gfv_drive_pic_fusion_flag` of 1 indicates that the currently decoded image can be input into `GenerativeNN()`. The currently decoded image corresponds to the driving image that can be used for fusion. A `gfv_drive_pic_fusion_flag` of 0 indicates that the currently decoded image should not be input into `GenerativeNN()`.

[0563] Note 3 - For example, a value of gfv_drive_pic_fusion_flag equal to 1 can be used to indicate whether the currently decoded image can enhance facial details or handle background changes.

[0564] Note 4 - Fusion uses three inputs to output the image: the base image, features from keypoints and / or matrices conveyed in the GFV SEI message, and the currently decoded image.

[0565] Note 5 - When the currently decoded image corresponds to the driving image, it should be marked as not intended for output.

[0566] A `optimized_for_machine_analysis_flag` value of 0 indicates that the generated / decoded result may not be optimized (may not be suitable) for machine analysis purposes. "Unsuitable" means the machine analysis result may differ from the expected result or may be inaccurate. A `optimized_for_machine_analysis_flag` value of 1 indicates that the generated / decoded result is suitable for machine analysis purposes. This syntax can be applied not only to GFV but also to other generated / decoded results.

[0567] A value of 0 for `gfv_optimized_for_personal_identification_flag` indicates that the results generated / decoded by GFV can be unoptimized (or unsuitable) for personal identification purposes. A value of 0 for `gfv_not_optimized_for_personal_identification_flag` indicates that the results generated / decoded by GFV can be suitable for personal identification purposes.

[0568] `gfv_matrix_property` indicates information about the properties of each matrix. Properties defined by `gfv_matrix_property` include transformations, removals, enhancements, etc. For example, when equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) are the same as the properties of the input data (e.g., a base image). When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. When equal to 2, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information. Additional information defining the properties of the matrix can be specified.

[0569] When gfv_matrix_transformed_flag is set to 1, the semantics of gfv_matrix_property defined under this condition can be specified as follows.

[0570] `gfv_matrix_property` indicates information about the properties of each matrix. Properties defined by `gfv_matrix_property` include transformations, removals, enhancements, etc. For example, when equal to 0, `gfv_matrix_property` indicates that the properties of the information represented by the matrix have been transformed differently from the properties of the input data. When equal to 1, `gfv_matrix_property` indicates that the properties of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed or partially removed for the purpose of protecting personally identifiable information. Additional information defining the properties of the matrix can be specified.

[0571] A value of 1 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix (e.g., mouth, eyes, hair, etc.) have been transformed (e.g., enhanced, removed, replaced, etc.) differently from the attributes of the input data (e.g., the base image). A value of 0 for `gfv_matrix_transformed_flag` indicates that the attributes of the information represented by the matrix are the same as the attributes of the input data.

[0572] A value of 1 for `gfv_coordinate_present_flag` indicates the existence of keypoint coordinate information. A value of 0 for `gfv_coordinate_present_flag` indicates the absence of keypoint coordinate information.

[0573] The requirement for bitstream consistency is that when gfv_matrix_type_idx[i] is equal to 0 or 1 for all i in the range from 0 to gfv_num_matrix_types_minus1 (inclusive), the value of gfv_coordinate_present_flag should be equal to 1.

[0574] Incrementing gfv_coordinate_precision_factor_minus1 by 1 indicates the bit length of gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i].

[0575] Incrementing `gfv_num_kps_minus1` by 1 indicates the number of keypoints. The value of `gfv_num_kp_minus1` should be between 0 and 2. 10 -1 range (including 0 and 2) 10 -1) inside.

[0576] A value of 1 for `gfv_kp_pred_flag` indicates the existence of the syntax elements `gfv_coordinate_dx_abs[i]`, `gfv_coordinate_dy_abs[i]`, and `gfv_coordinate_dz_abs[i]`, and may also contain the syntax elements `gfv_coordinate_dx_sign_flag[i]`, `gfv_coordinate_dy_sign_flag[i]`, and `gfv_coordinate_dz_sign_flag[i]`. A value of 0 for `gfv_kp_pred_flag` indicates the existence of `gfv_coordinate_x_abs[i]`, `gfv_coordinate_y_abs[i]`, and `gfv_coordinate_z_abs[i]`, and may also contain the syntax elements `gfv_coordinate_x_sign_flag[i]`, `gfv_coordinate_y_sign_flag[i]`, and `gfv_coordinate_z_sign_flag[i]`.

[0577] A value of 1 for `gfv_coordinate_z_present_flag` indicates the presence of z-axis coordinate information for keypoints. A value of 0 for `gfv_coordinate_z_present_flag` indicates the absence of z-axis coordinate information for keypoints.

[0578] The value of gfv_coordinate_z_max_value_minus1 plus 1 indicates the maximum absolute value of the z-axis coordinate of the key point.

[0579] gfv_coordinate_x_abs[i] indicates the normalized absolute value of the x-axis coordinate of the i-th key point.

[0580] `gfv_coordinate_x_sign_flag[i]` specifies the sign of the x-axis coordinate of the i-th keypoint. When `gfv_coordinate_x_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0581] gfv_coordinate_y_abs[i] indicates the normalized absolute value of the y-axis coordinate of the i-th key point.

[0582] `gfv_coordinate_y_sign_flag[i]` specifies the sign of the y-axis coordinate of the i-th keypoint. When `gfv_coordinate_y_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0583] gfv_coordinate_z_abs[i] indicates the normalized absolute value of the z-axis coordinate of the i-th key point.

[0584] `gfv_coordinate_z_sign_flag[i]` specifies the sign of the z-axis coordinate of the i-th keypoint. When `gfv_coordinate_z_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0585] gfv_coordinate_dx_abs[i] indicates the absolute difference of the normalized values ​​of the x-axis coordinates of the i-th key point.

[0586] `gfv_coordinate_dx_sign_flag[i]` specifies the sign of the difference in the x-axis coordinates of the i-th keypoint. When `gfv_coordinate_dx_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0587] gfv_coordinate_dy_abs[i] specifies the absolute difference of the normalized values ​​of the y-axis coordinates of the i-th keypoint.

[0588] `gfv_coordinate_dy_sign_flag[i]` specifies the sign of the difference in the y-axis coordinates of the i-th keypoint. When `gfv_coordinate_dy_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0589] gfv_coordinate_dz_abs[i] specifies the absolute difference of the normalized values ​​of the z-axis coordinates of the i-th keypoint.

[0590] `gfv_coordinate_dz_sign_flag[i]` specifies the sign of the difference in the z-axis coordinates of the i-th keypoint. When `gfv_coordinate_dz_sign_flag[i]` does not exist, it is inferred to be equal to 0.

[0591] As shown in Table 71 below, the variables coordinateDeltaX[i], coordinateDeltaY[i], and coordinateDeltaZ[i], representing the incremental x-axis coordinate, incremental y-axis coordinate, and incremental z-axis coordinate of the i-th key point, respectively, are derived.

[0592] [Table 71]

[0593] When gfv_kp_pred_flag equals 0, variables coordinateX[i], coordinateY[i], and coordinateZ[i] representing the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the i-th keypoint are derived as shown in Table 72 below. When gfv_kp_pred_flag equals 1, the variables are derived as shown in Table 73 below.

[0594] [Table 72]

[0595] [Table 73]

[0596] Here, as shown in Table 74 below, we derive BaseKpCoordinateX[i], BaseKpCoordinateY[i], and BaseKpCoordinateZ[i], which represent the x-axis coordinates, y-axis coordinates, and z-axis coordinates of the i-th keypoint of the base image, respectively.

[0597] [Table 74]

[0598] A value of 1 for `gfv_matrix_present_flag` indicates the presence of a matrix parameter. A value of 0 for `gfv_matrix_present_flag` indicates the absence of a matrix parameter.

[0599] Incrementing gfv_matrix_element_precision_factor_minus1 by 1 indicates the length of gfv_matrix_element_dec[i][j][k][m] in bits.

[0600] Incrementing `gfv_num_matrix_types_minus1` by 1 indicates the number of matrix types signaled in the SEI message. The value of `gfv_matrix_type_num_minus1` should be between 0 and 2. 6 -1 range (including 0 and 2) 6 -1) inside.

[0601] gfv_matrix_type_idx[i] indicates the index of the i-th matrix type as specified in Table 75.

[0602] [Table 75]

[0603] Note 2. An undefined matrix type is used to indicate a matrix type that is not an affine transformation matrix, covariance matrix, rotation matrix, translation matrix, or compact eigenma matrix. It can be used by the user to determine the matrix type.

[0604] A value of 1 for `gfv_num_matrices_equal_to_num_kps_flag[i]` indicates that the number of matrices of the i-th matrix type is equal to `gfv_num_kps_minus1 + 1`. A value of 0 for `gfv_num_matrices_equal_to_num_kps_flag[i]` indicates that the number of matrices of the i-th matrix type is not equal to `gfv_num_kps_minus1 + 1`.

[0605] gfv_num_matrices_info[i] provides information about the number of matrices of type i used to derive the matrix.

[0606] gfv_matrix_width_minus1[i] increments by 1 to indicate the width of the matrix of type i.

[0607] The increment of gfv_matrix_height_minus1[i] indicates the height of the matrix of type i.

[0608] A value of 1 for `gfv_matrix_for_3D_space_flag[i]` indicates that the matrix of type `i` is defined in three-dimensional space. A value of 0 for `gfv_matrix_for_3D_space_flag[i]` indicates that the matrix of type `i` is defined in two-dimensional space.

[0609] When gfv_matrix_width_minus1[i] does not exist, it is inferred as follows.

[0610] - When gfv_matrix_type_idx[i] is equal to 0, 1 or 4, and one of coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] exists and is equal to 1, gfv_matrix_width_minus1[i] is inferred to be equal to 2.

[0611] Otherwise, gfv_matrix_width_minus1[i] is inferred to be equal to 1 when matrix_type_idx[i] is equal to 0, 1 or 4, and one of coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] exists and is equal to 0.

[0612] Otherwise (when matrix_type_idx[i] equals 5 or 6), gfv_matrix_width_minus1[i] is inferred to be equal to 0.

[0613] When gfv_matrix_height_minus1[i] does not exist, it is inferred as follows.

[0614] - When matrix_type_idx is equal to 0, 1, 4, 5 or 6, and one of gfv_coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] exists and is equal to 1, gfv_matrix_height_minus1[i] is inferred to be equal to 2.

[0615] Otherwise (when gfv_matrix_type_idx is equal to 0, 1, 4, 5 or 6, and one of gfv_coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[i] is equal to 0), gfv_matrix_height_minus1[i] is inferred to be equal to 1.

[0616] As shown in Table 76 below, derive the variables matrixWidth[i] and matrixHeight[i], which represent the width and height of the matrix of type i, respectively.

[0617] [Table 76]

[0618] The increment of gfv_num_matrices_minus1[i] indicates the number of matrices of type i. Table 77 below derives the variable numMatrices[i] which represents the number of matrices of type i.

[0619] [Table 77]

[0620] gfv_matrix_element_int[i][j][k][m] indicates the integer part of the matrix element value at position (k, m) of the j-th matrix of type i.

[0621] gfv_matrix_element_dec[i][j][k][m] indicates the fractional part of the matrix element value at position (k, m) of the j-th matrix of type i.

[0622] `gfv_matrix_element_sign_flag[i][j][k][m]` indicates the sign of the matrix element at position (k, m) of the j-th matrix of type i. When `gfv_matrix_element_sign_flag[i][j][k][m]` does not exist, it is inferred to be equal to 0.

[0623] As shown in Table 72 below, the matrixElementVal[i][j][k][m] represents the value of the matrix element at position (k, m) of the i-th matrix of type j.

[0624] [Table 78]

[0625] GenerativeNN() is the process used to generate sample values ​​for the output image corresponding to the driving image. GenerativeNN() is only called when gfc_base_pic_flag equals 0.

[0626] The input to TranslatorNN() is: - sigKeyPoint and sigMatrix The output of TranslatorNN() is: - convKeyPoint and convMatrix The input to GenerativeNN() is: - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 0, and ChromaFormatIdc equals 0: inputBaseY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 0, and ChromaFormatIdc is not equal to 0: inputBaseY, inputBaseCb, inputBaseCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 1, and ChromaFormatIdc equals 0: inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix - When gfv_base_pic_flag equals 0, gfv_drive_pic_fusion_flag equals 1, and ChromaFormatIdc is not equal to 0: inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix The output of GenerativeNN() is: - Luminance sample array genY - When ChromaFormatIdc is not equal to 0, the two chromaticity sample arrays genCb and genCr Use the following procedure, as shown in Table 79 below, to generate video images.

[0627] [Table 79]

[0628] The procedure DeriveSigParam() is specified as follows to derive the input to TranslatorNN().

[0629] As shown in Table 80 below, export the keypoint coordinate array sigKeyPoint and the matrix sigMatrix.

[0630] [Table 80]

[0631] The input values ​​to GenerativeNN() are real numbers, and the functions InpY() and InpC() are specified as shown in Table 81 below.

[0632] [Table 81]

[0633] The output values ​​from GenerativeNN() are real numbers, and the functions OutY() and OutC() are specified as shown in Table 82 below.

[0634] [Table 82]

[0635] The procedure DeriveInputTensors() is specified as follows to derive the input to GenerativeNN().

[0636] When gfv_base_pic_flag equals 1, the BasePicture input tensors inputBaseY, inputBaseCb, and inputBaseCr are exported as shown in Table 83 below.

[0637] [Table 83]

[0638] When gfv_drive_pic_fusion_flag equals 1, the DrivePicture luminance sample arrays inputDriveY, inputDriveCb, and inputDriveCr are derived as shown in Table 84 below.

[0639] [Table 84]

[0640] When gfv_base_pic_flag equals 0, the keypoint coordinate array inputDriveKeyPoint and matrix inputDriveMatrix for the current image are exported as shown in Table 85 below.

[0641] [Table 85]

[0642] When gfv_base_pic_flag equals 1, the keypoint coordinate array inputBaseKeyPoint and matrix inputBaseMatrix for the base image are exported as shown in Table 86 below.

[0643] [Table 86]

[0644] The procedure StoreOutputTensors() is specified as follows to export output.

[0645] When gfv_base_pic_flag equals 0, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are exported as shown in Table 87 below.

[0646] [Table 87]

[0647] When gfv_base_pic_flag equals 1, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are exported as shown in Table 88 below.

[0648] [Table 88]

[0649] Sixth Implementation Method

[0650] Conventional video codecs aim to minimize pixel-level distortion for optimization to human viewing. Minimizing pixel-level distortion may not be optimal for machine analysis, as it relies on features rather than pixel-level details. Features represent information about various aspects of the input data (such as edges, texture, shape, or high-level semantic information), and features can also be the output of a neural network representing specific features of the input image. Research on feature-based optimizations for machine analysis has been reported below.

[0651] 1) Feature-based RDO. This method optimizes rate distortion by calculating distortion using extracted features. To extract features, a pre-trained model can be used, or a machine-aware perceptual metric model can be defined through training.

[0652] 2) Feature-based preprocessing. This method minimizes the loss of important features and reduces unnecessary information used for machine analysis. This makes compression more appropriate and allows for lower bit rates without adversely affecting machine analysis.

[0653] Encoder optimization information (SEI) messages indicate whether a video has been optimized for human viewing or machine analysis, and what types of optimization have been applied during preprocessing or encoding. The encoder can apply feature-based optimizations to improve the quality of regions with important features compared to regions in the decoded image that have fewer or no important features. The presence of important features can vary depending on the feature extraction method. Feature extraction methods can be defined based on the type and objective of the task affecting human and machine perception. Therefore, providing bitstream information about feature-based optimizations in the encoded video can be useful.

[0654] `eoi_cancel_flag` equal to 1 indicates that the persistence of encoder optimization information SEI messages included in any previous PU in the output order is canceled. `eoi_cancel_flag` equal to 0 indicates that information about optimizations provided in preprocessing or encoding follows.

[0655] `eoi_persistence_flag` specifies the persistence of the optimization information applied in this SEI message. `eoi_persistence_flag` equal to 0 indicates that the optimization information applies only to the current image. `eoi_persistence_flag` equal to 1 indicates that the optimization information applies to the current image and all subsequent images of the current layer in output order, until one or more of the following conditions are true.

[0656] - A new CLVS begins in the current layer.

[0657] - End of bitstream.

[0658] - Outputs the image in the current layer of the AU associated with the Content Protection Information (SEI) message, following the current image in the output order.

[0659] `eoi_for_human_viewing_flag` equal to 1 specifies that the purpose of the applied optimizations includes human viewing. `eoi_for_human_viewing_flag` equal to 0 specifies that the purpose of the applied optimizations does not include human viewing.

[0660] An `eoi_for_machine_analysis_flag` value of 1 indicates that the purpose of the applied optimization includes machine analysis. An `eoi_for_machine_analysis_flag` value of 0 indicates that the purpose of the applied optimization may or may not include machine analysis.

[0661] eoi_type indicates the type of optimization method specified in Table 89 below.

[0662] [Table 89]

[0663] Here, when (eoi_type & bitMask) is not equal to 0, this indicates that an optimization type with the bitmask value in Table 89 has been applied. When eoi_type is greater than 0 and (eoi_type & bitMask) is equal to 0, an optimization type with the bitmask value has not been applied. When eoi_type is equal to 0, the optimization determined by the application has been used.

[0664] The variables EoiObjectBasedFlag, EoiTemporalResamplingFlag, EoiSpatialResamplingFlag, EoiTemporalQualityFlag, EoiSpatialQualityFlag, EoiPixelBasedFlag, and EoiFeatureBasedFlag can be derived as shown in Table 90 below. These variables specify whether eoi_type indicates the type of optimization, including object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, pixel-based optimization, and feature-based optimization.

[0665] [Table 90]

[0666] Note - For example, when certain highest time sublayers have been encoded with this coarse quantization so that human viewers perceive quality fluctuations as uncomfortable but without compromising machine task performance, eoi_for_human_viewing_flag and eoi_for_machine_analysis_flag can be set to 0 and 1 respectively, and eoi_type can be set to a value that makes EoiTemporalQualityFlag equal to 1.

[0667] When eoi_persistence_flag equals 0, the bitstream consistency requirement is that EoiTemporalResamplingFlag should be equal to 0 and EoiTemporalQualityFlag should be equal to 0.

[0668] When present, eoi_object_based_idc indicates the type of object-based optimization as specified in Table 91 below.

[0669] [Table 91]

[0670] Here, (eoi_object_based_idc & bitMask) not equal to 0 indicates that the object-based optimization type associated with the bitmask value in Table 91 has been applied. When eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) equals 0, the object-based optimization type associated with the bitmask value has not been applied. When eoi_object_based_idc equals 0, the object-based optimization of the defined type has been applied. In bitstreams conforming to this specification, the value of eoi_object_based_idc should be in the range of 0 to 7 (inclusive). Values ​​of 8 to 65535 (inclusive) for eoi_object_based_idc are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this specification. When the value of eoi_object_based_idc is in the range of 8 to 65535 (inclusive), decoders following this specification will ignore eoi_object_based_idc.

[0671] When `eoi_temporal_resampling_type_flag` is equal to 0, the time-based resampling optimization is a subsampling operation. When `eoi_temporal_resampling_type_flag` is equal to 1, the time-based resampling optimization is an upsampling operation.

[0672] A value greater than 0 in `eoi_num_int_pics` indicates that the number of pictures the encoding system excludes between each pair of encoded pictures in output order (when `eoi_temporal_resampling_type_flag` equals 0) or adds between each pair of source pictures for encoding (when `eoi_temporal_resampling_type_flag` equals 1) is constant throughout the persistence of this SEI message. When `eoi_temporal_resampling_type_flag` equals 0 and `eoi_num_int_pics` is greater than 0, `eoi_num_int_pics` specifies the number of pictures the encoding system excludes between each pair of encoded pictures in output order. When `eoi_temporal_resampling_type_flag` equals 1 and `eoi_num_int_pics` is greater than 0, `eoi_num_int_pics` specifies the number of pictures the encoding system adds between each pair of source pictures for encoding.

[0673] An eoi_num_int_pics value of 0 indicates that the number of pictures excluded by the encoding system in the output order within this SEI message (when eoi_temporal_resampling_type_flag equals 0) or added between each pair of source pictures for encoding (when eoi_temporal_resampling_type_flag equals 1) is unknown or variable.

[0674] The value of eoi_num_int_pics should be in the range of 0 to 63 (inclusive).

[0675] When present, eoi_object_based_idc indicates the type of object-based optimization as specified in Table 91. Here, (eoi_object_based_idc & bitMask) not equal to 0 indicates that the object-based optimization type associated with the bitmask value in Table 91 has been applied.

[0676] eoi_feature_optimization_type_idc can be defined as identifying a single optimization method as shown in Table 92, or it can be defined as identifying multiple optimization methods through a bitmask as shown in Table 93.

[0677] [Table 92]

[0678] [Table 93]

[0679] When present, eoi_feature_optimization_type_idc indicates the type of feature-based optimization as specified in Table 92.

[0680] When present, eoi_feature_optimization_type_idc indicates the type of feature-based optimization as specified in Table 93. Here, (eoi_feature_optimization_type_idc & bitMask) not equal to 0 indicates that feature-based optimization associated with the bitmask value in Table 93 has been applied.

[0681] `eoi_partial_feature_use_flag` equal to 0 indicates that all features have been used for optimization. `eoi_partial_feature_use_flag` equal to 1 indicates that only some features are used for optimization.

[0682] Table 94 illustrates an example of the syntax for the Encoder Optimization Information (SEI) message.

[0683] [Table 94]

[0684] Seventh Implementation Method

[0685] Tables 95 and 96 illustrate examples of methods for representing optimization objectives, attributes, and application scope. This implementation includes a method of representing optimization objectives and states using a single identifier instead of flags defined for each optimization objective (e.g., optimization_for_machine_analysis_flag and optimization_for_human_viewing_flag in the second implementation).

[0686] [Table 95]

[0687] [Table 96]

[0688] When equal to 1, `optimization_cancel_flag` indicates that the persistence of previously applied optimizations is canceled. When equal to 0, `optimization_cancel_flag` indicates that optimizations and their scope can be applied according to the `optimization_persistence_flag`, `optimization_for_machine_analysis_flag`, and `optimization_type`.

[0689] `optimization_persistence_flag` indicates the persistence of the optimization indicated by `optimization_type`. When equal to 0, `optimization_persistence_flag` indicates that the optimization identified by `optimization_type` can be applied only to the current image. When equal to 1, `optimization_persistence_flag` indicates that the optimization identified by `optimization_type` can be applied to the current image and all subsequent images.

[0690] `optimization_purpose_idc` is information used to identify the purpose and state of optimization. The purpose of optimization can include human observation and machine analysis. The optimization state can be further subdivided and defined as: 1) It is unknown whether it is suitable for the corresponding purpose, 2) It is not suitable, 3) It is suitable but not optimized for the corresponding purpose, or 4) It is suitable and optimized for the corresponding purpose.

[0691] Tables 97, 98, 99, 100, and 101 illustrate examples of the `optimization_purpose_idc` definition. The defined optimization purpose and state are examples, and new optimization attributes can be defined using the corresponding structures.

[0692] [Table 97]

[0693] Table 97 illustrates the following example: the four optimization states defined above are each applied to two optimization purposes: human observation and machine analysis. In this case, 16 information items are identified, and optimization_purpose_idc is defined as 4 bits, as shown in Table 95.

[0694] [Table 98]

[0695] Table 98 provides an example of defining identification information under the condition that at least one of the defined optimization purposes must have a clearly defined optimization state. This is because, since the corresponding SEI represents optimization information, the presence of the corresponding SEI can imply that the optimization is applied to a specific purpose. In this case, seven information items are identified, and optimization_purpose_idc is defined as 3 bits, as shown in Table 96.

[0696] Table 98 is an example, and the identification information can vary depending on the preconditions and condition levels.

[0697] [Table 99]

[0698] Table 99 adds cases where the suitability of optimization is unknown (index 111) to Table 98. This can be defined when the impact and suitability of the applied optimization method for human viewing and machine analysis cannot be clearly determined. This can be defined when the viewing conditions on the receiving side or the type of machine analysis to be performed are not explicitly known. Alternatively, when the optimization method is defined by the application, the encoder or transmitting side may not be able to determine its impact, and therefore a corresponding identifier can be defined.

[0699] [Table 100]

[0700] Table 100 excludes cases from Table 99 where optimizations were made for human viewing and machine analysis purposes (index 000 in Table 99).

[0701] [Table 101]

[0702] Table 101 excludes cases where optimization is performed for human viewing and machine analysis purposes, and cases where it is unknown whether optimization is suitable for all purposes.

[0703] Tables 99, 100, and 101 are based on Table 98, where specific identifiers are excluded or included based on identifier definition conditions. The same applies to Table 97, and specific identifiers can be excluded or included individually or multiple identifiers can be included simultaneously. Furthermore, it is evident that the optimization purpose and status assigned to each identifier index can change. That is, the order of the indexes defined in Tables 97, 98, 99, 100, and 101 is an example and is subject to change. These can be defined based on the importance of the optimization purpose, utilization, etc.

[0704] The `optimization_type` attribute indicates the optimization method.

[0705] The optimization attributes defined in Tables 102 and 103 can be identified.

[0706] [Table 102]

[0707] [Table 103]

[0708] The optimization attributes defined in Tables 102 and 103 are examples, and new optimization attributes can be defined using the corresponding structures. Needless to say, optimization attributes and methods can be identified by including only any subset (not all) of the optimization attributes defined in Tables 102 and 103, or can be configured to include other optimization attributes and methods not defined in Tables 102 and 103. Tables 102 and 103 perform the same function but represent examples in different ways.

[0709] Table 103 represents another example of multiple optimization methods using a bitmask for optimization_type, and Tables 104 and 105 illustrate examples of them.

[0710] [Table 104]

[0711] [Table 105]

[0712] The `optimizationForPIIProtectionFlag` indicates whether the applied optimization method includes personal information protection optimization. A value of 1 indicates that the applied optimization method includes personal information protection optimization, and a value of 0 indicates that personal information protection optimization is not included.

[0713] Figure 21 This is a diagram illustrating a method for decoding image information according to an embodiment of the present disclosure.

[0714] Figure 21 The operations illustrated do not correspond to the necessary configuration of a decoding method according to one implementation and may be omitted. Figure 21 At least some of the operations shown in the examples, or you can add them. Figure 21 Other operations not illustrated in the example.

[0715] Figure 21 The terms or names described herein (e.g., names of syntactic elements or variable names) are merely examples, and the technical features of this disclosure are not limited to those described herein. Figure 21 The terms described in the text. For example... Figure 21 The image information described herein may include various information according to the embodiments described herein, and may include information described in at least one of Tables 1 to 105.

[0716] Figure 21 The operations illustrated herein can be performed by a decoding device including a memory and a processor electrically connected to the memory, and can be performed, for example, by a processor.

[0717] The decoding device can obtain encoder optimization information (S2110).

[0718] For example, the processor of the decoding device can obtain image information and obtain encoder optimization information from the image information.

[0719] Encoder optimization information can take various forms or have various names. Furthermore, encoder optimization information can have various forms of names or can be information with various names.

[0720] For example, encoder optimization information can be a syntactic element or a syntactic structure that includes one or more syntactic elements. For example, encoder optimization information can be represented in various ways, such as encoder_optimization_info(), but is not limited to this. In the following description, encoder optimization information is referred to as encoder_optimization_info, but is not limited to this.

[0721] Encoder optimization information (encoder_optimization_info) may include syntax elements such as optimization persistence cancellation information (e.g., eoi_cancel_flag), optimization persistence information (e.g., eoi_persistence_flag), optimization information for human viewing (e.g., eoi_for_human_viewing_flag), optimization information for machine analysis (e.g., eoi_for_machine_analysis_flag), optimization type information (e.g., eoi_type), feature-based optimization type information (e.g., eoi_feature_optimization_type), and / or partial feature use information (e.g., eoi_partial_feature_use_flag).

[0722] Syntax elements or information included in encoder optimization information (e.g., encoder_optimization_info) can be in various forms, such as 1-bit flags, 2-bit or more indicator (idc) or strings of characters / numbers.

[0723] The decoding device can determine the type of encoder optimization (S2120).

[0724] For example, the processor of the decoding device can determine the type of encoder optimization based on information related to encoder optimization, and in particular, optimization type information. The type of encoder optimization can include object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization.

[0725] Encoder optimization information may include optimization type information indicating the type of encoder optimization. Optimization type information may include information on whether object-based optimization is applied, whether temporal resampling is applied, whether spatial resampling is applied, whether temporal quality optimization is applied, whether spatial quality optimization is applied, and / or whether feature-based optimization is applied.

[0726] Optimization type information can be represented in various ways, such as `optimization_type` or `eoi_type`, but is not limited to these. In the following text, optimization type information will be represented as `eoi_type`, but is not limited to this. Furthermore, optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of characters / numbers.

[0727] The processor can identify the type of encoder optimization based on optimization type information. For example, the processor can identify whether object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and / or feature-based optimization is applied based on the logical product (AND) of optimization type information and a predefined bitmask.

[0728] Here, the type of encoder optimization can include feature-based optimization. Features represent information about various aspects of the input data (such as edges, texture, shape, or high-level semantic information), and features can also be the output of a neural network that represents specific features of the input image.

[0729] The processor can identify whether encoder optimization includes feature-based optimization based on optimization type information. As shown in Table 89 described above, the processor can identify whether encoder optimization includes feature-based optimization based on the logical product (AND) of optimization type information and a predefined bitmask of 0x20. Additionally, as shown in Table 90 described above, the processor can obtain the value of the variable `EoiFeatureBasedFlag` based on the logical product (AND) of optimization type information and a predefined bitmask of 0x20, and can identify whether encoder optimization includes feature-based optimization based on the value of the variable `EoiFeatureBasedFlag`.

[0730] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization, and partial feature usage information indicating whether some features or all features are used for optimization.

[0731] Feature-based optimization can include at least one of feature-based rate distortion optimization (RDO) for machine analysis, feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control.

[0732] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization. Feature-based optimization type information may include information about whether feature-based rate distortion optimization (RDO) is applied, information about whether feature-based preprocessing is applied, and / or information about whether feature-based quantization parameter (QP) adaptive / rate control is applied.

[0733] Feature-based optimization type information can be represented in various ways, such as eoi_feature_optimization_type, but is not limited to this. Furthermore, feature-based optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of characters / numbers.

[0734] The processor can identify the type of feature-based optimization based on feature-based optimization type information. For example, the processor can determine whether to apply feature-based rate distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control based on the value of the feature-based optimization type information. As another example, the processor can determine whether to apply feature-based rate distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control based on the logical product (AND) of the feature-based optimization type information and a predefined bitmask.

[0735] Feature-based optimization can be performed using some features or all features.

[0736] Encoder optimization information may include feature usage information indicating whether some features or all features were used for optimization. For example, encoder optimization information may include partial feature usage information or full feature usage information.

[0737] Partial feature usage information can indicate whether some features or all features are used for optimization in feature-based optimization.

[0738] Partial feature usage information can be represented in various ways, such as eoi_partial_feature_use_flag, but is not limited to this. In addition, partial feature usage information can be information in various forms such as a 1-bit flag, a 2-bit or more bit indicator (idc), or a string of characters / numbers.

[0739] A processor can identify whether some or all features are used for optimization based on partial feature usage information. For example, a processor can identify that all features are used for optimization based on a partial feature usage information value of 0. Conversely, a processor can identify that some features are used for optimization based on a partial feature usage information value of 1.

[0740] Unlike the example above, when the encoder optimization information includes all feature usage information, the processor can identify that all features were used for optimization based on a value of 1 for all feature usage information. Alternatively, the processor can identify that some features were used for optimization based on a value of 0 for all feature usage information.

[0741] The decoding device can process the image (S2130).

[0742] For example, the processor of a decoding device can process images based on the type of encoder optimization. For instance, the processor of a decoding device can perform optimizations for human viewing and / or optimizations for machine analysis. The processor can perform object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization. The processor can perform feature-based rate distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control. Additionally, the processor can use some or all features for feature-based optimization.

[0743] As described above, the decoding device can determine whether to apply feature-based optimization based on encoder optimization information, and can process the image according to the type of feature-based optimization and / or the use of optimization for some or all features. Thus, the VCM system can provide feature-based optimization for machine analysis.

[0744] Figure 22 This is a diagram illustrating a method for encoding image information according to an embodiment of the present disclosure.

[0745] Figure 22 The operations illustrated do not correspond to the necessary configuration of the encoding method according to one implementation and may be omitted. Figure 22 At least some of the operations shown in the examples, or you can add them. Figure 22 Other operations not illustrated in the example.

[0746] Figure 22 The terms or names described herein (e.g., names of syntactic elements or variable names) are merely examples, and the technical features of this disclosure are not limited to those described herein. Figure 22 The terms described in the text. For example... Figure 22 The image information described herein may include various information according to the embodiments described herein, and may include information described in at least one of Tables 1 to 105.

[0747] Figure 22 The operations illustrated herein can be performed by a decoding device including a memory and a processor electrically connected to the memory, and can be performed, for example, by a processor.

[0748] The encoding device can determine the type of encoder optimization (S2210).

[0749] For example, a processor included in an encoding device can determine the type of encoder optimization based on the characteristics and / or intended use of the image. For instance, the processor can determine whether to apply optimizations for human viewing and / or machine analysis based on the image's intended use. The processor can determine whether to apply object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization based on the image's characteristics. The processor can determine whether to apply feature-based rate distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control. Furthermore, the processor can determine whether to use some or all features for feature-based optimization.

[0750] The encoding device can process the image (S2220).

[0751] For example, the processor of the encoding device can perform optimized image processing based on the type of encoder optimization. Encoder optimization may include optimization for human viewing and / or optimization for machine analysis. Encoder optimization may include object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization. Feature-based optimization may include feature-based rate distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control. Additionally, some or all features may be used for feature-based optimization.

[0752] The processor of the encoding device can perform optimizations for human viewing and / or for machine analysis. The processor can perform object-based optimization, temporal resampling, spatial resampling, temporal quality optimization, spatial quality optimization, and / or feature-based optimization. The processor can perform feature-based rate distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control. Additionally, the processor can use some or all features for feature-based optimization.

[0753] The encoding device can generate encoder optimization information (S2230).

[0754] For example, the processor of an encoding device can generate encoder optimization information based on the type of encoder optimization.

[0755] Encoder optimization information may include optimization type information indicating the type of encoder optimization. Optimization type information may include information on whether object-based optimization is applied, whether temporal resampling is applied, whether spatial resampling is applied, whether temporal quality optimization is applied, whether spatial quality optimization is applied, and / or whether feature-based optimization is applied.

[0756] Optimization type information can be represented in various ways, such as `optimization_type` or `eoi_type`, but is not limited to these. In the following text, optimization type information will be represented as `eoi_type`, but is not limited to this. Furthermore, optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of characters / numbers.

[0757] The processor can generate optimization type information based on the type of encoder optimization. For example, the processor can generate optimization type information based on whether object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and / or feature-based optimization are applied.

[0758] Here, the type of encoder optimization can include feature-based optimization. Features represent information about various aspects of the input data (such as edges, texture, shape, or high-level semantic information), and features can also be the output of a neural network that represents specific features of the input image.

[0759] The processor can generate optimization type information based on whether the encoder optimization includes feature-based optimization.

[0760] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization, and partial feature usage information indicating whether some features or all features are used for optimization.

[0761] Feature-based optimization can include at least one of feature-based rate distortion optimization (RDO) for machine analysis, feature-based preprocessing, and / or feature-based quantization parameter (QP) adaptive / rate control.

[0762] Encoder optimization information may include feature-based optimization type information indicating the type of feature-based optimization. Feature-based optimization type information may include information about whether feature-based rate distortion optimization (RDO) is applied, information about whether feature-based preprocessing is applied, and / or information about whether feature-based quantization parameter (QP) adaptive / rate control is applied.

[0763] Feature-based optimization type information can be represented in various ways, such as eoi_feature_optimization_type, but is not limited to this. Furthermore, feature-based optimization type information can be in various forms, such as a 1-bit flag, a 2-bit or more indicator (idc), or a string of characters / numbers.

[0764] The processor can generate feature-based optimization type information based on the type of feature-based optimization. For example, the processor can generate feature-based optimization type information based on whether feature-based rate distortion optimization (RDO), feature-based preprocessing, and / or feature-based quantization parameters (QP) adaptive / rate control are applied.

[0765] Feature-based optimization can use some or all features for optimization.

[0766] Encoder optimization information may include feature usage information indicating whether some features or all features were used for optimization. For example, encoder optimization information may include partial feature usage information or full feature usage information.

[0767] Partial feature usage information can indicate whether some features or all features are used for optimization in feature-based optimization.

[0768] Partial feature usage information can be represented in various ways, such as eoi_partial_feature_use_flag, but is not limited to this. In addition, partial feature usage information can be information in various forms such as a 1-bit flag, a 2-bit or more bit indicator (idc), or a string of characters / numbers.

[0769] The processor can generate partial feature usage information based on whether some or all features are used for optimization. For example, the processor can set the value of partial feature usage information to 0 based on the determination that all features are used for optimization. Alternatively, the processor can set the value of partial feature usage information to 1 based on the determination that some features are used for optimization.

[0770] Unlike the example above, when the encoder optimization information includes all feature usage information, the processor can set the values ​​of all feature usage information to 1 based on the determination that all features were used for optimization. Alternatively, the processor can set the values ​​of all feature usage information to 0 based on the determination that some features were used for optimization.

[0771] The encoding device can encode image information (S2240).

[0772] For example, the processor of an encoding device can encode image information that includes encoder optimization information.

[0773] The encoded image information is converted into a bitstream, and the bitstream can be stored in a storage medium or sent through a transmission device.

[0774] As described above, the encoding device can optimize an image based on the type of feature-based optimization and / or the use of optimization for some or all features, and can generate encoder optimization information based on whether feature-based optimization is applied. Furthermore, the encoding device can encode the encoder optimization information. Thus, the VCM system can provide feature-based optimization for machine analysis.

[0775] Figure 23 This is a diagram illustrating an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0776] refer to Figure 23 The content streaming system using the embodiments of this disclosure can broadly include encoding servers, streaming servers, web servers, media storage, user equipment, and multimedia input devices.

[0777] An encoding server compresses content input from multimedia input devices such as smartphones, cameras, or camcorders into digital data, generates a bitstream, and sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, or camcorders directly generate bitstreams, the encoding server can be omitted.

[0778] The bitstream can be generated by the image encoding method and / or image encoding apparatus applied to the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.

[0779] A streaming server can send multimedia data to a user's device based on a user request via a web server, and the web server can act as a medium to notify the user of available services. When a user requests a service from the web server, the web server can send the request to the streaming server, and the streaming server can transmit the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which in this case can control the commands / responses between devices within the content streaming system.

[0780] A streaming server can receive content from media storage and / or encoding servers. For example, content can be received in real time when it is received from an encoding server. In this case, to provide a seamless streaming service, the streaming server can store a bitstream for a certain period of time.

[0781] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (i.e., smartwatches, smart glass, head-mounted displays (HMDs)), digital televisions, desktop computers, digital signage, etc.

[0782] In a content streaming system, each server can operate as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0783] Figure 24 This is a diagram illustrating another example of a content streaming system to which embodiments of the present disclosure can be applied.

[0784] refer to Figure 24 In implementations such as VCM, the task can be executed by a user terminal, or by an external device (e.g., a streaming server, analytics server, etc.) based on the device's performance, the user's request, the characteristics of the task to be executed, etc. In this way, in order to send the information necessary for executing the task to the external device, the user terminal can generate a bitstream directly or through an encoding server, which includes the information necessary for executing the task (e.g., information such as the task, neural network, and / or the information used).

[0785] An analysis server can execute tasks requested by the user after decoding encoded information sent from a user terminal (or an encoding server). The analysis server can then send the results obtained from executing the tasks back to the user terminal or another linked service server (e.g., a web server). For example, the analysis server can send the results obtained from executing a task to determine a fire to a fire-related server. The analysis server may include a separate control server, in which case the control server can play a role in controlling the commands / responses between the analysis server and each device associated with it. Additionally, the analysis server can request desired information from a web server based on information about the tasks the user device wants to perform and the tasks the user device can perform. When the analysis server requests the desired service from the web server, the web server can forward it to the analysis server, and the analysis server can then send its data to the user terminal. In this scenario, the control server of the content streaming system can play a role in controlling the commands / responses between each device within the streaming system.

[0786] In this disclosure, as an example, for clarity of description, the names of all the above-described syntax elements are arbitrarily specified, without limitation on the names of the corresponding syntax elements. Additionally, each syntax element may also be referred to as information. Furthermore, syntax elements can be obtained from a bitstream, but can also be derived from other syntax elements, and this may also be included in the embodiments of this disclosure.

[0787] In addition, bitstreams generated by image encoding methods can be stored in non-transitory computer-readable recording media.

[0788] Alternatively, as another example, a bitstream generated by an image encoding method can be sent to another device (e.g., an image decoding device). In this case, the method for sending the bitstream can include the process of sending the bitstream.

[0789] For clarity of description, the exemplary methods of this disclosure are represented as a series of operations, but this is not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order if necessary. To implement the methods according to this disclosure, additional steps may be included in addition to the steps shown; some steps may be excluded and remaining steps may be included; or some steps may be excluded and additional steps may be included.

[0790] In this disclosure, the image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) to check the conditions or circumstances for performing the corresponding operation (step). For example, when describing the execution of a predetermined operation under predetermined conditions, the image encoding device or image decoding device may perform an operation to check whether the predetermined conditions are met, and then perform the predetermined operation.

[0791] The various embodiments of this disclosure are not intended to enumerate all possible combinations, but rather to describe representative aspects of this disclosure, and the matters described in the various embodiments may be applied independently or in combination of both or more.

[0792] The embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional unit shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or the algorithm can be stored in a digital storage medium.

[0793] Furthermore, the decoders (decoding devices) and encoders (encoding devices) to which the embodiments of this disclosure are applied can be included in multimedia broadcasting transmitting / receiving devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT video (overhead video) devices, internet streaming service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle (including autonomous vehicles) terminals, robot terminals, aircraft terminals, ship terminals, etc.), and medical video devices, and can be used to process video signals or data signals. For example, OTT video (overhead video) devices can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (digital video recorders), etc.

[0794] Furthermore, the processing methods applying the embodiments of this disclosure can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data having data structures according to embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media implemented in the form of a carrier wave (e.g., transmission via the Internet). Additionally, bitstreams generated by encoding methods can be stored in a computer-readable recording medium or transmitted via wired or wireless communication networks.

[0795] Furthermore, the embodiments of this disclosure can be implemented as a computer program product using program code, and the program code can be executed on a computer using the embodiments of this disclosure. The program code can be stored on a computer-readable medium.

[0796] Industrial applicability

[0797] The embodiments of this disclosure can be used to encode / decode features / feature maps.

Claims

1. A method for decoding image information, the method comprising the following steps: Obtain the image information, including encoder optimization information; The type of encoder optimization is determined based on the encoder optimization information; as well as The image is processed based on the type optimized by the encoder. The encoder optimization information includes optimization type information indicating the type of encoder optimization, and The encoder optimization includes feature-based optimization for optimizing the image based on its features.

2. The method according to claim 1, wherein, Whether to apply the feature-based optimization is determined based on the logical product of the optimization type information and the predefined first-order mask.

3. The method according to claim 1, wherein, The encoder optimization information also includes feature-based optimization type information that indicates the type of feature-based optimization.

4. The method according to claim 3, wherein, Based on the value of the feature-based optimization type information, the feature-based optimization includes at least one of feature-based rate distortion optimization, feature-based preprocessing, or feature-based quantization parameter QP adaptive / rate control.

5. The method according to claim 3, wherein, Based on the logical product of the feature-based optimization type information and the predefined second bitmask, the feature-based optimization includes at least one of feature-based rate distortion optimization, feature-based preprocessing, or feature-based quality adaptation / rate control.

6. The method according to claim 1, wherein, The encoder optimization information also includes partial feature usage information indicating whether some features or all features are used for optimization.

7. The method according to claim 1, wherein, The encoder optimization information also includes optimization suitability information indicating whether the encoder optimization is suitable for human viewing or machine analysis.

8. The method according to claim 7, wherein, The optimized suitability information includes 3-bit or 4-bit indicators that indicate suitability for human viewing and suitability for machine analysis.

9. A method for encoding an image, the method comprising the following steps: Determine the type of encoder optimization; The image is processed based on the type optimized by the encoder; Generate encoder optimization information based on the type of encoder optimization; as well as The image information, including the encoder optimization information, is encoded. The encoder optimization information includes optimization type information indicating the type of encoder optimization, and The encoder optimization includes feature-based optimization for optimizing the image based on its features.

10. A method for storing a bitstream of image information in a computer-readable storage medium, the method comprising the following steps: The bitstream of the image information is obtained; as well as The data, including the bitstream, is stored in a storage medium. The image information includes encoder optimization information. The encoder optimization information is generated based on the type of encoder optimization determined from the image. The encoder optimization information includes optimization type information indicating the type of encoder optimization, and The encoder optimization includes feature-based optimization for optimizing the image based on its features.

11. A method for transmitting a bitstream of image information, the method comprising the following steps: The bitstream of the image information is obtained; as well as Send data including the bit stream. The image information includes encoder optimization information. The encoder optimization information is generated based on the type of encoder optimization determined from the image. The encoder optimization information includes optimization type information indicating the type of encoder optimization, and The encoder optimization includes feature-based optimization for optimizing the image based on its features.