Feature encoding / decoding method and apparatus, recording medium storing bit stream, and method for transmitting bit stream

By distinguishing the regions of features/feature maps, efficient encoding/decoding of the ROI and non-ROI of features is achieved, solving the problem that image compression in existing technologies is not suitable for artificial intelligence services, and realizing efficient feature encoding/decoding and image reconstruction.

CN120836155APending Publication Date: 2025-10-24LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480020282.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-22
Filing Date
2024-03-21
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing image compression technologies are not suitable for artificial intelligence services, cannot effectively handle large amounts of image data within limited resources, and lack efficient encoding/decoding methods for features/feature maps.

Method used

By distinguishing regions of features/feature maps, feature encoding/decoding methods and devices are used to efficiently encode/decode the regions of interest (ROI) and non-ROI of features, and generate a compressed bitstream for performing machine tasks.

Benefits of technology

It achieves efficient encoding/decoding of features/feature maps, improves encoding/decoding efficiency, generates compressed bitstreams for machine tasks, and supports image storage and reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120836155A_ABST
    Figure CN120836155A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a feature encoding / decoding method, a recording medium storing a bitstream, and a method for transmitting a bitstream. According to one embodiment of the present invention, a feature decoding method performed by a feature decoding device comprises the steps of: acquiring information on a channel within a feature from a bitstream; and reconstructing the channel based on the information on the channel, wherein the information on the channel may include information on whether encoding has been performed according to the importance level of each region of the channel.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a feature encoding / decoding method and apparatus, and more particularly, to a feature encoding / decoding method and apparatus for performing a machine task and a method of transmitting a bitstream generated by the feature encoding method / apparatus of the present application. BACKGROUND

[0002] With the development of machine learning technology, there is an increasing demand for artificial intelligence services based on image processing. In order to efficiently process a large amount of image data required for artificial intelligence services within limited resources, it is necessary to optimize image compression technology for performing a machine task. However, since existing image compression technology has been developed with the goal of high-resolution and high-quality image processing of human vision, there is a problem that they are not suitable for artificial intelligence services. Therefore, research and development of new machine-oriented image compression technology suitable for artificial intelligence services are actively conducted. SUMMARY

[0003] [PROBLEMS TO BE SOLVED BY THE INVENTION]

[0004] The present disclosure aims to provide a feature encoding / decoding method and apparatus having improved encoding / decoding efficiency.

[0005] The present application aims to generate a compressed bitstream for transmitting features / feature maps extracted from a deep neural network to perform a machine task.

[0006] The present application aims to provide a method and apparatus for performing partial encoding / decoding by distinguishing regions of features / feature maps.

[0007] The present application aims to provide a method and apparatus for efficiently encoding / decoding a region of interest (ROI) and a non-ROI of features.

[0008] In addition, the present disclosure provides a method of transmitting a bitstream generated by a feature encoding method or apparatus according to the present disclosure.

[0009] In addition, the present disclosure provides a recording medium storing a bitstream generated by a feature encoding method or apparatus according to the present disclosure.

[0010] In addition, the present disclosure provides a recording medium storing a bitstream received and decoded by an image decoding apparatus according to the present disclosure and used for reconstructing an image.

[0011] The technical problems to be solved by the present disclosure are not limited to the above-described technical problems, and other technical problems not described above can be clearly understood by those skilled in the art from the following description.

[0012] [TECHNICAL SOLUTION]

[0013] According to one aspect of the disclosure, a feature decoding method performed by a feature decoding apparatus includes obtaining information about a channel in a feature from a bitstream, and reconstructing the channel based on the information about the channel, wherein the information about the channel can include information about whether to encode according to a region-wise importance of the channel.

[0014] Meanwhile, based on the information about whether to encode indicating that only a ROI of the channel is encoded, a non-ROI can be reconstructed based on a representative value of channel data.

[0015] Meanwhile, the representative value of the channel data can be an average value of the channel data.

[0016] Meanwhile, the representative value of the channel data can be obtained from the bitstream.

[0017] Meanwhile, based on the information about whether to encode indicating that only a ROI of the channel is encoded, a non-ROI can be reconstructed further based on a difference between the representative value of channel data of the non-ROI and the representative value of the channel data.

[0018] Meanwhile, the representative value of the channel data of the non-ROI can be an average value of the channel data of the non-ROI.

[0019] Meanwhile, the difference can be obtained from the bitstream.

[0020] Meanwhile, the non-ROI can be reconstructed based on a value obtained by adding the difference to the representative value of the channel data of the non-ROI or the representative value of the channel data.

[0021] According to one aspect of the disclosure, a feature encoding method performed by a feature encoding apparatus includes determining information about a channel in a feature, and encoding the information about the channel into a bitstream, wherein the information about the channel can include information about whether to encode according to a region-wise importance of the channel.

[0022] Meanwhile, based on encoding only a ROI of the channel, a representative value of channel data can be encoded into the bitstream.

[0023] Meanwhile, based on encoding only a ROI of the channel, a difference between representative values of channel data can be encoded into the bitstream.

[0024] A recording medium according to another aspect of the disclosure can store a bitstream generated by the feature encoding method or the feature encoding apparatus of the disclosure.

[0025] A bitstream transmission method according to another aspect of the disclosure can transmit a bitstream generated by the feature encoding method or the feature encoding apparatus of the disclosure to a feature decoding apparatus.

[0026] The features briefly described above for the disclosure are merely exemplary aspects of the detailed description of the disclosure below and do not limit the scope of the disclosure.

[0027] [Advantageous Effects]

[0028] According to the disclosure, a feature encoding / decoding method and apparatus having improved encoding / decoding efficiency can be provided.

[0029] According to the disclosure, a compressed bitstream for features / feature maps extracted from a deep neural network can be generated to perform a machine task.

[0030] According to the disclosure, a feature / feature map can be partially encoded / decoded by distinguishing regions of the feature / feature map.

[0031] According to the disclosure, ROIs and non-ROIs of features can be efficiently encoded / decoded.

[0032] In addition, according to the disclosure, a VCM (Video Coding for Machines) bitstream for performing a machine task can be generated.

[0033] In addition, according to the disclosure, a method of transmitting a bitstream generated by a feature encoding method or apparatus according to the disclosure can be provided.

[0034] In addition, according to the disclosure, a recording medium storing a bitstream generated by a feature encoding method or apparatus according to the disclosure can be provided.

[0035] In addition, according to the disclosure, a recording medium storing a bitstream received and decoded by a feature decoding apparatus according to the disclosure and used for reconstruction of an image can be provided.

[0036] Effects obtainable according to the disclosure are not limited to the above-mentioned effects, and other effects not mentioned above will become apparent to those skilled in the art from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 FIG. 1 is a diagram schematically illustrating a VCM system to which embodiments of the disclosure can be applied.

[0038] Figure 2 FIG. 2 is a diagram schematically illustrating a VCM pipeline structure to which embodiments of the disclosure can be applied.

[0039] Figure 3 FIG. 3 is a diagram schematically illustrating an image / video encoder to which embodiments of the disclosure can be applied.

[0040] Figure 4 FIG. 4 is a diagram schematically illustrating an image / video decoder to which embodiments of the disclosure can be applied.

[0041] Figure 5 FIG. 1 is a flow chart schematically illustrating a feature / feature map encoding process to which embodiments of the present disclosure may be applied.

[0042] Figure 6 FIG. 1 is a flow chart schematically illustrating a feature / feature map decoding process to which embodiments of the present disclosure may be applied.

[0043] Figure 7 is a diagram illustrating an example of a feature extraction and reconstruction method to which an embodiment of the present disclosure can be applied.

[0044] Figure 8 is a diagram illustrating an example of an image segmentation method to which an embodiment of the present disclosure can be applied.

[0045] Figure 9 and Figure 10 is a diagram illustrating an example of an image encoding / decoding system including a VCM image encoder and decoder.

[0046] Figure 11 is a diagram showing an example of a VCM hierarchical structure.

[0047] Figure 12 is a diagram showing an example of a VCM bitstream consisting of encoded abstract features and NNAL information.

[0048] Figure 13 is a diagram that visualizes multiple channel data in features / feature maps according to an embodiment of the present disclosure.

[0049] Figure 14 is a diagram illustrating an example of regional division of feature / feature map channel data and regional data distribution according to an embodiment of the present disclosure.

[0050] Figure 15 FIG. 4 is an example showing a partial encoding / decoding method of ROI and non-ROI proposed in an embodiment of the present disclosure.

[0051] Figure 16 is an example illustrating a region-by-region partial encoding method according to another embodiment of the present disclosure.

[0052] Figure 17 is a diagram for explaining a feature decoding method that can be performed by a feature (image) decoding device according to an embodiment of the present disclosure.

[0053] Figure 18 is a diagram for explaining a feature encoding method that can be performed by a feature (image) encoding device according to an embodiment of the present disclosure.

[0054] Figure 19FIG. 1 is a diagram illustrating an example of a content streaming system to which embodiments of the disclosure can be applied.

[0055] Figure 20 FIG. 2 is a diagram illustrating another example of a content streaming system to which embodiments of the disclosure can be applied. DETAILED DESCRIPTION

[0056] Hereinafter, embodiments of the disclosure will be described in detail by referring to the accompanying drawings so that those of ordinary skill in the art can easily implement them. The disclosure, however, can be embodied in various different forms and is not limited to the embodiments described herein.

[0057] In describing embodiments of the disclosure, when well-known configurations or functions are deemed to make the gist of the disclosure unclear, their detailed explanations are omitted. Also, parts irrelevant to the description of the disclosure are omitted from the accompanying drawings, and like reference numerals have been assigned to like parts.

[0058] In the disclosure, when a certain component is described as being "connected," "coupled," or "linked" to another component, this can include not only a direct connection but also an indirect connection in which another component is present in between. Also, when a certain component is described as "including" or "having" another component, this means that it does not exclude other components unless explicitly described otherwise, but can further include additional components.

[0059] In the disclosure, unless explicitly described otherwise, the terms first, second, and the like are used only to distinguish one component from another component, and do not limit the order or importance of the components. Accordingly, within the scope of the disclosure, a first component in one embodiment can be referred to as a second component in another embodiment, and similarly, a second component in one embodiment can be referred to as a first component in another embodiment.

[0060] In the disclosure, distinguishable components are described to clearly explain their respective characteristics, and do not necessarily mean that the components are separated. In other words, a plurality of components can be integrated into a single hardware or software unit, or a single component can be distributed across a plurality of hardware or software units. Accordingly, such integrated or distributed embodiments are also included in the scope of the disclosure without explicitly describing them.

[0061] In the disclosure, the components described in various embodiments do not necessarily mean essential components, and some components can be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of the disclosure. Also, embodiments including additional components in addition to the components described in various embodiments are also included in the scope of the disclosure.

[0062] The present disclosure relates to encoding and decoding of images, and the terms used herein can have the ordinary meanings commonly used in the art to which the present disclosure belongs, unless the terms are newly defined in the present disclosure.

[0063] The present disclosure can be applied to methods disclosed in the Versatile Video Coding (VVC) standard and / or the Video Coding for Machines (VCM) standard. In addition, the present disclosure can be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation Audio Video Coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268, etc.).

[0064] The disclosure presents various embodiments related to video / image coding, and the embodiments can be executed in combination with each other unless otherwise specified. In the disclosure, “video” can refer to a set of images in time sequence. “Image” can be information generated by artificial intelligence (AI). Input information used in a process in which AI performs a series of tasks, information generated in an information processing process, and output information can be used as an image. In the disclosure, “picture” generally refers to a unit indicating a single image at a specific point in time, and a slice / tile is a coding unit that constitutes a part of a picture. One picture can consist of at least one slice / tile. In addition, a slice / tile can include at least one coding tree unit (CTU). A CTU can be partitioned into at least one unit. A tile is a rectangular region existing in a specific tile row and a specific tile column within a picture, and can consist of a plurality of CTUs. A tile column can be defined as a rectangular region of CTUs, and can have the same height as the height of a picture and a width specified by syntax elements signaled from a bitstream part such as a picture parameter set. A tile row can be defined as a rectangular region of CTUs, and can have the same width as the width of a picture and a height specified by syntax elements signaled from a bitstream part such as a picture parameter set. Tile scanning is a predetermined order sorting method for CTUs partitioned from a picture. Here, the CTUs can be sorted in tile raster scan order within a tile, and the tiles within a picture can be sequentially sorted in raster scan order of the tiles in the picture. A slice can include an integer number of complete tiles or an integer number of sequential complete CTU rows within a tile of a picture. A slice can be exclusively included in a single NAL unit. One picture can consist of at least one tile group. One tile group can include at least one tile. A brick can indicate a rectangular region of CTU rows within a tile in a picture. A tile can include at least one brick. A brick can indicate a rectangular region of CTU rows within a tile. One tile can be partitioned into a plurality of bricks, and each brick can include at least one CTU row belonging to the tile. A tile that is not partitioned into a plurality of bricks can also be regarded as a brick.

[0065] In the disclosure, “pixel” or “pel” can refer to the smallest unit constituting one picture (or image). In addition, the term “sample” can be used as a corresponding term for a pixel. A sample can generally indicate a pixel or a value of a pixel, and can indicate only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component.

[0066] In one embodiment, particularly when applied to VCM, when there are pictures composed of a set of components having different characteristics and meanings, a pixel / pixel value can indicate a pixel / pixel value of a component generated by independent information or combination, synthesis, and analysis of each component. For example, in an RGB input, it can indicate only a pixel / pixel value of R, can indicate only a pixel / pixel value of G, or can indicate only a pixel / pixel value of B. For example, it can indicate only a pixel / pixel value of a luminance component synthesized by using R, G, and B components. For example, it can indicate only a pixel / pixel value of information or an image extracted by analyzing R, G, and B components.

[0067] In the disclosure, a "unit" can refer to a basic unit of image processing. A unit can include at least one of a specific region of a picture or information related to the region. One unit can include one luminance block and two chroma (e.g., Cb, Cr) blocks. Depending on the context, the term "unit" can be used interchangeably with "sample array," "block," "region," etc. In general, an M x N block can include a set (or array) of samples (or sample array) or a set (or array) of transform coefficients consisting of M columns and N rows. In one embodiment, particularly when applied to VCM, a unit can indicate a basic unit including information for performing a specific task.

[0068] In the disclosure, the term "current block" can refer to one of a "current coding block," a "current coding unit," a "coding target block," a "decoding target block," or a "processing target block." When performing prediction, the "current block" can refer to a "current prediction block" or a "prediction target block." When performing transform (inverse transform) / quantization (dequantization), the "current block" can refer to a "current transform block" or a "transform target block." When performing filtering, the "current block" can refer to a "filtering target block."

[0069] In addition, in the disclosure, unless explicitly stated as a chroma block, the "current block" can refer to a "luminance block of the current block." The "chroma block of the current block" can be expressed by explicitly including an explicit description of the chroma block such as "chroma block" or "current chroma block."

[0070] In the disclosure, " / " and "," can refer to "and / or." For example, "A / B" and "A, B" can refer to "A and / or B." In addition, "A / B / C" and "A, B, C" can refer to "at least one of A, B, and / or C."

[0071] In the disclosure, "or" can refer to "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in the disclosure, "or" can also mean "additionally or alternatively."

[0072] The present disclosure relates to a video / image coding (VCM) for a machine.

[0073] The VCM refers to a compression technique of encoding / decoding a part of source image / video or information obtained from the source image / video for machine vision. In the VCM, an encoding / decoding target can be referred to as a feature. The feature can refer to information extracted from the source image / video based on a task purpose, a requirement, a neighboring environment, etc. The feature can have a different information form from the source image / video, and thus, a feature compression method and an expression format can also be different from a video source.

[0074] The VCM can be applied to various application fields. For example, in a surveillance system that recognizes and tracks an object or a person, the VCM can be used to store or transmit object recognition information. Also, in a smart transportation or intelligent transportation system, the VCM can be used to transmit vehicle location information collected from a GPS, sensing information collected from a LIDAR, a radar, etc., and various vehicle control information to other vehicles or infrastructure. Also, in the field of a smart city, the VCM can be used to perform individual tasks of interconnected sensor nodes or devices.

[0075] The present disclosure provides various embodiments regarding feature / feature map coding. Embodiments of the present disclosure can be implemented individually, or can be implemented in a combination of at least two, unless specifically stated otherwise.

[0076] Overview of VCM system

[0077] Figure 1 FIG. 1 is a diagram schematically illustrating a VCM system to which embodiments of the present disclosure can be applied.

[0078] Referring to Figure 1 , the VCM system can include an encoding apparatus 10 and a decoding apparatus 20.

[0079] The encoding apparatus 10 can compress / encode a feature / feature map extracted from a source image / video to generate a bitstream, and transmit the generated bitstream to the decoding apparatus 20 through a storage medium or a network. The encoding apparatus 10 can also be referred to as a feature encoding apparatus. In the VCM system, a feature / feature map can be generated in each hidden layer of a neural network. The size and number of channels of the generated feature map can vary depending on the type of the neural network or the position of the hidden layer. In the present disclosure, the feature map can be referred to as a feature set, and the feature or the feature map can be referred to as "feature information".

[0080] The encoding apparatus 10 can include a feature obtainer 11, an encoder 12, and a transmitter 13.

[0081] The feature obtainer 11 can obtain a feature / feature map of a source image / video. According to one embodiment, the feature obtainer 11 can obtain a feature / feature map from an external device (e.g., a feature extraction network). In this case, the feature obtainer 11 performs a feature reception interface function. Alternatively, the feature obtainer 11 can obtain a feature / feature map by executing a neural network (e.g., a CNN, a DNN, etc.) using a source image / video as input. In this case, the feature obtainer 11 performs a feature extraction network function.

[0082] According to one embodiment, the encoding apparatus 10 can further include a source image generator (not shown) for obtaining a source image / video, or can alternatively include the feature obtainer 11. The source image generator can be implemented by using an image sensor, a camera module, etc., and can obtain a source image / video through a process of capturing, synthesizing, or generating an image / video. In this case, the generated source image / video can be transmitted to a feature extraction network and used as input data for extracting a feature / feature map.

[0083] The encoder 12 can encode a feature / feature map obtained by the feature obtainer 11. The encoder 12 can perform a series of processes such as prediction, transformation, quantization, etc., to increase encoding efficiency. The encoded data (encoded feature / feature map information) can be output in the form of a bitstream. A bitstream including the encoded feature / feature map information can be referred to as a VCM bitstream.

[0084] The transmitter 13 can obtain feature / feature map information or data output in the form of a bitstream, and can transmit the obtained information or data in the form of a file or streaming to the decoding apparatus 20 or another external object through a digital storage medium or a network. Here, the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 can include an element for generating a media file having a predetermined file format, or an element for transmitting data through a broadcast / communication network. The transmitter 13 can be provided as a separate transmission device from the encoder 12, and in this case, the transmission device can include at least one processor for obtaining feature / feature map information or data output in the form of a bitstream and one transmitter for transmitting the same in the form of a file or streaming.

[0085] The decoding apparatus 20 can obtain feature / feature map information from the encoding apparatus 10, and reconstruct a feature / feature map based on the obtained information.

[0086] The decoding apparatus 20 can include a receiver 21 and a decoder 22.

[0087] The receiver 21 can receive a bitstream from the encoding apparatus 10, and obtain feature / feature map information from the received bitstream to transmit it to the decoder 22.

[0088] The decoder 22 can decode the feature / feature map based on the obtained feature / feature map information. The decoder 22 can perform a series of processes corresponding to the operations of the encoder 14, such as dequantization, inverse transform, prediction, etc., to increase decoding efficiency.

[0089] According to one embodiment, the decoding apparatus 20 can further include a task analysis / rendering unit 23.

[0090] The task analysis / rendering unit 23 can perform task analysis based on the decoded feature / feature map. In addition, the task analysis / rendering unit 23 can decode the feature / feature map to be rendered into a form suitable for performing a task. Based on the task analysis result and the rendered feature / feature map, various (machine-oriented) tasks can be performed.

[0091] Accordingly, the VCM system can encode / decode features extracted from a source image / video according to a user and / or machine request, a task purpose, and a neighboring environment, and perform various (machine-oriented) tasks based on the decoded features. The VCM system can also be implemented by extending / re-designing a video / image codec system, and can perform various encoding / decoding methods defined in the VCM standard.

[0092] VCM pipeline

[0093] Figure 2 FIG. 1 is a diagram schematically illustrating a VCM pipeline structure to which embodiments of the disclosure can be applied.

[0094] Referring to Figure 2 , the VCM pipeline 200 can include a first pipeline 210 for encoding / decoding an image / video and a second pipeline 220 for encoding / decoding a feature / feature map. In the disclosure, the first pipeline 210 can be referred to as a video codec pipeline, and the second pipeline 220 can be referred to as a feature codec pipeline.

[0095] The first pipeline 210 can include a first stage 211 for encoding an input image / video and a second stage 212 for decoding the encoded image / video to generate a reconstructed image / video. The reconstructed image / video can be used for human viewing, i.e., human vision.

[0096] The second pipeline 220 can include a third stage 221 for extracting features / feature maps from an input image / video, a fourth stage 222 for encoding the extracted features / feature maps, and a fifth stage 223 for decoding the encoded features / feature maps to generate reconstructed features / feature maps. The reconstructed features / feature maps can be used for machine (vision) tasks. Here, the machine (vision) tasks can refer to tasks in which machines consume images / videos. The machine vision tasks can be applied in service scenarios such as, for example, surveillance, intelligent transportation, smart city, smart industry, smart content, etc. According to an embodiment, the reconstructed features / feature maps can also be used for human vision.

[0097] According to an embodiment, the features / feature maps encoded in the fourth stage 222 can be sent to the first stage 221 and used for encoding the image / video. In this case, an additional bitstream can be generated based on the encoded features / feature maps, and the generated additional bitstream can be sent to the second stage 222 and used for decoding the image / video.

[0098] According to an embodiment, the features / feature maps decoded in the fifth stage 223 can be sent to the second stage 222 and used for decoding the image / video.

[0099] Although Figure 2 Although the case in which the VCM pipeline 200 includes the first pipeline 210 and the second pipeline 220 is illustrated, this is merely exemplary, and embodiments of the disclosure are not limited thereto. For example, the VCM pipeline 200 can include only the second pipeline 220, or the second pipeline 220 can be extended to a plurality of feature codec pipelines.

[0100] Meanwhile, in the first pipeline 210, the first stage 211 can be performed by an image / video encoder, and the second stage 212 can be performed by an image / video decoder. In addition, in the second pipeline 220, the third stage 221 can be performed by a VCM encoder (or a feature / feature map encoder), and the fourth stage 222 can be performed by a VCM decoder (or a feature / feature map decoder). Hereinafter, the encoder / decoder structure is described in detail.

[0101] Encoder

[0102] Figure 3 is a diagram schematically illustrating an image / video encoder to which embodiments of the disclosure can be applied.

[0103] Referring to Figure 3The image / video encoder 300 can include an image partitioner 310, a predictor 320, a residual processor 330, an entropy encoder 340, an adder 350, a filter 360, and a memory 370. The predictor 320 can include an inter-predictor 321 and an intra-predictor 322. The residual processor 330 can include a transformer 332, a quantizer 333, a dequantizer 334, and an inverse transformer 335. The residual processor 330 can further include a subtractor 331. The adder 350 can be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-described image partitioner 310, predictor 320, residual processor 330, entropy encoder 340, adder 350, and filter 360 can be configured by at least one hardware component (e.g., an encoder chipset or a processor). In addition, the memory 370 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The above-described hardware component can further include the memory 370 as an internal / external component.

[0104] The image partitioner 310 can partition an input image (or picture, frame) input to the image / video encoder 300 into at least one processing unit. As an example, the processing unit can be referred to as a coding unit (CU). The coding unit can be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU, L-unit) according to a quad-tree binary-tree ternary (QTBTTT) structure. For example, one coding unit can be partitioned into a plurality of coding units having a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary structure can be applied later. Alternatively, the binary-tree structure can be applied first. The image / video encoding process according to the disclosure can be performed based on a final coding unit that is no longer partitioned. In this case, the largest coding unit can be used as the final coding unit, etc., based on coding efficiency according to image characteristics, etc., or if necessary, the coding unit can be recursively divided into coding units of a deeper depth, and a coding unit of an optimal size can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, reconstruction, etc., described later. As another example, the processing unit can further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be divided or partitioned from the above-described final coding unit, respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.

[0105] In some cases, the term unit can be used interchangeably with terms such as block, region, and the like. In general, an MxN block can indicate a set of transform coefficients or a set of samples composed of M columns and N rows. A sample can generally indicate a pixel or a pixel value, or can indicate only a pixel / pixel value of a luma component, or can indicate only a pixel / pixel value of a chroma component. A sample can be used as a term corresponding to a pixel or a pel.

[0106] The image / video encoder 300 can generate a residual signal (a residual block, a residual sample array) by subtracting a prediction signal (a prediction block, a prediction sample array) output from the inter-predictor 321 or the intra-predictor 322 from an input image signal (an original block, an original sample array), and the generated residual signal is transmitted to the transformer 332. In this case, as illustrated, a unit that subtracts the prediction signal (the prediction block, the prediction sample array) from the input image signal (the original block, the original sample array) within the image / video encoder 300 can be referred to as a subtractor 331. The predictor can perform prediction on a block to be processed (hereinafter, referred to as a current block), and generate a prediction block including predicted samples of the current block. The predictor can determine whether intra-prediction or inter-prediction is applied in a unit of the current block or CU. The predictor can generate various information related to prediction, such as prediction mode information, and transmit the same to the entropy encoder 340. The information related to prediction can be encoded by the entropy encoder 340, and can be output in the form of a bitstream.

[0107] The intra-predictor 322 can predict the current block by referring to samples within the current picture. In this case, the reference samples can be located in a neighboring region of the current block, or can be located more distantly according to a prediction mode. In intra-prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, a DC mode and a planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the granularity of prediction direction. However, this is one example, and a greater or smaller number of directional prediction modes can be used according to configuration. The intra-predictor 322 can also determine a prediction mode applied to the current block by using a prediction mode applied to a neighboring block.

[0108] The inter predictor 321 can derive a prediction block of a current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted at a block, sub-block, or sample level based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block or a collocated CU (colCU), and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 321 can construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. The inter prediction can be performed based on various prediction modes, and for example, in the skip mode and the merge mode, the inter predictor 321 can use the motion information of the neighboring blocks as the motion information of the current block. In the skip mode, unlike the merge mode, a residual signal can not be transmitted. In a motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling a motion vector difference.

[0109] The predictor 320 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction for prediction of one block, and can also simultaneously apply both intra prediction and inter prediction. It can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding, for example, such as screen content coding (SCC) or the like. The IBC basically performs prediction within the current picture, but since it derives a reference block within the current picture, it can be similar to the operation of inter prediction. In other words, the IBC can use at least one of the inter prediction methods described in the present disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, a sample value within a picture can be signaled based on information related to a palette table and a palette index.

[0110] The prediction signal generated by the predictor 320 can be used to generate a reconstructed signal or to generate a residual signal. The transformer 332 can generate transform coefficients by applying a transform method to the residual signal. For example, the transform method can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when relationship information between pixels is indicated as a graph. The CNT refers to a transform obtained based on a prediction signal generated using all previous reconstructed pixels. In addition, the transform process can be applied to a pixel block of the same square size or a non-square variable size block.

[0111] The quantizer 333 can quantize the transform coefficients and transmit them to the entropy encoder 340, and the entropy encoder 340 can encode and output information about the quantized transform coefficients (information about the quantized transform coefficients) as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 333 can reorder the block-shaped quantized transform coefficients in the form of a one-dimensional vector based on a coefficient scan order, and can generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of a one-dimensional vector. The entropy encoder 340 can perform various encoding methods, such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and the like. The entropy encoder 340 can encode the quantized transform coefficients together with information necessary for video / image reconstruction, such as values of syntax elements, and the like, or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in a network abstraction layer (NAL) unit. The image / video information can further include information about various parameter sets, such as adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS), and the like. In addition, the video / image information can further include general constraint information. In addition, the image / video information can further include a method for generating and using the encoded information, a purpose thereof, and the like. In the disclosure, information and / or syntax elements transmitted / signalized from the image / video encoder to the image / video decoder can be included in the image / video information. The image / video information can be encoded through the encoding process described above and included in the bitstream. The bitstream can be transmitted through a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the like, and the digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. A transmitter (not shown) for transmission and / or a storage unit (not shown) for storing a signal output from the entropy encoder 340 can be constructed as an internal / external element of the image / video encoder 300, or the transmitter can be included in the entropy encoder 340.

[0112] The quantized transform coefficients output from the quantizer 333 can be used to generate a prediction signal. For example, a residual signal (a residual block or a residual sample array) can be reconstructed by the dequantizer 334 and the inverse transformer 335 by applying dequantization and inverse transformation to the quantized transform coefficients. The adder 350 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-predictor 321 or the intra-predictor 322. When there is no residual for the processing target block, for example, when a skip mode is applied, the prediction block can be used as the reconstructed block. The adder 350 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of the next processing target block within the current picture, and as described later, can also be used for inter-prediction of the next picture by filtering.

[0113] Meanwhile, the luma mapping with chroma scaling can be applied in the picture encoding and / or reconstruction process.

[0114] The filter 360 can apply filtering to the reconstructed signal to enhance subjective / objective image quality. For example, the filter 360 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and can store the modified reconstructed picture in the memory 370, specifically, in the DPB of the memory 370. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 360 can generate and transmit various filtering-related information to the entropy encoder 340. The filtering-related information can be encoded by the entropy encoder 340 and output in the form of a bitstream.

[0115] The modified reconstructed picture transmitted to the memory 370 can be used as a reference picture in the inter-predictor 321. By this, prediction mismatch at the encoder side and the decoder side can be avoided, and coding efficiency can be improved.

[0116] The DPB of the memory 370 can store the modified reconstructed picture to be used as a reference picture in the inter-predictor 321. The memory 370 can store motion information of a block in which motion information within the current picture and / or motion information of a block within the reconstructed picture is derived (or encoded). The stored motion information can be transmitted to the inter-predictor 321 to be used as motion information of a spatially or temporally neighboring block. The memory 370 can store reconstructed samples of a reconstructed block in the current picture, and transmit the stored reconstructed samples to the intra-predictor 322.

[0117] Meanwhile, the VCM encoder (or feature / feature map encoder) can have substantially the same structure as the basic encoder described above with reference to FIG. 1. Figure 3The described image / video encoder 300 is the same / similar structure as it performs a series of processes such as prediction, transform, quantization, etc. to encode features / feature maps. However, the VCM encoder is different from the image / video encoder 300 in that it targets features / feature maps for encoding, and thus, it can be different in the name of each unit (or component) (e.g., image partitioner 310, etc.) and its specific operational details from the image / video encoder 300. The specific operational details of the VCM encoder will be described later in detail.

[0118] Decoder

[0119] Figure 4 FIG. is schematically illustrates an image / video decoder to which embodiments of the disclosure can be applied.

[0120] Referring to Figure 4 , the image / video decoder 400 can include an entropy decoder 410, a residual processor 420, a predictor 430, an adder 440, a filter 450, and a memory 460. The predictor 430 can include an inter-predictor 431 and an intra-predictor 432. The residual processor 420 can include a dequantizer 421 and an inverse transformer 422. According to one embodiment, the above-described entropy decoder 410, residual processor 420, predictor 430, adder 440, and filter 450 can be configured by one hardware component (e.g., a decoder chipset or a processor). In addition, the memory 460 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component can further include the memory 460 as an internal / external component.

[0121] When a bitstream including video / image information is input, the image / video decoder 400 can reconstruct an image / video in response to a process of processing image / video information in Figure 3 , for example, the image / video decoder 400 can derive units / blocks based on block partitioning-related information obtained from the bitstream. The image / video decoder 400 can perform decoding by using a processing unit applied in the image / video encoder. Thus, the decoded processing unit can be, for example, a coding unit, and can be partitioned according to a quadtree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a largest coding unit. At least one transform unit can be derived from the coding unit. Also, a reconstructed image signal decoded and output by the image / video decoder 400 can be played by a playing device.

[0122] The image / video decoder 400 can obtain the bitstream from Figure 3The encoder in the image / video encoder receives the signal output and can decode the received signal through the entropy decoder 410. For example, the entropy decoder 410 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets, such as adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), video parameter set (VPS), etc. In addition, the video / image information can further include general constraint information. In addition, the image / video information can include a generation method, a usage method, a purpose, etc. of the decoded information. The image / video decoder 400 can further decode the picture based on the information on the parameter sets and / or the general constraint information. The signaled / received information and / or syntax elements can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 410 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and can output values of syntax elements necessary for image reconstruction and quantized values of transform coefficients related to a residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model by using information of a decoding target syntax element, decoding information of a neighboring block and a decoding target block, or information of a symbol / bin decoded in a previous step, and predict a probability of bin occurrence according to the determined context model and perform arithmetic decoding of the bin to generate a symbol of a value corresponding to each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using information of a decoded symbol / bin for a context model of a next symbol / bin. Among the information decoded by the entropy decoder 410, prediction-related information can be provided to the predictor (inter-predictor 432 and intra-predictor 431), and the residual values (i.e., quantized transform coefficients and related parameter information) entropy-decoded by the entropy decoder 410 can be input to the residual processor 420. The residual processor 420 can derive a residual signal (a residual block, a residual sample, a residual sample array). In addition, among the information decoded by the entropy decoder 410, filter-related information can be provided to the filter 450. Meanwhile, a receiver (not shown) that receives the signal output from the image / video encoder can be additionally constructed as an internal / external element of the image / video decoder 400, or the receiver can be a component of the entropy decoder 410. Meanwhile, the image / video decoder according to the disclosure can also be referred to as an image / video decoding apparatus, and the image / video decoder can be divided into an information decoder (image / video information decoder) and / or a sample decoder (image / video sample decoder).In this case, the information decoder can include the entropy decoder 410, and the sample decoder can include at least one of the dequantizer 321, the inverse transformer 322, the adder 440, the filter 450, the memory 460, the inter predictor 432, and the intra predictor 431.

[0123] The dequantizer 421 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 421 can re-order the quantized transform coefficients in the form of a two-dimensional block. In this case, the re-ordering can be performed based on a coefficient scanning order performed in the image / video encoder. The dequantizer 321 can perform dequantization on the quantized transform coefficients by using a quantization parameter (i.e., quantization step length information), and can obtain the transform coefficients.

[0124] The inverse transformer 422 can perform inverse transform on the transform coefficients to obtain a residual signal (a residual block, a residual sample array).

[0125] The predictor 430 can perform prediction for the current block and generate a prediction block including prediction samples of the current block. The predictor can determine whether intra prediction or inter prediction is applied to the current block based on prediction-related information output from the entropy decoder 410, and can determine a specific intra / inter prediction mode (prediction method).

[0126] The predictor 420 can generate a prediction signal based on various prediction methods. For example, the predictor can not only apply intra prediction or inter prediction, but also simultaneously apply intra prediction and inter prediction for prediction of one block. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding of a game, such as screen content coding (SCC), etc. The IBC basically performs prediction within a current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. In other words, the IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information related to a palette table and a palette index can be included in the image / video information and signaled.

[0127] The intra predictor 431 can predict the current block by referring to samples within the current picture. The reference samples can be located in the neighborhood of the current block, or can be located away from the current block according to the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 431 can determine the prediction mode applied to the current block by using the prediction mode applied to a neighboring block.

[0128] The inter predictor 432 can derive a prediction block of the current block based on reference blocks (reference sample arrays) specified by motion vectors on reference pictures. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted at a block, sub-block, or sample level based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In the inter prediction, the neighboring blocks can include spatial neighboring blocks within the current picture and temporal neighboring blocks in the reference pictures. For example, the inter predictor 432 can construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. The inter prediction can be performed based on various prediction modes, and the prediction-related information can include information indicating an inter prediction mode of the current block.

[0129] The adder 440 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (prediction block, prediction sample array) output from the predictors (including the inter predictor 432 and / or the intra predictor 431). When there is no residual for the processing target block, such as when a skip mode is applied, the prediction block can be used as the reconstructed block.

[0130] The adder 440 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of the next processing target block within the current picture, or can be output through filtering as described later, or can be used for inter prediction of the next picture.

[0131] Meanwhile, luma mapping with chroma scaling can be applied in the picture decoding process.

[0132] The filter 450 can apply filtering to the reconstructed signal to enhance subjective / objective image quality. For example, the filter 450 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and can transmit the modified reconstructed picture to the memory 460, specifically, to the DPB of the memory 460. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0133] The (modified) reconstructed pictures stored in the DPB of the memory 460 can be used as reference pictures in the inter prediction 432. The memory 460 can store motion information of blocks in which motion information of blocks in the already reconstructed pictures and / or the current picture is derived (or decoded). The stored motion information can be sent to the inter prediction 432 to be used as motion information of spatially or temporally neighboring blocks. The memory 460 can store reconstructed samples of reconstructed blocks in the current picture and send them to the intra prediction 431.

[0134] Meanwhile, the VCM decoder (or feature / map decoder) can have the same / similar structure as the image / video decoder 400 described above by referring to Figure 4 the above because it performs a series of processes such as prediction, inverse transform, dequantization, etc. to decode features / maps. However, the VCM decoder is different from the image / video decoder 400 because it targets features / maps for decoding, and thus, it can be different in the name of each unit (or component) (e.g., DPB, etc.) and its specific operational details from the image / video decoder 400. The operation of the VCM decoder can correspond to that of the VCM encoder, and its specific operational details will be described later in detail.

[0135] Feature / feature map encoding process

[0136] Figure 5 is a flowchart schematically illustrating a feature / map encoding process to which embodiments of the disclosure can be applied.

[0137] Referring to Figure 5 , the feature / map encoding process can include a prediction process S510, a residual processing process S520, and an information encoding process S530.

[0138] The prediction process S510 can be performed by the predictor 320 described above by referring to Figure 3 .

[0139] In particular, the intra predictor 322 can predict the current block (i.e., the current set of encoded feature elements) by referring to the feature elements in the current feature / feature map. The intra prediction can be performed based on the spatial similarity of the feature elements configuring the feature / feature map. For example, it can be estimated that the feature elements included in the same region of interest (RoI) within the image / video have similar data distribution characteristics. Thus, the intra predictor 322 can predict the current block by referring to the pre-reconstructed feature elements within the region of interest including the current block. In this case, the referred feature elements can be located in the vicinity of the current block, or can be positioned apart from the current block according to a prediction mode. The intra prediction modes for the feature / feature map encoding can include a plurality of non-directional prediction modes and a plurality of directional prediction modes. The non-directional prediction modes can include, for example, prediction modes corresponding to the DC mode and the planar mode of the image / video encoding process. In addition, the directional modes can include, for example, prediction modes corresponding to 33 directional modes or 65 directional modes of the image / video encoding process. However, this is merely an example, and according to one embodiment, the type and number of intra prediction modes can be configured / changed in various ways.

[0140] The inter predictor 321 can predict the current block based on a reference block (i.e., a set of reference feature elements) specified by motion information on the reference feature / feature map. The inter prediction can be performed based on temporal similarity of the feature elements configuring the feature / feature map. For example, temporally consecutive features can have similar data distribution characteristics. Accordingly, the inter predictor 321 can predict the current block by referring to the pre-reconstructed feature elements of the current feature and temporally neighboring features. In this case, the motion information for specifying the reference feature elements can include a motion vector and a reference feature / feature map index. The motion information can further include information related to an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). For the inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current feature / feature map and temporal neighboring blocks present in the reference feature / feature map. The reference feature / feature map including the reference block and the reference feature / feature map including the temporal neighboring blocks can be the same or different. The temporal neighboring blocks can be referred to as collocated reference blocks, etc., and the reference feature / feature map including the temporal neighboring blocks can be referred to as a collocated feature / feature map. The inter predictor 321 can configure a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or the reference feature / feature map index of the current block. The inter prediction can be performed based on various prediction modes, and, for example, for a skip mode and a merge mode, the inter predictor 321 can use motion information of the neighboring blocks as motion information of the current block. For the skip mode, unlike the merge mode, a residual signal can not be transmitted. For a motion vector prediction (MVP) mode, a motion vector of the neighboring blocks can be used as a motion vector predictor, and a motion vector of the current block can be indicated by signaling a motion vector difference. In addition to the above-described intra prediction and inter prediction, the predictor 320 can generate a prediction signal based on various prediction methods.

[0141] The prediction signal generated by the predictor 320 can be used to generate a residual signal (residual block, residual feature element) S520. The residual processing process S520 can be performed by the residual processor 330 described above by referring to Figure 3 In addition to the residual information in the bitstream, the entropy encoder 340 can encode information necessary for feature / feature map reconstruction, e.g., prediction information (e.g., prediction mode information, motion information, etc.).

[0142] Meanwhile, the feature / feature map encoding process can further include a process for generating a reconstructed feature / feature map for the current feature / feature map, and a process (optional) for applying in-loop filtering to the reconstructed feature / feature map, and a process S530 for encoding information (e.g., prediction information, residual information, partition information, etc.) for feature / feature map reconstruction and outputting it in the form of a bitstream.

[0143] The VCM encoder can derive (modified) residual features from the quantized transform coefficients through dequantization and inverse transform, and can generate reconstructed features / feature maps based on the predicted features as the output of S510 and the (modified) residual features. The reconstructed features / feature maps generated in this way can be the same as the reconstructed features / feature maps generated by the VCM decoder. When the in-loop filtering process is performed on the reconstructed features / feature maps, modified reconstructed features / feature maps can be generated through the in-loop filtering process on the reconstructed features / feature maps. The modified reconstructed features / feature maps can be stored in a decoded feature buffer (DFB) or a memory, and then used as reference features / feature maps in the prediction process of the features / feature maps. In addition, (in-loop) filtering-related information (parameters) can be encoded and output in the form of a bitstream. Through the in-loop filtering process, noise that can occur during feature / feature map coding can be removed, and task performance based on the features / feature maps can be improved. In addition, the in-loop filtering process can be performed on both the encoder side and the decoder side to guarantee the recognition of the prediction results, improve the reliability of the feature / feature map coding, and reduce the amount of data transmission for the feature / feature map coding.

[0144] Feature / feature map decoding process

[0145] Figure 6 FIG. 4 is a flowchart schematically illustrating a feature / feature map decoding process to which embodiments of the disclosure can be applied.

[0146] Reference Figure 6, the feature / feature map decoding process can include an image / video information acquisition process S610, a feature / feature map reconstruction process S620 to S640, and an in-loop filtering process S650 for reconstructing the feature / feature map. The feature / feature map reconstruction process can be performed based on the prediction signal and the residual signal obtained through the processes of inter / intra prediction S620, residual processing S630, and dequantization and inverse transform for the quantized transform coefficients described in the present disclosure. The modified reconstructed feature / feature map can be generated through the in-loop filtering process for reconstructing the feature / feature map, and the modified reconstructed feature / feature map can be output as a decoded feature / feature map. The decoded feature / feature map can be stored in a decoded feature buffer (DFB) or a memory and then used as a reference feature / feature map in an inter prediction process when decoding the feature / feature map. In some cases, the above-mentioned in-loop filtering process can be omitted. In this case, the reconstructed feature / feature map can be output as is as a decoded feature / feature map, and can be stored in a decoded feature buffer (DFB) or a memory and then used as a reference feature / feature map in an inter prediction process when decoding the feature / feature map.

[0147] Feature extraction method and data distribution feature

[0148] Embodiments of the present disclosure propose a method for generating a prediction process and a related bitstream required to compress an activation (feature) map generated in a hidden layer of a deep neural network.

[0149] Input data provided to a deep neural network undergoes a calculation process through a plurality of hidden layers, and according to the type of deep neural network in use and the position of the hidden layer within the corresponding deep neural network, the operation result from each hidden layer is output as a feature / feature map with various sizes and numbers of channels.

[0150] Figure 7 FIG. 1 is a diagram illustrating an example of a feature extraction and reconstruction method to which embodiments of the present disclosure can be applied.

[0151] Referring to Figure 7 The feature extraction network 710 can extract an intermediate layer activation (feature) map of a deep neural network from a source image / video and output the extracted feature map. The feature extraction network 710 can be a set of consecutive hidden layers from the input of the deep neural network.

[0152] The encoding device 720 can compress and output the output feature map in the form of a bitstream, and the decoding device 730 can reconstruct the (compressed) feature map from the output bitstream. The encoding device 720 can correspond to Figure 1The encoder 12 of FIG. 1 can correspond to the encoder 12 of FIG. 7, and the decoding device 730 can correspond to the decoder 22 of Table 1. The task network 740 can perform a task based on the reconstructed feature map.

[0153] The number of channels of the feature map to be compressed in the VCM can differ depending on the network used for feature extraction and the extraction location, and can be greater than the number of channels of the input data.

[0154] Figure 8 is a diagram illustrating an example of an image segmentation method to which embodiments of the disclosure can be applied. For example, it illustrates CTUs, slices, and tiles within an image.

[0155] A video / image encoding method according to the present document can be performed based on the following segmentation structure. The segmentation structure can be derived from a segmentation structure of a picture. Figure 8 Processes such as prediction, residual processing (transform / inverse transform, quantization / dequantization, etc.), syntax element encoding, filtering, etc. can be performed on CTUs, CUs (and / or TUs, PUs) derived from the segmentation structure of FIG. 1. The block segmentation process can be performed by an encoding device, and segmentation-related information can be encoded in the form of a bitstream and transmitted to a decoding device. The decoding device can derive the block segmentation structure of the current picture based on the segmentation-related information obtained from the bitstream, and can perform a series of processes for image decoding (e.g., prediction, residual processing, block / picture reconstruction, loop filtering, etc.) based on the block segmentation structure. The CU size can be equal to the TU size, and there can be multiple TUs within a CU region. Meanwhile, the CU size generally refers to a luma component (sample) CB size. The TU size generally refers to a luma component (sample) TB size. Depending on the component ratio of the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image, the chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size, and the transform / inverse transform can be performed in units of TUs (TBs).

[0156] Furthermore, in video / image coding according to the present document, a picture processing unit can have a hierarchical structure. A picture can be divided into one or more CUs, and one or more CUs can be grouped and distinguished as one or more tiles, slices, and / or tile groups. One slice can contain one or more tiles. One tile can include one or more CTU rows within a tile. A slice can include an integer number of tiles of a picture. One tile group can include one or more tiles. One tile can include one or more CTUs. A CTU can be divided into one or more CUs. A tile group can include an integer number of tiles according to a tile raster scan within a picture. A slice header can carry information / parameters applicable to a corresponding slice (i.e., blocks within a slice). A picture header can carry information / parameters applicable to a corresponding picture (or blocks within a picture). When a coding / decoding apparatus includes a multi-core processor, coding / decoding processes for tiles, slices, tiles, and / or tile groups can be performed in parallel. In the present document, a slice and a tile group can be used interchangeably. In other words, a tile group header can be referred to as a slice header. Here, a slice can have one of slice types including an I slice, a P slice, and a B slice.

[0157] In a coding apparatus, a tile / tile group, a slice, a tile, and maximum and minimum coding unit sizes can be determined depending on characteristics (e.g., resolution) of a video picture or considering coding efficiency or parallel processing, and information about them can be included in a bitstream or can be derived therefrom.

[0158] In a decoding apparatus, information indicating whether a tile / tile group, a tile, a slice, or a CTU within a tile of a current picture is divided into a plurality of coding units can be obtained. Such information can be obtained (or transmitted) only under certain conditions to improve efficiency.

[0159] A slice header (slice header syntax) can include information / parameters generally applicable to a slice. An APS (APS syntax) or a PPS (PPS syntax) can include information / parameters generally applicable to one or more pictures. An SPS (SPS syntax) can include information / parameters generally applicable to one or more sequences. A VPS (VPS syntax) can include information / parameters generally applicable to multiple layers. A DPS (DPS syntax) can include information / parameters generally applicable to an entire video. A DPS can include information / parameters related to a concatenation of coded video sequences (CVSs).

[0160] In the present document, the term high-level syntax can include at least one of an APS syntax, a PPS syntax, an SPS syntax, a VPS syntax, a DPS syntax, a picture header syntax, or a slice header syntax.

[0161] Also, for example, information about the partitioning and configuration of tiles / tile groups / tiling / slices can be configured at the encoding side by high-level syntax and can be transmitted to the decoding apparatus in the form of a bitstream.

[0162] Figure 9 is a diagram illustrating an example of a VCM image encoding / decoding system, and Figure 10 is a diagram illustrating another example of a VCM image encoding / decoding system. For example, Figure 9 and Figure 10 may show an extension / re-design of a video coding system (e.g., Figure 1 ) to use only a part of a video source or to obtain and use necessary parts / information from a video source depending on a user's or machine's request, purpose, and neighboring environment. In other words, Figure 9 and Figure 10 may relate to video coding for machines (VCM).

[0163] Video coding for machines (VCM) can refer to encoding / decoding of an entire image and / or a part of an image and / or necessary information (features) from an image depending on a user's and / or machine's request, purpose, and neighboring environment. The target of encoding in VCM can be an image itself, feature information extracted from an image according to a user's and / or machine's request, purpose, and neighboring environment, or a set of sequence information varying over time.

[0164] Referring to Figure 9 , a VCM system can include an image encoder and an image decoder for VCM. A source device Figure 1 may transmit encoded image information to a receiving device via a storage medium or a network. An entity using the device can be a human and / or a machine.

[0165] Referring to Figure 10 , a VCM system can include an image encoder and an image decoder for VCM. A source device can transmit encoded image information (features) to a receiving device through a storage medium or a network. An entity using the device can be a human and / or a machine.

[0166] For example, a process of extracting information (i.e., features) from an image can be referred to as feature extraction. Feature extraction can be performed in both a video / image capturing device and a video / image generating device. Features can be information extracted / processed from an image according to a user's and / or machine's request, purpose, and neighboring environment, and can represent a set of sequence information varying over time.

[0167] An image encoder for VCM can perform a series of processes such as prediction, transformation, quantization, etc. to effectively compress and encode the entire image and / or a part of the image and / or features. The encoded data can be output in the form of a bitstream.

[0168] An image decoder for VCM can decode a video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding apparatus (in other words, the encoder).

[0169] Decoded images and / or features can be rendered. Also, they can be used to perform a user's or machine's task. Examples of such tasks can include AI and computer vision tasks such as face recognition, behavior recognition, lane recognition, etc.

[0170] The present disclosure provides various embodiments related to acquisition and encoding of the entire image and / or a part of the image for VCM, and unless otherwise specified, these embodiments can be combined with and performed with each other. The methods / embodiments of the present disclosure can be applied to the methods disclosed in the video coding (VCM) standard for machines.

[0171] VCM layer structure

[0172] VCM can be based on a layer structure consisting of a feature coding layer, a neural network (feature) abstraction layer, and a feature extraction layer. Figure 11 is a diagram illustrating an example of a VCM layer structure, and Figure 12 is a diagram illustrating an example of a VCM bitstream consisting of encoded abstract features and NNAL information.

[0173] As an example, referring to Figure 11 , the VCM layer structure can include a feature extraction layer 1110, a neural network (feature) abstraction layer 1120, and a feature coding layer 1130.

[0174] The feature extraction layer 1110 can refer to a layer for extracting features from an input source, and can also include the result of extraction. The feature coding layer 1130 can refer to a layer for compressing the extracted features, and can also include the result of compression.

[0175] The neural network abstraction layer 1120 can abstract information generated from the feature extraction layer 1110 (e.g., information about extracted features / feature maps) and transmit the same to the feature coding layer 1130. The neural network abstraction layer 1120 can hide the internal structure of the feature extraction layer 1110 and provide a consistent feature interface function through abstraction of information. Accordingly, the feature coding layer 1130 can perform a consistent feature coding process even when a compression target is changed due to a change in a tool (e.g., a CNN, a DNN, etc.). In the disclosure, the neural network abstraction layer (NNAL) can also be referred to as a feature abstraction layer.

[0176] The interface between the feature extraction layer 1110 and the neural network abstraction layer 1120, and the interface between the feature coding layer 1130 and the neural network abstraction layer 1120 can be defined in advance, and the operation in the neural network abstraction layer 1120 can be configured to be modifiable thereafter.

[0177] Referring to Figure 12 A bitstream configured as illustrated can be referred to as a neural network abstraction layer (NNAL) unit. The NNAL unit can be an independent feature reconstruction unit. The input feature of a single NNAL unit can be extracted from the same layer within a neural network. Accordingly, the input feature of a single NNAL unit can be forced to have the same feature. For example, the same feature extraction method can be applied to the input feature of a single NNAL unit.

[0178] The NNAL unit can include an NNAL unit header and an NNAL unit payload. The NNAL unit header can include all information required to perform a task with a coded feature. The NNAL unit payload can include abstracted feature information. The NNAL unit payload can include a group header and group data. The group header can include configuration information of feature group data, such as a temporal order, a number, or a common attribute of feature channels constituting a feature group. The feature channel can refer to a unit of a coded feature. The group data can include a plurality of feature channels and a coding indicator, and each feature channel can include type information, prediction information, side information, and residual information. In this case, the type information can indicate a coding method, and the prediction information can indicate a prediction method. In addition, the side information can indicate additional information required for decoding (e.g., entropy coding, quantization-related information, etc.), and the residual information can include information about coded feature elements (i.e., a set of feature value information).

[0179] Embodiment

[0180] Embodiments of the present disclosure can relate to a feature / feature map encoding process and a feature / feature map decoding process, and can relate to a process of encoding / decoding a compressed bitstream to perform a machine task. For example, according to embodiments of the present disclosure, in order to generate a compressed bitstream for features / feature maps extracted from a deep neural network, each channel data of the encoded / decoded features / feature maps can be classified as a region of interest (ROI) and a non-ROI in performing a machine task. According to embodiments of the present disclosure, in order to reconstruct the classified ROIs and non-ROIs, region information can be required, and a partial coding method can be proposed in which encoding of the non-ROIs of the corresponding channels is omitted during encoding and the non-ROIs are reconstructed to a specific value during decoding. When encoding the channel data of the features / feature maps, this method can limit the encoding target region, thereby more effectively reducing the size of the bitstream while still being able to perform a machine task with a similar level of accuracy after encoding / decoding.

[0181] Meanwhile, in describing embodiments of the present disclosure, the term feature (or feature map) is used for explanation, but it is obvious that embodiments of the present disclosure can be applied to features, feature maps, or images, etc.

[0182] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0183] As an example, Figure 13 is an example of visualizing a plurality of channel data in a feature / feature map according to embodiments of the present disclosure, and relates to a process of generating a compressed bitstream for features / feature maps extracted from a deep neural network to perform a machine task. More specifically, Figure 13 (a) is an example of an input image of a deep neural network, and Figure 13 (b) is an example of visualizing a plurality of channel data in a feature / feature map extracted from the corresponding input image. As Figure 13As shown in the example of (b), the features / feature maps extracted from one input image can consist of multiple channels, and within each channel, there can be a difference in the distribution of channel data by region. In the present disclosure, in order to generate a compressed bitstream for features / feature maps extracted from a deep neural network to perform a machine task, the region-wise difference in the distribution of channel data can be used to distinguish each channel data of the encoded / decoded features / feature maps. For example, in performing a machine task, each channel data can be classified into a region of interest (ROI) and a non-ROI based on certain criteria such as whether an object is included and / or the frequency of regions, and by using region information, encoding of the non-ROI of the corresponding channel can be omitted during encoding, and the non-ROI can be reconstructed to a certain value during decoding. As an embodiment, in order to generate a compressed bitstream for features / feature maps extracted from a deep neural network to perform a machine task, the region-wise difference in the distribution of data within each channel of the features / feature maps can be used to classify each channel data of the encoded / decoded features / feature maps into an ROI and a non-ROI, and a partial encoding / decoding method can be proposed in which encoding of the non-ROI of the corresponding channel is omitted during encoding, and the non-ROI is reconstructed to a certain value during decoding by using region information.

[0184] Figure 14 is a diagram showing an example of region segmentation and region data distribution of feature / feature map channel data according to an embodiment of the present disclosure. More specifically, Figure 14 (a) is an example of visualizing one channel in a feature / feature map according to an embodiment of the present disclosure as a plurality of regions of the same size, and Figure 14(b) is an example of a histogram representing the data distribution of the entire channel data, where (b)-1 and (b)-2 are examples of histograms representing the channel data distribution of the regions as indicated in (b)-1 and (b)-2 in (a), respectively. In the case of (b)-1, it can be seen that the average of the data is similar to the average of the entire channel of (b), and the variance is relatively smaller than the variance of the entire channel data, while in the case of (b)-2, it can be seen that the average of the data is different from the average of the entire channel, and the variance is also relatively larger than the variance of the entire channel data. Therefore, the channel data of the (b)-1 region can have less impact on performing a machine task than the channel of the (b)-2 region. Therefore, the (b)-1 region can be classified as a non-ROI, the (b)-2 region can be classified as an ROI, and other regions not illustrated can also be classified according to the same criteria. Here, the channel data is divided into multiple regions of the same size for ease of illustration, and can be practically segmented in various forms of different sizes. In addition, here, the channel data is divided into ROIs and non-ROIs using the average and variance of each region for ease of illustration, and in practice, other criteria can be applied, and these embodiments are obviously included in the embodiments of the present disclosure.

[0185] According to the present embodiment, by effectively reducing the size of the compressed bitstream, the effect of reducing the encoding target region when encoding the channel data can be obtained, and the decoding process can also be more efficiently performed during decoding. In addition, it can perform a machine task with a similar level of accuracy after encoding / decoding, while further reducing the size of the compressed bitstream more efficiently.

[0186] Figure 15 is an example illustrating a partial encoding / decoding method of ROIs and non-ROIs proposed in embodiments according to the present disclosure.

[0187] For example, on the encoding side, there can be original channel data 1501 for features / feature maps of an image. In this regard, based on the entire channel data 1501, a region classification 1502 process for classifying ROIs and non-ROIs can be performed. For example, since the regions classified as non-ROIs according to the regional distribution of the channel data can have relatively less impact on performing a machine task, encoding of the channel data of the non-ROIs can be omitted during encoding, and only the ROIs can be encoded. For example, the ROIs 1506 can be encoded 1505. In other words, information about the ROIs can be encoded into the bitstream 1500. At the same time, based on the entire channel data 1501, a representative value (e.g., a statistical value such as an average) of the channel data can be derived 1503, and the average can be encoded 1504 into the bitstream 1500.

[0188] In this regard, at the decoding side, the encoded bitstream 1500 can be obtained, and the decoding for the channel data can be performed. For example, during the decoding, the ROI can be decoded based on the information about the encoded ROI, but the channel data for the non-ROI can be decoded in a different manner from the ROI. Based on the bitstream, the partial decoding 1510 of the channel data can be performed. For example, first, the ROI can be reconstructed 1511 based on the information about the ROI obtained from the bitstream. Meanwhile, the representative value (e.g., average value) for the channel data obtained from the bitstream can be decoded 1512, and the non-ROI can be reconstructed 1513 based on the representative value. Accordingly, the entire channel data can be reconstructed 1514 based on the reconstructed ROI and the reconstructed non-ROI.

[0189] For example, the representative value (e.g., average value) of the channel data to be encoded or decoded can be the representative value (e.g., average value) of the entire channel data, but can also be the representative value (e.g., average value) of the channel data for the non-ROI.

[0190] Meanwhile, Tables 1 and 2 below are examples of necessary information that can be signaled when applying the partial encoding / decoding method proposed in the present embodiment. Table 1 is an example of a syntax for signaling whether to use the partial encoding / decoding method of the channel data in the present example in a video sequence, and Table 2 is an example of a syntax indicating whether there is a coding unit (CU) within each channel to which the partial encoder is applied, a syntax for signaling a substitute value for reconstructing for a CU for which encoding is omitted, and a syntax indicating whether each CU is a CU for which encoding is omitted.

[0191] [Table 1]

[0192]

[0193] As an example, Table 1 is a table of examples for explaining a syntax that can be included in a parameter set (e.g., sequence parameter set) according to an embodiment of the present disclosure. For example, the parameter set (e.g., sequence parameter set SPS) can include information indicating whether the channel data is partially encoded according to an embodiment of the present disclosure as described above. For example, the information can be expressed as a syntax sps_ccs_enabled_flag (e.g., first syntax), and can indicate whether the partial encoding of the channel data is used in a video sequence. When the syntax has a first value (e.g., 0), it can indicate that the partial encoding is not used, and when the syntax has a second value (e.g., 1), it can indicate that the partial encoding is used.

[0194] Meanwhile, in Table 1, only an example in which the syntax is included in the sequence parameter set is described for clarity of description, but it is obvious that the syntax can also be included in a picture parameter set (PPS), a picture header, a slice header, etc., and these embodiments are also included in the present disclosure.

[0195] [Table 2]

[0196]

[0197] As an example, Table 2 is a table for explaining an example of syntax according to embodiments of the present disclosure. For example, syntax such as channel_ccs_enabled_flag[channel_idx], channel_ccs_val[channel_idx], cu_ccs_skip_flag[channel_idx][cu_idx], and / or channel_coding_unit(channel_idx, cu_idx) can be signaled. In addition, the syntax can be included in a parameter set of a specific unit (e.g., pic_channel_data) and signaled.

[0198] For example, there can be pic_channel_data, which is an example of a signaling unit in which the syntax disclosed in the table can be included. For example, the unit can be an example of a syntax unit for channel data in one picture of a video sequence. A plurality of channel data can be encoded within the unit.

[0199] As an example, syntax for indicating whether there is a coding unit (CU) to which a region-wise partial coding is applied and thus coding is omitted among coding units for channel data can be signaled. For example, the syntax can be referred to as channel_ccs_enabled_flag[channel_idx] (e.g., second syntax), and the variable channel_idx can be a variable for indicating an index of channel data, and can also be a variable for indicating whether a CU is used for channel data corresponding to a specific channel data index (channel_idx). When the syntax is a first value (e.g., 0), it can indicate that partial coding is not used for any CU in all channel data, and when the syntax is a second value (e.g., 1), it can indicate that there is a CU in which partial coding is used for specific channel data. In addition, the second syntax can be signaled based on the value of the first syntax. For example, the second syntax can be signaled only when the value of the first syntax is a specific value (e.g., 1).

[0200] For example, when there is a coding unit to which partial coding is applied and encoding is omitted, a representative value (e.g., an average value) of channel data for reconstructing the coding unit can be signaled. For example, the syntax can be referred to as channel_ccs_val[channel_idx] (e.g., a third syntax), and since the variable channel_idx is the same as described above, a repeated description is omitted. Thus, the third syntax can indicate a channel data representative value for reconstructing a coding unit in which partial coding is used and encoding is omitted among coding units for channel data corresponding to the channel data index (channel_idx). The third syntax can be signaled based on the value of the second syntax. For example, the third syntax can be signaled only when the value of the second syntax is a specific value (e.g., 1). Also, as an example, the representative value (e.g., an average value) to be encoded or decoded can be a representative value (e.g., an average value) of the entire channel data, but can also be a representative value of channel data of a non-ROI. In this case, the coding unit using partial coding and omitting encoding can be reconstructed based on the representative value of the entire channel data or the representative value of channel data within the non-ROI.

[0201] For example, information for specifying a coding unit to which partial coding is applied can be signaled. For example, the information can be referred to as a syntax cu_ccs_skip_flag[channel_idx][cu_idx] (e.g., a fourth syntax), and since the variable channel_idx is the same as described above, a repeated description is omitted, but the variable cu_idx can be a coding unit index for specifying a coding unit existing in a channel corresponding to the channel data index (channel_idx), and can indicate whether encoding of the coding unit is omitted by a partial coding method. In other words, the fourth syntax can be information indicating whether channel data partial coding is applied to a CU corresponding to an index syntax CU_idx specifying a coding unit existing in a channel corresponding to an index syntax channel_idx specifying channel data. When the value of the fourth syntax is a first value (e.g., 1), it can indicate that partial coding is applied to the CU, in other words, the CU is included in a non-ROI and encoding of the CU is omitted. Thus, on the decoding side, when the CU is reconstructed, a data value of the CU can be derived based on the value of the third syntax channel_ccs_val[channel_idx] described above. Meanwhile, the fourth syntax can be signaled based on the value of the first syntax and / or the second syntax. For example, the fourth syntax can be signaled only when the first syntax is a specific value (e.g., 1) and the second syntax is a specific value (e.g., 1).

[0202] For example, information for specifying a coding unit for which partial coding is not applied can be signaled. For example, the information can be referred to as a syntax channel_coding_unit (channel_idx, cu_idx) (e.g., a fifth syntax), and since the variable channel_idx / cu_idx is the same as described above, the repeated description is omitted. In other words, the fifth syntax can be an example of a coding unit of channel data of a CU corresponding to an index syntax CU_idx specifying a coding unit present in a channel corresponding to an index syntax channel_idx specifying channel data. For example, the fifth syntax can be signaled based on the values of the first syntax, the second syntax, and / or the fourth syntax. For example, when at least one of the values of the first syntax, the second syntax, or the fourth syntax is a specific value (e.g., 0), it can be signaled. As another example, when at least one of the values of the first syntax, the second syntax, or the fourth syntax is a specific value (e.g., 1), it can not be signaled.

[0203] Here, although an encoding unit (CU) is given as an example of a unit of a region for which coding is omitted, in other words, a unit included in a non-ROI in which coding can be omitted, this is only for ease of explanation, and the unit for which coding / decoding is omitted can also be any other unit such as a slice, a tile, a coding tree unit (CTU), a coding unit (CU), a sub-coding unit (sub-CU), or a region composed of a plurality of such units, and the like, and these cases are also obviously included in the embodiments of the present disclosure.

[0204] According to the embodiments of the present disclosure, by not coding the data of the non-ROI, the size of the bitstream can be effectively reduced while still being able to perform a machine task at a level similar to that on the encoding side, thereby improving coding efficiency and reducing overhead. Here, using the average value as a representative value of the entire channel data as a substitute value for the entire data of the non-ROI when coding / decoding the channel data is only for ease of explanation, and in some cases, another value can be coded / decoded and used, or a predetermined value can be replaced without separate coding / decoding, and these cases are also obviously included in the embodiments of the present disclosure.

[0205] Figure 16 is an example illustrating a region-by-region partial coding method according to another embodiment of the present disclosure.

[0206] For example, at the encoding side, there can be original channel data 1601 of features / feature maps for an image. In this regard, based on the entire channel data 1601, a region classification 1602 process for classification into ROI and non-ROI can be performed. For example, since a region classified as non-ROI according to the region distribution of the channel data can have a relatively small impact on performing a machine task, encoding of the channel data of the non-ROI can be omitted during encoding, and only the ROI can be encoded. For example, a region of interest (ROI) 1606 can be encoded 1605. In other words, information about the ROI can be encoded into a bitstream 1600. Meanwhile, based on the entire channel data 1601, a representative value (e.g., an average value) of the entire channel data can be derived 1603, and the average value can be encoded 1604 into the bitstream 1600. In addition, unlike Figure 15 the ROI, a representative value (e.g., an average value) of the channel data in the non-ROI can be derived 1607, and a difference value (e.g., the representative value of the channel data in the non-ROI - the representative value of the entire channel data, or the representative value of the entire channel data - the representative value of the channel data in the non-ROI) related to the representative value of the entire channel data can be derived 1608. The difference value can be encoded into the bitstream 1600.

[0207] In this regard, at the decoding side, the encoded bitstream 1600 can be obtained, and the decoding of the channel data can be performed. For example, during the decoding, the ROI can be decoded based on the information on the encoded ROI, but the channel data for the non-ROI can be decoded in a different manner from the ROI. Based on the bitstream, the partial decoding 1610 of the channel data can be performed. For example, first, the ROI can be reconstructed 1611 based on the information on the ROI obtained from the bitstream. Meanwhile, the representative value (e.g., average value) of the channel data obtained from the bitstream can be decoded 1612, and the difference value derived based on the representative value of the channel data can be decoded 1613, and the non-ROI can be reconstructed 1614 based on the representative value of the channel data and the difference value. For example, by adding the representative value and the difference value of the channel data, the representative value applied to the non-ROI can be derived. When the representative value is derived, the same representative value can be applied to the entire non-ROI. Meanwhile, as an example, the representative value (e.g., average value) to be encoded or decoded can be the representative value (e.g., average value) of the entire channel data, but can also be the representative value (e.g., average value) of the channel data within the non-ROI. Subsequently, the entire channel data can be reconstructed 1615 based on the reconstructed ROI and the reconstructed non-ROI. In this case, the coding unit to which the partial coding is used and the encoding is omitted can be reconstructed based on the representative value of the entire channel data or the representative value of the channel data within the non-ROI. For example, when the representative value of the channel data to be decoded is the representative value of the entire channel data, the difference value can be a value derived as the representative value of the channel data within the non-ROI representative value of the entire channel data. On the other hand, as another example, when the representative value of the channel data to be decoded is the representative value of the channel data within the non-ROI, the difference value can be a value derived as the representative value of the entire channel data - the representative value of the channel data within the non-ROI. In this case, the coding unit to which the partial coding is used and the encoding is omitted can be reconstructed based on the representative value of the entire channel data or the representative value of the channel data within the non-ROI, and can also be further reconstructed based on the difference value.

[0208] Meanwhile, Tables 3 and 4 below are examples of necessary information that can be signaled when applying the partial coding / decoding method proposed in the present embodiment. Table 3 is an example of a syntax for signaling whether the partial coding / decoding method of the channel data in the present example is used in a video sequence, and Table 4 is an example of a syntax indicating whether there is a coding unit (CU) to which the partial coding / decoding is applied within each channel, a syntax for signaling a value that can be used to reconstruct the CU to which the encoding is omitted, and a syntax indicating whether each CU is the CU to which the encoding is omitted.

[0209] [Table 3]

[0210]

[0211] As an example, Table 3 is a table for explaining an example of a syntax that can be included in a parameter set (e.g., sequence parameter set) according to an embodiment of the disclosure. For example, as described above, the parameter set (e.g., sequence parameter set SPS) can include information indicating whether channel data is partially coded according to an embodiment of the disclosure. For example, the information can be expressed as a syntax sps_ccs_enabled_flag (e.g., first syntax), and can indicate whether partial coding of channel data is used in a video sequence. In this regard, the same description as explained above with reference to Table 1 can be applied, and thus a repetitive description is omitted.

[0212] [Table 4]

[0213]

[0214] As an example, Table 4 is a table for explaining an example of a syntax according to an embodiment of the disclosure. For example, syntaxes such as channel_ccs_enabled_flag[channel_idx], channel_ccs_val[channel_idx], cu_ccs_skip_flag[channel_idx][cu_idx], cu_ccs_delta_val[channel_idx][cu_idx], and / or channel_coding_unit(channel_idx, cu_idx) can be signaled. In addition, the syntaxes can be included in a parameter set of a specific unit (e.g., pic_channel_data) and signaled.

[0215] For example, there can be pic_channel_data, which is an example of a signaling unit in which the syntaxes disclosed in the table can be included. In this regard, since the same as described above, a repetitive description is omitted.

[0216] As an example, a syntax for indicating whether there is a coding unit (CU) for channel data in which region-local coding is applied and coding is omitted can be signaled. For example, the syntax can be referred to as channel_ccs_enabled_flag[channel_idx] (e.g., second syntax). In this regard, since the same as described above, a repetitive description is omitted.

[0217] For example, when there is a coding unit to which partial coding is applied and encoding is omitted, a representative value (e.g., average value) of channel data for reconstructing the coding unit can be signaled. For example, the syntax can be referred to as channel_ccs_val[channel_idx] (e.g., third syntax). In this regard, since it is the same as described above, a repeated description is omitted.

[0218] For example, information for specifying a coding unit to which partial coding is applied can be signaled. The information can be referred to as a syntax cu_ccs_skip_flag[channel_idx][cu_idx] (e.g., fourth syntax). In this regard, since it is the same as described above, a repeated description is omitted.

[0219] As an example, information about a difference value for reconstructing a coding unit to which partial coding is applied can be signaled. For example, the information can be referred to as a syntax cu_ccs_delta_val[channel_idx][cu_idx] (e.g., sixth syntax). Since the variables channel_idx and cu_idx are the same as described above, a repeated description is omitted. In other words, the sixth syntax can represent a difference value derived based on a representative value (e.g., representative value, average value) of channel data signaled by the third syntax (e.g., channel_ccs_val[channel_idx]) of a CU, when encoding of a specific CU corresponding to CU_idx existing in a channel corresponding to channel_idx is omitted by channel data partial coding. For example, the representative value (e.g., average value) of channel data to be decoded, in other words, the value indicated by the third syntax, can be a representative value (e.g., average value) of the entire channel data, but can also be a representative value (e.g., average value) of channel data of a non-ROI. Accordingly, for example, when the value indicated by the third syntax is a representative value of the entire channel data, the difference value represented by the sixth syntax can be a value derived as a representative value of channel data in a non-ROI - a representative value of the entire channel data. However, as another example, when the value indicated by the third syntax is a representative value of channel data in a non-ROI, the difference value represented by the sixth syntax can be a value derived as a representative value of the entire channel data - a representative value of channel data in a non-ROI. In this case, a coding unit to which partial coding is applied and encoding is omitted can be reconstructed based on the representative value of the entire channel data or the representative value of channel data within a non-ROI. Furthermore, the sixth syntax can be signaled based on the value of the fourth syntax. For example, when the value of the fourth syntax is a specific value (e.g., 1), the sixth syntax can be obtained from a bitstream.

[0220] For example, information for specifying a coding unit for which partial coding is not used can be signaled. For example, the information can be referred to as a syntax channel_coding_unit (channel_idx, cu_idx) (e.g., a fifth syntax). In this regard, since it is the same as described above, a repeated description is omitted.

[0221] Here, although an example of a coding unit (CU) is given as a unit of a region for which coding is omitted, in other words, a unit included in a non-ROI for which coding is omitted, this is merely for convenience of explanation, and the unit for which coding / decoding is omitted can also be any other unit such as a slice, a tile, a coding tree unit (CTU), a coding unit (CU), a sub-coding unit (sub-CU), or a region composed of a plurality of such units, and these cases are also obviously included in the embodiments of the present disclosure.

[0222] According to the embodiments of the present disclosure, by not coding data of a non-ROI, the size of a bitstream can be effectively reduced while still being able to perform a machine task at a level similar to that on the encoding side, thereby improving coding efficiency and reducing overhead. Here, using the average value as a representative value of the entire channel data as a substitute value for the entire data of the non-ROI when coding / decoding channel data is merely for convenience of explanation, in some cases, another value can be coded / decoded and used, or a predetermined value can be replaced without separate coding / decoding, and such cases are also obviously included in the embodiments of the present disclosure.

[0223] Meanwhile, the syntax names, syntax forms (e.g., flags), and signaling order of syntaxes described above with reference to Figure 15 and Figure 16 and Tables 1 to 4 can be changed, and such cases are also obviously included in the present disclosure.

[0224] Figure 17 is a diagram for explaining a feature decoding method that a feature (image) decoding apparatus according to an embodiment of the present disclosure can perform. As an example, Figure 17 may be performed based on the embodiments described above with reference to other diagrams, or can be based on the syntaxes described with reference to other tables.

[0225] According to an example of a feature decoding method performed by a feature decoding apparatus, information about a channel in a feature can be obtained from a bitstream S1710, and the channel can be reconstructed based on the information about the channel S1720. For example, the information about the channel can include information about whether encoding is performed according to a region-wise importance of the channel. Here, whether encoding is performed according to the region-wise importance of the channel can mean the above-described region-wise partial encoding. For example, based on the information about whether encoding indicates that only a ROI of the channel is encoded, a non-ROI can be reconstructed based on a representative value of the channel data. The representative value of the channel data can be an average value of the channel data. For example, the representative value of the channel data can be a representative value (e.g., an average value) of the entire channel data or a representative value (e.g., an average value) of the channel data of the non-ROI. Meanwhile, the representative value of the channel data can be obtained from the bitstream. In other words, it can be explicitly signaled. Further, as an example, based on the information about whether encoding indicates that only the ROI of the channel is encoded, the non-ROI can be reconstructed based on a value obtained by adding a difference value to the representative value of the channel data of the non-ROI or the representative value of the (entire) channel data. Meanwhile, the representative value of the channel data of the non-ROI can be an average value of the channel data of the non-ROI. Further, as an example, the difference value can be obtained from the bitstream. Further, the non-ROI can also be reconstructed based on a value obtained by adding the representative value of the channel data of the non-ROI to the difference value. As another example, the non-ROI can be reconstructed based on a value obtained by adding the representative value of the entire channel data to the difference value.

[0226] According to an example of a feature decoding method performed by a feature decoding apparatus, information about a channel in a feature can be obtained from a bitstream S1710, and the channel can be reconstructed based on the information about the channel S1720. For example, the information about the channel can include information about whether encoding is performed according to a region-wise importance of the channel. Here, whether encoding is performed according to the region-wise importance of the channel can mean the above-described region-wise partial encoding. For example, based on the information about whether encoding indicates that only a ROI of the channel is encoded, a non-ROI can be reconstructed based on a representative value of the channel data. The representative value of the channel data can be an average value of the channel data. For example, the representative value of the channel data can be a representative value (e.g., an average value) of the entire channel data or a representative value (e.g., an average value) of the channel data of the non-ROI. Meanwhile, the representative value of the channel data can be obtained from the bitstream. In other words, it can be explicitly signaled. Further, as an example, based on the information about whether encoding indicates that only the ROI of the channel is encoded, the non-ROI can be reconstructed based on a value obtained by adding a difference value to the representative value of the channel data of the non-ROI or the representative value of the (entire) channel data. Meanwhile, the representative value of the channel data of the non-ROI can be an average value of the channel data of the non-ROI. Further, as an example, the difference value can be obtained from the bitstream. Further, the non-ROI can also be reconstructed based on a value obtained by adding the representative value of the channel data of the non-ROI to the difference value. As another example, the non-ROI can be reconstructed based on a value obtained by adding the representative value of the entire channel data to the difference value. Figure 17 According to an example of a feature decoding method performed by a feature decoding apparatus, information about a channel in a feature can be obtained from a bitstream S1710, and the channel can be reconstructed based on the information about the channel S1720. For example, the information about the channel can include information about whether encoding is performed according to a region-wise importance of the channel. Here, whether encoding is performed according to the region-wise importance of the channel can mean the above-described region-wise partial encoding. For example, based on the information about whether encoding indicates that only a ROI of the channel is encoded, a non-ROI can be reconstructed based on a representative value of the channel data. The representative value of the channel data can be an average value of the channel data. For example, the representative value of the channel data can be a representative value (e.g., an average value) of the entire channel data or a representative value (e.g., an average value) of the channel data of the non-ROI. Meanwhile, the representative value of the channel data can be obtained from the bitstream. In other words, it can be explicitly signaled. Further, as an example, based on the information about whether encoding indicates that only the ROI of the channel is encoded, the non-ROI can be reconstructed based on a value obtained by adding a difference value to the representative value of the channel data of the non-ROI or the representative value of the (entire) channel data. Meanwhile, the representative value of the channel data of the non-ROI can be an average value of the channel data of the non-ROI. Further, as an example, the difference value can be obtained from the bitstream. Further, the non-ROI can also be reconstructed based on a value obtained by adding the representative value of the channel data of the non-ROI to the difference value. As another example, the non-ROI can be reconstructed based on a value obtained by adding the representative value of the entire channel data to the difference value.

[0227] Meanwhile, since Figure 17 is a diagram for explaining an embodiment of the present disclosure, some steps can be changed or the order of some steps can be changed, and some steps can be deleted or some steps can be added. These cases are also obviously included in the present disclosure.

[0228] Figure 18 is a diagram for explaining a feature encoding method that a feature (image) encoding apparatus according to an embodiment of the present disclosure can perform. As an example, Figure 18 can be performed based on the above-described embodiments with reference to other diagrams or can be performed based on the syntax described with reference to other tables.

[0229] As an example, information about a channel in a feature S1810 can be determined, and information about a channel can be encoded into a bitstream S1820. The information about a channel can include information about whether to encode according to the region-wise importance of a channel. In addition, although not shown in the drawing, the step of determining information about a channel in a feature can also include a step of determining whether to encode according to the region-wise importance of a channel in a feature. Here, whether to encode according to the region-wise importance of a channel can mean the above-described region-wise partial encoding. Meanwhile, based on encoding only the ROI of a channel, a representative value of channel data can be encoded into a bitstream. Meanwhile, based on encoding only the ROI of a channel, a difference value between representative values of channel data can be encoded into a bitstream. Meanwhile, based on encoding only the ROI of a channel, a non-ROI can be reconstructed based on a representative value of channel data. The representative value of channel data can be an average value of channel data. For example, the representative value of channel data can be a representative value (e.g., an average value) of the entire channel data or a representative value (e.g., an average value) of channel data of a non-ROI. Meanwhile, the representative value of channel data can be encoded into a bitstream. In other words, it can be explicitly signaled. In addition, as an example, based on encoding only the ROI of a channel, a non-ROI can be reconstructed based on a value obtained by adding a difference value to a representative value of channel data of a non-ROI or a representative value of (entire) channel data. Meanwhile, the representative value of channel data of a non-ROI can be an average value of channel data of a non-ROI. In addition, as an example, the difference value can also be encoded into a bitstream. In addition, a non-ROI can be reconstructed based on a value obtained by adding the representative value of channel data of a non-ROI and the difference value. As another example, a non-ROI can be reconstructed based on a value obtained by adding the representative value of (entire) channel data and the difference value.

[0230] According to Figure 18 an embodiment, a machine task can be efficiently performed by compressing a bitstream, and encoding efficiency can be improved.

[0231] Meanwhile, since Figure 18 is a drawing for explaining an embodiment of the present disclosure, some steps can be changed or the order of some steps can be changed, and some steps can be deleted or some steps can be added. These cases are also obviously included in the present disclosure.

[0232] Meanwhile, although not shown in the drawing, according to an embodiment of the present disclosure, a medium recording a bitstream generated by a feature encoding method can be disclosed. In this case, the feature encoding method can include a step of determining information about a channel in a feature and a step of encoding information about a channel into a bitstream. The rest can be the same as described above, and repeated descriptions are omitted.

[0233] Further, although not shown in the drawings, according to an embodiment of the present disclosure, a method of transmitting a bitstream generated by a feature encoding method can be disclosed. In this case, the feature encoding method can include a step of determining information about a channel in a feature and a step of encoding the information about the channel into the bitstream. The rest can be the same as described above, and a repeated description is omitted.

[0234] For the sake of clarity of description, the names of all the syntax elements described above are arbitrarily assigned, and do not limit the names of the corresponding syntax elements. Further, the first syntax element to the sixth syntax element can be respectively referred to as first information to sixth information. Further, the first syntax element to the sixth syntax element can be obtained from a bitstream, but can also be derived from other syntax elements, and such a case can also be included in an embodiment of the present disclosure.

[0235] Further, the bitstream generated by the feature (image) encoding method can be stored in a non-transitory computer-readable recording medium.

[0236] Further, as another example, the bitstream generated by the feature (image) encoding method can be transmitted to another device (e.g., a feature (image) decoding device, etc.). In this case, the method of transmitting the bitstream can include a process of transmitting the bitstream.

[0237] For the sake of clarity of description, the names of all the syntax elements described above are arbitrarily assigned, and do not limit the names of the corresponding syntax elements. Further, the first syntax element to the sixth syntax element can be respectively referred to as first information to sixth information. Further, the first syntax element to the sixth syntax element can be obtained from a bitstream, but can also be derived from other syntax elements, and such a case can also be included in an embodiment of the present disclosure.

[0238] Further, the bitstream generated by the image encoding method can be stored in a non-transitory computer-readable recording medium.

[0239] Further, as another example, the bitstream generated by the image encoding method can be transmitted to another device (e.g., a feature (image) decoding device, etc.). In this case, the method of transmitting the bitstream can include a process of transmitting the bitstream.

[0240] Although the exemplary methods of the present disclosure are expressed as a series of acts for the sake of clarity of explanation, this is not intended to limit the order of execution of the steps, and each step can be executed simultaneously or in a different order if necessary. In order to implement the method according to the present disclosure, another step can be additionally included in the exemplary steps, or the remaining steps can be included excluding some steps, or another additional step can be included excluding some steps.

[0241] In the disclosure, the image encoding apparatus or the image decoding apparatus that performs a predetermined operation (step) can perform an operation (step) for checking a condition or a condition for performing a corresponding operation (step). For example, when it is stated that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding apparatus or the image decoding apparatus can perform an operation for checking whether the predetermined condition is satisfied, and then perform the predetermined operation.

[0242] Various embodiments of the disclosure do not list all possible combinations, but are intended to describe representative aspects of the disclosure, and matters described in various embodiments can be applied independently or in a combination of at least two.

[0243] Embodiments described in the disclosure can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information (for example, information on instructions) or an algorithm for implementation can be stored in a digital storage medium.

[0244] In addition, the decoders (decoding apparatuses) and encoders (encoding apparatuses) to which embodiments of the disclosure are applied can be included in a multimedia broadcast transmitting and receiving device, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video communication device, a real-time communication device such as video communication, etc., a mobile streaming device, a storage medium, a camcorder, a video on demand (VoD) service providing device, an OTT video (over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video phone video device, a transportation terminal (for example, a vehicle (including an autonomous vehicle) terminal, a robot terminal, an airplane terminal, a ship terminal, etc.), a medical video device, etc., and can be used to process a video signal or a data signal. For example, the OTT video (over the top video) device can include a game console, a Blu-ray player, an Internet-connected television, a home theater system, a smartphone, a tablet, a digital video recorder (DVR), etc.

[0245] In addition, a processing method to which the embodiments of the disclosure are applied can be generated in the form of a program that is executed by a computer, and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of the disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distribution devices that store computer-readable data. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device. In addition, the computer-readable recording medium includes a medium that is realized in a carrier wave form (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted through a wired or wireless communication network.

[0246] In addition, the embodiments of the disclosure can be implemented as a computer program product by program code, and the program code can be executed on a computer by the embodiments of the disclosure. The program code can be stored on a computer-readable carrier.

[0247] Figure 19 FIG. 1 is a diagram illustrating an example of a content streaming system to which the embodiments of the disclosure can be applied.

[0248] Referring to Figure 19 The content streaming system to which the embodiments of the disclosure are applied can widely include an encoding server, a streaming server, a web server, a media store, a user device, and a multimedia input device.

[0249] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data, generates a bitstream, and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server can be omitted.

[0250] The bitstream can be generated by an image encoding method and / or an image encoding apparatus to which the embodiments of the disclosure are applied, and the streaming server can temporarily store the bitstream during a process of transmitting or receiving the bitstream.

[0251] The streaming server can transmit multimedia data to the user device based on a user request through the web server, and the web server can act as a medium that informs the user of the available services. When the user requests a desired service from the web server, the web server can transmit the request to the streaming server, and the streaming server can transmit the multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server can function to control commands / responses between devices within the content streaming system.

[0252] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a seamless streaming service, the streaming server can store a bitstream for a certain period of time.

[0253] Examples of the user device can include a mobile phone, a smartphone, a notebook computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a tablet PC, a tablet, an ultrabook, a wearable device (i.e., a smart watch, smart glasses, a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc.

[0254] Each server in the content streaming system can operate as a distributed server, in which case data received by each server can be processed in a distributed manner.

[0255] Figure 20 FIG. 2 is a diagram illustrating another example of a content streaming system to which embodiments of the disclosure can be applied.

[0256] Referring to Figure 20 In embodiments such as the VCM, a task can be executed by the user terminal, or the task can be executed by an external device (e.g., a streaming server, an analysis server, etc.) according to the performance of the device, a user's request, the characteristics of the task to be executed, etc. In this way, in order to transmit information necessary to execute the task to the external device, the user terminal can generate a bitstream including information necessary to execute the task (e.g., information such as a task, a neural network, and / or usage) directly or through the encoding server.

[0257] The analysis server can perform a task requested by the user after decoding the encoded information transmitted from the user terminal (or from the encoding server). The analysis server can transmit a result obtained by performing the task to the user terminal or another linked service server (e.g., a web server) again. For example, the analysis server can transmit a result obtained by performing a task for determining a fire to a fire-related server. The analysis server can include a separate control server, in which case the control server can play a role in controlling commands / responses between each device associated with the analysis server and the server. In addition, the analysis server can request desired information from the web server based on information about a task that the user device wants to perform and a task that the user device can perform. When the analysis server requests the desired service from the web server, the web server can transmit it to the analysis server, and the analysis server can transmit its data to the user terminal. In this case, the control server of the content streaming system can play a role in controlling commands / responses between each device within the streaming system.

[0258] [INDUSTRIAL APPLICABILITY]

[0259] Embodiments according to the present disclosure can be used to encode / decode an image.

Claims

1. A feature decoding method performed by a feature decoding apparatus, comprising: obtaining information about a channel in a feature from a bitstream; and reconstructing the channel based on the information about the channel, wherein the information about the channel includes information about whether or not the channel is encoded according to a region-wise importance of the channel.

2. The feature decoding method of claim 1, wherein, based on the information about whether or not the channel is encoded indicating that only a region of interest of the channel is encoded, reconstructing a region of non-interest based on a representative value of channel data.

3. The feature decoding method of claim 2, wherein, the representative value of the channel data is an average value of the channel data.

4. The feature decoding method of claim 2, wherein, obtaining the representative value of the channel data from the bitstream.

5. The feature decoding method of claim 2, wherein, based on the information about whether or not the channel is encoded indicating that only the region of interest of the channel is encoded, further reconstructing the region of non-interest based on a difference between a representative value of channel data of the region of non-interest and the representative value of the channel data.

6. The feature decoding method of claim 5, wherein, the representative value of the channel data of the region of non-interest is an average value of the channel data of the region of non-interest.

7. The feature decoding method of claim 5, wherein, obtaining the difference from the bitstream.

8. The feature decoding method of claim 5, wherein, reconstructing the region of non-interest based on a value obtained by adding the difference to the representative value of the channel data of the region of non-interest or the representative value of the channel data.

9. A feature encoding method performed by a feature encoding apparatus, comprising: determining information about a channel in a feature; and encoding the information about the channel into a bitstream, wherein the information about the channel includes information about whether or not the channel is encoded according to a region-wise importance of the channel. based on only a region of interest of the channel being encoded, a representative value of channel data is encoded into the bitstream.

10. The feature encoding method of claim 9, wherein, based on only the region of interest of the channel being encoded, a difference between the representative values of channel data is encoded into the bitstream.

11. The feature encoding method of claim 10, wherein, the feature encoding method comprising:

12. A medium recording a bitstream generated by a feature encoding method, wherein, determining information about a channel in a feature; and encoding the information about the channel into a bitstream, wherein the information about the channel includes information about whether or not the channel is encoded according to a region-wise importance of the channel.

13. A method of transmitting a bitstream, comprising: transmitting a bitstream generated by a feature encoding method, wherein the feature encoding method comprises: determining information about a channel in a feature; and encoding the information about the channel into a bitstream, wherein the information about the channel includes information about whether or not the channel is encoded according to a region-wise importance of the channel. ​