Image encoding / decoding method and apparatus using optimized quantization according to image, and method of transmitting bitstream
By distinguishing between image data perceived by humans and machines, and adopting quantization optimization methods and machine video coding bitstreams, the problem that image compression in existing technologies is not suitable for artificial intelligence services is solved, and the encoding/decoding efficiency and machine perception performance are improved.
Patent Information
- Application Number
- CN202480014275.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-23
- Filing Date
- 2024-02-23
- Publication Date
- 2025-09-19
AI Technical Summary
Existing image compression technology has not been effectively applied to artificial intelligence services, resulting in low encoding/decoding efficiency and inability to meet the needs of machine tasks.
Through a quantization optimization method based on image usage, image data perceived by humans and machines is distinguished, machine video coding (VCM) bitstream is used for encoding/decoding, and a quantization method suitable for machine tasks is generated.
It improves the efficiency of image encoding/decoding, implements quantization optimization based on image usage, improves machine perception performance, and supports sending and storing bitstreams.
Smart Images

Figure CN120677698A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus that performs quantization optimization according to image usage, and a method of transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. Background Art
[0002] With the development of machine learning technology, the demand for AI services based on image processing is increasing. To efficiently process the large amounts of image data required for AI services within limited resources, optimized image compression technologies for machine tasks are essential. However, because existing image compression technologies were developed with the goal of processing high-resolution and high-quality images for human vision, they are not suitable for AI services. Therefore, research and development of new machine-oriented image compression technologies suitable for AI services is actively underway. Summary of the Invention
[0003] Technical issues
[0004] The present disclosure provides an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] The present disclosure provides an image encoding / decoding method and apparatus for performing quantization optimization based on image usage.
[0006] The present disclosure provides a method and apparatus for encoding / decoding an image by selecting a quantization method suitable for a machine task.
[0007] The present disclosure provides a method and apparatus for encoding / decoding an image based on image usage and a video coding machine (VCM) bit stream.
[0008] The present disclosure provides a method for defining attributes of necessary information according to machine task types.
[0009] The present disclosure provides a quantitative optimization method for improving machine perception performance.
[0010] The present disclosure provides a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0011] The present disclosure provides a recording medium storing a bit stream generated by the image encoding method or apparatus according to the present disclosure.
[0012] The present disclosure provides a recording medium storing a bit stream received and decoded by an image decoding apparatus according to the present disclosure and used to reconstruct an image.
[0013] Technical problems to be achieved in the present disclosure are not limited to the above-mentioned technical problems, and those having ordinary skill in the art can clearly understand other technical problems that are not described through the following description.
[0014] Technical Solution
[0015] According to one aspect of the present disclosure, an image data decoding method performed by an image data decoding device may include: obtaining image usage information for image data from a bit stream; and reconstructing the image data by performing inverse quantization based on the image usage information, wherein the image usage information may indicate whether the image data is for human perception or for machine perception.
[0016] According to one aspect of the present disclosure, an image data encoding method performed by an image data encoding device may include: performing quantization on image data by determining an image usage for the image data; and encoding image usage information indicating the image usage, wherein the image usage information may indicate whether the image data is for human perception or for machine perception.
[0017] A recording medium according to another aspect of the present disclosure may store a bit stream generated by the image encoding method or the image encoding device of the present disclosure.
[0018] A bitstream transmission method according to another aspect of the present disclosure may transmit a bitstream generated by the image encoding method or the image encoding device of the present disclosure to a decoding device.
[0019] The features briefly summarized above for the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.
[0020] Beneficial effects
[0021] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0022] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus that perform optimized quantization based on image usage.
[0023] According to the present disclosure, a method and apparatus for defining attributes of necessary information according to image usage can be provided.
[0024] According to the present disclosure, a method and apparatus for encoding / decoding an image may be provided, wherein image encoding performance is improved to achieve machine perception.
[0025] According to the present disclosure, a video coding machine (VCM) bitstream may be generated according to image usage.
[0026] According to the present invention, a method and apparatus for encoding / decoding an image by selecting a quantization method suitable for a machine task can be provided.
[0027] According to the present disclosure, a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure may be provided.
[0028] According to the present disclosure, a recording medium storing a bit stream generated by the image encoding method or apparatus according to the present disclosure can be provided.
[0029] According to the present disclosure, there can be provided a recording medium that stores a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used to reconstruct an image.
[0030] Effects obtainable from the present disclosure are not limited to the above-described effects, and other effects that are not described can be clearly understood from the following description by those having ordinary skill in the art. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a diagram schematically illustrating a VCM system to which embodiments of the present disclosure can be applied.
[0032] Figure 2 FIG2 is a diagram schematically showing a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0033] Figure 3 is a diagram schematically illustrating an image / video encoder to which embodiments of the present disclosure can be applied.
[0034] Figure 4 FIG2 is a diagram schematically illustrating an image / video decoder to which an embodiment of the present disclosure can be applied.
[0035] Figure 5 is a flow chart schematically illustrating a feature / feature map encoding process to which embodiments of the present disclosure may be applied.
[0036] Figure 6 is a flow chart schematically illustrating a feature / feature map decoding process to which embodiments of the present disclosure may be applied.
[0037] Figure 7 is a diagram illustrating an example of a feature extraction and reconstruction method to which an embodiment of the present disclosure can be applied.
[0038] Figure 8 A diagram illustrating an example of an image partitioning method to which an embodiment of the present disclosure can be applied.
[0039] Figure 9 and Figure 10 is a diagram illustrating an example of an image encoding / decoding system including a VCM image encoder and decoder.
[0040] Figure 11 is a diagram illustrating an example of a VCM hierarchical structure.
[0041] Figure 12 is a diagram illustrating an example of a VCM bitstream composed of encoded abstract features and NNAL information.
[0042] Figure 13 is a diagram illustrating an example of a quantization group to which an embodiment of the present disclosure can be applied.
[0043] Figure 14 is a diagram illustrating examples of images having different attributes according to an embodiment of the present disclosure.
[0044] Figure 15 is a diagram illustrating an example of a frequency sensitivity determination criterion according to an embodiment of the present disclosure.
[0045] Figure 16 is a diagram illustrating an example of the operation of an image encoder or decoder based on a frequency sensitivity determination standard according to an embodiment of the present disclosure.
[0046] Figure 17 is a diagram illustrating an example of the operation of an image encoder or decoder based on a frequency sensitivity determination criterion according to another embodiment of the present disclosure.
[0047] Figure 18 is a diagram illustrating an example of the operation of an image encoder or decoder based on a frequency sensitivity determination criterion according to another embodiment of the present disclosure.
[0048] Figure 19 is a diagram illustrating an example of defining frequency sensitivity levels according to an embodiment of the present disclosure.
[0049] Figure 20 is a diagram illustrating an image data decoding method according to an embodiment of the present disclosure.
[0050] Figure 21 is a diagram illustrating an image data encoding method according to an embodiment of the present disclosure.
[0051] Figure 22 is a diagram illustrating an example of a content streaming system to which an embodiment of the present disclosure can be applied.
[0052] Figure 23 is a diagram illustrating another example of a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION
[0053] Hereinafter, in order for those skilled in the art to easily implement them, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. However, the present disclosure can be implemented in various forms and is not limited to the embodiments described herein.
[0054] When describing the embodiments of the present disclosure, when well-known configurations or functions are considered to obscure the main points of the present disclosure, their detailed explanation is omitted. In addition, parts that are not related to the description of the present disclosure are omitted from the accompanying drawings, and similar reference numerals have been assigned to similar parts.
[0055] In the present disclosure, when a component is described as being “connected,” “coupled,” or “linked” to another component, this may include not only a direct connection but also an indirect connection with another component interposed therebetween. In addition, when a component is described as “including” or “having” another component, this means that, unless explicitly stated otherwise, it does not exclude other components but may further include additional components.
[0056] In this disclosure, unless otherwise expressly stated, the terms first, second, etc. are used only to distinguish one component from another and do not limit the order or importance of the components. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0057] In this disclosure, distinguishable components are described to clearly explain their respective characteristics and do not necessarily mean that the components are separate. In other words, multiple components can be integrated into a single hardware or software unit, or a single component can be distributed across multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure without explicitly describing them.
[0058] In the present disclosure, the components described in the various embodiments do not necessarily mean required components, and some components may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of the present disclosure. In addition, embodiments that include additional components in addition to the components described in the various embodiments are also included in the scope of the present disclosure.
[0059] The present disclosure relates to encoding and decoding of images, and terms used herein may have ordinary meanings commonly used in the technical field to which the present disclosure belongs, unless these terms are newly defined in the present disclosure.
[0060] The present disclosure may be applied to methods disclosed in the Versatile Video Coding (VVC) standard and / or the Video Coding for Machines (VCM) standard. Furthermore, the present disclosure may be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).
[0061] This disclosure presents various embodiments related to video / image coding, and unless otherwise specified, the embodiments may be implemented in combination with one another. In this disclosure, "video" may refer to a collection of images sequentially recorded over time. An "image" may be information generated by artificial intelligence (AI). Input information used in the process of AI performing a series of tasks, information generated during information processing, and output information may be used as images. In this disclosure, a "picture" generally refers to a unit representing a single image at a specific point in time, and a slice / tile is a coding unit that constitutes a portion of a picture. A picture may consist of at least one slice / tile. Furthermore, a slice / tile may include at least one coding tree unit (CTU). A CTU may be partitioned into at least one unit. A tile is a rectangular area within a specific tile row and tile column within a picture and may be composed of multiple CTUs. A tile column may be defined as a rectangular area of a CTU and may have the same height as the picture and a width specified by syntax elements signaled in the bitstream, such as a picture parameter set. A tile row can be defined as a rectangular area of a CTU and can have a width equal to the width of the picture and a height specified by a syntax element signaled from a bitstream portion such as a picture parameter set. Tile scan is a predetermined sequential ordering method for CTUs in a picture partition. Here, CTUs can be sequentially ordered according to a CTU raster scan within a tile, and tiles within a picture can be sequentially ordered according to a raster scan order of tiles in the picture. A slice can include an integer number of complete tiles or an integer number of sequential complete CTU rows within a tile of the picture. A slice can be exclusively included in a single NAL unit. A picture can be composed of at least one tile group. A tile group can include at least one tile. A brick can represent a rectangular area of a CTU row within a tile in a picture. A tile can include at least one brick. A brick can indicate a rectangular area of a CTU row within a tile. A tile can be partitioned into multiple tiles, and each tile can include at least one CTU row belonging to the tile. Tiles that are not partitioned into multiple tiles can also be considered tiles.
[0062] In this disclosure, "pixel" or "picture element (PEL)" may refer to the smallest unit constituting a picture (or image). Furthermore, the term "sample" may be used as a corresponding term for a pixel. A sample may generally represent a pixel or a pixel value, and may indicate only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.
[0063] In embodiments, particularly when applied to VCM, when there is an image composed of an integration of components with different characteristics and meanings, the pixel / pixel value may represent the pixel / pixel value of the component generated by combining, synthesizing, and analyzing independent information of each component. For example, in RGB input, it may represent only the pixel / pixel value of R, only the pixel / pixel value of G, or only the pixel / pixel value of B. For example, it may represent only the pixel / pixel value of the luminance component synthesized by using the R, G, and B components. For example, it may represent only the pixel / pixel value of the information or image extracted by analyzing the R, G, and B components.
[0064] In the present disclosure, "unit" may refer to a basic unit of image processing. A unit may include a specific area of a picture or at least one of information related to the area. One unit may include a luminance block and two chrominance (e.g., Cb, Cr) blocks. Depending on the context, the term "unit" may be used interchangeably with "sample array," "block," "area," and the like. Typically, an M×N block may include a set (or array) of samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows. In an embodiment, in particular, when it is applied to VCM, a unit may represent a basic unit that includes information for performing a specific task.
[0065] In the present disclosure, the term "current block" may refer to one of a "current coding block," a "current coding unit," a "coding target block," a "decoding target block," or a "processing target block." When prediction is performed, the "current block" may refer to a "current prediction block" or a "prediction target block." When transform (inverse transform) / quantization (dequantization) is performed, the "current block" may refer to a "current transform block" or a "transform target block." When filtering is performed, the "current block" may refer to a "filtering target block."
[0066] In addition, in the present disclosure, unless explicitly stated as a chroma block, "current block" may refer to "luminance block of the current block." "Chroma block of the current block" may be expressed by explicitly including an explicit description of the chroma block such as "chroma block" or "current chroma block."
[0067] In the present disclosure, " / " and "," may refer to "and / or". For example, "A / B" and "A, B" may refer to "A and / or B". In addition, "A / B / C" and "A, B, C" may refer to "at least one of A, B, and / or C".
[0068] In the present disclosure, "or" may mean "and / or". For example, "A or B" may mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" may also mean "in addition or alternatively".
[0069] The present disclosure relates to video / image coding for machines (VCM).
[0070] VCM refers to a compression technique that encodes / decodes a portion of a source image / video or information obtained from it for machine vision purposes. In VCM, the encoding / decoding target can be referred to as a feature. Features can refer to information extracted from the source image / video based on task objectives, requirements, and the surrounding environment. Features can have different information formats than the source image / video, and therefore, the feature compression method and representation format can also differ from the video source.
[0071] VCMs can be applied in a variety of applications. For example, in monitoring systems that identify and track objects or people, VCMs can be used to store or transmit object identification information. Furthermore, in intelligent transportation or smart traffic systems, VCMs can be used to transmit vehicle location information collected from GPS, sensor information collected from LIDAR, radar, and other sensors, and various vehicle control information to other vehicles or infrastructure. Furthermore, in the smart city sector, VCMs can be used to execute individual tasks for interconnected sensor nodes or devices.
[0072] The present disclosure provides various embodiments regarding feature / feature map encoding. Unless otherwise specified, the embodiments of the present disclosure can be implemented individually or in combination of at least two.
[0073] Overview of VCM System
[0074] Figure 1 is a diagram schematically illustrating a VCM system to which embodiments of the present disclosure can be applied.
[0075] refer to Figure 1 , the VCM system may include an encoding device 10 and a decoding device 20.
[0076] The encoding device 10 can compress / encode features / feature maps extracted from the source image / video to generate a bitstream, and transmit the generated bitstream to the decoding device 20 via a storage medium or a network. The encoding device 10 may also be referred to as a feature encoding device. In a VCM system, features / feature maps may be generated in each hidden layer of a neural network. The size and number of channels of the generated feature map may vary depending on the type of neural network or the position of the hidden layer. In the present disclosure, a feature map may be referred to as a feature set, and a feature or feature map may be referred to as "feature information."
[0077] The encoding device 10 may include a feature obtaining unit 11 , an encoder 12 , and a transmitter 13 .
[0078] The feature acquisition unit 11 can obtain features or feature maps for the source image / video. Depending on the embodiment, the feature acquisition unit 11 can obtain the features or feature maps from an external device (e.g., a feature extraction network). In this case, the feature acquisition unit 11 performs the function of a feature receiving interface. Alternatively, the feature acquisition unit 11 can obtain the features or feature maps by executing a neural network (e.g., a CNN, a DNN, etc.) using the source image / video as input. In this case, the feature acquisition unit 11 performs the function of a feature extraction network.
[0079] Depending on the embodiment, the encoding device 10 may further include a source image generator (not shown) for obtaining a source image / video, or may include it in place of the feature obtaining unit 11. The source image generator may be implemented using an image sensor, a camera module, etc., and may obtain the source image / video through a process of capturing, synthesizing, or generating an image / video. In this case, the generated source image / video may be sent to a feature extraction network and used as input data for extracting features / feature maps.
[0080] The encoder 12 may encode the features / feature maps obtained by the feature acquisition unit 11. The encoder 12 may perform a series of processes, such as prediction, transformation, and quantization, to increase encoding efficiency. The encoded data (encoded feature / feature map information) may be output in the form of a bitstream. The bitstream containing the encoded feature / feature map information may be referred to as a VCM bitstream.
[0081] The transmitter 13 can obtain the feature / feature map information or data output in the form of a bit stream, and can transmit the obtained information or data in the form of a file or streaming to the decoding device 20 or another external object via a digital storage medium or a network. Here, the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 may include an element for generating a media file having a predetermined file format, or an element for transmitting data via a broadcast / communication network. The transmitter 13 may be provided as a transmission device separate from the encoder 12, and in this case, the transmission device may include at least one processor for obtaining the feature / feature map information or data output in the form of a bit stream and a transmitter for transmitting it in the form of a file or streaming.
[0082] The decoding device 20 may obtain feature / feature map information from the encoding device 10 and reconstruct features / feature maps based on the obtained information.
[0083] The decoding device 20 may include a receiver 21 and a decoder 22 .
[0084] The receiver 21 may receive a bitstream from the encoding device 10 and obtain feature / feature map information from the received bitstream to transmit it to the decoder 22 .
[0085] The decoder 22 may decode the features / feature maps based on the obtained feature / feature map information. The decoder 22 may perform a series of processes corresponding to the operations of the encoder 14, such as inverse quantization, inverse transform, prediction, etc., to increase decoding efficiency.
[0086] According to an embodiment, the decoding device 20 may further include a task analysis / rendering unit 23 .
[0087] The task analysis / rendering unit 23 can perform task analysis based on the decoded features / feature maps. Furthermore, the task analysis / rendering unit 23 can render the decoded features / feature maps into a format suitable for performing the task. Based on the task analysis results and the rendered features / feature maps, various machine-oriented tasks can be performed.
[0088] Therefore, VCM systems can encode / decode features extracted from source images / videos based on user and / or machine requests, task objectives, and the surrounding environment, and perform various machine-oriented tasks based on the decoded features. VCM systems can also be implemented by extending / redesigning video / image coding systems and can implement various encoding / decoding methods defined in the VCM standard.
[0089] VCM pipeline
[0090] Figure 2 FIG2 is a diagram schematically showing a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0091] refer to Figure 2 , the VCM pipeline 200 may include a first pipeline 210 for encoding / decoding images / videos and a second pipeline 220 for encoding / decoding features / feature maps. In the present disclosure, the first pipeline 210 may be referred to as a video codec pipeline, and the second pipeline 220 may be referred to as a feature codec pipeline.
[0092] The first pipeline 210 may include a first stage 211 for encoding an input image / video and a second stage 212 for decoding the encoded image / video to generate a reconstructed image / video. The reconstructed image / video may be used for human viewing, ie, human vision.
[0093] The second pipeline 220 may include a third stage 221 for extracting features / feature maps from an input image / video, a fourth stage 222 for encoding the extracted features / feature maps, and a fifth stage 223 for decoding the encoded features / feature maps to generate reconstructed features / feature maps. The reconstructed features / feature maps may be used for machine (vision) tasks. Here, machine (vision) tasks may refer to tasks in which a machine consumes images / videos. Machine vision tasks may be applied to service scenarios such as, for example, surveillance, intelligent transportation, smart cities, intelligent industry, intelligent content, etc. According to an embodiment, the reconstructed features / feature maps may also be used for human vision.
[0094] According to an embodiment, the features / feature maps encoded in the fourth stage 222 may be sent to the first stage 221 and used to encode the image / video. In this case, an additional bitstream may be generated based on the encoded features / feature maps, and the generated additional bitstream may be sent to the second stage 222 and used to decode the image / video.
[0095] According to an embodiment, the features / feature maps decoded in the fifth stage 223 may be sent to the second stage 222 and used to decode the image / video.
[0096] although Figure 2 The VCM pipeline 200 is shown to include a first pipeline 210 and a second pipeline 220, but this is merely exemplary and the embodiments of the present disclosure are not limited thereto. For example, the VCM pipeline 200 may include only the second pipeline 220, or the second pipeline 220 may be extended to multiple feature codec pipelines.
[0097] Meanwhile, in the first pipeline 210, the first stage 211 may be performed by the image / video encoder, and the second stage 212 may be performed by the image / video decoder. Furthermore, in the second pipeline 220, the third stage 221 may be performed by the VCM encoder (or feature / feature map encoder), and the fourth stage 222 may be performed by the VCM decoder (or feature / feature map decoder). The encoder / decoder structure is described in detail below.
[0098] Encoder
[0099] Figure 3 is a diagram schematically illustrating an image / video encoder to which embodiments of the present disclosure can be applied.
[0100] refer to Figure 3 The image / video encoder 300 may include an image partitioner 310, a predictor 320, a residual processor 330, an entropy encoder 340, an adder 350, a filter 360, and a memory 370. The predictor 320 may include an inter-frame predictor 321 and an intra-frame predictor 322. The residual processor 330 may include a transformer 332, a quantizer 333, an inverse quantizer 334, and an inverse transformer 335. The residual processor 330 may further include a subtractor 331. The adder 350 may be referred to as a reconstructor or a reconstructed block generator. Depending on the embodiment, the image partitioner 310, the predictor 320, the residual processor 330, the entropy encoder 340, the adder 350, and the filter 360 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Furthermore, the memory 370 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware components may further include the memory 370 as an internal / external component.
[0101] The image partitioner 310 can partition the input image (or picture, or frame) input to the image / video encoder 300 into at least one processing unit. For example, a processing unit can be referred to as a coding unit (CU). Coding units can be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU, Lunit) based on a quadtree, binary tree, and / or ternary tree (QTBTTT) structure. For example, a coding unit can be partitioned into multiple coding units of a greater depth based on a quadtree, binary tree, and / or ternary structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree and / or ternary structure. Alternatively, a binary tree structure can be applied first. The image / video encoding process according to the present disclosure can be performed based on a final coding unit that is no longer partitioned. In this case, the largest coding unit can be used as the final coding unit based on factors such as coding efficiency according to image characteristics. Alternatively, if necessary, a coding unit can be recursively partitioned into coding units of a greater depth, with the optimally sized coding unit being used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit described above. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0102] In some cases, the term "unit" may be used interchangeably with terms such as "block," "region," and the like. In general, an MxN block may represent a set of transform coefficients or samples consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, or may represent only a pixel / pixel value of a luma component, or may represent only a pixel / pixel value of a chroma component. A sample may be used as a term corresponding to a pixel or a picture element.
[0103] The image / video encoder 300 can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 321 or the intra-frame predictor 322 from the input image signal (original block, original sample array). The generated residual signal is then sent to the transformer 332. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the image / video encoder 300 can be referred to as a subtractor 331. The predictor can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including prediction samples for the current block. The predictor can determine whether intra-frame prediction or inter-frame prediction is to be applied in units of the current block or unit. The predictor can generate various information related to prediction, such as prediction mode information, and send it to the entropy encoder 340. The information related to prediction can be encoded by the entropy encoder 340 and output in the form of a bitstream.
[0104] The intra-frame predictor 322 can predict the current block by referring to samples within the current picture. In this case, the reference samples can be located in a neighboring area of the current block, or can be located farther away depending on the prediction mode. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. The directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes based on the granularity of the prediction direction. However, this is an example, and a greater or lesser number of directional prediction modes can be used depending on the configuration. The intra-frame predictor 322 can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0105] The inter-frame predictor 321 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks can include spatially neighboring blocks within the current picture and temporally neighboring blocks within a reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks or collocated units (colUnits), and the reference picture containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 321 can construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 321 can use the motion information of neighboring blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.
[0106] The predictor 320 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction for a block, and can also apply both intra and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Furthermore, the predictor can use intra block copy (IBC) prediction mode or palette mode for block prediction. IBC prediction mode or palette mode can be used for content image / video coding, such as screen content coding (SCC). IBC essentially performs prediction within the current picture, but because it derives reference blocks within the current picture, it can operate similarly to inter prediction. In other words, IBC can use at least one of the inter prediction methods described in this disclosure. Palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values within the picture can be signaled based on information related to a palette table and palette index.
[0107] The prediction signal generated by the predictor 320 can be used to generate a reconstructed signal or a residual signal. The transformer 332 can generate transform coefficients by applying a transform method to the residual signal. For example, the transform method may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. In addition, the transform process can be applied to pixel blocks of the same square size or non-square, variable-sized blocks.
[0108] The quantizer 333 quantizes the transform coefficients and sends them to the entropy encoder 340. The entropy encoder 340 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. This information about the quantized transform coefficients can be referred to as residual information. The quantizer 333 can reorder the block-based quantized transform coefficients in the form of a one-dimensional vector based on the coefficient scanning order and generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of a one-dimensional vector. The entropy encoder 340 can implement various encoding methods, such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC). The entropy encoder 340 can encode not only the quantized transform coefficients but also information necessary for video / image reconstruction (e.g., syntax element values) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in a network abstraction layer (NAL) unit in the form of a bitstream. The image / video information may further include information about various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may further include general constraint information. Furthermore, the image / video information may further include methods for generating and using the coding information, its purpose, and the like. In the present disclosure, information and / or syntax elements transmitted / signaled from the image / video encoder to the image / video decoder may be included in the image / video information. The image / video information may be encoded using the encoding process described above and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. A transmitter (not shown) for transmission and / or a storage unit (not shown) for storing the signal output from the entropy encoder 340 may be configured as internal / external components of the image / video encoder 300, or the transmitter may be included in the entropy encoder 340.
[0109] The quantized transform coefficients output from the quantizer 333 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients through the inverse quantizer 334 and the inverse transformer 335. The adder 350 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-frame predictor 321 or the intra-frame predictor 322. When there is no residual for the processing target block, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 350 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current picture, and as described later, can also be used for inter-frame prediction of the next picture through filtering.
[0110] Meanwhile, luma mapping with chroma scaling can be applied during picture encoding and / or reconstruction.
[0111] The filter 360 can apply filtering to the reconstructed signal to enhance the subjective / objective image quality. For example, the filter 360 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture can be stored in the memory 370, specifically in the DPB of the memory 370. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 360 can generate various filter-related information and send it to the entropy encoder 340. The filter-related information can be encoded by the entropy encoder 340 and output in the form of a bitstream.
[0112] The modified reconstructed picture sent to the memory 370 may be used as a reference picture in the inter-frame predictor 321. Through this, prediction mismatch on the encoder side and the decoder side may be avoided, and encoding efficiency may be improved.
[0113] The DPB of the memory 370 may store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 321. The memory 370 may store motion information of the block from which motion information within the current picture and / or motion information of blocks within the reconstructed picture is derived (or encoded). The stored motion information may be sent to the inter-frame predictor 321 for use as motion information for spatially or temporally neighboring blocks. The memory 370 may store reconstructed samples of the reconstructed block in the current picture and send the stored reconstructed samples to the intra-frame predictor 322.
[0114] At the same time, the VCM encoder (or feature / feature map encoder) can have the same Figure 3The image / video encoder 300 described above has a similar / similar structure in that it performs a series of processes such as prediction, transform, quantization, etc. to encode features / feature maps. However, the VCM encoder differs from the image / video encoder 300 in that it targets features / feature maps for encoding, and therefore, may differ in the name of each unit (or component) (e.g., image partitioner 310, etc.) and its specific operating details from the image / video encoder 300. Specific operating details of the VCM encoder will be described in detail later.
[0115] Decoder
[0116] Figure 4 FIG2 is a diagram schematically illustrating an image / video decoder to which an embodiment of the present disclosure can be applied.
[0117] refer to Figure 4 , the image / video decoder 400 may include an entropy decoder 410, a residual processor 420, a predictor 430, an adder 440, a filter 450, and a memory 460. The predictor 430 may include an inter-frame predictor 431 and an intra-frame predictor 432. The residual processor 420 may include an inverse quantizer 421 and an inverse transformer 422. According to an embodiment, the entropy decoder 410, the residual processor 420, the predictor 430, the adder 440, and the filter 450 may be configured by a hardware component (e.g., a decoder chipset or processor). In addition, the memory 460 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 460 as an internal / external component.
[0118] When a bit stream including video / image information is input, the image / video decoder 400 may generate a signal in response to the bit stream in the image / video decoder. Figure 3 The image / video can be reconstructed by the process of processing the image / video information in the image / video encoder 300. For example, the image / video decoder 400 can derive the unit / block based on the block partition related information obtained from the bit stream. The image / video decoder 400 can perform decoding by using the processing unit applied in the image / video encoder. Therefore, the processing unit of decoding can be, for example, a coding unit, and the coding unit can be partitioned according to the quadtree structure, binary tree structure and / or ternary tree structure from the coding tree unit or the maximum coding unit. At least one transform unit can be derived from the coding unit. And, the reconstructed image signal decoded and output by the image / video decoder 400 can be played by a playback device.
[0119] The image / video decoder 400 can be obtained from the Figure 3The encoder in the decoder receives a signal output and can decode the received signal through the entropy decoder 410. For example, the entropy decoder 410 can parse the bitstream to derive information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), and video parameter set (VPS). Furthermore, the video / image information may further include general constraint information. Furthermore, the image / video information may include the generation method, usage method, and purpose of the decoding information. The image / video decoder 400 can further decode the picture based on the information about the parameter sets and / or general constraint information. The signaled / received information and / or syntax elements can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 410 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, or CABAC, and can output the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients associated with the residual. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information about the target syntax element, decoded information about neighboring and target blocks, or information about symbols / bins decoded in previous steps, predicts the probability of a bin occurring based on the determined context model, and performs arithmetic decoding on the bin to generate a symbol corresponding to the value of each syntax element. After determining the context model, the CABAC entropy decoding method updates the context model using information about the decoded symbol / bin for the next symbol / bin. Information decoded by the entropy decoder 410 includes prediction-related information that can be provided to the predictor (inter-frame predictor 432 and intra-frame predictor 431), and the residual values (i.e., quantized transform coefficients and related parameter information) entropy-decoded by the entropy decoder 410 can be input to the residual processor 420. The residual processor 420 can derive a residual signal (residual block, residual sample, residual sample array). Furthermore, information decoded by the entropy decoder 410 includes filtering-related information that can be provided to the filter 450. Meanwhile, a receiver (not shown) that receives a signal output from the image / video encoder may be additionally constructed as an internal / external element of the image / video decoder 400, or the receiver may be a component of the entropy decoder 410. Meanwhile, the image / video decoder according to the present disclosure may also be referred to as an image / video decoding device, and the image / video decoder may be divided into an information decoder (image / video information decoder) and / or a sample decoder (image / video sample decoder).In this case, the information decoder may include an entropy decoder 410 , and the sample decoder may include at least one of an inverse quantizer 321 , an inverse transformer 322 , an adder 440 , a filter 450 , a memory 460 , an inter predictor 432 , and an intra predictor 431 .
[0120] The inverse quantizer 421 may inverse quantize the quantized transform coefficients and the output transform coefficients. The inverse quantizer 421 may reorder the quantized transform coefficients in the form of two-dimensional blocks. In this case, the reordering may be performed based on the coefficient scanning order performed in the image / video encoder. The inverse quantizer 321 may inverse quantize the quantized transform coefficients using a quantization parameter (i.e., quantization step size information) to obtain the transform coefficients.
[0121] The inverse transformer 422 may perform an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0122] The predictor 430 may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether intra prediction or inter prediction is applied to the current block based on the prediction-related information output from the entropy decoder 410, and may determine a specific intra / inter prediction mode (prediction method).
[0123] The predictor 420 can generate prediction signals based on various prediction methods. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame and inter-frame prediction for a block. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). Furthermore, the predictor can use an intra-block copy (IBC) prediction mode or palette mode for block prediction. IBC prediction mode or palette mode can be used for content image / video coding for games, such as screen content coding (SCC). IBC essentially performs prediction within the current picture, but it can be performed similarly to inter-frame prediction in that it derives reference blocks within the current picture. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When palette mode is applied, information related to the palette table and palette index can be included in the image / video information and signaled.
[0124] The intra-frame predictor 431 can predict the current block by referencing samples within the current picture. The reference samples can be located in the neighborhood of the current block or can be located away from the current block depending on the prediction mode. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 431 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0125] The inter-frame predictor 432 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information regarding the inter-frame prediction direction (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks can include spatially neighboring blocks within the current picture and temporally neighboring blocks in a reference picture. For example, the inter-frame predictor 432 can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction-related information can include information indicating the inter-frame prediction mode used for the current block.
[0126] The adder 440 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 432 and / or the intra-frame predictor 431). When there is no residual for the processing target block, such as when skip mode is applied, the prediction block can be used as the reconstructed block.
[0127] The adder 440 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next processing target block in the current picture, or as described later, may be output through filtering, or may be used for inter-frame prediction of the next picture.
[0128] At the same time, luma mapping with chroma scaling can be applied during picture decoding.
[0129] The filter 450 may apply filtering to the reconstructed signal to enhance subjective / objective image quality. For example, the filter 450 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may send the modified reconstructed picture to the memory 460, specifically, to the DPB of the memory 460. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0130] The (modified) reconstructed picture stored in the DPB of the memory 460 can be used as a reference picture in the inter-frame predictor 432. The memory 460 can store the motion information of the block from which the motion information within the current picture and / or the motion information of the block in the reconstructed picture is derived (or decoded). The stored motion information can be sent to the inter-frame predictor 432 to be used as the motion information of the spatially or temporally neighboring blocks. The memory 460 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 431.
[0131] Meanwhile, the VCM decoder (or feature / feature map decoder) can have substantially the same Figure 4 The image / video decoder 400 described above has the same / similar structure as the one described above, as it performs a series of processes such as prediction, inverse transform, and inverse quantization to decode features / feature maps. However, the VCM decoder differs from the image / video decoder 400 in that it targets features / feature maps for decoding, and therefore, the name of each unit (or component) (e.g., DPB, etc.) and its specific operational details may differ from the image / video decoder 400. The operation of the VCM decoder may correspond to that of the VCM encoder, and its specific operational details will be described in detail later.
[0132] Feature / feature map encoding process
[0133] Figure 5 is a flow chart schematically illustrating a feature / feature map encoding process to which embodiments of the present disclosure may be applied.
[0134] refer to Figure 5 The feature / feature map encoding process may include a prediction process S510, a residual processing process S520, and an information encoding process S530.
[0135] The prediction process S510 can be obtained by referring to Figure 3 The predictor 320 described above is executed.
[0136] Specifically, the intra-frame predictor 322 can predict the current block (i.e., the set of feature elements currently to be encoded) by referencing feature elements in the current feature / feature map. Intra-frame prediction can be performed based on the spatial similarity of the feature elements that constitute the feature / feature map. For example, it can be estimated that feature elements included in the same region of interest (RoI) within an image / video have similar data distribution characteristics. Therefore, the intra-frame predictor 322 can predict the current block by referencing pre-reconstructed feature elements within the RoI that includes the current block. In this case, the referenced feature elements can be located adjacent to the current block or separately from the current block depending on the prediction mode. The intra-frame prediction modes used for feature / feature map encoding can include multiple non-directional prediction modes and multiple directional prediction modes. Non-directional prediction modes can include, for example, prediction modes corresponding to the DC mode and planar mode of the image / video encoding process. In addition, directional modes can include, for example, prediction modes corresponding to the 33 directional modes or the 65 directional modes of the image / video encoding process. However, these are merely examples, and the type and number of intra-frame prediction modes can be configured / changed in various ways depending on the embodiment.
[0137] The inter-frame predictor 321 can predict the current block based on a reference block (i.e., a set of reference feature elements) specified by motion information about a reference feature / feature map. Inter-frame prediction can be performed based on the temporal similarity of the feature elements that make up the feature / feature map. For example, temporally consecutive features may have similar data distribution characteristics. Therefore, the inter-frame predictor 321 can predict the current block by referencing pre-reconstructed feature elements of the current feature and temporally adjacent features. In this case, the motion information used to specify the reference feature elements may include a motion vector and a reference feature / feature map index. The motion information may further include information related to the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current feature / feature map and temporally neighboring blocks in the reference feature / feature map. The reference features / feature maps comprising the reference block and the reference features / feature maps comprising the temporally neighboring blocks may be the same or different. Temporally neighboring blocks may be referred to as collocated reference blocks, etc., and reference features / feature maps comprising temporally neighboring blocks may be referred to as collocated features / feature maps. The inter-frame predictor 321 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference feature / feature map index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip mode and merge mode, the inter-frame predictor 321 can use the motion information of the neighboring blocks as the motion information of the current block. For skip mode, unlike merge mode, the residual signal may not be sent. For motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference. In addition to the above-mentioned intra-frame prediction and inter-frame prediction, the predictor 320 can also generate a prediction signal based on various prediction methods.
[0138] The prediction signal generated by the predictor 320 can be used to generate a residual signal (residual block, residual feature element) S520. Figure 3 The residual processor 330 described above performs a residual processing process S520. Furthermore, (quantized) transform coefficients may be generated by the transform and / or quantization process for the residual signal, and the entropy encoder 340 may encode information related to the (quantized) transform coefficients as residual information in the bitstream S530. In addition, in addition to the residual information in the bitstream, the entropy encoder 340 may also encode information necessary for feature / feature map reconstruction, such as prediction information (e.g., prediction mode information, motion information, etc.).
[0139] At the same time, the feature / feature map encoding process may further include a process for generating a reconstructed feature / feature map for the current feature / feature map, a process for applying in-loop filtering to the reconstructed feature / feature map (optional), and a process S530 for encoding information for feature / feature map reconstruction (e.g., prediction information, residual information, partition information, etc.) and outputting it in the form of a bitstream.
[0140] The VCM encoder can derive (modified) residual features from the quantized transform coefficients through inverse quantization and inverse transformation, and can generate reconstructed features / feature maps based on the predicted features and the (modified) residual features as output in S510 . The reconstructed features / feature maps generated in this manner can be identical to those generated by the VCM decoder. When in-loop filtering is performed on the reconstructed features / feature maps, modified reconstructed features / feature maps can be generated through in-loop filtering. The modified reconstructed features / feature maps can be stored in a decoded feature buffer (DFB) or memory and subsequently used as reference features / feature maps in the feature / feature map prediction process. Furthermore, (in-loop) filtering-related information (parameters) can be encoded and output in the form of a bitstream. The in-loop filtering process can remove noise that may occur during feature / feature map encoding and improve the performance of tasks based on the feature / feature map. Furthermore, the in-loop filtering process can be performed on both the encoder and decoder sides to ensure the identity of the prediction results, improve the reliability of feature / feature map encoding, and reduce the amount of data transmitted for feature / feature map encoding.
[0141] Feature / feature map decoding process
[0142] Figure 6 is a flow chart schematically illustrating a feature / feature map decoding process to which embodiments of the present disclosure may be applied.
[0143] refer to Figure 6The feature / feature map decoding process may include an image / video information acquisition process S610, a feature / feature map reconstruction process S620 to S640, and an in-loop filtering process S650 for reconstructing the feature / feature map. The feature / feature map reconstruction process may be performed based on the prediction signal and residual signal obtained by the inter / intra prediction S620, residual processing S630, and inverse quantization and inverse transformation process for the quantized transform coefficients described in the present disclosure. Modified reconstructed features / feature maps may be generated by the in-loop filtering process for reconstructing features / feature maps, and the modified reconstructed features / feature maps may be output as decoded features / feature maps. The decoded features / feature maps may be stored in a decoded feature buffer (DFB) or a memory and then used as reference features / feature maps in the inter prediction process when decoding the features / feature maps. In some cases, the in-loop filtering process described above may be omitted. In this case, the reconstructed features / feature maps can be output as decoded features / feature maps as is and can be stored in a decoded feature buffer (DFB) or memory and then used as reference features / feature maps in the inter-frame prediction process when decoding the features / feature maps.
[0144] Feature extraction methods and data distribution characteristics
[0145] Embodiments of the present disclosure propose a method for generating the prediction process and the associated bitstream required for compressing activation (feature) maps generated in the hidden layers of a deep neural network.
[0146] Input data input to the deep neural network passes through the computation process of multiple hidden layers, and the computation results from each hidden layer are output as features / feature maps with different sizes and numbers of channels depending on the type of deep neural network being used and the position of the hidden layer in the corresponding deep neural network.
[0147] Figure 7 is a diagram illustrating an example of a feature extraction and reconstruction method to which an embodiment of the present disclosure can be applied.
[0148] refer to Figure 7 , the feature extraction network 710 can extract the intermediate layer activation (feature) map of the deep neural network from the source image / video and output the extracted feature map. The feature extraction network 710 can be a set of consecutive hidden layers from the input of the deep neural network.
[0149] The encoding device 720 may compress the output feature map and output it in the form of a bit stream, and the decoding device 730 may reconstruct the (compressed) feature map from the output bit stream. Figure 1 The encoder 12, and the decoding device 730 may correspond to Figure 1The decoder 22 in . The task network 740 can perform tasks based on the reconstructed feature maps.
[0150] The number of channels of a feature map to be compressed in VCM may vary depending on the network used for feature extraction and the extraction position, and may be larger than the number of channels of the input data.
[0151] Figure 8 FIG2 is a diagram illustrating an example of an image partitioning method to which an embodiment of the present disclosure can be applied, illustrating CTUs, slices, and tiles within an image as examples.
[0152] The video / image encoding method according to this document can be performed based on the following partition structure. Figure 8 Block partitioning is performed on CTUs, CUs (and / or TUs, PUs) derived from the block partition structure, including prediction, residual processing (transform / inverse transform, quantization / dequantization, etc.), syntax element encoding, and filtering. The encoding device can perform block partitioning, and the partition-related information can be encoded and transmitted to the decoding device in the form of a bitstream. The decoding device can derive the block partition structure of the current picture based on the partition-related information obtained from the bitstream and perform a series of image decoding processes (such as prediction, residual processing, block / picture reconstruction, and loop filtering) based on this block partition structure. The CU size can be equal to the TU size, and multiple TUs can exist within a CU area. The CU size generally refers to the luma component (sample) size. The TU size generally refers to the luma component (sample) size. The chroma component (sample) size can be derived from the luma component (sample) size based on the component ratio according to the color format (chroma format, such as 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. Transformation / inverse transformation can be performed on a per-TU (TB) basis.
[0153] In addition, in video / image coding according to this document, image processing units may have a hierarchical structure. A picture may be divided into one or more CUs, and one or more CUs may be grouped and differentiated into one or more tiles, patches, slices, and / or tile groups. A slice may include one or more patches. A patch may include one or more CTU rows within a tile. A slice may include an integer number of tiles of a picture. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be divided into one or more CUs. A tile group may include an integer number of tiles based on a tile raster scan within a picture. A slice header may carry information / parameters applicable to the corresponding slice (i.e., a block within a slice). A picture header may carry information / parameters applicable to the corresponding picture (or a block within a picture). When an encoding / decoding apparatus includes a multi-core processor, encoding / decoding processes for tiles, slices, patches, and / or tile groups may be performed in parallel. In this document, slice and tile group are used interchangeably. In other words, the tile group header may be referred to as a slice header. Here, a slice may have one of the slice types including I slice, P slice, and B slice.
[0154] In the encoding device, depending on the characteristics of the video image (e.g., resolution) or taking into account encoding efficiency or parallel processing, the tile / tile group, patch, slice, and maximum and minimum coding unit sizes can be determined, and information about them or information from which they can be derived can be included in the bitstream.
[0155] In the decoding device, information indicating whether a tile / tile group, block, slice, or CTU within a tile of the current picture is divided into multiple coding units can be obtained. Such information is only obtained (or sent) under certain conditions to improve efficiency.
[0156] The slice header (Slice Header Syntax) may include information / parameters that are generally applicable to the slice. The APS (APS Syntax) or PPS (PPS Syntax) may include information / parameters that are generally applicable to one or more pictures. The SPS (SPS Syntax) may include information / parameters that are generally applicable to one or more sequences. The VPS (VPS Syntax) may include information / parameters that are generally applicable to multiple layers. The DPS (DPS Syntax) may include information / parameters that are generally applicable to the entire video. The DPS may include information / parameters related to the concatenation of coded video sequences (CVS).
[0157] In this document, the term "high-level syntax" may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, or slice header syntax.
[0158] In addition, for example, information on partitioning and configuration of tiles / tile groups / tiles / slices may be configured at the encoding end through a high-level syntax and may be transmitted to a decoding device in the form of a bitstream.
[0159] Figure 9 is a diagram illustrating an example of a VCM image encoding / decoding system, and Figure 10 : is a diagram illustrating another example of a VCM image encoding / decoding system. As an example, Figure 9 and Figure 10 can be extended / redesigned to allow video coding systems (e.g. Figure 1 ) Use only a portion of the video source, or obtain and use the necessary portion / information from the video source depending on the user or machine's request and purpose and the surrounding environment. In other words, Figure 9 and Figure 10 May involve Video Coding Machine (VCM).
[0160] Video Coding for Machines (VCM) can refer to encoding and decoding an entire image, a portion of an image, and / or necessary information (features) derived from an image, depending on the user and / or machine's request, purpose, and surrounding environment. The encoding target in VCM can be the image itself, information called features extracted from an image based on the user and / or machine's request, purpose, and surrounding environment, or a continuous collection of information over time.
[0161] refer to Figure 9 , the VCM system may include an image encoder and an image decoder for VCM. Source device ( Figure 1 ) The encoded image information can be sent to a receiving device via a storage medium or a network. The entity using the device can be a human and / or a machine.
[0162] refer to Figure 10 The VCM system can include an image encoder and an image decoder for VCM. The source device can send the encoded image information (features) to the receiving device via a storage medium or a network. The entity using the device can be a human and / or a machine.
[0163] For example, the process of extracting information (i.e., features) from an image can be referred to as feature extraction. Feature extraction can be performed in both video / image capture devices and video / image generation devices. Features can be information extracted / processed from an image based on user and / or machine requests, objectives, and surrounding context, and can represent a continuous collection of information over time.
[0164] An image encoder for VCM can perform a series of processes such as prediction, transformation, quantization, etc. to efficiently compress and encode the entire image and / or a portion and / or features of the image. The encoded data can be output in the form of a bitstream.
[0165] The image decoder for VCM may decode a video / image by performing a series of processes such as inverse quantization, inverse transformation, prediction, etc., which correspond to the operation of an encoding device (in other words, an encoder).
[0166] The decoded images and / or features can be rendered. Additionally, they can be used to perform tasks for users or machines. Examples of such tasks include AI and computer vision tasks such as face recognition, behavior recognition, and lane recognition.
[0167] The present disclosure provides various embodiments related to the acquisition and encoding of entire images and / or portions of images for VCM, and unless otherwise specified, these embodiments can be combined and performed with each other. The methods / embodiments of the present disclosure can be applied to the methods disclosed in the Video Coding Machine (VCM) standard.
[0168] VCM layer structure
[0169] VCM can be based on a layer structure consisting of a feature encoding layer, a neural network (feature) abstraction layer, and a feature extraction layer. Figure 11 is a diagram illustrating an example of a VCM layer structure, and Figure 12 is a diagram illustrating an example of a VCM bitstream composed of coded abstract features and NNAL information.
[0170] As an example, refer to Figure 11 , the VCM layer structure may include a feature extraction layer 1110 , a neural network (feature) abstraction layer 1120 and a feature encoding layer 1130 .
[0171] The feature extraction layer 1110 may refer to a layer for extracting features from an input source and may also include the extracted results. The feature encoding layer 1130 may refer to a layer for compressing the extracted features and may also include the compressed results.
[0172] The neural network abstraction layer 1120 can abstract the information generated from the feature extraction layer 1110 (e.g., information about the extracted features / feature maps) and send it to the feature encoding layer 1130. The neural network abstraction layer 1120 can hide the internal structure of the feature extraction layer 1110 and provide a consistent feature interface function through information abstraction. Therefore, even when the compression target changes due to changes in tools (e.g., CNN, DNN, etc.), the feature encoding layer 1130 can perform a consistent feature encoding process. In this disclosure, the neural network abstraction layer (NNAL) may also be referred to as a feature abstraction layer.
[0173] The interface between the feature extraction layer 1110 and the neural network abstraction layer 1120, and the interface between the feature encoding layer 1130 and the neural network abstraction layer 1120 may be predefined, and the operations in the neural network abstraction layer 1120 may be configured to be modifiable later.
[0174] refer to Figure 12 , a bitstream configured as shown can be referred to as a neural network abstraction layer (NNAL) unit. NNAL units can be independent feature reconstruction units. Input features for a single NNAL unit can be extracted from the same layer in a neural network. Therefore, the input features for a single NNAL unit can be forced to have the same characteristics. For example, the same feature extraction method can be applied to the input features for a single NNAL unit.
[0175] An NNAL unit may include an NNAL unit header and an NNAL unit payload. The NNAL unit header may include all information necessary to use the encoded features for a task. The NNAL unit payload may include abstracted feature information. The NNAL unit payload may include a group header and group data. The group header may include configuration information for the feature group data, such as the temporal order, number, or common attributes of the feature channels that make up the feature group. A feature channel may refer to a unit of encoded features. The group data may include multiple feature channels and an encoding indicator, and each feature channel may include type information, prediction information, auxiliary information, and residual information. In this case, the type information may indicate the encoding method, and the prediction information may indicate the prediction method. In addition, the auxiliary information may indicate additional information required for decoding (e.g., entropy coding, quantization-related information, etc.), and the residual information may include information about the encoded feature elements (i.e., a collection of feature value information).
[0176] Overview of Quantization / Dequantization
[0177] As described above, the quantizer of the image / video encoder may apply quantization to the transform coefficients to derive the quantized transform coefficients, and the inverse quantizer of the image / video encoder or the inverse quantizer of the image / video decoder may apply inverse quantization to the quantized transform coefficients to derive the transform coefficients. Similarly, the VCM encoder may apply quantization to the transform coefficients to derive the quantized transform coefficients, and the inverse quantizer of the VCM encoder or the inverse quantizer of the VCM decoder may apply inverse quantization to the quantized transform coefficients to derive the transform coefficients. Generally, in feature / feature map encoding, the quantization rate may vary, and the varying quantization rate may be used to control compression efficiency. From an implementation perspective, a quantization parameter (QP) may be used instead of directly using the quantization rate in consideration of complexity. For example, a quantization parameter having an integer value from 0 to 63 may be used, and each quantization parameter value may correspond to an actual quantization rate. Quantization parameter (QP) for the luminance component (luminance sample) Y ) and the quantization parameter (QP) for the chroma components (chroma samples) C ) can be set differently.
[0178] During quantization, the transform coefficient (C) is taken as input and divided by the quantization rate (Qstep), thereby obtaining the quantized transform coefficient (C'). In this case, considering computational complexity, the scale may be multiplied by the quantization rate to convert it to integer form, and a shift operation may be performed by a value corresponding to the scale value. The quantization scale can be derived based on the product of the quantization rate and the scale value. In other words, the quantization scale can be derived based on the QP. The quantized transform coefficient (C') can be obtained by applying the quantization scale to the transform coefficient (C).
[0179] The inverse quantization process is the inverse of the quantization process, in which the quantized transform coefficient (C') is multiplied by the quantization rate (Qstep) to obtain the reconstructed transform coefficient (C''). In this case, the level scale can be derived from the quantization parameter, and the reconstructed transform coefficient (C'') can be derived by applying the level scale to the quantized transform coefficient (C''). Due to losses in the transformation and / or quantization process, the reconstructed transform coefficient (C'') may be different from the original transform coefficient (C). Therefore, the encoder also performs inverse quantization in the same manner as the decoder.
[0180] At the same time, an adaptive frequency-weighted quantization technique can be applied to adjust the quantization strength based on frequency. Adaptive frequency-weighted quantization is a method that applies different quantization strengths to each frequency. Adaptive frequency-weighted quantization can use a predefined quantization scaling matrix to apply different quantization strengths to each frequency. In other words, the quantization / inverse quantization process described above can be further performed based on the quantization scaling matrix. For example, different quantization scaling matrices can be used depending on the size of the current block and / or whether the prediction mode applied to the current block to generate the current block residual signal is inter-frame prediction or intra-frame prediction. A quantization scaling matrix can also be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix can be predefined. Furthermore, for frequency-adaptive scaling, quantization scaling information for each frequency of the quantization scaling matrix can be constructed / encoded in the encoder and signaled to the decoder. This quantization scaling information for each frequency can be referred to as quantized scaling information. This quantization scaling information for each frequency can include scaling table data. The quantization scaling matrix can be derived (modified) based on the scaling table data. This quantization scaling information for each frequency can also include current flag information indicating whether the scaling table data exists. Alternatively, when zoom list data is signaled at a higher level (e.g., sequence level, feature set group level, etc.), it may further include information indicating whether the zoom list data is modified at a lower level (e.g., feature set level, channel level, etc.).
[0181] Feature quantization / dequantization can be performed based on a predetermined quantization group. Specifically, multiple quantization intervals can be set based on the data distribution characteristics of the feature set. The set quantization intervals can have different data distribution ranges and can be defined as a quantization group. Feature quantization / dequantization operations can then be performed by converting the data distribution of each channel within the feature set into one of the data distribution ranges of the quantization interval.
[0182] Figure 13 is a diagram illustrating an example of a quantization group to which an embodiment of the present disclosure can be applied.
[0183] refer to Figure 13 , a quantization group may include four quantization intervals (A, B, C, or D) with different data distribution ranges. Each of the quantization intervals (A, B, C, or D) in a quantization group may be defined using a minimum value and a maximum value and may be set based on the data distribution characteristics of the current feature / feature map. For example, quantization interval A may be set to [-1, 3], quantization interval B may be set to [0, 2], quantization interval C may be set to [-2, 4], and quantization interval D may be set to [-2, 1].
[0184] Without having to individually calculate the maximum and minimum values of feature elements, the VCM encoder can quantize each channel (or feature) within a feature set based on predefined quantization intervals (A, B, C, or D). For example, channel 1 can be quantized based on quantization interval A, which has the most similar data distribution range. Channel 2 can be quantized based on quantization interval B, which has the most similar data distribution range. Channel 3 can be quantized based on quantization interval C, which has the most similar data distribution range. And channel 4 can be quantized based on quantization interval D, which has the most similar data distribution range. In this case, the VCM encoder can encode the number of quantization intervals, as well as their minimum and maximum values, into / signal information related to feature quantization. Furthermore, the VCM encoder can also signal quantization interval index information, indicating the quantization interval used to encode the current channel (or feature), as feature quantization information. Therefore, there is no need to signal the number of quantization bits for each channel and the maximum and minimum values of the feature elements, reducing the number of transmitted bits and further improving encoding / signaling efficiency.
[0185] The VCM decoder can construct a quantization group identical to the quantization group constructed by the VCM encoder based on the number of quantization intervals and the minimum and maximum values of each interval received from the feature encoding device. Alternatively, the VCM decoder can construct a quantization group based on the data distribution characteristics of the pre-reconstructed feature set. The VCM decoder can then perform inverse quantization of the current feature based on the quantization interval identified by the quantization interval index information received from the VCM encoder.
[0186] The feature quantization based on the quantization group can be calculated as follows.
[0187] [Equation 1]
[0188]
[0189] Among them, Fn can represent the feature set (Fset RxC ) (where n is an integer greater than or equal to 1). In addition, R may represent the width of the feature set, C may represent the height of the feature set, max(Fn) may represent the maximum value of the feature elements within the n-th channel (Fn), and min(Fn) may represent the minimum value of the feature elements within the n-th channel (Fn).
[0190] Referring to the above equation, the nth channel (Fn) may be normalized to have a feature element value between 0 and 1 based on the maximum and minimum values of the feature elements within the channel.
[0191] [Equation 2]
[0192]
[0193] Among them, Fn norm Can represent feature sets (Fset RxC ) in the normalized nth channel. In addition, R can represent the width of the feature set, C can represent the height of the feature set, max(F Gm ) can represent the quantization interval (F) applied to the nth channel (Fn) Gm ) and min(F Gm ) can represent the quantization interval (F) applied to the nth channel (Fn) Gm ) is the minimum value of .
[0194] Referring to the equation above, the quantization interval (F Gm )'s maximum value (max(F Gm )) and minimum value (min(F Gm )) for the normalized nth channel (Fn norm ) for quantification.
[0195] Meanwhile, for inverse quantization, quantization-related information may be encoded into a bitstream. According to an embodiment, the quantization-related information may include overall quantization information, activation function information, the number of quantization bits, and quantization interval index information. The overall quantization information may indicate whether the number of quantization bits is set individually for each quantization interval or is set equally for all quantization intervals. The activation function information may indicate the type of activation function applied to the current feature. For example, the activation function information may indicate whether the activation function applied to the current feature is a first activation function that requires both the maximum value and the minimum value of each quantization interval to be signaled, or a second activation function that only requires the maximum value of each quantization interval to be signaled. The number of quantization bits may be set for each quantization interval based on the above-mentioned overall quantization information, or may be set equally for all quantization intervals. When the number of quantization bits is set equally for all quantization intervals, the number of quantization bits may be signaled only once for all quantization intervals.
[0196] In addition, the quantization-related information may further include the number of quantization intervals within the quantization group and the minimum and maximum values of each quantization interval. In this case, the minimum and maximum values of each quantization interval can be adaptively signaled based on the type of activation function. For example, when the activation function applied to the current feature is the first activation function (e.g., Leaky ReLU), both the minimum and maximum values of each quantization interval can be signaled. Conversely, when the activation function applied to the current feature is the second activation function (e.g., ReLU), the minimum value of each quantization interval can be assumed to be 0, and only the maximum value of each quantization interval can be signaled.
[0197] Example
[0198] The present disclosure can be based on VCM, which is intended not only for conventional video coding techniques but also for performing machine tasks. Furthermore, in VCM, the attributes of the required information may vary depending on the type of machine task being performed. Conventional quantization processes assume human observation (in other words, processing by humans) and are optimized solely for this purpose, resulting in limitations in effectively processing the various attribute information required for machine tasks.
[0199] The present disclosure relates to a method for defining and representing the attributes and usage information of a VCM bitstream according to a quantization optimization method for a VCM bitstream, which can solve the problems of the above-mentioned prior art. In the present disclosure, a quantization optimization method is defined to improve the performance of machine perception and the attributes of information required for the optimization method, and an image encoding and decoding method based on the same is proposed. More specifically, as an embodiment, the present disclosure proposes a quantization method for machine perception, and a method for representing image usage and information attributes in a VCM bitstream according to its application. In addition, as an embodiment, the present disclosure proposes a method for optimizing quantization with the goal of improving machine perception performance, as well as a method for defining the attributes of encoded information and an applied optimization method capable of selecting a VCM bitstream and encoding method suitable for the machine task to be performed.
[0200] Meanwhile, although the terms feature and video are distinguished above, for the sake of clarity, the term video used in this disclosure may collectively refer to features / videos / images / image data, etc. In other words, the content described below that applies to videos can also be applied to features, images, image data, etc., and vice versa. In other words, the embodiments described below in terms of image data, features, etc. can also be applied to videos.
[0201] Meanwhile, the quantization (optimization) method according to an embodiment of the present disclosure as described below may be applied in reverse during the inverse quantization process.
[0202] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0203] Figure 14The diagram illustrates an example of images with different attributes according to an embodiment of the present disclosure. When the purpose of the video changes, in other words, for example, when the processing object such as machine perception or human perception changes, images with different attributes may be required. For example, unlike the VCM based on machine perception, the human visual system (HVS) may have such a characteristic that it is sensitive to quality changes in low-frequency components, but is insensitive to quality changes in high-frequency components compared to quality changes in low-frequency components. Based on this characteristic of the HVS, quantization can be optimized in a manner that enhances the quality in the low-frequency domain and reduces the quality in the high-frequency domain. As an example, the quantization optimization method based on the HVS can be referred to as frequency-based quantization optimization (FQO). The following equation 3 shows an example of applying FQO.
[0204] [Equation 3]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210] In the above equation, w k Can represent weight, QP can represent quantization parameter, and QP k It can represent the quantization parameter optimized after FQO. As an example, with w representing the weight k Becomes larger, QP k may become smaller, and therefore, the amount of coding bits and the quality may increase. Here, w k With a k becomes smaller and becomes larger structure, and a k The average value of the high frequency components can be expressed as the average value of values obtained by applying the Laplace filter F as a high-pass filter to the encoding target. k It can be called the region-by-region frequency sensitivity of the coding target. In other words, when w k Due to a k When the value of becomes greater than 1, QP k On the other hand, when w k Due to a k When the value of becomes less than 1, QP k In other words, the quantization parameter QP k Can be based on frequency a k and / or weight wk In this way, more bits can be allocated to the low-frequency regions that are sensitively sensed by the HVS, thereby improving the quality of these regions, and fewer bits can be allocated to the encoding of the high-frequency regions that are not sensitively sensed, thereby reducing their quality.
[0211] However, such FQO may be effective for human perception but inefficient for other purposes, for example, it may be inefficient for machine perception. In the case of images based on machine perception, performing a specific task (e.g., object detection, object tracking, etc.) may be the purpose, and the required information attributes may vary depending on the image usage. Therefore, an FQO based on a uniform application of region-by-region frequency sensitivity may be inefficient for machine perception. Figure 14 When the embodiment of the present disclosure is shown in FIG, it is assumed that the machine task to be performed on both images is face recognition. Figure 14 In (a), complex patterns are distributed in the background area, so the region-by-region frequency sensitivity indicating the amplitude of high-frequency components in the background may appear high. On the other hand, in Figure 14 In (b), the background is flat and does not vary, so the region-by-region frequency sensitivity may appear low. Assume that FQO is applied to Figure 14 The two images in Figure 14 In the case of (a), the bit amount of the background area that is not necessary for the machine task of face recognition can be reduced, and the bit amount of the face area that is necessary for performing the machine task can be increased, thereby improving the performance of the machine task and the coding efficiency. Figure 14 (b), it may increase the amount of bits in unnecessary background areas for machine tasks such as performing face recognition, thereby reducing coding efficiency. Therefore, when the purpose is to perform machine tasks, it may not be necessary to increase the amount of bits even in areas with low frequency sensitivity in terms of human perception. Therefore, it may be necessary to optimize and apply FQO taking into account the attributes of the information required for the machine task and the environment in which the machine task is performed (e.g., background, lighting, etc.). As an example, in order for the receiving end of the image to select an appropriate encoding method or encoded bit stream according to the purpose, it may be necessary to include the purpose and use of the bit stream (e.g., for machine tasks, for human perception, etc.) and information about the optimization method in the bit stream. In the present disclosure, the FQO optimized for the purpose of performing machine tasks is defined as machine perception quantization optimization (QO), and the FQO for human perception can be defined as human perception quantization optimization (QO).
[0212] Meanwhile, the quantization (optimization) method according to the embodiments of the present disclosure described above and below may be reversely applied in the inverse quantization process.
[0213] The embodiments described below with reference to Table 1 relate to a method for representing a quantization method optimized for a machine task and information about the applied quantization method in a bitstream. Table 1 below shows an example of machine task performance based on the application of FQO. In Table 1 below, machine task performance is expressed in bits per pixel (BPP) and mean average precision (mAP).
[0214] [Table 1]
[0215]
[0216] According to Table 1, when FQO is applied, BPP is increased at all QPs, and the performance of the machine task may also be improved except for QP22. When FQO is applied, the quality and bit amount of the low-frequency sensitivity area may increase, and the quality and bit amount of the high-frequency sensitivity area may decrease. In the case of QP22, the degradation in the performance of the machine task caused by the decrease in quality in the high-frequency sensitivity area may have a more significant impact on the overall performance than the improvement obtained from enhancing the quality of the low-frequency sensitivity area. However, because the information necessary for the performance of the machine task may still exist in the high-frequency sensitivity area, it is necessary to redefine the criteria for determining the high-frequency sensitivity area based on machine perception rather than HVS. Similarly, the criteria for determining the low-frequency sensitivity area may also need to be redefined based on machine perception rather than HBS to minimize the increase in bit amount.
[0217] Figure 15 The present invention illustrates an example of a frequency sensitivity determination standard according to an embodiment of the present disclosure. More specifically, the present invention illustrates an example of a general frequency sensitivity determination standard and a frequency sensitivity determination standard for machine perception. Figure 15 (a) shows an example of a frequency sensitivity determination standard in a human perception FQO according to an embodiment of the present disclosure, and Figure 15 (b) shows an example of the frequency sensitivity determination criteria in machine perception FQO. Figure 15 (c) illustrates an example of an area unnecessary for machine task performance based on defining a high-level threshold according to an embodiment of the present disclosure, and Figure 15 (d) shows an example of applying both high-level and low-level thresholds.
[0218] As an example, in reference Figure 15In the human perception FQO of (a), the entire image can be divided into a low-frequency sensitivity region 1501 and a high-frequency sensitivity region 1502. To divide these regions, a global frequency sensitivity 1503 can be used. In this case, based on the frequency sensitivity of the entire image, regions with frequency sensitivity greater than a specific value (e.g., a high-level threshold) can have their bit count reduced, while regions with frequency sensitivity less than this specific value can have their bit count increased. In other words, the bit count of the high-frequency sensitivity region 1502 can be reduced, while the bit count of the low-frequency sensitivity region 1501 can be increased. As an example, according to Equation 1, as the frequency sensitivity decreases, the bit count may increase.
[0219] According to an embodiment of the present disclosure, a standard for determining frequency sensitivity for avoiding an increase in the bit amount can be provided. Figure 15 (b), according to an embodiment of the present disclosure, an example of a method for determining an area with little information variation that is unnecessary for machine task performance may be presented by defining a low-level threshold. As an example, the entire image may be divided into a high-frequency sensitivity area 1511, a medium-frequency sensitivity area 1512, and a low-frequency sensitivity area 1513. In order to distinguish such areas, a global frequency sensitivity 1514 and a low-level threshold 1515 may be used. As an example, an area with very low frequency sensitivity (e.g., 1513) may be a smooth area with almost no information variation and may be considered unnecessary for machine task performance. Therefore, in this example, Figure 15 The low-frequency sensitivity region in (a) is divided into two regions—a medium-frequency sensitivity region and a low-frequency sensitivity region—so that it is possible to control the region with low-frequency sensitivity and the smaller amount of bits of information necessary for performing machine tasks.
[0220] According to another embodiment of the present disclosure, a criterion for determining frequency sensitivity may be presented in order to avoid degradation of machine mission performance. Figure 15 In the high-frequency sensitivity regions shown in (a), information necessary for executing the machine task may also be distributed, and uniformly reducing the bit amount in these regions may lead to degradation of the machine task performance. For example, in the case of QP22 in Table 1, it is shown that although the bit amount has increased, the machine task performance has degraded. Therefore, Figure 15 (c) shows an example of a method for determining an area necessary for performing a machine task by defining a high-level threshold according to an embodiment of the present disclosure. Figure 15(c) The entire image can be classified into ultra-high frequency sensitivity areas 1521, high frequency sensitivity areas 1522, and low frequency sensitivity areas 1523, and these areas can be distinguished using advanced thresholds 1524, global frequency sensitivity 1525, and the like. In this case, even when an area is determined by FQO to have high frequency sensitivity, when the frequency sensitivity is below the advanced threshold, in other words, in the case of a specific area 1522, the bit amount can be adjusted differently from the FQO. However, area 1521 having a frequency sensitivity higher than advanced threshold 1524 is likely to be information unnecessary for performing the machine task, for example, because it resembles noise.
[0221] at the same time, Figure 15 (d) shows an example of applying both high-level thresholds and low-level thresholds according to an embodiment of the present disclosure. Figure 15 (d), the entire image can be classified into a very high frequency sensitivity region 1531, a high frequency sensitivity region 1532, a medium frequency sensitivity region 1533, and a low frequency sensitivity region 1534. To classify such regions, a high level threshold 1535, a global frequency sensitivity 1536, and a low level threshold 1537 can be used. In this case, the low frequency sensitivity region 1534 and the very high frequency sensitivity region 1531 can be determined as regions of lesser importance from a machine perception perspective.
[0222] According to an embodiment of the present disclosure, by Figure 15 By subdividing the regions within the entire image based on frequency as shown in , the image can be classified based on image usage such as human perception, machine perception, etc., thereby improving encoding performance.
[0223] Figure 16 An example of the operation of an image encoder or decoder based on a frequency sensitivity determination standard according to an embodiment of the present disclosure is illustrated. More specifically, Figure 16 Show application Figure 15 An example of the operation of a standard image encoder or decoder. As an example, Figure 16 Can represent the operation of an image encoder or decoder based on a low-level threshold.
[0224] As an example, perceptual QO adaptation may be the beginning of process S1610. Perceptual QO adaptation may refer to Figure 16 A series of processes in the process, and may include steps S1620, S1630, S1640, S1650, S1660, S1670, etc.
[0225] For example, a L The calculation of can be performed S1620. Variable a Lcan represent the frequency sensitivity of the encoding target area and can correspond to a in Equation 1 k For example, S1620 may include a process of calculating the frequency sensitivity of the encoding target area based on Equation 1.
[0226] Subsequently, a determination may be made as to whether coding optimization for machine task execution is applied ( S1630 ). S1630 may be performed based on a specific syntax. For example, the encoder may include a process for determining the value of the specific syntax and / or a process for determining whether to apply optimization based on the value of the specific syntax; and the decoder may include a process for determining whether to apply optimization based on the value of the specific syntax. The specific syntax may be the syntax "machine_perceptual_QO_flag." For example, the syntax "machine_perceptual_QO_flag" may indicate whether coding optimization for machine task execution is applied. For example, when the value of this syntax is a first value (e.g., true, 1), it may indicate that optimization for machine task execution has been applied. In other words, it may indicate that machine perceptual QO has been applied, using frequency sensitivity determination criteria such as a low-level threshold and / or a high-level threshold. On the other hand, when the value of this syntax is a second value (e.g., false, 0), it may indicate that optimization for machine task execution has not been applied. In this case, conventional human perceptual QO may have been applied to the encoded bitstream. The decoder can identify the properties and usage of the encoded bitstream using the defined syntax "machine_perceptual_QO_flag."
[0227] When it is determined that encoding optimization for executing machine tasks is not applied, another type of QP optimization (S1640) may be performed. On the decoder side, inverse quantization may be performed based on the QP optimization method performed on the encoder side. In this case, the QP optimization may be based on human perception, FQO, or other methods.
[0228] Meanwhile, when determining to apply coding optimization for executing machine tasks, an operation may be performed based on a comparison result between the frequency sensitivity of each region and the frequency sensitivity of the entire image. First, as an example, a comparison is made between a T and a L The comparison between is performed S1650. As an example, it can be determined that a L Is it less than a T As an example, a T It can be a frequency sensitivity determination standard, such as the global frequency sensitivity mentioned above, Figure 15 The frequency sensitivity of the entire image in , or a in Equation 1 pic . variablea Lcan represent the frequency sensitivity of the encoding target area and can correspond to a in Equation 1 k The QP optimization method may vary depending on the results of S1650.
[0229] Thereafter, an operation may be performed based on the comparison result of the frequency sensitivity of each region and the frequency sensitivity of the entire image. L Less than a T When a LLT and a L A comparison between the two can be performed S1670, and a LLT It can represent a low-level threshold for machine perception. In this case, when the result of S1670 is true, in other words, a L More than a LLT , QP optimization S1680 may be applied. In this case, it may be determined that the region is a high frequency region, and the amount of bits assigned to the region may have been reduced. On the decoder side, inverse quantization may be performed based on the QP optimization method performed on the encoder side. For example, the QP optimization may be a machine-perceived QP optimization, and in this case, the amount of bits assigned to the region may be increased. On the other hand, when the result of S1670 is false, in other words, a L Equal to or less than a LLT When , this corresponds to the case where the regional frequency sensitivity is equal to or less than the low-level threshold, so QP optimization may not be applied. In this case, even if the region-by-region frequency sensitivity is less than the frequency sensitivity of the entire image, the bit amount may not increase.
[0230] At the same time, when the result of S1650 is false, in other words, a L Not less than a T When the QP optimization method is applied, S1690 can be applied. In this case, it can be determined that the region is unnecessary for the machine task, such as a high-frequency region (high activity region), and the bit amount of the region can be reduced. On the decoder side, inverse quantization can be performed based on the QP optimization method performed on the encoder side.
[0231] Figure 17 FIGURE 1 illustrates an example of the operation of an image encoder or decoder based on a frequency sensitivity determination standard according to another embodiment of the present disclosure. More specifically, Figure 17 Icon Application Figure 15 An example of the operation of a standard image encoder or decoder is given in FIG. As an example, Figure 17 An example of the operation of an advanced threshold-based image encoder or decoder may be illustrated.
[0232] As an example, perceptual QO adaptation may be the beginning of process S1710. Perceptual QO adaptation may refer to Figure 17 The series of processes shown in , and may include steps S1720, S1730, S1740, S1750, S1760, S1770, etc.
[0233] As an example, a L The calculation of may be performed S1720, and thereafter, it may be determined whether to apply encoding optimization for performing machine tasks S1730. Subsequently, when it is determined that encoding optimization for performing machine tasks is not applied, another type of QP optimization S1740 may be performed. The QP optimization performed here may include QP optimization based on human perception, in other words, FQO. In this case, S1720 may correspond to Figure 16 S1620 and S1730 in the Figure 16 S1630 in , and S1740 can correspond to Figure 16 S1640 in , so repeated description will be omitted.
[0234] Meanwhile, when determining to apply coding optimization for executing machine tasks, an operation may be performed based on a comparison result between the frequency sensitivity of each region and the frequency sensitivity of the entire image. First, as an example, a T and a L As an example, it is possible to determine a L Is it greater than a T As an example, a T Can be used for Figure 15 The frequency sensitivity of the entire image is used as a standard for determining the frequency sensitivity, that is, the threshold value, and can be a of Equation 1 pic As an example, the variable a L can represent the frequency sensitivity of the encoding target area and can correspond to a in Equation 1 k The QP optimization method may vary depending on the results of S1750.
[0235] Subsequently, an operation may be performed based on the comparison result between the region-by-region frequency sensitivity and the frequency sensitivity of the entire image. L Greater than a T When you can execute a HLT with a L Comparison S1770, where a HLTIt can represent an advanced threshold for machine perception. In this case, it can be determined as a UHF region, and QP optimization S1780 can be performed, and the amount of bits allocated to this region can be reduced. In this case, when the result of S1770 is false, in other words, when a L Less than a HLT , QP optimization for reducing the bit amount may not be applied. In other words, because the region-by-region frequency sensitivity is less than the high-level threshold, the region may be classified as a valid region, i.e., a region including valid data, when processed by the machine, and may be processed without reducing the bit amount and without QP optimization. Meanwhile, when the result of S1750 is false, in other words, when a L Less than or equal to a T When the region is not sufficiently large, QP optimization S1780 may be applied. In this case, the region may be determined to be unnecessary for the machine task, and the amount of bits may be reduced. The QP optimization may be machine-aware QP optimization, and in this case, the amount of bits assigned to the region may have been increased or decreased.
[0236] In the case of S1740 , S1780 , and S1790 , inverse quantization may be performed on the decoder side based on the QP optimization method performed on the encoder side.
[0237] Figure 18 FIGURE 1 illustrates an example of the operation of an image encoder or decoder based on a frequency sensitivity determination standard according to another embodiment of the present disclosure. More specifically, Figure 18 Icon Application Figure 15 An example of the operation of a standard image encoder or decoder is given in FIG. As an example, Figure 18 An example of the operation of an advanced threshold-based image encoder or decoder may be illustrated.
[0238] As an example, the perceptual QO adaptation may first start S1810. Perceptual QO adaptation may refer to Figure 18 The series of processes shown in , and may include steps S1820, S1830, S1840, S1850, S1860, S1870, etc.
[0239] As an example, you can execute a L , and thereafter, it may be determined whether to apply encoding optimization for performing machine tasks S1830. Subsequently, when it is determined that encoding optimization for performing machine tasks is not applied, another type of QP optimization S1840 may be performed. The QP optimization performed here may include QP optimization based on human perception, in other words, FQO. In this case, S1820 may correspond to Figure 16S1620 and S1830 in the Figure 16 S1630 in, and S1840 can correspond to Figure 16 S1640 in , so repeated description will be omitted.
[0240] Meanwhile, when determining to apply coding optimization for executing machine tasks, an operation may be performed based on a comparison result between the frequency sensitivity of each region and the frequency sensitivity of the entire image. First, as an example, a T and a L As an example, it is possible to determine a L Is it less than a T As an example, a T Can be Figure 15 The frequency sensitivity of the entire image shown in , serves as a frequency sensitivity determination standard, in other words, a threshold value, and can be a of Equation 1 pic As an example, the variable a L can represent the frequency sensitivity of the encoding target area and can correspond to a in Equation 1 k The QP optimization method may vary depending on the results of S1850.
[0241] Subsequently, an operation may be performed based on the comparison result between the frequency sensitivity of each region and the frequency sensitivity of the entire image. L Less than a T When you can execute a LLT with a L Comparison S1870 between, and value a LLT It can represent a low-level threshold for machine perception. When the result of S1870 is true, in other words, a LLT The value is less than a L When , QP optimization S1890 can be performed, and in this case, it can be determined as a region required for the machine task, and the bit amount can be increased. On the other hand, when the result of S1870 is false, in other words, when a LLT Greater than a L When , QP optimization may not be applied. In this case, although the frequency sensitivity of each region is less than the frequency sensitivity of the entire image, the bit amount may not be increased. This is consistent with the reference Figure 1 The explanation given is the same.
[0242] Meanwhile, when the result of S1850 is false, in other words, when a L Greater than or equal to a T When you can execute a HLT with a LComparison S1860 between, and when the result of S1860 is true, in other words, when a L Greater than or equal to a HLT , QP optimization S1880 for reducing the bit amount can be applied. In other words, when it corresponds to a case where the frequency sensitivity per region is higher than the high-level threshold, the region may not be classified as a valid region for machine processing, in other words, a region including valid data, and therefore QP optimization can be applied and the bit amount for the region can be reduced. On the other hand, when the result of S1860 is false, in other words, when a L Less than or equal to a HLT When , QP optimization may not be applied. In this case, although the region-by-region frequency sensitivity is greater than or equal to a T , but it is still less than or equal to the high-level threshold, so this region can be determined to be an area necessary for the machine task and can be processed without reducing the amount of bits. At the same time, the above-mentioned QP optimization can be a QP optimization based on machine perception, and in this case, the amount of bits assigned to this region may have been increased or decreased.
[0243] In S1840 , S1880 , and S1890 , the decoder may perform inverse quantization based on the QP optimization method performed by the encoder.
[0244] The following describes the syntax for quantization optimization methods according to embodiments of the present disclosure with reference to Tables 2 to 6. Tables 2 to 6 present examples of perceptual QO identification information. Perceptual QP identification information can be: 1) defined in a conventional NAL unit; 2) based on a NAL unit type defined for identifying perceptual quantization optimization (QO) methods; or 3) based on a supplemental enhancement information (SEI) configuration defined for perceptual QO identification.
[0245] [Table 2]
[0246]
[0247] According to the embodiment shown in Table 2, information regarding whether a machine-perceptual quantization optimization method is applied can be obtained. For example, this information can be referred to as a first syntax and can be referred to as machine_perceptual_QO_flag. For example, the first syntax can be obtained from the bitstream. In other words, the first syntax can indicate whether coding optimization for machine task execution is applied. For example, when the value of the first syntax is a first value (e.g., 1), it can indicate the application of machine-perceptual QO (machine-perceptual or machine-task-based quantization optimization). Conversely, when the value of the first syntax is a second value (e.g., 0), it can indicate that machine-perceptual QO is not applied.
[0248] [Table 3]
[0249]
[0250] [Table 4]
[0251]
[0252] Tables 3 and 4 illustrate examples of information used to indicate the application status and usage of perceptual QP. First, according to the embodiment shown in Table 3, information regarding whether perceptual quantization optimization (e.g., the second syntax) is applied, along with the first syntax. For example, the first syntax can be obtained based on the second syntax. For example, the second syntax can be called perceptual_QO_flag. When the value of the second syntax is a first value (e.g., 1) and the value of the first syntax is a second value (e.g., 1), this may indicate that the perceptual quantization QO applied during encoding is a quantization optimization method for machine tasks. Conversely, when the value of the second syntax is a first value but the value of the first syntax is a second value (e.g., 0), this may indicate that the perceptual QO applied during encoding is not an optimization method for machine tasks. In this case, for example, this may indicate the application of a perceptual QO method optimized for human perception.
[0253] According to the embodiment shown in Table 4, information regarding whether perceptual quantization optimization is applied (e.g., the second syntax) and information indicating whether quantization optimization based on human perception is applied (e.g., human_perceptual_QO_flag) can be obtained in the third syntax. For example, the third syntax can be obtained based on the second syntax. When the second syntax (e.g., perceptual_QO_flag) has a first value (e.g., 1) and the third syntax (e.g., human_perceptual_QO_flag) has a first value (e.g., 1), this can indicate that the perceptual QO applied during encoding is a quantization optimization method for human perception. Meanwhile, when the second syntax has a first value and the third syntax has a second value (e.g., 0), this can indicate that the perceptual QO applied during encoding is not a quantization optimization method for human perception. In this case, this can indicate that a perceptual QO optimized for machine tasks is applied.
[0254] Meanwhile, in the case of Table 3 and Table 4, for clarity of explanation, the information indicating whether machine perception is applied (e.g., the first syntax, machine_perceptual_QO_flag) and the information indicating whether human perception is applied (e.g., the third syntax, human_perceptual_QO_flag) are separately shown as being signaled, but this corresponds to an embodiment, so both the first syntax and the third syntax may be signaled together. In other words, the first syntax may be signaled as in Table 3, followed by the third syntax, or the third syntax may be signaled as in Table 4, followed by the first syntax, and this is also included in the embodiments of the present disclosure.
[0255] [Table 5]
[0256]
[0257] [Table 6]
[0258]
[0259] Table 5 illustrates an example of information indicating whether to apply perceptual QO and a quantization optimization method, and Table 6 illustrates an example of a perceptual QO method that can be indicated using specific syntax in Table 5. According to the embodiment of Table 5, information regarding a quantization optimization method (e.g., the fourth syntax) and a second syntax can be obtained. For example, the fourth syntax can be obtained based on the second syntax. For example, the fourth syntax can be called optimization_method_idc and can be an indicator of the optimization method used for quantization (or inverse quantization). For example, when the value of the second syntax is a first value (e.g., 1), the fourth syntax can be obtained. For example, when the value of the fourth syntax is 0, it can indicate the application of human-based QO. When the value is 1, it can indicate the application of machine-based QO based on a low-level threshold. When the value is 2, it can indicate the application of machine-based QO based on a high-level threshold. When the value is 3, it can indicate the application of machine-based QO based on both the low-level and high-level thresholds.
[0260] At the same time, depending on the machine task to be performed using the image, the characteristics of the required information may also vary. Therefore, the low-level threshold and high-level threshold that can be applied to the QO based on machine perception may vary depending on the properties of the machine task, and in order for the decoding end to select a bitstream suitable for the machine task intended to be performed for the user, the encoding end may need to encode the information required for such selection. In other words, it may be necessary to encode information that varies based on the properties of the machine task. Therefore, Figure 19 An example of defining frequency sensitivity levels according to an embodiment of the present disclosure may be illustrated. More specifically, Figure 19(a) illustrates an example of a process for defining frequency sensitivity levels, and Figure 19 (b) illustrates an example of frequency sensitivity statistical characteristics that may be used to define frequency sensitivity classes.
[0261] First, according to Figure 19 In the embodiment of (a), a data set may be first obtained, frequency sensitivity may be calculated based on the obtained data set, and the frequency sensitivity level may be defined by analyzing the frequency sensitivity distribution. Figure 19 (a), we can obtain values representing the distribution and magnitude of frequency sensitivity, such as mean and deviation. Figure 19 The statistical characteristics of frequency sensitivity are analyzed as shown in (b), and based on this, frequency sensitivity levels can be defined. As an example, the frequency sensitivity levels can be defined as shown in Table 7 below.
[0262] [Table 7]
[0263]
[0264] As shown in Table 7, a range can be defined for each frequency sensitivity level. Each range can be defined by applying the mean and deviation of the frequency sensitivity, and through this, the distribution coefficient of the information used for each level can be identified. For example, when the high-level threshold is determined to be level 9, QP optimization can be performed on areas with frequency sensitivity values of approximately 180 or higher, indicating that QP optimization is only applied to approximately 2.3% of the entire area. By defining each frequency sensitivity level based on this distribution characteristic, a threshold can also be defined based on the frequency sensitivity level and bit amount optimization requirements required for each machine task.
[0265] Meanwhile, the following Tables 8 to 15 show examples of perceptual QO identifiers and frequency sensitivity level information according to embodiments of the present disclosure. In the following embodiments, the QO-related information described with reference to the above tables will be extended by additional definitions for information related to frequency sensitivity levels, and the frequency sensitivity level information and perceptual QO identification information can be 1) defined in a traditional NAL unit, 2) based on a NAL unit type defined for identifying a perceptual quantization optimization (QO) method, or 3) configured based on supplementary enhancement information (SEI) defined for identifying perceptual QO.
[0266] The following Table 8 and Table 9 may illustrate an example of additionally defining frequency sensitivity level information for the first syntax (eg, machine_perceptual_QO_flag) defined in the other tables above.
[0267] [Table 8]
[0268]
[0269] According to Table 8, based on the first syntax (e.g., machine_perceptual_QO_flag), the fifth syntax (e.g., threshold_coded_flag), the sixth syntax (e.g., low_level_threshold_idc), and / or the seventh syntax (e.g., high_level_threshold_idc) can be obtained. As an example, the sixth and seventh syntaxes can be further obtained based on the fifth syntax.
[0270] The first syntax machine_perceptual_QO_flag may indicate whether to apply encoding optimization for executing machine tasks. The explanation for this is the same as that described above with reference to other tables, and redundant descriptions are omitted.
[0271] The fifth syntax, threshold_coded_flag, may indicate whether the frequency sensitivity level, which is a reference value for determining the frequency sensitivity level of the encoding target, is encoded, and may be defined in the bitstream based on the value of the first syntax (e.g., 1). When the value of the fifth syntax is a first value (e.g., 1), it may indicate that the frequency sensitivity level has been encoded in the bitstream, and when the value is a second value (e.g., 0), it may indicate that the frequency sensitivity level has not been encoded.
[0272] The sixth syntax, low_level_threshold_idc, may be an identifier indicating a low-level threshold value and may be defined in the bitstream based on the value (e.g., 1) of the fifth syntax. As an example, the sixth syntax may be encoded to represent a value corresponding to a low-level threshold value among the frequency sensitivity levels defined in Table 7. Meanwhile, in the example of Table 7, the range of frequency sensitivity levels is defined as 1 to 10, and the range of the low-level threshold value may be defined as a subset of the frequency sensitivity level range. For example, the range of the low-level threshold value may be defined within the frequency sensitivity level range, such as 1 to 4.
[0273] The seventh syntax, high_level_threshold_idc, is an identifier indicating a high-level threshold value and can be defined in the bitstream based on the value (e.g., 1) of the fifth syntax. For example, the seventh syntax can be encoded to represent a value corresponding to a high-level threshold value among the frequency sensitivity levels defined in Table 7. As an example, in the example of Table 7, the frequency sensitivity level range is 1 to 10, and the range of the high-level threshold value can be defined as a subset of the frequency sensitivity level range. For example, the range of the high-level threshold value can be defined within the frequency sensitivity level range, such as 7 to 10.
[0274] Meanwhile, the following Table 9 signals the first syntax and the fifth syntax similarly to Table 8, but relates to an example in which the threshold value itself is signaled.
[0275] [Table 9]
[0276]
[0277] Because the contents of the first and fifth syntaxes are the same as those in Table 8 above, redundant descriptions are omitted. Meanwhile, based on the value of the fifth syntax, a threshold value (e.g., a low-level threshold value and / or a high-level threshold value) can be obtained. Rather than obtaining the threshold value as an indicator indicating the threshold value as in Table 8, the threshold value itself can be signaled in the bitstream. For example, when the value of the fifth syntax is a first value (e.g., 1), an eighth syntax value (e.g., low_level_threshold) indicating a low-level threshold value and / or a ninth syntax value (e.g., high_level_threshold) indicating a high-level threshold value can be obtained. The eighth syntax value, low_level_threshold, can indicate a low-level threshold value. While the sixth syntax value, low_level_threshold_idc, is intended to identify a range corresponding to a coding target within a defined range, the eighth syntax value, low_level_threshold, can indicate the actual value of the low-frequency sensitivity determination criterion for the coding target.
[0278] The ninth syntax high_level_threshold may indicate a high-level threshold. While the seventh syntax high_level_threshold_idc is intended to identify a range to which a coding target corresponds in a defined range, the ninth syntax high_level_threshold may represent an actual value of a high-frequency sensitivity judgment standard of the coding target.
[0279] The following Table 10 and Table 11 represent examples of an expression method in which frequency sensitivity level information is added to whether the perceptual QO is applied and usage according to an embodiment of the present disclosure.
[0280] [Table 10]
[0281]
[0282] Table 10 relates to an embodiment in which, in addition to the syntax in Table 8, a second syntax (e.g., perceptual_QO_flag) may be further obtained. For example, the second syntax may be obtained before the first syntax (e.g., machine_perceptual_QO_flag) is obtained, and the first syntax may be obtained based on the value of the second syntax. Furthermore, the fifth syntax (e.g., threshold_coded_flag) may be obtained based on the value of the first syntax. Subsequently, the sixth syntax (e.g., low_level_threshold_idc) and the seventh syntax (e.g., high_level_threshold_idc) may be obtained based on the value of the fifth syntax. Because the descriptions of the first syntax and other syntaxes are the same as those of the syntaxes with the same names included in the other tables above, redundant descriptions are omitted.
[0283] Meanwhile, Table 11 below similarly signals the first syntax and the fifth syntax as in Table 10, but relates to an example in which the threshold value itself is signaled.
[0284] [Table 11]
[0285]
[0286] The descriptions of the first and fifth syntaxes are the same as those in Table 10 above, and therefore, redundant explanations are omitted. Furthermore, based on the value of the fifth syntax, thresholds (e.g., low-level threshold and / or high-level threshold) can be obtained, and the thresholds themselves can be directly signaled in the bitstream, rather than being obtained as information in the form of indicators indicating the thresholds as in Table 10. For example, when the value of the fifth syntax is the first value (e.g., 1), an eighth syntax (e.g., low_level_threshold) indicating the low-level threshold and / or a ninth syntax (e.g., high_level_threshold) indicating the high-level threshold can be obtained. The descriptions of the eighth through ninth syntaxes are the same as those described with reference to another table (e.g., Table 9), and therefore, redundant explanations are omitted. Tables 12 and 13 below respectively illustrate examples of additionally defining frequency sensitivity level information for the third syntax (e.g., human_perceptual_QO_flag), which is defined in another table instead of the first syntax (e.g., machine_perceptual_QO_flag) in the embodiments of Tables 10 and 11.
[0287] [Table 12]
[0288]
[0289] For example, the second syntax (e.g., perceptual_QO_flag) may be obtained before the third syntax (e.g., human_perceptual_QO_flag), and the third syntax may be obtained based on the value of the second syntax. Furthermore, the fifth syntax (e.g., threshold_coded_flag) may be obtained based on the value of the third syntax. Subsequently, the sixth syntax (e.g., low_level_threshold_idc) and the seventh syntax (e.g., high_level_threshold_idc) may be obtained based on the value of the fifth syntax. The description of the third syntax (e.g., human_perceptual_QO_flag) and other syntaxes is the same as that of the syntaxes with the same names described with reference to the other tables above, and therefore redundant explanations are omitted.
[0290] Meanwhile, Table 13 below illustrates an example similar to Table 12, in which the second syntax, the third syntax, and the fifth syntax are signaled, but the threshold value itself is signaled.
[0291] [Table 13]
[0292]
[0293] The descriptions of the second, third, and fifth syntaxes are the same as those described in Table 12 above, and therefore, redundant explanations are omitted. Furthermore, based on the value of the fifth syntax, a threshold value (e.g., a low-level threshold value and / or a high-level threshold value) can be obtained. These threshold values may not be obtained as indicators indicating the threshold values as in Table 12, but rather the threshold values themselves may be signaled in the bitstream. For example, when the value of the fifth syntax is a first value (e.g., 1), an eighth syntax value (e.g., low_level_threshold) indicating a low-level threshold value and / or a ninth syntax value (e.g., high_level_threshold) indicating a high-level threshold value may be obtained. The descriptions of the eighth through ninth syntaxes are the same as those described with reference to other tables (e.g., Table 9), and therefore, redundant explanations are omitted. Tables 14 and 15 below respectively illustrate examples in which, in addition to applying perceptual quantization optimization (perceptualQO) and the optimization method defined in another table, information related to frequency sensitivity levels is additionally defined.
[0294] Meanwhile, the following Table 15 illustrates an example similar to Table 14 in which the second syntax, the third syntax, and the fifth syntax are signaled, but the threshold value itself is signaled.
[0295] [Table 14]
[0296]
[0297] According to the embodiments shown in the table, the second syntax (e.g., perceptual_QO_flag) can be obtained first. Then, information about the quantization optimization method (e.g., the fourth syntax) and the second syntax can be obtained. For example, the fourth syntax can be obtained based on the second syntax. For example, the fourth syntax can be called optimization_method_idc and can be an optimization method indicator. Subsequently, based on the value of the fourth syntax, the fifth, sixth, and / or seventh syntax can be obtained. For example, when the value of the fourth syntax is a first value (e.g., 1), the sixth syntax (e.g., low_level_threshold_idc) can be obtained, and when the value of the fourth syntax is a second value (e.g., 2), the seventh syntax (e.g., high_level_threshold_idc) can be obtained. Conversely, when the value of the fourth syntax is a third value (e.g., 3), both the sixth and seventh syntaxes can be obtained. In other cases, when the value of the fourth syntax is a fourth value (e.g., 0), the fifth syntax may not be obtained. The descriptions of the fourth syntax and other syntaxes are the same as those of the syntaxes with the same names described in the other tables above, and therefore redundant explanations are omitted.
[0298] Meanwhile, the following Table 15 illustrates an example similar to Table 14, in which the second syntax, the fourth syntax, and the fifth syntax are signaled, but the threshold value itself is signaled.
[0299] [Table 15]
[0300]
[0301] The descriptions of the second, fourth, and fifth syntaxes are the same as those in Table 14 above, and therefore redundant explanations are omitted. Meanwhile, based on the value of the fifth syntax, a threshold value (e.g., a low-level threshold value and / or a high-level threshold value) can be obtained, and the threshold value may not be obtained in the form of an indicator indicating the threshold value as in Table 14, but the threshold value itself may be signaled in the bitstream. As an example, when the value of the fifth syntax is a first value (e.g., 1), an eighth syntax (e.g., low_level_threshold) indicating a low-level threshold value and / or a ninth syntax (e.g., high_level_threshold) indicating a high-level threshold value may be obtained. The descriptions of the eighth and ninth syntaxes are the same as those described with reference to other tables (e.g., Table 9), and therefore redundant explanations are omitted. According to the present disclosure, information that may vary depending on the type of task to be performed (e.g., a task based on machine perception, a task based on human perception, etc.) can be efficiently processed.
[0302] Figure 20 is a diagram illustrating an image (data) decoding method according to an embodiment of the present disclosure. Figure 20 The image (data) decoding method may be performed by the above-mentioned image (data) decoder, and may further include a process of performing inverse quantization based on the quantization optimization method described above with reference to the table, etc.
[0303] As an example, image usage information for image data can be obtained from a bitstream (S2010). Then, by performing inverse quantization based on the image usage information, the image data can be reconstructed (S2020). As an example, the image usage information can indicate at least one of whether the image data is intended for human perception or machine perception. The image usage information can include the first, second, and / or third syntaxes described above. Based on the image usage information indicating that the image data is intended for machine processing, a frequency sensitivity level reference value for the image data can be derived. As an example, the frequency sensitivity level reference value can be derived based on an inverse quantization method indicator, and the inverse quantization method indicator can include the fourth syntax described above. The frequency sensitivity level reference value can also include at least one of a low-level frequency sensitivity threshold or a high-level frequency sensitivity threshold, and the frequency sensitivity level reference value can include at least one of the sixth to ninth syntaxes described above. As an example, the frequency sensitivity level reference value can be signaled from the bitstream. Furthermore, the frequency sensitivity level reference value can be obtained based on a frequency sensitivity encoding flag indicating whether frequency sensitivity is encoded, and the frequency sensitivity encoding flag can include the fifth syntax described above. As an example, the frequency sensitivity encoding flag can be obtained based on the image usage information.
[0304] As an example, Figure 20 The embodiment may include an inverse quantization process based on the embodiment described with reference to the above table, and thus redundant explanations are omitted.
[0305] at the same time, Figure 20 The embodiments correspond to the embodiments of the present disclosure, and therefore, some steps may be modified or omitted, or the order of some steps may be changed, or additional steps may be included, and such modifications are also considered to fall within the scope of protection of the present disclosure.
[0306] Figure 21 is a diagram illustrating an image (data) encoding method according to an embodiment of the present disclosure. Figure 21The image data encoding method described above can be performed by the image (data) encoder and may further include performing a quantization process based on the quantization optimization method described above in the reference table. The quantization process can be performed on the image data by determining an image usage for the image data ( S2110 ). Image usage information indicating the image usage can then be encoded ( S2120 ). For example, the image usage information can indicate at least one of whether the image data is intended for human perception or machine perception and can include the first, second, and / or third syntaxes described above. Based on the image data being processed by the machine, a frequency sensitivity level reference value for the image data can be determined. For example, the frequency sensitivity level reference value can be derived based on a quantization method indicator, and the quantization method indicator can include the fourth syntax described above. Furthermore, quantization of the image data can be performed based on the frequency sensitivity level reference value for the image data. The frequency sensitivity reference level value can include at least one of a low-level frequency sensitivity threshold or a high-level frequency sensitivity threshold, and can include at least one of the sixth through ninth syntaxes described above. Furthermore, the frequency sensitivity reference level value can be encoded into the bitstream. For example, the frequency sensitivity reference level value can be signaled by encoding it into the bitstream. In addition, the frequency sensitivity level reference value can be determined based on whether frequency sensitivity is encoded, and the encoding status of the frequency sensitivity can be represented as a frequency sensitivity encoding flag, which can be encoded into the bitstream. The frequency sensitivity encoding flag can include the fifth syntax described above. As an example, whether frequency sensitivity is encoded can be determined based on the image usage.
[0307] Meanwhile, a computer-readable medium for storing a bit stream generated by the above-mentioned image data encoding method may be proposed.
[0308] In addition, the bit stream generated by the above-mentioned image data encoding method can be transmitted to an image (data) decoder, etc. The bit stream transmission method may include the step of transmitting the bit stream generated by the image data encoding method.
[0309] As an example, Figure 21 The embodiment may include the quantization process of the embodiment described based on the above reference table, and thus redundant explanation thereof is omitted.
[0310] At the same time, because Figure 21 The embodiments correspond to the embodiments of the present disclosure, so some steps may be changed or omitted, the order of some steps may be changed, or other steps may be further added, and this also corresponds to the embodiments of the present disclosure.
[0311] For ease of description, the names of all the above syntax elements are arbitrarily assigned and do not limit the names of the corresponding syntax elements. In addition, the syntax elements can be respectively referred to as information with different names. In addition, some syntax elements can be obtained from the bitstream, while other syntax elements can be derived from other syntax elements, and this can also be included in the embodiments of the present disclosure.
[0312] In addition, a bitstream generated by the image encoding method may be stored in a non-transitory computer-readable recording medium.
[0313] In addition, as another example, the bitstream generated by the image encoding method can be transmitted to other devices (eg, image decoding devices, etc.). In this case, the bitstream transmission method may include a process of transmitting the bitstream.
[0314] Although the exemplary methods of the present disclosure are expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order if necessary. To implement the methods according to the present disclosure, another step may be additionally included in the exemplary steps, or the remaining steps may be included while excluding some steps, or another additional step may be included while excluding some steps.
[0315] In the present disclosure, an image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) for checking a condition or perform a corresponding operation (step). For example, when a statement states that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or image decoding device may perform an operation for checking whether the predetermined condition is satisfied and then perform the predetermined operation.
[0316] The various embodiments of the present disclosure do not list all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments can be applied independently or in combination of at least two. For example, each of the embodiments described above with reference to the drawings and tables can be combined or used independently.
[0317] The embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information about instructions) or algorithms used for implementation can be stored in a digital storage medium.
[0318] In addition, the decoders (decoding devices) and encoders (encoding devices) to which the embodiments of the present disclosure are applied may be included in multimedia broadcast transmission and reception devices, mobile communication terminals, home theater video equipment, digital theater video equipment, surveillance cameras, video communication equipment, real-time communication equipment such as video communication, etc., mobile streaming devices, storage media, cameras, video on demand (VoD) service providers, OTT video (over-the-top) devices, Internet streaming service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony video devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, robot terminals, aircraft terminals, ship terminals, etc.), medical video equipment, etc., and may be used to process video signals or data signals. For example, OTT video (over-the-top video) devices may include game consoles, Blu-ray players, Internet-connected televisions, home theater systems, smartphones, tablet computers, digital video recorders (DVRs), etc.
[0319] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission via the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or can be sent via a wired or wireless communication network.
[0320] In addition, the embodiments of the present disclosure may be implemented as a computer program product through program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0321] Figure 22 is a diagram illustrating an example of a content streaming system to which an embodiment of the present disclosure can be applied.
[0322] refer to Figure 22 The content streaming system to which the embodiments of the present disclosure are applied may broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0323] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data, generates a bitstream, and sends it to the streaming server. As another example, when the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted.
[0324] A bitstream may be generated by applying the image encoding method and / or the image encoding apparatus according to the embodiments of the present disclosure, and a streaming server may temporarily store the bitstream during a process of transmitting or receiving the bitstream.
[0325] The streaming server can transmit multimedia data to a user device via a web server based on a user request, and the web server can act as an intermediary to inform the user of available services. When the user requests a desired service from the web server, the web server can transmit it to the streaming server, and the streaming server can transmit the multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server can play a role in controlling commands and responses between devices within the content streaming system.
[0326] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a seamless streaming service, the streaming server can store the bitstream for a certain period of time.
[0327] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (i.e., smart watches, smart glasses, head-mounted displays (HMDs)), digital televisions, desktop computers, digital signage, etc.
[0328] Each server within the content streaming system can operate as a distributed server, in which case data received by each server can be processed in a distributed manner.
[0329] Figure 23 is a diagram illustrating another example of a content streaming system to which an embodiment of the present disclosure can be applied.
[0330] refer to Figure 23In an embodiment such as VCM, a task may be performed by a user terminal, or the task may be performed by an external device (e.g., a streaming server, an analysis server, etc.) according to the performance of the device, a user's request, characteristics of the task to be performed, etc. In this way, in order to send information necessary for performing the task to the external device, the user terminal may generate a bitstream directly or through an encoding server, the bitstream including information necessary for performing the task (e.g., information such as the task, the neural network, and / or usage).
[0331] The analysis server can perform the task requested by the user after decoding the encoded information sent from the user terminal (or from the encoding server). The analysis server can then retransmit the results obtained from performing the task to the user terminal or another linked service server (e.g., a web server). For example, the analysis server can transmit the results obtained from performing a task to determine if a fire has occurred to a fire-related server. The analysis server may include a separate control server, in which case the control server can play a role in controlling the command / response between the analysis server and each device associated with the server. Furthermore, the analysis server can request desired information from the web server based on information about the task the user device wants to perform and the tasks that the user device can perform. When the analysis server requests a desired service from the web server, the web server can transmit the requested service to the analysis server, and the analysis server can transmit the requested service data to the user terminal. In this case, the control server of the content streaming system can play a role in controlling the command / response between each device within the streaming system.
[0332] [Industrial Applicability]
[0333] The embodiments according to the present disclosure can be used for encoding / decoding images.
Claims
1. A method for decoding image data performed by an image data decoding apparatus, comprising: obtaining image usage information for the image data from a bitstream; as well as reconstructing the image data by performing inverse quantization based on the image usage information, The image usage information indicates whether the image data is used for at least one of human perception and machine perception.
2. The image data decoding method according to claim 1, wherein: Based on the image usage information indicating that the image data is machine-processed, a frequency sensitivity level reference value of the image data is derived.
3. The image data decoding method according to claim 2, wherein: The frequency sensitivity level reference value is derived based on an inverse quantization method indicator.
4. The image data decoding method according to claim 2, wherein: The frequency sensitivity reference level value includes at least one of a low-level frequency sensitivity threshold or a high-level frequency sensitivity threshold.
5. The image data decoding method according to claim 2, wherein: The frequency sensitivity reference level value is signaled from a bitstream.
6. The image data decoding method according to claim 1, wherein: The frequency sensitivity level reference value is obtained based on a frequency sensitivity encoding flag indicating whether frequency sensitivity is encoded.
7. The image data decoding method according to claim 6, wherein: The frequency sensitivity encoding flag is obtained based on the image usage information.
8. A method for encoding image data performed by an image data encoding apparatus, comprising: performing quantization on the image data by determining an image usage for the image data; as well as encoding image usage information indicating the usage of the image, The image usage information indicates whether the image data is used for at least one of human perception and machine perception.
9. The image data encoding method according to claim 8, wherein: The quantization on the image data is further performed based on a frequency sensitivity level reference value of the image data.
10. The image data encoding method according to claim 9, wherein: The frequency sensitivity reference level value includes at least one of a low-level frequency sensitivity threshold or a high-level frequency sensitivity threshold.
11. The image data encoding method according to claim 9, wherein: The frequency sensitivity reference level value is encoded into a bit stream.
12. A computer-readable medium storing a bit stream generated by an image data encoding method, wherein: The image data encoding method comprises: performing quantization on the image data by determining an image usage for the image data; and encoding image usage information indicating the usage of the image, The image usage information indicates whether the image data is used for at least one of human perception and machine perception.
13. A bit stream transmission method, comprising: Sending the bit stream generated by the image data encoding method, The image data encoding method includes: performing quantization on the image data by determining an image usage for the image data; and encoding image usage information indicating the usage of the image, The image usage information indicates whether the image data is used for at least one of human perception and machine perception.