Channel skip and reconstruction method
The channel skipping and restoration method addresses the inefficiencies in existing image compression technologies for machine-dependent image analysis by using a VCM decoding device to restore omitted channels and a VCM encoding device to efficiently encode feature maps, resulting in effective image analysis and reduced resource consumption.
Patent Information
- Application Number
- PCT/KR2024/018163
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2024-11-18
- Publication Date
- 2025-06-12
AI Technical Summary
As the demand for higher-quality images and videos, such as 4K or 8K UHD, increases, existing image compression technologies are inefficient for machine-dependent image analysis, leading to issues with server load and power consumption.
A channel skipping and restoration method is introduced, which includes a VCM decoding device that decodes a bitstream to generate a feature map and restore omitted channels based on parsed information, and a VCM encoding device that performs temporal resampling, redundancy removal, truncation, and quantization to efficiently encode feature maps.
This method enables effective image analysis by machines, reducing server load and power consumption while maintaining high image quality, by efficiently compressing and restoring feature maps.
Smart Images

Figure KR2024018163_12062025_PF_FP_ABST
Abstract
Description
How to skip and restore channels
[0001] The present disclosure relates to a channel skipping and restoration method in an encoding / decoding method for a machine.
[0002] With the continuous development of the information and communication industry, broadcasting services with HD (High Definition) resolution have spread worldwide.
[0003] Through this proliferation, many users have become accustomed to high-resolution and high-quality images and / or videos, and the demand for higher-resolution and high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, has increased in various fields.
[0004] The technology for coding this UHD video data was completed in 2013 through HEVC (High Efficiency Video Coding), a standard technology.
[0005] HEVC is a next-generation video compression technology with a higher compression ratio and lower complexity than the previous H.264 / AVC technology, and is a key technology for effectively compressing the massive data of HD and UHD video.
[0006] HEVC performs block-by-block encoding, like previous compression standards.
[0007] However, unlike H.264 / AVC, there is only one profile. The core encoding technologies included in HEVC's sole profile are divided into eight areas: hierarchical encoding structure technology, transform technology, quantization technology, intra-frame prediction encoding technology, inter-frame motion prediction technology, entropy encoding technology, loop filter technology, and other technologies.
[0008] Since the establishment of the HEVC video codec in 2013, the Versatile Video Coding (VVC) standard, a next-generation video codec that aims to improve performance by more than twice that of HEVC, has been developed to address the expansion of realistic video and virtual reality services utilizing 4K and 8K video images. VVC is called H.266.
[0009] H.266 (VVC) was developed with the goal of being more than twice as efficient as the previous generation codec, H.265 (HEVC). VVC was initially developed with resolutions over 4K in mind, but it was also developed for ultra-high-resolution video processing at a whopping 16K level to support 360-degree videos due to the expansion of the VR market. In addition, as the HDR market is expanding due to the development of display technology, it supports 16-bit color depth as well as 10-bit color depth to respond to this, and supports brightness expressions of 1000 nits, 4000 nits, and 10000 nits. In addition, since it is being developed with the VR market and 360-degree video market in mind, it supports partial frame rates in the range of 0 to 120 FPS.
[0010] Advances in Artificial Intelligence
[0011] Artificial intelligence (AI) is also steadily developing. AI refers to the artificial imitation of human intelligence, including the ability to recognize, classify, infer, predict, and control / decision-making.
[0012] With the advancement of artificial intelligence technology and the increase in Internet of Things (IoT) devices, machine-to-machine traffic is expected to explode, and machine-dependent image analysis is expected to become widely used.
[0013] However, as the amount of images to be analyzed by machines is expected to increase exponentially, issues regarding server load and power consumption are expected to arise.
[0014] Accordingly, the present disclosure aims to provide a channel skipping and restoration method to enable effective machine-based image analysis.
[0015] In order to achieve the above-mentioned purpose, according to one disclosure of the present specification, a channel skipping and restoration method is disclosed.
[0016] A VCM decoding device according to one disclosure of the present specification comprises: an internal decoding performer that decodes a bitstream to generate a feature map, and parses information about the feature map by decoding the bitstream; and an image restoration performer that restores one or more channels omitted in the decoded feature map based on information about the feature map, wherein the information about the feature map may include at least one of group-specific coding information including whether to skip each channel of the feature map, quantization compensation parameters, feature map truncation information, temporal restoration information, and packing information.
[0017] A VCM encoding device according to one disclosure of the present specification may include a temporal resampling performer that changes a frame rate for a feature map; a feature map redundancy remover that removes redundancy for a spatial axis or a channel axis for each frame or a series of frames for the feature map; a feature map truncation performer that truncates a portion of the feature map for each frame or a series of frames for the feature map based on quantization information; a quantization performer that performs quantization for the feature map based on the quantization information; and an internal encoding performer that encodes at least one of temporal restoration information according to the temporal resampling process, group-wise coding information according to the feature map redundancy removal process, feature map truncation information according to the feature map truncation process, and quantization compensation parameters according to the quantization process of the feature map, and the feature map to generate a bitstream.
[0018] A non-volatile computer-readable storage medium having recorded thereon instructions according to one disclosure of the present specification, wherein the instructions, when executed by one or more processors, cause the one or more processors to: decode a bitstream to generate a feature map, decode the bitstream to parse information about the feature map, and restore one or more channels omitted in the decoded feature map based on the information about the feature map, wherein the information about the feature map may include at least one of group-specific coding information including whether to skip each channel of the feature map, quantization compensation parameters, feature map truncation information, temporal restoration information, and packing information.
[0019] According to the present disclosure, image analysis by a machine can be effectively performed.
[0020] Figure 1 schematically illustrates an example of a video / image coding system.
[0021] Figure 2 is a drawing schematically illustrating the configuration of a video / image encoding device.
[0022] Figure 3 is a drawing schematically illustrating the configuration of a video / image decoding device.
[0023] Figures 4a to 4d are exemplary diagrams showing a VCM encoder and a VCM decoder.
[0024] FIG. 5 is a block diagram of an encoding device according to an embodiment of the present disclosure.
[0025] FIG. 6 illustrates the detailed operation of a feature map deduplication remover according to an embodiment of the present disclosure.
[0026] FIG. 7 illustrates detailed operations of a quantization performer according to an embodiment of the present disclosure.
[0027] FIG. 8 illustrates detailed operations of an internal encoding performer according to an embodiment of the present disclosure.
[0028] FIG. 9 is a block diagram of a decoding device according to an embodiment of the present disclosure.
[0029] FIG. 10 illustrates detailed operations of an internal decryption performer according to one embodiment of the present disclosure.
[0030] FIG. 11 illustrates the detailed operation of a feature map redundancy restorer according to an embodiment of the present disclosure.
[0031] Specific structural or step-by-step descriptions of embodiments according to the concept of the present disclosure disclosed in this specification or application are merely illustrative for the purpose of explaining embodiments according to the concept of the present disclosure, and embodiments according to the concept of the present disclosure may be implemented in various forms, and embodiments according to the concept of the present disclosure may be implemented in various forms and should not be construed as being limited to the embodiments described in this specification or application.
[0032] Embodiments according to the concept of the present disclosure may have various modifications and take various forms. Therefore, specific embodiments are illustrated in the drawings and described in detail in this specification or application. However, this is not intended to limit embodiments according to the concept of the present disclosure to specific disclosed forms, and it should be understood that all modifications, equivalents, and alternatives included within the spirit and technical scope of the present disclosure are included.
[0033] While terms such as "first" and / or "second" may be used to describe various components, these components should not be limited by these terms. These terms are only intended to distinguish one component from another; for example, without departing from the scope of the present disclosure, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component."
[0034] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. Conversely, when a component is referred to as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Other expressions that describe the relationship between components, such as "between" and "directly between" or "adjacent to" and "directly adjacent to", should be interpreted similarly.
[0035] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the present disclosure. The singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0036] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0037] Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined herein.
[0038] In describing the embodiments, description of technical contents that are well known in the technical field to which the present disclosure belongs and are not directly related to the present disclosure will be omitted.
[0039] This is to convey the gist of the present disclosure more clearly without obscuring it by omitting unnecessary explanations.
[0040] This document relates to video / image coding. For example, the method / embodiment disclosed in this document may be related to the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), the essential video coding (EVC) standard, the AVS2 standard, etc.).
[0041] This document presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0042] In this document, "video" can refer to a series of images over time. "Picture" generally refers to a unit representing a single image from a specific time period, and "slice" / "tile" are units that constitute part of a picture in coding.
[0043] A slice / tile can contain one or more coding tree units (CTUs). A picture can consist of one or more slices / tiles. A picture can consist of one or more tile groups. A tile group can contain one or more tiles.
[0044] A pixel or pel can mean the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Alternatively, a sample can mean a pixel value in the spatial domain, or when such a pixel value is converted to the frequency domain, it can mean a transform coefficient in the frequency domain.
[0045] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region.
[0046] A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term "unit" may sometimes be used interchangeably with the terms "block" or "area." In general, an MxN block can contain a set (or array) of samples (or array of samples) or transform coefficients, each consisting of M columns and N rows.
[0047] Figure 1 schematically illustrates an example of a video / image coding system.
[0048] Referring to FIG. 1, a video / image coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or streaming.
[0049] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a reception unit, a decoding device, and a renderer.
[0050] The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0051] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0052] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0053] The transmission unit can transmit encoded video / image information or data output in bitstream form to the receiving unit of the receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file using a predetermined file format and an element for transmission via a broadcasting / communication network.
[0054] The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0055] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0056] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0057] Figure 2 is a drawing schematically illustrating the configuration of a video / image encoding device.
[0058] The term “video encoding device” hereinafter may include a video encoding device.
[0059] Referring to FIG. 2, the encoding device (10a) may be configured to include an image partitioner (10a-10), a prediction unit (predictor) (10a-20), a residual processor (residual processor) (10a-30), an entropy encoder (entropy encoder) (10a-40), an adder (adder) (10a-50), a filter (filter) (10a-60), and a memory (10a-70). The prediction unit (10a-20) may include an inter prediction unit (10a-21) and an intra prediction unit (10a-22). The residual processing unit (10a-30) may include a transformer (10a-32), a quantizer (10a-33), a dequantizer (10a-34), and an inverse transformer (10a-35). The residual processing unit (10a-30) may further include a subtractor (10a-31). The addition unit (10a-50) may be called a reconstructor or a reconstructed block generator. The above-described image segmentation unit (10a-10), prediction unit (10a-20), residual processing unit (10a-30), entropy encoding unit (10a-40), addition unit (10a-50), and filtering unit (10a-60) may be configured by one or more hardware components (e.g., encoder chipset or processor) according to an embodiment. In addition, the memory (10a-70) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (10a-70) as an internal / external component.
[0060] The image segmentation unit (10a-10) can segment an input image (or picture, frame) input to the encoding device (10a) into one or more processing units.
[0061] For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a Quad-tree binary-tree ternary-tree (QTBTTT) structure. For example, one coding unit may be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary structure may be applied later. Alternatively, the binary-tree structure may be applied first. The coding procedure according to the present document may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0062] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0063] The subtraction unit (10a-31) can subtract the prediction signal (predicted block, prediction samples, or prediction sample array) output from the prediction unit (10a-20) from the input image signal (original block, original samples, or original sample array) to generate a residual signal (residual block, residual samples, or residual sample array), and the generated residual signal is transmitted to the conversion unit (10a-32). The prediction unit (10a-20) can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block.
[0064] The prediction unit (10a-20) can determine whether intra-prediction or inter-prediction is applied to the current block or CU unit. As described later in the description of each prediction mode, the prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit (10a-40). The prediction-related information can be encoded by the entropy encoding unit (10a-40) and output in the form of a bitstream.
[0065] The intra prediction unit (10a-22) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode.
[0066] In intra prediction, prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes and planar modes. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the granularity of the prediction direction.
[0067] However, this is only an example; depending on the settings, a greater or lesser number of directional prediction modes may be used. The intra prediction unit (10a-22) may also determine the prediction mode to be applied to the current block by utilizing the prediction mode applied to the surrounding blocks.
[0068] The inter prediction unit (10a-21) can derive a predicted block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit (10a-21) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (10a-21) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0069] The prediction unit (10a-20) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction to predict a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) to predict a block. The intra block copy can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.
[0070] The prediction signal generated through the inter prediction unit (10a-21) and / or the intra prediction unit (10a-22) can be used to generate a reconstructed signal or a residual signal. The transform unit (10a-32) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transform process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size.
[0071] The quantization unit (10a-33) quantizes the transform coefficients and transmits them to the entropy encoding unit (10a-40), and the entropy encoding unit (10a-40) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information.
[0072] The quantization unit (10a-33) can rearrange the quantized transform coefficients in the form of a block into a one-dimensional vector based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of the one-dimensional vector. The entropy encoding unit (10a-40) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc.
[0073] The entropy encoding unit (10a-40) may encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted through a network or may be stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (10a-40) may be configured as an internal / external element of the encoding device (10a) by a transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal, or the transmitting unit may be included in the entropy encoding unit (10a-40).
[0074] The quantized transform coefficients output from the quantization unit (10a-33) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (10a-34) and the inverse transform unit (10a-35), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (10a-50) can add the reconstructed residual signal to the prediction signal output from the prediction unit (10a-20), thereby generating a reconstructed signal (reconstructed picture, reconstructed block, reconstructed samples, or reconstructed sample array). When there is no residual for the target block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0075] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0076] The filtering unit (10a-60) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (10a-60) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (10a-70), specifically, in the DPB of the memory (10a-70). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), an adaptive loop filter, a bilateral filter, etc. The filtering unit (10a-60) can generate various information regarding filtering and transmit the information to the entropy encoding unit (10a-90), as described below in the description of each filtering method. The information regarding filtering may be encoded by the entropy encoding unit (10a-90) and output in the form of a bitstream.
[0077] The modified restored picture transmitted to the memory (10a-70) can be used as a reference picture in the inter prediction unit (10a-80). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (10a) and the decoding device, and can also improve encoding efficiency.
[0078] The DPB of the memory (10a-70) can store the modified restored picture to be used as a reference picture in the inter prediction unit (10a-21). The memory (10a-70) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transmitted to the inter prediction unit (10a-21) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (10a-70) can store restored samples of restored blocks in the current picture and transmit them to the intra prediction unit (10a-22).
[0079] Figure 3 is a drawing schematically illustrating the configuration of a video / image decoding device.
[0080] Referring to FIG. 3, the decoding device (10b) may be configured to include an entropy decoder (10b-10), a residual processor (10b-20), a predictor (10b-30), an adder (10b-40), a filter (10b-50), and a memory (10b-60). The predictor (10b-30) may include an inter-prediction unit (10b-31) and an intra-prediction unit (10b-32). The residual processor (10b-20) may include a dequantizer (10b-21) and an inverse transformer (10b-21). The entropy decoding unit (10b-10), residual processing unit (10b-20), prediction unit (10b-30), addition unit (10b-40), and filtering unit (10b-50) described above may be configured by a single hardware component (e.g., decoder chipset or processor) according to an embodiment. In addition, the memory (10b-60) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (10b-60) as an internal / external component.
[0081] When a bitstream including video / image information is input, the decoding device (10b) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (10b) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (10b) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (10b) can be reproduced through a reproduction device.
[0082] The decoding device (10b) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (10b-10). For example, the entropy decoding unit (10b-10) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information.
[0083] The decoding device can further decode the picture based on information about the parameter set and / or the general restriction information. The signaling / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through the decoding procedure. For example, the entropy decoding unit (10b-10) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients for the residual.
[0084] In more detail, the CABAC entropy decoding method receives a bin corresponding to each syntax element in a bitstream, determines a context model using information of the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of symbols / bins decoded in a previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit (10b-10), information regarding prediction is provided to the prediction unit (10b-30), and information regarding the residual on which entropy decoding has been performed by the entropy decoding unit (10b-10), i.e., quantized transform coefficients and related parameter information, can be input to the inverse quantization unit (10b-21).
[0085] In addition, information regarding filtering among the information decoded by the entropy decoding unit (10b-10) may be provided to the filtering unit (10b-50). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (10b), or the receiving unit may be a component of the entropy decoding unit (10b-10). Meanwhile, the decoding device according to the present document may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The above information decoder may include the entropy decoding unit (10b-10), and the sample decoder may include at least one of the inverse quantization unit (10b-21), the inverse transformation unit (10b-22), the prediction unit (10b-30), the addition unit (10b-40), the filtering unit (10b-50), and the memory (10b-60).
[0086] The inverse quantization unit (10b-21) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (10b-21) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (10b-21) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0087] In the inverse transform unit (10b-22), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0088] The prediction unit can perform a prediction for the current block and generate a predicted block including prediction samples for the current block.
[0089] The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information about the prediction output from the entropy decoding unit (10b-10), and can determine a specific intra / inter prediction mode.
[0090] The prediction unit can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction to predict a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) to predict a block. The intra block copy can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.
[0091] The intra prediction unit (10b-32) can predict the current block by referencing samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode.
[0092] In intra prediction, prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit (10b-32) may determine the prediction mode to be applied to the current block by utilizing the prediction modes applied to the surrounding blocks.
[0093] The inter prediction unit (10b-31) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.).
[0094] In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (10b-31) may construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information regarding the prediction may include information indicating the mode of inter prediction for the current block.
[0095] The addition unit (10b-40) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (10b-30). In cases where there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.
[0096] The addition unit (10b-40) may be called a restoration unit or a restoration block generation unit.
[0097] The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, can be output after filtering as described below, or can be used for inter prediction of the next picture.
[0098] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0099] The filtering unit (10b-50) can improve subjective / objective image quality by applying filtering to the restoration signal. For example, the filtering unit (10b-50) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and transmit the modified restoration picture to the memory (60), specifically, the DPB of the memory (10b-60). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0100] The (corrected) reconstructed picture stored in the DPB of the memory (10b-60) can be used as a reference picture in the inter prediction unit (10b-31). The memory (10b-60) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transmitted to the inter prediction unit (10b-31) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (10b-60) can store reconstructed samples of reconstructed blocks within the current picture and transmit them to the intra prediction unit (10b-32).
[0101] In this specification, the embodiments described in the prediction unit (10b-30), the inverse quantization unit (10b-21), the inverse transformation unit (10b-22), and the filtering unit (10b-50) of the decoding device (10b) can be applied to the prediction unit (10a-20), the inverse quantization unit (10a-34), the inverse transformation unit (10a-35), and the filtering unit (10a-60) of the encoding device (10a) in the same manner or correspondingly.
[0102] As described above, prediction is performed to increase compression efficiency when performing video coding. Through this, a predicted block including prediction samples for a current block, which is a coding target block, can be generated. Here, the predicted block includes prediction samples in a spatial domain (or pixel domain). The predicted block is derived identically from an encoding device and a decoding device, and the encoding device can increase video coding efficiency by signaling information (residual information) about the residual between the original block and the predicted block, rather than the original sample value of the original block itself, to a decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by combining the residual block and the predicted block, and can generate a reconstructed picture including the reconstructed blocks.
[0103] The above residual information can be generated through transformation and quantization procedures.
[0104] For example, the encoding device can derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (a residual sample array) included in the residual block to derive transform coefficients, perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and signal related residual information to a decoding device (via a bitstream). Here, the residual information can include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device can perform an inverse quantization / inverse transform procedure based on the residual information to derive residual samples (or residual blocks). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also inverse quantize / inverse transform the quantized transform coefficients to derive a residual block for reference in inter prediction of a subsequent picture, and generate a reconstructed picture based on the residual block.
[0105] <VCM(Video coding for Machines)>
[0106] With the recent advancements in various industries such as surveillance, intelligent transportation, smart cities, intelligent industry, and intelligent content, the amount of image or feature map data consumed by machines is increasing. In contrast, traditional video compression methods currently in use were developed with human vision in mind, and therefore contain unnecessary information, making them inefficient for machine tasks. For example, the resolution of images from the viewer's perspective may be higher than that of images (e.g., feature maps) from the machine's perspective. Therefore, research on video codec technologies that efficiently compress feature maps for machine tasks is needed.
[0107] The Moving Picture Experts Group (MPEG), an international standardization group for multimedia encoding, is discussing Video Coding for Machines (VCM). VCM is an image or feature map encoding technology that targets machine vision, rather than human viewer vision. In this document, feature maps can be referred to as "feature maps," and features can be referred to as "features."
[0108] Figures 4a to 4d are exemplary diagrams showing a VCM encoder and a VCM decoder.
[0109] Referring to FIG. 4a, a VCM encoder (100a) and a VCM decoder (100b) are shown.
[0110] When a VCM encoder (100a) encodes a video and / or a feature map and transmits it as a bitstream, a VCM decoder (100b) can decode and output the bitstream. At this time, the VCM decoder (100b) can output one or more videos and / or feature maps. For example, the VCM decoder (100b) can output a first feature map for machine-based analysis and a first image for user viewing. The first image can have a higher resolution than the first feature map.
[0111] Referring to FIG. 4b, a feature extractor for extracting a feature map may be connected to the front end of the VCM encoder (100a).
[0112] The VCM encoder (100a) may include a feature encoder.
[0113] The VCM decoder (100b) may include a feature decoder and a video reconstructor. The feature decoder may decode a feature map from a bitstream and output a first feature map for machine-assisted analysis. The video reconstructor may regenerate and output a first video from the bitstream for viewing by a user.
[0114] Referring to Fig. 4c, a feature extractor for extracting a feature map is connected to the front end of the VCM encoder (100). The VCM encoder (100a) may include a feature encoder.
[0115] The VCM decoder (100b) may include a feature decoder. The feature decoder may decode a feature map from a bitstream and output a first feature map for machine-based analysis. That is, the bitstream may be encoded only as a feature map, not as an image. To elaborate, the feature map may be data containing information about features for processing a specific task of a machine based on an image.
[0116] Referring to FIG. 4d, a feature extractor may be connected to the front end of the VCM encoder (100a).
[0117] The VCM encoder (100a) may include a feature converter and a video encoder. The video encoder may be the encoding device (10a) illustrated in FIG. 2.
[0118] The VCM decoder (100b) illustrated in FIG. 4d may include a video decoder and an inverse converter. The video decoder may be the decoding device (10b) illustrated in FIG. 3.
[0119]
[0120] FIG. 5 is a block diagram of an encoding device according to an embodiment of the present disclosure.
[0121] Video Coding for Machines (VCM) technology is currently being validated only for multi-layer feature maps extracted from fixed networks and fixed segmentation points determined by the training dataset and task. Future FCM standard technology (Feature Coding for Machines) will require technology that considers the segmentation and decoding of single- or multi-layer feature maps at arbitrary segmentation points in deep learning networks. Furthermore, non-learning-based feature map compression methods also need to be considered. In VCM, methods such as omitting some information within frames (RoI-based processing) and omitting frames (Temporal Resampling) can be utilized in the FCM standard due to their high segmentation and decoding efficiency. For example, feature maps can typically have a significantly larger number of channels compared to the input data depending on the network. In this case, by omitting and encoding some of the channels, the decoder can restore them in various ways. Various embodiments of the present disclosure include examples of restoring omitted information in various ways at a decoder due to omission of some information within a frame (e.g., region of interest processing) at an encoder or frame omission (e.g., temporal resampling).
[0122] In various embodiments, detailed embodiments of omitting some information in a frame of a decoder / decoder or omitting frames, etc., are described, considering cases where an input feature map is a single-layer feature map and a multi-layer feature map, respectively. For example, for a feature map of dimension WxHxC in each frame, when redundant channels are skipped in order to compress each channel (WxH), there is an embodiment for restoring the corresponding channels in a decoder and a syntax definition thereof. According to various embodiments, in order to restore an encoded bitstream by skipping redundant channels, the decoder may restore the omitted channels by filling them with a specific value (e.g., 0), copy or filter the same-position channel value of a previous frame, interpolate the same-position channel value of another layer in the case of a multi-layer feature map, or filter other restored channel values in the current feature map.
[0123] In various embodiments, when using region of interest information during the encoding process, compensation parameters for the region of interest may be signaled to the decoder. In one embodiment, information for feature map compensation may be signaled for each period of the feature map. For example, when temporal resampling is performed during the encoding process, compensation parameters including the mean / variance values of the feature maps of the frames before / after the omission may be transmitted to the decoder. In one embodiment, when temporal resampling is performed periodically or aperiodic, various methods for performing temporal restoration may be performed.
[0124] An encoding device (10a) according to an embodiment of the present disclosure can receive an image (feature map), perform encoding, and output a bitstream. The encoding device (10a) can include an internal encoding preprocessing unit (500) and an internal encoding performer (560). The internal encoding preprocessing unit (500) can include at least one of a temporal resampling performer (510), a feature map redundancy remover (520), a feature map truncation performer (530), a quantization performer (540), and a packing performer (550). Each component included in the encoding device (10a) is only a distinction for explaining a logical operation, and physically, one processor can process all the components, or multiple processors can process multiple components by distinguishing them, and the hardware devices actually implemented can be diverse. In various embodiments, some of the components disclosed in FIG. 5 can be omitted, and the order between the components can be changed.
[0125] The input image may be a feature map having a size of HxWxC for integers H, W, and C greater than or equal to 1. The feature map may be one or more feature maps extracted from an intermediate layer of a deep neural network composed of one or more convolutional layers.
[0126] The internal encoding preprocessing unit (500) may perform various processing on the input image before encoding to efficiently perform encoding for the machine. In one embodiment, the encoding device (10a) may select an object to be encoded. Conversely, the encoding device (10a) may omit some frames or parts of frames of the input image when the decoding device (10b) can restore the frames without encoding. For example, the encoding device (10a) may perform temporal resampling to delete some of the frames included in the entire image. Alternatively, the encoding device (10a) may derive a region of interest to be encoded through region of interest processing, and may decide to fill in non-region of interest images with a specific value or not to encode them at all. However, the encoding device (10a) may signal information to the decoding device (10b) to restore the original image when necessary during the decoding process. In various embodiments, the encoding device (10a) can signal information defined by an agreement between the encoder / decoder into the bitstream and transmit it to the decoding device (10b).
[0127] The temporal resampling performer (510) can receive a feature map as input, perform sampling on a frame-by-frame basis, and output a sampled image in which the frame rate of some frames is changed. According to an embodiment, the temporal resampling performer (510) can perform temporal resampling through the same degree of sampling (the same frame rate) for the entire feature map sequence. According to an embodiment, the temporal resampling performer (510) can perform temporal resampling through different degrees of sampling (variable frame rates) for each group of frames in the feature map sequence.
[0128] According to an embodiment, the temporal resampling performer (510) may transmit information used in the temporal resampling process (for example, a temporal sampling rate, etc.) to the decoder.
[0129] The feature map redundancy remover (520) can receive a feature map or a feature map on which temporal resampling has been performed, and remove redundancy in a spatial axis or a channel axis within the feature map or between feature maps for each frame or a series of frames. According to an embodiment, the feature map redundancy remover (520) can perform feature map redundancy removal through a trained deep neural network composed of one or more convolutional layers. The feature map redundancy remover (520) can perform redundancy removal using information of an input image from which an input feature map is extracted. Alternatively, the feature map redundancy remover (520) can calculate redundancy within the feature map or between feature maps, classify redundant data within the feature map into a series of groups, and transmit information thereon to a decoder. In various embodiments, the encoding device (10a) can remove feature map redundancy by combining one or more of the above processes.
[0130] According to an embodiment, the feature map deduplication remover (520) may transmit information used for feature map deduplication (for example, deep neural network information, grouping information, redundancy information, etc.) to the decoder.
[0131] The feature map truncation performer (530) can receive a feature map or a feature map on which some processing has been performed, and output a feature map in which some information of the feature map has been truncation based on a quantization rate for each frame or a series of frames. The feature map truncation performer (530) can determine whether to truncate or not based on a quantization parameter of the feature map. The quantization parameter can be set as an input of the encoding device (10a).
[0132] According to an embodiment, the feature map pruning performer (530) may omit the pruning process for some of the grouped feature maps according to the degree of the quantization parameter, if grouping has been performed in the previous process for a feature map having a size of W×H×C. At this time, the omission of the pruning process may be performed in the order of the group having a large number of grouped channels according to the quantization parameter.
[0133] According to an embodiment, the feature map pruning performer (530) may omit some values within a feature map having a size of W×H×C. In this case, the threshold value of the omitted values may be determined by the input quantization parameter.
[0134] According to an embodiment, the feature map cutting performer (530) may transmit information used in the feature map cutting process (for example, the degree of cutting, etc.) to the decoder.
[0135] The quantization performer (540) can receive a feature map in floating point or integer form or a feature map on which a partial process has been performed and perform quantization into integer form. Depending on the embodiment, the quantization performer (540) can perform uniform quantization or non-uniform quantization based on the maximum / minimum value or average value of the input feature map. The quantization performer (540) can perform uniform or non-uniform quantization within a range after clipping a certain range.
[0136] According to an embodiment, the quantization performer (540) may transmit information about the quantization performed (e.g., maximum / minimum values, median values, clipping range, etc.).
[0137] The packing performer (550) can receive a feature map or a feature map having a size of W×H×C on which some process has been performed and perform packing. According to an embodiment, the packing performer (550) can use a method of packing each W×H in the feature map in a two-dimensional promised scanning method and order, and can perform the packing process through a method of packing W×H arranged in a time axis, etc.
[0138] In one embodiment, the packing performer (550) can perform two-dimensional packing in units of W'×H' for a feature map having a size of W'×H'×C'. At this time, the packing order may be one of the z-scan order, etc., and the number of columns / rows in units of W'×H' may be adaptively determined according to the value of C'.
[0139] In another embodiment, the packing performer (550) performs two-dimensional packing for a feature map having a size of W'×H'×C', some in units of W'×H', ((W'×X)×(H'×Y)), X×Y <C'), 이를 새로운 축 (일 예시로 시간 축)으로 패킹을 수행할 수 있다.
[0140] In another embodiment, when the encoder input feature map is a multi-layer feature map, the packing performer (550) can pack the multi-layer feature map into one two-dimensional frame, and can pack it into each two-dimensional frame.
[0141] According to an embodiment, the packing performer (550) can transmit information about the packing performed (e.g., the size of the packed image, the packing scan method, etc.).
[0142] The internal encoding performer (560) can receive a feature map or a feature map on which some processes have been performed and perform image encoding to generate a bitstream. According to an embodiment, the internal encoding performer (560) can use a 2D video encoder (AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, etc.), an entropy encoder (for example, DeepCABAC, etc.), or a 2D video encoder including one or more convolutional layers.
[0143] According to an embodiment, the internal encoding performer (560) may perform encoding after converting the color space of the input feature map to a color space such as YUV400, YUV420, YUV444, etc. At this time, according to an embodiment, the internal encoding performer (560) may perform the conversion through a conversion method defined by an agreement between the sub-decoder and the decoder, and may transmit color space conversion information to the decoder.
[0144] According to an embodiment, when the internal encoding performer (560) converts the color space to a format other than YUV400 (for example, YUV444 and YUV420), the values of U and V may be filled with the same value, and at this time, the same value may be one of the median value, the mode value, etc.
[0145] FIG. 6 illustrates the detailed operation of a feature map deduplication remover according to an embodiment of the present disclosure.
[0146] A feature map redundancy remover (520) of an encoding device (10a) according to an embodiment of the present disclosure can receive an image (feature map), group data within the image into a series of units, and output a coding method and a grouped image for each group. The feature map redundancy remover (520) can include a feature map grouping performer (610), a grouping information reordering performer (620), and a coding method mapping performer (630). The order of the components of the feature map redundancy remover (520) disclosed in FIG. 6 can be changed, and some components can be omitted depending on the embodiment.
[0147] The feature map grouping unit (610) can group feature maps having similar characteristics into a series of units (for example, W×H units, hereinafter each channel) for one or more feature maps having a size of W×H×C extracted from one frame of an image (video). The grouping method may vary depending on the embodiment, and the purpose is to perform grouping that maps channels having similar characteristics into one group. As an example, the feature map grouping unit (610) can perform grouping that calculates correlation between each channel existing in one feature map and groups channels with high correlations into one group. At this time, the process of calculating correlation may vary. As an example, the feature map grouping unit (610) can perform grouping based on average and variance values for channels in the feature map. The feature map grouping performing unit (610) can perform grouping through a combination of various methods, and these can be used differently in the encoder depending on the purpose.
[0148] The grouping information reordering performing unit (620) can perform reordering of groups according to grouping information on the grouped feature map resulting from the above process. Depending on the embodiment, the group reordering process can reorder in descending order of the number of grouped channels, or in ascending order. In this case, if the number of grouped channels is the same, reordering can be performed in order of the channel indexes, whether they are small or large.
[0149] The coding method mapping unit (630) receives a feature map in which grouping has been performed and the groups have been reordered, determines a coding method for each group, and signals the determined coding method for each group. Depending on the embodiment, the coding method may vary and may include the following coding methods.
[0150] In one embodiment, the coding method mapping performing unit (630) may use 1) a method of coding by performing quantization, packing, and internal encoding of representative channel values for channels of a grouped feature map, 2) a coding method of filling all channels with specific values, 3) a method of coding by performing copying and / or filtering of a specific channel group of a feature map extracted from a previously coded video frame by comparing it with the current video frame, 4) a coding method of copying and / or filtering values of another channel group extracted from the current video frame, etc. In this case, if the feature map extracted from the current video frame is a multi-layer feature map, the other channel group may be a channel group value of another layer. If it is a single-layer feature map, not a multi-layer feature map, the value of another channel group within the feature map may be used.
[0151] Depending on the embodiment, the coding method mapping performing unit (630) may determine the coding method in various ways. The coding method mapping performing unit (630) may use a combination of one or more of the following methods: a method for restoring a signal close to the original signal, a method for restoring channels within a feature map so that they are dependent on the target task and the deep neural network, and a method for minimizing the bit rate generated.
[0152] According to an embodiment, the coding method mapping performing unit (630) can signal the coding method mapped through the above process to the decoder.
[0153] FIG. 7 illustrates detailed operations of a quantization performer according to an embodiment of the present disclosure.
[0154] A quantization performer (540) of an encoding device (10a) according to an embodiment of the present disclosure may receive a feature map in a floating-point format or an integer format, perform quantization in an integer format, and output a quantized image (feature map). The quantization performer (540) may include a quantization method determination unit (710) and a quantization performer (720). The order of the components of the quantization performer (540) disclosed in FIG. 7 may be changed, and some components may be omitted depending on the embodiment.
[0155] The quantization method determining unit (710) may, according to an embodiment, receive an image (feature map) and determine a quantization method of the feature map. According to an embodiment, the quantization process of the feature map may be a method of receiving feature map data in a floating-point format and outputting it as feature map data in an integer format, or a method of receiving feature map data in an integer format and outputting it as feature map data in an integer format.
[0156] As an example, quantization methods can vary and may include the following:
[0157] A single quantization method may be 1) a uniform quantization method that performs quantization by evenly dividing the input feature map into a target integer number of bits based on the maximum and minimum values of the input feature map for each feature map, or each channel of each feature map, or each channel group of each feature map. Or, it may be a method that performs quantization by unevenly dividing the data according to the distribution of data values within the feature map. Or, 2) a method that performs clipping of data within the feature map to a specific range and performs quantization by evenly or unevenly dividing the clipped feature map data.
[0158] According to an embodiment, the quantization method determination unit (710) can signal the quantization method to the decoder for each quantization performance unit for the determined quantization method.
[0159] The quantization performing unit (720) may perform quantization on a feature map according to a method determined by the quantization method determining unit (710), depending on an embodiment. In the case where quantization is performed based on a specific value of a quantization performing unit, the specific value may be signaled to the decoder as a quantization performing unit. In the embodiment, the degree of quantization may be determined as an input of the encoding device (10a), and information about the degree of quantization may be signaled to the decoder.
[0160] According to an embodiment, the quantization performing unit (720) may perform quantization after adjusting the average of the distributed values to 0 for each quantization unit (e.g., a feature map or each channel within a feature map or a representative channel of each group). At this time, the average value for creating zero mean data may be signaled to the decoder.
[0161] Depending on the embodiment, the unit for transmitting additional data (e.g., mean, variance, scale value, etc.) may be a frame unit of the extracted feature map, or a series of frame units.
[0162] FIG. 8 illustrates detailed operations of an internal encoding performer according to an embodiment of the present disclosure.
[0163] An internal encoding performer (560) of an encoding device (10a) according to an embodiment of the present disclosure may perform encoding of an image (feature map) and group-specific coding information to output a bitstream. The internal encoding performer (560) may include a downsampling performer (810) and a feature map encoding performer (820). The order of the components of the internal encoding performer (560) disclosed in FIG. 8 may be changed, and some components may be omitted depending on the embodiment.
[0164] According to an embodiment, the image (feature map) may be a representative (channel) of each group on which direct coding is to be performed in the feature map redundancy removal process, and encoding may be performed after downsampling. According to an embodiment, the internal encoding performer (560) may perform MUXing on each bitstream output through each feature map encoding performer (820) to generate a single bitstream.
[0165] In the case of a multi-layer feature map, the downsampling unit (810) can perform downsampling for each layer at different sampling rates. Group-specific coding information may not be downsampled. In some embodiments, if downsampling is performed, sampling rate information may be signaled to the decoder.
[0166] The feature map encoding unit (820) may use a video encoder (AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, etc.), a 2D video encoder including one or more convolutional layers, or an entropy encoding method. Depending on the embodiment, encoding of each data may use different encoders. For example, an encoding method for a feature map and an encoding method for group-specific coding information may be different from each other.
[0167] FIG. 9 is a block diagram of a decoding device according to an embodiment of the present disclosure.
[0168] A decoding device (10b) according to an embodiment of the present disclosure can receive a bitstream, perform decoding, and output a restored image (feature map). The decoding device (10b) can include an internal decoding performer (910) and an image restoration performer (900). The image restoration performer (900) can include at least one of an unpacking performer (920), an inverse quantization performer (930), a feature map non-truncation performer (940), a feature map redundancy restoration performer (950), and a temporal restoration performer (960). The components included in the decoding device (10b) only mean logical operations, and physically, one processor can process all the components, or multiple processors can process multiple components separately, and the hardware devices actually implemented can be diverse. In various embodiments, some of the components disclosed in FIG. 9 can be omitted, and the order between the components can be changed. The decryption process may be performed in the reverse order of the encoding process of the encoding device (10a), if there is a process corresponding to the encoding process, or may be performed in a different order.
[0169] The internal decoding performer (910) can receive a bitstream as input and perform image decoding to generate a restored image. Depending on the embodiment, the decoding can use a 2D video decoder (AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, etc.), a 2D video decoder including one or more convolutional layers, or an entropy decoder. Depending on the embodiment, the color space of the decoded image (feature map) resulting from the internal decoding performer (910) can be implicitly or explicitly converted to another color space, and then the image decoding process can be performed.
[0170] The image restoration performer (900) can restore the original image based on the image and information decoded from the bitstream. In various embodiments, if temporal resampling is performed in the encoding device (10a), the bitstream may omit some frames of the original image, so the original image can be restored by performing temporal restoration. In one embodiment, if region-of-interest processing is performed in the encoding device (10a), the region-of-interest and information about the region-of-interest are encoded and transmitted, but the non-region-of-interest image may be encoded by replacing it with a specific value or deleted altogether, so restoration of the non-region-of-interest image is required based on the region-of-interest-based information. In response to a bitstream generated in response to at least one component included in the internal encoding preprocessing unit (500) of the encoding device (10a), the decoding device (10b) may perform various restorations on the decoded image by at least one component included in the image restoration performer (900).
[0171] The unpacking performer (920) may, according to an embodiment, restore an unpacked feature map using the decoded image and packing information (for example, a packing method, etc.) transmitted from the encoding device (10a). In one embodiment, the unpacking performer (920) may receive an image (feature map) decoded through internal decoding, and perform unpacking by receiving a scanning order and a packing method from the encoder. The unpacking performer (920) may receive a decoded image (feature map) as an input and create one or more images (feature maps) in a three-dimensional shape of W'×H'×C'. The unpacking performer (920) may perform unpacking according to the parsed packing method. For example, when performing packing in the encoder, if some information is cropped for each performing unit, the size of each performing unit may vary. In this way, when packing is performed in different units in the encoder, the unpacking unit (920) can perform uncropping according to the degree of cropping to restore to the original size. At this time, the value of the portion generated by the decoder can be used by copying the value adjacent to the corresponding pixel (or line) value, and can be restored in a form in which a specific value is filled in the same way. At this time, the specific value can be determined by an agreement between the unit / decoder, and can be implicitly determined by using a value in the feature map decoded by the decoder.
[0172] The inverse quantization performer (930) may, according to an embodiment, restore a dequantized feature map using a decoded image or a decoded image on which some processing has been performed and quantization information transmitted from an encoding device (10a). The information used in the quantization process may include, for example, information about a quantization method, a quantization degree, and a range (maximum / minimum value, mean / variance value, etc.). According to an embodiment, when a mean centering process is performed and quantization is performed in the quantization process, inverse quantization may be performed in response thereto in the inverse quantization process, and then mean value compensation may be performed using the mean value transmitted for the inverse quantized data.
[0173] According to an embodiment, in the quantization process of the encoding device (10a), one or more pieces of information including a mean value, a variance value, and a scale value may be signaled for feature map values prior to quantization. In one embodiment, the dequantization performer (930) may perform standardization by subtracting the mean and variance value of the dequantized feature map after dequantization in the decoding process and dividing by the variance value, and then perform addition and multiplication operations using the parsed original mean and variance values to perform compensation of the feature map. The mean and variance values used in the above process may be parsed by frame unit and level unit of the feature map, or may be parsed by a series of frame groups. Whether or not to perform the feature map compensation process by one or more series of frame groups may be signaled by the encoder as a flag.
[0174] According to an embodiment, the dequantization performer (930) restores and dequantizes each level (n) feature map. , the average value of each level feature map parsed from the bitstream is m n , the variance value is v n In this case, feature map compensation for each level can be performed as in mathematical equations 1 and 2 below.
[0175]
[0176]
[0177] Here, is a variable derived from the result of mathematical expression 1, and means the result of performing standardization by subtracting the average value from the restored feature map using the average value and variance value of the restored inverse quantized feature map of each level and dividing it by the variance value. Standardization refers to the process of converting the range of values to have a distribution with an average of 0 and a variance of 1. and refers to the mean value and variance value of the restored inverse quantized feature map of each level. represents the compensated feature map.
[0178] Alternatively, the dequantization performer (930) can perform compensation as in mathematical expression 3 below.
[0179]
[0180] Here, S n means the scale value for each feature map level n. The process as in mathematical expression 3 is to restore the feature For each element (for example, 1 pixel value of a WxHxC feature map) of S n Multiply by the scale value It may be a process of printing, and at this time S n can mean different scale values for each level. Also, the scale value S n can have a real number value greater than or equal to 0 and less than 1.
[0181] In various embodiments, unlike the methods of Equations 1 to 3, compensation is performed using variables other than the mean and variance, or scale values, to obtain a compensated feature map ( ) can be obtained. According to an embodiment, the average and variance values may be transmitted for each level from the encoding device (10a), and the average and variance values may be transmitted for one or more level groups. At this time, compensation parameters such as average, variance, and scale may be transmitted in the form of values, or / and an index of a parameter table defined by an agreement between an encoder and a decoder may be transmitted, and the decoder may parse this to use the parameters of the table in the compensation process.
[0182] In some embodiments, parameters for feature map compensation (e.g., mean and variance values) may be transmitted frame by frame or in units of a series of frame groups. In this case, the frame group unit may be one of GOP, Intra period, etc. When region-of-interest-based encoding is performed, information on the location and size of the region-of-interest within the feature map may be parsed in units of one or more frames and units of one or more levels. In this case, feature map compensation may be performed in units of region-of-interest, and whether or not it is performed may be signaled / parsed.
[0183] The feature map truncation performer (940) can receive a dequantized and decoded image (feature map) or a decoded image (feature map) as input and restore information truncation during the encoding process. According to an embodiment, the feature map truncation performer (940) can parse information used in the truncation process (for example, the degree of truncation, the size of the original data, etc.) and perform truncation on the decoded feature map so that it has the same resolution and dimension as the original feature map (the same as the data before performing feature map truncation in the encoding step). According to an embodiment, information on the degree of truncation can be transmitted from the encoder and implicitly determined based on the quantization parameter used in the encoding process. According to an embodiment, data generated through truncation can have a dimension size identical to that of the input data of the feature map truncation performer of the encoder, and values inside the data to be generated through truncation can be filled with specific values. Alternatively, it can be filled in by copying information (e.g., channel values) within another feature map that has already been restored. In this case, information for non-cutting (such as the degree of cutting and the non-cutting method) can be transmitted from the encoder, and non-cutting can be performed using a fixed method.
[0184] The feature map redundancy restorer (950) may, according to an embodiment, restore a feature map based on the redundancy removed by the encoder by parsing information used in the feature map redundancy removal process (for example, channel grouping information, each coding method of the grouped channels, etc.).
[0185] The temporal restoration performer (960) may, according to an embodiment, receive information on the number of frames to be restored (for example, temporal sampling rate, etc.) as a signal, and perform restoration of intermediate frames using the decoded image (feature map). According to an embodiment, for the decoded image (feature map) input to the temporal restoration performer (960), two temporally adjacent frames may be used for temporal restoration, and two temporally non-adjacent frames may be used for temporal restoration. At this time, more than two frames may be used. According to an embodiment, a method for determining a frame to be used for temporal restoration may be determined by an encoder, and the corresponding information may be transmitted, and the decoder may implicitly determine the condition through a comparison between decoded images. According to an embodiment, the above-mentioned condition may be determined based on a specific threshold value by comparing one or more values such as MSE (Mean squared error), SSIM (Structural similarity index measure), etc. between two frames. Depending on the embodiment, the above-described comparison method can also be used to implicitly determine and utilize a temporal restoration method. For example, if two frames used for temporal restoration are numerically similar, the temporally preceding or succeeding frame can be copied to create an intermediate frame.
[0186] In some embodiments, when temporal restoration of a frame is performed using two or more frames, input frames for restoration may be applied differently for each region within the frame to be restored. That is, when the input frames used in the process of performing temporal restoration of an arbitrary frame are A, B, and C, temporal restoration of a part of the frame to be restored may be performed using frames A and B, and of the remaining region of the frame to be restored may be performed using frames A and C. At this time, filtering may be performed to remove discontinuity between the two regions.
[0187] According to an embodiment, temporal restoration information (e.g., temporal_restoration(), etc.) may be transmitted from the encoding device (10a) at various levels to perform temporal restoration. In one embodiment, the encoding device (10a) may transmit temporal restoration information at at least one level among a sequence unit, a GOP unit, a frame unit, a slice unit, and a subsequence unit. In various embodiments, the encoding device (10a) may transmit necessary temporal restoration information at each of various levels.
[0188] FIG. 10 illustrates detailed operations of an internal decryption performer according to one embodiment of the present disclosure.
[0189] An internal decoding performer (910) of an encoding device (10a) according to an embodiment of the present disclosure may receive a bitstream and perform decoding on coding information for each image and group. The internal decoding performer (910) may include a feature map decoding performer (1010) and an upsampling performer (1020). Depending on the embodiment, the order of each process may be changed, and some components may be omitted.
[0190] In one embodiment, the input bitstream can be demuxed and applied as input to the feature map decoding unit (1010).
[0191] The upsampling unit (1020) can perform upsampling on each decoded image (feature map) at a parsed sampling rate. According to an embodiment, the upsampling method can be fixedly used among bilinear upsampling, bilateral upsampling, nearest-neighbor upsampling, and a deep neural network including one or more convolutional layers. Alternatively, the encoder can transmit the upsampling method, and the decoder can parse the upsampling method and perform upsampling according to the corresponding method and sampling rate. According to an embodiment, the restored group-specific coding information can include information (Skip_flag, etc.) on the coding mode for each channel of the original feature map.
[0192] FIG. 11 illustrates the detailed operation of a feature map redundancy restorer according to an embodiment of the present disclosure.
[0193] A feature map redundancy restorer (950) of a decoding device (10b) according to an embodiment of the present disclosure may receive a decoded image (feature map) and restored group-specific coding information as input, and output a redundancy-restored image (feature map). At this time, the decoded feature map may be a result of performing one or more of an unpacking, dequantization, and non-truncation process. According to an embodiment, a group of decoded group-specific coding information may mean a channel group composed of one or more channels in units of W×H for a feature map having a size of W×H×C. At this time, the feature map may be a feature map extracted for an input image. In one embodiment, for an extracted feature map (X×Y×Z), a group of decoded group-specific coding information may be a region (W×H×C) from which a region of interest is extracted. In this case, the location and the proportion of the region of interest in the feature map may be the same for the entire image, and may differ for some frame units. The following is a description based on the relevant embodiment. According to the embodiment, the feature map redundancy restoration process may be performed as illustrated in FIG. 11. In this case, each process may be omitted, and the order may be changed.
[0194] According to an embodiment, the feature map redundancy restorer (950) may receive a decoded image (feature map) and restored group-specific coding information, determine a feature map redundancy restorer method for each group, and perform feature map redundancy restorer according to the determined restorer method. Fig. 11 is an example for a case where the number of groups grouped for the feature map is n-1. The feature map redundancy restorer (950) may first initialize group_idx to 0, and check skip_flag for each of n-1 groups, and apply the feature map redundancy restorer method (coding method) determination module (1101) and the feature map redundancy restorer method (coding method) execution module (1102).
[0195] According to an embodiment, the feature map redundancy restorer (950) can check the skip_flag for each group or channel from the decoded group-specific coding information.
[0196] 1) At this time, if skip_flag is 1, the feature map redundancy restoration method execution module (1102) can use the channel value of the decrypted feature map as is for the corresponding channel group or channel.
[0197] 1-1) Obtain skip_flag for each group, and if the flag is 1, the feature map redundancy restoration method execution module (1102) can perform redundancy restoration by parsing the index information of the grouped channels of the corresponding group obtained from the restored group-specific coding information and copying the redundancy-restored channels to each channel location.
[0198] 1-2) Obtain skip_flag for each channel, and if the flag is 1, the feature map redundancy restoration method execution module (1102) can perform redundancy restoration through group information for each channel from the restored group-specific coding information.
[0199] 2) According to an embodiment, if the skip_flag parsed from the decoded group-specific coding information is 0, the feature map redundancy restoration method determination module (1101) can parse the feature map redundancy restoration method from the decoded group-specific coding information. Each method can be managed in the form of an index, and each index represents information about the coding method (redundancy restoration method).
[0200] According to an embodiment, if the skip_flag parsed from the decoded group-specific coding information is 0, the feature map redundancy reconstruction method determination module (1101) can parse the rrm_idx (redundancy reconstruction mode index). The decoded group-specific coding information includes mapping information between each rrm_idx and a reconstruction method (coding method), and the number of available reconstruction methods can be determined in advance by a sub / decoder agreement.
[0201] In one embodiment, the feature map redundancy restoration method performing module (1102) is configured to (1) when the restoration method is determined to be a method of filling in a specific value (for example, rrm_idx == 0), all channels of the corresponding group can be filled with the same specific value. That is, redundancy restoration can be performed with a restoration method in which the corresponding channels (W×H) are filled with the same value.
[0202] In one embodiment, the feature map redundancy restoration method performing module (1102) may parse information such as ① the index of the previous frame, ② the channel information index of the previous frame (channel group index or channel index), and ③ filtering information (filter coefficient or filtering index) from the restored group-specific coding information when (2) the restoration method is determined to be a method of restoration using channel information in a previously restored frame (for example, rrm_idx == 1). At this time, the feature map(s) of the previously restored frame may be managed in the form of a buffer according to temporal order. Redundancy restoration may be performed using the channel value accessed through the index of the previous frame and the channel information index of the previous frame. At this time, when the resolution of the channel to be currently restored and the resolution of the channel to be referenced are different, restoration may be performed by adjusting the resolution of the channel to be referenced. At this time, the resolution adjustment may use a filter derived from the parsed filtering information, or a fixed method may be used. For reference channels with matching resolution, final restoration can be performed by performing filtering using a filter obtained from the parsed filtering information.
[0203] In one embodiment, the feature map redundancy restoration method performing module (1102) may parse information such as ① previously restored channel information (layer index, channel index, channel group index, or index difference value between the current channel group and the channel group to be restored) and ② filtering information (filter coefficient or filtering index) from the restored group-specific coding information when (3) the restoration method is determined to be a method of restoration using previously restored channel information within the current frame (for example, rrm_idx == 2). At this time, the layer index may be parsed when restoring a multi-layer feature map. When restoring a multi-layer feature map and performing restoration of the current channel group using channel information in the feature map of another layer, resolution adjustment may be performed according to the resolution of the current channel group. At this time, the resolution adjustment may use a filter derived from the parsed filtering information, or may use a fixed method. For a reference channel with a matching resolution, filtering may be performed using a filter obtained from the parsed filtering information to perform final restoration. When the index difference value of the current channel group and the channel group to be used for restoration is transmitted from the encoder, restoration can be performed by filtering the channel information of the group obtained by subtracting the difference value from the current channel group.
[0204] <Description of Syntax and Semantics Used in Embodiments of the Present Disclosure>
[0205] Examples of syntax and semantics between an encoding device (10a) and a decoding device (10b) used in various embodiments of the present disclosure.
[0206] (1) skip_flag[i][j]: Indicates the skip flag of the jth channel of the i-layer feature map. If the flag is 1, the feature map channel can be coded as in the example of Fig. 11. In the case of a single-layer feature map, i can only have a value of 0.
[0207] (2) rrm_idx: Index information indicating the coding method when the skip flag is 0.
[0208] (3) fill_value: If the channel group is determined to be filled with a specific value, this information may be omitted to indicate the value to be filled with.
[0209] (4) ref_frame_idx: If the channel group is determined by a coding method in which the channel information of the previously coded frame is restored, index information for the previous frame, which may be a frame index according to time order, or an index indicating one of the frames stored in the buffer.
[0210] (5) channel_information()
[0211] - layer_idx: In the case of a multi-layer feature map, layer index information that determines which layer's feature map channel information to use. In the case of a single-layer feature map, this syntax can be omitted.
[0212] - channel_idx: Index information indicating which channel information of the layer selected from layer_idx to restore. Using layer_idx and channel_idx, you can determine which channel information of which layer to use.
[0213] - idx_diff_value: Information indicating the difference value between the current channel index information and the channel index information to be referenced. The decoder can perform restoration using the channel value of the index indicated by the difference value from the current channel index value.
[0214] (6) filter_information()
[0215] - filter_idx: Index information indicating one of the filters determined by the decoder / encryptor agreement. Filters determined by the agreement can be managed in table format.
[0216] - filter_coefficients_num: If filter_idx has a specific value (ex. x), it can be parsed. In this case, it means the number of filters to be parsed.
[0217] - coefficients: For each coefficient of the filter to be parsed, it means the filter coefficient value.
[0218] Examples of syntax for feature map headers according to various embodiments are shown in Table 1 below.
[0219] Descriptorfeature_map_header_rbsp() {… for (i=0;i <sps_num_layers; i++ ) {for ( j=0; j<sps_num_channels[i]; j++ ) {skip_flag[i][j]ue(1)if ( !skip_flag[i][j] ) {rrm_idxue(v)if ( rrm_idx == 0 ) {fill_valueue(v)} else if ( rrm_idx == 1 ) {ref_frame_idxue(v)channel_information ()filter_information ()} else if ( rrm_idx == 2 ) {channel_information ()filter_information ()}…}}}
[0220] Examples of syntax for channel information of feature maps according to various embodiments are as follows: Tables 2 and 3.
[0221] channel_information Example 1) Descriptorchannel_information () {layer_idxue(v)channel_idxue(v)}
[0222] channel_information Example 2) Descriptorchannel_information () {layer_idxue(v)idx_diff_valueue(v)}
[0223] Examples of syntax for filter information of feature maps according to various embodiments are as follows: Tables 4 and 5.
[0224] filter_information Example 1) Descriptorfilter_information () {filter_idxue(v)}
[0225] filter_information Example 2) Descriptorfilter_information () {filter_idxue(v)if (filter_idx == x) {filter_coefficients_numue(v)for (i=0; i <filter_coefficients_num; i++ ) {coefficientsue(v)}}}
[0226] In various embodiments, examples of syntax for feature map compensation related to the dequantization process are shown in Tables 6 and 7 below.
[0227] DescriptorSemanticfeature_restoration(){feature_restoration_flague(1)Flag for whether feature map compensation is performedif (feature_restoration_flag) {periodic_restoration_flague(1)Flag information for whether periodic feature map compensation is performed. A period can mean one or more frames, can have the same value within a sequence, or can have different periods for each specific section.if (periodic_restoration_flag) {restoration_periodue(v)In case periodic feature map compensation is performed, information about the period. The corresponding value can be transmitted as a value of the number of frames, and can be transmitted in the form of an index so that the periods indicated by each index are different. For example, if it is 1, it can mean period 4. If the periods in the sequence are different, information about the periods can also be transmitted for each section. if (poc % restoration_period == 0) {apply_restoration_flague(1) A flag indicating whether feature map compensation is performed for each period. Depending on the implementation example, if feature map compensation is performed for each period, compensation may not be performed for some periods. if (apply_restoration_flag) {for (i=0; i <num_of_levels; i++) {num_of_levels는 한 프레임에서 추출된 피처 맵이 다중 레벨 피처 맵 (일 예시로, P2, P3, P4, P5)일 때, 각 피처 맵의 인덱스를 의미한다. 단일 레벨 피처 맵인 경우, 1의 값을 가질 수 있다.layer_level_mean[i]ue(v)각 레벨 별 피처 맵 보상을 위한 평균 값, 이는 하나의 예시를 나타낸 것이며, 보상을 위한 값의 개수 및 종류는 실시 예에 따라 변경될 수 있다.layer_level_var[i]ue(v) Variance value for feature map compensation for each level. This is an example, and the number and type of values for compensation may change depending on the embodiment. …}}.
[0228] DescriptorSemanticfeature_restoration(){feature_restoration_flague(1)Flag for whether feature map compensation is performedif (feature_restoration_flag) {periodic_restoration_flague(1)Flag information for whether periodic feature map compensation is performed. A period can mean one or more frames, can have the same value within a sequence, or can have different periods for each specific section.if (periodic_restoration_flag) {restoration_periodue(v)In case periodic feature map compensation is performed, information about the period. The corresponding value can be transmitted as a value of the number of frames, and can be transmitted in the form of an index so that the periods indicated by each index are different. For example, if it is 1, it can mean period 4. If the periods in the sequence are different, information about the periods can also be transmitted for each section. if (poc % restoration_period == 0) {apply_restoration_flague(1) A flag indicating whether feature map compensation is performed for each period. Depending on the implementation example, if feature map compensation is performed for each period, compensation may not be performed for some periods. if (apply_restoration_flag) {if (roi_based_coding_flag) { This means that the current sequence has been subjected to region-of-interest-based coding. roi_based_restoration_flag A flag indicating whether region-of-interest-based feature map compensation is performed if (roi_based_restoration_flag) {for (i=0; i <num_of_levels; i++) {for (j=0; j<num_of_rois; j++) {num_of_rois는 관심 영역의 개수를 의미한다.layer_level_mean[i][j]ue(v) Mean value for compensation of the region corresponding to the region of interest within the feature map for each level. This is an example, and the number and types of values for compensation may vary depending on the embodiment. In this case, compensation may be performed only for the corresponding portion by parsing the location and size information of the region of interest. layer_level_var[i][j]ue(v) Variance value for compensation of the region corresponding to the region of interest within the feature map for each level. This is an example, and the number and types of values for compensation may vary depending on the embodiment. In this case, compensation may be performed only for the corresponding portion by parsing the location and size information of the region of interest. … If region-of-interest-based coding is performed but region-of-interest-based feature map compensation is not performed, compensation may be performed by parsing the flag indicating whether feature map compensation is performed and the compensation parameter information for the entire feature map.
[0229] In various embodiments, examples of syntax for temporal restoration during the decryption process are given in Table 8 below.
[0230] DescriptorSemantictemporal_restoration_data( ){temporal_restoration_flague(1)Flag for whether to apply temporal resampling, if 1, temporal resampling is appliedif( temporal_restoration_flag ) {temporal_resampling_ratio_idxue(v)Index of the list for temporal resampling ratios, since temporal resampling is indicated by temporal_restoration_flag, cases where the input signal and the output signal are the same are not included. For example, the sampling rate can be in the form of a ratio such as 1 / 2, 1 / 4, 1 / 8, etc., and can be the number of frames per second. The encoder and decoder can agree on the sampling rate list, and the encoder can transmit the index of the list. In some embodiments, if the resampling rate is fixed or the number of required frames is fixed for the transmitted application, the corresponding information may be omitted. same_period_flague(1) A flag for whether periodic resampling or aperiodic resampling is used when applying temporal resampling. In some embodiments, if the periodicity of temporal resampling is fixed, the corresponding information and related information may be omitted. If same_period_flag is 1, it means that temporal resampling is performed at the same period. if(!same_period_flag) {for(i=0; i <num_of_frames; i++ ) {delta_frame_idx[i]ue(v)프레임숫자에 따라 이전 프레임 인덱스와의 차분 값을 전송한다. 이전 인덱스와의 차분 값이므로 프레임율로 계산된 num_of_frames보다 1개 작은 개수를 전송하면 된다.i: 프레임 인덱스}}Remain_framesue(v)비디오 시퀀스의 시간적 복원을 수행할 마지막 주기의 잔여 프레임 수를 나타내는 값이며, 부호화기에서 복호화되야할 프레임의 전체 수가 시그널링된 경우, 생략될 수 있다.A variable that can have a value less than or equal to the temporal restoration cycle value. For example, if the cycle is 4, Remain_frames can have a value between 0 and 4.
[0231] The examples of the present disclosure presented in this specification and drawings are intended solely to facilitate the technical content of the present disclosure and to aid understanding thereof, and are not intended to limit the scope of the present disclosure. It will be apparent to those skilled in the art that other variations are possible in addition to the examples described above.
[0232] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
[0233] [Explanation of symbols]
[0234] 10a: Encoding device
[0235] 500: Internal encoding preprocessor
[0236] 510: Temporal Resampling Performer
[0237] 520: Feature Map Redundancy Remover
[0238] 530: Feature Map Pruning Performer
[0239] 540: Quantization performer
[0240] 550: Packing Performer
[0241] 560: Internal Encoding Performer
[0242] 10b: Decoding device
[0243] 900: Image Restoration Performer
[0244] 910: Internal decryption performer
[0245] 920: Unpacking Executor
[0246] 930: Dequantization performer
[0247] 940: Feature Map Non-Cutting Performer
[0248] 950: Feature Map Redundancy Restorer
[0249] 960: Temporal Restoration Performer
Claims
1. In a VCM (video coding for machines) decoding device, An internal decoding unit that decodes a bitstream to generate a feature map and parses information about the feature map by decoding the bitstream; and An image restoration performer configured to restore one or more channels omitted in the decoded feature map based on information about the feature map, A VCM decoding device, wherein information about the feature map includes at least one of group-wise coding information including whether to skip for each channel of the feature map, quantization compensation parameters, feature map truncation information, temporal restoration information, and packing information.
2. In paragraph 1, The above image restoration performer includes a dequantization performer, The above quantization compensation parameter includes at least one of the mean, variance, and scale values used in the quantization process, A VCM decoding device, wherein the inverse quantization performer performs inverse quantization on the feature map using the quantization compensation parameter.
3. In paragraph 1, A VCM decoding device, wherein the above quantization compensation parameters are transmitted in units of a series of frame groups.
4. In paragraph 1, The above image restoration performer includes a feature map de-slicing performer, A VCM decoding device, wherein the feature map truncation performer determines the size and data values of a channel to be generated using the feature map truncation information or the quantization compensation parameters, and generates one or more channels of the feature map to restore an image.
5. In paragraph 4, A VCM decoding device, wherein the feature map de-slicing performer copies channel values included in the feature map to generate the one or more channels.
6. In paragraph 1, The above image restoration performer includes a feature map redundancy restoration unit, The above feature map redundancy restorer determines a coding method for feature map redundancy restoration based on the group-specific coding information, A VCM decoding device for restoring redundancy for the feature map according to the above coding method.
7. In paragraph 6, The above feature map redundancy restorer is a VCM decoding device that, if the determined coding method is a method using channel information in a previously restored frame, performs restoration of the feature map by parsing the index of the previous frame, the channel information index of the previous frame, and filtering information from the group-specific coding information, and using the channel value identified through the information.
8. In paragraph 6, The above feature map redundancy restorer is a VCM decoding device that, if the determined coding method is a method using channel information restored within the current frame, performs restoration of the feature map by parsing the channel information and filtering information restored from the group-specific coding information and using the channel value identified through the information.
9. In paragraph 1, The above image restoration performer includes a temporal restoration performer, A VCM decoding device, wherein the temporal restoration performer performs temporal restoration by generating one or more intermediate frames for the feature map based on the temporal restoration information.
10. In paragraph 9, A VCM decoding device, wherein the temporal restoration performer determines at least one frame to be used for temporal restoration from among a plurality of frames of the feature map, and copies the at least one frame to generate the one or more intermediate frames.
11. In paragraph 9, A VCM decoding device, wherein the temporal restoration performer determines whether to use the first frame or the second frame for temporal restoration based on a comparison of specific metric values for the first frame and the second frame among a plurality of frames of the feature map.
12. In paragraph 11, A VCM decoding device, wherein the specific metric value is at least one of MSE (mean squared error) and SSIM (structural similarity index measure).
13. In paragraph 9, A VCM decoding device, wherein, when the temporal restoration performer generates the one or more intermediate frames using two or more frames, at least some of the frames used to generate a portion of the one or more intermediate frames and the frames used to generate the remaining portion of the one or more intermediate frames are different.
14. In paragraph 1, A VCM decoding device, wherein the temporal restoration information is transmitted in at least one of a group of pictures (GOP), a sequence, a subsequence, a frame, and a slice unit.
15. In paragraph 1, Including an unpacking performer, A VCM decoding device, wherein the unpacking performer performs unpacking for the feature map based on the scanning order and packing method included in the packing information.
16. In a VCM (video coding for machines) encoding device, A temporal resampling performer that changes the frame rate for feature maps; A feature map redundancy remover that removes redundancy along spatial or channel axes for each frame or a series of frames in the above feature map; A feature map cutting performer for cutting out a portion of the feature map for each frame or a series of frames based on quantization information; A quantization performer that performs quantization on the feature map based on the quantization information; and A VCM encoding device comprising: at least one of temporal restoration information according to the temporal resampling process, group-specific coding information according to the feature map redundancy removal process, feature map cutting information according to the feature map cutting process, and quantization compensation parameters according to the feature map quantization process, and an internal encoding performer that encodes the feature map to generate a bitstream.
17. In paragraph 16, A VCM encoding device, wherein the quantization compensation parameter includes at least one of a mean, variance, and scale value of the quantized feature map, and transmits the feature map in units of frames or in units of a series of frame groups.
18. In paragraph 16, A VCM encoding device, wherein the feature map redundancy remover groups the feature map into a series of units based on the similarity between the channels included in the feature map and determines a coding method for each grouped channel.
19. In Article 16, A VCM encoding device further comprising a packing performer that performs packing for changing a dimension or size of the feature map according to a predetermined packing order for the feature map.
20. A non-volatile computer-readable storage medium that records commands, The above instructions, when executed by one or more processors, cause the one or more processors to: A step of decoding a bitstream to generate a feature map, and decoding the bitstream to parse information about the feature map; and A step of restoring one or more channels omitted in the decrypted feature map based on information about the feature map, A non-transitory computer-readable storage medium, wherein information about the feature map includes at least one of group-wise coding information including whether to skip for each channel of the feature map, quantization compensation parameters, feature map truncation information, temporal restoration information, and packing information.
Citation Information
Patent Citations
Kit for stool testing of companion animals using multi-layer analysis algorithm
KR1020230173270A
Etching composition for silicon nitride layer and method for etching silicon nitride layer using the same
KR1020250042525A
Electronic apparatus and the controlling method thereof for a user access authorization method based on palm print and vein pattern information
KR102648877B1
KR20220136176A