Adaptive resampling and restoration method for video coding for machines

The adaptive resampling and restoration method addresses the challenge of efficiently compressing and analyzing high-quality images for machine-based applications by using an adaptive resampling and restoration method for machines, which includes region of interest processing and temporal restoration, effectively reducing server load and power consumption while maintaining high image quality.

WO2025121997A1PCT designated stage expired Publication Date: 2025-06-12HANWHA VISION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/096625
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-05
Filing Date
2024-11-18
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

As the demand for higher-quality images and videos, such as 4K or 8K, increases, existing video coding technologies face challenges in efficiently compressing and analyzing large amounts of image data for machine-based applications, leading to issues with server load and power consumption.

Method used

An adaptive resampling and restoration method for machines is introduced, which includes a VCM decoding device that decodes bitstreams to generate restored images with region of interest processing and temporal restoration, and a VCM encoding device that performs temporal resampling, region-of-interest-based processing, and internal encoding to generate a bitstream.

Benefits of technology

This method enables effective machine-based image analysis by efficiently compressing and restoring images, reducing the burden on server resources and power consumption while maintaining high image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096625_12062025_PF_FP_ABST
    Figure KR2024096625_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in one embodiment of the present disclosure is a video coding for machines (VCM) encoding apparatus. The encoding apparatus comprises: a temporal resampling performer that changes a frame rate for an input image; a region-of-interest-based processor that extracts one or more regions of interest included in each frame of the input image and generates a processed image on the basis of the one or more regions of interest; and an internal encoding performer that encodes the region-of-interest-based processed image, a non-region-of-interest image, and information on the one or more regions of interest so as to generate a bitstream, wherein the information on the region of interest may include at least one of whether region-of-interest-based processing has been performed, the number of groups of pictures (GOPs) in a sequence, the number of frames in a GOP, the number of regions of interest in a frame, the size of each region of interest, the position of each region of interest, the movement of each region of interest, the mapping relationship between regions of interest, and temporal resampling information on the region of interest.
Need to check novelty before this filing date? Find Prior Art

Description

Adaptive resampling and restoration methods for machine-readable video coding

[0001] The present disclosure relates to an adaptive resampling and restoration method in an encoding / decoding method for a machine.

[0002] With the continuous development of the information and communication industry, broadcasting services with HD (High Definition) resolution have spread worldwide.

[0003] Through this proliferation, many users have become accustomed to high-resolution and high-quality images and / or videos, and the demand for higher-resolution and high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, has increased in various fields.

[0004] The technology for coding this UHD video data was completed in 2013 through HEVC (High Efficiency Video Coding), a standard technology.

[0005] HEVC is a next-generation video compression technology with a higher compression ratio and lower complexity than the previous H.264 / AVC technology, and is a key technology for effectively compressing the massive data of HD and UHD video.

[0006] HEVC performs block-by-block encoding, like previous compression standards.

[0007] However, unlike H.264 / AVC, there is only one profile. The core encoding technologies included in HEVC's sole profile are divided into eight areas: hierarchical encoding structure technology, transform technology, quantization technology, intra-frame prediction encoding technology, inter-frame motion prediction technology, entropy encoding technology, loop filter technology, and other technologies.

[0008] Since the establishment of the HEVC video codec in 2013, the Versatile Video Coding (VVC) standard, a next-generation video codec that aims to improve performance by more than twice that of HEVC, has been developed to address the expansion of realistic video and virtual reality services utilizing 4K and 8K video images. VVC is called H.266.

[0009] H.266 (VVC) was developed with the goal of being more than twice as efficient as the previous generation codec, H.265 (HEVC). VVC was initially developed with resolutions over 4K in mind, but it was also developed for ultra-high-resolution video processing at a whopping 16K level to support 360-degree videos due to the expansion of the VR market. In addition, as the HDR market is expanding due to the development of display technology, it supports 16-bit color depth as well as 10-bit color depth to respond to this, and supports brightness expressions of 1000 nits, 4000 nits, and 10000 nits. In addition, since it is being developed with the VR market and 360-degree video market in mind, it supports partial frame rates in the range of 0 to 120 FPS.

[0010] Advances in Artificial Intelligence

[0011] Artificial intelligence (AI) is also steadily developing. AI refers to the artificial imitation of human intelligence, including the ability to recognize, classify, infer, predict, and control / decision-making.

[0012] With the advancement of artificial intelligence technology and the increase in Internet of Things (IoT) devices, machine-to-machine traffic is expected to explode, and machine-dependent image analysis is expected to become widely used.

[0013] However, as the amount of images to be analyzed by machines is expected to increase exponentially, issues with server load and power consumption are expected to arise.

[0014] Accordingly, the present disclosure aims to provide an adaptive resampling and restoration method for a machine so as to enable effective image analysis by the machine.

[0015] To achieve the above-mentioned object, according to one disclosure of the present specification, an adaptive resampling and restoration method for a machine is presented.

[0016] A VCM decoding device according to one disclosure of the present specification includes an internal decoding performer that decodes a bitstream to generate a restored image including a restored non-region of interest and a region of interest, and extracts information on a region of interest for the restored image; a region of interest-based restorer that restores a region of interest-based processed image from the restored region of interest based on the information on the region of interest; and a temporal restoration performer that performs temporal restoration on the restored image based on the information on the region of interest, wherein the information on the region of interest may include at least one of whether region of interest-based processing has been performed, the number of groups of pictures (GOPs) in a sequence, the number of frames in a GOP, the number of regions of interest in a frame, the size of each region of interest, the position of each region of interest, the movement of each region of interest, the mapping relationship between regions of interest, and temporal resampling information on the region of interest.

[0017] A VCM encoding device according to one disclosure of the present specification includes a temporal resampling performer that changes a frame rate for an input image; a region-of-interest-based processor that extracts one or more regions of interest included in each frame of the input image and generates a processed image based on the one or more regions of interest; and an internal encoding performer that generates a bitstream by encoding the region-of-interest-based processed image, a non-region-of-interest image, and information on the one or more regions of interest, wherein the information on the region of interest may include at least one of whether region-of-interest-based processing has been performed, the number of groups of pictures (GOPs) in a sequence, the number of frames in a GOP, the number of regions of interest in a frame, the size of each region of interest, the position of each region of interest, the movement of each region of interest, the mapping relationship between regions of interest, and temporal resampling information on the region of interest.

[0018] A non-volatile computer-readable storage medium having recorded thereon instructions according to one disclosure of the present specification, wherein the instructions, when executed by one or more processors, cause the one or more processors to: decode a bitstream to generate a restored image including a restored non-region of interest and a region of interest, and extract information on a region of interest for the restored image; based on the information on the region of interest, restoring a region of interest-based processed image from the restored region of interest; and performing temporal restoration on the restored image based on the information on the region of interest.

[0019] According to the present disclosure, image analysis by a machine can be effectively performed.

[0020] Figure 1 schematically illustrates an example of a video / image coding system.

[0021] Figure 2 is a drawing schematically illustrating the configuration of a video / image encoding device.

[0022] Figure 3 is a drawing schematically illustrating the configuration of a video / image decoding device.

[0023] Figures 4a to 4d are exemplary diagrams showing a VCM encoder and a VCM decoder.

[0024] FIG. 5 illustrates a block diagram of an encoding device according to an embodiment of the present disclosure.

[0025] FIG. 6 illustrates detailed operations of an area-of-interest-based processor according to one embodiment of the present disclosure.

[0026] FIG. 7 illustrates detailed operations of an internal encoding performer according to an embodiment of the present disclosure.

[0027] FIG. 8 illustrates a block diagram of a decoding device according to an embodiment of the present disclosure.

[0028] FIG. 9 illustrates detailed operations of an internal decryption performer according to an embodiment of the present disclosure.

[0029] FIG. 10 illustrates detailed operations of a region-of-interest-based restorer according to one embodiment of the present disclosure.

[0030] FIG. 11 illustrates detailed operations of a temporal restoration performer according to an embodiment of the present disclosure.

[0031] FIG. 12a is an example of temporal restoration that generates an intermediate frame using two adjacent frames according to one embodiment of the present disclosure.

[0032] FIG. 12b is an example of temporal restoration that generates an intermediate frame using two non-adjacent frames according to one embodiment of the present disclosure.

[0033] FIG. 13 is an example of temporal restoration that generates a frame of a future time point in time using some of the frames in a decrypted image according to one embodiment of the present disclosure.

[0034] Specific structural or step-by-step descriptions of embodiments according to the concept of the present disclosure disclosed in this specification or application are merely illustrative for the purpose of explaining embodiments according to the concept of the present disclosure, and embodiments according to the concept of the present disclosure may be implemented in various forms, and embodiments according to the concept of the present disclosure may be implemented in various forms and should not be construed as being limited to the embodiments described in this specification or application.

[0035] Embodiments according to the concept of the present disclosure may have various modifications and take various forms. Therefore, specific embodiments are illustrated in the drawings and described in detail in this specification or application. However, this is not intended to limit embodiments according to the concept of the present disclosure to specific disclosed forms, and it should be understood that all modifications, equivalents, and alternatives included within the spirit and technical scope of the present disclosure are included.

[0036] While terms such as "first" and / or "second" may be used to describe various components, these components should not be limited by these terms. These terms are only intended to distinguish one component from another; for example, without departing from the scope of the present disclosure, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component."

[0037] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. Conversely, when a component is referred to as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Other expressions that describe the relationship between components, such as "between" and "directly between" or "adjacent to" and "directly adjacent to", should be interpreted similarly.

[0038] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the present disclosure. The singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0039] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0040] Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined herein.

[0041] In describing the embodiments, description of technical contents that are well known in the technical field to which the present disclosure belongs and are not directly related to the present disclosure will be omitted.

[0042] This is to convey the gist of the present disclosure more clearly without obscuring it by omitting unnecessary explanations.

[0043] This document relates to video / image coding. For example, the method / embodiment disclosed in this document may be related to the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), the essential video coding (EVC) standard, the AVS2 standard, etc.).

[0044] This document presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.

[0045] In this document, "video" can refer to a series of images over time. "Picture" generally refers to a unit representing a single image from a specific time period, and "slice" / "tile" are units that constitute part of a picture in coding.

[0046] A slice / tile can contain one or more coding tree units (CTUs). A picture can consist of one or more slices / tiles. A picture can consist of one or more tile groups. A tile group can contain one or more tiles.

[0047] A pixel or pel can mean the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Alternatively, a sample can mean a pixel value in the spatial domain, or when such a pixel value is converted to the frequency domain, it can mean a transform coefficient in the frequency domain.

[0048] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region.

[0049] A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term "unit" may sometimes be used interchangeably with the terms "block" or "area." In general, an MxN block can contain a set (or array) of samples (or array of samples) or transform coefficients, each consisting of M columns and N rows.

[0050] Figure 1 schematically illustrates an example of a video / image coding system.

[0051] Referring to FIG. 1, a video / image coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or streaming.

[0052] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a reception unit, a decoding device, and a renderer.

[0053] The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.

[0054] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.

[0055] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0056] The transmission unit can transmit encoded video / image information or data output in bitstream form to the receiving unit of the receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file using a predetermined file format and an element for transmission via a broadcasting / communication network.

[0057] The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0058] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.

[0059] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.

[0060] Figure 2 is a drawing schematically illustrating the configuration of a video / image encoding device.

[0061] The term “video encoding device” hereinafter may include a video encoding device.

[0062] Referring to FIG. 2, the encoding device (10a) may be configured to include an image partitioner (10a-10), a prediction unit (predictor) (10a-20), a residual processor (residual processor) (10a-30), an entropy encoder (entropy encoder) (10a-40), an adder (adder) (10a-50), a filter (filter) (10a-60), and a memory (10a-70). The prediction unit (10a-20) may include an inter prediction unit (10a-21) and an intra prediction unit (10a-22). The residual processing unit (10a-30) may include a transformer (10a-32), a quantizer (10a-33), a dequantizer (10a-34), and an inverse transformer (10a-35). The residual processing unit (10a-30) may further include a subtractor (10a-31). The addition unit (10a-50) may be called a reconstructor or a reconstructed block generator. The above-described image segmentation unit (10a-10), prediction unit (10a-20), residual processing unit (10a-30), entropy encoding unit (10a-40), addition unit (10a-50), and filtering unit (10a-60) may be configured by one or more hardware components (e.g., encoder chipset or processor) according to an embodiment. In addition, the memory (10a-70) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (10a-70) as an internal / external component.

[0063] The image segmentation unit (10a-10) can segment an input image (or picture, frame) input to the encoding device (10a) into one or more processing units.

[0064] For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a Quad-tree binary-tree ternary-tree (QTBTTT) structure. For example, one coding unit may be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary structure may be applied later. Alternatively, the binary-tree structure may be applied first. The coding procedure according to the present document may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.

[0065] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).

[0066] The subtraction unit (10a-31) can subtract the prediction signal (predicted block, prediction samples, or prediction sample array) output from the prediction unit (10a-20) from the input image signal (original block, original samples, or original sample array) to generate a residual signal (residual block, residual samples, or residual sample array), and the generated residual signal is transmitted to the conversion unit (10a-32). The prediction unit (10a-20) can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block.

[0067] The prediction unit (10a-20) can determine whether intra-prediction or inter-prediction is applied to the current block or CU unit. As described later in the description of each prediction mode, the prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit (10a-40). The prediction-related information can be encoded by the entropy encoding unit (10a-40) and output in the form of a bitstream.

[0068] The intra prediction unit (10a-22) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode.

[0069] In intra prediction, prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes and planar modes. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the granularity of the prediction direction.

[0070] However, this is only an example; depending on the settings, a greater or lesser number of directional prediction modes may be used. The intra prediction unit (10a-22) may also determine the prediction mode to be applied to the current block by utilizing the prediction mode applied to the surrounding blocks.

[0071] The inter prediction unit (10a-21) can derive a predicted block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit (10a-21) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (10a-21) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0072] The prediction unit (10a-20) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction to predict a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) to predict a block. The intra block copy can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0073] The prediction signal generated through the inter prediction unit (10a-21) and / or the intra prediction unit (10a-22) can be used to generate a reconstructed signal or a residual signal. The transform unit (10a-32) can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transform process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size.

[0074] The quantization unit (10a-33) quantizes the transform coefficients and transmits them to the entropy encoding unit (10a-40), and the entropy encoding unit (10a-40) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information.

[0075] The quantization unit (10a-33) can rearrange the quantized transform coefficients in the form of a block into a one-dimensional vector based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of the one-dimensional vector. The entropy encoding unit (10a-40) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc.

[0076] The entropy encoding unit (10a-40) may encode, together or separately, information necessary for video / image restoration (e.g., values ​​of syntax elements, etc.) in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted through a network or may be stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (10a-40) may be configured as an internal / external element of the encoding device (10a) by a transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal, or the transmitting unit may be included in the entropy encoding unit (10a-40).

[0077] The quantized transform coefficients output from the quantization unit (10a-33) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (10a-34) and the inverse transform unit (10a-35), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (10a-50) can add the reconstructed residual signal to the prediction signal output from the prediction unit (10a-20), thereby generating a reconstructed signal (reconstructed picture, reconstructed block, reconstructed samples, or reconstructed sample array). When there is no residual for the target block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.

[0078] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0079] The filtering unit (10a-60) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (10a-60) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (10a-70), specifically, in the DPB of the memory (10a-70). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), an adaptive loop filter, a bilateral filter, etc. The filtering unit (10a-60) can generate various information regarding filtering and transmit the information to the entropy encoding unit (10a-90), as described below in the description of each filtering method. The information regarding filtering may be encoded by the entropy encoding unit (10a-90) and output in the form of a bitstream.

[0080] The modified restored picture transmitted to the memory (10a-70) can be used as a reference picture in the inter prediction unit (10a-80). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (10a) and the decoding device, and can also improve encoding efficiency.

[0081] The DPB of the memory (10a-70) can store the modified restored picture to be used as a reference picture in the inter prediction unit (10a-21). The memory (10a-70) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transmitted to the inter prediction unit (10a-21) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (10a-70) can store restored samples of restored blocks in the current picture and transmit them to the intra prediction unit (10a-22).

[0082] Figure 3 is a drawing schematically illustrating the configuration of a video / image decoding device.

[0083] Referring to FIG. 3, the decoding device (10b) may be configured to include an entropy decoder (10b-10), a residual processor (10b-20), a predictor (10b-30), an adder (10b-40), a filter (10b-50), and a memory (10b-60). The predictor (10b-30) may include an inter-prediction unit (10b-31) and an intra-prediction unit (10b-32). The residual processor (10b-20) may include a dequantizer (10b-21) and an inverse transformer (10b-21). The entropy decoding unit (10b-10), residual processing unit (10b-20), prediction unit (10b-30), addition unit (10b-40), and filtering unit (10b-50) described above may be configured by a single hardware component (e.g., decoder chipset or processor) according to an embodiment. In addition, the memory (10b-60) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (10b-60) as an internal / external component.

[0084] When a bitstream including video / image information is input, the decoding device (10b) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (10b) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (10b) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (10b) can be reproduced through a reproduction device.

[0085] The decoding device (10b) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (10b-10). For example, the entropy decoding unit (10b-10) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information.

[0086] The decoding device can further decode the picture based on information about the parameter set and / or the general restriction information. The signaling / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through the decoding procedure. For example, the entropy decoding unit (10b-10) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients for the residual.

[0087] In more detail, the CABAC entropy decoding method receives a bin corresponding to each syntax element in a bitstream, determines a context model using information of the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of symbols / bins decoded in a previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit (10b-10), information regarding prediction is provided to the prediction unit (10b-30), and information regarding the residual on which entropy decoding has been performed by the entropy decoding unit (10b-10), i.e., quantized transform coefficients and related parameter information, can be input to the inverse quantization unit (10b-21).

[0088] In addition, information regarding filtering among the information decoded by the entropy decoding unit (10b-10) may be provided to the filtering unit (10b-50). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (10b), or the receiving unit may be a component of the entropy decoding unit (10b-10). Meanwhile, the decoding device according to the present document may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The above information decoder may include the entropy decoding unit (10b-10), and the sample decoder may include at least one of the inverse quantization unit (10b-21), the inverse transformation unit (10b-22), the prediction unit (10b-30), the addition unit (10b-40), the filtering unit (10b-50), and the memory (10b-60).

[0089] The inverse quantization unit (10b-21) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (10b-21) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (10b-21) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0090] In the inverse transform unit (10b-22), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).

[0091] The prediction unit can perform a prediction for the current block and generate a predicted block including prediction samples for the current block.

[0092] The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information about the prediction output from the entropy decoding unit (10b-10), and can determine a specific intra / inter prediction mode.

[0093] The prediction unit can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction to predict a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) to predict a block. The intra block copy can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0094] The intra prediction unit (10b-32) can predict the current block by referencing samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode.

[0095] In intra prediction, prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit (10b-32) may determine the prediction mode to be applied to the current block by utilizing the prediction modes applied to the surrounding blocks.

[0096] The inter prediction unit (10b-31) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.).

[0097] In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (10b-31) may construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information regarding the prediction may include information indicating the mode of inter prediction for the current block.

[0098] The addition unit (10b-40) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (10b-30). In cases where there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.

[0099] The addition unit (10b-40) may be called a restoration unit or a restoration block generation unit.

[0100] The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, can be output after filtering as described below, or can be used for inter prediction of the next picture.

[0101] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0102] The filtering unit (10b-50) can improve subjective / objective image quality by applying filtering to the restoration signal. For example, the filtering unit (10b-50) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and transmit the modified restoration picture to the memory (60), specifically, the DPB of the memory (10b-60). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0103] The (corrected) reconstructed picture stored in the DPB of the memory (10b-60) can be used as a reference picture in the inter prediction unit (10b-31). The memory (10b-60) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transmitted to the inter prediction unit (10b-31) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (10b-60) can store reconstructed samples of reconstructed blocks within the current picture and transmit them to the intra prediction unit (10b-32).

[0104] In this specification, the embodiments described in the prediction unit (10b-30), the inverse quantization unit (10b-21), the inverse transformation unit (10b-22), and the filtering unit (10b-50) of the decoding device (10b) can be applied to the prediction unit (10a-20), the inverse quantization unit (10a-34), the inverse transformation unit (10a-35), and the filtering unit (10a-60) of the encoding device (10a) in the same manner or correspondingly.

[0105] As described above, prediction is performed to increase compression efficiency when performing video coding. Through this, a predicted block including prediction samples for a current block, which is a coding target block, can be generated. Here, the predicted block includes prediction samples in a spatial domain (or pixel domain). The predicted block is derived identically from an encoding device and a decoding device, and the encoding device can increase video coding efficiency by signaling information (residual information) about the residual between the original block and the predicted block, rather than the original sample value of the original block itself, to a decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by combining the residual block and the predicted block, and can generate a reconstructed picture including the reconstructed blocks.

[0106] The above residual information can be generated through transformation and quantization procedures.

[0107] For example, the encoding device can derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (a residual sample array) included in the residual block to derive transform coefficients, perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and signal related residual information to a decoding device (via a bitstream). Here, the residual information can include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device can perform an inverse quantization / inverse transform procedure based on the residual information to derive residual samples (or residual blocks). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also inverse quantize / inverse transform the quantized transform coefficients to derive a residual block for reference in inter prediction of a subsequent picture, and generate a reconstructed picture based on the residual block.

[0108] <VCM(Video coding for Machines)>

[0109] With the recent advancements in various industries such as surveillance, intelligent transportation, smart cities, intelligent industry, and intelligent content, the amount of image or feature map data consumed by machines is increasing. In contrast, traditional video compression methods currently in use were developed with human vision in mind, and therefore contain unnecessary information, making them inefficient for machine tasks. For example, the resolution of images from the viewer's perspective may be higher than that of images (e.g., feature maps) from the machine's perspective. Therefore, research on video codec technologies that efficiently compress feature maps for machine tasks is needed.

[0110] The Moving Picture Experts Group (MPEG), an international standardization group for multimedia encoding, is discussing Video Coding for Machines (VCM). VCM is an image or feature map encoding technology that targets machine vision, rather than human viewer vision. In this document, feature maps can be referred to as "feature maps," and features can be referred to as "features."

[0111] Figures 4a to 4d are exemplary diagrams showing a VCM encoder and a VCM decoder.

[0112] Referring to FIG. 4a, a VCM encoder (100a) and a VCM decoder (100b) are shown.

[0113] When a VCM encoder (100a) encodes a video and / or a feature map and transmits it as a bitstream, a VCM decoder (100b) can decode and output the bitstream. At this time, the VCM decoder (100b) can output one or more videos and / or feature maps. For example, the VCM decoder (100b) can output a first feature map for machine-based analysis and a first image for user viewing. The first image can have a higher resolution than the first feature map.

[0114] Referring to FIG. 4b, a feature extractor for extracting a feature map may be connected to the front end of the VCM encoder (100a).

[0115] The VCM encoder (100a) may include a feature encoder.

[0116] The VCM decoder (100b) may include a feature decoder and a video reconstructor. The feature decoder may decode a feature map from a bitstream and output a first feature map for machine-assisted analysis. The video reconstructor may regenerate and output a first video from the bitstream for viewing by a user.

[0117] Referring to Fig. 4c, a feature extractor for extracting a feature map is connected to the front end of the VCM encoder (100). The VCM encoder (100a) may include a feature encoder.

[0118] The VCM decoder (100b) may include a feature decoder. The feature decoder may decode a feature map from a bitstream and output a first feature map for machine-based analysis. That is, the bitstream may be encoded only as a feature map, not as an image. To elaborate, the feature map may be data containing information about features for processing a specific task of a machine based on an image.

[0119] Referring to FIG. 4d, a feature extractor may be connected to the front end of the VCM encoder (100a).

[0120] The VCM encoder (100a) may include a feature converter and a video encoder. The video encoder may be the encoding device (10a) illustrated in FIG. 2.

[0121] The VCM decoder (100b) illustrated in FIG. 4d may include a video decoder and an inverse converter. The video decoder may be the decoding device (10b) illustrated in FIG. 3.

[0122] FIG. 5 illustrates a block diagram of an encoding device according to an embodiment of the present disclosure.

[0123] An encoding device (10a) according to one embodiment of the present disclosure can receive an image (video), perform temporal and spatial resampling, process the image based on a region of interest, perform bit truncation, and generate and output a bit stream through internal encoding.

[0124] The encoding device (10a) according to various embodiments can adaptively perform temporal resampling / temporal restoration without separate additional signaling for temporal restoration by using the region of interest information used in region of interest-based processing (RoI-based processing) in the temporal restoration process. In various embodiments, by transmitting information about the region of interest, intermediate frames can be generated periodically / aperiodically for temporal restoration during the decoding process.

[0125] The encoding device (10a) may include a temporal resampling performer (510), a spatial resampling performer (520), a region-of-interest-based processor (530), a bit truncation performer (540), and an internal encoding performer (550).

[0126] The input image (video) to be encoded may be an original input image, and / or may be one or more feature maps extracted from the input image by a neural network. In the present disclosure, the term "image" may refer to the image (video) itself, and / or may refer to a "feature map."

[0127] The order of components included in the encoding device (10a) illustrated in FIG. 5 may be changed in various embodiments, and some components may be omitted. Depending on the embodiment, the temporal resampling performer (510) and / or the spatial resampling performer (520) may be omitted, and when the temporal resampling performer (510) and / or the spatial resampling performer (520) are omitted, an image processed based on a region of interest by a region of interest-based processor (530) may become an input to the internal encoding performer (550). For example, if the temporal resampling performer (510) and the spatial resampling performer (520) are omitted, one or more regions of interest are directly extracted from the input image by the region-of-interest-based processor (530), and the extracted regions of interest are input to the internal encoding performer (550), so that the internal encoding performer (550) can perform encoding to generate a bitstream.

[0128] For convenience of explanation in the present disclosure, the encoding device (10a) may be referred to as an encoder or encoder, and the decoding device (10b) may be referred to as a decoder or decoder.

[0129] The temporal resampling performer (510) can receive an image, perform frame-by-frame sampling, and output an image in which some frames are sampled. The temporal resampling performer (510) can change the frame rate of some frames. According to an embodiment, the temporal resampling performer (510) can perform the same degree of sampling on the entire image. That is, the temporal resampling performer (510) can resample the entire image at the same frame rate. According to another embodiment, the temporal resampling performer (510) can perform temporal resampling by applying different sampling degrees to each group of frames in the image. That is, the temporal resampling performer (510) can resample at a variable frame rate on a group-by-group basis.

[0130] In one embodiment, the temporal resampling performer (510) may perform sampling at different sampling rates for the region-of-interest-based processed image and the non-region-of-interest image. For example, the temporal resampling performer (510) may perform sampling at a first sampling rate for the region-of-interest-based processed image and at a second sampling rate for the non-region-of-interest image.

[0131] According to an embodiment, the temporal resampling performer (510) may transmit information used in the temporal resampling process (such as a temporal sampling rate, in one example) to the decoder. The decoder may perform decoding on the region-of-interest-based processed image and the non-region-of-interest image, and then parse the respective sampling rates corresponding to the region-of-interest-based processed image and the non-region-of-interest image to perform upsampling on each. At this time, the sampling rate applied to the region-of-interest-based processed image and the sampling rate applied to the non-region-of-interest image may be different.

[0132] The spatial resampling performer (520) can receive an input image or an image on which temporal resampling has been performed and output an image with a changed spatial resolution for each frame or a series of frames. The spatial resampling performer (520) can perform frame-by-frame or sequence-by-sequence sampling on the input image to output an image in which one or more frames in the sequence are spatially sampled (with a changed spatial resolution). The spatial resampling performer (520) can change the resolution for a region of interest included in each frame of the input image. In one embodiment, the spatial resampling performer (520) can change the resolution only for the region of interest. If there are multiple regions of interest in one frame, the spatial resampling performer (520) can change the resolution for each region of interest. Information on each region of interest can include changed resolution information. For example, a first scale factor can be applied to a first region of interest, and a second scale factor can be applied to a second region of interest. The scale factor can be stored using at least one of the following parameters: scale factor, scale factor nominator, scale factor, denominator, size(%), etc.

[0133] In one embodiment, the spatial resampling performer (520) can perform spatial sampling to have the same spatial resolution size through the same degree of sampling throughout the sequence. The resampling method for each input image can be any one of bilinear downsampling, bilateral downsampling, and a deep neural network including one or more convolutional layers. The spatial resampling performer (520) can select a method corresponding to each input image from among multiple resampling methods and signal it to the decoder.

[0134] In another embodiment, the spatial resampling performer (520) can perform spatial resampling at different sampling levels (variable spatial resolution) for each group of frames within an image.

[0135] The spatial resampling performer (520) can transmit information used in the spatial resampling process (e.g., a spatial sampling rate of each frame and / or a spatial sampling rate per series of frames, etc.) to the decoder. After decoding the bitstream, the decoder can parse the resolution information for the region-of-interest-based processed image and the non-region-of-interest image and change the resolution of each. At this time, the resolution applied to the region-of-interest-based processed image and the resolution applied to the non-region-of-interest may be different.

[0136] The region-of-interest-based processor (530) can receive an input image, an image on which temporal resampling has been performed, or an image on which temporal resampling and spatial resampling have been performed, extract a region of interest existing in each frame or a series of frames, and output a processed image based on the extracted region of interest. The region-of-interest-based processor (530) can transmit information used in the region-of-interest-based processing (in one example, region-of-interest ID information, region-of-interest location, region-of-interest size information, packing information, temporal sampling rate, spatial resolution, etc.) to a decoder.

[0137] The bit truncation performer (540) can output data in a form in which the bit depth of a component of an input image has been changed. Depending on the embodiment, the bit truncation performer (540) can change only the bit length of a specific component of the input data. The bit truncation performer (540) can transmit information used in the bit truncation process (such as the degree to which the bit depth has been changed, in one example) to the decoder.

[0138] The internal encoding performer (550) can receive an input image, or an image on which some or all of temporal / spatial resampling and region of interest processing have been performed, and perform image encoding to generate a bitstream. According to an embodiment, the internal encoding performer (550) can use a 2D video encoder (e.g., AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, ​​etc.) and can use a 2D video encoder including one or more convolution layers. According to an embodiment, encoding can be performed after converting the color space of the input image of the internal encoding performer (550) to a color space such as YUV420 or YUV444. At this time, according to an embodiment, the internal encoding performer (550) can convert through a conversion method defined by an agreement between the decoder and the decoder, and can transmit color space conversion information to the decoder.

[0139] In various embodiments of the present disclosure, when region-of-interest-based processing is performed on an input image and temporal resampling is performed, the encoding device (10a) can transmit only information about the region of interest without directly encoding and transmitting information necessary for temporal restoration in the decoding device (10b), thereby allowing the decoding device (10b) to extract the necessary information and perform temporal restoration.

[0140] In various embodiments, when the sizes of the regions of interest of two adjacent or non-adjacent frames are the same, the decoding device (10b) may compare one or more metric measurements of PSNR or SSIM for the regions of interest within each frame to determine a temporal frame restoration method. When region-of-interest-based processing is performed, the background, which is a region of no interest, may be filled with a specific value such as a median, and when temporal resampling is performed through SSIM-based comparison during the decoding process, a problem of deterioration of temporal restoration performance may occur. In other words, when the input image is divided into a region of interest and a region of no interest (such as a background) and the input image is changed through processing of the region of no interest, it may be difficult to use the existing method as it is to restore the original image from the changed input image in the decoder. To solve this problem, in various embodiments of the present disclosure, when region-of-interest-based processing is performed, temporal restoration may be performed using only information within the region of interest. In various embodiments, when the encoder processes the region of interest and the region of no interest included in the input image differently, the encoder may signal information about the region of interest of the input image to the decoder so that the decoder can restore the entire input image using the information about the region of interest of the input image.

[0141] In one embodiment, when comparing only the region of interest between frames, different temporal restoration methods can be applied depending on whether the distribution of objects within the region of interest is large or small. In one embodiment, the method for performing temporal restoration can be determined after comparing only the region of interest by considering cases where the size of the region of interest within the frames to be restored is the same or similar. In various embodiments, in addition to interpolation, extrapolation can also be applied to perform temporal restoration.

[0142] FIG. 6 illustrates detailed operations of an area-of-interest-based processor according to one embodiment of the present disclosure.

[0143] An area-of-interest-based processor (530) of an encoding device (10a) according to one embodiment of the present disclosure may include a frame analysis unit (531), an area-of-interest selection unit (532), and an area-of-interest-based image processing unit (533). In various embodiments, some of the detailed components of the area-of-interest-based processor (530) may be omitted, and the order of the components in FIG. 6 may be changed.

[0144] The frame analysis unit (531) can receive an input image and output information (for example, the location of the region of interest, the size of the region of interest, a score for the region of interest, etc.) of one or more region of interest candidates existing in each frame (picture) of the image. The input image may be an original image, or may be an image to which temporal and / or spatial resampling has been applied according to various embodiments.

[0145] The region of interest selection unit (532) can select a region of interest from among the region of interest candidates output from the frame analysis unit (531), determine a corresponding relationship among multiple regions of interest, and select a corresponding region of interest. The selected region of interest can be transmitted to the internal encoding performer (550) for image encoding. Depending on the embodiment, non-region of interest images of all frames and / or some frames can be transmitted to the internal encoding performer (550).

[0146] The region of interest-based image processing unit (533) may receive information on regions of interest selected in an image and generate a processed image based on the region of interest for each frame using the regions of interest selected in various embodiments. In one example, the region of interest-based image processing unit (533) may pack regions of interest and generate each frame in which the regions of interest are packed. In one embodiment, the region of interest-based image processing unit (533) may fill non-regions of interest excluding the region of interest in a frame with at least one specific value (e.g., an intermediate value, etc.) and generate a frame in which the non-regions of interest are filled with the specific values. In various embodiments, the region of interest-based image processing unit (533) may determine processing for non-regions of interest in various ways. For example, the region of interest-based image processing unit (533) may not encode all or part of the non-regions of interest, may fill the non-regions of interest with a specific value determined in a predetermined manner, or may adaptively process the non-regions of interest to the region of interest. Such embodiments are merely examples, and the region of interest-based image processing unit (533) can process the region of interest differently from the region of no interest, and can signal information about the region of interest so that the decoder can restore the region of no interest using only information about the region of interest.

[0147] According to an embodiment, the region of interest-based processor (530) may transmit information used in the region of interest extraction process (in one example, region of interest identification (ID) information, region of interest location, region of interest size information, packing information, etc.) to the decoder.

[0148] FIG. 7 illustrates detailed operations of an internal encoding performer according to an embodiment of the present disclosure.

[0149] An internal encoding performer (550) of an encoding device (10a) according to an embodiment of the present disclosure can generate a bitstream by encoding an image processed based on a region of interest, an image of a non-region of interest, and information on a region of interest. An internal encoding performer (550) according to an embodiment of the present disclosure can include a downsampling performer (541), an image encoding performer (542), and an area of ​​interest information encoding performer (543).

[0150] According to an embodiment, the internal encoding performer (550) may perform encoding for each frame for an image processed based on a region of interest, and may perform encoding for only some frames within a group of frames for an image of a non-region of interest.

[0151] The downsampling unit (541) can perform downsampling on the non-region of interest image and the region of interest-based processed image, respectively. According to an embodiment, the downsampling unit (541) can downsample the non-region of interest image and the region of interest-based processed image at different sampling rates. When downsampling is performed, the internal encoding unit (550) can signal sampling rate information to the decoder.

[0152] Depending on the embodiment, the input image may be subjected to both spatial resampling by the spatial resampling performer (520) and downsampling within the internal encoding performer (550), or only one of the two methods may be applied.

[0153] According to an embodiment, the downsampling method for each image may be one of bilinear downsampling, bilateral downsampling, and a deep neural network including one or more convolutional layers, and the internal encoding performer (550) may signal the method used.

[0154] The video encoding unit (542) can perform encoding on a non-region of interest image and a region of interest-based processed image. The video encoding unit (542) can use a video encoder (AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, ​​etc.) and can use a 2D video encoder including one or more convolution layers. Depending on the embodiment, the region of interest-based processed image and the non-region of interest image can use the same video encoder or different video encoders.

[0155] The image encoding performing unit (542) can perform MUXing on bitstreams generated for images processed based on the region of interest, images of non-region of interest, and information on the region of interest to generate one bitstream.

[0156] The region of interest information encoding unit (543) can encode information about a region of interest in an image processed based on a region of interest for which downsampling has been performed. The region of interest information encoding unit (543) can encode information (for example, location, movement, mapping information, etc.) of regions of interest of all frames or some frames in an input image. According to an embodiment, the region of interest information encoding unit (543) can utilize some processes (for example, entropy coding, etc.) of the image encoding unit (542).

[0157] FIG. 8 illustrates a block diagram of a decoding device according to an embodiment of the present disclosure.

[0158] A decoding device (10b) according to an embodiment of the present disclosure may receive a bitstream from an encoding device (10a), perform decoding, and output a restored image (video). The decoding device (10b) may include an internal decoding performer (810), a bit compensation performer (820), a region-of-interest-based restorer (830), a spatial restoration performer (840), a temporal restoration performer (850), and a post-processing filter performer (860). Depending on the embodiment, the order of the components in FIG. 8 may be changed. According to an embodiment, the order of the region-of-interest-based restorer (830), the spatial restoration performer (840), and the temporal restoration performer (850) may be the reverse order of the order in which each corresponding process (temporal sampling performer (510), spatial sampling performer (520), and region-of-interest-based processor (530)) is performed in the encoding process of the original image according to the encoding device (10a) of FIG. 5. Alternatively, the decoding order of the components of the decoding device (10b) may be different from the reverse order of the encoding order of the original image according to the encoding device (10a) of FIG. 5.

[0159] The internal decoding performer (810) can receive a bitstream as input and perform image decoding to generate a restored image. According to an embodiment, the image decoding can use a 2D video decoder (AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, ​​etc.) and a 2D video decoder including one or more convolution layers. According to an embodiment, the internal decoding performer (810) can implicitly or / and explicitly convert the color space of the restored image to another color space and then perform the image decoding process thereafter. As an example, when the color space of the restored image is not RGB444 but one of YUV420, YUV444, etc., the internal decoding performer (810) can implicitly or / and explicitly convert it to the RGB444 space and then perform the image decoding process thereafter. A detailed description of internal decryption is provided later in Fig. 9.

[0160] The bit compensation performer (820) can compensate for the bit depth of a specific component or all components of the restored data by using the information used in the bit truncation process transmitted from the encoding device (10a) for an image in which the above process has been partially or fully performed.

[0161] The region-of-interest-based restorer (830) can restore a region-of-interest-based processed image using the decoded image and region-of-interest-based processing information (for example, region-of-interest ID information, region-of-interest size information, packing information, etc.) transmitted from the encoder. The region-of-interest-based restorer (830) may include an image reconstruction unit (831), and the image reconstruction unit (831) may include a region boundary filtering performing module and a deblurring filtering module. A detailed description of region-of-interest-based restoration is described later with reference to FIG. 10.

[0162] The spatial restoration performer (840) can obtain an image on which spatial restoration has been performed using the decoded image or the decoded and region-of-interest-based restoration image and the information used in the spatial resampling process transmitted from the encoder (for example, the spatial sampling rate of each frame and / or the spatial sampling rate of a series of frames, etc.). During the encoding process, the resolution of the region of interest included in each frame of the input image may be changed. In one embodiment, the resolution may be changed only for the region of interest, and if there are multiple regions of interest in one frame, the resolution may be changed differently for each region of interest. The spatial restoration performer (840) can parse the resolution information included in the information on the region of interest and perform resolution restoration on each region of interest. For example, the information on the region of interest may include a scale factor for changing the resolution, and a first scale factor may be applied to a first region of interest, and a second scale factor may be applied to a second region of interest. At this time, the scale factor can be stored using at least one of the following parameters: scale factor, scale factor nominator, scale factor, denominator, size(%), etc.

[0163] The temporal restoration performer (850) can obtain an image on which temporal restoration has been performed using information (e.g., temporal sampling rate, etc.) used in the temporal resampling process transmitted from the encoder and an image on which the above process has been partially or fully performed. A detailed description of temporal restoration is provided below in FIG. 11.

[0164] The post-processing filter performer (860) may perform filtering on an image on which part or all of the above process has been performed, depending on the embodiment. Depending on the embodiment, the post-processing filter performer (860) may use a fixed filter, or may utilize a plurality of filters set by agreement between the encoding device (10a) and the decoding device (10b). The post-processing filter performer (860) may receive filter information agreed upon in advance from the encoding device (10a).

[0165] FIG. 9 illustrates detailed operations of an internal decryption performer according to an embodiment of the present disclosure.

[0166] An internal decoding performer (810) included in a decoding device (10b) according to an embodiment of the present disclosure may include an image encoding performer (811), an upsampling performer (812), and an area of ​​interest information decoding performer (813). According to an embodiment, a bitstream input from the internal decoding performer (810) may first undergo demuxing and then be applied as input to each of the decoding performers (811, 813) of FIG. 9.

[0167] The internal decoding performer (810) can receive a bitstream and perform decoding to restore image and region of interest information. Depending on the embodiment, the number of frames of the restored region of interest-based processed image and the number of frames of the restored non-region of interest image may be different.

[0168] The upsampling unit (812) may receive a sampling rate and perform upsampling on the region-of-interest-based processed image and the non-region-of-interest image after decoding. The sampling rates may be different for the region-of-interest-based processed image and the non-region-of-interest image. According to an embodiment, the upsampling unit (812) may select and use one of bilinear upsampling, bilateral upsampling, nearest-neighbor upsampling, and a deep neural network including one or more convolutional layers as an upsampling method. According to an embodiment, the upsampling unit (812) may perform upsampling on the region-of-interest-based processed image and the non-region-of-interest image using different upsampling methods, and may also variably select and apply the upsampling method within the region-of-interest-based processed image. For example, the upsampling performing unit (812) can determine the upsampling method according to the sampling rate, and if the sampling rate is different, the upsampling method may also be different.

[0169] According to an embodiment, when information on a downsampled region of interest is encoded in an encoder, an upsampling performing unit (812) may perform upsampling on each region of interest information after decoding the region of interest information.

[0170] The region of interest information decoding unit (813) can receive a bitstream, perform decoding, and restore the region of interest information. If the region of interest information down-sampled by the encoding device (10a) is encoded, the region of interest information decoding unit (813) can perform decoding on the region of interest information and then transmit it to the upsampling unit (812) to perform upsampling. According to an embodiment, if multiple regions of interest exist in one frame, the region of interest information decoding unit (813) can perform decoding on each region of interest.

[0171] FIG. 10 illustrates detailed operations of a region-of-interest-based restorer according to one embodiment of the present disclosure.

[0172] A region of interest-based restorer (830) included in a decoding device (10b) according to one embodiment of the present disclosure can receive a restored region of interest image and a restored non-region of interest image and reconstruct an internally decoded image using information about the region of interest.

[0173] The region of interest-based restorer (830) may include an image reconstruction unit (831), and the image reconstruction unit (831) may include a region boundary filtering module and a deblurring filtering module.

[0174] The image reconstruction unit (831) can perform filtering for resolution improvement on the restored image. The image reconstruction unit (831) can further reconstruct the image using restored region of interest information to generate a restored image. In various embodiments, the image reconstruction unit (831) can perform filtering on at least one of the boundary and region of interest for the region of interest-based processed image and the non-region of interest image. In one embodiment, the image reconstruction unit (831) can perform boundary filtering of the region of interest-based processed image and the non-region of interest using a region boundary filtering performing module. In another embodiment, the image reconstruction unit (831) can perform filtering, such as deblurring, on one or more regions of interest of the region of interest-based processed image.

[0175] In various embodiments, the image reconstruction unit (831) may perform filtering for resolution improvement to restore an image to which temporal resampling has been applied. In the temporal resampling process that deletes intermediate frames, intermediate frames with relatively low quality may be restored due to quality differences between frames. If deblurring filtering is performed on a region of interest of the restored intermediate frames, the quality of the intermediate frames may be improved. For example, the image reconstruction unit (831) may perform deblurring filtering on a region of interest of two frames input in the temporal resampling process to generate a restored image. Depending on the embodiment, the image reconstruction unit (831) may perform filtering using any one of several deep neural networks trained for various functions. For example, the image reconstruction unit (921) may perform deblurring within a region of interest using a deep neural network trained for the purpose of deblurring. When deblurring filtering is applied, restoration performance may be improved.

[0176] A restored non-region of interest image may exist for each frame of the restored region of interest, and one restored non-region of interest image may exist for each series of frame groups. For example, a restored non-region of interest image may exist only for the intra frame in each GOP unit. According to an embodiment, when one restored non-region of interest image exists for each series of frame groups, the image reconstruction unit (831) may generate a restored image using the same restored non-region of interest image for each frame of the restored region of interest. According to an embodiment, for each restored region of interest, different filtering may be applied to the boundary of the region of interest and the boundary of the restored non-region of interest image at the corresponding position. According to an embodiment, the image reconstruction unit (831) may implicitly determine a filtering method based on the difference in pixel values ​​between the boundary of the region of interest and the boundary of the restored non-region of interest image at the corresponding position for each restored region of interest. According to an embodiment, the image reconstruction unit (831) may receive a filtering method for a specific region of interest from the encoding device (10a) and perform filtering.

[0177] FIG. 11 illustrates detailed operations of a temporal restoration performer according to an embodiment of the present disclosure.

[0178] A temporal restoration performer (850) of a decoding device (800) according to an embodiment of the present disclosure may include a restoration target determination module (851) for temporal frame restoration, a restoration degree determination module (852), a restoration method determination module (853), and a temporal restoration execution module (854). In various embodiments, the detailed configuration of FIG. 11 may be omitted or the order may be changed.

[0179] The temporal restoration performer (850) can perform temporal restoration on the decoded image. In various embodiments, when temporal resampling is performed in the encoding device (10a), some frames of the original image may be omitted to generate a bitstream. In response, the decoding device (10b) can decode the bitstream and then perform temporal restoration to generate some of the omitted frames, thereby restoring the original image. In various embodiments, the decoded image may be an image that has been decoded by an internal decoding performer and then subjected to region-of-interest-based restoration and / or spatial restoration.

[0180] The temporal restoration performer (850) can perform temporal restoration by generating intermediate frames at arbitrary locations in the decoded image corresponding to some frames omitted from the original image based on information obtained by performing temporal resampling in the encoding device (10a). For example, the temporal restoration performer (850) can generate one or more new intermediate frames between a first frame and a second frame located in chronological order in the decoded image. At this time, the first frame and the second frame may be frames that are immediately adjacent in chronological order, or may not be immediately adjacent. In various embodiments, the temporal restoration performer (850) can generate intermediate frames based on the first frame and the second frame by an interpolation method, and can generate intermediate frames by an extrapolation method.

[0181] According to one embodiment, the temporal restoration performer (850) may determine whether temporal restoration has been performed periodically, and if so, may perform temporal restoration by generating the same number of intermediate frames for each period, and may determine a method for generating the intermediate frames. The method for generating the intermediate frames to perform temporal restoration may be one of the methods for generating one or more intermediate frames based on information about temporal resampling transmitted from the encoding device (10a). For example, the temporal restoration performer (850) may copy a frame immediately preceding the intermediate frame to be generated, and then perform interpolation using the immediately preceding frame and the immediately succeeding frame to generate the intermediate frame. As another example, the temporal restoration performer (850) may select a frame having the same size of the region of interest or the same number of regions of interest among the temporally preceding and succeeding frames to be generated, copy one of the selected frames, and then generate the intermediate frame through an interpolation method using the selected frames. For another example, the temporal restoration performer (850) may perform extrapolation using a temporally previous frame for an intermediate frame to be generated to generate the intermediate frame. In another embodiment, the temporal restoration performer (850) may generate an intermediate frame using the current frame based on information about a region of interest included in the decoded current frame. For example, the temporal restoration performer (850) may select an adjacent frame to be used to generate the intermediate frame based on at least one of the number, position, and size of the region of interest included in the decoded current frame and at least one metric measurement value of the peak signal-to-noise ratio (PSNR) and the structural similarity index measure (SSIM) of another frame adjacent to the current frame.In various embodiments, the temporal restoration performer (850) may apply various comparison methods other than PSNR and SSIM to select adjacent frames to use for temporal restoration. For example, various comparison methods may use metrics that can determine the degree of similarity between frames and / or between regions of interest.

[0182] The temporal restoration performer (850) can check whether temporal resampling has been performed periodically in the encoding device (10a), and if it has been performed aperiodically, determine the number of intermediate frames to be generated through temporal restoration using information on decoded frames, and determine a method for generating intermediate frames.

[0183] The temporal restoration performer (850) can check whether processing has been performed on the region of interest for the original image. For example, the temporal restoration performer (850) can decode the bitstream and, if the RoI_coding_flag is true, can check whether processing has been performed on the region of interest. The temporal restoration performer (850) can check whether temporal resampling has been performed on the original image. For example, the temporal restoration performer (850) can decode the bitstream and, if the Temp_Resamp_flag is true, can check whether temporal resampling has been performed on the original image and some frames have been omitted periodically / aperiodically.

[0184] In various embodiments, the temporal restoration performer (850) may determine whether to perform temporal restoration after checking whether temporal resampling has been performed regardless of whether processing has been performed on the region of interest.

[0185] The restoration target determination module (851) can determine one or more frames to be used for temporal restoration among the frames in the decoded image. The number of frames to be used for temporal restoration for generating intermediate frames may be q, for an integer q greater than or equal to 1, and q may vary within a sequence.

[0186] FIG. 12a is an example of temporal restoration that generates an intermediate frame using two adjacent frames according to one embodiment of the present disclosure.

[0187] In various embodiments, the restoration target determination module (851) may generate one or more intermediate frames that are temporally intermediate between two frames using only two adjacent frames in the decoded image.

[0188] Referring to Figure 12a, the current frame (F k ) and the next frame (F) which is the frame immediately following it in time. k+1 ) to get the current frame (F k ) and the next frame (F k+1 ) can create two new frames (nf1, nf2). The method to create a new frame (nf1, nf2) is to create the current frame (F k ) or next frame (F k+1 ) can be applied in various ways, such as copying, performing interpolation, etc.

[0189] FIG. 12b is an example of temporal restoration that generates an intermediate frame using two non-adjacent frames according to one embodiment of the present disclosure.

[0190] In various embodiments, the restoration target determination module (851) may determine a frame to be used for temporal restoration based on information about a region of interest of frames within a decoded image. The restoration target determination module (851) may generate one or more intermediate frames that are temporally intermediate between the current frame and the arbitrary frame by using any frame within the decoded image that is temporally subsequent to the decoded current frame, rather than the decoded frame that is closest to the current frame.

[0191] According to an embodiment, the restoration target determination module (851) may determine a frame to be used for temporal restoration by comparing the region of interest information of the decoded current frame with the region of interest information of the decoded frames in temporally subsequent order. The region of interest information may be the number of regions of interest in the frame, the size of the regions of interest, etc. For example, a restored frame having the same number of regions of interest as the number of regions of interest in the decoded current frame may be selected as a frame to be used for temporal restoration. A restored frame having a difference in the size of the region of interest in the decoded current frame and a size of the region of interest within a certain threshold may be selected as a frame to be used for temporal restoration. In this case, the size of the region of interest may be the area or ratio occupied by the region of interest existing in the frame. If the number of regions of interest is the same and the difference in the size of the region of interest between the two frames is within a certain threshold, the temporal frame restoration target determination module (851) may select the restored frame to perform temporal restoration.

[0192] According to an embodiment, the restoration target determination module (851) may perform temporal restoration by comparing one or more metric measurements of the PSNR (Peak signal-to-noise ratio) and / or SSIM (Structural similarity index measure) between the decoded current frame and subsequent frames and selecting the frame that is temporally closest to the current frame that satisfies a case where the metric is smaller than a specific threshold (or, in another embodiment, larger than the threshold). In the present disclosure, the average value may mean a representative value, and the representative value may be defined as at least one of a weighted average value, an average value, a median value, and a partial value.

[0193] The restoration target determination module (851) can independently perform the comparison of metric measurement and region of interest information. Alternatively, the restoration target determination module (851) can continuously perform the comparison of metric measurement and region of interest information to determine a temporal restoration target frame. The restoration target determination module (851) can perform the restoration target determination process only for frames within a certain temporal range based on the decoded current frame. According to an embodiment, the restoration target determination module (851) can perform temporal restoration using one or more frames among the frames in the decoded image and one or more frames among the frames that have already been temporally restored.

[0194] Referring to Figure 12b, the current frame (F k ) and the next frame (F) which is the frame immediately following it in time. k+1 ) in addition to the current frame (F k ) any frame after that (F k+m ) together to get the current frame (F k ) and the next frame (F k+1 ) can create two new frames (nf1, nf2). Any frame (F k+m ) is the current frame (F k) can be selected based on the degree of similarity with the current frame (F). For example, if the number of regions of interest is the same, the difference in the size of the regions of interest is within a threshold, or the metric measurement values ​​of PSNR and / or SSIM for the regions of interest are less than a preset threshold (or, in another embodiment, greater than the threshold), the frame can be selected as the frame to be used for temporal restoration. A method for generating a new frame (nf1, nf2) is to generate a current frame (F k ), next frame (F k+1 ), any frame (F k+m ) can be applied in various ways, such as copying one of them, performing interpolation, etc.

[0195] FIG. 13 is an example of temporal restoration that generates a frame of a future time point in time using some of the frames in a decrypted image according to one embodiment of the present disclosure.

[0196] The restoration target determination module (851) can generate a frame of a temporally future point in time using n frames from among the frames in the restored image. The restoration target determination module (851) can determine a frame to be used for temporal restoration based on region-of-interest information of frames in the decoded restored image, according to an embodiment. The restoration target determination module (851) can determine a frame to be used for temporal restoration by comparing region-of-interest information of the decoded restored current frame with region-of-interest information of restored frames in a temporally subsequent order, according to an embodiment. The region-of-interest information may be the number of regions of interest in the frame, the size of the region of interest, etc. The restoration target determination module (851) can select a restored frame having the same number of regions of interest as the number of regions of interest of the decoded restored current frame as the frame to be used for temporal restoration, according to an embodiment. The restoration target determination module (851) can select a restored frame having a difference value between the size of the region of interest of the decoded restored current frame and the size of the region of interest within a certain threshold value as the frame to be used for temporal restoration, according to an embodiment. The restoration target determination module (851) may, according to an embodiment, perform temporal restoration by selecting a restored frame when the number of regions of interest is the same and the difference in the sizes of the regions of interest of the two frames is within a certain threshold. The restoration target determination module (851) may, according to an embodiment, compare one or more metric measurements, such as PSNR and / or SSIM, between the decoded restored current frame and subsequent frames, and determine the frame that is temporally closest to the current frame and satisfies the case where the value is less than a certain threshold (or greater than the threshold in another embodiment), as a frame to be used for temporal restoration. The temporal frame restoration target determination module (851) may, according to an embodiment, independently perform the comparison of the above-described metric measurements and region of interest information.Alternatively, the temporal frame restoration target determination module (851) may determine a temporal restoration target frame by continuously performing a comparison of metric measurements and region of interest information. The restoration target determination module (851) may, according to an embodiment, perform the target determination process only for frames within a certain temporal range based on the decoded restored current frame. The temporal frame restoration target determination module (851) may, according to an embodiment, perform temporal restoration using one or more frames among the frames in the decoded restored image and one or more frames among the frames that have already been temporally restored.

[0197] In various embodiments, the temporal restoration performer (850) may determine the degree of restoration for temporal restoration. For example, the degree of restoration determination module (852) of the temporal restoration performer (850) may determine the number of frames to be restored. When temporal resampling is performed periodically or aperiodically on the original image, the number of frames to be restored may differ, and the number of frames to be restored may vary for each sequence.

[0198] In one embodiment, the restoration degree determination module (852) may receive the number of frames to be temporally restored from the encoding device (10a) through the decoded current frame when temporal restoration is performed periodically. The restoration degree determination module (852) may determine the degree of temporal restoration for each sequence within the image. In one example, the degree of restoration may be constant within a sequence, or the degree of restoration may vary within a specific range depending on conditions.

[0199] According to an embodiment, the temporal frame restoration degree determination module (852) may compare one or more of the information on the region of interest (e.g., the number of regions of interest, the size of the regions of interest, etc.) of the decoded current frame with the information on the region of interest of the decoded frame used for temporal restoration to determine the number of frames on which the temporal restoration will ultimately be performed. That is, the temporal frame restoration degree determination module (852) may determine the number of intermediate frames to be generated based on the information on the region of interest of the current frame and the target frame determined to be used for temporal restoration. The temporal frame restoration degree determination module (852) may determine the number of intermediate frames to be generated based on at least one of the number of regions of interest, the difference in the size of the regions of interest, the resolution of the restored image, and the frame rate of the restored image.

[0200] In an embodiment, when the number of intermediate frames to be generated according to a temporal restoration cycle is n, and the number of regions of interest between two frames is different, the temporal frame restoration degree determination module (852) may determine the number of frames on which temporal restoration of the corresponding temporal region is to be performed as m (m is an integer greater than or equal to 0 and less than or equal to n).

[0201] In an embodiment, when the number of intermediate frames to be generated according to a temporal restoration period is n, and when the difference in the size of the region of interest between two frames is less than a certain threshold (or greater than the threshold in another embodiment), the number of frames on which temporal restoration of the corresponding temporal region is to be performed can be determined as p (where p is an integer greater than or equal to 0 and less than or equal to n).

[0202] The restoration degree determination module (852) can determine n, m, and p based on at least one of the number of regions of interest, the difference in the sizes of the regions of interest, the resolution of the restored image, and the frame rate of the restored image.

[0203] In one embodiment, the restoration degree determination module (852) may receive from the encoding device (10a) the number of frames to be temporally restored through the decoded current frame when temporal restoration is performed aperiodically. When the number of frames to be restored is received, the restoration degree determination module (852) may determine the degree of restoration in the same manner as when temporal restoration is performed periodically.

[0204] In one embodiment, when temporal restoration is performed aperiodically, the restoration degree determination module (852) may determine the number of frames on which temporal restoration is to be performed implicitly using information of the decoded current frame and the decoded frame to be used for temporal restoration, without receiving the number of frames to be temporally restored from the encoding device (10a). The restoration degree determination module (852) may, according to an embodiment, compare one or more of the information on the region of interest (the number of regions of interest, the size of the region of interest, etc.) of the decoded current frame with the information on the region of interest of the decoded frame to be used for temporal restoration, to determine the number of frames on which temporal restoration is to be performed. According to an embodiment, when the maximum number of frames on which temporal restoration can be performed is t, and the number of regions of interest between two frames is different, the number of frames on which temporal restoration is to be performed may be determined as b (b is an integer greater than or equal to 0 and less than or equal to t). The restoration degree determination module (852) may determine, according to an embodiment, the number of frames on which temporal restoration is to be performed as v (v is an integer greater than or equal to 0 and less than or equal to t), when the maximum number of frames on which temporal restoration can be performed is t, and when the difference in the size of the region of interest between two frames is less than a specific threshold value (or greater than the threshold value in another embodiment). According to an embodiment, the information of t may be determined in units of sequences. The temporal frame restoration degree determination module (852) may determine t, b, and v based on at least one of the number of regions of interest, the difference in the size of the region of interest, the resolution of the restored image, and the frame rate of the restored image.

[0205] In various embodiments, the temporal restoration performer (850) may determine a method for performing temporal restoration. The temporal restoration performer (850) may first determine the target and extent of temporal restoration, and then determine the method for performing temporal restoration. In various embodiments, the order in which the target, extent, and method for temporal restoration are determined may be changed, and the target or extent may be determined by a method selected according to the embodiment.

[0206] In one embodiment, the restoration method determination module (853) may determine frames to be used for temporal restoration among frames in the decoded restored image, and then determine a temporal frame restoration method. According to an embodiment, the restoration method determination module (853) may determine a temporal frame restoration method by comparing one or more metric measurements, such as PSNR and / or SSIM, of two selected frames to be used for temporal restoration.

[0207] According to an embodiment, the restoration performing module (854) may perform temporal frame restoration by copying an adjacent frame among the two frames according to the temporal order of the final restored image, or by using an average value of the two frames or a weighted sum of the two frames, if the difference value of the metric measurement values ​​of the two frames is less than a certain threshold value (or greater than the threshold value in another embodiment). The restoration method determining module (853) may, according to an embodiment, implicitly determine the restoration method according to the degree of the difference value of the metric measurement values, may determine it using a fixed restoration method, or may determine it through parsing the syntax for the restoration method from the bitstream.

[0208] In an embodiment, when the size of the region of interest in the two frames is the same, the restoration method determination module (853) may determine a temporal frame restoration method by comparing one or more metric measurements, such as PSNR and / or SSIM, between the regions of interest in the two frames. In an embodiment, when the difference value of the metric measurements of the region of interest in the two frames is less than a specific threshold value (or greater than the threshold value in another embodiment), the restoration method determination module (853) may perform temporal frame restoration by copying an adjacent frame among the two frames in the temporal order of the final restored image, or by averaging the two frames, or by weighting the two frames.

[0209] According to an embodiment, the frame restoration method may be determined implicitly based on the degree of the difference value of the metric measurement value, may be determined using a fixed method, or may be determined through parsing the syntax for the method from the bitstream.

[0210] According to an embodiment, the restoration method determination module (853) may perform temporal frame restoration using a deep neural network composed of one or more learned convolutional layers if a condition less than a certain threshold value (or, in another embodiment, a condition greater than the threshold value) is not satisfied.

[0211] According to an embodiment, the restoration method determination module (853) may determine a temporal frame restoration method by comparing at least one metric measurement value, such as PSNR and / or SSIM, for the regions of interest within the two frames when the difference in size of the regions of interest within the two frames is less than a certain threshold. At this time, for the metric comparison, the restoration method determination module (853) may perform downsampling on the region of interest with a larger size among the regions of interest within the two frames to have the same size as the region of interest with a smaller size, and then perform the metric comparison, or vice versa, may perform the comparison through upsampling.

[0212] In some embodiments, the restoration method determination module (853) may, in order to compare regions of interest with different sizes between frames, move the upper left positions of the two regions to the same extent, and then perform a comparison on a region as small as the larger region or on a region as large as the smaller region. In this case, in an example where the restoration method determination module (853) performs a comparison on a larger region, the region outside the smaller region may be filled with a median value or padded using a boundary value.

[0213] According to an embodiment, the restoration method determination module (853) may perform temporal restoration by copying only the preceding frame or copying only the following frame, when the number of frames for which temporal restoration is to be performed using the two frames is n, if the difference in the size of the region of interest within the two frames is less than a certain threshold. Or, The second frame copies the frame that is temporally earlier than the other two frames, The second frame can perform temporal restoration by copying the frame that is temporally later among the two frames. Alternatively, the restoration method determination module (853) can create n intermediate frames by copying only the preceding frame or only the subsequent frame.

[0214] According to an embodiment, the restoration method determination module (853) may perform temporal frame restoration by copying an adjacent frame among the two frames in the temporal order of the final restored image, or by taking the average value of the two frames, or by taking the weighted sum of the two frames, if the difference value of the metric measurement value of the region of interest within the two frames is greater than or less than a specific threshold.

[0215] According to an embodiment, the restoration method determination module (853) may determine the frame restoration method implicitly based on the degree of the difference value of the metric measurement value, may determine it using a fixed method, or may determine it through parsing the syntax for the method from the bitstream.

[0216] According to an embodiment, the restoration method determination module (853) may perform temporal frame restoration using a deep neural network composed of one or more learned convolutional layers if a condition less than a certain threshold value (or, in another embodiment, a condition greater than the threshold value) is not satisfied.

[0217] According to an embodiment, the restoration method determination module (853) may not perform temporal frame restoration if the difference in the size of the region of interest within two frames exceeds a certain threshold.

[0218] According to an embodiment, if the difference in the size of the region of interest within two frames exceeds a certain threshold, the restoration method determination module (853) may perform temporal frame restoration by generating n frames by copying adjacent frames among the two frames according to the temporal order of the final restored image.

[0219] According to an embodiment, the restoration method determination module (853) can explicitly determine the restoration method by parsing it from the transmitted bitstream.

[0220] According to an embodiment, the restoration method determination module (853) can apply the same restoration method to the entire video sequence. The restoration method determination module (853) can apply different restoration methods to a series of frame groups within the video sequence, and at this time, the restoration method can be signaled by the encoder for each unit, and the decoder can parse and determine the restoration method. According to an embodiment, the restoration method determination module (853) can use different restoration methods according to conditions after parsing the restoration methods. For example, when the parsed restoration method is a restoration method using a deep neural network composed of one or more learned convolutional layers, the restoration method determination module (853) can use the parsed method or use another method depending on the relationship between the two frames to be restored. In this case, the relationship between the two frames can mean a case where a metric measurement value such as SSIM and / or PSNR between the two frames is smaller than a specific threshold value (or larger in another embodiment), or can be defined as the number of regions of interest between the two frames and the size of the regions.

[0221] According to an embodiment, when the restoration method determination module (853) generates an intermediate frame during a temporal frame restoration process, if two or more frames are used, the weights of each input frame may be different. For example, when generating an intermediate frame using two frames, if one of the two frames is a frame that only contains a region of interest and the other frame is a frame that contains both a region of interest and other regions, the restoration method determination module (853) may set the weight of the frame that contains all regions to be greater.

[0222] In various embodiments, the restoration performing module (854) can perform temporal restoration on the decrypted image according to the previously determined restoration target, restoration degree, and restoration method.

[0223] <Description of Semantics and Syntax Used in Embodiments of the Present Disclosure>

[0224] The semantics below are explained based on a series of frame group units, for example, GOP units.

[0225] 1. Semantics related to temporal restoration

[0226] (1)temporal_restoration_flag: Flag for whether to apply temporal resampling. If 1, temporal resampling is applied.

[0227] (2)Temporal_resampling_ratio_idx: Index of the list for temporal resampling ratios, and since it is indicated by temporal_restoration_flag for applying temporal resampling, it does not include cases where the input signal and the output signal are the same. For example, the sampling rate can be in the form of a ratio such as 1 / 2, 1 / 4, 1 / 8, etc., and can be the number of frames per second. The encoder and the decoder can agree on the list of sampling rates, and the encoder can transmit the index of the list. Depending on the embodiment, if the resampling rate is fixed for the transmitted application or the number of required frames is fixed, the information may be omitted.

[0228] (3)same_period_flag: A flag indicating whether periodic resampling or aperiodic resampling is used when applying temporal resampling. Depending on the embodiment, if the periodicity of temporal resampling is fixed, this information and related information may be omitted. If same_period_flag is 1, it means that temporal resampling is performed with the same period.

[0229] (4)delta_frame_idx[i]: Transmits the difference value from the previous frame index based on the frame number. Since it is the difference value from the previous index, it is sufficient to transmit a number that is one less than num_of_frames calculated by the frame rate. i represents the frame index.

[0230] (5)num_of_GOPs: Information indicating the number of GOPs in a sequence

[0231] (6) Restoration_method_idx: Index information for specifying the temporal restoration method. The temporal restoration method can be one of a learning-based network method, a non-learning-based interpolation method, a copy method, etc., and information about this can be signaled. The information can be transmitted in sequence units or in units of a series of frame groups. Alternatively, the decoder can explicitly select and use the temporal restoration method without signaling / parsing the information.

[0232] 2. Semantics related to images processed based on region of interest

[0233] (1)RoI_Processing_flag: 1-bit flag indicating whether the current sequence has undergone region-of-interest-based processing.

[0234] (2)num_of_GOPs: Information indicating the number of GOPs in a sequence

[0235] (3)num_of_frames: Information indicating the number of frames in one GOP, depending on the embodiment.

[0236] (4)num_of_RoIs: Information indicating the number of regions of interest in one frame

[0237] (5)upsamp_ratio_RoIs: Upsampling rate information for the region of interest, and a processed image based on the region of interest restored with the information can be obtained.

[0238] (6)upsamp_ratio_nonRoIs: Upsampling rate information for non-interest regions, and the restored non-interest region image can be obtained using the information.

[0239] (7)RoI_exist_region_LT(0), RoI_exist_region_LT(1): The upper left coordinate of the region where the region of interest exists in the picture.

[0240] (8)RoI_exist_region_RB(0), RoI_exist_region_RB(1): Coordinates of the lower right corner of the region where the region of interest exists within the picture

[0241] (9)frame_RoI_Information_flag[p]: Frame information for which region of interest information encoding is performed. If 1, encoding of region of interest information within the frame is performed. If 0, the frame is skipped. At this time, encoding of region of interest information between the two most adjacent frames with the flag set to 1 is performed.

[0242] (10)only_specific_RoI_flag and specific_RoI_flag[j]: When only_specific_RoI_flag is 1, information of all regions of interest within the frame where region of interest information encoding is performed is encoded. When only_specific_RoI_flag is 0, specific_RoI_flag[j] is parsed and information encoding is performed only for regions of interest for which the flag is 1. j represents the region of interest index.

[0243] 3. Semantics related to the region of interest (pos_diff_coding)

[0244] Hereinafter, n represents the GOP index, p represents the frame index, and m represents the region of interest index.

[0245] (1) pos_RoI(n)(p)(m)[0], pos_RoI(n)(p)(m)[1]: Location information of regions of interest existing in the first frame of each series of frame groups (GOPs).

[0246] (1-1)RoI_exist_region_LT(0), RoI_exist_region_LT(1), if a non-zero value is parsed, as an example, RoI_exist_region_LT(0) + pos_RoI(n)(m)[0], RoI_exist_region_LT(1) + pos_RoI(n)(m)[1] may be the final restored position of the m-th region of interest.

[0247] (1-2) For the position difference values ​​obtained through pos_diff_coding, the final restored position of the region of interest can be obtained in the same way.

[0248] (2)abs_pos_diff_greater0_flag_RoI(n)(p)(m)[0], abs_pos_diff_greater0_flag_RoI(n)(p)(m)[1]: 1-bit flag indicating whether the absolute value of the x-axis and y-axis position information difference value for the same region of interest in the 0th frame is 0.

[0249] (3)abs_pos_diff_greater1_flag_RoI(n)(p)(m)[0], abs_pos_diff_greater1_flag_RoI(n)(p)(m)[1]: 1-bit flag indicating whether the absolute value of the x-axis and y-axis position information difference value for the same region of interest in the 0th frame is 1.

[0250] (4)abs_pos_diff_minus2_RoI(i)(p)(j)[0], abs_pos_diff_minus2_RoI(i)(p)(j)[0]: Information indicating the corresponding value of -2 of the absolute value of the difference value of the x-axis and y-axis position information for the same region of interest in the 0th frame.

[0251] (5)pos_diff_sign_flag(i)(p)(j)[0], pos_diff_sign_flag(i)(p)(j)[1]: 1-bit flag indicating the x-axis and y-axis sign values ​​for the same region of interest in the 0th frame, 1 for + sign, 0 for - sign

[0252]

[0253] An example of the syntax of the above temporal restoration data is as shown in Table 1 below.

[0254] Descriptortemporal_restoration_data( ){temporal_restoration_flague(1)if( temporal_restoration_flag ) {temporal_resampling_ratio_idxue(v)same_period_flague(1)if( !same_period_flag ) {for( i=0; i <num_of_frames; i++ ) {delta_frame_idx[i]ue(v)}}for ( j=0; j<num_of_GOPs; j++ ) {Restoration_method_idxue(v)}}}Descriptortemporal_restoration_data( ) {temporal_restoration_flague(1)if( temporal_restoration_flag ) {temporal_resampling_ratio_idxue(v)same_period_flague(1)if( !same_period_flag ) {for( i=0; i<num_of_frames; i++ ) {delta_frame_idx[i]ue(v)}}Restoration_method_idx}}

[0255] An example of the syntax of the above area of ​​interest information is shown in Table 2 below.

[0256] DescriptorRoI_Processing_flagu(1)if( RoI_Processing_flag ) {RoI_exist_region_LT(0)ue(v)RoI_exist_region_LT(1)ue(v)RoI_exist_region_RB(0)ue(v)RoI_exist_region_RB(1)ue(v)for ( i=0; i<num_of_GOPs; i++ ) {upsamp_ratio_RoIsue(v)upsamp_ratio_nonRoIsue(v)num_of_framesue(v)num_of_RoIsue(v)only_specific_RoI_flague(1)for ( p=0; p<num_of_frames; p++ ) {frame_RoI_Information_flag[p]ue(1)}if ( only_specific_RoI_flag ) {for ( p=0; p<num_of_frames; p++ ) {for ( j=0; j<num_of_RoIs; j++ ) {if ( frame_RoI_Information_flag[p] ) {specific_RoI_flag[j]ue(1)}}}}int a = 0;for ( p=0; p<num_of_frames; p++ ) {if ( frame_RoI_Information_flag[p]) {if ( !only_specific_RoI_flag ) {for ( j=0; j<num_of_RoIs; j++ ) {if ( a == 0 ) {pos_RoI(i)(p)(j)[0]ue(v)pos_RoI(i)(p)(j)[1]ue(v)a++;} else {pos_diff_coding}}} else {for ( j=0; j<num_of_RoIs; j++ ) {if ( specific_RoI_flag[i] ) {if ( a == 0 ) {pos_RoI(i)(p)(j)[0]ue(v)pos_RoI(i)(p)(j)[1]ue(v)a++;} else {pos_diff_coding}}}}}}}

[0257] Table 3 below is an example syntax of restoration information for a region of interest according to various embodiments.

[0258] Descriptorpos_diff_coding( ){abs_pos_diff_greater0_flag_RoI(i)(p)(j)[0]ue(1)abs_pos_diff_greater0_flag_RoI(i)(p)(j)[1]ue(1)if( abs_pos_diff_greater0_flag_RoI(i)(p)(j)[0] )abs_pos_diff_greater1_flag_RoI(i)(p)(j)[0]ue(1)if( abs_pos_diff_greater0_flag_RoI(i)(p)(j)[1] )abs_pos_diff_greater1_flag_RoI(i)(p)(j)[1]ue(1)if( abs_pos_diff_greater0_flag_RoI(i)(p)(j)[0] ) {if( abs_pos_diff_greater1_flag_RoI(i)(p)(j)[0] )abs_pos_diff_minus2_RoI(i)(p)(j)[0]ue(v)pos_diff_sign_flag(i)(p)(j)[0]ue(1)}if( abs_pos_diff_greater0_flag_RoI(i)(p)(j)[1] ) {if( abs_pos_diff_greater1_flag_RoI(i)(p)(j)[1] )abs_pos_diff_minus2_RoI(i)(p)(j)[1]ue(v)pos_diff_sign_flag(i)(p)(j)[1]ue(1)}}

[0259] The examples of the present disclosure presented in this specification and drawings are intended solely to facilitate the technical content of the present disclosure and to aid understanding thereof, and are not intended to limit the scope of the present disclosure. It will be apparent to those skilled in the art that other variations are possible in addition to the examples described above.

[0260] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.

[0261] [Explanation of symbols]

[0262] 10a: Encoding device

[0263] 510: Temporal Resampling Performer

[0264] 520: Spatial Resampling Performer

[0265] 530: Region-of-interest-based processor

[0266] 540: Bit Cutting Machine

[0267] 550: Internal Encoding Performer

[0268] 10b: Decoding device

[0269] 810: Internal decryption performer

[0270] 820: Bit Compensation Performer

[0271] 830: Region-of-interest-based restorer

[0272] 840: Spatial Restoration Performer

[0273] 850: Temporal Restoration Performer

[0274] 860: Post-processing filter performer

Claims

1. In a VCM (video coding for machines) decoding device, An internal decoding unit that decodes a bitstream to generate a restored image including restored regions of no interest and regions of interest, and extracts information about the regions of interest for the restored image; A region-of-interest-based restorer that restores a region-of-interest-based processed image from the restored region-of-interest based on information about the region-of-interest; and A temporal restoration performer for performing temporal restoration on the restored image based on information about the region of interest, A decoding device, wherein the information on the region of interest includes at least one of: whether region-of-interest-based processing has been performed, the number of GOPs (groups of pictures) in a sequence, the number of frames in a GOP, the number of regions of interest in a frame, the size of each region of interest, the position of each region of interest, the movement of each region of interest, the mapping relationship between regions of interest, and temporal resampling information for the region of interest.

2. In paragraph 1, The above temporal restoration performer is a decoding device that performs temporal restoration by generating one or more intermediate frames between two frames included in the restored image.

3. In paragraph 2, The temporal restoration performer generates one or more intermediate frames by copying one of the two frames, A decoding device that performs interpolation on one or more intermediate frames using the two frames based on information about the region of interest.

4. In paragraph 2, The above temporal restoration performer is a decoding device that selects a frame to be used for temporal restoration based on information about a region of interest included in each frame among all frames included in the restored image.

5. In paragraph 4, A decoding device in which the temporal restoration performer performs extrapolation on one or more intermediate frames generated by using a frame temporally earlier than the two frames among the frames selected as frames to be used for the temporal restoration.

6. In paragraph 4, A decoding device, wherein the temporal restoration performer selects a frame to be used for temporal restoration by using at least one of the number of regions of interest, the size of each region of interest, the difference in the sizes of the regions of interest, the location, and the degree of similarity of the regions of interest included in the information about the regions of interest.

7. In paragraph 6, A decoding device, wherein the degree of similarity of the region of interest is determined by comparing at least one of the peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) metric measurements for the region of interest included in the current frame and the temporally subsequent frame among the entire frames.

8. In paragraph 1, The above temporal restoration performer is a decoding device that, when generating an intermediate frame using two or more frames in the restored image, generates the intermediate frame by applying different weights to the two or more frames.

9. In paragraph 1, A decoding device, wherein the temporal restoration performer determines the degree of temporal restoration for the restored image based on whether the temporal restoration is periodic.

10. In paragraph 1, A decoding device, wherein the temporal restoration performer determines at least one of a temporal restoration target, a temporal restoration period, a temporal restoration degree, and a temporal restoration method based on information about temporal restoration extracted from the bitstream.

11. In paragraph 1, A decoding device, wherein the temporal restoration method is at least one of a method using a learning-based neural network model, a non-learning-based interpolation method, and a copying method.

12. In paragraph 1, Further comprising a spatial restoration performer for restoring the resolution of the above restored image, A decoding device, wherein the spatial restoration performer restores the resolution of the restored image based on information about the region of interest.

13. In a VCM (video coding for machines) encoding device, A temporal resampling performer that changes the frame rate of an input image; A region-of-interest-based processor that extracts one or more regions of interest included in each frame of the input image and generates a processed image based on the one or more regions of interest; and It includes an internal encoding performer that generates a bitstream by encoding the processed image based on the region of interest, the image of the non-region of interest, and information about the one or more regions of interest, An encoding device, wherein the information on the region of interest includes at least one of: whether region-of-interest-based processing has been performed, the number of GOPs (groups of pictures) in a sequence, the number of frames in a GOP, the number of regions of interest in a frame, the size of each region of interest, the position of each region of interest, the movement of each region of interest, the mapping relationship between regions of interest, and temporal resampling information for the region of interest.

14. In paragraph 13, An encoding device in which the region-of-interest-based processor performs packing on regions of interest included in the input image, and fills regions of the input image excluding the region of interest with a specific value.

15. In paragraph 13, An encoding device in which the temporal resampling performer samples the region-of-interest-based image and the non-region-of-interest image at different frame rates.

16. In paragraph 13, An encoding device, wherein the temporal resampling information for the region of interest includes at least one of an index of a list for temporal resampling rates and whether temporal resampling is applied.

17. In paragraph 13, An encoding device in which information about temporal resampling is omitted when the periodicity of the above temporal resampling is fixed.

18. In paragraph 13, The temporal resampling information for the above region of interest includes a temporal restoration method for the frame on which temporal resampling was performed, An encoding device, wherein the temporal restoration method is at least one of a method using a learning-based neural network model, a non-learning-based interpolation method, and a copying method.

19. In paragraph 13, Further comprising a spatial resampling performer that changes the spatial resolution of each frame or a series of frames of the input image, An encoding device that changes the resolution of each frame of the input image based on information about the region of interest.

20. A non-volatile computer-readable storage medium that records commands, The above instructions, when executed by one or more processors, cause the one or more processors to: A step of decoding a bitstream to generate a restored image including a restored region of no interest and a region of interest, and extracting information on the region of interest for the restored image; A step of restoring a region-of-interest-based processed image from the restored region of interest based on information about the region of interest; and A non-transitory computer-readable storage medium comprising a step of performing temporal restoration on the restored image based on information about the region of interest.

Citation Information

Patent Citations

  • Region of interest based screen contents quality improving video encoding / decoding method and apparatus thereof

    KR1020130078569A

  • Machine learning algorithm using compression parameter for image reconstruction and image reconstruction method therewith

    KR102053242B1

  • System for compressing and restoring picture based on AI

    KR102190483B1

  • Compression apparatus of automatic supply terminal

    KR102542670B1