Coding and decoding method based on adaptive space-time resampling in VCM
By using adaptive spatiotemporal resampling technology, the server load and power consumption problems caused by machine analysis of high-resolution and high-frame-rate images are solved, the encoding and decoding efficiency is improved, and the effect of machine vision applications is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANWHA VISION CO LTD
- Filing Date
- 2024-09-20
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to effectively address the server load and power consumption issues caused by machine-analyzed images, especially when processing high-resolution and high-frame-rate images.
An adaptive spatiotemporal resampling technique is adopted, which changes the frame rate of the image through a temporal resampler and changes the resolution of the image through a spatial resampler. Combined with an internal encoder and decoder, an adaptive encoding and decoding method is achieved.
It improves the encoding and decoding efficiency of image analysis, enhances the effectiveness of machine vision applications, and reduces server load and power consumption.
Smart Images

Figure CN121890085A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to encoding and decoding methods based on adaptive spatiotemporal resampling in VCM. Background Technology
[0002] With the continuous development of the information and communication industry, broadcasting services with HD (high definition) resolution have spread globally.
[0003] This spread has made many users familiar with high-resolution and high-definition images and / or videos, and has increased the demand for high-resolution and high-quality images and videos, such as UHD (Ultra-High Definition) images / videos with higher definition, i.e., 4K or 8K or even higher definition, in various fields.
[0004] The technology used to encode this UHD image data was completed in 2013 using the standard technology HEVC (High Efficiency Video Coding).
[0005] HEVC is a next-generation image compression technology with higher compression ratio and lower complexity than the previous H.264 / AVC technology, and it is a core technology for effectively compressing large amounts of data in HD and UHD images.
[0006] Like previous compression standards, HEVC performs encoding on a block-by-block basis. However, unlike H.264 / AVC, it has only one configuration file. The core coding techniques included in HEVC's single configuration file encompass eight areas: layered coding structure, transform techniques, quantization techniques, intra-frame predictive coding, inter-frame motion prediction, entropy coding, loop filtering, and others.
[0007] Since the establishment of the HEVC video codec in 2013, with the expansion of services such as realistic images using 4K and 8K video images and virtual reality, a new standard, VVC (Variety Video Coding), has been developed as the next-generation video codec, aiming to improve performance by at least twice that of HEVC. VVC is also known as H.266.
[0008] H.266 (VVC) was developed with the goal of achieving twice the efficiency of the previous generation codec H.265 (HEVC). Initially developed with 4K resolution or higher in mind, VVC has been expanded to handle massive 16K ultra-high resolution images due to the growing VR market and its responsiveness to 360-degree imaging. Furthermore, as the HDR market has expanded along with display technology advancements, VVC supports not only 10-bit color depth but also 16-bit color depth to accommodate the HDR market, and supports brightness levels of 1000 nits, 4000 nits, and 10000 nits. Additionally, because VVC is being developed with the VR and 360-degree imaging markets in mind, it supports partial frame rates from 0 FPS to 120 FPS.
[0009] The Development of Artificial Intelligence
[0010] Artificial intelligence (AI) is also gradually developing. AI refers to the artificial imitation of human intelligence, that is, the intelligence that can recognize, classify, reason, predict, control / make decisions, etc.
[0011] With the development of artificial intelligence technology and the increase in Internet of Things (IoT) devices, machine-to-machine traffic is expected to increase dramatically, and machine-based image analytics is expected to be widely used.
[0012] However, server load and power consumption issues are expected as the number of images to be analyzed by machines is anticipated to increase exponentially. Summary of the Invention
[0013] Technical issues
[0014] Therefore, this disclosure provides an encoding and decoding method based on adaptive spatiotemporal resampling in VCM to enable machine-efficient image analysis.
[0015] One embodiment of this disclosure may provide an encoding and decoding method for adaptively changing the frame rate and resolution according to the application purpose.
[0016] Technical solutions
[0017] To achieve the objectives described above, according to a disclosure of this specification, an encoding apparatus for performing machine-oriented image encoding is provided.
[0018] An encoding apparatus for performing machine-oriented image encoding includes: a temporal resampler for changing the frame rate of an input image; a spatial resampler for changing the resolution of the input image; and an internal encoder for encoding the input image, and capable of transmitting temporally restored data including the frame rate of the input image changed by the temporal resampler and spatially restored data including the resolution of the input image changed by the spatial resampler, together with the encoded image, to a decoding apparatus.
[0019] According to one disclosure of this specification, a decoding apparatus for performing machine-oriented image decoding may include: an internal decoder for receiving an encoded input image as input from an encoding apparatus and performing decoding; a spatial restorer for changing the resolution of the decoded image based on spatial restoration data of the input image received from the encoding apparatus and restoring it to the resolution of the original image of the input image; and a temporal restorer for changing the frame rate of the decoded image based on temporal restoration data of the input image received from the encoding apparatus and restoring it to the frame rate of the original image of the input image.
[0020] According to one disclosure of this specification, as a non-volatile computer-readable storage medium for recording commands, the commands, when executed by at least one processor, can cause at least one processor to: change the frame rate of an input image; change the resolution of the input image; encode the input image; and transmit time-recovered data including the frame rate of the input image changed according to time resampling and spatial-recovered data including the resolution of the input image changed according to spatial resampling, together with the encoded image, to a decoding device.
[0021] Technical effect
[0022] Based on this disclosure, machine-based image analysis can be performed efficiently.
[0023] According to embodiments of this disclosure, the frame rate and resolution of the original image can be adaptively changed according to the application purpose, thereby not only improving encoding and decoding efficiency, but also enhancing the effectiveness of machine vision applications. Attached Figure Description
[0024] Figure 1 An example illustrating a video / image coding system.
[0025] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding device.
[0026] Figure 3 This is a diagram schematically illustrating the configuration of a video / image decoding device.
[0027] Figures 4a to 4d This is an exemplary diagram representing a VCM encoder and a VCM decoder.
[0028] Figure 5 It is a block diagram including components of an encoding device for machine vision according to embodiments of the present disclosure.
[0029] Figure 6 This is an example of changing the resolution by temporal resampling according to the embodiments of this disclosure.
[0030] Figure 7a , Figure 7b and Figure 7c These are examples of the syntax and semantics for temporal resampling according to embodiments of this disclosure.
[0031] Figure 8 This is an example of changing the resolution through spatial resampling in an embodiment of this disclosure.
[0032] Figure 9 This is an example of reordering frames by resolution after spatial resampling, according to an embodiment of this disclosure.
[0033] Figure 10 This is another example of changing the resolution through spatial resampling in an embodiment of this disclosure.
[0034] Figure 11a , Figure 11b and Figure 11c These are examples of syntax and semantics for spatial resampling based on embodiments of this disclosure.
[0035] Figure 12 It is a block diagram including components of an encoding device for machine vision according to embodiments of the present disclosure.
[0036] Figure 13 It is a block diagram including components of a decoding apparatus for machine vision according to an embodiment of the present disclosure.
[0037] Figure 14 This is a flowchart used to describe a time resampling method based on an example of this disclosure.
[0038] Figure 15 This is a flowchart used to describe another example of a time resampling method based on this disclosure.
[0039] Figure 16a and Figure 16b These are examples of the syntax and semantics for temporal resampling according to embodiments of this disclosure. Detailed Implementation
[0040] The specific structural or phased descriptions of embodiments of the concept of this disclosure disclosed in this specification or application are shown only for the purpose of describing embodiments of the concept of this disclosure, and embodiments of the concept of this disclosure may be implemented in various forms and should not be construed as limited to the embodiments described in this specification or application.
[0041] Because various modifications and forms can be made to the embodiments based on the concept of this disclosure, specific embodiments are shown in the accompanying drawings and described in detail in this specification or application. However, it is not intended to limit the embodiments based on the concept of this disclosure to the specific forms disclosed, and it should be understood that they include all changes, equivalents, or substitutions included within the spirit and scope of this disclosure.
[0042] Various components may be described using terms such as first, second, etc., but components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from other components, and, for example, without departing from the scope of the claims based on the concept of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component.
[0043] When a component is referred to as "linked" or "connected" to other components, it should be understood that the component can be directly linked or connected to other components, but other components may exist in between. On the other hand, when a component is referred to as "directly linked" or "directly connected" to other components, it should be understood that no other components exist in between. Other expressions describing the relationship between components, namely "between" and "directly between," or "adjacent to" and "directly adjacent to," should also be interpreted similarly.
[0044] Since the terminology used in this specification is for describing particular embodiments only, it is not intended to limit the scope of this disclosure. Singular expressions include plural expressions unless the singular expression clearly has a different meaning depending on the context. In this specification, it should be understood that terms such as "comprising" or "having" refer to the presence of the described features, numbers, steps, movements, components, portions, or combinations thereof, but do not preclude the presence or possibility of one or more other features, numbers, steps, movements, components, portions, or combinations thereof being present or added in advance.
[0045] Unless otherwise defined, all terms used herein, including technical or scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Terms defined in common dictionaries shall be interpreted as having the same meaning as in the context of the relevant art, and shall not be interpreted in an ideal or overly formalized sense unless expressly defined in this specification.
[0046] In describing the embodiments, technical details that are well-known in the art to which this disclosure pertains and are not directly related to this disclosure have been omitted. This is to more clearly convey the key points of this disclosure without obscuring it by omitting unnecessary descriptions.
[0047] This document relates to video / image coding. For example, the methods / implementations disclosed in this document may relate to the Multifunctional Video Coding (VVC) standard (ITU-T Rec.H.266), next-generation video / image coding standards that comply with VVC, or other video coding-related standards (e.g., High Efficiency Video Coding (HEVC) standard (ITU-T Rec.H.265), Basic Video Coding (EVC) standard, AVS2 standard, etc.).
[0048] This document presents various implementations related to video / image coding, and unless otherwise stated, these implementations can be combined with each other.
[0049] In this document, video can refer to a set of images over time. An image generally refers to a unit representing an image at a specific time interval, and a slice / tile is a unit configured as a portion of an image in coding. A slice / tile may include one or more coding tree units (CTUs). An image may consist of one or more slices / tiles. An image may consist of one or more groups of tiles. A group of tiles may include one or more tiles.
[0050] A pixel, or pel, can refer to the smallest unit that configures a picture (or image). Alternatively, a "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or pixel value, or it can represent only the pixel / pixel value of the luminance component, or it can represent only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when such a pixel value is transformed to the frequency domain, it can refer to the transform coefficients in the frequency domain.
[0051] A unit can represent a basic unit used in image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." Typically, A block may include a sample (or sample array) consisting of M columns and N rows or a set (or array) of transform coefficients.
[0052] Figure 1 An example illustrating a video / image coding system.
[0053] Reference Figure 1 A video / image encoding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0054] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may include a separate device or an external component.
[0055] Video sources can acquire video / images through processes such as video / image capture, compositing, and generation. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more camera devices, video / image files including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, smartphones, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated by computers, etc., and in this case, video / image capture processing can be replaced by processing that generates related data.
[0056] An encoding device can encode input video / images. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0057] A transmitter can transmit encoded video / image information or data, output in bitstream form, to a receiver via a digital storage medium or network in the form of a file or stream. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.
[0058] Decoding devices can decode video / images by performing a series of processes corresponding to the operations of encoding devices, such as dequantization, inverse transform, and prediction.
[0059] The renderer can render decoded video / images. The rendered video / images can then be displayed using a display unit.
[0060] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding device.
[0061] In the following text, a video encoding device may include an image encoding device.
[0062] Reference Figure 2 The encoding device 10a can be configured to include an image segmenter 10a-10, a predictor 10a-20, a residual processor 10a-30, an entropy encoder 10a-40, an adder 10a-50, a filter 10a-60, and a memory 10a-70. The predictor 10a-20 may include an inter-frame predictor 10a-21 and an intra-frame predictor 10a-22. The residual processor 10a-30 may include a transformer 10a-32, a quantizer 10a-33, a dequantizer 10a-34, and an inverse transformer 10a-35. The residual processor 10a-30 may also include a subtractor 10a-31. The adder 10a-50 may be referred to as a reconstructor or a reconstructed block generator. According to the implementation, the image segmenter 10a-10, predictor 10a-20, residual processor 10a-30, entropy encoder 10a-40, adder 10a-50, and filter 10a-60 described above can be configured by one or more hardware components (e.g., encoder chipsets or processors). Furthermore, the memory 10a-70 may include a decoded image buffer (DPB) and can be configured by a digital storage medium. The hardware components may also include the memory 10a-70 as an internal / external component.
[0063] Image segmenters 10a-10 can split an input image (or picture, frame) input to encoding device 10a into one or more processing units. As an example, a processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively split from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be split into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this document can be performed based on the final coding unit that is no longer split. In this case, based on encoding efficiency and other factors, and according to image characteristics, the maximum coding unit can be directly used as the final coding unit, or, if necessary, the coding unit can be recursively split into deeper coding units, and the optimally sized coding unit can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, as described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can be split or divided from the final encoding unit described above. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for obtaining transform coefficients and / or a unit for obtaining the residual signal from the transform coefficients.
[0064] In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." Generally, A block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, or they can represent only pixel / pixel values of the luminance component, or only pixel / pixel values of the chrominance component. Samples can be used as a term corresponding to pixels or pels in a picture (or image).
[0065] Subtractor 10a-31 generates a residual signal (residual block, residual sample, or residual sample array) by subtracting the prediction signal (prediction block, prediction sample, or prediction sample array) output from predictor 10a-20 from the input image signal (original block, original sample, or original sample array), and the generated residual signal is transmitted to converter 10a-32. Predictor 10a-20 performs prediction on the block to be processed (hereinafter referred to as the current block) and generates a prediction block including prediction samples for the current block. Predictor 10a-20 can determine whether to apply intra-frame prediction or inter-frame prediction on the basis of the current block or CU. As described below in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and transmit it to entropy encoder 10a-40. The prediction-related information can be encoded by entropy encoder 10a-40 and output as a bitstream.
[0066] Intra-predictor 10a-22 can predict the current block by referring to samples within the current image. Depending on the prediction mode, the reference sample can be located near or far from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the granularity of the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 10a-22 can also determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0067] Inter-frame predictors 10a-21 can obtain the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially adjacent blocks existing in the current image and temporally adjacent blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally adjacent block may be the same or different. The temporally adjacent block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally adjacent block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictors 10a-21 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to obtain the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip and merge modes, the inter-frame predictor 10a-21 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled to indicate the motion vector of the current block.
[0068] Predictors 10a-20 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-frame block copying (IBC) to predict blocks. Intra-frame block copying can be used, for example, for content image / video coding including games, such as Screen Content Coding (SCC). IBC essentially performs prediction within the current frame, but can be performed in a similar manner to inter-frame prediction because it obtains a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0069] The predicted signals generated by the inter-frame predictors 10a-21 and / or the intra-frame predictors 10a-22 can be used to generate the reconstructed signal or the residual signal. The transformer 10a-32 can apply transformation techniques to the residual signal to generate transform coefficients. For example, transform techniques may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Graph-Based Transform (GBT), Conditional Nonlinear Transform (CNT), etc. Here, GBT refers to the transform obtained from a graph when the relationship information between pixels is represented in the graph. CNT refers to the transform obtained by generating the predicted signal using all previously reconstructed pixels and based on it. Furthermore, the transform process can be applied to square pixel blocks of the same size or to non-square blocks of variable size.
[0070] The quantizer 10a-33 quantizes the transform coefficients and transmits them to the entropy encoder 10a-40, which encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 10a-33 can reorder the block-shaped quantized transform coefficients into a one-dimensional vector shape based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the one-dimensional vector shape. The entropy encoder 10a-40 can perform various encoding methods, such as Exponential Columbus, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC).
[0071] The entropy encoder 10a-40 can also encode information required for video / image reconstruction, excluding quantized transform coefficients (e.g., values of syntax elements, etc.), either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted as a bitstream or stored in units of the network abstraction layer. The video / image information can also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. Additionally, the video / image information can also include general constraint information. Information and / or syntax elements signaled / transmitted later in this document can be encoded and included in the bitstream using the encoding process described above. The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network can include broadcast networks and / or communication networks, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting signals output from the entropy encoder 10a-40 and / or a storage device (not shown) for storing signals output from the entropy encoder 10a-40 may be configured as an internal / external element of the encoding device 10a, or the transmitter may be included in the entropy encoder 10a-40.
[0072] The quantized transform coefficients output from quantizer 10a-33 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 10a-34 and inverse transform 10a-35. Adder 10a-50 can generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from predictor 10a-20. When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstructed block. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and can also be used for inter-frame prediction of the next image by filtering, as described below.
[0073] Simultaneously, during image encoding and / or reconstruction processing, luminance mapping with chroma scaling (LMCS) can be applied.
[0074] Filters 10a-60 can apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, filters 10a-60 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be stored in memory 10a-70, specifically in the DPB of memory 10a-70. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filtering, bilateral filtering, etc. Filters 10a-60 can generate various filter-related information and transmit it to entropy encoder 10a-90, as described later in the description of each filtering method. The filter-related information can be encoded by entropy encoder 10a-90 and output as a bitstream.
[0075] The modified reconstructed images transferred to the memories 10a-70 can be used as reference images in the inter-frame predictors 10a-80. When inter-frame prediction is applied in this way, the encoding device can avoid prediction mismatch between the encoding device 10a and the decoding device, and can also improve encoding efficiency.
[0076] The DPB of memory 10a-70 can store a modified reconstructed image that will be used as a reference image in inter-frame predictor 10a-21. Memory 10a-70 can store motion information of blocks from which motion information is derived (or encoded) within the current image, and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 10a-21 as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 10a-70 can store reconstructed samples of reconstructed blocks within the current image and transmit them to intra-frame predictor 10a-22.
[0077] Figure 3 This is a diagram schematically illustrating the configuration of a video / image decoding device.
[0078] Reference Figure 3The decoding device 10b can be configured including an entropy decoder 10b-10, a residual processor 10b-20, a predictor 10b-30, an adder 10b-40, a filter 10b-50, and a memory 10b-60. The predictor 10b-30 may include an inter-frame predictor 10b-31 and an intra-frame predictor 10b-32. The residual processor 10b-20 may include a dequantizer 10b-21 and an inverse transformer 10b-21. According to an embodiment, the entropy decoder 10b-10, residual processor 10b-20, predictor 10b-30, adder 10b-40, and filter 10b-50 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Furthermore, the memory 10b-60 may include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component may also include the memory 10b-60 as an internal / external component.
[0079] When a bitstream including video / image information is input, the decoding device 10b can respond to the video / image information in it. Figure 2 The image is reconstructed through processing within the encoding apparatus. For example, the decoding apparatus 10b can obtain units / blocks based on block splitting information obtained from the bitstream. The decoding apparatus 10b can perform decoding using processing units applied in the encoding apparatus. Therefore, the processing unit used for decoding can be, for example, an encoding unit, and the encoding unit can be split from encoding tree units or maximally encoded units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be obtained from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding apparatus 10b can be played back by a playback device.
[0080] Decoding device 10b can receive data in bitstream form from... Figure 2 The signal output by the encoding device, and the received signal, can be decoded by the entropy decoder 10b-10. For example, the entropy decoder 10b-10 can parse the bitstream to obtain the information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. In addition, the video / image information may also include general constraint information.
[0081] The decoding device can additionally decode the image based on information about the parameter set and / or general constraint information. Information and / or syntax elements notified / received by signals, as described later in this document, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 10b-10 can decode the information within the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements required for image reconstruction, as well as the values of the quantized transform coefficients with respect to the residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element from the bitstream, determine a context model by using information about the decoded target syntax element and decoding information of adjacent target blocks and decoded target blocks, or information about symbols / bins decoded in previous steps, and generate symbols corresponding to the values of each syntax element by performing arithmetic decoding of bins to predict the occurrence probability of bins based on the determined context model. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction information from the information decoded by the entropy decoder 10b-10 can be provided to the predictor 10b-30, and the residual information, i.e. the quantized transform coefficients and related parameter information, obtained by the entropy decoder 10b-10 through entropy decoding, can be input to the dequantizer 10b-21.
[0082] Additionally, filtering information from the information decoded by the entropy decoder 10b-10 can be provided to the filter 10b-50. Meanwhile, the receiver (not shown) receiving the signal output from the encoding device can be configured as an internal / external element of the decoding device 10b, or the receiver can be a component of the entropy decoder 10b-10. Furthermore, the decoding device according to this document can be referred to as a video / video / image decoding device, and this decoding device can be divided into an information decoder (video / video / image information decoder) and a sample decoder (video / video / image sample decoder). The information decoder can include the entropy decoder 10b-10, and the sample decoder can include at least one of a dequantizer 10b-21, an inverse transformer 10b-22, a predictor 10b-30, an adder 10b-40, a filter 10b-50, and a memory 10b-60.
[0083] Dequantizer 10b-21 can dequantize quantized transform coefficients to output transform coefficients. Dequantizer 10b-21 can reorder the quantized transform coefficients in the form of two-dimensional blocks. In this case, the reordering can be performed based on the coefficient scan order executed by the encoding device. Dequantizer 10b-21 can dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0084] The inverse transformer 10b-22 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0085] The predictor can perform prediction on the current block and generate a prediction block for the current block, including the prediction samples. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on prediction-related information output from the entropy decoder 10b-10, and can determine a specific intra-frame prediction mode / inter-frame prediction mode.
[0086] The predictor can generate a predicted signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-frame block copying (IBC) to predict blocks. Intra-frame block copying can be used, for example, for content image / video coding including games, such as Screen Content Coding (SCC). IBC essentially performs prediction within the current frame, but can be performed in a similar manner to inter-frame prediction because it obtains a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0087] The intra-frame predictor 10b-32 can predict the current block by referring to samples within the current image. Depending on the prediction mode, the reference sample can be located near or far from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 10b-32 can also determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0088] The inter-frame predictor 10b-31 can obtain the predicted block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.).
[0089] For inter-frame prediction, neighboring blocks can include spatially adjacent blocks existing in the current image and temporally adjacent blocks existing in the reference image. For example, inter-frame predictor 10b-31 can construct a motion information candidate list based on neighboring blocks and obtain the motion vector of the current block and / or the reference image index based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction-related information can include information indicating the inter-frame prediction mode of the current block.
[0090] Adder 10b-40 generates the reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from predictor 10b-30. When there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block.
[0091] Adder 10b-40 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and can also be output by filtering or used for inter-frame prediction of the next image, as described below.
[0092] Meanwhile, luminance mapping with chroma scaling (LMCS) can be applied during image decoding processing.
[0093] Filters 10b-50 can apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, filters 10b-50 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be transferred to memory 60, specifically to the DPB of memory 10b-60. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0094] The (modified) reconstructed image stored in the DPB of memory 10b-60 can be used as a reference image in inter-frame predictor 10b-31. Memory 10b-60 can store motion information of blocks from which motion information is derived (or decoded) within the current image, and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 10b-31 as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 10b-60 can store reconstructed samples of reconstructed blocks within the current image and transmit them to intra-frame predictor 10b-32.
[0095] In this specification, the embodiments described in the predictor 10b-30, dequantizer 10b-21, inverse transformer 10b-22, and filter 10b-50 of the decoding device 10b can also be applied to the predictor 10a-20, dequantizer 10a-34, inverse transformer 10a-35, and filter 10a-60 of the encoding device 10a in the same or corresponding manner, respectively.
[0096] As described above, prediction is performed during video encoding to improve compression efficiency. Thus, a prediction block including prediction samples for a current block that is a target block for encoding can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is obtained in the same manner by the encoding device and the decoding device, and the encoding device can improve image encoding efficiency by signaling information about the residual between the original block and the prediction block (residual information) to the decoding device instead of the original sample values of the original block itself. The decoding device can obtain a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by combining the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.
[0097] The residual information can be generated through a transformation process and a quantization process. For example, the encoding device can obtain a residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to obtain transform coefficients, and signal the relevant residual information to the decoding device (through a bitstream) by obtaining the transform coefficients quantized by performing a quantization process on the transform coefficients. Here, the residual information can include information such as the value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a dequantization process / inverse transformation process based on the residual information and obtain residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. The encoding device can also dequantize / inverse transform the quantized transform coefficients to obtain a residual block for reference in inter-frame prediction of subsequent pictures, and generate a reconstructed picture based on this.
[0098] <VCM (Video Coding for Machines)>
[0099] With the recent development of various industrial fields such as surveillance, intelligent transportation, smart cities, intelligent industries, and intelligent content, the amount of image or feature map data used by machines is increasing. In contrast, since the currently used traditional image compression methods are technologies developed by considering the human visual characteristics recognized by viewers and include unnecessary information, it is inefficient to perform machine tasks. Therefore, research on video codec technologies for efficiently compressing feature maps is needed to perform machine tasks.
[0100] Machine-oriented video coding (VCM) technology is being discussed by the Moving Picture Experts Group (MPEG), the international standards organization for multimedia coding. VCM is a picture or feature map coding technique that uses machine vision (machine vision) based on machine data, rather than the vision of a human viewer.
[0101] Figures 4a to 4d This is an exemplary diagram representing a VCM encoder and a VCM decoder.
[0102] Reference Figure 4a The diagram shows the VCM encoder 100a and the VCM decoder 100b.
[0103] When the VCM encoder 100a encodes video and / or feature maps and transmits them as a bitstream, the VCM decoder 100b can decode the bitstream and output it. In this case, the VCM decoder 100b can output one or more video and / or feature maps. For example, the VCM decoder 100b can output a first feature map for machine analysis or output a first image for user viewing. The first image may have a higher resolution than the first feature map.
[0104] Reference Figure 4b The feature extractor used to extract feature maps can be connected to the front end of the VCM encoder 100a.
[0105] VCM encoder 100a may include a feature encoder.
[0106] VCM decoder 100b may include a feature decoder and a video reconstructor. The feature decoder can decode feature maps from a bitstream and output a first feature map for machine analysis. The video reconstructor can reconstruct a first image from the bitstream and output the first image for user viewing.
[0107] Reference Figure 4c A feature extractor for extracting feature maps is connected to the front end of the VCM encoder 100. The VCM encoder 100a may include a feature encoder.
[0108] VCM decoder 100b may include a feature decoder. The feature decoder can decode feature maps from a bitstream and output a first feature map for machine analysis. In other words, the bitstream may be encoded only as feature maps, rather than images. To explain further, a feature map may be data that includes information about features used for a specific task of image-based machine processing.
[0109] Reference Figure 4d The feature extractor can be connected to the front end of the VCM encoder 100a.
[0110] The VCM encoder 100a may include a feature converter and a video encoder. The video encoder may be... Figure 2 The encoding device 10a shown is illustrated. Figure 4d The VCM decoder 100b shown may include a video decoder and an inverse converter. The video decoder may be... Figure 3 The decoding device 10b shown is illustrated.
[0111] Figure 5 It is a block diagram including components of an encoding device for machine vision according to embodiments of the present disclosure.
[0112] The encoding device 10a may include a temporal resampler 510, a spatial resampler 520, a region of interest extractor 530, and an internal encoder 540. Figure 5 In the structure of the encoding device, some components may be omitted in various embodiments, and the order of the components may be changed. In this disclosure, for ease of explanation, the encoding device may be referred to as an encoder or encoder, and the decoding device may be referred to as a decoder or decoder.
[0113] In VCM, the resolution and frame rate required for human vision are unnecessary for the practical application purposes of machine vision. The application purpose of the input image can be tracking, recognition, classification, or segmentation, but is not limited to these, and the purpose that can be used in machine vision can be defined in various ways. The required frame rate and resolution of the input image can be determined based on the application purpose of machine vision. The encoder and decoder can resample the input image based on the frame rate and resolution determined according to the application purpose, and can restore the resampled image to its original state if necessary. The encoder and decoder can transmit and receive resampled information (e.g., temporal_restoration_data and spatial_restoration_data) written according to a commonly defined format.
[0114] Reference Figure 5 The encoding device 10a can change the frame rate and resolution in the steps prior to internal encoding by performing temporal resampling and spatial resampling processing on the input signal in order to improve the encoding efficiency of machine vision, and perform encoding on the corresponding modified signal in the internal encoder.
[0115] When the frame rate and resolution are adaptively changed according to the application purpose, not only can the encoding and decoding efficiency be improved, but the effect of machine vision applications can also be enhanced. The encoding device 10a can transmit information related to the changes in frame rate and resolution to the decoding device 10b according to the application purpose and implementation method.
[0116] The input signal to the encoding device 10a for machine vision can be an image with video characteristics acquired by various sensors, such as images from general optical cameras, thermal cameras, and lidar sensors. Depending on the characteristics of the sensor, the corresponding image may include a single component or multiple components. For example, an image acquired by an optical camera may include multiple chromaticity components in the RGB and YUV domains.
[0117] The temporal resampler 510 is a step for changing the frame rate of the input image (or input signal), which can perform upsampling or downsampling on a frame-by-frame basis. The frame rate can be fixed depending on the application purpose of the machine vision, and when used adaptively, the corresponding frame rate can be transmitted to the decoder. According to the implementation, the frame rate can be variable, and for example, for an image with an input signal of 30 frames per second, some signals can be sampled at 5 frames per second, and other signals can be sampled at 1 frame per second.
[0118] Figure 6 This is an example of changing the resolution by temporal resampling according to the embodiments of this disclosure.
[0119] The time resampler 510 can determine the sampling period either fixedly or variably. The time resampler 510 can sample the entire input signal at a fixed period, or it can sample a portion of the input signal at a fixed period. Alternatively, the time resampler 510 can sample the input signal by variably determining the sampling period. For example, as... Figure 6 As shown in (a), sampling can be performed at a specific period, or as shown in (b), adaptive sampling can be performed for frames that are non-periodicly required by machine vision. The encoding device 10a can then transmit temporal resampling information (e.g., Figure 7a and Figure 7b The temporal resampler 510 transmits the temporal resoration data to the decoder. In an implementation, if necessary, the period or frame number (index) of the sampled frame can be transmitted to the decoder. The frame index can be an index sequentially assigned to all frames included in the input image. When the frame period is a fixed value predetermined by the decoder, the temporal resampler 510 may not transmit the period information separately to the decoder. When the frame period is variable, the temporal resampler 510 may transmit the period information separately. Alternatively, when the frame period is variable, the temporal resampler 510 may transmit the frame index number. According to an implementation, when the application purpose of the image decoded in the decoder and the frame rate according to the application purpose are fixed, information about the frame rate can be omitted. Information about the syntax transmitted from the encoder to the decoder is related to... Figure 7a and Figure 7bThe same applies as shown, which can be used for time recovery corresponding to time resampling in the decoder, and according to the implementation, when time recovery processing is not necessary in the decoder, the encoder can omit the transmission of information related to time recovery.
[0120] Figure 7a , Figure 7b and Figure 7c These are examples of the syntax and semantics for temporal resampling according to embodiments of this disclosure. Figure 7a This represents information related to temporal resampling (temporal_restoraion_data). Target_application_idx represents the information index (idx) for the machine vision target application to be used by the decoder. When not specified, information about the corresponding index of a list of target applications, such as tracking, recognition, classification, segmentation, etc., can be transmitted. In implementations, an idx value of 0 indicates no specification; 1 indicates tracking; 2 indicates recognition; 3 indicates classification; and 4 indicates segmentation.
[0121] Temporal_restoration_flag is a flag used to determine whether temporal resampling is applied.
[0122] The `Same_period_flag` flag indicates whether periodic or aperiodic resampling is used when applying time resampling. Depending on the implementation, when the time resampling period is fixed, this information can be omitted. When `Same_period_flag` is 1, it means that resampling is performed with the same period.
[0123] Since `Temporal_resampling_rete_idx` is an index to a list of resampling rates and is indicated by the `tempotal_restoration_flag` for temporal resampling applications, it does not include cases where the input and output signals are the same. For example, the sampling rate can be in the form of ratios such as 1 / 2, 1 / 4, 1 / 8, or it can be frames per second. The encoder and decoder can pre-agree on the corresponding list, and the encoder can transmit the index of the list. According to the implementation, when the resampling rate is fixed for the application being transmitted or the required number of frames is fixed, the corresponding information can be omitted.
[0124] When performing aperiodic resampling, the index information of the resampled frames (Same_period_flag) can be transmitted. When applying aperiodic resampling, the frame number can be calculated using temporal_resampling_rate_idx. For example, when the frame number is 30 and the resampling rate is 1 / 2, frame_num becomes 15. According to the implementation, when the frame number is fixed for the application, frame_num can be determined by a fixed number. Based on the frame number, the difference from the previous frame index (delta_frame_idx[i]) is transmitted. Since it is the difference from the previous index, a number that is 1 less than frame_num calculated based on the frame rate is transmitted.
[0125] According to the implementation, the time resampling period information can be transmitted via the difference between the previous reference frame index and the current reference frame index, without transmitting the aperiodic and periodic flags used for the time resampling period. For example, when time resampling is performed periodically, delta_frame_idx[i] has the same value. Figure 7b This is an example of the syntax used to transmit the difference between frame indices.
[0126] According to the implementation, information related to temporal restoration can be transmitted at various levels. For example, the temporal resampler 510 can transmit temporal restoration data at the GOP level, frame level, slice level, sequence level, subsequence level, etc. Figure 7a The disclosed implementation is based on sequence-by-sequence transmission, but does not include the actual details of the temporal_restoration of target_application_idx, which can be transmitted at various levels such as GOP, frame, slice, subsequence, etc. Figure 7a and Figure 7bThis describes the implementation of information that needs to be transmitted from the encoder to the decoder for temporal_restoration, but it does not imply that the corresponding information is transmitted at the same level all at once. Depending on the implementation, information related to temporal_restoration can be transmitted in a distributed manner at multiple levels, including sequences, subsequences, GOPs, frames, and slices. For example, target_application_idx can be transmitted at the sequence level, and all information after temporal_restoration_flag can be transmitted at the GOP level. Alternatively, target_application_idx can be transmitted at the sequence level, temporal_restoration_flag, same_period_flag, and temporal_resampling_rate_idx can be transmitted at the GOP level, and delta_frame_idx[i] can be transmitted at the frame or slice level.
[0127] Reference Figure 7c The time resampler 510 can transmit resample flags in slices, and when resampling is performed with the corresponding flag, it can transmit time resample-related information in a way that does not transmit the encoded information within the slice. From the HRD's perspective, trailing bits are transmitted to prevent underflow issues, so actually 1+a bits of information are transmitted.
[0128] Reference Figure 5 Spatial resampler 520 is a step for changing the resolution of the input image, which can upsample or downsample the resolution on a frame-by-frame basis. For example, spatial resampler 520 can downsample the resolution of the input image to half. Alternatively, according to an embodiment, spatial resampler 520 can upsample or downsample the input frame at different resolutions on a frame-by-frame basis.
[0129] Figure 8 This is an example of changing the resolution through spatial resampling in an embodiment of this disclosure. Figure 8 (a) is an example of resampling the input signal at the same resolution for each frame. For example, for all frames, an input signal with 4K resolution can be downsampled to half its original resolution and transformed to 2K resolution. As another example, Figure 8(b) The resolution of each frame of the input signal is resampled differently. For frames corresponding to a certain period (1st, 4th, 7th, ..., etc.), the resolution of the input signal can be downsampled to 1 / 2, and for the remaining frames (2nd, 3rd, 5th, 6th, 8th, 9th, ..., etc.), the resolution of the input signal can be downsampled to 1 / 4. The encoding device 10a can transmit the spatial resampling information (spatial_restoration_data) to the decoder.
[0130] Figure 9 This is an example of reordering frames by resolution after spatial resampling according to an embodiment of the present disclosure. According to the embodiment, in the step after spatial resampling, for efficiency and internal coding, signals resampled at different resolutions can be... Figure 9 The images are reordered by resolution for the next step.
[0131] like Figure 8 As shown, the spatial resampler 520 can upsample or downsample the entire input signal. Alternatively, as another embodiment, the spatial resampler 520 can upsample or downsample a portion of the input image.
[0132] Figure 10 This is another example of changing resolution through spatial resampling in embodiments of this disclosure. See also... Figure 10 (a) The spatial resampler 520 can downsample the resolution of the entire input signal. This is consistent with... Figure 8 Same. (Refer to...) Figure 10 (b) Spatial resampler 520 can change the resolution or size of an image by cropping only a portion of the input image to fit the desired resolution. (See reference...) Figure 10 (c) Spatial resampler 520 can crop a portion of the input image and then downsample the cropped image. See reference... Figure 10 (d) Spatial resampler 520 can crop a portion of the input image and then upsample the cropped image. Information about the syntax transmitted from the encoder to the decoder is related to... Figure 11a Same as shown. Figure 11a It can be used for spatial recovery corresponding to spatial resampling in the decoder. According to the implementation, when spatial recovery processing is not required in the decoder, the encoder can omit the transmission of information related to spatial recovery.
[0133] Figure 11a , Figure 11b and Figure 11c These are examples of syntax and semantics for spatial resampling based on embodiments of this disclosure. Figure 11aThis includes syntax related to spatial resampling of the entire input signal.
[0134] Target_application_idx represents the information index (idx) for the machine vision target application to be used by the decoder. When not specified, information about the corresponding index of a list of target applications, such as tracking, recognition, classification, segmentation, etc., can be transmitted. In an implementation, an idx value of 0 indicates that it is not specified; a value of 1 indicates tracking; a value of 2 indicates recognition; a value of 3 indicates classification; and a value of 4 indicates segmentation.
[0135] spatial_restoration_flag is a flag used to determine whether spatial resampling is applied.
[0136] The `same_ratio_flag` indicates whether each frame within a sequence is resampled at the same resolution. When `same_ratio_flag` is 1, it means that resampling is performed at different resolutions.
[0137] num_spatial_ratio_minus2 is the number of heterogeneous resolutions resampled, excluding the cases of 0 and 1. Therefore, when the number transmitted is 0, it can refer to two heterogeneous resolutions.
[0138] num_spatial_ration can be calculated by adding 2 to the transmitted num_spatial_ratio_minus2. When same_ratio_flag is 0 and no num_spatial_ratio_minus2 is transmitted, num_spatial_ratio becomes 1.
[0139] Spatial_resampling_rate_idx[i] represents an index of a list of resampling rates for each resolution, based on the type of heterogeneous resolution. For example, it can be a ratio such as 2, 4, 1 / 2, 1 / 4, etc., or it can be direct size information for horizontal and vertical resolutions. Since whether resampling is applied is transmitted by Spatial_resampling_flag, when num_spatial_ratio is 1, cases where the input and output signals have the same resolution are not included. According to the implementation, when num_spatial_ratio is greater than or equal to 2, it can have heterogeneous resolution after resampling, and therefore can include cases where the input and output signals have the same resolution.
[0140] `resampling_method_idx[i]` represents the index of the list of resampling methods by resolution. It can include information about methods such as downsampling, upsampling, cropping, etc., and information about the filters used in the corresponding resampling in the case of upsampling and downsampling.
[0141] When there are heterogeneous resolutions, the indexes of frames with the same resolution can be transmitted at a higher level, and in this case, the difference between the frame indices corresponding to each resolution can be transmitted as delta_frame_idx[i]. Alternatively, according to the implementation, resolution information can be transmitted to the decoder at the frame level for each frame. This may mean that information about the resolution of a frame can be transmitted directly from the encoder to the decoder frame by frame, or that the information can be transmitted by including indirect information that can determine the resolution. For example, the header of each frame may transmit the direct number of the horizontal and vertical resolutions of each frame, or an index of a list of resolution information may be transmitted. Alternatively, there may be an index of a list of direct values of ratios or ratio information to the resolution of previously transmitted frames, an index of a list of direct values of ratios or ratio information to the original resolution information previously transmitted, etc.
[0142] When clipping is used via a resampling method, the spatial resampler 520 can transmit information about the position of the clipped quadrilateral to the decoder. Figure 11b This includes syntax related to spatial resampling when cropping is applied (crop_location).
[0143] Location[x] represents the x-coordinate of the left corner of the cropped quadrilateral. The encoder can take the logarithm and transmit it for transmission efficiency, and the decoder can extract the coordinates within the actual image based on the values transmitted from the encoder in a pre-defined manner.
[0144] Location[y] represents the y-coordinate of the left corner of the cropped quadrilateral. The encoder can take the logarithm and transmit it for transmission efficiency, and the decoder can extract the coordinates within the actual image based on the values transmitted from the encoder in a pre-defined manner.
[0145] Restoration_ref_idx represents the reference index of the image used to recover the removed regions in the encoder. Depending on the implementation, one or more reference images may exist. The decoder can copy the corresponding reference image and use it to recover the adjacent cropped and removed regions.
[0146] The decoder can use location information when decoding a cropped image and recovering the cropped and removed regions during the restoration process. According to one implementation, location information may not be transmitted when the image is resampled in a cropped manner but the cropped and removed regions are not recovered separately. Figure 11c This includes spatial resampling-related syntax (crop_location) for cases where the position of the cropped quadrilateral is changed based on the location of the region of interest.
[0147] Location[x] represents the x-coordinate of the left corner of the cropped quadrilateral. The encoder can take the logarithm and transmit it for transmission efficiency, and the decoder can extract the coordinates within the actual image based on the values transmitted from the encoder in a pre-defined manner.
[0148] Location[y] represents the y-coordinate of the left corner of the cropped quadrilateral. The encoder can take the logarithm and transmit it for transmission efficiency, and the decoder can extract the coordinates within the actual image based on the values transmitted from the encoder in a pre-defined manner.
[0149] Restoration_ref_idx represents the reference index of the image used to recover the removed regions in the encoder. Depending on the implementation, one or more reference images may exist. The decoder can copy the corresponding reference image and use it to recover the adjacent cropped and removed regions.
[0150] The `loc_delta_flag` flag is used when the position of the cropped quadrilateral is changed in each frame, and it is necessary to transmit additional information about the changed position. This additional information can be transmitted via... Figure 11a The transmitted resampling_mehod_idx[i] and delta_frame_idx[i] are used for calculation.
[0151] delta_location[x] represents information about the position of the left corner of the cropped quadrilateral. It is the x-axis difference between the cropping position and the most recent frame of the same size, which is resampled to the same size by cropping in temporal order of the image.
[0152] delta_location[y] represents information about the position of the left corner of the cropped quadrilateral. It is the y-axis difference between the cropping position and the most recent frame of the same size, which is resampled to the same size by cropping in temporal order of the image.
[0153] When clipping is used during the resampling process, the spatial resampler 530 can change the position of the clipping quadrilateral based on the location of the region of interest, and in this case, it can be achieved through... Figure 11cInformation about the position of the cropped quadrilateral is transmitted to the decoder frame by frame. When transmitting frame by frame, it includes not only methods for directly transmitting the cropped quadrilateral by including a syntax corresponding to the header of the cropped frame (…). Figure 11c It also includes indirect methods for transmitting it at a higher level, but matching it with cropped frames to identify cropped locations within the corresponding frames. For example, by using the index of the cropped resampled frame and... Figure 11c Syntax information mapped to the corresponding index can be transmitted in units of sequences or GOPs.
[0154] Refer again Figure 5 The region of interest extractor 530 can distinguish regions predicted as objects from the background in the input signal for specific or general purposes required in machine vision applications. The region of interest extractor 530 can extract the location and size information of the regions predicted as objects through learning or in-image processing techniques. According to implementations, multiple regions of interest can be extracted from the input signal. The region of interest extractor 530 can transmit size and location information about each of the multiple regions of interest to an internal encoder 540. The internal encoder 540 can perform encoding by altering the image quality in a manner used to apply different quantizations to the object and the background.
[0155] The internal encoder 540 can perform encoding on the input signal. The internal encoder 540 can perform effective encoding by utilizing temporal resampling information from a pre-execution step against a reference structure or by modifying the reference method and quantization parameters of the object and background based on information about the region of interest. When... Figure 9 When the input includes input images of frames with various types of resolutions, the internal encoder 540 can perform encoding by supporting multiple subsequences within a sequence in the form of subsequences.
[0156] Figure 12 It is a block diagram including components of an encoding device for machine vision according to embodiments of the present disclosure. Figure 12 include Figure 5 All the components, but in a different order. Figure 12 Each component in corresponds to Figure 5 The components are listed, and the descriptions for each component are repeated and omitted.
[0157] Reference Figure 12 The encoding device 10a can first extract the region of interest (ROI) of the input signal, and then sequentially perform temporal resampling, spatial resampling, and internal encoding. According to... Figure 12The encoding apparatus 10a, with its encoding structure, can adaptively perform temporal and / or spatial resampling based on region-of-interest (ROI) information extracted from the input signal. For example, the encoding apparatus 10a can perform frame-centric aperiodic temporal resampling to extract ROIs suitable for the application purpose, or change the resolution through cropping. The processes performed in each step of the encoding apparatus 10a are similar to... Figure 5 The same applies as shown, but since the region of interest is extracted first, the temporal resampler 510 and the spatial resampler 520 can use information about the region of interest.
[0158] An encoding apparatus 10a for performing machine image encoding according to an embodiment of the present disclosure includes a temporal resampler 510 for changing the frame rate of an input image, a spatial resampler 520 for changing the resolution of an input image, and an internal encoder 540 for performing encoding of the input image. The temporal recovery data including the frame rate of the input image changed by the temporal resampler and the spatial recovery data including the resolution of the input image changed by the spatial resampler are transmitted together with the encoded image to a decoding apparatus 10b.
[0159] The temporal resampler 510 upsamples or downsamples all frames included in the input image on a frame-by-frame basis, and the sampling period can be determined according to the application purpose of the input image.
[0160] When the time resampler 510 periodically resamples the input image, it can transmit the sampling period to the decoding device 10b.
[0161] When the time resampler 510 performs resampling aperiodically, it can transmit the index number of the resampled frame.
[0162] The temporal resampler 510 transmits the difference between the previous frame index and the current frame index as information about the frame rate according to the resampling to the decoding device 10b, and the difference can be calculated as the difference between corresponding pixels between the two frames or the mean square error (MSE) of motion prediction between the two frames.
[0163] Time recovery data can be transmitted in whole or in part as at least one of sequence units, GOP units, frame units, slice units, and subsequence units.
[0164] The spatial resampler 520 upsamples or downsamples the resolution of the input image on a frame-by-frame basis, and the target resolution for resampling can be the same for all frames, or it can be determined differently on a frame-by-frame basis, depending on the application purpose.
[0165] Spatial resampler 520 can resample the resolution of all or a portion of each frame of the input image. A portion of a frame of the input image can be obtained by cropping a portion of the original frame. Spatial resampler 520 can upsample or downsample the resolution of the portion of the input image cropped by the cropping method.
[0166] The spatial resampler 520 can change the target resolution by cropping a portion of the input image, depending on the application purpose.
[0167] When the resolution of each frame of the input image is resampled differently, the spatial resampler 520 can reorder the resampled input image according to the resolution.
[0168] Temporal resampler 510 and spatial resampler 520 resample the frame rate and resolution based on the application purpose of the input image, respectively, and the application purpose of the input image can correspond to any of tracking, recognition, classification and segmentation.
[0169] The encoding apparatus 10a also includes a region of interest extractor 530 for extracting one or more regions of interest from the input image. The region of interest extractor 530 can distinguish between objects and background in the input image and extract one or more regions of interest by corresponding to one or more objects. The temporal resampler 510 and the spatial resampler 520 can perform temporal resampling and spatial resampling based on information about one or more regions of interest in the input image.
[0170] The region of interest extractor 530 can extract the region of interest from an input image in which at least one of temporal resampling and spatial resampling of the input image has been performed.
[0171] Figure 13 It is a block diagram including components of a decoding apparatus for machine vision according to an embodiment of the present disclosure.
[0172] The decoding device 10b may include an internal decoder 1310, a space restorer 1320, and a time restorer 1330. In various embodiments, some components of the decoding device 10b may be omitted, or the execution order between the components may be changed.
[0173] The internal decoder 1310 can perform decoding by receiving a bitstream as input. Depending on the implementation, various decoders can be used, and in this case, the internal decoder 1310 is determined by corresponding to the internal encoder 540 used in the encoder.
[0174] If necessary, the spatial restorer 1320 can perform spatial restoration using spatial restoration data (e.g., spatial_restoration_data) transmitted from the encoder. If not necessary, the spatial restorer 1320 may not need to perform spatial restoration for machine vision applications. For example, when an image is downsampled in the encoder's spatial resampler 520 and encoded and transmitted to the decoder, but the decoder determines for its application purposes that it does not need to be upsampled at the downsampled resolution, the next step can be performed without performing restoration processing using the spatial restoration data transmitted from the encoder. Alternatively, when the encoder determines that the application does not require spatial restoration, i.e., when spatial_restoration_flag is 0, the decoder does not perform spatial restoration because the encoder does not transmit information about spatial restoration. When the decoder determines that spatial restoration is necessary or when spatial_restoration_flag is 1, the decoder can extract information for spatial restoration of the decoded image from its internal decoder.
[0175] Spatial restorer 1320 can first check whether all frames have been resampled at the same resolution and extract information related to the resolution and the number of resampling methods as there are heterogeneous resolutions. When resampled at a single resolution, spatial restorer 1320 can extract information for restoring the corresponding resolution. Spatial restorer 1320 can determine the restoration method based on the resampling method. For example, when performing upsampling or downsampling at a specific ratio, the image decoded at the corresponding specific ratio can be upsampled or downsampled. In this case, when transmitting sampling filter information, spatial restorer 1320 can apply it to the restoration filter based on the corresponding information.
[0176] According to another embodiment, when resampling is performed by pruning, the spatial restorer 1320 can utilize the position information of the quadrilateral pruned by the decoder ( Figure 11b The decoded image is located to the internal decoder, and adjacent regions removed from the encoder and not decoded by the internal decoder are recovered. In this case, the spatial restorer 1320 can determine the frame referenced for recovery using restoration_ref_idx and perform recovery using the corresponding image. For example, the region other than the cropped region can be used by similarly copying the signal of the frame corresponding to the corresponding index. Alternatively, recovery can be performed by generating images using deep learning. According to the implementation, when the position in the cropped frame changes, it can be... Figure 11c The information is used to perform the recovery.
[0177] If necessary, the time restorer 1330 can perform time restoration by using time restoration data transmitted from the encoder. If not necessary, it can be performed without targeting the time restoration data. Figure 7a and Figure 7b The time recovery is performed according to the application's intended purpose. Alternatively, when the encoder determines that the application does not require time recovery, i.e., when `temporal_restoration_flag` is 0, the decoder may not perform time recovery because the encoder does not transmit information about time recovery. When the decoder determines that time recovery is required and `temporal_restoration_flag` is 1, the time recovery unit 1330 can extract the necessary information from the time recovery information (`temporal_restoration_data`) transmitted from the encoder. First, the time recovery unit 1330 can determine whether sampling is performed at the same period using `same_period_flag`. When sampling is performed at the same period, the signal decoded by the internal decoder 1310 and output by the spatial recovery unit 1320 is sampled at the same period, so the time recovery unit 1330 can recover the image corresponding to the intermediate frame in time, thereby recovering it at the frame rate before sampling. For example, the time recovery unit 1330 can perform recovery by interpolating the input signal, or by performing recovery by copying. Alternatively, it can also perform recovery by deep learning. Since the sampling interval differs when sampling is performed at different periods (not the same period), the time restorer 1330 can extract information about the time sampling period through delta_frame_idx[i] and perform restoration based on the interval between indices (sampling period). If the sampling is performed at the same period, the time restorer 1330 can perform restoration by using methods such as interpolation, copying, deep learning, etc.
[0178] Decoder 10b is the image signal that performs the time-recovery step, and filtering may be performed due to factors such as image quality degradation in the recovered signal during spatial and temporal recovery processing. Depending on the application purpose of machine vision, the output signal can become the input of the application. For example, the output signal can become the input value of a neural network model.
[0179] According to another embodiment of this disclosure, a decoding apparatus 10b for performing machine image decoding may include: an internal decoder for receiving an encoded input image from an encoding apparatus 10a as input and performing decoding; a spatial restorer for changing the resolution of the decoded image based on spatial restoration data of the input image received from the encoding apparatus and restoring it to the resolution of the original image of the input image; and a temporal restorer for changing the frame rate of the decoded image based on temporal restoration data of the input image received from the encoding apparatus and restoring it to the frame rate of the original image of the input image.
[0180] The spatial restorer can upsample or downsample the resolution of the input image on a frame-by-frame basis based on spatial restoration data.
[0181] Spatial restoration data may include at least one of the following: whether spatial resampling is applied, an index of a list of resampling methods by resolution, information about the resampling methods, information about the filters used by the corresponding resamplers in the case of upsampling and downsampling, location information of the cropped region, and image information referenced for restoring the region removed according to spatial resampling in the encoding device.
[0182] A time-recovery function can recover an input image based on time-recovery data at the frame rate of the original image.
[0183] The time recovery data may include at least one of the following: whether time resampling is applied, whether it is periodic resampling or non-periodic resampling, and, in the case of non-periodic resampling, the index information or frame period of the resampled frame.
[0184] As another embodiment of the present disclosure, a non-volatile computer-readable storage medium for recording commands, when executed by at least one processor, can cause the at least one processor to: change the frame rate of an input image; change the resolution of the input image; encode the input image; and transmit time-recovered data including the frame rate of the input image changed according to the time resampling and spatial-recovered data including the resolution of the input image changed according to spatial resampling, together with the encoded image, to a decoding device.
[0185] Figure 14 This is a flowchart used to describe a time resampling method based on an example of this disclosure.
[0186] In S1410, the encoding device 10a or the temporal resampler 510 can calculate the optimal resampling interval for encoding and decoding efficiency and transmit it to the decoder. The temporal resampler 510 can calculate delta_frame_idx[i] corresponding to the temporal resampling interval. The temporal resampler 510 can calculate the frame variation between input images. The variation can be calculated by the difference between corresponding pixels between two images or by the mean square error (MSE) of motion estimation (ME) between two images.
[0187] In S1420, the temporal resampler 510 can select a reference frame. The temporal resampler 510 can calculate the frame change between each frame in S1410 and determine the reference frame based on information about segments where the frame change exceeds a certain size or segments where the cumulative frame change from the previous reference frame exceeds a certain size.
[0188] In S1430, the time resampler 510 can transmit the interval between the current reference frame and the previous reference frame to the decoder via delta_frame_idx[i].
[0189] Figure 15 This is a flowchart used to describe another example of a time resampling method based on this disclosure.
[0190] In S1510, the encoding device 10a or the temporal resampler 510 can calculate the optimal resampling interval for encoding and decoding efficiency and transmit it to the decoder. The encoder can calculate delta_frame_idx[i] corresponding to the temporal resampling interval. The temporal resampler 510 calculates the frame variation between the input images. The variation can be calculated by the difference between corresponding pixels between two images or by the MSE of motion estimation between two images.
[0191] In S1520, the temporal resampler 510 determines candidates for reference frames based on information about segments where frame changes exceed a certain size or segments where cumulative frame changes from the previous reference frame exceed a certain size.
[0192] In steps S1530 to S1550, the temporal resampler 510 can select the optimal recovery method by simulating various recovery methods that can be executed by the decoder using the reference frames identified as candidate frames in the above steps. The temporal resampler 510 can determine the optimal recovery method by calculating the error between the frames recovered through restoration and the frames resampled through temporal resampling. Figure 16a and Figure 16b Information related to the best recovery method can be transmitted to the decoder. Figure 16a and Figure 16bThese are examples of the syntax and semantics for temporal resampling according to embodiments of this disclosure.
[0193] In S1550, the temporal resampler 510 can predict the recovery error of the frame to be recovered by the decoder through temporal recovery, and when the reference frame candidate is not suitable, the best reference frame can be selected by repeatedly performing verification via the reference frame candidate and temporal recovery.
[0194] In S1560, the temporal resampler 510 can transmit information about the interval between reference frames determined in the above steps (i.e., the number of temporally sampled frames between reference frames) as delta_frame_idx[i] to the decoder. According to the implementation, the optimal frame rate for the application can be determined, and the optimal delta_frame_idx[i] can be calculated by reflecting the frame changes between the corresponding frame rate and the reference frames, as well as the recovery error for temporal recovery.
[0195] In S1570, the time resampler 510 can be based on Figure 16a and Figure 16b Pass delta_frame_idx[i] to the decoder. Figure 16a and Figure 16b It represents the syntax and semantics of information that the time resampler 510 performs time resampling to correspond to the best recovery selected by simulating various recoveries.
[0196] Figure 16a and Figure 16b These are examples of the syntax and semantics for temporal resampling according to embodiments of this disclosure.
[0197] Figure 16a Includes the following syntax.
[0198] `target_application_idx` represents the information index (idx) for the machine vision target application to be used by the decoder. When not specified, information about the corresponding index of a list of target applications, such as tracking, recognition, classification, segmentation, etc., can be transmitted. In an implementation, an idx value of 0 indicates that it is not specified; a value of 1 indicates tracking; a value of 2 indicates recognition; a value of 3 indicates classification; and a value of 4 indicates segmentation.
[0199] temporal_restoration_flag is a flag used to determine whether temporal resampling is applied.
[0200] When applying aperiodic resampling, the frame number can be calculated using `temporal_resampling_rate_idx`. For example, when the frame number is 30 and the resampling rate is 1 / 2, `frame_num` becomes 15. According to the implementation, when the frame number is fixed for the application, `frame_num` can be determined by a fixed number. Based on the frame number, the difference between the previous frame index (`delta_frame_idx[i]`) is transmitted. Since it is the difference from the previous index, a number that is 1 less than `frame_num` calculated using the frame rate is transmitted.
[0201] `restoration_method[i]` represents the index of a list of restoration methods agreed upon between the decoder and encoder. For example, the index of the list of restoration methods agreed upon between the decoder and encoder is transmitted, such as 0 for temporal bilinear interpolation, 1 for copying the previous reference frame at display time, 2 for copying the previously restored image in decoding order, and 3 for deep learning-based image restoration, etc., and in the method corresponding to the index, restoration is performed in the same manner for the frame restored at time delta_frame_idx[i].
[0202] Figure 16b Includes the following syntax.
[0203] `target_application_idx` represents the information index (idx) for the machine vision target application to be used by the decoder. When not specified, information about the corresponding index of a list of target applications, such as tracking, recognition, classification, segmentation, etc., can be transmitted. In an implementation, an idx value of 0 indicates that it is not specified; a value of 1 indicates tracking; a value of 2 indicates recognition; a value of 3 indicates classification; and a value of 4 indicates segmentation.
[0204] temporal_restoration_flag is a flag used to determine whether temporal resampling is applied.
[0205] When applying aperiodic resampling, the frame number can be calculated using `temporal_resampling_rate_idx`. For example, when the frame number is 30 and the resampling rate is 1 / 2, `frame_num` becomes 15. According to the implementation, when the frame number is fixed for the application, `frame_num` can be determined by a fixed number. Based on the frame number, the difference between the previous frame index (`delta_frame_idx[i]`) is transmitted. Since it is the difference from the previous index, a number that is 1 less than `frame_num` calculated using the frame rate is transmitted.
[0206] `restoration_method[i][j]` represents the index of a list of restoration methods agreed upon between the decoder and encoder. For example, it transmits the index of a list of restoration methods agreed upon between the decoder and encoder, such as 0 for temporal bilinear interpolation, 1 for copying the previous reference frame at display time, 2 for copying the previously restored image in decoding order, and 3 for deep learning-based image restoration, etc. In this case, the index for the restoration method can be transmitted separately for each frame time-restored via `delta_frame_idx[i]`, and the decoder can perform temporal restoration by using the restoration method transmitted for each frame.
[0207] The examples and figures presented in this specification are merely specific examples to facilitate the explanation of the technical content of this disclosure and to aid in understanding it, and are not intended to limit the scope of this disclosure. It will be apparent to those skilled in the art that other variations besides the examples described above may be feasible.
[0208] The claims set forth in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in an apparatus, or the technical features of the apparatus claims in this specification can be combined and implemented in a method. Furthermore, the technical features of the method claims and the apparatus claims in this specification can be combined and implemented in an apparatus, or the technical features of the method claims and the apparatus claims in this specification can be combined and implemented in a method.
Claims
1. An encoding apparatus for performing machine-oriented image encoding, the encoding apparatus comprising: Temporal resampler, used to change the frame rate of the input image; A spatial resampler is used to change the resolution of the input image; as well as An internal encoder is used to perform encoding on the input image. At least one of the time-recovery data, which includes the frame rate of the input image changed by the time resampler, and the spatial recovery data, which includes the resolution of the input image changed by the spatial resampler, is transmitted to the decoding device along with the encoded image.
2. The apparatus according to claim 1, wherein: The temporal resampler upsamples or downsamples all frames included in the input image on a frame-by-frame basis, and The sampling period is determined based on the application purpose of the input image.
3. The apparatus according to claim 1, wherein: When periodic resampling is performed on the input image, the temporal resampler transmits the sampling period to the decoding device.
4. The apparatus according to claim 1, wherein: When performing aperiodic resampling, the time resampler transmits the index number of the resampled frame.
5. The apparatus according to claim 1, wherein: The temporal resampler transmits the difference between the previous frame index and the current frame index as information about the resampled frame rate to the decoding device, and The difference is calculated as the difference between pixels at corresponding positions between two frames or the mean square error (MSE) of motion prediction between the two frames.
6. The apparatus according to claim 1, wherein: The time recovery data is transmitted in whole or in part as at least one of sequence units, GOP units, frame units, slice units, and subsequence units.
7. The apparatus according to claim 1, wherein: The spatial resampler upsamples or downsamples the resolution of the input image on a frame-by-frame basis, and The target resolution for resampling may be the same for all frames, depending on the application purpose, or it may be determined differently on a frame-by-frame basis.
8. The apparatus according to claim 1, wherein: The spatial resampler resamples all or part of the resolution of each frame of the input image.
9. The apparatus according to claim 8, wherein: The portion of each frame of the input image is obtained by a cropping method used to crop a portion of the original frame.
10. The apparatus according to claim 9, wherein: The spatial resampler upsamples or downsamples the resolution of a portion of the original frame cropped by the cropping method.
11. The apparatus according to claim 1, wherein: The spatial resampler changes the target resolution according to the application purpose by using a cropping method to crop a portion of the input image.
12. The apparatus according to claim 1, wherein: When the resolution of each frame of the input image is resampled differently, the spatial resampler reorders the resampled input image according to the resolution.
13. The apparatus according to claim 1, wherein: The temporal resampler and the spatial resampler resample the frame rate and resolution based on the application purpose of the input image, respectively. The application purpose of the input image corresponds to any one of tracking, recognition, classification, and segmentation.
14. The apparatus according to claim 1, wherein: The apparatus further includes a region of interest extractor for extracting one or more regions of interest from the input image, and The region of interest extractor distinguishes between objects and background in the input image to extract one or more regions of interest in response to one or more objects.
15. The apparatus according to claim 14, wherein: The temporal resampler and the spatial resampler perform temporal resampling and spatial resampling based on information about one or more regions of interest in the input image.
16. The apparatus according to claim 14, wherein: The region of interest extractor extracts the region of interest for the input image for which at least one of temporal resampling and spatial resampling of the input image has been performed.
17. A decoding apparatus for performing machine-oriented image decoding, the decoding apparatus comprising: An internal decoder is used to receive encoded input images from the encoding device as input and perform decoding. A spatial restorer is used to change the resolution of a decoded image based on spatial restoration data of the input image received from the encoding device, and restore the resolution to the original resolution of the input image; as well as A time restorer is used to change the frame rate of the decoded image based on time-restored data of the input image received from the encoding device, and restore the frame rate to the frame rate of the original image of the input image.
18. The apparatus according to claim 17, wherein: The spatial restorer upsamples or downsamples the resolution of the input image frame by frame based on the spatial restoration data, and The spatial recovery data includes at least one of the following: whether spatial resampling is applied, an index of a list of resampling methods by resolution, information about the resampling methods, information about the filters used by the corresponding resamplers in the case of upsampling and downsampling, location information of the cropped region, and image information referenced for recovering the region removed according to spatial resampling in the encoding device.
19. The apparatus according to claim 17, wherein: The time restorer uses the time restoration data to restore the input image to the frame rate of the original image of the input image, and The time recovery data includes at least one of the following: whether time resampling is applied, whether it is periodic resampling or non-periodic resampling, and in the case of non-periodic resampling, the frame period or index information of the resampled frame.
20. A non-transitory computer readable storage medium for recording commands, wherein, The command causes the at least one processor to execute when it is executed by the at least one processor: The temporal resampling step is used to change the frame rate of the input image; A spatial resampling step is used to change the resolution of the input image; The step of encoding the input image; as well as The step of transmitting, together with the encoded image, temporal recovery data including the frame rate of the input image changed according to the temporal resampling and spatial recovery data including the resolution of the input image changed according to the spatial resampling, to a decoding device.