Encoding / decoding method and device based on region of interest distribution, and recording medium
The region of interest distribution-based encoding and decoding method addresses the challenges of server load and power consumption in machine-dependent image analysis by dividing images into specific regions, enhancing efficiency and reducing resource demands.
Patent Information
- Application Number
- PCT/KR2025/095367
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-08
- Filing Date
- 2025-05-23
- Publication Date
- 2025-11-27
AI Technical Summary
Existing video coding technologies face challenges in efficiently handling machine-dependent image analysis due to increasing server load and power consumption as the amount of images to be analyzed by machines grows exponentially, particularly with the rise of high-resolution and high-quality images/videos.
A region of interest distribution-based encoding and decoding method that divides input images into specific regions of interest, determines coding sections, and encodes region of interest distribution information to generate a bitstream, utilizing methods like neural networks for object detection and segmentation.
This approach enables effective machine-based image analysis by optimizing encoding and decoding processes, reducing server load and power consumption while maintaining high-quality image processing.
Smart Images

Figure KR2025095367_27112025_PF_FP_ABST
Abstract
Description
Encoding / decoding method, device, and recording medium based on region of interest distribution
[0001] The present disclosure relates to a method, device and recording medium for encoding / decoding an image based on region of interest distribution for a machine.
[0002] With the continuous development of the information and communication industry, broadcasting services with HD (High Definition) resolution have spread worldwide.
[0003] Through this proliferation, many users have become accustomed to high-resolution and high-quality images and / or videos, and the demand for higher-resolution and high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, has increased in various fields.
[0004] The technology for coding this UHD video data was completed in 2013 through HEVC (High Efficiency Video Coding), a standard technology.
[0005] HEVC is a next-generation video compression technology with a higher compression ratio and lower complexity than the previous H.264 / AVC technology, and is a key technology for effectively compressing the massive data of HD and UHD video.
[0006] HEVC performs block-by-block encoding, like previous compression standards.
[0007] However, unlike H.264 / AVC, there is only one profile. The core encoding technologies included in HEVC's sole profile are divided into eight areas: hierarchical encoding structure technology, transform technology, quantization technology, intra-frame prediction encoding technology, inter-frame motion prediction technology, entropy encoding technology, loop filter technology, and other technologies.
[0008] Since the establishment of the HEVC video codec in 2013, the Versatile Video Coding (VVC) standard, a next-generation video codec that aims to improve performance by more than twice that of HEVC, has been developed to address the expansion of realistic video and virtual reality services utilizing 4K and 8K video images. VVC is called H.266.
[0009] H.266 (VVC) was developed with the goal of being more than twice as efficient as the previous generation codec, H.265 (HEVC). VVC was initially developed with resolutions over 4K in mind, but it was also developed for ultra-high-resolution video processing at a whopping 16K level to support 360-degree videos due to the expansion of the VR market. In addition, as the HDR market is expanding due to the development of display technology, it supports 16-bit color depth as well as 10-bit color depth to respond to this, and supports brightness expressions of 1000 nits, 4000 nits, and 10000 nits. In addition, since it is being developed with the VR market and 360-degree video market in mind, it supports partial frame rates in the range of 0 to 120 FPS.
[0010] Advances in Artificial Intelligence
[0011] Artificial intelligence (AI) is also steadily developing. AI refers to the artificial imitation of human intelligence, including the ability to recognize, classify, infer, predict, and control / decision-making.
[0012] With the advancement of artificial intelligence technology and the increase in Internet of Things (IoT) devices, machine-to-machine traffic is expected to explode, and machine-dependent image analysis is expected to become widely used.
[0013] However, as the amount of images to be analyzed by machines is expected to increase exponentially, issues with server load and power consumption are expected to arise.
[0014] Accordingly, the present disclosure aims to provide a region of interest distribution-based encoding and decoding method to enable effective machine-based image analysis.
[0015] The region of interest distribution-based encoding / decoding method, device, and recording medium of the present disclosure may include the steps of: dividing at least one region of interest for each frame of an input image; determining a first region of interest distribution coding section from among two or more consecutive frame sections within a sequence of the input image; selecting at least one region of interest for each frame within the first region of interest distribution coding section as a region of interest distribution coding target, and distributing the region of interest to frames within the first region of interest distribution coding section; and encoding region of interest division information according to the division and the input image to generate a bitstream.
[0016] In the encoding / decoding method, device and recording medium based on region of interest distribution of the present disclosure, the at least one region of interest is divided into one or more types, and the region of interest division information may include a type for each of the at least one region of interest.
[0017] In the method, device, and recording medium for encoding / decoding based on region of interest distribution of the present disclosure, the step of distributing the at least one selected region of interest within the first region of interest distribution coding section may include the step of moving the first type of regions of interest included in the first frame included in the first region of interest distribution coding section to at least a portion within the first region of interest distribution coding section.
[0018] In the encoding / decoding method, device, and recording medium based on region of interest distribution of the present disclosure, the step of dividing the input image into at least one region of interest for each frame may include the step of deriving first types of regions of interest for each frame by a first region of interest derivation method; and the step of deriving second types of regions of interest for each frame by a second region of interest derivation method.
[0019] In the encoding / decoding method, device, and recording medium based on region of interest distribution of the present disclosure, the first region of interest derivation method independently derives a region of interest for each frame, and at least one of the positions or sizes of the derived regions of interest may be different, or at least some of the derived regions of interest may include a portion overlapping with each other.
[0020] In the encoding / decoding method, device, and recording medium based on region of interest distribution of the present disclosure, the first region of interest derivation method may use at least one of a method of deriving a region of interest by detecting movement through comparison with a current frame and adjacent frames, a method of using a neural network for object detection, or a method of using a segmentation lightweight neural network.
[0021] In the present disclosure, the encoding / decoding method, device and recording medium based on the distribution of the region of interest,
[0022] The above second method for deriving a region of interest can derive a region of interest by at least one of a position or a size specified by a sequence or frame group unit in all frames.
[0023] In the region of interest distribution-based encoding / decoding method, device, and recording medium of the present disclosure, the region of interest division information can be expressed by a method of directly specifying each region of interest unit or by a method of displaying it as a map based on basic units.
[0024] In the encoding / decoding method, device, and recording medium based on distribution of a region of interest of the present disclosure, the method of directly specifying the region unit includes, for each frame, at least one of the type of shape of each region of interest, a reference point of each region of interest, a width of each region of interest, or a height of each region of interest, to express each region of interest, and the method of displaying it as a map based on the basic unit includes, for the entire frame, a form in which two or more pixels are grouped in a rectangular shape, and each basic unit is listed so as not to overlap, and each region of interest can be expressed in correspondence to the basic units divided by a specified absolute size or relative size.
[0025] In the region of interest distribution-based encoding / decoding method, device, and recording medium of the present disclosure, in the step of selecting at least one region of interest as a region of interest distribution coding target for each frame within the first region of interest distribution coding section, and distributing the region of interest to frames within the first region of interest distribution coding section, the second type of regions of interest included in the first frame within the first region of interest distribution coding section may be divided into all frames within the first region of interest distribution coding section one by one, or the second type of regions of interest included in each frame within the first region of interest distribution coding section may be deleted except for one without overlapping.
[0026] In the region of interest distribution-based encoding / decoding method, device, and recording medium of the present disclosure, the region of interest segmentation information may include at least one of information on whether the period in which the region of interest distribution coding section appears within the sequence is fixed, or period information.
[0027] In the region of interest distribution-based encoding / decoding method, device, and recording medium of the present disclosure, the region of interest segmentation information may include at least one of an index indicating whether the first frame of the input image is a frame within a region of interest distribution coding section, an index indicating which frame the first frame is within the region of interest distribution coding section, an index of the type of the region of interest that is a target of distribution coding, or an index of the region of interest that is a target of distribution coding.
[0028] According to the present disclosure, image analysis by a machine can be effectively performed.
[0029] Figure 1 schematically illustrates an example of a video / image coding system.
[0030] Figure 2 is a drawing schematically illustrating the configuration of a video / image encoding device.
[0031] Figure 3 is a drawing schematically illustrating the configuration of a video / image decoding device.
[0032] Figures 4a to 4d are exemplary diagrams showing a VCM encoder and a VCM decoder.
[0033] FIG. 5 illustrates a block diagram of an encoding device according to one embodiment of the present disclosure.
[0034] FIG. 6 is a flowchart illustrating an operation of an encoding device for region-of-interest-based image processing according to one embodiment of the present disclosure.
[0035] FIG. 7 is an example of regions of interest segmented in various ways according to one embodiment of the present disclosure.
[0036] FIGS. 8a, 8b and 8c are examples of a template-based direct region of interest specification method according to one embodiment of the present disclosure.
[0037] FIG. 9 is an example of a direct region of interest specification method based on a free-form specification method according to an embodiment of the present disclosure.
[0038] FIG. 10 is an example of a method for directly specifying region segmentation information by region of interest unit according to one embodiment of the present disclosure.
[0039] FIG. 11 is an example of a method for displaying region of interest segmentation information as a basic unit map according to one embodiment of the present disclosure.
[0040] FIG. 12 is an example of a region of interest distribution coding section according to one embodiment of the present disclosure.
[0041] FIG. 13a and FIG. 13b are examples of setting a distribution coding target according to a first distribution coding target setting method according to one embodiment of the present disclosure.
[0042] FIGS. 14a to 14c are examples of setting a distribution coding target according to a second distribution coding target setting method according to one embodiment of the present disclosure.
[0043] FIG. 15 illustrates a block diagram of a decoding device according to one embodiment of the present disclosure.
[0044] FIG. 16 is a flowchart illustrating an operation of a decoding device to restore a region-of-interest-based image according to one embodiment of the present disclosure.
[0045] FIG. 17 is a flowchart illustrating a detailed operation of a region of interest-based resizing according to an embodiment of the present disclosure.
[0046] FIG. 18 is an example of a subframe division structure derived according to the resolution of a restored frame according to an embodiment of the present disclosure.
[0047] FIG. 19 is an example of deriving a subframe division structure from a frame resolution before resizing according to an embodiment of the present disclosure.
[0048] FIG. 20 is an example of a subframe unit resizing index according to one embodiment of the present disclosure.
[0049] FIG. 21 is an example of a result of subframe resizing according to one embodiment of the present disclosure.
[0050] Figure 22 is an example of a case where the area of interest distribution coding section is constant according to one embodiment of the present disclosure.
[0051] FIG. 23 is an example of a result of identifying a target of distribution coding of a region of interest according to one embodiment of the present disclosure.
[0052] FIG. 24 is an example of copying the region of interest of each frame identified within the region of interest distribution coding section to the first frame according to one embodiment of the present disclosure.
[0053] FIG. 25 is an example showing a filtering target according to one embodiment of the present disclosure.
[0054] Specific structural or step-by-step descriptions of embodiments according to the concept of the present disclosure disclosed in this specification or application are merely illustrative for the purpose of explaining embodiments according to the concept of the present disclosure, and embodiments according to the concept of the present disclosure may be implemented in various forms, and embodiments according to the concept of the present disclosure may be implemented in various forms and should not be construed as being limited to the embodiments described in this specification or application.
[0055] Embodiments according to the concept of the present disclosure may have various modifications and take various forms. Therefore, specific embodiments are illustrated in the drawings and described in detail in this specification or application. However, this is not intended to limit embodiments according to the concept of the present disclosure to specific disclosed forms, and it should be understood that all modifications, equivalents, and alternatives included within the spirit and technical scope of the present disclosure are included.
[0056] While terms such as "first" and / or "second" may be used to describe various components, these components should not be limited by these terms. These terms are only intended to distinguish one component from another; for example, without departing from the scope of the present disclosure, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component."
[0057] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. Conversely, when a component is referred to as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Other expressions that describe the relationship between components, such as "between" and "directly between" or "adjacent to" and "directly adjacent to", should be interpreted similarly.
[0058] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the present disclosure. The singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0059] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0060] Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined herein.
[0061] In describing the embodiments, description of technical contents that are well known in the technical field to which the present disclosure belongs and are not directly related to the present disclosure will be omitted.
[0062] This is to convey the gist of the present disclosure more clearly without obscuring it by omitting unnecessary explanations.
[0063] This document relates to video / image coding. For example, the method / embodiment disclosed in this document may be related to the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), the essential video coding (EVC) standard, the AVS2 standard, etc.).
[0064] This document presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0065] In this document, "video" can refer to a series of images over time. "Picture" generally refers to a unit representing a single image from a specific time period, and "slice" / "tile" are units that constitute part of a picture in coding.
[0066] A slice / tile can contain one or more coding tree units (CTUs). A picture can consist of one or more slices / tiles. A picture can consist of one or more tile groups. A tile group can contain one or more tiles.
[0067] A pixel or pel can mean the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Alternatively, a sample can mean a pixel value in the spatial domain, or when such a pixel value is converted to the frequency domain, it can mean a transform coefficient in the frequency domain.
[0068] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region.
[0069] A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term "unit" may sometimes be used interchangeably with the terms "block" or "area." In general, an MxN block can contain a set (or array) of samples (or array of samples) or transform coefficients, each consisting of M columns and N rows.
[0070] Figure 1 schematically illustrates an example of a video / image coding system.
[0071] Referring to FIG. 1, a video / image coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or streaming.
[0072] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a reception unit, a decoding device, and a renderer.
[0073] The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0074] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0075] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0076] The transmission unit can transmit encoded video / image information or data output in bitstream form to the receiving unit of the receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file using a predetermined file format and an element for transmission via a broadcasting / communication network.
[0077] The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0078] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0079] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0080] Figure 2 is a drawing schematically illustrating the configuration of a video / image encoding device.
[0081] The term “video encoding device” hereinafter may include a video encoding device.
[0082] Referring to FIG. 2, the encoding device (10a) may be configured to include an image partitioner (10a-10), a prediction unit (predictor) (10a-20), a residual processor (residual processor) (10a-30), an entropy encoder (entropy encoder) (10a-40), an adder (adder) (10a-50), a filter (filter) (10a-60), and a memory (10a-70). The prediction unit (10a-20) may include an inter prediction unit (10a-21) and an intra prediction unit (10a-22). The residual processing unit (10a-30) may include a transformer (10a-32), a quantizer (10a-33), a dequantizer (10a-34), and an inverse transformer (10a-35). The residual processing unit (10a-30) may further include a subtractor (10a-31). The addition unit (10a-50) may be called a reconstructor or a reconstructed block generator. The above-described image segmentation unit (10a-10), prediction unit (10a-20), residual processing unit (10a-30), entropy encoding unit (10a-40), addition unit (10a-50), and filtering unit (10a-60) may be configured by one or more hardware components (e.g., encoder chipset or processor) according to an embodiment. In addition, the memory (10a-70) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (10a-70) as an internal / external component.
[0083] The image segmentation unit (10a-10) can segment an input image (or picture, frame) input to the encoding device (10a) into one or more processing units.
[0084] For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a Quad-tree binary-tree ternary-tree (QTBTTT) structure. For example, one coding unit may be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary structure may be applied later. Alternatively, the binary-tree structure may be applied first. The coding procedure according to the present document may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0085] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0086] The subtraction unit (10a-31) can subtract the prediction signal (predicted block, prediction samples, or prediction sample array) output from the prediction unit (10a-20) from the input image signal (original block, original samples, or original sample array) to generate a residual signal (residual block, residual samples, or residual sample array), and the generated residual signal is transmitted to the conversion unit (10a-32). The prediction unit (10a-20) can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block.
[0087] The prediction unit (10a-20) can determine whether intra-prediction or inter-prediction is applied to the current block or CU unit. As described later in the description of each prediction mode, the prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit (10a-40). The prediction-related information can be encoded by the entropy encoding unit (10a-40) and output in the form of a bitstream.
[0088] The intra prediction unit (10a-22) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode.
[0089] In intra prediction, prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes and planar modes. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the granularity of the prediction direction.
[0090] However, this is only an example; depending on the settings, a greater or lesser number of directional prediction modes may be used. The intra prediction unit (10a-22) may also determine the prediction mode to be applied to the current block by utilizing the prediction mode applied to the surrounding blocks.
[0091] The inter prediction unit (10a-21) can derive a predicted block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit (10a-21) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (10a-21) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0092] The prediction unit (10a-20) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction to predict a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) to predict a block. The intra block copy can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.
[0093] The prediction signal generated through the inter prediction unit (10a-21) and / or the intra prediction unit (10a-22) can be used to generate a reconstructed signal or a residual signal. The transform unit (10a-32) can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transform process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size.
[0094] The quantization unit (10a-33) quantizes the transform coefficients and transmits them to the entropy encoding unit (10a-40), and the entropy encoding unit (10a-40) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information.
[0095] The quantization unit (10a-33) can rearrange the quantized transform coefficients in the form of a block into a one-dimensional vector based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of the one-dimensional vector. The entropy encoding unit (10a-40) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc.
[0096] The entropy encoding unit (10a-40) may encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted through a network or may be stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (10a-40) may be configured as an internal / external element of the encoding device (10a) by a transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal, or the transmitting unit may be included in the entropy encoding unit (10a-40).
[0097] The quantized transform coefficients output from the quantization unit (10a-33) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (10a-34) and the inverse transform unit (10a-35), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (10a-50) can add the reconstructed residual signal to the prediction signal output from the prediction unit (10a-20), thereby generating a reconstructed signal (reconstructed picture, reconstructed block, reconstructed samples, or reconstructed sample array). When there is no residual for the target block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0098] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0099] The filtering unit (10a-60) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (10a-60) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (10a-70), specifically, in the DPB of the memory (10a-70). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), an adaptive loop filter, a bilateral filter, etc. The filtering unit (10a-60) can generate various information regarding filtering and transmit the information to the entropy encoding unit (10a-90), as described below in the description of each filtering method. The information regarding filtering may be encoded by the entropy encoding unit (10a-90) and output in the form of a bitstream.
[0100] The modified restored picture transmitted to the memory (10a-70) can be used as a reference picture in the inter prediction unit (10a-80). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (10a) and the decoding device, and can also improve encoding efficiency.
[0101] The DPB of the memory (10a-70) can store the modified restored picture to be used as a reference picture in the inter prediction unit (10a-21). The memory (10a-70) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transmitted to the inter prediction unit (10a-21) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (10a-70) can store restored samples of restored blocks in the current picture and transmit them to the intra prediction unit (10a-22).
[0102] Figure 3 is a drawing schematically illustrating the configuration of a video / image decoding device.
[0103] Referring to FIG. 3, the decoding device (10b) may be configured to include an entropy decoder (10b-10), a residual processor (10b-20), a predictor (10b-30), an adder (10b-40), a filter (10b-50), and a memory (10b-60). The predictor (10b-30) may include an inter-prediction unit (10b-31) and an intra-prediction unit (10b-32). The residual processor (10b-20) may include a dequantizer (10b-21) and an inverse transformer (10b-21). The entropy decoding unit (10b-10), residual processing unit (10b-20), prediction unit (10b-30), addition unit (10b-40), and filtering unit (10b-50) described above may be configured by a single hardware component (e.g., decoder chipset or processor) according to an embodiment. In addition, the memory (10b-60) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (10b-60) as an internal / external component.
[0104] When a bitstream including video / image information is input, the decoding device (10b) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (10b) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (10b) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (10b) can be reproduced through a reproduction device.
[0105] The decoding device (10b) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (10b-10). For example, the entropy decoding unit (10b-10) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information.
[0106] The decoding device can further decode the picture based on information about the parameter set and / or the general restriction information. The signaling / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through the decoding procedure. For example, the entropy decoding unit (10b-10) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients for the residual.
[0107] In more detail, the CABAC entropy decoding method receives a bin corresponding to each syntax element in a bitstream, determines a context model using information of the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of symbols / bins decoded in a previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit (10b-10), information regarding prediction is provided to the prediction unit (10b-30), and information regarding the residual on which entropy decoding has been performed by the entropy decoding unit (10b-10), i.e., quantized transform coefficients and related parameter information, can be input to the inverse quantization unit (10b-21).
[0108] In addition, information regarding filtering among the information decoded by the entropy decoding unit (10b-10) may be provided to the filtering unit (10b-50). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (10b), or the receiving unit may be a component of the entropy decoding unit (10b-10). Meanwhile, the decoding device according to the present document may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The above information decoder may include the entropy decoding unit (10b-10), and the sample decoder may include at least one of the inverse quantization unit (10b-21), the inverse transformation unit (10b-22), the prediction unit (10b-30), the addition unit (10b-40), the filtering unit (10b-50), and the memory (10b-60).
[0109] The inverse quantization unit (10b-21) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (10b-21) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (10b-21) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0110] In the inverse transform unit (10b-22), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0111] The prediction unit can perform a prediction for the current block and generate a predicted block including prediction samples for the current block.
[0112] The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information about the prediction output from the entropy decoding unit (10b-10), and can determine a specific intra / inter prediction mode.
[0113] The prediction unit can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction to predict a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) to predict a block. The intra block copy can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.
[0114] The intra prediction unit (10b-32) can predict the current block by referencing samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode.
[0115] In intra prediction, prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit (10b-32) may determine the prediction mode to be applied to the current block by utilizing the prediction modes applied to the surrounding blocks.
[0116] The inter prediction unit (10b-31) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.).
[0117] In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (10b-31) may construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information regarding the prediction may include information indicating the mode of inter prediction for the current block.
[0118] The addition unit (10b-40) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (10b-30). In cases where there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.
[0119] The addition unit (10b-40) may be called a restoration unit or a restoration block generation unit.
[0120] The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, can be output after filtering as described below, or can be used for inter prediction of the next picture.
[0121] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0122] The filtering unit (10b-50) can improve subjective / objective image quality by applying filtering to the restoration signal. For example, the filtering unit (10b-50) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and transmit the modified restoration picture to the memory (60), specifically, the DPB of the memory (10b-60). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0123] The (corrected) reconstructed picture stored in the DPB of the memory (10b-60) can be used as a reference picture in the inter prediction unit (10b-31). The memory (10b-60) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transmitted to the inter prediction unit (10b-31) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (10b-60) can store reconstructed samples of reconstructed blocks within the current picture and transmit them to the intra prediction unit (10b-32).
[0124] In this specification, the embodiments described in the prediction unit (10b-30), the inverse quantization unit (10b-21), the inverse transformation unit (10b-22), and the filtering unit (10b-50) of the decoding device (10b) can be applied to the prediction unit (10a-20), the inverse quantization unit (10a-34), the inverse transformation unit (10a-35), and the filtering unit (10a-60) of the encoding device (10a) in the same manner or correspondingly.
[0125] As described above, prediction is performed to increase compression efficiency when performing video coding. Through this, a predicted block including prediction samples for a current block, which is a coding target block, can be generated. Here, the predicted block includes prediction samples in a spatial domain (or pixel domain). The predicted block is derived identically from an encoding device and a decoding device, and the encoding device can increase video coding efficiency by signaling information (residual information) about the residual between the original block and the predicted block, rather than the original sample value of the original block itself, to a decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by combining the residual block and the predicted block, and can generate a reconstructed picture including the reconstructed blocks.
[0126] The above residual information can be generated through transformation and quantization procedures.
[0127] For example, the encoding device can derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (a residual sample array) included in the residual block to derive transform coefficients, perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and signal related residual information to a decoding device (via a bitstream). Here, the residual information can include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device can perform an inverse quantization / inverse transform procedure based on the residual information to derive residual samples (or residual blocks). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also inverse quantize / inverse transform the quantized transform coefficients to derive a residual block for reference in inter prediction of a subsequent picture, and generate a reconstructed picture based on the residual block.
[0128] <VCM(Video coding for Machines)>
[0129] With the recent advancements in various industries such as surveillance, intelligent transportation, smart cities, intelligent industry, and intelligent content, the amount of image or feature map data consumed by machines is increasing. In contrast, traditional video compression methods currently in use were developed with human vision in mind, and therefore contain unnecessary information, making them inefficient for machine tasks. For example, the resolution of images from the viewer's perspective may be higher than that of images (e.g., feature maps) from the machine's perspective. Therefore, research on video codec technologies that efficiently compress feature maps for machine tasks is needed.
[0130] The Moving Picture Experts Group (MPEG), an international standardization group for multimedia encoding, is discussing Video Coding for Machines (VCM). VCM is an image or feature map encoding technology that targets machine vision, rather than human viewer vision. In this document, feature maps can be referred to as "feature maps," and features can be referred to as "features."
[0131] Figures 4a to 4d are exemplary diagrams showing a VCM encoder and a VCM decoder.
[0132] Referring to FIG. 4a, a VCM encoder (100a) and a VCM decoder (100b) are shown.
[0133] When a VCM encoder (100a) encodes a video and / or a feature map and transmits it as a bitstream, a VCM decoder (100b) can decode and output the bitstream. At this time, the VCM decoder (100b) can output one or more videos and / or feature maps. For example, the VCM decoder (100b) can output a first feature map for machine-based analysis and a first image for user viewing. The first image can have a higher resolution than the first feature map. The first image can be generated based on the first feature map.
[0134] Referring to FIG. 4b, a feature extractor (feature extractor) for extracting a feature map may be connected to the front end of the VCM encoder (100a).
[0135] The VCM encoder (100a) may include a feature encoder.
[0136] The VCM decoder (100b) may include a feature decoder and a video reconstructor. The feature decoder may decode a feature map from a bitstream and output a first feature map for machine-assisted analysis. The video reconstructor may regenerate and output a first video from the bitstream for viewing by a user.
[0137] Referring to Fig. 4c, a feature extractor for extracting a feature map is connected to the front end of the VCM encoder (100). The VCM encoder (100a) may include a feature encoder.
[0138] The VCM decoder (100b) may include a feature decoder. The feature decoder may decode a feature map from a bitstream and output a first feature map for machine-based analysis. That is, the bitstream may be encoded only as a feature map, not as an image. To elaborate, the feature map may be data containing information about features for processing a specific task of a machine based on an image.
[0139] Referring to FIG. 4d, a feature extractor may be connected to the front end of the VCM encoder (100a).
[0140] The VCM encoder (100a) may include a feature converter and a video encoder. The video encoder may be the encoding device (10a) illustrated in FIG. 2.
[0141] The VCM decoder (100b) illustrated in FIG. 4d may include a video decoder and an inverse converter. The video decoder may be the decoding device (10b) illustrated in FIG. 3.
[0142]
[0143] FIG. 5 illustrates a block diagram of an encoding device according to one embodiment of the present disclosure.
[0144] According to one embodiment of the present disclosure, an encoding device (10a) may receive a video as input and perform encoding to generate a bitstream. The bitstream may be generated for the purpose of storing or transmitting the video, and may be utilized for various purposes, such as machine vision or human vision, after being restored. Machine vision may mean that a computer, rather than a human, processes the restored video for purposes such as object detection, semantic segmentation, object tracking, and situation analysis. Human vision may mean that a human uses the restored video for purposes such as viewing or analyzing the situation.
[0145] According to one embodiment, the encoding device (10a) may include a temporal resampling performing unit (510), a spatial resampling performing unit (520), a region of interest-based processing unit (530), and an internal encoding unit (540). The above components are merely logically separated for the purpose of explaining the embodiments, and one processor within the encoding device (10a) may perform all operations of the components, or several physically separated processors within the encoding device (10a) may divide and perform some configurations of the components.
[0146] For convenience of explanation in the present disclosure, the encoding device (10a) may be referred to as an encoder or encoder, and the decoding device (10b) may be referred to as a decoder or decoder.
[0147] According to one embodiment, a temporal resampling performing unit (510) may receive an image, perform frame-by-frame sampling, and output an image in which some frames have been sampled. According to one embodiment, an encoding device (10a) may transmit information used in the temporal resampling process (for example, a temporal sampling rate, etc.) to a decoding device (10b).
[0148] According to one embodiment, a spatial resampling performing unit (520) may receive an image or an image on which temporal resampling has been performed and output an image with a changed spatial resolution for each frame or a series of frames. According to an embodiment, an encoding device (10a) may transmit information used in the spatial resampling process (for example, a spatial sampling rate of each frame or / and a spatial sampling rate of a series of frames, etc.) to a decoding device (10b).
[0149] According to one embodiment, a region of interest-based processing unit (530) may receive an image, or an image on which at least one of temporal resampling and spatial resampling has been performed on part or the entire image, extract a region of interest existing in each frame or a series of frames, and output an image processed based on the extracted region of interest. According to an embodiment, the encoding device (10a) may transmit information used in the region of interest extraction and processing process (for example, region of interest ID information, region of interest size information, packing information, etc.) to the decoding device (10b).
[0150] According to one embodiment, the internal encoder (540) may receive an image, or an image on which at least one of temporal resampling, spatial resampling, or region-of-interest-based processing has been performed on part or all of the image, and perform image encoding to generate a bitstream. According to an embodiment, the internal encoder (540) may use a 2D video encoder (AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, etc.) and may use a 2D video encoder including one or more convolution layers. According to an embodiment, the internal encoder (540) may convert a color space of an input image into a color space such as YUV420 or YUV444, and then perform encoding. According to one embodiment, the color space conversion can be performed by the internal encoding unit (540) through a conversion method defined by an agreement between the encoding device (10a) / decoding device (10b) (internal encoding unit (540) / internal decoding unit (1510)), and the color space conversion information can be omitted from encoding. According to another embodiment, the color space conversion information can be encoded and transmitted to the decoding device (10b).
[0151] According to one embodiment of the present disclosure, an encoding device (10a) can divide an input image into at least one region of interest for each frame.
[0152] The above encoding device (10a) can determine a first region of interest distribution coding section among two or more consecutive frame sections within an input image sequence.
[0153] The encoding device (10a) can select at least one region of interest as a region of interest distribution coding target for each frame within the first region of interest distribution coding section, and distribute the selected region of interest to frames within the first region of interest distribution coding section.
[0154] The above encoding device (10a) can generate a bitstream by encoding the region of interest segmentation information and the input image.
[0155] According to one embodiment, the at least one region of interest is divided into one or more types, and the region of interest segmentation information may include a type for each of the at least one region of interest.
[0156] The encoding device (10a) can move the first type of regions of interest included in the first frame included in the first region of interest distribution coding section to at least a portion within the first region of interest distribution coding section.
[0157] The encoding device (10a) can derive first types of regions of interest for each frame using a first region of interest derivation method, and derive second types of regions of interest for each frame using a second region of interest derivation method. The first region of interest derivation method and the second region of interest derivation method may be different from each other. That is, the first type of region of interest and the second type of region of interest can be distinguished by the region of interest derivation method.
[0158] According to one embodiment, the first region of interest derivation method independently derives a region of interest for each frame, and at least one of the positions or sizes of the derived regions of interest may differ. Furthermore, at least some of the derived regions of interest may include overlapping portions.
[0159] According to one embodiment, the first region of interest derivation method may utilize at least one of a method for detecting motion and deriving a region of interest by comparing the current frame with adjacent frames, a method using a neural network for object detection, or a method using a segmentation lightweight neural network. For example, when combining two or more methods, the region of interest can be derived by weighting the information of the obtained region of interest.
[0160] According to one embodiment, the second region of interest derivation method can derive a region of interest based on at least one of a position or a size specified in a sequence or frame group unit in all frames.
[0161] According to one embodiment, the region of interest segmentation information can be expressed by a method of directly specifying the region of interest unit or by a method of displaying it as a map based on basic units.
[0162] According to one embodiment, the method of directly specifying the region of interest unit may include, for each frame, at least one of the type of shape of each region of interest, a reference point of each region of interest, a width of each region of interest, or a height of each region of interest, to express each region of interest, and the method of displaying it as a map based on the basic unit may include, for the entire frame, a form in which two or more pixels are grouped in a rectangular shape, and each basic unit may be listed so as not to overlap, and each region of interest may be expressed in correspondence to the basic units divided into a specified absolute size or relative size.
[0163] According to one embodiment, the encoding device (10a) may select at least one region of interest as a region of interest distribution coding target for each frame within the first region of interest distribution coding section, and distribute the selected region of interest to frames within the first region of interest distribution coding section. For example, the distribution may be performed by dividing the second type of regions of interest included in the first frame within the first region of interest distribution coding section into one region of all frames within the first region of interest distribution coding section, or by deleting the second type of regions of interest included in each frame within the first region of interest distribution coding section, leaving only one region of interest that does not overlap with each other.
[0164] According to one embodiment, the region of interest segmentation information may include at least one of information on whether the period in which the region of interest distribution coding section appears within the sequence is fixed or period information.
[0165] According to one embodiment, the region of interest segmentation information may include at least one of an index indicating whether the first frame of the input image is a frame within a region of interest distribution coding section, an index indicating which frame the first frame is within the region of interest distribution coding section, an index of the type of the region of interest that is a target of distribution coding, or an index of the region of interest that is a target of distribution coding.
[0166]
[0167] FIG. 6 is a flowchart illustrating an operation of an encoding device for region-of-interest-based image processing according to one embodiment of the present disclosure.
[0168] According to one embodiment, the encoding device (10a) may extract a region of interest (ROI) existing within each frame or a series of frames of an input image or an image that has undergone some or all of temporal / spatial resampling, and output an image processed based on the ROI. According to one embodiment, the encoding device (10a) may transmit information used in the ROI processing process (for example, ROI ID information, ROI size information, packing information, etc.) to the decoding device (10b). According to one embodiment, the decoding device (10b) may derive region information from the information used in the ROI processing process during the encoding process, and perform frame-to-frame rearrangement of some ROIs within a frame or a group of frames.
[0169] In operation 610, the region-of-interest-based processing unit (530) according to one embodiment can divide a frame into multiple types of regions (e.g., a first region of interest, a second region of interest, ..., an Nth region of interest, a non-interest region). In one embodiment, the types by which the frame is divided can be 0 or more. For example, the region-of-interest-based processing unit (530) can divide one frame into four types (a first region of interest, a second region of interest, a third region of interest, a non-interest region), thereby dividing the frame into a total of seven regions (three first regions of interest, one second region of interest, two third regions of interest, and a non-interest region).
[0170] In one embodiment, in the “nth region of interest,” n may represent the intended use of the reconstructed region outside the decoder, for example, at the receiver (the device that uses the reconstructed image). For example, when n=1, this may mean that the reconstructed region is intended for machine vision purposes, and when n=2, this may mean that the reconstructed region is intended for human vision purposes.
[0171] In one embodiment, in the “nth region of interest”, n may represent an attention level depending on the content type of the sequence. For example, when n=1, it may mean an area requiring the most attention (attention level 1) depending on the content type of the sequence, and when n=2, it may mean an area requiring the second most attention (attention level 2). For example, when the sequence is an image from the viewpoint of a moving object, an area including other dynamic objects in the vicinity may be designated as a first region of interest (n=1), an area including static objects in the vicinity may be designated as a second region of interest (n=2), and an area including a road or the sky may be designated as a non-region of interest.
[0172] The region of interest-based processing unit (530) can derive regions of interest with fixed positions and sizes for each region of interest type in units of sequences or frame groups. Alternatively, the region of interest-based processing unit (530) can derive regions of interest with variable areas through analysis on a frame-by-frame basis. For example, the region of interest-based processing unit (530) can derive regions of new positions / sizes for each frame through frame analysis for the first region of interest. The region of interest-based processing unit (530) can designate regions of interest with fixed positions and sizes for the second region of interest in units of sequences or frame groups.
[0173] The region of interest-based processing unit (530) according to one embodiment can derive a region of interest with a fixed location and size for each sequence / frame unit, or derive a region of interest with a different location and size for each frame, for any type of region of interest.
[0174] First, a method for deriving regions of interest with fixed positions and sizes for a frame sequence or frame group for an arbitrary type of region of interest can be as follows.
[0175] The region-of-interest-based processing unit (530) can set the position and size of the region-of-interest within the frame to fixed values for the encoder / decoder. Alternatively, the region-of-interest-based processing unit (530) can divide the frame in a fixed manner and designate all or part of the divided region as a region-of-interest of any type.
[0176] Next, a method for deriving regions of interest whose positions and sizes are variable on a frame-by-frame basis for any type of region of interest can be as follows.
[0177] The region of interest-based processing unit (530) may derive regions of interest by using at least one of one or more methods for deriving regions of interest through image analysis (e.g., a motion detection-based method by comparing the current frame with adjacent frames, a method using an object detection lightweight neural network, a method using a segmentation lightweight neural network, etc.). The regions of interest derived in this way may have different locations or sizes. For example, the shape of each derived region of interest may be a rectangle or an arbitrary shape. The regions of interest may overlap each other. Regions other than those derived as regions of interest (e.g., the 0th region of interest to the nth region of interest) may be regarded as regions of no interest.
[0178] According to one embodiment, the region-of-interest-based processing unit (530) may determine a different method for deriving regions of interest depending on the type of region of interest for a sequence / group of frames / frames. For example, the region-of-interest-based processing unit (530) may derive variable regions of interest for a first region of interest, and derive regions of interest with fixed positions / sizes for a second region of interest. An example of segmenting regions of interest using various methods according to one embodiment will be described with reference to FIG. 7.
[0179] FIG. 7 is an example of regions of interest segmented in various ways according to one embodiment of the present disclosure.
[0180] According to one embodiment, the region of interest-based processing unit (530) can extract variable regions of interest for each frame, as shown in (a) of FIG. 7. The region of interest-based processing unit (530) can extract regions of interest variably for each type of region of interest, as shown in (b) of FIG. 7, or can extract regions of interest corresponding to fixed sizes and positions. Alternatively, the region of interest-based processing unit (530) can extract only one type of region of interest predefined by the encoder / decoder, with variable sizes, positions, and numbers.
[0181] Referring to (a) of Fig. 7, the results of independently extracting variable regions of interest for each frame (T), frame (T+1), and frame (T+2) in time order are exemplified. In frame (T), one first region of interest, one second region of interest, and one third region of interest are derived, and the remaining regions can be classified as regions of no interest. In frame (T+1), two first regions of interest, two second regions of interest, and the remaining regions can be classified as regions of no interest. In frame (T+2), three first regions of interest, one second region of interest, and the remaining regions can be classified as regions of no interest. The sizes and positions of each region of interest in each frame may be different.
[0182] Referring to (b) of FIG. 7, for the first region of interest, independently variable regions of interest are extracted in time order for each frame (T), frame (T+1), and frame (T+2), and for the second region of interest, the results of extracting regions of interest with fixed positions and sizes in all frames are exemplified. For the second region of interest, the entire frame can be divided into four to derive regions of interest of the same size in all frames. For the first region of interest, one first region of interest can be derived for frame (T), two first regions of interest can be derived for frame (T+1), and three first regions of interest can be derived for frame (T+2). In the case of (b), there is no region of non-interest.
[0183] Referring to (c) of Fig. 7, the type of region of interest to be derived is determined (limited) as one in the encoder or encoder / decoder, and the result of independently extracting the region of interest for each frame is exemplified. Regions of interest corresponding to the first region of interest can be derived for each frame (T), frame (T+1), and frame (T+2) in time order. One first region of interest can be derived for frame (T), two first regions of interest can be derived for frame (T+1), and three first regions of interest can be derived for frame (T+2). In each frame, the remaining regions except for the first regions of interest can be classified as regions of no interest.
[0184] According to one embodiment, the region of interest-based processing unit (530) can extract variable regions of interest for each frame, as in (a). The region of interest-based processing unit (530) can extract variable regions of interest for each type, as in (b), or extract regions of interest corresponding to fixed sizes and positions. The region of interest-based processing unit (530) can extract only one type of region of interest predefined by the encoder / decoder, with variable sizes, positions, and numbers.
[0185] In operation 620, the region-of-interest-based processing unit (530) according to one embodiment can derive and transmit information dividing a frame into multiple types of regions. The region-dividing information can be expressed by 1) directly specifying each region-of-interest unit or 2) using a basic unit-by-unit map method.
[0186] According to one embodiment, the region-of-interest-based processing unit (530) may express each region within a frame with information such as the type of shape, reference point of the region, width, and height, according to 1) a direct specification method of the region-of-interest unit. For example, the type of shape for the first region-of-interest may be a rectangle, the reference point of the first region-of-interest may be the upper left coordinate of the rectangle, and the width and height may be the width and height of the rectangle. Information such as the coordinates, width, and height of the shape may include and be set to include a minimum size or a certain margin that may include the first region-of-interest. An region-of-interest index may be assigned to each region-of-interest and transmitted to the decoding device (10b).
[0187] A base unit (or, may be referred to as a base unit) for expressing the coordinates, width, height, etc. of the region of interest may be specified. Information (e.g., size) of the base unit may be transmitted to the decoding device (10b) in units of sequence, frame, or frame group, or may be identically specified to the sub / decoding unit.
[0188] In one embodiment, a basic unit may be a rectangular shape grouping two or more pixels. Alternatively, a basic unit may be a square shape grouping two or more pixels. Each basic unit may be arranged in a non-overlapping raster scan order starting from the upper left corner of the frame.
[0189] In one embodiment, the size of the basic unit can be specified to the encoder / decoder as the same value, and the fixed value can be scaled based on the frame size.
[0190] Alternatively, the area of interest-based processing unit (530) can express and transmit in one of the following ways.
[0191] In one embodiment, the region-of-interest-based processing unit (530) can directly transmit information for expressing the shape and position of each region of interest according to a direct specification method for each region of interest. Detailed methods for the direct specification method for each region of interest may include a 'template-based method', a 'free-form specification method', etc., and an index indicating which of the multiple methods was used may be transmitted in units of sequence, frame, or frame group. In other words, the transmission unit of the index of the direct specification method for each region of interest may be different from the transmission unit of the information for expressing the shape and position of the region of interest.
[0192] FIGS. 8a, 8b and 8c are examples of a template-based direct region of interest specification method according to one embodiment of the present disclosure.
[0193] FIG. 9 is an example of a direct region of interest specification method based on a free-form specification method according to an embodiment of the present disclosure.
[0194] According to one embodiment, the region of interest-based processing unit (530) may express the region of interest by translating and scaling a specific form of template defined in the decoder / decoder when the expression method of the region of interest unit is a 'template-based method'. Referring to FIGS. 8a to 8c, the template may be defined in the form of a 'vertex expression method' or an 'occupancy map expression method'. T types of templates may be defined, and each template may be defined in a different method ('vertex expression method' or 'occupancy map expression method'). Here, T may be a natural number greater than or equal to 1, such as 1, 2, 3, or 4.
[0195] The 'vertex representation method' can define the coordinates of up to N vertices in a single template. The vertex coordinates defined for each template can be relative values based on (0, 0) of the picture. The vertices defined in a single template are sequentially connected between adjacent two vertices, and the first and last vertices are connected, so that each connection forms one edge of the template shape. The area inside the closed boundary defined by multiple edges can be the template area.
[0196] The 'Occupancy Map Representation Method' defines an occupancy map in the form of a two-dimensional matrix, and the internal and external areas of the template may be expressed with different values in the occupancy map. Information derived from one element in the occupancy map may correspond to a pixel of the frame or a basic unit of the frame. The upper left element of the occupancy map may correspond to the upper left coordinate of the frame.
[0197] You can select a template to use from among T types of templates and transmit the template index. The selected template can be scaled and then translated to finally express the region of interest. The scaling value of the template and the translation values along the x-axis and y-axis can be transmitted. If the template is a 'vertex expression method', you can define and store a formula for calculating the changed coordinates for each scale value and a table for the coordinate values changed by the formula. For example, Table 1 exemplifies a formula for calculating the changed coordinates of each vertex for each scale value, and Table 2 exemplifies the changed coordinate values of each vertex for each scale value.
[0198] lP0P1. . .P N-1 l unit ( )( ). . .( )l m ( , )( , ). . .( , )
[0199] l in Table 1 unit represents the template scale, and l m is the vertex of the region of interest scale (P0 to P N-1 ) can be used to express the equation for changing the coordinates.
[0200] lP0P1. . .P N-1 l0( )( ). . .( )l1( )( ). . .( ). . .. . .. . .. . .. .l M-1 ( )( ). . .( )
[0201] Table 2 can represent the scale values of M regions of interest changed according to the formula in Table 1. Figures 8a to 8c can represent examples of regions of interest expressed using different types of templates, respectively. Figures 8a and 8b can represent examples of regions of interest expressed based on templates defined using a vertex representation method, and Figure 8c can represent an example of regions of interest expressed based on templates defined using an occupancy map representation method.
[0202] Referring to FIG. 9, in an embodiment, when the method of expressing a region of interest unit is a 'free-form specification method', the shape of the region of interest may not be defined in the encoder / decoder, and information directly representing the shape may be transmitted. For example, an index indicating the type of the region of interest shape may be transmitted. The type of the region of interest shape may be a rectangle or an n-gon that is not a rectangle. When the shape of the region of interest is a polygon that is not a rectangle, the coordinates of each vertex constituting the polygon may be transmitted. In the order in which each vertex is transmitted, adjacent two vertices are connected, and the first and last vertices are connected, so that each connection forms one side of the template shape. An area within a closed boundary defined by multiple sides may be a template area, and FIG. 9 exemplifies vertex information for indicating an area within a closed boundary for a region of interest, such as a first frame, a second frame, and a third frame.
[0203] According to one embodiment, the region of interest-based processing unit (530) can express the region of interest according to the following method when the shape of the region of interest is a rectangle.
[0204] According to the first method, the region of interest-based processing unit (530) can express the width or height when the basic unit is a rectangular shape. When transmitting the width or height, the region of interest-based processing unit (530) can i) transmit a value in pixel units, ii) transmit an index that is one-to-one mapped to the value in pixel units, iii) transmit a scale value obtained by taking a logarithm (e.g., log2), and iv) transmit a value divided by a multiple of 2. The above method is exemplary, and can be pre-specified in various ways between the encoder and decoder.
[0205] According to the second method, the region of interest-based processing unit (530) can transmit the number of basic regions in the horizontal (horizontal) and vertical (vertical) directions. The width of the basic region of interest can be the value obtained by dividing the frame width by the number of vertical regions, and the height of the basic region of interest can be the value obtained by dividing the frame height by the number of horizontal regions.
[0206] According to the third method, the region-of-interest-based processing unit (530) can transmit an index containing the width and height of the basic region of interest. For example, index 0 can be specified between the encoder and decoder, indicating that the size of the basic region is 4x4, and index 1 can be specified that the size of the basic region is 8x8.
[0207] According to one embodiment, the region-of-interest-based processing unit (530) can transmit the type of the region of interest for each region of interest within the frame. The region-of-interest-based processing unit (530) can define between the encoder and decoder that a specific region of interest (e.g., the first region of interest) is present without transmitting the type of the region of interest. One embodiment of a method for directly specifying region-of-interest segmentation information for 1) rectangular regions of interest in units of regions of interest is described with reference to FIG. 10.
[0208] FIG. 10 is an example of a method for directly specifying region segmentation information by region of interest unit according to one embodiment of the present disclosure.
[0209] According to one embodiment, the region of interest-based processing unit (530) can transmit region of interest segmentation information to the decoding device (10b) by directly specifying it as in (a) and (b) of FIG. 10. The region of interest-based processing unit (530) can transmit ROI_location_x and ROI_location_y, which represent the upper left position of the region of interest in units of basic units, and can transmit ROI_width and ROI_height syntaxes, which represent the width and height of the region of interest.
[0210] Referring to (a) of Fig. 10, this is an example of a region of interest (ROI) in which the size of a basic unit of 16x16 pixels is expressed and transmitted as a mapped index. In (a) of Fig. 10, the basic unit of the region of interest (base_unit_of_ROI) is 2, and the location of the first region of interest is ROI_location_x=1, ROI_location_y=1, ROI_width=2, and ROI_height=3.
[0211] Referring to (b) of Fig. 10, for a region of interest (ROI), the width and height of the basic unit are expressed as the width of the picture / the number of basic units in the horizontal direction and the height of the picture / the number of basic units in the vertical direction, and this is an example of how the number of basic units in the horizontal direction and the number of basic units in the vertical direction are transmitted. In (b) of Fig. 10, the width of the basic unit is expressed as the width of the picture / the number of basic units in the horizontal direction ( ) and the height of the basic unit is the height of the picture / the number of basic units in the vertical direction ( ) in Fig. 10 (b), the location of the first region of interest is ROI_location_x=3, ROI_location_y=4, and the width and height of the first region of interest are ROI_width=4, ROI_height=3.
[0212] According to one embodiment, the region of interest-based processing unit (530) can express and transmit region segmentation information according to 2) the basic unit map method.
[0213] According to one embodiment, the region-of-interest-based processing unit (530) may divide the current frame into base units and assign an index of the region-of-interest with the largest area within each base unit to the corresponding base unit. The region-of-interest-based processing unit (530) may transmit the type of the region-of-interest for each region-of-interest. Alternatively, the type of the region-of-interest assigned an index may be transmitted for each base unit.
[0214] According to the basic unit unit map method, the region of interest-based processing unit (530) according to one embodiment can designate and transmit a region of interest index in basic unit units within a frame. The basic unit unit may be the same as the basic unit unit that directly specifies the region of interest. For example, the basic unit unit of FIG. 10 and the basic unit unit of FIG. 11 may be defined in the same manner. An embodiment of displaying region of interest segmentation information as a basic unit unit map will be described with reference to FIG. 11.
[0215] FIG. 11 is an example of a method for displaying region of interest segmentation information as a basic unit map according to one embodiment of the present disclosure.
[0216] According to the method for specifying the basic unit unit, the region of interest-based processing unit (530) according to one embodiment can designate and transmit the region of interest index in units of basic units within a frame. The basic unit unit may be the same as the basic unit of the region of interest unit of FIG. 10. The region of interest-based processing unit (530) can divide the current frame into basic unit units, and assign the index of the region of interest with the largest area of the region of interest within each basic unit unit to the corresponding basic unit unit. The region of interest-based processing unit (530) can transmit the type of the region of interest for each region of interest.
[0217] Referring to FIG. 11, the region of interest-based processing unit (530) can divide the frame into basic units of axb pixel size, and the entire frame can include a first region of interest #1, a first region of interest #2, a second region of interest #1, and a third region of interest #1. Here, the first, second, third… can indicate the type of the region of interest, and #1, #2… can indicate the numbering of the region of interest of the corresponding type.
[0218] In operation 630, the region of interest-based processing unit (530) according to one embodiment can set a region of interest distribution coding section.
[0219] According to one embodiment, the region-of-interest-based processing unit (530) may set two or more consecutive frame sections within a video sequence as a region-of-interest distribution coding section. The region-of-interest-based processing unit (530) may encode a portion of the region of the n-th region of interest by assigning the region to each frame within the region-of-interest distribution coding section for a specific n-th region of interest.
[0220] The region of interest-based processing unit (530) can set the region of interest distribution coding intervals to appear at fixed intervals within the sequence, or at irregular intervals. If the region of interest distribution coding intervals have fixed intervals, the region of interest-based processing unit (530) can transmit the value of the interval at the sequence level. The region of interest distribution intervals can appear at regular or irregular intervals. An example of the region of interest distribution intervals is described with reference to FIG. 12.
[0221] FIG. 12 is an example of a region of interest distribution coding section according to one embodiment of the present disclosure.
[0222] Referring to (a) of Figure 12, the distribution section of the area of interest may appear at regular intervals.
[0223] Referring to (b) of Figure 12, the distribution section of the area of interest may appear at irregular intervals.
[0224] In one embodiment, the number of consecutive frames included in each region-of-interest distribution coding interval may be a fixed value for the decoder / decoder, or may be adaptively calculated for each region-of-interest distribution coding interval. For example, it may be specified as the number of n-th regions-of-interest within the first frame of the region-of-interest distribution coding interval.
[0225] The region of interest-based processing unit (530) can transmit syntax indicating whether there is at least one region of interest distribution coding section in the current sequence in the sequence level parameter set. If there is a region of interest distribution coding section in the current sequence, the region of interest-based processing unit (530) can transmit syntax indicating whether the appearance period of the region of interest distribution coding section is fixed, and if it is fixed, a fixed period value can be transmitted together. Alternatively, the fixed period value can be predefined identically to the encoder / decoder without being transmitted.
[0226] In operation 640, the region of interest-based processing unit (530) according to one embodiment can set a region of interest distribution coding target.
[0227] According to one embodiment, a region of interest-based processing unit (530) may set a portion of the nth region of interest of each frame within the region of interest distribution coding section as a region of interest distribution coding target. The execution process may be as follows.
[0228] According to the first distribution coding target setting method, an n-th region of interest may exist in each frame within the region of interest distribution coding section. The region of interest-based processing unit (530) may select one n-th region of interest in each frame while traversing the frames in ascending order of frame index, and may select the n-th region of interest that is closest to the n-th region of interest assigned to previous frames and has the smallest overlap with one or more n-th regions of interest in each frame in the current frame. The n-th regions of interest excluding the n-th region of interest selected in the current frame may be changed to non-regions of interest. An embodiment according to the first distribution coding target setting method is described later in FIG. 13.
[0229] According to the second distribution coding target setting method, the n-th region of interest in the first frame within the region of interest distribution coding section can be copied to the following frame. In a more detailed execution process, all information on the n-th region of interest in the following frames can be deleted first. There may be A n-th regions of interest in the first frame and the number of following frames may be A-1, and the a-th n-th region of interest among the A can be copied to the frame after a based on the current frame. An embodiment according to the second distribution coding target setting method is described below in FIG. 13.
[0230] The region of interest-based processing unit (530) can transmit syntax related to the following region of interest distribution coding section in the frame level parameter set. The region of interest-based processing unit (530) can transmit at least one of ① a flag indicating whether the corresponding frame is a frame within the region of interest distribution coding section, ② an index indicating which frame the corresponding frame is within the region of interest distribution coding section, ③ an index of the type of region of interest that is a target of distribution coding, or ④ an index of the region of interest that is a target of distribution coding.
[0231] Within the region of interest distribution coding interval, the region of interest information of each frame can be transmitted in the same manner as frames not included in the region of interest distribution coding interval. Alternatively, the region of interest information from the first region of interest to the n-1th region of interest can be transmitted for each frame, and the information of the nth region of interest, which is the target of region of interest distribution coding, can be not transmitted, and the encoder can derive it using a predetermined method.
[0232]
[0233] FIGS. 13a and 13b are examples of setting a distribution coding target according to a first distribution coding target setting method according to one embodiment of the present disclosure.
[0234] According to one embodiment, the region of interest-based processing unit (530) can set the region of interest distribution coding target according to the first distribution coding target setting method.
[0235] Referring to Fig. 13a, a region of interest can be represented in four frames (frame (T), frame (T+1), frame (T+2), frame (T+3)) within an arbitrary region of interest distribution coding interval.
[0236] Referring to Fig. 13b, an example of the result of setting the target of the region of interest distribution coding can be shown. This may be the case where the target of the region of interest distribution coding is set to the second region of interest, and the result of deleting the remaining second regions of interest, leaving only the second regions of interest having the same index as the frame sequence among the second regions of interest of each frame of Fig. 13(a), can be shown as in Fig. 13(b).
[0237]
[0238] FIGS. 14a to 14c are examples of setting a distribution coding target according to a second distribution coding target setting method according to one embodiment of the present disclosure.
[0239] According to one embodiment, the region of interest-based processing unit (530) may set a region of interest distribution coding target according to a second distribution coding target setting method. The target of region of interest distribution coding may be, for example, a second region of interest.
[0240] Referring to Fig. 14a, a region of interest can be represented in four frames (frame (T), frame (T+1), frame (T+2), frame (T+3)) within an arbitrary region of interest distribution coding interval.
[0241] Referring to Figure 14b, the result of deleting all second regions of interest from the remaining frames except for the first frame within the region of interest distribution coding section can be shown.
[0242] Referring to Figure 14c, the results of moving some of the second regions of interest of the first frame to a subsequent frame can be shown. This can be the result of setting the region of interest distribution coding target.
[0243]
[0244] FIG. 15 illustrates a block diagram of a decoding device according to one embodiment of the present disclosure.
[0245] According to one embodiment of the present disclosure, a decoding device (10b) can receive a bitstream as input, perform decoding, and output a restored video.
[0246] According to one embodiment, the decoding device (10b) may include an internal decoding unit (15101510), a region-of-interest-based reconstruction unit (15201520), a spatial reconstruction performing unit (15301530), a temporal reconstruction performing unit (15401540), and a post-processing filter performing unit (15501550). The above components are merely logically separated for the purpose of explaining embodiments, and one processor within the decoding device (10b) may perform all operations of the components, or multiple physically separated processors within the decoding device (10b) may divide and perform some configurations of the components. The execution order between the components of FIG. 15 is exemplary, and the execution order may be changed depending on the embodiment. In addition, the execution order between the components of FIG. 15 may be performed in the reverse order of the order in which each corresponding process is performed during the encoding of the image.
[0247] According to one embodiment, the internal decoding unit (15101510) may receive a bitstream and perform image decoding to generate a restored image. According to an embodiment, the internal decoding unit (15101510) may use a 2D video decoder (AVC / H.264, HEVC / H.265, VVC / H.266, AV1, VP9, etc.) and may use a 2D video decoder including one or more convolution layers. According to an embodiment, when the color space of the restored image is not RGB444 but one of YUV420, YUV444, etc., the image decoding process may be performed after implicitly or / and explicitly converting to the RGB444 space. According to an embodiment, the color space of the restored image may be implicitly or / and explicitly converted to another color space, and then the image decoding process may be performed.
[0248] According to one embodiment, a region-of-interest-based restoration unit (15201520) can restore a region-of-interest-based processed image using a decoded image and region-of-interest-based processing information (for example, region-of-interest ID information, region-of-interest size information, packing information, etc.) transmitted from an encoder.
[0249] According to one embodiment, the spatial restoration performing unit (15301530) can obtain an image on which spatial restoration has been performed by using information used in the spatial resampling process transmitted from the decoded image or the decoded and region-of-interest-based restoration image and the encoder (for example, the spatial sampling rate of each frame and / or the spatial sampling rate of a series of frames, etc.).
[0250] According to one embodiment, the temporal restoration performing unit (1540) can obtain an image on which temporal restoration has been performed by using information (for example, a temporal sampling rate, etc.) used in the temporal resampling process transmitted from the encoder and an image on which part or all of the decoded image or region-of-interest-based restoration / spatial restoration process has been performed.
[0251] According to one embodiment, the post-processing filter performing unit (1550) may perform filtering on a decoded image or an image on which part or all of the region-of-interest-based restoration / spatial restoration / temporal restoration process has been performed. At this time, depending on the embodiment, a fixed filter may be used, or a plurality of filters may be defined by agreement between the encoder and decoder, and filter information may be received from the encoder to perform filtering.
[0252] According to one embodiment of the present disclosure, a decoding device (10b) can decode a bitstream to generate region of interest segmentation information and a decoded image.
[0253] The above decoding device (10b) can rearrange at least one region of interest in each frame of the decoded image based on the region of interest segmentation information.
[0254] According to one embodiment, the region of interest segmentation information may include at least one of an index indicating whether the first frame of the input image is a frame within a region of interest distribution coding section, an index indicating which frame the first frame is within the region of interest distribution coding section, an index of the type of the region of interest that is a target of distribution coding, or an index of the region of interest that is a target of distribution coding.
[0255] According to one embodiment, the at least one region of interest is divided into one or more types, and the region of interest segmentation information may include a type for each of the at least one region of interest.
[0256] According to one embodiment, the operation of repositioning at least one region of interest in each frame of the decoded image based on the region of interest segmentation information may include the following operations.
[0257] The decoding device (10b) may identify, based on the region of interest segmentation information, whether a region of interest distribution section exists in the decoded image at a sequence level or a frame level, and, in response to identifying that the region of interest distribution section exists, search for a first region of interest distribution section in the decoded image using period information included in the region of interest segmentation information. According to one embodiment, the first region of interest distribution section may include two or more consecutive frames.
[0258] The above decoding device (10b) can identify a target of region of interest distribution coding within the first region of interest distribution section based on the region of interest division information.
[0259] The above decoding device (10b) can relocate the identified region of interest distribution coding target distributed within the first region of interest distribution section to the original frame position based on the region of interest division information.
[0260] The above decoding device (10b) can copy the identified region of interest distribution coding target from all frames within the first region of interest distribution section to the corresponding position of the first frame within the first region of interest distribution section.
[0261] The above decoding device (10b) can perform filtering on pixels adjacent to the boundary of the region of interest copied from another frame in the first frame within the first region of interest distribution section.
[0262]
[0263] FIG. 16 is a flowchart illustrating an operation of a decoding device to restore a region-of-interest-based image according to one embodiment of the present disclosure.
[0264] According to one embodiment, the region of interest-based restoration unit (1520) of the decoding device (10b) can restore information related to the region of interest of each frame and rearrange the regions of interest within the region of interest distribution section. The information related to the region of interest can include information related to the division of the region of interest transmitted on a frame-by-frame basis, region of interest distribution section information, etc. The step of restoring the region of interest information can be specified in operation 1610. The step of rearranging the regions of interest within the region of interest distribution section can be specified in operations 1620 to 1650.
[0265] In operation 1610, according to an embodiment, a region-of-interest-based restoration unit (1520) can restore region-of-interest information. The region-of-interest-based restoration unit (1520) can restore information related to the region-of-interest in (1) a sequence-level parameter set or (2) a frame-level parameter set. Alternatively, (3) the region-of-interest-based restoration unit (1520) can restore region-of-interest segmentation information within a frame on a frame-by-frame basis in a manner independent of the type of the region-of-interest or in a manner specified according to the type of the region-of-interest.
[0266] (1) Information related to the region of interest in the sequence level parameter set can be restored through the following process.
[0267] The region-of-interest-based reconstruction unit (1520) can parse a syntax (use_ROI) indicating whether to use a region-of-interest coding mode in the current sequence. If the use_ROI syntax is true, the region-of-interest-based reconstruction unit (1520) can perform the region-of-interest coding mode in one or more frames in the current sequence.
[0268] The region-of-interest-based restoration unit (1520) can parse whether or not to use a region-of-interest distribution section (use_ROI_distribution) from the bitstream. If the use_ROI_distribution section is true, it may mean that one or more region-of-interest distribution sections exist within the frame section referencing the corresponding sequence-level parameter set.
[0269] The region of interest-based restoration unit (1520) can parse the syntax (is_ROI_distribution_static_period) that indicates whether the period of appearance of the region of interest distribution section is constant when the use_ROI_distribution syntax is true.
[0270] If the syntax is true, it can mean that the period of appearance of the region of interest distribution interval is constant, and the syntax (ROI_distribution_static_period) that means the period of appearance of the region of interest distribution interval (P) can be parsed. The region of interest distribution interval can appear every P number of frames starting from the first or any j-th frame within the current frame interval.
[0271] (2) Information related to the region of interest in the frame level parameter set can be restored through the following process.
[0272] use_ROI can indicate whether region-of-interest-based coding is performed. Use_ROI_distribution can indicate whether region-of-interest segmentation coding is performed. is_ROI_distribution_static_period can indicate whether the interval between region-of-interest segmentation coding intervals is constant. ROI_distribution_static_period can indicate a constant interval value between region-of-interest segmentation coding intervals.
[0273] A frame-level parameter set may include syntax indicating whether a frame referencing the frame-level parameter set is included in the region-of-interest distribution coding, and if so, related syntax. It may also include syntax related to region-of-interest segmentation information within the frame. Each syntax may or may not be transmitted, and may be transmitted only if certain conditions are met. The specific restoration process is as follows.
[0274] A syntax (is_ROI_used) can be parsed to indicate that there is at least one region of interest in the current frame. If the is_ROI_used syntax is true, it may indicate that there is at least one region of interest in the current frame, and syntaxes for the region of interest can be parsed as follows.
[0275] For each region of interest, a syntax (is_usage_of_ROI_present) may be parsed to indicate whether to parse the syntax indicating the purpose of the region of interest. The purpose of the region of interest may indicate the use for which the reconstructed region of interest will be utilized in a post-decoder procedure. For example, the types of purposes may include the first machine vision, the second machine vision, …, the i-th machine vision, human vision, etc., and each purpose may be mapped to an index greater than or equal to 0. According to the index n mapped to each purpose, it may be named as the n-th region of interest, and each n-th region of interest may be one or more in one frame.
[0276] In one embodiment, in the “nth region of interest,” n may represent the intended use of the reconstructed region outside the decoder, for example, at the receiver (the device that uses the reconstructed image). For example, when n=1, this may mean that the reconstructed region is intended for machine vision purposes, and when n=2, this may mean that the reconstructed region is intended for human vision purposes.
[0277] In one embodiment, in the “nth region of interest”, n may represent an attention level depending on the content type of the sequence. For example, when n=1, it may mean an area requiring the most attention (attention level 1) depending on the content type of the sequence, and when n=2, it may mean an area requiring the second most attention (attention level 2). For example, when the sequence is an image from the viewpoint of a moving object, an area including other dynamic objects in the vicinity may be designated as a first region of interest (n=1), an area including static objects in the vicinity may be designated as a second region of interest (n=2), and an area including a road or the sky may be designated as a non-region of interest.
[0278] In one embodiment, if is_usage_of_ROI_present is false or the syntax does not exist, the usage purpose of the region of interest may be considered machine vision.
[0279] Information related to region of interest distribution coding can be restored as follows. The syntax (has_distributed_ROI) indicating whether the current frame is included in the region of interest distribution coding section can be parsed. If the has_distributed_ROI syntax is true, the syntax (frame_idx_in_ROI_distribution_section) indicating which frame the current frame is in the region of interest distribution coding section can be parsed. If the has_distributed_ROI syntax is true, the syntax (distributed_ROI_idx) indicating which region of interest in the current frame is included in the region of interest distribution coding can be parsed. If there are M regions of interest in the current frame, distributed_ROI_idx can have a value in the range of 0 to M-1. Alternatively, the syntax (distributed_ROI_type) indicating which type of region of interest in the current frame is included in the region of interest distribution coding can be parsed. The distributed_ROI_type value can mean n of the nth region of interest, and if there are multiple nth regions of interest in the current frame, all of the multiple regions of interest can be included in the target of region of interest distribution coding. Alternatively, if distributed_ROI_idx and distributed_ROI_type are omitted for transmission, the region of interest type defined in the same way in the encoder / decoder (e.g., n of the nth region of interest) can be considered as the target of region of interest distribution coding. The region of interest corresponding to the distributed_ROI_idx or distributed_ROI_type can be the target to be copied to the first frame within the region of interest division section in the region of interest rearrangement step.
[0280] (3) According to one embodiment, the region of interest-based restoration unit (1520) can restore region of interest segmentation information within a frame on a frame-by-frame basis in a manner independent of the type of region of interest or in a manner specified according to the type of region of interest. The restoration process may be as follows.
[0281] 3-1) The process of restoring the region of interest segmentation information within a frame on a frame-by-frame basis regardless of the type of region of interest can be performed as follows.
[0282] You can parse the syntax (num_ROI) that indicates the number of regions of interest in the current frame and the syntax (ROI_representation) that indicates how the region of interest is represented. If ROI_representation is 0, it may indicate the direct representation method (DIRECT_REPRESENTATION), and if ROI_representation is 1, it may indicate the basic unit map representation method (MAP_REPRESENTATION).
[0283] You can parse the syntax (base_unit_of_ROI) that represents the basic unit of interest. The basic unit can indicate the precision of expressing the region of interest and can be a rectangular shape larger than one pixel. The basic unit can be applied to all ROI_representations. The value parsed in relation to the position or size of the ROI in each ROI_representation can have the unit as base_unit_of_ROI. base_unit_of_ROI can be an index value mapped to the size of the basic unit, and the size of the basic unit can be defined as, for example, 4x4 pixels, 8x8 pixels, 16x16 pixels, etc.
[0284] If ROI_representation is DIRECT_REPRESENTATION, information for deriving the location and shape of each region of interest as many as num_ROI can be parsed.
[0285] If ROI_representation is DIRECT_REPRESENTATION, the index (direct_representation_method) that indicates the detailed method of direct specification of the region of interest unit can be parsed, and the index can mean 'template-based method', 'free-form specification method', etc.
[0286] When the expression method of a region of interest unit is a 'template-based method', the region of interest can be expressed by translating and scaling a template of a specific form defined in the encoder / decoder. The template may be defined in a form of a 'vertex expression method' or an 'occupancy map expression method'. There may be T types of templates defined, and each template may be defined in a different method ('vertex expression method' or 'occupancy map expression method'). When the template expression method is a 'vertex expression method', the template area can be derived using the coordinates of up to N vertices defined in one template. The vertex coordinates defined for each template may be relative values based on (0, 0) of the picture. The vertices defined in one template are sequentially connected between adjacent two vertices, and the first and last vertices are connected, so that each connection forms one edge of the template shape. The area inside the closed boundary defined by multiple edges may be the template area. When the template expression method syntax refers to the 'occupancy map expression method', the template area can be derived using the occupancy map in the form of a two-dimensional matrix that is identically defined in the encoder / decoder. In the occupancy map, the area inside and outside the template can be expressed with different values. Information derived from one element in the occupancy map can correspond to a pixel of the frame or a basic unit of the frame. The upper left element of the occupancy map can correspond to the upper left coordinate of the frame. The template index to be used among T types of templates can be parsed (template_idx). The selected template can be scaled and then translated to finally express the area of interest. The scaling syntax of the template and the translation syntax for each of the x-axis and y-axis (template_scale, template_x_offset, template_y_offset) can be parsed.If the template is a 'vertex expression method', the changed coordinate values of each vertex for each scale value may be defined in the table.
[0287] When the ROI unit is expressed using the 'free-form specification method', the ROI shape may be another shape not defined in the encoder / decoder, and the index (polygon_type) indicating the ROI shape type can be parsed. The ROI shape type can be a rectangle or a non-rectangular polygon. If the ROI shape is a non-rectangular polygon, the number of vertices constituting the polygon (number_of_vertex_in_polygon) and the coordinates of each vertex (polygon_vertex_x, polygon_vertex_y) can be parsed. Each vertex is connected to each other in the order in which it is transmitted, and the first and last vertices are connected, and each connection forms one side of the template shape. The area inside the closed boundary defined by multiple sides can be a template area. If the shape of the region of interest is rectangular, the location (ROI_location_x, ROI_location_y), size (ROI_width, ROI_height), or type of region of interest (ROI_type) can be parsed. The value of the parsed location or size information can be in the base unit (base_unit_of_ROI). For example, if the region of interest is defined in the shape of a rectangle, the location of the region of interest can be the coordinates of the base unit located at the upper left corner within the region of interest, and the size can be a value expressed in the width and height of the rectangle in base unit units. The type of region of interest can mean n of the nth region of interest. The region of interest index can be assigned by increasing from 0 by 1 in the order in which the location and size information of the region of interest are restored first.
[0288] ROI_location_x, ROI_location_y, ROI_width, and ROI_height may represent values at the frame resolution restored by the decoding unit (1510), and in that case, the resized_image_size_width, resized_image_size_height, cropped_to_output_difference_width, and cropped_to_output_difference_height syntaxes may be parsed, and the cropped_image_size_width and cropped_image_size_height syntaxes may not be parsed. resized_image_size_width and resized_image_size_height may represent the frame width and height restored by the decoding unit (1510). cropped_image_size_width and cropped_image_size_height may represent the output frame width and height of the region-of-interest-based resizing (1660). cropped_to_output_difference_width and cropped_to_output_difference_height may indicate the length padded to the width and height of the output frame of interest region-based resizing (1660).
[0289] ROI_location_x, ROI_location_y, ROI_width, ROI_height may represent values at the output frame resolution of region-of-interest-based resizing (1660), in which case the resized_image_size_width, resized_image_size_height syntax may not be parsed, and the cropped_image_size_width, cropped_image_size_height, cropped_to_output_difference_width, cropped_to_output_difference_height syntax may be parsed.
[0290] When ROI_representation is MAP_REPRESENTATION, the base unit (base_unit_of_ROI) within the frame can be traversed in raster scan order and the region of interest index corresponding to each base unit can be parsed. A group of base units with the same region of interest index and continuous in at least one of the four directions, up, down, left, and right, can be defined as a single region of interest. The region of interest type can be parsed for each region of interest.
[0291] 3-2) For frames within the region of interest distribution coding section, position and size information of the region of interest can be restored using a restoration method determined according to the type of region of interest. For example, position and size information from the 1st region of interest to the n-1th region of interest can be parsed for each frame, and can be derived from the nth region of interest using a determined method. num_ROI can mean the number of regions of interest from the 1st region of interest to the n-1th region of interest. In this case, the 1st region of interest to the n-1th region of interest can be expressed using the direct representation method (DIRECT_REPRESENTATION) so that related syntaxes can be parsed. The process of deriving position and size information for the nth region of interest using a determined method can be as follows. After dividing the current frame using a method defined in the same way for the encoder / decoder, the region of interest corresponding to the frame_idx_in_ROI_distribution_section among the divided regions can be designated as the nth region of interest of the current frame. In this case, parsing of the distributed_ROI_idx or distributed_ROI_type syntax of the frame-level parameter set can be omitted. In this case, the n-th region of interest can be expressed using the direct representation method. The method defined identically for the encoder / decoder may be the same method defined for all frames within the region of interest distribution coding section.
[0292] By means of operations 1620 to 1650, the region-of-interest-based restoration unit (1520) can perform inter-frame re-distribution of regions of interest within the region-of-interest distribution coding section. The following process can be performed if the use_ROI_distribution syntax parsed from the sequence-level parameter set is true, and can be omitted if it is false.
[0293] FIG. 17 is a flowchart illustrating a detailed operation of a region of interest-based resizing according to an embodiment of the present disclosure.
[0294] FIG. 18 is an example of a subframe division structure derived according to the resolution of a restored frame according to an embodiment of the present disclosure.
[0295] FIG. 19 is an example of deriving a subframe division structure from a frame resolution before resizing according to an embodiment of the present disclosure.
[0296] FIG. 20 is an example of a subframe unit resizing index according to one embodiment of the present disclosure.
[0297] FIG. 21 is an example of a result of subframe resizing according to one embodiment of the present disclosure.
[0298] According to one embodiment, operation 1620 of the decoding device (10b) can resize the restored frame in the decoding unit (1510) into subframe units derived based on the region of interest according to the following detailed operations.
[0299] In operation 1621, a process of dividing a frame into subframes based on a region of interest may be performed, and may be performed in one of the following ways.
[0300] Method 1: A subframe division structure can be derived from the resolution of the restored frame in the decoding unit (1510). At this time, the upper left coordinates, width, and height (ROI_location_x, ROI_location_y, ROI_width, ROI_height) of the region of interest parsed from the frame level parameter set may be values at the frame resolution before resizing, and the width and height syntax (resized_image_size_width, resized_image_size_height) of the frame before resizing may have been parsed. The frame can be divided vertically at the x-axis coordinate position of the upper left coordinate and the upper right coordinate of all regions of interest in the frame, and the frame can be divided horizontally at the y-axis coordinate position. Fig. 18 can show the result of dividing a frame into subframes through Method 1.
[0301] Method 2: After deriving a subframe division structure from the resolution of the resized frame, which is the output of the current step, it is possible to derive a subframe division structure from the frame resolution before resizing (the resolution of the frame restored by the decoding unit (1510)) by scaling it. At this time, the upper left coordinates, width, and height (ROI_location_x, ROI_location_y, ROI_width, ROI_height) of the region of interest parsed from the frame level parameter set may be values from the frame resolution after resizing. The width / height of the frame after resizing (cropped_image_size_height, cropped_image_size_width) may have been parsed from the frame level parameter set. The frame can be vertically divided at the x-axis coordinate position of the upper left and upper right coordinates of all regions of interest in the frame at the frame resolution after resizing, and the frame can be horizontally divided at the y-axis coordinate position to derive subframes. The resizing index of a subframe can be derived using the method of the subframe resizing index derivation step (1920), and the subframe scaling value can be obtained according to the derived resizing index, and by performing scaling, the subframe division structure at the resolution before resizing can be derived. An example of the execution can be as shown in Fig. 19.
[0302] The subframe resizing index derivation step (1920) can derive horizontal / vertical resizing indices of each subframe. If the subframe derivation step (1910) is performed using Method 2, the subframe resizing indices may also have been derived while performing Method 2, and thus the current step may be omitted. The horizontal / vertical resizing indices of a subframe may be derived using the following methods. First, a resizing index may be assigned to each subframe. If a subframe is included in any region of interest, the subframe may be assigned a resizing index of the region of interest. The resizing index of the region of interest may have been parsed from a picture level parameter set (ROI_scale). If a subframe is not included in any region of interest, the subframe may be assigned a default value. The horizontal / vertical resizing indices of each subframe may be derived with reference to the resizing indices assigned to each subframe. Subframes with the same x-axis range within a frame may have the same horizontal resizing index, and the index may be derived in the following manner. The minimum value of the subframe-unit resizing indices of the subframes with the same x-axis range may be designated as the x-axis range resizing index of the corresponding subframes. Subframes with the same y-axis range within a frame may have the same vertical resizing indices, and the index may be derived in the following manner. The minimum value of the subframe-unit resizing indices of the subframes with the same y-axis range may be designated as the y-axis range resizing index of the corresponding subframes. Fig. 20 may show an example of a result in which a horizontal resizing index (2001) / vertical resizing index (2002) is assigned from a subframe-unit resizing index, and may be a case in which the default index n has a value greater than or equal to 4.
[0303] The subframe resizing performing step (1930) can resize the subframe with a scaling value indicated by the resizing index specified in each of the horizontal and vertical directions of the subframe. Resizing of each subframe can be performed in the horizontal / vertical direction in ascending order of the subframe index. Alternatively, horizontal resizing can be performed on the entire frame and then vertical resizing can be performed (or in the opposite order). Horizontal resizing and vertical resizing can be performed by up / downsampling based on one-dimensional filtering in each direction. Alternatively, up / downsampling can be performed based on two-dimensional filtering to derive the value of the resized pixel position. Fig. 21 can show an example of the result of subframe resizing.
[0304] Referring back to FIG. 6, at operation 1630, the region-of-interest-based reconstruction unit (1520) according to one embodiment may search for a frame section corresponding to a region-of-interest distribution coding section within a video sequence if the use_ROI_distribution syntax parsed from the sequence level parameter set is true.
[0305] If is_ROI_distribution_static_period parsed from the sequence level parameter set is true, it may mean that one or more region of interest distribution coding periods appear at a fixed period in the sequence, and it may be derived that the start frame of the region of interest distribution coding period exists at a frame period of ROI_distribution_static_period parsed from the sequence level parameter set. For example, FIG. 22 is an example of a case where the region of interest distribution coding period is constant according to an embodiment of the present disclosure.
[0306] Referring to Figure 22, an example of a case where the region of interest distribution coding section is constant can be shown.
[0307] It can indicate cases where the syntax values of is_ROI_distribution_static_period=1 and ROI_distribution_static_period=16 are parsed.
[0308] If the is_ROI_distribution_static_period parsed from the sequence level parameter set is false, the has_distributed_ROI values parsed from the frame level parameter set are scanned in frame index order within the sequence, and if the has_distributed_ROI value of any frame is true and the has_distributed_ROI value of the previous frame is false, it can be known that the arbitrary frame is the starting frame of the region of interest distribution coding period.
[0309] Alternatively, regardless of the is_ROI_distribution_static_period value, a continuous frame period for which the has_distributed_ROI value is true from the start frame can be derived as a region of interest distribution coding period.
[0310] In operation 1630, according to one embodiment, the region of interest-based restoration unit (1520) can identify a region of interest that is a target of region of interest distribution coding in each frame within an arbitrary region of interest distribution coding section.
[0311] When the distributed_ROI_idx syntax is parsed from the frame level parameter set of each frame, the region of interest-based restoration unit (1520) can identify a region of interest having a region of interest index equal to distributed_ROI_idx among multiple regions of interest within the frame as a target of region of interest distribution coding.
[0312] When the distributed_ROI_type syntax is parsed from the frame level parameter set of each frame, the region of interest-based restoration unit (1520) can identify a region of interest whose region of interest type is identical to distributed_ROI_type among multiple regions of interest within the frame as a target of region of interest distribution coding.
[0313] In the case where neither the distributed_ROI_idx nor the distributed_ROI_type syntax is parsed in the frame-level parameter set of each frame, the region of interest-based restoration unit (1520) can identify the nth region of interest in each frame as a target of region of interest distribution coding, and n can be specified as the same value for the decoder / decoder. An embodiment of identifying a target of region of interest distribution coding is described with reference to FIG. 23.
[0314] FIG. 23a and FIG. 23b are examples of results of identifying a target of distribution coding of an area of interest according to one embodiment of the present disclosure.
[0315] Referring to FIGS. 23a and 23b, when distributed_ROI_idx or distributed_ROI_type is transmitted, the result of identifying the target of region of interest distribution coding can be displayed using the transmitted syntax value.
[0316] Frame (T) exemplifies a case where the region of interest index (distributed_ROI_idx) is 0 and the region of interest type (distributed_ROI_type) is 2. Frame (T) includes a first region of interest #1 and a second region of interest #0, and the second region of interest #2 can be identified as a target of distributed coding by the transmitted syntax value.
[0317] Frame (T+1) exemplifies a case where the region of interest index (distributed_ROI_idx) is 1 and the region of interest type (distributed_ROI_type) is 2. Frame (T+1) includes the first region of interest #0 and the second region of interest #1, and the second region of interest #1 can be identified as a target of distributed coding by the transmitted syntax value.
[0318] Frame (T+2) exemplifies a case where the region of interest index (distributed_ROI_idx) is 2 and the region of interest type (distributed_ROI_type) is 2. Frame (T+2) includes the first region of interest #0, the first region of interest #1, and the second region of interest #2, and the second region of interest #2 can be identified as a target of distributed coding by the transmitted syntax value.
[0319] Frame (T+3) exemplifies a case where the region of interest index (distributed_ROI_idx) is 0 and the region of interest type (distributed_ROI_type) is 2. Frame (T+3) includes the first region of interest #0, the first region of interest #1, and the second region of interest #2, and the second region of interest #2 can be identified as a target of distributed coding by the transmitted syntax value.
[0320] In operation 1640, according to an embodiment, the region of interest-based restoration unit (1520) may copy the region of interest. Specifically, the region of interest-based restoration unit (1520) may copy the region of interest of each frame identified in the previous operation 1630 to the corresponding position of the first frame within any region of interest distribution coding section. However, if the corresponding position is included in the first region of interest to the n-1th region of interest, the copying process may be skipped for the corresponding position. If there are two or more region of interest pixels to be copied to any pixel of the first frame, the average value of the two or more region of interest values may be copied. Alternatively, the process of copying the region of interest may be performed in order from the last frame to the second frame within the region of interest distribution coding, and for pixel positions where a value has already been copied from the previous frame, the latest value of the corresponding pixel and the average value of the pixel to be copied from the current frame may be updated. An embodiment of copying the region of interest will be described with reference to FIG. 17.
[0321] FIG. 24 is an example of copying the region of interest of each frame identified within the region of interest distribution coding section to the first frame according to one embodiment of the present disclosure.
[0322] Referring to FIG. 24, within an arbitrary region of interest distribution coding interval consisting of four frames (frame (T), frame (T+1), frame (T+2), frame (T+3)), an example is shown of the result of copying a region of interest identified as a target of region of interest distribution coding in frames (frame (T+1), frame (T+2), frame (T+3)) excluding the first frame to the first frame (T). A basic unit indicated by an X notation may indicate an area where two or more regions of interest overlap.
[0323] In operation 1650, according to an embodiment, the region-of-interest-based restoration unit (1520) may filter the boundary of the region-of-interest. Specifically, in operation 1640, the region-of-interest-based restoration unit (1520) may perform filtering on pixels adjacent to the boundary of the copied region-of-interest. The type of filtering may be, for example, low-pass filtering, and a weighted sum may be performed on the values of surrounding pixels of the pixel to filter the pixel. The weight of the weighted sum may be determined by a specified low-pass filter coefficient. The filtering target is described with reference to FIG. 25.
[0324] FIG. 25 is an example showing a filtering target according to one embodiment of the present disclosure.
[0325] Fig. 25 may represent a state in which region of interest copying is performed as the first frame within an arbitrary region of interest distribution coding section. This may be identical to the frame (T) in which copying is completed in Fig. 24.
[0326] A basic unit marked with the O notation may represent a boundary of a region of interest and a target of filtering. Filtering may be performed on pixels adjacent to the boundary of two adjacent basic units marked with the O notation or located d pixels away from the boundary.
[0327] A basic unit marked with an X may represent an area where two or more areas of interest overlap.
[0328]
[0329] Examples of syntax and semantics related to area of interest distribution coding
[0330] The encoding device (10a) can transmit distribution coding information for a region of interest to the decoding device (10b) while encoding an image. The following examples of syntax and semantics can be shared by the encoding device (10a) and the decoding device (10b).
[0331] Table 3 is a syntax example for a sequence parameter set.
[0332] Sequence Parameter Set{use_ROIif(use_ROI){use_ROI_distributionif (use_ROI_distribution){is_ROI_distribution_static_periodif(is_ROI_distribution_static_period)ROI_distribution_static_period}}}
[0333] In Table 3, the semantics are as follows.
[0334] - use_ROI can indicate whether to perform region-of-interest based coding.
[0335] - Use_ROI_distribution can indicate whether region of interest segmentation coding is performed.
[0336] - is_ROI_distribution_static_period can indicate whether the interval between the region of interest segmentation coding sections is constant.
[0337] - ROI_distribution_static_period can represent a constant interval value between the coding sections of the region of interest segmentation.
[0338]
[0339] Tables 4 and 5 are syntax examples for the picture parameter set.
[0340]
[0341] Picture Parameter set{cropped_image_size_widthcropped_image_size_heightcropped_image_size_difference_flagif(resized_image_size_difference_flag){cropped_to_output_difference_widthcropped_to_output_difference_height}resized_image_size_widthresized_image_size_heightis_ROI_usedis_usage_of_area_presentif(is_ROI_used){has_distributed_ROIif(has_distributed_ROI){distributed_ROI_idxdistributed_ROI_typeframe_idx_in_ROI_distribution_section}num_ROIROI_representationbase_unit_of_ROIif(ROI_representation == DIRECT_REPRESENTATION){for(number_of_area){ROI_location_xROI_location_yROI_widthROI_heightROI_typeROI_scale}}else if(ROI_representation == MAP_REPRESENTATION){for(number of base unit)ROI_idxfor(number of ROI)ROI_type}}}
[0342] Picture Parameter set{cropped_image_size_widthcropped_image_size_heightcropped_image_size_difference_flagif(resized_image_size_difference_flag){cropped_to_output_difference_widthcropped_to_output_difference_height}resized_image_size_widthresized_image_size_heightis_ROI_usedis_usag e_of_area_presentif(is_ROI_used){has_distributed_ROIif(has_distributed_ROI){distributed_ROI_idxdistributed_ROI_typeframe_idx_in_ROI _distribution_section}num_ROIbase_unit_of_ROIfor(number_of_area){ROI_location_xROI_location_yROI_widthROI_heightROI_typeROI_scale}}}
[0343] In Tables 4 and 5, the semantics are as follows.
[0344] -cropped_image_size_width: You can indicate the width of the inversely scaled frame (output frame of 1660) based on the region of interest.
[0345] -cropped_image_size_height: You can indicate the height of the reverse-scaled frame (output frame of 1660) based on the region of interest.
[0346] -cropped_image_size_difference_flag: Indicates whether padding is performed on the edges of the inversely scaled frame (output frame of 1660) based on the region of interest.
[0347] -cropped_to_output_difference_width: This can indicate the number of pixels to be horizontally padded around the edges of the inversely scaled frame (output frame of 1660) based on the region of interest.
[0348] -cropped_to_output_difference_height: This can indicate the number of pixels to be padded vertically at the edges of the inversely scaled frame (output frame of 1660) based on the region of interest.
[0349] -resized_image_size_width: can indicate the width of the scaled frame (input frame of 1660) based on the region of interest.
[0350] -resized_image_size_height: can indicate the width of the scaled frame (input frame of 1660) based on the region of interest.
[0351] -is_ROI_used: Can indicate whether there is one or more regions of interest in the frame.
[0352] -is_usage_of_area_present: Can indicate whether the type of area of interest has been transmitted.
[0353] -has_distributed_ROI: Indicates whether the frame is included in the region of interest distribution coding interval.
[0354] -distributed_ROI_idx: This can indicate the index of the region of interest that is the target of the region of interest distribution coding.
[0355] -Distributed_ROI_type: This can indicate the type of region of interest that is the target of distributed coding.
[0356] -frame_idx_in_ROI_distribution_section: This can indicate the order of the current frame within the region of interest distribution coding section.
[0357] -num_ROI: Can indicate the number of regions of interest within the frame.
[0358] -ROI_representation: Can indicate how to express region of interest information.
[0359] -Base_unit_of_ROI: You can indicate an index that represents the size of the basic unit, which is the unit of region of interest information.
[0360] -direct_representation_method: This can be an index indicating the detailed method when the area of interest is expressed using a direct representation method.
[0361] -template_idx: If the method of directly specifying the region of interest is a template-based method, you can indicate the template index to use.
[0362] -template_scale: If the method of directly specifying the region of interest is a template-based method, it can be an index indicating the scale value for template scaling.
[0363] -template_x_offset: If the method of directly specifying the region of interest is a template-based method, it can be an index or value indicating the x-axis translation value of the template.
[0364] -template_y_offset: If the method of directly specifying the region of interest is a template-based method, it can be an index or value indicating the y-axis translation value of the template.
[0365] -polygon_type: If the direct specification method for the area of interest is a free-form specification method, it can mean the type of the shape.
[0366] -number_of_vertex_in_polygon: If the direct specification method for the region of interest is the free-form specification method and the shape is a polygon rather than a rectangle, it can mean the number of vertices of the polygon.
[0367] -polygon_vertex_x: If the direct specification method of the area of interest is a free-form specification method and the shape is a polygon rather than a rectangle, it can mean the x-axis coordinate of each vertex of the polygon.
[0368] -polygon_vertex_y: If the direct specification method of the area of interest is a free-form specification method and the shape is a polygon rather than a rectangle, it can mean the y-axis coordinate of each vertex of the polygon.
[0369] -ROI_location_x: You can display a value representing the x-axis location of the upper left basic unit of the region of interest in basic units.
[0370] -ROI_location_y: You can display a value representing the y-axis location of the upper left basic unit of the region of interest in basic units.
[0371] -ROI_width: You can display a value representing the width of the region of interest in basic units.
[0372] -ROI_height: You can display a value representing the height of the region of interest in basic units.
[0373] -ROI_idx: Can indicate the index of the region of interest.
[0374] -ROI_type: Can indicate the type of region of interest.
[0375] -ROI_scale: Can indicate the degree of scaling of the region of interest.
[0376]
[0377] The examples of the present disclosure presented in this specification and drawings are intended solely to facilitate the technical content of the present disclosure and to aid understanding thereof, and are not intended to limit the scope of the present disclosure. It will be apparent to those skilled in the art that other variations are possible in addition to the examples described above.
[0378] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
[0379] A non-transitory computer-readable recording medium having recorded thereon instructions according to one disclosure of the present specification, wherein the instructions, when executed by one or more processors, cause the one or more processors to: an internal decoding step of generating a decoded reconstructed image from a bitstream and parsing information about the reconstructed image; a region-of-interest-based reconstruction step of reconstructing the reconstructed image based on information related to a region of interest included in the information about the reconstructed image; a step of performing spatial reconstruction by changing a resolution of at least a portion of the reconstructed image based on information related to spatial reconstruction included in the information about the reconstructed image; and a step of performing temporal reconstruction by generating at least one frame of the reconstructed image based on information related to temporal reconstruction included in the information about the reconstructed image, wherein after performing the internal decoding step, an execution order of at least two of the region-of-interest-based reconstruction step, the spatial reconstruction step, and the temporal reconstruction step can be adaptively determined based on information about the reconstructed image.
[0380] The present disclosure can be used in the industrial field of an encoding / decoding method, device and recording medium based on region-of-interest distribution of an image for a machine.
Claims
1. A step of segmenting at least one region of interest for each frame of an input image; A step of determining a first region of interest distribution coding section among two or more consecutive frame sections within the sequence of the input image; Within the first region of interest distribution coding section, a step of selecting at least one region of interest for each frame as a target of the region of interest distribution coding and distributing the selected region to frames within the first region of interest distribution coding section; and An image encoding method comprising a step of generating a bitstream by encoding information on a region of interest segmentation according to the above segmentation and the input image.
2. In paragraph 1, The above at least one area of interest is classified into one or more types, An image encoding method, wherein the above region of interest segmentation information includes a type for each of the at least one region of interest.
3. In paragraph 2, The step of distributing at least one selected region of interest within the first region of interest distribution coding section is: A video encoding method, comprising a step of moving first types of regions of interest included in a first frame included in the first region of interest distribution coding section to at least a portion within the first region of interest distribution coding section.
4. In paragraph 2, For the above input image, the step of dividing each frame into at least one region of interest is as follows: A step of deriving first types of regions of interest by a first region of interest derivation method for each frame; and An image encoding method, comprising a step of deriving second types of regions of interest by a second region of interest derivation method for each frame.
5. In paragraph 4, The above first region of interest derivation method independently derives a region of interest for each frame, An image encoding method, wherein at least one of the positions or sizes of the derived regions of interest are different, or at least some of the derived regions of interest include a portion that overlaps with each other.
6. In paragraph 5, A method for deriving a region of interest, the first method comprising: a method for deriving a region of interest by detecting movement by comparing the current frame with adjacent frames; a method for using a neural network for object detection; or a method for using a segmentation lightweight neural network.
7. In paragraph 4, The second method for deriving a region of interest is an image encoding method for deriving a region of interest by at least one of a position or a size specified by a sequence or frame group unit in all frames.
8. In paragraph 1, An image encoding method in which the above region of interest segmentation information is expressed by a method of directly specifying the region of interest unit or by a method of displaying it as a map based on basic units.
9. In paragraph 8, The method of directly specifying the above region of interest units is to express each region of interest by including at least one of the type of shape of each region of interest, the reference point of each region of interest, the width of each region of interest, or the height of each region of interest for each frame, A method of displaying a map based on the above basic units is an image encoding method in which two or more pixels are grouped in a rectangular shape, each basic unit is listed so as not to overlap in the entire frame, and each region of interest is expressed in correspondence to the basic units divided by a specified absolute size or relative size.
10. In paragraph 2, In the first region of interest distribution coding section, the step of selecting at least one region of interest for each frame as a target of the region of interest distribution coding and distributing it to frames within the first region of interest distribution coding section is as follows. Divide the second type of regions of interest included in the first frame within the first region of interest distribution coding section into each frame within the first region of interest distribution coding section, or A video encoding method in which the second type of regions of interest included in each frame within the first region of interest distribution coding section is deleted, leaving only one region of interest that does not overlap with each other.
11. In paragraph 1, A method for encoding an image, wherein the above region of interest segmentation information includes at least one of information on whether the period in which the region of interest distribution coding section appears within the sequence is fixed or period information.
12. In paragraph 1, A method for encoding an image, wherein the above region of interest segmentation information includes at least one of an index indicating whether the first frame of the input image is a frame within a region of interest distribution coding section, an index indicating which frame the first frame is within the region of interest distribution coding section, an index of the type of the region of interest that is a target of distribution coding, or an index of the region of interest that is a target of distribution coding.
13. A step of decoding a bitstream to generate region of interest segmentation information and a decoded image; and A step of relocating at least one region of interest in each frame of the decoded image based on the region of interest segmentation information is included. A method for decoding an image, wherein the above region of interest segmentation information includes at least one of an index indicating whether the first frame of the input image is a frame within a region of interest distribution coding section, an index indicating which frame the first frame is within the region of interest distribution coding section, an index of the type of the region of interest that is a target of distribution coding, or an index of the region of interest that is a target of distribution coding.
14. In paragraph 13, The above at least one area of interest is classified into one or more types, An image decoding method, wherein the above region of interest segmentation information includes a type for each of the at least one region of interest.
15. In paragraph 13, Based on the above region of interest segmentation information, the step of relocating at least one region of interest in each frame of the decoded image is as follows: A step of identifying whether a region of interest distribution section exists at a sequence level or a frame level for the decoded image based on the region of interest segmentation information; and In response to identifying that the above region of interest distribution section exists, a step of searching for a first region of interest distribution section in the decoded image using period information included in the region of interest segmentation information is included. A method for decoding an image, wherein the first area of interest distribution section includes two or more consecutive frames.
16. In paragraph 15, Based on the above region of interest segmentation information, the step of relocating at least one region of interest in each frame of the decoded image is as follows: An image decoding method, comprising a step of identifying a region of interest distribution coding target within the first region of interest distribution section based on the region of interest segmentation information.
17. In paragraph 16, Based on the above region of interest segmentation information, the step of relocating at least one region of interest in each frame of the decoded image is as follows: An image decoding method, comprising a step of relocating the identified region of interest distribution coding target distributed within the first region of interest distribution section to the original frame position based on the region of interest segmentation information.
18. In paragraph 16, Based on the above region of interest segmentation information, the step of relocating at least one region of interest in each frame of the decoded image is as follows: An image decoding method comprising a step of copying the identified region of interest distribution coding target in all frames within the first region of interest distribution section to a corresponding position of the first frame within the first region of interest distribution section.
19. In paragraph 18, Based on the above region of interest segmentation information, the step of relocating at least one region of interest in each frame of the decoded image is as follows: An image decoding method, comprising a step of performing filtering on pixels adjacent to the boundary of a region of interest copied from another frame in the first frame within the first region of interest distribution section.
20. In a method of transmitting a bitstream, For an input image, a step of segmenting at least one region of interest for each frame; A step of determining a first region of interest distribution coding section among two or more consecutive frame sections within the sequence of the input image; Within the first region of interest distribution coding section, a step of selecting at least one region of interest for each frame as a target of the region of interest distribution coding and distributing the selected region to frames within the first region of interest distribution coding section; and A method for transmitting a bitstream, comprising a step of generating a bitstream by encoding information on a region of interest divided according to the above division and the input image and transmitting the generated bitstream.
Citation Information
Patent Citations
Novel compound and organic light emitting device comprising the same
KR1020240110345A
Robot control method and robot control device
KR1020240155796A
Method and System for Video Encoding Based on Object Detection Tracking
KR102550117B1
Docking system for aerial drone and ground robot
KR102619537B1
Distributed analysis of a multi-layer signal encoding
WO2024084248A1