Method and apparatus for encoding / decoding image data

JP2024157023A5Active Publication Date: 2025-05-19INST OF IMAGE TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024137918
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-07-17
Filing Date
2024-08-19
Publication Date
2025-05-19
Estimated Expiration
2037-10-10

AI Technical Summary

Technical Problem

Conventional image encoding/decoding methods struggle to efficiently process 360-degree images, leading to insufficient performance in handling the large data volumes generated by multi-view images for immersive media such as virtual and augmented reality.

Method used

A 360-degree image decoding method that includes receiving a bitstream, generating a predicted image by referring to syntax information, and reconstructing the image in various projection formats like ERP, CMP, and OHP, while utilizing motion vector candidates and image expansion techniques to improve compression performance.

Benefits of technology

Enhances compression performance for 360-degree images by optimizing the image setting process, addressing the inefficiencies in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method and an apparatus for encoding / decoding image data of a 360-degree image.SOLUTION: A decoding method includes: generating a prediction image by making reference to syntax information obtained from a bitstream that is formed by encoding a 360-degree image; combining the generated prediction image with a residual image obtained by dequantizing and inverse-transforming the bitstream, so as to obtain a decoded image; and reconstructing the decoded image into a 360-degree image according to a projection format. Here, generating the prediction image includes: obtaining, from motion information included in the syntax information, a motion vector candidate group including a motion vector of a block adjacent to a current block to be decoded; deriving a prediction motion vector from the motion vector candidate group, on the basis of selection information extracted from the motion information; and determining a prediction block for the current block to be decoded, using a final motion vector derived by adding the prediction motion vector to a differential motion vector extracted from the motion information.SELECTED DRAWING: Figure 44
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to image data encoding and decoding technology, and more particularly to a method and apparatus for processing 360-degree image encoding and decoding for immersive media services. [Background technology]

[0002] With the spread of the Internet and mobile terminals and the development of information and communication technology, the use of multimedia data is rapidly increasing. Recently, the demand for high-resolution images and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been generated in various fields, and the demand for immersive media services such as virtual reality and augmented reality is also rapidly increasing. In particular, in the case of 360-degree images for virtual reality and augmented reality, multi-view images taken by multiple cameras are processed, which generates a huge amount of data, but the performance of image processing systems to process this data is insufficient.

[0003] Thus, in the prior art image encoding / decoding methods and apparatus, there is a need for improved performance for image processing, particularly image encoding / decoding. Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention is directed to solving the above-mentioned problems, and has an object to provide a method for improving an image setting process in an early stage of encoding and decoding, and more particularly, to provide an encoding and decoding method and apparatus for improving an image setting process that takes into account the characteristics of a 360-degree image. [Means for solving the problem]

[0005] In order to achieve the above object, one aspect of the present invention provides a method for decoding a 360-degree image.

[0006] Here, the method for decoding a 360-degree image may include the steps of receiving a bitstream in which a 360-degree image is encoded, generating a predicted image by referring to syntax information obtained from the received bitstream, obtaining a decoded image by combining the generated predicted image with a residual image obtained by inverse quantizing and inverse transforming the bitstream, and reconstructing the decoded image into a 360-degree image in a projection format.

[0007] Here, the syntax information may include projection format information for the 360-degree image.

[0008] Here, the projection format information may be information indicating at least one of an ERP (Equi-Rectangular Projection) format in which the 360-degree image is projected onto a two-dimensional plane, a CMP (CubeMap Projection) format in which the 360-degree image is projected onto a cube, an OHP (OctaHedron Projection) format in which the 360-degree image is projected onto an octahedron, and an ISP (IcoSahedral Projection) format in which the 360-degree image is projected onto a polyhedron.

[0009] Here, the reconstructing step may include the steps of: obtaining placement information based on regional packing by referring to the syntax information; and rearranging each block of the decoded image based on the placement information.

[0010] Here, the step of generating the predicted image may include a step of performing image extension on a reference picture obtained by restoring the bitstream, and a step of generating the predicted image by referring to the reference picture that has been image extended.

[0011] Here, the step of performing the image extension may include performing image extension based on a division unit of the reference image.

[0012] Here, the step of performing image expansion on a division unit basis may generate an expanded region for each division unit individually using boundary pixels of the division unit.

[0013] Here, the expanded region can be generated using boundary pixels of division units that are spatially adjacent to the division unit to be expanded, or boundary pixels of division units that have image continuity with the division unit to be expanded.

[0014] Here, the step of performing image extension based on the division units may generate an extended image for a combined region by using boundary pixels of a region where two or more spatially adjacent division units among the division units are combined.

[0015] Here, the step of performing image expansion based on the division units may generate an expanded area between adjacent division units by using all adjacent pixel information of spatially adjacent division units among the division units.

[0016] Here, the step of performing image expansion on a division unit basis may generate the expanded region using an average value of adjacent pixels of each of the spatially adjacent division units.

[0017] Here, the step of generating the predicted image may include a step of obtaining a group of motion vector candidates including motion vectors of blocks adjacent to the current block to be decoded from the motion information included in the syntax information, a step of deriving a predicted motion vector from the group of motion vector candidates based on selection information extracted from the motion information, and a step of determining a predicted block of the current block to be decoded using a final motion vector derived by adding the predicted motion vector to a differential motion vector extracted from the motion information.

[0018] Here, when a block adjacent to the current block is different from the surface to which the current block belongs, the group of motion vector candidates can be composed of only motion vectors for blocks that belong to surfaces that have image continuity with the surface to which the current block belongs, among the adjacent blocks.

[0019] Here, the adjacent block may refer to a block adjacent to the current block in at least one direction among the upper left, upper, upper right, left, and lower left.

[0020] Here, the final motion vector may point to a reference area that belongs to at least one reference picture with respect to the current block and is set to an area where there is image continuity between surfaces according to the projection format.

[0021] Here, the reference picture can be expanded in the up, down, left, and right directions based on image continuity according to the projection format, and then the reference area can be set.

[0022] Here, the reference picture is extended in units of the surface, and the reference region can be set across the boundary of the surface.

[0023] Here, the motion information may include at least one of a reference picture list to which the reference picture belongs, an index of the reference picture, and a motion vector indicating the reference area.

[0024] Here, generating a predicted block for the current block may include dividing the current block into a plurality of sub-blocks and generating a predicted block for each of the divided sub-blocks. Effect of the Invention

[0025] When using the image encoding / decoding method and device according to the embodiment of the present invention as described above, it is possible to improve compression performance, particularly in the case of 360-degree images. [Brief description of the drawings]

[0026] [Figure 1] 1 is a block diagram of an image encoding device according to an embodiment of the present invention. [Diagram 2] 1 is a block diagram of an image decoding device according to an embodiment of the present invention. [Diagram 3] FIG. 2 is a diagram illustrating an example of image information divided into layers for image compression; [Figure 4] 1A-1D are conceptual diagrams illustrating various examples of image segmentation according to an embodiment of the present invention. [Diagram 5] 4 is a diagram illustrating another example of an image division method according to an embodiment of the present invention. [Figure 6] 1 is a diagram illustrating a typical method for adjusting the size of an image; [Figure 7] 1 is an exemplary diagram of image size adjustment according to an embodiment of the present invention; [Figure 8] 11 is an exemplary diagram illustrating a method for configuring an expanded area in an image resizing method according to an embodiment of the present invention. [Figure 9] 11 is an exemplary diagram illustrating a method for configuring an area to be deleted and an area to be generated by reducing in an image resizing method according to an embodiment of the present invention; [Figure 10] FIG. 2 is an exemplary diagram of image reconstruction according to an embodiment of the present invention. [Figure 11] 1A and 1B are exemplary diagrams illustrating images before and after an image setting process according to an embodiment of the present invention. [Figure 12] 11A and 11B are diagrams illustrating an example of size adjustment for each division unit in an image according to an embodiment of the present invention. [Figure 13] FIG. 13 is an exemplary diagram of a set of size adjustments or settings for division units within an image. [Figure 14] 11 is an exemplary diagram illustrating an image size adjustment process and a size adjustment process of a division unit within an image; [Figure 15] 1 is an exemplary diagram showing a three-dimensional space showing a three-dimensional image and a two-dimensional planar space. [Figure 16a]FIG. 2 is a conceptual diagram for explaining a projection format according to an embodiment of the present invention. [Figure 16b] FIG. 2 is a conceptual diagram for explaining a projection format according to an embodiment of the present invention. [Figure 16c] FIG. 2 is a conceptual diagram for explaining a projection format according to an embodiment of the present invention. [Figure 16d] FIG. 2 is a conceptual diagram for explaining a projection format according to an embodiment of the present invention. [Figure 17] FIG. 2 is a conceptual diagram of a projection format realized within a rectangular image according to an embodiment of the present invention. [Figure 18] FIG. 13 is a conceptual diagram of how to convert a projection format to a rectangular shape by rearranging surfaces to eliminate insignificant areas in accordance with an embodiment of the present invention. [Figure 19] 1 is a conceptual diagram illustrating a packing process performed by regions on a CMP projection format according to an embodiment of the present invention as a rectangular image. [Figure 20] FIG. 2 is a conceptual diagram of division of a 360-degree image according to an embodiment of the present invention. [Figure 21] 1 is an exemplary diagram of a 360-degree image division and image reconstruction according to an embodiment of the present invention; [Figure 22] FIG. 13 is an example diagram of an image projected or packed by CMP divided into tiles. [Diagram 23] FIG. 11 is a conceptual diagram for explaining an example of size adjustment of a 360-degree image according to an embodiment of the present invention. [Figure 24] FIG. 2 is a conceptual diagram for explaining the continuity between surfaces in a projection format (eg, CMP, OHP, ISP) according to an embodiment of the present invention. [Diagram 25] FIG. 21C is a conceptual diagram for explaining the continuity of the surface of FIG. 21C, which is an image acquired by an image reconstruction process or a regional packing process in a CMP projection format. [Figure 26] 11 is an exemplary diagram for explaining image size adjustment in a CMP projection format according to an embodiment of the present invention. FIG. [Figure 27] FIG. 13 is an exemplary diagram illustrating size adjustment for an image that has been converted into a CMP projection format and packed in accordance with an embodiment of the present invention. [Figure 28] 1 is an exemplary diagram illustrating a data processing method for adjusting the size of a 360-degree image according to an embodiment of the present invention. [Figure 29] FIG. 13 is an exemplary diagram showing a tree-based block shape. [Diagram 30] FIG. 13 is an exemplary diagram showing a type-based block shape. [Diagram 31] 4A to 4C are exemplary diagrams showing various block shapes that can be obtained by the block division unit of the present invention; [Diagram 32] FIG. 2 is an exemplary diagram illustrating tree-based partitioning according to an embodiment of the present invention. [Diagram 33] FIG. 2 is an exemplary diagram illustrating tree-based partitioning according to an embodiment of the present invention. [Diagram 34] 1 is an example diagram showing various cases in which a prediction block is obtained by inter-picture prediction; [Diagram 35] FIG. 2 is an exemplary diagram illustrating a method for constructing a reference picture list according to an embodiment of the present invention. [Diagram 36] FIG. 2 is a conceptual diagram illustrating a non-moving motion model according to an embodiment of the present invention. [Figure 37] 4 is an example diagram illustrating sub-block-based motion estimation according to an embodiment of the present invention; [Figure 38] 4 is an example diagram illustrating a block referenced in motion information prediction of a current block according to an embodiment of the present invention; [Figure 39] 2 is an exemplary diagram illustrating blocks referenced for motion information prediction of a current block in a non-motion motion model according to an embodiment of the present invention; [Diagram 40] 1 is an exemplary diagram illustrating inter-frame prediction using an extended picture according to an embodiment of the present invention; [Diagram 41] FIG. 13 is a conceptual diagram illustrating the expansion of a surface unit according to an embodiment of the present invention. [Diagram 42]1 is an example diagram illustrating inter-frame prediction using an extended image according to an embodiment of the present invention; [Diagram 43] 1 is an exemplary diagram illustrating inter-frame prediction using extended reference pictures according to an embodiment of the present invention; [Diagram 44] 1 is an illustrative diagram showing a configuration of a motion information prediction candidate group for inter-screen prediction in a 360-degree image according to an embodiment of the present invention. FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0027] Although the present invention can be modified in various ways and can have various embodiments, a specific embodiment will be illustrated in the drawings and described in detail herein. However, this does not limit the present invention to the specific embodiment, and it should be understood that the present invention includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention.

[0028] Terms such as first, second, A, B, etc. are used to describe various components, but the components should not be limited by these terms. These terms are used only to distinguish one component from another. For example, a first component can be named a second component and similarly, the second component can be named a first component without departing from the scope of the present invention. The term "and / or" includes any combination of two or more associated listed items or two or more associated listed items.

[0029] When an element is said to be "coupled" or "connected" to another element, it means that it is directly coupled or connected to the other element, but it should be understood that there may be other intervening elements between them. In contrast, when an element is said to be "directly coupled" or "directly connected" to another element, it should be understood that there are no other intervening elements between them.

[0030] The terms used in this specification are merely used to describe certain embodiments and are not intended to limit the present invention. A singular expression includes a plural expression unless the context clearly indicates otherwise. In this specification, the terms "include" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or additional possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0031] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention belongs. Terms defined in commonly used dictionaries should be interpreted in accordance with the contextual meaning of the relevant art, and should not be interpreted in an ideal or overly formal sense unless expressly defined in this specification.

[0032] The image encoding device and the decoding device may be a user terminal such as a personal computer (PC), a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a PlayStation Portable (PSP), a wireless communication terminal, a smartphone, a TV, a virtual reality device (VR), an augmented reality device (AR), a mixed reality device (MR), a head mounted display (HMD), or smart glasses, or may be a server terminal such as an application server or a service server. The image encoding device and the decoding device may include various devices including a communication device such as a communication modem for communicating with various devices or wired and wireless communication networks, a memory for storing various programs and data for encoding or decoding images or for intra-screen or inter-screen prediction for encoding or decoding, and a processor for executing programs for calculation and control. In addition, the image coded into a bitstream by the image coding device can be transmitted to an image decoding device in real time or non-real time via wired and wireless communication networks such as the Internet, a short-range wireless communication system, a wireless LAN network, a WiBro network, a mobile communication network, or via various communication interfaces such as a cable or a Universal Serial Bus (USB), and then decoded by the image decoding device to restore and play back the image.

[0033] Moreover, the image encoded into a bitstream by the image encoding device can also be transmitted from the encoding device to the decoding device via a computer-readable recording medium.

[0034] The image encoding device and the image decoding device may be separate devices, but may be implemented as a single image encoding / decoding device. In this case, some components of the image encoding device may be substantially the same technical elements as some components of the image decoding device, and may include at least the same structure or perform at least the same function.

[0035] Therefore, in the following detailed description of the technical elements and their operating principles, duplicated descriptions of the corresponding technical elements will be omitted.

[0036] The image decoding apparatus corresponds to a computer device that applies to decoding the image coding method performed by the image coding apparatus, and therefore the following description will focus on the image coding apparatus.

[0037] The computer device may include a memory for storing a program or software module for implementing the image encoding method and / or the image decoding method, and a processor connected to the memory for executing the program. The image encoding device may be called an encoder, and the image decoding device may be called a decoder.

[0038] Typically, an image can be composed of a series of still images, and these still images can be divided into GOP (Group of Pictures) units. Each still image can be called a picture. In this case, a picture can indicate either a frame or a field in a progressive signal or an interlaced signal, and an image can be represented as a "frame" when encoding / decoding is performed in units of frames, and as a "field" when encoding / decoding is performed in units of fields. In the present invention, a progressive signal is assumed, but the present invention can also be applied to an interlaced signal. As a higher-level concept, units such as GOP and sequence can exist. Also, each picture can be divided into predetermined regions such as slices, tiles, and blocks. Also, one GOP can include units such as I-pictures, P-pictures, and B-pictures. An I-picture may refer to a picture that is encoded / decoded by itself without using a reference picture, and a P-picture and a B-picture may refer to a picture that is encoded / decoded by performing processes such as motion estimation and motion compensation using a reference picture. In general, a P-picture may use an I-picture and a P-picture as reference pictures, and a B-picture may use an I-picture and a P-picture as reference pictures, but the above definitions may be changed depending on the encoding / decoding settings.

[0039] Here, a picture referred to in encoding / decoding is called a reference picture, and a block or pixel referred to is called a reference block or reference pixel. Also, the reference data may be not only pixel values ​​in the spatial domain, but also coefficient values ​​in the frequency domain, and various encoding / decoding information generated and determined during the encoding / decoding process. For example, the reference data may be intra-frame prediction-related information or motion-related information in a prediction unit, transformation-related information in a transform unit / inverse transform unit, quantization-related information in a quantizer / inverse quantizer, encoding / decoding-related information (context information) in an encoding unit / decoding unit, and filter-related information in an in-loop filter unit.

[0040] The smallest unit of an image may be a pixel. The number of bits used to express one pixel is called bit depth. Generally, the bit depth is 8 bits, and more bit depths may be supported depending on the encoding settings. At least one bit depth may be supported depending on the color space. Also, depending on the color format of the image, at least one color space may be configured. Depending on the color format, one or more pictures having a certain size or one or more pictures having different sizes may be configured. For example, in the case of YCbCr4:2:0, it may be configured with one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example). In this case, the composition ratio of the chrominance component to the luminance component may be 1:2 in the horizontal direction. As another example, in the case of 4:4:4, the horizontal and vertical directions may have the same composition ratio. When one or more color spaces are configured as in the above example, the picture may be divided into each color space.

[0041] In the present invention, a description will be given based on a certain color space (Y in this example) of a certain color format (YCbCr in this example). The same or similar application (specific color space dependent setting) can be made to another color space (Cb, Cr in this example) according to the color format. However, it is also possible to have a partial difference (specific color space independent setting) for each color space. That is, the setting dependent on each color space can mean having a setting proportional to or dependent on the composition ratio of each component (e.g., determined according to 4:2:0, 4:2:2, 4:4:4, etc.), and the setting independent on each color space can mean having a setting only for the corresponding color space independently of or regardless of the composition ratio of each component. In the present invention, depending on the encoder / decoder, some configurations can have an independent setting or a dependent setting.

[0042] The configuration information or syntax elements required in the image encoding process can be determined at the unit level such as video, sequence, picture, slice, tile, block, etc. This can be recorded in the bitstream in units such as VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), Slice Header, Tile Header, Block Header, etc. and transmitted to the decoder. The decoder can parse the configuration information transmitted from the encoder in the same level unit and use it in the image decoding process. Also, related information can be transmitted to the bitstream in the form of SEI (Supplement Enhancement Information) or metadata, etc., and parsed for use. Each parameter set has a unique ID value, and a lower parameter set can have the ID value of the upper parameter set it refers to. For example, a lower parameter set can refer to information of an upper parameter set having a matching ID value among one or more upper parameter sets. Among the various examples of units mentioned above, when any unit includes one or more other units, the relevant unit may be called a higher-level unit, and the included units may be called lower-level units.

[0043] In the case of the setting information generated in the unit, it may include contents for independent setting for each unit, or may include contents for dependent setting for the previous, subsequent, or higher unit. Here, dependent setting may be understood as indicating setting information for the unit by flag information indicating that the setting of the previous, subsequent, or higher unit is followed (for example, if a 1-bit flag is 1, the setting is followed; if it is 0, the setting is not followed). Although the setting information in the present invention will be described mainly with reference to an example of independent setting, examples of addition or replacement of contents for a dependent relationship to setting information of the previous, subsequent, or higher unit of the current unit may also be included.

[0044] Fig. 1 is a block diagram of an image encoding device according to an embodiment of the present invention, and Fig. 2 is a block diagram of an image decoding device according to an embodiment of the present invention.

[0045] Referring to FIG. 1, the image coding device may be configured to include a prediction unit, a subtraction unit, a transformation unit, a quantization unit, an inverse quantization unit, an inverse transformation unit, an addition unit, an in-loop filter unit, a memory and / or a coding unit, and some of the above components may not necessarily be included, and some or all of them may be selectively included depending on the implementation, and some additional components not shown may be included.

[0046] Referring to FIG. 2, the image decoding device may be configured to include a decoding unit, a prediction unit, an inverse quantization unit, an inverse transform unit, an adder unit, an in-loop filter unit, and / or a memory, and some of the above components may not necessarily be included, and some or all of them may be selectively included depending on the implementation, and some additional components not shown may be included.

[0047] The image encoding device and the image decoding device may be separate devices, but may be made into one image encoding / decoding device depending on the implementation. In that case, some configurations of the image encoding device are technical elements that are substantially the same as some configurations of the image decoding device, and can be implemented to include at least the same structure or perform at least the same function. Therefore, in the following detailed description of technical elements and their operating principles, duplicated descriptions of corresponding technical elements will be omitted. Since the image decoding device corresponds to a computer device that applies the image encoding method performed in the image encoding device to decoding, the following description will focus on the image encoding device. The image encoding device may be called an encoder, and the image decoding device may be called a decoder.

[0048] The prediction unit may be realized using a prediction module, which is a software module, and may generate a prediction block for a block to be coded using an intra prediction method or an inter prediction method. The prediction unit predicts a current block to be coded in an image to generate a prediction block. That is, the prediction unit predicts a pixel value of each pixel of a current block to be coded in an image through intra prediction or inter prediction, and generates a prediction block having a predicted pixel value of each pixel generated. In addition, the prediction unit may transmit information required for generating a prediction block to an encoding unit to encode information on a prediction mode, record the information on the prediction mode in a bitstream, and transmit the bitstream to a decoder, and a decoding unit of the decoder parses the information on the prediction mode to restore the information on the prediction mode, and then use the information for intra prediction or inter prediction.

[0049] The subtraction unit subtracts the predicted block from the current block to generate a residual block. That is, the subtraction unit calculates the difference between the pixel value of each pixel of the current block to be encoded and the predicted pixel value of each pixel of the predicted block generated by the prediction unit to generate the residual block, which is a residual signal in the form of a block.

[0050] The transform unit can transform a signal belonging to the spatial domain into a signal belonging to the frequency domain. At this time, a signal obtained through the transform process is called a transformed coefficient. For example, a residual block having a residual signal transmitted from the subtraction unit can be transformed to obtain a transformed block having a transform coefficient, and the input signal is determined according to the coding setting and is not limited to a residual signal.

[0051] The transform unit may transform the residual block using a transform technique such as a Hadamard transform, a discrete sine transform (DST based-transform), a discrete cosine transform (DCT based-transform), etc. However, the transform technique is not limited to these, and various transform techniques that are improvements or modifications of these may be used.

[0052] For example, at least one of the above conversion techniques may be supported, and at least one detailed conversion technique may be supported for each conversion technique. In this case, the at least one detailed conversion technique may be a conversion technique in which a part of the basis vector is different for each conversion technique. For example, a DST-based conversion and a DCT-based conversion may be supported as conversion techniques. In the case of DST, detailed conversion techniques such as DST-I, DST-II, DST-III, DST-V, DST-VI, DST-VII, and DST-VIII may be supported, and in the case of DCT, detailed conversion techniques such as DCT-I, DCT-II, DCT-III, DCT-V, DCT-VI, DCT-VII, and DCT-VIII may be supported.

[0053] Any of the above transformations (e.g., one transformation technique & one detailed transformation technique) can be set as a basic transformation technique, and additional transformation techniques (e.g., multiple transformation techniques || multiple detailed transformation techniques) can be supported. Whether to support additional transformation techniques can be determined in units such as sequences, pictures, slices, tiles, etc., and related information can be generated in the units, and when additional transformation techniques are supported, transformation technique selection information can be determined in units such as blocks, and related information can be generated.

[0054] The transformation may be performed in the k / vertical direction. For example, the spatial domain pixel values ​​may be transformed to the frequency domain using a one-dimensional transform in the horizontal direction and a one-dimensional transform in the vertical direction using the basis vectors in the transformation to perform a total two-dimensional transform.

[0055] Also, the conversion may be adaptively performed in the horizontal / vertical direction. In detail, whether or not the conversion is adaptively performed may be determined according to at least one encoding setting. For example, when the prediction mode in the intra prediction is a horizontal mode, DCT-I may be applied in the horizontal direction and DST-I may be applied in the vertical direction; when the prediction mode in the intra prediction is a vertical mode, DST-VI may be applied in the horizontal direction and DCT-VI may be applied in the vertical direction; when the prediction mode is a diagonal down left, DCT-II may be applied in the horizontal direction and DCT-V may be applied in the vertical direction; and when the prediction mode is a diagonal down right, DST-I may be applied in the horizontal direction and DST-VI may be applied in the vertical direction.

[0056] The size and shape of each transform block are determined according to the coding cost for each candidate size and shape of the transform block, and information such as image data of each determined transform block and the size and shape of each determined transform block can be coded.

[0057] Among the transformation shapes, a square transformation can be set as a basic transformation shape, and an additional transformation shape (e.g., a rectangular shape) can be supported for this. Whether or not to support an additional transformation shape is determined in units of a sequence, a picture, a slice, a tile, etc., and related information can be generated in the units, and transformation shape selection information can be determined in units of a block, etc., and related information can be generated.

[0058] Also, the support of the transform block shape can be determined according to the coding information. In this case, the coding information can correspond to a slice type, a coding mode, a size and shape of a block, a block division method, etc. That is, one transform shape can be supported according to at least one coding information, and multiple transform shapes can be supported according to at least one coding information. The former case can be an implicit situation, and the latter case can be an explicit situation. In the explicit case, adaptive selection information indicating an optimal candidate group among multiple candidates can be generated and recorded in a bitstream. In the present invention including this example, when the coding information is explicitly generated, the corresponding information can be recorded in a bitstream in various units, and the decoder can parse related information in various units and restore it to decoded information. Also, when the coding / decoding information is processed implicitly, it can be understood that the same process, rules, etc. are used in the encoder and the decoder.

[0059] For example, rectangular transformation support can be determined according to the slice type: for an I slice, the transformation shape supported is a square transformation, and for a P / B slice, the transformation shape supported may be a square or rectangular transformation.

[0060] For example, rectangular transformation support may be determined according to the coding mode, where the transformation shape supported in the case of Intra may be a square transformation, and the transformation shape supported in the case of Inter may be a square or rectangular transformation.

[0061] For example, rectangular transformation support may be determined according to the size and shape of the block. The transformation shape supported for blocks of a certain size or larger may be a square transformation, and the transformation shape supported for blocks of a certain size or smaller may be a square or rectangular transformation.

[0062] For example, rectangular transformation support may be determined according to a block division method: if a block to be transformed is a block obtained by a quad tree division method, the transformation shape supported may be a square transformation, and if the block is obtained by a binary tree division method, the transformation shape supported may be a square or rectangular transformation.

[0063] The above example is an example of supporting a transformation shape according to one piece of coding information, and multiple pieces of information may be combined to participate in additional transformation shape support settings. The above example is merely an example of supporting additional transformation shapes according to various coding settings, and is not limited to the above, and various modified examples are possible.

[0064] Depending on the encoding settings or image characteristics, the transform process may be omitted. For example, depending on the encoding settings (assuming a lossless compression environment in this example), the transform process (including the inverse process) may be omitted. As another example, if the compression performance by the transform is not exhibited depending on the image characteristics, the transform process may be omitted. In this case, the omitted transform may be in whole units or in either horizontal units or vertical units. Whether or not such omission is supported may be determined depending on the size and shape of the block.

[0065] For example, in a setting where the omission of horizontal and vertical transformations is grouped, when the transformation omission flag is 1, transformations in the horizontal and vertical directions may not be performed, and when the transformation omission flag is 0, transformations in the horizontal and vertical directions may be performed. In a setting where the omission of horizontal and vertical transformations operates independently, when the first transformation omission flag is 1, transformations in the horizontal direction are not performed, when the first transformation omission flag is 0, transformations in the horizontal direction are performed, when the second transformation omission flag is 1, transformations in the vertical direction are not performed, and when the second transformation omission flag is 0, transformations in the vertical direction are performed.

[0066] If the block size falls within range A, transformation skipping can be supported, and if the block size falls within range B, transformation skipping cannot be supported. For example, if the block width is larger than M or the block height is larger than N, the transformation skip flag cannot be supported, and if the block width is smaller than m or the block height is smaller than n, the transformation skip flag can be supported. M(m) and N(n) may be the same or different. The transformation-related settings can be determined in units of sequences, pictures, slices, etc.

[0067] If additional transformation techniques are supported, the configuration of the transformation techniques may be determined according to at least one coding information, which may include a slice type, a coding mode, a block size and shape, a prediction mode, and the like.

[0068] For example, the support of a transform technique may be determined according to the coding mode. In the case of Intra, the supported transform techniques may be DCT-I, DCT-III, DCT-VI, DST-II, and DST-III, and in the case of Inter, the supported transform techniques may be DCT-II, DCT-III, and DST-III.

[0069] For example, the support of a transform technique may be determined according to the slice type. The transform techniques supported for an I slice may be DCT-I, DCT-II, or DCT-III, the transform techniques supported for a P slice may be DCT-V, DST-V, or DST-VI, and the transform techniques supported for a B slice may be DCT-I, DCT-II, or DST-III.

[0070] For example, support of a transformation technique may be determined according to a prediction mode. The transformation techniques supported in prediction mode A may be DCT-I and DCT-II, the transformation techniques supported in prediction mode B may be DCT-I and DST-I, and the transformation technique supported in prediction mode C may be DCT-I. In this case, prediction modes A and B may be directional modes, and prediction mode C may be a non-directional mode.

[0071] For example, the support of a transform technique may be determined according to the size and shape of a block. The transform technique supported for blocks of a certain size or more may be DCT-II, the transform technique supported for blocks of less than a certain size may be DCT-II and DST-V, and the transform techniques supported for blocks of a certain size or more and less than a certain size may be DCT-I, DCT-II and DST-I. In addition, the transform techniques supported for a square shape may be DCT-I and DCT-II, and the transform techniques supported for a rectangular shape may be DCT-I and DST-I.

[0072] The above example is an example of supporting a transformation technique according to one piece of coding information, and multiple pieces of information may be combined to support additional transformation techniques. The present invention is not limited to the above example, and other modifications are possible. In addition, the transform unit may transmit information required to generate a transformation block to the encoding unit to encode the information, record the information in a bitstream, and transmit the information to the decoder, and the decoding unit of the decoder may parse the information and use it in the inverse transformation process.

[0073] The quantization unit may quantize an input signal. At this time, a signal obtained through the quantization process is called a quantized coefficient. For example, a residual block having a residual transform coefficient transmitted from the transform unit may be quantized to obtain a quantized block having a quantized coefficient. The input signal is determined according to a coding setting, and is not limited to the residual transform coefficient.

[0074] The quantization unit may quantize the transformed residual block using a quantization technique such as Dead Zone Uniform Threshold Quantization, Quantization Weighted Matrix, etc., and may use various quantization techniques that are improvements and modifications of the quantization technique, without being limited thereto. Whether or not to support an additional quantization technique may be determined in units such as a sequence, a picture, a slice, a tile, etc., and related information may be generated in the units, and if an additional quantization technique is supported, quantization technique selection information may be determined in units such as a block, and related information may be generated.

[0075] If an additional quantization technique is supported, the setting of the quantization technique may be determined according to at least one coding information, which may include a slice type, a coding mode, a block size and shape, a prediction mode, etc.

[0076] For example, the quantizer may set a quantization weight matrix according to a coding mode and a weight matrix applied according to inter prediction / intra prediction to be different from each other. Also, the weight matrix applied according to an intra prediction mode may be set to be different. In this case, the quantization weight matrix may be a quantization matrix having some different quantization components when it is assumed that the block size is the same as the quantization block size with a size of M×N.

[0077] The quantization process may be omitted depending on the encoding settings or image characteristics. For example, the quantization process (including the inverse process) may be omitted depending on the encoding settings (assuming a lossless compression environment in this example). As another example, the quantization process may be omitted when the compression performance by quantization is not exhibited depending on the image characteristics. In this case, the region to be omitted may be the entire region or a part of the region. Whether or not such omission is supported may be determined depending on the size and shape of the block.

[0078] Information about quantization parameters (QP) can be generated in units of sequences, pictures, slices, tiles, blocks, etc. For example, a basic QP can be set in the upper unit where QP information is generated first. <1> The lower the unit, the higher the QP can be set to a value that is the same as or different from the QP set in the higher unit. <2> Through this process, the QP can be finally determined by the quantization process performed on some units. <3> In this case, the units of sequences, pictures, etc. are <1> The units of slices, tiles, blocks, etc. are <2> The units of blocks are <3> This may be an example of the above.

[0079] The information on the QP may be generated based on the QP in each unit. Alternatively, a preset QP may be set as a predicted value to generate differential value information from the QP in each unit. Alternatively, a QP obtained based on at least one of the QP set in the higher unit, the QP previously set in the same unit, or the QP set in the adjacent unit may be set as a predicted value to generate differential value information from the QP in the current unit. Alternatively, a QP obtained based on the QP set in the higher unit and at least one piece of coding information may be set as a predicted value to generate differential value information from the QP in the current unit. In this case, the previous same unit may be a unit that can be defined according to the coding order of each unit, the adjacent unit may be a unit that is spatially adjacent, and the coding information may be a slice type, coding mode, prediction mode, position information, etc. of the corresponding unit.

[0080] As an example, the QP of the current unit may generate differential value information by setting the QP of the upper unit as a predicted value. Differential value information between a QP set in a slice and a QP set in a picture may be generated, or differential value information between a QP set in a tile and a QP set in a picture may be generated. Also, differential value information between a QP set in a block and a QP set in a slice or a tile may be generated. Also, differential value information between a QP set in a sub-block and a QP set in a block may be generated.

[0081] For example, the QP of the current unit may generate difference value information by setting a QP obtained based on the QP of at least one adjacent unit or a QP obtained based on the QP of at least one previous unit as a predicted value. Difference value information with respect to a QP obtained based on the QP of an adjacent block such as the left, upper left, lower left, upper, or upper right of the current block may be generated. Alternatively, difference value information with respect to a QP of a picture encoded before the current picture may be generated.

[0082] As an example, the QP of the current unit may generate difference value information by setting the QP of the upper unit and the QP obtained based on at least one encoding information as a predicted value. Difference value information between the QP of the current block and the QP of the slice corrected according to the slice type (I / P / B) may be generated. Alternatively, difference value information between the QP of the current block and the QP of the tile corrected according to the encoding mode (Intra / Inter) may be generated. Alternatively, difference value information between the QP of the current block and the QP of the picture corrected according to the prediction mode (directional / non-directional) may be generated. Alternatively, difference value information between the QP of the current block and the QP of the picture corrected according to the position information (x / y) may be generated. At this time, the meaning of the correction may mean that the QP of the upper unit used for prediction is added or subtracted in an offset form. At this time, at least one offset information may be supported according to the encoding setting, and the related information may be implicitly processed or explicitly generated according to a predetermined process. The present invention is not limited to the above example, and other modifications are possible.

[0083] The above examples may be possible examples when a signal indicating a QP variation is provided or activated. For example, when a signal indicating a QP variation is not provided or is deactivated, differential value information is not generated, and the predicted QP can be determined as the QP of each unit. As another example, when a signal indicating a QP variation is provided or activated, differential value information is generated, and when the value is 0, the predicted QP can be determined as the QP of each unit.

[0084] The quantization unit can transmit the information necessary to generate a quantization block to the encoding unit so that it can be encoded, and the resulting information can be recorded in a bitstream and transmitted to the decoder, and the decoding unit of the decoder can parse the corresponding information and use it in the inverse quantization process.

[0085] In the above example, the explanation is based on the assumption that the residual block is transformed and quantized through a transform unit and a quantization unit, but the residual signal may be transformed to generate a residual block having transform coefficients without performing the quantization process, the residual signal of the residual block may be transformed into transform coefficients without performing the quantization process, or both the transform and quantization processes may not be performed. This can be determined according to the settings of the encoder.

[0086] The encoding unit may generate a quantization coefficient sequence, a transform coefficient sequence, or a signal sequence by scanning the quantization coefficients, transform coefficients, or residual signals of the generated residual block according to at least one scanning order (e.g., zigzag scan, vertical scan, horizontal scan, etc.), and may encode the quantization coefficient sequence, transform coefficient sequence, or signal sequence using at least one entropy coding technique. In this case, information on the scan order may be determined according to an encoding setting (e.g., an encoding mode, a prediction mode, etc.), and may be implicitly determined or related information may be explicitly generated. For example, one of a plurality of scan orders may be selected according to an intra-frame prediction mode.

[0087] Also, it is possible to generate coded data including coding information transmitted from each component and output it to a bit stream, which can be realized by a multiplexer (MUX).In this case, coding techniques such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. can be used, but are not limited thereto, and various coding techniques that are improvements or modifications of these can be used.

[0088] When performing entropy coding (assumed to be CABAC in this example) on syntax elements such as the residual block data and information generated during the encoding / decoding process, the entropy coding device may include a binarizer, a context modeler, and a binary arithmetic coder. In this case, the binary arithmetic coder may include a regular coding engine and a bypass coding engine.

[0089] Since the syntax element input to the entropy coding device may not be binary, if the syntax element is not binary, a binarization unit binarizes the syntax element and can output a Bin String consisting of 0 or 1. In this case, Bin indicates a bit consisting of 0 or 1, and can be coded through a binary arithmetic coding unit. In this case, either a regular coding unit or a bypass coding unit can be selected based on the occurrence probability of 0 and 1. This can be determined according to the coding / decoding setting. If the syntax element is data with the same frequency of 0 and 1, the bypass coding unit can be used, and if not, the regular coding unit can be used.

[0090] Various methods can be used to binarize the syntax elements. For example, fixed length binarization, unary binarization, truncated rice binarization, K-th Exp-Golomb binarization, etc. can be used. Also, signed or unsigned binarization can be performed depending on the range of values ​​of the syntax elements. The binarization process for the syntax elements generated in the present invention can be performed not only by the binarization mentioned in the above examples, but also by other additional binarization methods.

[0091] The inverse quantization unit and the inverse transform unit can be realized by performing the processes in the transform unit and the quantization unit inversely. For example, the inverse quantization unit can inverse quantize the quantized transform coefficients generated by the quantization unit, and the inverse transform unit can inverse transform the inverse quantized transform coefficients to generate a reconstructed residual block.

[0092] The adder adds the predicted block and the reconstructed residual block to reconstruct a current block. The reconstructed block is stored in a memory and can be used as reference data (for the predictor and filter, etc.).

[0093] The in-loop filter unit may include at least one post-processing filter process such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter may remove block distortion occurring at the boundary between blocks from the restored image. The ALF may perform filtering based on a value obtained by comparing a restored image with an input image. In particular, filtering may be performed based on a value obtained by comparing an input image with a restored image after a block is filtered through the deblocking filter. Alternatively, filtering may be performed based on a value obtained by comparing an input image with a restored image after a block is filtered through the SAO. The SAO restores an offset difference based on a value obtained by comparing a restored image with an input image, and may be applied in the form of a band offset (BO), an edge offset (EO), or the like. In particular, the SAO may be applied in the form of a BO, an EO, or the like by adding an offset with respect to the original image in at least one pixel unit to the restored image to which the deblocking filter is applied. In detail, an offset from the original image is added to the restored image after the block is filtered through ALF, and can be applied in the form of BO, EO, etc., in pixel units.

[0094] As the filtering-related information, setting information on whether to support each post-processing filter may be generated in units of a sequence, a picture, a slice, a tile, a block, etc. Also, setting information on whether to execute each post-processing filter may be generated in units of a picture, a slice, a tile, a block, etc. The range in which the filter is applied may be divided into an inside of an image and a boundary of an image, and setting information taking this into consideration may be generated. Also, information related to the filtering operation may be generated in units of a picture, a slice, a tile, a block, etc. The information may be subjected to implicit or explicit processing, and the filtering may be applied as an independent filtering process or a dependent filtering process according to color components. This may be determined according to the encoding setting. The in-loop filter unit may transmit the filtering-related information to the encoding unit to encode it, and may include the information in a bitstream and transmit it to a decoder, and the decoding unit of the decoder may parse the information and apply it to the in-loop filter unit.

[0095] The memory may store the reconstructed block or picture. The reconstructed block or picture stored in the memory may be provided to a prediction unit that performs intra-frame prediction or inter-frame prediction. In particular, a storage space in the form of a queue of bitstreams compressed in an encoder may be processed as a coded picture buffer (CPB), and a space for storing decoded images in picture units may be processed as a decoded picture buffer (DPB). In the case of the CPB, decoding units are stored according to the decoding order, and a decoding operation may be emulated in the encoder, and a bitstream compressed in the emulation process may be stored. The bitstream output from the CPB is restored through a decoding process, the restored image is stored in the DPB, and the picture stored in the DPB may be referred to in subsequent image encoding and decoding processes.

[0096] The decoder can be realized by performing the process inversely to that of the encoder, for example, by receiving a quantization coefficient sequence, a transform coefficient sequence, or a signal sequence from a bitstream, decoding the received sequence, and parsing the decoded data including the decoding information to transmit the parsed data to each component.

[0097] Hereinafter, an image setting process applied to an image encoding / decoding device according to an embodiment of the present invention will be described. This may be an example applied to a stage before encoding / decoding (image initial setting), but some of the processes may be examples that can be applied to other stages (e.g., a stage after encoding / decoding is performed or an internal stage of encoding / decoding, etc.). The image setting process may be performed in consideration of network and user environments such as characteristics of multimedia content, bandwidth, performance and accessibility of a user terminal. For example, image division, image resizing, image reconstruction, etc. can be performed according to the encoding / decoding setting. The image setting process described below will be mainly described for rectangular images, but is not limited thereto and can also be applied to polygonal images. Regardless of the shape of the image, the same image setting or different image settings can be applied. This can be determined according to the encoding / decoding setting. For example, after confirming information on the shape of the image (e.g., rectangular or non-rectangular shape), information on the image setting can be configured accordingly.

[0098] In the following examples, it is assumed that a setting is dependent on the color space, but it is also possible to set an independent setting for the color space. In addition, in the following examples, the independent setting may include an example of setting encoding / decoding settings independently for each color space, and even if one color space is described, it is assumed that an example applied to another color space (e.g., when M is generated for the luminance component, N is generated for the chrominance component) is included, and this can be derived. In addition, in the case of dependent setting, it may include an example of setting proportional to the composition ratio of the color format (e.g., 4:4:4, 4:2:2, 4:2:0, etc.) (e.g., when M is generated for the luminance component, M / 2 is generated for the chrominance component in the case of 4:2:0), and it is assumed that an example applied to each color space is included, and this can be derived, even if there is no special description. This is not limited to the above example, and may be a description commonly applied to the present invention.

[0099] Some configurations in the examples described below may be applicable to various encoding techniques such as encoding in the spatial domain, encoding in the frequency domain, block-based encoding, and object-based encoding.

[0100] Although it is common to encode / decode an input image as is, there are also cases where an image is divided and then encoded / decoded. For example, an image can be divided for the purpose of error resilience to prevent damage caused by packet damage during transmission. Alternatively, an image can be divided for the purpose of classifying areas with different properties within the same image according to the characteristics and type of the image.

[0101] In the present invention, the image division process may include a division process and an inverse process thereof. In the following examples, the division process will be mainly described, but the contents of the inverse process of division can be derived from the division process.

[0102] FIG. 3 is a diagram showing an example of dividing image information into layers for image compression.

[0103] 3a is an example diagram of an image sequence composed of multiple GOPs. One GOP can be composed of I pictures, P pictures, and B pictures as in 3b. One picture can be composed of slices, tiles, etc. as in 3c. Slices, tiles, etc. are composed of multiple basic coding units as in 3d, and the basic coding unit can be composed of at least one sub-coding unit as in FIG. 3e. The image setting process of the present invention will be described based on an example applied to units such as pictures, slices, tiles, etc. as in 3b and 3c.

[0104] FIG. 4 is a conceptual diagram illustrating various examples of image segmentation according to an embodiment of the present invention.

[0105] 4a is a conceptual diagram of an image (e.g., a picture) divided horizontally and vertically at regular intervals. The divided regions can be called blocks, and each block is a basic coding unit (or a maximum coding unit) obtained through a picture division unit, and can also be a basic unit applied to division units described later.

[0106] 4b is a conceptual diagram of an image divided in at least one of the horizontal and vertical directions. The divided regions T0 to T3 can be called tiles, and each region can be encoded / decoded independently or dependently of the other regions.

[0107] 4c is a conceptual diagram of an image divided into groups of consecutive blocks. The divided regions S0 and S1 can be called slices, and each region can be an area for performing encoding / decoding independently or dependently on other regions. The groups of consecutive blocks can be defined according to a scan order, which generally follows a raster scan order but is not limited thereto, and can be determined according to the encoding / decoding settings.

[0108] 4d is a conceptual diagram of an image divided into groups of blocks with arbitrary settings defined by the user. The divided areas A0 to A2 can be called arbitrary partitions, and each area can be an area for performing encoding / decoding independently or dependently on other areas.

[0109] Independent encoding / decoding may mean that data of other units cannot be referenced when encoding / decoding some units (or regions). In particular, information used or generated in texture encoding and entropy encoding of some units is encoded independently without reference to each other, and the decoder may not need to reference parsing information and reconstruction information of other units for texture decoding and entropy decoding of some units. In this case, whether data of other units (or regions) can be referenced may be restricted in the spatial domain (e.g., between regions in one image), but may also be restricted in the temporal domain (e.g., between consecutive images or frames) depending on the encoding / decoding settings. For example, if some units of a current image and some units of other images have continuity or the same encoding environment, they can be referenced, and if not, the reference may be restricted.

[0110] Also, dependent encoding / decoding may mean that data of another unit can be referenced when encoding / decoding a certain unit. In particular, information used or generated in texture encoding and entropy encoding of a certain unit is referenced to each other and encoded dependently, and similarly, parsing information and reconstruction information of another unit can be referenced to each other for texture decoding and entropy decoding of the certain unit in a decoder. That is, the setting may be the same as or similar to general encoding / decoding. In this case, the area (surface area generated according to the projection format in this example) may be determined according to the characteristics and type of the image (e.g., a 360-degree image). <face>Such a division may be made in order to identify a particular

[0111] In the above example, some units (slices, tiles, etc.) may have independent encoding / decoding settings (e.g., independent slice segments), and some units may have dependent encoding / decoding settings (e.g., dependent slice segments). In the present invention, the description will focus on independent encoding / decoding settings.

[0112] The basic coding unit obtained through the picture division unit as in 4a is divided into basic coding blocks according to the color space, and the size and shape can be determined according to the characteristics and resolution of the image. The supported block size or shape is a block whose width and height are powers of 2 (2 n ) is an N×N square (2 n ×2 n 256×256, 128×128, 64×64, 32×32, 16×16, 8×8, etc. n is an integer between 3 and 8) or an M×N rectangle (2 m ×2 n For example, depending on the resolution, an input image can be divided into sizes such as 128x128 for an 8k UHD image, 64x64 for a 1080p HD image, and 16x16 for a WVGA image, and depending on the type of image, an input image can be divided into sizes such as 256x256 for a 360-degree image. A basic coding unit can be divided into sub-coding units and encoded / decoded, and information on the basic coding unit can be recorded in a bitstream and transmitted in units of sequence, picture, slice, tile, etc. This can be parsed by the decoder to restore related information.

[0113] An image encoding method and a decoding method according to an embodiment of the present invention may include the following image segmentation steps. In this case, the image segmentation process may include an image segmentation instruction step, an image segmentation type identification step, and an image segmentation execution step. Also, the image encoding device and the decoding device may be configured to include an image segmentation instruction unit, an image segmentation type identification unit, and an image segmentation execution unit that realize the image segmentation instruction step, the image segmentation type identification step, and the image segmentation execution step. In the case of encoding, an associated syntax element may be generated, and in the case of decoding, the associated syntax element may be parsed.

[0114] In each block division process of 4a, the image division instruction unit can be omitted, and the image division type identification unit is a process of checking information regarding the size and shape of the block, and the image division unit can perform division in basic coding units based on the identified division type information.

[0115] In the case of blocks, they may always be the units for division, but for other division units (tiles, slices, etc.), whether or not to divide may be determined depending on the encoding / decoding settings. The picture division unit may have a default setting of dividing a picture into blocks and then dividing the picture into other units. In this case, block division may be performed based on the size of the picture.

[0116] Alternatively, the image may be divided into other units (such as tiles or slices) and then divided into blocks. That is, block division may be performed based on the size of the division unit. This may be determined through an explicit or implicit process depending on the encoding / decoding settings. In the example described below, the former case will be assumed and units other than blocks will be mainly described.

[0117] In the image division instruction step, it is possible to determine whether to perform image division. For example, if a signal (e.g., tiles_enabled_flag) instructing image division is confirmed, division can be performed, and if the signal instructing image division is not confirmed, division is not performed or division can be performed by checking other encoding / decoding information.

[0118] In detail, when a signal instructing image division (e.g., tiles_enabled_flag) is confirmed, if the corresponding signal is activated (e.g., tiles_enabled_flag=1), division can be performed in multiple units, and if the corresponding signal is inactivated (e.g., tiles_enabled_flag=0), division can be not performed. Alternatively, if the signal instructing image division is not confirmed, it may mean that division is not performed or that division is performed in at least one unit, and whether division is performed in multiple units can be confirmed via another signal (e.g., first_slice_segment_in_pic_flag).

[0119] In summary, when a signal for instructing image division is provided, the corresponding signal is a signal for indicating whether to divide into a plurality of units, and whether the corresponding image is divided can be confirmed according to the signal. For example, when tiles_enabled_flag is a signal indicating whether an image is divided, if tiles_enabled_flag is 1, it may mean that the image is divided into a plurality of tiles, and if tiles_enabled_flag is 0, it may mean that the image is not divided.

[0120] In summary, if no signal instructing image division is provided, division is not performed, or whether the image is divided can be confirmed by other signals. For example, first_slice_segment_in_pic_flag is not a signal indicating whether the image is divided, but a signal indicating whether it is the first slice segment in the image, and this makes it possible to confirm whether the image is divided into two or more units (for example, if the flag is 0, it means that the image is divided into multiple slices).

[0121] The present invention is not limited to the above example, and other modifications are possible. For example, a signal instructing image division in a tile may not be provided, and a signal instructing image division in a slice may be provided. Alternatively, a signal instructing image division may be provided according to the type, characteristics, etc. of the image.

[0122] In the image segmentation type identification step, the image segmentation type can be identified. The image segmentation type can be defined by the segmentation method, segmentation information, etc.

[0123] In 4b, a tile can be defined as a unit obtained by dividing the image horizontally and vertically, and more specifically, as a group of adjacent blocks within a rectangular space defined by at least one horizontal or vertical dividing line that crosses the image.

[0124] The division information for the tiles may include information on the boundary positions between rows and columns, information on the number of tiles in the rows and columns, and information on the size of the tiles. The information on the number of tiles may include the number of rows (e.g., num_tile_columns) and the number of columns (e.g., num_tile_rows) of the tiles, and thus the tiles may be divided into tiles of (number of rows x number of columns). The size information for the tiles may be obtained based on the information on the number of tiles, and the width or height of the tiles may be uniform or non-uniform. This may be implicitly determined under a preset rule, or related information (e.g., uniform_spacing_flag) may be explicitly generated. In addition, the size information for the tiles may include size information for each row and column of the tiles (e.g., column_width_tile[i], row_height_tile[i]), or may include information on the width and height of each tile. Furthermore, the size information may be information that can be further generated depending on whether the tile size is uniform or not (for example, when uniform_spacing_flag is 0, which means non-uniform division).

[0125] In 4c, a slice can be defined as a group of consecutive blocks, specifically, a group of consecutive blocks based on a predetermined scan order (in this example, raster scan).

[0126] The division information for a slice may include slice number information, slice position information (e.g., slice_segment_address), etc. In this case, the slice position information may be position information of a predetermined block (e.g., the first block in the scan order within the slice). In this case, the position information may be scan order information of the block.

[0127] In 4d, various division settings are possible for any division area.

[0128] A division unit in 4d can be defined as a group of spatially adjacent blocks, and the division information for this can include the size, shape, position information, etc. of the division unit. This is only a partial example for an arbitrary division region, and various division shapes are possible as shown in FIG.

[0129] FIG. 5 is a diagram illustrating another example of an image dividing method according to an embodiment of the present invention.

[0130] In the cases of 5a and 5b, an image can be divided into a plurality of regions at least at one block interval in the horizontal or vertical direction, and the division can be performed based on the position information of the blocks. 5a shows examples A0 and A1 in which division is performed based on column information of each block in the horizontal direction, and 5b shows examples B0 to B3 in which division is performed based on row and column information of each block in the vertical and horizontal directions. The division information for this can include the number of division units, block interval information, division direction, etc., and if these are implicitly included according to a predetermined rule, some division information may not be generated.

[0131] In the cases of 5c and 5d, an image may be divided into a group of consecutive blocks based on the scan order. Additional scan orders other than the existing slice raster scan order may be applied to image division. 5c shows examples C0 and C1 in which a clockwise or counterclockwise scan (Box-Out) is performed around a start block, and 5d shows examples D0 and D1 in which a vertical scan (Vertical) is performed around a start block. The division information may include information on the number of division units, position information of the division units (e.g., the first order in the scan order within the division unit), information on the scan order, etc., and if this is implicitly included according to a predetermined rule, some division information may not be generated.

[0132] In the case of 5e, an image can be divided by horizontal and vertical division lines. Existing tiles can have a rectangular space division shape by dividing by horizontal or vertical division lines, but division by division lines may not be possible. For example, a division example that crosses an image along a partial division line of an image (e.g., a division line that forms the right boundary of E1, E3, and E4 and the left boundary of E5) is possible, but a division example that crosses an image along a partial division line of an image (e.g., a division line that forms the lower boundary of E2 and E3 and the upper boundary of E4) is not possible. In addition, division can be performed based on block units (e.g., division after block division is performed first), or division can be performed by the horizontal or vertical division lines (e.g., division by the division line regardless of block division), and therefore each division unit may not be composed of an integer multiple of a block. For this reason, division information different from existing tiles can be generated, and the division information for this can include information on the number of division units, position information of the division units, size information of the division units, etc. For example, position information of a division unit can be generated based on a predetermined position (e.g., the top left corner of an image) as position information (e.g., measured in pixel units or block units), and size information of a division unit can be generated based on horizontal and vertical size information of each division unit (e.g., measured in pixel units or block units).

[0133] As in the above example, the division having any setting defined by the user may be performed by applying a new division method, or by modifying and applying a part of the existing division configuration. That is, the existing division method may be replaced or supported with an additional division shape, or the existing division method (slicing, tile, etc.) may be supported with some settings modified (e.g., according to a different scan order, or another quadrilateral division method and generation of other division information according to the same, dependent encoding / decoding characteristics, etc.). In addition, a setting for configuring an additional division unit (e.g., a setting other than division according to the scan order or division according to a certain interval difference) may be supported, and an additional division unit shape (e.g., a polygonal shape such as a triangle other than division into a quadrilateral space) may be supported. In addition, an image division method may be supported based on the type, characteristics, etc. of the image. For example, some division methods (e.g., the surface of a 360-degree image) may be supported based on the type, characteristics, etc. of the image, and division information may be generated based on the method.

[0134] In the image segmentation execution step, the image can be segmented based on the identified segmentation type information, i.e., segmentation can be performed into a plurality of segmentation units based on the identified segmentation type, and encoding / decoding can be performed based on the acquired segmentation units.

[0135] In this case, it can be determined whether the partition unit has an encoding / decoding setting according to the partition type. That is, setting information required for the encoding / decoding process of each partition unit can be assigned to a higher unit (e.g., a picture) or the partition unit can have an independent encoding / decoding setting.

[0136] In general, a slice may have an independent encoding / decoding setting for each division unit (e.g., slice header), whereas a tile may not have an independent encoding / decoding setting for each division unit, but may have a setting that is dependent on a picture encoding / decoding setting (e.g., PPS). In this case, information generated in relation to the tile may be partition information, which may be included in the picture encoding / decoding setting. The present invention is not limited to the above-mentioned case, and other modified examples are possible.

[0137] Encoding / decoding setting information for a tile may be generated in units of video, sequence, picture, etc., or at least one encoding / decoding setting information may be generated in a higher level unit and any one of them may be referenced. Alternatively, independent encoding / decoding setting information (e.g., a tile header) may be generated in units of tiles. This differs from one encoding / decoding setting determined in a higher level unit in that encoding / decoding is performed by setting at least one encoding / decoding setting in units of tiles. That is, all tiles may be able to comply with one encoding / decoding setting, or at least one tile may be able to perform encoding / decoding according to an encoding / decoding setting different from that of another tile.

[0138] Although the above examples have been used to focus on various encoding / decoding settings in tiles, this is not limiting and similar or identical settings can be placed on other division types.

[0139] As an example, for some division types, division information can be generated in higher units, and encoding / decoding can be performed according to one encoding / decoding setting of the higher unit.

[0140] As an example, for some division types, division information may be generated in a higher-level unit, and independent encoding / decoding settings for each division unit may be generated in the higher-level unit, thereby performing encoding / decoding.

[0141] As an example, for some division types, division information can be generated in higher-level units, multiple encoding / decoding setting information can be supported in higher-level units, and encoding / decoding can be performed according to the encoding / decoding settings referenced in each division unit.

[0142] For example, for some division types, division information may be generated in a higher-level unit, and independent encoding / decoding settings may be generated in the corresponding division units, thereby performing encoding / decoding.

[0143] For example, for some partition types, an independent encoding / decoding setting including partition information may be generated for each partition unit, and encoding / decoding may be performed based on the setting.

[0144] The encoding / decoding setting information may include information required for encoding / decoding a tile, such as a tile type, information on a reference picture list, quantization parameter information, inter-picture prediction setting information, in-loop filtering setting information, in-loop filtering control information, a scan order, and whether or not encoding / decoding is performed. The encoding / decoding setting information may explicitly generate related information, or the setting for encoding / decoding may be implicitly determined according to the format, characteristics, etc. of an image determined in a higher unit. Also, the related information may be explicitly generated based on information obtained in the setting.

[0145] An example of image division performed by an encoding / decoding device according to an embodiment of the present invention will be described below.

[0146] A division process can be performed on an input image before the start of encoding. After division is performed using division information (e.g., image division information, division unit setting information, etc.), the image can be encoded in division units. After the encoding is completed, it can be stored in a memory, and the image encoding data can be included in a bitstream and transmitted.

[0147] A division process can be performed before the start of decoding. After division is performed using division information (e.g., image division information, division unit setting information, etc.), the image decoded data can be parsed and decoded in division units. After the decoding is completed, it can be stored in memory, and multiple division units can be merged into one to output the image.

[0148] The above example is used to explain the image segmentation process. In addition, multiple segmentation processes can be performed in the present invention.

[0149] For example, a segmentation may be performed on an image, or a segmentation unit of an image may be segmented. The segmentation may be the same segmentation process (e.g., slice / slice, tile / tile, etc.) or different segmentation processes (e.g., slice / tile, tile / slice, tile / surface, surface / tile, slice / surface, surface / slice, etc.). In this case, a subsequent segmentation process may be performed based on a previous segmentation result. The segmentation information generated in the subsequent segmentation process may be generated based on the previous segmentation result.

[0150] Also, a plurality of division processes A may be performed, and the division processes may be different division processes (e.g., slice / surface, tile / surface, etc.). In this case, the subsequent division process may be performed based on the previous division result, or the division process may be performed independently regardless of the previous division result. Division information generated in the subsequent division process may be generated based on the previous division result or may be generated independently.

[0151] The multiple division processes of the image can be determined according to the encoding / decoding settings, and are not limited to the above example, and various modified examples are also possible.

[0152] The encoder records information generated in the above process in at least one unit of a sequence, a picture, a slice, a tile, etc., into a bitstream, and the decoder parses the related information from the bitstream. That is, the information can be recorded in one unit, or can be recorded in multiple units in a duplicated manner. For example, a syntax element for supporting or not supporting some information or a syntax element for activating or not can be generated in some units (e.g., upper units), and the same or similar information can be generated in some units (e.g., lower units). That is, even if related information is supported and set in a higher unit, it can have an individual setting in a lower unit. This is not limited to the above example, and may be a description commonly applied in the present invention. Also, the information can be included in the bitstream in the form of SEI or metadata.

[0153] Meanwhile, although it may be common to perform encoding / decoding as is an input image, a case may occur where the size of an image is adjusted (enlarged or reduced, adjusting the resolution) before encoding / decoding. For example, an image size adjustment such as overall expansion or reduction of an image may be performed in a hierarchical coding method (Scalability Video Coding) for supporting spatial, temporal, and image quality scalability. Alternatively, an image size adjustment such as partial expansion or reduction of an image may be performed. Image size adjustment may be performed for various purposes, and may be performed for the purpose of adaptability to a coding environment, for the purpose of coding uniformity, for the purpose of coding efficiency, for the purpose of improving image quality, or according to the type and characteristics of an image.

[0154] As a first example, the size adjustment process may be performed in a process (eg, hierarchical encoding, 360-degree image encoding, etc.) that is performed according to the characteristics, type, etc. of the image.

[0155] As a second example, the resizing process may be performed in an early stage of encoding / decoding, or before encoding / decoding. The image to be resized may be encoded / decoded.

[0156] As a third example, a size adjustment process may be performed during a prediction step (intra prediction or inter prediction) or before prediction is performed. In the size adjustment process, image information in the prediction step (e.g., pixel information referred to in intra prediction, intra prediction mode related information, reference image information used in inter prediction, inter prediction mode related information, etc.) may be used.

[0157] As a fourth example, a size adjustment process may be performed in the filtering step or before filtering. In the size adjustment process, image information in the filtering step (e.g., pixel information applied to a deblocking filter, pixel information applied to SAO, SAO filtering-related information, pixel information applied to ALF, ALF filtering-related information, etc.) may be used.

[0158] In addition, after the size adjustment process, the image may or may not be changed to the image before the size adjustment (in terms of image size) through a size adjustment inverse process. This can be determined according to the encoding / decoding settings (e.g., the nature of the size adjustment). In this case, if the size adjustment process is expansion, the size adjustment inverse process may be reduction, and if the size adjustment process is reduction, the size adjustment inverse process may be expansion.

[0159] When the size adjustment process according to the first to fourth examples is performed, a size adjustment inverse process can be performed in a subsequent step to obtain an image before the size adjustment.

[0160] When the size adjustment process according to hierarchical coding or the third example is performed (or when the size of the reference picture is adjusted in inter prediction), the inverse size adjustment process may not be performed in the subsequent steps.

[0161] In one embodiment of the present invention, the image resizing process may be performed alone or an inverse process may be performed thereto, and in the following example, the resizing process will be mainly described. Since the inverse resizing process is the inverse process of the resizing process, the description of the inverse resizing process may be omitted to avoid repetition, but it is clear that a person skilled in the art can understand it in the same way as if it were literally described.

[0162] FIG. 6 is an example diagram of a general image resizing method.

[0163] Referring to 6a, an expanded image P0+P1 can be obtained by further including a partial region P1 from the initial image (or image before size adjustment; P0; thick solid line).

[0164] 6b, a reduced image S0 can be obtained by excluding a partial region S1 from the initial image S0+S1.

[0165] Referring to 6c, a resized image T0+T1 can be obtained by further including a partial region T1 in the initial image T0+T2 and excluding a partial region T2.

[0166] Hereinafter, the present invention will be described focusing on the size adjustment process by expansion and the size adjustment process by reduction, but it should be understood that the present invention is not limited thereto and also includes a case where size expansion and reduction are mixed and applied as in 6c.

[0167] FIG. 7 is a diagram illustrating an example of image size adjustment according to an embodiment of the present invention.

[0168] Referring to 7a, it can be seen how to expand an image during the resizing process, and referring to 7b, it can be seen how to shrink an image.

[0169] In 7a, the image before resizing is S0 and the image after resizing is S1, and in 7b, the image before resizing is T0 and the image after resizing is T1.

[0170] When an image is expanded as in 7a, it can be expanded in the up, down, left and right directions (ET, EL, EB, ER), and when an image is reduced as in 7b, it can be reduced in the up, down, left and right directions (RT, RL, RB, RR).

[0171] When comparing image expansion and image contraction, the up, down, left, and right directions in expansion correspond to the down, up, right, and left directions in contraction, respectively. Therefore, the following description will be based on image expansion, but it should be understood that a description of image contraction is also included.

[0172] Also, although the following describes image expansion or contraction in the top, bottom, left, and right directions, it should be understood that size adjustments can also be performed in the top-left, top-right, bottom-left, and bottom-right directions.

[0173] In this case, when extending to the lower right, the RC and BC regions are acquired, while the BR region may or may not be acquired depending on the encoding / decoding settings. That is, the TL, TR, BL, and BR regions may or may not be acquired, but for the sake of convenience, the following description will be given assuming that the corner regions (TL, TR, BL, and BR regions) can be acquired.

[0174] The image resizing process according to an embodiment of the present invention may be performed in at least one direction, for example, in all of the up, down, left and right directions, in two or more directions selected from the up, down, left and right directions (left+right, up+down, up+left, up+right, down+left, down+right, up+left+right, down+left+right, up+down+left, up+down+right, etc.), or in only one of the up, down, left and right directions.

[0175] For example, the size may be adjusted so that the image can be expanded symmetrically on both sides based on the center of the image, in the left+right, top+bottom, upper left+bottom right, or lower left+upper right directions; the size may be adjusted so that the image can be expanded symmetrically vertically, in the left+right, upper left+top right, or lower left+bottom right directions; the size may be adjusted so that the image can be expanded symmetrically horizontally, in the top+bottom, upper left+bottom left, or upper right+bottom right directions; and other size adjustments are also possible.

[0176] In 7a and 7b, the size of the image (S0, T0) before size adjustment is defined as P_Width (width) x P_Height (height), and the size of the image (S1, T1) after size adjustment is defined as P'_Width (width) x P'_Height (height). If the size adjustment values ​​for the left, right, top, and bottom directions are defined as Var_L, Var_R, Var_T, and Var_B (or collectively referred to as Var_x), the image size after size adjustment can be expressed as (P_Width+Var_L+Var_R) x (P_Height+Var_T+Var_B). In this case, the size adjustment values ​​Var_L, Var_R, Var_T, and Var_B in the left, right, top, and bottom directions can be Exp_L, Exp_R, Exp_T, and Exp_B (in this example, Exp_x is a positive number) in image expansion (FIG. 7a), and can be -Rec_L, -Rec_R, -Rec_T, and -Rec_B in image reduction (when Rec_L, Rec_R, Rec_T, and Rec_B are defined as positive numbers, they can be expressed as negative numbers depending on the image reduction). In addition, the coordinates of the top left, top right, bottom left, and bottom right of the image before size adjustment are (0, 0), (P_Width-1, 0), (0, P_Height-1), (P_Width-1, P_Height-1), and the coordinates of the image after size adjustment can be expressed as (0, 0), (P'_Width-1, 0), (0, P'_Height-1), (P'_Width-1, P'_Height-1). The size of the area (TL to BR in this example, i is an index that divides TL to BR) that is changed (or acquired, deleted) by size adjustment can be M[i]×N[i]. This can be expressed as Var_X×Var_Y (assuming that X is L or R and Y is T or B in this example). M and N can have various values ​​and can be the same regardless of i, or can have individual settings depending on i. Various cases regarding this will be described later.

[0177] Referring to 7a, S1 can be configured to include all or a part of TL-BR (upper left to lower right) generated by expanding S0 in various directions. Referring to 7b, T1 can be configured to exclude all or a part of TL-BR removed by shrinking T0 in various directions.

[0178] In 7a, when an existing image S0 is expanded in the upward, downward, left, and right directions, the image can be constructed to include the TC, BC, LC, and RC regions obtained through each size adjustment process, and can also include the TL, TR, BL, and BR regions.

[0179] As an example, when extension is performed in the upward (ET) direction, an image can be constructed by including a TC region in an existing image S0, and can include a TL or TR region depending on at least one different direction of extension (EL or ER).

[0180] As an example, when expanding in the downward (EB) direction, an image can be constructed by including a BC region in an existing image S0, and can include a BL or BR region depending on the expansion in at least one different direction (EL or ER).

[0181] As an example, when performing an extension in the left (EL) direction, an image can be constructed by including an LC region in an existing image S0, and can include a TL or BL region depending on at least one different direction of extension (ET or EB).

[0182] As an example, when extension is performed in the right (ER) direction, an image can be constructed by including an RC region in an existing image S0, and a TR or BR region can be included depending on at least one different direction of extension (ET or EB).

[0183] According to one embodiment of the present invention, a setting (e.g., spa_ref_enabled_flag or tem_ref_enabled_flag) can be placed that can spatially or temporally restrict the visibility of the region being resized (assumed to be extended in this example).

[0184] That is, it is possible to reference data of areas that are spatially or temporally sized depending on the encoding / decoding settings (e.g., spa_ref_enabled_flag=1 or tem_ref_enabled_flag=1) or to restrict references (e.g., spa_ref_enabled_flag=0 or tem_ref_enabled_flag=0).

[0185] The encoding / decoding of the image before resizing (S0, T1) and the areas added or deleted during resizing (TC, BC, LC, RC, TL, TR, BL, BR areas) may be performed as follows.

[0186] For example, when encoding / decoding an image before size adjustment and an area to be added or deleted, the data of the image before size adjustment and the data of the area to be added or deleted (encoded / decoded data, such as pixel values ​​or prediction-related information) can be referenced to each other spatially or temporally.

[0187] Alternatively, while the image before resizing and the data of the area to be added or deleted can be referenced spatially, the data of the image before resizing can be referenced temporally and the data of the area to be added or deleted cannot be referenced temporally.

[0188] That is, settings can be placed that limit the visibility of the region being added or deleted. The configuration information about the visibility of the region being added or deleted can be explicitly generated or implicitly determined.

[0189] An image size adjustment process according to an embodiment of the present invention may include an image size adjustment indication step, an image size adjustment type identification step, and / or an image size adjustment execution step. Also, an image encoding device and a decoding device may include an image size adjustment indication unit, an image size adjustment type identification unit, and an image size adjustment execution unit that realize the image size adjustment indication step, the image size adjustment type identification step, and the image size adjustment execution step. In the case of encoding, an associated syntax element may be generated, and in the case of decoding, the associated syntax element may be parsed.

[0190] In the image resizing instruction step, it is possible to determine whether to perform image resizing. For example, if a signal instructing image resizing (e.g., img_resizing_enabled_flag) is confirmed, resizing can be performed, and if the signal instructing image resizing is not confirmed, resizing may not be performed or resizing may be performed by checking other encoding / decoding information. Even if a signal instructing image resizing is not provided, the signal instructing size adjustment may be implicitly activated or inactivated according to encoding / decoding settings (e.g., image characteristics, type, etc.), and if size adjustment is performed, size adjustment related information may be generated accordingly, or size adjustment related information may be implicitly determined.

[0191] When a signal instructing image resizing is provided, the corresponding signal is a signal indicating whether or not to resize the image, and it is possible to check whether or not the corresponding image is to be resized according to the signal.

[0192] For example, a signal instructing image resizing (e.g., img_resizing_enabled_flag) is checked, and if the corresponding signal is activated (e.g., img_resizing_enabled_flag=1), it means that image resizing can be performed, and if the corresponding signal is deactivated (e.g., img_resizing_enabled_flag=0), it means that image resizing is not performed.

[0193] Also, if a signal instructing image resizing is not provided, the size may not be adjusted, or it may be possible to check whether the image needs resizing based on another signal.

[0194] For example, when an input image is divided into blocks, size adjustment (in this example, in the case of expansion; it is assumed that the size adjustment process is performed when the image size is not an integer multiple of the block size (e.g., width or height) can be performed depending on whether the image size (e.g., width or height) is an integer multiple of the block size (e.g., width or height). That is, size adjustment can be performed when the image width is not an integer multiple of the block width or when the image height is not an integer multiple of the block height. At this time, size adjustment information (e.g., size adjustment direction, size adjustment value, etc.) can be determined according to the encoding / decoding information (e.g., image size, block size, etc.). Alternatively, size adjustment can be performed according to the characteristics and type of image (e.g., 360-degree image), and the size adjustment information can be explicitly generated or assigned to a predetermined value. The present invention is not limited to the above example, and other modifications are also possible.

[0195] In the image size adjustment type identification step, the image size adjustment type can be identified. The image size adjustment type can be defined by a size adjustment method, size adjustment information, etc. For example, size adjustment using a scale factor, size adjustment using an offset factor, etc. can be performed. The present invention is not limited to this, and the above methods can be mixed and applied. For convenience of explanation, size adjustment using a scale factor and an offset factor will be mainly described.

[0196] In the case of a scale factor, size adjustment may be performed in a multiplication or division manner based on the size of the image. Information about the size adjustment operation (e.g., expansion or contraction) may be explicitly generated, and the expansion or contraction process may be performed according to the corresponding information. Also, the size adjustment process may be performed with a predetermined operation (e.g., either expansion or contraction) according to the encoding / decoding settings, in which case information about the size adjustment operation may be omitted. For example, if image size adjustment is activated in the image size adjustment instruction step, the image size adjustment may be performed with a predetermined operation.

[0197] The size adjustment direction may be at least one direction selected from the top, bottom, left, and right directions. At least one scale factor may be required depending on the size adjustment direction. That is, one scale factor may be required for each direction (unidirectional in this example), one scale factor may be required depending on the horizontal or vertical direction (bidirectional in this example), and one scale factor may be required depending on the overall direction of the image (all directions in this example). In addition, the size adjustment direction is not limited to the above example, and may be modified to other examples.

[0198] The scale factor can have a positive value and can be set with different range information depending on the encoding / decoding settings. For example, when generating information by mixing a resizing operation and a scale factor, the scale factor can be used as a multiplier value. If it is greater than 0 or less than 1, it can mean a shrinking operation, if it is greater than 1, it can mean an expanding operation, and if it is 1, it can mean no resizing. As another example, when generating scale factor information separately from a resizing operation, in the case of an expanding operation, the scale factor can be used as a multiplier value, and in the case of a shrinking operation, the scale factor can be used as a divisor value.

[0199] Referring again to FIGS. 7a and 7b, the process of using a scale factor to change an unscaled image (S0, T0) to a scaled image (S1, T1 in this example) can be explained.

[0200] As an example, if one scale factor (called sc) is used according to the overall orientation of the image, and the resizing direction is down+right, the resizing direction is ER, EB (or RR, RB), the resizing values ​​Var_L (Exp_L or Rec_L) and Var_T (Exp_T or Rec_T) are 0, and Var_R (Exp_R or Rec_R) and Var_B (Exp_B or Rec_B) can be expressed as P_Width×(sc-1), P_Height×(sc-1). Thus, the resizing image can be (P_Width×sc)×(P_Height×sc).

[0201] As an example, if a scale factor (in this example, sc_w, sc_h) is used according to the horizontal or vertical direction of the image, and the resizing direction is left+right, up+down (up+down+left+right when two are operated), the resizing direction is ET, EB, EL, ER, and the resizing values ​​Var_T and Var_B can be P_Height×(sc_h-1) / 2, and Var_L and Var_R can be P_Width×(sc_w-1) / 2. Therefore, the resizing image can be (P_Width×sc_w)×(P_Height×sc_h).

[0202] In the case of an offset factor, the size adjustment can be performed by adding or subtracting based on the size of the image, or by adding or subtracting based on the encoding / decoding information of the image, or by independently adding or subtracting. That is, the size adjustment process can be set dependently or independently.

[0203] Information about the resizing operation (e.g., expansion or contraction) can be explicitly generated, and the expansion or contraction process can be performed according to the corresponding information. Also, the resizing operation can be performed in a predetermined operation (e.g., either expansion or contraction) according to the encoding / decoding settings, in which case the information about the resizing operation can be omitted. For example, if image resizing is activated in the image resizing instruction step, the image resizing can be performed in a predetermined operation.

[0204] The size adjustment direction may be at least one of the up, down, left, and right directions. At least one offset factor may be required depending on the size adjustment direction. That is, one offset factor may be required for each direction (in this example, unidirectional), either offset factor may be required depending on the horizontal or vertical direction (in this example, symmetrical bidirectional), one offset factor may be required depending on a partial combination of each direction (in this example, asymmetrical bidirectional), and one offset factor may be required depending on the overall direction of the image (in this example, omnidirectional). In addition, the size adjustment direction is not limited to only the above example, and may be modified to other examples.

[0205] The offset factor can have a positive value or a positive and negative value, and can be set with different range information depending on the encoding / decoding settings. For example, when the size adjustment operation and the offset factor are mixed to generate information (assumed to have positive and negative values ​​in this example), the offset factor can be used as a value to be added or subtracted depending on the sign information of the offset factor. When the offset factor is larger than 0, it can mean an expansion operation, when it is smaller than 0, it can mean a reduction operation, and when it is 0, it can mean that no size adjustment is performed. As another example, when offset factor information is generated separately from the size adjustment operation (assumed to have a positive value in this example), the offset factor can be used as a value to be added or subtracted depending on the size adjustment operation. When it is larger than 0, it can mean that an expansion or reduction operation is performed depending on the size adjustment operation, and when it is 0, it can mean that no size adjustment is performed.

[0206] Referring again to FIGS. 7a and 7b, it can be seen how an offset factor is used to change from an unresized image (S0, T0) to a resized image (S1, T1).

[0207] As an example, if one offset factor (called os) is used according to the overall orientation of the image, and the resizing direction is up+down+left+right, the resizing direction may be ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values ​​Var_T, Var_B, Var_L, Var_R may be os. The resized image size may be (P_Width+os)×(P_Height+os).

[0208] As an example, if the offset factors (os_w, os_h) are used according to the horizontal or vertical direction of the image, and the resizing direction is left+right or up+down (up+down+left+right when both are active), the resizing direction is ET, EB, EL, ER (or RT, RB, RL, RR), the resizing values ​​Var_T, Var_B can be os_h, and Var_L, Var_R can be os_w. The resized image size can be {P_Width+(os_w×2)}×{P_Height+(os_h×2)}.

[0209] As an example, if the resize direction is down and right (down+right when working together), and the offset factors (os_b, os_r) are used according to the resize direction, the resize direction is EB, ER (or RB, RR), the resize value Var_B can be os_b, and Var_R can be os_r. The resized image size can be (P_Width+os_r)×(P_Height+os_b).

[0210] As an example, if the offset factors (os_t, os_b, os_l, os_r) are used for each direction of the image, and the resizing directions are up, down, left, and right (all working, up+down+left+right), the resizing directions are ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values ​​Var_T can be os_t, Var_B can be os_b, Var_L can be os_l, and Var_R can be os_r. The resized image size can be (P_Width+os_l+os_r)×(P_Height+os_t+os_b).

[0211] The above example shows a case where the offset factors are used as size adjustment values ​​(Var_T, Var_B, Var_L, Var_R) in the size adjustment process. That is, it means a case where the offset factors are used as size adjustment values ​​as they are. This may be an example of size adjustment performed independently. Alternatively, the offset factors may be used as input variables of the size adjustment values. In particular, the offset factors may be assigned as input variables, and the size adjustment values ​​may be obtained through a series of processes according to the encoding / decoding settings. This may be an example of size adjustment performed based on predetermined information (e.g., image size, encoding / decoding information, etc.), or an example of size adjustment performed dependently.

[0212] For example, the offset factor may be a multiple (e.g., 1, 2, 4, 6, 8, 16, etc.) or exponent (e.g., an exponent of 2 such as 1, 2, 4, 8, 16, 32, 64, 128, 256, etc.) of a predetermined value (integer in this example). Or, it may be a multiple or exponent of a value obtained based on the encoding / decoding setting (e.g., a value set based on the motion search range of inter prediction). Or, it may be a multiple or exponent of a unit obtained from the picture division unit (assumed to be A×B in this example). Or, it may be a multiple of a unit obtained from the picture division unit (assumed to be E×F in the case of tiles, etc. in this example).

[0213] Alternatively, the value may be smaller than or equal to the width and height of the unit obtained from the picture division unit. The multiples or exponents in the above examples may include a case where the value is 1, and are not limited to the above examples and may be modified to other examples. For example, if the offset factor is n, Var_x is 2×n or 2 n It could be.

[0214] Also, individual offset factors can be supported according to color components, and offset factors for some color components can be supported to derive offset factor information for other color components. For example, when an offset factor A for a luminance component (assuming that the composition ratio of the luminance component to the chrominance component is 2:1 in this example) is explicitly generated, an offset factor A / 2 for a chrominance component can be implicitly obtained. Or, when an offset factor A for a chrominance component is explicitly generated, an offset factor 2A for a luminance component can be implicitly obtained.

[0215] Information on the size adjustment direction and the size adjustment value may be explicitly generated, and the size adjustment process may be performed according to the corresponding information. Also, the size adjustment value may be implicitly determined according to the encoding / decoding setting, and the size adjustment process may be performed according to the information. At least one predetermined direction or adjustment value may be assigned, and in this case, the related information may be omitted. In this case, the encoding / decoding setting may be determined based on the characteristics, type, encoding information, etc. of the image. For example, at least one size adjustment direction according to at least one size adjustment operation may be predetermined, at least one size adjustment value according to at least one size adjustment operation may be predetermined, and at least one size adjustment value according to at least one size adjustment direction may be predetermined. Also, the size adjustment direction and size adjustment value in the size adjustment reverse process may be derived from the size adjustment direction and size adjustment value applied in the size adjustment process. In this case, the size adjustment value determined implicitly may be one of the above-mentioned examples (examples in which various size adjustment values ​​are obtained).

[0216] Also, in the above example, the case of multiplication or division has been described, but it can be realized by a shift operation depending on the implementation of the encoder / decoder. Multiplication can be realized by a left shift operation, and division can be realized by a right shift operation. This is not limited to the above example, and can be a description commonly applied to the present invention.

[0217] In the image resizing execution step, image resizing can be performed based on the identified resizing information, i.e., based on information such as resizing type, resizing operation, resizing direction, and resizing value, and encoding / decoding can be performed based on the obtained resizing-adjusted image.

[0218] In addition, in the image resizing step, the resizing may be performed using at least one data processing method. In particular, the resizing may be performed using at least one data processing method for the area to be resized according to the resizing type and the resizing operation. For example, according to the resizing type, it may be determined how to fill data when the resizing is an expansion, or it may be determined how to remove data when the resizing process is a reduction.

[0219] In summary, image size adjustment may be performed based on size adjustment information identified in the image size adjustment execution step. Alternatively, image size adjustment may be performed based on the size adjustment information and a data processing method in the image size adjustment execution step. The difference between the two cases is whether only the size of the image to be encoded / decoded is adjusted, or whether the size of the image and data processing of the area to be resized are also taken into consideration. Whether or not the data processing method is included in the image size adjustment execution step may be determined depending on the application step, position, etc. of the size adjustment process. In the following example, an example of performing size adjustment based on a data processing method will be mainly described, but is not limited to this.

[0220] When performing size adjustment using an offset factor, size adjustment can be performed using various methods in the cases of expansion and reduction. In the case of expansion, size adjustment can be performed using a method of filling at least one piece of data, and in the case of reduction, size adjustment can be performed using a method of removing at least one piece of data. In this case, in the case of size adjustment using an offset factor, new data or existing image data can be filled directly or after transformation into the size adjustment area (expansion), and data can be removed by applying simple removal or removal through a series of processes to the size adjustment area (reduction).

[0221] When performing resizing using a scale factor, in some cases (e.g., hierarchical coding, etc.), the expansion may perform resizing by applying upsampling, and the contraction may perform resizing by applying downsampling. For example, in the case of expansion, at least one upsampling filter may be used, and in the case of contraction, at least one downsampling filter may be used, and the filters applied horizontally and vertically may be the same or different. In this case, in the case of resizing using a scale factor, new data is not generated or removed in the resizing area, but existing image data may be rearranged using a method such as interpolation. Data processing methods related to the execution of resizing can be classified according to the filter used for the sampling. In addition, in some cases (e.g., similar to an offset factor), the expansion may perform resizing using a method of filling at least one data, and the contraction may perform resizing using a method of removing at least one data. In the present invention, a data processing method when performing resizing using an offset factor will be mainly described.

[0222] Generally, one predetermined data processing method can be used for a region to be resized, but as in the example described below, at least one data processing method can be used for a region to be resized, and selection information for the data processing method can be generated. In the former case, it can mean performing size resizing through a fixed data processing method, and in the latter case, it can mean performing size resizing through an adaptive data processing method.

[0223] In addition, a data processing method common to the entire area to be added or deleted during size adjustment (TL, TC, TR, ..., BR in Figures 7a and 7b) can be applied, or a data processing method can be applied to a portion of the area to be added or deleted during size adjustment (for example, each of TL to BR in Figures 7a and 7b or a partial combination thereof).

[0224] FIG. 8 is an exemplary diagram illustrating a method for configuring an area to be expanded in an image resizing method according to an embodiment of the present invention.

[0225] Referring to 8a, for convenience of explanation, an image can be divided into TL, TC, TR, LC, C, RC, BL, BC, and BR regions, which can correspond to the top left, top, top right, left, center, right, bottom left, bottom, and bottom right positions of the image, respectively. In the following, a case where an image is expanded in the bottom+right direction will be described, but it should be understood that the same can be applied to other directions.

[0226] The area added in response to the expansion of the image can be configured in various ways, for example it can be filled with an arbitrary value or it can be filled by referencing some data of the image.

[0227] Referring to FIG. 8b, it is possible to fill the extended region (A0, A2) to an arbitrary pixel value. The arbitrary pixel value can be determined using various methods.

[0228] As an example, a pixel value may be a pixel that belongs to a pixel value range {e.g., from 0 to 1<<(bit_depth)-1} that can be expressed by a bit depth, such as a minimum value, a maximum value, or a median value {e.g., 1<<(bit_depth-1), etc.} of the pixel value range (where bit_depth is the bit depth).

[0229] As an example, any pixel value can be expressed as a range of pixel values ​​{e.g., min P From max P Until. min P , max P are the minimum and maximum values ​​of the pixels in the image. P is greater than or equal to 0, max P may be a pixel in the range {1<<(bit_depth)-1 or less}. For example, a given pixel value may be a minimum, maximum, median, average (of at least two pixels), weighted sum, etc. of the range of pixel values.

[0230] As an example, the arbitrary pixel value may be a value determined within a range of pixel values ​​belonging to a partial region of an image. For example, when A0 is configured, the partial region may be TR+RC+BR. Also, the partial region may be 3×9 of TR, RC, and BR as the corresponding region, or 1×9 <assumed to be the rightmost line> as the corresponding region. This may depend on the encoding / decoding settings. In this case, the partial region may be a unit divided by the picture division unit. Specifically, the arbitrary pixel value may be the minimum value, maximum value, median value, average (of at least two pixels), weighted sum, etc. of the pixel value range.

[0231] Referring again to 8b, the area A1 added in response to the expansion of the image can be filled with pattern information generated using a plurality of pixel values ​​(for example, a pattern is assumed to use a plurality of pixels. It does not necessarily have to follow a certain rule.) In this case, the pattern information can be defined according to the encoding / decoding settings or related information can be generated, and at least one piece of pattern information can be used to fill the expanded area.

[0232] Referring to 8c, the area to be added according to the expansion of the image may be configured by referring to pixels of a part of the area belonging to the image. In particular, the area to be added may be configured by copying or padding pixels (hereinafter, reference pixels) of an area adjacent to the area to be added. In this case, the pixels of the area adjacent to the area to be added may be pixels before encoding or pixels after encoding (or decoding). For example, when performing size adjustment in a pre-encoding stage, the reference pixels may refer to pixels of an input image, and when performing size adjustment in an intra-screen prediction reference pixel generation stage, a reference image generation stage, a filtering stage, etc., the reference pixels may refer to pixels of a restored image. In this example, it is assumed that the pixels most adjacent to the area to be added are used, but this is not limited thereto.

[0233] The area A0 extended left or right in relation to the horizontal size adjustment of the image can be configured by padding (Z0) the outer pixels adjacent to the extended area A0 in the horizontal direction, the area A1 extended up or down in relation to the vertical size adjustment of the image can be configured by padding (Z1) the outer pixels adjacent to the extended area A1 in the vertical direction, and the area A2 extended to the lower right can be configured by padding (Z2) the outer pixels adjacent to the extended area A2 in the diagonal direction.

[0234] With reference to 8d, the data B0-B2 of a partial area belonging to the image can be referenced to configure the extended area B'0-B'2. 8d can be distinguished from 8c in that it can reference an area that is not adjacent to the extended area.

[0235] For example, if an area with high correlation with the area to be expanded exists in the image, the area to be expanded can be filled with reference to the pixels of the area with high correlation. At this time, position information, area size information, etc. of the area with high correlation can be generated. Alternatively, if an area with high correlation exists through encoding / decoding information such as image characteristics and type, and the position information, size information, etc. of the area with high correlation can be implicitly confirmed (for example, in the case of a 360-degree image), the data of the area can be filled into the area to be expanded. At this time, the position information, area size information, etc. of the area can be omitted.

[0236] As an example, in the case of a region B'2 that is expanded in the left or right direction associated with the horizontal resizing of an image, the expanded region can be filled by reference to pixels in the region B2 opposite the expanded region in the left or right direction associated with the horizontal resizing.

[0237] As an example, in the case of an area B'1 that is expanded in the upward or downward direction related to the vertical size adjustment of an image, the expanded area can be filled by referring to pixels in the area B1 opposite the expanded area in the upward or right direction related to the vertical size adjustment.

[0238] As an example, in the case of an area B'0 that is expanded by adjusting the size of a portion of an image (in this example, diagonally from the center of the image), the expanded area can be filled by referring to pixels in the opposite area B0, TL of the expanded area.

[0239] In the above example, we have described a case where data is obtained from an area where there is continuity at the boundaries at both ends of the image and which is located symmetrically in the size adjustment direction, but this is not limited to this and it is also possible for data to be obtained from other areas TL to BR.

[0240] When filling the area to be expanded with data of a part of an image, the data of the area can be filled by copying the data of the area as it is, or the data of the area can be filled after a conversion process based on the characteristics and type of the image. In this case, if the data is copied as it is, it may mean that the pixel values ​​of the area are used as they are, and if the conversion process is performed, it may mean that the pixel values ​​of the area are not used as they are. That is, through the conversion process, at least one pixel value of the area may change and be filled in the area to be expanded, or at least one acquisition position of some pixels may differ. That is, to fill the area A×B to be expanded, data of the area C×D may be used instead of data of A×B. In other words, at least one motion vector applied to the pixels to be filled may differ. The above example may be an example that occurs when a 360-degree image is composed of multiple surfaces according to the projection format and data of other surfaces is used to fill the area to be expanded. The data processing method for filling the area to be expanded due to image size adjustment is not limited to the above example, and an improved and modified data processing method or an additional data processing method may be used.

[0241] A set of candidates for a plurality of data processing methods can be supported according to the encoding / decoding settings, and data processing method selection information can be generated from the plurality of candidates and recorded in the bitstream. For example, one data processing method can be selected from a method of filling using a predetermined pixel value, a method of filling by copying outer pixels, a method of filling by copying a part of an image, a method of filling by transforming a part of an image, etc., and selection information for this method can be generated. Also, the data processing method can be determined implicitly.

[0242] For example, the data processing method applied to the entire area (TL to BR in FIG. 7a in this example) expanded by resizing the image may be any one of a method of filling using a predetermined pixel value, a method of filling by copying outer pixels, a method of filling by copying a partial area of ​​the image, a method of filling by transforming a partial area of ​​the image, and other methods, and selection information for the method may be generated. Also, a predetermined data processing method to be applied to the entire area may be determined.

[0243] Alternatively, the data processing method applied to the area expanded by resizing the image (in this example, each of the areas TL to BR in 7a of FIG. 7 or two or more areas therein) may be any one of a method of filling using a predetermined pixel value, a method of filling by copying outer pixels, a method of filling by copying a partial area of ​​the image, a method of filling by transforming a partial area of ​​the image, and other methods, and selection information for this method may be generated. Also, a predetermined data processing method to be applied to at least one area may be determined.

[0244] FIG. 9 is an exemplary diagram showing a method for configuring areas to be deleted and areas to be created by reducing an image size in an image size adjustment method according to an embodiment of the present invention.

[0245] The areas to be deleted during the image reduction process may not only be simply deleted, but may also be deleted after going through a series of exploitation processes.

[0246] 9a, some areas A0, A1, and A2 can be simply removed without additional processing during the image reduction process. In this case, the image A is divided into smaller areas and may be referred to as TL to BR, as in FIG. 8a.

[0247] Referring to 9b, a portion of the regions A0 to A2 are removed, but can be used as reference information when encoding / decoding image A. For example, the portion of the region A0 to A2 to be removed can be used in a restoration or correction process of a portion of the image A that is generated by reducing the size. The restoration or correction process can use a weighted sum or average of two regions (the region to be deleted and the region generated). Furthermore, the restoration or correction process can be a process that can be applied when two regions have a high correlation.

[0248] As an example, an area B'2 that is deleted by shrinking in the left or right direction associated with the horizontal size adjustment of an image can be used to restore or correct pixels in the area B2, LC opposite the area being reduced in the left or right direction associated with the horizontal size adjustment, and then the corresponding area can be removed from memory.

[0249] As an example, the area B'1 that is deleted in the upward or downward direction related to the vertical size adjustment of the image can be used in the encoding / decoding process (restoration or correction process) of the area B1, TR opposite the area to be reduced in the upward or downward direction related to the vertical size adjustment, and then the corresponding area can be removed from the memory.

[0250] As an example, when a portion of an image is resized (in this example, diagonally from the center of the image), an area B'0 is reduced, and the area B0 opposite the reduced area is used in the encoding / decoding process (such as the restoration or correction process) of TL, and then the corresponding area can be removed from the memory.

[0251] In the above example, we have explained the case where continuity exists at the boundaries at both ends of the image and the data is used to restore or correct an area that is located symmetrically to the size adjustment direction, but this is not limited to this, and the data may also be removed from memory after being used to restore or correct data of other areas TL to BR that are not located symmetrically.

[0252] The data processing method for removing the reduced area of ​​the present invention is not limited to the above example, and may be improved or modified, or additional data processing methods may be used.

[0253] A set of candidates for multiple data processing methods can be supported according to the encoding / decoding settings, and selection information for the candidates can be generated and recorded in the bitstream. For example, one data processing method can be selected from a method of simply removing the area to be resized, a method of using the area to be resized in a series of processes and then removing it, and selection information for the method can be generated. Also, the data processing method can be determined implicitly.

[0254] For example, the data processing method applied to the entire area (TL to BR in FIG. 7b in this example) that is deleted by reducing the size of the image can be one of a simple removal method, a removal method after using it in a series of processes, or other methods, and selection information for this can be generated. Also, the data processing method can be implicitly determined.

[0255] Alternatively, the data processing method applied to the individual regions (TL to BR in FIG. 7b in this example) that are reduced by resizing the image may be one of a simple removal method, a removal method after using it in a series of processes, or other methods, and selection information for this may be generated. Also, the data processing method may be implicitly determined.

[0256] In the above example, a case was described where size adjustment was performed by size adjustment operations (expansion, reduction), but in some cases, this may be an example that can be applied to a case where a size adjustment operation (in this example, expansion) is performed and then a size adjustment operation (in this example, reduction), which is the reverse process.

[0257] For example, a method of filling an area to be expanded with partial image data is selected, and a method of removing an area to be reduced in the inverse process after using the area to be reduced in a partial image data restoration or correction process is selected. Alternatively, a method of filling an area to be expanded with copies of outer pixels is selected, and a method of simply removing an area to be reduced in the inverse process is selected. In other words, the data processing method in the inverse process can be determined based on the data processing method selected in the image size adjustment process.

[0258] Unlike the above example, the data processing methods of the image resizing process and the inverse process can have an independent relationship. That is, the data processing method in the inverse process can be selected regardless of the data processing method selected in the image resizing process. For example, a method of filling an area to be expanded with partial image data can be selected, and a method of simply removing an area to be reduced in the inverse process can be selected.

[0259] In the present invention, the data processing method in the image size adjustment process can be implicitly determined according to the encoding / decoding settings, and the data processing method in the inverse process can be implicitly determined according to the encoding / decoding settings. Alternatively, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the inverse process can be explicitly generated. Alternatively, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the inverse process can be implicitly determined based on the data processing method.

[0260] Next, an example of performing image size adjustment in an encoding / decoding device according to an embodiment of the present invention will be described. In the following example, a case will be described in which the size adjustment process is expansion and the inverse size adjustment process is reduction. In addition, the difference between the "image before size adjustment" and the "image after size adjustment" may refer to the size of the image, and the size adjustment related information may be partially explicitly generated or partially determined implicitly depending on the encoding / decoding settings. In addition, the size adjustment related information may include information on the size adjustment process and the inverse size adjustment process.

[0261] As a first example, a size adjustment process can be performed on an input image before encoding starts. The size adjustment can be performed using size adjustment information (e.g., a size adjustment operation, a size adjustment direction, a size adjustment value, a data processing method, etc. The data processing method is used in the size adjustment process), and the size-adjusted image can be encoded. After encoding is completed, the image can be stored in a memory, and the image encoding data (meaning the size-adjusted image in this example) can be included in a bitstream and transmitted.

[0262] A size adjustment process can be performed before the start of decoding. After size adjustment is performed using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, etc.), the size-adjusted image decoded data can be parsed and decoded. After decoding is completed, the image can be stored in memory, and a size adjustment inverse process (in this example, using a data processing method, etc., which is used in the size adjustment inverse process) can be performed to change the output image to the image before size adjustment.

[0263] As a second example, before the start of encoding, a size adjustment process can be performed on a reference image. After size adjustment is performed using size adjustment information (e.g., a size adjustment operation, a size adjustment direction, a size adjustment value, a data processing method, etc.; the data processing method is the one used in the size adjustment process), the size-adjusted image (in this example, the size-adjusted reference image) can be stored in a memory, and an image can be encoded using this. After encoding is completed, image encoding data (in this example, meaning the image encoded using the reference image) can be recorded in a bitstream and transmitted. Also, when the encoded image is stored in a memory as a reference image, the size adjustment process can be performed as described above.

[0264] Before the start of decoding, a size adjustment process can be performed on the reference image. A size-adjusted image (in this example, a size-adjusted reference image) can be stored in a memory using size adjustment information (e.g., a size adjustment operation, a size adjustment direction, a size adjustment value, a data processing method, etc.; the data processing method is the one used in the size adjustment process), and image decoded data (in this example, the same as that encoded by the encoder using the reference image) can be parsed and decoded. After the decoding is completed, an output image can be generated, and if the decoded image is included in the reference image and stored in a memory, the size adjustment process can be performed as described above.

[0265] As a third example, after completion of encoding (specifically, meaning completion of encoding excluding the filtering process), a resizing process may be performed on an image before starting filtering of the image (assumed to be a deblocking filter in this example). After performing resizing using resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is the one used in the resizing process), a resized image may be generated, and filtering may be applied to the resized image. After completion of filtering, a resizing reverse process may be performed to change the image back to the image before resizing.

[0266] After completion of decoding (specifically, this means completion of decoding excluding the filtering process), a size adjustment process can be performed on the image before filtering of the image starts. After size adjustment is performed using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; the data processing method is the one used in the size adjustment process), a size-adjusted image can be generated, and filtering can be applied to the size-adjusted image. After filtering is completed, a size adjustment reverse process can be performed to change the image back to the image before size adjustment.

[0267] In the above examples, in some cases (first and third examples), a size adjustment process and a reverse size adjustment process may be performed, and in other cases (second example), only a size adjustment process may be performed.

[0268] Also, in some cases (examples 2 and 3), the size adjustment process in the encoder and the decoder may be the same, and in other cases (example 1), the size adjustment process in the encoder and the decoder may be the same or different. In this case, the difference between the size adjustment process in the encoder / decoder may be a size adjustment execution step. For example, in some cases (encoder in this example), a size adjustment execution step may be included that considers image size adjustment and data processing of the area to be resized, and in some cases (decoder in this example), a size adjustment execution step may be included that considers image size adjustment. In this case, the data processing in the former case may correspond to the data processing in the reverse size adjustment process in the latter case.

[0269] In addition, in some cases (e.g., the third example), the resizing process is a process applied only to the corresponding step, and the resizing area does not need to be stored in memory. For example, the resizing process can be performed by storing the resizing area in a temporary memory for use in the filtering process, filtering the area, and removing the corresponding area through a reverse resizing process. In this case, the resizing process does not change the size of the image. The above examples are not limiting, and other modifications are possible.

[0270] The size of the image can be changed through the size adjustment process, and therefore the coordinates of some pixels of the image can be changed through the size adjustment process. This can affect the operation of the picture division unit. In the present invention, block-based division can be performed based on the image before size adjustment through the above process, or block-based division can be performed based on the image after size adjustment. Also, some units (e.g., tiles, slices, etc.) can be divided based on the image before size adjustment, or some units can be divided based on the image after size adjustment. This can be determined according to the encoding / decoding settings. In the present invention, the case where the picture division unit operates based on the image after size adjustment (e.g., image division process after size adjustment process) will be mainly described, but other modifications are also possible. This will be described for the above example in multiple image settings described later.

[0271] The encoder records the information generated in the above process in a bitstream in at least one unit of a sequence, a picture, a slice, a tile, etc., and the decoder parses the relevant information from the bitstream. The information may also be included in the bitstream in the form of SEI or metadata.

[0272] Although it is common to encode / decode an input image as is, this may also occur when encoding / decoding an image by reconstructing it. For example, image reconstruction may be performed to improve image encoding efficiency, to take into account the network and user environment, or according to the type and characteristics of the image.

[0273] In the present invention, the image reconstruction process can be performed independently or can be an inverse process. In the following examples, the reconstruction process will be mainly described, but the contents of the inverse reconstruction process can be derived from the reconstruction process.

[0274] FIG. 10 is an exemplary diagram for image reconstruction according to an embodiment of the present invention.

[0275] Assuming that 10a is the first image input, 10a to 10d are example diagrams in which a rotation including 0 degrees is applied to the image (for example, a group of candidates can be generated by sampling 360 degrees into k intervals, where k can have values ​​such as 2, 4, 8, etc., and in this example, k is assumed to be 4), and 10e to 10h are example diagrams in which inversion (or symmetry) is applied based on 10a or based on 10b to 10d.

[0276] The start position or scan order of the image may be changed depending on the reconstruction of the image, or a predetermined start position and scan order may be followed regardless of the reconstruction. This can be determined depending on the encoding / decoding settings. In the embodiment described below, it is assumed that a predetermined start position (e.g., the upper left position of the image) and scan order (e.g., raster scan) are followed regardless of the reconstruction of the image.

[0277] An image encoding method and a decoding method according to an embodiment of the present invention may include the following image reconstruction steps. In this case, the image reconstruction process may include an image reconstruction instruction step, an image reconstruction type identification step, and an image reconstruction execution step. Also, the image encoding device and the decoding device may be configured to include an image reconstruction instruction unit, an image reconstruction type identification unit, and an image reconstruction execution unit that realize the image reconstruction instruction step, the image reconstruction type identification step, and the image reconstruction execution step. In the case of encoding, an associated syntax element may be generated, and in the case of decoding, the associated syntax element may be parsed.

[0278] In the image reconstruction instruction step, it can be determined whether to perform image reconstruction. For example, if a signal (e.g., convert_enabled_flag) instructing image reconstruction is confirmed, reconstruction can be performed, and if the signal instructing image reconstruction is not confirmed, reconstruction is not performed or reconstruction can be performed by checking other encoding / decoding information. Even if a signal instructing image reconstruction is not provided, the signal instructing reconstruction can be implicitly activated or inactivated according to encoding / decoding settings (e.g., image characteristics, type, etc.), and if reconstruction is performed, reconstruction-related information can be generated accordingly, or reconstruction-related information can be implicitly determined.

[0279] When a signal instructing image reconstruction is provided, the corresponding signal is a signal for indicating whether or not to perform image reconstruction, and whether or not to perform reconstruction of the corresponding image can be confirmed according to the signal. For example, when a signal instructing image reconstruction (e.g., convert_enabled_flag) is confirmed, if the corresponding signal is activated (e.g., convert_enabled_flag=1), reconstruction may be performed, and if the corresponding signal is deactivated (e.g., convert_enabled_flag=0), reconstruction may not be performed.

[0280] In addition, if a signal instructing image reconstruction is not provided, reconstruction may not be performed, or whether the image is reconstructed may be confirmed by other signals. For example, reconstruction may be performed according to the characteristics and type of image (e.g., 360-degree image), and reconstruction information may be explicitly generated or assigned to a predetermined value. The present invention is not limited to the above example, and other modifications may be possible.

[0281] In the image reconstruction type identification step, the image reconstruction type can be identified. The image reconstruction type can be defined by a reconstruction method, reconstruction mode information, etc. The reconstruction method (e.g., convert_type_flag) can include rotation, inversion, etc., and the reconstruction mode information can include a mode in the reconstruction method (e.g., convert_mode). In this case, the reconstruction related information can be composed of a reconstruction method and mode information. That is, it can be composed of at least one syntax element. In this case, the number of candidate groups of mode information according to each reconstruction method may be the same or different.

[0282] As an example, in the case of rotation, candidates with a fixed difference (90 degrees in this example) such as 10a to 10d may be included, and when 10a is 0 degree rotation, 10b to 10d may be examples of applying 90 degree, 180 degree, and 270 degree rotation, respectively (in this example, the angle is measured clockwise).

[0283] For example, in the case of inversion, candidates such as 10a, 10e, and 10f may be included, and when 10a is no inversion, 10e and 10f may be examples of applying left-right inversion and up-down inversion, respectively.

[0284] The above example describes a case where a setting for rotation with a certain interval and a setting for some inversion are described, but it is only one example for image reconstruction and is not limited to the above case, and other interval differences can include examples of other inversion operations, etc. This can be determined according to the encoding / decoding settings.

[0285] Alternatively, the reconstruction-related information may include integrated information (e.g., convert_com_flag) generated by mixing a reconstruction method and mode information corresponding thereto. In this case, the reconstruction-related information may be configured as information in which a reconstruction method and mode information are mixed.

[0286] For example, the integrated information may include candidates such as 10a to 10f, which may be examples of applying 0 degree rotation, 90 degree rotation, 180 degree rotation, 270 degree rotation, left-right flip, and up-down flip based on 10a.

[0287] Alternatively, the integrated information may include candidates such as 10a to 10h, which may be examples of applying 0 degree rotation, 90 degree rotation, 180 degree rotation, 270 degree rotation, left-right flip, up-down flip, left-right flip after 90 degree rotation (or 90 degree rotation after left-right flip), up-down flip after 90 degree rotation (or 90 degree rotation after up-down flip) based on 10a, or may be examples of applying 0 degree rotation, 90 degree rotation, 180 degree rotation, 270 degree rotation, left-right flip, left-right flip after 180 degree rotation (180 degree rotation after left-right flip), left-right flip after 90 degree rotation (90 degree rotation after left-right flip), left-right flip after 270 degree rotation (270 degree rotation after left-right flip).

[0288] The candidate group may include a mode to which rotation is applied, a mode to which inversion is applied, and a mode to which rotation and inversion are mixed. The mixed mode simply includes mode information of a method for performing reconstruction, and may include a mode generated by mixing mode information of each method. In this case, it may include a mode generated by mixing at least one mode of some methods (e.g., rotation) and at least one mode of some methods (e.g., inversion), and the above example includes a case where one mode of some methods and multiple modes of some methods are mixed (in this example, 90 degree rotation + multiple inversions / left-right inversions + multiple rotations). The mixed information may be configured to include {10a in this example} as a candidate group when reconstruction is not applied, and may include it as the first candidate group (e.g., 0 is assigned to the index) when reconstruction is not applied.

[0289] Alternatively, the reconstruction-related information may include mode information according to a predetermined reconstruction method. In this case, the reconstruction-related information may be configured with mode information according to a predetermined reconstruction method. That is, the information on the reconstruction method may be omitted, and may be configured with one syntax element related to the mode information.

[0290] For example, candidates such as 10a to 10d related to rotation can be included, or candidates such as 10a, 10e, and 10f related to inversion can be included.

[0291] The size of the image before and after the image reconstruction process may be the same or at least one length may be different. This can be determined according to the encoding / decoding settings. The image reconstruction process is a process of rearranging pixels in an image (in this example, the image reconstruction inverse process performs an inverse pixel rearrangement process, which can be inversely derived from the pixel rearrangement process), and the position of at least one pixel may be changed. The pixel rearrangement may be performed according to a rule based on the image reconstruction type information.

[0292] At this time, the pixel rearrangement process may be affected by the size and shape (e.g., square or rectangular) of the image, etc. In particular, the width and height of the image before the reconstruction process and the width and height of the image after the reconstruction process may act as variables in the pixel rearrangement process.

[0293] For example, at least one of ratio information (e.g., former / latter or latter / former, etc.) of the ratio between the width of the image before the reconstruction process and the width of the image after the reconstruction process, the ratio between the width of the image before the reconstruction process and the height of the image after the reconstruction process, the ratio between the height of the image before the reconstruction process and the width of the image after the reconstruction process, and the ratio between the height of the image before the reconstruction process and the height of the image after the reconstruction process can act as a variable in the pixel rearrangement process.

[0294] In the above example, if the size of the image before and after the reconstruction process is the same, the ratio of the width and height of the image can act as a variable in the pixel rearrangement process. Also, if the shape of the image is square, the ratio of the length of the image before the image reconstruction process to the length of the image after the reconstruction process can act as a variable in the pixel rearrangement process.

[0295] In the image reconstruction execution stage, image reconstruction can be performed based on the identified reconstruction information, i.e., image reconstruction can be performed based on information such as reconstruction type, reconstruction mode, etc., and encoding / decoding can be performed based on the acquired reconstructed image.

[0296] Next, an example of image reconstruction performed by the encoding / decoding device according to one embodiment of the present invention will be shown.

[0297] A reconstruction process can be performed on the input image before encoding begins. After reconstruction is performed using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), the reconstructed image can be coded. After coding is complete, it can be stored in memory, or the image coding data can be included in a bitstream and transmitted.

[0298] A reconstruction process can be performed before the start of decoding. After reconstruction is performed using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), the image decoded data can be parsed and decoded. After decoding is completed, the image can be stored in memory, and the image can be output after performing a reverse reconstruction process to change it to the image before reconstruction.

[0299] The encoder records the information generated in the above process in a bitstream in at least one unit of a sequence, a picture, a slice, a tile, etc., and the decoder parses the relevant information from the bitstream. The information may also be included in the bitstream in the form of SEI or metadata.

[0300] [Table 1]

[0301] Table 1 shows an example of syntax elements related to division during image settings. In the example described below, the added syntax elements will be mainly described. In addition, the syntax elements in the example described below are not limited to a specific unit, but may be syntax elements supported in various units such as a sequence, a picture, a slice, a tile, etc. Or, they may be syntax elements included in SEI, metadata, etc. In addition, the types of syntax elements, the order of syntax elements, conditions, etc. supported in the example described below are only limited to this example, and may be changed or determined according to the encoding / decoding settings.

[0302] In Table 1, tile_header_enabled_flag refers to a syntax element that indicates whether encoding / decoding settings are supported for tiles. If it is activated (tile_header_enabled_flag=1), it can have tile-level encoding / decoding settings, and if it is deactivated (tile_header_enabled_flag=0), it cannot have tile-level encoding / decoding settings and can be assigned the encoding / decoding settings of the upper unit.

[0303] The tile_coded_flag means a syntax element indicating whether or not to encode / decode a tile. When it is activated (tile_coded_flag=1), the corresponding tile can be encoded / decoded, and when it is deactivated (tile_coded_flag=0), the encoding / decoding cannot be performed. Here, not encoding may mean that encoded data in the corresponding tile is not generated (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc., which can be applied to a meaningless area in a partial projection format of a 360-degree image). Not decoding may mean that decoded data in the corresponding tile is no longer parsed (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc.). Also, not parsing decoded data may mean that no parsing is performed because no encoded data exists in the corresponding unit, but may also mean that no parsing is performed due to the flag even if encoded data exists. Depending on whether or not encoding / decoding of the tile is performed, header information for each tile may be supported.

[0304] Although the above example has been described with a focus on tiles, it is not limited to tiles and may be applicable to other division units in the present invention. In addition, an example of tile division setting is not limited to the above case and may be modified to other examples.

[0305] [Table 2]

[0306] Table 2 shows examples for syntax elements related to reconstruction during image configuration.

[0307] Referring to Table 2, convert_enabled_flag is a syntax element that indicates whether reconstruction is performed. When it is activated (convert_enabled_flag=1), it means that a reconstructed image is encoded / decoded, and additional reconstruction-related information can be confirmed. When it is deactivated (convert_enabled_flag=0), it means that an existing image is encoded / decoded.

[0308] The convert_type_flag indicates a combination of information regarding the reconstruction method and mode information. It can determine one of a number of candidate groups regarding a method in which rotation is applied, a method in which inversion is applied, and a method in which rotation and inversion are mixed.

[0309] [Table 3]

[0310] Table 3 shows examples of syntax elements related to size adjustment during image configuration.

[0311] Referring to Table 3, pic_width_in_samples and pic_height_in_samples refer to syntax elements related to the width and height of an image, and the size of an image can be confirmed by these syntax elements.

[0312] img_resizing_enabled_flag means a syntax element for indicating whether to adjust the size of an image. When it is activated (img_resizing_enabled_flag=1), it means that the image after resizing is encoded / decoded, and additional information related to resizing can be checked. When it is deactivated (img_resizing_enabled_flag=0), it means that the existing image is encoded / decoded. It may also be a syntax element that means resizing for intra prediction.

[0313] Resizing_met_flag means a syntax element for the resizing method. It can be determined to use one of a candidate group of resizing methods, such as using a scale factor (resizing_met_flag=0), using an offset factor (resizing_met_flag=1), or other resizing methods.

[0314] The resizing_mov_flag denotes a syntax element for the resizing operation. For example, it can decide to either expand or shrink.

[0315] The width_scale and height_scale refer to the scale factors related to the horizontal size adjustment and the vertical size adjustment among the size adjustments using the scale factors.

[0316] top_height_offset and bottom_height_offset refer to the upward and downward offset factors related to the horizontal size adjustment among the size adjustments using the offset factors, and left_width_offset and right_width_offset refer to the left and right offset factors related to the vertical size adjustment among the size adjustments using the offset factors.

[0317] The size of the image after resizing can be updated through the resizing related information and the image size information.

[0318] The resizing_type_flag indicates a syntax element for a data processing method of the area to be resized. Depending on the resizing method and the resizing operation, the number of candidates for the data processing method may be the same or different.

[0319] The image setting processes applied to the image encoding / decoding device described above may be performed individually or multiple image setting processes may be mixed together. In the example described below, a case where multiple image setting processes are mixed together will be described.

[0320] 11 is an exemplary diagram showing images before and after an image setting process according to an embodiment of the present invention. In detail, 11a shows an example before image reconstruction is performed on a divided image (e.g., an image projected by 360-degree image coding), and 11b shows an example after image reconstruction is performed on a divided image (e.g., an image packed by 360-degree image coding). That is, 11a can be understood as an exemplary diagram before the image setting process is performed, and 11b as an exemplary diagram after the image setting process is performed.

[0321] The image setup process in this example illustrates the case for image segmentation (assumed to be tiles in this example) and image reconstruction.

[0322] In the example described later, a case where an image is reconstructed after image division will be described, but it is also possible to perform image division after image reconstruction according to the encoding / decoding settings, and modifications to other cases are also possible. In addition, the above-mentioned image reconstruction process (including the inverse process) can be applied in the same or similar manner as the reconstruction process of the division unit in the image in this embodiment.

[0323] Image reconstruction may or may not be performed for all division units in an image, or may be performed for some of the division units. Therefore, the division units before reconstruction (e.g., some of P0 to P5) may or may not be the same as the division units after reconstruction (e.g., some of S0 to S5). Various cases related to the execution of image reconstruction will be described through examples described later. For convenience of explanation, it is assumed that the unit of an image is a picture, the unit of a divided image is a tile, and the division unit is a square shape.

[0324] As an example, whether to perform image reconstruction can be determined by some units (e.g., sps_convert_enabled_flag, SEI, metadata, etc.). Or, whether to perform image reconstruction can be determined by some units (e.g., pps_convert_enabled_flag). This is possible when it occurs first in the corresponding unit (picture in this example) or is activated in a higher unit (e.g., sps_convert_enabled_flag=1). Or, whether to perform image reconstruction can be determined by some units (e.g., tile_convert_flag[i], where i is a division unit index). This is possible when it occurs first in the corresponding unit (tile in this example) or is activated in a higher unit (e.g., pps_convert_enabled_flag=1). Also, whether to perform the partial image reconstruction can be implicitly determined according to the encoding / decoding settings, thereby allowing related information to be omitted.

[0325] As an example, it may be determined whether to perform reconstruction of a division unit in an image according to a signal (e.g., pps_convert_enabled_flag) instructing reconstruction of an image. In particular, it may be determined whether to perform reconstruction of all division units in an image according to the signal. At this time, a signal instructing reconstruction of one image may be generated in the image.

[0326] As an example, it may be determined whether to reconstruct a division unit in an image according to a signal (e.g., tile_convert_flag[i]) instructing image reconstruction. In particular, it may be determined whether to reconstruct a part of the division units in the image according to the signal. In this case, at least one signal instructing image reconstruction (e.g., generated for the number of division units) may be generated.

[0327] As an example, whether to reconstruct an image can be determined according to a signal (e.g., pps_convert_enabled_flag) instructing the reconstruction of an image, and whether to reconstruct a division unit within the image can be determined according to a signal (e.g., tile_convert_flag[i]) instructing the reconstruction of an image. In particular, when some signals are activated (e.g., pps_convert_enabled_flag=1), some signals (e.g., tile_convert_flag[i]) can be checked, and whether to reconstruct a division unit within the image can be determined according to the signal (tile_convert_flag[i] in this example). At this time, a signal instructing the reconstruction of multiple images can be generated.

[0328] When a signal instructing image reconstruction is activated, image reconstruction related information can be generated. Various cases of image reconstruction related information will be described in the examples below.

[0329] For example, reconstruction information that is applied to an image may be generated, and in particular, one reconstruction information may be used as reconstruction information for all division units within the image.

[0330] As an example, reconstruction information that is applied to a division unit in an image may be generated. In particular, at least one reconstruction information may be used as reconstruction information for a portion of the division units in an image. That is, one reconstruction information may be used as reconstruction information for one division unit, or one reconstruction information may be used as reconstruction information for multiple division units.

[0331] The example described below can be explained in combination with an example of image reconstruction.

[0332] For example, when a signal instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information that is commonly applied to division units within the image may be generated. Or, when a signal instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information that is individually applied to division units within the image may be generated. Or, when a signal instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information that is individually applied to division units within the image may be generated. Or, when a signal instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information that is commonly applied to division units within the image may be generated.

[0333] The reconfiguration information can be implicit or explicit depending on the encoding / decoding settings. In the implicit case, the reconfiguration information can be assigned a predetermined value depending on the characteristics, type, etc. of the image.

[0334] P0 to P5 of 11a may correspond to S0 to S5 of 11b, and a reconstruction process may be performed for each division unit. For example, P0 may be assigned to SO without reconstruction, P1 may be assigned to S1 after applying 90-degree rotation, P2 may be assigned to S2 after applying 180-degree rotation, P3 may be assigned to S3 after applying left-right flipping, P4 may be assigned to S4 after applying 90-degree rotation and left-right flipping, and P5 may be assigned to S5 after applying 180-degree rotation and left-right flipping.

[0335] However, the present invention is not limited to the above-mentioned examples, and various modified examples are possible. As in the above-mentioned examples, reconstruction may not be performed on the divided units of the image, or at least one of the reconstruction methods including a rotation-applied reconstruction, a flip-applied reconstruction, and a combination of a rotation and a flip-applied reconstruction may be performed.

[0336] When image reconstruction is applied to a division unit, an additional reconstruction process such as division unit rearrangement may be performed. That is, the image reconstruction process of the present invention may include an intra-image rearrangement of a division unit in addition to rearrangement of pixels in the image, and may be expressed by some syntax elements (e.g., part_top, part_left, part_width, part_height, etc.) as shown in Table 4. This means that the image division and image reconstruction processes can be understood in a mixed manner. The above description may be a possible example when an image is divided into multiple units.

[0337] P0 to P5 of 11a may correspond to S0 to S5 of 11b, and a reconstruction process may be performed for each division unit. For example, P0 may be assigned to S0 without reconstruction, P1 may be assigned to S2 without reconstruction, P2 may be rotated 90 degrees and assigned to S1, P3 may be flipped left and right and assigned to S4, P4 may be rotated 90 degrees and flipped left and right and assigned to S5, and P5 may be rotated 180 degrees and flipped left and right and assigned to S3, and is not limited thereto, and various modifications are also possible.

[0338] 7 can correspond to P_Width and P_Height in FIG. 11, and P'_Width and P'_Height in FIG. 7 can correspond to P'_Width and P'_Height in FIG. 11. The image size after size adjustment in FIG. 7 is P'_Width×P'_Height and can be expressed as (P_Width+Exp_L+Exp_R)×(P_Height+Exp_T+Exp_B), and the image size after size adjustment in FIG. 11 is P'_Width×P'_Height and can be expressed as (P_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R)×(P _Height+Var0_T+Var1_T+Var0_B+Var1_B) or it can be expressed as (Sub_P0_Width+Sub_P1_Width+Sub_P2_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R) x (Sub_P0_Height+Sub_P1_Height+Var0_T+Var1_T+Var0_B+Var1_B).

[0339] As in the above example, image reconstruction can perform pixel rearrangement within a division unit of an image, can perform rearrangement of a division unit within an image, and can perform not only pixel rearrangement within a division unit of an image but also rearrangement of a division unit within an image. In this case, intra-image rearrangement of the division unit can be performed after performing pixel rearrangement of the division unit, or intra-image rearrangement of the division unit can be performed after performing pixel rearrangement of the division unit.

[0340] The rearrangement of the division units in the image may be determined according to a signal instructing the reconstruction of the image. Alternatively, a signal for the rearrangement of the division units in the image may be generated. In particular, the signal may be generated when the signal instructing the reconstruction of the image is activated. Alternatively, the process may be implicit or explicit according to the encoding / decoding settings. If it is implicit, it may be determined according to the characteristics, type, etc. of the image.

[0341] In addition, information regarding the rearrangement of division units within an image can be processed implicitly or explicitly depending on the encoding / decoding settings, and can be determined depending on the characteristics, type, etc. of the image. That is, each division unit can be arranged according to the arrangement information of a given division unit.

[0342] Next, an example of reconstructing division units within an image in an encoding / decoding device according to an embodiment of the present invention will be shown.

[0343] Before the start of encoding, a division process can be performed on an input image using division information. A reconstruction process can be performed on a division unit basis using reconstruction information, and an image reconstructed on a division unit basis can be encoded. After encoding is completed, the image can be stored in a memory, or the image encoding data can be included in a bitstream and transmitted.

[0344] A division process can be performed using division information before the start of decoding. A reconstruction process can be performed using reconstruction information for the division units, and image decoded data can be parsed and decoded for the reconstructed division units. After decoding is completed, the data can be stored in memory, and after a reverse reconstruction process for the division units is performed, the division units can be merged into one and an image can be output.

[0345] Fig. 12 is an example diagram of size adjustment for each division unit in an image according to an embodiment of the present invention. P0 to P5 in Fig. 12 correspond to P0 to P5 in Fig. 11, and S0 to S5 in Fig. 12 correspond to S0 to S5 in Fig. 11.

[0346] In the examples described below, the case where the image size is adjusted after the image is divided will be mainly described, but the image can also be divided after the image size is adjusted according to the encoding / decoding settings, and other modifications are also possible. In addition, the above-mentioned image size adjustment process (including the inverse process) can be applied in the same or similar manner as the size adjustment process of the division unit in the image in this embodiment.

[0347] For example, TL to BR in Figure 7 may correspond to TL to BR of division unit SX (S0 to S5) in Figure 12, S0 and S1 in Figure 7 may correspond to PX and SX in Figure 12, P_Width and P_Height in Figure 7 may correspond to Sub_PX_Width and Sub_PX_Height in Figure 12, P'_Width and P'_Height in Figure 7 may correspond to Sub_SX_Width and Sub_SX_Height in Figure 12, Exp_L, Exp_R, Exp_T, and Exp_B in Figure 7 may correspond to VarX_L, VarX_R, VarX_T, and VarX_B in Figure 12, and other factors may also be corresponding.

[0348] The size adjustment process of division units within an image in 12a to 12f can be distinguished from the image size enlargement or reduction in 7a and 7b in Fig. 7 in that there may be a setting for image size enlargement or reduction in proportion to the number of division units. In addition, there may be a difference in that there is a setting commonly applied to division units within an image, or a setting individually applied to division units within an image. Various cases of size adjustment will be described in examples below, and the size adjustment process can be performed taking the above into consideration.

[0349] The image size adjustment in the present invention may or may not be performed on all division units in an image, and may be performed on some division units. Various cases of image size adjustment will be described through examples described later. For convenience of explanation, it is assumed that the size adjustment operation is expansion, the size adjustment method is an offset factor, the size adjustment direction is up, down, left, and right directions, the size adjustment direction is set according to size adjustment information, the image unit is a picture, and the divided image unit is a tile.

[0350] As an example, whether to resize an image can be determined by some units (e.g., sps_img_resizing_enabled_flag, or SEI, metadata, etc.). Or, whether to resize an image can be determined by some units (e.g., pps_img_resizing_enabled_flag). This is possible when it occurs first in the relevant unit (picture in this example) or is activated in a higher unit (e.g., sps_img_resizing_enabled_flag=1). Or, whether to resize an image can be determined by some units (e.g., tile_resizing_flag[i], where i is a division unit index). This is possible when it occurs first in the relevant unit (tile in this example) or is activated in a higher unit. Also, whether to resize the part of the images can be implicitly determined according to the encoding / decoding settings, whereby related information can be omitted.

[0351] As an example, it may be determined whether to resize the division units in the image in response to a signal (e.g., pps_img_resizing_enabled_flag) instructing the resizing of the image. In particular, it may be determined whether to resize all division units in the image in response to the signal. In this case, a signal instructing the resizing of one image may be generated.

[0352] As an example, it may be determined whether to resize a division unit in an image in response to a signal (e.g., tile_resizing_flag[i]) instructing resizing of an image. In particular, it may be determined whether to resize a part of a division unit in an image in response to the signal. In this case, at least one signal instructing resizing of an image may be generated (e.g., generated for the number of division units).

[0353] As an example, whether to adjust the size of an image can be determined according to a signal (e.g., pps_img_resizing_enabled_flag) instructing the size adjustment of an image, and whether to adjust the size of a division unit in the image can be determined according to a signal (e.g., tile_resizing_flag[i]) instructing the size adjustment of an image. In particular, when some signals are activated (e.g., pps_img_resizing_enabled_flag=1), some signals (e.g., tile_resizing_flag[i]) can be further checked, and whether to adjust the size of some division units in the image can be determined according to the signal (tile_resizing_flag[i] in this example). At this time, a signal instructing the size adjustment of a plurality of images can be generated.

[0354] When a signal instructing image resizing is activated, image resizing related information may be generated. Various cases of image resizing related information will be described in the following examples.

[0355] As an example, size adjustment information to be applied to an image may be generated. In particular, one size adjustment information or size adjustment information set may be used as size adjustment information for all division units in an image. For example, one size adjustment information commonly applied to the top, bottom, left, and right directions of a division unit in an image (or a size adjustment value applied to all size adjustment directions supported or allowed in the division unit; one piece of information in this example) or one size adjustment information set respectively applied to the top, bottom, left, and right directions (or the number of size adjustment directions supported or allowed in the division unit; up to four pieces of information in this example) may be generated.

[0356] As an example, size adjustment information applied to a division unit in an image may be generated. In particular, at least one size adjustment information or size adjustment information set may be used as size adjustment information for a portion of division units in an image. That is, one size adjustment information or size adjustment information set may be used as size adjustment information for one division unit, or may be used as size adjustment information for a plurality of division units. For example, one size adjustment information commonly applied to the top, bottom, left, and right directions of one division unit in an image, or one size adjustment information set applied to each of the top, bottom, left, and right directions, may be generated. Alternatively, one size adjustment information commonly applied to the top, bottom, left, and right directions of a plurality of division units in an image, or one size adjustment information set applied to each of the top, bottom, left, and right directions, may be generated. The configuration of a size adjustment set means size adjustment value information for at least one size adjustment direction.

[0357] In summary, size adjustment information that is commonly applied to division units within an image may be generated. Alternatively, size adjustment information that is individually applied to division units within an image may be generated. The following example can be explained in combination with an example of image size adjustment.

[0358] For example, when a signal indicating image resizing (e.g., pps_img_resizing_enabled_flag) is activated, resizing information that is commonly applied to division units within the image may be generated. Or, when a signal indicating image resizing (e.g., pps_img_resizing_enabled_flag) is activated, resizing information that is individually applied to division units within the image may be generated. Or, when a signal indicating image resizing (e.g., tile_resizing_flag[i]) is activated, resizing information that is individually applied to division units within the image may be generated. Or, when a signal indicating image resizing (e.g., tile_resizing_flag[i]) is activated, resizing information that is commonly applied to division units within the image may be generated.

[0359] Image sizing direction, sizing information, etc. can be handled implicitly or explicitly depending on the encoding / decoding settings. In the implicit case, sizing information can be assigned a predefined value depending on the image characteristics, type, etc.

[0360] As described above, the size adjustment direction in the size adjustment process of the present invention is at least one of the up, down, left, and right directions, and the size adjustment direction and size adjustment information can be processed explicitly or implicitly. That is, some directions are implicitly assigned size adjustment values ​​(including 0, i.e., no adjustment) and some directions are explicitly assigned size adjustment values ​​(including 0, i.e., no adjustment).

[0361] For each division unit in an image, the size adjustment direction and size adjustment information may have settings that can be processed implicitly or explicitly, and these settings may be applied to the division units in the image. For example, a setting may be applied to one division unit in an image (in this example, only the division unit is generated), or a setting may be applied to multiple division units in an image, or a setting may be applied to all division units in an image (in this example, one setting is generated), and at least one setting may be generated in an image (for example, settings as many as the number of division units can be generated from one setting). A setting set may be defined by collecting setting information applied to the division units in the image.

[0362] FIG. 13 is an exemplary diagram for adjusting or setting the size of a division unit in an image.

[0363] In detail, various examples of the size adjustment direction of the division unit in the image and the implicit or explicit processing of the size adjustment information are shown. In the examples described below, for convenience of explanation, it is assumed that the size adjustment value of some size adjustment directions is 0 in the implicit processing.

[0364] As in 13a, when the boundary of the division unit coincides with the boundary of the image (in this example, the thick solid line), the size adjustment can be performed explicitly, and when it does not coincide (thin solid line), the size adjustment can be performed implicitly. For example, P0 can be adjusted in size upward and leftward (a2, a0), P1 in the upward direction (a2), P2 in the upward and rightward direction (a2, a1), P3 in the downward and leftward (a3, a0), P4 in the downward direction (a3), and P5 in the downward and rightward (a3, a1), but size adjustment in other directions is not possible.

[0365] As shown in FIG. 13b, some directions of the division units (up and down in this example) can be explicitly processed for size adjustment, and some directions of the division units (left and right in this example) can be explicitly processed (thick solid lines in this example) when the boundaries of the division units match the boundaries of the image, and implicitly processed when they do not match (thin solid lines in this example). For example, P0 can be adjusted in size up, down, and left directions (b2, b3, b0), P1 can be adjusted in size up and down directions (b2, b3), P2 can be adjusted in size up, down, and right directions (b2, b3, b1), P3 can be adjusted in size up, down, and left directions (b3, b4, b0), P4 can be adjusted in size up and down directions (b3, b4), and P5 can be adjusted in size up, down, and right directions (b3, b4, b1), and size adjustment is not possible in other directions.

[0366] As shown in FIG. 13c, some directions of the division units (left, right in this example) can be explicitly processed for size adjustment, and some directions of the division units (up, down in this example) can be explicitly processed if the boundaries of the division units match the boundaries of the image (thick solid lines in this example), and implicitly processed if they do not match (thin solid lines in this example). For example, P0 can be adjusted in size in the up, left, right directions (c4, c0, c1), P1 in the up, left, right directions (c4, c1, c2), P2 in the up, left, right directions (c4, c2, c3), P3 in the down, left, right directions (c5, c0, c1), P4 in the down, left, right directions (c5, c1, c2), and P5 in the down, left, right directions (c5, c2, c3), and size adjustment is not possible in other directions.

[0367] As in the above example, the settings related to image resizing can have various cases. Multiple setting sets can be supported and setting set selection information can be generated explicitly, or a given setting set can be implicitly determined depending on the encoding / decoding settings (e.g., image characteristics, type, etc.).

[0368] FIG. 14 is an example diagram illustrating a process of adjusting the size of an image and a process of adjusting the size of a division unit within the image.

[0369] 14, the image size adjustment process and the inverse process can proceed in the directions of e and f, and the size adjustment process and the inverse process can proceed in the directions of d and g. That is, the size adjustment process can be performed on the image, and the size adjustment can be performed on the division units in the image, and the order of the size adjustment processes is not fixed. This means that multiple size adjustment processes are possible.

[0370] In summary, the image size adjustment process can be classified into image size adjustment (or image size adjustment before division) and size adjustment of division units within an image (or image size adjustment after division), and it is not necessary to perform both image size adjustment and size adjustment of division units within an image, or it is possible to perform either of both, or it is possible to perform both. This can be determined according to the encoding / decoding settings (e.g., characteristics, type, etc. of the image).

[0371] In the above example, when multiple size adjustment processes are performed, the size adjustment of the image may be performed in at least one of the top, bottom, left, and right directions of the image, and the size adjustment of at least one division unit among the division units in the image may be performed in at least one of the top, bottom, left, and right directions of the division unit to be adjusted.

[0372] Referring to FIG. 14, the size of image A before size adjustment can be defined as P_Width×P_Height, the size of image after the first size adjustment (or image before the second size adjustment, B) can be defined as P'_Width×P'_Height, and the size of image after the second size adjustment (or image after the final size adjustment, C) can be defined as P''_Width×P''_Height. Image A before size adjustment means an image without any size adjustment, image B after the first size adjustment means an image with some size adjustment, and image C after the second size adjustment means an image with all size adjustment. For example, image B after the first size adjustment means an image with size adjustment of the division unit in the image as shown in FIG. 13a to FIG. 13c, and image C after the second size adjustment can mean an image with size adjustment for the entire image B with the first size adjustment as shown in FIG. 7a, and vice versa. The above examples are not limited to the above examples, and various modified examples are possible.

[0373] In the size of image B after the first size adjustment, P'_Width can be obtained by P_Width and at least one size adjustment value in the left or right direction that can adjust the size horizontally, and P'_Height can be obtained by P_Height and at least one size adjustment value in the up or down direction that can adjust the size vertically. In this case, the size adjustment value may be a size adjustment value generated in a division unit.

[0374] In the size of image C after the secondary size adjustment, P″_Width can be obtained through P'_Width and at least one size adjustment value in the left or right direction that can be adjusted horizontally, and P″_Height can be obtained through P'_Height and at least one size adjustment value in the up or down direction that can be adjusted vertically. In this case, the size adjustment value may be a size adjustment value generated from the image.

[0375] In summary, the size of the resized image can be obtained via the size of the unresized image and at least one resizing value.

[0376] Information regarding a data processing method may be generated in the resized area of ​​the image. Various data processing methods will be described with reference to examples below, and the data processing method occurring in the resizing process may be the same or similar to the resizing process, and the data processing methods in the resizing process and the resizing process may be described with reference to various combinations of cases below.

[0377] As an example, a data processing method to be applied to an image may be generated. In particular, one data processing method or a set of data processing methods may be used as a data processing method for all division units in an image (assuming that all division units are resized in this example). For example, one data processing method commonly applied to the top, bottom, left, and right directions of a division unit in an image (or a data processing method applied to all size adjustment directions supported or allowed in the division unit; one piece of information in this example) or one data processing method set respectively applied to the top, bottom, left, and right directions (or the number of size adjustment directions supported or allowed in the division unit; up to four pieces of information in this example) may be generated.

[0378] As an example, a data processing method applied to a division unit in an image may be generated. In particular, at least one data processing method or a set of data processing methods may be used as a data processing method for a part of division units in an image (assumed to be a division unit to be resized in this example). That is, one data processing method or a set of data processing methods may be used as a data processing method for one division unit, or may be used as a data processing method for a plurality of division units. For example, one data processing method commonly applied to the top, bottom, left, and right directions of one division unit in an image, or one data processing method set respectively applied to the top, bottom, left, and right directions, may be generated. Alternatively, one data processing method commonly applied to the top, bottom, left, and right directions of a plurality of division units in an image, or one data processing method information set respectively applied to the top, bottom, left, and right directions may be generated. The configuration of a data processing method set means a data processing method for at least one size adjustment direction.

[0379] In summary, a data processing method commonly applied to division units in an image can be used. Or, a data processing method individually applied to division units in an image can be used. The data processing method can be a predetermined method. There can be at least one predetermined data processing method. This corresponds to an implicit case, and selection information regarding the data processing method can be explicitly generated. This can be determined according to the encoding / decoding settings (e.g., image characteristics, type, etc.).

[0380] That is, a data processing method commonly applied to division units in an image can be used, a predetermined method can be used, or one of a plurality of data processing methods can be selected, or a data processing method individually applied to division units in an image can be used, and a predetermined method can be used, or one of a plurality of data processing methods can be selected depending on the division unit.

[0381] The examples described below explain some cases of resizing (assuming expansion in this example) of division units within an image (in this example, partial image data is used to fill the resizing area).

[0382] A partial area TL to BR of a partial unit (e.g., S0 to S5 in Figs. 12a to 12f) can be resized using data of a partial area tl to br of a partial unit (P0 to P5 in Figs. 12a to 12f). At this time, the partial units can be the same (e.g., S0 and P0) or different areas (e.g., S0 and P1). That is, the areas TL to BR to be resized can be filled using partial data tl to br of the division unit, and the areas to be resized can be filled using partial data of a division unit different from the division unit.

[0383] As an example, the area TL-BR to be resized of the current division unit can be resized using the tl-br data of the current division unit. For example, the TL of S0 can be filled with the tl data of P0, the RC of S1 with the tr+rc+br data of P1, the BL+BC of S2 with the bl+bc+br data of P2, and the TL+LC+BL of S3 with the tl+lc+bl data of P3.

[0384] As an example, the area TL to BR to be resized of the current division unit can be resized using the tl to br data of the division unit spatially adjacent to the current division unit. For example, TL+TC+TR of S4 can be filled with b1+bc+br data of P1 in the upward direction, BL+BC of S2 can be filled with tl+tc+tr data of P5 in the downward direction, LC+BL of S2 can be filled with tl+rc+bl data of P1 in the left direction, RC of S3 can be filled with tl+lc+bl data of P4 in the right direction, and BR of S0 can be filled with tl data of P4 in the downward left direction.

[0385] As an example, the region TL-BR to be resized of the current division unit can be resized using the tl-br data of division units that are not spatially adjacent to the current division unit. For example, data of the regions at both ends of the image boundaries (e.g., left and right, top and bottom, etc.) can be obtained. LC of S3 can be obtained using the tr+rc+br data of S5, RC of S2 can be obtained using the tl+lc data of S0, BC of S4 can be obtained using the tc+tr data of S1, and TC of S1 can be obtained using the bc data of S4.

[0386] Alternatively, data of a partial area of ​​the image (an area that is not spatially adjacent but is determined to have a high correlation with the area to be resized) can be obtained. The BC of S1 can be obtained using the tl+lc+bl data of S3, the RC of S3 can be obtained using the tl+tc data of S1, and the RC of S5 can be obtained using the bc data of S0.

[0387] In addition, some cases (in this example, removal by restoration or correction using partial data of the image) related to size adjustment of division units in the image (assuming reduction in this example) are as follows.

[0388] The partial regions TL to BR of the partial units (e.g., S0 to S5 in Figs. 12a to 12f) can be used in the restoration or correction process of the partial regions tl to br of the partial units P0 to P5. In this case, the partial units may be the same (e.g., S0 and P0) or different regions (e.g., S0 and P2). That is, the region to be resized can be used to restore and remove some data of the corresponding division unit, and the region to be resized can be used to restore and remove some data of a division unit different from the corresponding division unit. A detailed example is omitted as it can be derived conversely from the extension process.

[0389] The above example is an example that is applied when there is data that is highly correlated with the area to be resized, and the information on the position referenced for resizing can be explicitly generated or implicitly obtained based on a predetermined rule, or the related information can be confirmed by mixing these. This may be an example that is applied when data is obtained from another area where continuity exists in encoding of a 360-degree image.

[0390] Next, an example of adjusting the size of a division unit within an image in an encoding / decoding device according to an embodiment of the present invention will be shown.

[0391] A division process can be performed on an input image before the start of encoding. A size adjustment process can be performed on the division unit using size adjustment information, and the image after size adjustment for the division unit can be encoded. After encoding is completed, it can be stored in a memory, and the image encoding data can be included in a bitstream and transmitted.

[0392] Before the start of decoding, a division process can be performed using division information. A size adjustment process can be performed using size adjustment information for the division units, and image decoded data can be parsed and decoded for the size-adjusted division units. After decoding is completed, the data can be stored in memory, and the division units can be merged into one after a reverse size adjustment process for the division units to output the image.

[0393] Other cases of the image size adjustment process described above can be modified and applied as in the above example, but the present invention is not limited thereto and can be modified to other examples.

[0394] In the image setting process, a combination of image size adjustment and image reconstruction is possible. Image size adjustment can be performed and then image reconstruction can be performed, or image size adjustment can be performed and then image reconstruction can be performed. Also, a combination of image division, image reconstruction, and image size adjustment is possible. Image size adjustment and image reconstruction can be performed after image division, and the order of image setting is not fixed but can be changed, which can be determined according to the encoding / decoding setting. In this example, the image setting process is described in the case where image division is performed, followed by image reconstruction and image size adjustment, but other orders are possible and can be changed according to the encoding / decoding setting.

[0395] For example, the image setting processes may be performed in the order of division → reconstruction, reconstruction → division, division → resize, resize → division, resize → reconstruct, resize → reconstruct, division → reconstruct → resize, division → resize → reconstruction, resize → division → reconstruct, resize → reconstruct → division, reconstruct → division → resize, resize → reconstruct → division, etc., and may be combined with additional image settings. As described above, the image setting processes may be performed sequentially, but all or some of the setting processes may be performed simultaneously. Also, some image setting processes may be performed multiple times depending on the encoding / decoding settings (e.g., image characteristics, type, etc.). Examples of various combinations of image setting processes are shown below.

[0396] As an example, P0 to P5 in FIG. 11a may correspond to S0 to S5 in FIG. 11b, and a reconfiguration process (in this example, rearrangement of pixels) and a size adjustment process (in this example, the same size adjustment for the division unit) may be performed for each division unit. For example, size adjustment using an offset may be applied to P0 to P5 and assigned to S0 to S5. Also, P0 may be assigned to S0 without reconfiguration, P1 may be assigned to S1 with a 90-degree rotation, P2 may be assigned to S2 with a 180-degree rotation, P3 may be assigned to S3 with a 270-degree rotation, P4 may be assigned to S4 with a horizontal inversion, and P5 may be assigned to S5 with a vertical inversion.

[0397] As an example, P0 to P5 in Fig. 11a may correspond to the same or different positions as S0 to S5 in Fig. 11b, and a reconfiguration process (in this example, rearrangement of pixels and division units) and a size adjustment process (in this example, the same size adjustment for the division units) may be performed for the division units. For example, P0 to P5 may be assigned to S0 to S5 by applying size adjustment using a scale. Also, P0 may be assigned to S0 without reconfiguration, P1 may be assigned to S2 without reconfiguration, P2 may be assigned to S1 by applying 90 degree rotation, P3 may be assigned to S4 by applying left-right flipping, P4 may be assigned to S5 by applying left-right flipping after 90 degree rotation, and P5 may be assigned to S3 by applying 180 degree rotation after left-right flipping.

[0398] As an example, P0 to P5 in Fig. 11a may correspond to E0 to E5 in Fig. 5e, and a reconfiguration process (in this example, rearrangement of pixels and division units) and a size adjustment process (in this example, size adjustment that is not the same for division units) may be performed for the division units. For example, P0 may be assigned to E0 without size adjustment and reconfiguration, P1 may be assigned to E1 without size adjustment using a scale and without reconfiguration, P2 may be assigned to E2 without size adjustment and reconfiguration, P3 may be assigned to E4 without size adjustment using an offset and without reconfiguration, P4 may be assigned to E5 without size adjustment and reconfiguration, and P5 may be assigned to E3 without size adjustment using an offset and reconfiguration.

[0399] As in the above example, the absolute or relative positions of the division units in the image before and after the image setting process may be maintained or changed. This can be determined according to the encoding / decoding settings (e.g., image characteristics, type, etc.). Also, various combinations of image setting processes are possible, and the present invention is not limited to the above example, and various modifications are also possible.

[0400] The encoder records the information generated in the above process in a bitstream in at least one unit of a sequence, a picture, a slice, a tile, etc., and the decoder parses the relevant information from the bitstream. The information may also be included in the bitstream in the form of SEI or metadata.

[0401] [Table 4]

[0402] Next, examples of syntax elements associated with multiple image settings will be described. In the examples described below, the added syntax elements will be mainly described. Furthermore, the syntax elements in the examples described below are not limited to a specific unit, and may be syntax elements supported in various units such as a sequence, a picture, a slice, a tile, etc. Alternatively, they may be syntax elements included in an SEI, metadata, etc.

[0403] Referring to Table 4, parts_enabled_flag is a syntax element for whether or not to divide into parts. When activated (parts_enabled_flag=1), it means that encoding / decoding is performed by dividing into multiple units, and additional division information can be confirmed. When deactivated (parts_enabled_flag=0), it means that an existing image is encoded / decoded. This example focuses on rectangular division units such as tiles, and can have different settings for existing tiles and division information.

[0404] num_partitions means a syntax element for the number of partition units, and the value added with 1 means the number of partition units.

[0405] part_top[i] and part_left[i] are syntax elements for position information of a division unit and mean the horizontal and vertical start positions of the division unit (e.g., the position of the upper left corner of the division unit). part_width[i] and part_height[i] are syntax elements for size information of a division unit and mean the horizontal and vertical widths of the division unit. In this case, the start position and size information can be set in pixel units or block units. In addition, the syntax elements may be syntax elements that can be generated in an image reconstruction process, or syntax elements that can be generated when an image division process and an image reconstruction process are mixed and configured.

[0406] The part_header_enabled_flag is a syntax element that indicates whether encoding / decoding settings are supported for each division unit. If it is activated (part_header_enabled_flag=1), it can have encoding / decoding settings for the division unit, and if it is deactivated (part_header_enabled_flag=0), it cannot have encoding / decoding settings and can be assigned the encoding / decoding settings of the upper unit.

[0407] The above example is an example of syntax elements associated with size adjustment and reconstruction in a division unit among image settings described later, and is not limited thereto, and other division units and settings of the present invention can be modified and applied. This example is described under the assumption that size adjustment and reconstruction are performed after division, but is not limited thereto, and can be modified and applied according to other image setting orders, etc. In addition, the types of syntax elements, the order of syntax elements, conditions, etc. supported in the example described later are only limited to this example, and can be changed and determined according to the encoding / decoding settings.

[0408] [Table 5]

[0409] Table 5 shows examples for syntax elements related to the reconstruction of segmentation units in an image setting.

[0410] Referring to Table 5, part_convert_flag[i] means a syntax element for whether or not a division unit is reconstructed. The syntax element may occur for each division unit, and when activated (part_convert_flag[i]=1), it means that the reconstructed division unit is encoded / decoded, and additional reconstruction-related information may be confirmed. When deactivated (part_convert_flag[i]=0), it means that the existing division unit is encoded / decoded. convert_type_flag[i] means mode information regarding the reconstruction of the division unit, and may be information regarding pixel rearrangement.

[0411] In addition, syntax elements for additional reconfiguration such as rearrangement of division units may occur. In this example, the rearrangement of division units may be performed via the syntax elements part_top and part_left related to the image division described above, or syntax elements related to the rearrangement of division units (e.g., index information, etc.) may occur.

[0412] [Table 6]

[0413] Table 6 shows examples of syntax elements related to adjusting the size of the division unit in the image setting.

[0414] Referring to Table 6, part_resizing_flag[i] is a syntax element for determining whether to resize an image of a division unit. The syntax element may occur for each division unit, and when activated (part_resizing_flag[i]=1), it indicates that the resized division unit is encoded / decoded, and additional size-related information may be confirmed. When deactivated (part_resiznig_flag[i]=0), it indicates that the existing division unit is encoded / decoded.

[0415] Width_scale[i] and height_scale[i] represent scale factors related to horizontal and vertical size adjustment in size adjustment using a scale factor in a division unit.

[0416] top_height_offset[i] and bottom_height_offset[i] refer to the upward and downward offset factors related to size adjustment using the offset factors in the division unit, and left_width_offset[i] and right_width_offset[i] refer to the left and right offset factors related to size adjustment using the offset factors in the division unit.

[0417] Resizing_type_flag[i][j] means a syntax element for a data processing method of an area resized in a division unit. The syntax element means an individual data processing method for a resized direction. For example, a syntax element for an individual data processing method of an area resized in the up, down, left, or right direction may be generated. This may also be generated based on resizing information (e.g., may only occur when resizing in some directions).

[0418] The image setting process described above may be applied depending on the characteristics, type, etc. of an image. In the following examples, the image setting process described above may be applied in the same manner or with modifications unless otherwise specified. In the following examples, the following examples will be described with a focus on additional or modified applications of the above examples.

[0419] For example, images generated through a 360-degree camera (360-degree video or omnidirectional video) have different characteristics from images acquired through a general camera, and require a different encoding environment than the compression of general images.

[0420] Unlike general images, 360-degree images do not have boundaries with discontinuous characteristics, and data in all areas can be continuous. In addition, in devices such as HMDs, images are played in front of the eyes through lenses, which can require high-quality images, and when images are acquired through a stereoscopic camera, the amount of image data to be processed can increase. In order to provide an efficient encoding environment, including the above examples, various image setting processes can be performed taking into account 360-degree images.

[0421] The 360-degree camera is a camera with multiple cameras or multiple lenses and sensors, where the cameras or lenses can cover all directions around an arbitrary central point captured by the camera.

[0422] A 360-degree image can be encoded using various methods. For example, it can be encoded using various image processing algorithms in a three-dimensional space, or it can be converted into a two-dimensional space and encoded using various image processing algorithms. In this invention, a method of converting a 360-degree image into a two-dimensional space and encoding / decoding it will be mainly described.

[0423] A 360-degree image encoding device according to an embodiment of the present invention may be configured to include all or part of the configuration shown in Fig. 1, and may further include a pre-processing unit that performs pre-processing (stitching, projection, region-wise packing) on ​​an input image. Meanwhile, a 360-degree image decoding device according to an embodiment of the present invention may include all or part of the configuration shown in Fig. 2, and may further include a post-processing unit that performs post-processing (rendering) before being decoded and reproduced as an output image.

[0424] To explain it again, the encoder can perform pre-processing on the input image, then encode it and transmit the corresponding bitstream, and the decoder can parse and decode the transmitted bitstream, then generate an output image after post-processing. At this time, the bitstream can include and transmit information generated in the pre-processing process and information generated in the encoding process, and the decoder can parse it and use it in the decoding process and post-processing process.

[0425] Next, the operation method of the 360-degree image encoder will be described in more detail. Since the operation method of the 360-degree image decoder is the inverse operation of the 360-degree image encoder, it can be easily derived by ordinary engineers and therefore a detailed description will be omitted.

[0426] The input image can be stitched and projected into a 3D projection structure in units of a sphere, through which image data on the 3D projection structure can be projected into a 2D image.

[0427] The projected image may be configured to include all or part of the 360-degree content depending on the encoding settings. In this case, the position information of the region (or pixel) located at the center of the projected image may be implicitly generated as a predetermined value, or the position information may be explicitly generated. In addition, when a projected image is configured to include a part of the 360-degree content, the range and position information of the included region may be generated. In addition, range information (e.g., vertical width, horizontal width) and position information (e.g., measured based on the upper left side of the image) for the region of interest (ROI) in the projected image may be generated. In this case, a part of the 360-degree content that has high importance may be set as the region of interest. Although all contents in the upper, lower, left, and right directions can be seen in a 360-degree image, the user's line of sight may be limited to a part of the image, and the region of interest may be set taking this into consideration. For efficient encoding, the region of interest may be set to have good quality and resolution, and other regions may be set to have lower quality and resolution than the region of interest.

[0428] Among the 360-degree image transmission methods, the single stream transmission method can transmit an entire image or a viewport image to a user in an individual single bit stream. The multi-stream transmission method can select the image quality according to the user's environment and communication conditions by transmitting a plurality of entire images with different image qualities in a multiple bit stream. The tiled stream transmission method can select a tile according to the user's environment and communication conditions by transmitting a partial image in tile units that is individually encoded in a multiple bit stream. Thus, the 360-degree image encoder generates and transmits bit streams with two or more qualities, and the 360-degree image decoder can set an area of ​​interest according to the user's gaze and selectively decode according to the area of ​​interest. That is, the area where the user's gaze is fixed can be set as an area of ​​interest through a head tracking or eye tracking system, and rendering can be performed only on the required part.

[0429] The projected image may be converted into a packed image by performing a region-wise packing process. The region-wise packing process may include a step of dividing the projected image into a plurality of regions. At this time, each divided region may be arranged (or rearranged) in the packed image according to a region-wise packing setting. The region-wise packing may be performed for the purpose of increasing spatial continuity when converting a 360-degree image into a 2D image (or a projected image). The size of the image may be reduced through the region-wise packing. In addition, the image quality degradation occurring during rendering may be reduced, and the region-wise packing may be performed for the purpose of enabling viewport-based projection and providing other types of projection formats. The region-wise packing may or may not be performed depending on the encoding setting, and may be determined based on a signal indicating whether or not to perform the region-wise packing (e.g., regionwise_packing_flag; in the example described below, the region-wise packing related information may be generated only when the regionwise_packing_flag is activated).

[0430] When regional packing is performed, setting information (or mapping information) that assigns (or places) a portion of the projected image as a portion of the packed image may be displayed (or generated). When regional packing is not performed, the projected image and the packed image may be the same image.

[0431] Although the stitching, projection, and regional packing processes are defined as separate processes above, some of the processes (e.g., stitching + projection, projection + regional packing) or all of the processes (e.g., stitching + projection + regional packing) can be defined as a single process.

[0432] At least one packed image can be generated for the same input image according to settings of the stitching, projection, and regional packing processes, etc. Also, at least one coded data can be generated for the same projected image according to settings of the regional packing process.

[0433] A packed image may be divided by performing a tiling process. In this case, tiling is a process of dividing an image into a plurality of regions and transmitting the image, and may be an example of the 360-degree image transmission method. As described above, tiling may be performed for the purpose of partial decoding in consideration of a user's environment, etc., and for the purpose of efficient processing of a huge amount of data of a 360-degree image. For example, when an image is composed of one unit, the entire image may be decoded for decoding the region of interest, but when an image is composed of a plurality of unit regions, it is efficient to decode only the region of interest. In this case, the division may be performed by dividing into tiles, which are division units according to an existing encoding method, or by dividing into various division units (square division, block, etc.) described in the present invention. In addition, the division units may be units for performing independent encoding / decoding. Tiling may be performed based on a projected image or a packed image, or may be performed independently. That is, division may be performed based on a surface boundary of a projected image, a surface boundary of a packed image, a packing setting, etc., and division may be performed independently for each division unit. This can affect the generation of partition information during the tiling process.

[0434] Next, the projected image or the packed image can be encoded. The encoded data and information generated in the pre-processing process can be recorded in a bitstream and transmitted to a 360-degree image decoder. The information generated in the pre-processing process can be recorded in the bitstream in the form of SEI or metadata. In this case, at least one encoded data and at least one pre-processing information that change a part of the encoding process setting or a part of the pre-processing process setting can be recorded in the bitstream. This may be for the purpose of forming a decoded image by mixing a plurality of encoded data (encoded data + pre-processing information) according to a user's environment in the decoder. In particular, a decoded image can be formed by selectively combining a plurality of encoded data. In addition, the above process may be performed in two parts for application in a stereoscopic system, or the above process may be performed for an additional depth image.

[0435] FIG. 15 is an exemplary diagram showing a three-dimensional space showing a three-dimensional image and a two-dimensional plane space.

[0436] Generally, 3DoF (Degree of Freedom) is required for a 360-degree 3D virtual space, and three rotations can be supported around the X (Pitch), Y (Yaw), and Z (Roll) axes. DoF refers to degrees of freedom in space, 3DoF refers to degrees of freedom including rotations around the X, Y, and Z axes as in 15a, and 6DoF refers to degrees of freedom that further allows movement along the X, Y, and Z axes in addition to 3DoF. The image encoding and decoding apparatus of the present invention will be described mainly for the case of 3DoF, and when supporting more than 3DoF (3DoF+), it may be combined with or modified with additional processes or devices not shown in the present invention.

[0437] Referring to 15a, Yaw can have a range from -π (-180 degrees) to π (180 degrees), Pitch can have a range from -π / 2rad (or -90 degrees) to π / 2rad (or 90 degrees), and Roll can have a range from -π / 2rad (or -90 degrees) to π / 2rad (or 90 degrees). In this case, assuming that ψ and θ are longitude and latitude in the map representation of the Earth, (x, y, z) in the three-dimensional space can be converted from (ψ, θ) in the two-dimensional space. For example, the three-dimensional space coordinates can be derived from the two-dimensional space coordinates based on the conversion formulas x = cos(θ) cos(ψ), y = sin(θ), z = -cos(θ) sin(ψ).

[0438] Also, (ψ, θ) can be transformed into (x, y, z). For example, ψ=tan -1 (-Z / X), θ=sin -1 (Y / (X 2 +Y 2 +Z 2 ) 1 / 2 ) Two-dimensional space coordinates can be derived from three-dimensional space coordinates based on the transformation formula.

[0439] When a pixel in a three-dimensional space is converted exactly into a two-dimensional space (e.g., an integer unit pixel in a two-dimensional space), the pixel in the three-dimensional space can be mapped to a pixel in a two-dimensional space. When a pixel in a three-dimensional space is not converted exactly into a two-dimensional space (e.g., a fractional unit pixel in a two-dimensional space), the two-dimensional pixel can be mapped to a pixel obtained by performing interpolation. In this case, the interpolation method used can be a nearest neighbor interpolation method, a bi-linear interpolation method, a B-spline interpolation method, a bi-cubic interpolation method, or the like. In this case, one of a plurality of interpolation candidates can be selected, and related information can be explicitly generated, or the interpolation method can be implicitly determined based on a predetermined rule. For example, a predetermined interpolation filter can be used according to a three-dimensional model, a projection format, a color format, a slice / tile type, and the like. In addition, when the interpolation information is explicitly generated, information about the filter information (e.g., a filter coefficient, etc.) can also be included.

[0440] 15b shows an example of transformation from 3D space to 2D space (2D plane coordinate system). (ψ, θ) can be sampled (i, j) based on the image size (width, height), where i can range from 0 to P_Width-1 and j can range from 0 to P_Height-1.

[0441] (ψ, θ) may be the center point {or reference point, the point indicated by C in FIG. 15, coordinates (ψ, θ)=(0, 0)} for positioning the 360-degree image at the center of the projected image. The setting for the center point may be specified in three-dimensional space, and position information for the center point may be explicitly generated or may be implicitly set to a value that has already been set. For example, center position information in Yaw, center position information in Pitch, center position information in Roll, etc. may be generated. When values ​​for the information are not specifically specified, each value may be assumed to be 0.

[0442] In the above example, an example of converting the entire 360-degree image from three-dimensional space to two-dimensional space has been described, but a partial area of ​​the 360-degree image can be targeted, and position information (e.g., a partial position belonging to the area. In this example, position information for the center point), range information, etc. for the partial area can be explicitly generated, or the position and range information already set implicitly can be followed. For example, center position information in Yaw, center position information in Pitch, center position information in Roll, range information in Yaw, range information in Pitch, range information in Roll, etc. can be generated, and in the case of a partial area, at least one area is generated, thereby allowing position information and range information for multiple areas to be processed. If a value for the information is not specifically specified, it can be assumed to be the entire 360-degree image.

[0443] H0 to H6 and W0 to W5 in 15a respectively indicate some latitudes and longitudes in 15b, and the coordinates of 15b can be expressed as (C, j) and (i, C) (C is a longitude or latitude component). Unlike a general image, a 360-degree image may be distorted or warped when converted into a two-dimensional space. This may vary depending on the area of ​​the image, and the encoding / decoding settings may be different for the position of the image or for the area partitioned according to the position. When the encoding / decoding settings are adaptively set based on the encoding / decoding information in the present invention, the position information (e.g., x, y components, or a range defined by x and y, etc.) may be included as an example of the encoding / decoding information.

[0444] The above description of three-dimensional and two-dimensional space is provided to aid in the description of the embodiments of the present invention, and is not intended to be limiting, and modifications of the details or applications in other cases are possible.

[0445] As described above, an image captured by a 360-degree camera can be converted into a two-dimensional space. At this time, a 360-degree image can be mapped using a three-dimensional model, and various three-dimensional models such as a sphere, cube, cylinder, pyramid, and polyhedron can be used. When a 360-degree image mapped based on the model is converted into a two-dimensional space, a projection process can be performed according to a projection format based on the model.

[0446] 16a to 16d are conceptual diagrams for explaining a projection format according to an embodiment of the present invention.

[0447] FIG. 16a shows an ERP (Equi-Rectangular Projection) format in which a 360-degree image is projected onto a two-dimensional plane. FIG. 16b shows a CMP CubeMap Projection format in which a 360-degree image is projected onto a cube. FIG. 16c shows an OHP (OctaHedron Projection) format in which a 360-degree image is projected onto an octahedron. FIG. 16d shows an ISP (IcoSahedral Projection) format in which a 360-degree image is projected onto a polyhedron. However, various projection formats can be used without being limited to these. The left side of FIG. 16a to FIG. 16d shows a three-dimensional model, and the right side shows an example converted into two-dimensional space through a projection process. There are various sizes and shapes depending on the projection format, and each shape can be composed of faces or surfaces, and the surfaces can be expressed as circles, triangles, rectangles, etc.

[0448] In the present invention, the projection format can be defined by the three-dimensional model, the surface settings (e.g., the number of surfaces, the surface shape, the surface shape configuration, etc.), the projection process settings, etc. If at least one element of the definition is different, it can be regarded as a different projection format. For example, in the case of ERP, if it is composed of a spherical model (three-dimensional model), one surface (the number of surfaces), and a square surface (the surface pattern), but if some of the settings in the projection process (e.g., the formula used when converting from three-dimensional space to two-dimensional space, etc., i.e., the elements that make at least one pixel difference in the projection image during the projection process while the remaining projection settings are the same) are different, it can be classified into different formats such as ERP1, EPR2, etc. As another example, in the case of CMP, if it is composed of a cube model, six surfaces, and a square surface, but if some of the settings in the projection process (e.g., the sampling method when converting from three-dimensional space to two-dimensional space, etc.) are different, it can be classified into different formats such as CMP1, CMP2, etc.

[0449] When multiple projection formats are used instead of one already set projection format, the projection format identification information (or projection format information) can be explicitly generated. The projection format identification information can be configured in various ways.

[0450] As an example, index information (e.g., proj_format_flag) can be assigned to multiple projection formats to identify the projection format. For example, ERP can be assigned number 0, CMP can be assigned number 1, OHP can be assigned number 2, ISP can be assigned number 3, ERP1 can be assigned number 4, CMP1 can be assigned number 5, OHP1 can be assigned number 6, ISP1 can be assigned number 7, CMP compact can be assigned number 8, OHP compact can be assigned number 9, ISP compact can be assigned number 10, and other formats can be assigned numbers 11 and above.

[0451] For example, the projection format may be identified from at least one element information constituting the projection format. The element information constituting the projection format may include 3D model information (e.g., 3d_model_flag, where 0 is a sphere, 1 is a cube, 2 is a cylinder, 3 is a pyramid, 4 is a polyhedron 1, and 5 is a polyhedron 2, etc.), number of surfaces (e.g., num_face_flag, which starts from 1 and increases by 1, or the number of surfaces generated in the projection format is assigned as index information, where 0 is 1, 1 is 3, 2 is 6, 3 is 8, and 4 is 20, etc.), shape information of the surface (e.g., shape_face_flag, where 0 is a rectangle, 1 is a circle, 2 is a triangle, 3 is a rectangle + circle, and 4 is a rectangle + triangle, etc.), and projection process setting information (e.g., 3d_2d_convert_idx, etc.).

[0452] As an example, the projection format can be identified by the projection format index information and element information constituting the projection format. For example, the projection format index information can assign 0 to ERP, 1 to CMP, 2 to OHP, 3 to ISP, and 4 or more to other formats, and the projection format (e.g., ERP, ERP1, CMP, CMP1, OHP, OHP1, ISP, ISP1, etc.) can be identified together with the element information constituting the projection format (in this example, projection process setting information). Or, the projection format (e.g., ERP, CMP, CMP compact, OHP, OHP compact, ISP, ISP compact, etc.) can be identified together with the element information constituting the projection format (in this example, whether it is regional packing or not).

[0453] In summary, the projection format can be identified by the projection format index information, can be identified by at least one projection format element information, or can be identified by the projection format index information and at least one projection format element information. This can be defined according to the encoding / decoding setting, and in the present invention, the case where it is identified by the projection format index will be described. In addition, in this example, the case of the projection format represented by surfaces having the same size and shape will be described, but the size and shape of each surface can be different. In addition, the configuration of each surface can be the same or different from that of FIG. 16a to FIG. 16d, and the numbers of each surface are used as symbols to identify each surface and are not limited to a specific order. For convenience of explanation, in the example described later, it is assumed that the projection format of ERP is one surface + rectangle, CMP is six surfaces + rectangle, OHP is eight surfaces + triangle, and ISP is 20 surfaces + triangle is a projection format based on the projection image, and the surfaces have the same size and shape, but the same or similar application is possible for other settings.

[0454] As shown in Figs. 16a to 16d, the projection format can be classified into one surface (e.g., ERP) or multiple surfaces (e.g., CMP, OHP, ISP, etc.). Also, each surface can be classified into a shape such as a rectangle or a triangle. The above classification may be an example of the type and characteristics of the image in the present invention that can be applied when the encoding / decoding setting is different according to the projection format. For example, the type of the image may be a 360-degree image, and the characteristics of the image may be any of the above classifications (e.g., each projection format, a projection format with one surface or multiple surfaces, a projection format with a rectangular or non-rectangular surface, etc.).

[0455] A two-dimensional plane coordinate system {e.g., (i, j)} can be defined for each surface of the two-dimensional projection image, and the characteristics of the coordinate system can vary depending on the projection format, the position of each surface, etc. In the case of ERP, one two-dimensional plane coordinate system, other projection formats can have multiple two-dimensional plane coordinate systems depending on the number of surfaces. In this case, the coordinate system can be expressed as (k, i, j), where k can be the index information of each surface.

[0456] FIG. 17 is a conceptual diagram of a projection format according to an embodiment of the present invention that is realized within a rectangular image.

[0457] That is, 17a to 17c can be understood as realizing the projection formats of Figs. 16b to 16d as rectangular images.

[0458] 17a to 17c, each image format can be configured in a rectangular shape for encoding / decoding a 360-degree image. In the case of ERP, one coordinate system can be used as it is, but in the case of another projection format, the coordinate systems of each surface can be integrated into one coordinate system, and a detailed description thereof will be omitted.

[0459] 17a to 17c, it can be seen that in the process of forming a rectangular image, there are areas filled with meaningless data such as blanks and backgrounds. That is, the rectangular image can be composed of an area containing actual data (in this example, the surface; active area) and a meaningless area filled to form the rectangular image (in this example, assumed to be filled with any pixel value; inactive area). This may cause a decrease in performance due to an increase in the amount of coded data caused by an increase in the size of the image due to the meaningless area, as well as the coding / decoding of actual image data.

[0460] Therefore, further steps may be taken to eliminate meaningless regions and compose the image with regions containing actual data.

[0461] FIG. 18 is a conceptual diagram of how to convert a projection format to a rectangular shape by rearranging surfaces to eliminate non-significant areas in accordance with an embodiment of the present invention.

[0462] Referring to 18a to 18c, an example of rearrangement of 17a to 17c can be seen, and such a process can be defined as a regional packing process (CMP compact, OHP compact, ISP compact, etc.). In this case, not only the rearrangement of the surface itself, but also the surface can be divided and rearranged (OHP compact, ISP compact, etc.). This can be performed for the purpose of not only removing meaningless areas but also improving coding performance through efficient arrangement of the surfaces. For example, when an arrangement is made in which images have continuity between surfaces (for example, B2-B3-B1, B5-B0-B4 in 18a), prediction accuracy during encoding is improved, and thus coding performance can be improved. Here, regional packing according to the projection format is merely an example in the present invention and is not limited thereto.

[0463] FIG. 19 is a conceptual diagram showing how a packing process is performed by converting a CMP projection format into a rectangular image according to an embodiment of the present invention.

[0464] Referring to 19a to 19c, the CMP projection format can be arranged as 6×1, 3×2, 2×3, 1×6. Also, when size adjustment is performed on some surfaces, the projection format can be arranged as shown in 19d to 19e. Although CMP is given as an example in 19a to 19e, the present invention is not limited to CMP and can be applied to other projection formats. The surface arrangement of the image obtained through the regional packing may follow a predetermined rule according to the projection format, or information regarding the arrangement may be explicitly generated.

[0465] A 360-degree image encoding / decoding device according to an embodiment of the present invention may be configured to include all or part of the image encoding / decoding device according to Figs. 1 and 2, and in particular, a format conversion unit and a format inverse conversion unit for converting and inversely converting a projection format may be further included in the image encoding device and the image decoding device, respectively. That is, in the image encoding device of Fig. 1, an input image can be encoded through a format conversion unit, and in the image decoding device of Fig. 2, a bitstream is decoded and then an output image can be generated through a format inverse conversion unit. Hereinafter, the encoder (in this example, "input image" to "encoding") for the above process will be mainly described, and the process in the decoder can be derived in reverse from the encoder. Also, descriptions that overlap with the above content will be omitted.

[0466] Next, the input image will be described on the assumption that it is the same as the 2D projected image or packed image obtained by performing a pre-processing process in the above-mentioned 360-degree encoding device. That is, the input image may be an image obtained by performing a projection process according to a certain projection format or a packing process by region. The projection format already applied to the input image may be any of various projection formats, and may be regarded as a common format or may be called a first format.

[0467] The format conversion unit can convert to a projection format other than the first format. In this case, the format to be converted can be called the second format. For example, ERP can be set as the first format and converted to the second format (for example, ERP2, CMP, OHP, ISP, etc.). In this case, ERP2 can be an EPR format having the same conditions such as the 3D model and surface configuration, but with some settings different. Or, it may be the same format (for example, ERP=ERP2) with the same projection format settings, and the image size or resolution may be different. Or, a part of the image setting process described later may be applied. For convenience of explanation, the above-mentioned examples are given, but the first format and the second format are one of various projection formats, and are not limited to the above examples, and can be changed to other cases.

[0468] In the process of converting between formats, due to the characteristics of different coordinate systems between projection formats, the pixels (integer pixels) of the converted image may be obtained not only from integer unit pixels in the image before conversion, but also from fractional unit pixels, so that interpolation can be performed. In this case, the interpolation filter used may be the same as or similar to the above-mentioned filter. The interpolation filter is selected from a plurality of interpolation filter candidates, and related information may be explicitly generated or may be implicitly determined by a rule that has already been set. For example, a predetermined interpolation filter may be used depending on the projection format, color format, slice / tile type, etc. In addition, when an interpolation filter is explicitly sent, information about the filter information (e.g., filter coefficients, etc.) may also be included.

[0469] The projection format in the format converter may be defined to include regional packing, etc. That is, projection and regional packing processes may be performed during the format conversion process, or regional packing processes may be performed after the format conversion and before encoding.

[0470] The encoder records the information generated in the above process in a bitstream in at least one unit of a sequence, a picture, a slice, a tile, etc., and the decoder parses the relevant information from the bitstream. The information may also be included in the bitstream in the form of SEI or metadata.

[0471] Next, an image setting process applied to a 360-degree image encoding / decoding device according to an embodiment of the present invention will be described. The image setting process of the present invention can be applied to a pre-processing process, a post-processing process, a format conversion process, a format reverse conversion process, etc. in a 360-degree image encoding / decoding device as well as a general encoding / decoding process. The image setting process described below will be described mainly with respect to a 360-degree image encoding device, and can be described including the contents of the image setting described above. A duplicated description of the image setting process described above will be omitted. Also, the example described below will be described mainly with respect to the image setting process, and the image setting reverse process can be derived inversely from the image setting process, and in some cases can be confirmed through various embodiments of the present invention described above.

[0472] The image setting process in the present invention may be performed at the stage of 360-degree image projection, at the stage of regional packing, at the stage of format conversion, or at other stages.

[0473] Fig. 20 is a conceptual diagram of 360-degree image division according to an embodiment of the present invention, in which an image projected by an ERP is assumed.

[0474] 20a shows an image projected by the ERP, which can be divided using various methods. In this example, slices and tiles are mainly described, and W0 to W2 and H0 and H1 are assumed to be slice or tile division boundaries, and are assumed to follow the raster scan order. The examples described below will mainly describe slices and tiles, but are not limited to this, and other division methods can be applied.

[0475] For example, division can be performed in slice units, and division boundaries of H0 and H1 can be provided, or division can be performed in tile units, and division boundaries of W0 to W2 and H0 and H1 can be provided.

[0476] 20b shows an example of dividing an image projected by the ERP into tiles {assuming that the same tile division boundaries (W0 to W2, H0, and H1 are all activated) as in FIG. 20a}. Assuming that the P region is the whole image and the V region is the region where the user's gaze is fixed or the viewport, there are various methods for providing an image corresponding to the viewport. For example, the whole image (e.g., tiles a to l) can be decoded to obtain the region corresponding to the viewport. In this case, the whole image can be decoded, and if it is divided, tiles a to l (A+B region in this example) can be decoded. Alternatively, the region corresponding to the viewport can be obtained by decoding the region belonging to the viewport. In this case, if it is divided, tiles f, g, j, and k (B region in this example) can be decoded to obtain the region corresponding to the viewport from the restored image. The former case can be called whole decoding (or Viewport Independent Coding), and the latter case can be called partial decoding (or Viewport Dependent Coding). The latter case is an example that can occur in a 360-degree image with a large amount of data, and since the division area can be obtained flexibly, a division method in units of tiles is more commonly used than division in units of slices. In the case of partial decoding, since it is not possible to know where the viewport originates, the referability of the division unit can be spatially or temporally restricted (implicitly processed in this example), and encoding / decoding can be performed taking this into consideration. The example described below will focus on the case of full decoding, but in order to prepare for the case of partial decoding, division of a 360-degree image will be described centering on tiles (or the rectangular division method of the present invention). The contents of the example described below can be applied to other division units in the same way or with modifications.

[0477] 21 is an exemplary diagram of 360-degree image division and image reconstruction according to an embodiment of the present invention, which will be described assuming the case of an image projected by CMP.

[0478] 21a shows the projected image in CMP, and the segmentation can be performed using various methods. We assume that W0-W2 and H0, H1 are the segmentation boundaries of the surface, slice, and tile, and follow the raster scan order.

[0479] For example, division can be performed in slice units, and division boundaries of H0 and H1 can be provided. Or, division can be performed in tile units, and division boundaries of W0 to W2 and H0 and H1 can be provided. Or, division can be performed in surface units, and division boundaries of W0 to W2 and H0 and H1 can be provided. In this example, the description will be given assuming that the surface is a part of the division unit.

[0480] In this case, the surface is a division unit (dependent encoding / decoding in this example) made for the purpose of classifying or dividing regions having different properties (e.g., plane coordinate systems of each surface, etc.) in the same image according to the characteristics and type of the image (360-degree image, projection format in this example), and the slice and tile may be division units (independent encoding / decoding in this example) made for the purpose of dividing an image according to a user's definition. In addition, the surface is a unit divided according to a predetermined definition (or induced from projection format information) during the projection process according to the projection format, and the slice and tile may be units divided by explicitly generating division information according to the user's definition. In addition, the surface may have a division shape in the shape of a polygon including a rectangle according to the projection format, the slice may have any division shape that cannot be defined as a rectangle or a polygon, and the tile may have a division shape of a square. The setting of the division unit may be limited and defined for the purpose of explaining this example.

[0481] In the above example, the surface has been described as a division unit classified for the purpose of region division, but depending on the encoding / decoding setting, at least one surface unit may be a unit for performing independent encoding / decoding, or may have a setting for performing independent encoding / decoding in combination with tiles, slices, etc. In this case, there may be cases where explicit information of tiles and slices is generated in combination with tiles, slices, etc., or there may be cases where tiles and slices are implicitly combined based on surface information. Alternatively, there may be cases where explicit information of tiles and slices is generated based on surface information.

[0482] As a first example, one image division process (in this example, the surface) is performed, and the image division can omit division information implicitly (obtaining division information from the projection format information). This example is an example for a dependent encoding / decoding setting, and may be an example corresponding to a case where the referentiality between surface units is not restricted.

[0483] As a second example, an image segmentation process (in this example, the surface) is performed, and the image segmentation can explicitly generate segmentation information. This example is an example for a dependent encoding / decoding setting, and may be an example corresponding to a case where the referentiality between surface units is not restricted.

[0484] As a third example, multiple image segmentation processes (in this example, surface, tile) are performed, some image segmentation processes (in this example, surface) may implicitly omit segmentation information or may explicitly generate segmentation information, and some image segmentation processes (in this example, tiles) may explicitly generate segmentation information. In this example, some image segmentation processes (in this example, surface) precede some image segmentation processes (in this example, tiles).

[0485] As a fourth example, multiple image division processes are performed, and some image divisions (in this example, surfaces) may implicitly omit or explicitly generate division information, and some image divisions (in this example, tiles) may explicitly generate division information based on some image divisions (in this example, surfaces). In this example, some image division processes (in this example, surfaces) precede some image division processes (in this example, tiles). This example is the same as in some cases (assuming the case of the second example) in which division information is explicitly generated, but there may be differences in the configuration of the division information.

[0486] As a fifth example, a plurality of image division processes are performed, and some image divisions (surfaces in this example) can implicitly omit division information, and some image divisions (tiles in this example) can implicitly omit division information based on some image divisions (surfaces in this example). For example, an individual surface unit can be set in tile units, or multiple surface units (in this example, adjacent surfaces are grouped if they are continuous surfaces, and are not grouped otherwise; B2-B3-B1 and B4-B0-B5 in 18a) can be set in tile units. According to a previously set rule, a surface unit can be set to a tile unit. This example is an example of independent encoding / decoding settings, and may be an example corresponding to a case where the reference possibility between surface units is limited. That is, in some cases {assuming the first example case}, the division information is implicitly processed in the same way, but there may be differences in encoding / decoding settings.

[0487] The above examples are explanations of cases where the division process may be performed in the projection step, the regional packing step, the initial encoding / decoding step, etc., but the image division process may occur within other encoders / decoders.

[0488] In 21a, a rectangular image can be constructed by including an area B that does not contain data in an area A that contains data. In this case, the positions, sizes, shapes, and numbers of areas A and B are information that can be confirmed by a projection format, or information that can be confirmed when information on a projected image is explicitly generated, and related information can be indicated by the image division information and image reconstruction information described above. For example, as shown in Tables 4 and 5, information on a part of an area of ​​a projected image (e.g., part_top, part_left, part_width, part_height, part_convert_flag, etc.) can be indicated, and the present invention is not limited to this example and may be an example that can be applied to other cases (e.g., other projection formats, other projection settings, etc.).

[0489] Region B can be formed into one image together with region A and encoded / decoded. Alternatively, different encoding / decoding settings can be set by performing division in consideration of the characteristics of each region. For example, it is not necessary to perform encoding / decoding for region B based on information on whether encoding / decoding is performed (e.g., tile_coded_flag if the division unit is assumed to be a tile). In this case, the corresponding region can be restored to certain data (in this example, an arbitrary pixel value) according to a previously set rule. Alternatively, the encoding / decoding settings in the above-mentioned image division process can be different for region B from region A. Alternatively, the corresponding region can be removed by performing a regional packing process.

[0490] 21b shows an example of dividing an image packed by CMP into tiles, slices, and surfaces. In this case, the packed image is an image that has undergone a surface rearrangement process or a region-based packing process, and may be an image obtained by performing image segmentation and image reconstruction according to the present invention.

[0491] In 21b, a rectangular shape can be formed by including an area including data. At this time, the position, size, shape, number, etc. of each area can be confirmed by a previously set setting, or can be confirmed when information on the packed image is explicitly generated, and related information can be indicated by the image division information, image reconstruction information, etc., described above. For example, as shown in Tables 4 and 5, information on some areas of the packed image (e.g., part_top, part_left, part_width, part_height, part_convert_flag, etc.) can be indicated.

[0492] The packed image can be partitioned using various partitioning methods, for example, the partition can be slice-based and have a partition boundary of H0, or the partition can be tile-based and have partition boundaries of W0, W1 and H0, or the partition can be surface-based and have partition boundaries of W0, W1 and H0.

[0493] The image segmentation and image reconstruction process of the present invention can be performed on the projected image. In this case, the reconstruction process can rearrange the surfaces in the image as well as the pixels in the surfaces. This may be a possible example when the image is segmented or composed into multiple surfaces. The following example will focus on the case where the image is segmented into tiles based on surface units.

[0494] SX,Y (S0,0 to S3,2) in 21a can correspond to S'U,V (S'0,0 to S'2,1. In this example, X,Y may be the same as or different from U,V) in 21b, and a reconstruction process can be performed on a surface-by-surface basis. For example, S2,1, S3,1, S0,1, S1,2, S1,1, and S1,0 can be assigned (or surface rearranged) to S'0,0, S'1,0, S'2,0, S'0,1, S'1,1, and S'2,1. Also, S2,1, S3,1, and S0,1 do not undergo reconstruction (or pixel rearrangement), and S1,2, S1,1, and S1,0 can undergo reconstruction by applying a 90-degree rotation, which can be shown as in FIG. 21c. The symbols displayed horizontally in 21c (S1,0, S1,1, S1,2) may be images laid flat to match the symbols in order to maintain image continuity.

[0495] The surface reconstruction can be an implicit or explicit process depending on the encoding / decoding settings: if implicit, it can be done according to pre-defined rules that take into account the type of image (in this example, a 360° image) and its characteristics (in this example, the projection format, etc.).

[0496] For example, in 21c, S'0,0 and S'1,0, S'1,0 and S'2,0, S'0,1 and S'1,1, and S'1,1 and S'2,1 have image continuity (or correlation) between the two surfaces based on the boundary of the surfaces, and 21c may be an example configured such that there is continuity between the upper three surfaces and the lower three surfaces. The surface is divided into multiple surfaces through a projection process from a three-dimensional space to a two-dimensional space, and reconstruction can be performed for the purpose of increasing image continuity between the surfaces for efficient surface reconstruction through a process of packing by region. Such surface reconstruction can be already set and processed.

[0497] Alternatively, the reconstruction process can be performed by explicit processing to generate reconstruction information for the same.

[0498] For example, when determining information (e.g., either information implicitly obtained or information explicitly generated) for an M×N configuration (e.g., 6×1, 3×2, 2×3, 1×6, etc. in the case of CMP compact; assumed to be a 3×2 configuration in this example) through a regional packing process, surface reconstruction can be performed according to the M×N configuration, and then information for the same can be generated. For example, in the case of intra-image rearrangement of surfaces, index information (or intra-image position information) can be assigned to each surface, and in the case of intra-surface pixel rearrangement, mode information for the reconstruction can be assigned.

[0499] The index information can already be defined as shown in 18a to 18c of Figure 18, and SX, Y or S'U, V in 21a to 21c can represent each surface as position information indicating the horizontal and vertical directions (e.g., S[i][j]) or one position information (e.g., assuming that the position information is assigned in raster scan order from the upper left surface of the image, S[i]), to which an index for each surface can be assigned.

[0500] For example, when indexes are assigned to position information indicating horizontal and vertical directions, in the case of FIG. 21c, S'0,0 can be assigned the index of the second surface, S'1,0 can be assigned the index of the third surface, S'2,0 can be assigned the index of the first surface, S'0,1 can be assigned the index of the fifth surface, S'1,1 can be assigned the index of the zeroth surface, and S'2,1 can be assigned the index of the fourth surface. Or, when indexes are assigned to one position information, S[0] can be assigned the index of the second surface, S[1] can be assigned the index of the third surface, S[2] can be assigned the index of the first surface, S[3] can be assigned the index of the fifth surface, S[4] can be assigned the index of the zeroth surface, and S[5] can be assigned the index of the fourth surface. For convenience of explanation, in the examples described later, S'0,0 to S'2,1 are referred to as a to f. Or, it can be expressed by position information indicating the horizontal and vertical directions of pixels or blocks based on the upper left side of the image.

[0501] In the case of packed images acquired through an image reconstruction process (or a regional packing process), the surface scan order may or may not be the same for the images depending on the reconstruction settings. For example, if one scan order (e.g., raster scan) is applied to 21a, the scan orders of a, b, and c may be the same, and the scan orders of d, e, and f may not be the same. For example, when the scan order of 21a, a, b, and c follows the order (0,0)→(1,0)→(0,1)→(1,1), the scan order of d, e, and f may follow the order (1,0)→(1,1)→(0,0)→(0,1). This can be determined depending on the reconstruction settings of the images, and other projection formats can have such settings as well.

[0502] The image segmentation process in 21b can set individual surface units to tiles. For example, surfaces a-f can each be set to a tile unit. Or, multiple surface units can be set to tiles. For example, surfaces a-c can be set to one tile, and surfaces d-f can be set to one tile. The configuration can be determined based on surface characteristics (e.g., continuity between surfaces, etc.), and surface tiling different from the above example is possible.

[0503] Next, an example of division information by a plurality of image division processes will be described. In this example, the division information for the surface is omitted, and the units other than the surface are tiles, and the division information is processed in various ways.

[0504] As a first example, image segmentation information can be obtained based on surface information and can be implicitly omitted. For example, individual surfaces can be set to tiles, or multiple surfaces can be set to tiles. In this case, when at least one surface is set to a tile can be determined by a predetermined rule based on surface information (e.g., continuity or correlation, etc.).

[0505] As a second example, image partitioning information can be explicitly generated regardless of surface information. For example, when partitioning information is generated by the number of tile rows (num_tile_columns in this example) and the number of columns (num_tile_rows in this example), the partitioning information can be generated by the method in the image partitioning process described above. For example, the range that the number of tile rows and the number of columns can have can be from 0 to the image width / block width (unit obtained from the picture partitioning unit in this example), or from 0 to the image height / block height. In addition, additional partitioning information (e.g., uniform_spacing_flag, etc.) can be generated. In this case, depending on the partitioning setting, the boundary of the surface and the boundary of the partitioning unit may or may not match.

[0506] As a third example, image division information can be explicitly generated based on surface information. For example, when division information is generated based on the number of rows and the number of columns of tiles, the division information can be generated based on surface information (in this example, the range of the number of rows is 0 to 2, and the range of the number of columns is 0, 1, since the surface configuration in the image is 3x2). For example, the range that the number of rows and the number of columns of tiles can have can be from 0 to 2, or from 0 to 1. In addition, additional division information (e.g., uniform_spacing_flag, etc.) may not be generated. In this case, the boundary of the surface and the boundary of the division unit can match.

[0507] In some cases {assuming the second and third examples}, the syntax elements of the division information are defined differently, or even if the same syntax elements are used, the settings of the syntax elements (for example, binarization settings, etc., when the range of candidates that the syntax elements have is limited and small, other binarizations can be used, etc.) can be made different. The above examples have described some of the various configurations of division information, but are not limited to these, and can be understood as examples in which different settings are possible depending on whether the division information is generated based on surface information.

[0508] FIG. 22 is an example diagram of an image projected or packed by CMP divided into tiles.

[0509] At this time, it is assumed that the tile division boundaries are the same as those in 21a of FIG. 21 (W0 to W2, H0, and H1 are all activated), and that the tile division boundaries are the same as those in 21b of FIG. 21 (W0, W1, and H0 are all activated). When it is assumed that the P region is the whole image and the V region is the viewport, full decoding or partial decoding can be performed. This example will mainly explain partial decoding. In 22a, the tiles e, f, and g are decoded in the case of CMP (left), and the tiles a, c, and e are decoded in the case of CMP compact (right), so that the area corresponding to the viewport can be obtained. In 22b, the tiles b, f, and i are decoded in the case of CMP, and the tiles d, e, and f are decoded in the case of CMP compact, so that the area corresponding to the viewport can be obtained.

[0510] In the above example, division into slices, tiles, etc. was described based on surface units (or surface boundaries), but it is also possible to divide within the surface (for example, ERP is composed of a single surface, whereas other projection formats are composed of multiple surfaces), or to divide including the surface boundaries, as shown in 20a of Figure 20.

[0511] Fig. 23 is a conceptual diagram for explaining an example of size adjustment of a 360-degree image according to an embodiment of the present invention. At this time, the explanation will be given assuming the case of an image projected by ERP. In addition, in the example described later, the explanation will be centered on the case of expansion.

[0512] Depending on the image resizing type, the projected image may be resized using a scale factor or an offset factor, where the image before resizing may be P_Width×P_Height and the image after resizing may be P'_Width×P'_Height.

[0513] In the case of a scale factor, after adjusting the size using the scale factor for the image's width and height (in this example, width a, height b), the image's width (P_Width x a) and height (P_Height x b) can be obtained. In the case of an offset factor, after adjusting the size using the offset factor for the image's width and height (in this example, width L, R, height T, B), the image's width (P_Width + L + R) and height (P_Height + T + B) can be obtained. Size adjustment can be performed using a pre-set method, or one of multiple methods can be selected for size adjustment.

[0514] In the following examples, the data processing method will be described mainly in the case of an offset factor. In the case of an offset factor, the data processing method may include a method of filling using a predetermined pixel value, a method of filling by copying outer pixels, a method of filling by copying a part of an image, a method of filling by transforming a part of an image, and the like.

[0515] In the case of a 360-degree image, the size can be adjusted taking into consideration the characteristic that there is continuity at the boundary of the image. In the case of ERP, there is no outer boundary in the 3D space, but when it is converted to 2D space through a projection process, an outer boundary area can exist. The data in the boundary area has continuous data outside the boundary, but can have a boundary due to spatial characteristics. The size can be adjusted taking into consideration such characteristics. At this time, the continuity can be confirmed according to the projection format, etc. For example, in the case of ERP, the image can have boundaries at both ends that are continuous with each other. In this example, the case where the left and right boundaries of the image are continuous and the case where the top and bottom boundaries of the image are continuous will be assumed and described, and the data processing method will be mainly described as a method of filling by copying a part of the image and a method of filling by converting a part of the image.

[0516] When resizing to the left of the image, the area being resized (in this example, LC or TL+LC+BL) can be filled with data from the right area of ​​the image that has continuity with the left side of the image (in this example, tr+rc+br). When resizing to the right of the image, the area being resized (in this example, RC or TR+RC+BR) can be filled with data from the left area of ​​the image that has continuity with the right side (in this example, tl+lc+bl). When resizing to the top of the image, the area being resized (in this example, TC or TL+TC+TR) can be filled with data from the bottom area of ​​the image that has continuity with the top (in this example, bl+bc+br). When resizing to the bottom of the image, the area being resized (in this example, BC or BL+BC+BR) can be filled with data.

[0517] If the size or length of the area to be resized is m, the area to be resized can have a range of (-m, y) to (-1, y) (resized to the left) or a range of (P_Width, y) to (P_Width+m-1, y) (resized to the right) based on the coordinate reference of the image before resizing (in this example, x is 0 to P_Width-1). The position x' of the area to obtain the data of the area to be resized can be derived by the formula x' = (x + P_Width) % P_Width. In this case, x means the coordinate of the area to be resized based on the image coordinate before resizing, and x' means the coordinate of the area referenced by the area to be resized based on the image coordinate before resizing. For example, if you adjust the size to the left, and m is 4 and the image width is 16, then (-4,y) can be obtained from (12,y), (-3,y) can be obtained from (13,y), (-2,y) can be obtained from (14,y), and (-1,y) can be obtained from (15,y). Or, if you adjust the size to the right, and m is 4 and the image width is 16, then (16,y) can be obtained from (0,y), (17,y) can be obtained from (1,y), (18,y) can be obtained from (2,y), and (19,y) can be obtained from (3,y).

[0518] If the size or length of the area to be resized is n, the area to be resized can have a range of (x, -n) to (x, -1) (resized upwards) or a range of (x, P_Height) to (x, P_Height+n-1) (resized downwards) based on the coordinates of the image before resizing (in this example, y is 0 to P_Height-1). The position y' of the area for obtaining the data of the area to be resized can be derived by a formula such as y'=(y+P_Height)%P_Height. In this case, y means the coordinate of the area to be resized based on the image coordinates before resizing, and y' means the coordinate of the area referenced by the area to be resized based on the image coordinates before resizing. For example, if resizing is done upwards, n is 4, and the image height is 16, then (x,-4) can get data from (x,12), (x,-3) can get data from (x,13), (x,-2) can get data from (x,14), and (x,-1) can get data from (x,15). Or, if resizing is done downwards, n is 4, and the image height is 16, then (x,16) can get data from (x,0), (x,17) can get data from (x,1), (x,18) can get data from (x,2), and (x,19) can get data from (x,3).

[0519] After filling the resizing area with data, the resizing image coordinates can be adjusted based on the standard (in this example, x is 0 to P'_Width-1, y is 0 to P'_Height-1). The above example may be an example that can be applied to a latitude and longitude coordinate system.

[0520] It can have various size adjustment combinations:

[0521] As an example, the image may be resized by m to the left, or by n to the right, or by o to the top, or by p to the bottom.

[0522] As an example, the image may be resized to the left by m and to the right by n, or the image may be resized to the top by o and to the bottom by p.

[0523] As an example, the image may be resized by m to the left, n to the right, and o to the top. Or the image may be resized by m to the left, n to the right, and p to the bottom. Or the image may be resized by m to the left, o to the top, and p to the bottom. Or the image may be resized by n to the right, o to the top, and p to the bottom.

[0524] As an example, the image may be resized by m to the left, by n to the right, by o to the top, and by p to the bottom.

[0525] As in the above example, at least one size adjustment is performed, and the size adjustment of the image may be performed implicitly according to the encoding / decoding settings, or size adjustment information may be explicitly generated and the size adjustment of the image may be performed based on the size adjustment information. That is, m, n, o, and p in the above example may be determined to be a predetermined value, or may be explicitly generated as size adjustment information, or some of them may be determined to be a predetermined value and some of them may be explicitly generated.

[0526] Although the above example has been described with a focus on the case where data is obtained from a part of an image, other methods are also applicable. The data may be pixels before encoding or pixels after encoding, and may be determined according to the characteristics of the image or stage in which the size adjustment is performed. For example, when size adjustment is performed in a pre-processing process or a pre-encoding stage, the data may refer to input pixels such as a projected image or a packed image, and when size adjustment is performed in a post-processing process, an intra-prediction reference pixel generation stage, a reference image generation stage, a filter stage, or the like, the data may refer to restored pixels. In addition, size adjustment may be performed using a data processing method for each area to be adjusted in size.

[0527] FIG. 24 is a conceptual diagram for explaining continuity between surfaces in a projection format (eg, CMP, OHP, ISP) according to an embodiment of the present invention.

[0528] In particular, it may be an example for an image consisting of multiple surfaces. Continuity is a characteristic that occurs in adjacent regions in a three-dimensional space, and when converted to a two-dimensional space through a projection process, Figures 24a to 24c can be classified into cases where they are spatially adjacent and have continuity (A), where they are spatially adjacent and do not have continuity (B), where they are not spatially adjacent and have continuity (C), and where they are not spatially adjacent and do not have continuity (D). There is a difference from the general images, which are classified into cases where they are spatially adjacent and have continuity (A) and where they are not spatially adjacent and do not have continuity (D). In this case, when continuity exists, the above-mentioned part of the examples (A or C) apply.

[0529] That is, referring to 24a to 24c, when there is spatial continuity (in this example, description will be given based on 24a), it is represented as b0 to b4, and when there is no spatial continuity, it is represented as B0 to B6. That is, it means the case of adjacent areas in a three-dimensional space, and by using b0 to b4 and B0 to B6 in the encoding process using the characteristic that they have continuity, it is possible to improve the encoding performance.

[0530] FIG. 25 is a conceptual diagram for explaining the continuity of the surface of FIG. 21c, which is an image acquired through an image reconstruction process or a regional packing process in a CMP projection format.

[0531] Here, 21c in Fig. 21 is a rearrangement of 21a in which the 360-degree image is expanded into a cube shape, so even at this time, the continuity of the surface in 21a in Fig. 21 is maintained. That is, as in 25a, surface S2,1 can be connected to S1,1 and S3,1 on the left and right, and can be connected to S1,0 rotated 90 degrees and S1,2 rotated -90 degrees on the top and bottom.

[0532] In a similar manner, continuity for surface S3,1, surface S0,1, surface S1,2, surface S1,1 and surface S1,0 can be confirmed from 25b to 25f.

[0533] The continuity between surfaces can be defined according to the projection format setting, etc., and is not limited to the above example, and other modified examples are possible. The examples described below will be described under the assumption that the continuity shown in Figures 24 and 25 exists.

[0534] FIG. 26 is an exemplary diagram for explaining image size adjustment in the CMP projection format according to an embodiment of the present invention.

[0535] 26a shows an example of resizing an image, 26b shows an example of resizing on a surface basis (or a division basis), and 26c shows an example of resizing (or multiple resizing) on ​​an image and surface basis.

[0536] The projected image can be resized using a scale factor or an offset factor depending on the image size adjustment type, and the image before size adjustment is P_Width×P_Height, the image after size adjustment is P'_Width×P'_Height, and the surface size can be F_Width×F_Height. The sizes may be the same or different depending on the surface, and the width and height of the surface may be the same or different, but in this example, for convenience of explanation, the explanation is given under the assumption that all surfaces in the image are the same size and have a square shape. Also, the explanation is given under the assumption that the size adjustment values ​​(WX, HY in this example) are the same. The data processing method in the example described later will be mainly explained in the case of the offset factor, and the data processing method will be mainly explained in the method of copying and filling a part of the image and the method of converting and filling a part of the image. The above settings can be applied to FIG. 27 as well.

[0537] In the cases of 26a to 26c, the boundary of the surface (assumed to have the continuity according to 24a in FIG. 24 in this example) can have continuity with the boundary of another surface. In this case, it can be divided into a case where the surfaces are spatially adjacent on a two-dimensional plane and have image continuity (first example) and a case where the surfaces are not spatially adjacent on a two-dimensional plane and have image continuity (second example).

[0538] For example, assuming the continuity of 24a in Figure 24, the upper, left, right and lower regions of S1,1 are spatially adjacent to the lower, right, left and upper regions of S1,0, S0,1, S2,1 and S1,2, and the images may also be continuous (first example case).

[0539] Alternatively, the left and right regions of S1,0 may not be spatially adjacent to the upper regions of S0,1 and S2,1, but their images may be continuous with each other (in the case of the second example). Also, the left and right regions of S0,1 and S3,1 may not be spatially adjacent to each other, but their images may be continuous with each other (in the case of the second example). Also, the left and right regions of S1,2 may be continuous with the lower regions of S0,1 and S2,1 (in the case of the second example). This is a limited example in this example, and may be configured differently from the above depending on the definition and setting of the projection format. For convenience of explanation, S0,0 to S3,2 in FIG. 26a are referred to as a to l.

[0540] 26a may be an example of filling with data of an area where continuity exists in the direction of the outer boundary of the image. Areas resized from area A where no data exists (in this example, a0 to a2, c0, d0 to d2, i0 to i2, k0, l0 to l2) can be filled with any predetermined value or via outer pixel padding, and areas resized from area B containing actual data (in this example, b0, e0, h0, j0) can be filled with data of an area (or surface) where image continuity exists. For example, b0 can be filled with data of the upper side of surface h, e0 can be filled with data of the right side of surface h, h0 can be filled with data of the left side of surface e, and j0 can be filled with data of the lower side of surface h.

[0541] In detail, b0 may be an example of filling using the lower surface data obtained by applying a 180 degree rotation to surface h, and j0 may be an example of filling using the upper surface data obtained by applying a 180 degree rotation to surface h, but in this example (and including the examples described below), only the position of the referenced surface is indicated, and the data obtained for the area to be sized can be obtained after an adjustment process (e.g., rotation) that takes into account the continuity between the surfaces, as shown in Figures 24 and 25.

[0542] 26b may be an example of filling using data of an area where continuity exists in the direction of the inner boundary of the image. In this example, the resizing operation performed along the surface may be different. Area A may undergo a shrinking process, and area B may undergo an expanding process. For example, in the case of surface a, resizing (in this example, shrinking) may be performed to the right by w0, and in the case of surface b, resizing (in this example, expanding) may be performed to the left by w0. Or, in the case of surface a, resizing (in this example, shrinking) may be performed to the bottom by h0, and in the case of surface e, resizing (in this example, expanding) may be performed to the top by h0. In this example, when looking at the change in the width of the image from surfaces a, b, c, and d, surface a is shrinking by w0, surface b is expanding by w0 and w1, and surface c is shrinking by w1, so the width of the image before resizing and the width of the image after resizing are the same. Looking at the change in the vertical width of the image from surfaces a, e, and i, surface a is reduced by h0, surface e is expanded by h0 and h1, and surface i is reduced by h1, so the vertical width of the image before and after resizing is the same.

[0543] The areas to be resized (in this example, b0, e0, be, b1, bg, g0, h0, e1, ej, j0, gi, g1, j1, h1) can either be simply removed, considering that they are shrinking from area A where no data exists, or they can be newly filled with data from areas where continuity exists, considering that they are expanding from area B which contains actual data.

[0544] For example, b0 can be filled using the data for the upper side of surface e, e0 can be filled using the data for the left side of surface b, be can be filled using the data for the left side of surface b, or the upper side of surface e, or the weighted sum of the left side of surface b and the upper side of surface e, b1 can be filled using the data for the upper side of surface g, bg can be filled using the data for the left side of surface b, or the upper side of surface g, or the weighted sum of the right side of surface b and the upper side of surface g, g0 can be filled using the data for the right side of surface b, h0 can be filled using the data for the upper side of surface b, e1 can be filled using the data for the left side of surface j, ej can be filled using the data for the underside of surface e, or the left side of surface j, or the weighted sum of the underside of surface e and the left side of surface j, j0 can be filled using the data for the underside of surface e, gj can be filled using the data for the underside of surface g, or the left side of surface j, or the weighted sum of the underside of surface g and the right side of surface j, g1 can be filled using the data for the right side of surface j, j1 can be filled using the data for the underside of surface g, and h1 can be filled using the data for the underside of surface j.

[0545] In the above example, when the area to be resized is filled with data of a part of the image, the data of the part can be copied and filled, or the data of the part can be filled with data obtained after a conversion process based on the characteristics and type of the image. For example, when a 360-degree image is converted into a two-dimensional space according to a projection format, a coordinate system (e.g., a two-dimensional plane coordinate system) for each surface can be defined. For convenience of explanation, it is assumed that (x, y, z) in the three-dimensional space is converted into (x, y, C) or (x, C, z) or (C, y, z) for each surface. The above example shows a case where data of a surface other than the part of the image is obtained in the area to be resized. That is, when the resizing is performed based on the current surface, if data of another surface having a different coordinate system characteristic is copied and filled as it is, there is a possibility that continuity will be distorted based on the boundary of the resizing. For this reason, the data of another surface obtained according to the coordinate system characteristic of the current surface can be converted and filled into the area to be resized. The conversion is merely one example of a data processing method, and is not limited thereto.

[0546] When filling the area to be resized by copying data from a part of the image, the boundary area between the area to be resized (e) and the area to be resized (e0) may contain distorted continuity (or continuity that changes abruptly). For example, a continuous feature may change based on the boundary, similar to an edge that had a straight shape becoming a bent shape based on the boundary.

[0547] When data of a partial area of ​​an image is converted and filled into the area to be resized, the boundary area between the area to be resized and the area to be resized may include a gradually changing continuity.

[0548] The above example may be one example of a data processing method of the present invention, in which data of a part of an image is converted based on the characteristics, type, etc. of the image during the size adjustment process (in this example, expansion), and the obtained data is filled into the area to be adjusted in size.

[0549] 26c may be an example of combining the image resizing processes of 26a and 26b to fill in the image using data of an area where continuity exists in the direction of the image boundaries (inner and outer boundaries). The resizing process of this example can be derived from 26a and 26b, so a detailed description will be omitted.

[0550] 26a may be an example of an image size adjustment process, 26b may be an example of a size adjustment process for a division unit within an image, and 26c may be an example of a plurality of size adjustment processes in which an image size adjustment process and a size adjustment for a division unit within an image are performed.

[0551] For example, size adjustment (in this example, area C) can be performed on an image (in this example, first format) acquired through a projection process, and size adjustment (in this example, area D) can be performed on an image (in this example, second format) acquired through a format conversion process. In this example, size adjustment (in this example, entire image) is performed on an image projected by an ERP, which may be an example of size adjustment (in this example, surface unit) after acquiring an image projected by a CMP through a format conversion unit. The above example is one example of performing multiple size adjustments, and is not limited thereto, and modifications to other cases are possible.

[0552] 27 is an exemplary diagram illustrating size adjustment for an image that has been converted into a CMP projection format and packed according to an embodiment of the present invention. Since FIG. 27 also assumes the continuity between surfaces according to FIG. 25, the boundaries of surfaces can have continuity with the boundaries of other surfaces.

[0553] In this example, the offset factors of W0 to W5 and H0 to H3 (assuming that the offset factors are used as size adjustment values ​​in this example) can have various values. For example, they can be derived from a predetermined value, a motion search range of inter prediction, a unit obtained from a picture division unit, or other values. In this case, the unit obtained from the picture division unit can include a surface. That is, the size adjustment value can be determined based on F_Width and F_Height.

[0554] 27a shows an example of filling each expanded area with data of an area having continuity by adjusting the size of each surface (in this example, in the upper, lower, left, and right directions of each surface). For example, for surface a, continuous data can be filled in its outer boundary a0 to a6, and continuous data can be filled in the outer boundary b0 to b6 of surface b.

[0555] 27b shows an example of filling the expanded area with data of an area where continuity exists by adjusting the size of the multiple surfaces (in this example, in the upper, lower, left, and right directions of the multiple surfaces). For example, the surfaces a, b, and c can be used as references to expand the contour to a0 to a4, b0 to b1, and c0 to c4.

[0556] 27c may be an example of filling the expanded area with data of an area where continuity exists by adjusting the size of the entire image (in this example, in the upper, lower, left, and right directions of the entire image). For example, the outer boundary of the entire image consisting of surfaces a to f can be expanded to a0 to a2, b0, c0 to c2, d0 to d2, e0, and f0 to f2.

[0557] That is, size adjustment can be performed in units of one surface, in units of multiple surfaces where continuity exists, or in units of the entire surface.

[0558] The resized regions in the above example (a0 to f7 in this example) can be filled using data from regions (or surfaces) where there is continuity as in Figure 24a, i.e. the data above, below, left and right of surfaces a to f can be used to fill the resized regions.

[0559] FIG. 28 is an exemplary diagram illustrating a data processing method for adjusting the size of a 360-degree image according to an embodiment of the present invention.

[0560] Referring to FIG. 28, the region B (a0 to a2, ad0, b0, c0 to c2, cf1, d0 to d2, e0, f0 to f2) which is a region to be adjusted in size can be filled with data of a region having continuity among pixel data belonging to a to f. In addition, the region C (ad1, be, cf0) which is another region to be adjusted in size can be filled by mixing data of the region to be adjusted with data of a region that is spatially adjacent and does not have continuity. Or, the region C can be filled by mixing data of the two regions since size adjustment is performed between two regions selected from a to f (e.g., a and d, b and e, c and f). For example, the surface b and the surface e can have a relationship that is spatially adjacent and does not have continuity. The region be between the surfaces b and e to be adjusted in size can be adjusted in size using data of the surface b and data of the surface e. For example, the region be can be filled with a value obtained by averaging the data of the surface b and the data of the surface e, or with a value obtained through a weighted sum according to distance. In this case, the pixels used for data to fill the regions sized by surface b and surface e may be boundary pixels of each surface, but may also be interior pixels of the surfaces.

[0561] In summary, regions that are sized between the division units of an image can be filled with data generated using a mix of data from both units.

[0562] The data processing method may be supported in some situations (in this example, when resizing in multiple regions).

[0563] In Figures 27a and 27b, the regions between division units whose sizes are adjusted are configured separately for each division unit (taking Figure 27a as an example, a6 and d1 are configured for a and d, respectively), but in Figure 28, the regions between division units whose sizes are adjusted can be configured one for each adjacent division unit (one ad1 for a and d). Of course, the above method can be included in the candidate group of data processing methods in Figures 27a to 27c, and size adjustment can be performed using a data processing method different from the above example in Figure 28.

[0564] In the image resizing process of the present invention, a predetermined data processing method may be implicitly used for the region to be resized, or related information may be explicitly generated using one of a plurality of data processing methods. The predetermined data processing method may be any of data processing methods such as a method of filling using an arbitrary pixel value, a method of filling by copying outer pixels, a method of filling by copying a part of an image, a method of filling by transforming a part of an image, and a method of filling with data derived from a plurality of regions of an image. For example, when the region to be resized is located inside an image (e.g., a packed image) and both sides of the region (e.g., a surface) are spatially adjacent and there is no continuity, a data processing method of filling with data derived from a plurality of regions may be applied to fill the region to be resized. In addition, one of the plurality of data processing methods may be selected to perform the resizing, and selection information for this may be explicitly generated. This may be an example that can be applied not only to 360-degree images but also to general images.

[0565] The encoder records the information generated in the above process in at least one unit of a sequence, a picture, a slice, a tile, etc., into a bitstream, and the decoder parses the relevant information from the bitstream. The information may also be included in the bitstream in the form of SEI or metadata. The division, reconstruction, and size adjustment processes for a 360-degree image have been described with a focus on some projection formats such as ERP and CMP, but are not limited thereto, and may be applied to other projection formats in the same or modified form.

[0566] It has been explained that the image setting process applied to the above-mentioned 360-degree image encoding / decoding device can be applied not only to the encoding / decoding process but also to the pre-processing process, post-processing process, format conversion process, format inverse conversion process, etc.

[0567] In summary, the projection process may be configured to include an image setting process. In particular, the projection process may be performed including at least one image setting process. Division may be performed in units of regions (or surfaces) based on the projected image. Division may be performed into one region or multiple regions depending on the projection format. Division information may be generated by the division. Also, the size of the projected image may be adjusted, or the size of the projected region may be adjusted. At this time, size adjustment may be performed for at least one region. Size adjustment information may be generated by the size adjustment. Also, reconstruction (or surface arrangement) of the projected image may be performed, or the projected region may be reconstructed. At this time, reconstruction may be performed for at least one region. Reconstruction information may be generated by the reconstruction.

[0568] In summary, the regional packing process may include an image setting process. In particular, the regional packing process may include at least one image setting process. A division process may be performed in units of regions (or surfaces) based on the packed image. The packed image may be divided into one region or a plurality of regions according to the regional packing setting. Division information base...

Claims

1. A method for decoding a 360-degree image performed by an image decoding device, comprising: receiving a bitstream in which the 360-degree image is encoded, the bitstream including data of an extended two-dimensional image, the extended two-dimensional image including a two-dimensional image and a predetermined extended region, the two-dimensional image being projected from an image having a three-dimensional projection structure selectively determined from a plurality of preset three-dimensional projection structures based on identification information obtained from the bitstream, the two-dimensional image including one or more surfaces; and reconstructing the extended two-dimensional image by decoding the data of the extended two-dimensional image, a size of the extension region is determined based on width information indicating a width of the extension region, the width information being obtained from the bitstream; The sample values ​​of the extension region are determined to be different according to a padding method selected from a plurality of padding methods; the sample values ​​of the extended region are determined based on a characteristic of the two-dimensional image by changing the sample values ​​of the surface to the sample values ​​of the extended region; Reconstructing the extended two-dimensional image includes generating a predicted image; The predicted image is generated by intra prediction; the width information includes at least one of a first width information of the extension region on a left side of the face or a second width information of the extension region on a right side of the face; A method for decoding a 360-degree image, wherein the projection format of the three-dimensional projection structure is determined based on selection information indicating one of a plurality of projection formats including an ERP format that projects the 360-degree image onto a two-dimensional plane and a CMP format that projects the 360-degree image onto a cube.

2. A method for encoding a 360-degree image performed by an image encoding device, comprising: obtaining a two-dimensional image projected from an image having a three-dimensional projection structure selectively determined from a plurality of preset three-dimensional projection structures, wherein identification information of the selected three-dimensional projection structure is encoded into a bitstream, and the two-dimensional image includes one or more surfaces; obtaining an extended two-dimensional image comprising the two-dimensional image and a predetermined extended region; encoding the extended 2D image data into a bitstream in which the 360 ​​degree image is encoded; a size of the extension region is encoded based on width information indicating a width of the extension region, the width information being encoded into the bitstream; The sample values ​​of the extension region are determined to be different according to a padding method selected from a plurality of padding methods; the sample values ​​of the extended region are determined based on a characteristic of the two-dimensional image by changing the sample values ​​of the surface to the sample values ​​of the extended region; Encoding the extended two-dimensional image includes generating a predicted image; The predicted image is generated by intra prediction; the width information includes at least one of a first width information of the extension region on a left side of the face or a second width information of the extension region on a right side of the face; A method for encoding a 360-degree image, wherein the projection format of the three-dimensional projection structure is encoded based on selection information indicating one of a plurality of projection formats including an ERP format for projecting the 360-degree image onto a two-dimensional plane and a CMP format for projecting the 360-degree image onto a cube.

3. A method for transmitting a bitstream, comprising the steps of: obtaining the bitstream generated by a 360-degree image encoding method; transmitting the bitstream to an image decoding device; Including, The method for encoding a 360-degree image includes the steps of: obtaining a projected two-dimensional image from an image having a three-dimensional projection structure selectively determined from a plurality of preset three-dimensional projection structures, wherein identification information of the selected three-dimensional projection structure is encoded in the bitstream, and the two-dimensional image includes one or more surfaces; obtaining an extended two-dimensional image comprising the two-dimensional image and a predetermined extended region; encoding the extended 2D image data into a bitstream in which the 360 ​​degree image is encoded; a size of the extension region is encoded based on width information indicating a width of the extension region, the width information being encoded into the bitstream; The sample values ​​of the extension region are determined to be different according to a padding method selected from a plurality of padding methods; the sample values ​​of the extended region are determined based on a characteristic of the two-dimensional image by changing the sample values ​​of the surface to the sample values ​​of the extended region; Encoding the extended two-dimensional image includes generating a predicted image; The predicted image is generated by intra prediction; the width information includes at least one of a first width information of the extension region on a left side of the face or a second width information of the extension region on a right side of the face; A projection format of the three-dimensional projection structure is encoded based on selection information indicating one of a plurality of projection formats including an ERP format for projecting the 360-degree image onto a two-dimensional plane and a CMP format for projecting the 360-degree image onto a cube. How the bitstream is transmitted.