Image encoding / decoding method, recording medium, and method for transmitting bitstream

Through an encoding and decoding method for 360-degree images, the problem of insufficient image processing performance in the prior art is solved, and more efficient image compression and processing is achieved.

CN119996686APending Publication Date: 2025-05-13INST OF IMAGE TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171912.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-07-17
Filing Date
2017-10-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has insufficient performance in the process of image encoding and decoding, especially when processing 360-degree images, and it is difficult to meet the needs of high-quality images and high-resolution images.

Method used

A method and device are provided for improved encoding and decoding of a 360-degree image, the specific steps include receiving an encoded 360-degree image bitstream, generating a predicted image with reference to syntax information, obtaining a residual image by inverse quantization and inverse transformation, combining the predicted and residual images to obtain a decoded image, and reconstructing the 360-degree image according to the projection format.

Benefits of technology

Through this method and device, the image compression performance can be significantly improved, especially when processing 360-degree images, and the performance of the image processing system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996686A_ABST
    Figure CN119996686A_ABST
Patent Text Reader

Abstract

Disclosed are an image encoding / decoding method, a recording medium, and a method of transmitting a bitstream. The method of decoding a 360-degree image includes: receiving a bitstream encoded with a 360-degree image, the bitstream including data of an extended two-dimensional image, the extended two-dimensional image including a two-dimensional image and a predetermined extension region, and the two-dimensional image being projected from an image having a three-dimensional projection structure and including one or more planes; generating a predicted image by performing prediction on the basis of the information on prediction included in the bitstream; and reconstructing an extended two-dimensional image based on the predicted image and the residual image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with application number 201780073662.3, application date October 10, 2017, national phase entry date May 28, 2019, and invention name “Image data encoding / decoding method and device”. Technical Field

[0002] The present invention relates to image data encoding and decoding technology, and more particularly, to a method and apparatus for encoding and decoding 360-degree images of reality media services. Background Art

[0003] With the popularization of the Internet and mobile terminals and the development of information and communication technology, the use of multimedia data is rapidly increasing. Recently, there has been a demand for high-quality images and high-resolution images such as high-definition (HD) images and ultra-high-definition (UHD) images in various fields, and the demand for reality media services such as virtual reality, augmented reality, etc. is also rapidly increasing. Specifically, since multi-view images captured with multiple cameras are processed for 360-degree images of virtual reality and augmented reality, the amount of data generated for processing has increased significantly, but the performance of the image processing system for processing large amounts of data is insufficient.

[0004] As described above, in the prior art image encoding and decoding methods and apparatuses, there is a need to improve the performance in image processing, especially the performance in image encoding / decoding. Summary of the invention

[0005] Technical issues

[0006] The object of the present invention is to provide a method for improving image setting processing in the initial step for encoding and decoding. More specifically, the present invention is to provide an encoding and decoding method and apparatus for improving image setting processing while taking into account the characteristics of 360-degree images.

[0007] Technical Solutions

[0008] According to one aspect of the present invention, a method for decoding a 360-degree image is provided.

[0009] Here, the method for decoding a 360-degree image may include: receiving a bit stream including an encoded 360-degree image; generating a predicted image with reference to syntax information obtained from the received bit stream; obtaining a decoded image by combining the generated predicted image with a residual image obtained by inverse quantizing and inverse transforming the bit stream; and reconstructing the decoded image into a 360-degree image according to a projection format.

[0010] Here, the syntax information may include projection format information of the 360-degree image.

[0011] Here, the projection format information may be information indicating at least one of the following: an equirectangular projection (ERP) format in which the 360-degree image is projected into a 2D plane; a cube map projection (CMP) format in which the 360-degree image is projected into a cube; an octahedron projection (OHP) format in which the 360-degree image is projected into an octahedron; and an icosahedral projection (ISP) format in which the 360-degree image is projected into a polyhedron.

[0012] Here, the reconstructing may include: acquiring arrangement information according to region-based packing with reference to syntax information; and rearranging blocks of the decoded image according to the arrangement information.

[0013] Here, the generating of the predicted image may include: performing image expansion on a reference picture acquired by restoring the bitstream; and generating the predicted image with reference to the reference picture on which the image expansion is performed.

[0014] Here, performing the image expansion may include performing the image expansion based on a division unit of the reference picture.

[0015] Here, performing the image extension based on the division units may include individually generating the extension area for each division unit by using reference pixels of the division units.

[0016] Here, the expansion area may be generated using boundary pixels of a division unit spatially adjacent to the division unit to be expanded or using boundary pixels of a division unit having image continuity with the division unit to be expanded.

[0017] Here, performing the image expansion based on the division unit may include generating an expanded image of the combined area using boundary pixels of an area combined with two or more division units spatially adjacent to each other among the division units.

[0018] Here, performing the image expansion based on the division units may include generating an expansion area between adjacent division units using all adjacent pixel information of division units that are spatially adjacent to each other among the division units.

[0019] Here, performing the image expansion based on the division unit may include generating the expansion area using an average value of adjacent pixels of spatially adjacent division units.

[0020] Beneficial Effects of the Invention

[0021] By using the image encoding / decoding method and apparatus according to the embodiments of the present invention, compression performance can be enhanced, especially for 360-degree images. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a block diagram of an image encoding apparatus according to an embodiment of the present invention.

[0023] Figure 2 is a block diagram of an image decoding apparatus according to an embodiment of the present invention.

[0024] Figure 3 An example diagram of dividing image information into multiple layers to compress the image.

[0025] Figure 4 is a conceptual diagram showing an example of image division according to an embodiment of the present invention.

[0026] Figure 5 is another exemplary diagram of the image segmentation method according to an embodiment of the present invention.

[0027] Figure 6 is an example diagram of a general image resizing method.

[0028] Figure 7 is an example diagram of image resizing according to an embodiment of the present invention.

[0029] Figure 8 is an exemplary diagram of a method of constructing a region generated by expansion in an image resizing method according to an embodiment of the present invention.

[0030] Fig. 9 is an exemplary diagram of a method of constructing a region to be deleted and a region to be generated in an image resizing method according to an embodiment of the present invention.

[0031] Fig.10 is an exemplary diagram of image reconstruction according to an embodiment of the present invention.

[0032] Fig.11 : are exemplary diagrams showing images before and after image setting processing according to an embodiment of the present invention.

[0033] Fig.12 2 is an exemplary diagram of adjusting the size of each division unit of an image according to an embodiment of the present invention.

[0034] Fig.13 This is an example diagram of a group of settings or size adjustments of division units in an image.

[0035] Fig.14is an exemplary diagram showing both a process of adjusting the size of an image and a process of adjusting the size of a division unit in the image.

[0036] Fig.15 is a diagram showing an example of a 3D space and a two-dimensional (2D) plane space in which a three-dimensional (3D) image is displayed.

[0037] Figures 16a to 16d is a conceptual diagram showing a projection format according to an embodiment of the present invention.

[0038] Fig.17 is a conceptual diagram showing that a projection format is included in a rectangular image according to an embodiment of the present invention.

[0039] Fig.18 is a conceptual diagram of a method of converting a projection format into a rectangular shape, that is, a method of performing rearrangement on a surface to exclude a meaningless area, according to an embodiment of the present invention.

[0040] Fig.19 is a conceptual diagram illustrating a region-wise packing process performed to convert a CMP projection format into a rectangular image according to an embodiment of the present invention.

[0041] Fig. 20 is a conceptual diagram of 360-degree image division according to an embodiment of the present invention.

[0042] Fig.21 360-degree image segmentation and image reconstruction according to an embodiment of the present invention.

[0043] Fig. 22 This is an example of how an image packed or projected by CMP is divided into tiles.

[0044] Fig.23 is a conceptual diagram illustrating an example of adjusting the size of a 360-degree image according to an embodiment of the present invention.

[0045] Fig.24 is a conceptual diagram illustrating continuity between surfaces in a projection format (eg, CHP, OHP, or ISP) according to an embodiment of the present invention.

[0046] Fig.25 21 is a conceptual diagram showing the continuity of the surface of the portion 21c, which is an image acquired by the image reconstruction process of the CMP projection format or the area packing process.

[0047] Fig.26 is an exemplary diagram showing image size adjustment in the CMP projection format according to an embodiment of the present invention.

[0048] Fig. 27 is an example diagram showing resizing of an image converted and packed in a CMP projection format according to an embodiment of the present invention.

[0049] Fig.28 is an exemplary diagram illustrating a data processing method for adjusting the size of a 360-degree image according to an embodiment of the present invention.

[0050] Fig.29 is an example diagram showing a tree-based block format.

[0051] Fig.30 is an example diagram showing a type-based block format.

[0052] Fig.31 is an exemplary diagram showing various types of blocks that can be acquired by the block dividing section of the present invention.

[0053] Fig.32 is an example diagram illustrating tree-based partitioning according to an embodiment of the present invention.

[0054] Fig.33 is an example diagram illustrating tree-based partitioning according to an embodiment of the present invention. DETAILED DESCRIPTION

[0055] Although the present invention may be subjected to various modifications and alternative forms, its specific embodiments are shown as examples in the accompanying drawings and will be described in detail herein. However, it should be understood that there is no intention to limit the present invention to the specific forms disclosed, but on the contrary, the present invention will cover all modifications, equivalents and alternatives that fall within the spirit and scope of the present invention.

[0056] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the scope of the present invention, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. As used herein, the term "and / or" includes any and all combinations of one or more of the items listed in the association.

[0057] It should be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements. Other words used to describe relationships between elements should be interpreted in a similar manner (i.e., "between" versus "directly between," "adjacent" versus "directly adjacent," etc.).

[0058] The professional terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" are intended to include plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "includes" and / or "including" when used herein specify the presence of stated features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or groups thereof.

[0059] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0060] The image encoding device and the image decoding device may be user terminals such as personal computers (PCs), laptop computers, personal digital assistants (PDAs), portable multimedia players (PMPs), portable game consoles (PSPs), wireless communication terminals, smart phones and televisions, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, head-mounted display (HMD) devices and smart glasses, or server terminals such as application servers and service servers, and may include various devices having the following: a communication device such as a communication modem for communicating with various devices or a wired / wireless communication network; a memory for storing various programs and data for encoding or decoding images or performing inter-frame prediction or intra-frame prediction for encoding or decoding; a processor for executing programs to perform calculation and control operations, etc. In addition, the image encoded into a bit stream by the image encoding device may be transmitted to the image decoding device in real time or non-real time through a wired / wireless communication network such as the Internet, a short-range wireless network, a wireless local area network (LAN), a WiBro network, a mobile communication network, or through various communication interfaces such as a cable, a universal serial bus (USB), etc. Then, the bit stream may be decoded by an image decoding device to restore and replay the bit stream as an image.

[0061] Furthermore, an image encoded into a bit stream by the image encoding device may be transmitted from the image encoding device to the image decoding device via a computer-readable recording medium.

[0062] The above-mentioned image encoding device and decoding device may be separate devices, but may be set as one image encoding / decoding device according to the implementation. In this case, some elements of the image encoding device may be substantially the same as some elements of the image decoding device, and may be implemented to include at least the same structure or perform the same function.

[0063] Therefore, in the following detailed description of the technical elements and their working principles, redundant descriptions of the corresponding technical elements will be omitted.

[0064] Furthermore, the image decoding device corresponds to a computing device that applies the image encoding method performed by the image encoding device to a decoding process, and therefore the following description will focus on the image encoding device.

[0065] The computing device may include: a memory configured to store a program or software mode for implementing the image encoding method and / or the image decoding method; and a processor connected to the memory to execute the program. In addition, the image encoding device may also be referred to as an encoder, and the image decoding device may also be referred to as a decoder.

[0066] Generally, an image may consist of a series of still images. Still images may be classified in units of a group of pictures (GOP), and each still image may be referred to as a picture. In this case, a picture may indicate one of a frame and a field in a progressive signal and an interlaced signal. When encoding / decoding is performed on a frame basis, a picture may be represented as a "frame", and when encoding / decoding is performed on a field basis, a picture may be represented as a "field". The present invention assumes a progressive signal, but may also be applied to an interlaced signal. As a higher concept, there may be units such as GOP and sequence, and each picture may also be divided into predetermined areas such as slices, tiles, blocks, etc. In addition, one GOP may include units such as I pictures, P pictures, and B pictures. I pictures may refer to pictures that are autonomously encoded / decoded without using a reference picture, and P pictures and B pictures may refer to pictures that are encoded / decoded by performing processes such as motion estimation and motion compensation using a reference picture. Generally, a P picture may use an I picture and a B picture as a reference picture, and a B picture may use an I picture and a P picture as a reference picture. However, the above definition may also be changed by the setting of encoding / decoding.

[0067] Here, a picture referenced in encoding / decoding is referred to as a reference picture, and a block or pixel referenced in encoding / decoding is referred to as a reference block or reference pixel. In addition, reference data may include various types of encoding / decoding information and frequency domain coefficients and spatial domain pixel values ​​generated and determined during the encoding / decoding process. For example, reference data may correspond to intra-frame prediction information or motion information in a prediction unit, transform information in a transform unit / inverse transform unit, quantization information in a quantization unit / inverse quantization unit, encoding / decoding information (context information) in an encoding unit / decoding unit, filter information in an in-loop filter unit, etc.

[0068] The smallest unit of an image may be a pixel, and the number of bits used to represent one pixel is called a bit depth. Typically, the bit depth may be 8 bits, and a bit depth of 8 bits or more may be supported according to the encoding setting. At least one bit depth may be supported according to the color space. In addition, at least one color space may be included according to the image color format. One or more pictures of the same size or one or more pictures of different sizes may be included according to the color format. For example, YCbCr 4:2:0 may be composed of one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example). At this time, the composition ratio of the chrominance component and the luminance component may be 1:2 in width and height. As another example, YCbCr 4:4:4 may have the same composition ratio in width and height. Similar to the above example, when one or more color spaces are included, the picture may be divided into the color space.

[0069] The present invention will be described based on any color space (Y in this example) of any color format (YCbCr in this example), and the description will apply to other color spaces (Cb and Cr in this example) of the color format in the same or similar manner (depending on the settings of the specific color space). However, each color space may be given a partial difference (independent of the settings of the specific color space). That is, the settings dependent on each color space may refer to settings that are proportional to or dependent on the composition ratio of each component (for example, 4:2:0, 4:2:2, or 4:4:4), and the settings independent of each color space may refer to settings of only the corresponding color space that are independent of the composition ratio of each component or are not related to the composition ratio of each component. In the present invention, some elements may have independent settings or dependent settings depending on the encoder / decoder.

[0070] The setting information or syntax elements required during the image coding process can be determined at the level of units such as video, sequence, picture, slice, tile, block, etc. These units include video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), slice header, tile header and block header. The encoder can add the unit to the bitstream and send the bitstream to the decoder. The decoder can parse the bitstream at the same level, restore the setting information sent by the encoder and use the setting information in the image decoding process. In addition, the relevant information can be sent through the bitstream in the form of supplementary enhancement information (SEI) or metadata, which can then be parsed and then used. Each parameter set has a unique ID value, and the lower parameter set can have an ID value of the upper parameter set to be referenced. For example, the lower parameter set can refer to the information of the upper parameter set with the corresponding ID value among one or more upper parameter sets. Among the various examples of the above units, when any one unit includes one or more different units, any one unit can be referred to as an upper unit, and the included unit can be referred to as a lower unit.

[0071] The setting information occurring in such a unit may include settings that are independent of each unit or settings that are dependent on the previous unit, the subsequent unit, or the upper unit. Here, it will be understood that the dependent setting uses the tag information corresponding to the setting of the previous unit, the subsequent unit, or the upper unit (for example, a 1-bit tag; 1 indicates following, and 0 indicates not following) to indicate the setting information of the corresponding unit. In the present invention, the setting information will be described focusing on the example of independent setting. However, it may also include an example in which the relationship of the setting information that is dependent on the previous unit, the subsequent unit, or the upper unit of the current unit is added to the independent setting or replaces the independent setting.

[0072] Figure 1 is a block diagram of an image encoding apparatus according to an embodiment of the present invention. Figure 2 is a block diagram of an image decoding apparatus according to an embodiment of the present invention.

[0073] Reference Figure 1 The image encoding device may be configured to include a prediction unit, a subtractor, a transform unit, a quantization unit, an inverse quantization unit, an inverse transform unit, an adder, an in-loop filter unit, a memory, and / or an encoding unit, some of which may not necessarily be included. Depending on the implementation, some or all of these elements may be selectively included, and some additional elements not shown herein may be included.

[0074] Reference Figure 2The image decoding device may be configured to include a decoding unit, a prediction unit, an inverse quantization unit, an inverse transformation unit, an adder, an in-loop filter unit, and / or a memory, some of which may not necessarily be included. Depending on the implementation, some or all of these elements may be selectively included, and some additional elements not shown herein may be included.

[0075] The image encoding device and the decoding device may be separate devices, but may be set as one image encoding / decoding device according to the implementation method. In this case, some elements of the image encoding device may be substantially the same as some elements of the image decoding device, and may be implemented to include at least the same structure or perform the same function. Therefore, in the following detailed description of the technical elements and their working principles, redundant descriptions of the corresponding technical elements will be omitted. The image decoding device corresponds to a computing device that applies the image encoding method performed by the image encoding device to the decoding process, so the following description will focus on the image encoding device. The image encoding device may also be referred to as an encoder, and the image decoding device may also be referred to as a decoder.

[0076] The prediction unit can be implemented using a prediction module, and a prediction block can be generated by performing intra-frame prediction or inter-frame prediction on the block to be encoded. The prediction unit generates a prediction block by predicting the current block to be encoded in the image. In other words, the prediction unit can predict the pixel values ​​of the pixels of the current block to be encoded in the image by intra-frame prediction or inter-frame prediction to generate a prediction block with the predicted pixel values ​​of the pixels. In addition, the prediction unit can deliver the information required to generate the prediction block to the encoding unit so that the prediction mode information is encoded. The encoding unit adds the corresponding information to the bitstream and transmits the bitstream to the decoder. The decoding unit of the decoder can parse the corresponding information, restore the prediction mode information, and then use the prediction mode information to perform intra-frame prediction or inter-frame prediction.

[0077] The subtractor subtracts the prediction block from the current block to generate a residual block. In other words, the subtractor can calculate the difference between the pixel value of each pixel of the current block to be encoded and the predicted pixel value of each pixel of the prediction block generated by the prediction unit to generate a residual block as a block type residual signal.

[0078] The transform unit may transform a signal belonging to the spatial domain into a signal belonging to the frequency domain. In this case, the signal obtained by the transform process is referred to as a transform coefficient. For example, a residual block having a residual signal delivered from a subtractor may be transformed into a transform block having a transform coefficient. In this case, the input signal is determined according to the encoding setting, and the input signal is not limited to the residual signal.

[0079] The transform part may perform transform on the residual block by using transform techniques such as Hadamard Transform, transform based on discrete sine transform (DST), and transform based on discrete cosine transform (DCT). However, the present invention is not limited thereto, and various enhanced and modified transform techniques may be used.

[0080] For example, at least one transform technology may be supported, and at least one detailed transform technology may be supported in each transform technology. In this case, at least one detailed transform technology may be a transform technology that constructs some base vectors differently in each transform technology. For example, as transform technologies, DST-based transforms and DCT-based transforms may be supported. For DST, detailed transform technologies such as DST-I, DST-II, DST-III, DST-V, DST-VI, DST-VII, and DST-VIII may be supported, and for DCT, detailed transform technologies such as DCT-I, DCT-II, DCT-III, DCT-V, DCT-VI, DCT-VII, and DCT-VIII may be supported.

[0081] One of the transform technologies may be set as a default transform technology (e.g., one transform technology && one detailed transform technology), and additional transform technologies may be supported (e.g., multiple transform technologies || multiple detailed transform technologies). Whether the additional transform technology is supported may be determined in units of a sequence, a picture, a slice, or a tile, and related information may be generated in units. When the additional transform technology is supported, the transform technology selection information may be determined in units of a block, and related information may be generated.

[0082] The transformation may be performed horizontally and / or vertically. For example, a two-dimensional (2D) transformation may be performed by performing a one-dimensional (1D) transformation horizontally and vertically using basis vectors so that pixel values ​​in the spatial domain may be transformed into the frequency domain.

[0083] In addition, the transformation can be performed horizontally and / or vertically in an adaptive manner. In detail, it can be determined whether to perform the transformation in an adaptive manner according to at least one encoding setting. For intra-frame prediction, for example, when the prediction mode is a horizontal mode, DCT-I can be applied horizontally and DST-I can be applied vertically, when the prediction mode is a vertical mode, DST-VI can be applied horizontally and DCT-VI can be applied vertically, when the prediction mode is a diagonal line to the lower left, DCT-II can be applied horizontally and DCT-V can be applied vertically, and when the prediction mode is a diagonal line to the lower right, DST-I can be applied horizontally and DST-VI can be applied vertically.

[0084] The size and shape of the transform block may be determined according to encoding costs of candidates for the size and shape of the transform block. Image data of the transform block and information on the determined size and shape of the transform block may be encoded.

[0085] Among the transform forms, a square transform may be set as a default transform form, and an additional transform form (e.g., a rectangular form) may be supported. Whether to support an additional transform form may be determined in units of a sequence, a picture, a slice, or a tile, and related information may be generated in units. Transform form selection information may be determined in units of a block, and related information may be generated.

[0086] In addition, it can be determined whether the transform block form is supported according to the coding information. In this case, the coding information may correspond to a slice type, a coding mode, a size and shape of a block, a block partitioning scheme, etc. That is, one transform form can be supported according to at least one piece of coding information, and multiple transform forms can be supported according to at least one piece of coding information. The former situation may be an implicit situation, and the latter situation may be an explicit situation. For the explicit situation, adaptive selection information indicating the best candidate group selected from multiple candidate groups can be generated, and the adaptive selection information is added to the bitstream. According to the present invention, in addition to this example, it will be understood that when the coding information is explicitly generated, the information is added to the bitstream in various units, and the relevant information is parsed in various units by the decoder and the relevant information is restored to the decoding information. In addition, it should be understood that when the coding / decoding information is implicitly processed, the processing is performed by the encoder and the decoder through the same processing, rules, etc.

[0087] As an example, support for rectangular transform may be determined according to the slice type. For I slices, the supported transform form may be square transform, and for P / B slices, the supported transform form may be square transform or rectangular transform.

[0088] As an example, support for rectangular transform may be determined according to the coding mode.For intra prediction, the supported transform form may be square transform, and for inter prediction, the supported transform form may be square transform and / or rectangular transform.

[0089] As an example, support for rectangular transform may be determined based on the size and shape of the block. The transform supported by a block of a certain size or larger may be a square transform, and the transform supported by a block smaller than a certain size may be a square transform and / or a rectangular transform.

[0090] As an example, support for rectangular transform may be determined according to a block partitioning scheme. When the block to be transformed is a block obtained by a quadtree partitioning scheme, the supported transform form may be a square transform. When the block to be transformed is a block obtained by a binary tree partitioning scheme, the supported transform form may be a square transform or a rectangular transform.

[0091] The above examples may be examples of support for transforms according to one piece of encoding information, and multiple pieces of information may be associated with additional transform support settings in combination. The above examples are merely examples of additional transform support according to various encoding settings. However, the present invention is not limited thereto, and various modifications may be made thereto.

[0092] The transform process may be omitted according to the encoding setting or the image characteristics. For example, the transform process (including the inverse process) may be omitted according to the encoding setting (for example, in this example, assuming a lossless compression environment). As another example, the transform process may be omitted when the compression performance by the transform is not shown according to the image characteristics. In this case, the transform may be omitted for all units or one of the horizontal unit and the vertical unit. Whether the omission is supported may be determined according to the size and shape of the block.

[0093] For example, it is assumed that the horizontal transform and the vertical transform are set to be omitted together. When the transform omission flag is 1, the transform may be performed neither horizontally nor vertically, and when the transform omission flag is 0, the transform may be performed horizontally and vertically. On the other hand, it is assumed that the horizontal transform and the vertical transform are set to be omitted independently. When the first transform omission flag is 1, the horizontal transform is not performed, and when the first transform omission flag is 0, the horizontal transform is performed. Moreover, when the second transform omission flag is 1, the vertical transform is not performed, and when the second transform omission flag is 0, the vertical transform is performed.

[0094] When the size of the block corresponds to range A, transform omission can be supported, and when the size of the block corresponds to range B, transform omission cannot be supported. For example, when the width of the block is greater than M or the height of the block is greater than N, the transform omission flag cannot be supported. When the width of the block is less than m or the height of the block is less than n, the transform omission flag can be supported. M(m) and N(n) can be the same as or different from each other. The settings associated with the transform can be determined in units of sequences, pictures, slices, etc.

[0095] When additional transform techniques are supported, the transform technique setting may be determined according to at least one piece of encoding information. In this case, the encoding information may correspond to a slice type, a coding mode, a size and shape of a block, a prediction mode, etc.

[0096] As an example, support for transform techniques may be determined based on the coding mode. For intra-frame prediction, supported transform techniques may include DCT-I, DCT-III, DCT-VI, DST-II, and DST-III, and for inter-frame prediction, supported transform techniques may include DCT-II, DCT-III, and DST-III.

[0097] As an example, support for transform techniques may be determined based on the slice type. For I slices, supported transform techniques may include DCT-I, DCT-II, and DCT-III, for P slices, supported transform techniques may include DCT-V, DST-V, and DST-VI, and for B slices, supported transform techniques may include DCT-I, DCT-II, and DST-III.

[0098] As an example, support for transform technology may be determined according to the prediction mode. Transform technologies supported by prediction mode A may include DCT-I and DCT-II, transform technologies supported by prediction mode B may include DCT-I and DST-I, and transform technologies supported by prediction mode C may include DCT-I. In this case, prediction mode A and prediction mode B may both be directional modes, and prediction mode C may be a non-directional mode.

[0099] As an example, support for transform technologies may be determined based on the size and shape of the block. Transform technologies supported by blocks of a specific size or larger may include DCT-II, transform technologies supported by blocks smaller than a specific size may include DCT-II and DST-V, and transform technologies supported by blocks of a specific size or larger than a specific size and blocks smaller than a specific size may include DCT-I, DCT-II, and DST-I. In addition, transform technologies supported in a square form may include DCT-I and DCT-II, and transform technologies supported in a rectangular shape may include DCT-I and DST-I.

[0100] The above examples can be examples of support for transform technology based on one piece of coding information, and multiple pieces of information can be associated with additional transform technology support settings in combination. The present invention is not limited to the above examples, and the above examples can be modified. In addition, the transform unit can deliver the information required to generate the transform block to the encoding unit so that the information is encoded. The encoding unit adds the corresponding information to the bit stream and transmits the bit stream to the decoder. The decoding unit of the decoder can parse the information and use the parsed information in the inverse transform process.

[0101] The quantization unit may quantize the input signal. In this case, the signal obtained by the quantization process is referred to as a quantization coefficient. For example, the quantization unit may quantize the residual block having the residual transform coefficient delivered from the transform unit, thereby obtaining a quantization block having the quantization coefficient. In this case, the input signal is determined according to the encoding setting and is not limited to the residual transform coefficient.

[0102] The quantization unit may use quantization techniques such as Dead Zone Uniform Threshold Quantization, quantization weighting matrix, etc. to quantize the transformed residual block. However, the present invention is not limited to this, and various improved and modified quantization techniques may be used. Whether additional quantization techniques are supported may be determined in units of sequences, pictures, slices, or tiles, and relevant information may be generated based on the units. When additional quantization techniques are supported, quantization technique selection information may be determined in units of blocks, and relevant information may be generated.

[0103] When additional quantization techniques are supported, the quantization technique setting may be determined according to at least one piece of encoding information. In this case, the encoding information may correspond to a slice type, an encoding mode, a size and shape of a block, a prediction mode, etc.

[0104] For example, the quantization section may differently set the quantization weighting matrix corresponding to the coding mode and the weighting matrix applied according to the inter-frame prediction / intra-frame prediction. In addition, the quantization section may differently set the weighting matrix applied according to the intra-frame prediction mode. In this case, when it is assumed that the quantization weighting matrix has the same size of M×N as the size of the quantization block, the quantization weighting matrix may be a quantization matrix in which some quantization components are constructed differently.

[0105] The quantization process may be omitted according to the encoding setting or the image characteristics. For example, the quantization process (including the inverse process) may be omitted according to the encoding setting (for example, such as, in this example, assuming a lossless compression environment). As another example, when the compression performance by quantization is not shown according to the image characteristics, the quantization process may be omitted. In this case, some or all of the regions may be omitted, and whether the omission is supported may be determined according to the size and shape of the block.

[0106] The information about the quantization parameter (QP) may be generated in units of a sequence, a picture, a slice, a tile, or a block. For example, the information about the quantization parameter (QP) may be generated in the upper unit where the QP information is first generated. <1> The default QP is set in the quantization process performed in some units by this process, and the QP can be set to a value that is the same as or different from the value of the QP set in the upper unit. In this case, the QP can be finally determined. In this case, the units such as sequences and pictures can be <1> Corresponding examples, units such as slices, tiles, and blocks may be <2> corresponding examples, and units such as blocks may be <3> The corresponding example.

[0107] Information about the QP may be generated based on the QP in each unit. Alternatively, a predetermined QP may be set as a prediction value, and information about the difference from the QP in the unit may be generated. Alternatively, a QP acquired based on at least one of the QP set in the upper unit, the QP set in the same and previous units, or the QP set in the adjacent unit may be set as a prediction value, and information about the difference from the QP in the current unit may be generated. Alternatively, the QP set in the upper unit and the QP acquired based on at least one piece of encoding information may be set as prediction values, and difference information from the QP in the current unit may be generated. In this case, the same and previous units may be units that may be defined in the order in which the units are encoded, the adjacent units may be spatially adjacent units, and the encoding information may be a slice type, encoding mode, prediction mode, position information, etc. of the corresponding unit.

[0108] As an example, the QP in the upper unit may be set as a predicted value using the QP in the current unit and difference information may be generated. Information on the difference between the QP set in the slice and the QP set in the picture may be generated, or information on the difference between the QP set in the tile and the QP set in the picture may be generated. In addition, information on the difference between the QP set in the block and the QP set in the slice or tile may be generated. In addition, information on the difference between the QP set in the sub-block and the QP set in the block may be generated.

[0109] As an example, a QP acquired based on a QP in at least one adjacent unit or a QP in at least one previous unit may be set as a prediction value using a QP in a current unit, and difference information may be generated. Information about a difference from a QP acquired based on a QP of an adjacent block such as a block on the left, upper left, lower left, upper side, upper right side, etc. of the current block may be generated. Alternatively, information about a difference from a QP of a coded picture before the current picture may be generated.

[0110] As an example, the QP in the upper unit and the QP acquired based on at least one piece of encoding information may be used to set the QP in the upper unit as a prediction value and generate difference information. In addition, information about the difference between the QP in the current block and the QP of the slice corrected according to the slice type (I / P / B) may be generated. Alternatively, information about the difference between the QP in the current block and the QP of the tile corrected according to the coding mode (intra / inter) may be generated. Alternatively, information about the difference between the QP in the current block and the QP of the picture corrected according to the prediction mode (directional / non-directional) may be generated. Alternatively, information about the difference between the QP in the current block and the QP of the picture corrected according to the position information (x / y) may be generated. In this case, the correction may refer to an operation of adding an offset to the QP in the upper unit for prediction or subtracting an offset from the QP. In this case, at least one piece of offset information may be supported according to the encoding setting, and implicitly processed information or explicitly associated information may be generated according to a predetermined process. The present invention is not limited to the above examples, and the above examples may be modified.

[0111] The above example may be an example that is allowed when a signal indicating a QP change is provided or activated. For example, when neither a signal indicating a QP change is provided nor a signal indicating a QP change is activated, difference information is not generated, and the predicted QP may be determined as the QP in each unit. As another example, when a signal indicating a QP change is provided or activated, difference information is generated, and when the value of the difference information is 0, the predicted QP may be determined as the QP in each unit.

[0112] The quantization unit can deliver the information required to generate the quantization block to the encoding unit so that the information is encoded. The encoding unit adds the corresponding information to the bit stream and transmits the bit stream to the decoder. The decoding unit of the decoder can parse the information and use the parsed information in the inverse quantization process.

[0113] The above example has been described under the assumption that the residual block is transformed and quantized by the transform section and the quantization section. However, the residual signal of the residual block may be transformed into a residual block having a transform coefficient without performing a quantization process. Alternatively, only a quantization process may be performed without transforming the residual signal of the residual block into a transform coefficient. Alternatively, neither a transform process nor a quantization process may be performed. This may be determined according to the encoding setting.

[0114] The encoding unit may scan the quantization coefficients, transform coefficients or residual signals of the generated residual block in at least one scanning order (e.g., zigzag scanning, vertical scanning, horizontal scanning, etc.), generate a quantization coefficient string, a transform coefficient string or a signal string and encode the quantization coefficient string, the transform coefficient string or the signal string using at least one entropy coding technique. In this case, information about the scanning order may be determined according to encoding settings (e.g., encoding mode, prediction mode, etc.), and information about the scanning order may be used to generate implicitly determined information or explicitly associated information. For example, a scanning order may be selected from a plurality of scanning orders according to an intra-frame prediction mode.

[0115] In addition, the encoding unit may generate encoded data including the encoding information delivered from each element, and may output the encoded data in a bit stream. This may be implemented with a multiplexer (MUX). In this case, encoding may be performed using methods such as exponential Golomb, context adaptive variable length coding (CAVLC), and context adaptive binary arithmetic coding (CABAC) as encoding techniques. However, the present invention is not limited thereto, and various encoding techniques obtained by improving and modifying the above encoding techniques may be used.

[0116] When entropy coding (e.g., CABAC in this example) is performed on syntax elements such as information generated by the encoding / decoding process and residual block data, the entropy coding device may include a binarizer, a context modeler, and a binary arithmetic encoder. In this case, the binary arithmetic encoder may include a conventional coding engine and a bypass coding engine.

[0117] The syntax elements input to the entropy encoding device may not be binary values. Therefore, when the syntax elements are not binary values, the binarizer may binarize the syntax elements and output a binary bit string (binstring) consisting of 0 or 1. In this case, the binary bit (bin) represents a bit consisting of 0 or 1, and can be encoded by a binary arithmetic encoder. In this case, one of the conventional encoding engine and the bypass encoding engine can be selected based on the probability of occurrence of 0 and 1, and this can be determined according to the encoding / decoding settings. When the syntax element is data where the frequency of 0 is equal to the frequency of 1, the bypass encoding engine can be used; otherwise, the conventional encoding engine can be used.

[0118] When binarizing a syntax element, various methods can be used. For example, fixed-length binarization, unary binarization, Truncated Rice Binarization, K-order exponential Columbus binarization, etc. can be used. In addition, signed binarization or unsigned binarization can be performed according to the range of the value of the syntax element. The binarization processing of the syntax element according to the present invention may include additional binarization methods as well as the binarization described in the above example.

[0119] The inverse quantization unit and the inverse transform unit can be implemented by inversely performing the processes performed in the transform unit and the quantization unit. For example, the inverse quantization unit can inversely quantize the transform coefficients quantized by the quantization unit, and the inverse transform unit can inversely transform the inversely quantized transform coefficients to generate a restored residual block.

[0120] The adder adds the predicted block and the restored residual block to restore the current block. The restored block can be stored in a memory and can be used as reference data (for the prediction part, the filter part, etc.).

[0121] The in-loop filter section may additionally perform post-processing filtering processing of one or more of a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), and the like. The deblocking filter may remove block distortion generated at the boundary between blocks from the restored image. The ALF may perform filtering based on a value obtained by comparing the input image with the restored image. In detail, the ALF may perform filtering based on a value obtained by comparing the input image with an image restored after filtering the block by the deblocking filter. Alternatively, the ALF may perform filtering based on a value obtained by comparing the input image with an image restored after filtering the block by the SAO. The SAO may restore the offset difference based on a value obtained by comparing the input image with the restored image, and may be applied in the form of a band offset (BO), an edge offset (EO), and the like. In detail, the SAO may add an offset relative to the original image to the restored image to which the deblocking filter is applied in units of at least one pixel, and may be applied in the form of BO, EO, and the like. In detail, SAO may add an offset with respect to an original image in units of pixels to an image restored after filtering a block through ALF, and may be applied in the form of BO, EO, and the like.

[0122] As filtering information, setting information about whether each post-processing filter is supported can be generated in units of sequence, picture, slice, tile, etc. In addition, setting information about whether each post-processing filter is executed can be generated in units of picture, slice, tile, block, etc. The scope of executing the filter can be classified into the interior of the image and the boundary of the image. Setting information considering classification can be generated. In addition, information about filtering operation can be generated in units of picture, slice, tile, block, etc. Information can be processed implicitly or explicitly, and independent filtering processing or dependent filtering processing can be applied to filtering according to color components, and this can be determined according to encoding settings. The in-loop filter unit can deliver the filtering information to the encoding unit so that the information is encoded. The encoding unit adds the corresponding information to the bit stream and transmits the bit stream to the decoder. The decoding unit of the decoder can parse the information and apply the parsed information to the in-loop filter unit.

[0123] The memory can store the restored blocks or pictures. The restored blocks or pictures stored in the memory can be provided to the prediction unit, which performs intra-frame prediction or inter-frame prediction. In detail, for processing, the space for storing the bit stream compressed by the encoder in the form of a queue can be set to a coded picture buffer (CPB), and the space for storing the decoded image in units of pictures can be set to a decoded picture buffer (DPB). The CPB can store the decoding unit in decoding order, simulate the decoding operation in the encoder, and store the compressed bit stream through the simulation process. The bit stream output from the CPB is restored through the decoding process, and the restored image is stored in the DPB, and the picture stored in the DPB can be referenced during the image encoding / decoding process.

[0124] The decoding unit can be implemented by reversely performing the processing of the encoding unit. For example, the decoding unit can receive a quantization coefficient string, a transformation coefficient string, or a signal string from a bit stream, decode the string, parse the decoded data including the decoding information, and deliver the parsed decoded data to each element.

[0125] Next, the image setting process applied to the image encoding / decoding device according to the embodiment of the present invention will be described. This is an example (initial image setting) applied before encoding / decoding, but some processes may be examples to be applied to other steps (e.g., steps after encoding / decoding or sub-steps of encoding / decoding). The image setting process can be performed in the case of a network and user environment taking into account, for example, multimedia content characteristics, bandwidth, user terminal performance, and accessibility. For example, image division, image resizing, image reconstruction, etc. can be performed according to the encoding / decoding settings. The following description of the image setting process focuses on rectangular images. However, the present invention is not limited thereto, and the image setting process can be applied to polygonal images. The same image setting can be applied regardless of the image form, or different image settings can be applied, which can be determined according to the encoding / decoding settings. For example, after checking information about the image shape (e.g., a rectangular shape or a non-rectangular shape), information about the corresponding image settings can be constructed.

[0126] The following examples will be described assuming that dependent settings are provided for color spaces. However, independent settings may be provided for color spaces. In addition, in the following examples, independent settings may include independently providing encoding / decoding settings for each color space. Although one color space is described, it is expected that examples in which the description is applied to other color spaces (for example, an example in which N is generated in a chrominance component when M is generated in a luminance component) are included, and the example can be derived. In addition, the dependent settings may include examples in which settings are made in proportion to a color format composition ratio (for example, 4:4:4, 4:2:2, 4:2:0, etc.) (for example, for 4:2:0, when the luminance component is M, the chrominance component is M / 2). It is expected that examples in which the description is applied to each color space are included, and the example can be derived. The description is not limited to the above examples, and can be applied to the present invention in common.

[0127] Some of the constructions in the following examples can be applied to various coding techniques, such as spatial domain coding, frequency domain coding, block-based coding, object-based coding, etc.

[0128] Typically, the input image can be encoded or decoded as is or after image segmentation. For example, segmentation can be performed for error robustness, etc., to prevent damage caused by packet loss during transmission. Alternatively, segmentation can be performed to classify regions with different properties in the same image according to the characteristics, type, etc. of the image.

[0129] According to the present invention, the image division process may include a division process and an inverse division process. The following example description will focus on the division process, but the inverse division process may be derived from the division process in reverse.

[0130] Figure 3 An example diagram of dividing image information into multiple layers to compress the image.

[0131] Part 3a is an example diagram in which an image sequence is composed of multiple GOPs. In addition, a GOP can be composed of an I picture, a P picture, and a B picture, as shown in part 3b. A picture can be composed of slices, tiles, etc., as shown in part 3c. As shown in part 3d, slices, tiles, etc. can be composed of multiple default coding parts, and as shown in part 3e, the default coding part can be composed of at least one coding sub-unit. The image setting process according to the present invention will be described based on examples to be applied to units such as pictures, slices, and tiles as shown in parts 3b and 3c.

[0132] Figure 4 is a conceptual diagram showing an example of image division according to an embodiment of the present invention.

[0133] Part 4a is a conceptual diagram in which an image (e.g., a picture) is divided horizontally and vertically at uniform intervals. The divided areas may be referred to as blocks. Each block may be a default encoding portion (or a maximum encoding portion) obtained by the picture division unit, and may be a basic unit to be applied to the division unit to be described below.

[0134] Part 4b is a conceptual diagram in which an image is divided in at least one direction selected from a horizontal direction and a vertical direction. The divided regions (T0 to T3) may be referred to as tiles, and each region may be encoded or decoded independently or dependently on other regions.

[0135] Part 4c is a conceptual diagram in which an image is divided into groups of continuous blocks. The divided areas (S0, S1) may be referred to as slices, and each area may be encoded or decoded independently or dependent on other areas. The groups of continuous blocks may be defined according to a scan order. Typically, the groups of continuous blocks conform to a raster scan order. However, the present invention is not limited thereto, and the groups of continuous blocks may be determined according to the encoding / decoding settings.

[0136] Part 4d is a conceptual diagram of dividing an image into groups of blocks according to any user-defined settings. The divided areas (A0 to A2) can be called arbitrary divisions, and each area can be encoded or decoded independently or dependent on other areas.

[0137] Independent encoding / decoding may mean that when encoding or decoding some units (or areas), data in other units cannot be referenced. In detail, multiple pieces of information used or generated during texture encoding and entropy encoding for some units may be encoded independently without reference to each other. Even in the decoder, for texture decoding and entropy decoding for some units, parsing information and recovery information in other units may not refer to each other. In this case, whether to refer to data in other units (or areas) may be limited to spatial areas (for example, between areas in one image), but may also be limited to time areas (for example, between consecutive images or between frames) according to the encoding / decoding settings. For example, when some units of the current image and some units of another image have continuity or have the same encoding environment, reference may be made; otherwise, the reference may be restricted.

[0138] In addition, dependent encoding / decoding may mean that when encoding or decoding some units, data in other units may be referenced. In detail, multiple pieces of information used or generated during texture encoding and entropy encoding for some units may be dependently encoded and referenced to each other. Even in the decoder, for texture decoding and entropy decoding for some units, parsing information and recovery information in other units may reference each other. That is, the above setting may be the same or similar to the setting of general encoding / decoding. In this case, in order to identify an area (here, a surface generated according to a projection format), <face>The regions may be divided according to the characteristics and types of the images (eg, 360-degree images).

[0139] In the above examples, independent encoding / decoding settings (e.g., independent slice segments) may be provided for some units (slices, tiles, etc.), and dependent encoding / decoding settings (e.g., dependent slice segments) may be provided for other units. According to the present invention, the following description will focus on independent encoding / decoding settings.

[0140] As shown in section 4a, the default coding portion obtained by the picture division unit may be divided into default coding blocks according to the color space, and may have a size and shape determined according to the characteristics and resolution of the image. The size or shape of the supported block may be an exponential multiple represented by 2 (2 n ) of width and height of an N×N square (2 n ×2 n ; 256×256, 128×128, 64×64, 32×32, 16×16, 8×8, etc.; n is an integer in the range of 3 to 8) or an M×N rectangle (2 m ×2 n ). For example, the input image may be divided according to resolution, into 128×128 for an 8k UHD image, into 64×64 for a 1080p HD image, or into 16×16 for a WVGA image, and the input image may be divided according to image type, into 256×256 for a 360-degree image. The default coding section may be divided into coding sub-units and then encoded or decoded. Information about the default coding section may be added to the bitstream in units of sequence, picture, slice, tile, etc., and may be parsed by the decoder to recover relevant information.

[0141] The image encoding method and the image decoding method according to the embodiment of the present invention may include the following image division step. In this case, the image division process may include an image division indication step, an image division type identification step, and an image division execution step. In addition, the image encoding device and the image decoding device may be configured to include an image division indication unit, an image division type identification unit, and an image division execution unit that respectively perform the image division indication step, the image division type identification step, and the image division execution step. For encoding, relevant syntax elements may be generated. For decoding, relevant syntax elements may be parsed.

[0142] In the block division process, as shown in part 4a, the image division instruction section may be omitted. The image division type identification section may check information on the size and shape of the block, and the image division section may perform division by the division type information identified in the default encoding section.

[0143] A block may be a unit to be always divided, but whether to divide other division units (tiles, slices, etc.) may be determined according to encoding / decoding settings. As a default setting, the picture division unit may perform division in units of blocks and then perform division in other units. In this case, block division may be performed based on the picture size.

[0144] In addition, partitioning can be performed in blocks after partitioning is performed in other units (tiles, slices, etc.). That is, block partitioning can be performed based on the size of the partitioning unit. This can be determined by explicit or implicit processing according to the encoding / decoding settings. The following example description assumes the former case and will also focus on units other than blocks.

[0145] In the image division indication step, it may be determined whether to perform image division. For example, when a signal indicating image division (eg, tiles_enabled_flag) is confirmed, division may be performed. When a signal indicating image division is not confirmed, division may not be performed, or division may be performed by confirming other encoding / decoding information.

[0146] In detail, it is assumed that a signal indicating image division (e.g., tiles_enabled_flag) is confirmed. When the signal is activated (e.g., tiles_enabled_flag=1), division may be performed in multiple units. When the signal is deactivated (e.g., tiles_enabled_flag=0), division may not be performed. Alternatively, the absence of confirmation of a signal indicating image division may mean that division is not performed or division is performed in at least one unit. Whether division is performed in multiple units may be confirmed by another signal (e.g., first_slice_segment_in_pic_flag).

[0147] In summary, when a signal indicating image division is provided, the corresponding signal is a signal for indicating whether to perform division in multiple units. Whether to divide the corresponding image can be determined based on the signal. For example, it is assumed that tiles_enabled_flag is a signal indicating whether to divide the image. Here, tiles_enabled_flag equal to 1 can indicate that the image is divided into multiple tiles, and tiles_enabled_flag equal to 0 can indicate that the image is not divided.

[0148] In short, when a signal indicating image division is not provided, division may not be performed, or whether the corresponding image is divided may be determined by another signal. For example, first_slice_segment_in_pic_flag is not a signal indicating whether image division is performed, but a signal indicating the first slice section in the image. Therefore, it can be confirmed whether division is performed in two or more units (for example, a flag of 0 indicates that the image is divided into a plurality of slices).

[0149] The present invention is not limited to the above examples, and the above examples may be modified. For example, a signal indicating image division may not be provided for each tile, and a signal indicating image division may be provided for each slice. Alternatively, a signal indicating image division may be provided based on the type, characteristics, etc. of the image.

[0150] In the image division type identification step, the image division type may be identified. The image division type may be defined by a division method, division information, and the like.

[0151] In section 4b, a tile may be defined as a unit obtained by horizontal and vertical division. In detail, a tile may be defined as a group of adjacent blocks in a quadrilateral space divided by at least one horizontal or vertical division line passing through an image.

[0152] The tile partition information may include the boundary position information of the columns and rows, the number of tiles of the columns and rows, the tile size information, etc. The tile number information may include the number of columns of the tile (e.g., num_tile_columns) and the number of rows of the tile (e.g., num_tile_rows). Therefore, the image may be divided into a number (=number of columns × number of rows) of tiles. The tile size information may be acquired based on the tile number information. The width or height of the tile may be uniform or non-uniform, so under a predetermined rule, the relevant information (e.g., uniform_spacing_flag) may be implicitly determined or explicitly generated. In addition, the tile size information may include the size information of each column and each row of the tile (e.g., column_width_tile[i] and row_height_tile[i]), or the size information of the width and height of each tile. In addition, the size information may be information that may be additionally generated depending on whether the tile size is uniform (e.g., in the case where the partitioning is non-uniform because uniform_spacing_flag is 0).

[0153] In section 4c, a slice may be defined as a unit for grouping consecutive blocks. In detail, a slice may be defined as a group of consecutive blocks in a predetermined scanning order (here, in a raster scan).

[0154] The slice partitioning information may include slice number information, slice position information (eg, slice_segment_address), etc. In this case, the slice position information may be position information of a predetermined block (eg, the first row in a scanning order in a slice). In this case, the position information may be block scanning order information.

[0155] In section 4d, various partition settings are allowed for arbitrary partitioning.

[0156] In section 4d, a division unit may be defined as a group of blocks that are spatially adjacent to each other, and information about the division may include information about the size, form, and position of the division unit. This is only an example of arbitrary division, and as Figure 5 Various partitioning forms are possible as shown.

[0157] Figure 5 is another exemplary diagram of the image segmentation method according to an embodiment of the present invention.

[0158] In parts 5a and 5b, the image may be divided into a plurality of regions horizontally or vertically at at least one block interval, and the division may be performed based on block position information. Part 5a shows an example (A0, A1) of performing division horizontally based on row information of each block, and part 5b shows an example (B0 to B3) of performing division horizontally and vertically based on column information and row information of each block. Information about the division may include the number of division units, block interval information, division direction, etc., and when the division information is implicitly included according to a predetermined rule, some division information may not be generated.

[0159] In parts 5c and 5d, the image may be divided into groups of continuous blocks in a scanning order. Additional scanning orders other than the conventional slice raster scanning order may be applied to the image division. Part 5c shows an example (C0, C1) of performing a scan clockwise or counterclockwise relative to a starting block (Box-Out), and part 5d shows an example (D0, D1) of performing a scan vertically relative to a starting block (vertical). Information about the division may include information about the number of division units, information about the position of the division unit (e.g., the first row in the division unit in the scanning order), information about the scanning order, etc., and when the division information is implicitly included according to a predetermined rule, some division information may not be generated.

[0160] In part 5e, the image can be divided using horizontal and vertical dividing lines. Existing tiles can be divided by horizontal or vertical dividing lines. Therefore, the division can be performed in the form of a quadrilateral space, but the image may not be divided using dividing lines. For example, it is possible to divide the image by some dividing lines of the image (e.g., the dividing line between the left boundary of E5 and the right boundary of E1, E3 and E4), and it is impossible to divide the image by some dividing lines of the image (e.g., the dividing line between the lower boundary of E2 and E3 and the upper boundary of E4). In addition, the division can be performed based on block units (e.g., after first performing block division), or the division can be performed by horizontal or vertical dividing lines (e.g., the division is performed by dividing lines regardless of block division). Therefore, each division unit may not be a multiple of a block. Therefore, division information different from the division information of the existing tiles can be generated, and the division information may include information about the number of division units, information about the position of the division unit, information about the size of the division unit, etc. For example, information about the position of the division unit can be generated as position information (e.g., measured in pixel units or block units) based on a predetermined position (e.g., in the upper left corner of the image), and information about the size of the division unit can be generated as information about the width and height of each division unit (e.g., measured in pixel units or block units).

[0161] Similar to the above example, the division according to any user-defined setting can be performed by applying a new division method or by changing some elements of the existing division. That is, the division method can be supported by replacing the traditional division method or adding to the traditional division method, and the division method can be supported by changing some settings (slices, tiles, etc.) of the traditional division method (for example, according to another scanning order, by using another division method of a quadrilateral shape to generate other division information, or according to dependent encoding / decoding characteristics). In addition, settings for configuring additional division units can be supported (for example, in addition to division according to the scanning order or division according to a specific interval difference), and additional division unit forms can be supported (for example, polygonal forms such as triangles other than division into quadrilateral spaces). In addition, image division methods can be supported based on the type, characteristics, etc. of the image. For example, partial division methods (for example, faces of 360-degree images) can be supported based on the type, characteristics, etc. of the image. Information about division can be generated based on support.

[0162] In the image division execution step, the image can be divided based on the identified division type information. That is, the image can be divided into a plurality of division units based on the identified division type, and the image can be encoded or decoded based on the acquired division units.

[0163] In this case, whether or not there is an encoding / decoding setting in each division unit can be determined according to the division type. That is, the setting information required during the encoding / decoding process for each division unit can be specified by the upper unit (e.g., picture), or independent encoding / decoding settings can be provided for each division unit.

[0164] Generally, a slice may have an independent encoding / decoding setting (e.g., a slice header) for each partition unit, and a tile may not have an independent encoding / decoding setting for each partition unit and may have a setting that depends on the encoding / decoding setting (e.g., PPS) of a picture. In this case, the information generated in association with the tile may be the partition information and may be included in the encoding / decoding setting of the picture. The present invention is not limited to the above examples, and the above examples may be modified.

[0165] Coding / decoding setting information for tiles can be generated in units of video, sequence, picture, etc. At least one piece of coding / decoding setting information is generated in the upper unit, and the generated piece of coding / decoding setting information can be referenced. Alternatively, independent coding / decoding setting information (for example, tile header) can be generated in units of tiles. This is different from the case of following one coding / decoding setting determined in the upper unit in that coding / decoding is performed while providing at least one coding / decoding setting in units of tiles. That is, all tiles can be encoded or decoded according to the same coding / decoding setting, or at least one tile can be encoded or decoded according to a coding / decoding setting that is different from the coding / decoding setting of other tiles.

[0166] The above examples focus on various encoding / decoding settings in tiles. However, the present invention is not limited thereto, and even the same or similar settings may be applied to other partition types.

[0167] As an example, in some division types, division information may be generated in a superordinate unit, and encoding or decoding may be performed according to a single encoding / decoding setting of the superordinate unit.

[0168] As an example, in some division types, division information may be generated in superordinate units, and independent encoding / decoding settings for each division unit in the superordinate unit may be generated, and encoding or decoding may be performed according to the generated encoding / decoding settings.

[0169] As an example, in some division types, division information may be generated in a superordinate unit, and a plurality of pieces of encoding / decoding setting information may be supported in the superordinate unit. Encoding or decoding may be performed according to the encoding / decoding setting referenced by each division unit.

[0170] As an example, in some division types, division information may be generated in a superordinate unit, and independent encoding / decoding settings may be generated in the corresponding division units. Encoding or decoding may be performed according to the generated encoding / decoding settings.

[0171] As an example, in some division types, independent encoding / decoding settings including division information may be generated in corresponding division units, and encoding or decoding may be performed according to the generated encoding / decoding settings.

[0172] The encoding / decoding setting information may include information required for encoding or decoding a tile, such as a tile type, information about a reference picture list, quantization parameter information, inter-frame prediction setting information, in-loop filter setting information, in-loop filter control information, scanning order, whether to perform encoding or decoding, etc. The encoding / decoding setting information may be used to explicitly generate relevant information, or may have encoding / decoding settings that are implicitly determined according to the format, characteristics, etc. of an image determined in an upper unit. In addition, relevant information may be explicitly generated based on information acquired through the setting.

[0173] Next, an example of performing image division in the encoding / decoding device according to the embodiment of the present invention will be described.

[0174] A division process may be performed on an input image before starting encoding. An image may be divided using division information (e.g., image division information, division unit setting information, etc.), and then the image may be encoded in the division units. Image encoding data may be stored in a memory after encoding is completed, and may be added to a bitstream and then transmitted.

[0175] The division process may be performed before starting decoding. The image may be divided using division information (e.g., image division information, division unit setting information, etc.), and then the image decoded data may be parsed and decoded in the division units. After decoding is completed, the image decoded data may be stored in a memory, and a plurality of division units may be merged into a single unit, and thus the image may be output.

[0176] Through the above examples, the image division process has been described. In addition, according to the present invention, a plurality of division processes can be performed.

[0177] For example, the image may be divided, and the divided units of the image may be divided. The division may be the same division process (e.g., slice / slice, tile / tile, etc.) or different division processes (e.g., slice / tile, tile / slice, tile / surface, surface / tile, slice / surface, surface / slice, etc.). In this case, the subsequent division process may be performed based on the previous division result, and the division information generated during the subsequent division process may be generated based on the previous division result.

[0178] In addition, a plurality of division processes (A) may be performed, and the division processes may be different division processes (e.g., slice / plane, tile / plane, etc.) In this case, the subsequent division process may be performed based on the previous division result or independently of the previous division result, and the division information generated during the subsequent division process may be generated based on the previous division result or independently of the previous division result.

[0179] A plurality of image division processes may be determined according to encoding / decoding settings. However, the present invention is not limited to the above-described examples, and various modifications may be made to the above-described examples.

[0180] The encoder can add the information generated during the above processing to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the relevant information from the bitstream. That is, information can be added to one unit, and the information can be copied and added to multiple units. For example, a syntax element indicating whether some information is supported or a syntax element indicating whether activation is performed can be generated in some units (e.g., a superordinate unit), and the same or similar information can be generated in some units (e.g., a subordinate unit). That is, even in the case where the relevant information is supported and set in the superordinate unit, the subordinate unit can have a separate setting. This description is not limited to the above examples and can be applied to the present invention in common. In addition, the information can be included in the bitstream in the form of SEI or metadata.

[0181] Generally, the input image may be encoded or decoded as it is, but the encoding or decoding may be performed after adjusting the image size (expansion or reduction; resolution adjustment). For example, in a layered coding scheme (adaptive video coding) for supporting spatial, temporal, and image quality adaptability, image size adjustment such as overall expansion and reduction of the image may be performed. Alternatively, image size adjustment such as partial expansion and reduction of the image may be performed. Image size adjustment may be performed differently, that is, image size adjustment may be performed for the purpose of adaptability to the encoding environment, for the purpose of encoding uniformity, for the purpose of encoding efficiency, for the purpose of image quality improvement, or according to the type, characteristics, etc. of the image.

[0182] As a first example, the resizing process may be performed during a process performed according to the characteristics, type, etc. of an image (eg, layered encoding, 360-degree image encoding, etc.).

[0183] As a second example, the resizing process may be performed at the initial encoding / decoding step. The resizing process may be performed before encoding or decoding is performed. The resized image may be encoded or decoded.

[0184] As a third example, resizing processing may be performed during a prediction step (intra-frame prediction or inter-frame prediction) or before prediction. During the resizing processing, image information (e.g., information about pixels referenced by intra-frame prediction, information about intra-frame prediction modes, information about reference pictures used for inter-frame prediction, information about inter-frame prediction modes, etc.) may be used at the prediction step.

[0185] As a fourth example, resizing processing may be performed during the filtering step or before filtering. In the resizing processing, image information in the filtering step (e.g., pixel information to be applied to a deblocking filter, pixel information to be applied to SAO, information about SAO filtering, pixel information applied to ALF, information about ALF filtering, etc.) may be used.

[0186] In addition, after the resizing process is performed, the image may be processed by the inverse resizing process and the image may be changed to the image before the resizing (in terms of the image size), or the image may not be changed. This may be determined according to the encoding / decoding settings (for example, the characteristics of performing resizing). In this case, the resizing process may be an expansion process and the inverse resizing process is a reduction process, and the resizing process may be a reduction process and the inverse resizing process is an expansion process.

[0187] When the resizing process is performed according to the first to fourth examples, the inverse resizing process is performed in the subsequent step so that the image before the resizing can be acquired.

[0188] When the resizing process is performed by layered encoding or according to the third example (or when the reference picture is resized in inter prediction), the inverse resizing process may not be performed in the subsequent step.

[0189] In an embodiment of the present invention, the image resizing process may be performed alone or together with the inverse process. The following example description will focus on the resizing process. In this case, since the inverse resizing process is the inverse process for the resizing process, the description of the inverse resizing process will be omitted to prevent redundant description. However, it is obvious that those skilled in the art can recognize the same content as the content described literally.

[0190] Figure 6 is an example diagram of a general image resizing method.

[0191] Referring to part 6 a , an extended image ( P0 + P1 ) may be acquired by adding a specific area ( P1 ) to an initial image P0 (or an image before resizing; which is indicated by a thick solid line).

[0192] Referring to section 6 b , the reduced image ( S0 ) may be acquired by removing a specific area ( S1 ) from an initial image ( S0 + S1 ).

[0193] Referring to part 6 c , the resized image ( T0 + T1 ) may be acquired by adding the specific region ( T1 ) to the initial image ( T0 + T2 ) and removing the specific region ( T2 ) from the entire image.

[0194] According to the present invention, the following description focuses on the resizing process for expansion and the resizing process for reduction. However, the present invention is not limited thereto, and should be understood to include a case where expansion and reduction are applied in combination as shown in section 6c.

[0195] Figure 7 is an example diagram of image resizing according to an embodiment of the present invention.

[0196] During the resizing process, an image expansion method will be described with reference to Section 7a, and an image reduction method will be described with reference to Section 7b.

[0197] In part 7a, the image before resizing is S0, and the image after resizing is S1. In part 7b, the image before resizing is T0, and the image after resizing is T1.

[0198] When the image is expanded as shown in part 7a, the image can be expanded in the "upward" direction, the "downward" direction, the "leftward" direction, or the "rightward" direction (ET, EL, EB, ER). When the image is reduced as shown in part 7b, the image can be reduced in the "upward" direction, the "downward" direction, the "leftward" direction, or the "rightward" direction (RT, RL, RB, RR).

[0199] Comparing image expansion and image reduction, the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction of expansion can correspond to the "downward" direction, "upward" direction, "rightward" direction, and "leftward" direction of reduction. Therefore, the following description focuses on image expansion, but it should be understood that the description of image reduction is included.

[0200] In the following description, image expansion or reduction is performed in the "upward" direction, the "downward" direction, the "leftward" direction, and the "rightward" direction. However, it should also be understood that resizing can be performed in the "upward and leftward" direction, the "upward and rightward" direction, the "downward and leftward" direction, or the "downward and rightward" direction.

[0201] In this case, when expansion is performed in the "downward and rightward" direction, regions RC and BC are acquired, and region BR may or may not be acquired according to encoding / decoding settings. That is, regions TL, TR, BL, and BR may or may not be acquired, but for ease of description, it will be described that the corner regions (i.e., regions TL, TR, BL, and BR) can be acquired.

[0202] The image resizing process according to the embodiment of the present invention may be performed in at least one direction. For example, the image resizing process may be performed in all directions such as upward, downward, left, and right, may be performed in two or more directions selected from upward, downward, left, and right (left+right, upward+downward, upward+left, upward+right, downward+left, downward+right, upward+left+right, downward+left+right, upward+downward+left, upward+downward+right, etc.), or may be performed only in one direction selected from upward, downward, left, and right.

[0203] For example, resizing may be performed in the "left + right" direction, the "up + down" direction, the "left and up + right and down" direction, and the "left and down + right and up" direction, which are symmetrically extendable to both ends relative to the center of the image, resizing may be performed in the "left + right" direction, the "left and up + right and up" direction, and the "left and down + right and down" direction, which are vertically symmetrically extendable relative to the image, and resizing may be performed in the "up + down" direction, the "left and up + left and down" direction, and the "right and up + right and down" direction, which are horizontally symmetrically extendable relative to the image. Other resizing may be performed.

[0204] In parts 7a and 7b, the size of the image (S0, T0) before resizing is defined as P_Width×P_Height, and the size of the image (S1, T1) after resizing is defined as P'_Width×P'_Height. Here, when the resizing values ​​in the "left", "right", "upward" and "downward" directions are defined as Var_L, Var_R, Var_T and Var_B (or collectively referred to as Var_x), the size of the image after resizing can be expressed as (P_Width+Var_L+Var_R)×(P_Height+Var_T+Var_B). In this case, Var_L, Var_R, Var_T, and Var_B, which are size adjustment values ​​in the "left", "right", "up", and "down" directions, may be Exp_L, Exp_R, Exp_T, and Exp_B (here, Exp_x is positive) for image expansion (in part 7a), and may be -Rec_L, -Rec_R, -Rec_T, and -Rec_B (which are expressed as negative values ​​for image reduction when Rec_L, Rec_R, Rec_T, and Rec_B are defined as positive values) for image reduction. In addition, the upper left corner coordinates, the upper right corner coordinates, the lower left corner coordinates, and the lower right corner coordinates of the image before resizing can be (0, 0), (P_Width-1, 0), (0, P_Height-1), and (P_Width-1, P_Height-1), and the upper left corner coordinates, the upper right corner coordinates, the lower left corner coordinates, and the lower right corner coordinates of the image after resizing can be expressed as (0, 0), (P'_Width-1, 0), (0, P'_Height-1), and (P'_Width-1, P'_Height-1). The size of the area changed (or acquired or removed) by resizing (here, TL to BR; i is an index for identifying TL to BR) can be M[i]×N[i] and can be expressed as Var_X×Var_Y (this example assumes that X is L or R, and Y is T or B). M and N can have various values ​​and can have the same settings regardless of i, or can have separate settings according to i. Various examples will be described below.

[0205] Referring to section 7a, S1 may be configured to include some or all of the regions TL to BR (upper left to lower right) generated by expanding S0 in several directions. Referring to section 7b, T1 may be configured to exclude all or some of the regions TL to BR removed by shrinking in several directions from T0.

[0206] In part 7a, when the existing image (S0) is expanded in the "upward", "downward", "leftward" and "rightward" directions, the image may include regions TC, BC, LC and RC obtained by the resizing process and may further include regions TL, TR, BL and BR.

[0207] As an example, when extension is performed in the "upward" direction (ET), an image may be constructed by adding regions TC to an existing image (S0), and the image may include regions TL or TR and extension in at least one different direction (EL or ER).

[0208] As an example, when extension is performed in the "downward" direction (EB), an image may be constructed by adding region BC to an existing image (S0), and the image may include regions BL or BR and extension in at least one different direction (EL or ER).

[0209] As an example, when extension is performed in the "left" direction (EL), an image may be constructed by adding region LC to an existing image (S0), and the image may include regions TL or BL and extension in at least one different direction (ET or EB).

[0210] As an example, when extension is performed in the "right" direction (ER), an image may be constructed by adding region RC to an existing image (S0), and the image may include region TR or BR and extension in at least one different direction (ET or EB).

[0211] According to an embodiment of the present invention, a setting (eg, spa_ref_enabled_flag or tem_ref_enabled_flag) for spatially or temporally limiting the referability of a resized region (this example assumes extension) may be provided.

[0212] That is, reference to data of a region spatially or temporally resized according to encoding / decoding settings may be allowed (eg, spa_ref_enabled_flag=1 or tem_ref_enabled_flag=1) or restricted (eg, spa_ref_enabled_flag=0 or tem_ref_enabled_flag=0).

[0213] Encoding / decoding of images before resizing (S0, T1) and regions added or deleted during resizing (TC, BC, LC, RC, TL, TR, BL, and BR) may be performed as follows.

[0214] For example, when encoding or decoding an image before resizing and an added or deleted area, data about the image before resizing and data about the added or deleted area (data after encoding or decoding is completed; pixel values ​​or prediction-related information) can reference each other spatially or temporally.

[0215] Alternatively, the image before resizing and the data about the added or deleted area may be spatially referenced, while the data about the image before resizing may be temporally referenced, and the data about the added or deleted area may not be temporally referenced.

[0216] That is, a setting for limiting the referability of an added or deleted area may be provided. The setting information on the referability of an added or deleted area may be explicitly generated or implicitly determined.

[0217] The image resizing process according to an embodiment of the present invention may include an image resizing indication step, an image resizing type identification step and / or an image resizing execution step. In addition, the image encoding device and the image decoding device may include an image resizing indication unit, an image resizing type identification unit and an image resizing execution unit respectively configured to execute the image resizing indication step, the image resizing type identification step and the image resizing execution step. For encoding, relevant syntax elements may be generated. For decoding, relevant syntax elements may be parsed.

[0218] In the image resizing indication step, it may be determined whether to perform image resizing. For example, when a signal indicating image resizing (e.g., img_resizing_enabled_flag) is confirmed, resizing may be performed. When a signal indicating image resizing is not confirmed, resizing may not be performed, or resizing may be performed by confirming other encoding / decoding information. In addition, although a signal indicating image resizing is not provided, a signal indicating image resizing may be implicitly activated or deactivated according to encoding / decoding settings (e.g., characteristics, type, etc. of an image). When resizing is performed, corresponding resizing-related information may be generated, or corresponding resizing-related information may be implicitly determined.

[0219] When a signal indicating image size adjustment is provided, the corresponding signal is a signal for indicating whether to perform image size adjustment. Whether to adjust the size of the corresponding image can be determined based on the signal.

[0220] For example, it is assumed that a signal indicating image resizing (eg, img_resizing_enabled_flag) is confirmed. When the corresponding signal is activated (eg, img_resizing_enabled_flag=1), image resizing may be performed. When the corresponding signal is deactivated (eg, img_resizing_enabled_flag=0), image resizing may not be performed.

[0221] Furthermore, when a signal indicating image resizing is not provided, resizing may not be performed, or whether to resize the corresponding image may be determined by another signal.

[0222] For example, when the input image is divided in units of blocks, resizing may be performed depending on whether the size of the image (e.g., width or height) is a multiple of the size of the block (e.g., width or height) (for the extension in this example, it is assumed that resizing processing is performed when the image size is not a multiple of the block size). That is, resizing may be performed when the width of the image is not a multiple of the width of the block or when the height of the image is not a multiple of the height of the block. In this case, resizing information (e.g., resizing direction, resizing value, etc.) may be determined based on encoding / decoding information (e.g., the size of the image, the size of the block, etc.). Alternatively, resizing may be performed based on the characteristics, type (e.g., 360-degree image), etc. of the image, and the resizing information may be explicitly generated or may be specified as a predetermined value. The present invention is not limited to the above examples, and the above examples may be modified.

[0223] In the image resizing type identification step, the image resizing type may be identified. The image resizing type may be defined by a resizing method, resizing information, and the like. For example, resizing based on a scale factor, resizing based on an offset factor, and the like may be performed. The present invention is not limited to the above examples, and these methods may be applied in combination. For ease of description, the following description will focus on resizing based on a scale factor and resizing based on an offset factor.

[0224] For the scale factor, resizing may be performed by multiplication or division based on the size of the image. Information about a resizing operation (e.g., expansion or reduction) may be explicitly generated, and an expansion or reduction process may be performed according to the corresponding information. In addition, the resizing process may be performed as a predetermined operation (e.g., one of an expansion operation and a reduction operation) according to the encoding / decoding setting. In this case, the information about the resizing operation is omitted. For example, when image resizing is activated in the image resizing instruction step, image resizing may be performed as a predetermined operation.

[0225] The resizing direction may be at least one direction selected from the following: upward, downward, leftward, and rightward. Depending on the resizing direction, at least one scaling factor may be required. That is, one scaling factor may be required for each direction (here unidirectional), one scaling factor may be required for the horizontal or vertical direction (here bidirectional), and one scaling factor may be required for all directions of the image (here omnidirectional). In addition, the resizing direction is not limited to the above examples, and the above examples may be modified.

[0226] The scale factor may have a positive value and may have range information that is different according to encoding / decoding settings. For example, when information is generated by combining a resizing operation and a scale factor, the scale factor may be used as a multiplicand. A scale factor greater than 0 or less than 1 may mean a reduction operation, a scale factor greater than 1 may mean an expansion operation, and a scale factor of 1 may mean that no resizing is performed. As another example, when the scale factor information is generated independently of the resizing operation, a scale factor for an expansion operation may be used as a multiplicand, and a scale factor for a reduction operation may be used as a dividend.

[0227] Will refer again Figure 7 Sections 7a and 7b of FIG. 7 describe a process of changing an image before resizing ( S0 , T0 ) into an image after resizing ( S1 and T1 here).

[0228] As an example, when one scale factor (called sc) is used in all directions of an image and the resizing direction is the "downward + right" direction, the resizing directions are ER and EB (or RR and RB), the resizing values ​​Var_L (Exp_L or Rec_L) and Var_T (Exp_T or Rec_T) are 0, and Var_R (Exp_R or Rec_R) and Var_B (Exp_B or Rec_B) can be expressed as P_Width×(sc-1) and P_Height×(sc-1). Therefore, the image after resizing can be (P_Width×sc)×(P_Height×sc).

[0229] As an example, when the corresponding scale factors (here sc_w and sc_h) are used in the horizontal direction or the vertical direction of the image and the resizing direction is the "left+right" direction and the "up+down" direction (up+down+left+right when both are operated), the resizing directions may be ET, EB, EL, and ER, the resizing values ​​Var_T and Var_B may be P_Height×(sc_h-1) / 2, and Var_L and Var_R may be P_Width×(sc_w-1) / 2. Therefore, the image after resizing may be (P_Width×sc_w)×(P_Height×sc_h).

[0230] For the offset factor, resizing may be performed by addition or subtraction based on the size of the image. Alternatively, resizing may be performed by addition or subtraction based on the encoding / decoding information of the image. Alternatively, resizing may be performed by independent addition or subtraction. That is, the resizing process may have a dependent setting or an independent setting.

[0231] Information about a resizing operation (e.g., expansion or reduction) may be explicitly generated, and expansion or reduction processing may be performed according to the corresponding information. In addition, the resizing operation may be performed as a predetermined operation (e.g., one of an expansion operation and a reduction operation) according to encoding / decoding settings. In this case, the information about the resizing operation may be omitted. For example, when image resizing is activated in the image resizing instruction step, image resizing may be performed as a predetermined operation.

[0232] The resizing direction may be at least one direction selected from the following: upward, downward, leftward, and rightward. Depending on the resizing direction, at least one offset factor may be required. That is, one offset factor may be required for each direction (here unidirectional), one offset factor may be required for the horizontal or vertical direction (here symmetrical bidirectional), one offset factor may be required according to a partial combination of directions (here asymmetrical bidirectional), and one offset factor may be required for all directions of the image (here omnidirectional). In addition, the resizing direction is not limited to the above examples, and the above examples may be modified.

[0233] The offset factor may have a positive value or both a positive value and a negative value, and may have range information that is different according to the encoding / decoding setting. For example, when generating information in conjunction with a resizing operation and an offset factor (here, it is assumed that the offset factor has both a positive value and a negative value), the offset factor may be used as a value to be added or subtracted according to the sign information of the offset factor. An offset factor greater than 0 may mean an expansion operation, an offset factor less than 0 may mean a reduction operation, and an offset factor of 0 may mean that no resizing is performed. As another example, when offset factor information is generated independently of a resizing operation (here, it is assumed that the offset factor has a positive value), the offset factor may be used as a value to be added or subtracted according to the resizing operation. An offset factor greater than 0 may mean that an expansion or reduction operation may be performed according to the resizing operation, and an offset factor of 0 may mean that no resizing is performed.

[0234] Will refer again Figure 7 Sections 7a and 7b of EMBODIMENT 3 describe a method of changing an image before resizing ( S0 , T0 ) into an image after resizing ( S1 , T1 ) using an offset factor.

[0235] As an example, when one offset factor (referred to as os) is used in all directions of an image and the resizing direction is the "up+down+left+right" direction, the resizing directions may be ET, EB, EL, and ER (or RT, RB, RL, and RR) and the resizing values ​​Var_T, Var_B, Var_L, and Var_R may be os. The size of the image after resizing may be (P_Width+os)×(P_Height+os).

[0236] As an example, when the offset factors (os_w, os_h) are used in the horizontal or vertical direction of the image and the resizing direction is the "left+right" direction and the "up+down" direction ("up+down+left+right" direction when both are operated), the resizing direction may be ET, EB, EL, and ER (or RT, RB, RL, and RR), the resizing values ​​Var_T and Var_B may be os_h, and the resizing values ​​Var_L and Var_R may be os_w. The size of the image after resizing may be {P_Width+(os_w×2)}×{P_Height+(os_h×2)}.

[0237] As an example, when the resizing directions are the "downward" direction and the "rightward" direction (the "downward+rightward" direction when operated together) and the offset factors (os_b, os_r) are used according to the resizing directions, the resizing directions may be EB and ER (or RB and RR), the resizing value Var_B may be os_b, and the resizing value Var_R may be os_r. The size of the image after resizing may be (P_Width + os_r)×(P_Height+os_b).

[0238] As an example, when the offset factors (os_t, os_b, os_l, os_r) are used according to the direction of the image and the resizing directions are the "up", "down", "left" and "right" directions (the "up+down+left+right" directions for all operations), the resizing directions may be ET, EB, EL and ER (or RT, RB, RL and RR), the resizing value Var_T may be os_t, the resizing value Var_B may be os_b, the resizing value Var_L may be os_l, and the resizing value Var_R may be os_r. The size of the image after resizing may be (P_Width+os_l+os_r)×(P_Height+os_t+os_b).

[0239] The above example indicates a case where an offset factor is used as a resizing value (Var_T, Var_B, Var_L, Var_R) during resizing processing. That is, this means that the offset factor is used as a resizing value without any change, which can be an example of resizing performed independently. Alternatively, the offset factor can be used as an input variable of the resizing value. In detail, the offset factor can be specified as an input variable, and the resizing value can be obtained through a series of processes according to the encoding / decoding setting, which can be an example of resizing performed based on predetermined information (e.g., image size, encoding / decoding information, etc.) or an example of resizing performed dependently.

[0240] For example, the offset factor may be a multiple (e.g., 1, 2, 4, 6, 8, and 16) or an exponential multiple (e.g., an exponential multiple of 2, such as 1, 2, 4, 8, 16, 32, 64, 128, and 256) of a predetermined value (here, an integer). Alternatively, the offset factor may be a multiple or exponential multiple of a value acquired based on encoding / decoding settings (e.g., a value set based on a motion search range for inter-frame prediction). Alternatively, the offset factor may be a multiple or an integer of a unit (here, assuming A×B) acquired from a picture partitioning unit. Alternatively, the offset factor may be a multiple of a unit (here, assuming E×F, such as a tile) acquired from a picture partitioning unit.

[0241] Alternatively, the offset factor may be a value less than or equal to the width and height of the unit obtained from the picture division unit. In the above example, the multiple or exponential multiple may have a value of 1. However, the present invention is not limited to the above example and may be modified. For example, when the offset factor is n, Var_x may be 2×n or 2 n .

[0242] In addition, separate offset factors may be supported according to color components. Offset factors for some color components may be supported, and thus offset factor information for other color components may be derived. For example, when an offset factor (A) for a luminance component is explicitly generated (here, assuming that the composition ratio of the luminance component to the chrominance component is 2:1), an offset factor (A / 2) for a chrominance component may be implicitly acquired. Alternatively, when an offset factor (A) for a chrominance component is explicitly generated, an offset factor (2A) for a luminance component may be implicitly acquired.

[0243] Information about the resizing direction and the resizing value may be explicitly generated, and the resizing process may be performed according to the corresponding information. In addition, the information may be implicitly determined according to the encoding / decoding settings, and the resizing process may be performed according to the determined information. At least one predetermined direction or resizing value may be assigned, and in this case, the relevant information may be omitted. In this case, the encoding / decoding settings may be determined based on the characteristics, type, encoding information, etc. of the image. For example, at least one resizing direction may be predetermined according to at least one resizing operation, at least one resizing value may be predetermined according to at least one resizing operation, and at least one resizing value may be predetermined according to at least one resizing direction. In addition, the resizing direction, resizing value, etc. during the inverse resizing process may be derived from the resizing direction, resizing value, etc. applied during the resizing process. In this case, the implicitly determined resizing value may be one of the above examples (examples of obtaining resizing values ​​differently).

[0244] In addition, multiplication or division has been described in the above examples, but a shift operation can be used according to the implementation of the encoder / decoder. Multiplication can be achieved by a left shift operation, and division can be achieved by a right shift operation. This description is not limited to the above examples and can be applied to the present invention in common.

[0245] In the image resizing execution step, image resizing may be performed based on the identified resizing information, that is, image resizing may be performed based on information about resizing type, resizing operation, resizing direction, resizing value, etc., and encoding / decoding may be performed based on the acquired resized image.

[0246] In addition, in the image resizing execution step, resizing may be performed using at least one data processing method. In detail, resizing may be performed on the area to be resized according to the resizing type and the resizing operation by using at least one data processing method. For example, according to the resizing type, how to fill data may be determined when resizing is for expansion, and how to remove data may be determined when resizing is for reduction.

[0247] In summary, in the image resizing execution step, image resizing may be performed based on the identified resizing information. Alternatively, in the image resizing execution step, image resizing may be performed based on the resizing information and the data processing method. The above two cases may differ from each other in that only the size of the image to be encoded or decoded is adjusted, or in that even the data processing of the image size and the area to be resized is considered. In the image resizing execution step, whether to execute the data processing method may be determined based on the step, position, etc. of applying the resizing processing. The following description focuses on an example of performing resizing based on a data processing method, but the present invention is not limited thereto.

[0248] When resizing based on the offset factor is performed, various methods may be used to perform resizing for expansion and resizing for reduction. For expansion, resizing may be performed using at least one data filling method. For reduction, resizing may be performed using at least one data removal method. In this case, when resizing based on the offset factor is performed, the resized area may be filled with new data or original image data directly or after modification (expansion), and the resized area may be removed simply or through a series of processes (reduction).

[0249] When resizing based on a scale factor is performed, in some cases (e.g., layered coding), resizing for expansion can be performed by applying upsampling, and resizing for reduction can be performed by applying downsampling. For example, at least one upsampling filter can be used for expansion, and at least one downsampling filter can be used for reduction. The filter applied horizontally can be the same as or different from the filter applied vertically. In this case, when resizing based on a scale factor is performed, new data is neither generated in the resized area nor removed from the resized area, but methods such as interpolation can be used to rearrange the original image data. Data processing methods associated with resizing can be classified according to the filter used for sampling. In addition, in some cases (e.g., similar to the case of an offset factor), resizing for expansion can be performed using a method of filling at least one data, and resizing for reduction can be performed using a method of removing at least one data. According to the present invention, the following description focuses on a data processing method corresponding to the case of performing resizing based on an offset factor.

[0250] Typically, a predetermined data processing method may be used in the area to be resized, but at least one data processing method may be used in the area to be resized, as in the following example. Selection information for the data processing method may be generated. The former may mean that resizing is performed by a fixed data processing method, while the latter may mean that resizing is performed by an adaptive data processing method.

[0251] Furthermore, the data processing method may be applied to all regions (TL, TC, TR, ..., BR in parts 7a and 7b) or some regions (e.g., each or a combination of TL to BR in parts 7a and 7b) among the regions to be added or deleted during resizing.

[0252] Figure 8 is an exemplary diagram of a method of constructing a region generated by expansion in an image resizing method according to an embodiment of the present invention.

[0253] Referring to part 8a, for ease of description, the image may be divided into regions TL, TC, TR, LC, C, RC, BL, BC, and BR corresponding to the upper left position, upper position, upper right position, left position, center position, right position, lower left position, lower position, and lower right position of the image. In the following description, the image is extended in the "downward + rightward" direction, but it should be understood that the description may be applied to other extension directions.

[0254] The region added according to the extension of the image can be constructed using various methods. For example, the region can be filled with an arbitrary value, or it can be filled with reference to some data of the image.

[0255] Referring to section 8b, the generated areas (A0, A2) may be filled with arbitrary pixel values. Various methods may be used to determine the arbitrary pixel values.

[0256] As an example, an arbitrary pixel value may be a pixel in a range of pixel values ​​that can be represented using a bit depth (e.g., from 0 to 1<<(bit_depth)-1). For example, an arbitrary pixel value may be a minimum value, a maximum value, a median value (e.g., 1<<(bit_depth-1) etc.) in the range of pixel values, etc. (here, bit_depth indicates the bit depth).

[0257] As an example, the arbitrary pixel value may be a pixel in a range of pixel values ​​of pixels belonging to the image (e.g., from min to P to max P ; min P and max P Indicates the minimum and maximum values ​​among the pixels belonging to the image; min P Greater than or equal to 0; max P less than or equal to 1 << (bit_depth) - 1). For example, any pixel value may be the minimum value, maximum value, median value, average value (of at least two pixels), weighted sum, etc. in the pixel value range.

[0258] As an example, an arbitrary pixel value may be a value determined in a pixel value range belonging to a specific area included in the image. For example, when constructing A0, the specific area may be TR+RC+BR. In addition, the specific area may be set to an area corresponding to 3×9 of TR, RC, and BR or an area corresponding to 1×9 (which is assumed to be the rightmost line). This may depend on the encoding / decoding settings. In this case, the specific area may be a unit to be divided by the picture division unit. In detail, an arbitrary pixel value may be a minimum value, a maximum value, a median value, an average value (of at least two pixels), a weighted sum, etc. in a pixel value range.

[0259] Referring again to section 8b, the area A1 to be added as the image is expanded may be filled with pattern information generated using a plurality of pixel values ​​(e.g., assuming that the pattern uses a plurality of pixels; no certain rules need to be followed). In this case, the pattern information may be defined according to encoding / decoding settings, or relevant information may be generated. The generated area may be filled with at least one piece of pattern information.

[0260] Referring to part 8c, the area added as the image is expanded can be constructed with reference to the pixels of a specific area included in the image. In detail, the added area can be constructed by copying or filling the pixels in the area adjacent to the added area (hereinafter referred to as reference pixels). In this case, the pixels in the area adjacent to the added area can be pixels before encoding or pixels after encoding (or decoding). For example, when resizing is performed in a pre-encoding step, the reference pixel may refer to a pixel of the input image, and when resizing is performed in an intra-frame prediction reference pixel generation step, a reference picture generation step, a filtering step, etc., the reference pixel may refer to a pixel of the restored image. In this example, it is assumed that the nearest pixel is used in the added area, but the present invention is not limited to this.

[0261] The region (A0) generated when the image is extended leftward or rightward in association with horizontal image resizing may be constructed by horizontally padding (Z0) external pixels adjacent to the generated region (A0), and the region (A1) generated when the image is extended upward or downward in association with vertical image resizing may be constructed by vertically padding (Z1) external pixels adjacent to the generated region (A1). Furthermore, the region (A2) generated when the image is extended downward and rightward may be constructed by diagonally padding (Z2) external pixels adjacent to the generated region (A2).

[0262] Referring to section 8d, the generated region (B'0 to B'2) may be constructed with reference to data of a specific region (B0 to B2) included in the image. In section 8d, unlike section 8c, a region not adjacent to the generated region may be referenced.

[0263] For example, when there is an area with high correlation with the generated area in the image, the generated area can be filled with reference to the pixels of the area with high correlation. In this case, the position information, size information, etc. of the area with high correlation can be generated. Alternatively, when there is an area with high correlation through encoding / decoding information of the characteristics, type, etc. of the image and the position information, size information, etc. of the area with high correlation can be implicitly checked (for example, for a 360-degree image), the generated area can be filled with the data of the corresponding area. In this case, the position information, size information, etc. of the corresponding area can be omitted.

[0264] As an example, pixels in an area (B2) opposite to the area generated when the image is expanded left or right in association with horizontal image resizing may be referenced to fill an area (B'2) generated when the image is expanded left or right in association with horizontal image resizing.

[0265] As an example, pixels in an area (B1) may be referenced to fill an area (B′1) generated when an image is expanded upward or downward in association with vertical image resizing, the area (B1) being opposite to the area generated when the image is expanded upward or downward in association with vertical image resizing.

[0266] As an example, pixels in region (B0, TL) may be referenced to fill region (B'0) generated when the image is extended by some image resizing (here, diagonally relative to the image center), to which region (B0, TL) is opposite.

[0267] An example has been described in which there is continuity at the boundary between both ends of the image and data of a region symmetrical with respect to the resizing direction is acquired. However, the present invention is not limited thereto, and data of other regions (TL to BR) may be acquired.

[0268] When the generated area is filled with data of a specific area of ​​an image, the data of the corresponding area can be copied as is and the data is used to fill the generated area, or the data of the corresponding area can be transformed based on the characteristics, type, etc. of the image and used to fill the generated area. In this case, copying the data as is can mean using the pixel values ​​of the corresponding area without any changes, and performing the transformation process can mean not using the pixel values ​​of the corresponding area without any changes. That is, at least one pixel value of the corresponding area can be changed by the transformation process. The generated area can be filled with the changed pixel value, or at least one of the positions where some pixels are obtained can be different from other positions. That is, in order to fill the generated area of ​​A×B, C×D data other than A×B data of the corresponding area can be used. In other words, at least one of the motion vectors applied to the pixels used to fill the generated area can be different from other pixels. In the above example, when the 360-degree image is composed of multiple faces according to the projection format, the generated area can be filled with data of other faces. The data processing method for filling the area generated when the image is expanded by image resizing is not limited to the above example. The data processing method can be improved or changed, or an additional data processing method can be used.

[0269] Multiple candidate groups for data processing methods can be supported according to encoding / decoding settings, and information about selecting a data processing method from the multiple candidate groups can be generated and added to the bitstream. For example, a data processing method can be selected from a filling method by using a predetermined pixel value, a filling method by copying an external pixel, a filling method by copying a specific area of ​​an image, a filling method by transforming a specific area of ​​an image, etc., and related selection information can be generated. In addition, the data processing method can be implicitly determined.

[0270] For example, the data processing method applied to all the areas to be generated with the expansion by image resizing (here, the areas TL to BR in the portion 7a) may be one of a filling method by using a predetermined pixel value, a filling method by copying external pixels, a filling method by copying a specific area of ​​an image, a filling method by transforming a specific area of ​​an image, etc., and the related selection information may be generated. In addition, one predetermined data processing method applied to the entire area may be determined.

[0271] Alternatively, applied to the area to be generated with the expansion by image resizing (here, Figure 7 The data processing method for each of the regions TL to BR in the portion 7a or two or more regions) may be one of a filling method by using a predetermined pixel value, a filling method by copying external pixels, a filling method by copying a specific region of an image, a filling method by transforming a specific region of an image, etc., and the related selection information may be generated. In addition, a predetermined data processing method applied to at least one region may be determined.

[0272] Fig. 9 is an exemplary diagram of a method of constructing a region to be deleted by reduction and a region to be generated in an image resizing method according to an embodiment of the present invention.

[0273] The area to be deleted in the image reduction process can be removed not only simply but also after a series of application processes.

[0274] Referring to part 9a, during the image reduction process, specific areas (A0, A1, A2) may be simply removed without additional application processing. In this case, the image (A) may be divided into areas TL to BR as shown in part 8a.

[0275] Referring to section 9b, the area (A0 to A2) may be removed, and the area (A0 to A2) may be used as reference information when encoding or decoding the image (A). For example, the deleted area (A0 to A2) may be utilized during a process of restoring or correcting a specific area of ​​the image (A) deleted by reduction. During the restoration or correction process, a weighted sum, an average, or the like of two areas (the deleted area and the generated area) may be used. Furthermore, the restoration or correction process may be a process that may be applied when the two areas have a high correlation.

[0276] As an example, an area (B'2) deleted when the image is reduced to the left or right in association with lateral image resizing can be used to restore or correct pixels in an area (B2, LC) opposite to the area deleted when the image is reduced to the left or right in association with lateral image resizing, and then the area (B'2) can be removed from the memory.

[0277] As an example, an area (B'1) deleted when an image is reduced upward or downward in association with longitudinal image resizing can be used for encoding / decoding processing (restoration or correction processing) of an area (B1, TR) opposite to the area deleted when an image is reduced upward or downward in association with longitudinal image resizing, and then the area (B'1) can be removed from the memory.

[0278] As an example, a region (B'0) that is deleted when an image is reduced by some image resizing (here, diagonally with respect to the image center) can be used for encoding / decoding processing (restoration or correction processing) of a region (B0, TL) opposite to the deleted region, and then the region (B'0) can be removed from the memory.

[0279] An example has been described in which data of an area that has continuity at the boundary between both ends of the image and is symmetrical about the resizing direction is used for restoration or correction. However, the present invention is not limited to this, and data of areas TL to BR other than the symmetrical area may be used for restoration or correction and then removed from the memory.

[0280] The data processing method for removing the area to be deleted is not limited to the above-mentioned example. The data processing method may be improved or changed, or an additional data processing method may be used.

[0281] A plurality of candidate groups for data processing methods may be supported according to encoding / decoding settings, and relevant selection information may be generated and added to a bitstream. For example, a data processing method may be selected from a method of simply removing a region to be deleted, a method of removing the region after using the region to be deleted in a series of processes, and the like, and relevant selection information may be generated. In addition, the data processing method may be determined implicitly.

[0282] For example, apply to all areas that are to be removed as they are reduced by image resizing (here, Figure 7 The data processing method of the area TL to BR in the portion 7b) may be one of a method of simply removing the area to be deleted, a method of removing the area to be deleted after using the area to be deleted in a series of processes, etc., and the related selection information may be generated. In addition, the data processing method may be implicitly determined.

[0283] Alternatively, it is applied to each area to be deleted as it is reduced by image resizing (here, Figure 7 The data processing method for each of the regions TL to BR in the portion 7b of the image processing unit 7a) may be one of a method of simply removing the region to be deleted, a method of removing the region to be deleted after using the region to be deleted in a series of processes, etc., and the related selection information may be generated. In addition, the data processing method may be implicitly determined.

[0284] An example in which resizing is performed according to a resizing (expansion or reduction) operation has been described. In some cases, the description may be applied to an example in which a resizing operation (here, expansion) is performed and then an inverse resizing operation (here, reduction) is performed.

[0285] For example, a method of filling the area generated with expansion with some data of the image may be selected, and then a method of removing the area to be deleted with reduction in the inverse process after using the area in the process of restoring or correcting some data of the image may be selected. Alternatively, a method of filling the area generated with expansion by copying external pixels may be selected, and then a method of simply removing the area to be deleted with reduction in the inverse process may be selected. That is, the data processing method in the inverse process may be determined based on the data processing method selected in the image resizing process.

[0286] Unlike the above example, the data processing method of the image resizing process and the data processing method of the inverse process may have an independent relationship. That is, the data processing method in the inverse process may be selected regardless of the data processing method selected in the image resizing process. For example, a method of filling an area generated with expansion by using some data of the image may be selected, and then a method of simply removing an area to be deleted with reduction in the inverse process may be selected.

[0287] According to the present invention, the data processing method during the image resizing process may be implicitly determined according to the encoding / decoding setting, and the data processing method during the inverse process may be implicitly determined according to the encoding / decoding setting. Alternatively, the data processing method during the image resizing process may be explicitly generated, and the data processing method during the inverse process may be explicitly generated. Alternatively, the data processing method during the image resizing process may be explicitly generated, and the data processing method during the inverse process may be implicitly determined based on the data processing method.

[0288] Next, an example of performing image resizing in an encoding / decoding device according to an embodiment of the present invention will be described. In the following description, as an example, the resizing process indicates expansion, and the inverse resizing process indicates reduction. In addition, the difference between the image before resizing and the image after resizing may refer to the image size, and the resizing-related information may have some segments that are explicitly generated and other segments that are implicitly determined according to the encoding / decoding settings. In addition, the resizing-related information may include information about the resizing process and the inverse resizing process.

[0289] As a first example, a process of resizing an input image may be performed before starting encoding. The input image may be resized using resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is used during the resizing process), and then the input image may be encoded. Image encoding data (here, the image after resizing) may be stored in a memory after encoding is completed, and may be added to a bitstream and then transmitted.

[0290] The resizing process may be performed before starting decoding. The image decoded data may be resized using resizing information (e.g., resizing operation, resizing direction, resizing value, etc.), and then the image decoded data may be parsed for decoding. The output image may be stored in a memory after decoding is completed, and may be changed to an image before resizing by performing an inverse resizing process (here, a data processing method, etc. is used; this is for the inverse resizing process).

[0291] As a second example, a process of resizing a reference picture may be performed before starting encoding. The size of the reference picture may be resized using resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is used during the resizing process), and then the reference picture (here, the resized reference picture) may be stored in a memory. The image may be encoded using the resized reference picture. After the encoding is completed, the image encoding data (here, data obtained by encoding using the reference picture) may be added to the bitstream and then transmitted. In addition, the above-mentioned resizing process may be performed when the encoded image is stored in the memory as a reference picture.

[0292] Before starting decoding, a resizing process may be performed on the reference picture. The resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is used during the resizing process) may be used to resize the reference picture, and then the reference picture (here, the resized reference picture) may be stored in a memory. The image decoding data (here, encoded using the reference picture by the encoder) may be parsed for decoding. After decoding is completed, an output image may be generated. The above-mentioned resizing process may be performed when the decoded image is stored in the memory as a reference picture.

[0293] As a third example, resizing processing may be performed on an image before filtering the image (here, assuming a deblocking filter) and after encoding (in detail, after encoding other than the filtering processing is completed). The image may be resized using resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is used during resizing), and then the resized image may be generated and then filtered. After filtering is completed, inverse resizing processing is performed so that the resized image may be changed to the image before resizing.

[0294] After decoding is completed (in detail, after decoding except for filtering is completed), and before filtering, resizing processing may be performed on the image. The image may be resized using resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is used during resizing), and then the resized image may be generated and then filtered. After filtering is completed, inverse resizing processing is performed so that the resized image may be changed to the image before resizing.

[0295] In some cases (the first example and the third example), the resizing process and the inverse resizing process may be performed. In other cases (the second example), only the resizing process may be performed.

[0296] In addition, in some cases (the second example and the third example), the same resizing process may be applied to the encoder and the decoder. In other cases (the first example), the same or different resizing processes may be applied to the encoder and the decoder. Here, the resizing processes of the encoder and the decoder may differ in the resizing execution steps. For example, in some cases (here, the encoder), a resizing execution step that considers image resizing and data processing for the resized area may be included. In other cases (here, the decoder), a resizing execution step that considers image resizing may be included. Here, the previous data processing may correspond to the latter data processing during the inverse resizing process.

[0297] In addition, in some cases (third example), the resizing process may be applied only to the corresponding steps, and the resized area may not be stored in the memory. For example, in order to use the resized area in the filtering process, the resized area may be stored in a temporary memory, filtered, and then removed by the inverse resizing process. In this case, there is no size change of the image due to the resizing. The present invention is not limited to the above examples, and the above examples may be modified.

[0298] The size of the image can be changed by the resizing process, so the coordinates of some pixels of the image can be changed by the resizing process. This may affect the operation of the picture division unit. According to the present invention, through this process, block-based division can be performed based on the image before resizing or the image after resizing. In addition, unit-based (e.g., tile, slice, etc.) division can be performed based on the image before resizing or the image after resizing, which can be determined according to the encoding / decoding settings. According to the present invention, the following description focuses on the case where the picture division unit operates based on the image after resizing (e.g., image division processing after resizing processing), but other modifications may be made. The above example will be described under multiple image settings to be described below.

[0299] The encoder can add the information generated during the above process to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the relevant information from the bitstream. In addition, the information can be included in the bitstream in the form of SEI or metadata.

[0300] Generally, the input image can be encoded or decoded as is or after image reconstruction. For example, image reconstruction can be performed to improve image coding efficiency, image reconstruction can be performed to consider network and user environments, and image reconstruction can be performed according to the type, characteristics, etc. of the image.

[0301] According to the present invention, the image reconstruction process may include the reconstruction process alone or may include the reconstruction process and the inverse reconstruction process. The following example description will focus on the reconstruction process, but the inverse reconstruction process may be derived inversely from the reconstruction process.

[0302] Fig.10 is an exemplary diagram of image reconstruction according to an embodiment of the present invention.

[0303] Assume that portion 10a shows an initial input image. Portions 10a to 10d are example diagrams of an image rotated by a predetermined angle including 0 degrees (for example, a candidate group can be generated by sampling 360 degrees into k portions; k can have values ​​of 2, 4, 8, etc.; in this example, it is assumed that k is 4). Portions 10e to 10h are example diagrams having an opposite (or symmetrical) relationship with respect to portion 10a or with respect to portions 10b to 10d.

[0304] The starting position or scanning order of the image may be changed according to image reconstruction, but the starting position and scanning order may be predetermined regardless of reconstruction, which may be determined according to encoding / decoding settings. The following embodiments assume that the starting position (e.g., the upper left position of the image) and the scanning order (e.g., raster scanning) are predetermined regardless of image reconstruction.

[0305] The image encoding method and the image decoding method according to the embodiment of the present invention may include the following image reconstruction steps. In this case, the image reconstruction process may include an image reconstruction indication step, an image reconstruction type identification step, and an image reconstruction execution step. In addition, the image encoding device and the image decoding device may be configured to include an image reconstruction indication unit, an image reconstruction type identification unit, and an image reconstruction execution unit that respectively perform the image reconstruction indication step, the image reconstruction type identification step, and the image reconstruction execution step. For encoding, relevant syntax elements may be generated. For decoding, relevant syntax elements may be parsed.

[0306] In the image reconstruction indication step, it may be determined whether to perform image reconstruction. For example, when a signal indicating image reconstruction (e.g., convert_enabled_flag) is confirmed, reconstruction may be performed. When a signal indicating image reconstruction is not confirmed, reconstruction may not be performed, or reconstruction may be performed by confirming other encoding / decoding information. In addition, although a signal indicating image reconstruction is not provided, a signal indicating image reconstruction may be implicitly activated or deactivated according to encoding / decoding settings (e.g., characteristics, type, etc. of an image). When reconstruction is performed, corresponding reconstruction-related information may be generated, or corresponding reconstruction-related information may be implicitly determined.

[0307] When a signal indicating image reconstruction is provided, the corresponding signal is a signal for indicating whether to perform image reconstruction. Whether to reconstruct the corresponding image can be determined according to the signal. For example, it is assumed that a signal indicating image reconstruction (e.g., convert_enabled_flag) is confirmed. When the corresponding signal is activated (e.g., convert_enabled_flag=1), reconstruction can be performed. When the corresponding signal is deactivated (e.g., convert_enabled_flag=0), reconstruction may not be performed.

[0308] In addition, when a signal indicating image reconstruction is not provided, reconstruction may not be performed, or whether the corresponding image is reconstructed may be determined by another signal. For example, reconstruction may be performed according to the characteristics, type, etc. of an image (e.g., a 360-degree image), and reconstruction information may be explicitly generated or may be specified as a predetermined value. The present invention is not limited to the above examples, and the above examples may be modified.

[0309] In the image reconstruction type identification step, the image reconstruction type may be identified. The image reconstruction type may be defined by a reconstruction method, reconstruction mode information, and the like. The reconstruction method (e.g., convert_type_flag) may include flipping, rotation, and the like, and the reconstruction mode information may include a mode of the reconstruction method (e.g., convert_mode). In this case, the reconstruction related information may consist of the reconstruction method and the mode information. That is, the reconstruction related information may consist of at least one syntax element. In this case, the number of candidate groups of the mode information may be the same or different depending on the reconstruction method.

[0310] As an example, the rotation may include candidates with uniform spacing (here 90 degrees), as shown in sections 10a to 10d. Section 10a shows a 0 degree rotation, section 10b shows a 90 degree rotation, section 10c shows a 180 degree rotation, and section 10d shows a 270 degree rotation (here, measured clockwise).

[0311] As an example, the flipping may include candidates as shown in parts 10a, 10e, and 10f. While part 10a shows no flipping, parts 10e and 10f show horizontal flipping and vertical flipping, respectively.

[0312] In the above examples, the setting for rotation with uniform intervals and the setting for flipping have been described. However, this is only an example of image reconstruction, and the present invention is not limited thereto, and the present invention may include other interval differences, other flipping operations, etc. that can be determined according to the encoding / decoding setting.

[0313] Alternatively, comprehensive information (eg, convert_com_flag) generated by mixing the reconstruction method and the corresponding mode information may be included. In this case, the reconstruction related information may be composed of a mixture of the reconstruction method and the mode information.

[0314] For example, the comprehensive information may include candidates as shown in parts 10a to 10f, which may be examples of 0 degree rotation, 90 degree rotation, 180 degree rotation, 270 degree rotation, horizontal flip, and vertical flip with respect to part 10a.

[0315] Alternatively, the comprehensive information may include candidates as shown in portions 10a to 10h, which may be examples of 0 degree rotation, 90 degree rotation, 180 degree rotation, 270 degree rotation, horizontal flip, vertical flip, 90 degree rotation and then horizontal flip (or horizontal flip and then 90 degree rotation), and 90 degree rotation and then vertical flip (or vertical flip and then 90 degree rotation) or examples of 0 degree rotation, 90 degree rotation, 180 degree rotation, 270 degree rotation, horizontal flip, 180 degree rotation and then horizontal flip (or horizontal flip and then 180 degree rotation), 90 degree rotation and then horizontal flip (or horizontal flip and then 90 degree rotation), and 270 degree rotation and then horizontal flip (or horizontal flip and then 270 degree rotation).

[0316] The candidate group can be configured to include a rotation mode, a flip mode, and a combination mode of rotation and flip. The combination mode may simply include the mode information in the reconstruction method, and may include a mode generated by mixing the mode information in each method. In this case, the combination mode may include a mode generated by mixing at least one mode of some methods (e.g., rotation) and at least one mode of other methods (e.g., flip). In the above example, the combination mode includes a case generated by combining one mode of some methods with multiple modes of some methods (here, 90 degree rotation + multiple flips / horizontal flips + multiple rotations). The mixed constructed information may include a case where reconstruction is not applied (here, part 10a) as a candidate group, and may include a case where reconstruction is not applied as the first candidate group (e.g., specifying #0 as the index).

[0317] Alternatively, the hybrid constructed information may include mode information corresponding to a predetermined reconstruction method. In this case, the reconstruction related information may consist of the mode information corresponding to the predetermined reconstruction method. That is, the information about the reconstruction method may be omitted, and the reconstruction related information may consist of one syntax element associated with the mode information.

[0318] For example, the reconstruction-related information may be configured to include rotation-specific candidates as shown in parts 10a to 10d. Alternatively, the reconstruction-related information may be configured to include flip-specific candidates as shown in parts 10a, 10e, and 10f.

[0319] The image before the image reconstruction process and the image after the image reconstruction process may have the same size or at least one different length, which may be determined according to the encoding / decoding setting. The image reconstruction process may be a process of rearranging pixels in the image (here, an inverse pixel rearrangement process is performed during an inverse image reconstruction process; this may be reversely derived from the pixel rearrangement process), and thus the position of at least one pixel may be changed. The pixel rearrangement may be performed according to a rule based on the image reconstruction type information.

[0320] In this case, the pixel rearrangement may be affected by the size and shape (eg, square or rectangular) of the image. In detail, the width and height of the image before the reconstruction process and the width and height of the image after the reconstruction process may serve as variables during the pixel rearrangement process.

[0321] For example, ratio information about at least one of the ratio of the width of the image before the reconstruction process to the width of the image after the reconstruction process, the ratio of the width of the image before the reconstruction process to the height of the image after the reconstruction process, the ratio of the height of the image before the reconstruction process to the width of the image after the reconstruction process, and the ratio of the height of the image before the reconstruction process to the height of the image after the reconstruction process (for example, the former / the latter or the latter / the former) can serve as a variable during the pixel rearrangement process.

[0322] In this example, when the image before the reconstruction process and the image after the reconstruction process have the same size, the ratio of the width of the image to the height of the image can serve as a variable during the pixel rearrangement process. In addition, when the image is a square shape, the ratio of the length of the image before the reconstruction process to the length of the image after the reconstruction process can serve as a variable during the pixel rearrangement process.

[0323] In the image reconstruction execution step, image reconstruction may be performed based on the identified reconstruction information. That is, image reconstruction may be performed based on information on the reconstruction type, reconstruction mode, etc., and encoding / decoding may be performed based on the acquired reconstructed image.

[0324] Next, an example in which image reconstruction is performed in the encoding / decoding device according to the embodiment of the present invention will be described.

[0325] The process of reconstructing the input image may be performed before starting encoding. Reconstruction may be performed using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), and the reconstructed image may be encoded. Image encoding data may be stored in a memory after encoding is completed, and may be added to a bitstream and then transmitted.

[0326] A reconstruction process may be performed before decoding is started. Reconstruction may be performed using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), and image decoded data may be parsed for decoding. The image may be stored in a memory after decoding is completed, and the image may be changed to an image before reconstruction by performing an inverse reconstruction process and then output.

[0327] The encoder can add the information generated during the above process to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the relevant information from the bitstream. In addition, the information can be included in the bitstream in the form of SEI or metadata.

[0328] [Table 1]

[0329]

[0330] Table 1 represents example syntax elements associated with partitioning in an image setting. The following description will focus on the additional syntax elements. In addition, in the following examples, the syntax elements are not limited to any specific unit and can be supported in various units such as sequences, pictures, slices, and tiles. Alternatively, the syntax elements may be included in SEI, metadata, etc. In addition, the types, order, conditions, etc. of the syntax elements supported in the following examples are limited to the examples, and therefore can be changed and determined according to the encoding / decoding settings.

[0331] In Table 1, tile_header_enabled_flag represents a syntax element indicating whether encoding / decoding settings for tiles are supported. When this syntax element is activated (tile_header_enabled_flag=1), encoding / decoding settings in tile units may be provided. When this syntax element is deactivated (tile_header_enabled_flag=0), encoding / decoding settings in tile units may not be provided, and encoding / decoding settings in upper units may be allocated.

[0332] In addition, tile_coded_flag represents a syntax element indicating whether a tile is encoded or decoded. When the syntax element is activated (tile_coded_flag=1), the corresponding tile can be encoded or decoded. When the syntax element is deactivated (tile_coded_flag=0), the corresponding tile cannot be encoded or decoded. Here, not performing encoding may mean that encoding data is not generated for the corresponding tile (here, it is assumed that the corresponding area is processed by a predetermined rule, etc.; meaningless areas in some projection formats applicable to 360-degree images). Not performing decoding means that the decoded data in the corresponding tile is no longer parsed (here, it is assumed that the corresponding area is processed by a predetermined rule). In addition, no longer parsing the decoded data may mean that there is no encoding data in the corresponding unit and therefore no parsing is performed, and may also mean that even if the encoding data exists, parsing is no longer performed by the flag. The header information of the tile unit can be supported according to whether the tile is encoded or decoded.

[0333] The above examples focus on tiles. However, the present invention is not limited to tiles, and the above description can be modified and then applied to other division units of the present invention. In addition, the examples of tile division settings are not limited to the above, and the above can be modified.

[0334] [Table 2]

[0335]

[0336] Table 2 shows example syntax elements associated with reconstruction in a picture setting.

[0337] Referring to Table 2, convert_enabled_flag indicates a syntax element indicating whether to perform reconstruction. When this syntax element is activated (convert_enabled_flag=1), the reconstructed image is encoded or decoded, and additional reconstruction related information can be checked. When this syntax element is deactivated (convert_enabled_flag=0), the original image is encoded or decoded.

[0338] Furthermore, convert_type_flag indicates mixed information about a reconstruction method and mode information. A method may be determined from a plurality of candidate groups for a method of applying rotation, a method of applying flipping, and a method of applying rotation and flipping.

[0339] [Table 3]

[0340]

[0341] Table 3 shows example syntax elements associated with resizing among image settings.

[0342] Referring to Table 3, pic_width_in_samples and pic_height_in_samples represent syntax elements indicating the width and height of an image. The size of an image can be checked through these syntax elements.

[0343] In addition, img_resizing_enabled_flag represents a syntax element indicating whether to perform image resizing. When this syntax element is activated (img_resizing_enabled_flag=1), the image is encoded or decoded after resizing, and additional resizing related information can be checked. When this syntax element is deactivated (img_resizing_enabled_flag=0), the original image is encoded or decoded. In addition, the syntax element can indicate resizing for intra prediction.

[0344] Furthermore, resizing_met_flag indicates a resizing method. One resizing method may be determined from a candidate group such as a resizing method based on a scale factor (resizing_met_flag=0), a resizing method based on an offset factor (resizing_met_flag=1), and the like.

[0345] In addition, resizing_mov_flag indicates a syntax element for a resizing operation. For example, one of expansion and reduction can be determined.

[0346] Furthermore, width_scale and height_scale indicate scale factors associated with horizontal resizing and vertical resizing of resizing based on the scale factor.

[0347] Furthermore, top_height_offset and bottom_height_offset represent the “upward” and “downward” offset factors associated with horizontal resizing based on the offset factor, and left_width_offset and right_width_offset represent the “left” and “right” offset factors associated with vertical resizing based on the offset factor.

[0348] The size of the image after resizing may be updated through the resizing related information and the image size information.

[0349] Furthermore, resizing_type_flag represents a syntax element indicating a data processing method for a resized region.The number of candidate groups of data processing methods may be the same or different depending on the resizing method and the resizing operation.

[0350] The image setting processes applied to the above-described image encoding / decoding apparatus may be executed individually or in combination. The following example description will focus on an example in which a plurality of image setting processes are executed in combination.

[0351] Fig.11 11 is an example diagram showing images before and after image setting processing according to an embodiment of the present invention. In detail, part 11a shows an example before image reconstruction is performed on the divided images (for example, an image projected during 360-degree image encoding), and part 11b shows an image after image reconstruction is performed on the divided images (for example, an image packed during 360-degree image encoding). That is, it can be understood that part 11a is an example diagram before performing image setting processing, and part 11b is an example diagram after performing image setting processing.

[0352] In this example, image division (here, tiles are assumed) and image reconstruction are described as image setting processing.

[0353] In the following example, image reconstruction is performed after image division is performed. However, depending on the encoding / decoding setting, image division may be performed after image reconstruction is performed, and it may be modified. In addition, the above-mentioned image reconstruction processing (including inverse processing) may be applied identically or similarly to the reconstruction processing in the division unit in the image in the present embodiment.

[0354] Image reconstruction may or may not be performed in all division units in the image, and may be performed in some division units. Therefore, the division units before reconstruction (e.g., some of P0 to P5) may be the same as or different from the division units after reconstruction (e.g., some of S0 to S5). Various image reconstruction situations will be described by the following examples. In addition, for ease of description, it is assumed that the unit of the image is a picture, the unit of dividing the image is a tile, and the division unit is a rectangular shape.

[0355] As an example, whether to perform image reconstruction may be determined in some units (e.g., sps_convert_enabled_flag or SEI or metadata, etc.). Alternatively, whether to perform image reconstruction may be determined in some units (e.g., pps_convert_enabled_flag). This may be allowed when it first appears in the corresponding unit (here, the picture) or when it is activated in the upper unit (e.g., sps_convert_enabled_flag=1). Alternatively, whether to perform image reconstruction may be determined in some units (e.g., tile_convert_flag[i]; i is the partition unit index). This may be allowed when it first appears in the corresponding unit (here, the tile) or when it is activated in the upper unit (e.g., pps_convert_enabled_flag=1). In addition, in part, whether to perform image reconstruction may be implicitly determined according to the encoding / decoding settings, so the relevant information may be omitted.

[0356] As an example, whether to reconstruct the division unit in the image can be determined according to a signal indicating image reconstruction (eg, pps_convert_enabled_flag). In detail, whether to reconstruct all division units in the image can be determined according to the signal. In this case, a single signal indicating image reconstruction can be generated in the image.

[0357] As an example, whether to reconstruct the division units in the image may be determined according to a signal indicating image reconstruction (e.g., tile_convert_flag[i]). In detail, whether to reconstruct some division units in the image may be determined according to the signal. In this case, at least one signal indicating image reconstruction (e.g., a number of signals equal to the number of division units) may be generated.

[0358] As an example, whether to reconstruct an image may be determined according to a signal indicating image reconstruction (e.g., pps_convert_enabled_flag), and whether to reconstruct a division unit in the image may be determined according to a signal indicating image reconstruction (e.g., tile_convert_flag[i]). In detail, when any signal is activated (e.g., pps_convert_enabled_flag=1), any other signal (e.g., tile_convert_flag[i]) may be additionally checked, and whether to reconstruct some division units in the image may be determined according to the signal (here, tile_convert_flag[i]). In this case, a plurality of signals indicating image reconstruction may be generated.

[0359] When the signal instructing image reconstruction is activated, image reconstruction related information may be generated. In the following examples, various image reconstruction related information will be described.

[0360] As an example, reconstruction information applied to an image may be generated. In detail, one piece of reconstruction information may be used as reconstruction information of all division units in the image.

[0361] As an example, reconstruction information applied to a division unit in an image may be generated. In detail, at least one piece of reconstruction information may be used as reconstruction information for some division units in the image. That is, one piece of reconstruction information may be used as reconstruction information for one division unit, or one piece of reconstruction information may be used as reconstruction information for multiple division units.

[0362] The following example will be described in conjunction with an example of performing image reconstruction.

[0363] For example, when a signal indicating image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information commonly applied to division units in an image may be generated. Alternatively, when a signal indicating image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information individually applied to division units in an image may be generated. Alternatively, when a signal indicating image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information individually applied to division units in an image may be generated. Alternatively, when a signal indicating image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information commonly applied to division units in an image may be generated.

[0364] The reconstruction information may be processed implicitly or explicitly according to the encoding / decoding settings. For implicit processing, the reconstruction information may be specified as a predetermined value according to the characteristics, type, etc. of the image.

[0365] P0 to P5 in part 11a may correspond to S0 to S5 in part 11b, and a reconstruction process may be performed on the division unit. For example, P0 may not be reconstructed, and then P0 may be assigned to S0. P1 may be rotated 90 degrees, and then may be assigned to S1. P2 may be rotated 180 degrees, and then may be assigned to S2. P3 may be flipped horizontally, and then may be assigned to S3. P4 may be rotated 90 degrees and flipped horizontally, and then may be assigned to S4. P5 may be rotated 180 degrees and flipped horizontally, and then may be assigned to S5.

[0366] However, the present invention is not limited to the above examples, and various modifications may be made to the above examples. Similar to the above examples, the division units in the image may not be reconstructed, or at least one of reconstruction using rotation, reconstruction using flipping, and reconstruction using rotation and flipping in combination may be performed.

[0367] When image reconstruction is applied to the division unit, additional reconstruction processing such as the rearrangement of the division unit can be performed. That is, the image reconstruction processing according to the present invention can be configured to include the rearrangement of the division units in the image and the rearrangement of the pixels in the image, and can be represented using some syntax elements in Table 4 (e.g., part_top, part_left, part_width, part_height, etc.). This means that the image division processing and the image reconstruction processing can be understood in combination. In the above example, it has been described that the image is divided into a plurality of units.

[0368] P0 to P5 in part 11a may correspond to S0 to S5 in part 11b, and a reconstruction process may be performed on the division unit. For example, P0 may not be reconstructed, and then P0 may be assigned to S0. P1 may not be reconstructed, and then P1 may be assigned to S2. P2 may be rotated 90 degrees, and then may be assigned to S1. P3 may be flipped horizontally, and then may be assigned to S4. P4 may be rotated 90 degrees and flipped horizontally, and then may be assigned to S5. P5 may be flipped horizontally and then rotated 180 degrees, and then may be assigned to S3. The present invention is not limited to this, and various modifications may also be made thereto.

[0369] also, Figure 7 The P_Width and P_Height can correspond to Fig.11 P_Width and P_Height, and Figure 7 The P'_Width and P'_Height can correspond to Fig.11 P'_Width and P'_Height. Figure 7 The size of the resized image after P'_Width × P'_Height can be expressed as (P_Width + Exp_L + Exp_R) × (P_Height + Exp_T + Exp_B), and Fig.11 The size P'_Width×P'_Height of the resized image can be expressed as (P_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R)×(P_Height+Var0_T+Var1_T+Var0_B+Var1_B) or (Sub_P0_Width+Sub_P1_Width+Sub_P2_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R)×(Sub_P0_Height+Sub_P1_Height+Var0_T+Var1_T+Var0_B+Var1_B).

[0370] Similar to the above example, for image reconstruction, the rearrangement of pixels in the division units of the image may be performed, the rearrangement of the division units in the image may be performed, and both the rearrangement of pixels in the division units of the image and the rearrangement of the division units in the image may be performed. In this case, the rearrangement of the division units in the image may be performed after the rearrangement of pixels in the division units is performed, or the rearrangement of pixels in the division units may be performed after the rearrangement of the division units in the image is performed.

[0371] Whether to perform rearrangement of the division units in the image may be determined based on a signal indicating image reconstruction. Alternatively, a signal for rearranging the division units in the image may be generated. In detail, the signal may be generated when the signal indicating image reconstruction is activated. Alternatively, the signal may be processed implicitly or explicitly based on encoding / decoding settings. For implicit processing, the signal may be determined based on characteristics, type, etc. of the image.

[0372] Furthermore, information on rearrangement of division units in an image may be performed implicitly or explicitly according to encoding / decoding settings and may be determined according to characteristics, types, etc. of the image. That is, each division unit may be arranged according to arrangement information predetermined for the division unit.

[0373] Next, an example of a division unit in reconstructing an image in the encoding / decoding device according to an embodiment of the present invention will be described.

[0374] The input image may be divided using the division information before starting encoding. The division unit may be reconstructed using the reconstruction information, and the image reconstructed for each division unit may be encoded. The image encoding data may be stored in a memory after encoding is completed, and may be added to a bit stream and then transmitted.

[0375] The division process may be performed using the division information before starting decoding. The reconstruction process may be performed on the division unit using the reconstruction information, and the image decoded data may be parsed to be decoded in the reconstructed division unit. After the decoding is completed, the image decoded data may be stored in a memory, and a plurality of division units may be merged into a single unit after performing an inverse reconstruction process in the division unit, so that an image may be output.

[0376] Fig.12 2 is an exemplary diagram of adjusting the size of each division unit of an image according to an embodiment of the present invention. Fig.12 P0 to P5 correspond to Fig.11 P0 to P5, and Fig.12 S0 to S5 correspond to Fig.11 S0 to S5.

[0377] In the following example, the description will focus on the case where image resizing is performed after image division is performed. However, depending on the encoding / decoding setting, image division may be performed after image resizing is performed, and it may be modified. In addition, the above-mentioned image resizing processing (including inverse processing) may be applied identically or similarly to the image division unit resizing processing in the present embodiment.

[0378] For example, Figure 7 The TL to BR can correspond to Fig.12 The division units SX (S0 to S5) are TL to BR. Figure 7 The S0 and S1 can correspond to Fig.12 PX and SX. Figure 7 The P_Width and P_Height can correspond to Fig.12 Sub_PX_Width and Sub_PX_Height. Figure 7 The P'_Width and P'_Height can correspond to Fig.12 Sub_SX_Width and Sub_SX_Height. Figure 7 Exp_L, Exp_R, Exp_T and Exp_B can correspond to Fig.12 VarX_L, VarX_R, VarX_T and VarX_B, and other factors may also correspond.

[0379] The processing of adjusting the size of the division unit in the image in parts 12a to 12f is similar to Figure 7 The image expansion or reduction in parts 7a and 7b may differ in that the setting for the image expansion or reduction may exist in proportion to the number of division units. In addition, the process of adjusting the size of the division units in the image may be different from the image expansion or reduction in that there are settings applied to the division units in the image in common or individually. In the following examples, various resizing cases will be described, and the resizing process may be performed in consideration of the above description.

[0380] According to the present invention, image resizing may or may not be performed on all division units in an image, and image resizing may be performed on some division units. Various image resizing situations will be described by the following examples. In addition, for ease of description, it is assumed that the resizing operation is used for expansion, the resizing operation is based on an offset factor, the resizing directions are an "upward" direction, a "downward" direction, a "leftward" direction, and a "rightward" direction, the resizing direction is set to operate through resizing information, the unit of the image is a picture, and the unit of dividing the image is a tile.

[0381] As an example, whether to perform image resizing may be determined in some units (e.g., sps_img_resizing_enabled_flag or SEI or metadata, etc.). Alternatively, whether to perform image resizing may be determined in some units (e.g., pps_img_resizing_enabled_flag). This may be allowed when it first appears in the corresponding unit (here, the picture) or when it is activated in the upper unit (e.g., sps_img_resizing_enabled_flag=1). Alternatively, whether to perform image resizing may be determined in some units (e.g., tile_resizing_flag[i]; i is the partition unit index). This may be allowed when it first appears in the corresponding unit (here, the tile) or when it is activated in the upper unit. In addition, in part, whether to perform image resizing may be implicitly determined according to the encoding / decoding settings, so the relevant information may be omitted.

[0382] As an example, whether to adjust the size of the division unit in the image can be determined according to a signal indicating image size adjustment (e.g., pps_img_resizing_enabled_flag). In detail, whether to adjust the size of all division units in the image can be determined according to the signal. In this case, a single signal indicating image size adjustment can be generated.

[0383] As an example, it may be determined whether to adjust the size of the division unit in the image according to a signal indicating image size adjustment (e.g., tile_resizing_flag[i]). In detail, it may be determined whether to adjust the size of some division units in the image according to the signal. In this case, at least one signal indicating image size adjustment (e.g., a number of signals equal to the number of division units) may be generated.

[0384] As an example, whether to adjust the image size may be determined according to a signal indicating image size adjustment (e.g., pps_img_resizing_enabled_flag), and whether to adjust the size of the division unit in the image may be determined according to a signal indicating image size adjustment (e.g., tile_resizing_flag[i]). In detail, when any signal is activated (e.g., pps_img_resizing_enabled_flag=1), any other signal (e.g., tile_resizing_flag[i]) may be additionally checked, and whether to adjust the size of some division units in the image may be performed according to the signal (here, tile_resizing_flag[i]). In this case, a plurality of signals indicating image size adjustment may be generated.

[0385] When a signal indicating image resizing is activated, image resizing related information may be generated. In the following examples, various image resizing related information will be described.

[0386] As an example, resizing information applied to an image may be generated. In detail, one piece of resizing information or a set of resizing information may be used as resizing information for all division units in the image. For example, one piece of resizing information commonly applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction of the division units in the image (or resizing values ​​applied to all resizing directions supported or allowed in the division units; in this example, one piece of information) or a set of resizing information individually applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction (or a number of resizing information equal to the number of resizing directions allowed or supported by the division units; in this example, up to four pieces of information) may be generated.

[0387] As an example, resizing information applied to a division unit in an image may be generated. In detail, at least one piece of resizing information or a set of resizing information may be used as resizing information for all division units in an image. That is, one piece of resizing information or a set of resizing information may be used as resizing information for one division unit or as resizing information for multiple division units. For example, resizing information for the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction commonly applied to one division unit in an image may be generated, or a set of resizing information applied individually to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction may be generated. Alternatively, resizing information for the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction commonly applied to multiple division units in an image may be generated, or a set of resizing information applied individually to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction may be generated. The configuration of the resizing group means resizing value information about at least one resizing direction.

[0388] In summary, resizing information commonly applied to division units in an image may be generated. Alternatively, resizing information individually applied to division units in an image may be generated. The following example will be described in conjunction with an example of performing image resizing.

[0389] For example, when a signal indicating image resizing (e.g., pps_img_resizing_enabled_flag) is activated, resizing information commonly applied to division units in an image may be generated. Alternatively, when a signal indicating image resizing (e.g., pps_img_resizing_enabled_flag) is activated, resizing information individually applied to division units in an image may be generated. Alternatively, when a signal indicating image resizing (e.g., tile_resizing_flag[i]) is activated, resizing information individually applied to division units in an image may be generated. Alternatively, when a signal indicating image resizing (e.g., tile_resizing_flag[i]) is activated, resizing information commonly applied to division units in an image may be generated.

[0390] The resizing direction, resizing information, etc. of an image may be processed implicitly or explicitly according to encoding / decoding settings. For implicit processing, the resizing information may be specified as a predetermined value according to characteristics, type, etc. of the image.

[0391] It has been described that the resizing direction in the resizing process of the present invention may be at least one of the "upward" direction, the "downward" direction, the "leftward" direction, and the "rightward" direction, and the resizing direction and the resizing information may be processed explicitly or implicitly. That is, the resizing value (including 0; which means no resizing) may be implicitly predetermined for some directions, and the resizing value (including 0; which means no resizing) may be explicitly specified for other directions.

[0392] Even in the division unit in the image, the resizing direction and the resizing information may be set to be implicitly or explicitly processed, and this may be applied to the division units in the image. For example, a setting applied to one division unit in the image may appear (here, a number of settings equal to the number of division units may appear), a setting applied to a plurality of division units in the image may appear, or a setting applied to all division units in the image may appear (here, one setting may appear), and at least one setting may appear in the image (for example, a number of settings from one to the number of division units may appear). Setting information applied to the division units in the image may be collected, and a single set of settings may be defined.

[0393] Fig.13 This is an example diagram of a group of settings or size adjustments of division units in an image.

[0394] In detail, Fig.13 Various examples of implicitly or explicitly processing resizing directions and resizing information for division units in an image are shown. In the following examples, for ease of description, the implicit processing assumes that resizing values ​​of some resizing directions are 0.

[0395] As shown in section 13a, resizing can be handled explicitly when the boundary of the division unit matches the boundary of the image (here, thick solid line), and resizing can be handled implicitly when the boundary of the division unit does not match the boundary of the image (thin solid line). For example, P0 can be resized in the "upward" direction and the "leftward" direction (a2, a0), P1 can be resized in the "upward" direction (a2), P2 can be resized in the "upward" direction and the "rightward" direction (a2, a1), P3 can be resized in the "downward" direction and the "leftward" direction (a3, a0), P4 can be resized in the "downward" direction (a3), and P5 can be resized in the "downward" direction and the "rightward" direction (a3, a1). In this case, resizing in other directions may not be allowed.

[0396] As shown in section 13b, some directions of the dividing unit (here, up and down) can allow explicit resizing, and some directions of the dividing unit (here, left and right) can allow explicit resizing when the boundary of the dividing unit matches the boundary of the image (here, thick solid line), and can allow implicit resizing when the boundary of the dividing unit does not match the boundary of the image (here, thin solid line). For example, P0 can be resized in the "upward" direction, the "downward" direction, and the "leftward" direction (b2, b3, b0), P1 can be resized in the "upward" direction and the "downward" direction (b2, b3), P2 can be resized in the "upward" direction, the "downward" direction, and the "rightward" direction (b2, b3, b1), P3 can be resized in the "upward" direction, the "downward" direction, and the "leftward" direction (b3, b4, b0), P4 can be resized in the "upward" direction and the "downward" direction (b3, b4), and P5 can be resized in the "upward" direction, the "downward" direction, and the "rightward" direction (b3, b4, b1). In this case, resizing in other directions may not be allowed.

[0397] As shown in section 13c, some directions of the dividing unit (here left and right) can allow explicit resizing, and some directions of the dividing unit (here up and down) can allow explicit resizing when the boundary of the dividing unit matches the boundary of the image (here thick solid line), and can allow implicit resizing when the boundary of the dividing unit does not match the boundary of the image (here thin solid line). For example, P0 can be resized in the "upward", "leftward" and "rightward" directions (c4, c0, c1), P1 can be resized in the "upward", "leftward" and "rightward" directions (c4, c1, c2), P2 can be resized in the "upward", "leftward" and "rightward" directions (c4, c2, c3), P3 can be resized in the "downward", "leftward" and "rightward" directions (c5, c0, c1), P4 can be resized in the "downward", "leftward" and "rightward" directions (c5, c1, c2), and P5 can be resized in the "downward", "leftward" and "rightward" directions (c5, c2, c3). In this case, resizing in other directions may not be allowed.

[0398] The settings related to image resizing similar to the above example can have various cases. Multiple groups of settings are supported so that setting group selection information can be explicitly generated, or a predetermined setting group can be implicitly determined according to encoding / decoding settings (eg, characteristics, type, etc. of an image).

[0399] Fig.14 is an exemplary diagram showing both a process of adjusting the size of an image and a process of adjusting the size of a division unit in the image.

[0400] Reference Fig.14 , the process of resizing the image and the inverse process can be performed in directions e and f, and the process of resizing the division unit in the image and the inverse process can be performed in directions d and g. That is, the resizing process can be performed on the image, and then the resizing process can be performed on the division unit in the image. The resizing order may not be fixed. This means that multiple resizing processes are possible.

[0401] In summary, the image resizing process can be classified into resizing of the image (or resizing the image before division) and resizing of the division unit in the image (or resizing the image after division). The resizing of the image and the resizing of the division unit in the image may not be performed, either one of the resizing of the image and the resizing of the division unit in the image may be performed, or both, which can be determined according to encoding / decoding settings (e.g., characteristics, type, etc. of the image).

[0402] When a plurality of resizing processes are performed in this example, resizing of the image may be performed in at least one of the "upward" direction, the "downward" direction, the "leftward" direction, and the "rightward" direction of the image, and the size of at least one division unit in the image may be adjusted. In this case, resizing may be performed in at least one of the "upward" direction, the "downward" direction, the "leftward" direction, and the "rightward" direction of the division unit to be resized.

[0403] Reference Fig.14 , the size of the image (A) before resizing may be defined as P_Width×P_Height, the size of the image after primary resizing (or the image before secondary resizing; B) may be defined as P'_Width×P'_Height, and the size of the image after secondary resizing (or the image after final resizing; C) may be defined as P''_Width×P''_Height. The image (A) before resizing represents an image on which no resizing is performed, the image (B) after primary resizing represents an image on which some resizing is performed, and the image (C) after secondary resizing represents an image on which all resizing is performed. For example, the image (B) after primary resizing may represent an image on which resizing is performed by division units of the image as shown in parts 13a to 13c, and the image (C) after secondary resizing may represent an image on which resizing is performed by division units of the image as shown in parts 13a to 13c. Figure 7 The image obtained by completely adjusting the size of the image (B) after the primary resizing as shown in part 7a of FIG. The opposite is also possible. However, the present invention is not limited to the above examples, and various modifications can be made to the above examples.

[0404] In the size of the image (B) after the primary resizing, P'_Width may be obtained by P_Width and at least one horizontal resizing value of the horizontal resizing, and P'_Height may be obtained by P_Height and at least one vertical resizing value of the vertical resizing. In this case, the resizing value may be a resizing value generated in a division unit.

[0405] In the size of the image (C) after the secondary resizing, P''_Width may be obtained by P'_Width and at least one horizontal resizing value of the horizontal resizing, and P''_Height may be obtained by P'_Height and at least one vertical resizing value of the vertical resizing. In this case, the resizing value may be a resizing value generated in the image.

[0406] In summary, the size of the image after resizing may be obtained through at least one resizing value and the size of the image before resizing.

[0407] In the resized area of ​​the image, information about the data processing method can be generated. Through the following examples, various data processing methods will be described. The data processing method generated during the inverse resizing process can be applied the same or similarly to the data processing method of the resizing process. The data processing methods in the resizing process and the inverse resizing process will be described through various combinations to be described below.

[0408] As an example, a data processing method applied to an image may be generated. In detail, one data processing method or a group of data processing methods may be used as the data processing method for all division units in the image (here, it is assumed that the sizes of all division units are to be adjusted). For example, one data processing method commonly applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction of the division units in the image (or a data processing method applied to all resizing directions supported or allowed in the division units, etc.; in this example, one piece of information) or a group of data processing methods applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction (or a number of data processing methods equal to the number of resizing directions supported or allowed in the division units; in this example, up to four pieces of information) may be generated.

[0409] As an example, a data processing method applied to a division unit in an image may be generated. In detail, at least one data processing method or a group of data processing methods may be used as a data processing method for some division units in an image (here, it is assumed that the size of the division unit is to be adjusted). That is, a data processing method or a group of data processing methods may be used as a data processing method for one division unit or a data processing method for multiple division units. For example, a data processing method commonly applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction of a division unit in an image may be generated, or a group of data processing methods individually applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction may be generated. Alternatively, a data processing method commonly applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction of multiple division units in an image may be generated, or a group of data processing methods individually applied to the "upward" direction, "downward" direction, "leftward" direction, and "rightward" direction may be generated. The configuration of the group of data processing methods means a data processing method for at least one size adjustment direction.

[0410] In summary, a data processing method commonly applied to the division units in the image can be used. Alternatively, a data processing method individually applied to the division units in the image can be used. The data processing method can use a predetermined method. A predetermined data processing method can be provided as at least one method. This corresponds to implicit processing, and selection information for the data processing method can be explicitly generated, which can be determined according to the encoding / decoding settings (e.g., characteristics, type, etc. of the image).

[0411] That is, a data processing method commonly applied to the division units in the image may be used. A predetermined method may be used, or one of a plurality of data processing methods may be selected. Alternatively, a data processing method individually applied to the division units in the image may be used. Depending on the division unit, a predetermined method may be used, or one of a plurality of data processing methods may be selected.

[0412] In the following examples, some cases of resizing a division unit in an image (here, assuming that the resizing is for expansion) will be described (here, the resized area is filled with some data of the image).

[0413] The size of a specific area TL to BR of some cells (e.g., S0 to S5 in the parts 12a to 12f) can be adjusted using data of a specific area tl to br of some cells P0 to P5 (in the parts 12a to 12f). In this case, some cells may be the same as each other (e.g., S0 and P0) or different (e.g., S0 and P1). That is, the area TL to BR to be resized may be filled with some data tl to br of the corresponding division unit, and may be filled with some data of the division unit other than the corresponding division unit.

[0414] As an example, the size of the area TL to BR whose size is adjusted by the current division unit can be adjusted using the data tl to br of the current division unit. For example, TL of S0 can be filled with data tl of P0, RC of S1 can be filled with data tr+rc+br of P1, BL+BC of S2 can be filled with data bl+bc+br of P2, and TL+LC+BL of S3 can be filled with data tl+lc+bl of P3.

[0415] As an example, the size of the area TL to BR whose size the current division unit is adjusted may be adjusted using data tl to br of a division unit spatially adjacent to the current division unit. For example, in the "upward" direction, TL+TC+TR of S4 may be filled with data bl+bc+br of P1, in the "downward" direction, BL+BC of S2 may be filled with data tl+tc+tr of P5, in the "leftward" direction, LC+BL of S2 may be filled with data tl+rc+bl of P1, in the "rightward" direction, RC of S3 may be filled with data tl+lc+bl of P4, and in the "downward+leftward" direction, BR of S0 may be filled with data tl of P4.

[0416] As an example, the size of the area TL to BR where the size of the current division unit is adjusted can be adjusted using data tl to br of a division unit that is not spatially adjacent to the current division unit. For example, data in a boundary area (e.g., horizontally, vertically, etc.) between both ends of the image can be acquired. The LC of S3 can be acquired using data tr+rc+br of S5, the RC of S2 can be acquired using data tl+lc of S0, the BC of S4 can be acquired using data tc+tr of S1, and the TC of S1 can be acquired using data bc of S4.

[0417] Alternatively, data of a specific region of the image (a region that is not spatially adjacent to the resized region but is determined to have a high correlation with the resized region) may be acquired. BC of S1 may be acquired using data tl+lc+bl of S3, RC of S3 may be acquired using data tl+tc of S1, and RC of S5 may be acquired using data bc of S0.

[0418] Furthermore, some cases of resizing the division unit in the image (here, assuming that the resizing is for reduction) are as follows (here, removal is performed by restoration or correction using some data of the image).

[0419] Specific areas TL to BR of some units (e.g., S0 to S5 in parts 12a to 12f) can be used for restoration or correction processing of specific areas tl to br of some units P0 to P5. In this case, some units may be the same as each other (e.g., S0 and P0) or different (e.g., S0 and P2). That is, the area to be resized can be used to restore some data of the corresponding divided unit and then removed, and the area to be resized can be used to restore some data of the divided unit other than the corresponding divided unit and then removed. Detailed examples can be reversely derived from the expansion process, so they will be omitted.

[0420] This example can be applied to the case where data with high correlation exists in the area to be resized, and information about the position to which the resizing is referred can be explicitly generated or implicitly acquired according to a predetermined rule. Alternatively, the relevant information can be checked in combination. This can be an example that can be applied in the case of acquiring data from another area with continuity in the encoding of a 360-degree image.

[0421] Next, an example of adjusting the size of a division unit in an image in an encoding / decoding device according to an embodiment of the present invention will be described.

[0422] The input image may be divided before the encoding is started. The division unit may be resized using the resizing information, and the image may be encoded after the division unit is resized. The image encoding data may be stored in a memory after the encoding is completed, and may be added to a bit stream and then transmitted.

[0423] The division process may be performed using the division information before starting decoding. The resizing process may be performed on the division unit using the resizing information, and the image decoded data may be parsed to be decoded in the resized division unit. After the decoding is completed, the image decoded data may be stored in a memory, and a plurality of division units may be merged into a single unit after performing an inverse resizing process on the division unit, so that the image may be output.

[0424] Other examples of the above-described image resizing process may be applied. The present invention is not limited thereto and modifications may be made thereto.

[0425] In the image setting process, it is allowed to combine image resizing and image reconstruction. Image reconstruction can be performed after image resizing. Alternatively, image resizing can be performed after image reconstruction is performed. In addition, it is allowed to combine image division, image reconstruction and image resizing. Image resizing and image reconstruction can be performed after image division is performed. The order of image setting is not fixed and can be changed, which can be determined according to the encoding / decoding settings. In this example, the image setting process is described as performing image reconstruction and image resizing after performing image division. However, depending on the encoding / decoding settings, other orders are possible and can also be modified.

[0426] For example, the image setting process may be performed in the following order: partition -> reconstruction; reconstruction -> partition; partition -> resizing; resizing -> partition; resizing -> reconstruction; reconstruction -> resizing; partition -> reconstruction -> resizing; partition -> resizing -> reconstruction; resizing -> partition -> reconstruction; resizing -> partition -> reconstruction; resizing -> reconstruction -> partition; reconstruction -> partition -> resizing; and reconstruction -> resizing -> partition, and combinations with additional image settings may be possible. As described above, the image setting process may be performed sequentially, but some or all of the setting processes may be performed simultaneously. In addition, as some of the image setting processes, multiple processes may be performed according to the encoding / decoding settings (e.g., characteristics, type, etc. of the image). The following examples indicate various combinations of image setting processes.

[0427] As an example, P0 to P5 in part 11a may correspond to S0 to S5 in part 11b, and reconstruction processing (here, rearrangement of pixels) and resizing processing (here, resizing the division units to have the same size) may be performed in the division units. For example, the sizes of P0 to P5 may be adjusted based on the offset, and P0 to P5 may be assigned to S0 to S5. In addition, P0 may not be reconstructed, and then P0 may be assigned to S0. P1 may be rotated 90 degrees, and then may be assigned to S1. P2 may be rotated 180 degrees, and then may be assigned to S2. P3 may be rotated 270 degrees, and then may be assigned to S3. P4 may be flipped horizontally, and then may be assigned to S4. P5 may be flipped vertically, and then may be assigned to S5.

[0428] As an example, P0 to P5 in part 11a may correspond to the same or different positions as S0 to S5 in part 11b, and reconstruction processing (here, rearrangement of pixels and division units) and resizing processing (here, resizing the division units to have the same size) may be performed in the division units. For example, the sizes of P0 to P5 may be adjusted based on the ratio, and P0 to P5 may be assigned to S0 to S5. In addition, P0 may not be reconstructed, and then P0 may be assigned to S0. P1 may not be reconstructed, and then P1 may be assigned to S2. P2 may be rotated 90 degrees, and then may be assigned to S1. P3 may be flipped horizontally, and then may be assigned to S4. P4 may be rotated 90 degrees and flipped horizontally, and then may be assigned to S5. P5 may be flipped horizontally and then rotated 180 degrees, and then may be assigned to S3.

[0429] As an example, P0 to P5 in portion 11a may correspond to E0 to E5 in portion 5e, and a reconstruction process (here, rearrangement of pixels and division units) and a resizing process (here, resizing the division units to have different sizes) may be performed in the division units. For example, P0 may not be resized and reconstructed and then may be assigned to E0, P1 may be resized but not reconstructed based on a ratio and then may be assigned to E1, P2 may not be resized but reconstructed and then may be assigned to E2, P3 may be resized but not reconstructed based on an offset and then may be assigned to E4, P4 may not be resized but reconstructed and may be assigned to E5, and P5 may be resized and reconstructed based on an offset and then may be assigned to E3.

[0430] Similar to the above example, the absolute position or relative position of the division unit in the image before and after the image setting process may be maintained or changed, which may be determined according to the encoding / decoding settings (e.g., the characteristics, type, etc. of the image). In addition, various combinations of image setting processes may be possible. The present invention is not limited thereto, and thus various modifications may be made thereto.

[0431] The encoder can add the information generated during the above process to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the relevant information from the bitstream. In addition, the information can be included in the bitstream in the form of SEI or metadata.

[0432] [Table 4]

[0433]

[0434] Table 4 represents example syntax elements associated with multiple image settings. The following description will focus on the additional syntax elements. In addition, in the following examples, the syntax elements are not limited to any specific unit and can be supported in various units such as sequences, pictures, slices, and tiles. Alternatively, the syntax elements can be included in SEI, metadata, etc.

[0435] Referring to Table 4, parts_enabled_flag represents a syntax element indicating whether to divide some units. When this syntax element is activated (parts_enabled_flag=1), the image can be divided into multiple units, and these multiple units can be encoded or decoded. In addition, additional division information can be checked. When this syntax element is deactivated (parts_enabled_flag=0), the original image is encoded or decoded. In this example, the description will focus on rectangular division units such as tiles, and different settings for existing tiles and division information can be provided.

[0436] Here, num_partitions refers to a syntax element indicating the number of partition units, and num_partitions plus 1 is equal to the number of partition units.

[0437] In addition, part_top[i] and part_left[i] refer to syntax elements indicating position information of a division unit, and represent the horizontal start position and the vertical start position of the division unit (for example, the upper left position of the division unit). In addition, part_width[i] and part_height[i] refer to syntax elements indicating size information of a division unit, and represent the width and height of the division unit. In this case, the start position and the size information may be set in units of pixels or in units of blocks. In addition, the syntax element may be a syntax element that may be generated during the image reconstruction process or a syntax element that may be generated when the image division process and the image reconstruction process are constructed in combination.

[0438] In addition, part_header_enabled_flag indicates a syntax element indicating whether encoding / decoding settings for partition units are supported. When this syntax element is activated (part_header_enabled_flag=1), encoding / decoding settings for partition units can be provided. When this syntax element is deactivated (part_header_enabled_flag=0), encoding / decoding settings cannot be provided, and encoding / decoding settings for higher-level units can be allocated.

[0439] The above example is not limited to the example of the syntax element associated with resizing and reconstruction in the division unit among the image settings, and it can be modified for other division units and settings of the present invention. This example has been described under the assumption that resizing and reconstruction are performed after division is performed, but the present invention is not limited thereto, and it can be modified in other image setting orders, etc. In addition, the types, order, conditions, etc. of the syntax elements supported in the following example are limited to this example, and therefore can be changed and determined according to the encoding / decoding settings.

[0440] [Table 5]

[0441]

[0442] Table 5 represents example syntax elements associated with reconstruction in partition units among picture settings.

[0443] Referring to Table 5, part_convert_flag[i] represents a syntax element indicating whether to reconstruct a partition unit. This syntax element can be generated for each partition unit. When this syntax element is activated (part_convert_flag[i]=1), the reconstructed partition unit can be encoded or decoded, and additional reconstruction-related information can be checked. When this syntax element is deactivated (part_convert_flag[i]=0), the original partition unit is encoded or decoded. Here, convert_type_flag[i] refers to mode information about the reconstruction of the partition unit, and can be information about pixel rearrangement.

[0444] In addition, a syntax element indicating additional reconstruction such as partition unit rearrangement may be generated. In this example, partition unit rearrangement may be performed by part_top and part_left, which are syntax elements indicating the above-mentioned image partitioning, or a syntax element associated with partition unit rearrangement (eg, index information) may be generated.

[0445] [Table 6]

[0446]

[0447] Table 6 represents example syntax elements associated with resizing in partition units among image settings.

[0448] Referring to Table 6, part_resizing_flag[i] represents a syntax element indicating whether to adjust the size of the partition unit in the image. This syntax element can be generated for each partition unit. When this syntax element is activated (part_resizing_flag[i]=1), the resized partition unit can be encoded or decoded after resizing, and additional resizing related information can be checked. When this syntax element is deactivated (part_resiznig_flag[i]=0), the original partition unit is encoded or decoded.

[0449] Furthermore, width_scale[i] and height_scale[i] represent scale factors associated with horizontal resizing and vertical resizing of scale factor-based resizing in the division unit.

[0450] In addition, top_height_offset[i] and bottom_height_offset[i] represent the offset factor for the "upward" direction and the offset factor for the "downward" direction associated with the offset factor-based resizing in the partition unit, and left_width_offset[i] and right_width_offset[i] represent the offset factor for the "left" direction and the offset factor for the "right" direction associated with the offset factor-based resizing in the partition unit.

[0451] In addition, resizing_type_flag[i][j] represents a syntax element indicating a data processing method for a resized region in a partition unit. This syntax element indicates a separate data processing method for a resizing direction. For example, a syntax element indicating separate data processing methods for a resized region in an "upward" direction, a "downward" direction, a "leftward" direction, and a "rightward" direction may be generated. This syntax element may be generated based on resizing information (for example, it may be generated only when resizing is performed in some directions).

[0452] The above-mentioned image setting processing may be processing applied according to the characteristics, type, etc. of the image. In the following examples, even if not particularly mentioned, the above-mentioned image setting processing may be applied without any change or with any change. In the following examples, the description will focus on the case of addition or change in the above-mentioned examples.

[0453] For example, a 360-degree image or an omnidirectional image generated by a 360-degree camera has characteristics different from those of an image acquired by a general camera, and has an encoding environment different from that of compression of a general image.

[0454] Unlike a general image, a 360-degree image may not have a boundary portion with discontinuity, and data of all regions of the 360-degree image may have continuity. In addition, a device such as an HMD may require a high-definition image because the image should be replayed in front of the eye through a lens. When an image is acquired by a stereo camera device, the amount of image data processed may increase. Various image setting processes considering a 360-degree image may be performed to provide an effective encoding environment including the above-mentioned examples.

[0455] A 360-degree camera may be multiple cameras or a camera with multiple lenses and sensors. The cameras or lenses may cover all directions around any center point captured by the camera.

[0456] Various methods can be used to encode a 360-degree image. For example, a 360-degree image can be encoded using various image processing algorithms in a 3D space, and a 360-degree image can be converted into a 2D space and encoded using various image processing algorithms. According to the present invention, the following description will focus on a method for converting a 360-degree image into a 2D space and encoding or decoding the converted image.

[0457] The 360-degree image encoding apparatus according to an embodiment of the present invention may include Figure 1 Some or all of the elements shown in , and may also include a pre-processing unit configured to pre-process (stitch, project, regional packing) the input image. At the same time, the 360-degree image decoding device according to an embodiment of the present invention may include Figure 2 The invention may include some or all of the elements shown in , and may further include a post-processing unit configured to post-process the encoded image before decoding the encoded image to reproduce the output image.

[0458] In other words, the encoder may preprocess an input image, encode the preprocessed image, and transmit a bitstream including the image, and the decoder may parse, decode, and post-process the transmitted bitstream to generate an output image. In this case, the transmitted bitstream may include information generated during the preprocessing process and information generated during the encoding process, and the bitstream may be parsed and used during the decoding process and the post-processing process.

[0459] Subsequently, the operating method for the 360-degree image encoder will be described in more detail, and a person skilled in the art can easily derive the operating method for the 360-degree image decoder. Since the operating method for the 360-degree image decoder is opposite to the operating method for the 360-degree image encoder, a detailed description of the operating method for the 360-degree image decoder will be omitted.

[0460] The input image may be subjected to stitching and projection processing performed on a sphere-based 3D projection structure, and the image data on the 3D projection structure may be projected into a 2D image through the processing.

[0461] The projected image may be configured to include some or all of the 360-degree content according to the encoding setting. In this case, the position information of the area (or pixel) to be placed at the center of the projected image may be implicitly generated as a predetermined value, or the position information of the area (or pixel) may be explicitly generated. In addition, when the projected image includes a specific area of ​​the 360-degree content, the range information and position information of the included area may be generated. In addition, the range information (e.g., width and height) and position information (e.g., it is measured based on the upper left end of the image) of the region of interest (ROI) may be generated according to the projected image. In this case, a specific area of ​​high importance in the 360-degree content may be set as the ROI. The 360-degree image may allow viewing of all contents in the "upward" direction, the "downward" direction, the "leftward" direction, and the "rightward" direction, but the user's gaze may be limited to a portion of the image, and a portion of the image may be set as the ROI in consideration of the limitation. For the purpose of efficient encoding, the ROI may be set to have good quality and high resolution, and other areas may be set to have lower quality and lower resolution than the ROI.

[0462] Among multiple 360-degree image transmission schemes, a single stream transmission scheme may allow a full image or a viewport image to be transmitted in a separate single bitstream for a user. A multi-stream transmission scheme may allow several full images with different image qualities to be transmitted in multiple bitstreams, so the image quality may be selected according to the user environment and communication conditions. A tile-based streaming scheme may allow separately encoded partial images based on tile units to be transmitted in multiple bitstreams, so tiles may be selected according to the user environment and communication conditions. Therefore, a 360-degree image encoder may generate and transmit a bitstream with two or more qualities, and a 360-degree image decoder may set an ROI according to the user's view, and may selectively decode the bitstream according to the ROI. That is, the place where the user's gaze is directed may be set as an ROI by a head tracking or eye tracking system, and only the required portion may be presented.

[0463] The projection image can be converted into a packed image obtained by performing a regional packing process. The regional packing process may include the step of dividing the projection image into a plurality of regions, and the divided regions may be arranged (or rearranged) in the packed image according to the regional packing setting. When a 360-degree image is converted into a 2D image (or a projection image), regional packing may be performed to increase spatial continuity. Therefore, the size of the image may be reduced by regional packing. In addition, regional packing may be performed to reduce the degradation of image quality caused during rendering, realize viewport-based projection, and provide other types of projection formats. Regional packing may be performed or not performed according to the encoding setting, which may be determined based on a signal indicating whether regional packing is performed (e.g., regionwise_packing_flag; information about regional packing may be generated only when regionwise_packing_flag is activated).

[0464] When regional packaging is performed, setting information (or mapping information) for allocating (or arranging) a specific area of ​​the projected image to a specific area of ​​the packaged image may be displayed (or generated). When regional packaging is not performed, the projected image and the packaged image may be the same image.

[0465] In the above description, the stitching process, the projection process, and the regional packing process are defined as separate processes, but some (e.g., stitching+projection, projection+regional packing) or all (e.g., stitching+projection+regional packing) of the processes may be defined as a single process.

[0466] At least one packed image may be generated from the same input image according to the settings of the stitching process, the projection process, and the region-based packing process. In addition, at least one coded data for the same projected image may be generated according to the settings of the region-based packing process.

[0467] The packed image may be divided by performing a tiling process. In this case, tiling, which is a process of dividing an image into a plurality of regions and then transmitting it, may be an example of a 360-degree image transmission scheme. As described above, tiling may be performed for partial decoding in consideration of the user environment, and tiling may also be performed for efficiently processing a large amount of data of a 360-degree image. For example, when an image consists of one unit, the entire image may be decoded to decode the ROI. On the other hand, when an image consists of a plurality of unit regions, decoding only the ROI may be efficient. In this case, the division may be performed in units of tiles that are division units according to a conventional encoding scheme, or may be performed in accordance with various division units (e.g., quadrilateral division blocks, etc.) that have been described according to the present invention. In addition, the division unit may be a unit for performing independent encoding / decoding. Tiling may be performed independently or based on a projected image or a packed image. That is, the division may be performed based on the face boundary of the projected image, the face boundary of the packed image, the packing setting, etc., and the division may be performed independently for each division unit. This may affect the generation of division information during the tiling process.

[0468] Next, the projected image or the packed image can be encoded. The encoded data and information generated during the preprocessing process can be added to the bitstream, and the bitstream can be transmitted to the 360-degree image decoder. The information generated during the preprocessing process can be added to the bitstream in the form of SEI or metadata. In this case, the bitstream may contain: at least one encoded data having partially different settings for the encoding process; and at least one preprocessing information having partially different settings for the preprocessing process. This is to combine multiple encoded data (encoded data + preprocessing information) to construct a decoded image according to the user environment. In detail, a decoded image can be constructed by selectively combining multiple encoded data. In addition, the process can be performed in two parts to apply to a binocular system, and the process can be performed on another depth image.

[0469] Fig.15 is an example diagram showing a 2D plane space and a 3D space showing a 3D image.

[0470] Typically, for the purpose of a 360-degree 3D virtual space, three degrees of freedom (3DoF) may be required, and three rotations may be supported relative to the X-axis (pitch), Y-axis (yaw), and Z-axis (roll). DoF refers to spatial degrees of freedom, 3DoF refers to degrees of freedom including rotations around the X-axis, Y-axis, and Z-axis as shown in section 15a, and 6DoF refers to degrees of freedom that additionally allow movement along the X-axis, Y-axis, and Z-axis as well as 3DoF. The following description will focus on the image encoding device and image decoding device of the present invention with 3DoF. When 3DoF or greater (3DoF+) is supported, the image encoding device and the image decoding device may be modified or combined with additional processing or devices not shown.

[0471] Referring to section 15a, the yaw may have a range from -π (-180 degrees) to π (180 degrees), the pitch may have a range from -π / 2 radians (or -90 degrees) to π / 2 radians (or 90 degrees), and the roll may have a range from -π / 2 radians (or -90 degrees) to π / 2 radians (or 90 degrees). In this case, when it is assumed that Φ and θ are longitude and latitude in a map representation of the earth, the 3D space coordinates (x, y, z) may be transformed from the 2D space coordinates (Φ, θ). For example, the 3D space coordinates may be derived from the 2D space coordinates according to the transformation formulas x=cos(θ)cos(Φ), y=sin(θ), and z=-cos(θ)sin(Φ).

[0472] In addition, (Φ, θ) can be transformed into (x, y, z). For example, according to the transformation formula Φ = tan -1 (-Z / X) and θ=sin -1 (Y / (X 2 +Y 2 +Z 2 ) 1 / 2 ) derives 2D space coordinates from 3D space coordinates.

[0473] When the pixels in the 3D space are accurately transformed into pixels in the 2D space (e.g., integer unit pixels in the 2D space), the pixels in the 3D space can be mapped to the pixels in the 2D space. When the pixels in the 3D space are not accurately transformed into pixels in the 2D space (e.g., fractional unit pixels in the 2D space), the pixels obtained by interpolation can be mapped to 2D pixels. In this case, as interpolation, nearest neighbor interpolation, bilinear interpolation, B-spline interpolation, bicubic interpolation, etc. can be used. In this case, the relevant information can be explicitly generated by selecting one of a plurality of interpolation candidates, or the interpolation method can be implicitly determined according to a predetermined rule. For example, a predetermined interpolation filter can be used according to the 3D model, projection format, color format, and slice / tile type. In addition, when the interpolation information is explicitly generated, information about filter information (e.g., filter coefficients) can be included.

[0474] Part 15b shows an image in which the 3D space is transformed into a 2D space (2D plane coordinate system). (Φ, θ) may be sampled (i, j) based on the size (width and height) of the image. Here, i may have a range from 0 to P_Width-1, and j may have a range from 0 to P_Height-1.

[0475] (Φ, θ) may be a center point (or reference point; depicted as Fig.15 The setting of the center point can be specified in the 3D space, and the position information of the center point can be explicitly generated or implicitly determined as a predetermined value. For example, the center position information in yaw, the center position information in pitch, the center position information in roll, etc. can be generated. When the value of the information is not individually specified, each value can be assumed to be zero.

[0476] An example of transforming the entire 360-degree image from a 3D space into a 2D space has been described above, but a specific area of ​​the 360-degree image may be transformed, and position information (e.g., some positions belonging to the area; in this example, position information about the center point), range information, etc. of the specific area may be explicitly generated, or the position information, range information, etc. of the specific area may implicitly follow predetermined position and range information. For example, center position information in yaw, center position information in pitch, center position information in roll, range information in yaw, range information in pitch, range information in roll, etc. may be generated, and the specific area may be at least one area. Therefore, position information, range information, etc. of multiple areas may be processed. When the value of information is not specified separately, the entire 360-degree image may be assumed.

[0477] H0 to H6 and W0 to W5 in part 15a indicate some latitudes and longitudes in part 15b, which can be expressed as coordinates (C, j) and (i, C) in part 15b (C is a longitude or latitude component). Unlike a general image, when a 360-degree image is converted into a 2D space, distortion may occur or warping of the content in the image may occur. This may depend on the area of ​​the image, and different encoding / decoding settings may be applied to the position of the image or the area divided according to the position. When the encoding / decoding setting is adaptively applied based on the encoding / decoding information in the present invention, position information (e.g., an x ​​component, a y component, or a range defined by x and y) may be included as an example of the encoding / decoding information.

[0478] The description of the 3D space and the 2D space is defined to help describe the embodiments of the present invention. However, the present invention is not limited thereto, and the above description may be modified in detail or may be applied to other situations.

[0479] As described above, the image acquired by the 360-degree camera device can be transformed into a 2D space. In this case, a 3D model can be used to map the 360-degree image, and various 3D models such as a sphere, a cube, a cylinder, a pyramid, and a polyhedron can be used. When the 360-degree image mapped based on the model is transformed into a 2D space, a projection process can be performed according to a projection format based on the model.

[0480] Figures 16a to 16d is a conceptual diagram showing a projection format according to an embodiment of the present invention.

[0481] Fig.16a An equirectangular projection (ERP) format showing a 360 degree image being projected into a 2D plane. Fig.16b A Cube Map Projection (CMP) format showing a 360 degree image being projected onto a cube. Fig.16c The octahedral projection (OHP) format shows that the 360-degree image is projected onto an octahedron. Fig.16d The 360-degree image is projected onto an icosahedral projection (ISP) format of a polyhedron. However, the present invention is not limited thereto, and various projection formats may be used. Figures 16a to 16d In the figure, the left side shows the 3D mode, and the right side shows an example of transformation into 2D space through projection processing. Various sizes and shapes can be provided according to the projection format. Each shape can be composed of surfaces or faces, and each face can be represented as a circle, triangle, quadrilateral, etc.

[0482] In the present invention, the projection format may be defined by a 3D mode, a face setting (e.g., the number of faces, the shape of the face, the shape configuration of the face, etc.), a projection processing setting, etc. When at least one element is different in definition, the projection format may be regarded as a different projection format. For example, ERP is composed of a spherical model (3D model), one face (the number of faces), and a quadrilateral face (the shape of the face). However, when some of the settings of the projection processing (e.g., the formula used during the transformation from 3D space to 2D space; that is, elements with the same remaining projection settings and producing a difference in at least one pixel of the projected image during the projection processing) are different, the format may be classified into different formats, such as ERP1 and ERP2. As another example, CMP is composed of a cube model, six faces, and a quadrilateral face. When some of the settings during the projection processing (e.g., the sampling method applied during the transformation from 3D space to 2D space) are different, the format may be classified into different formats, such as CMP1 and CMP2.

[0483] When a plurality of projection formats are used instead of one predetermined projection format, the projection format identification information (or projection format information) may be explicitly generated. The projection format identification information may be configured by various methods.

[0484] As an example, the projection format may be identified by assigning index information (e.g., proj_format_flag) to a plurality of projection formats. For example, #0 may be assigned to ERP, #1 may be assigned to CMP, #2 may be assigned to OHP, #3 may be assigned to ISP, #4 may be assigned to ERP1, #5 may be assigned to CMP1, #6 may be assigned to OHP1, #7 may be assigned to ISP1, #8 may be assigned to CMP compact, #9 may be assigned to OHP compact, #10 may be assigned to ISPcompact, and #11 or higher may be assigned to other formats.

[0485] As an example, the projection format may be identified using at least one element information constituting the projection format. In this case, as the element information constituting the projection format, 3D model information (e.g., 3d_model_flag; #0 indicates a sphere, #1 indicates a cube, #2 indicates a cylinder, #3 indicates a pyramid, #4 indicates a polyhedron 1, and #5 indicates a polyhedron 2), face number information (e.g., num_face_flag; a method of increasing by 1 starting from 1; the number of faces generated in the projection format is specified as index information, that is, #0 indicates one, #1 indicates three, #2 indicates six, #3 indicates eight, and #4 indicates twenty), face shape information (e.g., shape_face_flag; #0 indicates a quadrilateral, #1 indicates a circle, #2 indicates a triangle, #3 indicates a quadrilateral + a circle, and #4 indicates a quadrilateral + a triangle), projection processing setting information (e.g., 3d_2d_convert_idx), etc. may be included.

[0486] As an example, the projection format can be identified using the element information constituting the projection format and the projection format index information. For example, as the projection format index information, #0 can be assigned to ERP, #1 can be assigned to CMP, #2 can be assigned to OHP, #3 can be assigned to ISP, and #4 or higher can be assigned to other formats. The projection format (e.g., ERP, ERP1, CMP, CMP1, OHP, OHP1, ISP, and ISP1) can be identified together with the element information constituting the projection format (here, the projection processing setting information). Alternatively, the projection format (e.g., ERP, CMP, CMP compact, OHP, OHP compact, ISP, and ISP compact) can be identified together with the element information constituting the projection format (here, the regional packaging).

[0487] In summary, the projection format can be identified using projection format index information, the projection format can be identified using at least one projection format element information, and the projection format can be identified using projection format index information and at least one projection format element information. This can be defined according to the encoding / decoding setting. In the present invention, the following description assumes that the projection format is identified using the projection format index. In this example, the description will focus on the projection format expressed using faces with the same size and shape, but configurations with faces with different sizes and shapes are possible. In addition, the configuration of each face can be different from Figures 16a to 16d The configuration shown in the figure is the same or different, the number of each face is used as a symbol for identifying the corresponding face, and there is no limitation on a specific order. For ease of description, the following description assumes that, for the projection image, ERP is a projection format including one face + quadrilateral, CMP is a projection format including six faces + quadrilaterals, OHP is a projection format including eight faces + triangles, ISP is a projection format including twenty faces + triangles, and the faces have the same size and shape. However, the description can be applied identically or similarly even for different settings.

[0488] like Figures 16a to 16d As shown, the projection format can be classified into one face (e.g., ERP) or multiple faces (e.g., CMP, OHP, and ISP). In addition, the shape of each face can be classified into a quadrilateral, a triangle, etc. The classification can be an example of the type, characteristics, etc. of the image according to the present invention, which can be applied when different encoding / decoding settings are provided according to the projection format. For example, the type of the image can be a 360-degree image, and the characteristics of the image can be one of the classifications (e.g., each projection format, a projection format with one face or multiple faces, a projection format with a quadrilateral face or a non-quadrilateral face).

[0489] A 2D plane coordinate system (e.g., (I, j)) may be defined in each face of the 2D projection image, and the characteristics of the coordinate system may differ according to the projection format, the position of each face, etc. ERP may have one 2D plane coordinate system, and other projection formats may have multiple 2D plane coordinate systems according to the number of faces. In this case, the coordinate system may be expressed as (k, i, j), and k may indicate index information of each face.

[0490] Fig.17 is a conceptual diagram showing that a projection format according to an embodiment of the present invention is included in a rectangular image.

[0491] That is, it can be understood that parts 17a to 17c show Figures 16b to 16d The projection format is implemented as a rectangular image.

[0492] Referring to parts 17a to 17c, each image format may be configured in a rectangular shape to encode or decode a 360-degree image. For ERP, a single coordinate system may be used as is. However, for other projection formats, the coordinate systems of the faces may be integrated into a single coordinate system, and a detailed description thereof will be omitted.

[0493] Referring to parts 17a to 17c, when constructing a rectangular image, it can be confirmed that an area filled with meaningless data such as a blank or background is generated. That is, the rectangular image may be composed of an area including actual data (here, a face; a valid area) and a meaningless area added for constructing the rectangular image (here, it is assumed that the area is filled with any pixel value; an invalid area). This may reduce performance due to an increase in coded data, that is, an increase in image size caused by encoding / decoding of the meaningless area and the actual image data.

[0494] Therefore, a process of constructing an image by excluding a meaningless area and using an area including actual data may be additionally performed.

[0495] Fig.18 is a conceptual diagram of a method of converting a projection format into a rectangular shape, that is, a method of performing rearrangement on a surface to exclude a meaningless area according to an embodiment of the present invention.

[0496] With reference to parts 18a to 18c, an example for rearranging parts 17a to 17c can be confirmed, and the process can be defined as a regional packing process (CMP compact, OHP compact, ISP compact, etc.). In this case, the surface can be not only rearranged, but also divided and then rearranged (OHP compact, ISP compact, etc.). This can be performed to remove meaningless areas and improve encoding performance through efficient surface arrangement. For example, when images are arranged continuously between surfaces (for example, B2-B3-B1, B5-B0-B4, etc. in part 18a), the prediction accuracy at the time of encoding is enhanced, and thus the encoding performance can be enhanced. Here, regional packing according to the projection format is only an example, and the present invention is not limited to this.

[0497] Fig.19 is a conceptual diagram illustrating performing an area packing process to convert a CMP projection format into a rectangular image according to an embodiment of the present invention.

[0498] Referring to parts 19a to 19c, the CMP projection format may be arranged as 6×1, 3×2, 2×3, and 1×6. In addition, when the size of some faces is adjusted, the arrangement may be performed as shown in parts 19d and 19e. In parts 19a to 19e, CMP is applied as an example. However, the present invention is not limited thereto, and other projection formats may be applied. The arrangement of faces of an image acquired by regional packing may follow a predetermined rule corresponding to the projection format, or information on the arrangement may be explicitly generated.

[0499] The 360-degree image encoding and decoding apparatus according to an embodiment of the present invention may be configured to include Figure 1 and Figure 2 In particular, a format conversion unit configured to convert the projection format and an inverse format conversion unit configured to inversely convert the projection format may also be included in the image encoding device and the image decoding device, respectively. That is, the input image may be processed by the format conversion unit, and then Figure 1 The image encoding device is encoded, and the bit stream can be Figure 2 The image decoding device is decoded and then processed by the inverse format conversion unit to generate an output image. The following description will focus on the processing performed by the encoder (here, the input image, the encoding, etc.), and the processing performed by the decoder can be derived from the encoder in reverse. In addition, the redundant description of the foregoing will be omitted.

[0500] The following description assumes that the input image is the same as the packed image or 2D projection image obtained by performing a preprocessing process by the 360-degree encoding device. That is, the input image may be an image obtained by performing a projection process or a regional packing process according to some projection formats. The projection format pre-applied to the input image may be one of various projection formats, which may be regarded as a common format and is referred to as a first format.

[0501] The format conversion unit can perform a conversion to a projection format other than the first format. In this case, the projection format to which the conversion is to be performed can be referred to as a second format. For example, ERP can be set to a first format, and ERP can be converted into a second format (e.g., ERP2, CMP, OHP, and ISP). In this case, ERP2 includes the following types of EPR formats: the EPR formats have the same conditions such as 3D models and face configurations, but have some different settings. Alternatively, the projection format can be the same format with the same projection format settings (e.g., ERP=ERP2), and can have different image sizes or resolutions. Alternatively, some of the following image setting processes can be applied. For ease of description, such an example has been mentioned, but each of the first format and the second format can be one of various projection formats. However, the present invention is not limited to this and can be modified.

[0502] During the format conversion process, due to different coordinate system characteristics, pixels (integer pixels) of the image after the transformation can be obtained from the fractional unit pixels and integer unit pixels in the image before the transformation, so interpolation can be performed. The interpolation filter used in this case may be the same as or similar to the interpolation filter described above. In this case, the relevant information can be explicitly generated by selecting one of a plurality of interpolation filter candidates, or the interpolation filter can be implicitly determined according to a predetermined rule. For example, a predetermined interpolation filter can be used according to the projection format, color format, and slice / tile type. In addition, when the interpolation filter is explicitly provided, information about filter information (e.g., filter coefficients) may be included.

[0503] In the format conversion section, the projection format may be defined to include regional packaging, etc. That is, projection and regional packaging may be performed during the format conversion process. Alternatively, after the format conversion process, processing such as regional packaging may be performed before encoding is performed.

[0504] The encoder can add the information generated during the above process to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the relevant information from the bitstream. In addition, the information can be included in the bitstream in the form of SEI or metadata.

[0505] Next, the image setting process applied to the 360-degree image encoding / decoding device according to an embodiment of the present invention will be described. The image setting process according to the present invention can be applied to the preprocessing process, post-processing process, format conversion process, inverse format conversion process, etc. of the 360-degree image encoding / decoding device as well as general encoding / decoding process. The following description of the image setting process will focus on the 360-degree image encoding device and may include the above-mentioned image settings. The redundant description of the aforementioned image setting process will be omitted. In addition, the following examples will focus on the image setting process, and the inverse image setting process can be derived in reverse from the image setting process. Some situations can be confirmed through the aforementioned various embodiments of the present invention.

[0506] The image setting process according to the present invention may be performed in a 360-degree image projection step, a regional packaging step, a format conversion step, or other steps.

[0507] Fig. 20 360-degree image segmentation according to an embodiment of the present invention. Fig. 20 In , it is assumed that the image is projected via ERP.

[0508] Part 20a shows an image projected by ERP, and various methods can be used to divide the image. In this example, the description focuses on slices or tiles, and it is assumed that W0 to W2 and H0 and H1 are the dividing boundary lines of the slices or tiles and follow the raster scanning order. The following example focuses on slices and tiles. However, the present invention is not limited to this, and other division methods can be applied.

[0509] For example, the division may be performed in units of slices, and H0 and H1 may be set as division boundaries. Alternatively, the division may be performed in units of tiles, and W0 to W2, H0, and H1 may be set as division boundaries.

[0510] Part 20b shows an example in which an image projected by ERP is divided into tiles (assuming the same tile division boundaries as shown in part 20a (W0 to W2, H0 and H1 are all activated)). When it is assumed that region P is the entire image and region V is the region or viewport where the user's gaze rests, there may be various methods to provide an image corresponding to the viewport. For example, the region corresponding to the viewport may be acquired by decoding the entire image (e.g., tiles a to i). In this case, the entire image may be decoded, and tiles a to i (here, region A+region B) may be decoded in the case of dividing the image. Alternatively, the region corresponding to the viewport may be acquired by decoding the region belonging to the viewport. In this case, in the case of dividing the image, the region corresponding to the viewport may be acquired from the image restored by decoding tiles f, g, j and k (here, region B). The former case may be referred to as full decoding (or viewport-independent encoding), and the latter case may be referred to as partial decoding (or viewport-dependent encoding). The latter case may be an example that may occur in a 360-degree image with a large amount of data. Compared with the partitioning method based on slice units, the partitioning method based on tile units can be used more frequently because the partitioning area can be obtained flexibly. For partial decoding, the referentiality of the partitioning unit can be spatially or temporally restricted (here, implicitly processed) because it is impossible to find where the viewpoint will occur, and encoding / decoding can be performed considering this restriction. The following example will focus on full decoding for description, but will focus on tiles (or the rectangular partitioning method of the present invention) to describe 360-degree image partitioning in preparation for partial decoding. However, the following description can be applied to other partitioning units in the same way or in a modified manner.

[0511] Fig.21 360-degree image segmentation and image reconstruction according to an embodiment of the present invention. Fig.21 In FIG. 5 , it is assumed that the image is projected through a CMP.

[0512] Part 21a shows an image projected by CMP, and various methods can be used to divide the image. It is assumed that W0 to W2, H0 and H1 are division boundary lines of planes, slices and tiles and follow a raster scan order.

[0513] For example, the division may be performed in slice units, and H0 and H1 may be set as division boundaries. Alternatively, the division may be performed in tile units, and W0 to W2, H0, and H1 may be set as division boundaries. Alternatively, the division may be performed in face units, and W0 to W2, H0, and H1 may be set as division boundaries. In this example, it is assumed that a face is part of a division unit.

[0514] In this case, a face may be a division unit (here dependent encoding / decoding) executed to classify or distinguish regions with different properties in the same image (here a plane coordinate system for each face) according to characteristics, types (in this example, 360-degree images and projection formats), etc. of the image, while a slice or tile may be a division unit (here independent encoding / decoding) executed to divide an image according to a user definition. In addition, a face may be a unit divided by a predetermined definition (or induction from projection format information) during a projection process according to a projection format, while a slice or tile may be a unit divided by explicitly generating division information according to a user definition. In addition, a face may have a polygonal division shape including a quadrilateral according to a projection format, a slice may have any division shape that cannot be defined as a quadrilateral or a polygon, and a tile may have a quadrilateral division shape. The setting of the division unit may be defined only for the description of this example.

[0515] In this example, it has been described that the face is a division unit classified for area distinction. However, the face may be a unit for performing independent encoding / decoding according to the encoding / decoding setting as at least one face unit, and may have a setting for performing independent encoding / decoding in combination with a tile, a slice, etc. In this case, when the face is combined with a tile, a slice, etc., explicit information of the tile and the slice may be generated, or the tile and the slice may be implicitly combined based on the face information. Alternatively, explicit information of the tile and the slice may be generated based on the face information.

[0516] As a first example, one image division process (here, plane) is performed, and the image division can implicitly omit the division information (which is acquired based on the projection format information). This example is used to rely on the encoding / decoding setting, and can be an example corresponding to the case where the referenceability between plane units is not limited.

[0517] As a second example, one image division process (here, face) is performed, and the image division can explicitly generate division information. This example is used to rely on the encoding / decoding setting, and can be an example corresponding to the case where the referenceability between face units is not limited.

[0518] As a third example, multiple image division processes (here, faces and tiles) are performed, some image divisions (here, faces) may implicitly omit or explicitly generate division information, and other image divisions (here, tiles) may explicitly generate division information. In this example, one image division process (here, faces) precedes other image division processes (here, tiles).

[0519] As a fourth example, multiple image division processes are performed, some image divisions (here, faces) can implicitly omit or explicitly generate division information, and other image divisions (here, tiles) can explicitly generate division information based on these image divisions (here, faces). In this example, one image division process (here, face) precedes other image division processes (here, tiles). In some cases of this example (assuming the second example), it can be the same as explicitly generating division information, but there may be differences in the division information configuration.

[0520] As a fifth example, multiple image division processes are performed, some image divisions (here, faces) can implicitly omit division information, and other image divisions (here, tiles) can implicitly omit division information based on these some image divisions (here, faces). For example, a face unit can be set as a tile unit individually, or multiple face units (here, when adjacent faces have continuity, face units are grouped; otherwise, face units are not grouped; B2-B3-B1 and B4-B0-B5 in part 18a) can be set as tile units. According to a predetermined rule, the face unit can be set as a tile unit. This example is used for independent encoding / decoding settings, and can be an example corresponding to a case where the referenceability between face units is limited. That is, in some cases (assuming the first example), it can be the same as implicitly processing the division information, but there may be differences in the encoding / decoding settings.

[0521] This example may be a description of a case where division processing may be performed in a projection step, a region packing step, an initial encoding / decoding step, etc., and may be any other image division processing performed in an encoder / decoder.

[0522] In part 21a, a rectangular image can be constructed by adding an area (B) that does not include data to an area (A) that includes data. In this case, the position, size, shape, number, etc. of area A and area B can be information that can be checked by a projection format, etc. or can be checked when information about a projected image is explicitly generated, and the relevant information can be represented by the above-mentioned image division information, image reconstruction information, etc. For example, information about a specific area of ​​a projected image (e.g., part_top, part_left, part_width, part_height, and part_convert_flag) can be represented as shown in Tables 4 and 5. However, the present invention is not limited thereto and can be applied to other situations (e.g., another projection format, other projection settings, etc.).

[0523] Region B and region A may be constructed as a single image and then encoded or decoded. Alternatively, the division may be performed in consideration of regional characteristics, and different encoding / decoding settings may be applied. For example, encoding or decoding may not be performed on region B by using information about whether encoding or decoding is performed (e.g., tile_coded_flag when it is assumed that the division unit is a tile). In this case, the corresponding region may be restored to certain data (here any pixel value) according to a predetermined rule. Alternatively, in the above-mentioned image division process, region B may have a different encoding / decoding setting from region A. Alternatively, the corresponding region may be removed by performing regional packing processing.

[0524] Part 21b shows an example where an image packed by CMP is divided into tiles, slices or faces. In this case, the packed image is an image that is subjected to face rearrangement processing or regional packing processing, and can be an image obtained by performing image division and image reconstruction according to the present invention.

[0525] In part 21b, a rectangular shape may be constructed to include an area having data. In this case, the position, size, shape, number, etc. of the area may be information that can be checked by a predetermined setting or information that can be checked when information about a packed image is explicitly generated, and the relevant information may be represented by the above-mentioned image division information, image reconstruction information, etc. For example, information about a specific area of ​​a packed image (e.g., part_top, part_left, part_width, part_height, and part_convert_flag) may be represented as shown in Tables 4 and 5.

[0526] Various partitioning methods may be used to partition the packed image. For example, the partitioning may be performed in slice units, and H0 may be set as the partition boundary. Alternatively, the partitioning may be performed in tile units, and W0, W1, and H0 may be set as the partition boundaries. Alternatively, the partitioning may be performed in face units, and W0, W1, and H0 may be set as the partition boundaries.

[0527] The image division process and the image reconstruction process according to the present invention can be performed on the projected image. In this case, the reconstruction process can be used to rearrange the faces in the image and the pixels in the image. This can be a possible example when the image is divided into a plurality of faces or is composed of a plurality of faces. The following example will be described focusing on the case where the image is divided into tiles based on face units.

[0528] SX, Y (S0,0 to S3,2) in part 21a may correspond to S'U, V (S'0,0 to S'2,1) in part 21b (here, X and Y may be the same as or different from U and V), and reconstruction processing may be performed on a face-by-face basis. For example, S2,1, S3,1, S0,1, S1,2, S1,1, and S1,0 may be assigned to S'0,0, S'1,0, S'2,0, S'0,1, S'1,1, and S'2,1 (face rearrangement). In addition, S2,1, S3,1, and S0,1 may not be reconstructed (pixel rearrangement), and S1,2, S1,1, and S1,0 may be rotated 90 degrees and then reconstructed. This may be represented as shown in part 21c. In part 21c, the horizontally placed symbols (S1,0, S1,1, S1,2) may be images that are horizontally placed to maintain the continuity of the image.

[0529] The reconstruction of the face can be handled implicitly or explicitly depending on the encoding / decoding settings. The implicit processing can be performed according to predetermined rules taking into account the type (here 360 ​​degree image) and characteristics (here projection format, etc.) of the image.

[0530] For example, for S'0,0 and S'1,0, S'1,0 and S'2,0, S'0,1 and S'1,1, S'1,1 and S'2,1 in portion 21c, there is image continuity (or correlation) between two faces with respect to the face boundary, and portion 21c may be an example of continuity between three upper faces and three lower faces. In the case where an image is divided into a plurality of faces by a projection process from a 3D space to a 2D space and then the image is packed for each region, reconstruction may be performed to increase image continuity between faces so as to efficiently reconstruct the faces. Such face reconstruction may be predetermined and processed.

[0531] Alternatively, the reconstruction process may be performed by an explicit process, and the reconstruction information may be generated.

[0532] For example, when information (e.g., one of implicitly acquired information and explicitly generated information) about an M×N configuration (e.g., 6×1, 3×2, 2×3, 1×6, etc. for CMP compact; in this example, a 3×2 configuration is assumed) is checked by region packing processing, surface reconstruction can be performed according to the M×N configuration and then information about surface reconstruction can be generated. For example, when rearranging surfaces in an image, index information (or information about a position in an image) can be assigned to each surface. When rearranging pixels in a surface, pattern information for reconstruction can be assigned.

[0533] Index information can be Fig.18 In parts 21a to 21c, SX,Y or S'U,V represents each face using position information indicating width and height (e.g., S[i][j]) or using one piece of position information (e.g., S[i]; assuming that the position information is assigned starting from the upper left corner of the image in a raster scan order), and an index of each face can be assigned thereto.

[0534] For example, when the index is assigned using position information indicating the width and height, the surface index #2 may be assigned to S'0,0, the surface index #3 may be assigned to S'1,0, the surface index #1 may be assigned to S'2,0, the surface index #5 may be assigned to S'0,1, the surface index #0 may be assigned to S'1,1, and the surface index #4 may be assigned to S'2,1, as shown in section 21c. Alternatively, when the index is assigned using one piece of position information, the surface index #2 may be assigned to S[0], the surface index #3 may be assigned to S[1], the surface index #1 may be assigned to S[2], the surface index #5 may be assigned to S[3], the surface index #0 may be assigned to S[4] and the surface index #4 may be assigned to S[5]. For ease of description, in the following examples, S'0,0 to S'2,1 may be referred to as a to f. Alternatively, each surface may be represented using position information indicating the width and height of a pixel or block unit based on the upper left corner of the image.

[0535] For a packed image acquired by an image reconstruction process (or a region packing process), the plane scanning order is the same as or different from the image scanning order, depending on the reconstruction setting. For example, when a scanning order (e.g., raster scanning) is applied to the image shown in part 21a, a, b, and c may have the same scanning order, and d, e, and f may have different scanning orders. For example, when the scanning order of part 21a or the scanning order of a, b, and c follows the order of (0, 0) -> (1, 0) -> (0, 1) -> (1, 1), the scanning order of d, e, and f may follow the order of (1, 0) -> (1, 1) -> (0, 0) -> (0, 1). This can be determined according to the image reconstruction setting, and such a setting can be applied even to other projection formats.

[0536] In the image division process shown in part 21b, tiles can be individually set as face units. For example, each of faces a to f can be set as a tile unit. Alternatively, multiple face units can be set as tiles. For example, faces a to c can be set as one tile, and faces d to f can be set as one tile. Construction can be determined based on face characteristics (e.g., continuity between faces, etc.), and different tile settings of faces may be possible, unlike the above example.

[0537] The following is an example of division information processed according to a plurality of image divisions. In this example, it is assumed that division information of a face is omitted, units other than a face are tiles, and the division information is processed differently.

[0538] As a first example, the image division information may be acquired based on the surface information, and the image division information may be implicitly omitted. For example, a surface may be individually set as a tile, or a plurality of surfaces may be set as tiles. In this case, when at least one surface is set as a tile, this may be determined according to a predetermined rule based on the surface information (e.g., continuity or correlation).

[0539] As a second example, image division information may be explicitly generated independently of the face information. For example, when division information is generated using the number of columns (here num_tile_columns) and the number of rows (here num_tile_rows) of tiles, the division information may be generated in the above-mentioned method of image division processing. For example, the number of columns of tiles may be in the range from 0 to the width of the image or the width of the block (here the unit obtained from the picture division unit), and the number of rows of tiles may be in the range from 0 to the height of the image or the height of the block. In addition, additional division information (e.g., uniform_spacing_flag) may be generated. In this case, depending on the division setting, the boundaries of the face and the boundaries of the division unit may match or not match each other.

[0540] As a third example, image division information may be explicitly generated based on face information. For example, when division information is generated using the number of columns and rows of a tile, the division information may be generated based on face information (here, the number of columns is in the range from 0 to 2, and the number of rows is in the range from 0 to 1; because the configuration of the face in the image is 3x2). For example, the number of columns of a tile may be in the range from 0 to 2, and the number of rows of a tile may be in the range from 0 to 1. In addition, additional division information (e.g., uniform_spacing_flag) may not be generated. In this case, the boundary of the face and the boundary of the division unit may match each other.

[0541] In some cases (assuming the second example and the third example), the syntax elements of the partition information may be defined differently, or even if the same syntax element is used, the syntax element settings (e.g., binarization settings; when the range of the candidate group of the syntax element is limited and small, other binarization may be used) may be applied differently. The above examples have been described for some of the various elements of the partition information. However, the present invention is not limited thereto, and it is understood that other settings are possible depending on whether the partition information is generated based on the face information.

[0542] Fig. 22 An example of an image that is divided into tiles by CMP packing or projection.

[0543] In this case, assuming that Fig.21 The same tile partition boundaries as those shown in part 21a of FIG. 2 (W0 to W2, H0 and H1 are all activated) and have the same Fig.21 The same tile partition boundaries as those shown in part 21b of (W0, W1, and H0 are all activated). When it is assumed that area P indicates the entire image and area V indicates the viewport, full decoding or partial decoding can be performed. This example will be described focusing on partial decoding. In part 22a, tiles e, f, and g can be decoded for CMP (left side), and tiles a, c, and e can be decoded for CMP compact (right side) to obtain an area corresponding to the viewport. In part 22b, tiles b, f, and i can be decoded for CMP, and tiles d, e, and f can be decoded for CMP compact to obtain an area corresponding to the viewport.

[0544] The above example has been described for the case where the division of slices, tiles, etc. is performed based on surface units (or surface boundaries). Fig. 20 As shown in part 20a of , the division may be performed inside a surface (eg, an image consists of one surface under ERP and consists of multiple surfaces under other projection formats), or the division may be performed on the boundary of a surface as well as inside the surface.

[0545] Fig.23 360 degrees according to an embodiment of the present invention. In this case, it is assumed that the image is projected by ERP. In addition, the following example will be described focusing on the case of expansion.

[0546] The size of the projected image may be adjusted by a scale factor or by an offset factor according to the image resizing type. Here, the image before resizing may be P_Width×P_Height, and the image after resizing may be P'_Width×P'_Height.

[0547] For the scale factor, after adjusting the width and height of the image by the scale factor (here, a in terms of width and b in terms of height), the width (P_Width×a) and height (P_Height×b) of the image may be acquired. For the offset factor, after adjusting the width and height of the image by the offset factor (here, L and R in terms of width and T and B in terms of height), the width (P_Width+L+R) and height (P_Height+T+B) of the image may be acquired. The resizing may be performed using a predetermined method, or may be performed using one method selected from among a plurality of methods.

[0548] The data processing method in the following example will be described focusing on the case of the offset factor. For the offset factor, as the data processing method, there may be a filling method by using a predetermined pixel value, a filling method by copying external pixels, a filling method by copying a specific area of ​​an image, a filling method by transforming a specific area of ​​an image, and the like.

[0549] The size of a 360-degree image may be adjusted in consideration of a characteristic that there is continuity at the boundary of the image. For ERP, there is no outer boundary in the 3D space, but there may be an outer boundary when the 3D space is transformed into a 2D space through a projection process. The data in the boundary area includes data having outward continuity, but may have a boundary in terms of spatial characteristics. Size adjustment may be performed in consideration of these characteristics. In this case, continuity may be checked according to a projection format, etc. For example, an ERP image may be an image having a characteristic in which both end boundaries are continuous. The example will be described assuming that the left and right boundaries of the image are continuous with each other and the upper and lower boundaries of the image are continuous with each other. The data processing method will be described focusing on a filling method by copying a specific area of ​​an image and a filling method by transforming a specific area of ​​an image.

[0550] When the image is resized to the left, the resized area (here LC or TL+LC+BL) may be filled with data of the right area of ​​the image (here tr+rc+br) having continuity with the left part of the image. When the image is resized to the right, the resized area (here RC or TR+RC+BR) may be filled with data of the left area of ​​the image (here tl+lc+bl) having continuity with the right part of the image. When the image is resized upward, the resized area (here TC or TL+TC+TR) may be filled with data of the lower area of ​​the image (here bl+bc+br) having continuity with the upper side. When the image is resized downward, the resized area (here BC or BL+BC+BR) may be filled with data.

[0551] When the size or length of the resized area is m, the coordinates of the resized area relative to the image before resizing (here, x is in the range from 0 to P_Width-1) can have a range from (-m, y) to (-1, y) (resize to the left) or a range from (P_Width, y) to (P_Width+m-1, y) (resize to the right). The position (x') of the area for acquiring data of the resized area can be derived from the formula x'=(x+P_Width)%P_Width. In this case, x represents the coordinates of the resized area relative to the coordinates of the image before resizing, and x' represents the coordinates of the area referring to the resized area relative to the coordinates of the image before resizing. For example, when the image size is adjusted to the left, m is 4, and the image width is 16, the corresponding data (-4, y) can be obtained from (12, y), the corresponding data (-3, y) can be obtained from (13, y), the corresponding data (-2, y) can be obtained from (14, y), and the corresponding data (-1, y) can be obtained from (15, y). Alternatively, when the image size is adjusted to the right, m ​​is 4, and the image width is 16, the corresponding data (16, y) can be obtained from (0, y), the corresponding data (17, y) can be obtained from (1, y), the corresponding data (18, y) can be obtained from (2, y), and the corresponding data (19, y) can be obtained from (3, y).

[0552] When the size or length of the resized area is n, the coordinates of the resized area relative to the image before resizing (here, y is in the range from 0 to P_Height-1) can have a range from (x, -n) to (x, -1) (upward resizing) or a range from (x, P_Height) to (x, P_Height+n-1) (downward resizing). The position (y') of the area for obtaining data of the resized area can be derived from the formula y'=(y+P_Height)%P_Height. In this case, y represents the coordinates of the resized area relative to the coordinates of the image before resizing, and y' represents the coordinates of the area referring to the resized area relative to the coordinates of the image before resizing. For example, when the image size is adjusted upward, n is 4 and the height of the image is 16, the corresponding data (x, -4) can be obtained from (x, 12), the corresponding data (x, -3) can be obtained from (x, 13), the corresponding data (x, -2) can be obtained from (x, 14), and the corresponding data (x, -1) can be obtained from (x, 15). Alternatively, when the image size is resized downward, n is 4 and the height of the image is 16, the corresponding data (x, 16) can be obtained from (x, 0), the corresponding data (x, 17) can be obtained from (x, 1), the corresponding data (x, 18) can be obtained from (x, 2), and the corresponding data (x, 19) can be obtained from (x, 3).

[0553] After filling the resized area with data, resizing can be performed relative to the coordinates of the resized image (here, x ranges from 0 to P'_Width-1, and y ranges from 0 to P'_Height-1). This example can be applied to a coordinate system of latitude and longitude.

[0554] Various size adjustment combinations can be provided as follows.

[0555] As an example, the image may be resized to the left by m. Alternatively, the image may be resized to the right by n. Alternatively, the image may be resized upward by o. Alternatively, the image may be resized downward by p.

[0556] As an example, the image may be resized to the left by m and to the right by n. Alternatively, the image may be resized upward by o and downward by p.

[0557] As an example, the image may be resized to the left by m, to the right by n, and upward by o. Alternatively, the image may be resized to the left by m, to the right by n, and downward by p. Alternatively, the image may be resized to the left by m, upward by o, and downward by p. Alternatively, the image may be resized to the right by n, upward by o, and downward by p.

[0558] As an example, the image may be resized left by m, right by n, upward by o, and downward by p.

[0559] Similar to the above example, at least one resizing operation may be performed. Image resizing may be implicitly performed according to the encoding / decoding setting, or resizing information may be implicitly generated, and then image resizing may be performed based on the generated resizing information. That is, m, n, o, and p of the above example may be determined as predetermined values, or may be explicitly generated using resizing information. Alternatively, some of m, n, o, and p may be determined as predetermined values, and other values ​​may be explicitly generated.

[0560] The above examples have been described focusing on the case of acquiring data from a specific area of ​​an image, but other methods may also be applied. The data may be pixels before encoding or pixels after encoding, and may be determined according to the characteristics of the resizing step or the image to be resized. For example, when resizing is performed in a preprocessing process and a precoding step, the data may refer to input pixels of a projected image, a packed image, etc., and when resizing is performed in a postprocessing process, an intra-frame prediction reference pixel generation step, a reference picture generation step, a filtering step, etc., the data may refer to restored pixels. In addition, resizing may be performed by using a data processing method separately in each resized area.

[0561] Fig.24 is a conceptual diagram illustrating continuity between surfaces in a projection format (eg, CHP, OHP, or ISP) according to an embodiment of the present invention.

[0562] In detail, Fig.24 An example of an image composed of multiple faces can be shown. Continuity can be a characteristic generated in adjacent areas in a 3D space. Parts 24a to 24c variously show a case (A) with spatial adjacency and continuity, a case (B) with spatial adjacency but no continuity, a case (C) with no spatial adjacency but with continuity, and a case (D) with no spatial adjacency and no continuity when transformed into a 2D space by a projection process. Different from this, a general image is classified into a case (A) with both spatial adjacency and continuity and a case (D) with no spatial adjacency and no continuity. In this case, the case with continuity corresponds to some examples (A or C).

[0563] That is, referring to parts 24a to 24c, the case of having both spatial adjacency and continuity (here, described with reference to part 24a) can be shown as b0 to b4, and the case of having no spatial adjacency but having continuity can be shown as B0 to B6. That is, these cases indicate areas that are adjacent in 3D space, and the encoding performance can be enhanced by using the characteristic that b0 to b4 and B0 to B6 have continuity in the encoding process.

[0564] Fig.25 21c is a conceptual diagram showing the surface continuity in the portion 21c which is an image acquired by the image reconstruction process or the area packing process in the CMP projection format.

[0565] Here, Fig.21 Part 21c shows a rearrangement of the 360 ​​degree image in part 21a unfolded in the shape of a cube, thus maintaining the Fig.21 That is, as shown in part 25a, surface S2,1 can be horizontally continuous with surfaces S1,1 and S3,1, and can be vertically continuous with surface S1,0 rotated 90 degrees and surface S1,2 rotated -90 degrees.

[0566] In the same way, the continuity of the surfaces S3,1, S0,1, S1,2, S1,1 and S1,0 can be checked in the parts 25b to 25f.

[0567] The continuity between the surfaces can be defined according to the projection format setting, etc. However, the present invention is not limited thereto and can be modified. Fig.24 and Fig.25 The following example is described in the case where there is continuity as shown.

[0568] Fig.26 is an exemplary diagram showing image size adjustment in the CMP projection format according to an embodiment of the present invention.

[0569] Part 26a shows an example of resizing an image, part 26b shows an example of resizing a plane unit (or a division unit), and part 26c shows an example of resizing an image and a plane unit (or an example of performing resizing multiple times).

[0570] The size of the projected image can be adjusted by a scale factor or by an offset factor depending on the image resizing type. Here, the image before resizing can be P_Width×P_Height, the image after resizing can be P'_Width×P'_Height, and the size of the face can be F_Width×F_Height. The size can be the same or different depending on the face, and the width and height can be the same or different depending on the face. However, for ease of description, the example will be described assuming that all faces in the image have the same size and square shape. In addition, the description assumes that the resized values ​​(here WX and HY) are the same. In the following examples, the data processing method will be described focusing on the case of the offset factor and also focusing on a filling method by copying a specific area of ​​the image and a filling method by transforming a specific area of ​​the image. The above settings can even be applied to Fig. 27 The situation shown in .

[0571] For the portions 26a to 26c, the boundaries of the faces may be continuous with the boundaries of the other faces (here, it is assumed that they have a corresponding Fig.24 Here, continuity can be classified into a case where there is spatial adjacency and image continuity in a 2D plane (a first example) and a case where there is no spatial adjacency but image continuity in a 2D plane (a second example).

[0572] For example, when assuming Fig.24 When the continuity in part 24a is maintained, the upper, left, right and lower regions of S1,1 can be spatially adjacent to the lower, right, left and upper regions of S1,0, S0,1, S2,1 and S1,2 and have image continuity with the lower, right, left and upper regions of S1,0, S0,1, S2,1 and S1,2 (first example).

[0573] Alternatively, the left and right regions of S1,0 are not spatially adjacent to the upper regions of S0,1 and S2,1, but may have image continuity with the upper regions of S0,1 and S2,1 (second example). In addition, the left regions of S0,1 may not be spatially adjacent to each other, but may have image continuity with each other (second example). In addition, the left and right regions of S1,2 may be continuous with the lower regions of S0,1 and S2,1 (second example). This may be only a limited example, and other configurations may be applied depending on the definition and setting of the projection format. For ease of description, S0,0 to S3,2 in portion 26a are referred to as a to l.

[0574] Portion 26a may be an example of a filling method using data of an area having continuity toward the outer boundary of the image. The range from area A not including data to the resized area (here a0 to a2, c0, d0 to d2, i0 to i2, k0, and l0 to l2) may be filled with any predetermined value or filled by external pixel padding, and the range from area B including actual data to the resized area (here b0, e0, h0, and j0) may be filled with data of an area (or face) having image continuity. For example, b0 may be filled with data of the upper side of face h, e0 may be filled with data of the right side of face h, h0 may be filled with data of the left side of face e, and j0 may be filled with data of the lower side of face h.

[0575] In detail, as an example, b0 may be filled with data of the lower side of the surface obtained by rotating the surface h by 180 degrees, and j0 may be filled with data of the upper side of the surface obtained by rotating the surface h by 180 degrees. However, this example (including the following examples) may only represent the position of the reference surface, and may be considered as follows. Fig.24 and Fig.25 The continuity between the faces shown is obtained after resizing processing (eg, rotation, etc.) and data obtained from the resized area.

[0576] Part 26b may be an example of a filling method using data of an area having continuity toward the inner boundary of the image. In this example, different resizing operations may be performed for each face. A reduction process may be performed in region A, and an expansion process may be performed in region B. For example, the size of face a may be adjusted to the right (reduced here) by w0, and the size of face b may be adjusted to the left (expanded here) by w0. Alternatively, the size of face a may be adjusted downward (reduced here) by h0, and the size of face e may be adjusted upward (expanded here) by h0. In this example, when the width change of the image is observed through faces a, b, c, and d, face a is reduced by w0, face b is expanded by w0 and w1, and face c may be reduced by w1. Therefore, the width of the image before resizing is the same as the width of the image after resizing. When the height change of the image is observed through faces a, e, and i, face a is reduced by h0, face e is expanded by h0 and h1, and face i may be reduced by h1. Therefore, the height of the image before resizing is the same as the height of the image after resizing.

[0577] Considering that the region is reduced from region A which does not include data, the resized regions (here b0, e0, be, b1, bg, g0, h0, e1, ej, j0, gi, g1, j1 and h1) can be simply removed, and considering that the region is expanded from region B which includes actual data, the resized regions can be filled with data from regions with continuity.

[0578] For example, b0 can be filled with data from the upper side of face e; e0 can be filled with data from the left side of face b; be can be filled with data from the left side of face b, the upper side of face e, or the weighted sum of the left side and the upper side of face e; b1 can be filled with data from the upper side of face g; bg can be filled with data from the left side of face b, the upper side of face g, or the weighted sum of the right side and the upper side of face g; g0 can be filled with data from the right side of face b; h0 can be filled with data from the upper side of face b; e1 can be filled with data from the left side of face j; ej can be filled with data from the lower side of face e, the left side of face j, or the weighted sum of the lower side of face e and the left side of face j; j0 can be filled with data from the lower side of face e; gj can be filled with data from the lower side of face g, the left side of face j, or the weighted sum of the lower side of face g and the right side of face j; g1 can be filled with data from the right side of face j; j1 can be filled with data from the lower side of face g; and h1 can be filled with data from the lower side of face j.

[0579] In the above example, when the resized area is filled with data of a specific area of ​​the image, the data of the corresponding area may be copied and then used to fill the resized area, or may be transformed based on the characteristics, type, etc. of the image, and then used to fill the resized area. For example, when a 360-degree image can be transformed into a 2D space according to a projection format, a coordinate system (e.g., a 2D plane coordinate system) may be defined for each face. For ease of description, it is assumed that (x, y, z) in the 3D space is transformed into (x, y, C), (x, C, z) or (C, y, z) for each face. The above example indicates the following situation: data of faces other than the corresponding face are obtained from the resized area of ​​the face. That is, when resizing is performed on the current face, data of other faces with different coordinate system characteristics may be copied as is and then used. In this case, there is a possibility that continuity is distorted based on the resizing boundary. To this end, data of other faces acquired according to the coordinate system characteristics of the previous face may be transformed and used to fill the resized area. This transformation is also only an example of a data processing method, and the present invention is not limited thereto.

[0580] When data of a specific area of ​​the image is copied and used to fill the resized area, distorted continuity (or fundamentally changed continuity) may be included in the boundary area between the resized area (e) and the resized area (e0). For example, the continuity may be changed relative to the boundary, and a straight edge may be curved relative to the boundary.

[0581] When data of a particular region of an image is transformed and used to fill the resized regions, a gradually changing continuity may be included in the boundary region between the resized regions.

[0582] The above examples may be examples of a data processing method of the present invention to transform data of a specific region of an image based on characteristics, type, etc. of the image and fill a resized region with the transformed data.

[0583] Portion 26c may be an example of filling the resized area with data of an area having continuity toward the boundaries (inner and outer boundaries) of the image in combination with the image resizing processing corresponding to portions 26a and 26b. The resizing processing of this example may be derived from the resizing processing of portions 26a and 26b, and its detailed description will be omitted.

[0584] Portion 26a may be an example of a process of resizing an image, and portion 26b may be an example of resizing a division unit in an image. Portion 26c may be an example of a plurality of resizing processes including a process of resizing an image and a process of resizing a division unit in an image.

[0585] For example, the size of the image (here, the first format) acquired by the projection process can be adjusted (here, region C), and the size of the image (here, the second format) acquired by the format conversion process can be adjusted (here, region D). In this example, the size of the image projected by the ERP (here, the full image) can be adjusted and the image is converted into an image projected by the CMP by the format conversion unit, and the size of the image projected by the CMP (here, the surface unit) can be adjusted. The above example is an example of performing multiple size adjustment operations. However, the present invention is not limited to this, and modifications can be made thereto.

[0586] Fig. 27 is an example diagram showing resizing of an image converted and packed in a CMP projection format according to an embodiment of the present invention. Fig. 27 Also assume that Fig.25 Continuity between faces is shown, so the boundaries of a face can be continuous with the boundaries of another face.

[0587] In this example, the offset factors of W0 to W5 and H0 to H3 may have various values ​​(here, it is assumed that the offset factors are used as resizing values). For example, the offset factors may be derived from a predetermined value, a motion search range of inter prediction, a unit obtained from a picture division unit, etc., and other cases are also possible. In this case, the unit obtained from the pixel division unit may include a face. That is, the resizing value may be determined based on F_Width and F_Height.

[0588] Part 27a is an example of individually adjusting the size of a single face (here, upward, downward, left, and right relative to the face) and filling the extended area with data of the area having continuity. For example, the outer areas a0 to a6 of face a can be filled with continuous data, and the outer areas b0 to b6 of face b can be filled with continuous data.

[0589] Part 27b is an example of adjusting the size of multiple faces (here, upward, downward, left and right relative to the multiple faces) and filling the extended area with data of the area with continuity. For example, faces a, b and c can be extended to the outer areas a0 to a4, b0 and b1, and c0 to c4.

[0590] Portion 27c may be an example of adjusting the size of the full image (here, upward, downward, left, and right relative to the full image) and filling the extended area with data of the area having continuity. For example, the full image consisting of faces a to f may be extended to the outer areas a0 to a2, b0, c0 to c2, d0 to d2, and f0 to f2.

[0591] That is, resizing may be performed in a single surface unit, in a plurality of surface units having continuity with each other, and in a full image unit.

[0592] In the above example, the resized area (here, a0 to f7) can be filled with data of the area (or face) having continuity, as shown in part 24a. That is, the resized area can be filled with data of the upper side, lower side, left side and right side of faces a to f.

[0593] Fig.28 is an exemplary diagram illustrating a data processing method for adjusting the size of a 360-degree image according to an embodiment of the present invention.

[0594] Reference Fig.28 , area B (a0 to a2, ad0, b0, c0 to c2, cf1, d0 to d2, e0, f0 to f2) as a resized area can be filled with data of an area having continuity between pixel data belonging to a to f. In addition, area C (ad1, be, cf0) as another resized area can be filled in combination with data of an area to be resized and data of an area having spatial adjacency but no continuity. Alternatively, since resizing is performed between two areas selected from a to f (e.g., a and d, b and e, and c and f), area C can be filled in combination with multiple data of the two areas. For example, face b and face e can be adjacent to each other in space, but not continuous with each other. The size of the resized area located between face b and face e can be adjusted using data of face b and data of face e. For example, the area can be filled with a value obtained by averaging the data of face b and data of face e or with a value obtained by a weighted sum based on distance. In this case, the pixels of data used to fill the resized areas in the planes b and e may be boundary pixels of each plane or internal pixels of each plane.

[0595] In summary, the resized areas between the divided units of the image can be filled with data generated by using a plurality of data of two units in combination.

[0596] The data processing method may be supported under some conditions (here, when multiple regions are resized).

[0597] In parts 27a and 27b, the area to be resized between the division units is constructed separately for each division unit (in part 27a, a6 and d1 are constructed for a and d, respectively). Fig.28 In the example, a single area to be resized between the division units can be constructed for adjacent division units (ad1 for a and d). It should be understood that this method can be included in the candidate group of data processing methods in parts 27a and 27b, and even in Fig.28 In the example above, the resizing may be performed using a data processing method different from that in the above example.

[0598] In the process of adjusting the image size according to the present invention, a predetermined data processing method may be implicitly used in the resized area, or one of a plurality of data processing methods may be used for explicit related information. The predetermined data processing method may be one of a filling method using any pixel value, a filling method by copying external pixels, a filling method by copying a specific area of ​​an image, a filling method by transforming a specific area of ​​an image, a filling method using data derived from multiple areas of an image, and the like. For example, when the resized area is located inside an image (e.g., a packed image) and the areas (e.g., faces) on both sides have spatial adjacency but no continuity, a data processing method may be used to fill the resized area with data derived from multiple areas. In addition, resizing may be performed by a data processing method selected from a plurality of data processing methods, and related selection information may be explicitly generated. This may be an example applicable to general images as well as 360-degree images.

[0599] The encoder can add the information generated during the above processing to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the relevant information from the bitstream. In addition, the information can be included in the bitstream in the form of SEI or metadata. The division process, reconstruction process and size adjustment process for 360-degree images have been described focusing on some projection formats such as ERP and CMP. However, the present invention is not limited to this, and the above description can even be applied to other projection formats as is or after modification.

[0600] It has been described that the image setting process for the above-mentioned 360-degree image encoding / decoding device can be applied to the pre-processing process, the post-processing process, the format conversion process, the inverse format conversion process, etc. as well as the encoding / decoding process.

[0601] In summary, the projection process can be constructed to include an image setting process. In detail, in addition to at least one of the image setting processes, the projection process can also be performed. Division can be performed based on the projection image in units of regions (or faces). Depending on the projection format, division can be performed on a single region or multiple regions. For the division, division information can be generated. In addition, the size of the projection image can be adjusted, or the size of the projection area can be adjusted. In this case, size adjustment can be performed on at least one region. For size adjustment, size adjustment information can be generated. In addition, the projection image can be reconstructed (or face arrangement), or the projection area can be reconstructed. In this case, reconstruction can be performed on at least one region. For reconstruction, reconstruction information can be generated.

[0602] In summary, the regional packing process can be constructed to include the image setting process. In detail, in addition to at least one of the image setting processes, the regional packing projection process can also be performed. The divisio...

Claims

1. A method for decoding a 360-degree image, the method comprising: Receiving a bit stream encoded with the 360-degree image, the bit stream including data of an extended two-dimensional image, the extended two-dimensional image including a two-dimensional image and a predetermined extended area, and the two-dimensional image is projected from an image having a three-dimensional projection structure and includes one or more surfaces; generating a predicted image by performing prediction based on information about prediction included in the bitstream; as well as reconstructing the extended two-dimensional image based on the predicted image and the residual image, wherein the size of the extension region is determined based on one or more syntax elements obtained from the bitstream, wherein the number of the syntax elements is determined differently based on a projection format used for the three-dimensional projection structure, the projection format being one of a plurality of projection formats including an ERP format in which the 360-degree image is projected in a two-dimensional plane and a CMP format in which the 360-degree image is projected in a cube, wherein the predicted image is added to the residual image to reconstruct the extended two-dimensional image, wherein the residual image is obtained by decoding information about the residual image included in the bit stream, and The sample value of the extended area is determined by horizontally copying the sample value of the surface to the sample value of the extended area.

2. A method for encoding a 360-degree image, the method comprising: obtaining a two-dimensional image projected from an image having a three-dimensional projection structure and including at least one surface; obtaining an extended two-dimensional image including the two-dimensional image and a predetermined extended area; generating a predicted image by performing a prediction, information about the prediction being encoded into a bitstream; as well as encoding data of the extended two-dimensional image into the bitstream based on the predicted image and the residual image, The size of the extension area is encoded based on one or more syntax elements, and the one or more syntax elements are encoded into the bitstream, wherein the number of the syntax elements is determined differently based on a projection format used for the three-dimensional projection structure, the projection format being one of a plurality of projection formats including an ERP format in which the 360-degree image is projected in a two-dimensional plane and a CMP format in which the 360-degree image is projected in a cube, wherein the residual image is obtained based on the extended two-dimensional image and the predicted image, wherein the residual image is encoded by encoding information about the residual image into the bitstream, and The sample value of the extended area is determined by horizontally copying the sample value of the surface to the sample value of the extended area.

3. A non-transitory computer-readable recording medium storing a bitstream generated by a method for encoding a 360-degree image, the method comprising: obtaining a two-dimensional image projected from an image having a three-dimensional projection structure and including at least one surface; obtaining an extended two-dimensional image including the two-dimensional image and a predetermined extended area; generating a predicted image by performing a prediction, information about the prediction being encoded into the bitstream; as well as encoding data of the extended two-dimensional image into the bitstream based on the predicted image and the residual image, The size of the extension area is encoded based on one or more syntax elements, and the one or more syntax elements are encoded into the bitstream, wherein the number of the syntax elements is determined differently based on a projection format used for the three-dimensional projection structure, the projection format being one of a plurality of projection formats including an ERP format in which the 360-degree image is projected in a two-dimensional plane and a CMP format in which the 360-degree image is projected in a cube, wherein the residual image is obtained based on the extended two-dimensional image and the predicted image, wherein the residual image is encoded by encoding information about the residual image into the bitstream, and The sample value of the extended area is determined by horizontally copying the sample value of the surface to the sample value of the extended area.

4. A method for transmitting a bitstream, the bitstream being generated by a method for encoding a 360-degree image, the method for encoding a 360-degree image comprising: obtaining a two-dimensional image projected from an image having a three-dimensional projection structure and including at least one surface; obtaining an extended two-dimensional image including the two-dimensional image and a predetermined extended area; generating a predicted image by performing a prediction, information about the prediction being encoded into the bitstream; as well as encoding data of the extended two-dimensional image into the bitstream based on the predicted image and the residual image, The size of the extension area is encoded based on one or more syntax elements, and the one or more syntax elements are encoded into the bitstream, wherein the number of the syntax elements is determined differently based on a projection format used for the three-dimensional projection structure, the projection format being one of a plurality of projection formats including an ERP format in which the 360-degree image is projected in a two-dimensional plane and a CMP format in which the 360-degree image is projected in a cube, wherein the residual image is obtained based on the extended two-dimensional image and the predicted image, wherein the residual image is encoded by encoding information about the residual image into the bitstream, and The sample value of the extended area is determined by horizontally copying the sample value of the surface to the sample value of the extended area.