Method of processing image and method of transmitting bitstream
By receiving and reconstructing a 360-degree image bitstream, generating a predicted image, and combining it with a residual image obtained through inverse quantization and inverse transform, the problem of insufficient performance in existing image processing systems is solved, achieving efficient compression and decoding of 360-degree images.
Patent Information
- Application Number
- CN202511302499.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2017-07-17
- Filing Date
- 2017-10-10
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies have limitations in image processing systems when dealing with multi-view images, especially in encoding and decoding, where they struggle to effectively handle large amounts of 360-degree image data.
An image decoding method is provided, comprising receiving an encoded 360-degree image bitstream, generating a predicted image and combining it with a residual image of inverse quantization and inverse transform, reconstructing the image according to the projection format, and performing regional packing and rearrangement using grammatical information.
It improves the compression performance of 360-degree images and enhances the efficiency of image encoding and decoding.
Smart Images

Figure CN121099028A_ABST
Abstract
Description
[0001] This application is a divisional application of application No. 201780073662.3, filed on October 10, 2017, entered into the national phase on May 28, 2019, and having the title of "Image Data Encoding / Decoding Method and Apparatus". TECHNICAL FIELD
[0002] The present application relates to an image data encoding and decoding technology, and more particularly, to a method and apparatus for encoding and decoding a 360-degree image for a reality media service. BACKGROUND
[0003] With the popularization of the Internet and mobile terminals and the development of information and communication technology, the use of multimedia data is rapidly increasing. Recently, there is a demand for high-quality images and high-resolution images such as high-definition (HD) images and ultra-high-definition (UHD) images in various fields, and the demand for reality media services such as virtual reality, augmented reality, etc. is also rapidly increasing. Specifically, since multi-view images captured with a plurality of cameras are processed for 360-degree images for virtual reality and augmented reality, the amount of data generated for processing is greatly increased, but the performance of the image processing system for processing a large amount of data is insufficient.
[0004] As described above, in the image encoding and decoding method and apparatus of the related art, there is a need to improve the performance of image processing, particularly the performance of image encoding / decoding. SUMMARY
[0005] TECHNICAL PROBLEM
[0006] An object of the present application is to provide a method for improving image setting processing in an initial step for encoding and decoding. More specifically, the present application aims to provide an encoding and decoding method and apparatus for improving image setting processing in consideration of the characteristics of a 360-degree image.
[0007] TECHNICAL SOLUTION
[0008] According to an aspect of the present application, there is provided a method of decoding a 360-degree image.
[0009] Here, the method of decoding a 360-degree image can include receiving a bitstream including an encoded 360-degree image, generating a prediction image with reference to syntax information acquired from the received bitstream, acquiring a decoded image by combining the generated prediction image with a residual image acquired by inverse quantizing and inverse transforming the bitstream, and reconstructing the decoded image into a 360-degree image according to a projection format.
[0010] Here, the syntax information can include projection format information of the 360-degree image.
[0011] Here, the projection format information can be information indicating at least one of an Equi-Rectangular Projection (ERP) format in which the 360-degree image is projected into a 2D plane, a CubeMap Projection (CMP) format in which the 360-degree image projection is projected into a cube, an OctaHedron Projection (OHP) format in which the 360-degree image is projected into an octahedron, and an IcoSahedral Projection (ISP) format in which the 360-degree image is projected into a polyhedron.
[0012] Here, the reconstructing can include acquiring arrangement information according to the region-wise packing with reference to the syntax information, and rearranging the blocks of the decoded image according to the arrangement information.
[0013] Here, the generating of the prediction image can include performing image extension on a reference picture acquired by restoring the bitstream, and generating the prediction image with reference to the reference picture on which the image extension is performed.
[0014] Here, the performing of the image extension can include performing the image extension based on a division unit of the reference picture.
[0015] Here, the performing of the image extension based on the division unit can include generating an extension region for each division unit individually by using a reference pixel of the division unit.
[0016] Here, the extension region can be generated using a boundary pixel of a division unit that is spatially adjacent to the division unit to be extended or using a boundary pixel of a division unit that has image continuity with the division unit to be extended.
[0017] Here, the performing of the image extension based on the division unit can include generating an extension image of a combined region using a boundary pixel of the combined region that combines two or more division units that are spatially adjacent to each other in the division unit.
[0018] Here, the performing of the image extension based on the division unit can include generating an extension region between adjacent division units using all adjacent pixel information of the division units that are spatially adjacent to each other in the division unit.
[0019] Here, the performing of the image extension based on the division unit can include generating an extension region using an average value of adjacent pixels of the division units that are spatially adjacent to each other.
[0020] Advantages of the present application
[0021] With the image encoding / decoding method and apparatus according to the embodiments of the present application, compression performance can be enhanced. In particular, for a 360-degree image, compression performance can be enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a block diagram of an image encoding apparatus according to an embodiment of the present application.
[0023] Figure 2 is a block diagram of an image decoding apparatus according to an embodiment of the present application.
[0024] Figure 3 is an example diagram illustrating division of image information into a plurality of layers to compress an image.
[0025] Figure 4 is a conceptual diagram illustrating an example of image division according to an embodiment of the present application.
[0026] Figure 5 is another example diagram of an image division method according to an embodiment of the present application.
[0027] Figure 6 is an example diagram of a general image size adjustment method.
[0028] Figure 7 is an example diagram of image size adjustment according to an embodiment of the present application.
[0029] Figure 8 is an example diagram of a method of constructing a region generated by expansion in an image size adjustment method according to an embodiment of the present application.
[0030] Figure 9 is an example diagram of a method of constructing a region to be deleted and a region to be generated in an image size adjustment method according to an embodiment of the present application.
[0031] Figure 10 is an example diagram of image reconstruction according to an embodiment of the present application.
[0032] Figure 11 is an example diagram illustrating an image before and after image setting processing according to an embodiment of the present application.
[0033] Figure 12 is an example diagram of adjusting the size of each division unit of an image according to an embodiment of the present application.
[0034] Figure 13 is an example diagram of a group of setting or size adjustment of division units in an image.
[0035] Figure 14is an example diagram showing both a process of adjusting the size of an image and a process of adjusting the size of a division unit in the image.
[0036] Figure 15 is an example diagram showing a three-dimensional (3D) space and a two-dimensional (2D) plane space in which a 3D image is displayed.
[0037] Figures 16a to 16d is a conceptual diagram showing a projection format according to an embodiment of the present application.
[0038] Figure 17 is a conceptual diagram showing that a projection format according to an embodiment of the present application is included in a rectangular image.
[0039] Figure 18 is a conceptual diagram of a method of converting a projection format into a rectangular shape, i.e., a method of performing rearrangement on a surface to exclude meaningless areas, according to an embodiment of the present application.
[0040] Figure 19 is a conceptual diagram showing that a region-wise packing process is performed to convert a CMP projection format into a rectangular image, according to an embodiment of the present application.
[0041] Figure 20 is a conceptual diagram of 360-degree image division according to an embodiment of the present application.
[0042] Figure 21 is an example diagram of 360-degree image division and image reconstruction according to an embodiment of the present application.
[0043] Figure 22 is an example diagram in which an image packed or projected by a CMP is divided into tiles.
[0044] Figure 23 is a conceptual diagram showing an example of adjusting the size of a 360-degree image according to an embodiment of the present application.
[0045] Figure 24 is a conceptual diagram showing continuity between surfaces under a projection format (e.g., CHP, OHP, or ISP) according to an embodiment of the present application.
[0046] Figure 25 is a conceptual diagram showing continuity of surfaces of a portion 21c, which is an image obtained through an image reconstruction process or a region-wise packing process of a CMP projection format.
[0047] Figure 26 is an example diagram showing image size adjustment under a CMP projection format according to an embodiment of the present application.
[0048] Figure 27 is an example diagram illustrating resizing of images converted and packed in a CMP projection format according to an embodiment of the present application.
[0049] Figure 28 is an example diagram illustrating a data processing method for resizing a 360-degree image according to an embodiment of the present application.
[0050] Figure 29 is an example diagram illustrating a tree-based block form.
[0051] Figure 30 is an example diagram illustrating a type-based block form.
[0052] Figure 31 is an example diagram illustrating various types of blocks that can be obtained by a block dividing section of the present application.
[0053] Figure 32 is an example diagram illustrating tree-based partitioning according to an embodiment of the present application.
[0054] Figure 33 is an example diagram illustrating tree-based partitioning according to an embodiment of the present application. DETAILED DESCRIPTION
[0055] While the present application can be susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit the application to the particular form disclosed, but on the contrary, the application is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the application.
[0056] It should be understood that, although the terms “first,” “second,” etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the present application. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0057] It will be understood that when an element is referred to as being “connected” or “coupled” to another element, it can be directly connected or coupled to the other element or intervening elements can be present. In contrast, when an element is referred to as being “directly connected” or “directly coupled” to another element, there are no intervening elements present. Other words used to describe the relationship between elements should be interpreted in the like fashion (i.e., “between” versus “directly between,” “adjacent” versus “directly adjacent,” etc.).
[0058] The professional terms used herein are used only for the purpose of describing particular embodiments and are not intended to limit the present application. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including," when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0059] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0060] The image encoding apparatus and the image decoding apparatus can each be a user terminal such as a personal computer (PC), a laptop computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a portable game machine (PSP), a wireless communication terminal, a smart phone, and a television, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a head-mounted display (HMD) device, and smart glasses, or a server terminal such as an application server and a service server, and can include various devices having a communication device such as a communication modem for communication with various devices or a wired / wireless communication network, a memory for storing various programs and data for encoding or decoding an image or performing inter prediction or intra prediction for encoding or decoding, a processor for executing the programs to perform a calculation and a control operation, and the like. In addition, the image encoded into a bitstream by the image encoding apparatus can be transmitted to the image decoding apparatus in real time or non-real time through a wired / wireless communication network such as the Internet, a short-range wireless network, a wireless local area network (LAN), a WiBro network, a mobile communication network, or through various communication interfaces such as a cable, a universal serial bus (USB), and the like. Then, the bitstream can be decoded by the image decoding apparatus to restore and play back the image from the bitstream.
[0061] In addition, the image encoded into a bitstream by the image encoding apparatus can be transmitted to the image decoding apparatus through a computer-readable recording medium from the image encoding apparatus.
[0062] The above-described image encoding apparatus and decoding apparatus can be separate apparatuses, but can be set as one image encoding / decoding apparatus according to an implementation. In this case, some elements of the image encoding apparatus can be substantially the same as some elements of the image decoding apparatus, and can be implemented to include at least the same structure or perform the same function.
[0063] Therefore, in the detailed description of the following technical elements and their working principles, redundant descriptions of the corresponding technical elements will be omitted.
[0064] In addition, the image decoding apparatus corresponds to a computing apparatus to which an image encoding method to be performed by the image encoding apparatus is applied to a decoding process, and thus the following description will focus on the image encoding apparatus.
[0065] The computing apparatus can include a memory configured to store a program or a software pattern for implementing the image encoding method and / or the image decoding method, and a processor connected to the memory to execute the program. In addition, the image encoding apparatus can also be referred to as an encoder, and the image decoding apparatus can also be referred to as a decoder.
[0066] In general, an image can be composed of a series of still images. The still images can be classified in units of a group of pictures (GOP), and each still image can be referred to as a picture. In this case, the picture can indicate one of frames and fields in a progressive signal and an interlaced signal. When encoding / decoding is performed on a frame basis, the picture can be expressed as "frame", and when encoding / decoding is performed on a field basis, the picture can be expressed as "field". The present invention assumes a progressive signal, but can also be applied to an interlaced signal. As a higher concept, there can be units such as a GOP and a sequence, and each picture can also be divided into predetermined areas such as a slice, a tile, a block, etc. In addition, one GOP can include units such as an I picture, a P picture, and a B picture. The I picture can refer to a picture that is independently encoded / decoded without using a reference picture, and the P picture and the B picture can refer to pictures that are encoded / decoded by using a reference picture to perform processes such as motion estimation and motion compensation. In general, the P picture can use the I picture and the B picture as reference pictures, and the B picture can use the I picture and the P picture as reference pictures. However, the above definitions can be changed by a setting of encoding / decoding.
[0067] Here, a picture referred to in encoding / decoding is referred to as a reference picture, and a block or a pixel referred to in encoding / decoding is referred to as a reference block or a reference pixel. Also, the reference data can include various types of encoding / decoding information and frequency domain coefficients and spatial domain pixel values generated and determined during the encoding / decoding process. For example, the reference data can correspond to intra prediction information or motion information in the prediction section, transform information in the transform / inverse transform section, quantization information in the quantization / inverse quantization section, encoding / decoding information (context information) in the encoding / decoding section, filter information in the in-loop filter section, etc.
[0068] The minimum unit of an image can be a pixel (Pixel), and the number of bits used to represent one pixel is referred to as bit depth. Generally, the bit depth can be 8 bits, and a bit depth of 8 bits or more can be supported according to an encoding setting. At least one bit depth can be supported according to a color space. Also, at least one color space can be included according to an image color format. One or more pictures having the same size or one or more pictures having different sizes can be included according to a color format. For example, YCbCr 4:2:0 can consist of one luma component (Y in this example) and two chroma components (Cb / Cr in this example). At this time, the composition ratio of the chroma components and the luma component can be 1:2 in width and height. As another example, YCbCr 4:4:4 can have the same composition ratio in width and height. Similar to the above example, when one or more color spaces are included, the pictures can be divided into color spaces.
[0069] The present application will be described based on any color space (Y in this example) of any color format (YCbCr in this example), and the description will be applied in the same or similar manner (depending on the setting of the specific color space) to the other color spaces (Cb and Cr in this example) of the color format. However, each color space can be given a partial difference (independent of the setting of the specific color space). That is, depending on the setting of each color space can refer to a setting proportional to or dependent on the composition ratio of each component (for example, 4:2:0, 4:2:2, or 4:4:4), and independent of each color space can refer to a setting independent of or irrelevant to the composition ratio of each component, which is only the corresponding color space. In the present application, some elements can have independent settings or dependent settings depending on the encoder / decoder.
[0070] Setting information or syntax elements required during an image encoding process can be determined in a hierarchy of units such as a video, a sequence, a picture, a slice, a tile, a block, etc. These units include a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile header, and a block header. An encoder can add the units to a bitstream and transmit the bitstream to a decoder. The decoder can parse the bitstream in the same hierarchy, recover the setting information transmitted by the encoder, and use the setting information in an image decoding process. Also, related information can be transmitted through a bitstream in the form of supplemental enhancement information (SEI) or metadata, which can then be parsed and then used. Each parameter set has a unique ID value, and a lower parameter set can have an ID value of an upper parameter set to be referred to. For example, a lower parameter set can refer to information of an upper parameter set having a corresponding ID value among one or more upper parameter sets. Among the various examples of the above units, when any one unit includes one or more different units, any one unit can be referred to as an upper unit, and the included units can be referred to as lower units.
[0071] Setting information occurring in such units can include a setting independent of each unit or a setting dependent on a previous unit, a subsequent unit, or an upper unit. Here, it will be understood that a dependent setting indicates setting information of a corresponding unit using flag information corresponding to a setting of a previous unit, a subsequent unit, or an upper unit (e.g., 1-bit flag; 1 indicates follow, and 0 indicates not follow). In the present disclosure, setting information will be described focusing on an example of an independent setting. However, an example in which a relationship of setting information of a previous unit, a subsequent unit, or an upper unit dependent on a current unit is added to or replaces an independent setting can also be included.
[0072] Figure 1 is a block diagram of an image encoding apparatus according to an embodiment of the present disclosure. Figure 2 is a block diagram of an image decoding apparatus according to an embodiment of the present disclosure.
[0073] Referring to Figure 1 , the image encoding apparatus can be configured to include a prediction section, a subtracter, a transform section, a quantization section, an inverse quantization section, an inverse transform section, an adder, an in-loop filter section, a memory, and / or an encoding section, some of which can not necessarily be included. According to the implementation, some or all of these elements can be selectively included, and some additional elements not shown herein can be included.
[0074] Referring to Figure 2The image decoding apparatus can be configured to include a decoding section, a prediction section, an inverse quantization section, an inverse transform section, an adder, an in-loop filter section, and / or a memory, some of which can not necessarily be included. Depending on the implementation, some or all of these elements can be selectively included, and some additional elements not shown herein can be included.
[0075] The image encoding apparatus and the decoding apparatus can be separate apparatuses, but can be provided as one image encoding / decoding apparatus according to the implementation. In this case, some elements of the image encoding apparatus can be substantially the same as some elements of the image decoding apparatus, and can be implemented to include at least the same structure or perform the same function. Therefore, redundant descriptions of the corresponding technical elements will be omitted in the detailed description of the technical elements and their working principles below. The image decoding apparatus corresponds to a computing apparatus to which an image encoding method to be performed by the image encoding apparatus is applied to a decoding process, and thus the following description will focus on the image encoding apparatus. The image encoding apparatus can also be referred to as an encoder, and the image decoding apparatus can also be referred to as a decoder.
[0076] The prediction section can be implemented using a prediction module, and can generate a prediction block by performing intra prediction or inter prediction on a block to be encoded. The prediction section generates a prediction block by predicting a current block to be encoded in an image. In other words, the prediction section can predict pixel values of pixels of a current block to be encoded in an image through intra prediction or inter prediction to generate a prediction block having predicted pixel values of the pixels. In addition, the prediction section can deliver information required to generate the prediction block to the encoding section so that the prediction mode information is encoded. The encoding section adds the corresponding information to a bitstream and transmits the bitstream to a decoder. The decoding section of the decoder can parse the corresponding information, recover the prediction mode information, and then perform intra prediction or inter prediction using the prediction mode information.
[0077] The subtracter subtracts the prediction block from the current block to generate a residual block. In other words, the subtracter can calculate a difference between a pixel value of each pixel of a current block to be encoded and a predicted pixel value of each pixel of a prediction block generated by the prediction section to generate a residual block as a block-type residual signal.
[0078] The transform section can transform a signal belonging to a spatial domain into a signal belonging to a frequency domain. In this case, a signal obtained through the transform process is referred to as a transform coefficient. For example, the residual block having the residual signal delivered from the subtracter can be transformed into a transform block having transform coefficients. In this case, an input signal is determined according to an encoding setting, and the input signal is not limited to the residual signal.
[0079] The transform unit can perform a transform on the residual block by using a transform technique such as a Hadamard transform, a transform based on a discrete sine transform (DST), and a transform based on a discrete cosine transform (DCT). However, the present application is not limited thereto, and various enhanced and modified transform techniques can be used.
[0080] For example, at least one transform technique can be supported, and at least one detailed transform technique can be supported in each transform technique. In this case, at least one detailed transform technique can be a transform technique in which some basis vectors are differently constructed in each transform technique. For example, as a transform technique, a DST-based transform and a DCT-based transform can be supported. For the DST, detailed transform techniques such as DST-I, DST-II, DST-III, DST-V, DST-VI, DST-VII, and DST-VIII can be supported, and for the DCT, detailed transform techniques such as DCT-I, DCT-II, DCT-III, DCT-V, DCT-VI, DCT-VII, and DCT-VIII can be supported.
[0081] One of the transform techniques can be set as a default transform technique (e.g., one transform technique && one detailed transform technique), and additional transform techniques (e.g., multiple transform techniques || multiple detailed transform techniques) can be supported. Whether the additional transform techniques are supported can be determined in units of a sequence, a picture, a slice, or a tile, and related information can be generated according to the unit. When the additional transform techniques are supported, transform technique selection information can be determined in units of a block, and related information can be generated.
[0082] The transform can be performed horizontally and / or vertically. For example, a two-dimensional (2D) transform is performed by using basis vectors to perform a one-dimensional (1D) transform horizontally and vertically, so that pixel values in a spatial domain can be transformed into a frequency domain.
[0083] In addition, the transform can be performed horizontally and / or vertically in an adaptive manner. In detail, whether the transform is performed in an adaptive manner can be determined according to at least one encoding setting. For intra prediction, for example, DCT-I can be applied horizontally and DST-I can be applied vertically when a prediction mode is a horizontal mode, DST-VI can be applied horizontally and DCT-VI can be applied vertically when the prediction mode is a vertical mode, DCT-II can be applied horizontally and DCT-V can be applied vertically when the prediction mode is a diagonal left-down mode, and DST-I can be applied horizontally and DST-VI can be applied vertically when the prediction mode is a diagonal right-down mode.
[0084] The size and shape of the transform block can be determined according to the coding cost of the candidates for the size and shape of the transform block. The image data of the transform block and information about the determined size and shape of the transform block can be coded.
[0085] Among the transform forms, a square transform can be set as a default transform form, and an additional transform form (e.g., a rectangular form) can be supported. Whether to support the additional transform form can be determined in units of sequence, picture, slice, or tile, and related information can be generated according to the units. The transform form selection information can be determined in units of block, and related information can be generated.
[0086] In addition, whether to support the transform block form can be determined according to the coding information. In this case, the coding information can correspond to a slice type, a coding mode, a size and shape of a block, a block partitioning scheme, etc. That is, one transform form can be supported according to at least one piece of coding information, and multiple transform forms can be supported according to at least one piece of coding information. The former case can be an implicit situation, and the latter case can be an explicit situation. For the explicit situation, adaptive selection information indicating a best candidate group selected from among multiple candidate groups can be generated, and the adaptive selection information can be added to a bitstream. According to the present application, it will be understood that, in addition to this example, when the coding information is generated explicitly, the information is added to the bitstream in various units, and the related information is parsed in various units by a decoder and recovered into the decoding information. In addition, it will be understood that, when the coding / decoding information is processed implicitly, the processing is performed by the encoder and the decoder through the same processing, rules, etc.
[0087] As an example, support for a rectangular transform can be determined according to a slice type. For an I slice, the supported transform form can be a square transform, and for a P / B slice, the supported transform form can be a square transform or a rectangular transform.
[0088] As an example, support for a rectangular transform can be determined according to a coding mode. For intra prediction, the supported transform form can be a square transform, and for inter prediction, the supported transform form can be a square transform and / or a rectangular transform.
[0089] As an example, support for a rectangular transform can be determined according to a size and shape of a block. The supported transform form for a block of a specific size or greater can be a square transform, and the supported transform form for a block smaller than the specific size can be a square transform and / or a rectangular transform.
[0090] As an example, support for a rectangular transform can be determined according to a block partitioning scheme. When a block to be transformed is a block obtained through a quad-tree partitioning scheme, the supported transform form can be a square transform. When a block to be transformed is a block obtained through a binary-tree partitioning scheme, the supported transform form can be a square transform or a rectangular transform.
[0091] The above example can be an example of support for a transform form according to one piece of encoding information, and a plurality of pieces of information can be associated with additional transform form support settings in combination. The above example is merely an example of additional transform form support according to various encoding settings. However, the present application is not limited thereto, and various modifications can be made thereto.
[0092] Transform processing can be omitted according to an encoding setting or an image characteristic. For example, transform processing (including inverse processing) can be omitted according to an encoding setting (for example, it is assumed that a lossless compression environment in this example). As another example, transform processing can be omitted when compression performance through a transform is not shown according to an image characteristic. In this case, a transform can be omitted for all units or one of horizontal units and vertical units. Whether to support omission can be determined according to the size and shape of a block.
[0093] For example, it is assumed that horizontal transform and vertical transform are set to be omitted in common. When a transform omission flag is 1, a transform can not be performed horizontally and vertically, and when the transform omission flag is 0, a transform can be performed horizontally and vertically. On the other hand, it is assumed that horizontal transform and vertical transform are set to be omitted independently. When a first transform omission flag is 1, horizontal transform is not performed, and when the first transform omission flag is 0, horizontal transform is performed. Also, when a second transform omission flag is 1, vertical transform is not performed, and when the second transform omission flag is 0, vertical transform is performed.
[0094] When the size of a block corresponds to a range A, transform omission can be supported, and when the size of a block corresponds to a range B, transform omission cannot be supported. For example, when the width of a block is greater than M or the height of a block is greater than N, transform omission flag cannot be supported. When the width of a block is less than m or the height of a block is less than n, transform omission flag can be supported. M (m) and N (n) can be the same as or different from each other. Settings associated with a transform can be determined in units of a sequence, a picture, a slice, or the like.
[0095] When an additional transform technique is supported, a transform technique setting can be determined according to at least one piece of encoding information. In this case, the encoding information can correspond to a slice type, a coding mode, the size and shape of a block, a prediction mode, or the like.
[0096] As an example, support for transform techniques can be determined according to a coding mode. For intra prediction, supported transform techniques can include DCT-I, DCT-III, DCT-VI, DST-II, and DST-III, and for inter prediction, supported transform techniques can include DCT-II, DCT-III, and DST-III.
[0097] As an example, support for transform techniques can be determined according to a slice type. For I slices, supported transform techniques can include DCT-I, DCT-II, and DCT-III, for P slices, supported transform techniques can include DCT-V, DST-V, and DST-VI, and for B slices, supported transform techniques can include DCT-I, DCT-II, and DST-III.
[0098] As an example, support for transform techniques can be determined according to a prediction mode. Prediction mode A can support transform techniques including DCT-I and DCT-II, prediction mode B can support transform techniques including DCT-I and DST-I, and prediction mode C can support transform techniques including DCT-I. In this case, prediction mode A and prediction mode B can both be directional modes, and prediction mode C can be a non-directional mode.
[0099] As an example, support for transform techniques can be determined according to a size and shape of a block. Blocks of a certain size or larger can support transform techniques including DCT-II, blocks smaller than the certain size can support transform techniques including DCT-II and DST-V, and blocks of the certain size or larger than the certain size as well as blocks smaller than the certain size can support transform techniques including DCT-I, DCT-II, and DST-I. Further, blocks supported in a square form can support transform techniques including DCT-I and DCT-II, and blocks supported in a rectangular shape can support transform techniques including DCT-I and DST-I.
[0100] The above examples can be examples of support for transform techniques according to one piece of coding information, and multiple pieces of information can be associated in combination with additional transform technique support settings. The present application is not limited to the above examples, and the above examples can be modified. Further, the transform section can deliver information required to generate a transform block to the encoding section, so that the information is encoded. The encoding section adds the corresponding information to a bitstream and transmits the bitstream to a decoder. A decoding section of the decoder can parse the information and use the parsed information in an inverse transform process.
[0101] The quantization section can quantize the input signal. In this case, the signal obtained through the quantization process is referred to as a quantization coefficient. For example, the quantization section can quantize the residual block having the residual transform coefficient delivered from the transform section, thereby obtaining a quantization block having a quantization coefficient. In this case, the input signal is determined according to the encoding setting, and is not limited to the residual transform coefficient.
[0102] The quantization section can quantize the transformed residual block using a quantization technique such as dead zone uniform threshold quantization, quantization weighting matrix, etc. However, the present application is not limited thereto, and various quantization techniques improved and modified can be used. Whether to support an additional quantization technique can be determined in units of a sequence, a picture, a slice, or a tile, and related information can be generated according to the unit. When the additional quantization technique is supported, quantization technique selection information can be determined in units of a block, and related information can be generated.
[0103] When the additional quantization technique is supported, the quantization technique setting can be determined according to at least one piece of encoding information. In this case, the encoding information can correspond to a slice type, an encoding mode, a size and shape of a block, a prediction mode, etc.
[0104] For example, the quantization section can differently set a quantization weighting matrix corresponding to an encoding mode and a weighting matrix applied according to inter / intra prediction. In addition, the quantization section can differently set a weighting matrix applied according to an intra prediction mode. In this case, when it is assumed that the quantization weighting matrix has an M×N size identical to that of the quantization block, the quantization weighting matrix can be a quantization matrix in which some quantization components are differently constructed.
[0105] The quantization process can be omitted according to the encoding setting or the image characteristics. For example, the quantization process (including inverse processing) can be omitted according to the encoding setting (for example, such as in this example, it is assumed that a lossless compression environment). As another example, when the compression performance through quantization is not shown according to the image characteristics, the quantization process can be omitted. In this case, some or all of the regions can be omitted, and whether to support the omission can be determined according to the size and shape of the block.
[0106] Information on a quantization parameter (QP) can be generated in units of a sequence, a picture, a slice, a tile, or a block. For example, a default QP can be set in a higher unit <1> in which QP information is first generated, and the QP can be set to the same value as or a different value from the value of the QP set in the higher unit. In a quantization process performed in some units through this processing, the QP can be finally determined. In this case, units such as a sequence and a picture can be examples corresponding to <1>, units such as a slice, a tile, and a block can be examples corresponding to <2>, and units such as a block can be examples corresponding to <3>.
[0107] Information on a QP can be generated based on the QP in each unit. Alternatively, a predetermined QP can be set as a prediction value, and information on a difference from the QP in the unit can be generated. Alternatively, a QP acquired based on at least one of a QP set in a higher unit, a QP set in the same and previous unit, or a QP set in a neighboring unit can be set as a prediction value, and information on a difference from the QP in the current unit can be generated. Alternatively, a QP set in a higher unit and a QP acquired based on at least one piece of encoding information can be set as prediction values, and difference information from the QP in the current unit can be generated. In this case, the same and previous unit can be a unit that can be defined in the order in which the unit is encoded, the neighboring unit can be a spatially neighboring unit, and the encoding information can be a slice type, an encoding mode, a prediction mode, position information, or the like of the corresponding unit.
[0108] As an example, a QP in a higher unit can be set as a prediction value using a QP in a current unit and difference information can be generated. Information on a difference between a QP set in a slice and a QP set in a picture can be generated, or information on a difference between a QP set in a tile and a QP set in a picture can be generated. Further, information on a difference between a QP set in a block and a QP set in a slice or a tile can be generated. Further, information on a difference between a QP set in a sub-block and a QP set in a block can be generated.
[0109] As an example, a QP acquired based on a QP in at least one neighboring unit or a QP in at least one previous unit can be set as a prediction value using a QP in a current unit, and difference information can be generated. Information on a difference from a QP acquired based on a QP of a neighboring block such as a left side, an upper left side, a lower left side, an upper side, an upper right side, or the like of a current block can be generated. Alternatively, information on a difference from a QP of an encoded picture before a current picture can be generated.
[0110] As an example, the QP in the upper unit and the QP obtained based on at least one piece of encoding information can be set as a prediction value using the QP in the current unit and difference information can be generated. Further, information on a difference between the QP in the current block and the QP of the slice corrected according to the slice type (I / P / B) can be generated. Alternatively, information on a difference between the QP in the current block and the QP of the tile corrected according to the encoding mode (intra / inter) can be generated. Alternatively, information on a difference between the QP in the current block and the QP of the picture corrected according to the prediction mode (directional / non-directional) can be generated. Alternatively, information on a difference between the QP in the current block and the QP of the picture corrected according to the position information (x / y) can be generated. In this case, the correction can refer to an operation of adding an offset to or subtracting an offset from the QP in the upper unit used for prediction. In this case, at least one piece of offset information can be supported according to an encoding setting, and information processed implicitly or information associated explicitly can be generated according to a predetermined process. The present application is not limited to the above-described example, and the above-described example can be modified.
[0111] The above example can be an example allowed when a signal indicating a QP change is provided or activated. For example, when neither a signal indicating a QP change is provided nor a signal indicating a QP change is activated, difference information is not generated, and a predicted QP can be determined as the QP in each unit. As another example, when a signal indicating a QP change is provided or activated, difference information is generated, and when the value of the difference information is 0, a predicted QP can be determined as the QP in each unit.
[0112] The quantization section can deliver information required to generate a quantized block to the encoding section so that the information is encoded. The encoding section adds the corresponding information to a bitstream and transmits the bitstream to a decoder. The decoding section of the decoder can parse the information and use the parsed information in inverse quantization processing.
[0113] The above example has been described assuming that the residual block is transformed and quantized by the transform section and the quantization section. However, the residual signal of the residual block can be transformed into a residual block having transform coefficients without performing a quantization process. Alternatively, only a quantization process can be performed without transforming the residual signal of the residual block into transform coefficients. Alternatively, neither a transform process nor a quantization process can be performed. This can be determined according to an encoding setting.
[0114] The encoding section can scan the generated quantization coefficients, transform coefficients, or residual signal of the residual block in at least one scan order (e.g., zigzag scan, vertical scan, horizontal scan, etc.), generate a quantization coefficient string, a transform coefficient string, or a signal string, and encode the quantization coefficient string, the transform coefficient string, or the signal string using at least one entropy encoding technique. In this case, information about the scan order can be determined according to an encoding setting (e.g., an encoding mode, a prediction mode, etc.), and information implicitly determined or explicitly associated using the information about the scan order can be generated. For example, one scan order can be selected from among a plurality of scan orders according to an intra prediction mode.
[0115] In addition, the encoding section can generate encoded data including encoded information delivered from each element, and can output the encoded data in a bitstream. This can be implemented with a multiplexer (MUX). In this case, encoding can be performed using a method such as Exp-Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) as an encoding technique. However, the present application is not limited thereto, and various encoding techniques obtained by improving and modifying the above-described encoding techniques can be used.
[0116] When entropy encoding (e.g., CABAC in this example) is performed on syntax elements such as information and residual block data generated through encoding / decoding processing, the entropy encoding apparatus can include a binarizer, a context modeler, and a binary arithmetic encoder. In this case, the binary arithmetic encoder can include a regular encoding engine and a bypass encoding engine.
[0117] Syntax elements input to the entropy encoding apparatus can not be binary values. Accordingly, when the syntax elements are not binary values, the binarizer can binarize the syntax elements and output a bin string consisting of 0 or 1. In this case, a bin represents a bit consisting of 0 or 1, and can be encoded through the binary arithmetic encoder. In this case, one of the regular encoding engine and the bypass encoding engine can be selected based on the appearance probability of 0 and 1, and this can be determined according to an encoding / decoding setting. When the frequency of 0 is equal to the frequency of 1 in data, the bypass encoding engine can be used; otherwise, the regular encoding engine can be used.
[0118] When the syntax elements are binarized, various methods can be used. For example, fixed length binarization, unary binarization, Truncated Rice Binarization, K-th order exponential Golomb binarization, etc. can be used. In addition, signed binarization or unsigned binarization can be performed according to a range of values of the syntax elements. The binarization process of the syntax elements according to the present application can include additional binarization methods as well as the binarization described in the above examples.
[0119] The inverse quantization section and the inverse transform section can be implemented by performing the processes performed in the transform section and the quantization section in reverse. For example, the inverse quantization section can inverse quantize the transform coefficients quantized by the quantization section, and the inverse transform section can inverse transform the inverse quantized transform coefficients to generate a restored residual block.
[0120] The adder adds the prediction block and the restored residual block to restore the current block. The restored block can be stored in a memory and can be used as reference data (for the prediction section, the filter section, etc.).
[0121] The in-loop filter section can additionally perform post-processing filter processing of one or more of a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), etc. The deblocking filter can remove block distortion generated at boundaries between blocks from the restored image. The ALF can perform filtering based on a value obtained by comparing the input image with the restored image. In detail, the ALF can perform filtering based on a value obtained by comparing the input image with the restored image after filtering the blocks by the deblocking filter. Alternatively, the ALF can perform filtering based on a value obtained by comparing the input image with the restored image after filtering the blocks by the SAO. The SAO can restore offset differences based on a value obtained by comparing the input image with the restored image, and can be applied in the form of band offset (BO), edge offset (EO), etc. In detail, the SAO can add an offset with respect to the original image to the restored image to which the deblocking filter is applied in units of at least one pixel, and can be applied in the form of BO, EO, etc. In detail, the SAO can add an offset with respect to the original image to the restored image after filtering the blocks by the ALF in units of pixels, and can be applied in the form of BO, EO, etc.
[0122] As the filtering information, setting information on whether each post-processing filter is supported can be generated in units of sequence, picture, slice, tile, etc. Further, setting information on whether each post-processing filter is executed can be generated in units of picture, slice, tile, block, etc. The range in which the filter is executed can be classified into an inside of an image and a boundary of the image. Setting information considering the classification can be generated. Further, information on a filtering operation can be generated in units of picture, slice, tile, block, etc. The information can be processed implicitly or explicitly, and independent filtering processing or dependent filtering processing can be applied to filtering according to a color component, and this can be determined according to an encoding setting. The in-loop filter section can deliver the filtering information to the encoding section so that the information is encoded. The encoding section adds the corresponding information to a bitstream and transmits the bitstream to a decoder. The decoding section of the decoder can parse the information and apply the parsed information to the in-loop filter section.
[0123] The memory can store the recovered blocks or pictures. The recovered blocks or pictures stored in the memory can be provided to the prediction section which performs intra prediction or inter prediction. In detail, for processing, a space in which a bitstream compressed by an encoder is stored in a queue form can be set as a coded picture buffer (CPB), and a space in which a decoded image is stored in a picture unit can be set as a decoded picture buffer (DPB). The CPB can store a decoding section in a decoding order, a decoding operation is emulated in the encoder, and a compressed bitstream is stored through the emulation processing. The bitstream output from the CPB is recovered through a decoding process, and the recovered image is stored in the DPB, and the pictures stored in the DPB can be referred to during the image encoding / decoding processing.
[0124] The decoding section can be implemented by performing the processing of the encoding section in reverse. For example, the decoding section can receive a quantized coefficient string, a transform coefficient string, or a signal string from a bitstream, decode the string, parse the decoded data including the decoded information, and deliver the parsed decoded data to each element.
[0125] Next, an image setting process applied to the image encoding / decoding apparatus according to the embodiment of the present application will be described. This is an example applied before encoding / decoding (initial image setting), but some processes can be examples to be applied to other steps (e.g., steps after encoding / decoding or sub-steps of encoding / decoding). The image setting process can be performed in consideration of, for example, a multimedia content characteristic, a bandwidth, a user terminal performance, and a network and user environment of accessibility. For example, image partitioning, image size adjustment, image reconstruction, and the like can be performed according to an encoding / decoding setting. The following description of the image setting process will focus on a rectangular image. However, the present application is not limited thereto, and the image setting process can be applied to a polygonal image. The same image setting can be applied regardless of an image form, or different image settings can be applied, which can be determined according to an encoding / decoding setting. For example, after checking information on an image shape (e.g., a rectangular shape or a non-rectangular shape), information on a corresponding image setting can be constructed.
[0126] The following examples will be described assuming that a dependent setting is provided for a color space. However, an independent setting can be provided for a color space. Further, in the following examples, the independent setting can include an example in which an encoding / decoding setting is independently provided for each color space. Although one color space is described, examples including applying the description to other color spaces (e.g., an example in which N is generated in a chroma component in a case where M is generated in a luminance component) are contemplated and can be derived. Further, the dependent setting can include an example in which a setting is made in proportion to a color format composition ratio (e.g., 4:4:4, 4:2:2, 4:2:0, etc.) (e.g., for 4:2:0, a chroma component is M / 2 in a case where a luminance component is M). Examples including applying the description to each color space are contemplated and can be derived. The description is not limited to the above-described examples, and can be commonly applied to the present application.
[0127] Some of the constructions in the following examples can be applied to various encoding techniques, e.g., a spatial domain encoding, a frequency domain encoding, a block-based encoding, an object-based encoding, etc.
[0128] In general, an input image can be encoded or decoded as it is or after image partitioning. For example, partitioning can be performed for error robustness or the like to prevent damage caused by packet loss during transmission. Alternatively, partitioning can be performed to classify regions having different properties in the same image according to a characteristic, a type, or the like of the image.
[0129] According to the present application, the image partitioning process can include a partitioning process and an inverse partitioning process. The following examples will describe the partitioning process, but the inverse partitioning process can be derived in reverse from the partitioning process.
[0130] Figure 3 is an example diagram illustrating an image divided into a plurality of layers to compress the image.
[0131] Part 3a is an example diagram in which an image sequence is composed of a plurality of GOPs. Further, one GOP can be composed of an I picture, a P picture, and a B picture, as shown in part 3b. One picture can be composed of a slice, a tile, and the like, as shown in part 3c. As shown in part 3d, a slice, a tile, and the like can be composed of a plurality of default coding parts, and as shown in part 3e, a default coding part can be composed of at least one coding subunit. The image setting process according to the present application will be described based on an example to be applied to units such as pictures, slices, and tiles as shown in parts 3b and 3c.
[0132] Figure 4 is a conceptual diagram illustrating an example of image division according to an embodiment of the present application.
[0133] Part 4a is a conceptual diagram in which an image (e.g., a picture) is divided at uniform intervals in a horizontal direction and a vertical direction. The division region can be referred to as a block. Each block can be a default coding part (or a maximum coding part) obtained by a picture division section, and can be a basic unit to be applied to a division unit to be described below.
[0134] Part 4b is a conceptual diagram in which an image is divided in at least one direction selected from a horizontal direction and a vertical direction. The division region (T0 to T3) can be referred to as a tile, and each region can be encoded or decoded independently or in dependence on other regions.
[0135] Part 4c is a conceptual diagram in which an image is divided into a group of consecutive blocks. The division region (S0, S1) can be referred to as a slice, and each region can be encoded or decoded independently or in dependence on other regions. The group of consecutive blocks can be defined according to a scan order. Generally, the group of consecutive blocks conforms to a raster scan order. However, the present application is not limited thereto, and the group of consecutive blocks can be determined according to an encoding / decoding setting.
[0136] Part 4d is a conceptual diagram in which an image is divided into a group of blocks according to any user-defined setting. The division region (A0 to A2) can be referred to as an arbitrary division, and each region can be encoded or decoded independently or in dependence on other regions.
[0137] Independent encoding / decoding can mean that data in other units (or regions) cannot be referenced when some units are encoded or decoded. In detail, pieces of information used or generated during texture encoding and entropy encoding for some units can be independently encoded without reference to each other. Even in a decoder, for texture decoding and entropy decoding for some units, parsed information and restored information in other units can not reference each other. In this case, whether to reference data in other units (or regions) can be limited to a spatial region (e.g., between regions in one image), but can also be limited to a temporal region (e.g., between consecutive images or frames) according to an encoding / decoding setting. For example, when some units of a current image and some units of another image have continuity or have the same encoding environment, the reference can be made; otherwise, the reference can be limited.
[0138] Further, dependent encoding / decoding can mean that data in other units can be referenced when some units are encoded or decoded. In detail, pieces of information used or generated during texture encoding and entropy encoding for some units can be dependently encoded as well as reference to each other. Even in a decoder, for texture decoding and entropy decoding for some units, parsed information and restored information in other units can reference each other. That is, the above setting can be the same as or similar to that of general encoding / decoding. In this case, in order to identify a region (here, a face generated according to a projection format) to which the above setting is applied, a flag can be used. For example, a flag indicating whether the above setting is applied to a region can be used. In this case, the flag can be used in a slice header, a picture header, a tile group header, a tile group, a slice, a picture, a tile group, a tile, a region, or the like. <face>According to the present application, the image division processing can include an image division indication step, an image division type identification step, and an image division execution step. In this case, the image division indication step can include a step of indicating a division type of an image according to a characteristic, a type, or the like of the image. The image division type identification step can include a step of identifying a division type of an image according to a characteristic, a type, or the like of the image. The image division execution step can include a step of dividing an image according to a division type of the image.
[0139] In the above example, some units (slices, tiles, or the like) can be provided with independent encoding / decoding settings (for example, independent slice segments), and other units can be provided with dependent encoding / decoding settings (for example, dependent slice segments). According to the present application, the following description will focus on independent encoding / decoding settings.
[0140] As shown in section 4a, the default encoding section acquired by the picture division section can be divided into default encoding blocks according to the color space, and can have a size and a shape determined according to the characteristic and the resolution of the image. The size or the shape of the supported blocks can be an NxN square (2 n ) having a width and a height expressed as an exponential multiple of 2 (2 n × 2 n ; 256x256, 128x128, 64x64, 32x32, 16x16, 8x8, or the like; n is an integer in the range of 3 to 8) or an MxN rectangle (2 m × 2 n ). For example, the input image can be divided according to the resolution, can be divided into 128x128 for an 8k UHD image, can be divided into 64x64 for a 1080p HD image, or can be divided into 16x16 for a WVGA image, and can be divided according to the image type, can be divided into 256x256 for a 360-degree image. The default encoding section can be divided into encoding sub-units, and then encoded or decoded. Information about the default encoding section can be added to the bitstream in units of sequences, pictures, slices, tiles, or the like, and can be parsed by the decoder to recover the relevant information.
[0141] The image encoding method and the image decoding method according to the embodiment of the present application can include the following image division step. In this case, the image division processing can include an image division indication step, an image division type identification step, and an image division execution step. In addition, the image encoding apparatus and the image decoding apparatus can be configured to include an image division indication section, an image division type identification section, and an image division execution section that respectively perform the image division indication step, the image division type identification step, and the image division execution step. For encoding, the relevant syntax elements can be generated. For decoding, the relevant syntax elements can be parsed.
[0142] In the block division processing, as shown in section 4a, the image division indication section can be omitted. The image division type identification section can check information about the size and the shape of the block, and the image division section can perform the division by the division type information identified in the default encoding section.
[0143] The block can be a unit to be always divided, but whether to divide other division units (tile, slice, etc.) can be determined according to the encoding / decoding setting. As a default setting, the picture division section can perform division in units of blocks, and then perform division in other units. In this case, the block division can be performed based on the picture size.
[0144] Further, the division in units of blocks can be performed after the division in other units (tile, slice, etc.). That is, the block division can be performed based on the size of the division unit. This can be determined according to the encoding / decoding setting by explicit or implicit processing. The following examples describe the former case as an assumption, and will also focus on units other than blocks.
[0145] In the image division indication step, whether to perform image division can be determined. For example, when a signal (e.g., tiles_enabled_flag) indicating image division is confirmed, the division can be performed. When a signal indicating image division is not confirmed, the division can not be performed, or can be performed by confirming other encoding / decoding information.
[0146] In detail, it is assumed that a signal (e.g., tiles_enabled_flag) indicating image division is confirmed. When the signal is activated (e.g., tiles_enabled_flag = 1), the division can be performed in a plurality of units. When the signal is deactivated (e.g., tiles_enabled_flag = 0), the division can not be performed. Alternatively, a signal indicating image division not being confirmed can mean that the division is not performed or the division is performed in at least one unit. Whether to perform the division in a plurality of units can be confirmed by another signal (e.g., first_slice_segment_in_pic_flag).
[0147] In summary, when a signal indicating image division is provided, the corresponding signal is a signal for indicating whether to perform division in a plurality of units. Whether to divide the corresponding image can be determined according to the signal. For example, it is assumed that tiles_enabled_flag is a signal indicating whether to divide an image. Here, tiles_enabled_flag equal to 1 can mean that the image is divided into a plurality of tiles, and tiles_enabled_flag equal to 0 can mean that the image is not divided.
[0148] In summary, when a signal indicating image division is not provided, the division can not be performed, or it can be determined whether to divide the corresponding image through another signal. For example, first_slice_segment_in_pic_flag is not a signal indicating whether to perform image division, but a signal indicating a first slice segment in an image. Thus, it can be confirmed whether division is performed in two or more units (for example, marked as 0 indicates that the image is divided into a plurality of slices).
[0149] The present application is not limited to the above-described examples, and the above-described examples can be modified. For example, a signal indicating image division can not be provided for each tile, and a signal indicating image division can be provided for each slice. Alternatively, a signal indicating image division can be provided based on the type, characteristics, etc. of an image.
[0150] In the image division type identification step, an image division type can be identified. The image division type can be defined by a division method, division information, etc.
[0151] In section 4b, a tile can be defined as a unit obtained by horizontal and vertical division. In detail, a tile can be defined as a group of adjacent blocks in a quadrangular space divided by at least one horizontal or vertical division line passing through an image.
[0152] Tile division information can include column and row boundary position information, tile number information of columns and rows, tile size information, etc. Tile number information can include the number of columns of tiles (for example, num_tile_columns) and the number of rows of tiles (for example, num_tile_rows). Thus, an image can be divided into a number (= number of columns x number of rows) of tiles. Tile size information can be obtained based on tile number information. The width or height of a tile can be uniform or non-uniform, and thus, under a predetermined rule, related information (for example, uniform_spacing_flag) can be implicitly determined or explicitly generated. In addition, tile size information can include size information of each column and each row of a tile (for example, column_width_tile[i] and row_height_tile[i]), or size information of the width and height of each tile. In addition, size information can be information that can be additionally generated depending on whether tile size is uniform (for example, in the case of non-uniform division because uniform_spacing_flag is 0).
[0153] In section 4c, a slice can be defined as a unit grouping consecutive blocks. In detail, a slice can be defined as a group of consecutive blocks in a predetermined scan order (here, in a raster scan).
[0154] The slice division information can include slice number information, slice position information (e.g., slice_segment_address), etc. In this case, the slice position information can be position information of a predetermined block (e.g., the first row in the scan order in the slice). In this case, the position information can be block scan order information.
[0155] In section 4d, various division settings are allowed for arbitrary division.
[0156] In section 4d, the division unit can be defined as a group of blocks that are spatially adjacent to each other, and the information about the division can include information about the size, form, and position of the division unit. This is only an example of arbitrary division, and various division forms can be allowed as shown in Figure 5
[0157] Figure 5 is another example diagram of an image division method according to an embodiment of the present invention.
[0158] In sections 5a and 5b, an image can be divided into a plurality of regions with at least one block interval in the horizontal or vertical direction, and the division can be performed based on block position information. Section 5a shows an example in which division is performed horizontally based on row information of each block (A0, A1), and section 5b shows an example in which division is performed horizontally and vertically based on column information and row information of each block (B0 to B3). The information about the division can include the number of division units, block interval information, division direction, etc., and some division information can not be generated when the division information is implicitly included according to a predetermined rule.
[0159] In sections 5c and 5d, an image can be divided into a group of consecutive blocks in the scan order. An additional scan order other than the conventional slice raster scan order can be applied to image division. Section 5c shows an example in which scanning is performed clockwise or counterclockwise with respect to a starting block (Box-Out) (C0, C1), and section 5d shows an example in which scanning is performed vertically with respect to a starting block (vertical) (D0, D1). The information about the division can include information about the number of division units, information about the position of the division unit (e.g., the first row in the scan order in the division unit), information about the scan order, etc., and some division information can not be generated when the division information is implicitly included according to a predetermined rule.
[0160] In part 5e, the image can be divided using the horizontal and vertical division lines. The existing tile can be divided by the horizontal or vertical division line. Thus, the division can be performed in the form of a quadrangular space, but it can not be possible to divide the image using the division line. For example, an example of dividing the image through some division lines of the image (e.g., a division line between the left boundary of E5 and the right boundary of E1, E3, and E4) is possible, and an example of dividing the image through some division lines of the image (e.g., a division line between the lower boundary of E2 and E3 and the upper boundary of E4) is not possible. In addition, the division can be performed based on the block unit (e.g., after first performing the block division), or the division can be performed through the horizontal or vertical division line (e.g., the division is performed through the division line regardless of the block division). Thus, each division unit can not be a multiple of the block. Thus, division information different from the division information of the existing tile can be generated, and the division information can include information on the number of division units, information on the position of the division unit, information on the size of the division unit, etc. For example, the information on the position of the division unit can be generated as position information (e.g., measured in pixels or in blocks) based on a predetermined position (e.g., the upper left corner of the image), and the information on the size of the division unit can be generated as information on the width and height of each division unit (e.g., measured in pixels or in blocks).
[0161] Similar to the above example, the division according to any user-defined setting can be performed by applying a new division method or by changing some elements of the existing division. That is, the division method can be supported by replacing or adding to the conventional division method, and the division method can be supported by changing some settings (slices, tiles, etc.) of the conventional division method (e.g., according to another scan order, by using another division method of a quadrangular shape to generate other division information, or according to a dependent encoding / decoding characteristic). In addition, a setting for configuring an additional division unit (e.g., a setting other than division according to a scan order or division according to a specific interval difference) can be supported, and a form of an additional division unit (e.g., a polygonal form such as a triangle other than division into a quadrangular space) can be supported. In addition, the image division method can be supported based on the type, characteristic, etc. of the image. For example, a partial division method (e.g., a face of a 360-degree image) can be supported according to the type, characteristic, etc. of the image. The information on the division can be generated based on the support.
[0162] In the image division execution step, the image can be divided based on the identified division type information. That is, the image can be divided into a plurality of division units based on the identified division type, and the image can be encoded or decoded based on the obtained division unit.
[0163] In this case, it can be determined whether to have the encoding / decoding setting in each division unit according to the division type. That is, the setting information required during the encoding / decoding process for each division unit can be designated by a higher unit (e.g., a picture), or an independent encoding / decoding setting can be provided for each division unit.
[0164] Generally, a slice can have an independent encoding / decoding setting for each division unit (e.g., a slice header), and a tile cannot have an independent encoding / decoding setting for each division unit and can have a setting depending on a picture encoding / decoding setting (e.g., a PPS). In this case, the information generated in association with the tile can be division information, and can be included in the picture encoding / decoding setting. The present application is not limited to the above-described example, and the above-described example can be modified.
[0165] The encoding / decoding setting information for the tile can be generated in units of a video, a sequence, a picture, etc. At least one encoding / decoding setting information is generated in a higher unit, and the generated one encoding / decoding setting information can be referred to. Alternatively, independent encoding / decoding setting information (e.g., a tile header) can be generated in units of a tile. This is different from the case of following one encoding / decoding setting determined in a higher unit, in that encoding / decoding is performed in a case where at least one encoding / decoding setting is provided in units of a tile. That is, all tiles can be encoded or decoded according to the same encoding / decoding setting, or at least one tile can be encoded or decoded according to an encoding / decoding setting different from the encoding / decoding setting of the other tiles.
[0166] The above examples focus on various encoding / decoding settings in a tile. However, the present application is not limited thereto, and even the same or similar settings can be applied to other division types.
[0167] As an example, in some division types, division information can be generated in a higher unit, and encoding or decoding can be performed according to a single encoding / decoding setting of the higher unit.
[0168] As an example, in some division types, division information can be generated in a higher unit, and independent encoding / decoding settings for each division unit in the higher unit can be generated, and encoding or decoding can be performed according to the generated encoding / decoding settings.
[0169] As an example, in some division types, division information can be generated in a higher unit, and a plurality of encoding / decoding setting information can be supported in the higher unit. Encoding or decoding can be performed according to the encoding / decoding setting referred to by each division unit.
[0170] As an example, in some division types, division information can be generated in a higher unit, and independent encoding / decoding settings including the division information can be generated in a corresponding division unit, and encoding or decoding can be performed according to the generated encoding / decoding settings.
[0171] As an example, in some division types, division information can be generated in a higher unit, and independent encoding / decoding settings including the division information can be generated in a corresponding division unit, and encoding or decoding can be performed according to the generated encoding / decoding settings.
[0172] The encoding / decoding setting information can include information required to encode or decode a tile, such as a tile type, information on a reference picture list, quantization parameter information, inter prediction setting information, in-loop filtering setting information, in-loop filtering control information, a scan order, whether to perform encoding or decoding, etc. The encoding / decoding setting information can be used to explicitly generate related information, or can have encoding / decoding settings determined implicitly according to a format, characteristics, etc. of an image determined in a higher unit. In addition, related information can be explicitly generated based on information acquired through settings.
[0173] Next, an example of performing image division in an encoding / decoding apparatus according to an embodiment of the present application will be described.
[0174] Division processing can be performed on an input image before starting encoding. An image can be divided using division information (e.g., image division information, division unit setting information, etc.), and then the image can be encoded in a division unit. Image encoding data can be stored in a memory after encoding is completed, and can be added to a bitstream and then transmitted.
[0175] Division processing can be performed before starting decoding. An image can be divided using division information (e.g., image division information, division unit setting information, etc.), and then image decoding data can be parsed and decoded in a division unit. After decoding is completed, image decoding data can be stored in a memory, and a plurality of division units are merged into a single unit, and thus an image can be output.
[0176] Through the above examples, image division processing has been described. In addition, according to the present application, a plurality of division processes can be performed.
[0177] For example, the image can be divided, and the divided unit of the image can be divided. The division can be the same division process (e.g., slice / slice, tile / tile, etc.) or a different division process (e.g., slice / tile, tile / slice, tile / surface, surface / tile, slice / surface, surface / slice, etc.). In this case, the subsequent division process can be performed based on the previous division result, and the division information generated during the subsequent division process can be generated based on the previous division result.
[0178] Further, a plurality of division processes (A) can be performed, and the division processes can be different division processes (e.g., slice / surface, tile / surface, etc.). In this case, the subsequent division process can be performed based on or independently of the previous division result, and the division information generated during the subsequent division process can be generated based on or independently of the previous division result.
[0179] The plurality of image division processes can be determined according to the encoding / decoding setting. However, the present application is not limited to the above-described example, and various modifications can be made to the above-described example.
[0180] The encoder can add the information generated during the above-described process to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the related information from the bitstream. That is, the information can be added to one unit, and the information can be copied and added to a plurality of units. For example, a syntax element indicating whether some information is supported or a syntax element indicating whether activation is performed can be generated in some units (e.g., upper units), and the same or similar information can be generated in some units (e.g., lower units). That is, even in the case where the related information is supported and set in the upper units, the lower units can have separate settings. The description is not limited to the above-described example, and can be commonly applied to the present application. Further, the information can be included in the bitstream in the form of SEI or metadata.
[0181] In general, the input image can be encoded or decoded as it is, but the encoding or decoding can be performed after adjusting the image size (expansion or reduction; resolution adjustment). For example, in a layered encoding scheme (scalable video coding) for supporting spatial, temporal, and image quality scalability, image size adjustment such as overall expansion and reduction of the image can be performed. Alternatively, image size adjustment such as partial expansion and reduction of the image can be performed. The image size adjustment can be performed differently, that is, the image size adjustment can be performed for the purpose of adaptability to the encoding environment, for the purpose of encoding uniformity, for the purpose of encoding efficiency, for the purpose of image quality improvement, or according to the type, characteristics, etc. of the image.
[0182] As a first example, the size adjustment processing can be performed during processing performed in accordance with characteristics, types, and the like of an image (e.g., layered encoding, 360-degree image encoding, and the like).
[0183] As a second example, the size adjustment processing can be performed at an initial encoding / decoding step. The size adjustment processing can be performed before encoding or decoding is performed. The encoded or decoded image can be performed on the size-adjusted image.
[0184] As a third example, the size adjustment processing can be performed during a prediction step (intra prediction or inter prediction) or before prediction. During the size adjustment processing, image information (e.g., information on pixels of an intra prediction reference, information on an intra prediction mode, information on a reference picture for inter prediction, information on an inter prediction mode, and the like) can be used at the prediction step.
[0185] As a fourth example, the size adjustment processing can be performed during a filtering step or before filtering. During the size adjustment processing, image information in the filtering step (e.g., pixel information to be applied to a deblocking filter, pixel information to be applied to SAO, information on SAO filtering, pixel information applied to ALF, information on ALF filtering, and the like) can be used.
[0186] Further, after the size adjustment processing is performed, the image can be processed by inverse size adjustment processing and changed to the image before the size adjustment in terms of image size, or the image can not be changed. This can be determined in accordance with encoding / decoding settings (e.g., characteristics of performing size adjustment). In this case, the size adjustment processing can be expansion processing and the inverse size adjustment processing is reduction processing, and the size adjustment processing can be reduction processing and the inverse size adjustment processing is expansion processing.
[0187] When the size adjustment processing is performed in accordance with the first example to the fourth example, inverse size adjustment processing is performed in a subsequent step so that the image before the size adjustment can be acquired.
[0188] When the size adjustment processing is performed by layered encoding or in accordance with the third example (or when the size of a reference picture is adjusted in inter prediction), the inverse size adjustment processing can not be performed in a subsequent step.
[0189] In the embodiments of the present application, the image size adjustment processing can be performed alone or together with inverse processing. The following examples will be described with a focus on the size adjustment processing. In this case, since the inverse size adjustment processing is inverse processing for the size adjustment processing, the description of the inverse size adjustment processing will be omitted to prevent redundant description. However, it is obvious that the person skilled in the art can recognize the same content as that described literally.
[0190] Figure 6 is an example diagram of a general image resizing method.
[0191] Referring to portion 6a, an expanded image (P0+P1) can be obtained by adding a specific region (P1) to an initial image P0 (or an image before resizing; indicated by a thick solid line).
[0192] Referring to portion 6b, a reduced image (S0) can be obtained by removing a specific region (S1) from an initial image (S0+S1).
[0193] Referring to portion 6c, a resized image (T0+T1) can be obtained by adding a specific region (T1) to an initial image (T0+T2) and removing a specific region (T2) from the entire image.
[0194] According to the present application, the following description focuses on the resizing processing for expansion and the resizing processing for reduction. However, the present application is not limited to this, and should be understood to include the case where expansion and reduction are applied in combination as shown in portion 6c.
[0195] Figure 7 is an example diagram of image resizing according to an embodiment of the present application.
[0196] During the resizing processing, the image expansion method will be described with reference to portion 7a, and the image reduction method will be described with reference to portion 7b.
[0197] In portion 7a, the image before resizing is S0, and the image after resizing is S1. In portion 7b, the image before resizing is T0, and the image after resizing is T1.
[0198] When expanding an image as shown in portion 7a, the image can be expanded in the "upward” direction, the "downward” direction, the "leftward” direction, or the "rightward” direction (ET, EL, EB, ER). When reducing an image as shown in portion 7b, the image can be reduced in the "upward” direction, the "downward” direction, the "leftward” direction, or the "rightward” direction (RT, RL, RB, RR).
[0199] Comparing image expansion and image reduction, the "upward” direction, the "downward” direction, the "leftward” direction, and the "rightward” direction of expansion can correspond to the "downward” direction, the "upward” direction, the "rightward” direction, and the "leftward” direction of reduction. Therefore, the following description focuses on image expansion, but should be understood to include the description of image reduction.
[0200] In the following description, image expansion or reduction is performed in an "up" direction, a "down" direction, a "left" direction, and a "right" direction. However, it should also be understood that resizing can be performed in an "up and left" direction, an "up and right" direction, a "down and left" direction, or a "down and right" direction.
[0201] In this case, when expansion is performed in a "down and right" direction, the regions RC and BC are acquired, and the region BR can or can not be acquired depending on the encoding / decoding setting. That is, the regions TL, TR, BL, and BR can or can not be acquired, but for ease of description, it will be described that the corner regions (i.e., the regions TL, TR, BL, and BR) can be acquired.
[0202] The image resizing processing according to the embodiments of the present application can be performed in at least one direction. For example, the image resizing processing can be performed in all directions such as up, down, left, and right, can be performed in two or more directions selected from among up, down, left, and right (left+right, up+down, up+left, up+right, down+left, down+right, up+left+right, down+left+right, up+down+left, up+down+right, etc.), or can be performed in only one direction selected from among up, down, left, and right.
[0203] For example, resizing can be performed in a "left+right" direction, an "up+down" direction, a "left and up+right and down" direction, and a "left and down+right and up" direction, which symmetrically expand to both ends with respect to the center of the image, resizing can be performed in a "left+right" direction, a "left and up+right and up" direction, and a "left and down+right and down" direction, which symmetrically expand vertically with respect to the image, and resizing can be performed in an "up+down" direction, a "left and up+left and down" direction, and a "right and up+right and down" direction, which symmetrically expand horizontally with respect to the image. Other resizing can be performed.
[0204] In sections 7a and 7b, the size of the image before the resizing (S0, T0) is defined as P_Width x P_Height, and the size of the image after the resizing (S1, T1) is defined as P'_Width x P'_Height. Here, when the resizing values in the "left" direction, "right" direction, "top" direction, and "bottom" direction are defined as Var_L, Var_R, Var_T, and Var_B (or collectively as Var_x), the image size after the resizing can be expressed as (P_Width + Var_L + Var_R) x (P_Height + Var_T + Var_B). In this case, Var_L, Var_R, Var_T, and Var_B as the resizing values in the "left" direction, "right" direction, "top" direction, and "bottom" direction can be Exp_L, Exp_R, Exp_T, and Exp_B (here, Exp_x is positive) for image expansion (in section 7a), and can be -Rec_L, -Rec_R, -Rec_T, and -Rec_B (which represent negative values for image reduction when Rec_L, Rec_R, Rec_T, and Rec_B are defined as positive values). Further, the upper-left corner coordinate, upper-right corner coordinate, lower-left corner coordinate, and lower-right corner coordinate of the image before the resizing can be (0, 0), (P_Width - 1, 0), (0, P_Height - 1), and (P_Width - 1, P_Height - 1), and the upper-left corner coordinate, upper-right corner coordinate, lower-left corner coordinate, and lower-right corner coordinate of the image after the resizing can be expressed as (0, 0), (P'_Width - 1, 0), (0, P'_Height - 1), and (P'_Width - 1, P'_Height - 1). The size of the region changed (or acquired or removed) by the resizing (here, TL to BR; i is an index for identifying TL to BR) can be M[i] x N[i] and can be expressed as Var_X x Var_Y (this example assumes that X is L or R and Y is T or B). M and N can have various values and can have the same settings regardless of i, or can have separate settings according to i. Various examples will be described below.
[0205] Referring to section 7a, S1 can be configured to include some or all of the regions TL to BR (top-left to bottom-right) that would be generated by expanding S0 in several directions. Referring to section 7b, T1 can be configured to exclude all or some of the regions TL to BR that would be removed by reducing in several directions from T0.
[0206] In part 7a, when the existing image (S0) is expanded in the "up" direction, the "down" direction, the "left" direction, and the "right" direction, the image can include the regions TC, BC, LC, and RC obtained through the resizing process and can further include the regions TL, TR, BL, and BR.
[0207] As an example, when expansion is performed in the "up" direction (ET), the image can be constructed by adding the region TC to the existing image (S0), and the image can include the region TL or TR and expansion in at least one different direction (EL or ER).
[0208] As an example, when expansion is performed in the "down" direction (EB), the image can be constructed by adding the region BC to the existing image (S0), and the image can include the region BL or BR and expansion in at least one different direction (EL or ER).
[0209] As an example, when expansion is performed in the "left" direction (EL), the image can be constructed by adding the region LC to the existing image (S0), and the image can include the region TL or BL and expansion in at least one different direction (ET or EB).
[0210] As an example, when expansion is performed in the "right" direction (ER), the image can be constructed by adding the region RC to the existing image (S0), and the image can include the region TR or BR and expansion in at least one different direction (ET or EB).
[0211] According to embodiments of the present application, it is possible to provide a setting (e.g., spa_ref_enabled_flag or tem_ref_enabled_flag) for spatially or temporally limiting the referability of a resized region (this example assumes expansion).
[0212] That is, it is possible to allow (e.g., spa_ref_enabled_flag = 1 or tem_ref_enabled_flag = 1) or limit (e.g., spa_ref_enabled_flag = 0 or tem_ref_enabled_flag = 0) the reference to data of a region that is resized spatially or temporally according to an encoding / decoding setting.
[0213] Encoding / decoding of the image (S0, T1) before resizing and the regions (TC, BC, LC, RC, TL, TR, BL, and BR) added or deleted during resizing can be performed as follows.
[0214] For example, when encoding or decoding the image before resizing and the added or deleted region, data about the image before resizing and data about the added or deleted region (data after encoding or decoding is completed; pixel value or prediction-related information) can be spatially or temporally referenced to each other.
[0215] Alternatively, the image before resizing and data about the added or deleted region can be spatially referenced, while data about the image before resizing can be temporally referenced, and data about the added or deleted region cannot be temporally referenced.
[0216] That is, a setting for limiting the referenceability of the added or deleted region can be provided. Setting information about the referenceability of the added or deleted region can be explicitly generated or implicitly determined.
[0217] The image resizing process according to the embodiment of the present application can include an image resizing indication step, an image resizing type identification step, and / or an image resizing execution step. In addition, the image encoding apparatus and the image decoding apparatus can include an image resizing indication section, an image resizing type identification section, and an image resizing execution section configured to perform the image resizing indication step, the image resizing type identification step, and the image resizing execution step, respectively. For encoding, the relevant syntax elements can be generated. For decoding, the relevant syntax elements can be parsed.
[0218] In the image resizing indication step, whether to perform image resizing can be determined. For example, when a signal indicating image resizing (e.g., img_resizing_enabled_flag) is confirmed, resizing can be performed. When a signal indicating image resizing is not confirmed, resizing can not be performed, or resizing can be performed by confirming other encoding / decoding information. In addition, although a signal indicating image resizing is not provided, the signal indicating image resizing can be implicitly activated or deactivated according to the encoding / decoding setting (e.g., characteristics, type, etc. of the image). When resizing is performed, corresponding resizing-related information can be generated, or corresponding resizing-related information can be implicitly determined.
[0219] When a signal indicating image resizing is provided, the corresponding signal is a signal for indicating whether to perform image resizing. Whether to resize the corresponding image can be determined according to the signal.
[0220] For example, assume that a signal indicating image size adjustment (e.g., img_resizing_enabled_flag) is confirmed. When the corresponding signal (e.g., img_resizing_enabled_flag = 1) is activated, image size adjustment can be performed. When the corresponding signal (e.g., img_resizing_enabled_flag = 0) is deactivated, image size adjustment can not be performed.
[0221] In addition, when a signal indicating image size adjustment is not provided, size adjustment can not be performed, or whether to adjust the size of a corresponding image can be determined through another signal.
[0222] For example, when an input image is divided in units of blocks, size adjustment can be performed according to whether the size (e.g., width or height) of the image is a multiple of the size (e.g., width or height) of the block (for extension in this example, assume that size adjustment processing is performed when the image size is not a multiple of the block size). That is, when the width of the image is not a multiple of the width of the block or when the height of the image is not a multiple of the height of the block, size adjustment can be performed. In this case, size adjustment information (e.g., size adjustment direction, size adjustment value, etc.) can be determined according to encoding / decoding information (e.g., size of the image, size of the block, etc.). Alternatively, size adjustment can be performed according to the characteristics, type (e.g., 360-degree image), etc. of the image, and size adjustment information can be explicitly generated or can be designated as a predetermined value. The present application is not limited to the above-described example, and the above-described example can be modified.
[0223] In the image size adjustment type identification step, an image size adjustment type can be identified. The image size adjustment type can be defined by a size adjustment method, size adjustment information, etc. For example, size adjustment based on a scaling factor, size adjustment based on an offset factor, etc. can be performed. The present application is not limited to the above-described example, and these methods can be applied in combination. For ease of description, the following description will focus on size adjustment based on a scaling factor and size adjustment based on an offset factor.
[0224] For a scaling factor, size adjustment can be performed by multiplication or division based on the size of the image. Information about a size adjustment operation (e.g., expansion or reduction) can be explicitly generated, and expansion or reduction processing can be performed according to the corresponding information. In addition, size adjustment processing can be performed as a predetermined operation (e.g., one of an expansion operation and a reduction operation) according to encoding / decoding settings. In this case, information about the size adjustment operation will be omitted. For example, when image size adjustment is activated in the image size adjustment indication step, image size adjustment can be performed as a predetermined operation.
[0225] The resizing direction can be at least one direction selected from among: up, down, left, and right. Depending on the resizing direction, at least one scale factor can be required. That is, one scale factor can be required for each direction (here, unidirectional), one scale factor can be required for horizontal or vertical direction (here, bidirectional), and one scale factor can be required for all directions of the image (here, omnidirectional). In addition, the resizing direction is not limited to the above-described examples, and the above-described examples can be modified.
[0226] The scale factor can have a positive value, and can have range information that is different according to the encoding / decoding setting. For example, when the information is generated by combining the resizing operation and the scale factor, the scale factor can be used as a multiplier. The scale factor greater than 0 or less than 1 can mean a reduction operation, the scale factor greater than 1 can mean an expansion operation, and the scale factor being 1 can mean that the resizing is not performed. As another example, when the scale factor information is generated independently of the resizing operation, the scale factor for the expansion operation can be used as a multiplier, and the scale factor for the reduction operation can be used as a divisor.
[0227] The process of changing the image (S0, T0) before the resizing into the image (here, S1 and T1) after the resizing will be described again with reference to parts 7a and 7b of FIG. 7. Figure 7
[0228] As an example, when one scale factor (referred to as sc) is used in all directions of the image and the resizing direction is the "down + right" direction, the resizing direction is ER and EB (or RR and RB), the resizing values Var_L (Exp_L or Rec_L) and Var_T (Exp_T or Rec_T) are 0, and Var_R (Exp_R or Rec_R) and Var_B (Exp_B or Rec_B) can be expressed as P_Width×(sc-1) and P_Height×(sc-1). Thus, the image after the resizing can be (P_Width×sc)×(P_Height×sc).
[0229] As an example, when respective scale factors (here, sc_w and sc_h) are used in the landscape direction or the portrait direction of the image and the resizing directions are the "left+right" direction and the "up+down" direction (up+down+left+right when both are operated), the resizing directions can be ET, EB, EL, and ER, the resizing values Var_T and Var_B can be P_Height x (sc_h-1) / 2, and Var_L and Var_R can be P_Width x (sc_w-1) / 2. Thus, the image after resizing can be (P_Width x sc_w) x (P_Height x sc_h).
[0230] For the offset factor, resizing can be performed by addition or subtraction based on the size of the image. Alternatively, resizing can be performed by addition or subtraction based on the encoding / decoding information of the image. Alternatively, resizing can be performed by independent addition or subtraction. That is, the resizing process can have a dependent setting or an independent setting.
[0231] Information about the resizing operation (e.g., expansion or reduction) can be explicitly generated, and expansion or reduction processing can be performed according to the respective information. In addition, the resizing operation can be performed as a predetermined operation (e.g., one of the expansion operation and the reduction operation) according to the encoding / decoding setting. In this case, information about the resizing operation can be omitted. For example, when the image resizing is activated in the image resizing indication step, the image resizing can be performed as a predetermined operation.
[0232] The resizing direction can be at least one direction selected from among up, down, left, and right. Depending on the resizing direction, at least one offset factor can be required. That is, one offset factor can be required for each direction (here, unidirectional), one offset factor can be required for the landscape or portrait direction (here, symmetric bidirectional), one offset factor can be required according to a partial combination of directions (here, asymmetric bidirectional), and one offset factor can be required for all directions of the image (here, omnidirectional). In addition, the resizing direction is not limited to the above-described examples, and the above-described examples can be modified.
[0233] The offset factor can have a positive value or both positive and negative values, and can have range information that is different according to encoding / decoding settings. For example, when the offset factor is generated in conjunction with the size adjustment operation and the offset factor generation information (here, it is assumed that the offset factor has both positive and negative values), the offset factor can be used as a value to be added or subtracted according to the sign information of the offset factor. The offset factor greater than 0 can mean an expansion operation, the offset factor less than 0 can mean a reduction operation, and the offset factor being 0 can mean that no size adjustment is performed. As another example, when the offset factor information is generated independently of the size adjustment operation (here, it is assumed that the offset factor has a positive value), the offset factor can be used as a value to be added or subtracted according to the size adjustment operation. The offset factor greater than 0 can mean that an expansion or reduction operation can be performed according to the size adjustment operation, and the offset factor being 0 can mean that no size adjustment is performed.
[0234] The method of changing the images (S0, T0) before the size adjustment into the images (S1, T1) after the size adjustment using the offset factor will be described again with reference to parts 7a and 7b of FIG. 7. Figure 7
[0235] As an example, when one offset factor (referred to as os) is used in all directions of the image and the size adjustment direction is the "up + down + left + right" direction, the size adjustment direction can be ET, EB, EL, and ER (or RT, RB, RL, and RR) and the size adjustment values Var_T, Var_B, Var_L, and Var_R can be os. The size of the image after the size adjustment can be (P_Width + os) x (P_Height + os).
[0236] As an example, when offset factors (os_w, os_h) are used in the horizontal or vertical direction of the image and the size adjustment direction is the "left + right" direction and the "up + down" direction ("up + down + left + right" direction when both operations are performed), the size adjustment direction can be ET, EB, EL, and ER (or RT, RB, RL, and RR), the size adjustment values Var_T and Var_B can be os_h, and the size adjustment values Var_L and Var_R can be os_w. The size of the image after the size adjustment can be {P_Width + (os_w x 2)} x {P_Height + (os_h x 2)}.
[0237] As an example, when the resizing direction is the "down" direction and the "right" direction (the "down + right" direction when operated together) and the offset factors (os_b, os_r) are used according to the resizing direction, the resizing direction can be EB and ER (or RB and RR), the resizing value Var_B can be os_b, and the resizing value Var_R can be os_r. The size of the image after resizing can be (P_Width + os_r) x (P_Height + os_b).
[0238] As an example, when the offset factors (os_t, os_b, os_l, os_r) are used according to the direction of the image and the resizing direction is the "up" direction, the "down" direction, the "left" direction, and the "right" direction (the "up + down + left + right" direction when all are operated), the resizing direction can be ET, EB, EL, and ER (or RT, RB, RL, and RR), the resizing value Var_T can be os_t, the resizing value Var_B can be os_b, the resizing value Var_L can be os_l, and the resizing value Var_R can be os_r. The size of the image after resizing can be (P_Width + os_l + os_r) x (P_Height + os_t + os_b).
[0239] The above examples indicate a case where the offset factor is used as the resizing value (Var_T, Var_B, Var_L, Var_R) during the resizing process. That is, this means that the offset factor is used as the resizing value without any change, which can be an example of independently performed resizing. Alternatively, the offset factor can be used as an input variable of the resizing value. In detail, the offset factor can be designated as the input variable, and the resizing value can be acquired through a series of processes according to the encoding / decoding setting, which can be an example of resizing performed based on predetermined information (e.g., image size, encoding / decoding information, etc.) or an example of dependency-based resizing.
[0240] For example, the offset factor can be a multiple (e.g., 1, 2, 4, 6, 8, and 16) or an exponential multiple (e.g., an exponential multiple of 2, such as 1, 2, 4, 8, 16, 32, 64, 128, and 256) of a predetermined value (here, an integer). Alternatively, the offset factor can be a multiple or an exponential multiple of a value (e.g., a value set based on a motion search range for inter prediction) acquired based on the encoding / decoding setting. Alternatively, the offset factor can be a multiple or an integer of a unit (here, assumed to be AxB) acquired from the picture division section. Alternatively, the offset factor can be a multiple (here, assumed to be ExF, such as a tile) of a unit acquired from the picture division section.
[0241] Alternatively, the offset factor can be a value less than or equal to the width and height of the unit obtained from the picture division section. In the above example, the multiple or the exponential multiple can have a value of 1. However, the present application is not limited to the above example, and the above example can be modified. For example, when the offset factor is n, Var_x can be 2xn or 2 n .
[0242] Further, a separate offset factor can be supported according to color components. An offset factor can be supported for some color components, and thus an offset factor information for other color components can be derived. For example, when an offset factor (A) for a luminance component is explicitly generated (here, it is assumed that a composition ratio of the luminance component with respect to the chrominance component is 2:1), an offset factor (A / 2) for the chrominance component can be implicitly obtained. Alternatively, when an offset factor (A) for the chrominance component is explicitly generated, an offset factor (2A) for the luminance component can be implicitly obtained.
[0243] Information on a resizing direction and a resizing value can be explicitly generated, and a resizing process can be performed according to the corresponding information. Further, the information can be implicitly determined according to an encoding / decoding setting, and a resizing process can be performed according to the determined information. At least one predetermined direction or resizing value can be assigned, and in this case, the relevant information can be omitted. In this case, the encoding / decoding setting can be determined based on characteristics, types, encoding information, etc. of an image. For example, at least one resizing direction can be predetermined according to at least one resizing operation, at least one resizing value can be predetermined according to at least one resizing operation, and at least one resizing value can be predetermined according to at least one resizing direction. Further, a resizing direction, a resizing value, etc. during an inverse resizing process can be derived from a resizing direction, a resizing value, etc. applied during a resizing process. In this case, the implicitly determined resizing value can be one of the above examples (examples in which a resizing value is differently obtained).
[0244] Further, multiplication or division has been described in the above examples, but a shift operation can be used according to an implementation of an encoder / decoder. Multiplication can be implemented by a left shift operation, and division can be implemented by a right shift operation. The description is not limited to the above examples, and can be commonly applied to the present application.
[0245] In the image resizing execution step, an image resizing can be performed based on the identified resizing information. That is, an image resizing can be performed based on information on a resizing type, a resizing operation, a resizing direction, a resizing value, etc., and an encoding / decoding can be performed based on an image after the resizing is obtained.
[0246] Further, in the image size adjustment execution step, the size adjustment can be performed using at least one data processing method. In detail, the size adjustment can be performed on the region to be size-adjusted according to the size adjustment type and the size adjustment operation by using at least one data processing method. For example, according to the size adjustment type, it can be determined how to fill data when the size adjustment is for expansion, and it can be determined how to remove data when the size adjustment is for reduction.
[0247] In summary, in the image size adjustment execution step, the image size adjustment can be performed based on the identified size adjustment information. Alternatively, in the image size adjustment execution step, the image size adjustment can be performed based on the size adjustment information and a data processing method. The two cases can differ from each other in that only the size of the image to be encoded or decoded is adjusted, or in that even the data processing of the region to be size-adjusted and the image size is considered. In the image size adjustment execution step, it can be determined whether to perform the data processing method according to the step, position, etc., at which the size adjustment process is applied. The following description focuses on an example in which the size adjustment is performed based on the data processing method, but the present application is not limited thereto.
[0248] When the size adjustment based on the offset factor is performed, various methods can be used to perform the size adjustment for expansion and the size adjustment for reduction. For expansion, the size adjustment can be performed using at least one data filling method. For reduction, the size adjustment can be performed using at least one data removal method. In this case, when the size adjustment based on the offset factor is performed, the size-adjusted region can be filled with new data or original image data directly or after modification (expansion), and the size-adjusted region can be simply removed or through a series of processes (reduction).
[0249] When performing the size adjustment based on the scale factor, in some cases (e.g., hierarchical encoding), the size adjustment for expansion can be performed by applying upsampling, and the size adjustment for reduction can be performed by applying downsampling. For example, at least one up-sampling filter can be used for expansion, and at least one down-sampling filter can be used for reduction. The filter applied horizontally can be the same as or different from the filter applied vertically. In this case, when performing the size adjustment based on the scale factor, neither new data is generated in the size-adjusted region nor new data is removed from the size-adjusted region, but the original image data can be rearranged using a method such as interpolation. The data processing method associated with the size adjustment can be classified according to the filter used for sampling. Furthermore, in some cases (e.g., cases similar to the case of the offset factor), the size adjustment for expansion can be performed using a method of padding at least one data, and the size adjustment for reduction can be performed using a method of removing at least one data. According to the present application, the following description focuses on the data processing method corresponding to the case of performing the size adjustment based on the offset factor.
[0250] In general, a predetermined data processing method can be used in the region to be size-adjusted, but at least one data processing method can also be used in the region to be size-adjusted, as in the following examples. Selection information for the data processing method can be generated. The former can mean that the size adjustment is performed by a fixed data processing method, and the latter can mean that the size adjustment is performed by an adaptive data processing method.
[0251] Furthermore, the data processing method can be applied to all regions (TL, TC, TR,..., BR in parts 7a and 7b) in the region to be added or deleted during the size adjustment or some regions (e.g., each or a combination of TL to BR in parts 7a and 7b).
[0252] Figure 8 is an example diagram of a method of constructing a region generated by expansion in an image size adjustment method according to an embodiment of the present application.
[0253] Referring to part 8a, for ease of description, the image can be divided into regions TL, TC, TR, LC, C, RC, BL, BC, and BR corresponding to the upper left, upper, upper right, left, center, right, lower left, lower, and lower right positions of the image. In the following description, the image is expanded in the "down+right" direction, but it should be understood that the description can be applied to other expansion directions.
[0254] The region added according to the expansion of the image can be constructed using various methods. For example, the region can be padded with arbitrary values, or can be padded with reference to some data of the image.
[0255] Referring to the portion 8b, the generated area (A0, A2) can be filled with an arbitrary pixel value. Various methods can be used to determine the arbitrary pixel value.
[0256] As an example, the arbitrary pixel value can be one pixel in a pixel value range that can be represented using a bit depth (e.g., from 0 to 1 « (bit_depth) - 1). For example, the arbitrary pixel value can be a minimum value, a maximum value, a median value (e.g., 1 « (bit_depth - 1), etc.) in the pixel value range, etc. (Here, bit_depth indicates a bit depth.)
[0257] As an example, the arbitrary pixel value can be one pixel in a pixel value range of pixels belonging to the image (e.g., from min P to max P ; min P and max P indicate a minimum value and a maximum value in the pixels belonging to the image; min P is greater than or equal to 0; and max P is less than or equal to 1 « (bit_depth) - 1). For example, the arbitrary pixel value can be a minimum value, a maximum value, a median value, an average value of (at least two pixels), a weighted sum, etc. in the pixel value range.
[0258] As an example, the arbitrary pixel value can be a value determined in a pixel value range belonging to a specific area included in the image. For example, when A0 is constructed, the specific area can be TR + RC + BR. Further, the specific area can be set to an area corresponding to 3 x 9 of TR, RC, and BR or an area corresponding to 1 x 9 (which is assumed to be the rightmost line). This can depend on the encoding / decoding setting. In this case, the specific area can be a unit to be divided by the picture division section. In detail, the arbitrary pixel value can be a minimum value, a maximum value, a median value, an average value of (at least two pixels), a weighted sum, etc. in the pixel value range.
[0259] Referring again to the portion 8b, the area A1 to be added as the image expands can be filled with pattern information generated using a plurality of pixel values (e.g., assuming that a pattern uses a plurality of pixels; it is not necessary to follow certain rules). In this case, the pattern information can be defined according to the encoding / decoding setting, or the relevant information can be generated. The generated area can be filled with at least one piece of pattern information.
[0260] Referring to part 8c, a region added as the image is expanded can be constructed with reference to pixels of a specific region included in the image. In detail, the added region can be constructed by copying or padding pixels (hereinafter, referred to as reference pixels) in a region adjacent to the added region. In this case, the pixels in the region adjacent to the added region can be pixels before encoding or pixels after encoding (or decoding). For example, when resizing is performed in a pre-encoding step, the reference pixels can refer to pixels of an input image, and when resizing is performed in an intra prediction reference pixel generation step, a reference picture generation step, a filtering step, etc., the reference pixels can refer to pixels of a restored image. In this example, it is assumed that the nearest neighbor pixels are used in the added region, but the present application is not limited thereto.
[0261] A region (A0) generated when the image is expanded to the left or right in association with horizontal image resizing can be constructed by horizontally padding (Z0) outer pixels adjacent to the generated region (A0), and a region (A1) generated when the image is expanded upward or downward in association with vertical image resizing can be constructed by vertically padding (Z1) outer pixels adjacent to the generated region (A1). Further, a region (A2) generated when the image is expanded downward and to the right can be constructed by diagonally padding (Z2) outer pixels adjacent to the generated region (A2).
[0262] Referring to part 8d, the generated regions (B'0 to B'2) can be constructed with reference to data of specific regions (B0 to B2) included in the image. In part 8d, unlike part 8c, a region non-adjacent to the generated region can be referred to.
[0263] For example, when there is a region having high correlation with the generated region in the image, the generated region can be filled with pixels of the region having high correlation. In this case, position information, size information, etc. of the region having high correlation can be generated. Alternatively, when there is a region having high correlation by encoding / decoding information of characteristics, types, etc. of the image and position information, size information, etc. of the region having high correlation can be implicitly checked (for example, for a 360-degree image), the generated region can be filled with data of the corresponding region. In this case, the position information, size information, etc. of the corresponding region can be omitted.
[0264] As an example, a region (B'2) generated when the image is expanded to the left or right in association with horizontal image resizing can be filled with reference to pixels in a region (B2) opposite to the region generated when the image is expanded to the left or right in association with horizontal resizing.
[0265] As an example, a pixel in region (B1) can be referred to fill region (B'1) generated when expanding the image upward or downward in association with portrait image size adjustment, region (B1) being opposite to the generated region when expanding the image upward or downward in association with portrait size adjustment.
[0266] As an example, a pixel in region (B0, TL) can be referred to fill region (B'0) generated when expanding the image by some image size adjustment (here, diagonally with respect to the image center), region (B0, TL) being opposite to the generated region.
[0267] Examples in which continuity at a boundary between both ends of an image is described and data on a region symmetrical with respect to a size adjustment direction is acquired have been described. However, the present application is not limited to this, and data of other regions (TL to BR) can be acquired.
[0268] When filling the generated region with data of a specific region of the image, data of the corresponding region can be copied as is and used to fill the generated region, or data of the corresponding region can be transformed based on characteristics, types, etc. of the image and used to fill the generated region. In this case, copying data as is can mean using pixel values of the corresponding region without any change, and performing a transformation process can mean not using pixel values of the corresponding region without any change. That is, at least one pixel value of the corresponding region can be changed by a transformation process. The generated region can be filled with the changed pixel value, or at least one of the positions of some pixels can be different from the other positions. That is, in order to fill an AxB generated region, CxD data other than AxB data of the corresponding region can be used. In other words, at least one of the motion vectors applied to the pixels used to fill the generated region can be different from the other pixels. In the above example, when a 360-degree image is composed of a plurality of faces according to a projection format, the generated region can be filled with data of the other face. The data processing method for filling a region generated when expanding the image by image size adjustment is not limited to the above-described example. The data processing method can be improved or changed, or an additional data processing method can be used.
[0269] A plurality of candidate groups for the data processing method can be supported according to an encoding / decoding setting, and information on selecting a data processing method from the plurality of candidate groups can be generated and added to a bitstream. For example, one data processing method can be selected from among a padding method by using a predetermined pixel value, a padding method by copying an external pixel, a padding method by copying a specific region of an image, a padding method by transforming a specific region of an image, etc., and related selection information can be generated. Furthermore, the data processing method can be determined implicitly.
[0270] For example, the data processing method applied to all regions (here, regions TL to BR in the portion 7a) to be generated along with the expansion through the image size adjustment can be one of a padding method by using a predetermined pixel value, a padding method by copying an external pixel, a padding method by copying a specific region of an image, a padding method by transforming a specific region of an image, and the like, and the relevant selection information can be generated. Further, one predetermined data processing method applied to the entire region can be determined.
[0271] Alternatively, the data processing method applied to each of the regions TL to BR in the portion 7a or two or more regions to be generated along with the expansion through the image size adjustment can be one of a padding method by using a predetermined pixel value, a padding method by copying an external pixel, a padding method by copying a specific region of an image, a padding method by transforming a specific region of an image, and the like, and the relevant selection information can be generated. Further, one predetermined data processing method applied to at least one region can be determined. Figure 7
[0272] Figure 9 is an example diagram of a method of constructing a region to be deleted through reduction and a region to be generated in an image size adjustment method according to an embodiment of the present application.
[0273] The region to be deleted in the image reduction processing can not only be simply removed but also removed after a series of application processing.
[0274] Referring to the portion 9a, during the image reduction processing, a specific region (A0, A1, A2) can be simply removed without additional application processing. In this case, the image (A) can be divided into the regions TL to BR as shown in the portion 8a.
[0275] Referring to the portion 9b, the regions (A0 to A2) can be removed, and can be used as reference information when the image (A) is encoded or decoded. For example, the deleted regions (A0 to A2) can be utilized during a process of restoring or correcting a specific region of the image (A) deleted through reduction. During the restoration or correction processing, a weighted sum, an average, or the like of two regions (a deleted region and a generated region) can be used. Further, the restoration or correction processing can be a process that can be applied when the two regions have a high correlation.
[0276] As an example, a region (B'2) deleted when an image is reduced to the left or right in association with a horizontal image size adjustment can be used for recovery or correction of pixels in a region (B2, LC) opposite to the region deleted when the image is reduced to the left or right in association with a horizontal size adjustment, and then the region (B'2) can be removed from the memory.
[0277] As an example, a region (B'1) deleted when an image is reduced upward or downward in association with a vertical image size adjustment can be used for encoding / decoding processing (recovery or correction processing) of a region (B1, TR) opposite to the region deleted when the image is reduced upward or downward in association with a vertical size adjustment, and then the region (B'1) can be removed from the memory.
[0278] As an example, a region (B'0) deleted when an image is reduced by some image size adjustment (here, diagonally with respect to the center of the image) can be used for encoding / decoding processing (recovery or correction processing) of a region (B0, TL) opposite to the deletion region, and then the region (B'0) can be removed from the memory.
[0279] An example in which data of a region that exists continuously at a boundary between both ends of an image and is symmetrical with respect to a size adjustment direction is used for recovery or correction has been described. However, the present application is not limited to this, and data of regions TL to BR other than the symmetrical region can be used for recovery or correction, and then can be removed from the memory.
[0280] A data processing method for removing a region to be deleted is not limited to the above-described example. The data processing method can be improved or changed, or an additional data processing method can be used.
[0281] A plurality of candidate groups for the data processing method can be supported according to an encoding / decoding setting, and related selection information can be generated and added to a bitstream. For example, one data processing method can be selected from among a method of simply removing a region to be deleted, a method of removing the region to be deleted after using the region in a series of processes, and the like, and related selection information can be generated. Furthermore, the data processing method can be determined implicitly.
[0282] For example, a data processing method applied to regions TL to BR of part 7b of the entire region to be deleted as the image is reduced by image size adjustment (here, Figure 7 The data processing method can be one of a method of simply removing a region to be deleted, a method of removing the region to be deleted after using the region in a series of processes, and the like, and related selection information can be generated. Furthermore, the data processing method can be determined implicitly.
[0283] Alternatively, the data processing method applied to each of the regions TL to BR in the part 7b to be deleted with the reduction through the image size adjustment (here, Figure 7 may be one of a method of simply removing the region to be deleted, a method of removing the region to be deleted after using the region in a series of processes, and the like, and the related selection information can be generated. Further, the data processing method can be determined implicitly.
[0284] The example of performing the size adjustment according to the size adjustment (expansion or reduction) operation has been described. In some cases, the description can be applied to the example of performing the size adjustment operation (here, expansion) and then performing the inverse size adjustment operation (here, reduction).
[0285] For example, the method of filling the region generated with the expansion with some data of the image can be selected, and then the method of removing the region to be deleted with the reduction in the inverse process after using the region in the process of restoring or correcting some data of the image can be selected. Alternatively, the method of filling the region generated with the expansion by copying the outside pixels can be selected, and then the method of simply removing the region to be deleted with the reduction in the inverse process can be selected. That is, the data processing method in the inverse process can be determined based on the data processing method selected in the image size adjustment process.
[0286] Unlike the above-described example, the data processing method of the image size adjustment process and the data processing method of the inverse process can have independent relationships. That is, the data processing method in the inverse process can be selected regardless of the data processing method selected in the image size adjustment process. For example, the method of filling the region generated with the expansion by using some data of the image can be selected, and then the method of simply removing the region to be deleted with the reduction in the inverse process can be selected.
[0287] According to the present application, the data processing method during the image size adjustment process can be determined implicitly according to the encoding / decoding setting, and the data processing method during the inverse process can be determined implicitly according to the encoding / decoding setting. Alternatively, the data processing method during the image size adjustment process can be generated explicitly, and the data processing method during the inverse process can be generated explicitly. Alternatively, the data processing method during the image size adjustment process can be generated explicitly, and the data processing method during the inverse process can be determined implicitly based on the data processing method.
[0288] Next, an example of performing image size adjustment in the encoding / decoding apparatus according to the embodiment of the present application will be described. In the following description, as an example, the size adjustment process indicates expansion, and the inverse size adjustment process indicates reduction. Further, the difference between the image before the size adjustment and the image after the size adjustment can refer to the image size, and the size adjustment related information can have some pieces explicitly generated and other pieces implicitly determined according to the encoding / decoding settings. Further, the size adjustment related information can include information on the size adjustment process and the inverse size adjustment process.
[0289] As a first example, the process of adjusting the size of the input image can be performed before starting the encoding. The size of the input image can be adjusted using the size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; the data processing method is used during the size adjustment process), and then the input image can be encoded. The image encoded data (here, the image after the size adjustment) can be stored in the memory after the encoding is completed, and can be added to the bitstream and then transmitted.
[0290] The size adjustment process can be performed before starting the decoding. The size of the image decoded data can be adjusted using the size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, etc.), and then the image decoded data can be parsed for decoding. The output image can be stored in the memory after the decoding is completed, and can be changed to the image before the size adjustment (here, using the data processing method, etc.; this is used for the inverse size adjustment process) by performing the inverse size adjustment process.
[0291] As a second example, the process of adjusting the size of the reference picture can be performed before starting the encoding. The size of the reference picture can be adjusted using the size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; the data processing method is used during the size adjustment process), and then the reference picture (here, the size-adjusted reference picture) can be stored in the memory. The image can be encoded using the size-adjusted reference picture. After the encoding is completed, the image encoded data (here, the data obtained by encoding using the reference picture) can be added to the bitstream and then transmitted. Further, when the encoded image is stored in the memory as the reference picture, the above-described size adjustment process can be performed.
[0292] Before starting decoding, a size adjustment process can be performed on the reference picture. The size of the reference picture can be adjusted using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; data processing method is used during size adjustment process), and then the reference picture (here, the size-adjusted reference picture) can be stored in the memory. The image decoding data (here, encoded by using the reference picture by the encoder) can be parsed for decoding. After decoding is completed, the output image can be generated. When the decoded image is stored in the memory as the reference picture, the above-mentioned size adjustment process can be performed.
[0293] As a third example, a size adjustment process can be performed on the image before filtering the image (here, assuming a deblocking filter) and after encoding (in detail, after encoding is completed except for the filtering process). The size of the image can be adjusted using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; data processing method is used during size adjustment), and then the image after size adjustment can be generated and then filtered. After filtering is completed, inverse size adjustment process is performed so that the image after size adjustment can be changed to the image before size adjustment.
[0294] After decoding is completed (in detail, after decoding is completed except for the filtering process), and before filtering, a size adjustment process can be performed on the image. The size of the image can be adjusted using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; data processing method is used during size adjustment), and then the image after size adjustment can be generated and then filtered. After filtering is completed, inverse size adjustment process is performed so that the image after size adjustment can be changed to the image before size adjustment.
[0295] In some cases (first example and third example), size adjustment process and inverse size adjustment process can be performed. In other cases (second example), only size adjustment process can be performed.
[0296] Further, in some cases (second and third examples), the same resizing process can be applied to the encoder and the decoder. In other cases (first example), the same or different resizing process can be applied to the encoder and the decoder. Here, the resizing process of the encoder and the decoder can differ in terms of the resizing execution step. For example, in some cases (here, the encoder), the resizing execution step considering image resizing and data processing for the resized region can be included. In other cases (here, the decoder), the resizing execution step considering image resizing can be included. Here, the former data processing can correspond to the latter data processing during the inverse resizing process.
[0297] Further, in some cases (third example), the resizing process can be applied only to the corresponding step, and the resized region can not be stored in the memory. For example, in order to use the resized region in the filtering process, the resized region can be stored in a temporary memory, filtered, and then removed through the inverse resizing process. In this case, there is no change in the size of the image due to the resizing. The present application is not limited to the above-described examples, and modifications can be made to the above-described examples.
[0298] The size of the image can be changed by the resizing process, and thus the coordinates of some pixels of the image can be changed by the resizing process. This can affect the operation of the picture partitioning section. According to the present application, by this process, block-based partitioning can be performed based on the image before the resizing or the image after the resizing. Further, unit (e.g., tile, slice, etc.) based partitioning can be performed based on the image before the resizing or the image after the resizing, which can be determined according to the encoding / decoding setting. According to the present application, the following description focuses on the case where the picture partitioning section operates based on the image after the resizing (e.g., image partitioning process after the resizing process), but other modifications can be made. The above-described examples will be described under a plurality of image settings to be described below.
[0299] The encoder can add information generated during the above-described process to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the related information from the bitstream. Further, the information can be included in the bitstream in the form of SEI or metadata.
[0300] In general, the input image can be encoded or decoded as is or after image reconstruction. For example, image reconstruction can be performed to improve image encoding efficiency, image reconstruction can be performed to consider network and user environments, and image reconstruction can be performed according to the type, characteristics, etc. of the image.
[0301] According to the present application, the image reconstruction process can include a reconstruction process alone or can include a reconstruction process and an inverse reconstruction process. The following examples will describe focusing on the reconstruction process, but the inverse reconstruction process can be derived in reverse from the reconstruction process.
[0302] Figure 10 is an example diagram of image reconstruction according to an embodiment of the present application.
[0303] The portion 10a shows an initial input image. The portions 10a to 10d are example diagrams of image rotation including predetermined angles of 0 degrees (for example, a candidate group can be generated by sampling 360 degrees into k portions; k can have values of 2, 4, 8, etc.; in this example, it is assumed that k is 4). The portions 10e to 10h are example diagrams having an inverse (or symmetric) relationship with respect to the portion 10a or with respect to the portions 10b to 10d.
[0304] The start position or scan order of the image can be changed according to the image reconstruction, but the start position and scan order can be predetermined independently of the reconstruction, which can be determined according to the encoding / decoding settings. The following embodiments assume that the start position (for example, the top-left position of the image) and scan order (for example, raster scan) are predetermined independently of the image reconstruction.
[0305] The image encoding method and image decoding method according to the embodiment of the present application can include the following image reconstruction steps. In this case, the image reconstruction process can include an image reconstruction indication step, an image reconstruction type identification step, and an image reconstruction execution step. Furthermore, the image encoding apparatus and image decoding apparatus can be configured to include an image reconstruction indication section, an image reconstruction type identification section, and an image reconstruction execution section that respectively execute the image reconstruction indication step, the image reconstruction type identification step, and the image reconstruction execution step. For encoding, the relevant syntax elements can be generated. For decoding, the relevant syntax elements can be parsed.
[0306] In the image reconstruction indication step, it can be determined whether to perform image reconstruction. For example, when a signal indicating image reconstruction (for example, convert_enabled_flag) is confirmed, reconstruction can be performed. When a signal indicating image reconstruction is not confirmed, reconstruction can not be performed, or reconstruction can be performed by confirming other encoding / decoding information. Furthermore, although a signal indicating image reconstruction is not provided, the signal indicating image reconstruction can be implicitly activated or deactivated according to the encoding / decoding settings (for example, characteristics, types, etc. of the image). When reconstruction is performed, corresponding reconstruction-related information can be generated, or corresponding reconstruction-related information can be implicitly determined.
[0307] When a signal indicating image reconstruction is provided, the corresponding signal is a signal for indicating whether to perform image reconstruction. Whether to reconstruct the corresponding image can be determined according to the signal. For example, assume a signal (e.g., convert_enabled_flag) indicating image reconstruction is confirmed. When the corresponding signal is activated (e.g., convert_enabled_flag = 1), reconstruction can be performed. When the corresponding signal is deactivated (e.g., convert_enabled_flag = 0), reconstruction can not be performed.
[0308] In addition, when a signal indicating image reconstruction is not provided, reconstruction can not be performed, or whether to reconstruct the corresponding image can be determined through another signal. For example, reconstruction can be performed according to the characteristics, type, etc. of the image (e.g., 360-degree image), and reconstruction information can be explicitly generated or can be designated as a predetermined value. The present application is not limited to the above-described example, and the above-described example can be modified.
[0309] In the image reconstruction type identification step, the image reconstruction type can be identified. The image reconstruction type can be defined by a reconstruction method, reconstruction mode information, etc. The reconstruction method (e.g., convert_type_flag) can include flipping, rotation, etc., and the reconstruction mode information can include a mode of the reconstruction method (e.g., convert_mode). In this case, the reconstruction-related information can consist of the reconstruction method and the mode information. That is, the reconstruction-related information can consist of at least one syntax element. In this case, the number of candidate sets of the mode information can be the same or different depending on the reconstruction method.
[0310] As an example, rotation can include candidates having uniform intervals (here, 90 degrees), as shown in parts 10a to 10d. Part 10a shows 0-degree rotation, part 10b shows 90-degree rotation, part 10c shows 180-degree rotation, and part 10d shows 270-degree rotation (here, it is measured clockwise).
[0311] As an example, flipping can include candidates as shown in parts 10a, 10e, and 10f. When part 10a shows no flipping, parts 10e and 10f show horizontal flipping and vertical flipping, respectively.
[0312] In the above example, the setting for rotation having uniform intervals and the setting for flipping have been described. However, this is only an example of image reconstruction, and the present application is not limited thereto, and the present application can include another interval difference, another flipping operation, etc. that can be determined according to encoding / decoding settings.
[0313] Alternatively, comprehensive information generated by mixing a reconstruction method and corresponding mode information can be included (e.g., convert_com_flag). In this case, reconstruction-related information can consist of a reconstruction method and mode information mixing.
[0314] For example, the comprehensive information can include candidates as shown in parts 10a to 10f, which can be examples of 0-degree rotation, 90-degree rotation, 180-degree rotation, 270-degree rotation, horizontal flipping, and vertical flipping with respect to part 10a.
[0315] Alternatively, the comprehensive information can include candidates as shown in parts 10a to 10h, which can be examples of 0-degree rotation, 90-degree rotation, 180-degree rotation, 270-degree rotation, horizontal flipping, vertical flipping, 90-degree rotation and then horizontal flipping (or horizontal flipping and then 90-degree rotation), and 90-degree rotation and then vertical flipping (or vertical flipping and then 90-degree rotation) or 0-degree rotation, 90-degree rotation, 180-degree rotation, 270-degree rotation, horizontal flipping, 180-degree rotation and then horizontal flipping (or horizontal flipping and then 180-degree rotation), 90-degree rotation and then horizontal flipping (or horizontal flipping and then 90-degree rotation), and 270-degree rotation and then horizontal flipping (or horizontal flipping and then 270-degree rotation).
[0316] The candidate group can be configured to include a rotation mode, a flipping mode, and a combined mode of rotation and flipping. The combined mode can simply include mode information in a reconstruction method, and can include a mode generated by mixing mode information in each method. In this case, the combined mode can include a mode generated by mixing at least one mode (e.g., rotation) of some methods and at least one mode (e.g., flipping) of other methods. In the above example, the combined mode includes a case where one mode of some methods is combined with multiple modes of some methods (here, 90-degree rotation + multiple flipping / horizontal flipping + multiple rotation). The information of the mixed construction can include a case where reconstruction is not applied (here, part 10a) as a candidate group, and can include a case where reconstruction is not applied as a first candidate group (e.g., #0 is designated as an index).
[0317] Alternatively, the information of the mixed construction can include mode information corresponding to a predetermined reconstruction method. In this case, the reconstruction-related information can consist of mode information corresponding to the predetermined reconstruction method. That is, information about a reconstruction method can be omitted, and the reconstruction-related information can consist of one syntax element associated with the mode information.
[0318] For example, the reconstruction-related information can be configured to include rotation-specific candidates as shown in parts 10a to 10d. Alternatively, the reconstruction-related information can be configured to include flip-specific candidates as shown in parts 10a, 10e, and 10f.
[0319] The image before the image reconstruction processing and the image after the image reconstruction processing can have the same size or at least one different length, which can be determined according to the encoding / decoding setting. The image reconstruction processing can be a processing of rearranging pixels in the image (here, inverse pixel rearrangement processing is performed during the inverse image reconstruction processing; this can be derived inversely according to the pixel rearrangement processing), and thus can change the position of at least one pixel. The pixel rearrangement can be performed according to a rule based on the image reconstruction type information.
[0320] In this case, the pixel rearrangement can be affected by the size and shape (e.g., square or rectangular) of the image. In detail, the width and height of the image before the reconstruction processing and the width and height of the image after the reconstruction processing can serve as variables during the pixel rearrangement processing.
[0321] For example, at least one of the ratio information (e.g., former / latter or latter / former) regarding the ratio of the width of the image before the reconstruction processing to the width of the image after the reconstruction processing, the ratio of the width of the image before the reconstruction processing to the height of the image after the reconstruction processing, the ratio of the height of the image before the reconstruction processing to the width of the image after the reconstruction processing, and the ratio of the height of the image before the reconstruction processing to the height of the image after the reconstruction processing can serve as a variable during the pixel rearrangement processing.
[0322] In this example, when the image before the reconstruction processing and the image after the reconstruction processing have the same size, the ratio of the width of the image to the height of the image can serve as a variable during the pixel rearrangement processing. Further, when the image is in a square shape, the ratio of the length of the image before the reconstruction processing to the length of the image after the reconstruction processing can serve as a variable during the pixel rearrangement processing.
[0323] In the image reconstruction performing step, the image reconstruction can be performed based on the identified reconstruction information. That is, the image reconstruction can be performed based on the information regarding the reconstruction type, the reconstruction mode, and the like, and the encoding / decoding can be performed based on the reconstructed image.
[0324] Next, an example of performing the image reconstruction in the encoding / decoding apparatus according to the embodiment of the present application will be described.
[0325] The process of reconstructing the input image can be performed before starting encoding. The reconstruction can be performed using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), and the reconstructed image can be encoded. The image encoding data can be stored in a memory after the encoding is completed, and can be added to a bitstream and then transmitted.
[0326] The reconstruction process can be performed before starting decoding. The reconstruction can be performed using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), and the image decoding data can be parsed for decoding. The image can be stored in a memory after the decoding is completed, and can be changed into the image before the reconstruction and then output by performing an inverse reconstruction process.
[0327] The encoder can add information generated during the above-described processes to a bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the related information from the bitstream. Furthermore, the information can be included in the bitstream in the form of SEI or metadata.
[0328] [Table 1]
[0329]
[0330] Table 1 represents example syntax elements associated with partitioning among image settings. The following description will focus on additional syntax elements. Furthermore, in the following examples, the syntax elements are not limited to any particular unit, and can be supported in various units such as a sequence, a picture, a slice, and a tile. Alternatively, the syntax elements can be included in SEI, metadata, etc. Furthermore, the type, order, condition, etc. of the syntax elements supported in the following examples are limited to the example, and thus can be changed and determined according to encoding / decoding settings.
[0331] In Table 1, tile_header_enabled_flag represents a syntax element indicating whether encoding / decoding settings for a tile are supported. When the syntax element is activated (tile_header_enabled_flag = 1), encoding / decoding settings in a tile unit can be provided. When the syntax element is deactivated (tile_header_enabled_flag = 0), encoding / decoding settings in a tile unit cannot be provided, and encoding / decoding settings in an upper unit can be assigned.
[0332] Further, tile_coded_flag indicates a syntax element indicating whether a tile is encoded or decoded. When the syntax element is activated (tile_coded_flag = 1), the corresponding tile can be encoded or decoded. When the syntax element is deactivated (tile_coded_flag = 0), the corresponding tile cannot be encoded or decoded. Here, not performing encoding can mean that encoding data is not generated for the corresponding tile (here, it is assumed that the corresponding region is processed by a predetermined rule or the like; a meaningless region in some projection formats applicable to 360-degree images). Not performing decoding means that the decoded data in the corresponding tile is no longer parsed (here, it is assumed that the corresponding region is processed by a predetermined rule). Further, no longer parsing the decoded data can mean that there is no encoding data in the corresponding unit and thus parsing is no longer performed, and can also mean that even if the encoding data exists, parsing is no longer performed by the flag. The header information of the tile unit can be supported according to whether the tile is encoded or decoded.
[0333] The above examples focus on tiles. However, the present application is not limited to tiles, and the above description can be modified and then applied to other division units of the present application. Further, the examples of the tile division setting are not limited to the above-described cases, and the above-described cases can be modified.
[0334] [Table 2]
[0335]
[0336] Table 2 indicates example syntax elements associated with reconstruction among image settings.
[0337] Referring to Table 2, convert_enabled_flag indicates a syntax element indicating whether reconstruction is performed. When the syntax element is activated (convert_enabled_flag = 1), a reconstructed image is encoded or decoded, and additional reconstruction-related information can be checked. When the syntax element is deactivated (convert_enabled_flag = 0), an original image is encoded or decoded.
[0338] Further, convert_type_flag indicates mixed information about reconstruction methods and mode information. One method can be determined from a plurality of candidate groups of a method of applying rotation, a method of applying flipping, and a method of applying rotation and flipping.
[0339] [Table 3]
[0340]
[0341] Table 3 indicates example syntax elements associated with resizing among image settings.
[0342] Referring to Table 3, pic_width_in_samples and pic_height_in_samples denote syntax elements indicating the width and height of the picture. The size of the picture can be checked through the syntax elements.
[0343] In addition, img_resizing_enabled_flag denotes a syntax element indicating whether image size adjustment is performed. When the syntax element is activated (img_resizing_enabled_flag = 1), the picture is encoded or decoded after the size adjustment, and additional size adjustment-related information can be checked. When the syntax element is deactivated (img_resizing_enabled_flag = 0), the original picture is encoded or decoded. In addition, the syntax element can indicate size adjustment for intra prediction.
[0344] In addition, resizing_met_flag indicates a size adjustment method. One size adjustment method can be determined from a candidate group such as a scaling factor-based size adjustment method (resizing_met_flag = 0), an offset factor-based size adjustment method (resizing_met_flag = 1), etc.
[0345] In addition, resizing_mov_flag denotes a syntax element for a size adjustment operation. For example, one of expansion and reduction can be determined.
[0346] In addition, width_scale and height_scale denote scaling factors associated with horizontal size adjustment and vertical size adjustment of the scaling factor-based size adjustment.
[0347] In addition, top_height_offset and bottom_height_offset denote an "up" direction offset factor and a "down" direction offset factor associated with horizontal size adjustment of the offset factor-based size adjustment, and left_width_offset and right_width_offset denote a "left" direction offset factor and a "right" direction offset factor associated with vertical size adjustment of the offset factor-based size adjustment.
[0348] The size of the picture after the size adjustment can be updated through the size adjustment-related information and the picture size information.
[0349] Further, the resizing_type_flag indicates a syntax element indicating a data processing method for the resized region. The number of candidate groups of the data processing method can be the same or different according to the resizing method and the resizing operation.
[0350] The image setting processing applied to the above-described image encoding / decoding apparatus can be performed individually or in combination. The following examples will focus on examples in which a plurality of image setting processes are performed in combination.
[0351] Figure 11 FIG. 11 is an example diagram showing images before and after the image setting processing according to the embodiment of the present application. In detail, part 11a shows an example before performing image reconstruction on the divided image (e.g., an image projected during 360-degree image encoding), and part 11b shows an image after performing image reconstruction on the divided image (e.g., an image packed during 360-degree image encoding). That is, it can be understood that part 11a is an example diagram before performing the image setting processing, and part 11b is an example diagram after performing the image setting processing.
[0352] In this example, image division (here, it is assumed that tiles) and image reconstruction are described as the image setting processing.
[0353] In the following examples, image reconstruction is performed after performing image division. However, according to the encoding / decoding setting, image division can be performed after performing image reconstruction, and can be modified. Further, the above-described image reconstruction processing (including inverse processing) can be applied identically or similarly to the reconstruction processing in the division unit in the image in the present embodiment.
[0354] Image reconstruction can or can not be performed in all division units in the image, and image reconstruction can be performed in some division units. Therefore, the division unit before reconstruction (e.g., some of P0 to P5) can be the same as or different from the division unit after reconstruction (e.g., some of S0 to S5). Through the following examples, various image reconstruction cases will be described. Further, for ease of description, it is assumed that the unit of the image is a picture, the unit of the divided image is a tile, and the division unit is a rectangular shape.
[0355] As an example, it can be determined in some units whether to perform image reconstruction (e.g., sps_convert_enabled_flag or SEI or metadata or the like). Alternatively, it can be determined in some units whether to perform image reconstruction (e.g., pps_convert_enabled_flag). This can be allowed when first occurring in the respective unit (here: picture) or activated in a higher unit (e.g., sps_convert_enabled_flag = 1). Alternatively, it can be determined in some units whether to perform image reconstruction (e.g., tile_convert_flag[i]; i is a partition unit index). This can be allowed when first occurring in the respective unit (here: tile) or activated in a higher unit (e.g., pps_convert_enabled_flag = 1). Furthermore, it can be determined implicitly whether to perform image reconstruction depending on the encoding / decoding setup, in part, so that the related information can be omitted.
[0356] As an example, it can be determined whether to reconstruct a partition unit in an image depending on a signal indicating image reconstruction (e.g., pps_convert_enabled_flag). In detail, it can be determined whether to reconstruct all partition units in an image depending on the signal. In this case, a single signal indicating image reconstruction can be generated in the image.
[0357] As an example, it can be determined whether to reconstruct a partition unit in an image depending on a signal indicating image reconstruction (e.g., tile_convert_flag[i]). In detail, it can be determined whether to reconstruct some partition units in an image depending on the signal. In this case, at least one signal indicating image reconstruction (e.g., a number of signals equal to the number of partition units) can be generated.
[0358] As an example, it can be determined whether to reconstruct an image depending on a signal indicating image reconstruction (e.g., pps_convert_enabled_flag), and it can be determined whether to reconstruct a partition unit in an image depending on a signal indicating image reconstruction (e.g., tile_convert_flag[i]). In detail, when any signal is activated (e.g., pps_convert_enabled_flag = 1), any other signal (e.g., tile_convert_flag[i]) can be additionally checked, and it can be determined whether to reconstruct some partition units in an image depending on the signal (here: tile_convert_flag[i]). In this case, multiple signals indicating image reconstruction can be generated.
[0359] When a signal indicating image reconstruction is activated, image reconstruction-related information can be generated. In the following examples, various image reconstruction-related information will be described.
[0360] As an example, reconstruction information applied to an image can be generated. In detail, one piece of reconstruction information can be used as reconstruction information for all partition units in an image.
[0361] As an example, reconstruction information applied to a partition unit in an image can be generated. In detail, at least one piece of reconstruction information can be used as reconstruction information for some partition units in an image. That is, one piece of reconstruction information can be used as reconstruction information for one partition unit, or one piece of reconstruction information can be used as reconstruction information for a plurality of partition units.
[0362] The following examples will be described in connection with an example of performing image reconstruction.
[0363] For example, when a signal indicating image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information commonly applied to partition units in an image can be generated. Alternatively, when a signal indicating image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information individually applied to partition units of an image can be generated. Alternatively, when a signal indicating image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information individually applied to partition units in an image can be generated. Alternatively, when a signal indicating image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information commonly applied to partition units in an image can be generated.
[0364] Reconstruction information can be processed implicitly or explicitly according to encoding / decoding settings. For implicit processing, reconstruction information can be designated as a predetermined value according to characteristics, types, etc. of an image.
[0365] P0 to P5 in part 11a can correspond to S0 to S5 in part 11b, and reconstruction processing can be performed on partition units. For example, P0 can not be reconstructed, and then P0 can be assigned to S0. P1 can be rotated by 90 degrees, and then can be assigned to S1. P2 can be rotated by 180 degrees, and then can be assigned to S2. P3 can be horizontally flipped, and then can be assigned to S3. P4 can be rotated by 90 degrees and horizontally flipped, and then can be assigned to S4. P5 can be rotated by 180 degrees and horizontally flipped, and then can be assigned to S5.
[0366] However, the present invention is not limited to the above examples, and various modifications can be made to the above examples. Similar to the above examples, it is possible not to reconstruct the partitioning units in the image, or at least one of reconstruction using rotation, reconstruction using flipping, and reconstruction using a combination of rotation and flipping can be performed.
[0367] When image reconstruction is applied to partitioning units, additional reconstruction processing, such as rearranging the partitioning units, can be performed. That is, the image reconstruction processing according to the invention can be configured to include rearranging the partitioning units in the image as well as rearranging the pixels in the image, and can be represented using some of the syntax elements in Table 4 (e.g., part_top, part_left, part_width, part_height, etc.). This means that image partitioning processing and image reconstruction processing can be understood in combination. In the example above, the image has been described as being divided into multiple units.
[0368] P0 to P5 in part 11a can correspond to S0 to S5 in part 11b, and a reconstruction process can be performed on the partitioned units. For example, P0 can be left unreconstructed and then assigned to S0. P1 can be left unreconstructed and then assigned to S2. P2 can be rotated 90 degrees and then assigned to S1. P3 can be horizontally flipped and then assigned to S4. P4 can be rotated 90 degrees and horizontally flipped and then assigned to S5. P5 can be horizontally flipped and then rotated 180 degrees and then assigned to S3. The invention is not limited thereto and various modifications are possible.
[0369] also, Figure 7 P_Width and P_Height can correspond to Figure 11 P_Width and P_Height, and Figure 7 P'_Width and P'_Height can correspond to Figure 11 P'_Width and P'_Height. Figure 7 The dimensions of the resized image, P'_Width × P'_Height, can be expressed as (P_Width + Exp_L + Exp_R) × (P_Height + Exp_T + Exp_B), and Figure 11 The size P'_Width x P'_Height of the image after the medium size adjustment can be expressed as (P_Width + Var0_L + Var1_L + Var2_L + Var0_R + Var1_R + Var2_R) x (P_Height + Var0_T + Var1_T + Var0_B + Var1_B) or (Sub_P0_Width + Sub_P1_Width + Sub_P2_Width + Var0_L + Var1_L + Var2_L + Var0_R + Var1_R + Var2_R) x (Sub_P0_Height + Sub_P1_Height + Var0_T + Var1_T + Var0_B + Var1_B).
[0370] Similar to the above example, for the image reconstruction, the rearrangement of the pixels in the partition unit of the image can be performed, the rearrangement of the partition units in the image can be performed, and both the rearrangement of the pixels in the partition unit of the image and the rearrangement of the partition units in the image can be performed. In this case, the rearrangement of the partition units in the image can be performed after the rearrangement of the pixels in the partition unit is performed, or the rearrangement of the pixels in the partition unit can be performed after the rearrangement of the partition units in the image is performed.
[0371] Whether to perform the rearrangement of the partition units in the image can be determined according to the signal indicating the image reconstruction. Alternatively, the signal for rearranging the partition units in the image can be generated. In detail, when the signal indicating the image reconstruction is activated, the signal can be generated. Alternatively, the signal can be processed implicitly or explicitly according to the encoding / decoding setting. For the implicit processing, the signal can be determined according to the characteristics, type, etc. of the image.
[0372] Further, the information about the rearrangement of the partition units in the image can be performed implicitly or explicitly according to the encoding / decoding setting, and can be determined according to the characteristics, type, etc. of the image. That is, each partition unit can be arranged according to the arrangement information determined in advance for the partition unit.
[0373] Next, an example of reconstructing the partition units in the image in the encoding / decoding apparatus according to the embodiment of the present application will be described.
[0374] The partitioning processing can be performed on the input image using the partitioning information before the encoding is started. The reconstruction processing can be performed on the partition unit using the reconstruction information, and the image reconstructed for each partition unit can be encoded. The image encoding data can be stored in the memory after the encoding is completed, and can be added to the bitstream and then transmitted.
[0375] The partitioning information can be used to perform a partitioning process before starting decoding. The reconstruction information can be used to perform a reconstruction process on the partition units, and the image decoding data can be parsed to be decoded in the reconstructed partition units. After completing the decoding, the image decoding data can be stored in a memory, and the plurality of partition units can be merged into a single unit after performing an inverse reconstruction process in the partition units, so that the image can be output.
[0376] Figure 12 is an example of adjusting the size of each partition unit of an image according to an embodiment of the present application. Figure 12 P0 to P5 of Figure 11 P0 to P5 of Figure 12 S0 to S5 of Figure 11 S0 to S5 of
[0377] In the following examples, the description will focus on the case where image size adjustment is performed after performing image partitioning. However, depending on the encoding / decoding settings, image partitioning can be performed after image size adjustment, and can be modified. Furthermore, the above-mentioned image size adjustment process (including inverse process) can be applied identically or similarly to the image partition unit size adjustment process in the present embodiment.
[0378] For example, Figure 7 TL to BR of Figure 12 TL to BR of the partition units SX (S0 to S5) of Figure 7 S0 and S1 of Figure 12 PX and SX of Figure 7 P_Width and P_Height of Figure 12 Sub_PX_Width and Sub_PX_Height of Figure 7 P'_Width and P'_Height of Figure 12 Sub_SX_Width and Sub_SX_Height of Figure 7 Exp_L, Exp_R, Exp_T, and Exp_B of Figure 12 VarX_L, VarX_R, VarX_T, and VarX_B of
[0379] The process of adjusting the size of the partition units in the adjustment image in parts 12a to 12f corresponds to Figure 7 The difference between the image expansion or reduction in the parts 7a and 7b can be that the setting for the image expansion or reduction can exist in proportion to the number of division units. Further, the process of adjusting the size of the division unit in the image can be different from the image expansion or reduction in terms of having a setting that is commonly or individually applied to the division unit in the image. In the following examples, the cases of various size adjustment will be described, and the size adjustment process can be performed in consideration of the above-described cases.
[0380] According to the present application, the image size adjustment can or can not be performed on all division units in the image, and the image size adjustment can be performed on some division units. The various image size adjustment cases will be described through the following examples. Further, for convenience of description, it is assumed that the size adjustment operation is for expansion, the size adjustment operation is based on an offset factor, the size adjustment direction is the "up" direction, the "down" direction, the "left" direction, and the "right" direction, the size adjustment direction is set to operate by the size adjustment information, the unit of the image is a picture, and the unit that divides the image is a tile.
[0381] As an example, it can be determined whether to perform the image size adjustment in some units (e.g., sps_img_resizing_enabled_flag or SEI or metadata, etc.). Alternatively, it can be determined whether to perform the image size adjustment in some units (e.g., pps_img_resizing_enabled_flag). This can be allowed when it first occurs in the corresponding unit (here, a picture) or is activated in a higher unit (e.g., sps_img_resizing_enabled_flag = 1). Alternatively, it can be determined whether to perform the image size adjustment in some units (e.g., tile_resizing_flag[i]; i is a division unit index). This can be allowed when it first occurs in the corresponding unit (here, a tile) or is activated in a higher unit. Further, in part, it can be implicitly determined whether to perform the image size adjustment according to the encoding / decoding setting, and thus the related information can be omitted.
[0382] As an example, it can be determined whether to adjust the size of the division unit in the image according to a signal indicating the image size adjustment (e.g., pps_img_resizing_enabled_flag). In detail, it can be determined whether to adjust the size of all division units in the image according to the signal. In this case, a single signal indicating the image size adjustment can be generated.
[0383] As an example, it can be determined whether to resize the size of the partition units in the picture according to a signal indicating picture resizing (e.g., tile_resizing_flag[i]). In detail, it can be determined whether to resize the size of some of the partition units in the picture according to the signal. In this case, at least one signal indicating picture resizing (e.g., a number of signals equal to the number of partition units) can be generated.
[0384] As an example, it can be determined whether to resize the picture size according to a signal indicating picture resizing (e.g., pps_img_resizing_enabled_flag), and it can be determined whether to resize the size of the partition units in the picture according to a signal indicating picture resizing (e.g., tile_resizing_flag[i]). In detail, when any signal is activated (e.g., pps_img_resizing_enabled_flag = 1), any other signal (e.g., tile_resizing_flag[i]) can be additionally checked, and it can be determined whether to resize the size of some of the partition units in the picture according to the signal (here, tile_resizing_flag[i]). In this case, a plurality of signals indicating picture resizing can be generated.
[0385] When a signal indicating picture resizing is activated, picture resizing-related information can be generated. In the following examples, various picture resizing-related information will be described.
[0386] As an example, resizing information applied to the picture can be generated. In detail, one piece of resizing information or a set of resizing information can be used as the resizing information for all of the partition units in the picture. For example, one piece of resizing information commonly applied to the "up" direction, the "down" direction, the "left" direction, and the "right" direction of the partition units in the picture (or a resizing value applied to all of the resizing directions supported or allowed in the partition units; one piece of information in this example) or a set of resizing information separately applied to the "up" direction, the "down" direction, the "left" direction, and the "right" direction (or a number of pieces of resizing information equal to the number of resizing directions allowed or supported in the partition units; up to four pieces of information in this example) can be generated.
[0387] As an example, the size adjustment information applied to the division unit in the image can be generated. In detail, at least one piece of size adjustment information or a set of size adjustment information can be used as the size adjustment information of all division units in the image. That is, one piece of size adjustment information or a set of size adjustment information can be used as the size adjustment information of one division unit or as the size adjustment information of a plurality of division units. For example, the size adjustment information commonly applied to the "up" direction, "down" direction, "left" direction, and "right" direction of one division unit in the image can be generated, or a set of size adjustment information individually applied to the "up" direction, "down" direction, "left" direction, and "right" direction can be generated. Alternatively, the size adjustment information commonly applied to the "up" direction, "down" direction, "left" direction, and "right" direction of a plurality of division units in the image can be generated, or a set of size adjustment information individually applied to the "up" direction, "down" direction, "left" direction, and "right" direction can be generated. The configuration of the size adjustment group means the size adjustment value information with respect to at least one size adjustment direction.
[0388] In summary, the size adjustment information commonly applied to the division unit in the image can be generated. Alternatively, the size adjustment information individually applied to the division unit in the image can be generated. The following examples will be described in connection with an example of performing image size adjustment.
[0389] For example, when a signal (e.g., pps_img_resizing_enabled_flag) indicating image size adjustment is activated, the size adjustment information commonly applied to the division unit in the image can be generated. Alternatively, when a signal (e.g., pps_img_resizing_enabled_flag) indicating image size adjustment is activated, the size adjustment information individually applied to the division unit in the image can be generated. Alternatively, when a signal (e.g., tile_resizing_flag[i]) indicating image size adjustment is activated, the size adjustment information individually applied to the division unit in the image can be generated. Alternatively, when a signal (e.g., tile_resizing_flag[i]) indicating image size adjustment is activated, the size adjustment information commonly applied to the division unit in the image can be generated.
[0390] The size adjustment direction of the image, the size adjustment information, etc. can be processed implicitly or explicitly according to the encoding / decoding setting. For implicit processing, the size adjustment information can be designated as a predetermined value according to the characteristics, type, etc. of the image.
[0391] It has been described that the resizing direction in the resizing process of the present application can be at least one of an "up" direction, a "down" direction, a "left" direction, and a "right" direction, and the resizing direction and the resizing information can be processed explicitly or implicitly. That is, the resizing value (including 0; this means no resizing) can be predetermined implicitly for some directions, and the resizing value (including 0; this means no resizing) can be specified explicitly for other directions.
[0392] Even in the division unit in the image, the resizing direction and the resizing information can be set to be processed implicitly or explicitly, and this can be applied to the division unit in the image. For example, a setting applied to one division unit in the image (here, a setting equal to the number of division units can occur) can occur, a setting applied to a plurality of division units in the image can occur, or a setting applied to all division units in the image (here, one setting can occur) can occur, and at least one setting (for example, one to a number of settings equal to the number of division units) can occur in the image. The setting information applied to the division unit in the image can be collected, and one group of settings can be defined.
[0393] Figure 13 is an example diagram of a setting or a group of resizing of a division unit in an image.
[0394] In detail, Figure 13 Various examples of processing the resizing direction and the resizing information for the division unit in the image implicitly or explicitly are shown. In the following examples, for ease of description, the implicit processing assumes that the resizing value for some resizing directions is 0.
[0395] As shown in part 13a, when the boundary of the division unit matches the boundary of the image (here, the thick solid line), the resizing can be processed explicitly, and when the boundary of the division unit does not match the boundary of the image (thin solid line), the resizing can be processed implicitly. For example, P0 can be resized in the "up" direction and the "left" direction (a2, a0), P1 can be resized in the "up" direction (a2), P2 can be resized in the "up" direction and the "right" direction (a2, a1), P3 can be resized in the "down" direction and the "left" direction (a3, a0), P4 can be resized in the "down" direction (a3), and P5 can be resized in the "down" direction and the "right" direction (a3, a1). In this case, resizing in other directions can not be allowed.
[0396] As shown in part 13b, some directions of the partition unit (here, upward and downward) can allow explicit handling of resizing, and some directions of the partition unit (here, leftward and rightward) can allow explicit handling of resizing when the boundaries of the partition unit match the boundaries of the image (here, thick solid lines), and can allow implicit handling of resizing when the boundaries of the partition unit do not match the boundaries of the image (here, thin solid lines). For example, P0 can be resized in the "upward" direction, the "downward" direction, and the "leftward" direction (b2, b3, b0), P1 can be resized in the "upward" direction and the "downward" direction (b2, b3), P2 can be resized in the "upward" direction, the "downward" direction, and the "rightward" direction (b2, b3, b1), P3 can be resized in the "upward" direction, the "downward" direction, and the "leftward" direction (b3, b4, b0), P4 can be resized in the "upward" direction and the "downward" direction (b3, b4), and P5 can be resized in the "upward" direction, the "downward" direction, and the "rightward" direction (b3, b4, b1). In this case, resizing can not be allowed in other directions.
[0397] As shown in part 13c, some directions of the partition unit (here, leftward and rightward) can allow explicit handling of resizing, and some directions of the partition unit (here, upward and downward) can allow explicit handling of resizing when the boundaries of the partition unit match the boundaries of the image (here, thick solid lines), and can allow implicit handling of resizing when the boundaries of the partition unit do not match the boundaries of the image (here, thin solid lines). For example, P0 can be resized in the "upward" direction, the "leftward" direction, and the "rightward" direction (c4, c0, c1), P1 can be resized in the "upward" direction, the "leftward" direction, and the "rightward" direction (c4, c1, c2), P2 can be resized in the "upward" direction, the "leftward" direction, and the "rightward" direction (c4, c2, c3), P3 can be resized in the "downward" direction, the "leftward" direction, and the "rightward" direction (c5, c0, c1), P4 can be resized in the "downward" direction, the "leftward" direction, and the "rightward" direction (c5, c1, c2), and P5 can be resized in the "downward" direction, the "leftward" direction, and the "rightward" direction (c5, c2, c3). In this case, resizing can not be allowed in other directions.
[0398] Settings related to image resizing similar to the above-described example can have various cases. Multiple sets of settings are supported so that a set selection of settings can be explicitly generated, or a predetermined set of settings can be implicitly determined according to encoding / decoding settings (e.g., characteristics, types, etc. of the image).
[0399] Figure 14 is an example diagram showing both the processing of adjusting the size of an image and the processing of adjusting the size of a division unit in an image.
[0400] Referring to Figure 14 , the processing of adjusting the size of an image and the inverse processing can be performed in the directions e and f, and the processing of adjusting the size of a division unit in an image and the inverse processing can be performed in the directions d and g. That is, the size adjustment processing can be performed on an image, and then the size adjustment processing can be performed on a division unit in the image. The size adjustment order can not be fixed. This means that multiple size adjustment processes are possible.
[0401] In summary, the image size adjustment processing can be classified into the size adjustment of an image (or adjusting the size of an image before division) and the size adjustment of a division unit in an image (or adjusting the size of an image after division). Either or both of the size adjustment of an image and the size adjustment of a division unit in an image can not be performed, or can be performed, which can be determined according to the encoding / decoding settings (e.g., characteristics, types, etc. of an image).
[0402] When multiple size adjustment processes are performed in this example, the size adjustment of an image can be performed in at least one of the "upward” direction, the "downward” direction, the "leftward” direction, and the "rightward” direction of an image, and the size of at least one division unit in an image can be adjusted. In this case, the size adjustment can be performed in at least one of the "upward” direction, the "downward” direction, the "leftward” direction, and the "rightward” direction of a division unit whose size is to be adjusted.
[0403] Referring to Figure 14 , the size of an image before size adjustment (A) can be defined as P_Width x P_Height, the size of an image after primary size adjustment (or an image before secondary size adjustment; B) can be defined as P'_Width x P'_Height, and the size of an image after secondary size adjustment (or an image after final size adjustment; C) can be defined as P''_Width x P''_Height. The image before size adjustment (A) indicates an image on which no size adjustment is performed, the image after primary size adjustment (B) indicates an image on which some size adjustment is performed, and the image after secondary size adjustment (C) indicates an image on which all size adjustment is performed. For example, the image after primary size adjustment (B) can indicate an image on which size adjustment is performed by a division unit of an image as shown in parts 13a to 13c, and the image after secondary size adjustment (C) can indicate an image on which size adjustment is performed by a division unit of an image as shown in parts 14a to 14c. Figure 7 the image (B) after the primary size adjustment. The opposite case is also possible. However, the present application is not limited to the above-described example, and various modifications can be made to the above-described example.
[0404] In the size of the image (B) after the primary size adjustment, P'_Width can be acquired by at least one horizontal size adjustment value of the P_Width and the lateral adjustment size, and P'_Height can be acquired by at least one vertical size adjustment value of the P_Height and the longitudinal adjustment size. In this case, the size adjustment value can be a size adjustment value generated in the division unit.
[0405] In the size of the image (C) after the secondary size adjustment, P''_Width can be acquired by at least one horizontal size adjustment value of the P'_Width and the lateral adjustment size, and P''_Height can be acquired by at least one vertical size adjustment value of the P'_Height and the longitudinal adjustment size. In this case, the size adjustment value can be a size adjustment value generated in the image.
[0406] In summary, the size of the image after the size adjustment can be acquired by at least one size adjustment value and the size of the image before the size adjustment.
[0407] In the size-adjusted region of the image, information on a data processing method can be generated. Various data processing methods will be described through the following examples. The data processing method generated during the inverse size adjustment processing can be applied identically or similarly to the data processing method of the size adjustment processing. The data processing methods in the size adjustment processing and the inverse size adjustment processing will be described through various combinations to be described below.
[0408] As an example, a data processing method applied to the image can be generated. In detail, one data processing method or a set of data processing methods can be used as the data processing method of all division units in the image (here, it is assumed that the sizes of all division units are to be adjusted). For example, one data processing method commonly applied to the "up" direction, the "down" direction, the "left" direction, and the "right" direction of the division units in the image (or a data processing method applied to all size adjustment directions supported or allowed in the division units, etc.; in this example, one piece of information) or a set of data processing methods applied to the "up" direction, the "down" direction, the "left" direction, and the "right" direction (or a number of data processing methods equal to the number of size adjustment directions supported or allowed in the division units; in this example, up to four pieces of information) can be generated.
[0409] As an example, a data processing method that is applied to a partition unit in an image can be generated. In detail, at least one data processing method or a set of data processing methods can be used as a data processing method for some partition units in an image (here, it is assumed that the size of the partition units is to be adjusted). That is, one data processing method or a set of data processing methods can be used as a data processing method for one partition unit or a data processing method for a plurality of partition units. For example, one data processing method that is commonly applied to an "up" direction, a "down" direction, a "left" direction, and a "right" direction of one partition unit in an image can be generated, or a set of data processing methods that is individually applied to the "up" direction, the "down" direction, the "left" direction, and the "right" direction can be generated. Alternatively, one data processing method that is commonly applied to an "up" direction, a "down" direction, a "left" direction, and a "right" direction of a plurality of partition units in an image can be generated, or a set of data processing methods that is individually applied to the "up" direction, the "down" direction, the "left" direction, and the "right" direction can be generated. The configuration of the set of data processing methods means a data processing method for at least one size adjustment direction.
[0410] In summary, a data processing method that is commonly applied to a partition unit in an image can be used. Alternatively, a data processing method that is individually applied to a partition unit in an image can be used. The data processing method can use a predetermined method. A predetermined data processing method can be provided as at least one method. This corresponds to implicit processing, and selection information for the data processing method can be explicitly generated, which can be determined according to an encoding / decoding setting (e.g., characteristics, type, etc. of an image).
[0411] That is, a data processing method that is commonly applied to a partition unit in an image can be used. A predetermined method can be used, or one of a plurality of data processing methods can be selected. Alternatively, a data processing method that is individually applied to a partition unit in an image can be used. Depending on the partition unit, a predetermined method can be used, or one of a plurality of data processing methods can be selected.
[0412] In the following examples, some cases of adjusting the size of a partition unit in an image (here, it is assumed that the size adjustment is for expansion) will be described (here, the size-adjusted region is filled with some data of an image).
[0413] Data of specific regions tl to br of some units P0 to P5 (in parts 12a to 12f) can be used to adjust the size of specific regions TL to BR of some units (e.g., S0 to S5 in parts 12a to 12f). In this case, some units can be the same as each other (e.g., S0 and P0) or different from each other (e.g., S0 and P1). That is, regions TL to BR whose size is to be adjusted can be filled with some data tl to br of the corresponding partition unit, and can be filled with some data of a partition unit other than the corresponding partition unit.
[0414] As an example, data tl to br of the current partition unit can be used to adjust the size of regions TL to BR whose size of the current partition unit is adjusted. For example, TL of S0 can be filled with data tl of P0, RC of S1 can be filled with data tr+rc+br of P1, BL+BC of S2 can be filled with data bl+bc+br of P2, and TL+LC+BL of S3 can be filled with data tl+lc+bl of P3.
[0415] As an example, data tl to br of a partition unit that is spatially adjacent to the current partition unit can be used to adjust the size of regions TL to BR whose size of the current partition unit is adjusted. For example, in the "up" direction, TL+TC+TR of S4 can be filled with data bl+bc+br of P1, in the "down" direction, BL+BC of S2 can be filled with data tl+tc+tr of P5, in the "left" direction, LC+BL of S2 can be filled with data tl+rc+bl of P1, in the "right" direction, RC of S3 can be filled with data tl+lc+bl of P4, and in the "down+left" direction, BR of S0 can be filled with data tl of P4.
[0416] As an example, data tl to br of a partition unit that is not spatially adjacent to the current partition unit can be used to adjust the size of regions TL to BR whose size of the current partition unit is adjusted. For example, data in a boundary region (e.g., horizontal, vertical, etc.) between both ends of the image can be acquired. LC of S3 can be acquired using data tr+rc+br of S5, RC of S2 can be acquired using data tl+lc of S0, BC of S4 can be acquired using data tc+tr of S1, and TC of S1 can be acquired using data bc of S4.
[0417] Alternatively, data of a certain region of the image (a region that is not adjacent to the region to be resized in space but is determined to have a high correlation with the region to be resized) can be acquired. BC of S1 can be acquired using data tl+lc+bl of S3, RC of S3 can be acquired using data tl+tc of S1, and RC of S5 can be acquired using data bc of S0.
[0418] Further, some cases of adjusting the size of a divided unit in an image (here, assume that the size adjustment is for reduction) are as follows (here, removal is performed by using recovery or correction of some data of the image).
[0419] A certain region TL to BR of some units (for example, S0 to S5 in parts 12a to 12f) can be used for recovery or correction processing of a certain region tl to br of some units P0 to P5. In this case, some units can be the same as each other (for example, S0 and P0) or different (for example, S0 and P2). That is, a region to be resized can be used to recover some data of a corresponding divided unit and then removed, and a region to be resized can be used to recover some data of a divided unit other than the corresponding divided unit and then removed. Detailed examples can be derived in reverse from the extension processing, and thus will be omitted.
[0420] This example can be applied to a case where data with a high correlation exists in a region to be resized, and information on a position referred to for size adjustment can be generated explicitly or acquired implicitly according to a predetermined rule. Alternatively, the relevant information can be checked in combination. This can be an example that can be applied in a case where data is acquired from another region having continuity in encoding of a 360-degree image.
[0421] Next, an example of adjusting the size of a divided unit in an image in an encoding / decoding apparatus according to an embodiment of the present application will be described.
[0422] A division processing can be performed on an input image before starting encoding. A size adjustment processing can be performed on a divided unit using size adjustment information, and the image can be encoded after the size of the divided unit is adjusted. Image encoding data can be stored in a memory after encoding is completed, and can be added to a bitstream and then transmitted.
[0423] A division processing can be performed using division information before starting decoding. A size adjustment processing can be performed on a divided unit using size adjustment information, and the image decoding data can be parsed to be decoded in a divided unit that is resized. After decoding is completed, the image decoding data can be stored in a memory, and a plurality of divided units can be merged into a single unit after inverse size adjustment processing on the divided unit is performed, so that an image can be output.
[0424] Another example of the above-described image size adjustment processing can be applied. The present application is not limited thereto, and modifications can be made thereto.
[0425] In the image setting processing, it can be allowed to combine the image size adjustment and the image reconstruction. The image reconstruction can be performed after the image size adjustment is performed. Alternatively, the image size adjustment can be performed after the image reconstruction is performed. Further, it can be allowed to combine the image division, the image reconstruction, and the image size adjustment. The image size adjustment and the image reconstruction can be performed after the image division is performed. The order of the image setting is not fixed and can be changed, which can be determined according to the encoding / decoding setting. In this example, the image setting processing is described as performing the image reconstruction and the image size adjustment after the image division is performed. However, depending on the encoding / decoding setting, another order is possible, and modifications can be made thereto.
[0426] For example, the image setting processing can be performed in the following order: division -> reconstruction; reconstruction -> division; division -> size adjustment; size adjustment -> division; size adjustment -> reconstruction; reconstruction -> size adjustment; division -> reconstruction -> size adjustment; division -> size adjustment -> reconstruction; size adjustment -> division -> reconstruction; size adjustment -> reconstruction -> division; reconstruction -> division -> size adjustment; and reconstruction -> size adjustment -> division, and the combination with the additional image setting can be possible. As described above, the image setting processing can be sequentially performed, but some or all of the setting processing can be performed simultaneously. Further, as some of the image setting processing, a plurality of processing can be performed according to the encoding / decoding setting (e.g., the characteristics, the type, etc. of the image). The following examples indicate various combinations of the image setting processing.
[0427] As an example, P0 to P5 in the section 11a can correspond to S0 to S5 in the section 11b, and the reconstruction processing (here, rearrangement of pixels) and the size adjustment processing (here, adjusting the size of the division unit to have the same size) can be performed in the division unit. For example, the size of P0 to P5 can be adjusted based on the offset, and P0 to P5 can be allocated to S0 to S5. Further, P0 can not be reconstructed, and then P0 can be allocated to S0. P1 can be rotated by 90 degrees, and then can be allocated to S1. P2 can be rotated by 180 degrees, and then can be allocated to S2. P3 can be rotated by 270 degrees, and then can be allocated to S3. P4 can be horizontally flipped, and then can be allocated to S4. P5 can be vertically flipped, and then can be allocated to S5.
[0428] As an example, P0 to P5 in the section 11a can correspond to the same or different positions as S0 to S5 in the section 11b, and the reconstruction processing (here, rearrangement of pixels and division units) and the resizing processing (here, resizing of division units to have the same size) can be performed in the division units. For example, the sizes of P0 to P5 can be adjusted based on the scale, and P0 to P5 can be assigned to S0 to S5. Further, P0 can not be reconstructed, and then P0 can be assigned to S0. P1 can not be reconstructed, and then P1 can be assigned to S2. P2 can be rotated by 90 degrees, and then can be assigned to S1. P3 can be horizontally flipped, and then can be assigned to S4. P4 can be rotated by 90 degrees and horizontally flipped, and then can be assigned to S5. P5 can be horizontally flipped and then rotated by 180 degrees, and then can be assigned to S3.
[0429] As an example, P0 to P5 in the section 11a can correspond to E0 to E5 in the section 5e, and the reconstruction processing (here, rearrangement of pixels and division units) and the resizing processing (here, resizing of division units to have different sizes) can be performed in the division units. For example, P0 can not be resized and reconstructed and then can be assigned to E0, P1 can be resized based on the scale but not reconstructed and then can be assigned to E1, P2 can not be resized but reconstructed and then can be assigned to E2, P3 can be resized based on the offset but not reconstructed and then can be assigned to E4, P4 can not be resized but reconstructed and can be assigned to E5, and P5 can be resized based on the offset and reconstructed and then can be assigned to E3.
[0430] Similar to the above example, the absolute positions or the relative positions of the division units in the image before and after the image setting processing can be maintained or changed, which can be determined according to the encoding / decoding settings (e.g., characteristics, types, etc. of the image). Further, various combinations of the image setting processing can be possible. The present application is not limited thereto, and various modifications can be made thereto.
[0431] The encoder can add the information generated during the above processing to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder can parse the relevant information from the bitstream. Further, the information can be included in the bitstream in the form of SEI or metadata.
[0432] [Table 4]
[0433]
[0434] Table 4 represents example syntax elements associated with multiple image partitioning. The following description will focus on additional syntax elements. Also, in the following examples, syntax elements are not limited to any particular unit, and can be supported in various units such as sequence, picture, slice, and tile. Alternatively, syntax elements can be included in SEI, metadata, etc.
[0435] Referring to Table 4, parts_enabled_flag represents a syntax element indicating whether to partition some units. When the syntax element is activated (parts_enabled_flag = 1), an image can be partitioned into multiple units, and the multiple units can be encoded or decoded. Also, additional partitioning information can be checked. When the syntax element is deactivated (parts_enabled_flag = 0), an original image is encoded or decoded. In this example, the description will focus on rectangular partitioning units such as tiles, and different settings for existing tiles and partitioning information can be provided.
[0436] Here, num_partitions refers to a syntax element indicating the number of partitioned units, and num_partitions plus 1 is equal to the number of partitioned units.
[0437] Also, part_top[i] and part_left[i] refer to syntax elements indicating position information of partitioned units, and represent horizontal start positions and vertical start positions (e.g., top-left positions of partitioned units) of the partitioned units. Also, part_width[i] and part_height[i] refer to syntax elements indicating size information of partitioned units, and represent widths and heights of the partitioned units. In this case, start positions and size information can be set in units of pixels or in units of blocks. Also, the syntax elements can be syntax elements that can be generated during an image reconstruction process or syntax elements that can be generated when an image partitioning process and an image reconstruction process are combinedly constructed.
[0438] Also, part_header_enabled_flag represents a syntax element indicating whether to support encoding / decoding settings for partitioned units. When the syntax element is activated (part_header_enabled_flag = 1), encoding / decoding settings for partitioned units can be provided. When the syntax element is deactivated (part_header_enabled_flag = 0), encoding / decoding settings cannot be provided, and encoding / decoding settings for upper units can be assigned.
[0439] The above example is not limited to the example of the syntax elements associated with the resizing and reconstruction in the partition unit among the image settings, and can be modified with respect to other partition units and settings of the present application. The example has been described assuming that the resizing and reconstruction are performed after the partitioning is performed, but the present application is not limited thereto, and can be modified in other image setting orders, etc. Further, the type, order, condition, etc. of the syntax elements supported in the following example are limited to the example, and thus can be changed and determined according to the encoding / decoding settings.
[0440] [Table 5]
[0441]
[0442] Table 5 represents example syntax elements associated with the reconstruction in the partition unit among the image settings.
[0443] Referring to Table 5, part_convert_flag[i] represents a syntax element indicating whether to reconstruct the partition unit. The syntax element can be generated with respect to each partition unit. When the syntax element is activated (part_convert_flag[i]=1), the reconstructed partition unit can be encoded or decoded, and additional reconstruction-related information can be checked. When the syntax element is deactivated (part_convert_flag[i]=0), the original partition unit is encoded or decoded. Here, convert_type_flag[i] refers to mode information about the reconstruction of the partition unit, and can be information about the pixel rearrangement.
[0444] Further, a syntax element indicating additional reconstruction such as partition unit rearrangement can be generated. In the example, the partition unit rearrangement can be performed by part_top and part_left, which are syntax elements indicating the above-described image partitioning, or a syntax element (e.g., index information) associated with the partition unit rearrangement can be generated.
[0445] [Table 6]
[0446]
[0447] Table 6 represents example syntax elements associated with the resizing in the partition unit among the image settings.
[0448] Referring to Table 6, part_resizing_flag[i] indicates a syntax element indicating whether to adjust the size of a partition unit in an image. This syntax element can be generated for each partition unit. When this syntax element is activated (part_resizing_flag[i] = 1), the resized partition unit can be encoded or decoded after the size adjustment, and additional size adjustment-related information can be checked. When this syntax element is deactivated (part_resizing_flag[i] = 0), the original partition unit is encoded or decoded.
[0449] In addition, width_scale[i] and height_scale[i] indicate scale factors associated with horizontal and vertical size adjustment based on scale factors in a partition unit.
[0450] In addition, top_height_offset[i] and bottom_height_offset[i] indicate offset factors for "up" and "down" directions associated with offset factor-based size adjustment in a partition unit, and left_width_offset[i] and right_width_offset[i] indicate offset factors for "left" and "right" directions associated with offset factor-based size adjustment in a partition unit.
[0451] In addition, resizing_type_flag[i][j] indicates a syntax element indicating a data processing method for a resized region in a partition unit. This syntax element indicates a separate data processing method for a size adjustment direction. For example, syntax elements indicating separate data processing methods for resized regions in "up," "down," "left," and "right" directions can be generated. This syntax element can be generated based on size adjustment information (e.g., which can be generated only when size adjustment is performed in some directions).
[0452] The above-described image setting process can be a process applied according to the characteristics, types, etc. of an image. In the following examples, even if not specifically mentioned, the above-described image setting process can be applied without any change or with any change. In the following examples, the description will focus on cases of addition or change in the above-described examples.
[0453] For example, a 360-degree image or an omnidirectional image generated by a 360-degree camera has characteristics different from those of an image acquired by a general camera, and has an encoding environment different from that of compression encoding of a general image.
[0454] Unlike general images, a 360-degree image can have no boundary portion having discontinuity, and data of all regions of the 360-degree image can have continuity. In addition, a device such as an HMD can require a high-definition image because the image should be replayed in front of the eyes through a lens. When an image is acquired through a stereoscopic camera device, the amount of image data processed can increase. Various image setting processes considering a 360-degree image can be performed to provide an efficient encoding environment including the above-described examples.
[0455] A 360-degree camera device can be a plurality of camera devices or a camera device having a plurality of lenses and sensors. The camera device or the lens can cover all directions around any center point captured by the camera device.
[0456] A 360-degree image can be encoded using various methods. For example, a 360-degree image can be encoded using various image processing algorithms in a 3D space, and a 360-degree image can be converted into a 2D space and encoded using various image processing algorithms. According to the present application, the following description will focus on a method of converting a 360-degree image into a 2D space and encoding or decoding the converted image.
[0457] A 360-degree image encoding apparatus according to an embodiment of the present application can include some or all of the elements shown in FIG. 1, and can further include a pre-processing unit configured to pre-process an input image (stitching, projection, region-wise packing). Meanwhile, a 360-degree image decoding apparatus according to an embodiment of the present application can include some or all of the elements shown in FIG. 2, and can further include a post-processing unit configured to post-process an encoded image to reproduce an output image before decoding the encoded image. Figure 1 Figure 2
[0458] In other words, an encoder can pre-process an input image, encode the pre-processed image, and transmit a bitstream including the image, and a decoder can parse, decode, and post-process the transmitted bitstream to generate an output image. In this case, the transmitted bitstream can include information generated during a pre-processing process and information generated during an encoding process, and the bitstream can be parsed and used during a decoding process and a post-processing process.
[0459] Subsequently, an operation method for a 360-degree image encoder will be described in more detail, and an operation method for a 360-degree image decoder can be easily derived by those skilled in the art because the operation method for the 360-degree image decoder is opposite to the operation method for the 360-degree image encoder, and thus a detailed description of the operation method for the 360-degree image decoder will be omitted.
[0460] The input image can undergo stitching and projection processing performed on a sphere-based 3D projection structure, and can project image data on the 3D projection structure into a 2D image through the processing.
[0461] The projection image can be configured to include some or all of the 360-degree content according to an encoding setting. In this case, position information of an area (or pixel) to be placed at the center of the projection image can be generated as a predetermined value implicitly, or the position information of the area (or pixel) can be generated explicitly. Also, when the projection image includes a specific area of the 360-degree content, range information and position information of the included area can be generated. Also, range information (e.g., width and height) and position information (e.g., which is measured based on the upper left end of the image) of a region of interest (ROI) can be generated according to the projection image. In this case, a specific area having high importance in the 360-degree content can be set as the ROI. The 360-degree image can allow viewing of all content in the "up" direction, the "down" direction, the "left" direction, and the "right" direction, but the user's gaze can be limited to a portion of the image, which can be set as the ROI in consideration of the limitation. For the purpose of efficient encoding, the ROI can be set to have good quality and high resolution, and other areas can be set to have lower quality and lower resolution than the ROI.
[0462] Among a plurality of 360-degree image transmission schemes, a single stream transmission scheme can allow transmission of a full image or a viewport image in a single bitstream for a user alone. A multi-stream transmission scheme can allow transmission of several full images having different image qualities in a plurality of bitstreams, and thus can select an image quality according to a user environment and a communication condition. A tiled stream transmission scheme can allow transmission of separately encoded partial images based on a tile unit in a plurality of bitstreams, and thus can select a tile according to a user environment and a communication condition. Accordingly, the 360-degree image encoder can generate and transmit bitstreams having two or more qualities, and the 360-degree image decoder can set an ROI according to a user's view, and can selectively decode the bitstream according to the ROI. That is, a place pointed by a user's gaze can be set as the ROI through a head tracking or eye tracking system, and only a necessary portion can be presented.
[0463] The projection image can be converted into a packed image obtained by performing a region-wise packing process. The region-wise packing process can include a step of dividing the projection image into a plurality of regions, and can arrange (or rearrange) the divided regions in the packed image according to a region-wise packing setting. When the 360-degree image is converted into a 2D image (or a projection image), the region-wise packing can be performed to increase spatial continuity. Accordingly, the size of the image can be reduced by the region-wise packing. Further, the region-wise packing can be performed to reduce degradation of image quality caused during rendering, to implement viewport-based projection, and to provide other types of projection formats. The region-wise packing can or can not be performed according to an encoding setting, which can be determined based on a signal (e.g., regionwise_packing_flag; information about the region-wise packing can be generated only when the regionwise_packing_flag is activated) indicating whether the region-wise packing is performed.
[0464] When the region-wise packing is performed, setting information (or mapping information) assigning (or arranging) a certain region of the projection image to a certain region of the packed image can be displayed (or generated). When the region-wise packing is not performed, the projection image and the packed image can be the same image.
[0465] In the above description, the stitching process, the projection process, and the region-wise packing process are defined as separate processes, but some (e.g., stitching + projection, projection + region-wise packing) or all (e.g., stitching + projection + region-wise packing) of the processes can be defined as a single process.
[0466] At least one packed image can be generated from the same input image according to settings of the stitching process, the projection process, and the region-wise packing process. Further, at least one encoding data for the same projection image can be generated according to a setting of the region-wise packing process.
[0467] The packed image can be divided by performing a tiling process. In this case, the tiling, which is a process of dividing an image into a plurality of regions and then transmitting, can be an example of a 360-degree image transmission scheme. As described above, the tiling can be performed for partial decoding in consideration of a user environment, and the tiling can also be performed for efficiently processing a large amount of data of a 360-degree image. For example, when an image is composed of one unit, the entire image can be decoded to decode an ROI. On the other hand, when an image is composed of a plurality of unit regions, it can be efficient to decode only an ROI. In this case, the division can be performed in units of tiles, which are division units according to a conventional encoding scheme, or can be performed in units of various division units (e.g., a quadrangular division tile, etc.) that have been described according to the present application. Further, the division unit can be a unit for performing independent encoding / decoding. The tiling can be performed independently or based on a projection image or a packed image. That is, the division can be performed based on a face boundary of a projection image, a face boundary of a packed image, a packing setting, etc., and the division can be independently performed for each division unit. This can affect the generation of division information during the tiling process.
[0468] Next, the projection image or the packed image can be encoded. The encoding data and information generated during the pre-processing process can be added to a bitstream, and the bitstream can be transmitted to a 360-degree image decoder. The information generated during the pre-processing process can be added to the bitstream in the form of SEI or metadata. In this case, the bitstream can include at least one encoding data having a different setting for a part of an encoding process, and at least one piece of pre-processing information having a different setting for a part of the pre-processing process. This is to construct a decoded image in combination with a plurality of encoding data (encoding data + pre-processing information) according to a user environment. In detail, the decoded image can be constructed by selectively combining a plurality of encoding data. Further, the process can be performed in a case where the process is divided into two parts to be applied to a binocular system, and the process can be performed on another depth image.
[0469] Figure 15 is an example diagram showing a 2D planar space and a 3D space showing a 3D image.
[0470] In general, three degrees of freedom (3DoF) can be required for the purpose of 360-degree 3D virtual space, and three rotations can be supported with respect to an X-axis (pitch), a Y-axis (yaw), and a Z-axis (roll). DoF refers to a degree of freedom in space, 3DoF refers to a degree of freedom including rotations about the X-axis, the Y-axis, and the Z-axis as shown in part 15a, and 6DoF refers to a degree of freedom that additionally allows movements along the X-axis, the Y-axis, and the Z-axis as well as 3DoF. The following description will focus on the image encoding apparatus and the image decoding apparatus of the present application having 3DoF. When 3DoF or more (3DoF+) is supported, the image encoding apparatus and the image decoding apparatus can be modified or combined with additional processes or apparatuses not shown.
[0471] Referring to part 15a, yaw can have a range from -π (-180 degrees) to π (180 degrees), pitch can have a range from -π / 2 radian (or -90 degrees) to π / 2 radian (or 90 degrees), and roll can have a range from -π / 2 radian (or -90 degrees) to π / 2 radian (or 90 degrees). In this case, when Φ and θ are assumed to be longitude and latitude in a map representation of the Earth, 3D space coordinates (x, y, z) can be transformed from 2D space coordinates (Φ, θ). For example, 3D space coordinates can be derived from 2D space coordinates according to the transformation formulas x = cos(θ)cos(Φ), y = sin(θ), and z = -cos(θ)sin(Φ).
[0472] Further, (Φ, θ) can be transformed into (x, y, z). For example, 3D space coordinates can be derived from 2D space coordinates according to the transformation formulas Φ = tan -1 (-Z / X) and θ = sin -1 (Y / (X 2 +Y 2 +Z 2 ) 1 / 2 .
[0473] When a pixel in the 3D space is accurately transformed into a pixel in the 2D space (e.g., an integer unit pixel in the 2D space), the pixel in the 3D space can be mapped to the pixel in the 2D space. When a pixel in the 3D space is not accurately transformed into a pixel in the 2D space (e.g., a decimal unit pixel in the 2D space), a pixel obtained by interpolation can be mapped to the 2D pixel. In this case, as the interpolation, a nearest neighbor interpolation, a bilinear interpolation, a B-spline interpolation, a bicubic interpolation, or the like can be used. In this case, the relevant information can be explicitly generated by selecting one of a plurality of interpolation candidates, or the interpolation method can be implicitly determined according to a predetermined rule. For example, a predetermined interpolation filter can be used according to the 3D model, the projection format, the color format, and the slice / tile type. Also, when the interpolation information is explicitly generated, information on filter information (e.g., filter coefficients) can be included.
[0474] Part 15b illustrates an image in which the 3D space is transformed into the 2D space (2D plane coordinate system). (Φ, θ) can be sampled (i, j) based on the size (width and height) of the image. Here, i can have a range from 0 to P_Width-1, and j can have a range from 0 to P_Height-1.
[0475] (Φ, θ) can be a center point (or reference point; depicted as Figure 15 the point of C; coordinates (Φ, θ) = (0, 0)) for arranging the 360-degree image with respect to the projection image. The setting of the center point can be specified in the 3D space, and the position information of the center point can be explicitly generated or implicitly determined as a predetermined value. For example, the center position information in yaw, the center position information in pitch, the center position information in roll, or the like can be generated. When the value of the information is not separately specified, it can be assumed that each value is zero.
[0476] The above has described an example of transforming the entire 360-degree image from the 3D space to the 2D space, but a specific region of the 360-degree image can be transformed, and the position information (e.g., some positions belonging to the region; in this example, the position information about the center point), the range information, or the like of the specific region can be explicitly generated or the position information, the range information, or the like of the specific region can implicitly follow the predetermined position and range information. For example, the center position information in yaw, the center position information in pitch, the center position information in roll, the range information in yaw, the range information in pitch, the range information in roll, or the like can be generated, and the specific region can be at least one region. Thus, the position information, the range information, or the like of a plurality of regions can be processed. When the value of the information is not separately specified, it can be assumed that the entire 360-degree image.
[0477] H0 to H6 and W0 to W5 in part 15a indicate some latitudes and longitudes in part 15b, which can be expressed as coordinates (C, j) and (i, C) in part 15b (C is a longitude or latitude component). Unlike a general image, when a 360-degree image is converted into a 2D space, distortion can occur or warping of content in the image can occur. This can depend on the area of the image, and different encoding / decoding settings can be applied to the location of the image or areas divided according to the location. When encoding / decoding settings are adaptively applied based on encoding / decoding information in the present application, location information (for example, an x component, a y component, or a range defined by x and y) can be included as an example of the encoding / decoding information.
[0478] The descriptions of 3D space and 2D space are defined to help describe embodiments of the present application. However, the present application is not limited thereto, and the above descriptions can be modified in details or can be applied to other cases.
[0479] As described above, an image acquired through a 360-degree camera can be transformed into a 2D space. In this case, a 360-degree image can be mapped using a 3D model, and various 3D models such as a sphere, a cube, a cylinder, a pyramid, and a polyhedron can be used. When a 360-degree image mapped based on a model is transformed into a 2D space, projection processing can be performed according to a model-based projection format.
[0480] Figures 16a to 16d is a conceptual diagram showing a projection format according to an embodiment of the present application.
[0481] Figure 16a An equirectangular projection (ERP) format in which a 360-degree image is projected into a 2D plane is shown. Figure 16b A cube map projection (CMP) format in which a 360-degree image is projected into a cube is shown. Figure 16c An octahedron projection (OHP) format in which a 360-degree image is projected into an octahedron is shown. Figure 16d An icosahedron projection (ISP) format in which a 360-degree image is projected into a polyhedron is shown. However, the present application is not limited thereto, and various projection formats can be used. In Figures 16a to 16d In, the left side shows a 3D model, and the right side shows an example transformed into a 2D space through projection processing. Various sizes and shapes can be provided according to a projection format. Each shape can be composed of a surface or a face, and each face can be expressed as a circle, a triangle, a quadrangle, or the like.
[0482] In the present application, a projection format can be defined by a 3D mode, a face setting (e.g., a number of faces, a shape of faces, a shape configuration of faces, etc.), a projection processing setting, etc. When at least one element is different in definition, a projection format can be regarded as a different projection format. For example, ERP consists of a sphere model (3D model), one face (a number of faces), and a quadrangular face (a shape of faces). However, when some of settings of a projection processing (e.g., a formula used during a transformation from a 3D space to a 2D space; i.e., an element having the same remaining projection settings and producing a difference in at least one pixel of a projection image in a projection processing) are different, the format can be classified as different formats, e.g., ERP1 and ERP2. As another example, CMP consists of a cube model, six faces, and a quadrangular face. When some of settings during a projection processing (e.g., a sampling method applied during a transformation from a 3D space to a 2D space) are different, the format can be classified as different formats, e.g., CMP1 and CMP2.
[0483] When a plurality of projection formats is used instead of one predetermined projection format, projection format identification information (or projection format information) can be explicitly generated. The projection format identification information can be configured by various methods.
[0484] As an example, a projection format can be identified by assigning index information (e.g., proj_format_flag) to a plurality of projection formats. For example, #0 can be assigned to ERP, #1 can be assigned to CMP, #2 can be assigned to OHP, #3 can be assigned to ISP, #4 can be assigned to ERP1, #5 can be assigned to CMP1, #6 can be assigned to OHP1, #7 can be assigned to ISP1, #8 can be assigned to CMP compact, #9 can be assigned to OHP compact, #10 can be assigned to ISP compact, and #11 or higher can be assigned to other formats.
[0485] As an example, a projection format can be identified using at least one piece of element information constituting the projection format. In this case, as the element information constituting the projection format, 3D model information (e.g., 3d_model_flag; #0 indicates a sphere, #1 indicates a cube, #2 indicates a cylinder, #3 indicates a pyramid, #4 indicates a polyhedron 1, and #5 indicates a polyhedron 2), face number information (e.g., num_face_flag; a method of increasing 1 from 1; the number of faces generated in the projection format is designated as index information, i.e., #0 indicates one, #1 indicates three, #2 indicates six, #3 indicates eight, and #4 indicates twenty), face shape information (e.g., shape_face_flag; #0 indicates a quadrangle, #1 indicates a circle, #2 indicates a triangle, #3 indicates a quadrangle+circle, and #4 indicates a quadrangle+triangle), projection processing setting information (e.g., 3d_2d_convert_idx), and the like can be included.
[0486] As an example, a projection format can be identified using element information constituting the projection format and projection format index information. For example, as the projection format index information, #0 can be assigned to ERP, #1 can be assigned to CMP, #2 can be assigned to OHP, #3 can be assigned to ISP, and #4 or more can be assigned to other formats. The projection format (e.g., ERP, ERP1, CMP, CMP1, OHP, OHP1, ISP, and ISP1) can be identified together with the element information constituting the projection format (here, projection processing setting information). Alternatively, the projection format (e.g., ERP, CMP, CMP compact, OHP, OHP compact, ISP, and ISP compact) can be identified together with the element information constituting the projection format (here, region-wise packing).
[0487] In summary, a projection format can be identified using projection format index information, a projection format can be identified using at least one piece of projection format element information, and a projection format can be identified using projection format index information and at least one piece of projection format element information. This can be defined according to the encoding / decoding setting. In the present disclosure, the following description assumes that a projection format is identified using a projection format index. In this example, the description will focus on a projection format expressed using faces having the same size and shape, but a configuration of faces having different sizes and shapes is possible. In addition, the configuration of each face can be designated as index information, and the configuration of each face can be designated as element information constituting the projection format. Figures 16a to 16d The number of each face is used as a symbol for identifying the corresponding face, and there is no limitation to a particular order, regardless of the configuration shown in the middle. For convenience of description, the following description assumes that, for a projection image, the ERP is a projection format including one face + a quadrangle, the CMP is a projection format including six faces + a quadrangle, the OHP is a projection format including eight faces + a triangle, the ISP is a projection format including twenty faces + a triangle, and the faces have the same size and shape. However, the description can be applied identically or similarly even for different settings.
[0488] As shown in Figures 16a to 16d , the projection format can be classified as one face (e.g., ERP) or multiple faces (e.g., CMP, OHP, and ISP). In addition, the shape of each face can be classified as a quadrangle, a triangle, etc. The classification can be an example according to the type, characteristics, etc. of the image according to the present application, which can be applied when providing different encoding / decoding settings according to the projection format. For example, the type of the image can be a 360-degree image, and the characteristics of the image can be one of the classifications (e.g., each projection format, a projection format having one face or multiple faces, a projection format having a quadrangular face or a non-quadrangular face).
[0489] A 2D plane coordinate system (e.g., (i, j)) can be defined in each face of the 2D projection image, and the characteristics of the coordinate system can differ according to the projection format, the position of each face, etc. The ERP can have one 2D plane coordinate system, and the other projection formats can have multiple 2D plane coordinate systems according to the number of faces. In this case, the coordinate system can be expressed as (k, i, j), and k can indicate index information of each face.
[0490] Figure 17 is a conceptual diagram showing that the projection format according to the embodiment of the present application is included in a rectangular image.
[0491] That is, it can be understood that parts 17a to 17c show Figures 16b to 16d that the projection format of
[0492] Referring to parts 17a to 17c, each image format can be configured in a rectangular shape to encode or decode a 360-degree image. For the ERP, a single coordinate system can be used as it is. However, for the other projection formats, the coordinate systems of the faces can be integrated into a single coordinate system, and a detailed description thereof will be omitted.
[0493] Referring to parts 17a to 17c, it can be confirmed that a region filled with meaningless data such as white or background is generated when a rectangular image is constructed. That is, the rectangular image can be composed of a region including actual data (here, a face; a valid region) and a meaningless region (here, it is assumed that the region is filled with any pixel value; an invalid region) added to construct the rectangular image. This can degrade performance due to an increase in coded data, that is, an increase in image size caused by coding / decoding of the meaningless region as well as the actual image data.
[0494] Accordingly, a process of constructing an image by excluding the meaningless region and using the region including the actual data can be additionally performed.
[0495] Figure 18 is a conceptual diagram of a method of converting a projection format into a rectangular shape, that is, a method of performing rearrangement on a face to exclude a meaningless region, according to an embodiment of the present application.
[0496] Referring to parts 18a to 18c, examples for rearranging parts 17a to 17c can be confirmed, and the process can be defined as a region-wise packing process (CMP compact, OHP compact, ISP compact, etc.). In this case, a face can not only be rearranged but also divided and then rearranged (OHP compact, ISP compact, etc.). This can be performed to remove a meaningless region and improve coding performance through efficient face arrangement. For example, when images are arranged continuously between faces (for example, B2-B3-B1, B5-B0-B4, etc. in part 18a), prediction accuracy at the time of coding is enhanced, and thus coding performance can be enhanced. Here, the region-wise packing according to the projection format is merely an example, and the present application is not limited thereto.
[0497] Figure 19 is a conceptual diagram illustrating a region-wise packing process performed to convert a CMP projection format into a rectangular image, according to an embodiment of the present application.
[0498] Referring to parts 19a to 19c, a CMP projection format can be arranged as 6x1, 3x2, 2x3, and 1x6. Further, when the size of some faces is adjusted, arrangement can be made as shown in parts 19d and 19e. In parts 19a to 19e, CMP is applied as an example. However, the present application is not limited thereto, and other projection formats can be applied. Arrangement of faces of an image obtained through region-wise packing can follow a predetermined rule corresponding to the projection format, or information about the arrangement can be explicitly generated.
[0499] A 360-degree image encoding and decoding apparatus according to an embodiment of the present application can be configured to include Figure 1 and Figure 2 Some or all elements of the image encoding and decoding apparatuses shown in FIGS. 1 and 2. In particular, the format conversion section configured to convert the projection format and the inverse format conversion section configured to inversely convert the projection format can also be included in the image encoding apparatus and the image decoding apparatus, respectively. That is, the input image can be processed by the format conversion section, and then encoded by the image encoding apparatus, and the bitstream can be decoded by the image decoding apparatus and then processed by the inverse format conversion section to generate the output image. The following description will focus on the processing performed by the encoder (here, the input image, encoding, etc.), and the processing performed by the decoder can be derived in reverse from the encoder. In addition, the redundant description of the foregoing will be omitted. Figure 1 Figure 2 The following description assumes that the input image is the same as the packed image or the 2D projection image acquired by performing a pre-processing procedure by the 360-degree encoding apparatus. That is, the input image can be an image acquired by performing projection processing or area-wise packing processing according to some projection format. The projection format applied in advance to the input image can be one of various projection formats, which can be regarded as a common format and referred to as a first format.
[0500] The format conversion section can perform conversion to a projection format other than the first format. In this case, the projection format to which conversion is to be performed can be referred to as a second format. For example, ERP can be set as the first format, and ERP can be converted to a second format (e.g., ERP2, CMP, OHP, and ISP). In this case, ERP2 includes an EPR format of the following kind: the EPR format has the same conditions such as a 3D model and a face configuration, but has some different settings. Alternatively, the projection format can be the same format having the same projection format settings (e.g., ERP = ERP2), and can have different image sizes or resolutions. Alternatively, some of the following image setting processing can be applied. For ease of description, such examples have been mentioned, but each of the first format and the second format can be one of various projection formats. However, the present application is not limited thereto, and modifications can be made thereto.
[0501] The format conversion section can perform conversion to a projection format other than the first format. In this case, the projection format to which conversion is to be performed can be referred to as a second format. For example, ERP can be set as the first format, and ERP can be converted to a second format (e.g., ERP2, CMP, OHP, and ISP). In this case, ERP2 includes an EPR format of the following kind: the EPR format has the same conditions such as a 3D model and a face configuration, but has some different settings. Alternatively, the projection format can be the same format having the same projection format settings (e.g., ERP = ERP2), and can have different image sizes or resolutions. Alternatively, some of the following image setting processing can be applied. For ease of description, such examples have been mentioned, but each of the first format and the second format can be one of various projection formats. However, the present application is not limited thereto, and modifications can be made thereto.
[0502] During the format conversion processing, an integer pixel of the image after conversion can be acquired from a fractional unit pixel and an integer unit pixel in the image before conversion due to different coordinate system characteristics, and thus interpolation can be performed. The interpolation filter used in this case can be the same as or similar to the interpolation filter described above. In this case, the related information can be explicitly generated by selecting one of a plurality of interpolation filter candidates, or the interpolation filter can be implicitly determined according to a predetermined rule. For example, a predetermined interpolation filter can be used according to the projection format, the color format, and the slice / tile type. In addition, when the interpolation filter is explicitly provided, information on the filter information (e.g., filter coefficients) can be included.
[0503] In the format conversion section, the projection format can be defined to include region-wise packing or the like. That is, projection and region-wise packing can be performed during the format conversion processing. Alternatively, processing such as region-wise packing can be performed before encoding is performed after the format conversion processing.
[0504] The encoder can add the information generated during the above processing to the bitstream in units of at least one of a sequence, a picture, a slice, a tile, or the like, and the decoder can parse the related information from the bitstream. In addition, the information can be included in the bitstream in the form of SEI or metadata.
[0505] Next, the image setting processing applied to the 360-degree image encoding / decoding apparatus according to the embodiment of the present application will be described. The image setting processing according to the present application can be applied to the preprocessing process, the post-processing process, the format conversion processing, the inverse format conversion processing, and the like of the 360-degree image encoding / decoding apparatus as well as the general encoding / decoding processing. The following description of the image setting processing will focus on the 360-degree image encoding apparatus, and can include the above-described image setting. Redundant description of the foregoing image setting processing will be omitted. In addition, the following examples will focus on the image setting processing, and the inverse image setting processing can be derived inversely from the image setting processing. Some cases can be confirmed through the foregoing various embodiments of the present application.
[0506] The image setting processing according to the present application can be performed in the 360-degree image projection step, the region-wise packing step, the format conversion step, or other steps.
[0507] Figure 20 is a conceptual diagram of 360-degree image division according to the embodiment of the present application. In Figure 20 In, it is assumed that the image is projected through ERP.
[0508] Part 20a shows an image projected through ERP, and various methods can be used to divide the image. In this example, the description focuses on slices or tiles, and it is assumed that W0 to W2 and H0 and H1 are division boundaries of the slices or tiles and follow a raster scan order. The following examples focus on slices and tiles. However, the present application is not limited thereto, and other division methods can be applied.
[0509] For example, division can be performed in units of slices, and H0 and H1 can be set as division boundaries. Alternatively, division can be performed in units of tiles, and W0 to W2, H0, and H1 can be set as division boundaries.
[0510] Part 20b shows an example in which the image projected through ERP is divided into tiles (assuming the same tile division boundaries (W0 to W2, H0, and H1 are all activated) as shown in part 20a). When it is assumed that the area P is the entire image and the area V is the area in which the user's gaze stays or the viewport, various methods can exist to provide an image corresponding to the viewport. For example, the area corresponding to the viewport can be acquired by decoding the entire image (e.g., tiles a to i). In this case, the entire image can be decoded, and tiles a to i (here, area A + area B) can be decoded in the case of dividing the image. Alternatively, the area corresponding to the viewport can be acquired by decoding the area belonging to the viewport. In this case, in the case of dividing the image, the area corresponding to the viewport can be acquired from the image recovered by decoding tiles f, g, j, and k (here, area B). The former case can be referred to as full decoding (or viewport-independent encoding), and the latter case can be referred to as partial decoding (or viewport-dependent encoding). The latter case can be an example that can occur in a 360-degree image having a large amount of data. The division method based on the tile unit can be used more frequently than the division method based on the slice unit because the division area can be acquired flexibly. For partial decoding, the referability of the division unit can be spatially or temporally limited (here, implicitly handled) because it is not possible to find where the viewpoint will occur, and encoding / decoding can be performed considering the limitation. The following examples will be described focusing on full decoding, but 360-degree image division will be described focusing on tiles (or the rectangular division method of the present application) in preparation for partial decoding. However, the following description can be applied to other division units in the same manner or with modifications.
[0511] Figure 21 is an example diagram of 360-degree image division and image reconstruction according to an embodiment of the present application. In Figure 21 , it is assumed that the image is projected through CMP.
[0512] Part 21a shows an image projected by the CMP, and various methods can be used to divide the image. It is assumed that W0 to W2, H0, and H1 are division boundary lines of the faces, slices, and tiles and follow a raster scan order.
[0513] For example, division can be performed in units of slices, and H0 and H1 can be set as division boundaries. Alternatively, division can be performed in units of tiles, and W0 to W2, H0, and H1 can be set as division boundaries. Alternatively, division can be performed in units of faces, and W0 to W2, H0, and H1 can be set as division boundaries. In this example, it is assumed that a face is a part of the division unit.
[0514] In this case, a face can be a division unit (here, dependent encoding / decoding) performed to classify or distinguish regions (here, the plane coordinate system of each face) having different properties in the same image according to the characteristics, type (in this example, 360-degree image and projection format), and the like of the image, and a slice or tile can be a division unit (here, independent encoding / decoding) performed to divide the image according to a user definition. Furthermore, a face can be a unit divided by a predetermined definition (or induced from projection format information) during projection processing according to a projection format, and a slice or tile can be a unit divided by explicitly generating division information according to a user definition. Furthermore, a face can have a polygon division shape including a quadrangle according to a projection format, a slice can have any division shape that cannot be defined as a quadrangle or polygon, and a tile can have a quadrangle division shape. The setting of the division unit can be defined only for the description of this example.
[0515] In this example, it has been described that a face is a division unit classified for region distinction. However, a face can be a unit for performing independent encoding / decoding according to an encoding / decoding setting as at least one face unit, and can have a setting for performing independent encoding / decoding in combination with a tile, slice, or the like. In this case, when a face is combined with a tile, slice, or the like, explicit information of the tile and slice can be generated, or the tile and slice can be implicitly combined based on face information. Alternatively, explicit information of the tile and slice can be generated based on face information.
[0516] As a first example, one image division processing (here, a face) is performed, and image division can implicitly omit division information (which is acquired according to projection format information). This example is for a dependent encoding / decoding setting, and can be an example corresponding to a case where referability between face units is not limited.
[0517] As a second example, one picture partitioning process (here, faces) is performed, and the picture partitioning can explicitly generate the partition information. This example is for a dependent coding / decoding setup, and can be an example corresponding to the case where the referenceability between face units is not restricted.
[0518] As a third example, multiple picture partitioning processes (here, faces and tiles) are performed, some picture partitioning (here, faces) can implicitly omit or explicitly generate the partition information, and other picture partitioning (here, tiles) can explicitly generate the partition information. In this example, one picture partitioning process (here, faces) precedes the other picture partitioning process (here, tiles).
[0519] As a fourth example, multiple picture partitioning processes are performed, some picture partitioning (here, faces) can implicitly omit or explicitly generate the partition information, and other picture partitioning (here, tiles) can explicitly generate the partition information based on the some picture partitioning (here, faces). In this example, one picture partitioning process (here, faces) precedes the other picture partitioning process (here, tiles). In some cases of this example (assuming the second example), it can be the same as explicitly generating the partition information, but there can be a difference in the partition information configuration.
[0520] As a fifth example, multiple picture partitioning processes are performed, some picture partitioning (here, faces) can implicitly omit the partition information, and other picture partitioning (here, tiles) can implicitly omit the partition information based on the some picture partitioning (here, faces). For example, face units can be set individually as tile units, or multiple face units (here, face units are grouped when adjacent faces have continuity; otherwise, face units are not grouped; B2-B3-B1 and B4-B0-B5 in part 18a) can be set as tile units. According to a predetermined rule, face units can be set as tile units. This example is for an independent coding / decoding setup, and can be an example corresponding to the case where the referenceability between face units is restricted. That is, in some cases (assuming the first example), it can be the same as implicitly processing the partition information, but there can be a difference in the coding / decoding setup.
[0521] This example can be a description of the case where the partitioning process can be performed in the projection step, the region-wise packing step, the initial coding / decoding step, and the like, and can be any other picture partitioning process performed in the encoder / decoder.
[0522] In part 21a, a rectangular image can be constructed by adding a region (B) not including data to a region (A) including data. In this case, the position, size, shape, number, and the like of the region A and the region B can be information that can be checked by a projection format or the like or information that can be checked when information about a projection image is explicitly generated, and the relevant information can be expressed with the above-described image division information, image reconstruction information, and the like. For example, information about a specific region of a projection image (e.g., part_top, part_left, part_width, part_height, and part_convert_flag) can be expressed as shown in Table 4 and Table 5. However, the present application is not limited to this and can be applied to other cases (e.g., another projection format, other projection settings, and the like).
[0523] The region B and the region A can be constructed as a single image and then encoded or decoded. Alternatively, division can be performed taking into account the region-wise characteristics, and different encoding / decoding settings can be applied. For example, the region B can not be encoded or decoded by using information about whether encoding or decoding is performed (e.g., tile_coded_flag when a division unit is assumed to be a tile). In this case, the corresponding region can be restored to certain data (here, any pixel value) according to a predetermined rule. Alternatively, in the above-described image division process, the region B can have different encoding / decoding settings from the region A. Alternatively, the corresponding region can be removed by performing a region-wise packing process.
[0524] Part 21b shows an example in which an image packed by CMP is divided into tiles, slices, or faces. In this case, the packed image is an image for which a face rearrangement process or a region-wise packing process is performed, and can be an image obtained by performing image division and image reconstruction according to the present application.
[0525] In part 21b, a rectangular shape can be constructed to include a region having data. In this case, the position, size, shape, number, and the like of the region can be information that can be checked by a predetermined setting or information that can be checked when information about a packed image is explicitly generated, and the relevant information can be expressed with the above-described image division information, image reconstruction information, and the like. For example, information about a specific region of a packed image (e.g., part_top, part_left, part_width, part_height, and part_convert_flag) can be expressed as shown in Table 4 and Table 5.
[0526] The packing image can be divided using various division methods. For example, division can be performed in units of slices, and H0 settings can be set as division boundaries. Alternatively, division can be performed in units of tiles, and W0, W1, and H0 settings can be set as division boundaries. Alternatively, division can be performed in units of faces, and W0, W1, and H0 settings can be set as division boundaries.
[0527] The image division processing and the image reconstruction processing according to the present application can be performed on a projection image. In this case, the reconstruction processing can be used to rearrange faces in the image and pixels in the image. This can be a possible example when the image is divided into or constituted by a plurality of faces. The following example will be described focusing on a case where the image is divided into tiles based on face units.
[0528] SX, Y (S0, 0 to S3, 2) in the portion 21a can correspond to S'U, V (S'0, 0 to S'2, 1) in the portion 21b (here, X and Y can be the same as or different from U and V), and the reconstruction processing can be performed in units of faces. For example, S2, 1, S3, 1, S0, 1, S1, 2, S1, 1, and S1, 0 can be allocated to S'0, 0, S'1, 0, S'2, 0, S'0, 1, S'1, 1, and S'2, 1 (face rearrangement). In addition, S2, 1, S3, 1, and S0, 1 can not be reconstructed (pixel rearrangement), and S1, 2, S1, 1, and S1, 0 can be rotated by 90 degrees and then reconstructed. This can be represented as shown in the portion 21c. In the portion 21c, the horizontally placed symbols (S1, 0, S1, 1, S1, 2) can be images that are horizontally placed to maintain continuity of the image.
[0529] The reconstruction of the faces can be processed implicitly or explicitly according to the encoding / decoding settings. The implicit processing can be performed according to predetermined rules in consideration of the type of the image (here, a 360-degree image) and the characteristics (here, a projection format, and the like).
[0530] For example, for S'0, 0 and S'1, 0, S'1, 0 and S'2, 0, S'0, 1 and S'1, 1, S'1, 1 and S'2, 1 in the portion 21c, there is image continuity (or correlation) between two faces with respect to the face boundary, and the portion 21c can be an example in which there is continuity between three upper faces and three lower faces. In a case where the image is divided into a plurality of faces through projection processing from a 3D space to a 2D space and then packed for each region, the reconstruction can be performed to increase image continuity between the faces in order to efficiently reconstruct the faces. Such face reconstruction can be determined and processed in advance.
[0531] Alternatively, the reconstruction processing can be performed through explicit processing, and reconstruction information can be generated.
[0532] For example, when examining information about an M×N build (e.g., 6×1, 3×2, 2×3, 1×6, etc. for CMP compaction; in this example, a 3×2 configuration is assumed) via region-based packing processing (e.g., one of implicitly acquired and explicitly generated information), face reconstruction can be performed based on the M×N build, and information about the face reconstruction can then be generated. For example, when rearranging faces in an image, index information (or information about their position in the image) can be assigned to each face. When rearranging pixels within a face, pattern information for reconstruction can be assigned.
[0533] Index information can be as follows Figure 18 As shown in sections 18a to 18c, they are predefined. In sections 21a to 21c, SX,Y or S'U,V represent each face using positional information indicating width and height (e.g., S[i][j]) or using a single positional information (e.g., S[i]; assuming the positional information is assigned in raster scan order starting from the top left of the image), and each face can be assigned an index.
[0534] For example, when indexing is assigned using positional information indicating width and height, face index #2 can be assigned to S'0,0, face index #3 can be assigned to S'1,0, face index #1 can be assigned to S'2,0, face index #5 can be assigned to S'0,1, face index #0 can be assigned to S'1,1, and face index #4 can be assigned to S'2,1, as shown in section 21c. Alternatively, when indexing is assigned using a single positional information, face index #2 can be assigned to S[0], face index #3 can be assigned to S[1], face index #1 can be assigned to S[2], face index #5 can be assigned to S[3], face index #0 can be assigned to S[4], and face index #4 can be assigned to S[5]. For ease of description, in the following examples, S'0,0 to S'2,1 can be referred to as a to f. Alternatively, each face can be represented based on positional information indicating the width and height of a pixel or block unit at the top left corner of the image.
[0535] For the packed images acquired through the image reconstruction process (or the region-wise packing process), depending on the reconstruction setting, the face scan order is the same as or different from the image scan order. For example, when one scan order (e.g., raster scan) is applied to the images shown in part 21a, a, b, and c can have the same scan order, and d, e, and f can have a different scan order. For example, when the scan order of part 21a or the scan order of a, b, and c follows the order of (0, 0) -> (1, 0) -> (0, 1) -> (1, 1), the scan order of d, e, and f can follow the order of (1, 0) -> (1, 1) -> (0, 0) -> (0, 1). This can be determined according to the image reconstruction setting, and such a setting can even be applied to other projection formats.
[0536] In the image division process shown in part 21b, the tiles can be individually set as face units. For example, each of faces a to f can be set as a tile unit. Alternatively, a plurality of face units can be set as a tile. For example, faces a to c can be set as one tile, and faces d to f can be set as one tile. The construction can be determined based on face characteristics (e.g., continuity between faces, etc.), and unlike the above-described example, different tile settings of faces can be possible.
[0537] The following is an example of division information according to a plurality of image division processes. In this example, it is assumed that the division information of the faces is omitted, the unit other than the face is the tile, and the division information is differently processed.
[0538] As a first example, the image division information can be acquired based on the face information, and the image division information can be implicitly omitted. For example, the face can be individually set as a tile, or a plurality of faces can be set as a tile. In this case, when at least one face is set as a tile, this can be determined according to a predetermined rule based on the face information (e.g., continuity or correlation).
[0539] As a second example, the image division information can be explicitly generated regardless of the face information. For example, when the division information is generated using the number of columns (here, num_tile_columns) and the number of rows (here, num_tile_rows) of the tile, the division information can be generated in the above-described method of the image division process. For example, the number of columns of the tile can be in the range from 0 to the width of the image or the width of the block (here, the unit acquired from the picture division section), and the number of rows of the tile can be in the range from 0 to the height of the image or the height of the block. In addition, additional division information (e.g., uniform_spacing_flag) can be generated. In this case, depending on the division setting, the boundaries of the faces and the boundaries of the division units can match or not match each other.
[0540] As a third example, image partition information can be explicitly generated based on face information. For example, when the number of columns and the number of rows of tiles are used to generate the partition information, the partition information can be generated based on the face information (here, the number of columns is in the range of 0 to 2, and the number of rows is in the range of 0 to 1; since the configuration of faces in the image is 3x2). For example, the number of columns of tiles can be in the range of 0 to 2, and the number of rows of tiles can be in the range of 0 to 1. In addition, additional partition information (e.g., uniform_spacing_flag) can not be generated. In this case, the boundaries of faces and the boundaries of partition units can match each other.
[0541] In some cases (assuming the second example and the third example), the syntax elements of the partition information can be defined differently, or even if the same syntax elements are used, the syntax element settings (e.g., binarization settings; when the range of candidate groups of a syntax element is limited and small, other binarization can be used) can be applied differently. The above examples have been described for some of the various elements of the partition information. However, the present application is not limited thereto, and it can be understood that other settings are possible depending on whether the partition information is generated based on face information.
[0542] Figure 22 An example image in which the image packed or projected by the CMP is partitioned into tiles.
[0543] In this case, it is assumed that the tile partition boundaries (W0 to W2, H0 and H1 are all activated) are the same as those shown in the tile partition boundaries of the portion 21a of FIG. 20, and that the tile partition boundaries (W0, W1 and H0 are all activated) are the same as those shown in the tile partition boundaries of the portion 21b of FIG. 20. Figure 21 Figure 21 When it is assumed that the region P indicates the entire image and the region V indicates the viewport, full decoding or partial decoding can be performed. This example will be described focusing on partial decoding. In the portion 22a, tiles e, f and g can be decoded for the CMP (left side), and tiles a, c and e can be decoded for the CMP compact (right side) to acquire the region corresponding to the viewport. In the portion 22b, tiles b, f and i can be decoded for the CMP, and tiles d, e and f can be decoded for the CMP compact to acquire the region corresponding to the viewport.
[0544] The above examples have been described for the case where the partitioning of slices, tiles, etc. is performed based on face units (or face boundaries). However, as shown in the portion 20a of FIG. 20, the partitioning can be performed inside a face (e.g., the image consists of one face under the ERP, and consists of multiple faces under other projection formats), or the partitioning can be performed on the boundaries of faces as well as inside. Figure 20
[0545] Figure 23 is a conceptual diagram illustrating an example of adjusting the size of a 360-degree image according to an embodiment of the present application. In this case, it is assumed that an image is projected through an ERP. Also, the following example will be described focusing on the case of expansion.
[0546] The size of a projected image can be adjusted according to an image size adjustment type, through a scale factor or through an offset factor. Here, an image before size adjustment can be P_Width x P_Height, and an image after size adjustment can be P'_Width x P'_Height.
[0547] For a scale factor, after adjusting the width and height of an image through a scale factor (here, a in the case of width and b in the case of height), the width (P_Width x a) and height (P_Height x b) of the image can be obtained. For an offset factor, after adjusting the width and height of an image through an offset factor (here, L and R in the case of width and T and B in the case of height), the width (P_Width + L + R) and height (P_Height + T + B) of the image can be obtained. Size adjustment can be performed using a predetermined method, or can be performed using a method selected from among a plurality of methods.
[0548] The data processing method in the following example will be described focusing on the case of an offset factor. For an offset factor, as a data processing method, there can be a padding method through the use of a predetermined pixel value, a padding method through the copying of external pixels, a padding method through the copying of a specific region of an image, a padding method through the transformation of a specific region of an image, etc.
[0549] The size of a 360-degree image can be adjusted considering the characteristic that continuity exists at the boundary of an image. For an ERP, there is no outer boundary in a 3D space, but an outer boundary can exist when a 3D space is transformed into a 2D space through a projection process. Data in a boundary region includes data having outward continuity, but can have a boundary in terms of spatial characteristics. Size adjustment can be performed considering these characteristics. In this case, continuity can be checked according to a projection format, etc. For example, an ERP image can be an image having the characteristic in which both end boundaries are continuous. This example will be described assuming that the left boundary and right boundary of an image are continuous with each other and the upper boundary and lower boundary of an image are continuous with each other. The data processing method will be described focusing on a padding method through the copying of a specific region of an image and a padding method through the transformation of a specific region of an image.
[0550] When the image size is adjusted to the left, the resized area (here, LC or TL+LC+BL) can be padded with data of the right area (here, tr+rc+br) of the image having continuity with the left of the image. When the image size is adjusted to the right, the resized area (here, RC or TR+RC+BR) can be padded with data of the left area (here, tl+lc+bl) of the image having continuity with the right of the image. When the image size is adjusted upward, the resized area (here, TC or TL+TC+TR) can be padded with data of the lower area (here, bl+bc+br) of the image having continuity with the upper side. When the image size is adjusted downward, the resized area (here, BC or BL+BC+BR) can be padded with data.
[0551] When the size or length of the resized area is m, the coordinates (here, x ranges from 0 to P_Width-1) of the resized area with respect to the image before the size adjustment can have a range from (-m, y) to (-1, y) (size adjustment to the left) or a range from (P_Width, y) to (P_Width+m-1, y) (size adjustment to the right). The position (x') of the area for obtaining data of the resized area can be derived from the equation x'=(x+P_Width)%P_Width. In this case, x denotes the coordinates of the coordinates of the resized area with respect to the image before the size adjustment, and x' denotes the coordinates of the coordinates of the area with respect to the image before the size adjustment with reference to the resized area. For example, when the image size is adjusted to the left, m is 4, and the image width is 16, corresponding data (-4, y) can be obtained from (12, y), corresponding data (-3, y) can be obtained from (13, y), corresponding data (-2, y) can be obtained from (14, y), and corresponding data (-1, y) can be obtained from (15, y). Alternatively, when the image size is adjusted to the right, m is 4, and the image width is 16, corresponding data (16, y) can be obtained from (0, y), corresponding data (17, y) can be obtained from (1, y), corresponding data (18, y) can be obtained from (2, y), and corresponding data (19, y) can be obtained from (3, y).
[0552] When the size or length of the resized region is n, the resized region can have a range from (x, -n) to (x, -1) (upward resizing) or a range from (x, P_Height) to (x, P_Height+n-1) (downward resizing) with respect to the coordinates of the image before resizing (here, y ranges from 0 to P_Height-1). The position (y') of the region from which data of the resized region is acquired can be derived from the equation y' = (y+P_Height)%P_Height. In this case, y denotes a coordinate of a coordinate of the resized region with respect to the image before resizing, and y' denotes a coordinate of a region with respect to the coordinates of the image before resizing with reference to the resized region. For example, when the image size is resized upward, n is 4, and the height of the image is 16, corresponding data (x, -4) can be acquired from (x, 12), corresponding data (x, -3) can be acquired from (x, 13), corresponding data (x, -2) can be acquired from (x, 14), and corresponding data (x, -1) can be acquired from (x, 15). Alternatively, when the image size is resized downward, n is 4, and the height of the image is 16, corresponding data (x, 16) can be acquired from (x, 0), corresponding data (x, 17) can be acquired from (x, 1), corresponding data (x, 18) can be acquired from (x, 2), and corresponding data (x, 19) can be acquired from (x, 3).
[0553] After filling the resized region with data, resizing can be performed with respect to the coordinates of the image after resizing (here, x ranges from 0 to P'_Width-1, and y ranges from 0 to P'_Height-1). This example can be applied to a coordinate system of latitude and longitude.
[0554] Various combinations of resizing can be provided as follows.
[0555] As an example, the size of the image can be resized leftward by m. Alternatively, the size of the image can be resized rightward by n. Alternatively, the size of the image can be resized upward by o. Alternatively, the size of the image can be resized downward by p.
[0556] As an example, the size of the image can be resized leftward by m and rightward by n. Alternatively, the size of the image can be resized upward by o and downward by p.
[0557] As an example, the size of the image can be resized leftward by m, rightward by n, and upward by o. Alternatively, the size of the image can be resized leftward by m, rightward by n, and downward by p. Alternatively, the size of the image can be resized leftward by m, upward by o, and downward by p. Alternatively, the size of the image can be resized rightward by n, upward by o, and downward by p.
[0558] As an example, the size of the image can be adjusted by m to the left, n to the right, o upward, and p downward.
[0559] Similar to the above example, at least one size adjustment operation can be performed. The image size adjustment can be performed implicitly according to the encoding / decoding setting, or size adjustment information can be generated implicitly, and then the image size adjustment can be performed based on the generated size adjustment information. That is, m, n, o, and p of the above example can be determined as predetermined values, or can be generated explicitly using size adjustment information. Alternatively, some of m, n, o, and p can be determined as predetermined values, and other values can be generated explicitly.
[0560] The above example has been described focusing on a case where data is acquired from a specific region of an image, but other methods can also be applied. The data can be a pixel before encoding or a pixel after encoding, and can be determined according to the size adjustment step or the characteristics of the image to be adjusted in size. For example, when the size adjustment is performed in a pre-processing process and a pre-encoding step, the data can refer to an input pixel of a projected image, a packed image, or the like, and when the size adjustment is performed in a post-processing process, an intra prediction reference pixel generation step, a reference picture generation step, a filtering step, or the like, the data can refer to a restored pixel. Furthermore, the size adjustment can be performed in each of the regions adjusted in size by separately using a data processing method.
[0561] Figure 24 is a conceptual diagram showing continuity between faces under a projection format (e.g., CHP, OHP, or ISP) according to an embodiment of the present application.
[0562] In detail, Figure 24 An example of an image composed of a plurality of faces can be shown. The continuity can be a characteristic generated in adjacent regions in a 3D space. Parts 24a to 24c differently show a case (A) having spatial adjacency and continuity, a case (B) having spatial adjacency but no continuity, a case (C) having no spatial adjacency but having continuity, and a case (D) having neither spatial adjacency nor continuity, when transformed to a 2D space by a projection process. Unlike this, a general image is classified into a case (A) having both spatial adjacency and continuity and a case (D) having neither spatial adjacency nor continuity. In this case, the case having continuity corresponds to some examples (A or C).
[0563] That is, with reference to the portions 24a to 24c, the case where both spatial adjacency and continuity exist (here, the case described with reference to the portion 24a) can be shown as b0 to b4, and the case where spatial adjacency does not exist but continuity exists can be shown as B0 to B6. That is, these cases indicate regions that are adjacent in the 3D space, and it is possible to enhance the encoding performance by using the properties of b0 to b4 and B0 to B6 having continuity in the encoding process.
[0564] Figure 25 is a conceptual diagram showing face continuity in the portion 21c, which is an image acquired through the image reconstruction process or the region-wise packing process under the CMP projection format.
[0565] Here, Figure 21 The portion 21c of shows rearrangement of the 360-degree image expanded in the shape of a cube in the portion 21a, and thus the face continuity applied to the portion 21a is maintained. Figure 21 That is, as shown in the portion 25a, the face S2,1 can be horizontally continuous with the faces S1,1 and S3,1, and can be vertically continuous with the face S1,0 rotated by 90 degrees and the face S1,2 rotated by -90 degrees.
[0566] In the same manner, the continuity of the faces S3,1, S0,1, S1,2, S1,1, and S1,0 can be checked in the portions 25b to 25f.
[0567] The continuity between faces can be defined in accordance with the projection format setting or the like. However, the present application is not limited to this, and modifications can be made thereto. The following examples will be described assuming that the continuity exists as shown in Figure 24 and Figure 25
[0568] Figure 26 is an example diagram showing image size adjustment under the CMP projection format according to an embodiment of the present application.
[0569] The portion 26a shows an example of adjusting the image size, the portion 26b shows an example of adjusting the size of the face unit (or division unit), and the portion 26c shows an example of adjusting the sizes of the image and the face unit (or an example of performing multiple size adjustments).
[0570] The size of the projection image can be adjusted by a scale factor or by an offset factor according to the image size adjustment type. Here, the image before the size adjustment can be P_Width x P_Height, the image after the size adjustment can be P'_Width x P'_Height, and the size of the face can be F_Width x F_Height. The size can be the same or different according to the face, and the width and height can be the same or different according to the face. However, for ease of description, the example will be described assuming that all the faces in the image have the same size and a square shape. In addition, the description assumes that the size adjustment values (here, WX and HY) are the same. In the following example, the case of the offset factor will be focused on and the data processing method will also be described focusing on the padding method by copying a certain region of the image and the padding method by transforming a certain region of the image. Even the above-described setting can be applied to Figure 27 the case shown in FIG. 24.
[0571] For the portions 26a to 26c, the boundary of the face can have continuity with the boundary of another face (here, it is assumed to have continuity corresponding to Figure 24 the portion 24a of FIG. 23). Here, the continuity can be classified into the case of having spatial adjacency and image continuity in the 2D plane (first example) and the case of not having spatial adjacency but having image continuity in the 2D plane (second example).
[0572] For example, when the continuity in the portion 24a of FIG. 23 is assumed, the upper, left, right, and lower regions of S1,1 can be spatially adjacent to and have image continuity with the lower, right, left, and upper regions of S1,0, S0,1, S2,1, and S1,2 (first example). Figure 24 Alternatively, the left and right regions of S1,0 can not be spatially adjacent to the upper regions of S0,1 and S2,1, but can have image continuity with the upper regions of S0,1 and S2,1 (second example). In addition, the left regions of S0,1 can not be spatially adjacent to each other, but can have image continuity with each other (second example). In addition, the left and right regions of S1,2 can be continuous with the lower regions of S0,1 and S2,1 (second example). This can be only a limited example, and other configurations can be applied according to the definition and setting of the projection format. For ease of description, S0,0 to S3,2 in the portion 26a are referred to as a to l.
[0573]
[0574] The portion 26a can be an example of a padding method using data of a region having continuity toward an outer boundary of an image. A range from a region A not including data to a resized region (here, a0 to a2, c0, d0 to d2, i0 to i2, k0, and l0 to l2) can be padded with any predetermined value or padded by an external pixel padding, and a range from a region B including actual data to a resized region (here, b0, e0, h0, and j0) can be padded with data of a region (or a face) having continuity of the image. For example, b0 can be padded with data of an upper side of a face h, e0 can be padded with data of a right side of the face h, h0 can be padded with data of a left side of a face e, and j0 can be padded with data of a lower side of the face h.
[0575] In detail, as an example, b0 can be padded with data of a lower side of a face acquired by rotating the face h by 180 degrees, and j0 can be padded with data of an upper side of a face acquired by rotating the face h by 180 degrees. However, this example (including the following examples) can only indicate a position of a reference face, and data acquired from a resized region can be acquired after a resizing process (e.g., rotation, etc.) considering continuity between faces as illustrated in Figure 24 and Figure 25
[0576] The portion 26b can be an example of a padding method using data of a region having continuity toward an inner boundary of an image. In this example, different resizing operations can be performed for each face. A reduction process can be performed in a region A, and an expansion process can be performed in a region B. For example, a size of a face a can be adjusted (here, reduced) rightward by w0, and a size of a face b can be adjusted (here, expanded) leftward by w0. Alternatively, a size of the face a can be adjusted (here, reduced) downward by h0, and a size of a face e can be adjusted (here, expanded) upward by h0. In this example, when a width change of the image is observed through faces a, b, c, and d, the face a is reduced by w0, the face b is expanded by w0 and w1, and the face c can be reduced by w1. Accordingly, a width of the image before the resizing is the same as a width of the image after the resizing. When a height change of the image is observed through faces a, e, and i, the face a is reduced by h0, the face e is expanded by h0 and h1, and the face i can be reduced by h1. Accordingly, a height of the image before the resizing is the same as a height of the image after the resizing.
[0577] Considering that a region is reduced from a region A not including data, resized regions (here, b0, e0, be, b1, bg, g0, h0, e1, ej, j0, gi, g1, j1, and h1) can be simply removed, and considering that a region is expanded from a region B including actual data, the resized regions can be padded with data of a region having continuity.
[0578] For example, b0 can be padded with data of the upper side of the face e; e0 can be padded with data of the left side of the face b; be can be padded with data of the left side of the face b, the upper side of the face e, or a weighted sum of the left side of the face b and the upper side of the face e; b1 can be padded with data of the upper side of the face g; bg can be padded with data of the left side of the face b, the upper side of the face g, or a weighted sum of the right side of the face b and the upper side of the face g; g0 can be padded with data of the right side of the face b; h0 can be padded with data of the upper side of the face b; e1 can be padded with data of the left side of the face j; ej can be padded with data of the lower side of the face e, the left side of the face j, or a weighted sum of the lower side of the face e and the left side of the face j; j0 can be padded with data of the lower side of the face e; gj can be padded with data of the lower side of the face g, the left side of the face j, or a weighted sum of the lower side of the face g and the right side of the face j; g1 can be padded with data of the right side of the face j; j1 can be padded with data of the lower side of the face g; and h1 can be padded with data of the lower side of the face j.
[0579] In the above example, when the resized region is padded with data of a specific region of the image, the data of the corresponding region can be copied and then used to pad the resized region, or can be transformed based on characteristics, types, etc. of the image, and then used to pad the resized region. For example, when a 360-degree image can be transformed into a 2D space according to a projection format, a coordinate system (for example, a 2D plane coordinate system) can be defined for each face. For ease of description, it is assumed that (x, y, z) in a 3D space is transformed into (x, y, C), (x, C, z), or (C, y, z) for each face. The above example indicates a case where data of a face other than the corresponding face is acquired from the resized region of the face. That is, when resizing is performed on the current face, data of another face having a different coordinate system characteristic can be copied as is and then used. In this case, there is a possibility that continuity is distorted based on the resizing boundary. For this reason, data of another face acquired according to the coordinate system characteristic of the current face can be transformed and used to pad the resized region. This transformation is also merely an example of a data processing method, and the present invention is not limited thereto.
[0580] When data of a specific region of the image is copied and used to pad the resized region, there can be included distorted continuity (or fundamentally changed continuity) in a boundary region between the resized region (e) and the resized region (e0). For example, continuity can change with respect to the boundary, and a straight line edge can be bent with respect to the boundary.
[0581] When data of a specific region of the image is transformed and used to pad the resized region, there can be included gradually changed continuity in a boundary region between the resized regions.
[0582] The above example can be an example of a data processing method of the present application to transform data of a specific region of an image based on a characteristic, a type, or the like of the image, and to fill the resized region with the transformed data.
[0583] The portion 26c can be an example of filling the resized region with data of a region having continuity toward the boundaries (inner and outer boundaries) of the image, corresponding to the image resizing processing of the portions 26a and 26b. The resizing processing of this example can be derived from the resizing processing of the portions 26a and 26b, and detailed description thereof will be omitted.
[0584] The portion 26a can be an example of a processing of resizing an image, and the portion 26b can be an example of a processing of resizing a division unit in an image. The portion 26c can be an example of a plurality of resizing processes including the processing of resizing an image and the processing of resizing a division unit in an image.
[0585] For example, the size (here, the region C) of an image (here, a first format) acquired through a projection process can be resized, and the size (here, the region D) of an image (here, a second format) acquired through a format conversion process can be resized. In this example, the size of an image (here, a full image) acquired through ERP projection can be resized and the image is converted into an image acquired through CMP projection by a format conversion section, and the size of the image (here, a face unit) acquired through CMP projection can be resized. The above example is an example of performing a plurality of resizing operations. However, the present application is not limited to this, and modifications thereof can be made.
[0586] Figure 27 is an example diagram illustrating resizing of an image converted and packed in a CMP projection format according to an embodiment of the present application. Figure 27 It is also assumed that continuity between faces as shown in Figure 25 Thus, the boundary of a face can have continuity with the boundary of another face.
[0587] In this example, the offset factors of W0 to W5 and H0 to H3 can have various values (here, it is assumed that the offset factors are used as the resizing values). For example, the offset factors can be derived from a predetermined value, a motion search range of inter prediction, a unit acquired from a picture division section, and the like, and other cases are also possible. In this case, the unit acquired from the pixel division unit can include a face. That is, the resizing values can be determined based on F_Width and F_Height.
[0588] Part 27a is an example of adjusting the size of a single face (here, with respect to the upward, downward, leftward, and rightward faces) and filling...
Claims
1. A method for processing an image, the method comprising: Receive the bitstream of the image; Information regarding image size adjustment for the image is obtained based on the bitstream; The image is reconstructed by decoding the bitstream; as well as Based on the information regarding image size adjustment, perform image size adjustment on the reconstructed image. The information regarding image resizing includes an offset factor for each direction of the reconstructed image, and The bitstream includes information about the rotation of the image.
2. The method of claim 1, wherein the image size adjustment is performed taking into account both horizontal and vertical scaling factors, and the horizontal scaling factor and the vertical scaling factor are obtained independently of each other.
3. The method of claim 1, wherein the image resizing is performed based on a resizing value, and the resizing value is obtained based on an offset factor and decoding settings included in the information regarding the image resizing.
4. The method according to claim 3, wherein, According to the decoding settings, the adjusted size value is calculated to be equal to the offset factor multiplied by 2.
5. The method of claim 1, wherein the image size adjustment of the chromaticity component is performed based on the image size adjustment of the luminance component.
6. A method for processing an image, the method comprising: A bitstream is generated by encoding the image; as well as Information regarding image size adjustment for the image is encoded into the bitstream; The information regarding image resizing is used to perform image resizing on the image during reconstruction. The information regarding image resizing includes an offset factor for each direction of the image, and The bitstream includes information about the rotation of the image.
7. A method for transmitting a bit stream, the method comprising: A bitstream is generated by encoding the image; Information regarding image size adjustment for the image is encoded into the bitstream; as well as Transmit the bit stream, The information regarding image resizing is used to perform image resizing on the image during reconstruction. The information regarding image resizing includes an offset factor for each direction of the image, and The bitstream includes information about the rotation of the image.