Image data encoding / decoding method and device
The method addresses the inefficiencies in processing 360-degree images by generating and reconstructing images in specific projection formats and extending image data, leading to improved compression performance for 360-degree images.
Patent Information
- Application Number
- JP2019518973
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-07-17
- Filing Date
- 2017-10-10
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2037-10-10
AI Technical Summary
Existing image processing systems struggle to efficiently handle the large data volumes generated by 360-degree images for virtual and augmented reality, necessitating improved encoding and decoding methods to enhance performance.
A method for decoding 360-degree images involves generating a predicted image by referring to syntax information, combining it with a residual image, and reconstructing the image in a projection format, utilizing projection formats like ERP, CMP, OHP, and ISP, and performing image extension based on division units to improve compression performance.
The proposed method enhances compression performance for 360-degree images, improving the efficiency of image processing systems in handling these high-data-volume images.
Smart Images

Figure 0007224280000007 
Figure 0007224280000008 
Figure 0007224280000009
Abstract
Description
[Technical Field]
[0001] The present invention relates to image data encoding and decoding technology, and more particularly to a method and apparatus for processing encoding and decoding of 360-degree images for immersive media services. [Background technology]
[0002] With the spread of the Internet and mobile devices and the development of information and communication technology, the use of multimedia data is rapidly increasing. Recently, demand for high-resolution and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been rising in various fields, and demand for immersive media services such as virtual reality and augmented reality is also rapidly increasing. In particular, in the case of 360-degree images for virtual reality and augmented reality, multi-view images captured by multiple cameras are processed, which generates a huge amount of data, but the performance of image processing systems to process this data is currently insufficient.
[0003] Thus, in the prior art image encoding / decoding methods and apparatuses, there is a need for improved performance for image processing, particularly image encoding / decoding. Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention has been made to solve the above-mentioned problems, and its object is to provide a method for improving an image setting process in an early stage of encoding and decoding, and more particularly, to provide an encoding and decoding method and apparatus for improving the image setting process by taking into account the characteristics of 360-degree images. [Means for solving the problem]
[0005] To achieve the above object, one aspect of the present invention provides a method for decoding a 360-degree image.
[0006] Here, the 360-degree image decoding method may include the steps of receiving a bitstream in which a 360-degree image is encoded, generating a predicted image by referring to syntax information obtained from the received bitstream, obtaining a decoded image by combining the generated predicted image with a residual image obtained by inverse quantizing and inverse transforming the bitstream, and reconstructing the decoded image into a 360-degree image in a projection format.
[0007] Here, the syntax information may include projection format information for the 360-degree image.
[0008] Here, the projection format information may be information indicating at least one of an ERP (Equi-Rectangular Projection) format in which the 360-degree image is projected onto a two-dimensional plane, a CMP (CubeMap Projection) format in which the 360-degree image is projected onto a cube, an OHP (OctaHedron Projection) format in which the 360-degree image is projected onto an octahedron, and an ISP (IcoSahedral Projection) format in which the 360-degree image is projected onto a polyhedron.
[0009] Here, the reconstructing step may include obtaining placement information based on regional packing by referring to the syntax information, and rearranging each block of the decoded image based on the placement information.
[0010] Here, the step of generating the predicted image may include a step of performing image extension on a reference picture obtained by restoring the bitstream, and a step of generating the predicted image by referring to the reference picture that has been image extended.
[0011] Here, the step of performing the image extension may include performing the image extension based on a division unit of the reference image.
[0012] Here, the step of performing image expansion based on the division unit may generate an expanded region individually for each division unit using boundary pixels of the division unit.
[0013] Here, the expanded region can be generated using boundary pixels of division units spatially adjacent to the division unit to be expanded, or boundary pixels of division units having image continuity with the division unit to be expanded.
[0014] Here, the step of performing image extension based on the division units may generate an extended image for the combined region using boundary pixels of a region where two or more spatially adjacent division units among the division units are combined.
[0015] Here, the step of performing image extension based on the division units may generate an extended area between the adjacent division units by using all adjacent pixel information of the division units that are spatially adjacent among the division units.
[0016] Here, the step of expanding the image based on the division unit may generate the expanded region using an average value of adjacent pixels of each of the spatially adjacent division units.
[0017] Here, the step of generating the predicted image may include the steps of: obtaining a group of motion vector candidates including motion vectors of blocks adjacent to the current block to be decoded from the motion information included in the syntax information; deriving a predicted motion vector from the group of motion vector candidates based on selection information extracted from the motion information; and determining a predicted block of the current block to be decoded using a final motion vector derived by adding the predicted motion vector to a differential motion vector extracted from the motion information.
[0018] Here, when a block adjacent to the current block is different from the surface to which the current block belongs, the group of motion vector candidates can be composed of only motion vectors for blocks belonging to surfaces that have image continuity with the surface to which the current block belongs, among the adjacent blocks.
[0019] Here, the adjacent block may refer to a block adjacent to the current block in at least one direction among the upper left, upper, upper right, left, and lower left.
[0020] Here, the final motion vector may point to a reference area that belongs to at least one reference picture based on the current block and is set to an area where there is image continuity between surfaces according to the projection format.
[0021] Here, the reference picture can be extended in the up, down, left, and right directions based on image continuity according to the projection format, and then the reference area can be set.
[0022] Here, the reference picture is extended in units of the surface, and the reference area can be set across the boundary of the surface.
[0023] Here, the motion information may include at least one of a reference picture list to which the reference picture belongs, an index of the reference picture, and a motion vector indicating the reference area.
[0024] Here, generating a predicted block for the current block may include dividing the current block into a plurality of sub-blocks and generating a predicted block for each of the divided sub-blocks. [Effects of the Invention]
[0025] When the image encoding / decoding method and apparatus according to the embodiment of the present invention are used as described above, it is possible to improve compression performance, especially in the case of 360-degree images. [Brief explanation of the drawings]
[0026] [Figure 1] 1 is a block diagram of an image encoding device according to an embodiment of the present invention. [Figure 2] 1 is a block diagram of an image decoding device according to an embodiment of the present invention. [Figure 3] 1 is a diagram illustrating an example in which image information is divided into layers for image compression; [Figure 4] 1A-1C are conceptual diagrams illustrating various examples of image segmentation according to an embodiment of the present invention. [Figure 5] FIG. 10 is another exemplary diagram of an image division method according to an embodiment of the present invention. [Figure 6] 1 is a diagram illustrating a general method for adjusting the size of an image. [Figure 7] 10A and 10B are diagrams illustrating image size adjustment according to an embodiment of the present invention. [Figure 8] 10A and 10B are diagrams illustrating a method for configuring an expanded area in an image resizing method according to an embodiment of the present invention; [Figure 9] 10A and 10B are diagrams illustrating a method for configuring an area to be deleted and an area to be generated by reducing the size in an image resizing method according to an embodiment of the present invention; [Figure 10] FIG. 1 is an exemplary illustration of image reconstruction according to an embodiment of the present invention. [Figure 11] 1A and 1B are exemplary diagrams illustrating images before and after an image setting process according to an embodiment of the present invention. [Figure 12] 10A and 10B are diagrams illustrating size adjustment for each division unit in an image according to an embodiment of the present invention. [Figure 13] 10A and 10B are diagrams illustrating an example of a set of size adjustments or settings for division units within an image. [Figure 14] 10 is an exemplary diagram illustrating an image size adjustment process and a size adjustment process for division units within an image together; [Figure 15] 1 is an exemplary diagram showing a three-dimensional space showing a three-dimensional image and a two-dimensional plane space. [Figure 16]1A to 1D are conceptual diagrams for explaining a projection format according to one embodiment of the present invention. [Figure 17] FIG. 1 is a conceptual diagram illustrating a projection format realized within a rectangular image according to an embodiment of the present invention. [Figure 18] FIG. 10 is a conceptual diagram of a method for converting a projection format to a rectangular shape by rearranging surfaces to eliminate insignificant areas according to an embodiment of the present invention. [Figure 19] 10 is a conceptual diagram illustrating a process of packing a CMP projection format into a rectangular image according to an embodiment of the present invention. [Figure 20] FIG. 1 is a conceptual diagram illustrating division of a 360-degree image according to an embodiment of the present invention. [Figure 21] 1 is an exemplary diagram of 360-degree image division and image reconstruction according to an embodiment of the present invention; [Figure 22] FIG. 10 is an exemplary diagram showing an image projected or packed by CMP divided into tiles. [Figure 23] FIG. 10 is a conceptual diagram illustrating an example of size adjustment of a 360-degree image according to an embodiment of the present invention. [Figure 24] 1 is a conceptual diagram illustrating continuity between surfaces in a projection format (eg, CMP, OHP, ISP) according to an embodiment of the present invention. [Figure 25] FIG. 21C is a conceptual diagram for explaining the continuity of the surface of FIG. 21C, which is an image acquired by the image reconstruction process or the regional packing process in the CMP projection format. [Figure 26] 10A and 10B are exemplary diagrams illustrating image size adjustment in a CMP projection format according to an embodiment of the present invention. [Figure 27] 10A and 10B are exemplary diagrams illustrating size adjustment for an image that has been converted into a CMP projection format and packed according to an embodiment of the present invention. [Figure 28] 1 is an exemplary diagram illustrating a data processing method for adjusting the size of a 360-degree image according to an embodiment of the present invention. [Figure 29] FIG. 10 is an exemplary diagram showing a tree-based block shape. [Figure 30] FIG. 10 is an exemplary diagram showing a type-based block shape. [Figure 31] 10A to 10C are exemplary diagrams showing various block shapes that can be obtained by the block dividing unit of the present invention. [Figure 32] FIG. 1 is an exemplary diagram illustrating tree-based partitioning according to an embodiment of the present invention. [Figure 33] FIG. 1 is an exemplary diagram illustrating tree-based partitioning according to an embodiment of the present invention. [Figure 34] 10A and 10B are diagrams illustrating various cases in which a prediction block is obtained by inter-picture prediction; [Figure 35] FIG. 10 is an exemplary diagram illustrating a reference picture list according to an embodiment of the present invention. [Figure 36] FIG. 2 is a conceptual diagram illustrating a non-moving motion model according to an embodiment of the present invention. [Figure 37] 1 is an example diagram illustrating sub-block-based motion estimation according to an embodiment of the present invention; [Figure 38] 1 is an exemplary diagram illustrating blocks referenced in motion information prediction of a current block according to an embodiment of the present invention; [Figure 39] 1 is an exemplary diagram illustrating blocks referenced for motion information prediction of a current block in a non-motion motion model according to an embodiment of the present invention; [Figure 40] 1 is an exemplary diagram illustrating inter-frame prediction using an extended picture according to an embodiment of the present invention; [Figure 41] FIG. 10 is a conceptual diagram illustrating the expansion of surface units according to an embodiment of the present invention. [Figure 42] 1 is a diagram illustrating an example of performing inter-frame prediction using an extended image according to an embodiment of the present invention; [Figure 43] 1 is an exemplary diagram illustrating inter-frame prediction using extended reference pictures according to an embodiment of the present invention; [Figure 44] FIG. 10 is an illustrative diagram showing a configuration of a motion information prediction candidate group for inter-picture prediction in a 360-degree image according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] Although the present invention can be modified in various ways and can have various embodiments, specific embodiments are illustrated in the drawings and described in detail herein. However, it should be understood that this does not limit the present invention to the specific embodiments, and that it includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention.
[0028] Terms such as "first," "second," "A," and "B" are used to describe various elements, but the elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element can be designated a second element, and similarly, a second element can be designated a first element, without departing from the scope of the present invention. The term "and / or" includes a combination of two or more related listed items or any of two or more related listed items.
[0029] When a component is said to be "coupled" or "connected" to another component, it means that the component is directly coupled or connected to the other component, but it should be understood that there may be other components between them. In contrast, when a component is said to be "directly coupled" or "directly connected" to another component, it should be understood that there are no other components between them.
[0030] The terms used in this specification are merely used to describe specific embodiments and do not limit the present invention. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or additional possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0031] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which this invention belongs. Terms defined in commonly used dictionaries should be interpreted in accordance with the contextual meaning of the relevant art, and should not be interpreted in an ideal or overly formal sense unless expressly defined herein.
[0032] The image encoding and decoding devices may be user terminals such as personal computers (PCs), notebook computers, personal digital assistants (PDAs), portable multimedia players (PMPs), PlayStation Portables (PSPs), wireless communication terminals, smartphones, TVs, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, head mounted displays (HMDs), or smart glasses, or may be server terminals such as application servers or service servers. The image encoding and decoding devices may include various devices including a communication device such as a communication modem for communicating with various devices or wired or wireless communication networks, a memory for storing various programs and data for encoding or decoding images or for performing intra-screen or inter-screen prediction for encoding or decoding, and a processor for executing programs, performing calculations, and control. In addition, the image coded into a bitstream by the image coding device can be transmitted to the image decoding device in real time or non-real time via wired or wireless communication networks such as the Internet, a short-range wireless communication system, a wireless LAN network, a WiBro network, or a mobile communication network, or via various communication interfaces such as a cable or a universal serial bus (USB), and can be decoded by the image decoding device to restore and play back the image.
[0033] Furthermore, an image coded into a bitstream by an image coding device can also be transmitted from the coding device to a decoding device via a computer-readable recording medium.
[0034] The image encoding device and the image decoding device may be separate devices, but may be implemented as a single image encoding / decoding device. In this case, some components of the image encoding device may be substantially the same technical elements as some components of the image decoding device, and may be implemented to include at least the same structure or perform at least the same function.
[0035] Therefore, in the detailed description of the technical elements and their operating principles below, redundant descriptions of the corresponding technical elements will be omitted.
[0036] The image decoding apparatus corresponds to a computer device that applies the image coding method performed by the image coding apparatus to decoding, and therefore the following description will focus on the image coding apparatus.
[0037] The computer device may include a memory for storing a program or software module for implementing the image encoding method and / or the image decoding method, and a processor connected to the memory for executing the program. An image encoding device may be called an encoder, and an image decoding device may be called a decoder.
[0038] Typically, an image can be composed of a series of still images, which can be divided into GOP (Group of Pictures) units. Each still image can be called a picture. A picture can refer to either a frame or a field in a progressive or interlaced signal. An image can be referred to as a "frame" when encoding / decoding is performed in frame units, or as a "field" when encoding / decoding is performed in field units. While the present invention will be described assuming a progressive signal, it can also be applied to an interlaced signal. Higher-level units such as GOP and sequence can exist. Each picture can be divided into predetermined regions such as slices, tiles, and blocks. One GOP can include units such as I-pictures, P-pictures, and B-pictures. An I picture may refer to a picture that is encoded / decoded by itself without using a reference picture, and P and B pictures may refer to pictures that are encoded / decoded using a reference picture through processes such as motion estimation and motion compensation. Generally, a P picture can use an I picture or a P picture as a reference picture, and a B picture can use an I picture or a P picture as a reference picture, but the above definitions can be changed depending on the encoding / decoding settings.
[0039] Here, a picture referenced in encoding / decoding is called a reference picture, and a block or pixel referenced is called a reference block or reference pixel. The reference data may be not only pixel values in the spatial domain but also coefficient values in the frequency domain, and various encoding / decoding information generated or determined during the encoding / decoding process. For example, this may include intra-frame prediction-related information or motion-related information in a predictor, transform-related information in a transformer / inverse transformer, quantization-related information in a quantizer / inverse quantizer, encoding / decoding-related information (context information) in an encoder / decoder, and filter-related information in an in-loop filter.
[0040] The smallest unit of an image may be a pixel. The number of bits used to represent one pixel is called the bit depth. Generally, the bit depth is 8 bits, but higher bit depths may be supported depending on the encoding settings. At least one bit depth may be supported depending on the color space. Furthermore, depending on the color format of the image, an image may be configured with at least one color space. Depending on the color format, an image may be configured with one or more pictures of a uniform size or one or more pictures of different sizes. For example, in the case of YCbCr 4:2:0, an image may be configured with one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example). In this case, the chrominance component to luminance component ratio may be 1:2 horizontally:vertically. As another example, in the case of 4:4:4, the horizontal and vertical ratios may be equal. When an image is configured with one or more color spaces as in the above example, the picture may be divided into each color space.
[0041] The present invention will be described based on a certain color space (Y in this example) of a certain color format (YCbCr in this example). The same or similar application (specific color space-dependent setting) can be made to other color spaces (Cb, Cr in this example) according to the color format. However, partial differences can also be made for each color space (specific color space-independent setting). That is, a setting that is dependent on each color space can mean that the setting is proportional to or dependent on the component ratio of each component (e.g., determined according to 4:2:0, 4:2:2, 4:4:4, etc.), and a setting that is independent on each color space can mean that the setting is independent of or independent of the component ratio of each component and is only for the corresponding color space. In the present invention, depending on the encoder / decoder, some configurations can have independent or dependent settings.
[0042] The configuration information or syntax elements required for the image coding process can be determined at the unit level, such as video, sequence, picture, slice, tile, or block. This can be recorded in the bitstream in units such as video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), slice header, tile header, or block header and transmitted to the decoder. The decoder parses the configuration information transmitted from the encoder at the same level to restore it and use it in the image decoding process. In addition, related information can be transmitted to the bitstream in the form of supplemental enhancement information (SEI) or metadata, and parsed for use. Each parameter set has a unique ID value, and lower-level parameter sets can have the ID value of the higher-level parameter set they reference. For example, a lower-level parameter set can reference information from one or more higher-level parameter sets with a matching ID value. Among the various examples of units mentioned above, when one unit contains one or more other units, the corresponding unit is sometimes called a higher-order unit, and the contained units are sometimes called lower-order units.
[0043] In the case of the setting information generated in the unit, it may include content for independent setting for each corresponding unit, or content for dependent setting for the previous, subsequent, or higher unit. Here, dependent setting can be understood as indicating setting information for the corresponding unit by flag information indicating that the setting of the previous, subsequent, or higher unit is to be followed (for example, if a 1-bit flag is 1, the setting is followed; if it is 0, the setting is not followed). Although the setting information in the present invention will be described mainly as an example of independent setting, examples of addition or replacement of content for a dependent relationship with setting information of the previous, subsequent, or higher unit of the current unit may also be included.
[0044] Fig. 1 is a block diagram of an image encoding device according to an embodiment of the present invention, and Fig. 2 is a block diagram of an image decoding device according to an embodiment of the present invention.
[0045] Referring to FIG. 1, the image coding device may be configured to include a prediction unit, a subtraction unit, a transformation unit, a quantization unit, an inverse quantization unit, an inverse transformation unit, an addition unit, an in-loop filter unit, a memory, and / or a coding unit, and some of the above components may not necessarily be included, and some or all of them may be selectively included depending on the implementation, or additional components not shown may be included.
[0046] Referring to FIG. 2, the image decoding device may be configured to include a decoding unit, a prediction unit, an inverse quantization unit, an inverse transform unit, an addition unit, an in-loop filter unit, and / or a memory, and some of the above components may not necessarily be included, and depending on the implementation, some or all of them may be selectively included, and some additional components not shown may be included.
[0047] The image encoding device and the image decoding device may be separate devices, but may be implemented as a single image encoding / decoding device. In this case, some components of the image encoding device may be substantially the same technical elements as some components of the image decoding device, and may be implemented to include at least the same structure or perform at least the same functions. Therefore, in the following detailed description of the technical elements and their operating principles, redundant descriptions of corresponding technical elements will be omitted. Since the image decoding device corresponds to a computer device that applies the image encoding method performed in the image encoding device to decoding, the following description will focus on the image encoding device. The image encoding device may be called an encoder, and the image decoding device may be called a decoder.
[0048] The prediction unit may be implemented using a prediction module, which is a software module, and may generate a predicted block for a block to be coded using an intra-frame prediction method or an inter-frame prediction method. The prediction unit generates a predicted block by predicting a current block to be coded in an image. That is, the prediction unit predicts pixel values of each pixel of the current block to be coded in an image using intra-frame prediction or inter-frame prediction, and generates a predicted block having predicted pixel values of each generated pixel. In addition, the prediction unit may transmit information required for generating the predicted block to the coding unit to encode information regarding a prediction mode, and record the information in a bitstream and transmit it to a decoder. The decoding unit of the decoder parses the information to restore information regarding the prediction mode and then uses the information for intra-frame prediction or inter-frame prediction.
[0049] The subtractor subtracts the predicted block from the current block to generate a residual block. That is, the subtractor calculates the difference between the pixel value of each pixel of the current block to be encoded and the predicted pixel value of each pixel of the predicted block generated by the predictor, and generates the residual block, which is a residual signal in the form of a block.
[0050] The transform unit can transform a signal belonging to the spatial domain into a signal belonging to the frequency domain. At this time, the signal obtained through the transform process is called a transformed coefficient. For example, a transform block having transform coefficients can be obtained by transforming a residual block having a residual signal transmitted from the subtraction unit. The input signal is determined according to coding settings and is not limited to a residual signal.
[0051] The transform unit may transform the residual block using a transform technique such as a Hadamard transform, a discrete sine transform (DST-based transform), a discrete cosine transform (DCT-based transform), etc. However, the transform is not limited to these, and various improved and modified transform techniques may be used.
[0052] For example, at least one of the above transforms may be supported, and each transform may support at least one detailed transform. In this case, the at least one detailed transform may be a transform in which some of the basis vectors are different for each transform. For example, DST-based transforms and DCT-based transforms may be supported as transform techniques. For the DST, detailed transform techniques such as DST-I, DST-II, DST-III, DST-V, DST-VI, DST-VII, and DST-VIII may be supported. For the DCT, detailed transform techniques such as DCT-I, DCT-II, DCT-III, DCT-V, DCT-VI, DCT-VII, and DCT-VIII may be supported.
[0053] Any of the above transformations (e.g., one transformation method and one detailed transformation method) can be set as a basic transformation method, and additional transformation methods (e.g., multiple transformation methods || multiple detailed transformation methods) can be supported. Whether to support additional transformation methods can be determined in units such as a sequence, a picture, a slice, or a tile, and related information can be generated for the units. When additional transformation methods are supported, transformation method selection information can be determined in units such as a block, and related information can be generated.
[0054] Transformations can be performed horizontally or vertically. For example, spatial pixel values can be transformed into the frequency domain using a 1-D transform in the horizontal direction and a 1-D transform in the vertical direction, resulting in a total 2-D transform using the basis vectors in the transform.
[0055] In addition, the transform may be adaptively performed in the horizontal / vertical directions. In particular, whether the transform is adaptively performed or not may be determined according to at least one encoding setting. For example, when the prediction mode in intra prediction is a horizontal mode, DCT-I may be applied in the horizontal direction and DST-I may be applied in the vertical direction; when the prediction mode in intra prediction is a vertical mode, DST-VI may be applied in the horizontal direction and DCT-VI may be applied in the vertical direction; when the prediction mode is a diagonal down left mode, DCT-II may be applied in the horizontal direction and DCT-V may be applied in the vertical direction; and when the prediction mode is a diagonal down right mode, DST-I may be applied in the horizontal direction and DST-VI may be applied in the vertical direction.
[0056] The size and shape of each transform block are determined according to the candidate encoding cost for the size and shape of the transform block, and information such as the image data of each determined transform block and the size and shape of each determined transform block can be encoded.
[0057] Among the transformation shapes, a square transformation can be set as a basic transformation shape, and additional transformation shapes (e.g., rectangular shapes) can be supported. Whether additional transformation shapes are supported is determined in units such as a sequence, a picture, a slice, or a tile, and related information can be generated in these units. Transformation shape selection information can be determined in units such as a block, and related information can be generated.
[0058] In addition, the support of transform block shapes can be determined according to coding information. The coding information can include slice type, coding mode, block size and shape, block division method, etc. That is, one transform shape can be supported according to at least one piece of coding information, and multiple transform shapes can be supported according to at least one piece of coding information. The former can be an implicit situation, while the latter can be an explicit situation. In the explicit case, adaptive selection information indicating the optimal candidate group among multiple candidates can be generated and recorded in the bitstream. In the present invention, including this example, when explicitly generating coding information, the corresponding information can be recorded in the bitstream in various units, and the decoder can parse the related information in various units and restore it to decoded information. In addition, when processing encoding / decoding information implicitly, the encoder and decoder can be processed according to the same process, rules, etc.
[0059] For example, rectangular transformation support can be determined according to the slice type: for an I slice, the transformation shape supported is a square transformation, and for a P / B slice, the transformation shape supported may be a square or rectangular transformation.
[0060] For example, rectangular transformation support can be determined depending on the encoding mode. In the case of Intra, the transformation shape supported is square, and in the case of Inter, the transformation shape supported may be square or rectangular.
[0061] For example, rectangular transformation support can be determined according to the size and shape of the block. The transformation shape supported for blocks of a certain size or larger is a square transformation, and the transformation shape supported for blocks of a certain size or smaller can be a square or rectangular transformation.
[0062] For example, rectangular transformation support may be determined depending on the block division method: if the block to be transformed is a block obtained by a quad tree division method, the supported transformation shape may be a square transformation, and if the block is obtained by a binary tree division method, the supported transformation shape may be a square or rectangular transformation.
[0063] The above example is an example of supporting a transformation shape according to one piece of coding information, and multiple pieces of information may be combined to support additional transformation shapes. The above example is merely an example of supporting additional transformation shapes according to various coding settings, and is not limited to the above, and various modified examples are possible.
[0064] Depending on the encoding settings or image characteristics, the transform process may be omitted. For example, depending on the encoding settings (assuming a lossless compression environment in this example), the transform process (including the inverse process) may be omitted. As another example, if the transform does not provide sufficient compression performance depending on the image characteristics, the transform process may be omitted. In this case, the omitted transform may be for the entire block or for either a horizontal or vertical block. Whether or not to support such omission may be determined depending on the size and shape of the block.
[0065] For example, in a setting where the omission of horizontal and vertical transformations is grouped, if the transformation omission flag is 1, transformations in the horizontal and vertical directions are not performed, and if the transformation omission flag is 0, transformations in the horizontal and vertical directions may be performed. In a setting where the omission of horizontal and vertical transformations operates independently, if the first transformation omission flag is 1, transformations in the horizontal direction are not performed, if the first transformation omission flag is 0, transformations in the horizontal direction are performed, if the second transformation omission flag is 1, transformations in the vertical direction are not performed, and if the second transformation omission flag is 0, transformations in the vertical direction are performed.
[0066] If the block size falls within range A, transformation skipping can be supported, and if the block size falls within range B, transformation skipping cannot be supported. For example, if the block width is larger than M or the block height is larger than N, the transformation skip flag cannot be supported, and if the block width is smaller than m or the block height is smaller than n, the transformation skip flag can be supported. M(m) and N(n) may be the same or different. The transformation-related settings can be determined in units of sequences, pictures, slices, etc.
[0067] If additional transform techniques are supported, the setting of the transform technique may be determined according to at least one coding information, which may include a slice type, a coding mode, a block size and shape, a prediction mode, etc.
[0068] For example, the support of transform techniques may be determined depending on the coding mode. In the case of Intra, the supported transform techniques may be DCT-I, DCT-III, DCT-VI, DST-II, and DST-III, and in the case of Inter, the supported transform techniques may be DCT-II, DCT-III, and DST-III.
[0069] For example, the supported transform techniques may be determined according to the slice type. For an I slice, the supported transform techniques may be DCT-I, DCT-II, and DCT-III; for a P slice, the supported transform techniques may be DCT-V, DST-V, and DST-VI; and for a B slice, the supported transform techniques may be DCT-I, DCT-II, and DST-III.
[0070] For example, support of a transform technique may be determined according to a prediction mode. The transform techniques supported in prediction mode A may be DCT-I and DCT-II, the transform techniques supported in prediction mode B may be DCT-I and DST-I, and the transform technique supported in prediction mode C may be DCT-I. In this case, prediction modes A and B may be directional modes, and prediction mode C may be a non-directional mode.
[0071] For example, the support of a transform technique may be determined according to the size and shape of the block. The transform technique supported for blocks of a certain size or larger may be DCT-II, the transform technique supported for blocks of a certain size or smaller may be DCT-II and DST-V, and the transform techniques supported for blocks of a certain size or larger and smaller may be DCT-I, DCT-II, and DST-I. Also, the transform techniques supported for square shapes may be DCT-I and DCT-II, and the transform techniques supported for rectangular shapes may be DCT-I and DST-I.
[0072] The above example is an example of supporting a transform technique according to one piece of coding information, and multiple pieces of information may be combined to support additional transform techniques. The present invention is not limited to the above example, and other modifications are possible. Furthermore, the transform unit may transmit information necessary to generate a transform block to the encoder to encode the same, record the resulting information in a bitstream, and transmit the bitstream to the decoder. The decoder of the decoder may parse the information and use it in the inverse transform process.
[0073] The quantization unit may quantize an input signal. At this time, a signal obtained through the quantization process is called a quantized coefficient. For example, a residual block having a residual transform coefficient transmitted from the transform unit may be quantized to obtain a quantized block having a quantized coefficient. However, the input signal is determined according to a coding setting, and is not limited to a residual transform coefficient.
[0074] The quantization unit may quantize the transformed residual block using a quantization technique such as Dead Zone Uniform Threshold Quantization, Quantization Weighted Matrix, etc., but is not limited thereto, and various improved and modified quantization techniques may be used. Whether to support an additional quantization technique may be determined in units such as a sequence, a picture, a slice, a tile, etc., and related information may be generated in the units. If an additional quantization technique is supported, quantization technique selection information may be determined in units such as a block, and related information may be generated.
[0075] If additional quantization techniques are supported, the setting of the quantization technique may be determined according to at least one piece of coding information, which may include a slice type, a coding mode, a block size and shape, a prediction mode, etc.
[0076] For example, the quantizer may set a quantization weight matrix according to a coding mode and a weight matrix applied according to inter prediction / intra prediction to be different from each other. Also, the quantizer may set a weight matrix applied according to an intra prediction mode to be different. In this case, the quantization weight matrix may be a quantization matrix with some different quantization components, assuming that the block size is the same as the quantization block size, with a size of M×N.
[0077] The quantization process can be omitted depending on the encoding settings or image characteristics. For example, the quantization process (including the inverse process) can be omitted depending on the encoding settings (assuming a lossless compression environment in this example). As another example, the quantization process can be omitted if the compression performance through quantization is not achieved depending on the image characteristics. In this case, the omitted region can be the entire region or a part of the region. Whether or not such omission is supported can be determined depending on the size and shape of the block.
[0078] Information about quantization parameters (QP) can be generated in units such as sequences, pictures, slices, tiles, and blocks. For example, a basic QP can be set in the higher unit where QP information is first generated. <1> The lower the unit, the higher the QP can be set to a value that is the same as or different from the QP set in the higher unit. <2> Through this process, the QP can be finally determined in the quantization process performed in some units. <3> In this case, the units of sequences, pictures, etc. are <1> The units of slices, tiles, blocks, etc. are <2> The units of blocks are <3> This may be an example of the above.
[0079] The information about the QP may be generated based on the QP for each unit. Alternatively, a predetermined QP may be set as a predicted value to generate difference value information from the QP for each unit. Alternatively, a QP obtained based on at least one of a QP set for a higher unit, a QP previously set for the same unit, or a QP set for an adjacent unit may be set as a predicted value to generate difference value information from the QP for the current unit. Alternatively, a QP set for a higher unit and a QP obtained based on at least one of coding information may be set as a predicted value to generate difference value information from the QP for the current unit. In this case, the same previous unit may be a unit that can be defined according to the coding order of each unit, the adjacent unit may be a unit that is spatially adjacent, and the coding information may be a slice type, coding mode, prediction mode, position information, etc. of the corresponding unit.
[0080] For example, the QP of the current unit may generate differential value information by setting the QP of the upper unit as a predicted value. Differential value information may be generated between a QP set in a slice and a QP set in a picture, or between a QP set in a tile and a QP set in a picture. Furthermore, differential value information may be generated between a QP set in a block and a QP set in a slice or a tile. Furthermore, differential value information may be generated between a QP set in a sub-block and a QP set in a block.
[0081] For example, the QP of the current unit may generate differential value information by setting a QP obtained based on the QP of at least one adjacent unit or a QP obtained based on the QP of at least one previous unit as a predicted value. Differential value information from a QP obtained based on the QP of a neighboring block, such as the left, upper left, lower left, upper, or upper right of the current block, may be generated. Alternatively, differential value information from a QP of a picture coded before the current picture may be generated.
[0082] For example, the QP of the current unit may generate difference value information by setting the QP of the upper unit and a QP obtained based on at least one encoding information as a predicted value. Difference value information between the QP of the current block and the QP of a slice that is corrected according to a slice type (I / P / B) may be generated. Alternatively, difference value information between the QP of the current block and the QP of a tile that is corrected according to a coding mode (Intra / Inter) may be generated. Alternatively, difference value information between the QP of the current block and the QP of a picture that is corrected according to a prediction mode (directional / non-directional) may be generated. Alternatively, difference value information between the QP of the current block and the QP of a picture that is corrected according to position information (x / y) may be generated. In this case, the correction may mean that an offset is added or subtracted from the QP of the upper unit used for prediction. At least one offset information may be supported according to the encoding setting, and the offset information may be implicitly processed or explicitly generated according to a predetermined process. The present invention is not limited to the above example, and other modifications are possible.
[0083] The above example may be a possible example when a signal indicating a QP variation is provided or activated. For example, when a signal indicating a QP variation is not provided or is inactivated, differential value information is not generated, and the predicted QP can be determined as the QP of each unit. As another example, when a signal indicating a QP variation is provided or activated, differential value information is generated, and when the value is 0, the predicted QP can be determined as the QP of each unit.
[0084] The quantization unit can transmit the information necessary to generate a quantization block to the encoding unit to encode it, and then record the information into a bitstream and transmit it to the decoder, and the decoding unit of the decoder can parse the information and use it in the inverse quantization process.
[0085] In the above example, the explanation is based on the assumption that the residual block is transformed and quantized through a transform unit and a quantization unit, but the residual signal may be transformed to generate a residual block having transform coefficients without performing the quantization process, or the residual signal of the residual block may be quantized without being transformed into transform coefficients, or both the transform and quantization processes may not be performed. This can be determined according to the settings of the encoder.
[0086] The encoding unit may scan quantized coefficients, transform coefficients, or residual signals of the generated residual block according to at least one scan order (e.g., zigzag scan, vertical scan, horizontal scan, etc.) to generate a quantized coefficient sequence, a transform coefficient sequence, or a signal sequence, and encode the quantized coefficient sequence, transform coefficient sequence, or signal sequence using at least one entropy coding technique. In this case, information about the scan order may be determined according to an encoding setting (e.g., a coding mode, a prediction mode, etc.), and may be implicitly determined or related information may be explicitly generated. For example, one of a plurality of scan orders may be selected according to an intra-frame prediction mode.
[0087] In addition, coded data including coding information transmitted from each component can be generated and output as a bitstream, which can be achieved by a multiplexer (MUX).In this case, coding techniques such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc., can be used, but are not limited to these, and various coding techniques that are improvements or modifications of these can also be used.
[0088] When performing entropy coding (assumed to be CABAC in this example) on syntax elements such as the residual block data and information generated during the encoding / decoding process, the entropy coding device may include a binarizer, a context modeler, and a binary arithmetic coder. In this case, the binary arithmetic coder may include a regular coding engine and a bypass coding engine.
[0089] Since the syntax elements input to the entropy coding device may not be binary, if the syntax elements are not binary, a binarization unit binarizes the syntax elements and outputs a Bin string consisting of 0 or 1. Here, Bin indicates a bit consisting of 0 or 1 and can be coded through a binary arithmetic coding unit. Here, either a regular coding unit or a bypass coding unit can be selected based on the occurrence probability of 0 and 1. This can be determined according to coding / decoding settings. If the syntax elements are data in which the frequency of 0 and the frequency of 1 are the same, the bypass coding unit can be used, and if not, the regular coding unit can be used.
[0090] Various methods can be used to binarize the syntax elements. For example, fixed length binarization, unary binarization, truncated Rice binarization, K-th Exp-Golomb binarization, etc. can be used. Also, signed or unsigned binarization can be performed depending on the value range of the syntax element. The binarization process for the syntax elements generated in the present invention can be performed not only by the binarization methods mentioned in the above examples, but also by other additional binarization methods.
[0091] The inverse quantization unit and the inverse transform unit can be implemented by reversing the processes of the transform unit and the quantization unit. For example, the inverse quantization unit can inversely quantize the quantized transform coefficients generated by the quantization unit, and the inverse transform unit can inversely transform the inversely quantized transform coefficients to generate a reconstructed residual block.
[0092] The adder adds the predicted block and the reconstructed residual block to reconstruct the current block. The reconstructed block is stored in a memory and can be used as reference data (for the predictor and filter, etc.).
[0093] The in-loop filter unit may include at least one post-processing filter process, such as a deblocking filter, a sample adaptive offset (SAO), or an adaptive loop filter (ALF). The deblocking filter may remove block artifacts that occur at boundaries between blocks from a restored image. The ALF may perform filtering based on a value obtained by comparing a restored image with an input image. Specifically, the deblocking filter may perform filtering based on a value obtained by comparing a restored image with an input image after filtering blocks through the deblocking filter. Alternatively, the SAO may perform filtering based on a value obtained by comparing a restored image with an input image after filtering blocks through the SAO. The SAO restores an offset difference based on a value obtained by comparing a restored image with an input image, and may be applied in the form of a band offset (BO), edge offset (EO), etc. Specifically, the SAO may add an offset from the original image in at least one pixel unit to the restored image to which the deblocking filter has been applied, and may be applied in the form of a BO, an EO, etc. In detail, an offset from the original image is added to the restored image after the block is filtered through ALF, and can be applied in the form of BO, EO, etc., in pixel units.
[0094] As the filtering-related information, setting information on whether to support each post-processing filter may be generated in units of a sequence, a picture, a slice, a tile, a block, etc. In addition, setting information on whether to execute each post-processing filter may be generated in units of a picture, a slice, a tile, a block, etc. The range to which the filter is applied may be divided into the inside of an image and an image boundary, and setting information may be generated taking this into consideration. In addition, information related to the filtering operation may be generated in units of a picture, a slice, a tile, a block, etc. The information may perform implicit or explicit processing, and the filtering may apply an independent filtering process or a dependent filtering process depending on color components. This may be determined according to encoding settings. The in-loop filter unit may transmit the filtering-related information to the encoder to encode it, and may include the resulting information in a bitstream and transmit it to a decoder. The decoding unit of the decoder may parse the corresponding information and apply it to the in-loop filter unit.
[0095] The memory may store reconstructed blocks or pictures. The reconstructed blocks or pictures stored in the memory may be provided to a prediction unit that performs intra-frame prediction or inter-frame prediction. In particular, a queue-type storage space for compressed bitstreams in an encoder may be stored and processed as a coded picture buffer (CPB), and a space for storing decoded images in picture units may be stored and processed as a decoded picture buffer (DPB). In the case of the CPB, decoding units are stored in the decoding order, and the decoding operation may be emulated in the encoder and the compressed bitstream may be stored during the emulation process. The bitstream output from the CPB is restored through a decoding process, and the restored images are stored in the DPB. The pictures stored in the DPB may be referenced in subsequent image encoding and decoding processes.
[0096] The decoding unit can be realized by performing the process of the encoding unit in reverse. For example, it can receive and decode a quantized coefficient sequence, a transform coefficient sequence, or a signal sequence from a bitstream, parse decoded data including decoding information, and transmit the parsed data to each component.
[0097] An image setting process applied to an image encoding / decoding device according to an embodiment of the present invention will now be described. This may be an example applied to a stage before encoding / decoding (image initial setting), but some processes may also be applicable to other stages (e.g., a stage after encoding / decoding or an internal stage of encoding / decoding). The image setting process may be performed taking into consideration network and user environments such as characteristics of multimedia content, bandwidth, and performance and accessibility of a user terminal. For example, image division, image resizing, image reconstruction, etc. may be performed according to the encoding / decoding setting. The image setting process described below will be described mainly for rectangular images, but is not limited thereto and may also be applied to polygonal images. Regardless of the shape of the image, the same image setting or different image settings may be applied. This can be determined according to the encoding / decoding setting. For example, after information on the shape of the image (e.g., rectangular or non-rectangular shape) is confirmed, information on the image setting may be configured accordingly.
[0098] In the examples described below, it is assumed that color space-dependent settings are provided; however, color space-independent settings are also possible. Furthermore, in the examples described below, independent settings may include examples in which encoding / decoding settings are provided independently for each color space, and even if a description is given for one color space, it is assumed that examples applied to other color spaces (e.g., when M is generated for the luminance component, N is generated for the chrominance component) are included and derived. Furthermore, dependent settings may include examples in which settings proportional to the color format composition ratio (e.g., 4:4:4, 4:2:2, 4:2:0, etc.) are provided (e.g., in the case of 4:2:0, when M is generated for the luminance component, M / 2 is generated for the chrominance component). Even without a special description, it is assumed that examples applied to each color space are included and derived. This is not limited to the above examples, and may be a description commonly applicable to the present invention.
[0099] Some of the configurations in the examples described below may be applicable to various encoding techniques such as encoding in the spatial domain, encoding in the frequency domain, block-based encoding, and object-based encoding.
[0100] Although it is common to encode / decode an input image as is, there are also cases where an image is divided and then coded / decoded. For example, division can be performed for error resilience purposes to prevent damage caused by packet corruption during transmission. Alternatively, division can be performed to classify regions with different properties within the same image according to the characteristics or type of the image.
[0101] In the present invention, the image division process can include a division process and its inverse process. In the examples described below, the division process will be mainly described, but the content of the inverse division process can be derived from the division process.
[0102] FIG. 3 is an example diagram showing image information divided into layers for image compression.
[0103] 3a is an example diagram showing an image sequence composed of multiple GOPs. One GOP can be composed of I pictures, P pictures, and B pictures as in 3b. One picture can be composed of slices, tiles, etc. as in 3c. Slices, tiles, etc. are composed of multiple basic coding units as in 3d, and the basic coding unit can be composed of at least one sub-coding unit as in 3e. The image setting process of the present invention will be described based on examples applied to units such as pictures, slices, and tiles as in 3b and 3c.
[0104] FIG. 4 is a conceptual diagram illustrating various examples of image segmentation according to an embodiment of the present invention.
[0105] 4a is a conceptual diagram of an image (e.g., a picture) divided horizontally and vertically at regular intervals. The divided areas can be called blocks, and each block is a basic coding unit (or maximum coding unit) obtained through a picture dividing unit, and can also be a basic unit applied to division units, which will be described later.
[0106] 4b is a conceptual diagram of an image divided in at least one of the horizontal and vertical directions. The divided regions T0 to T3 can be called tiles, and each region can be coded / decoded independently or dependently of the other regions.
[0107] 4c is a conceptual diagram of dividing an image into groups of consecutive blocks. The divided regions S0 and S1 can be called slices, and each region can be an area for encoding / decoding independently or dependently from other regions. The group of consecutive blocks can be defined according to a scan order, which generally follows a raster scan order, but is not limited to this, and can be determined according to the encoding / decoding settings.
[0108] 4d is a conceptual diagram of an image divided into groups of blocks in an arbitrary setting defined by the user. The divided areas A0 to A2 can be called arbitrary partitions, and each area can be an area for encoding / decoding independently or dependently of other areas.
[0109] Independent encoding / decoding may mean that data of other units cannot be referenced when encoding / decoding a certain unit (or region). Specifically, information used or generated in texture encoding and entropy encoding of a certain unit is encoded independently without reference to each other, and a decoder may also not reference parsing information and reconstruction information of other units for texture decoding and entropy decoding of a certain unit. In this case, whether data of other units (or regions) can be referenced may be restricted in the spatial domain (e.g., between regions within an image), but restrictions may also be set in the temporal domain (e.g., between consecutive images or frames) depending on the encoding / decoding settings. For example, if a certain unit of a current image and a certain unit of another image have continuity or the same encoding environment, they can be referenced, and if not, reference may be restricted.
[0110] Also, dependent encoding / decoding may mean that data of other units can be referenced when encoding / decoding a certain unit. In particular, information used or generated in texture encoding and entropy encoding of a certain unit is referenced to each other and coded dependently, and similarly, parsing information and reconstruction information of other units can be referenced to each other for texture decoding and entropy decoding of a certain unit in a decoder. That is, the setting may be the same as or similar to general encoding / decoding. In this case, depending on the characteristics and type of the image (e.g., a 360-degree image), a region (in this example, a surface generated according to a projection format) may be coded. <face>This may be the case when the data is split for the purpose of identifying a specific entity (e.g., a specific entity).
[0111] In the above example, some units (slices, tiles, etc.) can have independent encoding / decoding settings (e.g., independent slice segments), and some units can have dependent encoding / decoding settings (e.g., dependent slice segments). In this invention, we will mainly describe independent encoding / decoding settings.
[0112] The basic coding unit obtained through the picture division unit as shown in 4a is divided into basic coding blocks according to the color space, and the size and shape can be determined according to the characteristics and resolution of the image. The supported block size or shape is one in which the width and height are exponential powers of 2 (2 n ) is an N × N square (2 n ×2 n 256x256, 128x128, 64x64, 32x32, 16x16, 8x8, etc. n is an integer between 3 and 8), or an MxN rectangle (2 m ×2 n For example, depending on the resolution, an input image can be divided into sizes such as 128x128 for an 8k UHD image, 64x64 for a 1080p HD image, and 16x16 for a WVGA image. Depending on the type of image, an input image can be divided into sizes such as 256x256 for a 360-degree image. A basic coding unit can be divided into sub-coding units for encoding / decoding, and information about the basic coding unit can be recorded in a bitstream and transmitted in units such as sequence, picture, slice, and tile. This can be parsed by the decoder to restore related information.
[0113] An image encoding method and a decoding method according to an embodiment of the present invention may include the following image segmentation steps. In this case, the image segmentation process may include an image segmentation instruction step, an image segmentation type identification step, and an image segmentation execution step. Also, the image encoding device and the decoding device may be configured to include an image segmentation instruction unit, an image segmentation type identification unit, and an image segmentation execution unit that realize the image segmentation instruction step, the image segmentation type identification step, and the image segmentation execution step. In the case of encoding, associated syntax elements may be generated, and in the case of decoding, the associated syntax elements may be parsed.
[0114] In each block division process of 4a, the image division instruction unit can be omitted, and the image division type identification unit is a process of checking information about the size and shape of the block, and the image division unit can perform division in basic coding units based on the identified division type information.
[0115] While blocks may always be the unit of division, whether or not to divide other division units (tiles, slices, etc.) can be determined depending on the encoding / decoding settings. The picture division unit may be set to perform division in block units first, followed by division in other units. In this case, block division may be performed based on the size of the picture.
[0116] Alternatively, the image may be divided into other units (tiles, slices, etc.) and then divided into blocks. That is, block division may be performed based on the size of the division unit. This may be determined through explicit or implicit processing depending on the encoding / decoding settings. In the following example, the former case will be assumed, and units other than blocks will be mainly described.
[0117] In the image segmentation instruction step, it is possible to determine whether to perform image segmentation. For example, if a signal instructing image segmentation (e.g., tiles_enabled_flag) is checked, segmentation can be performed, and if the signal instructing image segmentation is not checked, segmentation can be omitted or segmentation can be performed by checking other encoding / decoding information.
[0118] In detail, when a signal instructing image division (e.g., tiles_enabled_flag) is checked, if the corresponding signal is activated (e.g., tiles_enabled_flag=1), division can be performed in multiple units, and if the corresponding signal is deactivated (e.g., tiles_enabled_flag=0), division can be not performed. Alternatively, if the signal instructing image division is not checked, it may mean that division is not performed or that division is performed in at least one unit, and whether division is performed in multiple units can be checked via another signal (e.g., first_slice_segment_in_pic_flag).
[0119] In summary, when a signal indicating image division is provided, the signal indicates whether to divide the image into multiple units, and whether the image is divided can be confirmed according to the signal. For example, when tiles_enabled_flag is a signal indicating whether to divide the image, if tiles_enabled_flag is 1, it may indicate that the image is divided into multiple tiles, and if it is 0, it may indicate that the image is not divided.
[0120] In summary, if a signal instructing image division is not provided, division is not performed, or whether the image is divided can be confirmed by other signals. For example, first_slice_segment_in_pic_flag is not a signal indicating whether the image is divided, but a signal indicating whether it is the first slice segment in the image, and this can be used to confirm whether the image is divided into two or more units (for example, if the flag is 0, it means that the image is divided into multiple slices).
[0121] The present invention is not limited to the above example, and other modifications are possible. For example, a signal instructing image division in tiles may not be provided, and a signal instructing image division in slices may be provided. Alternatively, a signal instructing image division may be provided depending on the type, characteristics, etc. of an image.
[0122] In the image segmentation type identification step, the image segmentation type can be identified. The image segmentation type can be defined by the segmentation method, segmentation information, etc.
[0123] In 4b, a tile can be defined as a unit obtained by dividing the image horizontally and vertically, and more specifically, as a group of adjacent blocks within a rectangular space defined by at least one horizontal or vertical dividing line that crosses the image.
[0124] The division information for tiles may include information on the boundary positions between rows and columns, information on the number of tiles in each row and column, and tile size information. The tile number information may include the number of tile rows (e.g., num_tile_columns) and the number of tile columns (e.g., num_tile_rows), thereby enabling division into tiles with a number equal to (number of rows x number of columns). The tile size information may be obtained based on the tile number information, and the horizontal and vertical widths of the tiles may be uniform or non-uniform. This may be implicitly determined based on a preset rule, or related information (e.g., uniform_spacing_flag) may be explicitly generated. The tile size information may also include size information for each row and column of tiles (e.g., column_width_tile[i], row_height_tile[i]) or may include information on the vertical and horizontal widths of each tile. Furthermore, the size information may be information that can be further generated depending on whether the tile size is uniform or not (for example, when uniform_spacing_flag is 0, meaning non-uniform division).
[0125] In 4c, a slice can be defined as a group of consecutive blocks, specifically, a group of consecutive blocks based on a predetermined scan order (raster scan in this example).
[0126] The division information for a slice may include information on the number of slices, slice position information (e.g., slice_segment_address), etc. In this case, the slice position information may be position information of a predetermined block (e.g., the first block in the scan order within the slice). In this case, the position information may be scan order information of the block.
[0127] In 4d, various division settings are possible for any division area.
[0128] A division unit in 4d can be defined as a group of spatially adjacent blocks, and the division information for this can include the size, shape, position information, etc. This is only a partial example of an arbitrary division area, and various division shapes are possible, as shown in Figure 5.
[0129] FIG. 5 is a diagram illustrating another example of an image division method according to an embodiment of the present invention.
[0130] In the cases of 5a and 5b, an image can be divided into a plurality of regions at intervals of at least one block in the horizontal or vertical direction, and the division can be performed based on block position information. 5a shows examples A0 and A1 in which division is performed horizontally based on column information of each block, and 5b shows examples B0 to B3 in which division is performed vertically and horizontally based on row and column information of each block. The division information can include the number of division units, block interval information, division direction, etc., and if these are implicitly included according to a predetermined rule, some division information may not be generated.
[0131] In cases 5c and 5d, an image can be divided into groups of consecutive blocks based on the scan order. Additional scan orders other than the existing slice raster scan order can be applied to image division. 5c shows examples C0 and C1 in which a clockwise or counterclockwise scan (box-out) is performed around a starting block, and 5d shows examples D0 and D1 in which a vertical scan (vertical) is performed around a starting block. The corresponding division information can include information on the number of division units, position information of the division units (e.g., the first order in the scan order within the division unit), information on the scan order, etc. If this information is implicitly included according to a predetermined rule, some division information may not be generated.
[0132] In the case of 5e, an image can be divided along horizontal and vertical division lines. Existing tiles can be divided along horizontal or vertical division lines, thereby creating a rectangular spatial division shape, but division along division lines may not be possible. For example, a division along a partial division line of the image that crosses the image (e.g., a division line forming the right boundary of E1, E3, and E4 and the left boundary of E5) is possible, but a division along a partial division line of the image that crosses the image (e.g., a division line forming the bottom boundary of E2 and E3 and the top boundary of E4) is not possible. In addition, division can be performed based on block units (e.g., block division is performed first, followed by division), or division can be performed along the horizontal or vertical division line (e.g., division along the division line regardless of block division), and therefore each division unit may not be composed of an integer multiple of a block. Therefore, division information different from existing tiles can be generated, and the corresponding division information can include information on the number of division units, position information of the division units, size information of the division units, etc. For example, position information of a division unit can be generated based on a predetermined position (e.g., the top left corner of the image) as position information (e.g., measured in pixel units or block units), and size information of a division unit can be generated based on the horizontal and vertical size information of each division unit (e.g., measured in pixel units or block units).
[0133] As in the above example, segmentation with user-defined settings may be performed by applying a new segmentation method or by modifying and applying a portion of an existing segmentation configuration. That is, an existing segmentation method may be replaced or supported as an additional segmentation shape, or an existing segmentation method (e.g., slice, tile, etc.) may be supported with some modified settings applied (e.g., according to a different scan order, or a different rectangular segmentation method and the generation of other segmentation information accordingly, dependent encoding / decoding characteristics, etc.). Additionally, settings for configuring additional segmentation units (e.g., settings other than segmentation according to the scan order or segmentation according to a fixed interval) may be supported, and additional segmentation unit shapes (e.g., polygonal shapes such as triangles other than segmentation into rectangular spaces) may be supported. Furthermore, an image segmentation method may be supported based on the type and characteristics of the image. For example, some segmentation methods (e.g., the surface of a 360-degree image) may be supported depending on the type and characteristics of the image, and segmentation information may be generated based on the supported method.
[0134] In the image segmentation step, the image can be segmented based on the identified segmentation type information, i.e., segmentation can be performed into a plurality of segmentation units based on the identified segmentation type, and encoding / decoding can be performed based on the acquired segmentation units.
[0135] In this case, it can be determined whether the partition unit has encoding / decoding configuration depending on the partition type. That is, configuration information required for the encoding / decoding process of each partition unit can be assigned to a higher unit (e.g., a picture), or each partition unit can have its own independent encoding / decoding configuration.
[0136] Generally, a slice may have an independent encoding / decoding setting for each division unit (e.g., a slice header), whereas a tile may not have an independent encoding / decoding setting for each division unit, but may have a setting that is dependent on the encoding / decoding setting for the picture (e.g., PPS). In this case, information generated in relation to the tile may be partition information, which may be included in the encoding / decoding setting for the picture. The present invention is not limited to the above-described example, and other variations are possible.
[0137] Encoding / decoding setting information for a tile can be generated in units such as video, sequence, or picture, or at least one encoding / decoding setting information can be generated in a higher-level unit and any one of them can be referenced. Alternatively, independent encoding / decoding setting information (e.g., a tile header) can be generated in units of tiles. This differs from a single encoding / decoding setting determined in a higher-level unit in that encoding / decoding is performed by setting at least one encoding / decoding setting in units of tiles. That is, all tiles can be encoded / decoded in accordance with a single encoding / decoding setting, or at least one tile can be encoded / decoded in accordance with an encoding / decoding setting that is different from that of another tile.
[0138] Although the above example has been used to mainly describe various encoding / decoding settings in tiles, the present invention is not limited to this, and similar or identical settings can be set for other division types.
[0139] For example, for some division types, division information may be generated for each higher unit, and encoding / decoding may be performed according to one encoding / decoding setting for the higher unit.
[0140] For example, for some division types, division information is generated in higher units, and independent encoding / decoding settings for each division unit are generated in the higher units, and encoding / decoding can be performed accordingly.
[0141] As an example, for some division types, division information can be generated in higher units, multiple encoding / decoding setting information can be supported in higher units, and encoding / decoding can be performed according to the encoding / decoding settings referenced in each division unit.
[0142] For example, for some division types, division information may be generated in higher units, and independent encoding / decoding settings may be generated in the corresponding division units, thereby performing encoding / decoding.
[0143] For example, for some division types, an independent encoding / decoding setting including division information may be generated for each division unit, and encoding / decoding may be performed accordingly.
[0144] The encoding / decoding setting information may include information necessary for encoding / decoding of tiles, such as a tile type, information on a reference picture list, quantization parameter information, inter-picture prediction setting information, in-loop filtering setting information, in-loop filtering control information, a scan order, whether encoding / decoding is performed, etc. As for the encoding / decoding setting information, related information may be explicitly generated, or settings for encoding / decoding may be implicitly determined according to an image format, characteristics, etc. determined in a higher unit. Also, related information may be explicitly generated based on information obtained in the setting.
[0145] An example of image division performed by an encoding / decoding device according to one embodiment of the present invention will be described below.
[0146] Before encoding begins, an input image can be segmented. After segmentation is performed using segmentation information (e.g., image segmentation information, segmentation unit setting information, etc.), the image can be encoded in segments. After encoding is complete, the image can be stored in memory, or the encoded image data can be recorded in a bitstream and transmitted.
[0147] A division process can be performed before the start of decoding. After division is performed using division information (e.g., image division information, division unit setting information, etc.), the image decoded data can be parsed and decoded in division units. After decoding is complete, the data can be stored in memory, and multiple division units can be merged into one to output the image.
[0148] The image segmentation process has been described using the above example, and multiple segmentation processes can be performed in the present invention.
[0149] For example, an image may be segmented, or segmentation may be performed on a division unit of the image. The segmentation may be the same division process (e.g., slice / slice, tile / tile, etc.) or different division processes (e.g., slice / tile, tile / slice, tile / surface, surface / tile, slice / surface, surface / slice, etc.). In this case, a subsequent segmentation process may be performed based on a previous segmentation result. Division information generated in a subsequent segmentation process may be generated based on the previous segmentation result.
[0150] In addition, multiple division processes A can be performed, and the division processes can be different division processes (e.g., slice / surface, tile / surface, etc.). In this case, the subsequent division process can be performed based on the previous division result, or can be performed independently regardless of the previous division result. Division information generated in the subsequent division process can be generated based on the previous division result or can be generated independently.
[0151] The multiple division processes of the image can be determined according to the encoding / decoding settings, and are not limited to the above example, and various modified examples are also possible.
[0152] The encoder records information generated in the above process in at least one unit, such as a sequence, a picture, a slice, or a tile, into a bitstream, and the decoder parses the related information from the bitstream. That is, the information can be recorded in one unit or can be recorded redundantly in multiple units. For example, a syntax element indicating whether or not some information is supported or activated can be generated in some units (e.g., upper units), and the same or similar information can be generated in some units (e.g., lower units). That is, even if related information is supported and configured in an upper unit, it can have individual settings in the lower units. This is not limited to the above example and can be a description commonly applicable to the present invention. The information can also be included in the bitstream in the form of SEI or metadata.
[0153] Meanwhile, although it may be common to encode / decode an input image as is, it may also be possible to adjust the size of the image (enlargement or reduction, resolution adjustment) before encoding / decoding. For example, image size adjustment, such as overall image enlargement or reduction, may be performed in a hierarchical coding method (scalability video coding) that supports spatial, temporal, and image quality scalability. Alternatively, image size adjustment, such as partial image enlargement or reduction, may also be performed. Image size adjustment can be performed for various purposes, including adaptability to the coding environment, coding uniformity, coding efficiency, image quality improvement, or depending on the type and characteristics of the image.
[0154] As a first example, a size adjustment process may be performed in a process (eg, hierarchical encoding, 360-degree image encoding, etc.) that is performed according to the characteristics, type, etc. of an image.
[0155] As a second example, the resizing process may be performed in the early stage of encoding / decoding, or before encoding / decoding, or the image to be resized may be encoded / decoded.
[0156] As a third example, a size adjustment process may be performed during a prediction step (intra prediction or inter prediction) or before prediction is performed. The size adjustment process may use image information during the prediction step (e.g., pixel information referenced in intra prediction, information related to an intra prediction mode, reference image information used in inter prediction, information related to an inter prediction mode, etc.).
[0157] As a fourth example, a resizing process may be performed in the filtering step or before filtering. In the resizing process, image information in the filtering step (e.g., pixel information applied to a deblocking filter, pixel information applied to an SAO, information related to SAO filtering, pixel information applied to an ALF, information related to ALF filtering, etc.) may be used.
[0158] Furthermore, after the resizing process, the image may or may not be changed to the image before resizing (in terms of image size) through an inverse resizing process. This can be determined depending on the encoding / decoding settings (e.g., the nature of the resizing process). In this case, if the resizing process is expansion, the inverse resizing process may be reduction, and if the resizing process is reduction, the inverse resizing process may be expansion.
[0159] When the size adjustment process according to the first to fourth examples is performed, a reverse size adjustment process can be performed in a subsequent step to obtain an image before size adjustment.
[0160] When the resizing process according to the hierarchical coding or the third example is performed (or when the size of the reference image is adjusted in inter prediction), the reverse resizing process may not be performed in the subsequent steps.
[0161] In one embodiment of the present invention, the image resizing process may be performed alone or inversely, and the following examples will focus on the resizing process. Since the inverse resizing process is the reverse process of the resizing process, a description of the inverse resizing process will be omitted to avoid redundancy, but it is clear that a person skilled in the art would understand the process in the same way as if it were literally described.
[0162] FIG. 6 is an example of a general image resizing method.
[0163] Referring to 6a, an expanded image P0+P1 can be obtained by further including a partial region P1 from the initial image (or image before size adjustment; P0; thick solid line).
[0164] 6b, a reduced image S0 can be obtained by excluding a partial region S1 from the initial image S0+S1.
[0165] Referring to 6c, the resized image T0+T1 can be obtained by further including the partial region T1 in the initial image T0+T2 and excluding the partial region T2.
[0166] Hereinafter, the present invention will be described focusing on the size adjustment process by expansion and the size adjustment process by reduction, but it should be understood that the present invention is not limited to this and also includes cases where size expansion and reduction are mixed and applied, as in 6c.
[0167] FIG. 7 is an example diagram of image size adjustment according to an embodiment of the present invention.
[0168] Referring to 7a, it can be seen how to expand an image during the resizing process, and referring to 7b, it can be seen how to shrink an image.
[0169] In 7a, the image before resizing is S0 and the image after resizing is S1, and in 7b, the image before resizing is T0 and the image after resizing is T1.
[0170] When expanding an image as in 7a, it can be expanded in the up, down, left, and right directions (ET, EL, EB, ER), and when shrinking an image as in 7b, it can be shrinked in the up, down, left, and right directions (RT, RL, RB, RR).
[0171] When comparing image expansion and image reduction, the up, down, left, and right directions in expansion correspond to the down, up, right, and left directions in reduction, respectively. Therefore, the following explanation will be based on image expansion, but it should be understood that the explanation also includes image reduction.
[0172] Also, although the following describes expanding or shrinking an image in the up, down, left, and right directions, it should be understood that size adjustments can also be made in the top-left, top-right, bottom-left, and bottom-right directions.
[0173] In this case, when extending in the lower right direction, the RC and BC regions are acquired, and depending on the encoding / decoding settings, the BR region may or may not be acquired. That is, the TL, TR, BL, and BR regions may or may not be acquired, but for the sake of convenience, the following description will be given assuming that the corner regions (TL, TR, BL, and BR regions) can be acquired.
[0174] The image resizing process according to an embodiment of the present invention may be performed in at least one direction, for example, in all of the up, down, left, and right directions, in two or more directions selected from the up, down, left, and right directions (e.g., left+right, up+down, up+left, up+right, down+left, down+right, up+left+right, down+left+right, up+down+left, up+down+right), or in only one of the up, down, left, and right directions.
[0175] For example, the size may be adjusted in the left+right, top+bottom, top left+bottom right, or bottom left+top right directions, so that the image can be expanded symmetrically on both sides based on the center of the image; the size may be adjusted in the left+right, top left+top right, or bottom left+bottom right directions, so that the image can be expanded symmetrically vertically; the size may be adjusted in the top+bottom, top left+bottom left, or top right+bottom right directions, so that the image can be expanded symmetrically horizontally; and other size adjustments are also possible.
[0176] In 7a and 7b, the size of the image (S0, T0) before size adjustment is defined as P_Width (width) × P_Height (height), and the size of the image (S1, T1) after size adjustment is defined as P'_Width (width) × P'_Height (height). Here, if the size adjustment values for the left, right, top, and bottom directions are defined as Var_L, Var_R, Var_T, and Var_B (or collectively referred to as Var_x), the image size after size adjustment can be expressed as (P_Width + Var_L + Var_R) × (P_Height + Var_T + Var_B). In this case, the size adjustment values Var_L, Var_R, Var_T, and Var_B in the left, right, top, and bottom directions can be Exp_L, Exp_R, Exp_T, and Exp_B (in this example, Exp_x is a positive number) in image expansion (FIG. 7a), and can be −Rec_L, −Rec_R, −Rec_T, and −Rec_B in image contraction (when Rec_L, Rec_R, Rec_T, and Rec_B are defined as positive numbers, they can be expressed as negative numbers according to the contraction of the image). Furthermore, the coordinates of the top left, top right, bottom left, and bottom right of the image before resizing are (0, 0), (P_Width-1, 0), (0, P_Height-1), (P_Width-1, P_Height-1), and the coordinates of the image after resizing can be expressed as (0, 0), (P'_Width-1, 0), (0, P'_Height-1), (P'_Width-1, P'_Height-1). The size of the area (in this example, TL to BR, where i is an index that distinguishes TL to BR) that is changed (or acquired or deleted) during resizing can be M[i] x N[i]. This can be expressed as Var_X x Var_Y (in this example, X is assumed to be L or R, and Y is assumed to be T or B). M and N can have various values and can be the same regardless of i, or can have individual settings depending on i. Various cases regarding this will be described later.
[0177] Referring to 7a, S1 can be configured to include all or part of TL-BR (upper left to lower right) generated by expanding S0 in various directions. Referring to 7b, T1 can be configured to exclude all or part of TL-BR removed by contracting T0 in various directions.
[0178] In 7a, when an existing image S0 is expanded in the up, down, left, and right directions, the image can be constructed including the TC, BC, LC, and RC regions obtained through each size adjustment process, and can also include the TL, TR, BL, and BR regions.
[0179] As an example, when extension is performed in the upward (ET) direction, an image can be constructed by including a TC region in an existing image S0, and can include a TL or TR region depending on extension in at least one different direction (EL or ER).
[0180] As an example, when expanding in the downward (EB) direction, an image can be constructed by including a BC region in an existing image S0, and can include a BL or BR region depending on the expansion in at least one different direction (EL or ER).
[0181] As an example, when expanding in the left (EL) direction, an image can be constructed by including an LC region in the existing image S0, and can include a TL or BL region depending on the expansion in at least one different direction (ET or EB).
[0182] As an example, when extending in the right (ER) direction, an image can be constructed by including an RC region in an existing image S0, and can include a TR or BR region depending on at least one different direction of extension (ET or EB).
[0183] According to one embodiment of the present invention, a setting (e.g., spa_ref_enabled_flag or tem_ref_enabled_flag) can be placed that can spatially or temporally restrict the visibility of the region being resized (assumed to be extended in this example).
[0184] That is, you can either reference data in areas that are spatially or temporally sized depending on the encoding / decoding settings (e.g., spa_ref_enabled_flag=1 or tem_ref_enabled_flag=1), or restrict references (e.g., spa_ref_enabled_flag=0 or tem_ref_enabled_flag=0).
[0185] The encoding / decoding of the image (S0, T1) before resizing and the areas added or deleted during resizing (TC, BC, LC, RC, TL, TR, BL, BR areas) can be performed as follows.
[0186] For example, when encoding / decoding an image before resizing and an area to be added or deleted, the data of the image before resizing and the data of the area to be added or deleted (encoded / decoded data, such as pixel values or prediction-related information) can be referenced to each other spatially or temporally.
[0187] Alternatively, the image before resizing and the data of the area to be added or deleted can be referenced spatially, while the data of the image before resizing can be referenced temporally, and the data of the area to be added or deleted cannot be referenced temporally.
[0188] That is, a setting can be set that limits the visibility of the region being added or deleted. The setting information about the visibility of the region being added or deleted can be explicitly created or implicitly determined.
[0189] An image resizing process according to an embodiment of the present invention may include an image resizing instruction step, an image resizing type identification step, and / or an image resizing execution step. Furthermore, an image encoding device and a decoding device may include an image resizing instruction unit, an image resizing type identification unit, and an image resizing execution unit that implement the image resizing instruction step, the image resizing type identification step, and the image resizing execution step. In the case of encoding, associated syntax elements may be generated, and in the case of decoding, associated syntax elements may be parsed.
[0190] In the image resizing instruction step, it can be determined whether to perform image resizing. For example, if a signal instructing image resizing (e.g., img_resizing_enabled_flag) is checked, resizing can be performed. If the signal instructing image resizing is not checked, resizing may not be performed or resizing may be performed by checking other encoding / decoding information. Also, even if a signal instructing image resizing is not provided, the signal instructing resizing may be implicitly activated or deactivated depending on encoding / decoding settings (e.g., image characteristics, type, etc.). If resizing is performed, resizing-related information may be generated accordingly, or the resizing-related information may be implicitly determined.
[0191] When a signal instructing image resizing is provided, the signal is a signal indicating whether or not to resize the image, and it is possible to check whether or not the size of the corresponding image is to be resized according to the signal.
[0192] For example, if a signal instructing image resizing (e.g., img_resizing_enabled_flag) is checked and the corresponding signal is activated (e.g., img_resizing_enabled_flag=1), image resizing can be performed, and if the corresponding signal is deactivated (e.g., img_resizing_enabled_flag=0), image resizing will not be performed.
[0193] Also, if a signal instructing image size adjustment is not provided, size adjustment may not be performed, or it may be possible to check whether the image size adjustment is required or not based on another signal.
[0194] For example, when an input image is divided into blocks, size adjustment (in this example, in the case of expansion; it is assumed that the size adjustment process is performed when the image size is not an integer multiple of the block size) can be performed depending on whether the image size (e.g., width or height) is an integer multiple of the block size (e.g., width or height). That is, size adjustment can be performed when the image width is not an integer multiple of the block width or when the image height is not an integer multiple of the block height. In this case, size adjustment information (e.g., size adjustment direction, size adjustment value, etc.) can be determined according to the encoding / decoding information (e.g., image size, block size, etc.). Alternatively, size adjustment can be performed according to the characteristics or type of image (e.g., 360-degree image), and the size adjustment information can be explicitly generated or assigned a predetermined value. The above example is not limited to the above example, and other modifications are also possible.
[0195] In the image size adjustment type identification step, an image size adjustment type can be identified. The image size adjustment type can be defined by a size adjustment method, size adjustment information, etc. For example, size adjustment using a scale factor, size adjustment using an offset factor, etc. can be performed. However, the present invention is not limited to this, and a combination of the above methods can also be applied. For convenience of explanation, size adjustment using a scale factor and an offset factor will be mainly described.
[0196] In the case of a scale factor, size adjustment can be performed by multiplying or dividing the scale factor based on the image size. Information about the size adjustment operation (e.g., expansion or contraction) can be explicitly generated, and the expansion or contraction process can be performed according to the information. Also, the size adjustment process can be performed using a predetermined operation (e.g., either expansion or contraction) depending on the encoding / decoding settings, in which case the information about the size adjustment operation can be omitted. For example, if image size adjustment is activated in the image size adjustment instruction step, the image size can be adjusted using a predetermined operation.
[0197] The size adjustment direction may be at least one direction selected from the top, bottom, left, and right directions. At least one scale factor may be required depending on the size adjustment direction. That is, one scale factor (in this example, unidirectional) may be required for each direction, one scale factor (in this example, bidirectional) may be required for the horizontal or vertical direction, and one scale factor (in this example, omnidirectional) may be required depending on the overall direction of the image. Furthermore, the size adjustment direction is not limited to the above example, and other variations are possible.
[0198] The scale factor can have a positive value and can be set to different range information depending on the encoding / decoding settings. For example, when generating information by combining a resizing operation and a scale factor, the scale factor can be used as a multiplier. A value greater than 0 or less than 1 can indicate a shrinking operation, a value greater than 1 can indicate an expanding operation, and a value of 1 can indicate no resizing. As another example, when generating scale factor information separately from a resizing operation, the scale factor can be used as a multiplier for an expanding operation and as a divisor for a shrinking operation.
[0199] Referring again to 7a and 7b of FIG. 7, the process of using a scale factor to change from an unscaled image (S0, T0) to a scaled image (S1, T1 in this example) can be explained.
[0200] For example, if one scale factor (called sc) is used according to the overall orientation of the image and the resizing direction is down+right, the resizing direction is ER, EB (or RR, RB), the resizing values Var_L (Exp_L or Rec_L) and Var_T (Exp_T or Rec_T) are 0, and Var_R (Exp_R or Rec_R) and Var_B (Exp_B or Rec_B) can be expressed as P_Width×(sc-1) and P_Height×(sc-1). Therefore, the resized image can be (P_Width×sc)×(P_Height×sc).
[0201] As an example, if a scale factor (in this example, sc_w, sc_h) is used depending on the horizontal or vertical direction of the image, and the resizing direction is left+right or up+down (when two are operated, it becomes up+down+left+right), the resizing direction is ET, EB, EL, ER, and the resizing values Var_T and Var_B can be P_Height×(sc_h−1) / 2, and Var_L and Var_R can be P_Width×(sc_w−1) / 2. Therefore, the resizing image can be (P_Width×sc_w)×(P_Height×sc_h).
[0202] In the case of an offset factor, size adjustment can be performed by adding or subtracting based on the size of the image, or by adding or subtracting based on the encoding / decoding information of the image, or by independently adding or subtracting. That is, the size adjustment process can be set dependently or independently.
[0203] Information about the resizing operation (e.g., expansion or contraction) can be explicitly generated, and the expansion or contraction process can be performed according to the information. Also, the resizing operation can be performed by a predetermined operation (e.g., either expansion or contraction) according to the encoding / decoding settings, in which case the information about the resizing operation can be omitted. For example, if image resizing is activated in the image resizing instruction step, the image resizing can be performed by a predetermined operation.
[0204] The size adjustment direction may be at least one of the up, down, left, and right directions. At least one offset factor may be required depending on the size adjustment direction. That is, one offset factor (in this example, unidirectional) may be required for each direction, either a horizontal or vertical offset factor (in this example, symmetrical bidirectional) may be required depending on the horizontal or vertical direction, one offset factor (in this example, asymmetrical bidirectional) may be required depending on a partial combination of each direction, or one offset factor (in this example, omnidirectional) may be required depending on the overall direction of the image. Furthermore, the size adjustment direction is not limited to the above example, and other variations are possible.
[0205] The offset factor can have a positive value or a combination of positive and negative values, and can be set with different range information depending on the encoding / decoding settings. For example, when generating information by mixing a resizing operation and an offset factor (assumed to have positive and negative values in this example), the offset factor can be used as a value to be added or subtracted depending on the sign information of the offset factor. If the offset factor is greater than 0, it can indicate an expansion operation, if it is less than 0, it can indicate a reduction operation, and if it is 0, it can indicate no resizing. As another example, when generating offset factor information separately from a resizing operation (assumed to have a positive value in this example), the offset factor can be used as a value to be added or subtracted depending on the resizing operation. If it is greater than 0, it can indicate an expansion or reduction operation depending on the resizing operation, and if it is 0, it can indicate no resizing.
[0206] Referring again to FIGS. 7a and 7b, it can be seen how an offset factor is used to change from an unscaled image (S0, T0) to a scaled image (S1, T1).
[0207] For example, if one offset factor (called os) is used according to the overall orientation of the image and the resizing direction is up+down+left+right, the resizing direction is ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values Var_T, Var_B, Var_L, and Var_R can be os. The resized image size can be (P_Width+os) x (P_Height+os).
[0208] For example, if the offset factors (os_w, os_h) are used depending on the horizontal or vertical direction of the image, and the resizing direction is left+right or up+down (up+down+left+right when both are combined), the resizing direction is ET, EB, EL, ER (or RT, RB, RL, RR), the resizing values Var_T and Var_B can be os_h, and Var_L and Var_R can be os_w. The resized image size can be {P_Width+(os_w×2)}×{P_Height+(os_h×2)}.
[0209] For example, if the resizing direction is down and right (down + right when working together) and offset factors (os_b, os_r) are used according to the resizing direction, the resizing direction is EB, ER (or RB, RR), the resizing value Var_B can be os_b, and Var_R can be os_r. The resizing image size can be (P_Width + os_r) x (P_Height + os_b).
[0210] For example, if offset factors (os_t, os_b, os_l, os_r) are used for each direction of the image, and the resizing directions are up, down, left, and right (up + down + left + right when all are active), the resizing directions are ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values Var_T can be os_t, Var_B can be os_b, Var_L can be os_l, and Var_R can be os_r. The resized image size can be (P_Width + os_l + os_r) x (P_Height + os_t + os_b).
[0211] The above example shows a case where offset factors are used as size adjustment values (Var_T, Var_B, Var_L, Var_R) in the size adjustment process. That is, the offset factors are used as size adjustment values as they are. This may be an example of an independent size adjustment. Alternatively, the offset factors may be used as input variables for the size adjustment values. In particular, the offset factors may be assigned as input variables, and size adjustment values may be obtained through a series of processes according to encoding / decoding settings. This may be an example of a size adjustment performed based on predetermined information (e.g., image size, encoding / decoding information, etc.), or an example of a dependent size adjustment.
[0212] For example, the offset factor may be a multiple (e.g., 1, 2, 4, 6, 8, 16, etc.) or exponent (e.g., an exponent of two, such as 1, 2, 4, 8, 16, 32, 64, 128, 256, etc.) of a predetermined value (an integer in this example). Alternatively, the offset factor may be a multiple or exponent of a value obtained based on encoding / decoding settings (e.g., a value set based on the motion search range of inter prediction). Alternatively, the offset factor may be a multiple or exponent of a unit (assumed to be A×B in this example) obtained from the picture divider. Alternatively, the offset factor may be a multiple of a unit (assumed to be E×F in the case of a tile, etc.) obtained from the picture divider.
[0213] Alternatively, the value may be smaller than or in the same range as the width and height of the unit obtained from the picture division unit. The multiples or exponents in the above examples may also include a case where the value is 1, and are not limited to the above examples, and modifications to other examples are also possible. For example, if the offset factor is n, Var_x is 2×n or 2 n It could be.
[0214] In addition, individual offset factors can be supported for each color component, and offset factors for some color components can be supported to derive offset factor information for other color components. For example, if an offset factor A for a luminance component (assuming the luminance component and chrominance component have a composition ratio of 2:1 in this example) is explicitly generated, an offset factor A / 2 for a chrominance component can be implicitly obtained. Alternatively, if an offset factor A for a chrominance component is explicitly generated, an offset factor 2A for the luminance component can be implicitly obtained.
[0215] Information regarding the resizing direction and resizing value can be explicitly generated, and the resizing process can be performed according to the corresponding information. Alternatively, the resizing process can be performed implicitly according to the encoding / decoding settings. At least one predetermined direction or resizing value can be assigned, in which case related information can be omitted. In this case, the encoding / decoding settings can be determined based on the characteristics, type, encoding information, etc. of the image. For example, at least one resizing direction according to at least one resizing operation can be predetermined, at least one resizing value according to at least one resizing operation can be predetermined, and at least one resizing value according to at least one resizing direction can be predetermined. In addition, the resizing direction and resizing value in the reverse resizing process can be derived from the resizing direction and resizing value applied in the resizing process. In this case, the implicitly determined resizing value can be one of the examples described above (examples in which various resizing values are obtained).
[0216] Also, although the above examples have been described with respect to cases of multiplication or division, they can be realized by shift operations depending on the implementation of the encoder / decoder. Multiplication can be realized by using a left shift operation, and division can be realized by using a right shift operation. This is not limited to the above examples, but can be a description commonly applied to the present invention.
[0217] In the image resizing execution step, image resizing can be performed based on the identified resizing information, i.e., based on information such as resizing type, resizing operation, resizing direction, and resizing value, and encoding / decoding can be performed based on the obtained resizing image.
[0218] In addition, in the image resizing step, the resizing may be performed using at least one data processing method. Specifically, the resizing may be performed using at least one data processing method for the area to be resized depending on the resizing type and the resizing operation. For example, depending on the resizing type, it may be possible to determine how to fill data when the resizing is expansion, or how to remove data when the resizing process is reduction.
[0219] In summary, image size adjustment can be performed based on size adjustment information identified in the image size adjustment execution step. Alternatively, image size adjustment can be performed based on the size adjustment information and a data processing method in the image size adjustment execution step. The difference between the two cases is whether only the size of the image to be encoded / decoded is adjusted, or whether the size of the image and data processing of the area to be resized are also taken into consideration. Whether a data processing method is included in the image size adjustment execution step can be determined depending on the application step, position, etc. of the size adjustment process. In the following example, size adjustment based on a data processing method will be mainly described, but is not limited to this.
[0220] When performing size adjustment using an offset factor, various methods can be used for expansion and reduction. In the case of expansion, size adjustment can be performed using a method of filling at least one piece of data, and in the case of reduction, size adjustment can be performed using a method of removing at least one piece of data. In this case, when performing size adjustment using an offset factor, new data or existing image data can be filled directly or by modifying it into the size adjustment area (expansion), and data can be removed from the size adjustment area (reduction) by applying simple removal or removal through a series of processes.
[0221] When performing resizing using a scale factor, in some cases (e.g., hierarchical coding), expansion may be performed by applying upsampling, and contraction may be performed by applying downsampling. For example, expansion may use at least one upsampling filter, and contraction may use at least one downsampling filter. The horizontal and vertical filters may be the same or different. In this case, resizing using a scale factor may involve relocating existing image data using a method such as interpolation, rather than generating or removing new data in the resizing region. Data processing methods related to resizing can be categorized by the filter used for sampling. In addition, in some cases (e.g., similar to an offset factor), expansion may be performed by filling in at least one piece of data, and contraction may be performed by removing at least one piece of data. This invention will mainly describe a data processing method when performing resizing using an offset factor.
[0222] Generally, one predetermined data processing method can be used for a region to be resized, but as in the example described below, at least one data processing method can be used for a region to be resized, and selection information for the data processing method can be generated. In the former case, it can mean performing size resizing through a fixed data processing method, and in the latter case, it can mean performing size resizing through an adaptive data processing method.
[0223] In addition, a data processing method common to the entire area to be added or deleted during size adjustment (TL, TC, TR, ..., BR in Figures 7a and 7b) can be applied, or a data processing method can be applied to a portion of the area to be added or deleted during size adjustment (for example, each of TL to BR in Figures 7a and 7b or a partial combination thereof).
[0224] FIG. 8 is an exemplary diagram illustrating a method for configuring an area to be expanded in an image resizing method according to an embodiment of the present invention.
[0225] 8a, for convenience of explanation, an image can be divided into TL, TC, TR, LC, C, RC, BL, BC, and BR regions, which can correspond to the top left, top, top right, left, center, right, bottom left, bottom, and bottom right positions of the image, respectively. Hereinafter, a case where an image is expanded in the bottom+right direction will be described, but it should be understood that the same can be applied to other directions.
[0226] The area added in response to the expansion of the image can be configured in various ways, for example, it can be filled with an arbitrary value or can be filled by referencing partial data of the image.
[0227] 8b, the area (A0, A2) can be filled to any pixel value. The pixel value can be determined using various methods.
[0228] For example, a given pixel value may be a pixel that belongs to a range of pixel values that can be expressed by a bit depth {e.g., from 0 to 1<<(bit_depth)-1}, such as the minimum value, maximum value, or median value {e.g., 1<<(bit_depth-1)} of the range of pixel values (where bit_depth is the bit depth).
[0229] As an example, any pixel value can be set within a range of pixel values {e.g., min P From max P Until . min P , max P are the minimum and maximum values of the pixels in the image. P is greater than or equal to 0, max P may be a pixel belonging to the range {1<<(bit_depth)-1 or less}. For example, any pixel value may be the minimum, maximum, median, average (of at least two pixels), weighted sum, etc. of the range of pixel values.
[0230] For example, the given pixel value may be a value determined within a range of pixel values belonging to a partial region of an image. For example, when constructing A0, the partial region may be TR+RC+BR. The partial region may be a 3x9 region of TR, RC, and BR, or a 1x9 region (assuming the rightmost line). This may depend on the encoding / decoding settings. In this case, the partial region may be a unit divided by a picture divider. Specifically, the given pixel value may be the minimum value, maximum value, median value, average (of at least two pixels), weighted sum, etc., of the pixel value range.
[0231] Referring again to 8b, the area A1 added in response to the expansion of the image can be filled with pattern information generated using multiple pixel values (for example, a pattern is assumed to be one using multiple pixels, and does not necessarily have to follow a specific rule). In this case, the pattern information can be defined according to the encoding / decoding settings or related information can be generated, and at least one pattern information can be used to fill the expanded area.
[0232] Referring to 8c, a region to be added in response to image expansion may be configured by referring to pixels of a partial region belonging to the image. Specifically, the region to be added may be configured by copying or padding pixels (hereinafter, referred to as reference pixels) of a region adjacent to the region to be added. In this case, the pixels of the region adjacent to the region to be added may be pixels before encoding or pixels after encoding (or decoding). For example, when performing size adjustment in a pre-encoding step, the reference pixels may refer to pixels of an input image, and when performing size adjustment in an intra-frame prediction reference pixel generation step, a reference image generation step, a filtering step, or the like, the reference pixels may refer to pixels of a restored image. In this example, it is assumed that pixels closest to the region to be added are used, but this is not limiting.
[0233] The area A0 expanded to the left or right related to the horizontal size adjustment of the image can be configured by padding (Z0) the outer pixels adjacent to the expanded area A0 in the horizontal direction, the area A1 expanded to the top or bottom related to the vertical size adjustment of the image can be configured by padding (Z1) the outer pixels adjacent to the expanded area A1 in the vertical direction, and the area A2 expanded to the bottom right can be configured by padding (Z2) the outer pixels adjacent to the expanded area A2 in the diagonal direction.
[0234] Referring to 8d, the expanded region B'0-B'2 can be configured by referencing data B0-B2 of a partial region belonging to the image. 8d can be distinguished from 8c in that it can reference regions that are not adjacent to the expanded region.
[0235] For example, if an image contains an area that is highly correlated with the area to be expanded, the area to be expanded can be filled by referencing the pixels of the highly correlated area. At this time, position information, area size information, etc. of the highly correlated area can be generated. Alternatively, if a highly correlated area exists through encoding / decoding information such as image characteristics and type, and the position information, size information, etc. of the highly correlated area can be implicitly confirmed (for example, in the case of a 360-degree image), the data of the relevant area can be filled into the area to be expanded. At this time, the position information, area size information, etc. of the relevant area can be omitted.
[0236] As an example, in the case of an area B'2 that is expanded in the left or right direction related to the horizontal size adjustment of an image, the expanded area can be filled by referring to pixels in the area B2 opposite the expanded area in the left or right direction related to the horizontal size adjustment.
[0237] As an example, in the case of an area B'1 that is expanded in the upward or downward direction related to the vertical size adjustment of an image, the expanded area can be filled by referring to pixels in the area B1 opposite the expanded area in the upward or right direction related to the vertical size adjustment.
[0238] As an example, in the case of an area B'0 that is expanded by adjusting the size of a portion of an image (in this example, diagonally from the center of the image), the expanded area can be filled by referring to the pixels of the area B0, TL opposite the expanded area.
[0239] In the above example, we have explained the case where data is obtained from an area where there is continuity at the boundaries on both ends of the image and where the area is located symmetrically to the size adjustment direction, but this is not limited to this and it is also possible to obtain data from other areas TL to BR.
[0240] When filling an expanded area with data from a portion of an image, the data can be copied as is or can be filled after undergoing a transformation process based on the image characteristics and type. Copying as is can mean using the pixel values of the area as is, while transforming the data can mean not using the pixel values of the area as is. That is, the transformation process can cause at least one pixel value change in the area to be filled in the expanded area, or at least one pixel acquisition position can be different. That is, to fill the expanded area A×B, data C×D of the area can be used instead of data A×B. In other words, at least one motion vector applied to the pixels to be filled can be different. The above example may occur when a 360-degree image is composed of multiple surfaces depending on the projection format and data from other surfaces is used to fill the expanded area. The data processing method for filling an expanded area due to image resizing is not limited to the above example, and improvements, modifications, or additional data processing methods can be used.
[0241] A set of candidates for multiple data processing methods can be supported depending on the encoding / decoding settings, and data processing method selection information can be generated from the set of candidates and recorded in the bitstream. For example, one data processing method can be selected from a method of filling using a predetermined pixel value, a method of filling by copying outer pixels, a method of filling by copying a partial area of an image, a method of filling by transforming a partial area of an image, etc., and selection information for this method can be generated. Alternatively, the data processing method can be determined implicitly.
[0242] For example, the data processing method applied to the entire area (TL to BR in FIG. 7a in this example) expanded by resizing the image may be one of filling using predetermined pixel values, filling by copying outer pixels, filling by copying a partial area of the image, filling by transforming a partial area of the image, or other methods, and selection information for the method may be generated. Also, a predetermined data processing method to be applied to the entire area may be determined.
[0243] Alternatively, the data processing method applied to the area expanded by resizing the image (in this example, each of the areas TL to BR in 7a of FIG. 7 or two or more of them) may be one of a method of filling using predetermined pixel values, a method of filling by copying outer pixels, a method of filling by copying a partial area of the image, a method of filling by transforming a partial area of the image, and other methods, and selection information for this method may be generated. Also, one predetermined data processing method to be applied to at least one area may be determined.
[0244] FIG. 9 is an exemplary diagram illustrating a method for configuring areas to be deleted and areas to be created by reducing an image size in an image size adjustment method according to an embodiment of the present invention.
[0245] The area to be deleted in the image reduction process may not only be simply removed, but may be removed after going through a series of exploitation processes.
[0246] Referring to FIG. 9a, during the image reduction process, some areas A0, A1, and A2 can be simply removed without any additional processing. In this case, image A can be subdivided and referred to as TL to BR, as in FIG. 8a.
[0247] Referring to Figure 9b, partial regions A0 to A2 are removed, but can be used as reference information when encoding / decoding image A. For example, the removed partial regions A0 to A2 can be used in a process of restoring or correcting a partial region of image A generated by reducing it. The restoration or correction process can use a weighted sum or average of two regions (the deleted region and the generated region). Furthermore, the restoration or correction process can be a process that can be applied when two regions have a high correlation.
[0248] As an example, an area B'2 that is deleted by shrinking to the left or right in relation to the horizontal size adjustment of an image can be used to restore or correct pixels in the area B2, LC opposite the area being shrunk in the left or right direction in relation to the horizontal size adjustment, and then the corresponding area can be removed from memory.
[0249] As an example, the area B'1 that is deleted in the upward or downward direction related to the vertical size adjustment of the image can be used in the encoding / decoding process (restoration or correction process) of the area B1, TR, opposite the area to be reduced in the upward or downward direction related to the vertical size adjustment, and then the corresponding area can be removed from memory.
[0250] As an example, when a portion of an image is reduced in size (in this example, diagonally from the center of the image), the area B'0 opposite the reduced area can be used in the encoding / decoding process (such as restoration or correction process) of the TL, and then the corresponding area can be removed from memory.
[0251] In the above example, we have explained the case where the data is used to restore or correct an area where there is continuity at the boundaries on both ends of the image and which is located symmetrically to the size adjustment direction, but this is not limited to this, and the data can also be removed from memory after being used to restore or correct data in other areas TL to BR that are not located symmetrically.
[0252] The data processing method for removing the reduced area of the present invention is not limited to the above example, and it may be improved or modified, or additional data processing methods may be used.
[0253] Depending on the encoding / decoding settings, a group of candidates for multiple data processing methods can be supported, and selection information for the candidate can be generated and recorded in the bitstream. For example, one data processing method can be selected from a method of simply removing the resized area, a method of using the resized area in a series of processes and then removing it, and selection information for the candidate can be generated. Alternatively, the data processing method can be determined implicitly.
[0254] For example, the data processing method applied to the entire area (TL to BR in 7b of FIG. 7 in this example) that is deleted due to reduction in image size adjustment can be one of a simple removal method, a removal method after using it in a series of processes, or other methods, and selection information for this method can be generated. Also, the data processing method can be implicitly determined.
[0255] Alternatively, the data processing method applied to the individual regions (TL to BR in FIG. 7b in this example) that are reduced by resizing the image may be one of a simple removal method, a removal method after using it in a series of processes, or other methods, and selection information for this may be generated. Alternatively, the data processing method may be implicitly determined.
[0256] In the above example, we have explained the case of performing size adjustment by size adjustment operations (expansion, reduction), but in some cases, this may be an example that can be applied when performing a size adjustment operation (in this example, expansion) and then performing a size adjustment operation (in this example, reduction), which is the reverse process.
[0257] For example, a method of filling the expanded area with partial image data may be selected, and a method of using the reduced area in a partial image data restoration or correction process and then removing it in the inverse process may be selected. Alternatively, a method of filling the expanded area with copies of outer pixels may be selected, and a method of simply removing the reduced area in the inverse process may be selected. In other words, the data processing method in the inverse process can be determined based on the data processing method selected in the image resizing process.
[0258] Unlike the above example, the data processing methods for the image resizing process and the inverse process can be independent of each other. That is, the data processing method for the inverse process can be selected independently of the data processing method selected for the image resizing process. For example, a method can be selected in which the expanded area is filled with partial image data, and a method can be selected in which the reduced area is simply removed in the inverse process.
[0259] In the present invention, the data processing method in the image size adjustment process can be implicitly determined according to the encoding / decoding settings, and the data processing method in the inverse process can be implicitly determined according to the encoding / decoding settings. Alternatively, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the inverse process can be explicitly generated. Alternatively, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the inverse process can be implicitly determined based on the data processing method.
[0260] Next, an example of image resizing in an encoding / decoding device according to an embodiment of the present invention will be described. In the following example, the resizing process is expansion and the inverse resizing process is reduction. The difference between the "image before resizing" and the "image after resizing" may refer to the size of the image. The resizing-related information may be partially explicitly generated or partially implicitly determined depending on the encoding / decoding settings. The resizing-related information may include information on the resizing process and the inverse resizing process.
[0261] As a first example, a size adjustment process can be performed on an input image before encoding begins. After size adjustment is performed using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; the data processing method is used in the size adjustment process), the size-adjusted image can be encoded. After encoding is completed, it can be stored in memory, and the image encoding data (meaning the size-adjusted image in this example) can be included in a bitstream and transmitted.
[0262] A resizing process can be performed before decoding begins. After resizing using resizing information (e.g., resizing operation, resizing direction, resizing value, etc.), the resized decoded image data can be parsed and decoded. After decoding is complete, the data can be stored in memory, and a resizing inverse process (using a data processing method, etc., used in the resizing inverse process) can be performed to change the output image back to the image before resizing.
[0263] As a second example, a size adjustment process can be performed on a reference image before encoding begins. After performing size adjustment using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; the data processing method is the one used in the size adjustment process), the size-adjusted image (in this example, the size-adjusted reference image) can be stored in memory, and an image can be encoded using this. After encoding is completed, image encoding data (in this example, meaning the image encoded using the reference image) can be recorded in a bitstream and transmitted. Furthermore, if the encoded image is stored in memory as a reference image, the size adjustment process can be performed as described above.
[0264] Before decoding begins, a size adjustment process can be performed on the reference image. The size-adjusted image (in this example, the size-adjusted reference image) can be stored in memory using size adjustment information (e.g., size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc.; the data processing method is the one used in the size adjustment process). The image decoded data (in this example, the same as that encoded using the reference image in the encoder) can be parsed and decoded. After decoding is complete, an output image can be generated. If the decoded image is included in the reference image and stored in memory, the size adjustment process can be performed as described above.
[0265] As a third example, after completion of encoding (specifically, meaning completion of encoding excluding filtering), a resizing process can be performed on an image before starting filtering of the image (assumed to be a deblocking filter in this example). After performing resizing using resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is the one used in the resizing process), a resized image can be generated, and filtering can be applied to the resized image. After filtering is completed, a resizing reverse process can be performed to return the image to its pre-resize state.
[0266] (More specifically, this means completion of decoding excluding the filtering process.) After completion of decoding, a resizing process can be performed on the image before starting filtering of the image. After performing resizing using resizing information (e.g., resizing operation, resizing direction, resizing value, data processing method, etc.; the data processing method is the one used in the resizing process), a resized image can be generated, and filtering can be applied to the resized image. After filtering is completed, a resizing reverse process can be performed to return the image to its pre-resize state.
[0267] In the above examples, in some cases (first and third examples), the size adjustment process and the reverse size adjustment process may be performed, and in other cases (second example), only the size adjustment process may be performed.
[0268] In addition, in some cases (examples 2 and 3), the size adjustment process in the encoder and decoder may be the same, while in other cases (example 1), the size adjustment process in the encoder and decoder may or may not be the same. In this case, the difference between the size adjustment process in the encoder / decoder may be a size adjustment execution step. For example, in some cases (in this example, the encoder), a size adjustment execution step that considers image size adjustment and data processing of the resized area may be included, while in other cases (in this example, the decoder), a size adjustment execution step that considers image size adjustment may be included. In this case, the data processing in the former may correspond to the data processing in the reverse size adjustment process in the latter.
[0269] In addition, in some cases (e.g., the third example), the resizing process is a process applied only to the relevant step, and the resizing area does not need to be stored in memory. For example, the resizing area can be stored in temporary memory for use in the filtering process, filtered, and then removed through the reverse resizing process. In this case, the resizing process does not change the size of the image. The above example is not limited to the above, and other variations are possible.
[0270] The size of the image can be changed through the resizing process, and therefore, the coordinates of some pixels of the image can be changed through the resizing process. This can affect the operation of the picture divider. In the present invention, block-based division can be performed based on the image before resizing through the above process, or block-based division can be performed based on the image after resizing. Also, division into some units (e.g., tiles, slices, etc.) can be performed based on the image before resizing, or block-based division can be performed based on the image after resizing. This can be determined depending on the encoding / decoding settings. In the present invention, the case where the picture divider operates based on the image after resizing (e.g., the image division process after the resizing process) will be mainly described, but other variations are also possible. This will be described for the above example in multiple image settings described below.
[0271] The encoder records the information generated in the above process into a bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder parses the related information from the bitstream, and the information may be included in the bitstream in the form of SEI or metadata.
[0272] Although it is common to encode / decode an input image as is, this problem may also occur when encoding / decoding an image by reconstructing it. For example, image reconstruction can be performed to improve image coding efficiency, to take into account the network and user environment, or according to the type and characteristics of the image.
[0273] In the present invention, the image reconstruction process can be performed independently or can be an inverse process. In the examples below, the reconstruction process will be mainly described, but the content of the inverse reconstruction process can be derived from the reconstruction process.
[0274] FIG. 10 is an exemplary diagram for image reconstruction according to an embodiment of the present invention.
[0275] When 10a is the first image input, 10a to 10d are illustrative diagrams in which the image is rotated by 0 degrees (for example, a group of candidates can be generated by sampling 360 degrees into k intervals. k can have values such as 2, 4, 8, etc., and in this example, 4 is assumed), and 10e to 10h are illustrative diagrams in which inversion (or symmetry) is applied based on 10a or 10b to 10d.
[0276] The starting position or scan order of an image can be changed depending on the image reconstruction, or a predetermined starting position and scan order can be followed regardless of whether the image is reconstructed. This can be determined depending on the encoding / decoding settings. In the following embodiments, it is assumed that a predetermined starting position (e.g., the upper left position of the image) and scan order (e.g., raster scan) are followed regardless of whether the image is reconstructed.
[0277] An image encoding method and a decoding method according to an embodiment of the present invention may include the following image reconstruction steps. In this case, the image reconstruction process may include an image reconstruction instruction step, an image reconstruction type identification step, and an image reconstruction execution step. Also, the image encoding device and the decoding device may be configured to include an image reconstruction instruction unit, an image reconstruction type identification unit, and an image reconstruction execution unit that realize the image reconstruction instruction step, the image reconstruction type identification step, and the image reconstruction execution step. In the case of encoding, an associated syntax element may be generated, and in the case of decoding, the associated syntax element may be parsed.
[0278] In the image reconstruction instruction step, it can be determined whether to perform image reconstruction. For example, if a signal instructing image reconstruction (e.g., convert_enabled_flag) is confirmed, reconstruction can be performed. If the signal instructing image reconstruction is not confirmed, reconstruction is not performed, or reconstruction can be performed by checking other encoding / decoding information. Also, even if a signal instructing image reconstruction is not provided, the signal instructing reconstruction can be implicitly activated or deactivated depending on the encoding / decoding settings (e.g., image characteristics, type, etc.). If reconstruction is performed, reconstruction-related information can be generated accordingly, or the reconstruction-related information can be implicitly determined.
[0279] When a signal instructing image reconstruction is provided, the corresponding signal indicates whether or not to perform image reconstruction, and whether or not to perform image reconstruction can be confirmed according to the signal. For example, when a signal instructing image reconstruction (e.g., convert_enabled_flag) is confirmed, if the corresponding signal is activated (e.g., convert_enabled_flag=1), reconstruction may be performed, and if the corresponding signal is deactivated (e.g., convert_enabled_flag=0), reconstruction may not be performed.
[0280] In addition, if a signal instructing image reconstruction is not provided, reconstruction may not be performed, or whether the image is reconstructed may be confirmed by another signal. For example, reconstruction may be performed according to the characteristics or type of image (e.g., a 360-degree image), and reconstruction information may be explicitly generated or assigned a predetermined value. The present invention is not limited to the above example, and other modifications are also possible.
[0281] In the image reconstruction type identification step, the image reconstruction type can be identified. The image reconstruction type can be defined by a reconstruction method, reconstruction mode information, etc. The reconstruction method (e.g., convert_type_flag) can include rotation, inversion, etc., and the reconstruction mode information can include a mode in the reconstruction method (e.g., convert_mode). In this case, the reconstruction-related information can be composed of a reconstruction method and mode information. That is, it can be composed of at least one syntax element. In this case, the number of mode information candidates for each reconstruction method can be the same or different.
[0282] As an example, in the case of rotation, candidates with a fixed difference (90 degrees in this example) such as 10a to 10d can be included, and if 10a is a 0-degree rotation, 10b to 10d can be examples of applying a 90-degree, 180-degree, and 270-degree rotation, respectively (in this example, the angle is measured clockwise).
[0283] For example, in the case of inversion, candidates such as 10a, 10e, and 10f may be included, and when 10a is no inversion, 10e and 10f may be examples of applying horizontal inversion and vertical inversion, respectively.
[0284] The above example describes a case where a setting is made for rotation with a fixed interval and a setting is made for partial inversion, but this is only one example for image reconstruction and is not limited to the above case, and other interval differences can include examples such as other inversion operations, etc. This can be determined depending on the encoding / decoding settings.
[0285] Alternatively, the information may include integrated information (e.g., convert_com_flag) generated by combining the reconfiguration method and the mode information associated therewith. In this case, the reconfiguration-related information may be configured as a combination of the reconfiguration method and the mode information.
[0286] For example, the integrated information may include candidates such as 10a to 10f, which may be examples of applying 0-degree rotation, 90-degree rotation, 180-degree rotation, 270-degree rotation, horizontal flip, and vertical flip based on 10a.
[0287] Alternatively, the integrated information may include candidates such as 10a to 10h, which may be examples of applying 0-degree rotation, 90-degree rotation, 180-degree rotation, 270-degree rotation, horizontal flip, vertical flip, horizontal flip after 90-degree rotation (or 90-degree rotation after horizontal flip), and vertical flip after 90-degree rotation (or 90-degree rotation after vertical flip) based on 10a, or examples of applying 0-degree rotation, 90-degree rotation, 180-degree rotation, 270-degree rotation, horizontal flip, horizontal flip after 180-degree rotation (180-degree rotation after horizontal flip), horizontal flip after 90-degree rotation (90-degree rotation after horizontal flip), and horizontal flip after 270-degree rotation (270-degree rotation after horizontal flip).
[0288] The candidate group may include a rotated mode, a flipped mode, and a mixed rotation and flip mode. The mixed mode simply includes mode information of the reconstruction method and may include a mode generated by mixing mode information of each method. In this case, it may include a mode generated by mixing at least one mode of some methods (e.g., rotation) and at least one mode of some methods (e.g., flip). The above example includes a case where one mode of some methods and multiple modes of some methods are mixed (in this example, 90-degree rotation + multiple flips / left-right flips + multiple rotations). The mixed information may include a candidate group (in this example, 10a) when reconstruction is not applied, and may include a first candidate group (e.g., index 0) when reconstruction is not applied.
[0289] Alternatively, mode information according to a predetermined reconfiguration method may be included. In this case, the reconfiguration-related information may be configured with mode information according to a predetermined reconfiguration method. That is, information about the reconfiguration method may be omitted, and the reconfiguration-related information may be configured with one syntax element related to the mode information.
[0290] For example, it can be configured to include candidates such as 10a to 10d related to rotation, or candidates such as 10a, 10e, and 10f related to inversion.
[0291] The size of the image before and after the image reconstruction process may be the same, or at least one length may be different. This can be determined according to the encoding / decoding settings. The image reconstruction process is a process of rearranging pixels within an image (in this example, the image reconstruction inverse process performs an inverse pixel rearrangement process, which can be inversely derived from the pixel rearrangement process), and the position of at least one pixel may be changed. The pixel rearrangement may be performed according to a rule based on the image reconstruction type information.
[0292] In this case, the pixel rearrangement process may be affected by the size and shape (e.g., square or rectangular) of the image. In particular, the width and height of the image before the reconstruction process and the width and height of the image after the reconstruction process may act as variables in the pixel rearrangement process.
[0293] For example, at least one of ratio information (e.g., former / latter or latter / former) of the ratio between the width of the image before the reconstruction process and the width of the image after the reconstruction process, the ratio between the width of the image before the reconstruction process and the height of the image after the reconstruction process, the ratio between the height of the image before the reconstruction process and the width of the image after the reconstruction process, and the ratio between the height of the image before the reconstruction process and the height of the image after the reconstruction process can act as a variable in the pixel rearrangement process.
[0294] In the above example, if the size of the image before and after the reconstruction process is the same, the ratio of the width to the height of the image can act as a variable in the pixel rearrangement process. Also, if the image shape is square, the ratio of the length of the image before the image reconstruction process to the length of the image after the reconstruction process can act as a variable in the pixel rearrangement process.
[0295] In the image reconstruction execution stage, an image can be reconstructed based on the identified reconstruction information, i.e., based on information such as a reconstruction type and a reconstruction mode, and encoding / decoding can be performed based on the acquired reconstructed image.
[0296] Next, an example of image reconstruction performed by the encoding / decoding device according to one embodiment of the present invention will be shown.
[0297] A reconstruction process can be performed on the input image before encoding begins. After reconstruction is performed using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), the reconstructed image can be coded. After coding is complete, it can be stored in memory, or the coded image data can be included in a bitstream and transmitted.
[0298] A reconstruction process can be performed before the start of decoding. After reconstruction using reconstruction information (e.g., image reconstruction type, reconstruction mode, etc.), the image decoded data can be parsed and decoded. After decoding is complete, the image can be stored in memory, and after performing a reverse reconstruction process to change it to the image before reconstruction, the image can be output.
[0299] The encoder records the information generated in the above process into a bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder parses the related information from the bitstream, and the information may be included in the bitstream in the form of SEI or metadata.
[0300] [Table 1]
[0301] Table 1 shows examples of syntax elements related to division during image configuration. In the examples described below, the added syntax elements will be mainly described. Furthermore, the syntax elements in the examples described below are not limited to a specific unit, but may be syntax elements supported in various units such as a sequence, a picture, a slice, or a tile. Alternatively, they may be syntax elements included in SEI, metadata, etc. Furthermore, the types, order, and conditions of syntax elements supported in the examples described below are limited only to this example, and may be changed or determined depending on the encoding / decoding settings.
[0302] In Table 1, tile_header_enabled_flag is a syntax element that indicates whether encoding / decoding settings are supported for tiles. When activated (tile_header_enabled_flag=1), a tile can have encoding / decoding settings. When deactivated (tile_header_enabled_flag=0), a tile cannot have encoding / decoding settings and can be assigned the encoding / decoding settings of a higher unit.
[0303] The tile_coded_flag is a syntax element that indicates whether to encode / decode a tile. When activated (tile_coded_flag=1), the tile can be encoded / decoded, and when deactivated (tile_coded_flag=0), the tile cannot be encoded / decoded. Here, "not encoding" may mean that encoded data for the tile is not generated (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc., which may be applicable to meaningless areas in a partial projection format of a 360-degree image). "Not decoding" may mean that decoded data for the corresponding tile is no longer parsed (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc.). "No longer parsing decoded data" may mean that parsing is no longer performed because no encoded data exists in the corresponding unit, or that parsing is no longer performed due to the flag, even if encoded data exists. Header information for each tile may be supported depending on whether tile encoding / decoding is performed.
[0304] Although the above example has been described with a focus on tiles, it is not limited to tiles and may be applicable to other division units in the present invention. Furthermore, the example of tile division setting is not limited to the above case and may be modified to other examples.
[0305] [Table 2]
[0306] Table 2 shows examples of syntax elements related to reconstruction during image configuration.
[0307] Referring to Table 2, convert_enabled_flag is a syntax element that indicates whether reconstruction is performed. When activated (convert_enabled_flag=1), it means that a reconstructed image is encoded / decoded, and additional reconstruction-related information can be confirmed. When deactivated (convert_enabled_flag=0), it means that an existing image is encoded / decoded.
[0308] The convert_type_flag indicates the mixed information about the reconstruction method and mode information. It can determine one of several candidate groups of the method that applies rotation, the method that applies inversion, and the method that applies rotation and inversion together.
[0309] [Table 3]
[0310] Table 3 shows examples of syntax elements related to size adjustment during image configuration.
[0311] Referring to Table 3, pic_width_in_samples and pic_height_in_samples refer to syntax elements relating to the width and height of an image, and the size of the image can be confirmed by these syntax elements.
[0312] img_resizing_enabled_flag is a syntax element that indicates whether to resize an image. When activated (img_resizing_enabled_flag=1), it indicates that the resized image is to be encoded / decoded, and additional resizing-related information can be checked. When deactivated (img_resizing_enabled_flag=0), it indicates that the existing image is to be encoded / decoded. It may also be a syntax element that indicates resizing for intra-frame prediction.
[0313] The resizing_met_flag is a syntax element for the resizing method. It can be determined from a group of candidates, such as resizing using a scale factor (resizing_met_flag=0), resizing using an offset factor (resizing_met_flag=1), or other resizing methods.
[0314] The resizing_mov_flag indicates the syntax element for the resizing operation, for example, it can decide between expanding and shrinking.
[0315] The width_scale and height_scale refer to scale factors relating to the horizontal size adjustment and the vertical size adjustment among the size adjustments using the scale factors.
[0316] The top_height_offset and bottom_height_offset refer to the upward and downward offset factors related to the horizontal size adjustment among the size adjustments using the offset factors, and the left_width_offset and right_width_offset refer to the left and right offset factors related to the vertical size adjustment among the size adjustments using the offset factors.
[0317] The size of the image after resizing can be updated through the size adjustment related information and the image size information.
[0318] The resizing_type_flag is a syntax element for the data processing method of the area to be resized. Depending on the resizing method and resizing operation, the number of candidates for the data processing method may be the same or different.
[0319] The image setting processes applied to the image encoding / decoding device described above may be performed individually or multiple image setting processes may be mixed together. In the example described below, a case where multiple image setting processes are mixed together will be described.
[0320] 11 is an exemplary diagram showing images before and after an image setting process according to an embodiment of the present invention. In particular, 11a shows an example before image reconstruction is performed on the divided images (e.g., an image projected by 360-degree image coding), and 11b shows an example after image reconstruction is performed on the divided images (e.g., an image packed by 360-degree image coding). That is, 11a can be understood as an exemplary diagram before the image setting process is performed, and 11b as an exemplary diagram after the image setting process is performed.
[0321] The image setup process in this example is explained in terms of image segmentation (assumed to be tiles in this example) and image reconstruction.
[0322] In the example described below, a case where an image is reconstructed after image division will be described, but it is also possible to perform image division after image reconstruction according to encoding / decoding settings, and modifications to other cases are also possible. Furthermore, the above-mentioned image reconstruction process (including the inverse process) can be applied in the same or similar manner as the reconstruction process of the division unit within the image in this embodiment.
[0323] Image reconstruction may or may not be performed for all division units in the image, or may be performed for some of the division units. Therefore, the division units before reconstruction (e.g., some of P0 to P5) may or may not be identical to the division units after reconstruction (e.g., some of S0 to S5). Various cases related to the execution of image reconstruction will be described through examples described below. For convenience of explanation, it is assumed that the unit of an image is a picture, the unit of a divided image is a tile, and the division units are square in shape.
[0324] As an example, whether to perform image reconstruction can be determined by a certain unit (e.g., sps_convert_enabled_flag, SEI, metadata, etc.). Alternatively, whether to perform image reconstruction can be determined by a certain unit (e.g., pps_convert_enabled_flag). This is possible when it occurs first in the corresponding unit (in this example, a picture) or is activated in a higher unit (e.g., sps_convert_enabled_flag=1). Alternatively, whether to perform image reconstruction can be determined by a certain unit (e.g., tile_convert_flag[i], where i is a division unit index). This is possible when it occurs first in the corresponding unit (in this example, a tile) or is activated in a higher unit (e.g., pps_convert_enabled_flag=1). Furthermore, whether to perform the partial image reconstruction can be implicitly determined according to the encoding / decoding settings, thereby omitting related information.
[0325] For example, whether to reconstruct a division unit within an image can be determined according to a signal (e.g., pps_convert_enabled_flag) instructing image reconstruction. Specifically, whether to reconstruct all division units within an image can be determined according to the signal. In this case, a signal instructing image reconstruction can be generated for the image.
[0326] For example, whether to reconstruct a division unit within an image can be determined in response to a signal (e.g., tile_convert_flag[i]) instructing image reconstruction. In particular, whether to reconstruct a partial division unit within an image can be determined in response to the signal. In this case, at least one signal instructing image reconstruction (e.g., generated as many times as the number of division units) can be generated.
[0327] As an example, whether to reconstruct an image can be determined according to a signal (e.g., pps_convert_enabled_flag) instructing image reconstruction, and whether to reconstruct a division unit within the image can be determined according to a signal (e.g., tile_convert_flag[i]) instructing image reconstruction. In particular, when some signals are activated (e.g., pps_convert_enabled_flag=1), some other signals (e.g., tile_convert_flag[i]) can be checked, and whether to reconstruct a division unit within the image can be determined according to the signal (in this example, tile_convert_flag[i]). At this time, signals instructing the reconstruction of multiple images can be generated.
[0328] When a signal instructing image reconstruction is activated, image reconstruction related information can be generated. Various cases of image reconstruction related information will be described in the examples below.
[0329] For example, reconstruction information that is applied to an image can be generated. Specifically, one reconstruction information can be used as reconstruction information for all division units within the image.
[0330] For example, reconstruction information that is applied to division units within an image may be generated. Specifically, at least one reconstruction information may be used as reconstruction information for a portion of the division units within the image. That is, one reconstruction information may be used as reconstruction information for one division unit, or one reconstruction information may be used as reconstruction information for multiple division units.
[0331] The example described below can be explained in combination with an example of image reconstruction.
[0332] For example, when a signal instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information commonly applied to division units within the image may be generated. Alternatively, when a signal instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information individually applied to division units within the image may be generated. Alternatively, when a signal instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information individually applied to division units within the image may be generated. Alternatively, when a signal instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information commonly applied to division units within the image may be generated.
[0333] The reconfiguration information can be processed implicitly or explicitly depending on the encoding / decoding settings. In the implicit case, the reconfiguration information can be assigned a predetermined value depending on the characteristics, type, etc. of the image.
[0334] P0 to P5 of 11a may correspond to S0 to S5 of 11b, and a reconfiguration process may be performed for each division unit. For example, P0 may be assigned to S0 without reconfiguration, P1 may be rotated 90 degrees and assigned to S1, P2 may be rotated 180 degrees and assigned to S2, P3 may be flipped left and right and assigned to S3, P4 may be rotated 90 degrees and flipped left and right and assigned to S4, and P5 may be rotated 180 degrees and flipped left and right and assigned to S5.
[0335] However, the present invention is not limited to the above-described examples, and various modifications are possible. As in the above-described example, reconstruction may not be performed for each division unit of an image, or at least one of the following reconstruction methods may be performed: reconstruction with rotation, reconstruction with inversion, and reconstruction with a combination of rotation and inversion.
[0336] When image reconstruction is applied to a division unit, an additional reconstruction process such as division unit rearrangement may be performed. That is, the image reconstruction process of the present invention may include rearrangement of pixels within an image as well as division unit image rearrangement, and may be expressed by some syntax elements (e.g., part_top, part_left, part_width, part_height, etc.) as shown in Table 4. This means that the image division and image reconstruction processes may be understood as being mixed. The above description may be a possible example when an image is divided into multiple units.
[0337] P0 to P5 of 11a may correspond to S0 to S5 of 11b, and a reconstruction process may be performed for each division unit. For example, P0 may be assigned to S0 without reconstruction, P1 may be assigned to S2 without reconstruction, P2 may be rotated 90 degrees and assigned to S1, P3 may be flipped horizontally and assigned to S4, P4 may be rotated 90 degrees and flipped horizontally and assigned to S5, and P5 may be flipped horizontally and rotated 180 degrees and assigned to S3, and various other variations are possible.
[0338] 7 can correspond to P_Width and P_Height in FIG. 11, and P'_Width and P'_Height in FIG. 7 can correspond to P'_Width and P'_Height in FIG. 11. The image size after size adjustment in FIG. 7 is P'_Width×P'_Height, which can be expressed as (P_Width+Exp_L+Exp_R)×(P_Height+Exp_T+Exp_B), and the image size after size adjustment in FIG. 11 is P'_Width×P'_Height, which can be expressed as (P_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R)×(P _Height+Var0_T+Var1_T+Var0_B+Var1_B) or (Sub_P0_Width+Sub_P1_Width+Sub_P2_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R) x (Sub_P0_Height+Sub_P1_Height+Var0_T+Var1_T+Var0_B+Var1_B).
[0339] As in the above example, image reconstruction can involve pixel rearrangement within a division unit of an image, or pixel rearrangement within a division unit within an image, and can involve not only pixel rearrangement within a division unit of an image but also pixel rearrangement within a division unit within an image. In this case, pixel rearrangement within a division unit can be performed after pixel rearrangement within the division unit, or pixel rearrangement within the division unit can be performed after pixel rearrangement within the division unit.
[0340] The rearrangement of division units within an image may be determined according to a signal instructing image reconstruction. Alternatively, a signal for rearrangement of division units within an image may be generated. In particular, the signal may be generated when a signal instructing image reconstruction is activated. Alternatively, the processing may be implicit or explicit according to encoding / decoding settings. If it is implicit, it may be determined according to the characteristics, type, etc. of the image.
[0341] Furthermore, information regarding the rearrangement of division units within an image can be processed implicitly or explicitly depending on the encoding / decoding settings, and can be determined depending on the characteristics, type, etc. of the image. In other words, each division unit can be arranged according to the arrangement information of a predetermined division unit.
[0342] Next, an example of reconstructing division units within an image using an encoding / decoding device according to an embodiment of the present invention will be shown.
[0343] Before encoding begins, a division process can be performed on an input image using division information. A reconstruction process can be performed on a division unit using reconstruction information, and the image reconstructed on a division unit can be coded. After coding is complete, the image can be stored in memory, or the coded image data can be recorded in a bitstream and transmitted.
[0344] Before decoding begins, a division process can be performed using division information. A reconstruction process can be performed using reconstruction information for the division units, and image decoded data can be parsed and decoded for the reconstructed division units. After decoding is complete, the data can be stored in memory, and after a reverse reconstruction process for the division units is performed, the division units can be merged into one and the image can be output.
[0345] Fig. 12 is a diagram illustrating size adjustment for each division unit in an image according to an embodiment of the present invention. P0 to P5 in Fig. 12 correspond to P0 to P5 in Fig. 11, and S0 to S5 in Fig. 12 correspond to S0 to S5 in Fig. 11.
[0346] In the examples described below, the case where image size adjustment is performed after image division will be mainly described, but it is also possible to perform image division after image size adjustment according to encoding / decoding settings, and modifications to other cases are also possible. Furthermore, the above-mentioned image size adjustment process (including the inverse process) can be applied in the same or similar manner as the process of adjusting the size of the division unit within the image in this embodiment.
[0347] For example, TL to BR in Figure 7 can correspond to TL to BR of the division unit SX (S0 to S5) in Figure 12, S0 and S1 in Figure 7 can correspond to PX and SX in Figure 12, P_Width and P_Height in Figure 7 can correspond to Sub_PX_Width and Sub_PX_Height in Figure 12, P'_Width and P'_Height in Figure 7 can correspond to Sub_SX_Width and Sub_SX_Height in Figure 12, Exp_L, Exp_R, Exp_T, and Exp_B in Figure 7 can correspond to VarX_L, VarX_R, VarX_T, and VarX_B in Figure 12, and other factors can also be corresponded.
[0348] The size adjustment process for division units within an image in 12a to 12f can be distinguished from the image size enlargement or reduction in 7a and 7b of Figures 7 in that there may be settings for image size enlargement or reduction in proportion to the number of division units. Also, there may be differences in that there may be settings commonly applied to division units within an image, or settings individually applied to division units within an image. Various size adjustment cases will be described in examples below, and the size adjustment process can be performed taking the above into consideration.
[0349] In the present invention, image size adjustment may be performed on all division units within an image, or may be performed on some division units. Various image size adjustment cases will be described using examples below. For convenience of explanation, it is assumed that the size adjustment operation is expansion, the size adjustment method is an offset factor, the size adjustment direction is up, down, left, and right, the size adjustment direction is set according to size adjustment information, the image unit is a picture, and the divided image unit is a tile.
[0350] As an example, whether to resize an image can be determined by some units (e.g., sps_img_resizing_enabled_flag, or SEI, metadata, etc.). Alternatively, whether to resize an image can be determined by some units (e.g., pps_img_resizing_enabled_flag). This is possible when it occurs first in the corresponding unit (picture in this example) or is activated in a higher unit (e.g., sps_img_resizing_enabled_flag=1). Alternatively, whether to resize an image can be determined by some units (e.g., tile_resizing_flag[i], where i is a division unit index). This is possible when it occurs first in the corresponding unit (tile in this example) or is activated in a higher unit. Furthermore, whether to resize the partial images can be implicitly determined according to encoding / decoding settings, thereby omitting related information.
[0351] For example, whether to resize division units within an image can be determined according to a signal instructing image resizing (e.g., pps_img_resizing_enabled_flag). More specifically, whether to resize all division units within an image can be determined according to the signal. In this case, a signal instructing image resizing can be generated.
[0352] For example, whether to resize a division unit within an image may be determined in response to a signal (e.g., tile_resizing_flag[i]) instructing image resizing. Specifically, whether to resize a partial division unit within an image may be determined in response to the signal. In this case, at least one signal instructing image resizing (e.g., generated as many times as the number of division units) may be generated.
[0353] As an example, whether to resize an image can be determined according to a signal instructing image resizing (e.g., pps_img_resizing_enabled_flag), and whether to resize a division unit within the image can be determined according to a signal instructing image resizing (e.g., tile_resizing_flag[i]). In particular, when some signals are activated (e.g., pps_img_resizing_enabled_flag=1), some other signals (e.g., tile_resizing_flag[i]) can be checked, and whether to resize some division units within the image can be determined according to the signal (in this example, tile_resizing_flag[i]). In this case, signals instructing image resizing can be generated.
[0354] When a signal instructing image resizing is activated, information related to image resizing may be generated. Various cases of information related to image resizing will be described in the following examples.
[0355] As an example, size adjustment information to be applied to an image may be generated. Specifically, one size adjustment information or size adjustment information set may be used as size adjustment information for all division units within an image. For example, one size adjustment information (or size adjustment values applied to all size adjustment directions supported or allowed in the division units; in this example, one information) commonly applied to the top, bottom, left, and right directions of the division units within the image may be generated, or one size adjustment information set (or as many size adjustment directions supported or allowed in the division units; in this example, up to four pieces of information) individually applied to the top, bottom, left, and right directions may be generated.
[0356] As an example, size adjustment information to be applied to division units within an image may be generated. Specifically, at least one size adjustment information or size adjustment information set may be used as size adjustment information for a portion of division units within an image. That is, one size adjustment information or size adjustment information set may be used as size adjustment information for one division unit or for multiple division units. For example, one size adjustment information commonly applied to the top, bottom, left, and right directions of one division unit within an image, or one size adjustment information set individually applied to the top, bottom, left, and right directions, may be generated. Alternatively, one size adjustment information commonly applied to the top, bottom, left, and right directions of multiple division units within an image, or one size adjustment information set individually applied to the top, bottom, left, and right directions, may be generated. The configuration of a size adjustment set refers to size adjustment value information for at least one size adjustment direction.
[0357] In summary, size adjustment information that is commonly applied to division units within an image may be generated, or size adjustment information that is individually applied to division units within an image may be generated. The following examples can be explained in combination with examples of image size adjustment.
[0358] For example, when a signal indicating image resizing (e.g., pps_img_resizing_enabled_flag) is activated, resizing information commonly applied to division units within the image may be generated. Alternatively, when a signal indicating image resizing (e.g., pps_img_resizing_enabled_flag) is activated, resizing information individually applied to division units within the image may be generated. Alternatively, when a signal indicating image resizing (e.g., tile_resizing_flag[i]) is activated, resizing information individually applied to division units within the image may be generated. Alternatively, when a signal indicating image resizing (e.g., tile_resizing_flag[i]) is activated, resizing information commonly applied to division units within the image may be generated.
[0359] The image resizing direction and resizing information can be implicitly or explicitly processed depending on the encoding / decoding settings. In the implicit case, the resizing information can be assigned a predetermined value depending on the characteristics, type, etc. of the image.
[0360] As described above, the resizing direction in the resizing process of the present invention is at least one of up, down, left, and right directions, and the resizing direction and resizing information can be processed explicitly or implicitly. That is, some directions are implicitly assigned resizing values (including 0, i.e., no resizing) and some directions are explicitly assigned resizing values (including 0, i.e., no resizing).
[0361] For each division unit within an image, the resizing direction and resizing information can have settings that can be processed implicitly or explicitly, and these can be applied to the division units within the image. For example, a setting that is applied to one division unit within the image (in this example, only the division unit is generated) can be generated, or a setting that is applied to multiple division units within the image can be generated, or a setting that is applied to all division units within the image (in this example, one setting is generated) can be generated, and at least one setting can be generated for the image (for example, as many settings as there are division units can be generated from one setting). A setting set can be defined by collecting the setting information that is applied to the division units within the image.
[0362] FIG. 13 is an example diagram for adjusting or setting the size of division units within an image.
[0363] In detail, various examples of implicit or explicit processing of resizing directions and resizing information of division units within an image are shown. In the examples described below, for convenience of explanation, the implicit processing is explained assuming that the resizing values of some resizing directions are 0.
[0364] As in 13a, when the boundary of the division unit matches the boundary of the image (in this example, the thick solid line), explicit size adjustment processing is possible, and when they do not match (thin solid line), implicit size adjustment processing is possible. For example, P0 can be adjusted in size upward and leftward (a2, a0), P1 upward (a2), P2 upward and rightward (a2, a1), P3 downward and leftward (a3, a0), P4 downward (a3), and P5 downward and rightward (a3, a1), but size adjustment in other directions is not possible.
[0365] As shown in 13b, some directions of the division units (up and down in this example) can be explicitly adjusted for size, and some directions of the division units (left and right in this example) can be explicitly adjusted (thick solid lines in this example) if the boundaries of the division units match the boundaries of the image, and implicitly adjusted if they do not match (thin solid lines in this example). For example, P0 can be adjusted in size up, down, and left directions (b2, b3, b0), P1 in up and down directions (b2, b3), P2 in up, down, and right directions (b2, b3, b1), P3 in up, down, and left directions (b3, b4, b0), P4 in up and down directions (b3, b4), and P5 in up, down, and right directions (b3, b4, b1), but size adjustment is not possible in other directions.
[0366] As shown in 13c, some directions of the division units (left, right in this example) can be explicitly adjusted for size, and some directions of the division units (up, down in this example) can be explicitly adjusted if the boundaries of the division units match the boundaries of the image (thick solid lines in this example) and implicitly adjusted if they do not match (thin solid lines in this example). For example, P0 can be adjusted in size in the up, left, right directions (c4, c0, c1), P1 in the up, left, right directions (c4, c1, c2), P2 in the up, left, right directions (c4, c2, c3), P3 in the down, left, right directions (c5, c0, c1), P4 in the down, left, right directions (c5, c1, c2), and P5 in the down, left, right directions (c5, c2, c3), but size adjustment is not possible in other directions.
[0367] As in the above example, the settings related to image resizing can have various cases. Multiple setting sets can be supported and setting set selection information can be generated explicitly, or a predetermined setting set can be implicitly determined depending on the encoding / decoding settings (e.g., image characteristics, type, etc.).
[0368] FIG. 14 is an exemplary diagram illustrating both the image size adjustment process and the size adjustment process of the division unit within the image.
[0369] 14, the image size adjustment process and the inverse process can proceed in the directions of e and f, and the size adjustment process and the inverse process of the division units within the image can proceed in the directions of d and g. That is, the size adjustment process can be performed on the image, and the size adjustment of the division units within the image can be performed, and the order of the size adjustment processes is not fixed. This means that multiple size adjustment processes are possible.
[0370] In summary, the image size adjustment process can be classified into image size adjustment (or image size adjustment before division) and size adjustment of division units within an image (or image size adjustment after division). Neither image size adjustment nor size adjustment of division units within an image may be performed, or either one of both may be performed, or both may be performed. This can be determined depending on the encoding / decoding settings (e.g., image characteristics, type, etc.).
[0371] In the above example, when multiple size adjustment processes are performed, the image size adjustment may be performed in at least one of the directions of the top, bottom, left, and right of the image, or at least one division unit within the image may be adjusted in at least one of the directions of the top, bottom, left, and right of the division unit to be adjusted.
[0372] Referring to FIG. 14, the size of image A before size adjustment can be defined as P_Width×P_Height, the size of image B after primary size adjustment (or image B before secondary size adjustment) can be defined as P'_Width×P'_Height, and the size of image C after secondary size adjustment (or image C after final size adjustment) can be defined as P"_Width×P"_Height. Image A before size adjustment means an image that has not been subjected to any size adjustment, image B after primary size adjustment means an image that has been subjected to partial size adjustment, and image C after secondary size adjustment means an image that has been subjected to all size adjustments. For example, image B after primary size adjustment can mean an image that has been subjected to size adjustment for each division unit within the image as shown in FIGS. 13a to 13c, and image C after secondary size adjustment can mean an image that has been subjected to size adjustment for the entire image B that has been subjected to primary size adjustment as shown in FIG. 7a, or vice versa. The above examples are not limiting, and various modifications are possible.
[0373] Regarding the size of image B after the primary size adjustment, P'_Width can be obtained by P_Width and at least one size adjustment value in the left or right direction for horizontal size adjustment, and P'_Height can be obtained by P_Height and at least one size adjustment value in the up or down direction for vertical size adjustment, where the size adjustment value may be generated in a division unit.
[0374] Regarding the size of image C after secondary size adjustment, P"_Width can be obtained by P'_Width and at least one size adjustment value in the left or right direction that can be adjusted horizontally, and P"_Height can be obtained by P'_Height and at least one size adjustment value in the up or down direction that can be adjusted vertically. In this case, the size adjustment value may be a size adjustment value generated from the image.
[0375] In summary, the size of the image after resizing can be obtained via the size of the image before resizing and at least one resizing value.
[0376] Information regarding a data processing method may be generated in the resized area of the image. Various data processing methods will be described using examples below. The data processing method occurring during the resizing process may be the same as or similar to the resizing process, and the data processing methods during the resizing process and the resizing process may be described using various combinations of examples below.
[0377] As an example, a data processing method to be applied to an image may be generated. Specifically, one data processing method or a set of data processing methods may be used as the data processing method for all division units in the image (assuming that all division units are resized in this example). For example, one data processing method may be commonly applied to the top, bottom, left, and right directions of the division units in the image (or data processing methods applied to all size adjustment directions supported or allowed in the division units; one piece of information in this example) or one data processing method set may be individually applied to the top, bottom, left, and right directions (or as many pieces of information as the number of size adjustment directions supported or allowed in the division units; up to four pieces of information in this example).
[0378] As an example, a data processing method to be applied to a division unit within an image may be generated. Specifically, at least one data processing method or a set of data processing methods may be used as a data processing method for some division units within an image (assumed to be the division units to be resized in this example). That is, one data processing method or a set of data processing methods may be used as a data processing method for one division unit or for multiple division units. For example, one data processing method may be commonly applied to the top, bottom, left, and right directions of one division unit within an image, or one data processing method set may be individually applied to the top, bottom, left, and right directions of each division unit within an image. Alternatively, one data processing method may be commonly applied to the top, bottom, left, and right directions of multiple division units within an image, or one data processing method set may be individually applied to the top, bottom, left, and right directions of each division unit within an image. The configuration of a data processing method set refers to a data processing method for at least one resizing direction.
[0379] In summary, a data processing method commonly applied to division units within an image can be used. Alternatively, a data processing method individually applied to division units within an image can be used. The data processing method can be a predetermined method. There can be at least one predetermined data processing method. This corresponds to an implicit case, and explicit selection information regarding the data processing method can be generated. This can be determined according to encoding / decoding settings (e.g., image characteristics, type, etc.).
[0380] That is, a data processing method commonly applied to division units within an image can be used, a predetermined method can be used, or one of a plurality of data processing methods can be selected, or a data processing method individually applied to division units within an image can be used, and a predetermined method can be used or one of a plurality of data processing methods can be selected depending on the division unit.
[0381] The examples described below will explain some cases (in this example, partial image data is used to fill the resizing area) regarding resizing (assuming expansion in this example) of division units within an image.
[0382] The size of partial regions TL to BR of some units (e.g., S0 to S5 in FIGS. 12a to 12f) can be adjusted using data of partial regions tl to br of some units (P0 to P5 in FIGS. 12a to 12f). At this time, the partial units may be the same (e.g., S0 and P0) or different regions (e.g., S0 and P1). That is, the regions TL to BR to be adjusted in size can be filled using partial data tl to br of the division unit, and the regions to be adjusted in size can be filled using partial data of a division unit different from the division unit.
[0383] For example, the regions TL to BR to be resized in the current division unit can be resized using the tl to br data of the current division unit. For example, TL of S0 can be filled with the tl data of P0, RC of S1 with the tr+rc+br data of P1, BL+BC of S2 with the bl+bc+br data of P2, and TL+LC+BL of S3 with the tl+lc+bl data of P3.
[0384] As an example, the regions TL to BR of the current division unit to be resized can be resized using the tl to br data of division units spatially adjacent to the current division unit. For example, TL+TC+TR of S4 can be filled with the b1+bc+br data of P1 in the upward direction, BL+BC of S2 can be filled with the tl+tc+tr data of P5 in the downward direction, LC+BL of S2 can be filled with the tl+rc+bl data of P1 in the left direction, RC of S3 can be filled with the tl+lc+bl data of P4 in the right direction, and BR of S0 can be filled with the tl data of P4 in the downward left direction.
[0385] As an example, the region TL-BR to be resized in the current division unit can be resized using the tl-br data of division units that are not spatially adjacent to the current division unit. For example, data on the boundary regions at both ends of the image (e.g., left and right, top and bottom, etc.) can be obtained. LC of S3 can be obtained using the tr+rc+br data of S5, RC of S2 can be obtained using the tl+lc data of S0, BC of S4 can be obtained using the tc+tr data of S1, and TC of S1 can be obtained using the bc data of S4.
[0386] Alternatively, data on a partial region of the image (a region that is not spatially adjacent but is determined to have a high correlation with the region to be resized) can be obtained. The BC of S1 can be obtained using the tl+lc+bl data of S3, the RC of S3 can be obtained using the tl+tc data of S1, and the RC of S5 can be obtained using the bc data of S0.
[0387] In addition, some cases (in this example, removal by restoration or correction using partial image data) related to size adjustment of division units within an image (in this example, assumed to be reduction) are as follows.
[0388] The partial regions TL-BR of the partial units (e.g., S0 to S5 in Figures 12a to 12f) can be used in the restoration or correction process of the partial regions tl-br of the partial units P0 to P5. In this case, the partial units may be the same (e.g., S0 and P0) or different regions (e.g., S0 and P2). That is, the region to be resized can be used to restore and remove some data of the corresponding division unit, and the region to be resized can be used to restore and remove some data of a division unit different from the corresponding division unit. A detailed example is omitted as it can be derived inversely from the extension process.
[0389] The above example is an example that is applied when there is data that is highly correlated with the area to be resized, and the information on the position referenced for resizing can be explicitly generated, implicitly acquired based on a predetermined rule, or a combination of these can be used to confirm related information. This may be an example that is applied when acquiring data from other areas where continuity exists in encoding a 360-degree image.
[0390] Next, an example of adjusting the size of a division unit within an image in an encoding / decoding device according to an embodiment of the present invention will be shown.
[0391] Before encoding begins, an input image can be divided. A size adjustment process can be performed using size adjustment information for each division unit, and the image after size adjustment for each division unit can be coded. After coding is complete, the image can be stored in memory, or the coded image data can be recorded in a bitstream and transmitted.
[0392] Before decoding begins, a division process can be performed using division information. A size adjustment process can be performed using size adjustment information for the division units, and the image decoded data can be parsed and decoded for the size-adjusted division units. After decoding is complete, the data can be stored in memory, and after performing a reverse size adjustment process for the division units, the division units can be merged into one and the image can be output.
[0393] Other cases in the image size adjustment process described above can be modified and applied as in the above example, but the present invention is not limited to this and can be modified to other examples.
[0394] In the image setting process, a combination of image size adjustment and image reconstruction is possible. Image size adjustment can be performed first, followed by image reconstruction, or image size adjustment can be performed first, followed by image reconstruction. Also, a combination of image segmentation, image reconstruction, and image size adjustment is possible. Image segmentation can be performed first, followed by image size adjustment and image reconstruction, and the order of image setting is not fixed but can be changed, which can be determined according to the encoding / decoding settings. In this example, the image setting process is described assuming that image segmentation is performed first, followed by image reconstruction, and then image size adjustment. However, other orders are possible according to the encoding / decoding settings, and changes to other cases are also possible.
[0395] For example, the image setting processes may be performed in the following order: division → reconstruction, reconstruction → division, division → resize, resize → division, resize → reconstruction, resize → division, resize → reconstruction, division → resize → reconstruction, resize → division → reconstruction, resize → reconstruct → division, resize → division → resize, resize → division, reconstruction → division → resize, resize → division, etc., and may also be combined with additional image settings. As described above, the image setting processes may be performed sequentially, but all or some of the setting processes may be performed simultaneously. Also, some image setting processes may involve multiple processes depending on the encoding / decoding settings (e.g., image characteristics, type, etc.). Examples of various combinations of image setting processes are shown below.
[0396] For example, P0 to P5 in FIG. 11a may correspond to S0 to S5 in FIG. 11b, and a reconfiguration process (in this example, pixel rearrangement) and a size adjustment process (in this example, the same size adjustment for each division unit) may be performed for each division unit. For example, P0 to P5 may be assigned to S0 to S5 after size adjustment using an offset. Also, P0 may be assigned to S0 without reconfiguration, P1 may be rotated 90 degrees and assigned to S1, P2 may be rotated 180 degrees and assigned to S2, P3 may be rotated 270 degrees and assigned to S3, P4 may be flipped left to right and assigned to S4, and P5 may be flipped upside down and assigned to S5.
[0397] For example, P0 to P5 in FIG. 11a may correspond to the same or different positions as S0 to S5 in FIG. 11b, and a reconfiguration process (in this example, rearrangement of pixels and division units) and a size adjustment process (in this example, the same size adjustment for each division unit) may be performed for each division unit. For example, P0 to P5 may be assigned to S0 to S5 after applying size adjustment using a scale. Also, P0 may be assigned to S0 without reconfiguration, P1 may be assigned to S2 without reconfiguration, P2 may be assigned to S1 after applying a 90-degree rotation, P3 may be assigned to S4 after applying a horizontal flip, P4 may be assigned to S5 after applying a 90-degree rotation and a horizontal flip, and P5 may be assigned to S3 after applying a 180-degree rotation and a horizontal flip.
[0398] 11a may correspond to E0 to E5 in FIG. 5e, and a reconfiguration process (in this example, rearrangement of pixels and division units) and a resizing process (in this example, resizing that is not the same for division units) may be performed for each division unit. For example, P0 may be assigned to E0 without resizing or reconfiguration, P1 may be assigned to E1 with resizing using a scale but without reconfiguration, P2 may be assigned to E2 without resizing and reconfiguration, P3 may be assigned to E4 with resizing using an offset but without reconfiguration, P4 may be assigned to E5 without resizing and reconfiguration, and P5 may be assigned to E3 with resizing using an offset and reconfiguration.
[0399] As in the above example, the absolute or relative positions of the division units within the image before and after the image setting process may be maintained or changed. This can be determined according to the encoding / decoding settings (e.g., image characteristics, type, etc.). Also, various combinations of image setting processes are possible, and the present invention is not limited to the above example, and various modifications are also possible.
[0400] The encoder records the information generated in the above process into a bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder parses the related information from the bitstream, and the information may be included in the bitstream in the form of SEI or metadata.
[0401] [Table 4]
[0402] Next, examples of syntax elements associated with multiple image settings will be described. In the examples described below, additional syntax elements will be mainly described. Furthermore, the syntax elements in the examples described below are not limited to a specific unit, but may be syntax elements supported in various units such as a sequence, a picture, a slice, a tile, etc. Alternatively, they may be syntax elements included in an SEI, metadata, etc.
[0403] Referring to Table 4, parts_enabled_flag is a syntax element that indicates whether or not a part of the image is to be divided. When activated (parts_enabled_flag = 1), it indicates that the image is divided into multiple units for encoding / decoding, and additional division information can be confirmed. When deactivated (parts_enabled_flag = 0), it indicates that the existing image is to be encoded / decoded. This example focuses on rectangular division units such as tiles, and different settings for existing tiles and division information can be configured.
[0404] num_partitions means a syntax element for the number of division units, and the value obtained by adding 1 to it means the number of division units.
[0405] part_top[i] and part_left[i] are syntax elements for position information of a division unit, and represent the horizontal and vertical start positions of the division unit (e.g., the position of the upper left corner of the division unit). part_width[i] and part_height[i] are syntax elements for size information of the division unit, and represent the horizontal and vertical widths of the division unit. In this case, the start position and size information can be set in pixel units or block units. In addition, the syntax elements may be syntax elements that can be generated in an image reconstruction process, or syntax elements that can be generated when an image division process and an image reconstruction process are mixed and configured.
[0406] The part_header_enabled_flag is a syntax element that indicates whether encoding / decoding settings are supported for each division unit. If it is activated (part_header_enabled_flag=1), the division unit can have its own encoding / decoding settings. If it is deactivated (part_header_enabled_flag=0), the division unit cannot have its own encoding / decoding settings and can be assigned the encoding / decoding settings of the upper unit.
[0407] The above example is one example of syntax elements associated with resizing and reconstructing in a division unit among image settings described below, and is not limited thereto, and modifications and application of other division units and settings of the present invention are possible. This example is described under the assumption that resizing and reconstructing are performed after division, but is not limited thereto, and modifications and application are possible according to other image setting orders, etc. In addition, the types of syntax elements, the order of syntax elements, conditions, etc. supported in the example described below are limited only to this example, and may be changed and determined according to encoding / decoding settings.
[0408] [Table 5]
[0409] Table 5 shows examples of syntax elements related to the reconstruction of segmentation units in an image setting.
[0410] Referring to Table 5, part_convert_flag[i] indicates a syntax element for determining whether a division unit is to be reconstructed. This syntax element can be generated for each division unit, and when activated (part_convert_flag[i]=1), it indicates that a reconstructed division unit is to be encoded / decoded, and additional reconstruction-related information can be confirmed. When deactivated (part_convert_flag[i]=0), it indicates that an existing division unit is to be encoded / decoded. convert_type_flag[i] indicates mode information for the reconstruction of a division unit, and may be information regarding pixel rearrangement.
[0411] In addition, syntax elements for additional reconfiguration such as rearrangement of division units may be generated. In this example, rearrangement of division units may be performed using the syntax elements part_top and part_left related to image division described above, or syntax elements related to rearrangement of division units (e.g., index information, etc.) may be generated.
[0412] [Table 6]
[0413] Table 6 shows examples of syntax elements related to adjusting the size of the division unit in the image setting.
[0414] Referring to Table 6, part_resizing_flag[i] is a syntax element indicating whether to resize an image for a division unit. This syntax element can occur for each division unit, and when activated (part_resizing_flag[i]=1), it indicates that the resized division unit is to be encoded / decoded and additional size-related information can be confirmed. When deactivated (part_resizing_flag[i]=0), it indicates that the existing division unit is to be encoded / decoded.
[0415] The width_scale[i] and height_scale[i] represent scale factors for horizontal and vertical size adjustment in size adjustment using a scale factor in a division unit.
[0416] top_height_offset[i] and bottom_height_offset[i] mean the upward and downward offset factors related to size adjustment using offset factors in the division unit, and left_width_offset[i] and right_width_offset[i] mean the left and right offset factors related to size adjustment using offset factors in the division unit.
[0417] Resizing_type_flag[i][j] indicates a syntax element for a data processing method for a region resized in units of division. The syntax element indicates an individual data processing method for a region resized in the direction of resizing. For example, a syntax element for an individual data processing method for a region resized in the up, down, left, or right direction can be generated. This can also be generated based on resizing information (e.g., it can only be generated when resizing in some directions).
[0418] The image setting process described above may be applied depending on the characteristics, type, etc. of the image. In the examples described below, the image setting process described above may be applied in the same manner or with modifications unless otherwise specified. In the examples described below, the description will focus on cases where the image setting process described above is applied in an additional or modified manner.
[0419] For example, images generated through a 360-degree camera (360-degree video or omnidirectional video) have different characteristics from images acquired through a general camera and require a different encoding environment than the compression of general images.
[0420] Unlike ordinary images, 360-degree images have no boundary areas with discontinuous characteristics, and data in all areas can be continuous. Furthermore, in devices such as HMDs, images are played in front of the eyes through lenses, which can require high-quality images, and when images are acquired through a stereoscopic camera, the amount of image data to be processed can increase. To provide an efficient encoding environment, including the above examples, various image setting processes can be performed taking 360-degree images into consideration.
[0421] The 360-degree camera is a camera with multiple cameras or multiple lenses and sensors, and the cameras or lenses can cover all directions around an arbitrary central point captured by the camera.
[0422] 360-degree images can be encoded using various methods. For example, they can be encoded using various image processing algorithms in three-dimensional space, or they can be converted into two-dimensional space and encoded using various image processing algorithms. This invention will mainly describe a method of converting a 360-degree image into two-dimensional space and encoding / decoding it.
[0423] A 360-degree image encoding device according to an embodiment of the present invention may be configured to include all or part of the components shown in Fig. 1 and may further include a pre-processing unit that performs pre-processing (stitching, projection, region-wise packing) on input images. Meanwhile, a 360-degree image decoding device according to an embodiment of the present invention may be configured to include all or part of the components shown in Fig. 2 and may further include a post-processing unit that performs post-processing (rendering) before decoding and reproducing as output images.
[0424] To explain it again, the encoder can perform pre-processing on the input image, then encode it, and transmit the corresponding bitstream, and the decoder can parse and decode the transmitted bitstream, then perform post-processing to generate the output image. At this time, the bitstream can contain and transmit information generated in the pre-processing process and information generated in the encoding process, and can be parsed by the decoder and used in the decoding process and post-processing process.
[0425] Next, the operation method of the 360-degree image encoder will be described in more detail. Since the operation method of the 360-degree image decoder is the inverse operation of the 360-degree image encoder, it can be easily derived by ordinary engineers, and therefore a detailed description will be omitted.
[0426] The input image can be stitched and projected into a 3D projection structure in units of spheres, through which the image data on the 3D projection structure can be projected into a 2D image.
[0427] A projected image can be configured to include all or part of the 360-degree content depending on the encoding settings. In this case, the position information of the region (or pixel) located at the center of the projected image can be implicitly generated as a predetermined value, or the position information can be explicitly generated. When a projected image is configured to include a portion of the 360-degree content, the range and position information of the included region can be generated. Range information (e.g., vertical and horizontal width) and position information (e.g., measured based on the upper left corner of the image) for a region of interest (ROI) in the projected image can be generated. In this case, a portion of the 360-degree content that is highly important can be set as the ROI. While a 360-degree image allows viewing of all content in the above, below, left, and right directions, a user's line of sight can be limited to a portion of the image, and the ROI can be set taking this into consideration. For efficient encoding, the ROI can be set to have high quality and resolution, while other regions can be set to have lower quality and resolution than the ROI.
[0428] Among 360-degree image transmission methods, the single stream transmission method can transmit a full image or a viewport image to a user using a single individual bitstream. The multi-stream transmission method transmits multiple full images with different image qualities using multiple bitstreams, allowing the user to select the image quality according to their environment and communication conditions. The tiled stream transmission method transmits individually coded partial images in tile units using multiple bitstreams, allowing the user to select tiles according to their environment and communication conditions. Therefore, a 360-degree image encoder generates and transmits bitstreams with two or more qualities, and a 360-degree image decoder can set a region of interest according to the user's gaze and selectively decode the image according to the region of interest. That is, a head tracking or eye tracking system can be used to set the area where the user's gaze is fixed as the region of interest and render only the required portion.
[0429] The projected image may be converted into a packed image through a region-wise packing process. The region-wise packing process may include dividing the projected image into multiple regions. Each divided region may then be arranged (or rearranged) in the packed image according to a region-wise packing setting. Region-wise packing may be performed to improve spatial continuity when converting a 360-degree image into a 2D image (or projected image). Region-wise packing may reduce the image size. It may also reduce image quality degradation during rendering, enable viewport-based projection, and provide other types of projection formats. Region-wise packing may or may not be performed depending on the encoding setting, and may be determined based on a signal indicating whether or not to perform it (e.g., regionwise_packing_flag; in the example described below, region-wise packing-related information can be generated only when regionwise_packing_flag is activated).
[0430] When regional packing is performed, setting information (or mapping information) that assigns (or arranges) a portion of the projected image as a portion of the packed image can be displayed (or generated). When regional packing is not performed, the projected image and the packed image can be the same image.
[0431] In the above, the stitching, projection, and regional packing processes are defined as separate processes, but some (e.g., stitching + projection, projection + regional packing) or all (e.g., stitching + projection + regional packing) of these processes can be defined as a single process.
[0432] At least one packed image can be generated for the same input image according to settings of the stitching, projection, and regional packing processes, etc. Also, at least one coded data can be generated for the same projected image according to settings of the regional packing process.
[0433] A packed image can be divided by performing a tiling process. Tiling is a process of dividing an image into multiple regions and transmitting them, and can be an example of the 360-degree image transmission method. As described above, tiling can be performed for the purpose of partial decoding taking into account the user's environment, etc., or for the purpose of efficiently processing a large amount of data of a 360-degree image. For example, if an image is composed of a single unit, the entire image can be decoded to decode the region of interest. However, if the image is composed of multiple unit regions, it is more efficient to decode only the region of interest. In this case, the division can be performed by dividing into tiles, which are division units in existing encoding methods, or by dividing into various division units (square divisions, blocks, etc.) described in the present invention. Furthermore, the division units can be units for performing independent encoding / decoding. Tiling can be performed based on the projected image or the packed image, or can be performed independently. That is, division can be performed based on the surface boundary of the projected image, the surface boundary of the packed image, packing settings, etc., and each division unit can be divided independently. This can affect the generation of partition information during the tiling process.
[0434] Next, the projected image or packed image can be encoded. The encoded data and information generated during the preprocessing process can be recorded in a bitstream and transmitted to a 360-degree image decoder. The information generated during the preprocessing process can be recorded in the bitstream in the form of SEI or metadata. At least one piece of encoded data and at least one piece of preprocessing information that change some settings of the encoding process or some settings of the preprocessing process can be recorded in the bitstream. This may be for the purpose of allowing the decoder to combine multiple pieces of encoded data (encoded data + preprocessing information) according to the user's environment to form a decoded image. In particular, multiple pieces of encoded data can be selectively combined to form a decoded image. Furthermore, for application in a stereoscopic system, the above process may be performed separately in two, or for an additional depth image.
[0435] FIG. 15 is an exemplary diagram showing a three-dimensional space representing a three-dimensional image and a two-dimensional plane space.
[0436] Generally, a 360-degree three-dimensional virtual space requires 3DoF (Degree of Freedom), which can support three rotations around the X (Pitch), Y (Yaw), and Z (Roll) axes. DoF refers to degrees of freedom in space, and 3DoF refers to degrees of freedom including rotations around the X, Y, and Z axes as in 15a, while 6DoF refers to degrees of freedom that further allow movement along the X, Y, and Z axes in addition to 3DoF. The image encoding and decoding apparatus of the present invention will be described mainly for the case of 3DoF, and when supporting more than 3DoF (3DoF+), the apparatus may be combined with or modified with additional processes or devices not shown in the present invention.
[0437] Referring to 15a, Yaw can range from -π (-180 degrees) to π (180 degrees), Pitch can range from -π / 2 rad (or -90 degrees) to π / 2 rad (or 90 degrees), and Roll can range from -π / 2 rad (or -90 degrees) to π / 2 rad (or 90 degrees). If we assume that ψ and θ are the longitude and latitude of the Earth's map representation, then (x, y, z) in the three-dimensional space can be transformed from (ψ, θ) in the two-dimensional space. For example, the three-dimensional coordinates can be derived from the two-dimensional coordinates based on the transformation formulas x = cos(θ) cos(ψ), y = sin(θ), z = -cos(θ) sin(ψ).
[0438] Also, (ψ, θ) can be transformed into (x, y, z). For example, ψ=tan -1 (-Z / X), θ=sin -1 (Y / (X 2 +Y 2 +Z 2 ) 1 / 2 ) based on the transformation formula, two-dimensional space coordinates can be derived from three-dimensional space coordinates.
[0439] When a pixel in 3D space is accurately converted to 2D space (e.g., an integer unit pixel in 2D space), the pixel in 3D space can be mapped to a pixel in 2D space. When a pixel in 3D space is not accurately converted to 2D space (e.g., a fractional unit pixel in 2D space), the 2D pixel can be mapped to the pixel obtained by interpolation. The interpolation method used can be nearest neighbor interpolation, bilinear interpolation, B-spline interpolation, bicubic interpolation, or the like. In this case, related information can be explicitly generated by selecting one of multiple interpolation candidates, or the interpolation method can be implicitly determined based on a predetermined rule. For example, a predetermined interpolation filter can be used depending on the 3D model, projection format, color format, slice / tile type, etc. When explicitly generating interpolation information, information about the filter (e.g., filter coefficients) can also be included.
[0440] 15b shows an example of transformation from 3D space to 2D space (2D plane coordinate system). (ψ, θ) can be sampled (i, j) based on the image size (width, height), where i can range from 0 to P_Width-1 and j can range from 0 to P_Height-1.
[0441] (ψ, θ) can be the center point (or reference point, the point marked C in FIG. 15, coordinates (ψ, θ) = (0, 0)) for positioning a 360-degree image at the center of the projected image. The setting for the center point can be specified in three-dimensional space, and position information for the center point can be explicitly generated or implicitly set to a previously set value. For example, center position information for Yaw, center position information for Pitch, center position information for Roll, etc. can be generated. If values for the information are not specifically specified, each value can be assumed to be 0.
[0442] In the above example, an example was described in which the entire 360-degree image was converted from three-dimensional space to two-dimensional space. However, a partial region of the 360-degree image can also be targeted, and position information (e.g., a partial position belonging to the region; in this example, position information relative to the center point), range information, etc. for the partial region can be explicitly generated, or position and range information that has already been implicitly set can be used. For example, center position information in Yaw, center position information in Pitch, center position information in Roll, range information in Yaw, range information in Pitch, range information in Roll, etc. can be generated. In the case of a partial region, at least one region is required, allowing position information and range information for multiple regions to be processed. If a value for the information is not specifically specified, it can be assumed to be the entire 360-degree image.
[0443] H0 to H6 and W0 to W5 in 15a respectively indicate latitude and longitude of a portion of 15b, and the coordinates of 15b can be expressed as (C, j) and (i, C) (C is the longitude or latitude component). Unlike ordinary images, 360-degree images can be distorted or warped when converted to two-dimensional space. This can vary depending on the region of the image, and different encoding / decoding settings can be set for different positions on the image or for regions partitioned according to the position. When adaptively setting encoding / decoding settings based on encoding / decoding information in the present invention, the position information (e.g., x and y components, or a range defined by x and y, etc.) can be included as an example of the encoding / decoding information.
[0444] The above description of three-dimensional and two-dimensional space is provided to assist in explaining the embodiments of the present invention, and is not intended to be limiting, and modifications of the details or applications in other cases are possible.
[0445] As mentioned above, images captured by a 360-degree camera can be converted into two-dimensional space. At this time, the 360-degree image can be mapped using a three-dimensional model, and various three-dimensional models such as a sphere, cube, cylinder, pyramid, and polyhedron can be used. When the 360-degree image mapped based on the model is converted into two-dimensional space, a projection process can be performed using a projection format based on the model.
[0446] 16a to 16d are conceptual diagrams for explaining a projection format according to an embodiment of the present invention.
[0447] FIG. 16a shows the ERP (Equi-Rectangular Projection) format, in which a 360-degree image is projected onto a two-dimensional plane. FIG. 16b shows the CMP CubeMap Projection format, in which a 360-degree image is projected onto a cube. FIG. 16c shows the OHP (OctaHedron Projection) format, in which a 360-degree image is projected onto an octahedron. FIG. 16d shows the ISP (IcoSahedral Projection) format, in which a 360-degree image is projected onto a polyhedron. However, various projection formats can be used without being limited to these. The left side of FIGS. 16a to 16d shows a 3D model, and the right side shows an example converted into two-dimensional space through a projection process. Depending on the projection format, various sizes and shapes are available, and each shape can be composed of faces, with faces being represented as circles, triangles, rectangles, etc.
[0448] In the present invention, a projection format can be defined by a 3D model, surface settings (e.g., number of surfaces, surface shape, surface shape configuration, etc.), projection process settings, etc. If at least one element of the definition is different, it can be considered a different projection format. For example, in the case of ERP, if it is composed of a spherical model (3D model), one surface (number of surfaces), and a square surface (surface pattern), but if some of the settings in the projection process (e.g., the mathematical formula used when converting from 3D space to 2D space; i.e., the element that causes a difference in at least one pixel of the projected image during the projection process) are different, it can be classified as a different format, such as ERP1 or EPR2. As another example, in the case of CMP, if it is composed of a cubic model, six surfaces, and square surfaces, but if some of the settings in the projection process (e.g., the sampling method when converting from 3D space to 2D) are different, it can be classified as a different format, such as CMP1 or CMP2.
[0449] When multiple projection formats are used instead of one pre-defined projection format, projection format identification information (or projection format information) can be explicitly generated. The projection format identification information can be configured in various ways.
[0450] As an example, index information (e.g., proj_format_flag) can be assigned to multiple projection formats to identify them. For example, ERP can be assigned number 0, CMP can be assigned number 1, OHP can be assigned number 2, ISP can be assigned number 3, ERP1 can be assigned number 4, CMP1 can be assigned number 5, OHP1 can be assigned number 6, ISP1 can be assigned number 7, CMP compact can be assigned number 8, OHP compact can be assigned number 9, ISP compact can be assigned number 10, and other formats can be assigned numbers 11 and above.
[0451] For example, the projection format can be identified from at least one element information constituting the projection format. The element information constituting the projection format may include 3D model information (e.g., 3d_model_flag, where 0 is a sphere, 1 is a cube, 2 is a cylinder, 3 is a pyramid, 4 is a polyhedron 1, and 5 is a polyhedron 2, etc.), surface number information (e.g., num_face_flag, which starts from 1 and increases by 1, or the number of surfaces generated in the projection format is assigned as index information, where 0 is 1, 1 is 3, 2 is 6, 3 is 8, and 4 is 20, etc.), surface shape information (e.g., shape_face_flag, where 0 is a rectangle, 1 is a circle, 2 is a triangle, 3 is a rectangle + circle, and 4 is a rectangle + triangle, etc.), and projection process setting information (e.g., 3d_2d_convert_idx, etc.).
[0452] As an example, a projection format can be identified by projection format index information and element information constituting the projection format. For example, the projection format index information can assign 0 to ERP, 1 to CMP, 2 to OHP, 3 to ISP, and 4 or more to other formats, and the projection format (e.g., ERP, ERP1, CMP, CMP1, OHP, OHP1, ISP, ISP1, etc.) can be identified together with the element information constituting the projection format (in this example, projection process setting information). Alternatively, the projection format (e.g., ERP, CMP, CMP compact, OHP, OHP compact, ISP, ISP compact, etc.) can be identified together with the element information constituting the projection format (in this example, whether it is regional packing or not).
[0453] In summary, a projection format can be identified by projection format index information, at least one projection format element information, or both projection format index information and at least one projection format element information. This can be defined according to the encoding / decoding settings, and the present invention will be described assuming that a projection format is identified by a projection format index. While this example focuses on a projection format represented by surfaces of the same size and shape, configurations in which the sizes and shapes of the surfaces are not identical are also possible. The configurations of the surfaces may be the same or different from those shown in FIGS. 16a to 16d, and the numbers on the surfaces are used as symbols to identify each surface and are not limited to a specific order. For convenience of explanation, in the examples described below, the projection format will be described assuming that the surfaces have the same size and shape, based on the projection image, for ERP: one surface + rectangle, CMP: six surfaces + rectangle, OHP: eight surfaces + triangle, and ISP: 20 surfaces + triangle. However, the same or similar application can be applied to other configurations.
[0454] As shown in Figures 16a to 16d, the projection format can be classified into one surface (e.g., ERP) or multiple surfaces (e.g., CMP, OHP, ISP, etc.). Furthermore, each surface can be classified into a rectangular shape, a triangular shape, etc. The above classification may be an example of the image type, characteristics, etc. in the present invention that can be applied when encoding / decoding settings are different depending on the projection format. For example, the image type may be a 360-degree image, and the image characteristics may be one of the above classifications (e.g., each projection format, a projection format with one surface or multiple surfaces, a projection format with a rectangular or non-rectangular surface, etc.).
[0455] A 2D plane coordinate system {e.g., (i, j)} can be defined for each surface of a 2D projection image, and the characteristics of the coordinate system may vary depending on the projection format, the position of each surface, etc. In the case of ERP, there is one 2D plane coordinate system, and other projection formats can have multiple 2D plane coordinate systems depending on the number of surfaces. In this case, the coordinate system can be expressed as (k, i, j), where k can be the index information of each surface.
[0456] FIG. 17 is a conceptual diagram illustrating a projection format realized within a rectangular image according to an embodiment of the present invention.
[0457] That is, 17a to 17c can be understood as the projection formats of Figs. 16b to 16d realized as rectangular images.
[0458] 17a to 17c, each image format can be configured in a rectangular shape for encoding / decoding a 360-degree image. In the case of ERP, one coordinate system can be used as is, but in the case of another projection format, the coordinate systems of each surface can be integrated into one coordinate system, and a detailed description thereof will be omitted.
[0459] 17a to 17c, it can be seen that in the process of constructing a rectangular image, areas filled with meaningless data such as blank spaces or backgrounds are generated. That is, the rectangular image can be composed of an area containing actual data (in this example, the surface; active area) and a meaningless area filled to construct the rectangular image (in this example, assumed to be filled with arbitrary pixel values; inactive area). This may cause a decrease in performance not only in the encoding / decoding of actual image data but also due to an increase in the amount of coded data caused by the increase in image size due to the meaningless area.
[0460] Therefore, further steps can be taken to eliminate meaningless regions and construct the image from regions containing actual data.
[0461] FIG. 18 is a conceptual diagram of a method for converting a projection format to a rectangular shape by rearranging surfaces to eliminate insignificant areas according to an embodiment of the present invention.
[0462] 18a to 18c show an example of rearrangement of 17a to 17c, and such a process can be defined as a regional packing process (e.g., CMP compact, OHP compact, ISP compact). In this case, not only can the surface itself be rearranged, but the surface can also be divided and rearranged (e.g., OHP compact, ISP compact). This can be performed not only to remove meaningless areas but also to improve coding performance through efficient surface arrangement. For example, when an arrangement is made in which images are contiguous between surfaces (e.g., B2-B3-B1, B5-B0-B4 in 18a), prediction accuracy during coding is improved, thereby improving coding performance. Regional packing according to a projection format is merely an example of the present invention and is not limited thereto.
[0463] FIG. 19 is a conceptual diagram illustrating a process of packing by region into a rectangular image in a CMP projection format according to an embodiment of the present invention.
[0464] Referring to 19a to 19c, the CMP projection format can be arranged as 6x1, 3x2, 2x3, or 1x6. Furthermore, when size adjustment is performed on some surfaces, the CMP can be arranged as shown in 19d to 19e. While 19a to 19e use CMP as an example, the present invention is not limited to CMP and can be applied to other projection formats. The surface arrangement of the images obtained through the regional packing may follow a predetermined rule depending on the projection format, or information regarding the arrangement may be explicitly generated.
[0465] A 360-degree image encoding / decoding device according to an embodiment of the present invention may be configured to include all or part of the image encoding / decoding device shown in Figures 1 and 2. In particular, a format converter and a format inverse converter for converting and inversely converting a projection format may be further included in the image encoding device and the image decoding device, respectively. That is, in the image encoding device of Figure 1, an input image can be encoded through a format converter, and in the image decoding device of Figure 2, a bitstream is decoded and then an output image can be generated through a format inverse converter. Hereinafter, the encoder (in this example, "input image" to "encoding") for the above process will be mainly described, and the process in the decoder can be reversely derived from the encoder. Also, a description that overlaps with the above content will be omitted.
[0466] Next, the following description is based on the premise that the input image is the same as the 2D projected image or packed image obtained by performing a preprocessing process in the 360-degree encoding device described above. That is, the input image may be an image obtained by performing a projection process using a certain projection format or a regional packing process. The projection format already applied to the input image may be any of various projection formats and may be considered a common format or may be referred to as a first format.
[0467] The format conversion unit can convert the image data into a projection format other than the first format. In this case, the format to be converted can be referred to as the second format. For example, ERP can be set as the first format and converted into a second format (e.g., ERP2, CMP, OHP, ISP, etc.). In this case, ERP2 can be an EPR format that has the same conditions, such as the 3D model and surface configuration, but with some different settings. Alternatively, the two formats can be the same format (e.g., ERP=ERP2) with the same projection format settings, but with different image sizes or resolutions. Alternatively, a portion of the image setting process described below can be applied. While the above examples are given for convenience of explanation, the first and second formats are merely examples of various projection formats, and are not limited to the above examples and can be changed to other formats.
[0468] During the conversion process between formats, due to the different characteristics of the coordinate systems between projection formats, the pixels (integer pixels) of the converted image may be obtained not only from the integer unit pixels of the image before conversion but also from fractional unit pixels, so interpolation is possible. In this case, the interpolation filter used may be the same or similar to the filters described above. The interpolation filter is selected from multiple interpolation filter candidates, and related information can be generated explicitly or can be implicitly determined by pre-defined rules. For example, a predetermined interpolation filter can be used depending on the projection format, color format, slice / tile type, etc. Furthermore, when an interpolation filter is explicitly sent, information about the filter (e.g., filter coefficients) may also be included.
[0469] The projection format in the format converter may be defined to include regional packing, etc. That is, projection and regional packing processes may be performed during the format conversion process, or regional packing processes may be performed after the format conversion and before encoding.
[0470] The encoder records the information generated in the above process into a bitstream in units of at least one of a sequence, a picture, a slice, a tile, etc., and the decoder parses the related information from the bitstream, and the information may be included in the bitstream in the form of SEI or metadata.
[0471] Next, an image setting process applied to a 360-degree image encoding / decoding device according to an embodiment of the present invention will be described. The image setting process of the present invention can be applied not only to general encoding / decoding processes but also to pre-processing processes, post-processing processes, format conversion processes, and format inverse conversion processes in a 360-degree image encoding / decoding device. The image setting process described below will be described mainly with respect to a 360-degree image encoding device, and can be described including the content of the image setting process described above. A redundant description of the image setting process described above will be omitted. Furthermore, the examples described below will be mainly with respect to the image setting process, and the image setting inverse process can be derived from the image setting process in reverse, and in some cases can be confirmed through various embodiments of the present invention described above.
[0472] The image setting process in the present invention may be performed at the stage of 360-degree image projection, at the stage of regional packing, at the stage of format conversion, or at other stages.
[0473] 20 is a conceptual diagram of 360-degree image division according to an embodiment of the present invention, which will be explained assuming the case of an image projected by ERP.
[0474] 20a shows an image projected by the ERP, which can be segmented using various methods. In this example, slices and tiles are mainly described, and W0-W2 and H0 and H1 are assumed to be the dividing boundaries of the slices or tiles, and are assumed to follow the raster scan order. The examples below will mainly describe slices and tiles, but are not limited to this, and other segmentation methods can be applied.
[0475] For example, division can be performed in slice units, resulting in division boundaries of H0 and H1, or division can be performed in tile units, resulting in division boundaries of W0 to W2 and H0 and H1.
[0476] 20b shows an example of an image projected by the ERP divided into tiles (assuming the same tile division boundaries (W0-W2, H0, and H1) as in FIG. 20a). Assuming that the P region is the entire image and the V region is the region where the user's gaze is fixed, or the viewport, various methods can be used to provide an image corresponding to the viewport. For example, the entire image (e.g., tiles a-l) can be decoded to obtain the region corresponding to the viewport. In this case, the entire image can be decoded, and if the image is divided, tiles a-l (in this example, regions A+B) can be decoded. Alternatively, the region corresponding to the viewport can be obtained by decoding the region belonging to the viewport. In this case, if the image is divided, tiles f, g, j, and k (in this example, region B) can be decoded to obtain the region corresponding to the viewport from the restored image. The former case can be called full decoding (or viewport-independent coding), and the latter case can be called partial decoding (or viewport-dependent coding). The latter case is an example that can occur in a 360-degree image with a large amount of data, and since it allows for flexible acquisition of division regions, a tile-based division method is more commonly used than a slice-based division method. In the case of partial decoding, since it is not known where the viewport originates, the referability of the division unit can be spatially or temporally limited (implicitly processed in this example), and encoding / decoding can be performed taking this into consideration. The following examples will focus on the case of full decoding, but in order to prepare for the case of partial decoding, division of a 360-degree image will be described centering on tiles (or the rectangular division method of the present invention). The contents of the following examples can be applied to other division units in the same way or with modifications.
[0477] 21 is an exemplary diagram of 360-degree image division and image reconstruction according to an embodiment of the present invention, which will be described assuming the case of an image projected by CMP.
[0478] Figure 21a shows the image projected by CMP, and segmentation can be performed using various methods. W0-W2 and H0, H1 are assumed to be the segmentation boundaries of the surface, slice, and tile, and are assumed to follow the raster scan order.
[0479] For example, division can be performed in slice units, resulting in division boundaries of H0 and H1. Or, division can be performed in tile units, resulting in division boundaries of W0 to W2 and H0 and H1. Or, division can be performed in surface units, resulting in division boundaries of W0 to W2 and H0 and H1. In this example, the description will be made assuming that the surface is a part of the division unit.
[0480] In this case, the surface is a division unit (dependent encoding / decoding in this example) created to classify or separate regions having different properties (e.g., the plane coordinate system of each surface) within the same image according to the image characteristics and type (in this example, a 360-degree image, a projection format), and the like. The slice and tile may be division units (independent encoding / decoding in this example) created to divide an image according to a user's definition. The surface is a unit divided according to a predetermined definition (or derived from projection format information) during the projection process according to the projection format, and the slice and tile may be units divided by explicitly generating division information according to a user's definition. The surface may have a polygonal division shape, including a rectangle, according to the projection format, the slice may have any division shape that cannot be defined as a rectangle or a polygon, and the tile may have a square division shape. The division unit settings may be limited and defined for the purposes of describing this example.
[0481] In the above example, the surface has been described as a division unit classified for the purpose of region division, but depending on the encoding / decoding setting, it may be a unit for performing independent encoding / decoding on at least one surface unit, or may have a setting for performing independent encoding / decoding in combination with tiles, slices, etc. In this case, there may be cases where explicit information of tiles and slices is generated in combination with tiles, slices, etc., or there may be cases where tiles and slices are implicitly combined based on surface information. There may also be cases where explicit information of tiles and slices is generated based on surface information.
[0482] As a first example, one image division process (in this example, the surface) is performed, and the image division can implicitly omit division information (divide information is obtained from projection format information). This example is an example of dependent encoding / decoding settings, and may be an example corresponding to a case where referenceability between surface units is not restricted.
[0483] As a second example, an image segmentation process (in this example, the surface) is performed, and the image segmentation can explicitly generate segmentation information. This example is an example of a dependent encoding / decoding setting, and may be an example corresponding to a case where referenceability between surface units is not restricted.
[0484] As a third example, multiple image segmentation processes (in this example, surfaces, tiles) are performed, and some image segmentations (in this example, surfaces) can implicitly omit or explicitly generate segmentation information, and some image segmentations (in this example, tiles) can explicitly generate segmentation information. In this example, some image segmentation processes (in this example, surfaces) precede some image segmentation processes (in this example, tiles).
[0485] In a fourth example, multiple image segmentation processes are performed, and some image segmentations (in this example, surfaces) can implicitly omit or explicitly generate segmentation information, and some image segmentations (in this example, tiles) can explicitly generate segmentation information based on some image segmentations (in this example, surfaces). In this example, some image segmentation processes (in this example, surfaces) precede some image segmentation processes (in this example, tiles). This example is the same as in the second example, where segmentation information is explicitly generated, but there may be differences in the configuration of the segmentation information.
[0486] As a fifth example, multiple image division processes are performed, and some image divisions (surfaces in this example) can implicitly omit division information, while some image divisions (tiles in this example) can implicitly omit division information based on some image divisions (surfaces in this example). For example, individual surface units can be set as tiles, or multiple surface units (in this example, adjacent surfaces are grouped if they are continuous surfaces and not grouped otherwise; B2-B3-B1 and B4-B0-B5 in 18a) can be set as tiles. Surface units can be set as tiles according to pre-defined rules. This example is an example of independent encoding / decoding settings, and may correspond to a case where reference between surface units is limited. That is, in some cases (assuming the first example), division information is implicitly processed, but there may be differences in encoding / decoding settings.
[0487] The above example is a description of a case where the division process can be performed in the projection step, the regional packing step, the initial encoding / decoding step, etc., but the image division process can also occur within other encoders / decoders.
[0488] In 21a, a rectangular image can be constructed by including an area A containing data and an area B not containing data. In this case, the position, size, shape, number, etc. of areas A and B are information that can be confirmed by the projection format, etc., or information that can be confirmed when information on the projected image is explicitly generated, and related information can be indicated by the image division information, image reconstruction information, etc., described above. For example, as shown in Tables 4 and 5, information on a portion of the projected image (e.g., part_top, part_left, part_width, part_height, part_convert_flag, etc.) can be indicated. This is not limited to this example and may be an example that can be applied to other cases (e.g., other projection formats, other projection settings, etc.).
[0489] Region B can be configured together with region A into a single image and then encoded / decoded. Alternatively, different encoding / decoding settings can be set by dividing the image taking into account the characteristics of each region. For example, encoding / decoding for region B can be omitted based on information on whether encoding / decoding is performed (e.g., tile_coded_flag if the division unit is a tile). In this case, the corresponding region can be restored to certain data (in this example, any pixel value) according to a pre-set rule. Alternatively, the encoding / decoding settings for region B during the image division process can be set differently from those for region A. Alternatively, the corresponding region can be removed by performing a regional packing process.
[0490] 21b shows an example of dividing a packed image into tiles, slices, and surfaces using CMP. In this case, the packed image is an image that has undergone a surface rearrangement process or a regional packing process, and may be an image obtained by performing image segmentation and image reconstruction according to the present invention.
[0491] In 21b, a rectangular shape can be formed by including regions containing data. In this case, the position, size, shape, number, etc. of each region can be confirmed by pre-configured settings or can be confirmed when information on the packed image is explicitly generated, and related information can be indicated by the image division information, image reconstruction information, etc. For example, as shown in Tables 4 and 5, information on some regions of the packed image (e.g., part_top, part_left, part_width, part_height, part_convert_flag, etc.) can be indicated.
[0492] The packed image can be segmented using various segmentation methods, for example, by slice, with a segmentation boundary of H0, or by tile, with a segmentation boundary of W0, W1, and H0, or by surface, with a segmentation boundary of W0, W1, and H0.
[0493] The image segmentation and image reconstruction processes of the present invention can be performed on a projected image. In this case, the reconstruction process can rearrange not only pixels within a surface but also surfaces within the image. This can be a possible example when an image is segmented or constructed into multiple surfaces. The following example will focus on the case where the image is segmented into tiles based on surface units.
[0494] SX,Y (S0,0 to S3,2) in 21a can correspond to S'U,V (S'0,0 to S'2,1. In this example, X and Y may be the same as or different from U and V), and a reconstruction process can be performed on a surface-by-surface basis. For example, S2,1, S3,1, S0,1, S1,2, S1,1, and S1,0 can be assigned (or surface relocated) to S'0,0, S'1,0, S'2,0, S'0,1, S'1,1, and S'2,1. Furthermore, S2,1, S3,1, and S0,1 can be reconstructed (or pixel relocated) without reconstructing, while S1,2, S1,1, and S1,0 can be reconstructed by applying a 90-degree rotation, as shown in FIG. 21c. The symbols displayed horizontally in 21c (S1,0, S1,1, S1,2) may be images laid horizontally to match the symbols in order to maintain image continuity.
[0495] The surface reconstruction can be implicit or explicit depending on the encoding / decoding settings: if implicit, it can be done according to predefined rules that take into account the type of image (in this example, a 360-degree image) and its characteristics (in this example, the projection format, etc.).
[0496] For example, in 21c, S'0,0 and S'1,0, S'1,0 and S'2,0, S'0,1 and S'1,1, and S'1,1 and S'2,1 have image continuity (or correlation) between the two surfaces based on the surface boundaries, and 21c may be an example configured such that there is continuity between the upper three surfaces and the lower three surfaces. The surface is divided into multiple surfaces through a projection process from 3D space to 2D space, and then reconstructed through a regional packing process to enhance image continuity between the surfaces for efficient surface reconstruction. Such surface reconstruction can be pre-configured and processed.
[0497] Alternatively, the reconstruction process can be performed by explicit processing and reconstruction information generated for it.
[0498] For example, when information (e.g., either implicitly acquired information or explicitly generated information) for an M×N configuration (e.g., 6×1, 3×2, 2×3, 1×6, etc. in the case of CMP compact; in this example, a 3×2 configuration is assumed) is confirmed through a regional packing process, the surface can be reconstructed to match the M×N configuration, and then information for the M×N configuration can be generated. For example, in the case of surface relocation within an image, index information (or position information within an image) can be assigned to each surface, and in the case of pixel relocation within a surface, mode information for the reconstruction can be assigned.
[0499] The index information can be already defined as shown in 18a to 18c of Figure 18, and SX, Y or S'U, V in 21a to 21c can represent each surface as position information indicating the horizontal and vertical directions (e.g., S[i][j]) or one position information (e.g., assuming that the position information is assigned in raster scan order from the upper left surface of the image, S[i]), to which an index of each surface can be assigned.
[0500] For example, when assigning indexes to position information indicating width and height, in the case of Figure 21c, S'0,0 can be assigned the index of surface 2, S'1,0 can be assigned the index of surface 3, S'2,0 can be assigned the index of surface 1, S'0,1 can be assigned the index of surface 5, S'1,1 can be assigned the index of surface 0, and S'2,1 can be assigned the index of surface 4. Alternatively, when assigning indexes to one piece of position information, S[0] can be assigned the index of surface 2, S[1] can be assigned the index of surface 3, S[2] can be assigned the index of surface 1, S[3] can be assigned the index of surface 5, S[4] can be assigned the index of surface 0, and S[5] can be assigned the index of surface 4. For convenience of explanation, in the examples described below, S'0,0 to S'2,1 will be referred to as a to f. Alternatively, it can be expressed as position information indicating width and height in pixel or block units based on the upper left corner of the image.
[0501] In the case of packed images acquired through an image reconstruction process (or a regional packing process), the surface scanning order may or may not be the same for the images depending on the reconstruction settings. For example, if a single scanning order (e.g., raster scanning) is applied to 21a, the scanning order of a, b, and c may be the same, but the scanning order of d, e, and f may not be the same. For example, for 21a, a, b, and c, the scanning order may follow the order (0,0) → (1,0) → (0,1) → (1,1), while for d, e, and f, the scanning order may follow the order (1,0) → (1,1) → (0,0) → (0,1). This can be determined depending on the image reconstruction settings, and such settings can also be used for different projection formats.
[0502] The image division process in 21b can set individual surface units to tiles. For example, surfaces a to f can each be set to a tile unit. Or, multiple surface units can be set to tiles. For example, surfaces a to c can be set to one tile, and surfaces d to f can be set to one tile. The configuration can be determined based on surface characteristics (e.g., continuity between surfaces), and surface tiling different from the above example is possible.
[0503] Next, an example of division information obtained by a plurality of image division processes will be described. In this example, the division information for the surface is omitted, and the units other than the surface are tiles, and the division information is processed in various ways.
[0504] As a first example, image segmentation information can be implicitly obtained based on surface information. For example, individual surfaces can be set to tiles, or multiple surfaces can be set to tiles. In this case, whether at least one surface is set to a tile can be determined by a predetermined rule based on surface information (e.g., continuity or correlation).
[0505] As a second example, image partitioning information can be explicitly generated regardless of surface information. For example, when partitioning information is generated based on the number of tile rows (in this example, num_tile_columns) and the number of tile columns (in this example, num_tile_rows), the partitioning information can be generated using the method of the image partitioning process described above. For example, the range of the number of tile rows and the number of tile columns can be from 0 to the image width / block width (in this example, the unit obtained from the picture partitioning unit), or from 0 to the image height / block height. In addition, additional partitioning information (e.g., uniform_spacing_flag, etc.) can be generated. In this case, depending on the partitioning setting, the boundary of the surface and the boundary of the partitioning unit may or may not coincide with each other.
[0506] As a third example, image division information can be explicitly generated based on surface information. For example, when division information is generated based on the number of rows and columns of tiles, the division information can be generated based on surface information (in this example, the range of the number of rows is 0 to 2, and the range of the number of columns is 0, 1, because the surface configuration in the image is 3x2). For example, the ranges that the number of rows and columns of tiles can have are 0 to 2 and 0 to 1, respectively. In addition, additional division information (e.g., uniform_spacing_flag, etc.) may not be generated. In this case, the boundary of the surface and the boundary of the division unit may coincide.
[0507] In some cases (assuming the cases of the second and third examples), the syntax elements of the division information are defined differently, or even if the same syntax elements are used, the settings of the syntax elements (for example, binarization settings, etc. If the range of candidates that a syntax element has is limited and narrow, other binarizations can be used, etc.) can be different. The above examples have described some of the various configurations of division information, but are not limited to these and can be understood as examples in which different settings are possible depending on whether the division information is generated based on surface information.
[0508] FIG. 22 is an example diagram showing an image projected or packed by CMP divided into tiles.
[0509] In this case, it is assumed that the tile division boundaries are the same as those in 21a of FIG. 21 (W0 to W2, H0, and H1 are all active), and that the tile division boundaries are the same as those in 21b of FIG. 21 (W0, W1, and H0 are all active). Assuming that the P region is the entire image and the V region is the viewport, full decoding or partial decoding can be performed. This example focuses on partial decoding. In 22a, tiles e, f, and g are decoded in the case of CMP (left), and tiles a, c, and e are decoded in the case of CMP compact (right), to obtain the area corresponding to the viewport. In 22b, tiles b, f, and i are decoded in the case of CMP, and tiles d, e, and f are decoded in the case of CMP compact, to obtain the area corresponding to the viewport.
[0510] In the above example, we have explained the case where division into slices, tiles, etc. is performed based on surface units (or surface boundaries), but it is also possible to perform division within the surface (for example, ERP has an image composed of one surface, while other projection formats have multiple surfaces), or to perform division including the surface boundaries, as shown in 20a of Figure 20.
[0511] 23 is a conceptual diagram for explaining an example of size adjustment of a 360-degree image according to an embodiment of the present invention. Here, the explanation will be given assuming the case of an image projected by ERP. Furthermore, the example described below will mainly focus on the case of expansion.
[0512] Depending on the image resizing type, the projected image can be resized using a scale factor or an offset factor, where the image before resizing is P_Width x P_Height and the image after resizing can be P'_Width x P'_Height.
[0513] In the case of a scale factor, after adjusting the size using the scale factors for the image's width and height (in this example, width a, height b), the image's width (P_Width x a) and height (P_Height x b) can be obtained. In the case of an offset factor, after adjusting the size using the offset factors for the image's width and height (in this example, width L, R, height T, B), the image's width (P_Width + L + R) and height (P_Height + T + B) can be obtained. Size adjustment can be performed using a pre-set method, or one of multiple methods can be selected.
[0514] In the examples below, the data processing method will be mainly explained for the case of an offset factor. In the case of an offset factor, the data processing method may include filling using a predetermined pixel value, filling by copying outer pixels, filling by copying a partial area of an image, or filling by transforming a partial area of an image.
[0515] In the case of 360-degree images, size adjustment can be performed taking into account the characteristic of continuity at the image boundaries. In the case of ERP, there is no outer boundary in 3D space, but when converted to 2D space through a projection process, an outer boundary area can exist. Data in the boundary area has continuous data outside the boundary, but can have a boundary due to spatial characteristics. Size adjustment can be performed taking these characteristics into account. In this case, continuity can be confirmed depending on the projection format, etc. For example, in the case of ERP, an image may have continuous boundaries at both ends. This example will be explained assuming that the left and right boundaries of the image are continuous, and the top and bottom boundaries of the image are continuous. Data processing methods will be explained, focusing on methods of copying and filling a portion of the image and methods of converting and filling a portion of the image.
[0516] When resizing to the left of the image, the resized area (in this example, LC or TL+LC+BL) can be filled with data from the right area of the image (in this example, tr+rc+br) that is contiguous with the left side of the image. When resizing to the right of the image, the resized area (in this example, RC or TR+RC+BR) can be filled with data from the left area of the image (in this example, tl+lc+bl) that is contiguous with the right side. When resizing to the top of the image, the resized area (in this example, TC or TL+TC+TR) can be filled with data from the bottom area of the image (in this example, bl+bc+br) that is contiguous with the top. When resizing to the bottom of the image, the resized area (in this example, BC or BL+BC+BR) can be filled with data.
[0517] If the size or length of the area to be resized is m, the area to be resized can have a range of (-m, y) to (-1, y) (resize to the left) or a range of (P_Width, y) to (P_Width+m-1, y) (resize to the right) based on the coordinates of the image before resizing (in this example, x is 0 to P_Width-1). The position x' of the area to obtain the data of the area to be resized can be derived using the formula x' = (x + P_Width) % P_Width. Here, x means the coordinate of the area to be resized based on the image coordinates before resizing, and x' means the coordinate of the area referenced by the area to be resized based on the image coordinates before resizing. For example, if you adjust the size to the left, m is 4, and the image width is 16, then (-4,y) will get the corresponding data from (12,y), (-3,y) will get the corresponding data from (13,y), (-2,y) will get the corresponding data from (14,y), and (-1,y) will get the corresponding data from (15,y). Or, if you adjust the size to the right, m is 4, and the image width is 16, then (16,y) will get the corresponding data from (0,y), (17,y) will get the corresponding data from (1,y), (18,y) will get the corresponding data from (2,y), and (19,y) will get the corresponding data from (3,y).
[0518] If the size or length of the region to be resized is n, the region to be resized can have a range of (x, -n) to (x, -1) (resize upward) or a range of (x, P_Height) to (x, P_Height + n-1) (resize downward) based on the coordinates of the image before resizing (in this example, y is 0 to P_Height - 1). The position y' of the region for obtaining the data of the region to be resized can be derived using a formula such as y' = (y + P_Height) % P_Height. Here, y means the coordinate of the region to be resized based on the image coordinates before resizing, and y' means the coordinate of the region referenced by the region to be resized based on the image coordinates before resizing. For example, if resizing is performed upwards, n is 4, and the image height is 16, then (x,-4) can obtain data from (x,12), (x,-3) from (x,13), (x,-2) from (x,14), and (x,-1) from (x,15). Or, if resizing is performed downwards, n is 4, and the image height is 16, then (x,16) can obtain data from (x,0), (x,17) from (x,1), (x,18) from (x,2), and (x,19) from (x,3).
[0519] After filling the data in the resizing area, the image coordinates after resizing can be adjusted based on the reference (in this example, x is 0 to P'_Width-1, y is 0 to P'_Height-1). The above example can be applied to a latitude and longitude coordinate system.
[0520] You can have various size adjustment combinations:
[0521] As an example, the image may be resized by m to the left, or by n to the right, or by o to the top, or by p to the bottom.
[0522] As an example, the image may be resized by m to the left and n to the right, or the image may be resized by o to the top and p to the bottom.
[0523] As an example, the image may be resized by m to the left, n to the right, and o to the top. Or the image may be resized by m to the left, n to the right, and p to the bottom. Or the image may be resized by m to the left, o to the top, and p to the bottom. Or the image may be resized by n to the right, o to the top, and p to the bottom.
[0524] As an example, the image may be resized by m to the left, n to the right, o to the top, and p to the bottom.
[0525] As in the above example, at least one size adjustment is performed, and the size adjustment of the image may be performed implicitly depending on the encoding / decoding settings, or size adjustment information may be explicitly generated and the size adjustment of the image may be performed based on the size adjustment information. That is, m, n, o, and p in the above example may be determined to predetermined values, or may be explicitly generated as size adjustment information, or some may be determined to predetermined values and some may be explicitly generated.
[0526] Although the above example has been described focusing on the case where data is obtained from a partial region of an image, other methods are also applicable. The data may be pixels before encoding or pixels after encoding, and may be determined according to the characteristics of the image or stage where resizing is performed. For example, when resizing is performed in a pre-processing process or a pre-encoding process, the data may refer to input pixels such as a projected image or a packed image. When resizing is performed in a post-processing process, an intra-frame prediction reference pixel generation process, a reference image generation process, a filtering process, or the like, the data may refer to reconstructed pixels. In addition, resizing may be performed using a data processing method for each region to be resized.
[0527] FIG. 24 is a conceptual diagram illustrating continuity between surfaces in a projection format (eg, CMP, OHP, ISP) according to an embodiment of the present invention.
[0528] In particular, this may be an example of an image consisting of multiple surfaces. Continuity is a characteristic that occurs in adjacent regions in three-dimensional space, and when Figures 24a to 24c are converted into two-dimensional space through a projection process, they can be classified into (A) where they are spatially adjacent and continuity exists, (B) where they are spatially adjacent but continuity does not exist, (C) where they are not spatially adjacent but continuity exists, and (D) where they are not spatially adjacent but continuity does not exist. This is different from general images, which are classified into (A) where they are spatially adjacent and continuity exists, and (D) where they are not spatially adjacent and continuity does not exist. In this case, the presence of continuity corresponds to some of the examples (A or C) above.
[0529] That is, referring to 24a to 24c, when there is spatial continuity and adjacent regions (in this example, the description is based on 24a), they are represented as b0 to b4, and when there is spatial continuity and adjacent regions, they are represented as B0 to B6. That is, this refers to the case of adjacent regions in a three-dimensional space, and by using b0 to b4 and B0 to B6 in the encoding process based on their continuity, encoding performance can be improved.
[0530] FIG. 25 is a conceptual diagram for explaining the continuity of the surface of FIG. 21c, which is an image acquired through an image reconstruction process or a regional packing process in a CMP projection format.
[0531] Here, 21c in Figure 21 is a rearrangement of 21a, which is a 360-degree image expanded into a cube, so the continuity of the surfaces in 21a in Figure 21 is maintained. That is, as in 25a, surface S2,1 is contiguous with S1,1 and S3,1 to the left and right, and is contiguous with surfaces S1,0 rotated 90 degrees and S1,2 rotated -90 degrees to the top and bottom.
[0532] In a similar manner, continuity for surfaces S3,1, S0,1, S1,2, S1,1 and S1,0 can be confirmed from 25b to 25f.
[0533] The continuity between surfaces can be defined according to the projection format setting, etc., and is not limited to the above example, and other modified examples are possible. The examples described below will be explained under the assumption that the continuity shown in Figures 24 and 25 exists.
[0534] FIG. 26 is an exemplary diagram illustrating image size adjustment in the CMP projection format according to an embodiment of the present invention.
[0535] 26a shows an example of adjusting the size of an image, 26b shows an example of adjusting the size in surface units (or division units), and 26c shows an example of adjusting the size (or multiple size adjustments) in image and surface units.
[0536] The projected image can be resized using a scale factor or an offset factor depending on the image size adjustment type. The image before size adjustment is P_Width × P_Height, the image after size adjustment is P'_Width × P'_Height, and the surface size can be F_Width × F_Height. Surfaces may be the same or different sizes, and the width and height of surfaces may be the same or different. However, for convenience of explanation, this example will be described assuming that all surfaces in the image are the same size and have a square shape. Also, the size adjustment values (WX, HY in this example) are the same. The data processing methods in the examples described below will focus on the offset factor, and will focus on methods of copying and filling a partial area of the image, and methods of converting and filling a partial area of the image. The above settings can also be applied to Figure 27.
[0537] In the cases of 26a to 26c, the boundary of a surface (assumed to have continuity according to 24a in FIG. 24 in this example) can have continuity with the boundary of another surface. In this case, they can be classified into a case where they are spatially adjacent on a two-dimensional plane and have image continuity (first example) and a case where they are not spatially adjacent on a two-dimensional plane and have image continuity (second example).
[0538] For example, assuming the continuity of 24a in Figure 24, the upper, left, right and lower regions of S1,1 are spatially adjacent to the lower, right, left and upper regions of S1,0, S0,1, S2,1 and S1,2, and the images may also be continuous (in the first example case).
[0539] Alternatively, the left and right regions of S1,0 may not be spatially adjacent to the upper regions of S0,1 and S2,1, but their images may be continuous (as in the second example). Also, the left and right regions of S0,1 and S3,1 may not be spatially adjacent to each other, but their images may be continuous (as in the second example). Also, the left and right regions of S1,2 may be continuous to the lower regions of S0,1 and S2,1 (as in the second example). This is a limited example, and configurations may differ depending on the definition and settings of the projection format. For ease of explanation, S0,0 to S3,2 in FIG. 26a will be referred to as a to l.
[0540] 26a may be an example of filling using data from areas where continuity exists toward the outer boundary of the image. Areas resized from area A where no data exists (in this example, a0-a2, c0, d0-d2, i0-i2, k0, l0-l2) can be filled using predetermined arbitrary values or outer pixel padding, and areas resized from area B containing actual data (in this example, b0, e0, h0, j0) can be filled using data from areas (or surfaces) where image continuity exists. For example, b0 can be filled using data from the upper side of surface h, e0 can be filled using data from the right side of surface h, h0 can be filled using data from the left side of surface e, and j0 can be filled using data from the lower side of surface h.
[0541] In detail, b0 may be an example of filling using the lower surface data obtained by applying a 180-degree rotation to surface h, and j0 may be an example of filling using the upper surface data obtained by applying a 180-degree rotation to surface h, but in this example (and in the examples described below), only the position of the referenced surface is indicated, and the data obtained in the area to be sized can be obtained after an adjustment process (e.g., rotation) that takes into account the continuity between surfaces, as shown in Figures 24 and 25.
[0542] 26b may be an example of filling using data from an area where continuity exists along the internal boundary of the image. In this example, the size adjustment operations performed along the surfaces may differ. Area A may undergo a shrinking process, while area B may undergo an expanding process. For example, surface a may undergo a size adjustment to the right by w0 (in this example, shrinking), and surface b may undergo a size adjustment to the left by w0 (in this example, expanding). Alternatively, surface a may undergo a size adjustment to the bottom by h0 (in this example, shrinking), and surface e may undergo a size adjustment to the top by h0 (in this example, expanding). In this example, looking at the change in the width of the image from surfaces a, b, c, and d, surface a shrinks by w0, surface b expands by w0 and w1, and surface c shrinks by w1. Therefore, the width of the image before and after size adjustment is the same. Looking at the change in the vertical width of the image from surfaces a, e, and i, surface a is reduced by h0, surface e is expanded by h0 and h1, and surface i is reduced by h1, so the vertical width of the image before and after size adjustment is the same.
[0543] The areas to be resized (in this example, b0, e0, be, b1, bg, g0, h0, e1, ej, j0, gi, g1, j1, h1) can either be simply removed, considering that they are shrinking from area A where no data exists, or they can be newly filled with data from areas where continuity exists, considering that they are expanding from area B which contains actual data.
[0544] For example, b0 can be filled using the data for the upper side of surface e, e0 can be filled using the data for the left side of surface b, be can be filled using the data for the left side of surface b, or the upper side of surface e, or the weighted sum of the left side of surface b and the upper side of surface e, b1 can be filled using the data for the upper side of surface g, bg can be filled using the data for the left side of surface b, or the upper side of surface g, or the weighted sum of the right side of surface b and the upper side of surface g, g0 can be filled using the data for the right side of surface b, h0 can be filled using the data for the upper side of surface b, e1 can be filled using the data for the left side of surface j, ej can be filled using the data for the underside of surface e, or the left side of surface j, or the weighted sum of the underside of surface e and the left side of surface j, j0 can be filled using the data for the underside of surface e, gj can be filled using the data for the underside of surface g, or the left side of surface j, or the weighted sum of the underside of surface g and the right side of surface j, g1 can be filled using the data for the right side of surface j, j1 can be filled using the data for the underside of surface g, and h1 can be filled using the data for the underside of surface j.
[0545] In the above example, when filling the resized area with data from a portion of an image, the data from the corresponding area can be copied and filled, or the data from the corresponding area can be filled with data obtained after a conversion process based on the image characteristics and type. For example, when a 360-degree image is converted into a two-dimensional space according to a projection format, a coordinate system (e.g., a two-dimensional plane coordinate system) for each surface can be defined. For convenience of explanation, it is assumed that (x, y, z) in three-dimensional space is converted to (x, y, C) or (x, C, z) or (C, y, z) for each surface. The above example illustrates a case where data from a surface other than the corresponding surface is obtained in the resized area. That is, while resizing is performed around the current surface, if data from another surface with different coordinate system characteristics is copied and filled as is, continuity may be distorted based on the resizing boundary. Therefore, the data from the other surface obtained according to the coordinate system characteristics of the current surface can be converted and filled into the resized area. This conversion is merely an example of a data processing method and is not limiting.
[0546] When filling the area to be resized by copying data from a portion of the image, the boundary between the area to be resized (e) and the area to be resized (e0) may contain distorted continuity (or continuity that changes suddenly). For example, continuous features may change based on the boundary, similar to when a straight edge becomes bent based on the boundary.
[0547] When the data of a partial area of the image is converted and filled into the area to be resized, the boundary area between the area to be resized and the area to be resized may include a gradually changing continuity.
[0548] The above example may be an example of a data processing method of the present invention, in which data in a portion of an image is converted based on the characteristics, type, etc. of the image during the size adjustment process (in this example, expansion), and the acquired data is filled into the area to be resized.
[0549] 26c may be an example in which the image size adjustment processes of 26a and 26b are combined to fill in the image using data from an area where continuity exists along the image boundaries (inner and outer boundaries). The size adjustment process of this example can be derived from 26a and 26b, so a detailed description will be omitted.
[0550] 26a may be an example of an image size adjustment process, 26b may be an example of a size adjustment process for division units within an image, and 26c may be an example of a plurality of size adjustment processes in which the size adjustment process for an image and the size adjustment for division units within an image are performed.
[0551] For example, size adjustment (area C in this example) can be performed on an image (first format in this example) acquired through a projection process, and size adjustment (area D in this example) can be performed on an image (second format in this example) acquired through a format conversion process. In this example, size adjustment (the entire image in this example) is performed on an image projected by an ERP, which may be an example of size adjustment (per surface in this example) performed on an image projected by a CMP via a format conversion unit after acquisition. The above example is one example of performing multiple size adjustments, and is not limited to this, and modifications to other cases are possible.
[0552] 27 is an exemplary diagram illustrating size adjustment for an image that has been converted into a CMP projection format and packed according to an embodiment of the present invention. Since FIG. 27 also assumes the continuity between surfaces as shown in FIG. 25, the boundaries of surfaces can have continuity with the boundaries of other surfaces.
[0553] In this example, the offset factors of W0 to W5 and H0 to H3 (assuming that the offset factors are used as size adjustment values in this example) can have various values. For example, they can be derived from a predetermined value, a motion search range for inter prediction, a unit obtained from a picture divider, or other values. In this case, the unit obtained from the picture divider can include a surface. That is, the size adjustment value can be determined based on F_Width and F_Height.
[0554] 27a shows an example in which data of areas where continuity exists is filled in each expanded area by adjusting the size of each surface (in this example, in the up, down, left, and right directions of each surface). For example, for surface a, continuous data can be filled in its outer boundary a0 to a6, and continuous data can be filled in the outer boundary b0 to b6 of surface b.
[0555] 27b shows an example of filling the expanded area with data of areas where continuity exists by adjusting the size of multiple surfaces (in this example, in the up, down, left, and right directions of the multiple surfaces). For example, using surfaces a, b, and c as references, the contours can be expanded to a0-a4, b0-b1, and c0-c4.
[0556] 27c is an example of filling the expanded area with data of areas where continuity exists by adjusting the size of the entire image (in this example, in the up, down, left, and right directions of the entire image). For example, the outline of the entire image consisting of surfaces a to f can be expanded to a0 to a2, b0, c0 to c2, d0 to d2, e0, and f0 to f2.
[0557] That is, size adjustment can be performed in units of one surface, in units of multiple surfaces where continuity exists, or in units of the entire surface.
[0558] The resized areas in the above example (a0 to f7 in this example) can be filled using data from areas (or surfaces) where there is continuity, as shown in Figure 24a. That is, the resized areas can be filled using data from the top, bottom, left, and right sides of surfaces a to f.
[0559] FIG. 28 is an exemplary diagram illustrating a data processing method for adjusting the size of a 360-degree image according to an embodiment of the present invention.
[0560] Referring to FIG. 28, region B (a0-a2, ad0, b0, c0-c2, cf1, d0-d2, e0, f0-f2), which is a region to be resized, can be filled with data from a region where continuity exists among pixel data belonging to a to f. Region C (ad1, be, cf0), another region to be resized, can be filled with a mixture of data from a region where the size is adjusted and data from a region that is spatially adjacent but not contiguous. Region C can be filled with a mixture of data from two regions selected from a to f (e.g., a and d, b and e, c and f), since size adjustment is performed between the two regions. For example, surfaces b and e may be spatially adjacent but not contiguous. Region be between surfaces b and e can be resized using data from surfaces b and e. For example, region be can be filled with a value obtained by averaging the data from surfaces b and e, or with a value obtained through a distance-weighted sum. In this case, the pixels used for data filling the area sized by surface b and surface e may be boundary pixels of each surface, but may also be internal pixels of the surface.
[0561] In summary, regions sized between image division units can be filled with data generated using a mix of data from both units.
[0562] The data processing method may be supported in some situations (in this example, when adjusting the size in multiple regions).
[0563] In Figures 27a and 27b, the areas whose size is adjusted between division units are configured separately for each division unit (taking Figure 27a as an example, a6 and d1 are configured for a and d, respectively), but in Figure 28, the areas whose size is adjusted between division units can be configured one for each adjacent division unit (one ad1 for a and d). Of course, the above method can be included in the candidate group of data processing methods in Figures 27a to 27c, and size adjustment can be performed using a data processing method different from the above example in Figure 28.
[0564] In the image resizing process of the present invention, a predetermined data processing method may be implicitly applied to the region to be resized, or related information may be explicitly generated using one of several data processing methods. The predetermined data processing method may be one of data processing methods such as filling using arbitrary pixel values, filling by copying outer pixels, filling by copying a partial region of an image, filling by transforming a partial region of an image, and filling with data derived from multiple regions of an image. For example, if the region to be resized is located inside an image (e.g., a packed image) and the regions on both sides (e.g., surfaces) are spatially adjacent but not contiguous, a data processing method for filling with data derived from multiple regions may be applied to fill the region to be resized. In addition, one of the several data processing methods may be selected to perform the resizing, and the selection information for this may be explicitly generated. This example may be applicable not only to 360-degree images but also to general images.
[0565] The encoder records the information generated in the above process into a bitstream in at least one unit selected from the group consisting of a sequence, a picture, a slice, and a tile, and the decoder parses the related information from the bitstream. The information may also be included in the bitstream in the form of SEI or metadata. The division, reconstruction, and resizing processes for a 360-degree image have been described with reference to some projection formats such as ERP and CMP, but are not limited thereto, and may be applied in the same manner or with modifications to other projection formats.
[0566] It has been explained that the image setting process applied to the 360-degree image encoding / decoding device described above can be applied not only to the encoding / decoding process but also to the pre-processing process, post-processing process, format conversion process, format inverse conversion process, etc.
[0567] In summary, the projection process may be configured to include an image setting process. In particular, the projection process may be performed including at least one image setting process. Division may be performed in units of regions (or surfaces) based on the projected image. Division may be performed into one region or multiple regions depending on the projection format. Division information may be generated based on the division. Furthermore, the size of the projected image may be adjusted, or the size of the projected region may be adjusted. At this time, the size adjustment may be performed for at least one region. Size adjustment information may be generated based on the size adjustment. Furthermore, reconstruction (or surface placement) of the projected image may be performed, or the projected region may be adjusted. At this time, reconstruction may be performed for at least one region. Reconstruction information may be generated based on the reconstruction.
[0568] In summary, the regional packing process may be configured to include an image setting process. More specifically, the regional packing process may be performed by including at least one image setting process. A division process may be performed in units of regions (or surfaces) based on the packed image. The packed image may be divided into one region or multiple regions according to the regional packing setting. Division information may be generated based on the division. Furthermore, the packed image may be resized, or the packed region may be resized. At this time, the size may be resized for at least one region. Size adjustment information may be generated based on the size adjustment. Furthermore, the packed image may be reconstructed, or the packed region may be reconstructed. At this time, the reconstruction may be performed for at least one region. Reconstruction information may be generated based on the reconstruction.
[0569] The projection process may include all or part of an image setting process and may include image setting information, which may be setting information for the projected image, or more specifically, setting information for an area within the projected image.
[0570] The regional packing process may include all or part of the image setting process and may include image setting information. This may be setting information of a packed image. Specifically, it may be setting information of an area in a packed image. Or it may be mapping information between a projected image and a packed image (see, for example, the description related to FIG. 11 . This can be understood assuming that P0 to P1 are projected images and S0 to S5 are packed images). Specifically, it may be mapping information between a partial area in a projected image and a partial area in a packed image. That is, it may be setting information assigned from a partial area in a projected image to a partial area in a packed image.
[0571] The information may be represented by information acquired through the various embodiments described above during the image configuration process of the present invention. For example, when related information is represented using at least one syntax element in Tables 1 to 6, the configuration information of the projected image may include pic_width_in_samples, pic_height_in_samples, part_top[i], part_left[i], part_width[i], part_height[i], etc., and the configuration information of the packed image may include pic_width_in_samples, pic_height_in_samples, part_top[i], part_left[i], part_width[i], part_height[i], convert_type_flag[i], part_resizing_flag[i], top_height_offset[i], bottom_height_offset[i], left_width_offset[i], right_width_offset[i], resizing_type_flag[i], etc. The above example may be an example of explicitly generating information for a surface (for example, part_top[i], part_left[i], part_width[i], and part_height[i] among the setting information of a projected image).
[0572] Part of the image setting process may be included in the projection process or the area packing process according to the projection format in a given operation.
[0573] For example, in the case of ERP, a process of adjusting the size by copying and filling data from an area in the opposite direction of the image size adjustment direction (in this example, the right and left directions) into an area expanded by m and n in the left and right directions, respectively, may be implicitly included. Alternatively, in the case of CMP, a process of adjusting the size by converting and filling data from an area that has continuity with the area to be adjusted into an area expanded by m, n, o, and p in the up, down, left, and right directions, respectively, may be implicitly included.
[0574] The projection formats in the above examples may be examples that replace existing projection formats or examples of projection formats (e.g., ERP1, CMP1) that are additional to existing projection formats. The above examples are not limiting, and various examples of the image setting process of the present invention may be alternatively combined. Similar applications are possible for other formats.
[0575] Meanwhile, the image encoding apparatus and the image decoding apparatus of FIGS. 1 and 2 may further include a block division unit, although not shown. Information regarding a basic coding unit may be obtained from the picture division unit, and the basic coding unit may refer to a basic (or starting) unit for prediction, transformation, quantization, etc. in the image encoding / decoding process. In this case, the coding unit may be composed of one luminance coding block and two chrominance coding blocks according to a color format (YCbCr in this example), and the size of each block may be determined according to the color format. In the following example, the description will be based on a block (luminance component in this example). Here, it is assumed that a block is a unit that can be obtained after each unit is determined, and that similar settings can be applied to other types of blocks.
[0576] The block division unit can be set in relation to each component of the image encoding device and decoding device, and the size and shape of the block can be determined through this process. At this time, the set block can be defined differently depending on the component, and can correspond to a prediction block in the case of a prediction unit, a transformation block in the case of a transformation unit, a quantization block in the case of a quantization unit, etc. Without being limited thereto, a block unit can be additionally defined according to other components. The size and shape of the block can be defined by the width and height of the block.
[0577] In the block division section, blocks can be expressed as MxN, and the maximum and minimum values of each block can be obtained within the range. For example, if the block shape is supported as a square and the maximum value of the block is set to 256x256 and the minimum value is set to 8x8, the size of 2 m ×2 m (In this example, m is an integer between 3 and 8, for example, 8x8, 16x16, 32x32, 64x64, 128x128, 256x256) or a block of size 2mx2m (In this example, m is an integer between 4 and 128) or a block of size mxm (In this example, m is an integer between 8 and 256). Alternatively, the block shape supports square and rectangular shapes, and if they have the same range as the above example, a block of size 2mx2m can be obtained. m ×2 n (In this example, m and n are integers from 3 to 8. Assuming that the ratio of width to height is a maximum of 2...
Claims
1. 1. A method for decoding a picture signal including a current picture, comprising: receiving a bitstream in which the current picture is encoded; Dividing the current picture based on at least one partitioning structure of tiles or slices; Dividing a first coding block in the current picture into a plurality of second coding blocks; obtaining a prediction block of a current block, the current block being one of the second coded blocks, based on syntax information from the bitstream; reconstructing the second coded block based on the predicted block; The first coding block is divided based on a plurality of division schemes, including four-way division, two-way division, three-way division, and index-based division; the quadrant is a division scheme in which one coding block is divided into four coding blocks based on both one horizontal line and one vertical line; the bisection is a division method of dividing one coding block into two coding blocks based on either one horizontal line or one vertical line; the third division is a division method of dividing one coding block into three coding blocks based on either two horizontal lines or two vertical lines; The index-based partitioning is a partitioning method for partitioning a target block into a plurality of sub-blocks based on a partitioning shape indicated by an index from a plurality of partitioning shape candidates; the plurality of division shape candidates are set to differ based on a size of the target block; the slice is made up of a group of tiles; A decoding method, wherein the slice forms a rectangular area in the current picture.
2. when the size of the first coded block is greater than a predetermined threshold size, only the quadrant is available for the first coded block; 2. The decoding method of claim 1, wherein both the quadrant and bipartition are available for the first coded block when the size of the first coded block is less than or equal to the threshold size.
3. 3. The decoding method of claim 2, wherein the bisection is allowed only when the block division based on the quadruple division is no longer performed.
4. 2. The decoding method of claim 1, wherein when the first coding block is divided into three second coding blocks based on the third division, one of the three second coding blocks has a size larger than the sizes of the other two of the three second coding blocks, and the other two of the three second coding blocks have the same size.
5. The decoding method according to claim 4 , wherein a size of one of the three second coding blocks is equal to a sum of sizes of the other two of the three second coding blocks.
6. The decoding method of claim 4 , wherein one of the three second coding blocks is disposed between two other of the three second coding blocks.
7. dividing the first coded block includes determining a division direction based on a flag obtained from the bitstream; 2. The method of claim 1, wherein the flag equal to a first value indicates a vertical orientation and the flag equal to a second value indicates a horizontal orientation.
8. 1. A method for encoding an image, comprising: Dividing the current picture based on at least one partitioning structure of tiles or slices; Dividing a first coding block in the current picture into a plurality of second coding blocks; obtaining a predicted block of a current block, information of the predicted block being coded into a bitstream, and the current block being one of the second coded blocks; encoding the second coded block based on the predicted block into the bitstream; The first coding block is divided based on a plurality of division schemes, including four-way division, two-way division, three-way division, and index-based division; the quadrant is a division scheme in which one coding block is divided into four coding blocks based on both one horizontal line and one vertical line; the bisection is a division method of dividing one coding block into two coding blocks based on either one horizontal line or one vertical line; the third division is a division method of dividing one coding block into three coding blocks based on either two horizontal lines or two vertical lines; The index-based partitioning is a partitioning method for partitioning a target block into a plurality of sub-blocks based on a partitioning shape indicated by an index from a plurality of partitioning shape candidates; the plurality of division shape candidates are set to differ based on a size of the target block; the slice is made up of a group of tiles; A method of encoding wherein the slice forms a rectangular area in the current picture.
9. 1. A method for transmitting a bitstream, comprising: Dividing the current picture based on at least one partitioning structure of tiles or slices; Dividing a first coding block in the current picture into a plurality of second coding blocks; obtaining a predicted block of a current block, information of the predicted block being coded into a bitstream, and the current block being one of the second coded blocks; encoding the second coding block based on the predicted block into the bitstream; transmitting the bitstream to an image decoding device; The first coding block is divided based on a plurality of division schemes, including four-way division, two-way division, three-way division, and index-based division; the quadrant is a division scheme in which one coding block is divided into four coding blocks based on both one horizontal line and one vertical line; the bisection is a division method of dividing one coding block into two coding blocks based on either one horizontal line or one vertical line; the third division is a division method of dividing one coding block into three coding blocks based on either two horizontal lines or two vertical lines; The index-based partitioning is a partitioning method for partitioning a target block into a plurality of sub-blocks based on a partitioning shape indicated by an index from a plurality of partitioning shape candidates; the plurality of division shape candidates are set to differ based on a size of the target block; the slice is made up of a group of tiles; The method, wherein the slices form rectangular regions in the current picture.
Citation Information
Patent Citations
Image coding device, image decoding device, image coding method, image decoding method, image coding program, and image decoding program
JP2013229674A
Moving image encoding device, moving image decoding device, moving image encoding method, moving image decoding method, program, and recording medium
JP2017112639A
Multi-type-tree framework for video coding
WO2017123980A1