Image data encoding / decoding method and device
The proposed method enhances the compression performance of 360-degree images by decoding the images through a predicted image generation and reconstruction process, addressing the performance limitations of conventional methods.
Patent Information
- Application Number
- JP2025040508
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-07-17
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2037-10-10
AI Technical Summary
Conventional image encoding/decoding methods struggle with efficiently processing 360-degree images due to the large amount of data generated, leading to insufficient performance in image processing systems.
A method for decoding 360-degree images that involves receiving a bitstream, generating a predicted image using syntax information, combining it with a residual image obtained through inverse quantization and inverse transformation, and reconstructing the decoded image into a 360-degree image in a projection format.
This approach improves the compression performance of 360-degree images by optimizing the image setting process during encoding and decoding, effectively addressing the performance limitations of existing systems.
Smart Images

Figure 2025090777000007 
Figure 2025090777000008 
Figure 2025090777000009
Abstract
Description
Technical Field
[0001] The present invention relates to image data encoding and decoding technologies, and more particularly, to a method and apparatus for processing encoding and decoding of 360-degree images for immersive media services.
Background Art
[0002] With the spread of the Internet and mobile terminals and the development of information and communication technologies, the use of multimedia data has increased rapidly. Recently, demands for high-resolution images and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, have arisen in various fields, and demands for immersive media services such as virtual reality and augmented reality have increased rapidly. In particular, in the case of 360-degree images for virtual reality and augmented reality, since multi-view images captured by a plurality of cameras are processed, the amount of data generated thereby increases enormously, but the performance of the image processing system for processing this is insufficient.
[0003] As described above, in the conventional image encoding / decoding method and apparatus, improvement in performance for image processing, particularly image encoding / decoding, is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present invention is for solving the problems as described above, and an object thereof is to provide a method for improving an image setting process in an initial stage of encoding and decoding. More specifically, it is to provide an encoding and decoding method and apparatus for improving an image setting process considering characteristics of 360-degree images.
Means for Solving the Problems
[0005] One aspect of the present invention for achieving the above object provides a method for decoding a 360-degree image.
[0006] Here, the method for decoding a 360-degree image may include receiving a bitstream in which the 360-degree image is encoded, generating a predicted image with reference to syntax information obtained from the received bitstream, obtaining a decoded image by combining the generated predicted image with a residual image obtained by inverse quantization and inverse transformation of the bitstream, and reconstructing the decoded image into a 360-degree image in a projection format.
[0007] Here, the syntax information may include projection format information for the 360-degree image.
[0008] Here, the projection format information may indicate at least one of an ERP (Equi-Rectangular Projection) format in which the 360-degree image is projected onto a two-dimensional plane, a CMP (CubeMap Projection) format in which the 360-degree image is projected onto a cube, an OHP (OctaHedron Projection) format in which the 360-degree image is projected onto an octahedron, and an ISP (IcoSahedral Projection) format in which the 360-degree image is projected onto an icosahedron.
[0009] Here, the reconstructing step may include obtaining arrangement information by regional packing with reference to the syntax information, and rearranging each block of the decoded image based on the arrangement information.
[0010] Here, the step of generating the predicted image may include performing image expansion on a reference picture obtained by restoring the bitstream, and generating a predicted image with reference to the reference picture on which the image expansion has been performed.
[0011] Here, the step of performing the image expansion may include performing image expansion based on a division unit of the reference picture.
[0012] Here, in the step of performing image expansion based on the division unit, an area individually expanded for each division unit can be generated using the boundary pixels of the division unit.
[0013] Here, the expanded area can be generated using the boundary pixels of a division unit that is spatially adjacent to the division unit to be expanded, or the boundary pixels of a division unit having image continuity with the division unit to be expanded.
[0014] Here, in the step of performing image expansion based on the division unit, an expanded image for the combined area can be generated using the boundary pixels of an area where two or more spatially adjacent division units among the division units are combined.
[0015] Here, in the step of performing image expansion based on the division unit, an expanded area can be generated between the adjacent division units using all of the adjacent pixel information of the spatially adjacent division units among the division units.
[0016] Here, in the step of performing image expansion based on the division unit, the expanded area can be generated using the average value of the adjacent pixels of each of the spatially adjacent division units.
[0017] Here, the step of generating the predicted image can include: obtaining a group of motion vector candidates including the motion vectors of the blocks adjacent to the current block to be decoded from the motion information included in the syntax information; deriving a predicted motion vector from among the group of motion vector candidates based on the selection information extracted from the motion information; and determining a predicted block of the current block to be decoded using the final motion vector derived by adding the predicted motion vector and the differential motion vector extracted from the motion information.
[0018] Here, when the motion vector candidate group is such that a block adjacent to the current block is different from the surface to which the current block belongs, the motion vector candidate group can be composed only of motion vectors for blocks belonging to a surface having image continuity with the surface to which the current block belongs among the adjacent blocks.
[0019] Here, the adjacent block can mean a block adjacent to the current block in at least one of the upper left, upper, upper right, left, and lower left directions of the current block.
[0020] Here, the final motion vector can indicate a reference region that belongs to at least one reference picture based on the current block and is set in a region having image continuity between surfaces according to the projection format.
[0021] Here, the reference picture can be extended based on image continuity according to the projection format in the up, down, left, and right directions, and then the reference region can be set.
[0022] Here, the reference picture is extended in units of the surface, and the reference region can be set across the boundary of the surface.
[0023] Here, the motion information can include at least one of a reference picture list to which the reference picture belongs, an index of the reference picture, and a motion vector indicating the reference region.
[0024] Here, the step of generating a predicted block of the current block can include dividing the current block into a plurality of sub-blocks and generating a predicted block for each of the plurality of divided sub-blocks.
Advantages of the Invention
[0025] When using the image encoding / decoding method and apparatus according to the embodiment of the present invention as described above, the compression performance can be improved. In particular, in the case of a 360-degree image, the compression performance can be improved.
Brief Description of Drawings
[0026]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16a
Figure 16b
Figure 16c
Figure 16d
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Embodiments for Carrying Out the Invention
[0027] The present invention can be subjected to various modifications and can have various embodiments. Here, specific embodiments are illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiments, and it should be understood to include all modifications, equivalents, or alternatives included in the spirit and technical scope of the present invention.
[0028] Terms such as first, second, A, B, etc. are used to describe various components, but the components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present invention, the first component can be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of the plurality of related described items.
[0029] When it is said that a certain component is "connected to" or "coupled to" another component, it means that it is directly connected or coupled to the other component, but it should be understood that another component may be interposed therebetween. On the other hand, when it is said that a certain component is "directly connected to" or "directly coupled to" another component, it should be understood that no other component is interposed therebetween.
[0030] The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "including" or "having" are intended to specify the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0031] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Terms defined in commonly used dictionaries should be interpreted as consistent with the meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless clearly defined herein.
[0032] The image encoding device and the decoding device can be user terminals such as a personal computer (PC), a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a PlayStation Portable (PSP), a wireless communication terminal, a smartphone, a TV, a virtual reality device (VR), an augmented reality device (AR), a mixed reality device (MR), a head-mounted display (HMD), smart glasses, etc., or server terminals such as an application server or a service server. It can include various devices equipped with a communication device such as a communication modem for communicating with various devices or a wired / wireless communication network, a memory for storing various programs and data for encoding or decoding an image or for predicting within or between screens for encoding or decoding, a processor for executing programs for calculation and control, etc. Also, the image encoded into a bitstream by the image encoding device can be transmitted to the image decoding device in real-time or non-real-time via a wired / wireless communication network such as the Internet, a short-range wireless communication system, a wireless LAN network, a WiBro network, a mobile communication network, etc., or via various communication interfaces such as a cable or a universal serial bus (USB), and decoded by the image decoding device to be restored and played back as an image.
[0033] Also, the image encoded into a bitstream by the image encoding device can be transmitted from the encoding device to the decoding device via a computer-readable recording medium.
[0034] The above-described image encoding device and image decoding device can be separate devices, respectively, but depending on the implementation, they can be made into one image encoding / decoding device. In that case, some components of the image encoding device are technical elements substantially the same as some components of the image decoding device, including at least the same structure or being realizable to perform at least the same function.
[0035] Therefore, in the following detailed descriptions of the following technical elements and their operating principles, etc., duplicate descriptions of corresponding technical elements will be omitted.
[0036] Since the image decoding device corresponds to a computer device that applies the image encoding method performed by the image encoding device to decoding, the following description will focus on the image encoding device.
[0037] The computer device can include a memory that stores a program and / or software module for implementing the image encoding method and / or image decoding method, and a processor connected to the memory to execute the program. The image encoding device may sometimes be called an encoder, and the image decoding device may sometimes be called a decoder.
[0038] Generally, an image can be composed of a series of still images, and these still images can be divided into units of GOP (Group of Pictures). Each still image may be referred to as a picture. At this time, a picture can indicate any one of a progressive signal, a frame and a field in an interlace signal, and when the encoding / decoding is performed in units of frames, the image can be represented by "frame", and when it is performed in units of fields, it can be represented by "field". In the present invention, the description is made assuming a progressive signal, but it is also applicable to an interlace signal. As higher-level concepts, there can be units such as GOP and sequence. In addition, each picture can be divided into predetermined regions such as slices, tiles, and blocks. Further, one GOP may include units such as I pictures, P pictures, and B pictures. An I picture can mean a picture that is encoded / decoded by itself without using a reference picture, and P pictures and B pictures can mean pictures that are encoded / decoded by performing processes such as motion estimation and motion compensation using a reference picture. Generally, in the case of a P picture, I pictures and P pictures can be used as reference pictures, and in the case of a B picture, I pictures and P pictures can be used as reference pictures, but this definition can also be changed depending on the encoding / decoding settings.
[0039] Here, the picture referred to in encoding / decoding is called a reference picture, and the referred block or pixel is called a reference block or a reference pixel. Also, the reference data can be not only the pixel values in the spatial domain but also the coefficient values in the frequency domain and various encoding / decoding information generated and determined during the encoding / decoding process. For example, in the prediction unit, it can be intra-prediction related information or motion related information, in the conversion unit / inverse conversion unit, it can be conversion related information, in the quantization unit / inverse quantization unit, it can be quantization related information, in the encoding unit / decoding unit, it can be encoding / decoding related information (context information), and in the in-loop filter unit, it can be filter related information, etc.
[0040] The minimum unit forming an image can be a pixel. The number of bits used to represent one pixel is called the bit depth. Generally, the bit depth is 8 bits, and depending on the encoding settings, a greater bit depth can be supported. At least one bit depth can be supported according to the color space. Also, according to the color format of the image, it can be composed of at least one color space. According to the color format, it can be composed of one or more pictures having a certain size or one or more pictures having different sizes. For example, in the case of YCbCr4:2:0, it can be composed of one luminance component (in this example, Y) and two color difference components (in this example, Cb / Cr). At this time, the composition ratio of the color difference component and the luminance component can have a horizontal ratio of 1: vertical ratio of 2. As another example, in the case of 4:4:4, the horizontal and vertical can have the same composition ratio. When composed of one or more color spaces as in the above examples, the picture can be divided into each color space.
[0041] In the present invention, the description is based on a part of the color space (in this example, Y) of a part of the color format (in this example, YCbCr). The same or similar application (settings dependent on a specific color space) can also be made to another color space (in this example, Cb, Cr) according to the color format. However, it is also possible to make partial differences (settings independent of a specific color space) in each color space. That is, the settings dependent on each color space can mean being proportional to the composition ratio of each component (determined according to, for example, 4:2:0, 4:2:2, 4:4:4, etc.) or having settings dependent thereon, and the settings independent of each color space can mean having settings only for the corresponding color space regardless of or independently of the composition ratio of each component. In the present invention, depending on the encoder / decoder, some configurations can have settings independent of or dependent on them.
[0042] The setting information or syntax elements required in the image encoding process can be determined at unit levels such as video, sequence, picture, slice, tile, block, etc. This can be recorded in the bitstream in units such as VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), Slice Header, Tile Header, Block Header, etc. and transmitted to the decoder. In the decoder, it can be parsed at the same level of unit to restore the setting information transmitted from the encoder and used in the image decoding process. Also, related information can be transmitted to the bitstream in the form of SEI (Supplement Enhancement Information) or metadata, parsed, and used. Each parameter set has a unique ID value, and a lower-level parameter set can have the ID value of the upper-level parameter set it references. For example, a lower-level parameter set can reference the information of the upper-level parameter set with a matching ID value among one or more upper-level parameter sets. Among the various unit examples described above, if any unit contains one or more other units, the corresponding unit may be called the upper unit, and the contained units may be called the lower units.
[0043] In the case of the setting information generated in the said unit, it can include content for an independent setting for each corresponding unit, or content related to a dependent setting that depends on previous, subsequent, or upper units, etc. Here, the dependent setting can be understood to indicate the setting information of the corresponding unit with flag information (for example, if a 1-bit flag is 1, it follows the setting; if it is 0, it does not follow the setting) that follows the settings of previous, subsequent, or upper units. Although the setting information in the present invention is mainly described with examples of independent settings, examples of addition or substitution for content regarding the dependent relationship with the setting information of previous, subsequent units, or upper units of the current unit may also be included.
[0044] FIG. 1 is a block diagram of an image encoding apparatus according to an embodiment of the present invention. FIG. 2 is a block diagram of an image decoding apparatus according to an embodiment of the present invention.
[0045] Referring to FIG. 1, the image encoding apparatus can be configured to include a prediction unit, a subtraction unit, a conversion unit, a quantization unit, an inverse quantization unit, an inverse conversion unit, an addition unit, an in-loop filter unit, a memory, and / or an encoding unit. Among the above configurations, some may not necessarily be included, and some or all of them may be selectively included according to the implementation, and additional configurations not shown may also be included.
[0046] Referring to FIG. 2, the image decoding apparatus can be configured to include a decoding unit, a prediction unit, an inverse quantization unit, an inverse conversion unit, an addition unit, an in-loop filter unit, and / or a memory. Among the above configurations, some may not necessarily be included, and some or all of them may be selectively included depending on the implementation, and additional configurations not shown may also be included.
[0047] The image encoding apparatus and the image decoding apparatus can be separate apparatuses, but depending on the implementation, they may be made into one image encoding / decoding apparatus. In that case, some configurations of the image encoding apparatus are technical elements substantially the same as some configurations of the image decoding apparatus, and can be realized so as to include at least the same structure or perform at least the same function. Therefore, in the following detailed description of these technical elements and their operating principles, etc., duplicate descriptions of corresponding technical elements will be omitted. Since the image decoding apparatus corresponds to a computer apparatus that applies the image encoding method performed by the image encoding apparatus to decoding, the following description will focus on the image encoding apparatus. The image encoding apparatus may sometimes be called an encoder, and the image decoding apparatus may sometimes be called a decoder.
[0048] The prediction unit can be realized by using a prediction module, which is a software module, and can generate a prediction block for the block to be encoded using an intra prediction method or an inter prediction method. The prediction unit predicts the current block to be currently encoded in the image to generate a prediction block. That is, the prediction unit predicts the pixel value of each pixel of the current block to be encoded in the image through intra prediction or inter prediction, and generates a prediction block having the predicted pixel value of each generated pixel. In addition, the prediction unit can transmit the information necessary to generate the prediction block to the encoding unit so as to encode the information for the prediction mode, record the information thereby in the bitstream, and transmit this to the decoder. The decoding unit of the decoder can parse the information for this, restore the information for the prediction mode, and then use this for intra prediction or inter prediction.
[0049] The subtraction unit subtracts the prediction block from the current block to generate a residual block. That is, the subtraction unit calculates the difference between the pixel value of each pixel of the current block to be encoded and the predicted pixel value of each pixel of the prediction block generated through the prediction unit, and generates a residual block, which is a residual signal in block form.
[0050] The conversion unit can convert a signal belonging to the spatial domain into a signal belonging to the frequency domain. At this time, the signal obtained through the conversion process is called a transformed coefficient. For example, a residual block having a residual signal transmitted from the subtraction unit can be converted to obtain a conversion block having transformed coefficients, but the input signal is determined according to the encoding settings, and this is not limited to the residual signal.
[0051] The conversion unit can perform conversion on the residual block using conversion techniques such as Hadamard Transform, DST Based-Transform (Discrete Sine Transform), DCT Based-Transform (Discrete Cosine Transform), etc. However, it is not limited to this, and various conversion techniques obtained by improving and modifying this can be used.
[0052] For example, at least one of the above conversions can be supported, and at least one detailed conversion technique can be supported for each conversion technique. At this time, at least one detailed conversion technique can be a conversion technique configured such that a part of the basis vectors is different for each conversion technique. For example, as conversion techniques, DST-based conversion and DCT-based conversion can be supported. In the case of DST, detailed conversion techniques such as DST-I, DST-II, DST-III, DST-V, DST-VI, DST-VII, DST-VIII can be supported, and in the case of DCT, detailed conversion techniques such as DCT-I, DCT-II, DCT-III, DCT-V, DCT-VI, DCT-VII, DCT-VIII can be supported.
[0053] Any one of the above conversions (for example, one conversion technique & one detailed conversion technique) can be set as the basic conversion technique, and additional conversion techniques (for example, multiple conversion techniques || multiple detailed conversion techniques) can be supported. Whether to support additional conversion techniques is determined in units such as sequences, pictures, slices, tiles, etc., and relevant information can be generated in the unit. When additional conversion techniques are supported, the conversion technique selection information is determined in units such as blocks, and relevant information can be generated.
[0054] The conversion can be performed in the k / vertical direction. For example, by performing one-dimensional conversion in the horizontal direction using the basis vectors in the conversion and one-dimensional conversion in the vertical direction to perform a total two-dimensional conversion, the pixel values in the spatial domain can be converted into the frequency domain.
[0055] In addition, the conversion can be adaptively performed in the horizontal / vertical direction. Specifically, it can be determined whether to perform the conversion adaptively according to at least one encoding setting. For example, when the prediction mode in intra-frame prediction is the horizontal mode, DCT-I can be applied in the horizontal direction and DST-I can be applied in the vertical direction; when the prediction mode in intra-frame prediction is the vertical mode, DST-VI can be applied in the horizontal direction and DCT-VI can be applied in the vertical direction; when the prediction mode is Diagonal down left, DCT-II can be applied in the horizontal direction and DCT-V can be applied in the vertical direction; when the prediction mode is Diagonal down right, DST-I can be applied in the horizontal direction and DST-VI can be applied in the vertical direction.
[0056] According to the encoding cost for each candidate of the size and shape of the conversion block, the size and shape of each conversion block are determined, and information such as the image data of each determined conversion block and the size and shape of each determined conversion block can be encoded.
[0057] Among the conversion shapes, a square conversion can be set as the basic conversion shape, and additional conversion shapes (for example, rectangular shapes) can be supported. Whether to support additional conversion shapes is determined in units such as sequence, picture, slice, and tile, related information can be generated in these units, and the conversion shape selection information is determined in units such as blocks and related information can be generated.
[0058] Also, the support for the conversion block shape can be determined according to the encoding information. At this time, the encoding information can include slice type, encoding mode, block size and shape, block splitting method, etc. That is, one conversion shape can be supported according to at least one piece of encoding information, and multiple conversion shapes can be supported according to at least one piece of encoding information. The former case is an implicit situation, and the latter case can be an explicit situation. In the explicit case, adaptive selection information indicating the optimal candidate group among multiple candidate groups can be generated and recorded in the bitstream. Including this example, in the present invention, when explicitly generating encoding information, it can be understood that the corresponding information is recorded in the bitstream in various units, and the decoder parses the relevant information in various units and restores it to the decoded information. Also, when implicitly processing the encoding / decoding information, it can be understood that the encoder and the decoder process it according to the same process, rules, etc.
[0059] As an example, the support for rectangular conversion can be determined according to the slice type. The conversion shape supported in the case of an I slice is square conversion, and the conversion shape supported in the case of a P / B slice can be square or rectangular conversion.
[0060] As an example, the support for rectangular conversion can be determined according to the encoding mode. The conversion shape supported in the Intra case is square conversion, and the conversion shape supported in the Inter case can be square or rectangular conversion.
[0061] As an example, the support for rectangular conversion can be determined according to the block size and shape. The conversion shape supported by blocks of a certain size or more is square conversion, and the conversion shape supported by blocks smaller than a certain size can be square or rectangular conversion.
[0062] As an example, the conversion assistance for a rectangle can be determined according to the block division method. When the block to be converted is a block obtained by a quad tree division method, the supported conversion shape is a square conversion. When the block is a block obtained by a binary tree division method, the supported conversion shape can be a square or a rectangle conversion.
[0063] The above example is an example of the conversion shape assistance according to one piece of coding information, and a plurality of pieces of information can also be combined to participate in the additional conversion shape assistance setting. The above example is only an example of the additional conversion shape assistance according to various coding settings, and is not limited to the above, and various deformation examples are possible.
[0064] According to the coding setting or the characteristics of the image, the conversion process can be omitted. For example, according to the coding setting (assuming a lossless compression environment in this example), the conversion process (including the reverse process) can be omitted. As another example, when the compression performance by conversion is not exhibited according to the characteristics of the image, the conversion process can be omitted. At this time, the conversion to be omitted can be in the unit of the whole, or in either the horizontal unit or the vertical unit. It can be determined whether to support such an omission according to the block size and shape, etc.
[0065] For example, in a setting where the omission of horizontal and vertical conversions is grouped, when the conversion omission flag is 1, the conversions in the horizontal and vertical directions are not performed. When the conversion omission flag is 0, the conversions in the horizontal and vertical directions can be performed. In a setting where the omission of horizontal and vertical conversions operates independently, when the first conversion omission flag is 1, the conversion in the horizontal direction is not performed. When the first conversion omission flag is 0, the conversion in the horizontal direction is performed. When the second conversion omission flag is 1, the conversion in the vertical direction is not performed. When the second conversion omission flag is 0, the conversion in the vertical direction is performed.
[0066] When the block size falls within range A, conversion omission can be supported; when the block size falls within range B, conversion omission cannot be supported. For example, when the horizontal width of the block is greater than M or the vertical height of the block is greater than N, the conversion omission flag cannot be supported; when the horizontal width of the block is less than m or the vertical height of the block is less than n, the conversion omission flag can be supported. M(m) and N(n) may be the same or different. The conversion-related settings can be determined in units such as sequences, pictures, slices, etc.
[0067] When additional conversion techniques are supported, the settings of the conversion techniques can be determined according to at least one piece of coding information. At this time, the coding information can include slice type, coding mode, block size and shape, prediction mode, etc.
[0068] As an example, the support for conversion techniques can be determined according to the coding mode. The conversion techniques supported in the Intra case are DCT-I, DCT-III, DCT-VI, DST-II, DST-III, and the conversion techniques supported in the Inter case can be DCT-II, DCT-III, DST-III.
[0069] As an example, the support for conversion techniques can be determined according to the slice type. The conversion techniques supported in the I slice case are DCT-I, DCT-II, DCT-III, the conversion techniques supported in the P slice case are DCT-V, DST-V, DST-VI, and the conversion techniques supported in the B slice case can be DCT-I, DCT-II, DST-III.
[0070] As an example, the support for conversion techniques can be determined according to the prediction mode. The conversion techniques supported in prediction mode A are DCT-I, DCT-II, the conversion techniques supported in prediction mode B are DCT-I, DST-I, and the conversion techniques supported in prediction mode C can be DCT-I. At this time, prediction modes A and B are directional modes, and prediction mode C can be a non-directional mode.
[0071] As an example, the support for the conversion technique can be determined according to the size and shape of the block. The conversion technique supported for blocks with a certain size or more is DCT-II, the conversion techniques supported for blocks with a size less than a certain size are DCT-II and DST-V, and the conversion techniques supported for blocks with a size of a certain size or more and less than a certain size can be DCT-I, DCT-II, and DST-I. Also, the conversion techniques supported for a square shape are DCT-I and DCT-II, and the conversion techniques supported for a rectangular shape can be DCT-I and DST-I.
[0072] The above example is an example of the support for the conversion technique according to one piece of encoding information, and a plurality of pieces of information can be combined and involved in the support setting of additional conversion techniques. It is not limited to only the case of the above example, and deformation to other examples is also possible. Also, the conversion unit can transmit information necessary for generating a conversion block to the encoding unit to encode it, record the information obtained thereby in a bit stream, and transmit it to the decoder. The decoding unit of the decoder can parse the information for this and use it in the inverse conversion process.
[0073] The quantization unit can quantize the input signal. At this time, the signal obtained through the quantization process is called a quantized coefficient. For example, a residual block having residual transform coefficients transmitted from the conversion unit can be quantized to obtain a quantized block having quantized coefficients, but the input signal is determined according to the encoding setting, and this is not limited to residual transform coefficients.
[0074] The quantization unit can quantize the transformed residual block using quantization techniques such as Dead Zone Uniform Threshold Quantization and Quantization Weighted Matrix, and is not limited thereto. Various quantization techniques obtained by improving and modifying this can be used. Whether to support additional quantization techniques is determined in units such as sequences, pictures, slices, and tiles, and relevant information can be generated in these units. When additional quantization techniques are supported, the quantization technique selection information can be determined in units such as blocks, and relevant information can be generated.
[0075] When additional quantization techniques are supported, the setting of the quantization technique can be determined according to at least one piece of encoding information. At this time, the encoding information can include slice type, encoding mode, block size and shape, prediction mode, etc.
[0076] For example, the quantization unit can be set such that the quantization weighted matrix according to the encoding mode is different from the weighted matrix applied according to inter-picture prediction / intra-picture prediction. Also, the weighted matrix applied according to the intra-picture prediction mode can be set to be different. At this time, assuming that the quantization weighted matrix is of size M×N and the block size is the same as the quantization block size, some of the quantization components can be different quantization matrices.
[0077] Depending on the encoding settings or image characteristics, the quantization process can be omitted. For example, depending on the encoding settings (assuming a lossless compression environment in this example), the quantization process (including the inverse process) can be omitted. As another example, when the compression performance by quantization is not exhibited depending on the image characteristics, the quantization process can be omitted. At this time, the area to be omitted can be the entire area or a partial area. Whether to support such omission can be determined according to the block size and shape, etc.
[0078] Information about the quantization parameter (QP) can be generated in units such as sequences, pictures, slices, tiles, and blocks. For example, the basic QP can be set in the higher-level unit where the QP information is first generated <1>, and the QP can be set to the same or a different value as the QP set in the higher-level unit as it goes to the lower-level unit <2>. Through such a process, in the quantization process performed in some units, the QP can be finally determined <3>. At this time, units such as sequences and pictures correspond to <1>, units such as slices, tiles, and blocks correspond to <2>, and units such as blocks correspond to <3>.
[0079] Information about the QP can be generated based on the QP in each unit. Or, a preset QP can be set as a predicted value, and the difference value information from the QP in each unit can be generated. Or, a QP obtained based on at least one of the QP set in the higher-level unit, the QP set in the same unit previously, or the QP set in the adjacent unit can be set as a predicted value, and the difference value information from the QP in the current unit can be generated. Or, a QP obtained based on the QP set in the higher-level unit and at least one piece of coding information can be set as a predicted value, and the difference value information from the QP in the current unit can be generated. At this time, the previous same unit is a unit that can be defined according to the coding order of each unit, the adjacent unit is a spatially adjacent unit, and the coding information can be the slice type, coding mode, prediction mode, position information, etc. of the corresponding unit.
[0080] As an example, the QP of the current unit can set the QP of the higher-level unit as a predicted value and generate the difference value information. The difference value information between the QP set in the slice and the QP set in the picture can be generated, or the difference value information between the QP set in the tile and the QP set in the picture can be generated. Also, the difference value information between the QP set in the block and the QP set in the slice or tile can be generated. Also, the difference value information between the QP set in the sub-block and the QP set in the block can be generated.
[0081] As an example, the QP of the current unit can set, as a predicted value, the QP obtained based on at least one adjacent unit's QP or the QP obtained based on at least one previous unit's QP, and generate difference value information. Difference value information can be generated with respect to the QP obtained based on the QPs of adjacent blocks such as the left, upper left, lower left, upper, and upper right of the current block. Or, difference value information can be generated with respect to the QP of the encoded picture before the current picture.
[0082] As an example, the QP of the current unit can set, as a predicted value, the QP obtained based on the QP of the upper unit and at least one piece of encoding information, and generate difference value information. Difference value information can be generated with respect to the QP of the slice corrected according to the QP of the current block and the slice type (I / P / B). Or, difference value information can be generated with respect to the QP of the tile corrected according to the QP of the current block and the encoding mode (Intra / Inter). Or, difference value information can be generated with respect to the QP of the picture corrected according to the QP of the current block and the prediction mode (directional / non-directional). Or, difference value information can be generated with respect to the QP of the picture corrected according to the QP of the current block and the position information (x / y). At this time, the meaning of the correction can mean being added or subtracted in an offset form to the QP of the upper unit used for prediction. At this time, at least one offset information can be supported according to the encoding setting, and can be implicitly processed according to a predetermined process or related information can be explicitly generated. It is not limited only to the case of the above example, and variations to other examples are also possible.
[0083] The above example can be a possible example when a signal indicating QP variation is provided or activated. For example, when a signal indicating QP variation is not provided or deactivated, difference value information is not generated, and the predicted QP can be determined as the QP of each unit. As another example, when a signal indicating QP variation is provided or activated, difference value information is generated, and when the value is 0, the predicted QP can be determined as the QP of each unit.
[0084] The quantization unit can transmit the information necessary to generate quantization blocks to the encoding unit so as to encode the same, record the information thereby in a bit stream, and transmit the bit stream to a decoder. The decoding unit of the decoder can parse the information therefor and use the same in an inverse quantization process.
[0085] In the above example, the description was made under the assumption that the residual block is converted and quantized via the conversion unit and the quantization unit. However, it is not necessary to perform the quantization process by converting the residual signal to generate a residual block having conversion coefficients, and it is possible not only to perform only the quantization process without converting the residual signal of the residual block to conversion coefficients, but also not to perform both the conversion and the quantization processes. This can be determined according to the settings of the encoder.
[0086] The encoding unit can scan at least one of the quantization coefficients, conversion coefficients, or residual signals of the generated residual block according to at least one scan order (for example, zigzag scan, vertical scan, horizontal scan, etc.) to generate a quantization coefficient sequence, a conversion coefficient sequence, or a signal sequence, and can encode the same using at least one entropy coding technique. At this time, the information about the scan order can be determined according to the encoding settings (for example, encoding mode, prediction mode, etc.), and can be determined implicitly or explicitly generate related information. For example, according to the in-picture prediction mode, any one of a plurality of scan orders can be selected.
[0087] In addition, encoded data including the encoded information transmitted from each component can be generated and output to a bit stream, which can be realized by a multiplexer (MUX). At this time, as encoding techniques, methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC) can be used for encoding, and it is not limited to this, and various encoding techniques obtained by improving and modifying this can be used.
[0088] When performing entropy coding (assuming CABAC in this example) on syntax elements such as the residual block data and the information generated in the encoding / decoding process, the entropy coding device can include a binarizer, a context modeler, and a binary arithmetic coder. At this time, the binary arithmetic coder can include a regular coding engine and a bypass coding engine.
[0089] Since the syntax elements input to the entropy coding device may not be binary, when the syntax elements are not binary, the binarizer can binarize the syntax elements and output a Bin String consisting of 0 or 1. At this time, Bin indicates a bit consisting of 0 or 1 and can be encoded through the binary arithmetic coder. At this time, either the regular coding engine or the bypass coding engine can be selected based on the occurrence probabilities of 0 and 1. This can be determined according to the encoding / decoding settings. When the data has the same frequency of 0 and 1 for the syntax elements, the bypass coding engine can be used, and otherwise, the regular coding engine can be used.
[0090] When performing binarization on the syntax element, various methods can be used. For example, Fixed Length Binarization, Unary Binarization, Truncated Rice Binarization, K-th Exp-Golomb Binarization, etc. can be used. Also, depending on the range of values that the syntax element has, signed binarization or unsigned binarization can be performed. The binarization process for the syntax element generated in the present invention can be performed including not only the binarization mentioned in the above examples but also other additional binarization methods.
[0091] The inverse quantization unit and the inverse transformation unit can be realized by performing the processes in the transformation unit and the quantization unit in reverse. For example, the inverse quantization unit can inverse-quantize the quantization transformation coefficients generated by the quantization unit, and the inverse transformation unit can inverse-transform the inverse-quantized transformation coefficients to generate a restored residual block.
[0092] The addition unit adds the prediction block and the restored residual block to restore the current block. The restored block is stored in the memory and can be used as reference data (such as the prediction unit and the filter unit).
[0093] The in-loop filter section can include at least one post-processing filter process such as a deblocking filter, sample adaptive offset (SAO), adaptive loop filter (ALF), etc. The deblocking filter can remove block distortion generated at the boundary between blocks from the restored image. The ALF can perform filtering based on a value obtained by comparing the restored image with the input image. Specifically, filtering can be performed based on a value obtained by comparing the restored image with the input image after the block has been filtered through the deblocking filter. Alternatively, filtering can be performed based on a value obtained by comparing the restored image with the input image after the block has been filtered through the SAO. The SAO can restore an offset difference based on a value obtained by comparing the restored image with the input image and can be applied in forms such as band offset (BO) and edge offset (EO). Specifically, the SAO can add an offset from the original image to the restored image to which the deblocking filter has been applied in at least one pixel unit and can be applied in forms such as BO and EO. Specifically, an offset from the original image can be added to the restored image after the block has been filtered through the ALF in pixel units and can be applied in forms such as BO and EO.
[0094] As filtering-related information, setting information regarding whether to support each post-processing filter can be generated in units such as sequence, picture, slice, tile, etc. Also, setting information regarding whether to execute each post-processing filter can be generated in units such as picture, slice, tile, block, etc. The range to which the execution of the filter is applied can be divided into the inside of the image and the boundary of the image, and setting information considering this can be generated. Also, information related to the filtering operation can be generated in units such as picture, slice, tile, block, etc. The said information can be subjected to implicit or explicit processing, and for the said filtering, an independent filtering process or a dependent filtering process can be applied according to the color component. This can be determined according to the encoding setting. The in-loop filter unit can transmit the filtering-related information to the encoding unit to encode it, record the information thereby in the bitstream and transmit it to the decoder, and the decoding unit of the decoder can parse the information for this and apply it to the in-loop filter unit.
[0095] The memory can store the restored block or picture. The restored block or picture stored in the memory can be provided to a prediction unit that performs intra-prediction or inter-prediction. Specifically, a storage space in the form of a queue of the bitstream compressed by the encoder can be processed as a coded picture buffer (CPB), and a space for storing the decoded image in picture units can be processed as a decoded picture buffer (DPB). In the case of CPB, the decoding units are stored in the decoding order, the decoding operation can be emulated in the encoder, and the bitstream compressed in the emulation process can be stored. The bitstream output from the CPB is restored through the decoding process, the restored image is stored in the DPB, and the picture stored in the DPB can be referred to in subsequent image encoding and decoding processes.
[0096] The decoding unit can be realized by reversing the process in the encoding unit. For example, it can receive a quantization coefficient sequence, a transform coefficient sequence, or a signal sequence from a bit stream, decode it, and parse the decoded data including decoded information to transmit it to each component unit.
[0097] Hereinafter, an image setting process applied to an image encoding / decoding apparatus according to an embodiment of the present invention will be described. This can be an example (image initial setting) applied at a stage before encoding / decoding, but some processes can also be examples applicable at other stages (for example, a stage after encoding / decoding or an internal stage of encoding / decoding, etc.). The image setting process can be performed in consideration of network and user environments such as characteristics of multimedia content, bandwidth, performance and accessibility of a user terminal. For example, according to the settings of encoding / decoding, image segmentation, image size adjustment, image reconstruction, etc. can be performed. The image setting process to be described later will be mainly described with a rectangular image as the center, but is not limited thereto and is also applicable to polygonal images. Regardless of the shape of the image, the same image setting or different image settings can be applied. This can be determined according to the settings of encoding / decoding. For example, after confirming information about the shape of the image (for example, rectangular or non-rectangular shape), information for the image setting based thereon can be configured.
[0098] In the examples described below, it is assumed that settings dependent on the color space are made for the explanation, but it is also possible to make settings independent of the color space. Further, in the case of independent settings in the examples described below, examples including encoding / decoding settings independently in each color space can be included, and even if an explanation is given for one color space, it is assumed that examples applicable to other color spaces (for example, when generating M for the luminance component, generating N for the color difference component) can be included and can be induced therefrom. Further, in the case of dependent settings, examples including settings proportional to the composition ratio of the color format (for example, 4:4:4, 4:2:2, 4:2:0, etc.) (for example, in the case of 4:2:0, when generating M for the luminance component, generating M / 2 for the color difference component) can be included, and it is assumed that examples applicable to each color space can be included and can be induced therefrom without special explanation. This is not limited to the above examples and can be an explanation commonly applicable to the present invention.
[0099] Some configurations in the examples described below may be applicable to various encoding techniques such as encoding in the spatial domain, encoding in the frequency domain, block-based encoding, and object-based encoding.
[0100] It may be common to perform encoding / decoding as the input image is, but there may also be cases where the image is divided for encoding / decoding. For example, it can be divided for error tolerance or the like for the purpose of preventing damage due to packet loss or the like during transmission. Or, it can be divided for the purpose of classifying regions having different properties within the same image according to the characteristics, types, etc. of the image.
[0101] In the present invention, the image division process can include the division process and the reverse process thereof. In the examples described below, the division process will be mainly explained, but the content regarding the division reverse process can be induced inversely from the division process.
[0102] FIG. 3 is an exemplary diagram showing hierarchical separation of image information for compressing an image.
[0103] 3a is an exemplary diagram showing a sequence of images composed of a number of GOPs. One GOP can be composed of an I picture, a P picture, and a B picture as shown in 3b. One picture can be composed of slices, tiles, etc. as shown in 3c. Slices, tiles, etc. are composed of a number of basic encoding units as shown in 3d, and a basic encoding unit can be composed of at least one sub-encoding unit as shown in FIG. 3e. The image setting process in the present invention will be described based on examples applied to units such as pictures, slices, and tiles as shown in 3b and 3c.
[0104] FIG. 4 is a conceptual diagram showing various examples of image segmentation according to an embodiment of the present invention.
[0105] 4a is a conceptual diagram showing an image (for example, a picture) divided at regular intervals in the horizontal and vertical directions. The divided regions can be called blocks, and each block is a basic encoding unit (or maximum encoding unit) obtained through a picture division unit and can also be a basic unit applied in the division units described later.
[0106] 4b is a conceptual diagram showing an image divided in at least one of the horizontal and vertical directions. The divided regions T0 to T3 can be called tiles, and each region can perform independent or dependent encoding / decoding with respect to other regions.
[0107] 4c is a conceptual diagram showing an image divided into groups of consecutive blocks. The divided regions S0 and S1 can be called slices, and each region can be a region that performs independent or dependent encoding / decoding with respect to other regions. The group of consecutive blocks can be defined according to the scan order, and generally follows the raster scan order, but is not limited thereto and can be determined according to the settings of encoding / decoding.
[0108] 4d is a conceptual diagram that divides an image into groups of blocks with arbitrary settings defined by the user. The divided regions A0 to A2 can be called arbitrary partition regions, and each region can be a region that performs independent or dependent encoding / decoding with respect to other regions.
[0109] Independent encoding / decoding can mean that when encoding / decoding some units (or regions), it is not possible to refer to the data of other units. Specifically, the information used or generated in the texture encoding and entropy encoding of some units is encoded independently without mutual reference, and in the decoder, for the texture decoding and entropy decoding of some units, the parsing information and restoration information of other units do not have to refer to each other. At this time, whether it is possible to refer to the data of other units (or regions) can be restricted in the spatial region (for example, between regions within one image), but depending on the encoding / decoding settings, restricted settings can also be placed in the temporal region (for example, between consecutive images or frames). For example, if a partial unit of the current image and a partial unit of another image have continuity or the same encoding environment, they can be referred to, and if not, the reference can be restricted.
[0110] Also, dependent encoding / decoding can mean that when encoding / decoding some units, it is possible to refer to the data of other units. Specifically, the information used or generated in the texture encoding and entropy encoding of some units is mutually referred to and encoded dependently, and in the decoder, similarly, for the texture decoding and entropy decoding of some units, the parsing information and restoration information of other units can be referred to each other. That is, it can be the same or a similar setting as general encoding / decoding. In this case, depending on the characteristics, types, etc. of the image (for example, 360-degree image), the region (in this example, the surface generated according to the projection format) <face>It may be the case when it is divided for the purpose of identifying (such as).
[0111] Independent encoding / decoding settings (e.g., independent slice segments) can be placed on some units (such as slices, tiles) in the above example, and dependent encoding / decoding settings (e.g., dependent slice segments) can be placed on some units. In the present invention, the description will be centered around independent encoding / decoding settings.
[0112] The basic encoding unit obtained through the picture division unit as in 4a is divided into basic encoding blocks according to the color space, and its size and shape can be determined according to the characteristics and resolution of the image, etc. The size or shape of the supported block is an N×N square (2 n ) expressed by an exponent of 2 (2 n ×2 n . Such as 256×256, 128×128, 64×64, 32×32, 16×16, 8×8, etc. n is an integer between 3 and 8), or it can be a rectangle of M×N (2 m ×2 n ). For example, according to the resolution, in the case of an 8k UHD-class image, the input image can be divided into sizes such as 128×128, in the case of a 1080p HD-class image, 64×64, and in the case of a WVGA-class image, 16×16. According to the type of image, in the case of a 360-degree image, the input image can be divided into a size of 256×256. The basic encoding unit can be divided into lower-level encoding units for encoding / decoding, and the information for the basic encoding unit can be recorded and transmitted in a bitstream in units such as sequences, pictures, slices, tiles, etc. This can be parsed by a decoder to restore relevant information.
[0113] An image encoding method and a decoding method according to an embodiment of the present invention may include the following image segmentation stages. At this time, the image segmentation process may include an image segmentation instruction stage, an image segmentation type identification stage, and an image segmentation execution stage. Further, an image encoding apparatus and a decoding apparatus may be configured to include an image segmentation instruction unit, an image segmentation type identification unit, and an image segmentation execution unit that realize the image segmentation instruction stage, the image segmentation type identification stage, and the image segmentation execution stage. In the case of encoding, associated syntax elements can be generated, and in the case of decoding, associated syntax elements can be parsed.
[0114] In each block division process of 4a, the image segmentation instruction unit may be omitted, and the image segmentation type identification unit is a process of checking information regarding the size and shape of the block, and division can be performed in a basic encoding unit by the image segmentation unit via the identified division type information.
[0115] In the case of a block, it can always be a unit in which division is performed, but for other division units (such as tiles, slices, etc.), it can be determined whether to divide according to the encoding / decoding settings. The picture segmentation unit can be basically set to perform division of other units after performing division in block units. At this time, block division can be performed based on the size of the picture.
[0116] Also, division can be performed in block units after division in other units (such as tiles, slices, etc.). That is, block division can be performed based on the size of the division unit. This can be determined through explicit or implicit processing according to the encoding / decoding settings. In the example described later, the former case is assumed and the description is centered on units other than blocks.
[0117] In the image segmentation instruction stage, it can be determined whether to perform image segmentation. For example, when a signal for instructing image segmentation (for example, tiles_enabled_flag) is confirmed, division can be performed, and when a signal for instructing image segmentation is not confirmed, division may not be performed or other encoding / decoding information may be confirmed to perform division.
[0118] Specifically, a signal for instructing image segmentation (e.g., tiles_enabled_flag) is checked. When the corresponding signal is activated (e.g., tiles_enabled_flag = 1), segmentation can be performed in multiple units. When the corresponding signal is deactivated (e.g., tiles_enabled_flag = 0), segmentation cannot be performed. Or, when the signal for instructing image segmentation is not checked, it can mean not performing segmentation or performing segmentation in at least one unit. Whether to perform segmentation in multiple units can be checked via another signal (e.g., first_slice_segment_in_pic_flag).
[0119] In summary, when a signal for instructing image segmentation is provided, the corresponding signal is a signal for indicating whether to perform segmentation in multiple units, and it is possible to check whether to segment the corresponding image according to the signal. For example, when tiles_enabled_flag is a signal indicating whether to perform image segmentation, when tiles_enabled_flag is 1, it can mean that the image is segmented into multiple tiles, and when it is 0, it can mean that the image is not segmented.
[0120] In summary, when a signal for instructing image segmentation is not provided, segmentation is not performed, or whether to segment the corresponding image can be checked by another signal. For example, first_slice_segment_in_pic_flag is not a signal indicating whether to perform image segmentation, but a signal indicating whether it is the first slice segment in the image. However, based on this, it is possible to check whether to perform segmentation into two or more units (e.g., when the flag is 0, it means that the image is segmented into multiple slices).
[0121] It is not limited to only the case of the above example, and deformation to other examples is also possible. For example, a signal instructing image segmentation may not be provided in the tile, or a signal instructing image segmentation may be provided in the slice. Or, according to the type, characteristics, etc. of the image, a signal instructing image segmentation may be provided.
[0122] In the image segmentation type identification stage, the image segmentation type can be identified. The image segmentation type can be defined by the method of performing the segmentation, segmentation information, etc.
[0123] In 4b, the tile can be defined as a unit obtained by dividing in the horizontal and vertical directions. Specifically, it can be defined as a group of adjacent blocks within a quadrilateral space partitioned by at least one horizontal or vertical division line across the image.
[0124] The segmentation information for the tile can include boundary position information between rows and columns, the number of tiles in rows and columns, tile size information, etc. The number information of the tiles can include the number of tiles in the horizontal row (for example, num_tile_columns) and the number of tiles in the vertical column (for example, num_tile_rows), and thus can be divided into (the number of horizontal rows × the number of vertical columns) tiles. The tile size information can be obtained based on the number information of the tiles, but the horizontal width or vertical height of the tile can be equal or unequal. This can be implicitly determined under a preset rule, or relevant information (for example, uniform_spacing_flag) can be explicitly generated. Also, the tile size information can include the size information of each horizontal row and vertical column of the tile (for example, column_width_tile[i], row_height_tile[i]), or can include the vertical height and horizontal width information of each tile. Also, the size information may be information that can be further generated according to whether the tile size is equal or not (for example, when uniform_spacing_flag is 0, meaning non-uniform division).
[0125] In 4c, a slice can be defined as a group unit of consecutive blocks. Specifically, it can be defined as a group of consecutive blocks based on a predetermined scan order (in this example, a raster scan).
[0126] The splitting information for a slice can include the number information of the slices, the position information of the slices (e.g., slice_segment_address), etc. At this time, the position information of the slice can be the position information of a predetermined (e.g., the first order in the scan order within the slice) block. At this time, the position information can be the scan order information of the block.
[0127] In 4d, various splitting settings are possible for any splitting area.
[0128] The splitting unit in 4d can be defined as a group of spatially adjacent blocks, and the splitting information for this can include the size, shape, position information, etc. of the splitting unit. This is some examples for any splitting area, and various splitting shapes are possible as shown in FIG. 5.
[0129] FIG. 5 is another exemplary diagram of the image splitting method according to an embodiment of the present invention.
[0130] In the cases of 5a and 5b, the image can be split into a plurality of regions with at least one block interval in the horizontal or vertical direction, and the splitting can be performed based on the position information of the blocks. 5a shows examples A0 and A1 where the splitting is performed based on the vertical column information of each block in the horizontal direction, and 5b shows examples B0 to B3 where the splitting is performed based on the horizontal row and vertical column information of each block in the vertical and horizontal directions. The splitting information for this can include the number of splitting units, block interval information, splitting direction, etc., and when this is implicitly included according to a predetermined rule, some splitting information may not be generated.
[0131] In the cases of 5c and 5d, based on the scanning order, the image can be divided into groups of consecutive blocks. Additional scanning orders other than the existing raster scan order of slices can be applied to image division. 5c shows examples C0 and C1 where scanning (Box-Out) is performed clockwise or counterclockwise around the starting block, and 5d shows examples D0 and D1 where vertical scanning (Vertical) is performed around the starting block. The division information for this can include the number information of the division units, the position information of the division units (for example, the first order in the scanning order within the division unit), information regarding the scanning order, etc. When some division information is implicitly included according to a predetermined rule, some division information may not be generated.
[0132] In the case of 5e, the image can be divided by horizontal and vertical dividing lines. The existing tiles can be divided by horizontal or vertical dividing lines, and thereby can have a divided shape of a rectangular space, but there may be cases where the division across the image by the dividing lines is not possible. For example, an example of dividing across the image along a partial dividing line of the image (for example, the dividing line formed by the right boundary of E1, E3, E4 and the left boundary of E5) is possible, while an example of dividing across the image along a partial dividing line of the image (for example, the dividing line formed by the lower boundary of E2 and E3 and the upper boundary of E4) is not possible. Also, division can be performed based on block units (for example, after block division is first performed and then division), or division can be performed by the said horizontal or vertical dividing lines, etc. (for example, division by the said dividing lines regardless of block division), and thereby each division unit may not be composed of an integer multiple of blocks. Therefore, division information different from the existing tiles can be generated, and the division information for this can include the number information of the division units, the position information of the division units, the size information of the division units, etc. For example, the position information of the division units can generate position information (measured in pixel units or block units) based on a predetermined position (for example, the upper left corner of the image), and the size information of the division units can generate the horizontal and vertical size information of each division unit (measured in pixel units or block units).
[0133] The splitting with arbitrary settings defined by the user as in the above example can be performed by applying a new splitting method or by applying a change to a partial configuration of an existing split. That is, it may be supported by replacing the existing splitting method or by an additional splitting shape, or it may be supported in a form where some settings are applied with changes to the existing splitting method (such as slicing, tiling, etc.) (for example, following other scan orders, other splitting methods with rectangular shapes and generation of other splitting information, dependent encoding / decoding characteristics, etc.). Also, settings for constituting additional splitting units (for example, settings other than splitting according to the scan order or splitting according to a difference at regular intervals) can be supported, and it is also possible to support additional splitting unit shapes (for example, polygonal shapes such as triangles other than splitting into a rectangular space). Also, it is possible to support an image splitting method based on the type, characteristics, etc. of the image. For example, some splitting methods (such as the surface of a 360-degree image) can be supported according to the type, characteristics, etc. of the image, and splitting information can be generated based on this.
[0134] In the image splitting execution stage, the image can be split based on the identified split type information. That is, splitting can be performed into a plurality of splitting units based on the identified split type, and encoding / decoding can be performed based on the obtained splitting units.
[0135] At this time, it is possible to determine whether to have encoding / decoding settings for the splitting units according to the split type. That is, the setting information required for the encoding / decoding process of each splitting unit can receive an assignment at a higher-level unit (for example, a picture), or can have an independent encoding / decoding setting for the splitting unit.
[0136] Generally, in the case of slices, it can have independent encoding / decoding settings (e.g., slice headers) for the division units. In the case of tiles, it cannot have independent encoding / decoding settings for the division units and can have settings that are dependent on the encoding / decoding settings of the picture (e.g., PPS). At this time, the information generated in relation to the tile can be division information. This can be included in the encoding / decoding settings of the picture. In the present invention, it is not limited only to the cases described above, and other modified examples are possible.
[0137] Encoding / decoding setting information for tiles can be generated in units such as video, sequence, and picture. At least one encoding / decoding setting information can be generated at a higher unit, and any one of them can be referred to. Or, independent encoding / decoding setting information (e.g., tile header) can be generated in tile units. This is different from conforming to one encoding / decoding setting determined at a higher unit in that encoding / decoding is performed with at least one encoding / decoding setting in tile units. That is, it is possible to conform to one encoding / decoding setting for all tiles, or encoding / decoding can be performed according to an encoding / decoding setting different from that of other tiles for at least one tile.
[0138] Although the above examples have been mainly described with respect to various encoding / decoding settings for tiles, it is not limited thereto, and similar or identical settings can be made for other division types.
[0139] As an example, for some division types, division information can be generated at a higher unit, and encoding / decoding can be performed according to one encoding / decoding setting of the higher unit.
[0140] As an example, for some division types, division information can be generated at a higher unit, and independent encoding / decoding settings for each division unit can be generated at the higher unit, whereby encoding / decoding can be performed.
[0141] As an example, for some split types, split information can be generated at a higher unit, a plurality of encoding / decoding setting information can be supported at the higher unit, and encoding / decoding can be performed according to the encoding / decoding settings referred to in each split unit.
[0142] As an example, for some split types, split information can be generated at a higher unit, independent encoding / decoding settings can be generated in the corresponding split unit, and thereby encoding / decoding can be performed.
[0143] As an example, for some split types, independent encoding / decoding settings including split information can be generated in the corresponding split unit, and thereby encoding / decoding can be performed.
[0144] The encoding / decoding setting information can include information necessary for encoding / decoding of tiles such as the type of tile, information regarding the reference picture list, quantization parameter information, inter-picture prediction setting information, in-loop filtering setting information, in-loop filtering control information, scan order, and whether to perform encoding / decoding. The encoding / decoding setting information can explicitly generate related information, or it is also possible that the settings for encoding / decoding are implicitly determined according to the format, characteristics, etc. of the image determined at a higher unit. Also, related information can be explicitly generated based on the information obtained in the above settings.
[0145] Hereinafter, an example of performing image splitting by an encoding / decoding apparatus according to an embodiment of the present invention will be shown.
[0146] Before the start of encoding, the splitting process for the input image can be performed. After splitting using split information (for example, image splitting information, split unit setting information, etc.), the image can be encoded in split units. After the completion of encoding, it can be stored in a memory, and the image encoded data can be recorded in a bitstream and transmitted.
[0147] The splitting process can be performed before the start of decryption. After splitting using splitting information (e.g., image splitting information, splitting unit setting information, etc.), the encrypted image data can be parsed and decrypted in units of splitting. After the completion of decryption, it can be stored in memory, and a plurality of splitting units can be merged into one to output an image.
[0148] The splitting process of the image has been described using the above example. Also, in the present invention, a plurality of splitting processes can be performed.
[0149] For example, an image can be split, and the splitting unit of the image can be split. The splitting can be the same splitting process (e.g., slice / slice, tile / tile, etc.) or different splitting processes (e.g., slice / tile, tile / slice, tile / surface, surface / tile, slice / surface, surface / slice, etc.). At this time, based on the previous splitting result, the subsequent splitting process can be performed. The splitting information generated in the subsequent splitting process can be generated based on the previous splitting result.
[0150] Also, a plurality of splitting processes A can be performed, and the splitting processes can be different splitting processes (e.g., slice / surface, tile / surface, etc.). At this time, based on the previous splitting result, the subsequent splitting process can be performed, or the splitting process can be performed independently regardless of the previous splitting result. The splitting information generated in the subsequent splitting process can be generated based on the previous splitting result or independently.
[0151] The plurality of splitting processes of the image can be determined according to the encoding / decoding settings, and are not limited to the above examples, and various modified examples are also possible.
[0152] The symbolizer records the information generated in the above process in a bitstream in at least one unit among units such as sequences, pictures, slices, tiles, etc., and the decoder parses the relevant information from the bitstream. That is, it can be recorded in one unit and can be recorded repeatedly in multiple units. For example, syntax elements regarding whether or not to support some information, or syntax elements regarding activation or not, etc. can be generated in some units (for example, upper-level units), and the same or similar information as in the above case can be generated in some units (for example, lower-level units). That is, even when relevant information is supported and set in the upper-level unit, individual settings in the lower-level unit can be held. This is not limited to the above example and can be an explanation commonly applied in the present invention. Also, it can be included in the bitstream in the form of SEI or metadata.
[0153] On the other hand, although it is common to perform encoding / decoding as the input image is, it can also occur that encoding / decoding is performed after adjusting the size of the image (expanding or shrinking. Adjusting the resolution). For example, in a hierarchical coding method (Scalability Video Coding) that supports spatial, temporal, and picture quality scalability, image size adjustment such as overall expansion or contraction of the image can be performed. Or, image size adjustment such as partial expansion or contraction of the image can also be performed. Image size adjustment is possible for various purposes, and it may be performed for the purpose of adaptability to the encoding environment, for the purpose of encoding uniformity, for the purpose of encoding efficiency, for the purpose of image quality improvement, or may be performed according to the type, characteristics, etc. of the image.
[0154] As a first example, a size adjustment process can be performed in a process (for example, hierarchical coding, 360-degree image coding, etc.) performed according to the characteristics, type, etc. of the image.
[0155] As a second example, a size adjustment process can be performed at the initial stage of encoding / decoding. A size adjustment process can be performed before encoding / decoding. The image whose size is adjusted can be encoded / decoded.
[0156] As a third example, a size adjustment process may be performed before the prediction stage (intra-picture prediction or inter-picture prediction) or before the prediction is executed. In the size adjustment process, image information in the prediction stage (for example, pixel information referred to in intra-picture prediction, intra-picture prediction mode related information, reference image information used in inter-picture prediction, inter-picture prediction mode related information, etc.) can be used.
[0157] As a fourth example, a size adjustment process may be performed before the filtering stage or before filtering is performed. In the size adjustment process, image information in the filtering stage (for example, pixel information applied to the deblocking filter, pixel information applied to SAO, SAO filtering related information, pixel information applied to ALF, ALF filtering related information, etc.) can be used.
[0158] Also, after the size adjustment process is performed, the image may or may not be changed back to the image before size adjustment (from the perspective of the image size) through the reverse size adjustment process. This can be determined according to the encoding / decoding settings (for example, the nature of size adjustment). At this time, if the size adjustment process is an expansion, the reverse size adjustment process is a reduction, and if the size adjustment process is a reduction, the reverse size adjustment process may be an expansion.
[0159] When the size adjustment process according to the first to fourth examples is performed, the image before size adjustment can be obtained by performing the reverse size adjustment process at a later stage.
[0160] When hierarchical encoding or the size adjustment process according to the third example is performed (or when the size of the reference image is adjusted in inter-picture prediction), the reverse size adjustment process may not be performed at a later stage.
[0161] In one embodiment of the present invention, the image size adjustment process can be performed independently, or the reverse process thereof can be carried out. In the examples described later, the description will focus on the size adjustment process. At this time, since the size adjustment reverse process is the opposite process of the size adjustment process, in order to avoid duplicate explanations, the description of the size adjustment reverse process can be omitted, but it is obvious that an ordinary technician can recognize it in the same way as described verbally.
[0162] FIG. 6 is an exemplary diagram of a general image size adjustment method.
[0163] Referring to 6a, an expanded image P0 + P1 can be obtained by further including a partial region P1 from the initial image (or the image before size adjustment. P0. thick solid line).
[0164] Referring to 6b, a reduced image S0 can be obtained by excluding a partial region S1 from the initial image S0 + S1.
[0165] Referring to 6c, a size-adjusted image T0 + T1 can be obtained by further including a partial region T1 in the initial image T0 + T2 and excluding a partial region T2.
[0166] Hereinafter, in the present invention, the size adjustment process by expansion and the size adjustment process by reduction will be mainly described, but it is not limited thereto, and it should be understood that cases where size expansion and reduction are mixed and applied as in 6c are also included.
[0167] FIG. 7 is an exemplary diagram of image size adjustment according to an embodiment of the present invention.
[0168] Referring to 7a, the method of expanding an image in the size adjustment process can be described, and referring to 7b, the method of reducing an image can be described.
[0169] In 7a, the image before size adjustment is S0, and the image after size adjustment is S1. In 7b, the image before size adjustment is T0, and the image after size adjustment is T1.
[0170] When expanding the image as in 7a, it can be expanded in the up, down, left, and right directions (ET, EL, EB, ER). When shrinking the image as in 7b, it can be shrunk in the up, down, left, and right directions (RT, RL, RB, RR).
[0171] Comparing the expansion and shrinking of the image, since the up, down, left, and right directions in expansion can correspond to the respective down, up, right, and left directions in shrinking, hereinafter, the explanation will be based on the expansion of the image, but it should be understood that the explanation of the shrinking of the image is also included.
[0172] Also, hereinafter, the expansion or shrinking of the image in the up, down, left, and right directions will be described, but it should be understood that the size adjustment can be performed in the upper left, upper right, lower left, and lower right directions.
[0173] At this time, when expanding in the lower right direction, the RC and BC regions are acquired. Depending on the encoding / decoding settings, it may or may not be possible to acquire the BR region. That is, it may or may not be possible to acquire the TL, TR, BL, and BR regions. Hereinafter, for the sake of convenience of explanation, it will be described that the corner regions (TL, TR, BL, BR regions) can be acquired.
[0174] The process of adjusting the size of the image according to an embodiment of the present invention can be performed in at least one direction. For example, it may be performed in all of the up, down, left, and right directions, or in two or more selected directions (left + right, up + down, up + left, up + right, down + left, down + right, up + left + right, down + left + right, up + down + left, up + down + right, etc.) selected from the up, down, left, and right directions, or it may be performed in only any one of the up, down, left, and right directions.
[0175] For example, it may be possible to adjust the size in the left + right, up + down, upper left + lower right, and lower left + upper right directions that can be symmetrically extended to both ends based on the center of the image. It may also be possible to adjust the size in the left + right, upper left + upper right, and lower left + lower right directions where the image can be vertically symmetrically extended. It may also be possible to adjust the size in the up + down, upper left + lower left, and upper right + lower right directions where the image can be horizontally symmetrically extended. Other size adjustments are also possible.
[0176] In 7a and 7b, the size of the image (S0, T0) before size adjustment is defined as P_Width (width) × P_Height (height), and the size of the image after size adjustment (S1, T1) is defined as P’_Width (width) × P’_Height (height). Here, if the size adjustment values in the left, right, up, and down directions are defined as Var_L, Var_R, Var_T, Var_B (or collectively called Var_x), the size of the image after size adjustment can be expressed as (P_Width + Var_L + Var_R) × (P_Height + Var_T + Var_B). At this time, Var_L, Var_R, Var_T, and Var_B, which are the size adjustment values in the left, right, up, and down directions, are Exp_L, Exp_R, Exp_T, Exp_B (in this example, Exp_x is a positive number) in image expansion (Figure 7a), and can be -Rec_L, -Rec_R, -Rec_T, -Rec_B (when Rec_L, Rec_R, Rec_T, Rec_B are defined as positive numbers, they are expressed as negative numbers according to the reduction of the image) in image reduction. Also, the coordinates of the upper left, upper right, lower left, and lower right of the image before size adjustment are (0, 0), (P_Width - 1, 0), (0, P_Height - 1), (P_Width - 1, P_Height - 1), and the coordinates of the image after size adjustment can be expressed as (0, 0), (P’_Width - 1, 0), (0, P’_Height - 1), (P’_Width - 1, P’_Height - 1). The size of the area (in this example, TL~BR. i is the index for distinguishing TL~BR) changed (or acquired, deleted) by size adjustment can be M[i] × N[i]. This can be expressed as Var_X × Var_Y (in this example, it is assumed that X is L or R, and Y is T or B). M and N can have various values, and can be the same regardless of i, or can have individual settings according to i. Various cases of this will be described later.
[0177] Referring to 7a, S1 can be composed of including all or part of TL~BR (upper left to lower right) generated by expanding S0 in various directions. Referring to 7b, T1 can be composed of excluding all or part of TL~BR removed by reducing T0 in various directions.
[0178] In 7a, when expanding the existing image S0 in the up, down, left, and right directions, the image can be composed including the TC, BC, LC, and RC regions obtained through each resizing process, and further the TL, TR, BL, and BR regions can also be included.
[0179] As an example, when expanding in the up (ET) direction, the image can be composed including the TC region in the existing image S0, and the TL or TR region can be included according to the expansion in at least one different direction (EL or ER).
[0180] As an example, when expanding in the down (EB) direction, the image can be composed including the BC region in the existing image S0, and the BL or BR region can be included according to the expansion in at least one different direction (EL or ER).
[0181] As an example, when expanding in the left (EL) direction, the image can be composed including the LC region in the existing image S0, and the TL or BL region can be included according to the expansion in at least one different direction (ET or EB).
[0182] As an example, when expanding in the right (ER) direction, the image can be composed including the RC region in the existing image S0, and the TR or BR region can be included according to the expansion in at least one different direction (ET or EB).
[0183] According to an embodiment of the present invention, a setting (for example, spa_ref_enabled_flag or tem_ref_enabled_flag) can be provided to restrict the referability of the region to be resized (assumed to be expanded in this example) spatially or temporally.
[0184] That is, it is possible to refer to the data of the region whose size is adjusted spatially or temporally according to the encoding / decoding setting (for example, spa_ref_enabled_flag = 1 or tem_ref_enabled_flag = 1), or to limit the reference (for example, spa_ref_enabled_flag = 0 or tem_ref_enabled_flag = 0).
[0185] The encoding / decoding of the image before size adjustment (S0, T1) and the regions added or deleted during size adjustment (TC, BC, LC, RC, TL, TR, BL, BR regions) can be performed as follows.
[0186] For example, in the encoding / decoding of the image before size adjustment and the regions added or deleted, the data of the image before size adjustment and the data of the regions added or deleted (encoded / decoded data, such as pixel values or prediction-related information) can be referred to each other spatially or temporally.
[0187] Alternatively, while the data of the image before size adjustment and the data of the regions added or deleted can be referred to spatially, the data of the image before size adjustment can be referred to temporally, and the data of the regions added or deleted cannot be referred to temporally.
[0188] That is, it is possible to set a setting that limits the referability of the regions added or deleted. The setting information regarding the referability of the regions added or deleted can be generated explicitly or determined implicitly.
[0189] The image size adjustment process according to an embodiment of the present invention may include an image size adjustment instruction stage, an image size adjustment type identification stage, and / or an image size adjustment execution stage. Further, the image encoding device and the decoding device may include an image size adjustment instruction unit, an image size adjustment type identification unit, and an image size adjustment execution unit that realize the image size adjustment instruction stage, the image size adjustment type identification stage, and the image size adjustment execution stage. In the case of encoding, associated syntax elements can be generated, and in the case of decoding, the associated syntax elements can be parsed.
[0190] In the image size adjustment instruction stage, it is possible to determine whether to perform image size adjustment. For example, when a signal indicating image size adjustment (e.g., img_resizing_enabled_flag) is confirmed, size adjustment can be performed. When the signal indicating image size adjustment is not confirmed, size adjustment may not be performed or other encoding / decoding information may be confirmed to perform size adjustment. Also, even if a signal indicating image size adjustment is not provided, depending on the encoding / decoding settings (e.g., image characteristics, type, etc.), the signal indicating size adjustment may be implicitly activated or deactivated. When performing size adjustment, size adjustment related information can be generated thereby, or it is also possible that the size adjustment related information is implicitly determined.
[0191] When a signal indicating image size adjustment is provided, the corresponding signal is a signal for indicating whether to perform image size adjustment of the image, and it is possible to confirm whether to perform size adjustment of the corresponding image according to the signal.
[0192] For example, when a signal indicating image size adjustment (e.g., img_resizing_enabled_flag) is confirmed and the corresponding signal is activated (e.g., img_resizing_enabled_flag = 1), image size adjustment can be performed. When the corresponding signal is deactivated (e.g., img_resizing_enabled_flag = 0), it can be meant that image size adjustment is not performed.
[0193] Also, when a signal instructing image size adjustment is not provided, size adjustment is not performed, or whether to perform size adjustment on the image can be confirmed by other signals regarding whether to perform size adjustment.
[0194] For example, when dividing the input image in block units, size adjustment (in this example, in the case of expansion. Assume that the size adjustment process is performed when it is not an integer multiple) can be performed according to whether the size of the image (e.g., width or height) is an integer multiple of the size of the block (e.g., width or height). That is, when the width of the image is not an integer multiple of the width of the block, or when the height of the image is not an integer multiple of the height of the block, size adjustment can be performed. At this time, the size adjustment information (e.g., size adjustment direction, size adjustment value, etc.) can be determined according to the encoding / decoding information (e.g., size of the image, size of the block, etc.). Or, size adjustment can be performed according to the characteristics, type of the image (e.g., 360-degree image, etc.), and the size adjustment information can be explicitly generated or assigned to a predetermined value. It is not limited to only the above examples, and deformation to other examples is also possible.
[0195] In the image size adjustment type identification stage, the image size adjustment type can be identified. The image size adjustment type can be defined by the method of performing size adjustment, size adjustment information, etc. For example, size adjustment using a scale factor, size adjustment using an offset factor, etc. can be performed. It is not limited to this, and mixed application of the above methods is also possible. For the convenience of explanation, size adjustment using a scale factor and an offset factor will be mainly described.
[0196] In the case of a scale factor, size adjustment can be performed in a manner of multiplication or division based on the size of the image. Information about the size adjustment operation (e.g., expansion or reduction) can be explicitly generated, and the expansion or reduction process can be performed according to the corresponding information. Also, according to the encoding / decoding settings, the size adjustment process can be performed with a predetermined operation (e.g., either expansion or reduction). In this case, information about the size adjustment operation can be omitted. For example, when image size adjustment is activated at the image size adjustment instruction stage, the size adjustment of the image can be performed with a predetermined operation.
[0197] The size adjustment direction can be at least one direction selected from the up, down, left, and right directions. Depending on the size adjustment direction, at least one scale factor may be required. That is, one scale factor (in this example, unidirectional) is required for each direction, one scale factor (in this example, bidirectional) is required according to the horizontal or vertical direction, and one scale factor (in this example, omni-directional) may be required according to the overall direction of the image. Also, the size adjustment direction is not limited to the case of the above example, and deformation to other examples is also possible.
[0198] The scale factor can have a positive value, and the range information can be set differently according to the encoding / decoding settings. For example, when generating information by mixing the size adjustment operation and the scale factor, the scale factor can be used as the value to be multiplied. When it is greater than 0 or less than 1, it can mean a reduction operation. When it is greater than 1, it can mean an expansion operation. When it is 1, it can mean not performing size adjustment. As another example, when generating scale factor information separately from the size adjustment operation, in the case of an expansion operation, the scale factor can be used as the value to be multiplied, and in the case of a reduction operation, the scale factor can be used as the value to be divided.
[0199] Referring back to FIGS. 7a and 7b, the process of changing from the pre-sizing image (S0, T0) to the post-sizing image (in this example, S1, T1) using the scale factor can be described.
[0200] As an example, when using one scale factor (referred to as sc) according to the overall direction of the image and the sizing direction is the down + right direction, the sizing directions are ER, EB (or RR, RB), and the sizing values Var_L (Exp_L or Rec_L) and Var_T (Exp_T or Rec_T) are 0, and Var_R (Exp_R or Rec_R) and Var_B (Exp_B or Rec_B) can be expressed as P_Width×(sc - 1), P_Height×(sc - 1). Therefore, the post-sizing image can be (P_Width×sc)×(P_Height×sc).
[0201] As an example, when using respective scale factors (in this example, sc_w, sc_h) according to the horizontal or vertical direction of the image and the sizing directions are the left + right, up + down directions (when both operate, it is up + down + left + right), the sizing directions are ET, EB, EL, ER, and the sizing values Var_T and Var_B are P_Height×(sc_h - 1) / 2, and Var_L and Var_R can be P_Width×(sc_w - 1) / 2. Therefore, the post-sizing image can be (P_Width×sc_w)×(P_Height×sc_h).
[0202] In the case of the offset factor, sizing can be performed by adding or subtracting based on the size of the image. Or, sizing can be performed by adding or subtracting based on the encoding / decoding information of the image. Or, sizing can be performed by adding or subtracting independently. That is, the sizing process can be set dependently or independently.
[0203] Information about the sizing operation (e.g., expansion or reduction) can be explicitly generated, and the expansion or reduction process can be performed according to the corresponding information. Also, according to the encoding / decoding settings, the sizing operation can be performed with a predetermined operation (e.g., either expansion or reduction). In this case, the information about the sizing operation can be omitted. For example, when image sizing is activated at the image sizing instruction stage, the sizing of the image can be performed with a predetermined operation.
[0204] The sizing direction can be at least one of the up, down, left, and right directions. At least one offset factor may be required according to the sizing direction. That is, one offset factor (in this example, unidirectional) is required for each direction, one of the offset factors (in this example, symmetric bidirectional) is required according to the horizontal or vertical direction, one offset factor (in this example, asymmetric bidirectional) is required according to the partial combination of each direction, and one offset factor (in this example, omnidirectional) may be required according to the overall direction of the image. Also, the sizing direction is not limited only to the case of the above example, and variations to other examples are also possible.
[0205] The offset factor can have a positive value or can have both positive and negative values, and can be set with different range information according to the encoding / decoding settings. For example, when generating information by mixing the sizing operation and the offset factor (assuming in this example that it has both positive and negative values), the offset factor can be used as a value to be added or subtracted according to the sign information of the offset factor. When the offset factor is greater than 0, it can mean an expansion operation; when it is less than 0, it can mean a reduction operation; and when it is 0, it can mean not performing a sizing operation. As another example, when generating offset factor information separately from the sizing operation (assuming in this example that it has a positive value), the offset factor can be used as a value to be added or subtracted according to the sizing operation. When it is greater than 0, an expansion or reduction operation can be performed according to the sizing operation; and when it is 0, it can mean not performing a sizing operation.
[0206] Referring again to FIGS. 7a and 7b of FIG. 7, a method of changing from the pre-sized image (S0, T0) to the post-sized image (S1, T1) using the offset factor can be described.
[0207]
[0208] As an example, respective offset factors (os_w, os_h) are used according to the horizontal or vertical direction of the image. When the resizing directions are left + right, up + down directions (when both operate, up + down + left + right), the resizing directions are ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values Var_T, Var_B can be os_h, and Var_L, Var_R can be os_w. The image size after resizing can be {P_Width+(os_w×2)}×{P_Height+(os_h×2)}.
[0209] As an example, when the resizing directions are down, right directions (when operating together, down + right), and respective offset factors (os_b, os_r) are used according to the resizing directions, the resizing directions are EB, ER (or RB, RR), and the resizing value Var_B can be os_b, and Var_R can be os_r. The image size after resizing can be (P_Width+os_r)×(P_Height+os_b).
[0210] As an example, respective offset factors (os_t, os_b, os_l, os_r) are used according to each direction of the image. When the resizing directions are up, down, left, right directions (when all operate, up + down + left + right), the resizing directions are ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values Var_T can be os_t, Var_B can be os_b, Var_L can be os_l, and Var_R can be os_r. The image size after resizing can be (P_Width+os_l+os_r)×(P_Height+os_t+os_b).
[0211] The above example shows the case where the offset factor is used as a sizing value (Var_T, Var_B, Var_L, Var_R) in the sizing process. That is, the offset factor means the case where it is used as is as a sizing value. This can be an example of sizing performed independently. Or, the offset factor can also be used as an input variable of the sizing value. Specifically, the offset factor can be assigned as an input variable, and the sizing value can be obtained through a series of processes according to the encoding / decoding settings. This can be an example of sizing performed based on predetermined information (such as the size of an image, encoding / decoding information, etc.) or an example of sizing performed dependently.
[0212] For example, the offset factor can be a multiple (such as 1, 2, 4, 6, 8, 16, etc.) or an exponent (such as a power of 2 like 1, 2, 4, 8, 16, 32, 64, 128, 256, etc.) of a predetermined value (in this example, an integer). Or, it can be a multiple or an exponent of a value obtained based on the encoding / decoding settings (such as a value set based on the motion search range of inter-picture prediction). Or, it can be a multiple or an exponent of a unit (assumed to be A×B in this example) obtained from the picture partitioning section. Or, it can be a multiple of a unit (in the case of a tile, etc., assumed to be E×F in this example) obtained from the picture partitioning section.
[0213] Or, it can be a value in a range smaller than or equal to the horizontal and vertical of the unit obtained from the picture partitioning section. The multiples or exponents in the above examples can include the case where the value is 1, are not limited to the above examples, and can also be deformed to other examples. For example, when the offset factor is n, Var_x can be 2×n or 2 n and can be.
[0214] In addition, individual offset factors can be supported according to color components, and offset factors for some color components can be supported to derive offset factor information for other color components. For example, when an offset factor A for a luminance component (assuming that the composition ratio of the luminance component and the color difference component is 2:1 in this example) is explicitly generated, an offset factor A / 2 for the color difference component can be implicitly obtained. Or, when an offset factor A for the color difference component is explicitly generated, an offset factor 2A for the luminance component can be implicitly obtained.
[0215] Information about the size adjustment direction and size adjustment value can be explicitly generated, and the size adjustment process can be performed according to the corresponding information. Also, it can be implicitly determined according to the encoding / decoding settings, thereby enabling the size adjustment process. At least one predetermined direction or adjustment value can be assigned, in which case the relevant information can be omitted. At this time, the encoding / decoding settings can be determined based on the characteristics, types, encoding information, etc. of the image. For example, at least one size adjustment direction by at least one size adjustment operation can be predetermined, at least one size adjustment value by at least one size adjustment operation can be predetermined, and at least one size adjustment value by at least one size adjustment direction can be predetermined. Also, the size adjustment direction and size adjustment value, etc. in the reverse size adjustment process can be derived from the size adjustment direction and size adjustment value, etc. applied in the size adjustment process. At this time, the implicitly determined size adjustment value can be one of the above-mentioned examples (examples where various size adjustment values are obtained).
[0216] In addition, although the cases of multiplication or division in the above examples have been described, it can be realized by a shift operation according to the implementation of the encoder / decoder. When multiplying, it can be realized by using a left shift operation, and when dividing, it can be realized by using a right shift operation. This is not limited to the above examples and can be an explanation commonly applicable in the present invention.
[0217] In the image size adjustment execution stage, image size adjustment can be performed based on the identified size adjustment information. That is, image size adjustment can be performed based on information such as the size adjustment type, size adjustment operation, size adjustment direction, and size adjustment value, and encoding / decoding can be performed based on the obtained image after size adjustment.
[0218] Also, in the image size adjustment execution stage, size adjustment can be performed using at least one data processing method. Specifically, size adjustment can be performed using at least one data processing method on the area to be size-adjusted according to the size adjustment type and size adjustment operation. For example, according to the size adjustment type, it can be determined how to fill the data when the size adjustment is an expansion, or how to remove the data when the size adjustment process is a reduction.
[0219] In summary, image size adjustment can be performed based on the size adjustment information identified in the image size adjustment execution stage. Or, in the image size adjustment execution stage, image size adjustment can be performed based on the size adjustment information and the data processing method. The difference between the two cases lies in whether only the size of the image for encoding / decoding is adjusted, or whether the data processing of the size of the image and the area to be size-adjusted is also considered. Whether to include the data processing method in the image size adjustment execution stage can be determined according to the application stage, position, etc. of the size adjustment process. In the examples described later, examples of performing size adjustment based on the data processing method will be mainly described, but it is not limited to this.
[0220] When performing size adjustment using an offset factor, size adjustment can be performed using various methods in the case of expansion and reduction. In the case of expansion, size adjustment can be performed using a method of filling at least one piece of data. In the case of reduction, size adjustment can be performed using a method of removing at least one piece of data. At this time, in the case of size adjustment using an offset factor, new data or data of an existing image can be directly filled or filled after being deformed in the size adjustment area (expansion), and simple removal or removal through a series of processes can be applied and removed in the size adjustment area (reduction).
[0221] When performing size adjustment using a scale factor, in some cases (for example, hierarchical coding, etc.), expansion can be performed by applying upsampling for size adjustment, and reduction can be performed by applying downsampling for size adjustment. For example, in the case of expansion, at least one upsampling filter can be used. In the case of reduction, at least one downsampling filter can be used. The filters applied horizontally and vertically may be the same or different. At this time, in the case of size adjustment using a scale factor, rather than new data being generated or removed in the size adjustment area, it may be possible to relocate the data of the existing image using methods such as interpolation. The data processing method related to the execution of size adjustment can be classified by the filter used for the sampling. Also, in some cases (for example, in cases similar to the offset factor), expansion can be performed by using a method of filling at least one piece of data, and reduction can be performed by using a method of removing at least one piece of data. In the present invention, the data processing method in the case of performing size adjustment using an offset factor will be mainly described.
[0222] In general, a predetermined single data processing method can be used for the region to be resized, but at least one data processing method can also be used for the region to be resized as in the examples described later, and selection information for the data processing method can also be generated. In the former case, it can be meant that resizing is performed through a fixed data processing method, and in the latter case, through an adaptive data processing method.
[0223] Also, a data processing method common to the entire region (TL, TC, TR,..., BR in FIGS. 7a and 7b) added or deleted during resizing can be applied, or a data processing method can be applied in units of a part of the region added or deleted during resizing (for example, each of TL to BR in FIGS. 7a and 7b or a combination of a part thereof).
[0224] FIG. 8 is an exemplary diagram of a method for constructing an expanded region in an image resizing method according to an embodiment of the present invention.
[0225] Referring to FIG. 8a, the image can be divided into regions TL, TC, TR, LC, C, RC, BL, BC, BR for the sake of convenience of explanation, and can correspond to the upper left, upper, upper right, left, center, right, lower left, lower, and lower right positions of the image, respectively. Hereinafter, the case where the image is expanded in the lower + right direction will be described, but it should be understood that the same applies to other directions.
[0226] The region added according to the expansion of the image can be configured in various ways. For example, it can be filled with an arbitrary value or filled by referring to a part of the image data.
[0227] Referring to FIG. 8b, the regions (A0, A2) expanded to arbitrary pixel values can be filled. The arbitrary pixel value can be determined using various methods.
[0228] As an example, any pixel value can be one pixel belonging to the range of pixel values that can be represented by the bit depth {for example, from 0 to 1<<(bit_depth)-1}. For example, it can be the minimum value, maximum value, median value {such as 1<<(bit_depth-1), etc.} of the range of the pixel values (where bit_depth is the bit depth).
[0229] As an example, any pixel value can be one pixel belonging to the range of pixel values of the pixels belonging to the image {for example, from min P to max P up to. min P and max P are the minimum value and maximum value of the pixels belonging to the image. min P is equal to or greater than 0, and max P is equal to or smaller than 1<<(bit_depth)-1}. For example, any pixel value can be the minimum value, maximum value, median value, average (of at least two pixels), weighted sum, etc. of the range of the pixel values.
[0230] As an example, any pixel value can be a value determined by the range of pixel values belonging to a partial region belonging to the image. For example, when configuring A0, the partial region can be TR+RC+BR. Also, the partial region can be set as the corresponding region of 3×9 of TR, RC, BR, or can be set as the corresponding region of 1×9 <assuming the rightmost line>. This can be adjusted according to the encoding / decoding settings. At this time, the partial region may also be a unit divided from the picture division section. Specifically, any pixel value can be the minimum value, maximum value, median value, average (of at least two pixels), weighted sum, etc. of the range of the pixel values.
[0231] Referring back to 8b, the area A1 added according to the expansion of the image can be filled using pattern information (for example, assuming that what uses a plurality of pixels is a pattern. It is not necessarily required to follow a certain rule.) generated using a plurality of pixel values. At this time, the pattern information can be defined according to the encoding / decoding settings or can generate related information, and at least one pattern information can be used to fill the expanded area.
[0232] Referring to 8c, the area added according to the expansion of the image can be configured by referring to the pixels of a partial area belonging to the image. Specifically, the area to be added can be configured by copying or padding the pixels of the area adjacent to the area to be added (hereinafter, referred to as reference pixels). At this time, the pixels of the area adjacent to the area to be added can be the pixels before encoding or the pixels after encoding (or decoding). For example, when performing size adjustment in the pre-encoding stage, the reference pixels can mean the pixels of the input image, and when performing size adjustment in the in-picture prediction reference pixel generation stage, the reference image generation stage, the filtering stage, etc., the reference pixels can mean the pixels of the restored image. In this example, it is assumed that the pixels closest to the area to be added are used, but it is not limited thereto.
[0233] The area A0 expanded in the left or right direction related to the horizontal size adjustment of the image can be configured by horizontally padding (Z0) the outer pixels adjacent to the area A0 to be expanded, and the area A1 expanded in the up or down direction related to the vertical size adjustment of the image can be configured by vertically padding (Z1) the outer pixels adjacent to the area A1 to be expanded. Also, the area A2 expanded in the lower right direction can be configured by diagonally padding (Z2) the outer pixels adjacent to the area A2 to be expanded.
[0234] Referring to 8d, the expanded areas B'0 to B'2 can be configured by referring to the data B0 to B2 of a partial area belonging to the image. In 8d, it can be distinguished from 8c in that areas not adjacent to the expanded area can be referred to.
[0235] For example, when there is an area with a high correlation with the area to be expanded in the image, the area to be expanded can be filled by referring to the pixels of the area with a high correlation. At this time, the position information, area size information, etc. of the area with a high correlation can be generated. Or, when there is an area with a high correlation through encoding / decoding information such as the characteristics and types of the image, and the position information, size information, etc. of the area with a high correlation can be implicitly confirmed (for example, in the case of a 360-degree image), the data of the corresponding area can be filled into the area to be expanded. At this time, the position information, area size information, etc. of the area can be omitted.
[0236] As an example, in the case of the area B'2 expanded in the left or right direction related to the horizontal size adjustment of the image, the area to be expanded can be filled by referring to the pixels of the area B2 on the opposite side of the area to be expanded in the left or right direction related to the horizontal size adjustment.
[0237] As an example, in the case of the area B'1 expanded in the up or down direction related to the vertical size adjustment of the image, the area to be expanded can be filled by referring to the pixels of the area B1 on the opposite side of the area to be expanded in the up or right direction related to the vertical size adjustment.
[0238] As an example, in the case of the area B'0 expanded in the partial size adjustment of the image (in this example, in the diagonal direction based on the center of the image), the area to be expanded can be filled by referring to the pixels of the area B0, TL on the opposite side of the area to be expanded.
[0239] In the above example, the case of obtaining the data of the area with continuity at the boundaries at both ends of the image and symmetric to the size adjustment direction has been described, but it is not limited to this, and it is also possible to obtain it from the data of other areas TL~BR.
[0240] When filling the data of a partial area of an image into the area to be expanded, can the data of the corresponding area be directly copied and filled, or can it be filled after a conversion process based on the characteristics, type, etc. of the image? At this time, if it is directly copied, it can mean using the pixel values of the corresponding area as they are. If it goes through the conversion process, it can mean not using the pixel values of the corresponding area as they are. That is, through the conversion process, at least one pixel value of the area changes, and it may be filled into the area to be expanded, or the acquisition positions of some pixels may be at least one different. That is, to fill the area A×B to be expanded, instead of using the A×B data of the corresponding area, C×D data can be used. In other words, at least one of the motion vectors applied to the pixels to be filled may be different. The above example can be an example that occurs when filling the area to be expanded using the data of other surfaces when the 360-degree image is composed of multiple surfaces according to the projection format. The data processing method for filling the area expanded by resizing the image is not limited to the above example, and this can be improved and deformed, or additional data processing methods can be used.
[0241] According to the encoding / decoding settings, a candidate group for a plurality of data processing methods can be supported, and data processing method selection information can be generated from among the plurality of candidate groups and recorded in the bitstream. For example, one data processing method can be selected from methods such as filling using a predetermined pixel value, copying and filling the outer pixels, copying and filling a partial area of the image, and converting and filling a partial area of the image, and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0242] For example, the data processing method applied to the entire area (in this example, TL to BR in FIG. 7a) expanded by image size adjustment is one of the methods of filling with a predetermined pixel value, copying and filling the outer pixels, copying and filling a partial area of the image, converting and filling a partial area of the image, and other methods, and selection information for this can be generated. Also, one predetermined data processing method applied to the entire area can be determined.
[0243] Alternatively, the data processing method applied to the area expanded by image size adjustment (in this example, each area of TL to BR in FIG. 7a of FIG. 7 or two or more of them) is one of the methods of filling with a predetermined pixel value, copying and filling the outer pixels, copying and filling a partial area of the image, converting and filling a partial area of the image, and other methods, and selection information for this can be generated. Also, one predetermined data processing method applied to at least one area can be determined.
[0244] FIG. 9 is an exemplary diagram of a method of configuring an area to be deleted and an area to be generated by reducing the image size in an image size adjustment method according to an embodiment of the present invention.
[0245] The area deleted in the image reduction process can be not only simply removed but also removed after a series of utilization processes.
[0246] Referring to 9a, in the image reduction process, partial areas A0, A1, and A2 can be simply removed without an additional utilization process. At this time, the image A may be subdivided and called TL to BR as in FIG. 8a.
[0247] Referring to 9b, although some regions A0 to A2 are removed, they can be utilized as reference information during the encoding / decoding of image A. For example, in the restoration process or correction process of a partial region of the reduced and generated image A, the partially removed regions A0 to A2 can be utilized. In the said restoration or correction process, a weighted sum, average, etc. of two regions (the removed region and the generated region) can be used. Also, the said restoration or correction process can be a process applicable when the two regions have a high correlation.
[0248] As an example, the region B’2 removed by being reduced in the left or right direction related to the horizontal size adjustment of the image can be used to restore or correct the pixels of the region B2, LC on the opposite side of the reduced region among the left or right directions related to the horizontal size adjustment, and then the corresponding region can be removed from the memory.
[0249] As an example, the region B’1 removed in the up or down direction related to the vertical size adjustment of the image can be used in the encoding / decoding process (restoration or correction process) of the region B1, TR on the opposite side of the reduced region among the up or down directions related to the vertical size adjustment, and then the corresponding region can be removed from the memory.
[0250] As an example, the region B’0 reduced in the diagonal direction with respect to the center of the image in a partial size adjustment of the image (in this example) can be used in the encoding / decoding process (such as restoration or correction process) of the region B0, TL on the opposite side of the reduced region, and then the corresponding region can be removed from the memory.
[0251] In the above example, an explanation has been given regarding the case of using it for data restoration or correction of a region having continuity at the boundaries at both ends of the image and being at a position symmetric to the size adjustment direction. However, it is not limited to this, and it can also be removed from the memory after being used for data restoration or correction of other regions TL to BR other than the symmetric positions.
[0252] The data processing method for removing the reduced region of the present invention is not limited to the above example, and it can be improved and changed, or additional data processing methods can be used.
[0253] Candidate groups for a plurality of data processing methods can be supported according to the settings of symbolization / decryption, and selection information for this can be generated and recorded in the bitstream. For example, one data processing method can be selected from methods such as simply removing the area to be resized, removing the area to be resized after using it in a series of processes, etc., and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0254] For example, the data processing method applied to the entire area (in this example, TL~BR in FIG. 7b) that is deleted by being reduced by image resizing is any one of simply removing it, removing it after using it in a series of processes, and other methods, and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0255] Alternatively, the data processing method applied to the individual areas (in this example, each of TL~BR in FIG. 7b) that are reduced by image resizing is any one of simply removing it, removing it after using it in a series of processes, and other methods, and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0256] In the above example, the case regarding the execution of resizing by a resizing operation (expansion, reduction) has been described. However, in some cases, it can be an example applicable to the case where, after performing a resizing operation (in this example, expansion), a resizing operation (in this example, reduction), which is the reverse process thereof, is performed.
[0257] For example, a method of filling an expanded area with partial image data is selected, and in the reverse process, a method of using the area to be reduced for restoring or correcting partial image data and then removing it is selected. Or, a method of filling an expanded area by using a copy of outline pixels is selected, and in the reverse process, a method of simply removing the area to be reduced is selected. That is, based on the data processing method selected in the image size adjustment process, the data processing method in the reverse process can be determined.
[0258] Unlike the above example, the data processing methods in the image size adjustment process and its reverse process can also have an independent relationship. That is, the data processing method in the reverse process can be selected regardless of the data processing method selected in the image size adjustment process. For example, a method of filling an expanded area with partial image data is selected, and a method of simply removing the area to be reduced in the reverse process can be selected.
[0259] In the present invention, the data processing method in the image size adjustment process can be implicitly determined according to the encoding / decoding setting, and the data processing method in the reverse process can be implicitly determined according to the encoding / decoding setting. Or, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the reverse process can be explicitly generated. Or, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the reverse process can be implicitly determined based on the data processing method.
[0260] Next, an example of performing image size adjustment by an encoding / decoding device according to an embodiment of the present invention is shown. In the example described below, the case where the size adjustment process is expansion and the reverse size adjustment process is reduction is taken as an example. Also, the difference between the "image before size adjustment" and the "image after size adjustment" can mean the size of the image, and the size adjustment related information can be partially explicitly generated or partially implicitly determined according to the encoding / decoding setting. Also, the size adjustment related information can include information for the size adjustment process and the reverse size adjustment process.
[0261] As a first example, a size adjustment process for the input image can be performed before the start of encoding. After performing the size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is used in the size adjustment process), the image after size adjustment can be encoded. After the completion of encoding, it can be stored in the memory, and the image encoding data (in this example, meaning the image after size adjustment) can be recorded in a bitstream and transmitted.
[0262] A size adjustment process can be performed before the start of decoding. After performing the size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, etc.), the image decoding data after size adjustment can be parsed and decoded. After the completion of decoding, it can be stored in the memory, and the inverse size adjustment process (in this example, using data processing methods, etc. This is used in the inverse size adjustment process) can be performed to change the output image to the image before size adjustment.
[0263] As a second example, a size adjustment process for the reference image can be performed before the start of encoding. After performing the size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is used in the size adjustment process), the image after size adjustment (in this example, the reference image after size adjustment) can be stored in the memory, and this can be used to encode the image. After the completion of encoding, the image encoding data (in this example, meaning the data encoded using the reference image) can be recorded in a bitstream and transmitted. Also, when the encoded image is stored in the memory as a reference image, the size adjustment process can be performed as in the above process.
[0264] Before the start of decoding, the size adjustment process for the reference image can be performed. The size-adjusted image (in this example, the size-adjusted reference image) can be stored in the memory using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is the one used in the size adjustment process), and the image decoding data (in this example, the same as the one encoded using the reference image by the encoder) can be parsed and decoded. After the completion of decoding, it can be generated as an output image. When the decoded image is included in the reference image and stored in the memory, the size adjustment process can be performed as in the above process.
[0265] As a third example, (specifically, it means the completion of encoding excluding the filtering process) before the start of image filtering (assumed to be a deblocking filter in this example) after the completion of encoding, the size adjustment process for the image can be performed. After performing the size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is the one used in the size adjustment process), the size-adjusted image can be generated, and filtering can be applied to the size-adjusted image. After the completion of filtering, the inverse size adjustment process can be performed to change it back to the image before size adjustment.
[0266] (Specifically, it means the completion of decoding excluding the filtering process) Before the start of image filtering after the completion of decoding, the size adjustment process for the image can be performed. After performing the size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is the one used in the size adjustment process), the size-adjusted image can be generated, and filtering can be applied to the size-adjusted image. After the completion of filtering, the inverse size adjustment process can be performed to change it back to the image before size adjustment.
[0267] In the above examples, in some cases (the first exemplary case and the third exemplary case), the resizing process and the reverse resizing process may be performed, and in some other cases (the second exemplary case), only the resizing process may be performed.
[0268] Also, in some cases (the second exemplary case and the third exemplary case), the resizing processes in the encoder and the decoder may be the same, and in some other cases (the first exemplary case), the resizing processes in the encoder and the decoder may or may not be the same. At this time, the difference in the resizing process in the encoder / decoder can be at the resizing execution stage. For example, in some cases (in this example, the encoder), it can include the resizing execution stage that considers the resizing of the image and the data processing of the area to be resized, and in some cases (in this example, the decoder), it can include the resizing execution stage that considers the resizing of the image. At this time, the former data processing can correspond to the data processing of the reverse resizing process of the latter.
[0269] Also, the resizing process in some cases (the third exemplary case) is a process applied only at the corresponding stage, and it may not be necessary to store the resizing area in the memory. For example, for the purpose of use in the filtering process, it can be stored in temporary memory for filtering, and the corresponding area can be removed through the reverse resizing process. In this case, it can be said that there is no change in the size of the image due to resizing. The present invention is not limited to the above examples, and modifications to other examples are also possible.
[0270] Through the above size adjustment process, the size of the image can be changed, and thereby, the coordinates of some pixels of the image can be changed through the size adjustment process. This can affect the operation of the picture division unit. In the present invention, block-by-block division can be performed based on the image before size adjustment through the above process, or block-by-block division can be performed based on the image after size adjustment. Also, division of some units (for example, tiles, slices, etc.) can be performed based on the image before size adjustment, or division of some units can be performed based on the image after size adjustment. This can be determined according to the encoding / decoding settings. In the present invention, the case where the picture division unit operates based on the image after size adjustment (for example, the image division process after the size adjustment process) will be mainly described, but other variations are also possible. This will be described in the case of the above example in a plurality of image settings described later.
[0271] In the encoder, the information generated in the above process is recorded in the bit stream in at least one unit among units such as sequence, picture, slice, tile, etc., and in the decoder, the related information is parsed from the bit stream. Also, it may be included in the bit stream in the form of SEI or metadata.
[0272] Although it may be common to perform encoding / decoding with the input image as it is, it may also occur when the image is reconstructed and then encoding / decoding is performed. For example, the image can be reconstructed for the purpose of improving the encoding efficiency of the image, the image can be reconstructed for the purpose of considering the network and user environments, and the image can be reconstructed according to the type and characteristics of the image.
[0273] In the present invention, the image reconstruction process can perform the reconstruction process alone, or can perform the reverse process thereof. In the example described later, the reconstruction process will be mainly described, but the content regarding the reconstruction reverse process can be induced conversely from the reconstruction process.
[0274] FIG. 10 is an exemplary diagram for image reconstruction according to an embodiment of the present invention.
[0275] When taking the first input image as 10a, 10a to 10d are exemplary diagrams to which a rotation including 0 degrees (for example, a candidate group generated by sampling 360 degrees into k intervals can be formed. k can have values such as 2, 4, 8, etc., and in this example, 4 is assumed for explanation) is applied, and 10e to 10h show exemplary diagrams to which inversion (or symmetry) is applied based on 10a or based on 10b to 10d.
[0276] Depending on the reconstruction of the image, can the start position or scan order of the image be changed, or can it follow a predetermined start position and scan order regardless of whether it is reconstructed or not. This can be determined according to the encoding / decoding settings. In the embodiments described later, it is assumed for explanation that, regardless of whether the image is reconstructed or not, it follows a predetermined start position (for example, the upper left position of the image) and scan order (for example, raster scan).
[0277] In the image encoding method and decoding method according to an embodiment of the present invention, the following image reconstruction steps can be included. At this time, the image reconstruction process can include an image reconstruction instruction step, an image reconstruction type identification step, and an image reconstruction execution step. Also, the image encoding device and decoding device can be configured to include an image reconstruction instruction unit, an image reconstruction type identification unit, and an image reconstruction execution unit that realize the image reconstruction instruction step, the image reconstruction type identification step, and the image reconstruction execution step. In the case of encoding, associated syntax elements can be generated, and in the case of decoding, the associated syntax elements can be parsed.
[0278] In the image reconstruction instruction stage, it is possible to determine whether to perform image reconstruction. For example, when a signal indicating image reconstruction (e.g., convert_enabled_flag) is confirmed, reconstruction can be performed. When the signal indicating image reconstruction is not confirmed, reconstruction may not be performed, or other encoding / decoding information can be confirmed to perform reconstruction. Also, even if a signal indicating image reconstruction is not provided, a signal indicating reconstruction may be implicitly activated or deactivated according to the encoding / decoding settings (e.g., image characteristics, type, etc.). When performing reconstruction, reconstruction-related information can be generated thereby, or reconstruction-related information can be implicitly determined.
[0279] When a signal indicating image reconstruction is provided, the corresponding signal is a signal for indicating whether to perform image reconstruction, and it is possible to confirm whether to reconstruct the corresponding image according to the signal. For example, when a signal indicating image reconstruction (e.g., convert_enabled_flag) is confirmed and the corresponding signal is activated (e.g., convert_enabled_flag = 1), reconstruction may be performed. When the corresponding signal is deactivated (e.g., convert_enabled_flag = 0), reconstruction may not be performed.
[0280] Also, when a signal indicating image reconstruction is not provided, reconstruction may not be performed, or whether to reconstruct the corresponding image can be confirmed by other signals. For example, reconstruction can be performed according to image characteristics, type, etc. (e.g., 360-degree image), and reconstruction information can be explicitly generated or assigned to a predetermined value. It is not limited to only the above example, and variations to other examples are also possible.
[0281] In the image reconstruction type identification stage, the image reconstruction type can be identified. The image reconstruction type can be defined by the method of performing reconstruction, reconstruction mode information, etc. The method of performing reconstruction (for example, convert_type_flag) can include rotation, inversion, etc., and the reconstruction mode information can include the mode (for example, convert_mode) in the method of performing reconstruction. In this case, the reconstruction-related information can be composed of the method of performing reconstruction and the mode information. That is, it can be composed of at least one syntax element. At this time, the number of candidate groups of mode information for each method of performing reconstruction may be the same or different.
[0282] As an example, in the case of rotation, it can include candidates having a difference at regular intervals (in this example, 90 degrees) such as 10a to 10d. When 10a is set as 0-degree rotation, 10b to 10d can be examples where 90-degree, 180-degree, and 270-degree rotations are applied respectively (in this example, the angle is measured clockwise).
[0283] As an example, in the case of inversion, it can include candidates such as 10a, 10e, and 10f. When 10a is set without inversion, 10e and 10f can be examples where left-right inversion and up-down inversion are applied respectively.
[0284] The above examples illustrate the cases of settings for rotation with a certain interval and settings for some inversions, but they are only examples for image reconstruction and are not limited to the above cases. Other differences in intervals can include examples of other inversion operations, etc. This can be determined according to the encoding / decoding settings.
[0285] Or, it can include integrated information (for example, convert_com_flag) generated by mixing the method of performing reconstruction and the mode information thereby. In this case, the reconstruction-related information can be composed of information in which the method of performing reconstruction and the mode information are mixed.
[0286] For example, the integrated information can include candidates such as 10a to 10f. This can be an example where rotations of 0 degrees, 90 degrees, 180 degrees, 270 degrees, horizontal flipping, and vertical flipping are applied based on 10a.
[0287] Alternatively, the integrated information can include candidates such as 10a to 10h. This can be an example where rotations of 0 degrees, 90 degrees, 180 degrees, 270 degrees, horizontal flipping, vertical flipping, horizontal flipping after 90 - degree rotation (or 90 - degree rotation after horizontal flipping), vertical flipping after 90 - degree rotation (or 90 - degree rotation after vertical flipping) are applied, or it can be an example where rotations of 0 degrees, 90 degrees, 180 degrees, 270 degrees, horizontal flipping, horizontal flipping after 180 - degree rotation (180 - degree rotation after horizontal flipping), horizontal flipping after 90 - degree rotation (90 - degree rotation after horizontal flipping), horizontal flipping after 270 - degree rotation (270 - degree rotation after horizontal flipping) are applied.
[0288] The candidate group can be configured to include modes where rotation is applied, modes where inversion is applied, and modes where rotation and inversion are mixed. The mixed - configured mode simply includes mode information in the method of performing reconstruction and can include modes generated by mixing mode information of each method. At this time, it can include modes generated by mixing at least one mode of some methods (for example, rotation) and at least one mode of some methods (for example, inversion). The above example includes cases where one mode of some methods and multiple modes of some methods are mixed (in this example, 90 - degree rotation + multiple inversions / horizontal flipping + multiple rotations). The mixed - configured information can be configured to include, as a candidate group, the case where no reconstruction is applied {in this example, 10a}, and when no reconstruction is applied, it can be included as the first candidate group (for example, assigned an index of 0).
[0289] Alternatively, it can include mode information according to a method of performing a predetermined reconstruction. In this case, the reconstruction - related information can be composed of mode information according to a method of performing a predetermined reconstruction. That is, information about the method of performing reconstruction can be omitted and can be composed of one syntactic element related to mode information.
[0290] For example, it can be configured to include candidates such as 10a to 10d related to rotation. Or, it can be configured to include candidates such as 10a, 10e, and 10f related to inversion.
[0291] The sizes of the images before and after the image reconstruction process may be the same, or at least one length may be different. This can be determined according to the encoding / decoding settings. The image reconstruction process is a process of rearranging the pixels in the image (in this example, in the reverse process of image reconstruction, the reverse process of pixel rearrangement is performed. This can be induced inversely from the pixel rearrangement process), and the position of at least one pixel can be changed. The rearrangement of the pixels can be performed according to rules based on the image reconstruction type information.
[0292] At this time, the pixel rearrangement process can be affected by the size and shape of the image (for example, square or rectangle), etc. Specifically, the horizontal width and vertical height of the image before the reconstruction process and the horizontal width and vertical height of the image after the reconstruction process can act as variables in the pixel rearrangement process.
[0293] For example, at least one ratio information of the ratio of the horizontal width of the image before the reconstruction process to the horizontal width of the image after the reconstruction process, the ratio of the horizontal width of the image before the reconstruction process to the vertical height of the image after the reconstruction process, the ratio of the vertical height of the image before the reconstruction process to the horizontal width of the image after the reconstruction process, and the ratio of the vertical height of the image before the reconstruction process to the vertical height of the image after the reconstruction process (for example, former / latter or latter / former, etc.) can act as variables in the pixel rearrangement process.
[0294] In the above example, when the sizes of the images before and after the reconstruction process are the same, the ratio of the horizontal width to the vertical height of the image can act as a variable in the pixel rearrangement process. Also, when the shape of the image is square, the ratio of the length of the image before the image reconstruction process to the length of the image after the reconstruction process can act as a variable in the pixel rearrangement process.
[0295] In the image reconstruction execution stage, image reconstruction can be performed based on the identified reconstruction information. That is, image reconstruction can be performed based on information such as the reconstruction type and the reconstruction mode, and encoding / decoding can be performed based on the obtained reconstructed image.
[0296] Next, an example of performing image reconstruction by an encoding / decoding apparatus according to an embodiment of the present invention will be shown.
[0297] Before the start of encoding, a reconstruction process for the input image can be performed. After performing reconstruction using the reconstruction information (for example, the image reconstruction type, the reconstruction mode, etc.), the reconstructed image can be encoded. After the completion of encoding, it can be stored in the memory, and the image encoding data can be recorded in the bitstream and transmitted.
[0298] Before the start of decoding, a reconstruction process can be performed. After performing reconstruction using the reconstruction information (for example, the image reconstruction type, the reconstruction mode, etc.), the image decoding data can be parsed and decoded. After the completion of decoding, it can be stored in the memory, and after performing the reverse reconstruction process to change it to the image before reconstruction, the image can be output.
[0299] The encoder records the information generated in the above process in the bitstream in at least one unit among units such as sequences, pictures, slices, and tiles, and the decoder parses the relevant information from the bitstream. Also, it may be included in the bitstream in the form of SEI or metadata.
[0300] [Table 1]
[0301] Table 1 shows examples of syntax elements related to the partitioning during image setting. In the examples described later, the explanation will focus on the added syntax elements. Also, the syntax elements in the examples described later are not limited to a specific unit, and can be syntax elements supported in various units such as sequences, pictures, slices, tiles, etc. Or they can be syntax elements included in SEI, metadata, etc. Further, the types, order, conditions, etc. of the syntax elements supported in the examples described later are only limited in this example, and can be changed and determined according to the encoding / decoding settings.
[0302] In Table 1, tile_header_enabled_flag means a syntax element regarding whether to support encoding / decoding settings for a tile. When activated (tile_header_enabled_flag = 1), it can have encoding / decoding settings at the tile unit. When deactivated (tile_header_enabled_flag = 0), it cannot have encoding / decoding settings at the tile unit and can receive the assignment of encoding / decoding settings of a higher unit.
[0303] The tile_coded_flag means a syntax element indicating whether to perform tile encoding / decoding. When activated (tile_coded_flag = 1), the encoding / decoding of the corresponding tile can be performed. When deactivated (tile_coded_flag = 0), the encoding / decoding cannot be performed. Here, not performing encoding means not generating encoded data for the tile (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc., and is applicable to meaningless areas in a partial projection format of a 360-degree image). Not performing decoding means no longer parsing the decoded data for the corresponding tile (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc.). Also, no longer parsing the decoded data can mean that there is no encoded data in the corresponding unit, so no longer parsing, but it can also mean no longer parsing even if there is encoded data, due to the said flag. Depending on whether to perform tile encoding / decoding, header information at the tile unit can be supported.
[0304] Although the above example is described centered around tiles, it is not an example limited to tiles, and can be an example applicable to other division units in the present invention. Also, as an example of tile division settings, it is not limited to the above case, and deformation to other examples is also possible.
[0305] [Table 2]
[0306] Table 2 shows examples of syntax elements related to reconstruction during image setting.
[0307] Referring to Table 2, convert_enabled_flag means a syntax element regarding whether to perform reconstruction. When activated (convert_enabled_flag = 1), it means encoding / decoding the reconstructed image, and additional reconstruction-related information can be checked. When deactivated (convert_enabled_flag = 0), it means encoding / decoding the existing image.
[0308] convert_type_flag means mixed information regarding the method and mode information for performing reconstruction. It can be determined to be any one of a plurality of candidate groups regarding the method applying rotation, the method applying inversion, and the method applying a mixture of rotation and inversion.
[0309]
Table 3
[0310] Table 3 shows an example regarding the syntax elements related to resizing during image setting.
[0311] Referring to Table 3, pic_width_in_samples and pic_height_in_samples mean syntax elements regarding the horizontal and vertical widths of the image, and the size of the image can be checked by said syntax elements.
[0312] img_resizing_enabled_flag means a syntax element regarding whether to resize the image. When activated (img_resizing_enabled_flag = 1), it means encoding / decoding the resized image, and additional resizing-related information can be checked. When deactivated (img_resizing_enabled_flag = 0), it means encoding / decoding the existing image. Also, it may be a syntax element meaning resizing for intra prediction.
[0313] The resizing_met_flag means a syntax element for the resizing method. When performing resizing using a scale factor (resizing_met_flag = 0), when performing resizing using an offset factor (resizing_met_flag = 1), it can be determined to be any one of a candidate group such as other resizing methods.
[0314] The resizing_mov_flag means a syntax element for the resizing operation. For example, it can be determined to be either an expansion or a reduction.
[0315] width_scale and height_scale mean scale factors related to horizontal resizing and vertical resizing in resizing using a scale factor.
[0316] top_height_offset and bottom_height_offset mean the upward and downward offset factors related to horizontal resizing in resizing using an offset factor, and left_width_offset and right_width_offset mean the leftward and rightward offset factors related to vertical resizing in resizing using an offset factor.
[0317] Through the resizing-related information and the image size information, the size of the image after resizing can be updated.
[0318] The resizing_type_flag means a syntax element for the data processing method of the area to be resized. Depending on the resizing method and the resizing operation, the number of candidate groups for the data processing method may be the same or different.
[0319] The image setting process applied to the aforementioned image encoding / decoding device may be performed individually, or a plurality of image setting processes may be performed in combination. In the example described later, the case where a plurality of image setting processes are performed in combination will be described.
[0320] FIG. 11 is an exemplary diagram showing images before and after an image setting process according to an embodiment of the present invention. Specifically, 11a is an example before reconstructing the image in the divided image (for example, an image projected by 360-degree image encoding), and 11b shows an example after performing image reconstruction on the divided image (for example, an image packed by 360-degree image encoding). That is, 11a can be understood as an exemplary diagram before performing the image setting process, and 11b can be understood as an exemplary diagram after performing the image setting process.
[0321] The image setting process in this example will be described for the cases of image segmentation (assumed as tiles in this example) and image reconstruction.
[0322] In the example described later, the case where image reconstruction is performed after image segmentation will be explained. However, it is also possible that image segmentation is performed after image reconstruction according to the encoding / decoding settings, and variations to other cases are also possible. Also, the above-described image reconstruction process (including the reverse process) can be applied in the same or similar manner to the reconstruction process of the divided units within the image in this embodiment.
[0323] Image reconstruction may or may not be performed on all divided units within the image, and may be performed on some of the divided units. Therefore, the divided units before reconstruction (for example, some of P0 to P5) may or may not be the same as the divided units after reconstruction (for example, some of S0 to S5). The cases related to the execution of various image reconstructions will be explained through the examples described later. Also, for the sake of convenience of explanation, it is assumed that the unit of the image is a picture, the unit of the divided image is a tile, and the divided unit is in a square shape.
[0324] As an example, whether to perform image reconstruction can be determined by some units (e.g., sps_convert_enabled_flag or SEI, metadata, etc.). Or, whether to perform image reconstruction can be determined by some units (e.g., pps_convert_enabled_flag). This is possible when it first occurs in the corresponding unit (in this example, the picture) or when it is activated by a higher-level unit (e.g., sps_convert_enabled_flag = 1). Or, whether to perform image reconstruction can be determined by some units (e.g., tile_convert_flag[i]. i is the index of the division unit). This is possible when it first occurs in the corresponding unit (in this example, the tile) or when it is activated by a higher-level unit (e.g., pps_convert_enabled_flag = 1). Also, whether to perform the reconstruction of the partial image can be implicitly determined according to the encoding / decoding settings, whereby the related information can be omitted.
[0325] As an example, according to a signal for instructing image reconstruction (e.g., pps_convert_enabled_flag), it can be determined whether to perform the reconstruction of the division units within the image. Specifically, according to the signal, it can be determined whether to perform the reconstruction of all the division units within the image. At this time, a signal for instructing the reconstruction of one image can be generated.
[0326] As an example, according to a signal for instructing image reconstruction (e.g., tile_convert_flag[i]), it can be determined whether to perform the reconstruction of the division units within the image. Specifically, according to the signal, it can be determined whether to perform the reconstruction of a partial division unit within the image. At this time, at least one signal for instructing the reconstruction of an image (e.g., generated as many as the number of division units) can be generated.
[0327] As an example, it is possible to determine whether to perform image reconstruction according to a signal for instructing image reconstruction (e.g., pps_convert_enabled_flag), and it is possible to determine whether to perform reconstruction of a divided unit within the image according to a signal for instructing image reconstruction (e.g., tile_convert_flag[i]). Specifically, when some signals are activated (e.g., pps_convert_enabled_flag = 1), some other signals (e.g., tile_convert_flag[i]) can be further checked, and according to the said signal (in this example, tile_convert_flag[i]), it is possible to determine whether to perform reconstruction of a divided unit within the image. At this time, signals for instructing reconstruction of a plurality of images can be generated.
[0328] When a signal for instructing image reconstruction is activated, image reconstruction-related information can be generated. In the examples described later, cases regarding various image reconstruction-related information will be explained.
[0329] As an example, reconstruction information applied to the image can be generated. Specifically, one piece of reconstruction information can be used as the reconstruction information for all divided units within the image.
[0330] As an example, reconstruction information applied to a divided unit within the image can be generated. Specifically, at least one piece of reconstruction information can be used as the reconstruction information for a divided unit within the image. That is, one piece of reconstruction information can be used as the reconstruction information for one divided unit, or one piece of reconstruction information can be used as the reconstruction information for a plurality of divided units.
[0331] The examples described later can be explained by combinations with examples of performing image reconstruction.
[0332] For example, when a signal for instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information that is commonly applied to the divided units within the image can be generated. Or, when a signal for instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information that is individually applied to the divided units within the image can be generated. Or, when a signal for instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information that is individually applied to the divided units within the image can be generated. Or, when a signal for instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information that is commonly applied to the divided units within the image can be generated.
[0333] In the case of the aforementioned reconstruction information, implicit or explicit processing can be performed according to the encoding / decoding settings. In the case of implicit processing, the reconstruction information can be reassigned to a predetermined value according to the characteristics, types, etc. of the image.
[0334] P0 to P5 in 11a can correspond to S0 to S5 in 11b, and a reconstruction process can be performed on the divided unit. For example, it can be assigned to S0 without performing reconstruction on P0, it can be assigned to S1 by applying a 90-degree rotation to P1, it can be assigned to S2 by applying a 180-degree rotation to P2, it can be assigned to S3 by applying a horizontal flip to P3, it can be assigned to S4 by applying a horizontal flip after a 90-degree rotation to P4, and it can be assigned to S5 by applying a horizontal flip after a 180-degree rotation to P5.
[0335] However, it is not limited to the above-described examples, and various modified examples are also possible. It is possible to perform no reconstruction on the divided unit of the image as in the above example, or at least one of the reconstruction methods such as reconstruction with rotation applied, reconstruction with flip applied, or reconstruction with a mixture of rotation and flip applied.
[0336] When the image reconstruction is applied to the division unit, additional reconstruction processes such as division unit rearrangement can be performed. That is, the image reconstruction process of the present invention can be configured to include rearrangement within the division unit of the image in addition to rearranging the pixels in the image, and can be expressed by some syntax elements such as those in Table 4 (for example, part_top, part_left, part_width, part_height, etc.). This means that the image division and the image reconstruction process can be understood as being mixed. The above description can be an example possible when the image is divided into a plurality of units.
[0337] P0 to P5 of 11a can correspond to S0 to S5 of 11b, and the reconstruction process can be performed on the division unit. For example, it can be assigned to S0 without performing reconstruction on P0, assigned to S2 without performing reconstruction on P1, assigned to S1 by applying a 90-degree rotation to P2, assigned to S4 by applying a horizontal flip to P3, assigned to S5 by applying a horizontal flip after a 90-degree rotation to P4, and assigned to S3 by applying a 180-degree rotation after a horizontal flip to P5, and is not limited thereto, and examples of various deformations are also possible.
[0338] In addition, P_Width and P_Height in FIG. 7 can correspond to P_Width and P_Height in FIG. 11, and P’_Width and P’_Height in FIG. 7 can correspond to P’_Width and P’_Height in FIG. 11. The image size after size adjustment in FIG. 7 is P’_Width×P’_Height, which can be expressed as (P_Width+Exp_L+Exp_R)×(P_Height+Exp_T+Exp_B), and the image size after size adjustment in FIG. 11 is P’_Width×P’_Height, which can be expressed as (P_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R)×(P_Height+Var0_T+Var1_T+Var0_B+Var1_B), or (Sub_P0_Width+Sub_P1_Width+Sub_P2_Width+Var0_L+Var1_L+Var2_L+Var0_R+Var1_R+Var2_R)×(Sub_P0_Height+Sub_P1_Height+Var0_T+Var1_T+Var0_B+Var1_B).
[0339] As in the above example, the reconstruction of the image can perform pixel rearrangement within the division unit of the image, can perform rearrangement of the division units within the image, and can perform not only pixel rearrangement within the division unit of the image but also rearrangement of the division units within the image. At this time, after performing pixel rearrangement of the division unit, rearrangement within the image of the division unit can be performed, or after performing rearrangement within the image of the division unit, pixel rearrangement of the division unit can be performed.
[0340] The rearrangement of the division units within the image can be determined whether to be performed according to a signal instructing the reconstruction of the image. Alternatively, a signal for the rearrangement of the division units within the image can be generated. Specifically, the signal can be generated when the signal instructing the reconstruction of the image is activated. Alternatively, implicit or explicit processing can be performed according to the encoding / decoding settings. In the case of implicit, it can be determined according to the characteristics, types, etc. of the image.
[0341] In addition, information about the rearrangement of the divided units within the image can be processed implicitly or explicitly according to the encoding / decoding settings, and can be determined according to the characteristics, types, etc. of the image. That is, each divided unit can be arranged according to the arrangement information of a predetermined divided unit.
[0342] Next, an example of reconstructing the divided units within the image by the encoding / decoding apparatus according to an embodiment of the present invention is shown.
[0343] Before the start of encoding, a division process can be performed on the input image using division information. A reconstruction process can be performed using the reconstruction information for the divided units, and the image reconstructed for the divided units can be encoded. After the completion of encoding, it can be stored in a memory, and the image encoding data can be recorded in a bit stream and transmitted.
[0344] Before the start of decoding, a division process can be performed using division information. A reconstruction process is performed using the reconstruction information for the divided units, and the image decoding data can be parsed and decoded with the divided units for which the reconstruction has been performed. After the completion of decoding, it can be stored in a memory, and after performing the reverse process of reconstructing the divided units, the divided units can be merged into one to output an image.
[0345] FIG. 12 is an exemplary diagram of size adjustment for each of the divided units within the image according to an embodiment of the present invention. P0 to P5 in FIG. 12 correspond to P0 to P5 in FIG. 11, and S0 to S5 in FIG. 12 correspond to S0 to S5 in FIG. 11.
[0346] In the example described below, the case where image size adjustment is performed after image division will be mainly described, but it is also possible to perform image division after image size adjustment according to the encoding / decoding settings, and variations to other cases are also possible. Also, the above-described image size adjustment process (including the reverse process) can be applied in the same or similar manner to the size adjustment process of the divided units within the image in the present embodiment.
[0347] For example, TL to BR in FIG. 7 can correspond to TL to BR of the divided units SX (S0 to S5) in FIG. 12, S0 and S1 in FIG. 7 can correspond to PX and SX in FIG. 12, P_Width and P_Height in FIG. 7 can correspond to Sub_PX_Width and Sub_PX_Height in FIG. 12, P’_Width and P’_Height in FIG. 7 can correspond to Sub_SX_Width and Sub_SX_Height in FIG. 12, Exp_L, Exp_R, Exp_T, Exp_B in FIG. 7 can correspond to VarX_L, VarX_R, VarX_T, VarX_B in FIG. 12, and other factors can also correspond.
[0348] In the process of adjusting the size of the divided units in the images of 12a to 12f, it can be distinguished from the image size enlargement or reduction in FIGS. 7a and 7b in that there may be settings regarding image size enlargement or reduction in proportion to the number of divided units. Also, there can be differences in having settings commonly applied to the divided units in the image or having settings individually applied to the divided units in the image. Cases of size adjustment in various situations will be described in the examples to be described later, and the size adjustment process can be performed considering the above matters.
[0349] The image size adjustment in the present invention may or may not be performed on all the divided units in the image, and may be performed on a part of the divided units. Cases regarding various image size adjustments will be described through the examples to be described later. Also, for the sake of convenience of explanation, it is assumed that the size adjustment operation is expansion, the method of performing the size adjustment is the offset factor, the size adjustment direction is the up, down, left, and right directions, the size adjustment direction operates according to the size adjustment information, the unit of the image is a picture, and the unit of the divided image is a tile for explanation.
[0350] As an example, whether to perform image size adjustment can be determined by some units (for example, sps_img_resizing_enabled_flag, or SEI, metadata, etc.). Or, whether to perform image size adjustment can be determined by some units (for example, pps_img_resizing_enabled_flag). This is possible when it first occurs in the corresponding unit (in this example, the picture), or when it is activated by a higher-level unit (for example, sps_img_resizing_enabled_flag = 1). Or, whether to perform image size adjustment can be determined by some units (for example, tile_resizing_flag[i]. i is the split unit index). This is possible when it first occurs in the corresponding unit (in this example, the tile), or when it is activated by a higher-level unit. Also, whether to perform size adjustment on some of the images can be implicitly determined according to the encoding / decoding settings, whereby the related information can be omitted.
[0351] As an example, according to a signal for instructing image size adjustment (for example, pps_img_resizing_enabled_flag), it can be determined whether to perform size adjustment on the split units within the image. Specifically, according to the signal, it can be determined whether to perform size adjustment on all the split units within the image. At this time, a signal for instructing size adjustment of one image can be generated.
[0352] As an example, according to a signal for instructing image size adjustment (for example, tile_resizing_flag[i]), it can be determined whether to perform size adjustment on the split units within the image. Specifically, according to the signal, it can be determined whether to perform size adjustment on a part of the split units within the image. At this time, at least one signal for instructing size adjustment of an image (for example, generated as many as the number of split units) can be generated.
[0353] As an example, according to a signal for instructing image size adjustment (e.g., pps_img_resizing_enabled_flag), it can be determined whether to perform image size adjustment, and according to a signal for instructing image size adjustment (e.g., tile_resizing_flag[i]), it can be determined whether to perform size adjustment of the divided units within the image. Specifically, when some signals are activated (e.g., pps_img_resizing_enabled_flag = 1), some other signals (e.g., tile_resizing_flag[i]) can be further checked, and according to the said signal (in this example, tile_resizing_flag[i]), it can be determined whether to perform size adjustment of some of the divided units within the image. At this time, signals for instructing size adjustment of multiple images can be generated.
[0354] When the signal for instructing image size adjustment is activated, image size adjustment related information can be generated. Cases regarding various image size adjustment related information will be described in the examples below.
[0355] As an example, size adjustment information applicable to the image can be generated. Specifically, one piece of size adjustment information or a set of size adjustment information can be used as the size adjustment information for all the divided units within the image. For example, one piece of size adjustment information (or a size adjustment value applicable to all the size adjustment directions supported or allowed by the divided unit, etc. One piece of information in this example) commonly applicable to the up, down, left, and right directions of the divided units within the image, or a set of one piece of size adjustment information respectively applicable to the up, down, left, and right directions (or the number of size adjustment directions supported or allowed by the divided unit. At most 4 pieces of information in this example) can be generated.
[0356] As an example, size adjustment information applicable to a division unit within an image can be generated. Specifically, at least one piece of size adjustment information or a set of size adjustment information can be used as the size adjustment information for a division unit within the image. That is, one piece of size adjustment information or a set of size adjustment information can be used as the size adjustment information for one division unit, or can be used as the size adjustment information for multiple division units. For example, one piece of size adjustment information that is commonly applied in the up, down, left, and right directions of one division unit within the image, or a set of size adjustment information that is respectively applied in the up, down, left, and right directions can be generated. Or, one piece of size adjustment information that is commonly applied in the up, down, left, and right directions of multiple division units within the image, or a set of size adjustment information that is respectively applied in the up, down, left, and right directions can be generated. The composition of the size adjustment set means size adjustment value information for at least one size adjustment direction.
[0357] In summary, size adjustment information that is commonly applied to the division units within the image can be generated. Or, size adjustment information that is individually applied to the division units within the image can be generated. The examples described later can be explained by combinations with examples of performing image size adjustment.
[0358] For example, when a signal for instructing image size adjustment (e.g., pps_img_resizing_enabled_flag) is activated, size adjustment information that is commonly applied to the division units within the image can be generated. Or, when a signal for instructing image size adjustment (e.g., pps_img_resizing_enabled_flag) is activated, size adjustment information that is individually applied to the division units within the image can be generated. Or, when a signal for instructing image size adjustment (e.g., tile_resizing_flag[i]) is activated, size adjustment information that is individually applied to the division units within the image can be generated. Or, when a signal for instructing image size adjustment (e.g., tile_resizing_flag[i]) is activated, size adjustment information that is commonly applied to the division units within the image can be generated.
[0359] The size adjustment direction and size adjustment information of the image can be processed implicitly or explicitly according to the encoding / decoding settings. In the case of implicit processing, the size adjustment information can be assigned to a predetermined value according to the characteristics and types of the image.
[0360] The size adjustment direction in the size adjustment process of the present invention described above is at least one of the up, down, left, and right directions, and it has been explained that the size adjustment direction and size adjustment information can be processed implicitly or explicitly. That is, for some directions, the size adjustment value (including 0. That is, no adjustment) is determined in advance implicitly, and for some directions, the size adjustment value (including 0. That is, no adjustment) is assigned explicitly.
[0361] Even for the divided units within the image, the size adjustment direction and size adjustment information can be set to be capable of implicit or explicit processing, and this can be applied to the divided units within the image. For example, a setting (in this example, only the divided unit occurs) applied to one divided unit within the image can occur, or a setting applied to a plurality of divided units within the image can occur, or a setting applied to all divided units within the image (in this example, one setting occurs) can occur, and at least one setting can occur in the image (for example, the number of divided units of settings can occur from one setting). The setting information applied to the divided units within the image can be collected to define one set of settings.
[0362] FIG. 13 is an exemplary diagram for the size adjustment or set of settings of the divided units within the image.
[0363] Specifically, various examples of implicit or explicit processing of the size adjustment direction and size adjustment information of the divided units within the image are shown. In the examples described later, for the sake of convenience of explanation, the implicit processing is explained assuming that the size adjustment value of some size adjustment directions is 0.
[0364] When the boundary of the division unit coincides with the boundary of the image as in 13a (in this example, the thick solid line), explicit processing for size adjustment can be performed, and when they do not coincide (the thin solid line), implicit processing can be performed. For example, P0 can be adjusted in the upward and leftward directions (a2, a0), P1 in the upward direction (a2), P2 in the upward and rightward directions (a2, a1), P3 in the downward and leftward directions (a3, a0), P4 in the downward direction (a3), P5 in the downward and rightward directions (a3, a1), and size adjustment is not possible in other directions.
[0365] As in 13b, for some directions of the division unit (in this example, up and down), explicit processing for size adjustment can be performed, and for some directions of the division unit (in this example, left and right), when the boundary of the division unit coincides with the boundary of the image, explicit processing (in this example, the thick solid line) can be performed, and when they do not coincide (in this example, the thin solid line), implicit processing can be performed. For example, P0 can be adjusted in the up, down, and left directions (b2, b3, b0), P1 in the up and down directions (b2, b3), P2 in the up, down, and right directions (b2, b3, b1), P3 in the up, down, and left directions (b3, b4, b0), P4 in the up and down directions (b3, b4), P5 in the up, down, and right directions (b3, b4, b1), and size adjustment is not possible in other directions.
[0366] As in 13c, for some directions of the division unit (in this example, left and right), explicit processing for size adjustment can be performed, and for some directions of the division unit (in this example, up and down), when the boundary of the division unit coincides with the boundary of the image (in this example, the thick solid line), explicit processing can be performed, and when they do not coincide (in this example, the thin solid line), implicit processing can be performed. For example, P0 can be adjusted in the up, left, and right directions (c4, c0, c1), P1 in the up, left, and right directions (c4, c1, c2), P2 in the up, left, and right directions (c4, c2, c3), P3 in the down, left, and right directions (c5, c0, c1), P4 in the down, left, and right directions (c5, c1, c2), P5 in the down, left, and right directions (c5, c2, c3), and size adjustment is not possible in other directions.
[0367] As in the above example, the settings related to image size adjustment can have various cases. Whether multiple setting sets are supported and explicit setting set selection information can be generated, or a predetermined setting set can be implicitly determined according to the encoding / decoding settings (for example, image characteristics, types, etc.).
[0368] FIG. 14 is an exemplary diagram showing the image size adjustment process and the size adjustment process of the divided units within the image together.
[0369] Referring to FIG. 14, the image size adjustment process and the reverse process can proceed in the directions of e and f, and the size adjustment process and the reverse process of the divided units within the image can proceed in the directions of d and g. That is, the image size adjustment process can be performed on the image, the size adjustment of the divided units within the image can be performed, and the order of the size adjustment process is not fixed. This means that multiple size adjustment processes are possible.
[0370] In summary, the image size adjustment process can be classified into image size adjustment (or image size adjustment before division) and size adjustment of the divided units within the image (or image size adjustment after division). It is not necessary to perform both the image size adjustment and the size adjustment of the divided units within the image. Either one of them can be performed, or both can be performed. This can be determined according to the encoding / decoding settings (for example, image characteristics, types, etc.).
[0371] When performing multiple size adjustment processes in the above example, the image size adjustment can be performed in at least one of the up, down, left, and right directions of the image, and the size adjustment of at least one of the divided units within the image can be performed. At this time, the size adjustment can be performed in at least one of the up, down, left, and right directions of the divided unit where the size adjustment is performed.
[0372] Referring to FIG. 14, the size of the image A before size adjustment can be defined as P_Width×P_Height, the size of the image after the first size adjustment (or the image before the second size adjustment, B) is P’_Width×P’_Height, and the size of the image after the second size adjustment (or the image after the final size adjustment, C) is P’’_Width×P’’_Height. The image A before size adjustment means an image without any size adjustment, the image B after the first size adjustment means an image with some size adjustments, and the image C after the second size adjustment means an image with all size adjustments. For example, the image B after the first size adjustment means an image in which the size of the division unit within the image has been adjusted as shown in FIGS. 13a to 13c, and the image C after the second size adjustment can mean an image in which the entire image B that has been adjusted in size for the first time has been adjusted in size as shown in FIG. 7a, and the reverse case is also possible. It is not limited to the above examples, and various modified examples are possible.
[0373] In the size of the image B after the first size adjustment, P’_Width can be obtained via at least one size adjustment value in the left or right direction that is horizontally size-adjustable with respect to P_Width, and P’_Height can be obtained via at least one size adjustment value in the up or down direction that is vertically size-adjustable with respect to P_Height. At this time, the size adjustment value can be a size adjustment value that occurs in the division unit.
[0374] In the size of the image C after the second size adjustment, P’’_Width can be obtained via at least one size adjustment value in the left or right direction that is horizontally size-adjustable with respect to P’_Width, and P’’_Height can be obtained via at least one size adjustment value in the up or down direction that is vertically size-adjustable with respect to P’_Height. At this time, the size adjustment value can be a size adjustment value that occurs from the image.
[0375] In summary, the size of the image after size adjustment can be obtained via at least one size adjustment value and the size of the image before size adjustment.
[0376] Information regarding a data processing method can be generated in the area where the size of the image is adjusted. Through the examples described below, cases regarding various data processing methods will be explained. For the case of the data processing method generated in the reverse process of size adjustment, the same or similar application as in the case of the size adjustment process is possible, and the data processing methods in the size adjustment process and the reverse process of size adjustment can be explained through various combinations described below.
[0377] As an example, a data processing method applied to an image can be generated. Specifically, one data processing method or a set of data processing methods can be used as the data processing method for all the divided units in the image (assuming that all the divided units are adjusted in size in this example). For example, one data processing method commonly applied in the up, down, left, and right directions of the divided units in the image (or a data processing method applied to all the size adjustment directions supported or allowed by the divided units, etc., which is one piece of information in this example) or a set of one data processing method applied to the up, down, left, and right directions respectively (or the number of size adjustment directions supported or allowed by the divided units, which is at most four pieces of information in this example) can be generated.
[0378] As an example, a data processing method applied to the divided units in the image can be generated. Specifically, at least one data processing method or a set of data processing methods can be used as the data processing method for some of the divided units in the image (assuming the divided units to be adjusted in size in this example). That is, one data processing method or a set of data processing methods can be used as the data processing method for one divided unit or for multiple divided units. For example, one data processing method commonly applied in the up, down, left, and right directions of one divided unit in the image or a set of one data processing method applied to the up, down, left, and right directions respectively can be generated. Or, one data processing method commonly applied in the up, down, left, and right directions of multiple divided units in the image or a set of one data processing method information applied to the up, down, left, and right directions respectively can be generated. The composition of the set of data processing methods means the data processing method for at least one size adjustment direction.
[0379] In summary, a data processing method that is commonly applied to the divided units within an image can be used. Alternatively, a data processing method that is individually applied to the divided units within an image can be used. The data processing method can use a predetermined method. The predetermined data processing method can include at least one method. This is applicable in an implicit case, and selection information regarding the data processing method can be explicitly generated. This can be determined according to the encoding / decoding settings (such as the characteristics and types of the image).
[0380] That is, a data processing method that is commonly applied to the divided units within an image can be used, using a predetermined method or selecting any one of a plurality of data processing methods. Alternatively, a data processing method that is individually applied to the divided units within an image can be used, using a predetermined method according to the divided unit or selecting any one of a plurality of data processing methods.
[0381] The examples described below explain some cases regarding the size adjustment (assuming expansion in this example) of the divided units within an image (in this example, filling the size adjustment area using a part of the image data).
[0382] For some regions TL~BR of some units (such as S0 to S5 in FIGS. 12a to 12f), size adjustment can be performed using the data of some regions tl~br of some units (P0 to P5 in FIGS. 12a to 12f). At this time, the said some units can be the same (such as S0 and P0) or different regions (such as S0 and P1). That is, the region TL to BR to be size-adjusted can be filled using some data tl to br of the said divided unit, and the region to be size-adjusted can be filled using some data of a divided unit different from the said divided unit.
[0383] As an example, the area TL~BR to be size-adjusted in the current divided unit can be size-adjusted using the tl~br data of the current divided unit. For example, the TL of S0 can be filled using the tl data of P0, the RC of S1 can be filled using the tr+rc+br data of P1, the BL+BC of S2 can be filled using the bl+bc+br data of P2, and the TL+LC+BL of S3 can be filled using the tl+lc+bl data of P3.
[0384] As an example, the area TL~BR to be size-adjusted in the current divided unit can be size-adjusted using the tl~br data of the divided unit that is spatially adjacent to the current divided unit. For example, the TL+TC+TR of S4 can be filled using the b1+bc+br data of P1 in the upward direction, the BL+BC of S2 can be filled using the tl+tc+tr data of P5 in the downward direction, the LC+BL of S2 can be filled using the tl+rc+bl data of P1 in the leftward direction, the RC of S3 can be filled using the tl+lc+bl data of P4 in the rightward direction, and the BR of S0 can be filled using the tl data of P4 in the lower left direction.
[0385] As an example, the area TL~BR to be size-adjusted in the current divided unit can be size-adjusted using the tl~br data of the divided unit that is not spatially adjacent to the current divided unit. For example, the data of the boundary areas (such as left and right, top and bottom, etc.) at both ends of the image can be obtained. The LC of S3 can be obtained using the tr+rc+br data of S5, the RC of S2 can be obtained using the tl+lc data of S0, the BC of S4 can be obtained using the tc+tr data of S1, and the TC of S1 can be obtained using the bc data of S4.
[0386] Alternatively, the data of a partial area of the image (an area that is not spatially adjacent but is determined to have a high correlation with the area to be size-adjusted) can be obtained. The BC of S1 can be obtained using the tl+lc+bl data of S3, the RC of S3 can be obtained using the tl+tc data of S1, and the RC of S5 can be obtained using the bc data of S0.
[0387] Also, in some cases regarding the size adjustment (assuming reduction in this example) of the divided units within the image (in this example, restoration or correction and removal using partial data of the image), it is as follows.
[0388] A partial area TL to BR of a unit (for example, S0 to S5 in FIGS. 12a to 12f) can be used in the restoration or correction process of the partial area tl to br of the partial units P0 to P5. At this time, the partial units can be the same (for example, S0 and P0) or different areas (for example, S0 and P2). That is, the area to be size-adjusted can be used and removed for the restoration of part of the data of the corresponding divided unit, and the area to be size-adjusted can be used and removed for the restoration of part of the data of a divided unit different from the corresponding divided unit. Since detailed examples can be induced inversely from the expansion process, they are omitted.
[0389] The above example is an example applicable when there is data highly correlated with the area to be size-adjusted, and the information on the position referred to for size adjustment can be explicitly generated, or obtained implicitly based on a predetermined rule, or a combination of these can be used to confirm relevant information. This can be an example applicable when obtaining data from other areas where continuity exists in the encoding of 360-degree images.
[0390] Next, an example of performing size adjustment of a divided unit in an image by an encoding / decoding device according to an embodiment of the present invention is shown.
[0391] Before the start of encoding, a division process for the input image can be performed. A size adjustment process can be performed using the size adjustment information for the divided unit, and the image after size adjustment of the divided unit can be encoded. After the completion of encoding, it can be saved in the memory, and the image encoding data can be recorded in a bit stream and transmitted.
[0392] Before the start of decoding, a division process can be performed using the division information. A size adjustment process can be performed using the size adjustment information for the divided unit, and the image decoding data can be parsed and decoded with the divided unit after size adjustment. After the completion of decoding, it can be saved in the memory, and after performing the reverse process of size adjustment of the divided unit, the divided units can be merged into one to output an image.
[0393] In other cases during the above-described image size adjustment process, the changes can be applied as in the above example, and are not limited thereto. Changes to other examples are also possible.
[0394] In the above image setting process, a combination of image size adjustment and image reconstruction is possible. Image reconstruction can be executed after image size adjustment, or image size adjustment can be executed after image reconstruction. Also, a combination of image segmentation, image reconstruction, and image size adjustment is possible. After image segmentation, image size adjustment and image reconstruction can be executed. The order of image setting is not fixed and can be changed, which can be determined according to the encoding / decoding settings. In this example, the image setting process will describe the case where image reconstruction is performed after image segmentation and then image size adjustment is performed. However, other orders are possible according to the encoding / decoding settings, and changes to other cases are also possible.
[0395] For example, it may be performed in the order of segmentation → reconstruction, reconstruction → segmentation, segmentation → size adjustment, size adjustment → segmentation, size adjustment → reconstruction, reconstruction → size adjustment, segmentation → reconstruction → size adjustment, segmentation → size adjustment → reconstruction, size adjustment → segmentation → reconstruction, size adjustment → reconstruction → segmentation, reconstruction → segmentation → size adjustment, reconstruction → size adjustment → segmentation, etc. Combinations with additional image settings are also possible. As described above, the image setting process may be performed sequentially, but all or part of the setting processes can also be performed simultaneously. Also, for some image setting processes, multiple processes can be performed according to the encoding / decoding settings (e.g., image characteristics, types, etc.). Next, examples of various combinations of the image setting process will be shown.
[0396] As an example, P0 to P5 in FIG. 11a can correspond to S0 to S5 in FIG. 11b, and a reconstruction process (in this example, rearrangement of pixels), a size adjustment process (in this example, the same size adjustment for the divided units) can be performed on the divided units. For example, size adjustment using an offset can be applied to P0 to P5 and assigned to S0 to S5. Also, it can be assigned to S0 without reconstructing P0, it can be assigned to S1 by applying a 90-degree rotation to P1, it can be assigned to S2 by applying a 180-degree rotation to P2, it can be assigned to S3 by applying a 270-degree rotation to P3, it can be assigned to S4 by applying a horizontal flip to P4, and it can be assigned to S5 by applying a vertical flip to P5.
[0397] As an example, P0 to P5 in FIG. 11a can correspond to positions that are the same as or different from those of S0 to S5 in FIG. 11b, and a reconstruction process (in this example, rearrangement of pixels and divided units), a size adjustment process (in this example, the same size adjustment for the divided units) can be performed on the divided units. For example, size adjustment using a scale can be applied to P0 to P5 and assigned to S0 to S5. Also, it can be assigned to S0 without reconstructing P0, it can be assigned to S2 without reconstructing P1, it can be assigned to S1 by applying a 90-degree rotation to P2, it can be assigned to S4 by applying a horizontal flip to P3, it can be assigned to S5 by applying a horizontal flip after a 90-degree rotation to P4, and it can be assigned to S3 by applying a 180-degree rotation after a horizontal flip to P5.
[0398] As an example, P0 to P5 in FIG. 11a can correspond to E0 to E5 in FIG. 5e, and a reconstruction process (in this example, rearrangement of pixels and divided units), and a size adjustment process (in this example, size adjustment not identical for divided units) can be performed on the divided units. For example, it can be assigned to E0 without performing size adjustment and reconstruction on P0, it can be assigned to E1 by performing size adjustment using a scale on P1 without performing reconstruction, it can be assigned to E2 without performing size adjustment and by performing reconstruction on P2, it can be assigned to E4 by performing size adjustment using an offset on P3 without performing reconstruction, it can be assigned to E5 without performing size adjustment and by performing reconstruction on P4, and it can be assigned to E3 by performing size adjustment using an offset and by performing reconstruction on P5.
[0399] As in the above example, the absolute or relative position of the divided units within the image before and after the image setting process may be maintained or may be changed. This can be determined according to the encoding / decoding settings (for example, characteristics, types, etc. of the image). Also, various combinations of image setting processes are possible, not limited to the above example, and variations to various examples are also possible.
[0400] The encoder records the information generated in the above process in a bitstream in at least one unit among units such as sequences, pictures, slices, tiles, etc., and the decoder parses the relevant information from the bitstream. Also, it may be included in the bitstream in the form of SEI or metadata.
[0401]
Table 4
[0402] Next, examples of syntax elements associated with multiple image settings are shown. In the examples described below, the description will focus on the added syntax elements. Also, the syntax elements in the examples described below are not limited to a specific unit, and can be syntax elements supported by various units such as sequences, pictures, slices, tiles, etc. Or, they can be syntax elements included in SEI, metadata, etc.
[0403] Referring to Table 4, parts_enabled_flag means a syntax element for whether or not a partial unit is divided. When activated (parts_enabled_flag = 1), it means that encoding / decoding is performed by dividing into a plurality of units, and additional division information can be confirmed. When deactivated (parts_enabled_flag = 0), it means encoding / decoding an existing image. This example is mainly described with a rectangular division unit such as a tile, and different settings for the existing tile and division information can be provided.
[0404] num_partitions means a syntax element for the number of division units, and the value obtained by adding 1 means the number of division units.
[0405] part_top[i] and part_left[i] mean syntax elements for the position information of the division unit, and mean the horizontal and vertical start positions of the division unit (for example, the position of the upper left end of the division unit). part_width[i] and part_height[i] mean syntax elements for the size information of the division unit, and mean the horizontal width and vertical height of the division unit. At this time, the start position and size information can be set in pixel units or block units. Further, the syntax elements may be syntax elements that can occur in the image reconstruction process, or may be syntax elements that can occur when the image division process and the image reconstruction process are mixed.
[0406] part_header_enabled_flag means a syntax element for whether or not to support the encoding / decoding setting for the division unit. When activated (part_header_enabled_flag = 1), it can have an encoding / decoding setting for the division unit. When deactivated (part_header_enabled_flag = 0), it cannot have an encoding / decoding setting and can receive the assignment of the encoding / decoding setting of the upper unit.
[0407] The above example is an example of a syntax element associated with size adjustment and reconstruction in units of division among the image settings described later, and is not limited thereto. Other division units and settings of the present invention can be applied with modifications. This example is described under the assumption that size adjustment and reconstruction are performed after division, but is not limited thereto, and can be applied with modifications depending on other image setting orders and the like. Also, the types of syntax elements, the order of syntax elements, conditions, etc. supported in the examples described later are only limited in this example, and can be changed and determined according to the encoding / decoding settings.
[0408]
Table 5
[0409] Table 5 shows an example of syntax elements related to the reconstruction of division units in image settings.
[0410] Referring to Table 5, part_convert_flag[i] means a syntax element for whether or not to reconstruct the division unit. The syntax element can occur for each division unit, and when activated (part_convert_flag[i]=1), it means encoding / decoding the reconstructed division unit, and additional reconstruction-related information can be confirmed. When deactivated (part_convert_flag[i]=0), it means encoding / decoding the existing division unit. convert_type_flag[i] means mode information related to the reconstruction of the division unit and can be information related to the rearrangement of pixels.
[0411] Also, syntax elements for additional reconstruction such as rearrangement of division units can occur. In this example, rearrangement of division units can also be performed via part_top and part_left, which are the syntax elements related to the above-described image division, or syntax elements related to the rearrangement of division units (for example, index information, etc.) can occur.
[0412]
Table 6
[0413] Table 6 shows an example of syntax elements related to the size adjustment of the division unit in image setting.
[0414] Referring to Table 6, part_resizing_flag[i] means a syntax element for whether to perform image size adjustment of the division unit. The syntax element can occur for each division unit, and when activated (part_resizing_flag[i]=1), it means encoding / decoding the division unit after size adjustment, and additional size-related information can be checked. When deactivated (part_resiznig_flag[i]=0), it means encoding / decoding the existing division unit.
[0415] width_scale[i] and height_scale[i] mean scale factors for horizontal size adjustment and vertical size adjustment in size adjustment using scale factors in the division unit.
[0416] top_height_offset[i] and bottom_height_offset[i] mean upward and downward offset factors related to size adjustment using offset factors in the division unit, and left_width_offset[i] and right_width_offset[i] mean leftward and rightward offset factors related to size adjustment using offset factors in the division unit.
[0417] resizing_type_flag[i][j] means a syntax element for the data processing method of the area to be size-adjusted in the division unit. The syntax element means an individual data processing method in the direction of size adjustment. For example, syntax elements for individual data processing methods of areas size-adjusted in the up, down, left, and right directions can occur. This can also be generated based on size adjustment information (for example, it can only occur when size-adjusted in some directions).
[0418] The above-described image setting process can be a process applied according to the characteristics, types, etc. of the image. In the examples described below, even without special mention, can the above-described image setting process be applied in the same way, or can a modified application be possible? In the examples described below, the description will focus on cases where it is additional or with a modified application in the above-described examples.
[0419] For example, in the case of an image generated via a 360-degree camera {360-degree Video or Omnidirectional Video}, it has characteristics different from those of an image acquired via a general camera and has an encoding environment different from that of general image compression.
[0420] Unlike general images, a 360-degree image has no boundary parts with discontinuous characteristics, and the data in all regions can have continuity. Also, in devices such as an HMD, an image is reproduced in front of the eyes via a lens, and high-quality images can be required. When an image is acquired via a stereoscopic camera, the processed image data can increase. For the purpose of providing an efficient encoding environment including the above examples, various image setting processes considering 360-degree images can be performed.
[0421] The 360-degree camera is a camera having a plurality of cameras or a plurality of lenses and sensors, and the camera or lens can handle all directions around an arbitrary central point captured by the camera.
[0422] A 360-degree image can be encoded using various methods. For example, it can be encoded using various image processing algorithms in a three-dimensional space, or it can also be converted into a two-dimensional space and encoded using various image processing algorithms. In the present invention, a method of converting a 360-degree image into a two-dimensional space for encoding / decoding will be mainly described.
[0423] The 360-degree image encoding device according to an embodiment of the present invention can be configured to include all or part of the configuration shown in FIG. 1, and can further include a preprocessing unit that performs preprocessing (Stitching, Projection, Region-wise Packing) on the input image. On the other hand, the 360-degree image decoding device according to an embodiment of the present invention can include all or part of the configuration shown in FIG. 2, and can further include a postprocessing unit that performs postprocessing (Rendering) before being decoded and reproduced as an output image.
[0424] To explain again, after the preprocessing process (Pre-processing) on the input image in the encoder, encoding can be performed and the bitstream for this can be transmitted. The bitstream transmitted from the decoder can be parsed for decoding, and an output image can be generated after going through the postprocessing process (Post-processing). At this time, the bitstream can record and transmit the information generated in the preprocessing process and the information generated in the encoding process, and the decoder can parse this and use it in the decoding process and the postprocessing process.
[0425] Next, the operation method of the 360-degree image encoder will be described in more detail. Since the operation method of the 360-degree image decoder is the reverse operation of the 360-degree image encoder, it can be easily derived by an ordinary technician and a detailed description will be omitted.
[0426] The input image can undergo stitching and projection processes in a three-dimensional projection structure (Projection Structure) in units of spheres (Sphere). Through the above processes, the image data on the three-dimensional projection structure can be projected onto a two-dimensional image.
[0427] The projected image can be configured to include all or part of the 360-degree content according to the encoding settings. At this time, the position information of the area (or pixel) arranged at the center of the projected image can be implicitly generated as a predetermined value, or the position information can be explicitly generated. Also, when configuring a projected image that includes a partial area of the 360-degree content, the range and position information of the included area can be generated. Further, range information (e.g., vertical width, horizontal width) and position information (e.g., measured based on the upper left side of the image) for the region of interest (ROI) in the projected image can be generated. At this time, a partial area with high importance among the 360-degree content can be set as the region of interest. Although the 360-degree image can view all the content in the up, down, left, and right directions, the user's line of sight can be limited to a part of the image, and this can be considered and set as the region of interest. For efficient encoding, the region of interest can be set to have good quality and resolution, and the other regions can be set to have lower quality and resolution than the region of interest.
[0428] Among 360-degree image transmission methods, in the single stream transmission method (Single Stream), the entire image or viewport image can be transmitted to the user as an individual single bit stream. In the multi stream transmission method (Multi Stream), by transmitting a plurality of entire images with different image qualities as multi-bit streams, the image quality can be selected according to the user's environment and communication situation. In the tiled stream transmission method, by transmitting individually encoded partial images in tile units as multi-bit streams, tiles can be selected according to the user's environment and communication situation. Therefore, the 360-degree image encoder can generate and transmit bit streams with two or more qualities, and the 360-degree image decoder can set the region of interest according to the user's line of sight and selectively decode according to the region of interest. That is, the place where the user's line of sight stays through a head tracking or eye tracking system can be set as the region of interest, and only the necessary part can be rendered.
[0429] The projected image can be converted into a packed image by performing a region-wise packing process. The region-wise packing process can include the step of dividing the projected image into a plurality of regions. At this time, each divided region can be arranged (or rearranged) in the packed image according to the settings of the region-wise packing. Region-wise packing can be performed for the purpose of enhancing spatial continuity when converting a 360-degree image into a two-dimensional image (or a projected image). The size of the image can be reduced through region-wise packing. Also, it can reduce the image quality degradation that occurs during rendering, enable viewport-based projection, and can be performed for the purpose of providing other types of projection formats. Region-wise packing may or may not be performed according to the encoding settings, and can be determined based on a signal (e.g., regionwise_packing_flag, in the example described later, region-wise packing related information can only occur when the regionwise_packing_flag is activated) that indicates whether to perform it.
[0430] When region-wise packing is performed, it is possible to display (or generate) setting information (or mapping information) such that a partial region of the projected image is assigned (or arranged) as a partial region of the packed image. When region-wise packing is not performed, the projected image and the packed image can be the same image.
[0431] In the above, the stitching, projection, and region-wise packing processes were defined as individual processes, but a part (e.g., stitching + projection, projection + region-wise packing) or all (e.g., stitching + projection + region-wise packing) of the above processes can be defined as one process.
[0432] According to the settings of the stitching, projection, regional packing process, etc., at least one packed image can be generated for the same input image. Also, according to the settings of the regional packing process, at least one encoded data for the same projected image can be generated.
[0433] A tiling process can be performed to divide the packed image. At this time, tiling is a process of dividing and transmitting an image into a plurality of regions, and can be an example of the 360-degree image transmission method. As described above, tiling can be performed for the purpose of partial decoding in consideration of the user's environment, etc., and can be performed for the purpose of efficient processing of the huge data of the 360-degree image. For example, when an image is composed of one unit, all of the image can be decoded for decoding the region of interest, but when the image is composed of a plurality of unit regions, it is efficient to decode only the region of interest. At this time, the division can be performed by dividing into tiles which are the division units by the existing encoding method, or by dividing into various division units (square division, blocks, etc.) described in the present invention. Also, the division unit can be a unit for performing independent encoding / decoding. Tiling can be performed based on the projected image or the packed image, or can be performed independently. That is, it can be divided based on the surface boundary of the projected image, the surface boundary of the packed image, the packing settings, etc., and can be divided independently for each division unit. This can affect the generation of division information in the tiling process.
[0434] Next, the projected image or the packed image can be encoded. The encoded data and the information generated in the preprocessing process can be recorded in a bitstream and transmitted to a 360-degree image decoder. The information generated in the preprocessing process may be recorded in the bitstream in the form of SEI or metadata. At this time, the bitstream can include at least one encoded data and at least one preprocessing information that vary some settings of the encoding process or some settings of the preprocessing process. This can be for the purpose of mixing multiple encoded data (encoded data + preprocessing information) according to the user's environment in the decoder to construct a decoded image. Specifically, multiple encoded data can be selectively combined to construct a decoded image. Also, the process may be performed separately into two for application in a stereoscopic system, or the process may be performed on an additional depth image.
[0435] FIG. 15 is an exemplary diagram showing a three-dimensional space indicating a three-dimensional image and a two-dimensional planar space.
[0436] Generally, for a 360-degree three-dimensional virtual space, 3DoF (Degree of Freedom) is required, and three rotations can be assisted around the X (Pitch), Y (Yaw), and Z (Roll) axes. DoF means the degree of freedom in space, 3DoF means the degree of freedom including rotations around the X, Y, and Z axes as shown in 15a, and 6DoF means the degree of freedom that further allows movement along the X, Y, and Z axes in addition to 3DoF. The image encoding device and the decoding device of the present invention will be mainly described for the case of 3DoF. When assisting 3DoF or more (3DoF+), it can be combined with additional processes or devices not shown in the present invention or applied with modifications.
[0437] Referring to 15a, Yaw can have a range from -π (-180 degrees) to π (180 degrees), Pitch can have a range from -π / 2 rad (or -90 degrees) to π / 2 rad (or 90 degrees), and Roll can have a range from -π / 2 rad (or -90 degrees) to π / 2 rad (or 90 degrees). At this time, assuming that ψ and θ are Longitude and Latitude in the geographical representation of the earth, the (x, y, z) in three-dimensional space can be converted from the (ψ, θ) in two-dimensional space. For example, the coordinates in three-dimensional space can be derived from the two-dimensional space coordinates based on the conversion formula of x = cos(θ)cos(ψ), y = sin(θ), and z = -cos(θ)sin(ψ).
[0438] Also, (ψ, θ) can be converted to (x, y, z). For example, ψ = tan -1 (-Z / X), θ = sin -1 (Y / (X 2 +Y 2 +Z 2 ) 1 / 2 ) Based on the conversion formula, the two-dimensional space coordinates can be derived from the three-dimensional space coordinates.
[0439] When the pixels in three-dimensional space are accurately converted to two-dimensional space (for example, integer unit pixels in two-dimensional space), the pixels in three-dimensional space can be mapped to the pixels in two-dimensional space. When the pixels in three-dimensional space are not accurately converted to two-dimensional space (for example, fractional unit pixels in two-dimensional space), the two-dimensional pixels can be mapped to the pixels obtained by interpolation. At this time, the interpolation methods that can be used include the Nearest neighbor interpolation method, Bi-linear interpolation method, B-spline interpolation method, Bi-cubic interpolation method, etc. At this time, any one can be selected from multiple interpolation candidates, and relevant information can be explicitly generated, or the interpolation method can be implicitly determined based on a predetermined rule. For example, a predetermined interpolation filter can be used according to the three-dimensional model, projection format, color format, slice / tile type, etc. Also, when explicitly generating interpolation information, information about the filter information (for example, filter coefficients, etc.) can also be included.
[0440] 15b shows an example of conversion from a three-dimensional space to a two-dimensional space (two-dimensional plane coordinate system). (ψ, θ) can be sampled at (i, j) based on the size (width, height) of the image, where i can range from 0 to P_Width - 1 and j can range from 0 to P_Height - 1.
[0441] (ψ, θ) can be the central point {or reference point, the point denoted as C in FIG. 15, with coordinates (ψ, θ) = (0, 0)} for the 360-degree image arrangement at the center of the projected image. The setting for the central point can be specified in the three-dimensional space, and the position information for the central point can be explicitly generated or determined as an implicitly already set value. For example, the central position information in Yaw, the central position information in Pitch, the central position information in Roll, etc. can be generated. If the value for the said information is not specifically specified, each value can be assumed to be 0.
[0442] In the above example, an example of converting the entire 360-degree image from a three-dimensional space to a two-dimensional space was described. However, a partial region of the 360-degree image can be targeted, and the position information for the partial region (for example, some positions belonging to the region. In this example, the position information relative to the central point), range information, etc. can be explicitly generated or follow the implicitly already set position and range information. For example, the central position information in Yaw, the central position information in Pitch, the central position information in Roll, the range information in Yaw, the range information in Pitch, the range information in Roll, etc. can be generated. In the case of a partial region, it is at least one region, whereby the position information, range information, etc. of a plurality of regions can be processed. If the value for the said information is not specifically specified, it can be assumed to be the entire 360-degree image.
[0443] H0 to H6 and W0 to W5 in 15a respectively indicate some latitudes and longitudes in 15b, and the coordinates of 15b can be expressed as (C, j) and (i, C) (C is a longitude or latitude component). Different from general images, when a 360-degree image is converted into a two-dimensional space, distortion or warping of the content in the image can occur. This varies according to the region of the image, and the encoding / decoding settings can be made different for the position of the image or the regions partitioned according to the position. When adaptively setting the encoding / decoding based on the encoding / decoding information in the present invention, the position information (for example, x, y components, or the range defined by x and y, etc.) can be included as an example of the encoding / decoding information.
[0444] The description of the three-dimensional and two-dimensional spaces is defined to assist in the description of the embodiments of the present invention, and is not limited thereto, and variations of the detailed content or applications in other cases are possible.
[0445] As described above, the image acquired by a 360-degree camera can be converted into a two-dimensional space. At this time, the 360-degree image can be mapped using a three-dimensional model, and various three-dimensional models such as a sphere, a cube, a cylinder, a pyramid, and a polyhedron can be used. When converting the 360-degree image mapped based on the model into a two-dimensional space, a projection process according to the projection format based on the model can be performed.
[0446] Figures 16a to 16d are conceptual diagrams for explaining a projection format according to an embodiment of the present invention.
[0447] FIG. 16a shows an ERP (Equi-Rectangular Projection) format in which a 360-degree image is projected onto a two-dimensional plane. FIG. 16b shows a (CMP CubeMap Projection) format in which a 360-degree image is projected onto a cube. FIG. 16c shows an OHP (OctaHedron Projection) format in which a 360-degree image is projected onto an octahedron. FIG. 16d shows an ISP (IcoSahedral Projection) format in which a 360-degree image is projected onto an icosahedron. However, it is not limited to this, and various projection formats can be used. The left sides of FIGS. 16a to 16d show a 3D model, and the right sides show examples converted into a two-dimensional space through the projection process. Depending on the projection format, they have various sizes and shapes, and each shape can be composed of faces or surfaces, and the surfaces can be represented by circles, triangles, quadrilaterals, etc.
[0448] In the present invention, the projection format can be defined by a 3D model, settings of the surface (for example, the number of surfaces, the form of the surface, the form composition of the surface, etc.), settings of the projection process, etc. When at least one of the elements of the above definition is different, it can be regarded as a different projection format. For example, in the case of ERP, it is composed of a spherical model (3D model), one surface (the number of surfaces), and a quadrilateral surface (the pattern of the surface). However, if a part of the settings in the projection process (for example, the mathematical formula used when converting from a three-dimensional space to a two-dimensional space, that is, the remaining projection settings are the same, and the element that creates a difference in at least one pixel of the projected image in the projection process) is different, it can be classified into different formats such as ERP1 and EPR2. As another example, in the case of CMP, it is composed of a cube model, six surfaces, and a square surface. However, if a part of the settings in the projection process (for example, the sampling method when converting from a three-dimensional space to a two-dimensional space) is different, it can be classified into different formats such as CMP1 and CMP2.
[0449] When using multiple projection formats instead of a single already-set projection format, projection format identification information (or projection format information) can be explicitly generated. The projection format identification information can be configured in various ways.
[0450] As an example, index information (e.g., proj_format_flag) can be assigned to multiple projection formats to identify the projection formats. For example, 0 can be assigned to ERP, 1 to CMP, 2 to OHP, 3 to ISP, 4 to ERP1, 5 to CMP1, 6 to OHP1, 7 to ISP1, 8 to CMP compact, 9 to OHP compact, 10 to ISP compact, and 11 or higher to other formats.
[0451] As an example, the projection format can be identified from at least one element information that constitutes the projection format. At this time, the element information that constitutes the projection format may include 3D model information (e.g., 3d_model_flag. 0 is a sphere, 1 is a cube, 2 is a cylinder, 3 is a pyramid, 4 is polyhedron 1, 5 is polyhedron 2, etc.), the number of surface information (e.g., num_face_flag. Starting from 1 and increasing by 1 each time, or assigning the number of surfaces generated in the projection format as index information, 0 is 1, 1 is 3, 2 is 6, 3 is 8, 4 is 20, etc.), the surface form information (e.g., shape_face_flag. 0 is a quadrilateral, 1 is a circle, 2 is a triangle, 3 is a quadrilateral + circle, 4 is a quadrilateral + triangle, etc.), the projection process setting information (e.g., 3d_2d_convert_idx, etc.).
[0452] As an example, the projection format can be identified by the projection format index information and the element information that constitutes the projection format. For example, the projection format index information can assign 0 to ERP, 1 to CMP, 2 to OHP, 3 to ISP, and 4 or more to other formats. Together with the element information that constitutes the projection format (in this example, the projection process setting information), the projection format (for example, ERP, ERP1, CMP, CMP1, OHP, OHP1, ISP, ISP1, etc.) can be identified. Or, together with the element information that constitutes the projection format (in this example, whether it is regional packing or not), the projection format (for example, ERP, CMP, CMP compact, OHP, OHP compact, ISP, ISP compact, etc.) can be identified.
[0453] In summary, the projection format can be identified by the projection format index information, can be identified by at least one projection format element information, and can be identified by the projection format index information and at least one projection format element information. This can be defined according to the encoding / decoding settings. In the present invention, the case of being identified by the projection format index will be assumed for explanation. Also, in this example, the explanation will focus on the case of the projection format represented by surfaces having the same size and shape, but a configuration in which the sizes and shapes of each surface are not the same is also possible. Also, the configuration of each surface may be the same as or different from FIGS. 16a to 16d. The numbers of each surface are used as symbols for identifying each surface and are not limited to a specific order. For the sake of convenience of explanation, in the examples described later, based on the projection image, ERP is a projection format of one surface + a quadrilateral, CMP is a projection format of six surfaces + a quadrilateral, OHP is a projection format of eight surfaces + a triangle, and ISP is a projection format of twenty surfaces + a triangle. The case where the surfaces have the same size and shape will be assumed for explanation, but the same or similar application is also possible for other settings.
[0454] As shown in FIGS. 16a to 16d, the projection format can be divided into one surface (e.g., ERP) or multiple surfaces (e.g., CMP, OHP, ISP, etc.). Also, each surface can be divided into shapes such as a quadrilateral and a triangle. The division can be an example of the type, characteristics, etc. of the image in the present invention when it is made different from the encoding / decoding setting according to the projection format. For example, the type of the image can be a 360-degree image, and the characteristic of the image can be any of the above divisions (e.g., each projection format, a projection format of one surface or multiple surfaces, a projection format where the surface is a quadrilateral or not a quadrilateral, etc.).
[0455] The two-dimensional plane coordinate system {e.g., (i, j)} can be defined on each surface of the two-dimensional projection image, and the characteristics of the coordinate system can vary depending on the projection format, the position of each surface, etc. In the case of ERP, there can be one two-dimensional plane coordinate system, and for other projection formats, multiple two-dimensional plane coordinate systems can be possessed according to the number of surfaces. At this time, the coordinate system can be expressed as (k, i, j), where k can be the index information of each surface.
[0456] FIG. 17 is a conceptual diagram showing that the projection format according to an embodiment of the present invention is included in a rectangular image.
[0457] That is, FIGS. 17a to 17c can be understood as realizing the projection formats of FIGS. 16b to 16d as a rectangular image.
[0458] Referring to FIGS. 17a to 17c, for the encoding / decoding of the 360-degree image, each image format can be configured in a rectangular shape. In the case of ERP, it can be used as it is in one coordinate system, but in the case of other projection formats, the coordinate systems of each surface can be integrated into one coordinate system, and detailed description thereof will be omitted.
[0459] Referring to FIGS. 17a to 17c, it can be confirmed that in the process of constructing a rectangular image, areas filled with meaningless data such as blanks and backgrounds are generated. That is, it can be composed of an area containing actual data (in this example, the surface, Active Area) and a meaningless area filled to form a rectangular image (in this example, assumed to be filled with arbitrary pixel values, Inactive Area). This may result in a performance degradation due to not only the encoding / decoding of actual image data but also an increase in the amount of encoded data caused by an increase in the image size due to the meaningless area.
[0460] Therefore, a process for excluding meaningless areas and constructing an image with an area containing actual data can be further performed.
[0461] FIG. 18 is a conceptual diagram of a method for converting a projection format according to an embodiment of the present invention into a rectangular shape, the method of rearranging the surface so as to exclude meaningless areas.
[0462] Referring to FIGS. 18a to 18c, an example of rearranging FIGS. 17a to 17c can be confirmed, and such a process can be defined as a regional packing process (such as CMP compact, OHP compact, ISP compact, etc.). At this time, not only the rearrangement of the surface itself but also the surface can be divided and rearranged (such as OHP compact, ISP compact, etc.). This can be done for the purpose of not only removing meaningless areas but also improving encoding performance through efficient placement of the surface. For example, when arranging images with continuity between surfaces (for example, B2 - B3 - B1, B5 - B0 - B4 in FIG. 18a), the encoding performance can be improved by improving the prediction accuracy during encoding. Here, the regional packing according to the projection format is only an example in the present invention and is not limited thereto.
[0463] FIG. 19 is a conceptual diagram showing the process of performing regional packing by converting the CMP projection format according to an embodiment of the present invention into a rectangular image.
[0464] Referring to 19a to 19c, the CMP projection formats can be arranged as 6×1, 3×2, 2×3, 1×6. Also, when size adjustment is performed on some surfaces, they can be arranged as 19d to 19e. Although CMP is taken as an example in 19a to 19e, it is not limited to CMP and can be applied to other projection formats. The surface arrangement of the image obtained through the regional packing follows a predetermined rule according to the projection format, or information regarding the arrangement can be explicitly generated.
[0465] The 360-degree image encoding / decoding device according to an embodiment of the present invention can be configured to include all or part of the image encoding / decoding device according to FIGS. 1 and 2. In particular, a format conversion unit and a format inverse conversion unit for converting and inversely converting the projection format can be further included in the image encoding device and the image decoding device, respectively. That is, in the image encoding device of FIG. 1, the input image can be encoded through the format conversion unit, and after the bit stream is decoded in the image decoding device of FIG. 2, the output image can be generated through the format inverse conversion unit. Hereinafter, the encoder for the above process (in this example, "input image" to "encoding") will be mainly described, and the process in the decoder can be derived inversely from the encoder. Also, descriptions overlapping with the above will be omitted.
[0466] Next, the input image will be described on the premise that it is the same as the two-dimensional projection image or the packing image obtained through the preprocessing process in the above-described 360-degree encoding device. That is, the input image can be an image obtained through a projection process or a regional packing process according to some projection formats. The projection format already applied to the input image can be any of various projection formats, and may be regarded as a common format or may be called the first format.
[0467] The format conversion unit can perform conversion to other projection formats other than the first format. At this time, the format to be converted can be called the second format. For example, if ERP is set as the first format, it can be converted to the second format (for example, ERP2, CMP, OHP, ISP, etc.). At this time, ERP2 may be an EPR format that has the same conditions such as the same 3D model and surface configuration but has some different settings. Or, it may be the same format with the same projection format settings (for example, ERP = ERP2), but the image size or resolution may be different. Or, a part of the image setting process described later may be applied. For the convenience of explanation, the above examples are given, but the first format and the second format are one of various projection formats, not limited to the above examples, and can be changed to other cases.
[0468] Due to the characteristics of different coordinate systems between projection formats in the conversion process between formats, the pixels (integer pixels) of the converted image may be obtained not only from integer unit pixels in the pre-conversion image but also from fractional unit pixels, so interpolation can be performed. At this time, the interpolation filter used can be the same or similar filter as described above. The interpolation filter can select any one from a plurality of interpolation filter candidates, and the relevant information can be explicitly generated or implicitly determined according to the preset rules. For example, a predetermined interpolation filter can be used according to the projection format, color format, slice / tile type, etc. Also, when explicitly sending an interpolation filter, information about the filter information (for example, filter coefficients, etc.) may also be included.
[0469] The projection format in the format conversion unit may be defined including regional packing, etc. That is, a projection and regional packing process may be performed in the format conversion process. Or, after the format conversion and before the encoding, a process such as regional packing may be performed.
[0470] In the symbolizer, the information generated in the above process is recorded in a bitstream in at least one unit among units such as sequences, pictures, slices, tiles, etc., and in the decoder, the relevant information is parsed from the bitstream. Also, it may be included in the bitstream in the form of SEI or metadata.
[0471] Next, an image setting process applied to a 360-degree image encoding / decoding apparatus according to an embodiment of the present invention will be described. The image setting process in the present invention can be applied not only to a general encoding / decoding process, but also to a preprocessing process, a postprocessing process, a format conversion process, a format inverse conversion process, etc. in a 360-degree image encoding / decoding apparatus. The image setting process to be described later will be described centering on the 360-degree image encoder, and can be described including the content in the above-described image setting. Duplicate descriptions in the above-described image setting process will be omitted. Also, the examples to be described later will be described centering on the image setting process, and the inverse image setting process can be induced inversely from the image setting process, and in some cases, can be confirmed through various embodiments of the present invention described above.
[0472] The image setting process in the present invention may be performed in the projection stage of the 360-degree image, may be performed in the regional packing stage, may be performed in the format conversion stage, or may be performed in other stages.
[0473] FIG. 20 is a conceptual diagram of 360-degree image division according to an embodiment of the present invention. In FIG. 20, the case of an image projected by ERP will be described by way of assumption.
[0474] 20a shows an image projected by ERP, and division can be performed using various methods. In this example, the description will be centered on slices and tiles, and it is assumed that W0 to W2 and H0, H1 are the division boundary lines of slices or tiles, and it is assumed that they follow the raster scan order. The examples to be described later will be described centering on slices and tiles, but are not limited thereto, and other division methods can be applied.
[0475] For example, it can be divided in slice units and can have division boundaries of H0 and H1. Or, it can be divided in tile units and can have division boundaries of W0 to W2 and H0 and H1.
[0476] 20b shows an example in which the image projected by ERP is divided into tiles {assuming the same tile division boundaries (all of W0 to W2 and H0 and H1 are activated) as in Fig. 20a}. Assuming that the P region is the entire image and the V region is the region where the user's line of sight stays or the viewport, there can be various methods to provide the image corresponding to the viewport. For example, the entire image (e.g., tiles a to l) can be decoded to obtain the region corresponding to the viewport. At this time, the entire image can be decoded, and if it is divided, tiles a to l (in this example, the A + B region) can be decoded. Or, by decoding the region belonging to the viewport, the region corresponding to the viewport can be obtained. At this time, if it is divided, by decoding tiles f, g, j, and k (in this example, the B region), the region corresponding to the viewport can be obtained from the restored image. The former case is called whole decoding (or Viewport Independent Coding), and the latter case is called partial decoding (or Viewport Dependent Coding). The latter case is an example that can occur in a 360-degree image with a large amount of data, and since the divided region can be obtained flexibly, the tile unit division method can be used more frequently than the slice unit division. In the case of partial decoding, since it is not possible to know where the viewport is generated, the referability of the division unit can be limited spatially or temporally (impliedly processed in this example), and encoding / decoding can be performed considering this. The examples described later will be mainly explained for the case of whole decoding, but for the purpose of preparing for the case of partial decoding, the division of the 360-degree image will be explained centering on tiles (or the square division method of the present invention). The content of the examples described later can be similarly applied or modified to other division units.
[0477] Figure 21 is an exemplary diagram of 360-degree image segmentation and image reconstruction according to an embodiment of the present invention. In Figure 21, the case of an image projected by CMP will be described by way of assumption.
[0478] 21a shows an image projected by CMP, and the segmentation can be performed using various methods. Assume that W0 to W2 and H0, H1 are the segmentation boundary lines of the surface, slice, and tile, and assume that they follow the raster scan order.
[0479] For example, the segmentation can be performed in units of slices and can have the segmentation boundaries of H0 and H1. Or, the segmentation can be performed in units of tiles and can have the segmentation boundaries of W0 to W2 and H0, H1. Or, the segmentation can be performed in units of surfaces and can have the segmentation boundaries of W0 to W2 and H0, H1. In this example, the surface will be described by assuming it is part of the segmentation unit.
[0480] At this time, the surface is a segmentation unit (in this example, dependent encoding / decoding) performed for the purpose of classifying or dividing regions having different properties (for example, the plane coordinate system of each surface, etc.) within the same image according to the characteristics and types of the image (in this example, 360-degree image, projection format, etc.). The slice and tile can be segmentation units (in this example, independent encoding / decoding) performed for the purpose of dividing the image according to the user's definition. Also, the surface is a unit divided by a predetermined definition (or derived from projection format information) in the projection process according to the projection format, and the slice and tile can be units that explicitly generate segmentation information according to the user's definition and are divided. Also, the surface can have a segmentation shape in the form of a polygon including a quadrilateral according to the projection format, the slice can have an arbitrary segmentation shape that cannot be defined as a quadrilateral or polygon, and the tile can have a square segmentation shape. The setting of the segmentation unit can be content defined and limited for the purpose of explaining this example.
[0481] In the above example, the surface was described as a segmentation unit classified for the purpose of area division. However, depending on the encoding / decoding settings, it can be a unit that performs independent encoding / decoding in at least one surface unit, or it can have a setting that combines with tiles, slices, etc. to perform independent encoding / decoding. At this time, it can occur that explicit information of tiles, slices, etc. is generated in the combination with tiles, slices, etc., or it can occur that an implicit case where tiles, slices are combined based on surface information. Or, it may occur that explicit information of tiles, slices is generated based on surface information.
[0482] As a first example, one image segmentation process (in this example, the surface) is performed, and the image segmentation can implicitly omit the segmentation information (obtain the segmentation information from the projection format information). This example is an example for the setting of dependent encoding / decoding, and can be an example corresponding to the case where the referability between surface units is not restricted.
[0483] As a second example, one image segmentation process (in this example, the surface) is performed, and the image segmentation can explicitly generate the segmentation information. This example is an example for the setting of dependent encoding / decoding, and can be an example corresponding to the case where the referability between surface units is not restricted.
[0484] As a third example, a plurality of image segmentation processes (in this example, the surface, tiles) are performed. Some of the image segmentations (in this example, the surface) can implicitly omit the segmentation information or explicitly generate it, and some of the image segmentations (in this example, the tiles) can explicitly generate the segmentation information. In this example, some of the image segmentation processes (in this example, the surface) precede some of the image segmentation processes (in this example, the tiles).
[0485] As a fourth example, multiple image segmentation processes are performed. Some image segmentations (in this example, the surface) can implicitly omit segmentation information or explicitly generate it, and some image segmentations (in this example, the tile) can explicitly generate segmentation information based on some image segmentations (in this example, the surface). In this example, some image segmentation processes (in this example, the surface) precede some image segmentation processes (in this example, the tile). Although it is the same that segmentation information is explicitly generated in some cases {assuming the case of the second example}, there may be differences in the segmentation information configuration.
[0486] As a fifth example, multiple image segmentation processes are performed. Some image segmentations (in this example, the surface) can implicitly omit segmentation information, and some image segmentations (in this example, the tile) can implicitly omit segmentation information based on some image segmentations (in this example, the surface). For example, individual surface units can be set in tile units, or multiple surface units (in this example, when adjacent surfaces are continuous, they are grouped, and if not, they are not grouped. B2 - B3 - B1 and B4 - B0 - B5 in 18a) can be set in tile units. According to the already set rules, the setting of surface units to tile units is possible. This example is an example for independent encoding / decoding settings and may be an example applicable when the referability between surface units is restricted. That is, although it is the same that segmentation information is implicitly processed in some cases {assuming the case of the first example}, there may be differences in the encoding / decoding settings.
[0487] The above examples are explanations for cases where the segmentation process can be performed in the projection stage, the regional packing stage, the encoding / decoding initial stage, etc., and can be image segmentation processes occurring in other encoding / decoders.
[0488] In 21a, by including an area B that does not contain data in an area A that contains data, a rectangular image can be formed. At this time, the position, size, shape, number, etc. of areas A and B are information that can be confirmed by a projection format or the like, or information that can be confirmed when explicitly generating information for the projected image, and related information can be indicated by the above-described image segmentation information, image reconstruction information, etc. For example, information for a partial area of the projected image (e.g., part_top, part_left, part_width, part_height, part_convert_flag, etc.) can be indicated as shown in Tables 4 and 5, and this is not limited to this example and can be an example applicable to other cases (e.g., other projection formats, other projection settings, etc.).
[0489] The B area can be combined with the A area into one image for encoding / decoding. Or, by performing division considering the characteristics of each area, different encoding / decoding settings can be set. For example, encoding / decoding for the B area may not be performed via information on whether to perform encoding / decoding (e.g., tile_coded_flag assuming that the division unit is a tile). At this time, the corresponding area can be restored to certain data (in this example, an arbitrary pixel value) according to the set rules. Or, the encoding / decoding settings in the above-described image segmentation process can be made different for the B area and the A area. Or, a region-by-region packing process can be performed to remove the corresponding area.
[0490] 21b shows an example in which an image packed by CMP is divided into tiles, slices, and surfaces. At this time, the packed image is an image in which a surface rearrangement process or a region-by-region packing process has been performed, and can be an image obtained by performing image segmentation and image reconstruction of the present invention.
[0491] It is possible to form a rectangular shape including an area containing data in 21b. At this time, the position, size, shape, number, etc. of each area are information that can be confirmed by settings that have already been set, or information that can be confirmed when explicitly generating information for the packed image, and can indicate related information with the above-described image segmentation information, image reconstruction information, etc. For example, information (such as part_top, part_left, part_width, part_height, part_convert_flag, etc.) for a partial area of the packed image as shown in Tables 4 and 5 can be indicated.
[0492] The packed image can be divided using various division methods. For example, it can be divided in slice units and can have a division boundary of H0. Or, it can be divided in tile units and can have division boundaries of W0, W1, and H0. Or, it can be divided in surface units and can have division boundaries of W0, W1, and H0.
[0493] The image segmentation and image reconstruction processes of the present invention can be performed on the projected image. At this time, the reconstruction process can rearrange not only the pixels within the surface but also the surfaces within the image. This can be an example possible when the image is divided or composed of a plurality of surfaces. The examples described later will mainly explain the case where it is divided into tiles based on surface units.
[0494] SX,Y (S0,0 to S3,2) of 21a can correspond to S’U,V (S’0,0 to S’2,1. In this example, X, Y may be the same as or different from U, V.), and the reconstruction process can be performed on a surface unit. For example, S2,1, S3,1, S0,1, S1,2, S1,1, S1,0 can be assigned (or surface rearrangement) to S’0,0, S’1,0, S’2,0, S’0,1, S’1,1, S’2,1. Also, S2,1, S3,1, S0,1 can be reconstructed without performing reconstruction (or pixel rearrangement), and S1,2, S1,1, S1,0 can be reconstructed by applying a 90-degree rotation, which can be shown as in Figure 21c. The symbols (S1,0, S1,1, S1,2) displayed horizontally in 21c can be images laid horizontally according to the symbols to maintain the continuity of the image.
[0495] The reconstruction of the surface can perform implicit or explicit processing according to the encoding / decoding settings. In the case of implicit processing, it can be performed according to the already set rules considering the type of the image (in this example, a 360-degree image), characteristics (in this example, projection format, etc.).
[0496] For example, between S’0,0 and S’1,0, S’1,0 and S’2,0, S’0,1 and S’1,1, S’1,1 and S’2,1 in 21c, there is image continuity (or correlation) between the two surfaces based on the surface boundary, and 21c can be an example configured such that there is continuity between the upper three surfaces and the lower three surfaces. It is divided into a plurality of surfaces through the projection process from a three-dimensional space to a two-dimensional space, and reconstruction can be performed for the purpose of enhancing the image continuity between the surfaces for efficient surface reconstruction in the process through a regional packing process. Such surface reconstruction can be already set and processed.
[0497] Or, the reconstruction process can be performed by explicit processing, and reconstruction information about this can be generated.
[0498] For example, when checking information (either implicitly obtained information or explicitly generated information) for an M×N configuration (e.g., in the case of CMP compact, 6×1, 3×2, 2×3, 1×6, etc. Assume a 3×2 configuration in this example) through a regional packing process, after reconstructing the surface according to the M×N configuration, information regarding it can be generated. For example, in the case of in-surface image rearrangement, index information (or position information within the image) can be assigned to each surface, and in the case of in-surface pixel rearrangement, mode information for the reconstruction can be assigned.
[0499] The index information can already be defined as shown in 18a to 18c of FIG. 18, and SX,Y or S’U,V in 21a to 21c can represent each surface with position information indicating horizontal and vertical (e.g., S[i][j]) or one position information (e.g., assuming the position information is assigned in raster scan order from the upper left surface of the image. S[i]), and an index for each surface can be assigned to this.
[0500] For example, when assigning an index to position information indicating horizontal and vertical, in the case of FIG. 21c, S’0,0 can be assigned the index of the 2nd surface, S’1,0 can be assigned the index of the 3rd surface, S’2,0 can be assigned the index of the 1st surface, S’0,1 can be assigned the index of the 5th surface, S’1,1 can be assigned the index of the 0th surface, and S’2,1 can be assigned the index of the 4th surface. Or, when assigning an index to one position information, S[0] can be assigned the index of the 2nd surface, S[1] can be assigned the index of the 3rd surface, S[2] can be assigned the index of the 1st surface, S[3] can be assigned the index of the 5th surface, S[4] can be assigned the index of the 0th surface, and S[5] can be assigned the index of the 4th surface. For the sake of convenience in explanation, in the examples described later, S’0,0 to S’2,1 are referred to as a to f. Or, it can also be expressed with position information indicating horizontal and vertical in terms of pixels or blocks based on the upper left side of the image.
[0501] In the case of a packed image obtained through an image reconstruction process (or a regional packing process), depending on the reconstruction settings, the surface scan order may or may not be the same in the image. For example, when one scan order (e.g., raster scan) is applied to 21a, the scan orders of a, b, and c may be the same, and the scan orders of d, e, and f may not be the same. For example, in the case of 21a, a, b, and c, when the scan order follows the order of (0, 0) → (1, 0) → (0, 1) → (1, 1), in the case of d, e, and f, the scan order can follow the order of (1, 0) → (1, 1) → (0, 0) → (0, 1). This can be determined according to the image reconstruction settings, and such settings can also be applied to other projection formats.
[0502] The image segmentation process in 21b can set individual surface units as tiles. For example, surfaces a to f can be set in tile units respectively. Or, units of multiple surfaces can be set as tiles. For example, surfaces a to c can be set as one tile, and d to f can be set as one tile. The above configuration can be determined based on surface characteristics (e.g., continuity between surfaces, etc.), and different tile settings for surfaces are possible compared to the above examples.
[0503] Next, an example of the segmentation information by multiple image segmentation processes is shown. In this example, the segmentation information for the surface is omitted, and it is assumed that units other than the surface are tiles and the segmentation information is processed in various ways for explanation.
[0504] As a first example, the image segmentation information can be obtained based on the surface information and implicitly omitted. For example, individual surfaces can be set as tiles, or multiple surfaces can be set as tiles. At this time, when at least one surface is set as a tile, it can be determined according to a predetermined rule based on the surface information (e.g., continuity or correlation, etc.).
[0505] As a second example, the image segmentation information can be explicitly generated regardless of the surface information. For example, when generating the segmentation information based on the number of horizontal rows of tiles (in this example, num_tile_columns) and the number of vertical rows (in this example, num_tile_rows), the segmentation information can be generated by the method in the aforementioned image segmentation process. For example, the possible range that the number of horizontal rows and the number of vertical rows of tiles can have can be from 0 to the width of the image / the width of the block (in this example, the unit obtained from the picture segmentation unit), and from 0 to the height of the image / the height of the block. Also, additional segmentation information (such as uniform_spacing_flag, etc.) can be generated. At this time, depending on the segmentation setting, it may occur that the boundary of the surface coincides with or does not coincide with the boundary of the segmentation unit.
[0506] As a third example, the image segmentation information can be explicitly generated based on the surface information. For example, when generating the segmentation information based on the number of horizontal rows of tiles and the number of vertical rows, the segmentation information can be generated based on the surface information (in this example, the range of the number of horizontal rows is 0 to 2, and the range of the number of vertical rows is 0, 1. Since the configuration of the surface in the image is 3x2). For example, the possible range that the number of horizontal rows and the number of vertical rows of tiles can have can be from 0 to 2, and from 0 to 1. Also, additional segmentation information (such as uniform_spacing_flag, etc.) may not be generated. At this time, the boundary of the surface can coincide with the boundary of the segmentation unit.
[0507] In some cases {assuming the cases of the second example and the third example}, it can be defined that the syntax elements of the segmentation information are different, or even if the same syntax elements are used, the settings of the syntax elements (such as binarization settings, etc. When the range of the candidate group that the syntax element has is limited and small, other binarizations can be used, etc.) can be made different. The above examples have explained a part of various configurations of the segmentation information, but it is not limited to this, and it can be understood as an example where different settings are possible depending on whether the segmentation information is generated based on the surface information.
[0508] FIG. 22 is an exemplary diagram showing a CMP-projected image or a packed image divided into tiles.
[0509] At this time, assume that it has the same tile division boundaries (all of W0 to W2, H0, and H1 are activated) as 21a in FIG. 21, and assume that it has the same tile division boundaries (all of W0, W1, and H0 are activated) as 21b in FIG. 21. When the P region is the entire image and the V region is the viewport, whole decoding or partial decoding can be performed. This example will be mainly described with partial decoding. In 22a, in the case of CMP (left), tiles e, f, and g are decoded, and in the case of CMP compact (right), tiles a, c, and e are decoded, so that the region corresponding to the viewport can be obtained. In 22b, in the case of CMP, tiles b, f, and i are decoded, and in the case of CMP compact, tiles d, e, and f are decoded, so that the region corresponding to the viewport can be obtained.
[0510] In the above example, the case of performing division such as slicing and tiling based on the surface unit (or surface boundary) has been described. However, as shown in 20a of FIG. 20, it is also possible to perform division inside the surface (for example, in ERP, the image is composed of one surface; another projection format is composed of multiple surfaces), or to perform division including the surface boundary.
[0511] FIG. 23 is a conceptual diagram for explaining an example of resizing a 360-degree image according to an embodiment of the present invention. At this time, the case of an image projected by ERP will be assumed for explanation. Also, in the examples described later, the case of expansion will be mainly described.
[0512] The projected image can be resized using a scale factor or an offset factor according to the image resizing type. The image before resizing is P_Width×P_Height, and the image after resizing can be P’_Width×P’_Height.
[0513] In the case of the scale factor, after size adjustment using the scale factors for the horizontal and vertical widths of the image (in this example, horizontal a and vertical b), the horizontal width (P_Width × a) and vertical width (P_Height × b) of the image can be obtained. In the case of the offset factor, after size adjustment using the offset factors for the horizontal and vertical widths of the image (in this example, horizontal L, R, vertical T, B), the horizontal width (P_Width + L + R) and vertical width (P_Height + T + B) of the image can be obtained. Size adjustment can be performed using the already set method, or any one of a plurality of methods can be selected for size adjustment.
[0514] The data processing method in the example described later will be mainly explained for the case of the offset factor. In the case of the offset factor, in the data processing method, there may be methods such as filling using a predetermined pixel value, filling by copying the outer contour pixels, filling by copying a partial area of the image, and filling by converting a partial area of the image.
[0515] In the case of a 360-degree image, size adjustment can be performed considering the characteristic of continuity existing at the boundary of the image. In the case of ERP, although there is no outer boundary in three-dimensional space, when it is converted to two-dimensional space through the projection process, an outer boundary region can exist. The data in the boundary region has continuous data outside the boundary, but due to spatial characteristics, it can have a boundary. Size adjustment can be performed considering such characteristics. At this time, the continuity can be confirmed according to the projection format, etc. For example, in the case of ERP, it can be an image in which the boundaries at both ends have continuous characteristics with each other. In this example, the case where the left and right boundaries of the image are continuous and the case where the upper and lower boundaries of the image are continuous are assumed for explanation, and the data processing method will be mainly explained with methods such as filling by copying a partial area of the image and filling by converting a partial area of the image.
[0516] When resizing to the left side of the image, the resized area (in this example, LC or TL+LC+BL) can be filled using the data from the right side area of the image (in this example, tr+rc+br) that is continuous with the left side of the image. When resizing to the right side of the image, the resized area (in this example, RC or TR+RC+BR) can be filled using the data from the left side area of the image (in this example, tl+lc+bl) that is continuous with the right side. When resizing to the upper side of the image, the resized area (in this example, TC or TL+TC+TR) can be filled using the data from the lower side area of the image (in this example, bl+bc+br) that is continuous with the upper side. When resizing to the lower side of the image, the data of the resized area (in this example, BC or BL+BC+BR) can be used for filling.
[0517] When the size or length of the resized area is m, the resized area can have a range of (-m, y) to (-1, y) (resizing to the left side) or (P_Width, y) to (P_Width+m-1, y) (resizing to the right side) based on the coordinate reference of the image before resizing (in this example, x is from 0 to P_Width-1). The position x' of the area for obtaining the data of the resized area can be derived by the formula x'=(x+P_Width)%P_Width. At this time, x means the coordinate of the resized area based on the image coordinates before resizing, and x' means the coordinate of the area referred to the resized area based on the image coordinates before resizing. For example, when resizing to the left side and m is 4 and the width of the image is 16, the corresponding data can be obtained from (12, y) for (-4, y), (13, y) for (-3, y), (14, y) for (-2, y), and (15, y) for (-1, y). Or when resizing to the right side and m is 4 and the width of the image is 16, the corresponding data can be obtained from (0, y) for (16, y), (1, y) for (17, y), (2, y) for (18, y), and (3, y) for (19, y).
[0518] When the size or length of the resized area is n, the resized area can have a range of (x, -n) to (x, -1) (resized upward) or (x, P_Height) to (x, P_Height + n - 1) (resized downward) with reference to the coordinates of the pre-resized image (in this example, y ranges from 0 to P_Height - 1). The position y' of the area for obtaining the data of the resized area can be derived by an equation such as y' = (y + P_Height) % P_Height. At this time, y means the coordinates of the resized area with reference to the pre-resized image coordinates, and y' means the coordinates of the area referred to by the resized area with reference to the pre-resized image coordinates. For example, when resizing upward and n is 4 and the vertical width of the image is 16, data can be obtained from (x, -4) corresponding to (x, 12), (x, -3) corresponding to (x, 13), (x, -2) corresponding to (x, 14), and (x, -1) corresponding to (x, 15). Or, when resizing downward and n is 4 and the vertical width of the image is 16, data can be obtained from (x, 16) corresponding to (x, 0), (x, 17) corresponding to (x, 1), (x, 18) corresponding to (x, 2), and (x, 19) corresponding to (x, 3).
[0519] After filling the data of the resized area, it can be adjusted with reference to the coordinates of the post-resized image (in this example, x ranges from 0 to P'_Width - 1 and y ranges from 0 to P'_Height - 1). The above example may be an example applicable to a latitude and longitude coordinate system.
[0520] The following various combinations of resizing can be available.
[0521] As an example, resizing can be performed by m to the left side of the image. Or, resizing can be performed by n to the right side of the image. Or, resizing can be performed by o to the upper side of the image. Or, resizing can be performed by p to the lower side of the image.
[0522] As an example, resizing can be performed by m to the left side and by n to the right side of the image. Or, resizing can be performed by o to the upper side and by p to the lower side of the image.
[0523] As an example, the size of the image can be adjusted by m to the left side, n to the right side, and o to the upper side. Or, the size of the image can be adjusted by m to the left side, n to the right side, and p to the lower side. Or, the size of the image can be adjusted by m to the left side, o to the upper side, and p to the lower side. Or, the size of the image can be adjusted by n to the right side, o to the upper side, and p to the lower side.
[0524] As an example, the size of the image can be adjusted by m to the left side, n to the right side, o to the upper side, and p to the lower side.
[0525] At least one size adjustment is performed as in the above example, and the size of the image can be implicitly adjusted according to the encoding / decoding settings, or size adjustment information can be explicitly generated, and the size of the image can be adjusted based on it. That is, m, n, o, and p in the above example can be determined to be predetermined values, or can be explicitly generated as size adjustment information, or some can be determined to be predetermined values and some can be explicitly generated.
[0526] The above example has been mainly described for the case of obtaining data from a partial area of the image, but other methods are also applicable. The data is either the pixel before encoding or the pixel after encoding, and can be determined according to the characteristics of the image or stage for which the size adjustment is performed. For example, when performing size adjustment in the preprocessing process or the stage before encoding, the data means input pixels such as a projection image and a packing image, and when performing size adjustment in the postprocessing process, the stage of generating in-picture prediction reference pixels, the stage of generating a reference image, the filter stage, etc., the data can mean restored pixels. Also, size adjustment can be performed using a data processing method for each area to be size-adjusted.
[0527] FIG. 24 is a conceptual diagram for explaining the continuity between surfaces in a projection format (for example, CMP, OHP, ISP) according to an embodiment of the present invention.
[0528] Specifically, it can be an example for an image composed of a plurality of surfaces. Continuity is a characteristic that occurs in adjacent regions in three-dimensional space. FIGS. 24a to 24c can be classified into cases where they are spatially adjacent and continuous (A) when converted into two-dimensional space through the projection process, cases where they are spatially adjacent but not continuous (B), cases where they are not spatially adjacent but continuous (C), and cases where they are not spatially adjacent and not continuous (D). General images are different from those classified into cases where they are spatially adjacent and continuous (A) and cases where they are not spatially adjacent and not continuous (D). At this time, when continuity exists, the above-mentioned partial examples (A or C) apply.
[0529] That is, referring to FIGS. 24a to 24c, cases where they are spatially adjacent and continuous (described with reference to 24a in this example) are displayed as b0 to b4, and cases where they are not spatially adjacent and continuous can be displayed as B0 to B6. That is, it means cases for regions adjacent in three-dimensional space. By using the characteristics of b0 to b4 and B0 to B6 having continuity in the encoding process, the encoding performance can be improved.
[0530] FIG. 25 is a conceptual diagram for explaining the continuity of the surface of FIG. 21c, which is an image obtained through the image reconstruction process or the regional packing process in the CMP projection format.
[0531] Here, since 21c in FIG. 21 is a rearrangement of 21a in which a 360-degree image is expanded in the shape of a cube, the continuity of the surface of 21a in FIG. 21 is also maintained at this time. That is, as in 25a, the surface S2,1 can be continuous with S1,1 and S3,1 on the left and right, and can be continuous with the surface S1,0 rotated 90 degrees and the surface S1,2 rotated -90 degrees above and below.
[0532] In the same way, the continuity of the surfaces S3,1, S0,1, S1,2, S1,1, and S1,0 can be confirmed from FIGS. 25b to 25f.
[0533] The continuity between surfaces can be defined according to the projection format setting or the like, and is not limited to the above example, and other modified examples are possible. The examples described later will be explained under the assumption that there is continuity as shown in FIGS. 24 and 25.
[0534] FIG. 26 is an exemplary diagram for explaining the size adjustment of an image in the CMP projection format according to an embodiment of the present invention.
[0535] 26a shows an example of performing image size adjustment, 26b shows an example of performing size adjustment in surface units (or divided units), and 26c shows an example of performing size adjustment (or multiple size adjustments) in image and surface units.
[0536] The projected image can perform size adjustment using a scale factor or size adjustment using an offset factor according to the image size adjustment type. The image before size adjustment is P_Width×P_Height, the image after size adjustment is P’_Width×P’_Height, and the size of the surface can be F_Width×F_Height. The sizes of the surfaces may be the same or different depending on the surface, and the horizontal and vertical widths of the surface may be the same or different. However, in this example, for the sake of convenience of explanation, it is assumed that the sizes of all surfaces in the image are the same and have a square shape. Also, it is explained under the assumption that the size adjustment values (in this example, WX, HY) are the same. The data processing method in the examples described later will be mainly explained for the case of the offset factor, and the data processing method will be mainly explained for the method of copying and filling a partial area of the image and the method of converting and filling a partial area of the image. The above settings can be similarly applied to FIG. 27.
[0537] In the cases of 26a to 26c, the boundary of the surface (assuming in this example that it has continuity according to 24a in FIG. 24) can have continuity with the boundaries of other surfaces. At this time, it can be classified into the case where they are spatially adjacent in the two-dimensional plane and there is image continuity (the first exemplification) and the case where they are not spatially adjacent in the two-dimensional plane and there is image continuity (the second exemplification).
[0538] For example, assuming the continuity of 24a in FIG. 24, the upper, left, right, and lower regions of S1,1 are spatially adjacent to the lower, right, left, and upper regions of S1,0, S0,1, S2,1, and S1,2, respectively, and the images can also be continuous (in the case of the first illustration).
[0539] Alternatively, the left and right regions of S1,0 are not spatially adjacent to the upper regions of S0,1 and S2,1, but their images can be continuous (in the case of the second illustration). Also, the left region of S0,1 and the right region of S3,1 are not spatially adjacent, but their images can be continuous (in the case of the second illustration). Also, the left and right regions of S1,2 and the lower regions of S0,1 and S2,1 can be continuous with each other (in the case of the second illustration). This is a limited example in this case, and depending on the definition and setting of the projection format, it can have a configuration different from the above. For the sake of convenience of explanation, S0,0 to S3,2 in FIG. 26a are referred to as a to l.
[0540] 26a can be an example of filling using data in a region where continuity exists in the outer boundary direction of the image. Regions that are resized from the A region where no data exists (in this case, a0 to a2, c0, d0 to d2, i0 to i2, k0, l0 to l2) can be filled with any predetermined value or via outer contour pixel padding, and regions that are resized from the B region containing actual data (in this case, b0, e0, h0, j0) can be filled using data in a region (or surface) where image continuity exists. For example, b0 can be filled using the lower data of the surface obtained by applying a 180-degree rotation to surface h, and j0 can be filled using the upper data of the surface obtained by applying a 180-degree rotation to surface h. However, in this case (including the examples described later), only the position of the surface to be referred to is shown, and the data obtained for the resized region can be obtained after an adjustment process (such as rotation) considering the continuity between surfaces, as shown in FIGS. 24 and 25.
[0541] Specifically, b0 can be an example of filling using the lower data of the surface obtained by applying a 180-degree rotation to surface h, and j0 can be an example of filling using the upper data of the surface obtained by applying a 180-degree rotation to surface h. However, in this case (including the examples described later), only the position of the surface to be referred to is shown, and the data obtained for the resized region can be obtained after an adjustment process (such as rotation) considering the continuity between surfaces, as shown in FIGS. 24 and 25.
[0542] 26b may be an example of filling using data of a region where continuity exists in the internal boundary direction of the image. In this example, the resizing operations performed along the surface may be different. Region A can perform a shrinking process, and region B can perform an expanding process. For example, in the case of surface a, resizing (shrinking in this example) by w0 to the right is performed, and in the case of surface b, resizing (expanding in this example) by w0 to the left may be performed. Or, in the case of surface a, resizing (shrinking in this example) by h0 downward is performed, and in the case of surface e, resizing (expanding in this example) by h0 upward may be performed. In this example, looking at the change in the horizontal width of the image from surfaces a, b, c, d, surface a is shrunk by w0, surface b is expanded by w0 and w1, and surface c is shrunk by w1, so the horizontal width of the image before resizing is the same as the horizontal width of the image after resizing. Looking at the change in the vertical height of the image from surfaces a, e, i, surface a is shrunk by h0, surface e is expanded by h0 and h1, and surface i is shrunk by h1, so the vertical height of the image before resizing is the same as the vertical height of the image after resizing.
[0543] Considering that the resized regions (in this example, b0, e0, be, b1, bg, g0, h0, e1, ej, j0, gi, g1, j1, h1) are shrunk from region A where there is no data, they can be simply removed. Considering that they are expanded from region B containing actual data, they can be newly filled with data of the region where continuity exists.
[0544] For example, b0 is the upper side of surface e, e0 is the left side of surface b, be is the left side of surface b, or the upper side of surface e, or the weighted sum of the left side of surface b and the upper side of surface e, b1 is the upper side of surface g, bg is the left side of surface b, or the upper side of surface g, or the weighted sum of the right side of surface b and the upper side of surface g, g0 is the right side of surface b, h0 is the upper side of surface b, e1 is the left side of surface j, ej is the lower side of surface e, or the left side of surface j, or the weighted sum of the lower side of surface e and the left side of surface j, j0 is the lower side of surface e, gj is the lower side of surface g, or the left side of surface j, or the weighted sum of the lower side of surface g and the right side of surface j, g1 is the right side of surface j, j1 is the lower side of surface g, and h1 is the data on the lower side of surface j can be used for filling.
[0545] When filling a part of the image data into the area to be resized in the above example, the data of the corresponding area can be copied and filled, or the data obtained after going through a conversion process based on the characteristics and types of the image can be filled. For example, when a 360-degree image is converted according to the projection format in a two-dimensional space, the coordinate system of each surface (for example, a two-dimensional plane coordinate system) can be defined. For the convenience of explanation, it is assumed that in the three-dimensional space (x, y, z), each surface is converted to (x, y, C) or (x, C, z) or (C, y, z). The above example shows the case of obtaining data of a surface different from the current surface for the area to be resized on some surfaces. That is, when resizing around the current surface and directly copying and filling the data of other surfaces with other coordinate system characteristics, there is a possibility that the continuity will be distorted based on the boundary of the resizing. Therefore, it is also possible to convert the data of other surfaces obtained according to the coordinate system characteristics of the current surface and fill it into the area to be resized. The case of conversion is only an example of the data processing method and is not limited to this.
[0546] When copying and filling the data of a partial area of an image into the area to be resized, the boundary area between the area to be resized (e) and the area being resized (e0) can include distorted continuity (or, rapidly changing continuity). For example, there may be cases where continuous features change based on the boundary. This is similar to an edge that had a straight shape becoming folded based on the boundary.
[0547] When converting and filling the data of a partial area of an image into the area to be resized, the boundary area between the area to be resized and the area being resized can include gradually changing continuity.
[0548] The above example can be an example of the data processing method of the present invention in which, during the resizing process (in this example, expansion), the data of a partial area of an image is subjected to conversion processing based on the characteristics, type, etc. of the image, and the obtained data is filled into the area to be resized.
[0549] 26c can be an example of filling using the data of an area where continuity exists in the direction of the image boundary (inner boundary and outer boundary) by combining the image resizing processes by 26a and 26b. Since the resizing process of this example can be derived from 26a and 26b, detailed description is omitted.
[0550] 26a can be an example of an image resizing process, and 26b can be an example of a resizing process of a division unit within the image. 26c can be an example of a plurality of resizing processes in which image resizing and resizing of division units within the image are performed.
[0551] For example, size adjustment (in this example, C area) can be performed on an image obtained through the projection process (in this example, the first format), and size adjustment (in this example, D area) can be performed on an image obtained through the format conversion process. In this example, size adjustment (in this example, the entire image) is performed on the image projected by the ERP, which can be an example where size adjustment (in this example, surface unit) is performed after obtaining the image projected by the CMP through the format conversion unit. The above example is an example of performing multiple size adjustments, and is not limited thereto, and modifications to other cases are also possible.
[0552] FIG. 27 is an exemplary diagram for explaining size adjustment for an image converted into a CMP projection format and packed according to an embodiment of the present invention. Since FIG. 27 is also explained on the premise of the continuity between surfaces according to FIG. 25, the boundary of the surface can have continuity with the boundaries of other surfaces.
[0553] In this example, the offset factors of W0 to W5 and H0 to H3 (assuming in this example that the offset factor is used as the size adjustment value) can have various values. For example, they can be derived from a predetermined value, the motion search range for prediction between screens, the unit obtained from the picture division unit, et...
Claims
1. A method for decoding a 360-degree image, comprising the steps of: receiving a bitstream in which a 360-degree image is encoded; generating a predicted image by referring to syntax information obtained from a received bitstream; combining the generated prediction image with a residual image obtained by inverse quantizing and inverse transforming the bitstream to obtain a decoded image; and reconstructing the decoded image into a 360 degree image in a projection format; The step of generating a predicted image includes: obtaining a motion vector candidate set including motion vectors of blocks adjacent to a current block to be decoded from motion information included in the syntax information; deriving a predicted motion vector from among a group of motion vector candidates based on selection information extracted from the motion information; and determining a prediction block of a current block to be decoded using a final motion vector derived by adding the predicted motion vector to a differential motion vector extracted from the motion information.
2. The group of motion vector candidates includes: If the block adjacent to the current block is different from the surface to which the current block belongs, The method for decoding a 360-degree image according to claim 1 , wherein the motion vectors are composed only of motion vectors for blocks belonging to a surface that has image continuity with a surface to which the current block belongs, among the adjacent blocks.
3. The adjacent blocks are The method of claim 1 , wherein the block adjacent to the current block is a block adjacent to the current block in at least one of an upper left corner, an upper corner, an upper right corner, a left corner, and a lower left corner.
Citation Information
Patent Citations
Image decoding apparatus, image decoding method, image encoding apparatus and image encoding method
JP2009247019A
Image decoding device, image encoding device, and data structure of encoded data
JP2013118424A
Image coding device, image decoding device, image coding method, image decoding method, image coding program, and image decoding program
JP2013229674A
Apparatus and method for image encoding and decoding
US20060268982A1
Motion compensation apparatus, video coding apparatus, video decoding apparatus, motion compensation method, program, and integrated circuit
US20130044816A1