Image data encoding / decoding method and device
The method enhances the decoding of 360-degree images by using syntax information to generate predicted images and combine them with residual images, improving compression performance and addressing the inefficiencies of conventional methods.
Patent Information
- Application Number
- JP2025038596
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-07-17
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
AI Technical Summary
Conventional image encoding/decoding methods struggle with efficiently processing 360-degree images due to the large amount of data generated, leading to insufficient performance in image processing systems.
A method for decoding 360-degree images that involves receiving a bitstream, generating a predicted image using syntax information, combining it with a residual image obtained through inverse quantization and transformation, and reconstructing the decoded image into a 360-degree image in a projection format, specifically utilizing projection format information such as ERP, CMP, OHP, and ISP.
This approach improves compression performance, particularly for 360-degree images, by enhancing the image setting process during encoding and decoding, thereby addressing the limitations of existing technologies.
Smart Images

Figure 2025090734000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image data encoding and decoding techniques, and more particularly, to a method and apparatus for processing the encoding and decoding of 360-degree images for immersive media services.
Background Art
[0002] With the spread of the Internet and mobile terminals and the development of information and communication technologies, the use of multimedia data has increased rapidly. Recently, demands for high-resolution images and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, have arisen in various fields, and the demand for immersive media services such as virtual reality and augmented reality has increased rapidly. In particular, in the case of 360-degree images for virtual reality and augmented reality, since multi-view images captured by a plurality of cameras are processed, the amount of data generated thereby increases enormously, but the performance of the image processing system for processing this is insufficient.
[0003] As described above, in the conventional image encoding / decoding method and apparatus, improvement in performance for image processing, particularly image encoding / decoding, is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present invention is for solving the above-described problems, and an object thereof is to provide a method for improving an image setting process in the initial stages of encoding and decoding. More specifically, it is to provide an encoding and decoding method and apparatus for improving an image setting process that takes into account the characteristics of 360-degree images.
Means for Solving the Problems
[0005] One aspect of the present invention for achieving the above object provides a method for decoding a 360-degree image.
[0006] Here, the method for decoding a 360-degree image may include receiving a bitstream in which the 360-degree image is encoded, generating a predicted image by referring to syntax information obtained from the received bitstream, obtaining a decoded image by combining the generated predicted image with a residual image obtained by inverse quantization and inverse transformation of the bitstream, and reconstructing the decoded image into a 360-degree image in a projection format.
[0007] Here, the syntax information may include projection format information for the 360-degree image.
[0008] Here, the projection format information may indicate at least one of an ERP (Equi-Rectangular Projection) format in which the 360-degree image is projected onto a two-dimensional plane, a CMP (CubeMap Projection) format in which the 360-degree image is projected onto a cube, an OHP (OctaHedron Projection) format in which the 360-degree image is projected onto an octahedron, and an ISP (IcoSahedral Projection) format in which the 360-degree image is projected onto an icosahedron.
[0009] Here, the reconstructing step may include obtaining arrangement information by regional packing by referring to the syntax information, and rearranging each block of the decoded image based on the arrangement information.
[0010] Here, the step of generating the predicted image may include performing image expansion on a reference picture obtained by restoring the bitstream, and generating a predicted image by referring to the reference picture on which the image expansion has been performed.
[0011] Here, the step of performing the image expansion may include performing image expansion based on a division unit of the reference picture.
[0012] Here, in the step of performing image expansion based on the division unit, an area individually expanded for each division unit can be generated using the boundary pixels of the division unit.
[0013] Here, the expanded area can be generated using the boundary pixels of the division unit that is spatially adjacent to the division unit to be expanded or the boundary pixels of the division unit having image continuity with the division unit to be expanded.
[0014] Here, in the step of performing image expansion based on the division unit, an expanded image for the combined area can be generated using the boundary pixels of the area where two or more spatially adjacent division units among the division units are combined.
[0015] Here, in the step of performing image expansion based on the division unit, an expanded area can be generated between the adjacent division units using all the adjacent pixel information of the spatially adjacent division units among the division units.
[0016] Here, in the step of performing image expansion based on the division unit, the expanded area can be generated using the average value of the adjacent pixels of each of the spatially adjacent division units.
[0017] Here, the step of generating the predicted image can include: obtaining a group of motion vector candidates including the motion vectors of the blocks adjacent to the current block to be decoded from the motion information included in the syntax information; deriving a predicted motion vector from the group of motion vector candidates based on the selection information extracted from the motion information; and determining the predicted block of the current block to be decoded using the final motion vector derived by adding the predicted motion vector and the differential motion vector extracted from the motion information.
[0018] Here, when the motion vector candidate group is such that a block adjacent to the current block is different from the surface to which the current block belongs, the motion vector candidate group can be composed only of motion vectors for blocks belonging to a surface having image continuity with the surface to which the current block belongs among the adjacent blocks.
[0019] Here, the adjacent block can mean a block adjacent to the current block in at least one of the upper left, upper, upper right, left, and lower left directions of the current block.
[0020] Here, the final motion vector can indicate a reference region that belongs to at least one reference picture based on the current block and is set in a region having image continuity between surfaces according to the projection format.
[0021] Here, the reference picture can be extended based on image continuity according to the projection format in the up, down, left, and right directions, and then the reference region can be set.
[0022] Here, the reference picture is extended in units of the surface, and the reference region can be set across the boundary of the surface.
[0023] Here, the motion information can include at least one of a reference picture list to which the reference picture belongs, an index of the reference picture, and a motion vector indicating the reference region.
[0024] Here, the step of generating a predicted block of the current block can include the step of dividing the current block into a plurality of sub-blocks and generating a predicted block for each of the plurality of divided sub-blocks.
Advantages of the Invention
[0025] When using the image encoding / decoding method and apparatus according to the embodiment of the present invention as described above, the compression performance can be improved. In particular, in the case of a 360-degree image, the compression performance can be improved.
Brief Description of the Drawings
[0026]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16a
Figure 16b
Figure 16c
Figure 16d
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Mode for Carrying Out the Invention
[0027] The present invention can be subjected to various modifications and can have various embodiments. Here, specific embodiments are illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiments, and it should be understood to include any modifications, equivalents, or alternatives included in the spirit and technical scope of the present invention.
[0028] Terms such as first, second, A, B, etc. are used to describe various components, but the components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present invention, the first component can be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of the plurality of related described items.
[0029] When it is said that a certain component is "connected to" or "connected with" another component, it means that it is directly connected or connected to the other component, but it should be understood that another component may be interposed therebetween. On the other hand, when it is said that a certain component is "directly connected to" or "directly connected with" another component, it should be understood that no other component is interposed therebetween.
[0030] The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "including" or "having" are intended to specify the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0031] Unless otherwise defined, all terms used herein, including technical or scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Terms defined in commonly used dictionaries shall be interpreted as consistent with the meaning in the context of the relevant art and shall not be interpreted in an idealized or overly formal sense unless clearly defined herein.
[0032] The image encoding device and the decoding device can be user terminals such as a personal computer (PC), a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a PlayStation Portable (PSP), a wireless communication terminal, a smart phone, a TV, a virtual reality device (VR), an augmented reality device (AR), a mixed reality device (MR), a head-mounted display (HMD), smart glasses, etc., or server terminals such as an application server or a service server. It can include various devices equipped with a communication device such as a communication modem for communicating with various devices or a wired / wireless communication network, a memory for storing various programs and data for encoding or decoding an image or for predicting within a screen or between screens for encoding or decoding, a processor for executing a program for arithmetic and control, etc. Further, the image encoded into a bitstream by the image encoding device can be transmitted to the image decoding device in real-time or non-real-time via a wired / wireless communication network such as the Internet, a short-range wireless communication system, a wireless LAN network, a WiBro network, a mobile communication network, etc., or via various communication interfaces such as a cable or a universal serial bus (USB), and can be decoded by the image decoding device and restored and played back as an image.
[0033] Also, the image encoded into a bitstream by the image encoding device can be transmitted from the encoding device to the decoding device via a computer-readable recording medium.
[0034] The above-described image encoding device and image decoding device can be separate devices, respectively, but depending on the implementation, they can be made into one image encoding / decoding device. In that case, some of the components of the image encoding device are technical elements that are substantially the same as some of the components of the image decoding device, and include at least the same structure or can be realized to perform at least the same function.
[0035] Therefore, in the following detailed description of the following technical elements and their operating principles, etc., duplicate descriptions of corresponding technical elements will be omitted.
[0036] Since the image decoding device corresponds to a computer device that applies the image encoding method performed by the image encoding device to decoding, the following description will focus on the image encoding device.
[0037] The computer device can include a memory that stores programs and software modules for implementing the image encoding method and / or the image decoding method, and a processor connected to the memory to execute the programs. The image encoding device may sometimes be called an encoder, and the image decoding device may sometimes be called a decoder.
[0038] Generally, an image can be composed of a series of still images, and these still images can be divided into units of GOP (Group of Pictures). Each still image may be referred to as a picture. At this time, a picture can indicate any one of a progressive signal, a frame and a field in an interlace signal, and when the encoding / decoding is performed in units of frames, the image can be represented by "frame", and when it is performed in units of fields, it can be represented by "field". In the present invention, the description will be made assuming a progressive signal, but it is also applicable to an interlace signal. As upper-level concepts, units such as GOP and sequence can exist. In addition, each picture can be divided into predetermined regions such as slices, tiles, and blocks. Also, one GOP may include units such as I pictures, P pictures, and B pictures. An I picture can mean a picture that is encoded / decoded by itself without using a reference picture, and a P picture and a B picture can mean pictures that are encoded / decoded by performing processes such as motion estimation and motion compensation using a reference picture. Generally, in the case of a P picture, I pictures and P pictures can be used as reference pictures, and in the case of a B picture, I pictures and P pictures can be used as reference pictures, but this definition can also be changed depending on the encoding / decoding settings.
[0039] Here, the pictures referred to in encoding / decoding are called reference pictures, and the referred blocks or pixels are called reference blocks and reference pixels, respectively. Also, the reference data can be not only the pixel values in the spatial domain but also the coefficient values in the frequency domain and various encoding / decoding information generated and determined during the encoding / decoding process. For example, in the prediction unit, it can be intra-prediction related information or motion related information, in the conversion unit / inverse conversion unit, it can be conversion related information, in the quantization unit / inverse quantization unit, it can be quantization related information, in the encoding unit / decoding unit, it can be encoding / decoding related information (context information), and in the in-loop filter unit, it can be filter related information, etc.
[0040] The minimum unit forming an image can be a pixel. The number of bits used to represent one pixel is called the bit depth. Generally, the bit depth is 8 bits, and depending on the encoding settings, a higher bit depth can be supported. At least one bit depth can be supported according to the color space. Also, according to the color format of the image, it can be composed of at least one color space. According to the color format, it can be composed of one or more pictures having a certain size or one or more pictures having different sizes. For example, in the case of YCbCr4:2:0, it can be composed of one luminance component (in this example, Y) and two chrominance components (in this example, Cb / Cr). At this time, the composition ratio of the chrominance components to the luminance component can have a horizontal ratio of 1: vertical ratio of 2. As another example, in the case of 4:4:4, the horizontal and vertical can have the same composition ratio. When composed of one or more color spaces as in the above examples, the picture can be divided into each color space.
[0041] In the present invention, the description is based on a part of a color space (in this example, Y) of a part of a color format (in this example, YCbCr). The same or similar applications (settings dependent on a specific color space) can also be made for another color space (in this example, Cb, Cr) according to the color format. However, it is also possible to make partial differences (settings independent of a specific color space) for each color space. That is, the settings dependent on each color space can mean being proportional to the composition ratio of each component (determined according to, for example, 4:2:0, 4:2:2, 4:4:4, etc.) or having settings dependent thereon, and the settings independent of each color space can mean having settings that are not related to the composition ratio of each component or having settings only for the corresponding color space independently. In the present invention, depending on the encoder / decoder, some configurations can have independent settings or dependent settings.
[0042] The setting information or syntax elements required in the image encoding process can be determined at unit levels such as video, sequence, picture, slice, tile, block, etc. This can be recorded in the bitstream in units such as VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), Slice Header, Tile Header, Block Header, etc. and transmitted to the decoder. The decoder can parse at the same level of unit to restore the setting information transmitted from the encoder and use it in the image decoding process. Also, related information can be transmitted to the bitstream in the form of SEI (Supplement Enhancement Information) or metadata and parsed for use. Each parameter set has a unique ID value, and a lower-level parameter set can have the ID value of the upper-level parameter set it references. For example, a lower-level parameter set can reference the information of the upper-level parameter set with a matching ID value among one or more upper-level parameter sets. Among the various unit examples described above, when any unit contains one or more other units, the corresponding unit may be called the upper unit, and the contained units may be called the lower units.
[0043] In the case of the setting information generated in the said unit, it can include content for an independent setting for each corresponding unit, or content related to a dependent setting that depends on previous, subsequent, or upper units, etc. Here, the dependent setting can be understood to indicate the setting information of the corresponding unit with flag information (for example, if a 1-bit flag is 1, it follows the setting; if it is 0, it does not follow the setting) that follows the settings of previous, subsequent, or upper units. The setting information in the present invention will be mainly described with examples of independent settings, but examples of addition or substitution for content regarding the dependent relationship with the setting information of previous, subsequent units, or upper units of the current unit may also be included.
[0044] FIG. 1 is a block diagram of an image encoding apparatus according to an embodiment of the present invention. FIG. 2 is a block diagram of an image decoding apparatus according to an embodiment of the present invention.
[0045] Referring to FIG. 1, the image encoding apparatus can be configured to include a prediction unit, a subtraction unit, a conversion unit, a quantization unit, an inverse quantization unit, an inverse conversion unit, an addition unit, an in-loop filter unit, a memory and / or an encoding unit. Among the above configurations, some may not necessarily be included, and some or all may be selectively included according to the implementation, and additional configurations not shown may also be included.
[0046] Referring to FIG. 2, the image decoding apparatus can be configured to include a decoding unit, a prediction unit, an inverse quantization unit, an inverse conversion unit, an addition unit, an in-loop filter unit and / or a memory. Among the above configurations, some may not necessarily be included, and some or all may be selectively included depending on the implementation, and additional configurations not shown may also be included.
[0047] The image encoding apparatus and the image decoding apparatus may each be separate apparatuses, but depending on the implementation, they may be formed as a single image encoding / decoding apparatus. In that case, some configurations of the image encoding apparatus are technical elements that are substantially the same as some configurations of the image decoding apparatus and can be realized to include at least the same structure or perform at least the same function. Therefore, in the following detailed description of these technical elements and their operating principles, etc., duplicate descriptions of corresponding technical elements will be omitted. Since the image decoding apparatus corresponds to a computer apparatus that applies the image encoding method performed by the image encoding apparatus to decoding, the following description will focus on the image encoding apparatus. The image encoding apparatus may sometimes be called an encoder, and the image decoding apparatus may sometimes be called a decoder.
[0048] The prediction unit can be realized by using a prediction module, which is a software module, and can generate a prediction block for a block to be encoded using an intra prediction method or an inter prediction method within the screen. The prediction unit predicts the current block to be currently encoded in the image to generate a prediction block. That is, the prediction unit predicts the pixel value of each pixel of the current block to be encoded in the image through intra prediction or inter prediction, and generates a prediction block having the predicted pixel value of each generated pixel. Further, the prediction unit can transmit the information necessary for generating the prediction block to the encoding unit so as to encode the information for the prediction mode, record the information thereby in the bitstream, and transmit this to the decoder. The decoding unit of the decoder can parse the information for this, restore the information for the prediction mode, and then use this for intra prediction or inter prediction.
[0049] The subtraction unit subtracts the prediction block from the current block to generate a residual block. That is, the subtraction unit calculates the difference between the pixel value of each pixel of the current block to be encoded and the predicted pixel value of each pixel of the prediction block generated through the prediction unit, and generates a residual block, which is a residual signal in block form.
[0050] The conversion unit can convert a signal belonging to the spatial domain into a signal belonging to the frequency domain. At this time, the signal obtained through the conversion process is called a transformed coefficient. For example, it is possible to convert a residual block having a residual signal transmitted from the subtraction unit to obtain a conversion block having transformed coefficients, but the input signal is determined according to the encoding settings, and this is not limited to the residual signal.
[0051] The conversion unit can perform conversion on the residual block using conversion techniques such as Hadamard Transform, DST Based-Transform (Discrete Sine Transform), and DCT Based-Transform (Discrete Cosine Transform). However, it is not limited to this, and various conversion techniques obtained by improving and modifying this can be used.
[0052] For example, at least one of the above conversions can be supported, and at least one detailed conversion technique can be supported for each conversion technique. At this time, at least one detailed conversion technique can be a conversion technique configured such that a part of the basis vectors is different for each conversion technique. For example, as conversion techniques, DST-based conversion and DCT-based conversion can be supported. In the case of DST, detailed conversion techniques such as DST-I, DST-II, DST-III, DST-V, DST-VI, DST-VII, and DST-VIII can be supported, and in the case of DCT, detailed conversion techniques such as DCT-I, DCT-II, DCT-III, DCT-V, DCT-VI, DCT-VII, and DCT-VIII can be supported.
[0053] Any one of the above conversions (for example, one conversion technique & one detailed conversion technique) can be set as the basic conversion technique, and additional conversion techniques (for example, multiple conversion techniques || multiple detailed conversion techniques) can be supported. Whether to support additional conversion techniques is determined in units such as sequences, pictures, slices, and tiles, and related information can be generated in these units. When additional conversion techniques are supported, the conversion technique selection information is determined in units such as blocks, and related information can be generated.
[0054] The conversion can be performed in the k / vertical direction. For example, by performing one-dimensional conversion in the horizontal direction and one-dimensional conversion in the vertical direction using the basis vectors in the conversion to perform a total two-dimensional conversion, the pixel values in the spatial domain can be converted into the frequency domain.
[0055] In addition, the horizontal / vertical conversion can be adaptively performed. Specifically, it can be determined whether to perform the conversion adaptively according to at least one encoding setting. For example, when the prediction mode in intra-frame prediction is the horizontal mode, DCT-I can be applied in the horizontal direction and DST-I can be applied in the vertical direction. When the prediction mode in intra-frame prediction is the vertical mode, DST-VI can be applied in the horizontal direction and DCT-VI can be applied in the vertical direction. When the prediction mode is Diagonal down left, DCT-II can be applied in the horizontal direction and DCT-V can be applied in the vertical direction. When the prediction mode is Diagonal down right, DST-I can be applied in the horizontal direction and DST-VI can be applied in the vertical direction.
[0056] According to the encoding cost for each candidate of the size and shape of the conversion block, the size and shape of each conversion block are determined, and information such as the image data of each determined conversion block and the size and shape of each determined conversion block can be encoded.
[0057] Among the conversion shapes, a square conversion can be set as the basic conversion shape, and additional conversion shapes (for example, rectangular shapes) can be supported. Whether to support additional conversion shapes is determined in units such as sequence, picture, slice, and tile. Related information can be generated in these units, and the conversion shape selection information is determined in units such as blocks, and related information can be generated.
[0058] In addition, the support for the conversion block shape can be determined according to the encoding information. At this time, the encoding information can include slice type, encoding mode, block size and shape, block partitioning method, etc. That is, one conversion shape can be supported according to at least one piece of encoding information, and multiple conversion shapes can be supported according to at least one piece of encoding information. The former case is an implicit situation, and the latter case can be an explicit situation. In the explicit case, adaptive selection information indicating the optimal candidate group among multiple candidate groups can be generated and recorded in the bitstream. Including this example, in the present invention, when explicitly generating encoding information, the corresponding information is recorded in the bitstream in various units, and the decoder parses the relevant information in various units and restores it to the decoded information. Also, when implicitly processing the encoding / decoding information, it can be understood that the encoder and the decoder process it according to the same process, rules, etc.
[0059] As an example, the support for rectangular conversion can be determined according to the slice type. The conversion shape supported in the case of an I slice is square conversion, and the conversion shape supported in the case of a P / B slice can be square or rectangular conversion.
[0060] As an example, the support for rectangular conversion can be determined according to the encoding mode. The conversion shape supported in the Intra case is square conversion, and the conversion shape supported in the Inter case can be square or rectangular conversion.
[0061] As an example, the support for rectangular conversion can be determined according to the block size and shape. The conversion shape supported for blocks of a certain size or more is square conversion, and the conversion shape supported for blocks smaller than a certain size can be square or rectangular conversion.
[0062] As an example, conversion assistance for a rectangle can be determined according to the block division method. When the block to be converted is a block obtained by a quad tree division method, the supported conversion shape is a square conversion. When the block is a block obtained by a binary tree division method, the supported conversion shape can be a square or a rectangle conversion.
[0063] The above example is an example of conversion shape assistance according to one piece of coding information, and a plurality of pieces of information can be combined and involved in additional conversion shape assistance settings. The above example is only an example of additional conversion shape assistance according to various coding settings, and is not limited to the above, and various deformation examples are possible.
[0064] Depending on the coding setting or the characteristics of the image, the conversion process can be omitted. For example, according to the coding setting (assuming a lossless compression environment in this example), the conversion process (including the reverse process) can be omitted. As another example, when the compression performance by conversion is not exhibited according to the characteristics of the image, the conversion process can be omitted. At this time, the conversion to be omitted can be in the whole unit or in either the horizontal unit or the vertical unit. It can be determined whether to support such omission according to the block size and shape, etc.
[0065] For example, in a setting where the omission of horizontal and vertical conversions is grouped, when the conversion omission flag is 1, the conversions in the horizontal and vertical directions are not performed. When the conversion omission flag is 0, the conversions in the horizontal and vertical directions can be performed. In a setting where the omission of horizontal and vertical conversions operates independently, when the first conversion omission flag is 1, the conversion in the horizontal direction is not performed. When the first conversion omission flag is 0, the conversion in the horizontal direction is performed. When the second conversion omission flag is 1, the conversion in the vertical direction is not performed. When the second conversion omission flag is 0, the conversion in the vertical direction is performed.
[0066] When the block size falls within range A, conversion omission can be supported; when the block size falls within range B, conversion omission cannot be supported. For example, when the horizontal width of the block is greater than M or the vertical height of the block is greater than N, the conversion omission flag cannot be supported; when the horizontal width of the block is less than m or the vertical height of the block is less than n, the conversion omission flag can be supported. M(m) and N(n) may be the same or different. The conversion-related settings can be determined in units such as sequence, picture, slice, etc.
[0067] When additional conversion techniques are supported, the settings of the conversion techniques can be determined according to at least one piece of coding information. At this time, the coding information can include slice type, coding mode, block size and shape, prediction mode, etc.
[0068] As an example, the support for conversion techniques can be determined according to the coding mode. The conversion techniques supported in the Intra case are DCT-I, DCT-III, DCT-VI, DST-II, DST-III, and the conversion techniques supported in the Inter case can be DCT-II, DCT-III, DST-III.
[0069] As an example, the support for conversion techniques can be determined according to the slice type. The conversion techniques supported in the I slice case are DCT-I, DCT-II, DCT-III, the conversion techniques supported in the P slice case are DCT-V, DST-V, DST-VI, and the conversion techniques supported in the B slice case can be DCT-I, DCT-II, DST-III.
[0070] As an example, the support for conversion techniques can be determined according to the prediction mode. The conversion techniques supported in the prediction mode A are DCT-I, DCT-II, the conversion techniques supported in the prediction mode B are DCT-I, DST-I, and the conversion techniques supported in the prediction mode C can be DCT-I. At this time, the prediction modes A and B are directional modes, and the prediction mode C can be a non-directional mode.
[0071] As an example, the support for the conversion technique can be determined according to the size and shape of the block. The conversion technique supported for blocks of a certain size or larger is DCT-II, the conversion techniques supported for blocks smaller than a certain size are DCT-II and DST-V, and the conversion techniques supported for blocks of a certain size or larger and smaller than a certain size can be DCT-I, DCT-II, and DST-I. Also, the conversion techniques supported for a square shape are DCT-I and DCT-II, and the conversion techniques supported for a rectangular shape can be DCT-I and DST-I.
[0072] The above example is an example of the support for the conversion technique according to one piece of encoding information, and a plurality of pieces of information can be combined and involved in the support setting of additional conversion techniques. It is not limited to only the case of the above example, and deformation to other examples is also possible. Also, the conversion unit can transmit information necessary for generating a conversion block to the encoding unit to encode it, record the resulting information in a bitstream, and transmit it to the decoder. The decoding unit of the decoder can parse the information for this and use it in the inverse conversion process.
[0073] The quantization unit can quantize the input signal. At this time, the signal obtained through the quantization process is called a quantized coefficient. For example, it is possible to quantize a residual block having residual conversion coefficients transmitted from the conversion unit to obtain a quantized block having quantized coefficients, but the input signal is determined according to the encoding setting, and this is not limited to residual conversion coefficients.
[0074] The quantization unit can quantize the transformed residual block using quantization techniques such as Dead Zone Uniform Threshold Quantization and Quantization Weighted Matrix, and is not limited thereto. Various quantization techniques obtained by improving and modifying this can be used. Whether to support additional quantization techniques is determined in units such as sequences, pictures, slices, and tiles, and related information can be generated in these units. When additional quantization techniques are supported, the quantization technique selection information is determined in units such as blocks, and related information can be generated.
[0075] When additional quantization techniques are supported, the setting of the quantization technique can be determined according to at least one piece of coding information. At this time, the coding information can include slice type, coding mode, block size and shape, prediction mode, etc.
[0076] For example, the quantization unit can be set so that the quantization weight matrix according to the coding mode and the weight matrix applied according to inter-picture prediction / intra-picture prediction are different from each other. Also, the weight matrix applied according to the intra-picture prediction mode can be set to be different. At this time, assuming that the quantization weight matrix is of size M×N and the block size is the same as the quantization block size, some of the quantization components can be different quantization matrices.
[0077] The quantization process can be omitted according to the coding setting or the characteristics of the image. For example, the quantization process (including the inverse process) can be omitted according to the coding setting (assuming a lossless compression environment in this example). As another example, the quantization process can be omitted when the compression performance by quantization is not exhibited according to the characteristics of the image. At this time, the area to be omitted can be the entire area or a partial area. Whether to support such omission can be determined according to the block size and shape, etc.
[0078] Information about the quantization parameter (QP) can be generated in units such as sequences, pictures, slices, tiles, and blocks. For example, the basic QP can be set in the higher-level unit where the QP information is first generated <1>, and the QP can be set to the same or a different value from the QP set in the higher-level unit as it goes to the lower-level unit <2>. Through such a process, in the quantization process performed in some units, the QP can be finally determined <3>. At this time, units such as sequences and pictures correspond to <1>, units such as slices, tiles, and blocks correspond to <2>, and units such as blocks correspond to <3>.
[0079] Information about the QP can be generated based on the QP in each unit. Or, a preset QP can be set as a predicted value, and the difference value information from the QP in each unit can be generated. Or, a QP obtained based on at least one of the QP set in the higher-level unit, the QP set in the same unit previously, or the QP set in the adjacent unit can be set as a predicted value, and the difference value information from the QP in the current unit can be generated. Or, a QP obtained based on the QP set in the higher-level unit and at least one piece of coding information can be set as a predicted value, and the difference value information from the QP in the current unit can be generated. At this time, the previous same unit is a unit that can be defined according to the coding order of each unit, the adjacent unit is a spatially adjacent unit, and the coding information can be the slice type, coding mode, prediction mode, position information, etc. of the corresponding unit.
[0080] As an example, the QP of the current unit can set the QP of the higher-level unit as a predicted value and generate the difference value information. The difference value information between the QP set in the slice and the QP set in the picture can be generated, or the difference value information between the QP set in the tile and the QP set in the picture can be generated. Also, the difference value information between the QP set in the block and the QP set in the slice or tile can be generated. Also, the difference value information between the QP set in the sub-block and the QP set in the block can be generated.
[0081] As an example, for the QP of the current unit, the QP obtained based on at least one adjacent unit's QP or the QP obtained based on at least one previous unit's QP can be set as a predicted value to generate differential value information. Differential value information can be generated with respect to the QP obtained based on the QPs of adjacent blocks such as the left, upper left, lower left, upper, and upper right of the current block. Alternatively, differential value information can be generated with respect to the QP of the encoded picture before the current picture.
[0082] As an example, for the QP of the current unit, the QP obtained based on the QP of the upper unit and at least one piece of encoding information can be set as a predicted value to generate differential value information. Differential value information can be generated with respect to the QP of the slice corrected according to the QP of the current block and the slice type (I / P / B). Alternatively, differential value information can be generated with respect to the QP of the tile corrected according to the QP of the current block and the encoding mode (Intra / Inter). Alternatively, differential value information can be generated with respect to the QP of the picture corrected according to the QP of the current block and the prediction mode (directional / non-directional). Alternatively, differential value information can be generated with respect to the QP of the picture corrected according to the QP of the current block and the position information (x / y). At this time, the meaning of the correction can mean that it is added or subtracted in an offset form to the QP of the upper unit used for prediction. At this time, at least one offset information can be supported according to the encoding setting, and it can be implicitly processed according to a predetermined process or relevant information can be explicitly generated. It is not limited only to the case of the above example, and variations to other examples are also possible.
[0083] The above example can be a possible example when a signal indicating QP variation is provided or activated. For example, when a signal indicating QP variation is not provided or deactivated, differential value information is not generated, and the predicted QP can be determined as the QP of each unit. As another example, when a signal indicating QP variation is provided or activated, differential value information is generated, and when the value is 0, the predicted QP can be determined as the QP of each unit.
[0084] The quantization unit can transmit information necessary for generating quantization blocks to the encoding unit to cause the encoding unit to perform encoding, record the information thereby obtained in a bit stream, and transmit the bit stream to a decoder. The decoding unit of the decoder can parse the information for this purpose and use the information in an inverse quantization process.
[0085] In the above example, the description was made under the assumption that the residual block is transformed and quantized via the transformation unit and the quantization unit. However, it is not necessary to perform the quantization process by transforming the residual signal to generate a residual block having transformation coefficients. Instead, only the quantization process can be performed without transforming the residual signal of the residual block into transformation coefficients, and moreover, it is not necessary to perform both the transformation and the quantization processes. This can be determined according to the settings of the encoder.
[0086] The encoding unit can scan at least one of the quantization coefficients, transformation coefficients, or residual signals of the generated residual block according to at least one scanning order (for example, zigzag scan, vertical scan, horizontal scan, etc.) to generate a quantization coefficient sequence, a transformation coefficient sequence, or a signal sequence, and can perform encoding using at least one entropy coding technique. At this time, the information about the scanning order can be determined according to the encoding settings (for example, encoding mode, prediction mode, etc.), and can be determined implicitly or explicitly generate related information. For example, according to the in-picture prediction mode, any one of a plurality of scanning orders can be selected.
[0087] Also, encoded data including the encoded information transmitted from each component can be generated and output as a bitstream, which can be realized by a multiplexer (MUX). At this time, as encoding techniques, methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC) can be used for encoding, and it is not limited to this, and various encoding techniques obtained by improving and modifying this can be used.
[0088] When performing entropy coding (assuming CABAC in this example) on syntax elements such as the residual block data and the information generated in the encoding / decoding process, the entropy coding device can include a binarizer, a context modeler, and a binary arithmetic coder. At this time, the binary arithmetic coder can include a regular coding engine and a bypass coding engine.
[0089] Since the syntax elements input to the entropy coding device may not be binary, when the syntax elements are not binary, the binarizer can binarize the syntax elements and output a Bin String consisting of 0 or 1. At this time, Bin indicates a bit consisting of 0 or 1 and can be encoded through the binary arithmetic coder. At this time, either the regular coding engine or the bypass coding engine can be selected based on the occurrence probabilities of 0 and 1. This can be determined according to the encoding / decoding settings. When the data has the same frequency of 0 and 1 for the syntax elements, the bypass coding engine can be used, and otherwise, the regular coding engine can be used.
[0090] When performing binarization on the syntax element, various methods can be used. For example, Fixed Length Binarization, Unary Binarization, Truncated Rice Binarization, K-th Exp-Golomb Binarization, etc. can be used. Also, depending on the range of values that the syntax element has, signed binarization or unsigned binarization can be performed. The binarization process for the syntax element generated in the present invention can be performed including not only the binarization mentioned in the above examples but also other additional binarization methods.
[0091] The inverse quantization unit and the inverse transformation unit can be realized by performing the processes in the transformation unit and the quantization unit in reverse. For example, the inverse quantization unit can inverse-quantize the quantization transformation coefficient generated by the quantization unit, and the inverse transformation unit can inverse-transform the inverse-quantized transformation coefficient to generate a restored residual block.
[0092] The addition unit adds the prediction block and the restored residual block to restore the current block. The restored block is stored in the memory and can be used as reference data (such as the prediction unit and the filter unit).
[0093] The in-loop filter section can include at least one post-processing filter process such as a deblocking filter, sample adaptive offset (SAO), adaptive loop filter (ALF), etc. The deblocking filter can remove block distortion generated at the boundary between blocks from the restored image. The ALF can perform filtering based on a value obtained by comparing the restored image and the input image. Specifically, filtering can be performed based on a value obtained by comparing the restored image after a block has been filtered through the deblocking filter with the input image. Or, filtering can be performed based on a value obtained by comparing the restored image after a block has been filtered through the SAO with the input image. The SAO can restore an offset difference based on a value obtained by comparing the restored image with the input image and can be applied in forms such as band offset (BO) and edge offset (EO). Specifically, the SAO can add an offset with respect to the original image in at least one pixel unit to the restored image to which the deblocking filter has been applied and can be applied in forms such as BO and EO. Specifically, an offset with respect to the original image can be added to the restored image after a block has been filtered through the ALF in pixel units and can be applied in forms such as BO and EO.
[0094] As filtering-related information, setting information regarding whether to support each post-processing filter can be generated in units such as sequence, picture, slice, and tile. Also, setting information regarding whether to execute each post-processing filter can be generated in units such as picture, slice, tile, and block. The range to which the execution of the filter is applied can be divided into the inside of the image and the boundary of the image, and setting information considering this can be generated. Also, information related to the filtering operation can be generated in units such as picture, slice, tile, and block. The information can be subjected to implicit or explicit processing, and for the filtering, an independent filtering process or a dependent filtering process can be applied according to the color component. This can be determined according to the encoding setting. The in-loop filter unit can transmit the filtering-related information to the encoding unit to encode it, record the information thereby in the bitstream, and transmit it to the decoder. The decoding unit of the decoder can parse the information for this and apply it to the in-loop filter unit.
[0095] The memory can store the restored block or picture. The restored block or picture stored in the memory can be provided to a prediction unit that performs intra-prediction or inter-prediction. Specifically, a storage space in the form of a queue of the bitstream compressed by the encoder can be processed as a coded picture buffer (CPB), and a space for storing the decoded image in picture units can be processed as a decoded picture buffer (DPB). In the case of the CPB, the decoding units are stored in the decoding order, the decoding operation is emulated in the encoder, and the bitstream compressed in the emulation process can be stored. The bitstream output from the CPB is restored through the decoding process, the restored image is stored in the DPB, and the picture stored in the DPB can be referred to in subsequent image encoding and decoding processes.
[0096] The decoding unit can be realized by performing the process in the encoding unit in reverse. For example, it can receive a quantization coefficient sequence, a transform coefficient sequence, or a signal sequence from a bit stream, decode the same, and parse the decoded data including decoded information to transmit it to each component unit.
[0097] Hereinafter, an image setting process applied to an image encoding / decoding apparatus according to an embodiment of the present invention will be described. This may be an example (image initial setting) applied at a stage before encoding / decoding, but some processes may also be examples applicable at other stages (for example, a stage after encoding / decoding or an internal stage of encoding / decoding). The image setting process can be performed in consideration of network and user environments such as characteristics of multimedia content, bandwidth, performance of a user terminal, and accessibility. For example, depending on the encoding / decoding settings, image segmentation, image size adjustment, image reconstruction, etc. can be performed. The image setting process described below will be described centering on a rectangular image, but is not limited thereto and is also applicable to a polygonal image. Regardless of the shape of the image, the same image setting can be applied or different image settings can be applied. This can be determined according to the encoding / decoding settings. For example, after confirming information on the shape of the image (for example, a rectangular or non-rectangular shape), information on the image setting based thereon can be configured.
[0098] In the examples described below, an explanation will be given assuming settings that are dependent on the color space. However, it is also possible to make settings that are independent of the color space. Further, in the case of independent settings in the examples described below, examples including encoding / decoding settings independently in each color space can be included. Even if an explanation is given for one color space, it is assumed that examples applicable to other color spaces (for example, when generating M for the luminance component, generating N for the color difference component) can be included and can be derived from this. Also, in the case of dependent settings, examples including settings proportional to the composition ratio of the color format (for example, 4:4:4, 4:2:2, 4:2:0, etc.) (for example, in the case of 4:2:0, when generating M for the luminance component, generating M / 2 for the color difference component) can be included. It is assumed that examples applicable to each color space can be included without special explanation and can be derived from this. This is not limited to the above examples and can be an explanation commonly applicable to the present invention.
[0099] Some of the configurations in the examples described below may be content applicable to various encoding techniques such as encoding in the spatial domain, encoding in the frequency domain, block-based encoding, and object-based encoding.
[0100] It may be common to perform encoding / decoding as the input image is, but there may also be cases where the image is divided for encoding / decoding. For example, division can be performed for error tolerance and the like for the purpose of preventing damage due to packet loss or the like during transmission. Or, division can be performed for the purpose of classifying regions having different properties within the same image according to the characteristics, types, etc. of the image.
[0101] In the present invention, the image division process can include the division process and the reverse process thereof. In the examples described below, the division process will be mainly explained, but the content regarding the division reverse process can be derived inversely from the division process.
[0102] FIG. 3 is an exemplary diagram showing hierarchical separation of image information for compressing an image.
[0103] 3a is an exemplary diagram showing a sequence of images composed of a number of GOPs. One GOP can be composed of an I picture, a P picture, and a B picture as shown in 3b. One picture can be composed of slices, tiles, etc. as shown in 3c. Slices, tiles, etc. are composed of a number of basic encoding units as shown in 3d, and a basic encoding unit can be composed of at least one sub-encoding unit as shown in Fig. 3e. The image setting process in the present invention will be described based on examples applied to units such as pictures, slices, and tiles as shown in 3b and 3c.
[0104] Figure 4 is a conceptual diagram showing various examples of image segmentation according to an embodiment of the present invention.
[0105] 4a is a conceptual diagram showing an image (e.g., a picture) divided at regular intervals in the horizontal and vertical directions. The divided regions can be called blocks, and each block is a basic encoding unit (or maximum encoding unit) obtained through a picture division unit and can also be a basic unit applied in the division units described later.
[0106] 4b is a conceptual diagram showing an image divided in at least one of the horizontal and vertical directions. The divided regions T0 to T3 can be called tiles, and each region can perform independent or dependent encoding / decoding with respect to other regions.
[0107] 4c is a conceptual diagram showing an image divided into groups of consecutive blocks. The divided regions S0 and S1 can be called slices, and each region can be a region that performs independent or dependent encoding / decoding with respect to other regions. The group of consecutive blocks can be defined according to a scan order, and generally follows a raster scan order, but is not limited thereto and can be determined according to the settings of encoding / decoding.
[0108] 4d is a conceptual diagram that divides an image into groups of blocks with arbitrary settings defined by the user. The divided regions A0 to A2 can be called arbitrary partition regions, and each region can be a region that performs independent or dependent encoding / decoding from other regions.
[0109] Independent encoding / decoding can mean that when encoding / decoding some units (or regions), data from other units cannot be referenced. Specifically, the information used or generated in the texture encoding and entropy encoding of some units is encoded independently without mutual reference, and in the decoder, for the texture decoding and entropy decoding of some units, the parsing information and restoration information of other units do not have to reference each other. At this time, whether data from other units (or regions) can be referenced can be restricted in a spatial region (for example, between regions within an image), but depending on the encoding / decoding settings, restricted settings can also be placed in a temporal region (for example, between consecutive images or frames). For example, if a partial unit of the current image and a partial unit of another image have continuity or the same encoding environment, they can be referenced; otherwise, the reference can be restricted.
[0110] Also, dependent encoding / decoding can mean that when encoding / decoding some units, data from other units can be referenced. Specifically, the information used or generated in the texture encoding and entropy encoding of some units is referenced mutually and encoded dependently, and in the decoder, for the texture decoding and entropy decoding of some units, the parsing information and restoration information of other units can reference each other. That is, it can be the same or a similar setting as general encoding / decoding. In this case, depending on the characteristics, types, etc. of the image (for example, a 360-degree image), the region (in this example, the surface generated according to the projection format) <face>It may be the case when it is divided for the purpose of identifying (such as).
[0111] Independent encoding / decoding settings (e.g., independent slice segments) can be placed on some units (such as slices, tiles) in the above example, and dependent encoding / decoding settings (e.g., dependent slice segments) can be placed on some units. In the present invention, the description will be centered around independent encoding / decoding settings.
[0112] The basic encoding unit obtained through the picture division unit as in 4a is divided into basic encoding blocks according to the color space, and its size and shape can be determined according to the characteristics and resolution of the image, etc. The size or shape of the supported block is an N×N square (2 n ) expressed by an exponent of 2 (2 n ×2 n . Such as 256×256, 128×128, 64×64, 32×32, 16×16, 8×8. n is an integer between 3 and 8), or it can be a rectangle of M×N (2 m ×2 n ). For example, according to the resolution, the input image can be divided into sizes such as 128×128 for 8k UHD-class images, 64×64 for 1080p HD-class images, 16×16 for WVGA-class images, etc. According to the type of image, the input image can be divided into a size of 256×256 for 360-degree images. The basic encoding unit can be divided into lower-level encoding units for encoding / decoding, and the information for the basic encoding unit can be recorded and transmitted in the bitstream in units such as sequences, pictures, slices, tiles, etc. This can be parsed by the decoder to restore the relevant information.
[0113] An image encoding method and a decoding method according to an embodiment of the present invention may include the following image segmentation stages. At this time, the image segmentation process may include an image segmentation instruction stage, an image segmentation type identification stage, and an image segmentation execution stage. Further, an image encoding device and a decoding device may be configured to include an image segmentation instruction unit, an image segmentation type identification unit, and an image segmentation execution unit that realize the image segmentation instruction stage, the image segmentation type identification stage, and the image segmentation execution stage. In the case of encoding, associated syntax elements can be generated, and in the case of decoding, associated syntax elements can be parsed.
[0114] In each block segmentation process of 4a, the image segmentation instruction unit may be omitted, and the image segmentation type identification unit is a process of checking information regarding the size and shape of the block, and based on the identified segmentation type information, the image segmentation unit can perform segmentation in basic encoding units.
[0115] In the case of a block, it can always be a unit for which segmentation is performed, but for other segmentation units (such as tiles, slices, etc.), it can be determined whether to perform segmentation according to the encoding / decoding settings. The picture segmentation unit can be basically set to perform segmentation of other units after performing block-based segmentation. At this time, block segmentation can be performed based on the size of the picture.
[0116] Also, it is also possible to perform block-based segmentation after segmentation in other units (such as tiles, slices, etc.). That is, block segmentation can be performed based on the size of the segmentation unit. This can be determined through explicit or implicit processing according to the encoding / decoding settings. In the example described later, the former case is assumed and the description is centered on units other than blocks.
[0117] In the image segmentation instruction stage, it is possible to determine whether to perform image segmentation. For example, when a signal for instructing image segmentation (for example, tiles_enabled_flag) is confirmed, segmentation can be performed, and when a signal for instructing image segmentation is not confirmed, segmentation may not be performed or other encoding / decoding information can be confirmed to perform segmentation.
[0118] Specifically, a signal for instructing image segmentation (e.g., tiles_enabled_flag) is checked. When the corresponding signal is activated (e.g., tiles_enabled_flag = 1), division can be performed in multiple units. When the corresponding signal is deactivated (e.g., tiles_enabled_flag = 0), division cannot be performed. Or, when the signal for instructing image segmentation is not checked, it can mean not performing division or performing division in at least one unit. Whether to perform division in multiple units can be checked via another signal (e.g., first_slice_segment_in_pic_flag).
[0119] In summary, when a signal for instructing image segmentation is provided, the corresponding signal is a signal for indicating whether to perform division in multiple units, and it is possible to check whether to divide the corresponding image according to the signal. For example, when tiles_enabled_flag is a signal indicating whether to perform image segmentation, when tiles_enabled_flag is 1, it can mean that the image is divided into multiple tiles, and when it is 0, it can mean that the image is not divided.
[0120] In summary, when a signal for instructing image segmentation is not provided, division is not performed, or whether to divide the corresponding image can be checked by another signal. For example, first_slice_segment_in_pic_flag is not a signal indicating whether to perform image segmentation, but a signal indicating whether it is the first slice segment in the image. However, thereby, it is possible to check whether to divide into two or more units (e.g., when the flag is 0, it means that the image is divided into multiple slices).
[0121] It is not limited to only the above examples, and variations to other examples are also possible. For example, a signal instructing image segmentation may not be provided for tiles, or a signal instructing image segmentation may be provided for slices. Or, depending on the type, characteristics, etc. of the image, a signal instructing image segmentation may be provided.
[0122] In the image segmentation type identification stage, the image segmentation type can be identified. The image segmentation type can be defined by the method of performing the segmentation, segmentation information, etc.
[0123] In 4b, a tile can be defined as a unit obtained by dividing in the horizontal and vertical directions. Specifically, it can be defined as a group of adjacent blocks within a quadrilateral space partitioned by at least one horizontal or vertical division line crossing the image.
[0124] The segmentation information for tiles can include boundary position information between rows and columns, the number of tiles in rows and columns, tile size information, etc. The number information of tiles can include the number of rows of tiles (e.g., num_tile_columns) and the number of columns of tiles (e.g., num_tile_rows), and thus can be divided into (the number of rows × the number of columns) tiles. The tile size information can be obtained based on the tile number information, but the horizontal width or vertical height of the tiles can be equal or unequal. This can be implicitly determined under a preset rule, or relevant information (e.g., uniform_spacing_flag) can be explicitly generated. Also, the tile size information can include the size information of each row and column of tiles (e.g., column_width_tile[i], row_height_tile[i]), or can include the vertical height and horizontal width information of each tile. Also, the size information may be further generated information depending on whether the tile size is equal or not (e.g., when uniform_spacing_flag is 0, meaning non-uniform division).
[0125] In 4c, a slice can be defined as a group unit of consecutive blocks. Specifically, it can be defined as a group of consecutive blocks based on a predetermined scan order (in this example, raster scan).
[0126] The division information for a slice can include the number information of the slice, the position information of the slice (e.g., slice_segment_address), etc. At this time, the position information of the slice can be the position information of a predetermined (e.g., the first order in the scan order within the slice) block. At this time, the position information can be the scan order information of the block.
[0127] In 4d, various division settings are possible for any division area.
[0128] The division unit in 4d can be defined as a group of spatially adjacent blocks, and the division information for this can include the size, shape, position information, etc. of the division unit. This is some examples for any division area, and various division shapes are possible as shown in FIG. 5.
[0129] FIG. 5 is another exemplary diagram of the image division method according to an embodiment of the present invention.
[0130] In the cases of 5a and 5b, the image can be divided into a plurality of regions with at least one block interval in the horizontal or vertical direction, and the division can be performed based on the position information of the blocks. 5a shows examples A0 and A1 where the division is performed based on the vertical column information of each block in the horizontal direction, and 5b shows examples B0 to B3 where the division is performed based on the horizontal row and vertical column information of each block in the vertical and horizontal directions. The division information for this can include the number of division units, block interval information, division direction, etc., and when some of this is implicitly included according to a predetermined rule, some division information may not be generated.
[0131] In the cases of 5c and 5d, based on the scan order, the image can be divided into groups of consecutive blocks. Additional scan orders other than the existing raster scan order of slices can be applied to image division. 5c shows examples C0 and C1 where scanning (Box-Out) is performed clockwise or counterclockwise around the start block, and 5d shows examples D0 and D1 where vertical scanning (Vertical) is performed around the start block. The division information for this can include the number information of the division units, the position information of the division units (for example, the first order in the scan order within the division unit), the information regarding the scan order, etc. When some division information is implicitly included according to a predetermined rule, some division information may not be generated.
[0132] In the case of 5e, the image can be divided by horizontal and vertical dividing lines. The existing tiles can be divided by horizontal or vertical dividing lines, thereby having a divided shape of a rectangular space, but there may be cases where the division across the image by the dividing lines is not possible. For example, an example of dividing across the image along a partial dividing line of the image (for example, the dividing line formed by the right boundary of E1, E3, E4 and the left boundary of E5) is possible, while an example of dividing across the image along a partial dividing line of the image (for example, the dividing line formed by the lower boundary of E2 and E3 and the upper boundary of E4) is not possible. Also, division can be performed based on block units (for example, after block division is first performed and then division), or division can be performed by the horizontal or vertical dividing lines, etc. (for example, division by the dividing lines regardless of block division), whereby each division unit may not be composed of an integer multiple of blocks. Therefore, division information different from the existing tiles can be generated, and the division information for this can include the number information of the division units, the position information of the division units, the size information of the division units, etc. For example, the position information of the division units can generate position information (measured in pixel units or block units) based on a predetermined position (for example, the upper left corner of the image), and the size information of the division units can generate the horizontal and vertical size information of each division unit (measured in pixel units or block units).
[0133] As in the above example, the splitting with arbitrary settings defined by the user can be performed by applying a new splitting method or by applying a change to a part of the existing splitting. That is, it may be supported by replacing the existing splitting method or by the additional splitting shape, or it may be supported in a form where some settings are applied with changes to the existing splitting method (such as slicing, tiling, etc.) (for example, following other scanning orders, other splitting methods with rectangular shapes and generation of other splitting information, dependent encoding / decoding characteristics, etc.). Also, settings for constituting additional splitting units (for example, settings other than splitting according to the scanning order or splitting according to a certain interval difference) can be supported, and it is also possible to support additional splitting unit shapes (for example, polygonal shapes such as triangles other than splitting into rectangular spaces). Also, it is possible to support an image splitting method based on the type, characteristics, etc. of the image. For example, some splitting methods (such as the surface of a 360-degree image) can be supported according to the type, characteristics, etc. of the image, and splitting information can be generated based on this.
[0134] At the image splitting execution stage, the image can be split based on the identified splitting type information. That is, splitting can be performed into a plurality of splitting units based on the identified splitting type, and encoding / decoding can be performed based on the obtained splitting units.
[0135] At this time, it is possible to determine whether to have encoding / decoding settings for the splitting units according to the splitting type. That is, the setting information required for the encoding / decoding process of each splitting unit can receive an assignment at a higher-level unit (for example, a picture), or can have independent encoding / decoding settings for the splitting units.
[0136] Generally, in the case of slices, it can have independent encoding / decoding settings (e.g., slice headers) for the division units. In the case of tiles, it cannot have independent encoding / decoding settings for the division units and can have settings that are dependent on the encoding / decoding settings of the picture (e.g., PPS). At this time, the information generated in relation to the tile can be division information. This can be included in the encoding / decoding settings of the picture. In the present invention, it is not limited only to the cases described above, and other variations are possible.
[0137] The encoding / decoding setting information for tiles can be generated in units such as video, sequence, and picture. At least one encoding / decoding setting information can be generated at a higher-level unit, and any one of them can be referred to. Or, independent encoding / decoding setting information (e.g., tile headers) can be generated in tile units. This is different from conforming to one encoding / decoding setting determined at a higher-level unit in that encoding / decoding is performed with at least one encoding / decoding setting in tile units. That is, it is possible to conform to one encoding / decoding setting for all tiles, or encoding / decoding can be performed according to a different encoding / decoding setting from other tiles for at least one tile.
[0138] Although the above examples have been mainly described with respect to various encoding / decoding settings for tiles, it is not limited thereto, and similar or identical settings can be made for other division types.
[0139] As an example, for some division types, division information can be generated at a higher-level unit, and encoding / decoding can be performed according to one encoding / decoding setting of the higher-level unit.
[0140] As an example, for some division types, division information can be generated at a higher-level unit, and independent encoding / decoding settings for each division unit can be generated at the higher-level unit, thereby performing encoding / decoding.
[0141] As an example, for some split types, split information can be generated at a higher unit, a plurality of encoding / decoding setting information can be supported at the higher unit, and encoding / decoding can be performed according to the encoding / decoding settings referred to in each split unit.
[0142] As an example, for some split types, split information can be generated at a higher unit, independent encoding / decoding settings can be generated in the corresponding split unit, and thereby encoding / decoding can be performed.
[0143] As an example, for some split types, independent encoding / decoding settings including split information can be generated in the corresponding split unit, and thereby encoding / decoding can be performed.
[0144] The encoding / decoding setting information can include information necessary for encoding / decoding of tiles such as tile type, information regarding the reference picture list, quantization parameter information, inter-picture prediction setting information, in-loop filtering setting information, in-loop filtering control information, scan order, whether to perform encoding / decoding, etc. The encoding / decoding setting information can explicitly generate relevant information, or it is also possible that the settings for encoding / decoding are implicitly determined according to the format, characteristics, etc. of the image determined at a higher unit. Also, relevant information can be explicitly generated based on the information obtained in the above settings.
[0145] Hereinafter, an example of performing image splitting by an encoding / decoding apparatus according to an embodiment of the present invention will be shown.
[0146] Before the start of encoding, the splitting process for the input image can be performed. After splitting using split information (for example, image splitting information, split unit setting information, etc.), the image can be encoded in split units. After the completion of encoding, it can be saved in the memory, and the image encoded data can be recorded in a bitstream and transmitted.
[0147] The splitting process can be performed before the start of decryption. After splitting using splitting information (e.g., image splitting information, splitting unit setting information, etc.), the encrypted image data can be parsed and decrypted in units of splits. After the completion of decryption, it can be stored in memory, and a plurality of split units can be merged into one to output an image.
[0148] The splitting process of the image has been described using the above example. Also, in the present invention, a plurality of splitting processes can be performed.
[0149] For example, an image can be split, and the split unit of the image can be split. The splitting can be the same splitting process (e.g., slice / slice, tile / tile, etc.) or different splitting processes (e.g., slice / tile, tile / slice, tile / surface, surface / tile, slice / surface, surface / slice, etc.). At this time, based on the previous splitting result, the subsequent splitting process can be performed. The splitting information generated in the subsequent splitting process can be generated based on the previous splitting result.
[0150] Also, a plurality of splitting processes A can be performed, and the splitting processes can be different splitting processes (e.g., slice / surface, tile / surface, etc.). At this time, based on the previous splitting result, the subsequent splitting process can be performed, or the splitting process can be performed independently regardless of the previous splitting result. The splitting information generated in the subsequent splitting process can be generated based on the previous splitting result or independently.
[0151] The plurality of splitting processes of the image can be determined according to the encoding / decoding settings, and are not limited to the above examples, and various modified examples are also possible.
[0152] The symbolizer records the information generated in the above process in a bitstream in at least one unit among units such as sequences, pictures, slices, tiles, etc., and the decoder parses the relevant information from the bitstream. That is, it can be recorded in one unit and can be redundantly recorded in multiple units. For example, syntax elements regarding whether or not to support some information, or syntax elements regarding activation or the like can be generated in some units (for example, upper-level units), and the same or similar information as in the above case can be generated in some units (for example, lower-level units). That is, even when relevant information is supported and set in the upper-level unit, individual settings in the lower-level unit can be provided. This is not limited to the above example and can be an explanation commonly applied in the present invention. Also, it may be included in the bitstream in the form of SEI or metadata.
[0153] On the other hand, although it is common to perform encoding / decoding on the input image as it is, it can also occur to perform encoding / decoding after adjusting the size of the image (expanding or shrinking. Adjusting the resolution). For example, in a hierarchical coding method (Scalability Video Coding) that supports spatial, temporal, and picture quality scalability, image size adjustment such as overall expansion or contraction of the image can be performed. Or, image size adjustment such as partial expansion or contraction of the image can also be performed. Image size adjustment can be performed for various purposes, but it may be performed for the purpose of adaptability to the encoding environment, for the purpose of encoding uniformity, for the purpose of encoding efficiency, for the purpose of image quality improvement, or may be performed according to the type, characteristics, etc. of the image.
[0154] As a first example, a size adjustment process can be performed in a process (for example, hierarchical coding, 360-degree image coding, etc.) performed according to the characteristics, type, etc. of the image.
[0155] As a second example, a size adjustment process can be performed at the initial stage of encoding / decoding. A size adjustment process can be performed before encoding / decoding. The image whose size has been adjusted can be encoded / decoded.
[0156] As a third example, a resizing process may be performed before the prediction stage (intra-picture prediction or inter-picture prediction) or before the prediction is executed. In the resizing process, the image information in the prediction stage (for example, pixel information referred to in intra-picture prediction, intra-picture prediction mode related information, reference image information used in inter-picture prediction, inter-picture prediction mode related information, etc.) can be used.
[0157] As a fourth example, a resizing process may be performed before the filtering stage or before filtering is performed. In the resizing process, the image information in the filtering stage (for example, pixel information applied to the deblocking filter, pixel information applied to SAO, SAO filtering related information, pixel information applied to ALF, ALF filtering related information, etc.) can be used.
[0158] Also, after the resizing process is performed, the image may or may not be changed back to the pre-resizing image (in terms of image size) through the reverse resizing process. This can be determined according to the encoding / decoding settings (for example, the nature of resizing being performed, etc.). At this time, if the resizing process is an expansion, the reverse resizing process is a reduction, and if the resizing process is a reduction, the reverse resizing process can be an expansion.
[0159] When the resizing process according to the first to fourth examples is performed, the pre-resizing image can be obtained by performing the reverse resizing process at a later stage.
[0160] When hierarchical encoding or the resizing process according to the third example is performed (or when the size of the reference image is adjusted in inter-picture prediction), the reverse resizing process may not be performed at a later stage.
[0161] In one embodiment of the present invention, the image size adjustment process can be performed alone or the reverse process thereof can be carried out. In the examples described later, the description will focus on the size adjustment process. At this time, since the size adjustment reverse process is the opposite process of the size adjustment process, in order to avoid redundant explanations, the description of the size adjustment reverse process can be omitted, but it is obvious that an ordinary technician can recognize it in the same way as described in words.
[0162] FIG. 6 is an exemplary diagram of a general image size adjustment method.
[0163] Referring to 6a, an enlarged image P0+P1 can be obtained by further including a partial region P1 from the initial image (or the image before size adjustment. P0. thick solid line).
[0164] Referring to 6b, a reduced image S0 can be obtained by excluding a partial region S1 from the initial image S0+S1.
[0165] Referring to 6c, a size-adjusted image T0+T1 can be obtained by further including a partial region T1 in the initial image T0+T2 and excluding a partial region T2.
[0166] Hereinafter, in the present invention, the size adjustment process by expansion and the size adjustment process by reduction will be mainly described, but it is not limited thereto, and it should be understood that cases where size expansion and reduction are mixed and applied as in 6c are also included.
[0167] FIG. 7 is an exemplary diagram of image size adjustment according to an embodiment of the present invention.
[0168] Referring to 7a, the method of expanding an image in the size adjustment process can be described, and referring to 7b, the method of reducing an image can be described.
[0169] In 7a, the image before size adjustment is S0, and the image after size adjustment is S1. In 7b, the image before size adjustment is T0, and the image after size adjustment is T1.
[0170] When expanding the image as in 7a, it can be expanded in the up, down, left, and right directions (ET, EL, EB, ER). When shrinking the image as in 7b, it can be shrunk in the up, down, left, and right directions (RT, RL, RB, RR).
[0171] Comparing the expansion and shrinking of the image, since the up, down, left, and right directions in expansion can correspond to the respective down, up, right, and left directions in shrinking, the following will explain based on the expansion of the image, but it should be understood that the explanation of the shrinking of the image is also included.
[0172] Also, in the following, the expansion or shrinking of the image in the up, down, left, and right directions will be explained, but it should be understood that size adjustment can be performed in the upper left, upper right, lower left, and lower right directions.
[0173] At this time, when expanding in the lower right direction, the RC and BC regions are obtained. Depending on the encoding / decoding settings, it may or may not be possible to obtain the BR region. That is, it may or may not be possible to obtain the TL, TR, BL, and BR regions. For the sake of convenience in explanation, the following will explain that the corner regions (TL, TR, BL, BR regions) can be obtained.
[0174] The process of adjusting the size of the image according to an embodiment of the present invention can be performed in at least one direction. For example, it may be performed in all of the up, down, left, and right directions, or in two or more selected directions (left + right, up + down, up + left, up + right, down + left, down + right, up + left + right, down + left + right, up + down + left, up + down + right, etc.) selected from the up, down, left, and right directions, or it may be performed in only any one of the up, down, left, and right directions.
[0175] For example, it may be possible to adjust the size in the left + right, up + down, upper left + lower right, and lower left + upper right directions that can be symmetrically extended to both ends based on the center of the image, or it may be possible to adjust the size in the left + right, upper left + upper right, and lower left + lower right directions that can be vertically symmetrically extended for the image, or it may be possible to adjust the size in the up + down, upper left + lower left, and upper right + lower right directions that can be horizontally symmetrically extended for the image, and other size adjustments are also possible.
[0176] In 7a and 7b, the size of the pre - resized image (S0, T0) is defined as P_Width (width) × P_Height (height), and the size of the resized image (S1, T1) is defined as P’_Width (width) × P’_Height (height). Here, if the resizing values in the left, right, up, and down directions are defined as Var_L, Var_R, Var_T, Var_B (or generically called Var_x), the size of the resized image can be expressed as (P_Width + Var_L + Var_R) × (P_Height + Var_T + Var_B). At this time, Var_L, Var_R, Var_T, Var_B, which are the resizing values in the left, right, up, and down directions, are Exp_L, Exp_R, Exp_T, Exp_B (in this example, Exp_x is a positive number) in image expansion (Figure 7a), and can be -Rec_L, -Rec_R, -Rec_T, -Rec_B (when Rec_L, Rec_R, Rec_T, Rec_B are defined as positive numbers, expressed as negative numbers according to the reduction of the image) in image reduction. Also, the coordinates of the upper - left, upper - right, lower - left, and lower - right of the pre - resized image are (0, 0), (P_Width - 1, 0), (0, P_Height - 1), (P_Width - 1, P_Height - 1), and the coordinates of the resized image can be expressed as (0, 0), (P’_Width - 1, 0), (0, P’_Height - 1), (P’_Width - 1, P’_Height - 1). The size of the area changed (or acquired, deleted) by resizing (in this example, TL~BR. i is the index for distinguishing TL~BR) can be M[i] × N[i]. This can be expressed as Var_X × Var_Y (in this example, assume X is L or R, and Y is T or B). M and N can have various values, and can be the same regardless of i, or can have individual settings according to i. Various cases regarding this will be described later.
[0177] Referring to 7a, S1 can be composed of including all or part of TL~BR (upper - left to lower - right) generated by expanding S0 in various directions. Referring to 7b, T1 can be composed of excluding all or part of TL~BR removed by reducing T0 in various directions.
[0178] In 7a, when expanding the existing image S0 in the up, down, left, and right directions, the image can be composed including the TC, BC, LC, and RC regions obtained through each size adjustment process, and further can also include the TL, TR, BL, and BR regions.
[0179] As an example, when expanding in the up (ET) direction, the image can be composed including the TC region in the existing image S0, and can include the TL or TR region according to the expansion in at least one different direction (EL or ER).
[0180] As an example, when expanding in the down (EB) direction, the image can be composed including the BC region in the existing image S0, and can include the BL or BR region according to the expansion in at least one different direction (EL or ER).
[0181] As an example, when expanding in the left (EL) direction, the image can be composed including the LC region in the existing image S0, and can include the TL or BL region according to the expansion in at least one different direction (ET or EB).
[0182] As an example, when expanding in the right (ER) direction, the image can be composed including the RC region in the existing image S0, and can include the TR or BR region according to the expansion in at least one different direction (ET or EB).
[0183] According to an embodiment of the present invention, a setting (for example, spa_ref_enabled_flag or tem_ref_enabled_flag) can be placed to spatially or temporally limit the referability of the region to be size-adjusted (assumed to be expansion in this example).
[0184] That is, it is possible to refer to the data of the region whose size is adjusted spatially or temporally according to the encoding / decoding setting (for example, spa_ref_enabled_flag = 1 or tem_ref_enabled_flag = 1), or to restrict the reference (for example, spa_ref_enabled_flag = 0 or tem_ref_enabled_flag = 0).
[0185] Encoding / decoding of the image before size adjustment (S0, T1) and the regions added or deleted during size adjustment (TC, BC, LC, RC, TL, TR, BL, BR regions) can be performed as follows.
[0186] For example, in the encoding / decoding of the image before size adjustment and the regions added or deleted, the data of the image before size adjustment and the data of the regions added or deleted (encoded / decoded data, such as pixel values or prediction-related information) can be referred to each other spatially or temporally.
[0187] Alternatively, while the data of the image before size adjustment and the data of the regions added or deleted can be referred to spatially, the data of the image before size adjustment can be referred to temporally, and the data of the regions added or deleted cannot be referred to temporally.
[0188] That is, a setting can be made to restrict the referability of the regions added or deleted. The setting information regarding the referability of the regions added or deleted can be generated explicitly or determined implicitly.
[0189] The image size adjustment process according to an embodiment of the present invention may include an image size adjustment instruction stage, an image size adjustment type identification stage, and / or an image size adjustment execution stage. Further, the image encoding device and the decoding device may include an image size adjustment instruction unit, an image size adjustment type identification unit, and an image size adjustment execution unit that realize the image size adjustment instruction stage, the image size adjustment type identification stage, and the image size adjustment execution stage. In the case of encoding, associated syntax elements can be generated, and in the case of decoding, the associated syntax elements can be parsed.
[0190] In the image size adjustment instruction stage, it is possible to determine whether to perform image size adjustment. For example, when a signal indicating image size adjustment (e.g., img_resizing_enabled_flag) is confirmed, size adjustment can be performed. When the signal indicating image size adjustment is not confirmed, size adjustment may not be performed or other encoding / decoding information may be confirmed to perform size adjustment. Also, even if a signal indicating image size adjustment is not provided, the signal indicating image size adjustment may be implicitly activated or deactivated according to the encoding / decoding settings (e.g., characteristics and types of images). When performing size adjustment, size adjustment related information can be generated thereby, or it is also possible that the size adjustment related information is implicitly determined.
[0191] When a signal indicating image size adjustment is provided, the corresponding signal is a signal for indicating whether to perform image size adjustment of the image, and it is possible to confirm whether to perform size adjustment of the corresponding image according to the signal.
[0192] For example, when a signal indicating image size adjustment (e.g., img_resizing_enabled_flag) is confirmed and the corresponding signal is activated (e.g., img_resizing_enabled_flag = 1), image size adjustment can be performed. When the corresponding signal is deactivated (e.g., img_resizing_enabled_flag = 0), it can be meant that image size adjustment is not performed.
[0193] Also, when a signal for instructing image size adjustment is not provided, size adjustment is not performed, or whether to perform size adjustment on the image can be confirmed by other signals.
[0194] For example, when dividing an input image in block units, size adjustment (in this example, in the case of expansion. Assume that the size adjustment process is performed when it is not an integer multiple) can be performed according to whether the size of the image (e.g., width or height) is an integer multiple of the size of the block (e.g., width or height). That is, when the width of the image is not an integer multiple of the width of the block, or when the height of the image is not an integer multiple of the height of the block, size adjustment can be performed. At this time, the size adjustment information (e.g., size adjustment direction, size adjustment value, etc.) can be determined according to the encoding / decoding information (e.g., size of the image, size of the block, etc.). Or, size adjustment can be performed according to the characteristics, type of the image (e.g., 360-degree image, etc.), and the size adjustment information can be explicitly generated or assigned to a predetermined value. It is not limited to only the case of the above example, and deformation to other examples is also possible.
[0195] In the image size adjustment type identification stage, the image size adjustment type can be identified. The image size adjustment type can be defined by the method of performing size adjustment, size adjustment information, etc. For example, size adjustment using a scale factor, size adjustment using an offset factor, etc. can be performed. It is not limited to this, and mixed application of the above methods is also possible. For the convenience of explanation, size adjustment using a scale factor and an offset factor will be mainly described.
[0196] In the case of a scale factor, resizing can be performed in a manner of multiplying or dividing based on the size of the image. Information about the resizing operation (e.g., expansion or reduction) can be explicitly generated, and the expansion or reduction process can be performed according to the corresponding information. Also, according to the encoding / decoding settings, the resizing process can be performed with a predetermined operation (e.g., either expansion or reduction). In this case, the information about the resizing operation can be omitted. For example, when image resizing is activated at the image size adjustment instruction stage, the resizing of the image can be performed with a predetermined operation.
[0197] The resizing direction can be at least one direction selected from the up, down, left, and right directions. Depending on the resizing direction, at least one scale factor may be required. That is, one scale factor (in this example, unidirectional) is required for each direction, one scale factor (in this example, bidirectional) is required according to the horizontal or vertical direction, and one scale factor (in this example, omnidirectional) may be required according to the overall direction of the image. Also, the resizing direction is not limited only to the case of the above example, and variations to other examples are also possible.
[0198] The scale factor can have a positive value, and the range information can be set differently according to the encoding / decoding settings. For example, when generating information by mixing the resizing operation and the scale factor, the scale factor can be used as the value to be multiplied. When it is greater than 0 or less than 1, it can mean a reduction operation. When it is greater than 1, it can mean an expansion operation. When it is 1, it can mean not performing resizing. As another example, when generating scale factor information separately from the resizing operation, in the case of an expansion operation, the scale factor can be used as the value to be multiplied, and in the case of a reduction operation, the scale factor can be used as the value to be divided.
[0199] Referring back to FIGS. 7a and 7b of FIG. 7, the process of changing from the pre-resized images (S0, T0) to the resized images (in this example, S1, T1) using the scale factor can be described.
[0200] As an example, when using one scale factor (referred to as sc) according to the overall direction of the image, and the resizing direction is the down + right direction, the resizing direction is ER, EB (or RR, RB), and the resizing values Var_L (Exp_L or Rec_L) and Var_T (Exp_T or Rec_T) are 0, and Var_R (Exp_R or Rec_R) and Var_B (Exp_B or Rec_B) can be expressed as P_Width×(sc - 1) and P_Height×(sc - 1). Therefore, the resized image can be (P_Width×sc)×(P_Height×sc).
[0201] As an example, when using respective scale factors (in this example, sc_w, sc_h) according to the horizontal or vertical direction of the image, and the resizing direction is the left + right, up + down direction (when both operate, it is up + down + left + right), the resizing direction is ET, EB, EL, ER, and the resizing values Var_T and Var_B are P_Height×(sc_h - 1) / 2, and Var_L and Var_R can be P_Width×(sc_w - 1) / 2. Therefore, the resized image can be (P_Width×sc_w)×(P_Height×sc_h).
[0202] In the case of the offset factor, resizing can be performed by adding or subtracting based on the size of the image. Or, resizing can be performed by adding or subtracting based on the encoding / decoding information of the image. Or, resizing can be performed by adding or subtracting independently. That is, the resizing process can be set dependently or independently.
[0203] Information about the resizing operation (e.g., expansion or contraction) can be explicitly generated, and the expansion or contraction process can be performed according to the corresponding information. Also, according to the encoding / decoding settings, a resizing operation can be performed with a predetermined operation (e.g., either expansion or contraction). In this case, the information about the resizing operation can be omitted. For example, when image resizing is activated at the image resizing instruction stage, the resizing of the image can be performed with a predetermined operation.
[0204] The resizing direction can be at least one of the up, down, left, and right directions. At least one offset factor may be required according to the resizing direction. That is, one offset factor (in this example, unidirectional) is required for each direction, either offset factor (in this example, symmetric bidirectional) is required according to the horizontal or vertical direction, one offset factor (in this example, asymmetric bidirectional) is required according to the partial combination of each direction, and one offset factor (in this example, omnidirectional) may be required according to the overall direction of the image. Also, the resizing direction is not limited only to the case of the above example, and deformation to other examples is also possible.
[0205] The offset factor can have a positive value or can have both positive and negative values, and can be set with different range information according to the encoding / decoding settings. For example, when generating information by mixing a sizing operation and an offset factor (assuming both positive and negative values in this example), the offset factor can be used as a value to be added or subtracted according to the sign information of the offset factor. When the offset factor is greater than 0, it can mean an expansion operation; when it is less than 0, it can mean a reduction operation; and when it is 0, it can mean no sizing operation. As another example, when generating offset factor information separately from the sizing operation (assuming a positive value in this example), the offset factor can be used as a value to be added or subtracted according to the sizing operation. When it is greater than 0, an expansion or reduction operation can be performed according to the sizing operation; when it is 0, it can mean no sizing operation.
[0206] Referring again to FIGS. 7a and 7b of FIG. 7, a method of changing from an image (S0, T0) before sizing to an image (S1, T1) after sizing using an offset factor can be described.
[0207]
[0208] As an example, respective offset factors (os_w, os_h) are used according to the horizontal or vertical direction of the image. When the resizing directions are left + right, up + down directions (when both operate, up + down + left + right), the resizing directions are ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values Var_T and Var_B can be os_h, and Var_L and Var_R can be os_w. The image size after resizing can be {P_Width+(os_w×2)}×{P_Height+(os_h×2)}.
[0209] As an example, when the resizing directions are down and right directions (when operating together, down + right), and respective offset factors (os_b, os_r) are used according to the resizing directions, the resizing directions are EB, ER (or RB, RR), and the resizing value Var_B can be os_b, and Var_R can be os_r. The image size after resizing can be (P_Width+os_r)×(P_Height+os_b).
[0210] As an example, respective offset factors (os_t, os_b, os_l, os_r) are used according to each direction of the image. When the resizing directions are up, down, left, right directions (when all operate, up + down + left + right), the resizing directions are ET, EB, EL, ER (or RT, RB, RL, RR), and the resizing values Var_T can be os_t, Var_B can be os_b, Var_L can be os_l, and Var_R can be os_r. The image size after resizing can be (P_Width+os_l+os_r)×(P_Height+os_t+os_b).
[0211] The above example shows the case where the offset factor is used as the sizing value (Var_T, Var_B, Var_L, Var_R) in the sizing process. That is, the offset factor means the case where it is used as the sizing value as it is. This can be an example of sizing performed independently. Or, it is also possible that the offset factor is used as an input variable of the sizing value. Specifically, the offset factor can be assigned as an input variable, and the sizing value can be obtained through a series of processes according to the encoding / decoding settings. This can be an example of sizing performed based on predetermined information (such as the size of an image, encoding / decoding information, etc.) or an example of sizing performed dependently.
[0212] For example, the offset factor can be a multiple (such as 1, 2, 4, 6, 8, 16, etc.) or an exponent (such as a power of 2 like 1, 2, 4, 8, 16, 32, 64, 128, 256, etc.) of a predetermined value (in this example, an integer). Or, it can be a multiple or an exponent of a value obtained based on the encoding / decoding settings (such as a value set based on the motion search range of inter-picture prediction). Or, it can be a multiple or an exponent of a unit (assumed to be A×B in this example) obtained from the picture partitioning section. Or, it can be a multiple of a unit (in the case of a tile, etc., assumed to be E×F in this example) obtained from the picture partitioning section.
[0213] Or, it can be a value in a range smaller than or equal to the horizontal and vertical of the unit obtained from the picture partitioning section. The multiples or exponents in the above examples can include the case where the value is 1, and are not limited to the above examples, and variations to other examples are also possible. For example, when the offset factor is n, Var_x can be 2×n or 2 n and can be.
[0214] In addition, individual offset factors can be supported according to color components, and offset factors for some color components can be supported to derive offset factor information for other color components. For example, when an offset factor A for a luminance component (assuming that in this example, the composition ratio of the luminance component to the color difference component is 2:1) is explicitly generated, an offset factor A / 2 for the color difference component can be implicitly obtained. Or, when an offset factor A for the color difference component is explicitly generated, an offset factor 2A for the luminance component can be implicitly obtained.
[0215] Information about the size adjustment direction and the size adjustment value can be explicitly generated, and the size adjustment process can be performed according to the corresponding information. Also, it can be implicitly determined according to the encoding / decoding settings, thereby enabling the size adjustment process. At least one predetermined direction or adjustment value can be assigned, in which case the related information can be omitted. At this time, the encoding / decoding settings can be determined based on the characteristics, types, encoding information, etc. of the image. For example, at least one size adjustment direction by at least one size adjustment operation can be predetermined, at least one size adjustment value by at least one size adjustment operation can be predetermined, and at least one size adjustment value by at least one size adjustment direction can be predetermined. Also, the size adjustment direction and the size adjustment value in the reverse size adjustment process can be derived from the size adjustment direction and the size adjustment value applied in the size adjustment process. At this time, the implicitly determined size adjustment value can be one of the examples described above (examples where various size adjustment values are obtained).
[0216] In addition, although the cases of multiplication or division in the above example have been described, it can be realized by a shift operation according to the implementation of the encoder / decoder. When multiplying, it can be realized by using a left shift operation, and when dividing, it can be realized by using a right shift operation. This is not limited to the above example and can be an explanation commonly applicable in the present invention.
[0217] In the image size adjustment execution stage, image size adjustment can be performed based on the identified size adjustment information. That is, image size adjustment can be performed based on information such as the size adjustment type, size adjustment operation, size adjustment direction, and size adjustment value, and encoding / decoding can be performed based on the obtained image after size adjustment.
[0218] Also, in the image size adjustment execution stage, size adjustment can be performed using at least one data processing method. Specifically, size adjustment can be performed using at least one data processing method on the area to be size-adjusted according to the size adjustment type and size adjustment operation. For example, according to the size adjustment type, it can be determined how to fill the data when the size adjustment is an expansion, or how to remove the data when the size adjustment process is a reduction.
[0219] In summary, image size adjustment can be performed based on the size adjustment information identified in the image size adjustment execution stage. Or, in the image size adjustment execution stage, image size adjustment can be performed based on the size adjustment information and the data processing method. The difference between the two cases lies in whether only the size of the image for encoding / decoding is adjusted, or whether the data processing of the image size and the area to be size-adjusted is also considered. Whether to include the data processing method in the image size adjustment execution stage can be determined according to the application stage, position, etc. of the size adjustment process. In the examples described later, examples of size adjustment based on the data processing method will be mainly described, but it is not limited to this.
[0220] When performing size adjustment using an offset factor, size adjustment can be performed using various methods in the case of expansion and reduction. In the case of expansion, size adjustment can be performed using a method of filling at least one piece of data, and in the case of reduction, size adjustment can be performed using a method of removing at least one piece of data. At this time, in the case of size adjustment using an offset factor, new data or data of an existing image can be directly or after deformation filled into the size adjustment area (expansion), and simple removal or removal through a series of processes can be applied to the size adjustment area (reduction) for removal.
[0221] When performing size adjustment using a scale factor, in some cases (for example, hierarchical coding, etc.), expansion can perform size adjustment by applying upsampling, and reduction can perform size adjustment by applying downsampling. For example, in the case of expansion, at least one upsampling filter can be used, and in the case of reduction, at least one downsampling filter can be used. The filters applied horizontally and vertically may be the same or different. At this time, in the case of size adjustment using a scale factor, new data is not generated or removed from the size adjustment area, but rather the data of the existing image can be rearranged using methods such as interpolation. The data processing method related to the execution of size adjustment can be classified by the filter used for the sampling. Also, in some cases (for example, cases similar to the offset factor), expansion can perform size adjustment using a method of filling at least one piece of data, and reduction can perform size adjustment using a method of removing at least one piece of data. In the present invention, the data processing method in the case of performing size adjustment using an offset factor will be mainly described.
[0222] Generally, a predetermined single data processing method can be used for the area to be resized. However, as in the examples described later, at least one data processing method can also be used for the area to be resized, and selection information for the data processing method can also be generated. In the former case, it can be meant that resizing is performed via a fixed data processing method, and in the latter case, via an adaptive data processing method.
[0223] Also, a data processing method common to the entire area (TL, TC, TR,..., BR in FIGS. 7a and 7b) added or deleted during resizing can be applied, or a data processing method can be applied in units of a part of the area added or deleted during resizing (for example, each of TL to BR in FIGS. 7a and 7b or a combination of a part thereof).
[0224] FIG. 8 is an exemplary diagram of a method for constructing an area to be expanded in an image resizing method according to an embodiment of the present invention.
[0225] Referring to FIG. 8a, for the sake of convenience of explanation, the image can be divided into TL, TC, TR, LC, C, RC, BL, BC, and BR areas, which can respectively correspond to the upper left, upper, upper right, left, center, right, lower left, lower, and lower right positions of the image. Hereinafter, the case where the image is expanded in the downward + rightward direction will be described, but it should be understood that the same applies to other directions.
[0226] The area added according to the expansion of the image can be configured in various ways. For example, it can be filled with an arbitrary value or filled by referring to a part of the image data.
[0227] Referring to FIG. 8b, the areas (A0, A2) expanded to arbitrary pixel values can be filled. The arbitrary pixel value can be determined using various methods.
[0228] As an example, any pixel value can be one pixel belonging to the range of pixel values that can be represented by the bit depth {for example, from 0 to 1<<(bit_depth)-1}. For example, it can be the minimum value, maximum value, median value {such as 1<<(bit_depth-1), etc.} of the range of the pixel values (where bit_depth is the bit depth).
[0229] As an example, any pixel value can be one pixel belonging to the range of pixel values of the pixels belonging to the image {for example, from min P to max P inclusive. min P , max P are the minimum value and maximum value of the pixels belonging to the image. min P is equal to or greater than 0, and max P is equal to or less than 1<<(bit_depth)-1}. For example, any pixel value can be the minimum value, maximum value, median value, average (of at least two pixels), weighted sum, etc. of the range of the pixel values.
[0230] As an example, any pixel value can be a value determined by the range of pixel values of a partial region belonging to the image. For example, when configuring A0, the partial region can be TR+RC+BR. Also, the partial region can be set as the corresponding region of 3×9 of TR, RC, BR, or can be set as the corresponding region of 1×9 <assuming the rightmost line>. This can be adjusted according to the encoding / decoding settings. At this time, the partial region may also be a unit divided from the picture division unit. Specifically, any pixel value can be the minimum value, maximum value, median value, average (of at least two pixels), weighted sum, etc. of the range of the pixel values.
[0231] Referring back to 8b, the region A1 added according to the expansion of the image can be filled using pattern information (for example, assuming that what uses a plurality of pixels is a pattern. It is not necessarily required to follow a certain rule.) generated using a plurality of pixel values. At this time, the pattern information can be defined according to the encoding / decoding settings or can generate related information, and at least one pattern information can be used to fill the expanded region.
[0232] Referring to 8c, the region added according to the expansion of the image can be configured by referring to the pixels of a partial region belonging to the image. Specifically, the region to be added can be configured by copying or padding the pixels of the region adjacent to the region to be added (hereinafter, referred to as reference pixels). At this time, the pixels of the region adjacent to the region to be added can be pixels before encoding or pixels after encoding (or decoding). For example, when performing size adjustment in the pre-encoding stage, the reference pixels can mean the pixels of the input image, and when performing size adjustment in the in-screen prediction reference pixel generation stage, the reference image generation stage, the filtering stage, etc., the reference pixels can mean the pixels of the restored image. In this example, it is assumed that the pixels closest to the region to be added are used, but it is not limited thereto.
[0233] The region A0 expanded in the left or right direction related to the horizontal size adjustment of the image can be configured by horizontally padding (Z0) the outer pixels adjacent to the region A0 to be expanded, and the region A1 expanded in the up or down direction related to the vertical size adjustment of the image can be configured by vertically padding (Z1) the outer pixels adjacent to the region A1 to be expanded. Also, the region A2 expanded in the lower right direction can be configured by diagonally padding (Z2) the outer pixels adjacent to the region A2 to be expanded.
[0234] Referring to 8d, the expanded regions B'0 to B'2 can be configured by referring to the data B0 to B2 of a partial region belonging to the image. In 8d, it can be distinguished from 8c in that regions not adjacent to the expanded region can be referred to.
[0235] For example, when there is an area with a high correlation with the area to be extended within the image, the area to be extended can be filled by referring to the pixels of the area with a high correlation. At this time, the position information, area size information, etc. of the area with a high correlation can be generated. Or, when there is an area with a high correlation through encoding / decoding information such as the characteristics and types of the image, and the position information, size information, etc. of the area with a high correlation can be implicitly confirmed (for example, in the case of a 360-degree image, etc.), the data of the corresponding area can be filled into the area to be extended. At this time, the position information, area size information, etc. of the area can be omitted.
[0236] As an example, in the case of the area B'2 extended in the left or right direction related to the horizontal size adjustment of the image, the area to be extended can be filled by referring to the pixels of the area B2 on the opposite side of the area to be extended in the left or right direction related to the horizontal size adjustment.
[0237] As an example, in the case of the area B'1 extended in the up or down direction related to the vertical size adjustment of the image, the area to be extended can be filled by referring to the pixels of the area B1 on the opposite side of the area to be extended in the up or right direction related to the vertical size adjustment.
[0238] As an example, in the case of the area B'0 extended in the diagonal direction (in this example, with the center of the image as the reference) for partial size adjustment of the image, the area to be extended can be filled by referring to the pixels of the area B0, TL on the opposite side of the area to be extended.
[0239] In the above example, the case of obtaining the data of the area that has continuity at the boundaries at both ends of the image and is symmetric to the size adjustment direction has been described, but it is not limited to this, and it is also possible to obtain it from the data of other areas TL~BR.
[0240] When filling the data of a partial area of an image into the area to be expanded, can the data of the corresponding area be directly copied and filled, or can it be filled after a conversion process based on the characteristics, type, etc. of the image? At this time, if it is directly copied, it can mean using the pixel values of the corresponding area as they are. If it goes through the conversion process, it can mean not using the pixel values of the corresponding area as they are. That is, through the conversion process, at least one pixel value of the area changes, and it is filled into the area to be expanded, or the acquisition positions of some pixels may be at least one different. That is, to fill the area A×B to be expanded, instead of using the A×B data of the corresponding area, C×D data can be used. In other words, at least one of the motion vectors applied to the pixels to be filled may be different. The above example can be an example that occurs when filling the area to be expanded using the data of other surfaces when the 360-degree image is composed of multiple surfaces according to the projection format. The data processing method for filling the area expanded by resizing the image is not limited to the above example, and this can be improved and deformed, or additional data processing methods can be used.
[0241] According to the encoding / decoding settings, candidate groups for multiple data processing methods can be supported, and data processing method selection information can be generated from among the multiple candidate groups and recorded in the bitstream. For example, one data processing method can be selected from methods such as filling using a predetermined pixel value, copying and filling the outer pixels, copying and filling a partial area of the image, and converting and filling a partial area of the image, and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0242] For example, the data processing method applied to the entire area expanded by image size adjustment (in this example, TL to BR in FIG. 7a) can be any one of a method of filling with a predetermined pixel value, a method of copying and filling the outer pixels, a method of copying and filling a partial area of the image, a method of converting and filling a partial area of the image, and other methods, and selection information for this can be generated. Also, a predetermined data processing method applied to the entire area can be determined.
[0243] Alternatively, the data processing method applied to the area expanded by image size adjustment (in this example, each area of TL to BR in FIG. 7a of FIG. 7 or two or more of them) can be any one of a method of filling with a predetermined pixel value, a method of copying and filling the outer pixels, a method of copying and filling a partial area of the image, a method of converting and filling a partial area of the image, and other methods, and selection information for this can be generated. Also, a predetermined data processing method applied to at least one area can be determined.
[0244] FIG. 9 is an exemplary diagram of a method of configuring an area to be deleted and an area to be generated by reducing the image size in an image size adjustment method according to an embodiment of the present invention.
[0245] The area deleted in the image reduction process can be simply removed, or can be removed after a series of utilization processes.
[0246] Referring to FIG. 9a, in the image reduction process, partial areas A0, A1, and A2 can be simply removed without an additional utilization process. At this time, image A may be subdivided and called TL to BR as in FIG. 8a.
[0247] Referring to 9b, although some regions A0 to A2 are removed, they can be utilized as reference information during the encoding / decoding of image A. For example, in the restoration process or correction process of a partial region of the reduced and generated image A, the partially removed regions A0 to A2 can be utilized. In the said restoration or correction process, weighted sums, averages, etc. of two regions (the removed region and the generated region) can be used. Also, the said restoration or correction process can be a process applicable when the two regions have a high correlation.
[0248] As an example, the region B’2 removed by being reduced in the left or right direction related to the horizontal size adjustment of the image can be used to restore or correct the pixels of the region B2, LC on the opposite side of the reduced region in the left or right direction related to the horizontal size adjustment, and then the corresponding region can be removed from the memory.
[0249] As an example, the region B’1 removed in the upward or downward direction related to the vertical size adjustment of the image can be used in the encoding / decoding process (restoration or correction process) of the region B1, TR on the opposite side of the reduced region in the upward or downward direction related to the vertical size adjustment, and then the corresponding region can be removed from the memory.
[0250] As an example, the region B’0 reduced in the diagonal direction with respect to the center of the image (in this example) can be used in the encoding / decoding process (such as restoration or correction process) of the region B0, TL on the opposite side of the reduced region, and then the corresponding region can be removed from the memory.
[0251] In the above example, an explanation has been given regarding the case of using it for data restoration or correction of regions with continuity at the boundaries of both ends of the image and located at positions symmetric to the size adjustment direction. However, it is not limited to this, and it can also be removed from the memory after being used for data restoration or correction of other regions TL to BR other than the symmetric positions.
[0252] The data processing method for removing the reduced regions of the present invention is not limited to the above example, and it can be improved and changed, or additional data processing methods can be used.
[0253] Candidate groups for a plurality of data processing methods can be supported according to the settings of symbolization / decryption, and selection information for this can be generated and recorded in the bit stream. For example, one data processing method can be selected from methods such as simply removing the area to be resized, removing the area to be resized after using it in a series of processes, etc., and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0254] For example, the data processing method applied to the entire area (in this example, TL~BR in FIG. 7b) that is deleted by being reduced by image resizing is any one of the method of simply removing it, the method of removing it after using it in a series of processes, and other methods, and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0255] Alternatively, the data processing method applied to the individual areas (in this example, each of TL~BR in FIG. 7b) that are reduced by image resizing is any one of the method of simply removing it, the method of removing it after using it in a series of processes, and other methods, and selection information for this can be generated. Also, the data processing method can be implicitly determined.
[0256] In the above example, the case regarding the execution of resizing by the resizing operation (expansion, reduction) has been described. However, in some cases, it can be an example applicable to the case where, after performing the resizing operation (in this example, expansion), the resizing operation that is the reverse process thereof (in this example, reduction) is performed.
[0257] For example, a method of filling an expanded area using partial data of an image is selected, and in the reverse process, a method of removing the reduced area after using it in the process of restoring or correcting partial data of the image is selected. Or, a method of filling an expanded area using a copy of the outline pixels is selected, and in the reverse process, a method of simply removing the reduced area is selected. That is, based on the data processing method selected in the image size adjustment process, the data processing method in the reverse process can be determined.
[0258] Different from the above example, the data processing methods in the image size adjustment process and its reverse process can also have an independent relationship. That is, the data processing method in the reverse process can be selected regardless of the data processing method selected in the image size adjustment process. For example, a method of filling an expanded area using partial data of an image is selected, and a method of simply removing the reduced area in the reverse process can be selected.
[0259] In the present invention, the data processing method in the image size adjustment process can be implicitly determined according to the encoding / decoding setting, and the data processing method in the reverse process can be implicitly determined according to the encoding / decoding setting. Or, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the reverse process can be explicitly generated. Or, the data processing method in the image size adjustment process can be explicitly generated, and the data processing method in the reverse process can be implicitly determined based on the data processing method.
[0260] Next, an example of performing image size adjustment with an encoding / decoding device according to an embodiment of the present invention is shown. In the example described later, the case where the size adjustment process is expansion and the reverse size adjustment process is reduction is taken as an example. Also, the difference between the "image before size adjustment" and the "image after size adjustment" can mean the size of the image, and the size adjustment related information can be partially explicitly generated or partially implicitly determined according to the encoding / decoding setting. Also, the size adjustment related information can include information for the size adjustment process and the reverse size adjustment process.
[0261] As a first example, a size adjustment process for the input image can be performed before the start of encoding. After performing size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is what is used in the size adjustment process), the image after size adjustment can be encoded. After the completion of encoding, it can be stored in memory, and the image encoding data (in this example, meaning the image after size adjustment) can be recorded in a bitstream and transmitted.
[0262] A size adjustment process can be performed before the start of decoding. After performing size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, etc.), the image decoding data after size adjustment can be parsed and decoded. After the completion of decoding, it can be stored in memory, and a reverse size adjustment process (using, for example, a data processing method in this example. This is what is used in the reverse size adjustment process) can be performed to change the output image to the image before size adjustment.
[0263] As a second example, a size adjustment process for the reference image can be performed before the start of encoding. After performing size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is what is used in the size adjustment process), the image after size adjustment (in this example, the reference image after size adjustment) can be stored in memory, and this can be used to encode the image. After the completion of encoding, the image encoding data (in this example, meaning what is encoded using the reference image) can be recorded in a bitstream and transmitted. Also, when the encoded image is stored in memory as a reference image, a size adjustment process can be performed as in the above process.
[0264] Before the start of decoding, a size adjustment process for the reference image can be performed. Using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is what is used in the size adjustment process), the image after size adjustment (in this example, the reference image after size adjustment) can be stored in memory, and the image decoding data (in this example, the same as what was encoded using the reference image by the encoder) can be parsed and decoded. After the completion of decoding, it can be generated as an output image. When the decoded image is included in the reference image and stored in memory, the size adjustment process can be performed as in the above process.
[0265] As a third example, a size adjustment process for the image can be performed before the start of image filtering (assumed to be a deblocking filter in this example) after the completion of encoding (specifically, meaning the completion of encoding excluding the filtering process). After performing size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is what is used in the size adjustment process), an image after size adjustment can be generated, and filtering can be applied to the image after size adjustment. After the completion of filtering, a reverse size adjustment process can be performed to change it back to the image before size adjustment.
[0266] (Specifically, meaning the completion of decoding excluding the filtering process) A size adjustment process for the image can be performed before the start of image filtering after the completion of decoding. After performing size adjustment using size adjustment information (for example, size adjustment operation, size adjustment direction, size adjustment value, data processing method, etc. The data processing method is what is used in the size adjustment process), an image after size adjustment can be generated, and filtering can be applied to the image after size adjustment. After the completion of filtering, a reverse size adjustment process can be performed to change it back to the image before size adjustment.
[0267] In the above examples, in some cases (the first exemplary case and the third exemplary case), the resizing process and the reverse resizing process may be performed, and in some other cases (the second exemplary case), only the resizing process may be performed.
[0268] Also, in some cases (the second exemplary case and the third exemplary case), the resizing processes in the encoder and the decoder may be the same, and in some other cases (the first exemplary case), the resizing processes in the encoder and the decoder may or may not be the same. At this time, the difference in the resizing process in the encoder / decoder can be at the resizing execution stage. For example, in some cases (in this example, the encoder), it can include the resizing execution stage that considers the resizing of the image and the data processing of the area to be resized, and in some cases (in this example, the decoder), it can include the resizing execution stage that considers the resizing of the image. At this time, the former data processing can correspond to the data processing of the reverse resizing process of the latter.
[0269] Also, the resizing process in some cases (the third exemplary case) is a process that is applied only at the corresponding stage, and it may not be necessary to store the resizing area in the memory. For example, it can be stored in a temporary memory for the purpose of use in the filtering process for filtering, and the corresponding area can be removed through the reverse resizing process. In this case, it can be said that there is no change in the size of the image due to resizing. The present invention is not limited to the above examples, and modifications to other examples are also possible.
[0270] Through the size adjustment process, the size of the image can be changed, and thereby, the coordinates of some pixels of the image can be changed through the size adjustment process. This can affect the operation of the picture division unit. In the present invention, block-based division can be performed based on the image before size adjustment through the above process, or block-based division can be performed based on the image after size adjustment. Also, division of some units (for example, tiles, slices, etc.) can be performed based on the image before size adjustment, or division of some units can be performed based on the image after size adjustment. This can be determined according to the encoding / decoding settings. In the present invention, the case where the picture division unit operates based on the image after size adjustment (for example, the image division process after the size adjustment process) will be mainly described, but other variations are also possible. This will be described for the cases corresponding to the above example in a plurality of image settings described later.
[0271] In the encoder, the information generated in the above process is recorded in the bitstream in at least one unit among units such as sequence, picture, slice, tile, etc., and in the decoder, the related information is parsed from the bitstream. Also, it may be included in the bitstream in the form of SEI or metadata.
[0272] It may be common to perform encoding / decoding on the input image as it is, but it may also occur when the image is reconstructed and then encoding / decoding is performed. For example, the image can be reconstructed for the purpose of improving the encoding efficiency of the image, the image can be reconstructed for the purpose of considering the network and user environment, and the image can be reconstructed according to the type and characteristics of the image.
[0273] In the present invention, the image reconstruction process can perform the reconstruction process alone, or can perform the reverse process thereof. In the example described later, the reconstruction process will be mainly described, but the content regarding the reconstruction reverse process can be induced inversely from the reconstruction process.
[0274] FIG. 10 is an exemplary diagram for image reconstruction according to an embodiment of the present invention.
[0275] When taking the first input image as 10a, 10a to 10d are exemplary diagrams to which a rotation including 0 degrees (for example, a candidate group generated by sampling 360 degrees into k intervals can be formed. k can have values such as 2, 4, 8, etc., and in this example, it is assumed to be 4 for explanation) is applied to the image, and 10e to 10h show exemplary diagrams to which inversion (or symmetry) is applied based on 10a or based on 10b to 10d.
[0276] Depending on the reconstruction of the image, the start position or scan order of the image can be changed, or it can follow a predetermined start position and scan order regardless of whether it is reconstructed or not. This can be determined according to the encoding / decoding settings. In the embodiments described later, it is assumed for explanation that it follows a predetermined start position (for example, the upper left position of the image) and scan order (for example, raster scan) regardless of whether the image is reconstructed or not.
[0277] In the image encoding method and decoding method according to an embodiment of the present invention, the following image reconstruction stages can be included. At this time, the image reconstruction process can include an image reconstruction instruction stage, an image reconstruction type identification stage, and an image reconstruction execution stage. Also, the image encoding device and decoding device can be configured to include an image reconstruction instruction unit, an image reconstruction type identification unit, and an image reconstruction execution unit that realize the image reconstruction instruction stage, the image reconstruction type identification stage, and the image reconstruction execution stage. In the case of encoding, associated syntax elements can be generated, and in the case of decoding, the associated syntax elements can be parsed.
[0278] In the image reconstruction instruction stage, it is possible to determine whether to perform image reconstruction. For example, when a signal indicating image reconstruction (e.g., convert_enabled_flag) is confirmed, reconstruction can be performed. When the signal indicating image reconstruction is not confirmed, reconstruction may not be performed, or other encoding / decoding information can be confirmed to perform reconstruction. Also, even if a signal indicating image reconstruction is not provided, the signal indicating reconstruction may be implicitly activated or deactivated according to the encoding / decoding settings (e.g., image characteristics, type, etc.). When performing reconstruction, reconstruction-related information can be generated thereby, or the reconstruction-related information can be implicitly determined.
[0279] When a signal indicating image reconstruction is provided, the corresponding signal is a signal for indicating whether to perform image reconstruction, and it is possible to confirm whether to reconstruct the corresponding image according to the signal. For example, when a signal indicating image reconstruction (e.g., convert_enabled_flag) is confirmed and the corresponding signal is activated (e.g., convert_enabled_flag = 1), reconstruction may be performed. When the corresponding signal is deactivated (e.g., convert_enabled_flag = 0), reconstruction may not be performed.
[0280] Also, when a signal indicating image reconstruction is not provided, reconstruction may not be performed, or whether to reconstruct the corresponding image can be confirmed by other signals. For example, reconstruction can be performed according to image characteristics, type, etc. (e.g., 360-degree image), and the reconstruction information can be explicitly generated or assigned to a predetermined value. It is not limited to only the above example, and deformation to other examples is also possible.
[0281] In the image reconstruction type identification stage, the image reconstruction type can be identified. The image reconstruction type can be defined by the method of performing reconstruction, reconstruction mode information, etc. The method of performing reconstruction (e.g., convert_type_flag) can include rotation, inversion, etc., and the reconstruction mode information can include the mode (e.g., convert_mode) in the method of performing reconstruction. In this case, the reconstruction-related information can be composed of the method of performing reconstruction and the mode information. That is, it can be composed of at least one syntax element. At this time, the number of candidate groups of mode information for each method of performing reconstruction may be the same or different.
[0282] As an example, in the case of rotation, it can include candidates having a difference at regular intervals (in this example, 90 degrees) such as 10a to 10d. When 10a is a 0-degree rotation, 10b to 10d can be examples where 90-degree, 180-degree, and 270-degree rotations are applied respectively (in this example, the angle is measured clockwise).
[0283] As an example, in the case of inversion, it can include candidates such as 10a, 10e, and 10f. When 10a has no inversion, 10e and 10f can be examples where left-right inversion and up-down inversion are applied respectively.
[0284] The above examples illustrate the cases of settings for rotation with a certain interval and settings for some inversions, but they are only examples for image reconstruction and are not limited to the above cases. Other differences in intervals can include examples of other inversion operations, etc. This can be determined according to the encoding / decoding settings.
[0285] Or, it can include integrated information (e.g., convert_com_flag) generated by mixing the method of performing reconstruction and the mode information thereby. In this case, the reconstruction-related information can be composed of information in which the method of performing reconstruction and the mode information are mixed.
[0286] For example, the integrated information can include candidates such as 10a to 10f. This can be an example where rotations of 0 degrees, 90 degrees, 180 degrees, 270 degrees, horizontal flipping, and vertical flipping are applied based on 10a.
[0287] Alternatively, the integrated information can include candidates such as 10a to 10h. This can be an example where rotations of 0 degrees, 90 degrees, 180 degrees, 270 degrees, horizontal flipping, vertical flipping, horizontal flipping after 90 - degree rotation (or 90 - degree rotation after horizontal flipping), vertical flipping after 90 - degree rotation (or 90 - degree rotation after vertical flipping) are applied, or it can be an example where rotations of 0 degrees, 90 degrees, 180 degrees, 270 degrees, horizontal flipping, horizontal flipping after 180 - degree rotation (180 - degree rotation after horizontal flipping), horizontal flipping after 90 - degree rotation (90 - degree rotation after horizontal flipping), horizontal flipping after 270 - degree rotation (270 - degree rotation after horizontal flipping) are applied.
[0288] The candidate group can be configured to include modes where rotation is applied, modes where inversion is applied, and modes where rotation and inversion are mixed. The mixed - configured mode simply includes mode information in the method of reconstruction, and can include modes generated by mixing mode information of each method. At this time, it can include modes generated by mixing at least one mode of some methods (for example, rotation) and at least one mode of some methods (for example, inversion). The above example includes cases where one mode of some methods and multiple modes of some methods are mixed (in this example, 90 - degree rotation + multiple inversions / horizontal flipping + multiple rotations). The mixed - configured information can be configured to include, as a candidate group, the case where no reconstruction is applied {in this example, 10a}, and when no reconstruction is applied, it can be included as the first candidate group (for example, assigned an index of 0).
[0289] Alternatively, it can include mode information according to a method of performing a predetermined reconstruction. In this case, the reconstruction - related information can be composed of mode information according to a method of performing a predetermined reconstruction. That is, information about the method of performing reconstruction can be omitted and can be composed of one syntactic element related to mode information.
[0290] For example, it can be configured to include candidates such as 10a to 10d related to rotation. Or, it can be configured to include candidates such as 10a, 10e, and 10f related to inversion.
[0291] The sizes of the images before and after the image reconstruction process may be the same, or at least one length may be different. This can be determined according to the encoding / decoding settings. The image reconstruction process is a process of rearranging the pixels in the image (in this example, in the reverse process of image reconstruction, the reverse process of pixel rearrangement is performed. This can be derived inversely from the pixel rearrangement process), and the position of at least one pixel can be changed. The rearrangement of the pixels can be performed according to rules based on the image reconstruction type information.
[0292] At this time, the pixel rearrangement process can be affected by the size and shape of the image (for example, square or rectangle), etc. Specifically, the horizontal width and vertical height of the image before the reconstruction process and the horizontal width and vertical height of the image after the reconstruction process can act as variables in the pixel rearrangement process.
[0293] For example, at least one ratio information of the ratio of the horizontal width of the image before the reconstruction process to the horizontal width of the image after the reconstruction process, the ratio of the horizontal width of the image before the reconstruction process to the vertical height of the image after the reconstruction process, the ratio of the vertical height of the image before the reconstruction process to the horizontal width of the image after the reconstruction process, and the ratio of the vertical height of the image before the reconstruction process to the vertical height of the image after the reconstruction process (for example, former / latter or latter / former, etc.) can act as variables in the pixel rearrangement process.
[0294] In the above example, when the sizes of the images before and after the reconstruction process are the same, the ratio of the horizontal width to the vertical height of the image can act as a variable in the pixel rearrangement process. Also, when the shape of the image is square, the ratio of the length of the image before the image reconstruction process to the length of the image after the reconstruction process can act as a variable in the pixel rearrangement process.
[0295] In the image reconstruction execution stage, image reconstruction can be performed based on the identified reconstruction information. That is, image reconstruction can be performed based on information such as the reconstruction type and the reconstruction mode, and encoding / decoding can be performed based on the obtained reconstructed image.
[0296] Next, an example of performing image reconstruction by an encoding / decoding apparatus according to an embodiment of the present invention will be shown.
[0297] Before the start of encoding, a reconstruction process for the input image can be performed. After performing the reconstruction using the reconstruction information (for example, the image reconstruction type, the reconstruction mode, etc.), the reconstructed image can be encoded. After the completion of encoding, it can be stored in the memory, and the image encoded data can be recorded in a bitstream and transmitted.
[0298] Before the start of decoding, a reconstruction process can be performed. After performing the reconstruction using the reconstruction information (for example, the image reconstruction type, the reconstruction mode, etc.), the image decoded data can be parsed and decoded. After the completion of decoding, it can be stored in the memory, and after performing the reverse reconstruction process to change it to the image before reconstruction, the image can be output.
[0299] In the encoder, the information generated in the above process is recorded in a bitstream in at least one unit among units such as a sequence, a picture, a slice, and a tile, and in the decoder, the related information is parsed from the bitstream. Also, it may be included in the bitstream in the form of SEI or metadata.
[0300] [Table 1]
[0301] Table 1 shows examples of syntax elements related to the division during image setting. In the examples described later, the explanation will focus on the added syntax elements. Also, the syntax elements in the examples described later are not limited to a specific unit, and can be syntax elements supported in various units such as sequence, picture, slice, tile, etc. Or, they can be syntax elements included in SEI, metadata, etc. Also, the types of syntax elements supported, the order of syntax elements, conditions, etc. in the examples described later are only limited in this example, and can be changed and determined according to the encoding / decoding settings.
[0302] In Table 1, tile_header_enabled_flag means a syntax element for whether to support encoding / decoding settings for a tile. When activated (tile_header_enabled_flag = 1), it can have encoding / decoding settings at the tile unit. When deactivated (tile_header_enabled_flag = 0), it cannot have encoding / decoding settings at the tile unit and can receive the assignment of encoding / decoding settings of the upper unit.
[0303] The tile_coded_flag represents a syntax element indicating whether to perform tile encoding / decoding. When activated (tile_coded_flag = 1), the encoding / decoding of the corresponding tile can be performed. When deactivated (tile_coded_flag = 0), encoding / decoding cannot be performed. Here, not performing encoding means not generating encoded data for the tile (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc., and is applicable to meaningless areas in a partial projection format of a 360-degree image). Not performing decoding means no longer parsing the decoded data for the corresponding tile (in this example, it is assumed that the corresponding area is processed according to a predetermined rule, etc.). Also, no longer parsing the decoded data can mean that there is no encoded data in the corresponding unit, so it is no longer parsed, but it can also mean that even if there is encoded data, it is no longer parsed according to the flag. Depending on whether to perform tile encoding / decoding, header information at the tile unit can be supported.
[0304] Although the above example is described centered around tiles, it is not an example limited to tiles and can be an example applicable to other division units in the present invention. Also, as an example of tile division settings, it is not limited to the above case, and deformation to other examples is also possible.
[0305]
Table 2
[0306] Table 2 shows examples of syntax elements related to reconstruction during image setting.
[0307] Referring to Table 2, convert_enabled_flag means a syntax element regarding whether to perform reconstruction. When activated (convert_enabled_flag = 1), it means encoding / decoding the reconstructed image, and additional reconstruction-related information can be checked. When deactivated (convert_enabled_flag = 0), it means encoding / decoding the existing image.
[0308] convert_type_flag means hybrid information regarding the method and mode information for performing reconstruction. It can be determined to be any one of a plurality of candidate groups regarding the method applying rotation, the method applying inversion, and the method applying a mixture of rotation and inversion.
[0309]
Table 3
[0310] Table 3 shows an example regarding syntax elements related to size adjustment during image setting.
[0311] Referring to Table 3, pic_width_in_samples and pic_height_in_samples mean syntax elements regarding the horizontal and vertical widths of the image, and the size of the image can be checked by said syntax elements.
[0312] img_resizing_enabled_flag means a syntax element regarding whether to adjust the size of the image. When activated (img_resizing_enabled_flag = 1), it means encoding / decoding the image after size adjustment, and additional size-adjustment-related information can be checked. When deactivated (img_resizing_enabled_flag = 0), it means encoding / decoding the existing image. Also, it can be a syntax element meaning size adjustment for intra prediction.
[0313] The resizing_met_flag means a syntax element for the resizing method. When performing resizing using a scale factor (resizing_met_flag = 0), when performing resizing using an offset factor (resizing_met_flag = 1), it can be determined to be any one of a candidate group such as other resizing methods.
[0314] The resizing_mov_flag means a syntax element for the resizing operation. For example, it can be determined to be either an expansion or a reduction.
[0315] width_scale and height_scale mean scale factors related to horizontal resizing and vertical resizing in resizing using a scale factor.
[0316] top_height_offset and bottom_height_offset mean upward and downward offset factors related to horizontal resizing in resizing using an offset factor, and left_width_offset and right_width_offset mean leftward and rightward offset factors related to vertical resizing in resizing using an offset factor.
[0317] Through the resizing-related information and the image size information, the size of the image after resizing can be updated.
[0318] The resizing_type_flag means a syntax element for the data processing method of the area to be resized. Depending on the resizing method and the resizing operation, the number of candidate groups for the data processing method may be the same or different.
[0319] The image setting process applied to the aforementioned image encoding / decoding device may be performed individually, or a plurality of image setting processes may be mixed and performed. In the example described later, the case where a plurality of image setting processes are mixed and performed will be described.
[0320] FIG. 11 is an exemplary diagram showing images before and after an image setting process according to an embodiment of the present invention. Specifically, 11a is an example before performing image reconstruction on the divided image (for example, an image projected by 360-degree image encoding), and 11b shows an example after performing image reconstruction on the divided image (for example, an image packed by 360-degree image encoding). That is, 11a can be understood as an exemplary diagram before performing the image setting process, and 11b can be understood as an exemplary diagram after performing the image setting process.
[0321] The image setting process in this example will be described for the cases of image segmentation (assumed to be tiles in this example) and image reconstruction.
[0322] In the example described below, the case where image reconstruction is performed after image segmentation will be described. However, it is also possible that image segmentation is performed after image reconstruction according to the encoding / decoding settings, and variations to other cases are also possible. Also, the above-described image reconstruction process (including the reverse process) can be applied in the same or similar manner to the reconstruction process of the divided units within the image in this embodiment.
[0323] Image reconstruction may or may not be performed on all divided units within the image, and may be performed on some of the divided units. Therefore, the divided units before reconstruction (for example, a part of P0 to P5) may or may not be the same as the divided units after reconstruction (for example, a part of S0 to S5). The cases regarding the execution of various image reconstructions will be described through the examples described below. Also, for the sake of convenience of explanation, it is assumed that the unit of the image is a picture, the unit of the divided image is a tile, and the divided unit is in a square shape for the explanation.
[0324] As an example, whether to perform image reconstruction can be determined by some units (for example, sps_convert_enabled_flag or SEI, metadata, etc.). Or, whether to perform image reconstruction can be determined by some units (for example, pps_convert_enabled_flag). This is possible when it first occurs in the corresponding unit (in this example, the picture) or when it is activated in a higher-level unit (for example, sps_convert_enabled_flag = 1). Or, whether to perform image reconstruction can be determined by some units (for example, tile_convert_flag[i]. i is the index of the divided unit). This is possible when it first occurs in the corresponding unit (in this example, the tile) or when it is activated in a higher-level unit (for example, pps_convert_enabled_flag = 1). Also, whether to perform the reconstruction of the partial image can be implicitly determined according to the encoding / decoding settings, and thus the related information can be omitted.
[0325] As an example, according to a signal for instructing image reconstruction (for example, pps_convert_enabled_flag), it can be determined whether to perform the reconstruction of the divided units in the image. Specifically, according to the signal, it can be determined whether to perform the reconstruction of all the divided units in the image. At this time, a signal for instructing the reconstruction of one image can be generated for the image.
[0326] As an example, according to a signal for instructing image reconstruction (for example, tile_convert_flag[i]), it can be determined whether to perform the reconstruction of the divided units in the image. Specifically, according to the signal, it can be determined whether to perform the reconstruction of a partial divided unit in the image. At this time, at least one signal for instructing the reconstruction of an image (for example, generated by the number of divided units) can be generated.
[0327] As an example, it is possible to determine whether to perform image reconstruction according to a signal for instructing image reconstruction (for example, pps_convert_enabled_flag), and it is possible to determine whether to perform reconstruction of a divided unit within the image according to a signal for instructing image reconstruction (for example, tile_convert_flag[i]). Specifically, when some signals are activated (for example, pps_convert_enabled_flag = 1), some other signals (for example, tile_convert_flag[i]) can be further checked, and according to the said signal (in this example, tile_convert_flag[i]), it is possible to determine whether to perform reconstruction of a divided unit within the image. At this time, signals for instructing reconstruction of a plurality of images can be generated.
[0328] When a signal for instructing image reconstruction is activated, image reconstruction related information can be generated. Cases regarding various image reconstruction related information will be described in the examples described later.
[0329] As an example, reconstruction information applied to the image can be generated. Specifically, one piece of reconstruction information can be used as the reconstruction information for all divided units within the image.
[0330] As an example, reconstruction information applied to a divided unit within the image can be generated. Specifically, at least one piece of reconstruction information can be used as the reconstruction information for a divided unit within the image. That is, one piece of reconstruction information can be used as the reconstruction information for one divided unit, or one piece of reconstruction information can be used as the reconstruction information for a plurality of divided units.
[0331] The examples described later can be explained by combinations with examples of performing image reconstruction.
[0332] For example, when a signal for instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information that is commonly applied to the divided units within the image can be generated. Or, when a signal for instructing image reconstruction (e.g., pps_convert_enabled_flag) is activated, reconstruction information that is individually applied to the divided units within the image can be generated. Or, when a signal for instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information that is individually applied to the divided units within the image can be generated. Or, when a signal for instructing image reconstruction (e.g., tile_convert_flag[i]) is activated, reconstruction information that is commonly applied to the divided units within the image can be generated.
[0333] In the case of the aforementioned reconstruction information, implicit or explicit processing can be performed according to the encoding / decoding settings. In the case of implicit processing, the reset information can be assigned a predetermined value according to the characteristics, types, etc. of the image.
[0334] P0 to P5 in 11a can correspond to S0 to S5 in 11b, and a reconstruction process can be performed on the divided units. For example, it can be assigned to S0 without performing reconstruction on P0, it can be assigned to S1 by applying a 90-degree rotation to P1, it can be assigned to S2 by applying a 180-degree rotation to P2, it can be assigned to S3 by applying a left-right flip to P3, it can be assigned to S4 by applying a left-right flip after a 90-degree rotation to P4, and it can be assigned to S5 by applying a left-right flip after a 180-degree rotation to P5.
[0335] However, it is not limited to the above-described examples, and various modified examples are also possible. It is possible to perform no reconstruction on the divided units of the image as in the above example, or at least one of the reconstruction methods such as reconstruction with rotation applied, reconstruction with inversion applied, or reconstruction with a mixed application of rotation and inversion.
[0336] When the image reconstruction is applied to the division unit, additional reconstruction processes such as division unit rearrangement can be performed. That is, the image reconstruction process of the present invention can be configured to include the rearrangement within the image of the division unit in addition to rearranging the pixels in the image, and can be expressed by some syntax elements such as those in Table 4 (for example, part_top, part_left, part_width, part_height, etc.). This means that the image division and the image reconstruction process can be understood as being mixed. The above description may be an example possible when the image is divided into a plurality of units.
[0337] P0 to P5 of 11a can correspond to S0 to S5 of 11b, and the reconstruction process can be performed on the division unit. For example, it can be assigned to S0 without performing reconstruction on P0, assigned to S2 without performing reconstruction on P1, assigned to S1 by applying a 90-degree rotation to P2, assigned to S4 by applying a horizontal flip to P3, assigned to S5 by applying a horizontal flip after a 90-degree rotation to P4, and assigned to S3 by applying a 180-degree rotation after a horizontal flip to P5, and is not limited thereto, and examples of various deformations are also possible.
[0338] Also, P_Width and P_Height in FIG. 7 can correspond to P_Width and P_Height in FIG. 11, and P'_Width and P'_Height in FIG. 7 can correspond to P'_Width and P'_Height in FIG. 11. The image size after size adjustment in FIG. 7 is P'_Width×P'_Height, which can be expressed as (P_Width + Exp_L + Exp_R)×(P_Height + Exp_T + Exp_B). The image size after size adjustment in FIG. 11 is P'_Width×P'_Height, which can be expressed as (P_Width + Var0_L + Var1_L + Var2_L + Var0_R + Var1_R + Var2_R)×(P_Height + Var0_T + Var1_T + Var0_B + Var1_B), or (Sub_P0_Width + Sub_P1_Width + Sub_P2_Width + Var0_L + Var1_L + Var2_L + Var0_R + Var1_R + Var2_R)×(Sub_P0_Height + Sub_P1_Height + Var0_T + Var1_T + Var0_B + Var1_B).
[0339] As in the above example, the reconstruction of the image can perform pixel rearrangement within the division unit of the image, can perform rearrangement of the division units within the image, and can perform not only pixel rearrangement within the division unit of the image but also rearrangement of the division units within the image. At this time, after performing pixel rearrangement of the division unit, rearrangement within the image of the division unit can be performed, or after performing rearrangement within the image of the division unit, pixel rearrangement of the division unit can be performed.
[0340] The rearrangement of the division units within the image can be determined whether to be performed according to a signal instructing the reconstruction of the image. Or, a signal for the rearrangement of the division units within the image can be generated. Specifically, the signal can be generated when the signal instructing the reconstruction of the image is activated. Or, implicit or explicit processing can be performed according to the encoding / decoding settings. In the case of implicit, it can be determined according to the characteristics, types, etc. of the image.
[0341] In addition, information regarding the rearrangement of divided units within an image can be processed implicitly or explicitly according to the encoding / decoding settings, and can be determined according to the characteristics, types, etc. of the image. That is, each divided unit can be arranged according to the arrangement information of a predetermined divided unit.
[0342] Next, an example of reconstructing divided units within an image by an encoding / decoding apparatus according to an embodiment of the present invention is shown.
[0343] Before the start of encoding, a division process can be performed on the input image using division information. A reconstruction process can be performed using reconstruction information for the divided units, and the image reconstructed in the divided units can be encoded. After the completion of encoding, it can be stored in a memory, and the image encoding data can be recorded in a bit stream and transmitted.
[0344] Before the start of decoding, a division process can be performed using division information. A reconstruction process is performed using reconstruction information for the divided units, and the image decoding data can be parsed and decoded in the divided units where the reconstruction has been performed. After the completion of decoding, it can be stored in a memory, and after performing the reverse process of reconstructing the divided units, the divided units can be merged into one to output an image.
[0345] FIG. 12 is an exemplary diagram of size adjustment for each divided unit within an image according to an embodiment of the present invention. P0 to P5 in FIG. 12 correspond to P0 to P5 in FIG. 11, and S0 to S5 in FIG. 12 correspond to S0 to S5 in FIG. 11.
[0346] In the examples described below, the case where image size adjustment is performed after image division will be mainly described. However, it is also possible to perform image division after image size adjustment according to the encoding / decoding settings, and variations to other cases are also possible. In addition, the above-described image size adjustment process (including the reverse process) can be applied in the same or similar manner to the size adjustment process of the divided units within the image in this embodiment.
[0347] For example, TL to BR in FIG. 7 can correspond to TL to BR of the divided units SX (S0 to S5) in FIG. 12, S0 and S1 in FIG. 7 can correspond to PX and SX in FIG. 12, P_Width and P_Height in FIG. 7 can correspond to Sub_PX_Width and Sub_PX_Height in FIG. 12, P’_Width and P’_Height in FIG. 7 can correspond to Sub_SX_Width and Sub_SX_Height in FIG. 12, Exp_L, Exp_R, Exp_T, Exp_B in FIG. 7 can correspond to VarX_L, VarX_R, VarX_T, VarX_B in FIG. 12, and other factors can also correspond.
[0348] In the process of adjusting the size of the divided units within the images of 12a to 12f, it can be distinguished from the image size enlargement or reduction in FIGS. 7a and 7b in that there may be settings regarding the enlargement or reduction of the image size in proportion to the number of divided units. Also, there can be differences in having settings commonly applied to the divided units within the image or having settings individually applied to the divided units within the image. The cases of size adjustment in various situations will be described in the examples below, and the size adjustment process can be performed considering the above matters.
[0349] The image size adjustment in the present invention may or may not be performed on all the divided units within the image, and may also be performed on a part of the divided units. The cases regarding various image size adjustments will be described through the examples below. Also, for the sake of convenience in explanation, it is assumed that the size adjustment operation is expansion, the method of size adjustment is the offset factor, the size adjustment directions are the up, down, left, and right directions, the size adjustment direction operates according to the size adjustment information, the unit of the image is the picture, and the unit of the divided image is the tile for the explanation.
[0350] As an example, whether to perform image size adjustment can be determined in some units (for example, sps_img_resizing_enabled_flag, or SEI, metadata, etc.). Or, whether to perform image size adjustment can be determined in some units (for example, pps_img_resizing_enabled_flag). This is possible when it first occurs in the corresponding unit (in this example, the picture), or when it is activated in a higher-level unit (for example, sps_img_resizing_enabled_flag = 1). Or, whether to perform image size adjustment can be determined in some units (for example, tile_resizing_flag[i]. i is the partition unit index). This is possible when it first occurs in the corresponding unit (in this example, the tile), or when it is activated in a higher-level unit. Also, whether to perform size adjustment of the part of the image can be implicitly determined according to the encoding / decoding settings, whereby the related information can be omitted.
[0351] As an example, according to a signal for instructing image size adjustment (for example, pps_img_resizing_enabled_flag), it can be determined whether to perform size adjustment of the partition units in the image. Specifically, according to the signal, it can be determined whether to perform size adjustment of all the partition units in the image. At this time, a signal for instructing size adjustment of one image can be generated.
[0352] As an example, according to a signal for instructing image size adjustment (for example, tile_resizing_flag[i]), it can be determined whether to perform size adjustment of the partition units in the image. Specifically, according to the signal, it can be determined whether to perform size adjustment of a part of the partition units in the image. At this time, at least one signal for instructing size adjustment of an image (for example, generated as many as the number of partition units) can be generated.
[0353] As an example, according to a signal for instructing image size adjustment (for example, pps_img_resizing_enabled_flag), it can be determined whether to perform image size adjustment, and according to a signal for instructing image size adjustment (for example, tile_resizing_flag[i]), it can be determined whether to perform size adjustment of a divided unit within the image. Specifically, when some signals are activated (for example, pps_img_resizing_enabled_flag = 1), some other signals (for example, tile_resizing_flag[i]) can be further checked, and according to the said signal (in this example, tile_resizing_flag[i]), it can be determined whether to perform size adjustment of some divided units within the image. At this time, signals for instructing size adjustment of a plurality of images can be generated.
[0354] When a signal for instructing image size adjustment is activated, image size adjustment related information can be generated. Cases regarding various image size adjustment related information will be described in the examples below.
[0355] As an example, size adjustment information applied to the image can be generated. Specifically, one size adjustment information or a set of size adjustment information can be used as size adjustment information for all divided units within the image. For example, one size adjustment information (or a size adjustment value applied commonly in the up, down, left, and right directions of the divided unit within the image, or all applied to the size adjustment directions supported or allowed by the divided unit. One piece of information in this example) applied commonly in the up, down, left, and right directions of the divided unit within the image, or a set of size adjustment information applied to the up, down, left, and right directions respectively (or the number of size adjustment directions supported or allowed by the divided unit. At most 4 pieces of information in this example) can be generated.
[0356] As an example, size adjustment information applicable to a divided unit within an image can be generated. Specifically, at least one size adjustment information or size adjustment information set can be used as the size adjustment information for a divided unit within the image. That is, one size adjustment information or size adjustment information set can be used as the size adjustment information for one divided unit, or can be used as the size adjustment information for a plurality of divided units. For example, one size adjustment information commonly applied in the up, down, left, and right directions of one divided unit within the image, or one size adjustment information set respectively applied in the up, down, left, and right directions can be generated. Or, one size adjustment information commonly applied in the up, down, left, and right directions of a plurality of divided units within the image, or one size adjustment information set respectively applied in the up, down, left, and right directions can be generated. The composition of the size adjustment set means size adjustment value information for at least one size adjustment direction.
[0357] In summary, size adjustment information commonly applied to the divided units within the image can be generated. Or, size adjustment information individually applied to the divided units within the image can be generated. The examples described later can be explained by combinations with examples of performing image size adjustment.
[0358] For example, when a signal for instructing image size adjustment (e.g., pps_img_resizing_enabled_flag) is activated, size adjustment information commonly applied to the divided units within the image can be generated. Or, when a signal for instructing image size adjustment (e.g., pps_img_resizing_enabled_flag) is activated, size adjustment information individually applied to the divided units within the image can be generated. Or, when a signal for instructing image size adjustment (e.g., tile_resizing_flag[i]) is activated, size adjustment information individually applied to the divided units within the image can be generated. Or, when a signal for instructing image size adjustment (e.g., tile_resizing_flag[i]) is activated, size adjustment information commonly applied to the divided units within the image can be generated.
[0359] The image size adjustment direction, size adjustment information, etc. can be processed implicitly or explicitly according to the encoding / decoding settings. In the case of implicit processing, the size adjustment information can be assigned to a predetermined value according to the characteristics, types, etc. of the image.
[0360] The size adjustment direction in the size adjustment process of the present invention described above is at least one of the up, down, left, and right directions, and it has been explained that the size adjustment direction and the size adjustment information can be processed implicitly or explicitly. That is, for some directions, the size adjustment value (including 0, that is, no adjustment) is determined in advance implicitly, and for some directions, the size adjustment value (including 0, that is, no adjustment) is assigned explicitly.
[0361] Even for the divided units within the image, the size adjustment direction and the size adjustment information can be set to be capable of implicit or explicit processing, and this can be applied to the divided units within the image. For example, a setting (in this example, only the divided unit occurs) applied to one divided unit within the image can occur, or a setting applied to a plurality of divided units within the image can occur, or a setting applied to all the divided units within the image (in this example, one setting occurs) can occur, and at least one setting can occur in the image (for example, the number of divided units of settings can occur from one setting). The setting information applied to the divided units within the image can be collected to define one set of settings.
[0362] FIG. 13 is an exemplary diagram of the size adjustment or set of settings for the divided units within the image.
[0363] Specifically, various examples of the implicit or explicit processing of the size adjustment direction and the size adjustment information of the divided units within the image are shown. In the examples described later, for the sake of convenience of explanation, the implicit processing is described assuming that the size adjustment value of some size adjustment directions is 0.
[0364] When the boundary of the division unit coincides with the boundary of the image as shown in 13a (in this example, the thick solid line), explicit processing for size adjustment can be performed, and when they do not coincide (the thin solid line), implicit processing can be performed. For example, P0 can be adjusted in the upward and leftward directions (a2, a0), P1 in the upward direction (a2), P2 in the upward and rightward directions (a2, a1), P3 in the downward and leftward directions (a3, a0), P4 in the downward direction (a3), P5 in the downward and rightward directions (a3, a1), and size adjustment is not possible in other directions.
[0365] As shown in 13b, for some directions of the division unit (in this example, up and down), explicit processing for size adjustment can be performed, and for some directions of the division unit (in this example, left and right), when the boundary of the division unit coincides with the boundary of the image, explicit processing (in this example, the thick solid line) can be performed, and when they do not coincide (in this example, the thin solid line), implicit processing can be performed. For example, P0 can be adjusted in the upward, downward, and leftward directions (b2, b3, b0), P1 in the upward and downward directions (b2, b3), P2 in the upward, downward, and rightward directions (b2, b3, b1), P3 in the upward, downward, and leftward directions (b3, b4, b0), P4 in the upward and downward directions (b3, b4), P5 in the upward, downward, and rightward directions (b3, b4, b1), and size adjustment is not possible in other directions.
[0366] As shown in 13c, for some directions of the division unit (in this example, left and right), explicit processing for size adjustment can be performed, and for some directions of the division unit (in this example, up and down), when the boundary of the division unit coincides with the boundary of the image (in this example, the thick solid line), explicit processing can be performed, and when they do not coincide (in this example, the thin solid line), implicit processing can be performed. For example, P0 can be adjusted in the upward, leftward, and rightward directions (c4, c0, c1), P1 in the upward, leftward, and rightward directions (c4, c1, c2), P2 in the upward, leftward, and rightward directions (c4, c2, c3), P3 in the downward, leftward, and rightward directions (c5, c0, c1), P4 in the downward, leftward, and rightward directions (c5, c1, c2), P5 in the downward, leftward, and rightward directions (c5, c2, c3), and size adjustment is not possible in other directions.
[0367] Settings related to image size adjustment, as in the above example, can have various cases. Multiple setting sets can be supported and explicit setting set selection information can be generated, or a predetermined setting set can be implicitly determined according to encoding / decoding settings (e.g., image characteristics, types, etc.).
[0368] FIG. 14 is an exemplary diagram showing the image size adjustment process and the size adjustment process of the divided units within the image together.
[0369] Referring to FIG. 14, the image size adjustment process and the reverse process can proceed in the directions of e and f, and the size adjustment process and the reverse process of the divided units within the image can proceed in the directions of d and g. That is, the image size adjustment process can be performed on the image, the size adjustment of the divided units within the image can be performed, and the order of the size adjustment process is not fixed. This means that multiple size adjustment processes are possible.
[0370] In summary, the image size adjustment process can be classified into image size adjustment (or image size adjustment before division) and size adjustment of the divided units within the image (or image size adjustment after division). It is not necessary to perform both the image size adjustment and the size adjustment of the divided units within the image. Either one of them can be performed, or both can be performed. This can be determined according to encoding / decoding settings (e.g., image characteristics, types, etc.).
[0371] When performing multiple size adjustment processes in the above example, the image size adjustment can be performed in at least one of the up, down, left, and right directions of the image, and the size adjustment of at least one of the divided units within the image can be performed. At this time, the size adjustment can be performed in at least one of the up, down, left, and right directions of the divided unit for which the size adjustment is performed.
[0372] Referring to FIG. 14, the size of the image A before size adjustment can be defined as P_Width×P_Height, the size of the image after the first size adjustment (or the image before the second size adjustment, B) can be defined as P’_Width×P’_Height, and the size of the image after the second size adjustment (or the image after the final size adjustment, C) can be defined as P’’_Width×P’’_Height. The image A before size adjustment means an image without any size adjustment, the image B after the first size adjustment means an image with some size adjustments, and the image C after the second size adjustment means an image with all size adjustments. For example, the image B after the first size adjustment means an image in which the size of the division units within the image has been adjusted as shown in FIGS. 13a to 13c, and the image C after the second size adjustment can mean an image in which the entire image B that has been adjusted in size in the first stage has been adjusted in size as shown in FIG. 7a, and the reverse case is also possible. It is not limited to the above examples, and various modified examples are possible.
[0373] In the size of the image B after the first size adjustment, P’_Width can be obtained via at least one size adjustment value in the left or right direction that can be adjusted horizontally with respect to P_Width, and P’_Height can be obtained via at least one size adjustment value in the up or down direction that can be adjusted vertically with respect to P_Height. At this time, the size adjustment value can be a size adjustment value that occurs in division units.
[0374] In the size of the image C after the second size adjustment, P’’_Width can be obtained via at least one size adjustment value in the left or right direction that can be adjusted horizontally with respect to P’_Width, and P’’_Height can be obtained via at least one size adjustment value in the up or down direction that can be adjusted vertically with respect to P’_Height. At this time, the size adjustment value can be a size adjustment value that occurs from the image.
[0375] In summary, the size of the image after size adjustment can be obtained via at least one size adjustment value and the size of the image before size adjustment.
[0376] Information regarding a data processing method can be generated in the area where the size of the image is adjusted. Through the examples described below, cases regarding various data processing methods will be explained. For the case of the data processing method generated in the reverse process of size adjustment, the same or similar application as in the case of the size adjustment process is possible, and the data processing methods in the size adjustment process and the reverse process of size adjustment can be explained through various combinations described below.
[0377] As an example, a data processing method applied to an image can be generated. Specifically, one data processing method or a set of data processing methods can be used as the data processing method for all the divided units in the image (assuming that all the divided units are adjusted in size in this example). For example, one data processing method commonly applied in the up, down, left, and right directions of the divided units in the image (or a data processing method applied to all the size adjustment directions supported or allowed by the divided units, etc., which is one piece of information in this example) or a set of data processing methods respectively applied in the up, down, left, and right directions (or the number of size adjustment directions supported or allowed by the divided units, which is a maximum of 4 pieces of information in this example) can be generated.
[0378] As an example, a data processing method applied to the divided units in the image can be generated. Specifically, at least one data processing method or a set of data processing methods can be used as the data processing method for some of the divided units in the image (assuming the divided units to be adjusted in size in this example). That is, one data processing method or a set of data processing methods can be used as the data processing method for one divided unit or for multiple divided units. For example, one data processing method commonly applied in the up, down, left, and right directions of one divided unit in the image or a set of data processing methods respectively applied in the up, down, left, and right directions can be generated. Or, one data processing method commonly applied in the up, down, left, and right directions of multiple divided units in the image or a set of information on one data processing method respectively applied in the up, down, left, and right directions can be generated. The composition of the set of data processing methods means the data processing method for at least one size adjustment direction.
[0379] In summary, a data processing method that is commonly applied to the divided units within an image can be used. Alternatively, a data processing method that is individually applied to the divided units within an image can be used. The data processing method can use a predetermined method. The predetermined data processing method can include at least one method. This applies in implicit cases, and selection information regarding the data processing method can be explicitly generated. This can be determined according to the encoding / decoding settings (e.g., image characteristics, types, etc.).
[0380] That is, a data processing method that is commonly applied to the divided units within an image can be used, using a predetermined method or selecting any one of a plurality of data processing methods. Alternatively, a data processing method that is individually applied to the divided units within an image can be used, using a predetermined method according to the divided unit or selecting any one of a plurality of data processing methods.
[0381] The following example describes some cases regarding the size adjustment (assumed to be expansion in this example) of the divided units within an image. (In this example, the size adjustment area is filled using a part of the image data.)
[0382] For some regions TL~BR of some units (e.g., S0 to S5 in FIGS. 12a to 12f), size adjustment can be performed using the data of some regions tl~br of some units (P0 to P5 in FIGS. 12a to 12f). At this time, the said some units can be the same (e.g., S0 and P0) or different regions (e.g., S0 and P1). That is, the region TL to BR to be size-adjusted can be filled using some data tl to br of the said divided unit, and the region to be size-adjusted can be filled using some data of a divided unit different from the said divided unit.
[0383] As an example, the area TL~BR to be resized in the current division unit can be resized using the tl~br data of the current division unit. For example, the TL of S0 can be filled using the tl data of P0, the RC of S1 can be filled using the tr+rc+br data of P1, the BL+BC of S2 can be filled using the bl+bc+br data of P2, and the TL+LC+BL of S3 can be filled using the tl+lc+bl data of P3.
[0384] As an example, the area TL~BR to be resized in the current division unit can be resized using the tl~br data of the division unit that is spatially adjacent to the current division unit. For example, the TL+TC+TR of S4 can be filled using the b1+bc+br data of P1 in the upward direction, the BL+BC of S2 can be filled using the tl+tc+tr data of P5 in the downward direction, the LC+BL of S2 can be filled using the tl+rc+bl data of P1 in the leftward direction, the RC of S3 can be filled using the tl+lc+bl data of P4 in the rightward direction, and the BR of S0 can be filled using the tl data of P4 in the lower left direction.
[0385] As an example, the area TL~BR to be resized in the current division unit can be resized using the tl~br data of the division unit that is not spatially adjacent to the current division unit. For example, the data of the boundary areas (such as left and right, up and down, etc.) at both ends of the image can be obtained. The LC of S3 can be obtained using the tr+rc+br data of S5, the RC of S2 can be obtained using the tl+lc data of S0, the BC of S4 can be obtained using the tc+tr data of S1, and the TC of S1 can be obtained using the bc data of S4.
[0386] Alternatively, the data of a partial area of the image (an area that is not spatially adjacent but is determined to have a high correlation with the area to be resized) can be obtained. The BC of S1 can be obtained using the tl+lc+bl data of S3, the RC of S3 can be obtained using the tl+tc data of S1, and the RC of S5 can be obtained using the bc data of S0.
[0387] Also, in some cases regarding the resizing (assuming reduction in this example) of the division units within the image (in this example, restoring or correcting and removing using partial data of the image), it is as follows.
[0388] A partial area TL to BR of a partial unit (for example, S0 to S5 in FIGS. 12a to 12f) can be used in the restoration or correction process of a partial area tl to br of partial units P0 to P5. At this time, the partial units can be the same (for example, S0 and P0) or different areas (for example, S0 and P2). That is, the area to be size-adjusted can be used and removed for the restoration of partial data of the corresponding divided unit, and the area to be size-adjusted can be used and removed for the restoration of partial data of a divided unit different from the corresponding divided unit. Since detailed examples can be inversely derived from the expansion process, they are omitted.
[0389] The above example is an example applied when there is data highly correlated with the area to be size-adjusted, and the information on the position referred to for size adjustment can be explicitly generated, or implicitly obtained based on a predetermined rule, or a combination of these can be used to confirm relevant information. This can be an example applied when obtaining data from other areas where there is continuity in the encoding of 360-degree images.
[0390] Next, an example of performing size adjustment of a divided unit in an image by an encoding / decoding apparatus according to an embodiment of the present invention is shown.
[0391] Before the start of encoding, a division process for the input image can be performed. A size adjustment process can be performed using size adjustment information for the divided unit, and the image after size adjustment of the divided unit can be encoded. After the completion of encoding, it can be stored in a memory, and the image encoding data can be recorded in a bit stream and transmitted.
[0392] Before the start of decoding, a division process can be performed using division information. A size adjustment process can be performed using size adjustment information for the divided unit, and the image decoding data can be parsed and decoded with the divided unit having been size-adjusted. After the completion of decoding, it can be stored in a memory, and after performing the reverse process of size adjustment of the divided unit, the divided units can be merged into one to output an image.
[0393] In other cases during the above-described image size adjustment process, the changes can be applied as in the above example and are not limited thereto. Changes to other examples are also possible.
[0394] In the above image setting process, it is possible to combine image size adjustment and image reconstruction. Image reconstruction can be executed after image size adjustment, or image size adjustment can be executed after image reconstruction. Also, combinations of image segmentation, image reconstruction, and image size adjustment are possible. After image segmentation, image size adjustment and image reconstruction can be executed, and the order of image setting is not fixed and can be changed, which can be determined according to the encoding / decoding settings. In this example, the image setting process will describe the case where image reconstruction is performed after image segmentation and then image size adjustment is performed. However, other orders are possible according to the encoding / decoding settings, and changes to other cases are also possible.
[0395] For example, it may be performed in the order of segmentation → reconstruction, reconstruction → segmentation, segmentation → size adjustment, size adjustment → segmentation, size adjustment → reconstruction, reconstruction → size adjustment, segmentation → reconstruction → size adjustment, segmentation → size adjustment → reconstruction, size adjustment → segmentation → reconstruction, size adjustment → reconstruction → segmentation, reconstruction → segmentation → size adjustment, reconstruction → size adjustment → segmentation, etc., and combinations with additional image settings are also possible. As described above, the image setting process may be performed sequentially, but all or part of the setting processes can also be performed simultaneously. Also, for some image setting processes, multiple processes can be performed according to the encoding / decoding settings (e.g., image characteristics, types, etc.). Next, examples of various combinations of the image setting process will be shown.
[0396] As an example, P0 to P5 in FIG. 11a can correspond to S0 to S5 in FIG. 11b, and a reconstruction process (in this example, rearrangement of pixels), a size adjustment process (in this example, the same size adjustment for the divided units) can be performed for the divided units. For example, size adjustment using an offset can be applied to P0 to P5 and assigned to S0 to S5. Also, it can be assigned to S0 without reconstructing P0, can be assigned to S2 without reconstructing P1, can be assigned to S1 by applying a 90-degree rotation to P2, can be assigned to S4 by applying a horizontal flip to P3, can be assigned to S5 by applying a horizontal flip after a 90-degree rotation to P4, and can be assigned to S3 by applying a 180-degree rotation after a horizontal flip to P5.
[0397] As an example, P0 to P5 in FIG. 11a can correspond to positions that are the same as or different from each other in S0 to S5 in FIG. 11b, and a reconstruction process (in this example, rearrangement of pixels and divided units), a size adjustment process (in this example, the same size adjustment for the divided units) can be performed for the divided units. For example, size adjustment using a scale can be applied to P0 to P5 and assigned to S0 to S5. Also, it can be assigned to S0 without reconstructing P0, can be assigned to S2 without reconstructing P1, can be assigned to S1 by applying a 90-degree rotation to P2, can be assigned to S4 by applying a horizontal flip to P3, can be assigned to S5 by applying a horizontal flip after a 90-degree rotation to P4, and can be assigned to S3 by applying a 180-degree rotation after a horizontal flip to P5.
[0398] As an example, P0 to P5 in FIG. 11a can correspond to E0 to E5 in FIG. 5e, and a reconstruction process (in this example, rearrangement of pixels and division units), and a size adjustment process (in this example, size adjustment not identical for division units) can be performed for the division units. For example, it can be assigned to E0 without performing size adjustment and reconstruction for P0, it can be assigned to E1 by performing size adjustment using a scale for P1 without performing reconstruction, it can be assigned to E2 without performing size adjustment and by performing reconstruction for P2, it can be assigned to E4 by performing size adjustment using an offset for P3 without performing reconstruction, it can be assigned to E5 without performing size adjustment and by performing reconstruction for P4, and it can be assigned to E3 by performing size adjustment using an offset and by performing reconstruction for P5.
[0399] As in the above example, the absolute or relative position within the image of the division units before and after the image setting process may be maintained or may be changed. This can be determined according to the settings of encoding / decoding (for example, characteristics, types of images, etc.). Also, various combinations of image setting processes are possible, not limited to the above example, and variations to various examples are also possible.
[0400] In the encoder, the information generated in the above process is recorded in a bitstream in at least one unit among units such as sequence, picture, slice, tile, etc., and in the decoder, the related information is parsed from the bitstream. Also, it may be included in the bitstream in the form of SEI or metadata.
[0401] [Table 4]
[0402] Next, examples for syntax elements associated with a plurality of image settings are shown. In the examples described below, the description will focus on the added syntax elements. Also, the syntax elements in the examples described below are not limited to a specific unit, and can be syntax elements supported by various units such as sequence, picture, slice, tile, etc. Or, they can be syntax elements included in SEI, metadata, etc.
[0403] Referring to Table 4, parts_enabled_flag means a syntax element for whether or not to perform splitting of some units. When activated (parts_enabled_flag = 1), it means that encoding / decoding is performed by splitting into a plurality of units, and additional splitting information can be confirmed. When deactivated (parts_enabled_flag = 0), it means encoding / decoding an existing image. This example is centered on rectangular splitting units such as tiles and can have different settings for existing tiles and splitting information.
[0404] num_partitions means a syntax element for the number of splitting units, and the value obtained by adding 1 means the number of splitting units.
[0405] part_top[i] and part_left[i] mean syntax elements for the position information of the splitting unit, and mean the horizontal and vertical start positions of the splitting unit (for example, the position of the upper left end of the splitting unit). part_width[i] and part_height[i] mean syntax elements for the size information of the splitting unit, and mean the horizontal width and vertical height of the splitting unit. At this time, the start position and size information can be set in pixel units or block units. Further, the syntax elements can be syntax elements that can occur in the image reconstruction process, or can be syntax elements that can occur when the image splitting process and the image reconstruction process are mixed.
[0406] part_header_enabled_flag means a syntax element for whether or not to support encoding / decoding settings for the splitting unit. When activated (part_header_enabled_flag = 1), it can have encoding / decoding settings for the splitting unit. When deactivated (part_header_enabled_flag = 0), it cannot have encoding / decoding settings and can receive the assignment of encoding / decoding settings of the upper unit.
[0407] The above example is an example of a syntax element associated with size adjustment and reconstruction in units of division among the image settings described later, but is not limited thereto, and other division units and settings of the present invention can be applied with modifications. This example is described under the assumption that size adjustment and reconstruction are performed after division, but is not limited thereto, and can be applied with modifications depending on other image setting orders and the like. Also, the types of syntax elements, the order of syntax elements, conditions, etc. supported in the examples described later are only limited in this example, and can be changed and determined according to the encoding / decoding settings.
[0408]
Table 5
[0409] Table 5 shows an example of syntax elements related to the reconstruction of division units in image settings.
[0410] Referring to Table 5, part_convert_flag[i] means a syntax element for whether or not to reconstruct a division unit. The syntax element can occur for each division unit, and when activated (part_convert_flag[i]=1), it means encoding / decoding the reconstructed division unit, and additional reconstruction-related information can be confirmed. When deactivated (part_convert_flag[i]=0), it means encoding / decoding the existing division unit. convert_type_flag[i] means mode information related to the reconstruction of the division unit and can be information related to the rearrangement of pixels.
[0411] Also, syntax elements for additional reconstructions such as rearrangement of division units can occur. In this example, rearrangement of division units can also be performed via part_top and part_left, which are the syntax elements related to the above-described image division, or syntax elements related to the rearrangement of division units (for example, index information, etc.) can occur.
[0412]
Table 6
[0413] Table 6 shows an example of syntax elements related to the size adjustment of the division unit in image setting.
[0414] Referring to Table 6, part_resizing_flag[i] means a syntax element for whether to perform image size adjustment of the division unit. The syntax element can occur for each division unit. When activated (part_resizing_flag[i]=1), it means encoding / decoding the division unit after size adjustment, and additional size-related information can be checked. When deactivated (part_resiznig_flag[i]=0), it means encoding / decoding the existing division unit.
[0415] width_scale[i] and height_scale[i] mean the scale factors for horizontal size adjustment and vertical size adjustment in size adjustment using scale factors in the division unit.
[0416] top_height_offset[i] and bottom_height_offset[i] mean the upward and downward offset factors related to size adjustment using offset factors in the division unit, and left_width_offset[i] and right_width_offset[i] mean the leftward and rightward offset factors related to size adjustment using offset factors in the division unit.
[0417] resizing_type_flag[i][j] means a syntax element for the data processing method of the area to be size-adjusted in the division unit. The syntax element means an individual data processing method in the direction of size adjustment. For example, syntax elements for individual data processing methods of areas size-adjusted in the up, down, left, and right directions can occur. This can also be generated based on size adjustment information (which can only occur when size-adjusted in some directions).
[0418] The above-described image setting process may be a process applied according to the characteristics, types, etc. of the image. In the examples described below, without special mention, can the above-described image setting process be applied in the same way, or can a modified application be possible? In the examples described below, the description will focus on cases where it is additional or with modified application in the above-described examples.
[0419] For example, in the case of an image generated through a 360-degree camera {360-degree Video or Omnidirectional Video}, it has characteristics different from those of an image obtained through a general camera and has an encoding environment different from that of general image compression.
[0420] Different from a general image, a 360-degree image has no boundary part with discontinuous characteristics, and the data in all regions can have continuity. Also, in devices such as an HMD, an image is reproduced in front of the eyes through a lens, and high-quality images can be required. When an image is obtained through a stereoscopic camera, the processed image data can increase. For the purpose of providing an efficient encoding environment including the above examples, various image setting processes considering 360-degree images can be performed.
[0421] The 360-degree camera is a camera having a plurality of cameras or a plurality of lenses and sensors, and the camera or lens can handle all directions around an arbitrary central point captured by the camera.
[0422] The 360-degree image can be encoded using various methods. For example, it can be encoded using various image processing algorithms in a three-dimensional space, or it can also be converted into a two-dimensional space and encoded using various image processing algorithms. In the present invention, a method of converting a 360-degree image into a two-dimensional space for encoding / decoding will be mainly described.
[0423] The 360-degree image encoding device according to an embodiment of the present invention can be configured to include all or part of the configuration shown in FIG. 1, and can further include a preprocessing unit that performs preprocessing (Stitching, Projection, Region-wise Packing) on the input image. On the other hand, the 360-degree image decoding device according to an embodiment of the present invention can include all or part of the configuration shown in FIG. 2, and can further include a postprocessing unit that performs postprocessing (Rendering) before being decoded and reproduced as an output image.
[0424] To explain again, after the preprocessing process (Pre-processing) on the input image in the encoder, encoding can be performed and the bitstream for this can be transmitted. The bitstream transmitted from the decoder can be parsed for decoding, and an output image can be generated after going through the postprocessing process (Post-processing). At this time, the bitstream can record and transmit the information generated in the preprocessing process and the information generated in the encoding process, and the decoder can parse this and use it in the decoding process and the postprocessing process.
[0425] Next, the operation method of the 360-degree image encoder will be described in more detail. Since the operation method of the 360-degree image decoder is the reverse operation of the 360-degree image encoder, it can be easily derived by an ordinary technician and a detailed description will be omitted.
[0426] The input image can undergo the stitching and projection processes in a three-dimensional projection structure (Projection Structure) in units of spheres (Spheres). Through the above processes, the image data on the three-dimensional projection structure can be projected onto a two-dimensional image.
[0427] The projected image can be configured to include all or part of the 360-degree content according to the encoding settings. At this time, the position information of the area (or pixel) arranged at the center of the projected image can be implicitly generated as a predetermined value, or the position information can be explicitly generated. Also, when configuring a projected image that includes a partial area of the 360-degree content, the range and position information of the included area can be generated. Further, range information (e.g., vertical width, horizontal width) and position information (e.g., measured based on the upper left side of the image) for the region of interest (ROI) in the projected image can be generated. At this time, a partial area with high importance among the 360-degree content can be set as the region of interest. Although the 360-degree image can view all the content in the up, down, left, and right directions, the user's line of sight can be limited to a part of the image, and this can be considered and set as the region of interest. For efficient encoding, the region of interest can be set to have good quality and resolution, and the other regions can be set to have lower quality and resolution than the region of interest.
[0428] Among 360-degree image transmission methods, in the single stream transmission method (Single Stream), the entire image or viewport image can be transmitted to a user as an individual single bit stream. In the multi stream transmission method (Multi Stream), by transmitting a plurality of entire images with different qualities as multi-bit streams, the image quality can be selected according to the user's environment and communication situation. In the tiled stream transmission method, by transmitting individually encoded partial images in tile units as multi-bit streams, tiles can be selected according to the user's environment and communication situation. Therefore, the 360-degree image encoder can generate and transmit bit streams with two or more qualities, and the 360-degree image decoder can set the region of interest according to the user's line of sight and selectively decode according to the region of interest. That is, the place where the user's line of sight stays through a head tracking or eye tracking system can be set as the region of interest, and only the necessary part can be rendered.
[0429] The projected image can be converted into a packed image through a region-wise packing process. The region-wise packing process can include the step of dividing the projected image into a plurality of regions. At this time, each divided region can be arranged (or rearranged) in the packed image according to the settings of the region-wise packing. Region-wise packing can be performed for the purpose of enhancing spatial continuity when converting a 360-degree image into a two-dimensional image (or a projected image). The size of the image can be reduced through region-wise packing. Also, it can reduce the image quality degradation that occurs during rendering, enable viewport-based projection, and be performed for the purpose of providing other types of projection formats. Region-wise packing may or may not be performed according to the encoding settings, and can be determined based on a signal indicating whether to perform it (for example, regionwise_packing_flag, in the example described later, region-wise packing related information can only be generated when regionwise_packing_flag is activated).
[0430] When region-wise packing is performed, setting information (or mapping information) such as which partial region of the projected image is assigned (or arranged) as a partial region of the packed image can be displayed (or generated). When region-wise packing is not performed, the projected image and the packed image can be the same image.
[0431] In the above, the stitching, projection, and region-wise packing processes were defined as individual processes, but a part (for example, stitching + projection, projection + region-wise packing) or all (for example, stitching + projection + region-wise packing) of the processes can be defined as one process.
[0432] At least one packed image can be generated for the same input image according to settings such as the stitching, projection, and region - specific packing processes. Also, at least one encoded data for the same projected image can be generated according to the settings of the region - specific packing process.
[0433] A tiling process can be performed to divide the packed image. At this time, tiling is a process of dividing and transmitting an image into a plurality of regions, and can be an example of the 360 - degree image transmission method. As described above, tiling can be performed for the purpose of partial decoding in consideration of the user's environment, etc., and can be performed for the purpose of efficient processing of the huge data of the 360 - degree image. For example, when an image is composed of one unit, all of the image can be decoded for decoding of the region of interest, but when an image is composed of a plurality of unit regions, it is efficient to decode only the region of interest. At this time, the division can be performed by dividing into tiles, which are the division units according to the existing encoding method, or by dividing into various division units (square division, blocks, etc.) described in the present invention. Also, the division unit can be a unit for performing independent encoding / decoding. Tiling can be performed based on the projected image or the packed image, or can be performed independently. That is, it can be divided based on the surface boundary of the projected image, the surface boundary of the packed image, the packing settings, etc., and can be divided independently for each division unit. This can affect the generation of division information in the tiling process.
[0434] Next, the projected image or the packed image can be encoded. The encoded data and the information generated in the preprocessing process can be recorded in a bitstream and transmitted to a 360-degree image decoder. The information generated in the preprocessing process may be recorded in the bitstream in the form of SEI or metadata. At this time, the bitstream can include at least one encoded data and at least one preprocessing information that vary some settings of the encoding process or some settings of the preprocessing process, and record them in the bitstream. This can be for the purpose of mixing a plurality of encoded data (encoded data + preprocessing information) according to the user's environment in the decoder to construct a decoded image. Specifically, a plurality of encoded data can be selectively combined to construct a decoded image. Also, the process may be performed separately into two for application in a stereoscopic system, or the process may be performed on an additional depth image.
[0435] FIG. 15 is an exemplary diagram showing a three-dimensional space showing a three-dimensional image and a two-dimensional plane space.
[0436] Generally, for a 360-degree three-dimensional virtual space, 3DoF (Degree of Freedom) is required, and three rotations can be assisted around the X (Pitch), Y (Yaw), and Z (Roll) axes. DoF means the degree of freedom in space, 3DoF means the degree of freedom including rotations around the X, Y, and Z axes as shown in 15a, and 6DoF means the degree of freedom that further allows movement along the X, Y, and Z axes in addition to 3DoF. The image encoding device and the decoding device of the present invention will be mainly described for the case of 3DoF. When assisting 3DoF or more (3DoF+), it can be combined with additional processes or devices not illustrated in the present invention or applied with modifications.
[0437] Referring to 15a, Yaw can have a range from -π (-180 degrees) to π (180 degrees), Pitch can have a range from -π / 2 rad (or -90 degrees) to π / 2 rad (or 90 degrees), and Roll can have a range from -π / 2 rad (or -90 degrees) to π / 2 rad (or 90 degrees). At this time, assuming that ψ and θ are Longitude and Latitude in the map representation of the earth, the (x, y, z) in three-dimensional space can be converted from the (ψ, θ) in two-dimensional space. For example, the coordinates in three-dimensional space can be derived from the two-dimensional space coordinates based on the conversion formula of x = cos(θ)cos(ψ), y = sin(θ), and z = -cos(θ)sin(ψ).
[0438] Also, (ψ, θ) can be converted to (x, y, z). For example, ψ = tan -1 (-Z / X), θ = sin -1 (Y / (X 2 +Y 2 +Z 2 ) 1 / 2 ) based on the conversion formula, the two-dimensional space coordinates can be derived from the three-dimensional space coordinates.
[0439] When the pixels in three-dimensional space are accurately converted to two-dimensional space (for example, the integer unit pixels in two-dimensional space), the pixels in three-dimensional space can be mapped to the pixels in two-dimensional space. When the pixels in three-dimensional space are not accurately converted to two-dimensional space (for example, the fractional unit pixels in two-dimensional space), the two-dimensional pixels can be mapped to the pixels obtained by interpolation. At this time, the interpolation methods that can be used include the Nearest neighbor interpolation method, Bi-linear interpolation method, B-spline interpolation method, Bi-cubic interpolation method, etc. At this time, any one can be selected from a plurality of interpolation candidates, and the relevant information can be explicitly generated, or the interpolation method can be implicitly determined based on a predetermined rule. For example, a predetermined interpolation filter can be used according to the three-dimensional model, projection format, color format, slice / tile type, etc. Also, when explicitly generating interpolation information, information about the filter information (for example, filter coefficients, etc.) may also be included.
[0440] 15b shows an example of conversion from a three-dimensional space to a two-dimensional space (two-dimensional plane coordinate system). (ψ, θ) can be sampled at (i, j) based on the size (width, height) of the image, where i can range from 0 to P_Width - 1 and j can range from 0 to P_Height - 1.
[0441] (ψ, θ) can be the central point {or reference point, the point denoted by C in FIG. 15, with coordinates (ψ, θ) = (0, 0)} for the 360-degree image arrangement at the center of the projected image. The setting for the central point can be specified in the three-dimensional space, and the position information relative to the central point can be explicitly generated or determined as an implicitly set value. For example, the central position information in Yaw, the central position information in Pitch, the central position information in Roll, etc. can be generated. When the value for the said information is not specifically specified, each value can be assumed to be 0.
[0442] In the above example, an example of converting the entire 360-degree image from a three-dimensional space to a two-dimensional space was described. However, a partial region of the 360-degree image can be targeted, and the position information (e.g., some positions belonging to the region. In this example, the position information relative to the central point), range information, etc. for the partial region can be explicitly generated or follow the implicitly set position and range information. For example, the central position information in Yaw, the central position information in Pitch, the central position information in Roll, the range information in Yaw, the range information in Pitch, the range information in Roll, etc. can be generated. In the case of a partial region, it is at least one region, whereby the position information, range information, etc. of multiple regions can be processed. When the value for the said information is not specifically specified, it can be assumed to be the entire 360-degree image.
[0443] H0 to H6 and W0 to W5 in 15a respectively indicate some latitudes and longitudes in 15b, and the coordinates of 15b can be expressed as (C, j) and (i, C) (C is a longitude or latitude component). Different from general images, when a 360-degree image is converted into a two-dimensional space, distortion or warping of the content in the image can occur. This varies according to the area of the image, and the encoding / decoding settings can be made different for the position of the image or the area partitioned according to the position. When adaptively setting the encoding / decoding based on the encoding / decoding information in the present invention, the position information (for example, x, y components, or the range defined by x and y, etc.) can be included as an example of the encoding / decoding information.
[0444] The description of the three-dimensional and two-dimensional spaces is the content defined to assist in the description of the embodiments in the present invention, and is not limited thereto, and variations of the detailed content or application in other cases are possible.
[0445] As described above, the image obtained by a 360-degree camera can be converted into a two-dimensional space. At this time, the 360-degree image can be mapped using a three-dimensional model, and various three-dimensional models such as a sphere, a cube, a cylinder, a pyramid, and a polyhedron can be used. When converting the 360-degree image mapped based on the model into a two-dimensional space, a projection process according to the projection format based on the model can be performed.
[0446] Figures 16a to 16d are conceptual diagrams for explaining the projection format according to an embodiment of the present invention.
[0447] FIG. 16a shows an ERP (Equi-Rectangular Projection) format in which a 360-degree image is projected onto a two-dimensional plane. FIG. 16b shows a (CMP CubeMap Projection) format in which a 360-degree image is projected onto a cube. FIG. 16c shows an OHP (OctaHedron Projection) format in which a 360-degree image is projected onto an octahedron. FIG. 16d shows an ISP (IcoSahedral Projection) format in which a 360-degree image is projected onto an icosahedron. However, it is not limited to this, and various projection formats can be used. The left sides of FIGS. 16a to 16d show a 3D model, and the right sides show examples converted into a two-dimensional space through the projection process. Depending on the projection format, they have various sizes and shapes, and each shape can be composed of faces or surfaces, and the surfaces can be represented by circles, triangles, quadrilaterals, etc.
[0448] In the present invention, the projection format can be defined by a 3D model, settings of surfaces (e.g., the number of surfaces, the form of the surfaces, the morphological composition of the surfaces, etc.), settings of the projection process, etc. When at least one of the elements of the above definitions is different, it can be regarded as a different projection format. For example, in the case of ERP, it is composed of a spherical model (3D model), one surface (the number of surfaces), and a quadrilateral surface (the pattern of the surface). However, if a part of the settings in the projection process (e.g., the mathematical formula used when converting from a three-dimensional space to a two-dimensional space, that is, the remaining projection settings are the same, and the element that creates a difference in at least one pixel of the projected image during the projection process) is different, it can be classified into different formats such as ERP1 and EPR2. As another example, in the case of CMP, it is composed of a cube model, six surfaces, and a square surface. However, if a part of the settings in the projection process (e.g., the sampling method when converting from a three-dimensional space to a two-dimensional space, etc.) is different, it can be classified into different formats such as CMP1 and CMP2.
[0449] When using multiple projection formats instead of a single already-set projection format, projection format identification information (or projection format information) can be explicitly generated. The projection format identification information can be configured in various ways.
[0450] As an example, index information (e.g., proj_format_flag) can be assigned to multiple projection formats to identify the projection formats. For example, 0 can be assigned to ERP, 1 to CMP, 2 to OHP, 3 to ISP, 4 to ERP1, 5 to CMP1, 6 to OHP1, 7 to ISP1, 8 to CMP compact, 9 to OHP compact, 10 to ISP compact, and 11 or higher to other formats.
[0451] As an example, the projection format can be identified from at least one element information that constitutes the projection format. At this time, the element information that constitutes the projection format may include 3D model information (e.g., 3d_model_flag. 0 represents a sphere, 1 represents a cube, 2 represents a cylinder, 3 represents a pyramid, 4 represents polyhedron 1, 5 represents polyhedron 2, etc.), the number of surface information (e.g., num_face_flag. Starting from 1 and increasing by 1 each time, or assigning the number of surfaces generated in the projection format as index information, 0 represents 1, 1 represents 3, 2 represents 6, 3 represents 8, 4 represents 20, etc.), the form information of the surface (e.g., shape_face_flag. 0 represents a quadrilateral, 1 represents a circle, 2 represents a triangle, 3 represents a quadrilateral + circle, 4 represents a quadrilateral + triangle, etc.), projection process setting information (e.g., 3d_2d_convert_idx, etc.).
[0452] As an example, the projection format can be identified by the projection format index information and the element information that constitutes the projection format. For example, the projection format index information can be assigned 0 for ERP, 1 for CMP, 2 for OHP, 3 for ISP, and 4 or higher for other formats. Together with the element information that constitutes the projection format (in this example, the projection process setting information), the projection format (for example, ERP, ERP1, CMP, CMP1, OHP, OHP1, ISP, ISP1, etc.) can be identified. Or, together with the element information that constitutes the projection format (in this example, whether it is regional packing or not), the projection format (for example, ERP, CMP, CMP compact, OHP, OHP compact, ISP, ISP compact, etc.) can be identified.
[0453] In summary, the projection format can be identified by the projection format index information, can be identified by at least one projection format element information, and can be identified by the projection format index information and at least one projection format element information. This can be defined according to the encoding / decoding settings. In the present invention, the case of being identified by the projection format index will be assumed for explanation. Also, in this example, the case of the projection format represented by a surface having the same size and shape will be mainly explained, but a configuration in which the sizes and shapes of each surface are not the same is also possible. Also, the configuration of each surface may be the same as or different from FIGS. 16a to 16d. The numbers of each surface are used as symbols for identifying each surface and are not limited to a specific order. For the convenience of explanation, in the examples described later, based on the projection image, ERP is a projection format of one surface + a quadrilateral, CMP is a projection format of six surfaces + a quadrilateral, OHP is a projection format of eight surfaces + a triangle, and ISP is a projection format of twenty surfaces + a triangle. The case where the surfaces have the same size and shape will be assumed for explanation, but the same or similar application is also possible for other settings.
[0454] As shown in FIGS. 16a to 16d, the projection format can be divided into one surface (e.g., ERP) or multiple surfaces (e.g., CMP, OHP, ISP, etc.). Also, each surface can be divided into shapes such as rectangles and triangles. The division can be an example of the type, characteristics, etc. of the image in the present invention when it is made different from the encoding / decoding settings according to the projection format. For example, the type of the image can be a 360-degree image, and the characteristics of the image can be any of the above divisions (e.g., each projection format, a projection format of one surface or multiple surfaces, a projection format where the surface is a rectangle or not a rectangle, etc.).
[0455] The two-dimensional plane coordinate system {e.g., (i, j)} can be defined on each surface of the two-dimensional projection image, and the characteristics of the coordinate system can vary according to the projection format, the position of each surface, etc. In the case of ERP, there can be one two-dimensional plane coordinate system, and for other projection formats, multiple two-dimensional plane coordinate systems can be possessed according to the number of surfaces. At this time, the coordinate system can be expressed as (k, i, j), where k can be the index information of each surface.
[0456] FIG. 17 is a conceptual diagram showing that the projection format according to an embodiment of the present invention is included in a rectangular image.
[0457] That is, FIGS. 17a to 17c can be understood as realizing the projection formats of FIGS. 16b to 16d as a rectangular image.
[0458] Referring to FIGS. 17a to 17c, for the encoding / decoding of the 360-degree image, each image format can be configured in a rectangular shape. In the case of ERP, it can be directly used in one coordinate system, but in the case of other projection formats, the coordinate systems of each surface can be integrated into one coordinate system, and the detailed description thereof is omitted.
[0459] Referring to FIGS. 17a through 17c, it can be confirmed that in the process of forming a rectangular image, regions filled with meaningless data such as blanks and backgrounds are generated. That is, it can be composed of a region containing actual data (in this example, the surface. Active Area) and a meaningless region filled to form a rectangular image (in this example, assumed to be filled with arbitrary pixel values. Inactive Area). This may result in a performance degradation due to not only the encoding / decoding of actual image data but also an increase in the amount of encoded data caused by an increase in the size of the image due to the meaningless region.
[0460] Therefore, a process for excluding meaningless regions and forming an image with a region containing actual data can be further performed.
[0461] FIG. 18 is a conceptual diagram of a method for converting a projection format according to an embodiment of the present invention into a rectangular shape, the method of rearranging the surface so as to exclude meaningless regions.
[0462] Referring to FIGS. 18a through 18c, an example of rearranging FIGS. 17a through 17c can be confirmed, and such a process can be defined as a region-by-region packing process (such as CMP compact, OHP compact, ISP compact, etc.). At this time, not only the rearrangement of the surface itself but also the surface can be divided and rearranged (such as OHP compact, ISP compact, etc.). This can be done for the purpose of improving the encoding performance not only by removing meaningless regions but also through an efficient arrangement of the surface. For example, when arranging the images to have continuity between surfaces (such as B2 - B3 - B1, B5 - B0 - B4 in FIG. 18a), the encoding performance can be improved by improving the prediction accuracy during encoding. Here, the region-by-region packing according to the projection format is only an example in the present invention and is not limited thereto.
[0463] FIG. 19 is a conceptual diagram showing the process of performing region-by-region packing by converting a CMP projection format according to an embodiment of the present invention into a rectangular image.
[0464] Referring to FIGS. 19a to 19c, the CMP projection format can be arranged as 6×1, 3×2, 2×3, 1×6. Also, when size adjustment is performed on some surfaces, it can be arranged as shown in FIGS. 19d to 19e. Although FIGS. 19a to 19e take CMP as an example, it is not limited to CMP and can be applied to other projection formats. The surface arrangement of the image obtained through the regional packing can follow a predetermined rule according to the projection format, or information regarding the arrangement can be explicitly generated.
[0465] The 360-degree image encoding / decoding apparatus according to an embodiment of the present invention can be configured to include all or part of the image encoding / decoding apparatus shown in FIGS. 1 and 2. In particular, a format conversion unit and a format inverse conversion unit for converting and inversely converting the projection format can be further included in the image encoding apparatus and the image decoding apparatus, respectively. That is, in the image encoding apparatus of FIG. 1, the input image can be encoded through the format conversion unit, and in the image decoding apparatus of FIG. 2, after the bit stream is decoded, the output image can be generated through the format inverse conversion unit. Hereinafter, the encoder for the above process (in this example, "input image" to "encoding") will be mainly described, and the process in the decoder can be derived conversely from the encoder. Also, descriptions overlapping with the above will be omitted.
[0466] Next, the input image will be described on the premise that it is the same image as the two-dimensional projection image or the packing image obtained by performing the preprocessing process in the above-described 360-degree encoding apparatus. That is, the input image can be an image obtained by performing a projection process or a regional packing process according to some projection formats. The projection format already applied to the input image can be any of various projection formats, and may be regarded as a common format or may be called the first format.
[0467] The format conversion unit can perform conversion to other projection formats other than the first format. At this time, the format to be converted can be called the second format. For example, if ERP is set as the first format, it can be converted to the second format (for example, ERP2, CMP, OHP, ISP, etc.). At this time, ERP2 may be an EPR format that has the same conditions such as the same 3D model and surface configuration but has some different settings. Or, it may be the same format with the same projection format settings (for example, ERP = ERP2), and there may be cases where the image size or resolution is different. Or, a part of the image setting process described later may be applied. For the sake of convenience of explanation, examples as described above have been given, but the first format and the second format are one of various projection formats, not limited to the above examples, and changes to other cases are also possible.
[0468] During the conversion process between formats, due to the characteristics of different coordinate systems between projection formats, the pixels (integer pixels) of the converted image may be obtained not only from the integer unit pixels in the pre-conversion image but also from fractional unit pixels, so interpolation can be performed. At this time, the interpolation filter used can be the same or similar filter as described above. The interpolation filter can select any one from a plurality of interpolation filter candidates, and relevant information can be explicitly generated or implicitly determined according to the already set rules. For example, a predetermined interpolation filter can be used according to the projection format, color format, slice / tile type, etc. Also, when explicitly sending an interpolation filter, information about the filter information (for example, filter coefficients, etc.) may also be included.
[0469] The projection format in the format conversion unit may be defined including region-by-region packing, etc. That is, a projection and region-by-region packing process may be performed during the format conversion process. Or, a process such as region-by-region packing may be performed after the format conversion and before the encoding.
[0470] In the symbolizer, the information generated in the above process is recorded in a bitstream in at least one unit among units such as sequences, pictures, slices, and tiles, and in the decoder, the relevant information is parsed from the bitstream. Also, it may be included in the bitstream in the form of SEI or metadata.
[0471] Next, an image setting process applied to a 360-degree image encoding / decoding apparatus according to an embodiment of the present invention will be described. The image setting process in the present invention can be applied not only to a general encoding / decoding process, but also to a preprocessing process, a postprocessing process, a format conversion process, a format inverse conversion process, etc. in a 360-degree image encoding / decoding apparatus. The image setting process to be described later will be described centering on the 360-degree image encoding apparatus, and can be described including the content in the above-described image setting. Duplicate descriptions in the above-described image setting process will be omitted. Also, the examples to be described later will be described centering on the image setting process, and the inverse image setting process can be induced inversely from the image setting process, and in some cases, can be confirmed through various embodiments of the present invention described above.
[0472] The image setting process in the present invention may be performed in the projection stage of the 360-degree image, may be performed in the regional packing stage, may be performed in the format conversion stage, or may be performed in other stages.
[0473] FIG. 20 is a conceptual diagram of 360-degree image segmentation according to an embodiment of the present invention. In FIG. 20, the case of an image projected by ERP will be described by way of assumption.
[0474] 20a shows an image projected by ERP, and can be divided using various methods. In this example, the description will be centered on slices and tiles, and it is assumed that W0 to W2 and H0, H1 are the division boundary lines of slices or tiles, and it is assumed that they follow the raster scan order. The examples to be described later will be described centering on slices and tiles, but are not limited thereto, and other division methods can be applied.
[0475] For example, it can be divided in slice units and can have division boundaries of H0 and H1. Or, it can be divided in tile units and can have division boundaries of W0 to W2 and H0 and H1.
[0476] 20b shows an example in which the image projected by ERP is divided into tiles {assuming the same tile division boundaries (all of W0 to W2 and H0 and H1 are activated) as in Fig. 20a}. Assuming that the P region is the entire image and the V region is the region where the user's line of sight stays or the viewport, there can be various methods to provide the image corresponding to the viewport. For example, the entire image (e.g., tiles a to l) can be decoded to obtain the region corresponding to the viewport. At this time, the entire image can be decoded, and if it is divided, tiles a to l (in this example, the A + B region) can be decoded. Or, by decoding the region belonging to the viewport, the region corresponding to the viewport can be obtained. At this time, if it is divided, by decoding tiles f, g, j, and k (in this example, the B region), the region corresponding to the viewport can be obtained from the restored image. The former case is called overall decoding (or Viewport Independent Coding), and the latter case is called partial decoding (or Viewport Dependent Coding). The latter case is an example that can occur in a 360-degree image with a large amount of data, and since the divided region can be obtained flexibly, the tile-based division method can be used more frequently than the slice-based division. In the case of partial decoding, since it is not possible to know where the viewport is generated, the referability of the division unit can be restricted spatially or temporally (impliedly processed in this example), and encoding / decoding can be performed considering this. The examples described later will mainly explain the case of overall decoding, but for the purpose of preparing for the case of partial decoding, the division of the 360-degree image will be explained centering on tiles (or the square division method of the present invention). The content of the examples described later can be similarly applied or modified to other division units.
[0477] FIG. 21 is an exemplary diagram of 360-degree image segmentation and image reconstruction according to an embodiment of the present invention. In FIG. 21, the case of an image projected by CMP will be described by way of assumption.
[0478] 21a shows an image projected by CMP, and can be divided using various methods. W0 to W2 and H0, H1 are assumed to be the division boundary lines of the surface, slice, and tile, and it is assumed that they follow the raster scan order.
[0479] For example, it can be divided in units of slices and can have the division boundaries of H0 and H1. Or, it can be divided in units of tiles and can have the division boundaries of W0 to W2 and H0, H1. Or, it can be divided in units of the surface and can have the division boundaries of W0 to W2 and H0, H1. In this example, the surface will be described by assuming that it is part of the division unit.
[0480] At this time, the surface is a division unit (in this example, dependent encoding / decoding) performed for the purpose of classifying or distinguishing regions having different properties (for example, the plane coordinate system of each surface, etc.) within the same image according to the characteristics and types of the image (in this example, 360-degree image, projection format, etc.). Slices and tiles can be division units (in this example, independent encoding / decoding) performed for the purpose of dividing the image according to the user's definition. Also, the surface is a unit divided by a predetermined definition (or derived from projection format information) in the projection process by the projection format, and slices and tiles can be units that explicitly generate division information according to the user's definition and are divided. Also, the surface can have a division shape in the form of a polygon including a quadrilateral according to the projection format, slices can have any division shape that cannot be defined as a quadrilateral or polygon, and tiles can have a square division shape. The setting of the division unit can be content defined and limited for the description of this example.
[0481] In the above example, the surface was described as a division unit classified for the purpose of area division. However, depending on the encoding / decoding settings, it can be a unit that performs independent encoding / decoding in at least one surface unit, or it can have a setting that combines with tiles, slices, etc. to perform independent encoding / decoding. At this time, it can occur that explicit information of tiles, slices, etc. is generated in the combination with tiles, slices, etc., or an implicit case where tiles, slices are combined based on surface information can occur. Or, it may occur that explicit information of tiles, slices is generated based on surface information.
[0482] As a first example, one image division process (in this example, the surface) is performed, and the image division can implicitly omit the division information (obtain the division information from the projection format information). This example is an example for the setting of dependent encoding / decoding, and can be an example corresponding to the case where the referability between surface units is not restricted.
[0483] As a second example, one image division process (in this example, the surface) is performed, and the image division can explicitly generate the division information. This example is an example for the setting of dependent encoding / decoding, and can be an example corresponding to the case where the referability between surface units is not restricted.
[0484] As a third example, a plurality of image division processes (in this example, the surface, tiles) are performed. Some of the image divisions (in this example, the surface) can implicitly omit the division information or explicitly generate it, and some of the image divisions (in this example, the tiles) can explicitly generate the division information. In this example, some of the image division processes (in this example, the surface) precede some of the image division processes (in this example, the tiles).
[0485] As a fourth example, a plurality of image segmentation processes are performed. Some of the image segmentations (in this example, the surface) can implicitly omit the segmentation information or explicitly generate it, and some of the image segmentations (in this example, the tile) can explicitly generate the segmentation information based on some of the image segmentations (in this example, the surface). In this example, some of the image segmentation processes (in this example, the surface) precede some of the image segmentation processes (in this example, the tile). Although it is the same that the segmentation information is explicitly generated in some cases {assuming the case of the second example}, there may be differences in the segmentation information configuration.
[0486] As a fifth example, a plurality of image segmentation processes are performed. Some of the image segmentations (in this example, the surface) can implicitly omit the segmentation information, and some of the image segmentations (in this example, the tile) can implicitly omit the segmentation information based on some of the image segmentations (in this example, the surface). For example, individual surface units can be set in tile units, or a plurality of surface units (in this example, when adjacent surfaces are continuous surfaces, they are grouped, and otherwise they are not grouped. B2 - B3 - B1 and B4 - B0 - B5 in 18a) can be set in tile units. According to the already set rules, the setting of surface units to tile units is possible. This example is an example for independent encoding / decoding settings and may be an example applicable when the referability between surface units is restricted. That is, although it is the same that the segmentation information is implicitly processed in some cases {assuming the case of the first example}, there may be differences in the encoding / decoding settings.
[0487] The above examples are explanations for cases where the segmentation process can be performed in the projection stage, the regional packing stage, the encoding / decoding initial stage, etc., and can be image segmentation processes occurring in other encoding / decoders.
[0488] In 21a, by including an area B that does not contain data in an area A that contains data, a rectangular image can be formed. At this time, the position, size, shape, number, etc. of areas A and B are information that can be confirmed by a projection format or the like, or information that can be confirmed when explicitly generating information for the projected image, and related information can be indicated by the above-described image segmentation information, image reconstruction information, etc. For example, information for a partial area of the projected image (e.g., part_top, part_left, part_width, part_height, part_convert_flag, etc.) can be indicated as shown in Tables 4 and 5, and this is not limited to this example and can be an applicable example in other cases (e.g., other projection formats, other projection settings, etc.).
[0489] The B area can be combined with the A area into one image for encoding / decoding. Or, by performing a division considering the characteristics of each area, different encoding / decoding settings can be set. For example, encoding / decoding for the B area may not be performed via information on whether to perform encoding / decoding (e.g., tile_coded_flag assuming that the division unit is a tile). At this time, the corresponding area can be restored to certain data (in this example, an arbitrary pixel value) according to the set rules. Or, the encoding / decoding settings in the above-described image segmentation process can be made different for the B area and the A area. Or, a region-by-region packing process can be performed to remove the corresponding area.
[0490] 21b shows an example in which an image packed by CMP is divided into tiles, slices, and surfaces. At this time, the packed image is an image in which a surface rearrangement process or a region-by-region packing process has been performed, and can be an image obtained by performing image segmentation and image reconstruction of the present invention.
[0491] In 21b, a rectangular shape can be formed including a region containing data. At this time, the position, size, shape, number, etc. of each region are information that can be confirmed according to the already set settings, or information that can be confirmed when explicitly generating information for the packed image, and related information can be indicated by the above-described image segmentation information, image reconstruction information, etc. For example, information (such as part_top, part_left, part_width, part_height, part_convert_flag, etc.) for a partial region of the packed image as shown in Table 4 and Table 5 can be indicated.
[0492] The packed image can be divided using various division methods. For example, it can be divided in slice units and can have a division boundary of H0. Or, it can be divided in tile units and can have division boundaries of W0, W1, and H0. Or, it can be divided in surface units and can have division boundaries of W0, W1, and H0.
[0493] The image segmentation and image reconstruction processes of the present invention can be performed on the projected image. At this time, the reconstruction process can rearrange not only the pixels within the surface but also the surfaces within the image. This can be an example possible when the image is divided or composed of a plurality of surfaces. The examples described later will mainly explain the case where it is divided into tiles based on surface units.
[0494] SX,Y (S0,0 to S3,2) of 21a can correspond to S’U,V (S’0,0 to S’2,1. In this example, X, Y may be the same as or different from U, V.), and the reconstruction process can be performed on the surface unit. For example, S2,1, S3,1, S0,1, S1,2, S1,1, S1,0 can be assigned (or surface rearrangement) to S’0,0, S’1,0, S’2,0, S’0,1, S’1,1, S’2,1. Also, S2,1, S3,1, S0,1 can be reconstructed without performing reconstruction (or pixel rearrangement), and S1,2, S1,1, S1,0 can be reconstructed by applying a 90-degree rotation, which can be shown as in Figure 21c. The symbols (S1,0, S1,1, S1,2) displayed horizontally in 21c can be images laid horizontally according to the symbols to maintain the continuity of the image.
[0495] The reconstruction of the surface can perform implicit or explicit processing according to the encoding / decoding settings. In the case of implicit processing, it can be performed according to the already set rules considering the type of the image (in this example, a 360-degree image), characteristics (in this example, projection format, etc.).
[0496] For example, there is image continuity (or correlation) between the two surfaces based on the surface boundary for S’0,0 and S’1,0, S’1,0 and S’2,0, S’0,1 and S’1,1, S’1,1 and S’2,1 in 21c, and 21c can be an example configured such that there is continuity between the upper three surfaces and the lower three surfaces. It is divided into multiple surfaces through the projection process from three-dimensional space to two-dimensional space, and reconstruction can be performed for the purpose of enhancing the image continuity between the surfaces for efficient surface reconstruction in the process through the regional packing process. Such surface reconstruction can be already set and processed.
[0497] Alternatively, the reconstruction process can be performed by explicit processing, and reconstruction information about this can be generated.
[0498] For example, when checking information (either implicitly acquired information or explicitly generated information) for an M×N configuration (e.g., in the case of CMP compact, 6×1, 3×2, 2×3, 1×6, etc. Assume a 3×2 configuration in this example) via a regional packing process, after reconstructing the surface according to the M×N configuration, information regarding it can be generated. For example, in the case of in-image rearrangement of the surface, index information (or position information within the image) can be assigned to each surface, and in the case of in-surface pixel rearrangement, mode information for the reconstruction can be assigned.
[0499] The index information can already be defined as shown in 18a to 18c of FIG. 18, and SX,Y or S’U,V in 21a to 21c can represent each surface with position information indicating horizontal and vertical (e.g., S[i][j]) or one position information (e.g., assuming the position information is assigned in raster scan order from the upper left surface of the image. S[i]), and an index for each surface can be assigned to this.
[0500] For example, when assigning an index to position information indicating horizontal and vertical, in the case of FIG. 21c, S’0,0 can be assigned the index of the 2nd surface, S’1,0 can be assigned the index of the 3rd surface, S’2,0 can be assigned the index of the 1st surface, S’0,1 can be assigned the index of the 5th surface, S’1,1 can be assigned the index of the 0th surface, and S’2,1 can be assigned the index of the 4th surface. Or, when assigning an index to one position information, S[0] can be assigned the index of the 2nd surface, S[1] can be assigned the index of the 3rd surface, S[2] can be assigned the index of the 1st surface, S[3] can be assigned the index of the 5th surface, S[4] can be assigned the index of the 0th surface, and S[5] can be assigned the index of the 4th surface. For the sake of convenience in explanation, in the examples described later, S’0,0 to S’2,1 are referred to as a to f. Or, it can also be expressed with position information indicating horizontal and vertical in terms of pixels or blocks based on the upper left side of the image.
[0501] In the case of a packed image obtained through an image reconstruction process (or a regional packing process), depending on the reconstruction settings, the surface scan order may or may not be the same in the image. For example, when one scan order (e.g., raster scan) is applied to 21a, the scan orders of a, b, and c may be the same, and the scan orders of d, e, and f may not be the same. For example, in the case of 21a, a, b, and c, when the scan order follows the order of (0, 0) → (1, 0) → (0, 1) → (1, 1), in the case of d, e, and f, the scan order can follow the order of (1, 0) → (1, 1) → (0, 0) → (0, 1). This can be determined according to the image reconstruction settings, and such settings can also be applied to other projection formats.
[0502] The image segmentation process in 21b can set individual surface units as tiles. For example, surfaces a to f can each be set in tile units. Or, units of multiple surfaces can be set as tiles. For example, surfaces a to c can be set as one tile, and d to f can be set as one tile. The above configuration can be determined based on surface characteristics (e.g., continuity between surfaces, etc.), and different surface tile settings are possible compared to the above examples.
[0503] Next, an example of the segmentation information obtained through multiple image segmentation processes is shown. In this example, the segmentation information for the surface is omitted, and it is assumed that units other than the surface are tiles and the segmentation information is processed in various ways for explanation.
[0504] As a first example, the image segmentation information can be obtained based on surface information and implicitly omitted. For example, individual surfaces can be set as tiles, or multiple surfaces can be set as tiles. At this time, when at least one surface is set as a tile, it can be determined according to a predetermined rule based on surface information (e.g., continuity or correlation, etc.).
[0505] As a second example, the image segmentation information can be explicitly generated regardless of the surface information. For example, when generating the segmentation information based on the number of horizontal rows of tiles (in this example, num_tile_columns) and the number of vertical rows (in this example, num_tile_rows), the segmentation information can be generated by the method in the image segmentation process described above. For example, the possible range that the number of horizontal rows and the number of vertical rows of tiles can have can be from 0 to the width of the image / the width of the block (in this example, the unit obtained from the picture segmentation unit), and from 0 to the height of the image / the height of the block. Also, additional segmentation information (for example, uniform_spacing_flag, etc.) can be generated. At this time, depending on the segmentation setting, it may occur that the boundary of the surface coincides with or does not coincide with the boundary of the segmentation unit.
[0506] As a third example, the image segmentation information can be explicitly generated based on the surface information. For example, when generating the segmentation information based on the number of horizontal rows of tiles and the number of vertical rows, the segmentation information can be generated based on the surface information (in this example, the range of the number of horizontal rows is 0 to 2, and the range of the number of vertical rows is 0, 1. Since the configuration of the surface in the image is 3x2). For example, the possible range that the number of horizontal rows and the number of vertical rows of tiles can have can be from 0 to 2, and from 0 to 1. Also, additional segmentation information (for example, uniform_spacing_flag, etc.) may not be generated. At this time, the boundary of the surface can coincide with the boundary of the segmentation unit.
[0507] In some cases {assuming the cases of the second example and the third example}, it can be defined that the syntax elements of the segmentation information are different, or even if the same syntax elements are used, the settings of the syntax elements (for example, binarization settings, etc. When the range of the candidate group that the syntax element has is limited and small, other binarizations can be used, etc.) can be made different. The above examples have described a part of various configurations of the segmentation information, but it is not limited to this, and it can be understood as an example where different settings are possible depending on whether the segmentation information is generated based on the surface information.
[0508] FIG. 22 is an exemplary diagram showing the division of an image projected by CMP or a packed image into tiles.
[0509] At this time, assume that it has the same tile division boundaries (all of W0 to W2, H0, and H1 are activated) as 21a in FIG. 21, and assume that it has the same tile division boundaries (all of W0, W1, and H0 are activated) as 21b in FIG. 21. When assuming that the P region is the entire image and the V region is the viewport, whole decoding or partial decoding can be performed. This example will be mainly described with partial decoding. In 22a, in the case of CMP (left), tiles e, f, and g are decoded, and in the case of CMP compact (right), tiles a, c, and e are decoded, so that the region corresponding to the viewport can be obtained. In 22b, in the case of CMP, tiles b, f, and i are decoded, and in the case of CMP compact, tiles d, e, and f are decoded, so that the region corresponding to the viewport can be obtained.
[0510] In the above example, the case of performing division such as slicing and tiling based on the surface unit (or surface boundary) has been described. However, as in 20a of FIG. 20, when dividing inside the surface (for example, ERP has an image composed of one surface, and another projection format is composed of a plurality of surfaces), or when dividing including the surface boundary, it is also possible.
[0511] FIG. 23 is a conceptual diagram for explaining an example of size adjustment of a 360-degree image according to an embodiment of the present invention. At this time, an example of an image projected by ERP will be described. Also, in the examples described later, the case of expansion will be mainly described.
[0512] The projected image can be size-adjusted using a scale factor or an offset factor according to the image size adjustment type. The image before size adjustment is P_Width×P_Height, and the image after size adjustment can be P’_Width×P’_Height.
[0513] In the case of the scale factor, after resizing using the scale factors for the width and height of the image (in this example, horizontal a and vertical b), the width (P_Width × a) and height (P_Height × b) of the image can be obtained. In the case of the offset factor, after resizing using the offset factors for the width and height of the image (in this example, horizontal L, R, vertical T, B), the width (P_Width + L + R) and height (P_Height + T + B) of the image can be obtained. Resizing can be performed using the already set method, or any one of a plurality of methods can be selected for resizing.
[0514] The data processing method in the example described later will be mainly explained for the case of the offset factor. In the case of the offset factor, in the data processing method, there may be methods such as filling using a predetermined pixel value, filling by copying the outer pixels, filling by copying a partial area of the image, and filling by converting a partial area of the image.
[0515] In the case of a 360-degree image, resizing can be performed considering the characteristic that there is continuity at the boundary of the image. In the case of ERP, although there is no outer boundary in three-dimensional space, when it is converted to two-dimensional space through the projection process, an outer boundary region can exist. The data in the boundary region has continuous data outside the boundary, but due to spatial characteristics, it can have a boundary. Resizing can be performed considering such characteristics. At this time, the continuity can be confirmed according to the projection format, etc. For example, in the case of ERP, it can be an image in which the boundaries at both ends have continuous characteristics. In this example, the case where the left and right boundaries of the image are continuous and the case where the upper and lower boundaries of the image are continuous are assumed for explanation, and the data processing method will be mainly explained for the methods of filling by copying a partial area of the image and filling by converting a partial area of the image.
[0516] When resizing to the left side of the image, the resized area (in this example, LC or TL+LC+BL) can be filled using the data of the right-side area of the image (in this example, tr+rc+br) that is continuous with the left side of the image. When resizing to the right side of the image, the resized area (in this example, RC or TR+RC+BR) can be filled using the data of the left-side area of the image (in this example, tl+lc+bl) that is continuous with the right side. When resizing to the upper side of the image, the resized area (in this example, TC or TL+TC+TR) can be filled using the data of the lower-side area of the image (in this example, bl+bc+br) that is continuous with the upper side. When resizing to the lower side of the image, the data of the resized area (in this example, BC or BL+BC+BR) can be used for filling.
[0517] When the size or length of the resized area is m, the resized area can have a range of (-m, y) to (-1, y) (resizing to the left side) or (P_Width, y) to (P_Width+m-1, y) (resizing to the right side) based on the coordinate system of the image before resizing (in this example, x ranges from 0 to P_Width-1). The position x' of the area for obtaining the data of the resized area can be derived by the formula x'=(x+P_Width)%P_Width. At this time, x represents the coordinates of the resized area based on the image coordinates before resizing, and x' represents the coordinates of the area referred to by the resized area based on the image coordinates before resizing. For example, when resizing to the left and m is 4 and the width of the image is 16, the corresponding data can be obtained from (12, y) for (-4, y), (13, y) for (-3, y), (14, y) for (-2, y), and (15, y) for (-1, y). Or, when resizing to the right and m is 4 and the width of the image is 16, the corresponding data can be obtained from (0, y) for (16, y), (1, y) for (17, y), (2, y) for (18, y), and (3, y) for (19, y).
[0518] When the size or length of the resized area is n, the resized area can have a range of (x, -n) to (x, -1) (resized upward) or (x, P_Height) to (x, P_Height + n - 1) (resized downward) with reference to the coordinates of the pre - resized image (in this example, y ranges from 0 to P_Height - 1). The position y' of the area for obtaining the data of the resized area can be derived by an equation such as y'=(y + P_Height)%P_Height. At this time, y means the coordinates of the resized area with reference to the pre - resized image coordinates, and y' means the coordinates of the area referred to the resized area with reference to the pre - resized image coordinates. For example, when resizing upward and n is 4 and the vertical width of the image is 16, data can be obtained from (x, -4) corresponding to (x, 12), (x, -3) corresponding to (x, 13), (x, -2) corresponding to (x, 14), and (x, -1) corresponding to (x, 15). Or, when resizing downward and n is 4 and the vertical width of the image is 16, data can be obtained from (x, 16) corresponding to (x, 0), (x, 17) corresponding to (x, 1), (x, 18) corresponding to (x, 2), and (x, 19) corresponding to (x, 3).
[0519] After filling the data of the resized area, it can be adjusted with reference to the coordinates of the post - resized image (in this example, x ranges from 0 to P'_Width - 1, y ranges from 0 to P'_Height - 1). The above example may be an example applicable to a latitude - longitude coordinate system.
[0520] The following various combinations of resizing can be available.
[0521] As an example, resizing can be performed by m to the left side of the image. Or, resizing can be performed by n to the right side of the image. Or, resizing can be performed by o to the upper side of the image. Or, resizing can be performed by p to the lower side of the image.
[0522] As an example, resizing can be performed by m to the left side and by n to the right side of the image. Or, resizing can be performed by o to the upper side and by p to the lower side of the image.
[0523] As an example, the size of the image can be adjusted by m to the left side, n to the right side, and o to the upper side. Or, the size of the image can be adjusted by m to the left side, n to the right side, and p to the lower side. Or, the size of the image can be adjusted by m to the left side, o to the upper side, and p to the lower side. Or, the size of the image can be adjusted by n to the right side, o to the upper side, and p to the lower side.
[0524] As an example, the size of the image can be adjusted by m to the left side, n to the right side, o to the upper side, and p to the lower side.
[0525] At least one size adjustment is performed as in the above example, and the size of the image can be implicitly adjusted according to the encoding / decoding settings, or size adjustment information can be explicitly generated, and the size of the image can be adjusted based on it. That is, m, n, o, and p in the above example can be determined to be predetermined values, or can be explicitly generated as size adjustment information, or some can be determined to be predetermined values and some can be explicitly generated.
[0526] The above example has been mainly described for the case of obtaining data from a partial area of the image, but other methods are also applicable. The data is either a pixel before encoding or a pixel after encoding, and can be determined according to the characteristics of the image or stage for which the size adjustment is performed. For example, when performing size adjustment in the preprocessing process or the stage before encoding, the data means input pixels such as a projection image or a packing image, and when performing size adjustment in the postprocessing process, the stage of generating in-picture prediction reference pixels, the stage of generating a reference image, the filter stage, etc., the data can mean restored pixels. Also, size adjustment can be performed using a data processing method for each area to be size-adjusted.
[0527] FIG. 24 is a conceptual diagram for explaining the continuity between surfaces in a projection format (e.g., CMP, OHP, ISP) according to an embodiment of the present invention.
[0528] Specifically, it can be an example for an image composed of a plurality of surfaces. Continuity is a characteristic that occurs in adjacent regions in three-dimensional space. FIGS. 24a to 24c show the cases where, when converted into two-dimensional space through the projection process, there is spatial adjacency and continuity (A), spatial adjacency but no continuity (B), no spatial adjacency but continuity (C), and no spatial adjacency and no continuity (D). A general image is different from being classified into the cases where there is spatial adjacency and continuity (A) and where there is no spatial adjacency and no continuity (D). At this time, when there is continuity, the said partial examples (A or C) are applicable.
[0529] That is, referring to FIGS. 24a to 24c, the case where there is spatial adjacency and continuity (explained with reference to 24a in this example) can be displayed as b0 to b4, and the case where there is no spatial adjacency and there is continuity can be displayed as B0 to B6. That is, it means the case for adjacent regions in three-dimensional space. By using the characteristics of continuity of b0 to b4 and B0 to B6 in the encoding process, the encoding performance can be improved.
[0530] FIG. 25 is a conceptual diagram for explaining the continuity of the surface of FIG. 21c, which is an image obtained through the image reconstruction process or the regional packing process in the CMP projection format.
[0531] Here, since 21c in FIG. 21 is a rearrangement of 21a that spreads a 360-degree image in the shape of a cube, at this time as well, the continuity of the surface possessed by 21a in FIG. 21 is maintained. That is, as in 25a, the surface S2,1 can be continuous with S1,1 and S3,1 on the left and right, and can be continuous with the surface S1,0 rotated 90 degrees and the surface S1,2 rotated -90 degrees above and below.
[0532] In a similar way, the continuity for the surface S3,1, the surface S0,1, the surface S1,2, the surface S1,1, and the surface S1,0 can be confirmed from FIGS. 25b to 25f.
[0533] The continuity between surfaces can be defined according to projection format settings etc., and is not limited to the above examples, and other modified examples are possible. The examples described later will be explained under the assumption that there is continuity as shown in FIGS. 24 and 25.
[0534] FIG. 26 is an exemplary diagram for explaining the size adjustment of an image in a CMP projection format according to an embodiment of the present invention.
[0535] 26a shows an example of performing image size adjustment, 26b shows an example of performing size adjustment in surface units (or divided units), and 26c shows an example of performing size adjustment (or multiple size adjustments) in image and surface units.
[0536] The projected image can perform size adjustment using a scale factor and size adjustment using an offset factor according to the image size adjustment type. The image before size adjustment is P_Width×P_Height, the image after size adjustment is P’_Width×P’_Height, and the size of the surface can be F_Width×F_Height. The size of the surface may be the same or different depending on the surface, and the horizontal width and vertical width of the surface may be the same or different. However, in this example, for the sake of convenience of explanation, it will be explained under the assumption that the sizes of all surfaces in the image are the same and have a square shape. Also, it will be explained under the assumption that the size adjustment values (in this example, WX, HY) are the same. The data processing method in the examples described later will be mainly explained for the case of the offset factor, and the data processing method will be mainly explained for the method of copying and filling a partial area of the image and the method of converting and filling a partial area of the image. The above settings can be similarly applied to FIG. 27.
[0537] In the cases of 26a to 26c, the boundary of the surface (assuming in this example that it has continuity according to 24a in FIG. 24) can have continuity with the boundaries of other surfaces. At this time, it can be classified into the case where they are spatially adjacent in the two-dimensional plane and there is image continuity (the first exemplification) and the case where they are not spatially adjacent in the two-dimensional plane and there is image continuity (the second exemplification).
[0538] For example, assuming the continuity of 24a in FIG. 24, the upper, left, right, and lower regions of S1,1 are spatially adjacent to the lower, right, left, and upper regions of S1,0, S0,1, S2,1, and S1,2, and the images can also be continuous (in the case of the first illustration).
[0539] Alternatively, the left and right regions of S1,0 are not spatially adjacent to the upper regions of S0,1 and S2,1, but their images can be continuous (in the case of the second illustration). Also, the left region of S0,1 and the right region of S3,1 are not spatially adjacent, but their images can be continuous (in the case of the second illustration). Also, the left and right regions of S1,2 and the lower regions of S0,1 and S2,1 can be continuous with each other (in the case of the second illustration). This is a limited example in this case, and depending on the definition and setting of the projection format, it can have a configuration different from the above. For convenience of explanation, S0,0 to S3,2 in FIG. 26a are referred to as a to l.
[0540] 26a can be an example of filling using data of regions where continuity exists in the outer boundary direction of the image. Regions that are resized from the A region where no data exists (in this example, a0 to a2, c0, d0 to d2, i0 to i2, k0, l0 to l2) can be filled via a predetermined arbitrary value or outer contour pixel padding, and regions that are resized from the B region containing actual data (in this example, b0, e0, h0, j0) can be filled using data of regions (or surfaces) where image continuity exists. For example, b0 can be filled using the lower data of the surface obtained by applying a 180-degree rotation to surface h, and j0 can be filled using the upper data of the surface obtained by applying a 180-degree rotation to surface h. However, in this case (including the examples described later), only the position of the surface being referred to is shown, and the data obtained for the resized region can be obtained after an adjustment process (such as rotation) considering the continuity between surfaces, as shown in FIGS. 24 and 25.
[0541] Specifically, b0 is an example of being filled using the lower data of the surface obtained by applying a 180-degree rotation to surface h, and j0 can be an example of being filled using the upper data of the surface obtained by applying a 180-degree rotation to surface h. However, in this case (including the examples described later), only the position of the surface being referred to is shown, and the data obtained for the resized region can be obtained after an adjustment process (such as rotation) considering the continuity between surfaces, as shown in FIGS. 24 and 25.
[0542] 26b may be an example of filling using data of a region where continuity exists in the direction of the inner boundary of the image. In this example, the resizing operations performed along the surface may be different. Region A can perform a shrinking process, and region B can perform an expanding process. For example, in the case of surface a, resizing (shrinking in this example) by w0 to the right is performed, and in the case of surface b, resizing (expanding in this example) by w0 to the left may be performed. Or, in the case of surface a, resizing (shrinking in this example) by h0 downward is performed, and in the case of surface e, resizing (expanding in this example) by h0 upward may be performed. In this example, looking at the change in the horizontal width of the image from surfaces a, b, c, d, surface a is shrunk by w0, surface b is expanded by w0 and w1, and surface c is shrunk by w1, so the horizontal width of the image before resizing is the same as the horizontal width of the image after resizing. Looking at the change in the vertical height of the image from surfaces a, e, i, surface a is shrunk by h0, surface e is expanded by h0 and h1, and surface i is shrunk by h1, so the vertical height of the image before resizing is the same as the vertical height of the image after resizing.
[0543] Considering that the regions to be resized (in this example, b0, e0, be, b1, bg, g0, h0, e1, ej, j0, gi, g1, j1, h1) are shrunk from region A where there is no data, they can also be simply removed. Considering that they are expanded from region B containing actual data, they can be newly filled with data of a region where continuity exists.
[0544] For example, b0 is above the surface e, e0 is to the left of the surface b, be is to the left of the surface b, or above the surface e, or a weighted sum of to the left of the surface b and above the surface e, b1 is above the surface g, bg is to the left of the surface b, or above the surface g, or a weighted sum of to the right of the surface b and above the surface g, g0 is to the right of the surface b, h0 is above the surface b, e1 is to the left of the surface j, ej is below the surface e, or to the left of the surface j, or a weighted sum of below the surface e and to the left of the surface j, j0 is below the surface e, gj is below the surface g, or to the left of the surface j, or a weighted sum of below the surface g and to the right of the surface j, g1 is to the right of the surface j, j1 is below the surface g, and h1 is data below the surface j can be used for filling.
[0545] When filling a resized area in the above example with data from a partial area of an image, the data of the corresponding area can be copied and filled, or the data obtained after going through a conversion process based on the characteristics, type, etc. of the image can be used to fill the data of the corresponding area. For example, when a 360-degree image is converted according to a projection format in a two-dimensional space, a coordinate system for each surface (for example, a two-dimensional plane coordinate system) can be defined. For the sake of convenience of explanation, it is assumed that in the three-dimensional space (x, y, z), each surface is converted to (x, y, C) or (x, C, z) or (C, y, z). The above example shows the case of obtaining data of a surface different from the corresponding surface in a resized area of some surfaces. That is, when resizing with the current surface as the center and directly copying and filling the data of other surfaces with other coordinate system characteristics, there is a possibility that the continuity is distorted based on the boundary of the resizing. Therefore, it is also possible to convert the data of other surfaces obtained according to the coordinate system characteristics of the current surface and fill it into the resized area. The case of conversion is only an example of a data processing method and is not limited to this.
[0546] When copying and filling the data of a partial area of an image into the area to be resized, the boundary area between the area to be resized (e) and the area being resized (e0) can include distorted continuity (or, rapidly changing continuity). For example, there may be cases where continuous features change based on the boundary. This is similar to an edge that had a straight shape becoming folded based on the boundary.
[0547] When converting and filling the data of a partial area of an image into the area to be resized, the boundary area between the area to be resized and the area being resized can include gradually changing continuity.
[0548] The above example can be an example of the data processing method of the present invention in which, in the resizing process (in this example, expansion), the data of a partial area of an image is subjected to conversion processing based on the characteristics, type, etc. of the image, and the obtained data is filled into the area to be resized.
[0549] 26c can be an example of filling using the data of the area where continuity exists in the direction of the boundary (inner boundary and outer boundary) of the image by combining the image resizing processes by 26a and 26b. Since the resizing process of this example can be derived from 26a and 26b, detailed description is omitted.
[0550] 26a can be an example of an image resizing process, and 26b can be an example of a resizing process of a division unit within the image. 26c can be an example of a plurality of resizing processes in which the image resizing process and the resizing of the division unit within the image are performed.
[0551] For example, size adjustment (in this example, C area) can be performed on an image obtained through a projection process (in this example, the first format), and size adjustment (in this example, D area) can be performed on an image obtained through a format conversion process (in this example, the second format). In this example, size adjustment (in this example, the entire image) is performed on the image projected by ERP, which may be an example where size adjustment (in this example, per surface unit) is performed after obtaining the image projected by CMP through the format conversion unit. The above example is an example of performing multiple size adjustments, and is not limited thereto, and modifications to other cases are also possible.
[0552] FIG. 27 is an exemplary diagram for explaining size adjustment for an image converted into a CMP projection format and packed according to an embodiment of the present invention. Since FIG. 27 is also explained on the premise of the continuity between surfaces as shown in FIG. 25, the boundary of the surface can have continuity with the boundaries of other surfaces.
[0553] In this example, the offset factors of W0...
Claims
1. A method for decoding a 360-degree image, comprising the steps of: receiving a bitstream in which a 360-degree image is encoded; generating a predicted image by referring to syntax information obtained from a received bitstream; combining the generated prediction image with a residual image obtained by inverse quantizing and inverse transforming the bitstream to obtain a decoded image; and reconstructing the decoded image into a 360 degree image in a projection format; The step of generating a predicted image includes: obtaining a motion vector candidate set including motion vectors of blocks adjacent to a current block to be decoded from motion information included in the syntax information; deriving a predicted motion vector from among a group of motion vector candidates based on selection information extracted from the motion information; and determining a prediction block of a current block to be decoded using a final motion vector derived by adding the predicted motion vector to a differential motion vector extracted from the motion information.
2. The group of motion vector candidates includes: If the block adjacent to the current block is different from the surface to which the current block belongs, The method for decoding a 360-degree image according to claim 1 , wherein the motion vectors are composed only of motion vectors for blocks belonging to a surface that has image continuity with a surface to which the current block belongs, among the adjacent blocks.
3. The adjacent blocks are The method of claim 1 , wherein the block adjacent to the current block is a block adjacent to the current block in at least one of an upper left corner, an upper corner, an upper right corner, a left corner, and a lower left corner.