Image decoding method and apparatus using a partition unit including an additional area

By introducing additional regions and multiple reference pixel levels in image decoding, the problem of low image coding efficiency in existing technologies is solved, achieving more efficient image compression and accurate in-frame prediction.

CN116248867BActive Publication Date: 2025-11-25INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310305081.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-16
Filing Date
2018-07-03
Publication Date
2025-11-25
Estimated Expiration
2038-07-03

AI Technical Summary

Technical Problem

In existing image coding methods, parallel blocks coded independently cannot effectively utilize temporally and spatially adjacent image data as references, resulting in low coding efficiency. Furthermore, in-frame prediction relies on the nearest neighbor pixel reference, which may not be applicable.

Method used

The image decoding method employs segmented units that include appended regions. It obtains syntax elements through bitstream, sets appended regions, and references blocks within the appended regions during decoding. It supports multiple reference pixel levels and adaptive reference pixel filtering, thereby improving image compression efficiency and in-frame prediction accuracy.

Benefits of technology

It improves image compression efficiency and in-frame prediction accuracy, enhances compression efficiency during image encoding/decoding, and adapts to different image characteristics for optimal reference pixel filtering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116248867B_ABST
    Figure CN116248867B_ABST
Patent Text Reader

Abstract

Disclosed is an image decoding method and apparatus using a partition unit including an additional area. The image decoding method using a partition unit including an additional area includes the steps of partitioning an encoded image included in a received bitstream into at least one partition unit by referring to syntax elements acquired from the bitstream; setting an additional area for the at least one partition unit; and decoding the encoded image based on the partition unit after the additional area is set. Thus, the image coding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application No. 2018800450230, titled "Image decoding method and apparatus using partition unit including additional area", filed on July 3, 2018. TECHNICAL FIELD

[0002] The present application relates to an image decoding method and apparatus using a partition unit including an additional area, and more particularly, to a technology for improving encoding efficiency by setting an additional area on the upper side, lower side, left side, and right side of a partition unit such as a tile within an image and simultaneously referring to image data in the additional area when encoding. BACKGROUND

[0003] In recent years, the demand for multimedia data such as video in the Internet is rapidly increasing. However, the development speed of the channel bandwidth is still behind the rapidly increasing amount of multimedia data, and thus there is a need for a method capable of effectively compressing the rapidly increasing amount of multimedia data. For this reason, the Moving Picture Expert Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) and the Video Coding Expert Group (VCEG) of the Telecommunication Standardization Sector (ITU-T) are making efforts to develop a more efficient video compression standard through persistent cooperative research.

[0004] In addition, when performing independent encoding on an image, independent encoding is usually performed on individual partition units including a tile, and thus there is a problem in that image data of other partition units adjacent in time and space cannot be used as a reference.

[0005] Therefore, there is a need for a method capable of using adjacent image data as a reference while maintaining the existing parallel processing based on independent encoding.

[0006] In addition, the intra prediction based on the existing image encoding / decoding method uses the most adjacent pixel to the current block as a reference pixel, and depending on the type of the image, the method of using the most adjacent pixel as a reference pixel can not be desirable.

[0007] Therefore, there is a need for a method capable of improving the intra prediction efficiency by adopting a different reference pixel configuration method than the existing method. SUMMARY

[0008] TECHNICAL PROBLEM

[0009] To solve the existing problems as described above, the present application aims to provide an image decoding apparatus and method using a partition unit including an additional area.

[0010] To solve the existing problems as described above, another object of the present application is to provide an image encoding apparatus and method using a partition unit including an additional area.

[0011] To solve the existing problems as described above, the present application aims to provide an image decoding method supporting multiple reference pixel levels.

[0012] To solve the existing problems as described above, another object of the present application is to provide an image decoding apparatus supporting multiple reference pixel levels.

[0013] Technical Solution

[0014] To achieve the above objects, in one aspect of the present application, there is provided an image decoding method using a partition unit including an additional area.

[0015] The image decoding method using a partition unit including an additional area can include the steps of partitioning an encoded image included in a received bitstream into at least one partition unit by referring to a syntax element acquired from the bitstream, setting an additional area for the at least one partition unit, and decoding the encoded image based on the partition unit after the additional area is set.

[0016] The step of decoding the encoded image can include the step of determining a reference block related to a current block to be decoded in the encoded image, according to information indicating reference possibility included in the bitstream.

[0017] The reference block can be a block included in a position overlapping the additional area set in a partition unit to which the reference block belongs.

[0018] To achieve the above objects, in another aspect of the present application, there is provided an image decoding method supporting multiple reference pixel levels.

[0019] The image decoding method supporting multiple reference pixel levels can include a step of confirming whether multiple reference pixel levels are supported through a bitstream, a step of determining a reference pixel level to be used in a current block by referring to syntax information included in the bitstream when the multiple reference pixel levels are supported, a step of constituting a reference pixel using a pixel included in the determined reference pixel level, and a step of performing an intra prediction of the current block using the constituted reference pixel.

[0020] The image decoding method supporting multiple reference pixel levels can include a step of confirming whether multiple reference pixel levels are supported through a bitstream, a step of determining a reference pixel level to be used in a current block by referring to syntax information included in the bitstream when the multiple reference pixel levels are supported, a step of constituting a reference pixel using a pixel included in the determined reference pixel level, and a step of performing an intra prediction of the current block using the constituted reference pixel.

[0021] The image decoding method supporting multiple reference pixel levels can include a step of confirming whether multiple reference pixel levels are supported through a bitstream, a step of determining a reference pixel level to be used in a current block by referring to syntax information included in the bitstream when the multiple reference pixel levels are supported, a step of constituting a reference pixel using a pixel included in the determined reference pixel level, and a step of performing an intra prediction of the current block using the constituted reference pixel.

[0022] Technical Effects

[0023] When the image decoding method and apparatus using a partition unit including an additional region according to the present application are used, the image compression efficiency can be improved because more image data can be used as a reference.

[0024] When the image decoding method and apparatus supporting multiple reference pixel levels according to the present application are used, the accuracy of intra prediction can be improved because multiple reference pixels can be used.

[0025] In addition, in the present application, because adaptive reference pixel filtering is supported, optimal reference pixel filtering can be performed according to the characteristics of an image.

[0026] In addition, the compression efficiency during image encoding / decoding can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 is a conceptual diagram illustrating an image encoding and decoding system according to an embodiment of the present application;

[0028] Figure 2 is a block diagram illustrating an image encoding apparatus according to an embodiment of the present application;

[0029] Figure 3 is a block diagram illustrating an image decoding apparatus according to an embodiment of the present application;

[0030] Figures 4a to 4d is a conceptual diagram for explaining a projection format according to an embodiment of the present application;

[0031] Figures 5a to 5c is a conceptual diagram for explaining a configuration of a surface to which an embodiment of the present application is applied;

[0032] Figures 6a to 6b is an explanatory diagram for explaining a division section to which an embodiment of the present application is applied;

[0033] Figure 7 is an explanatory diagram for dividing one image into a plurality of parallel blocks;

[0034] Figures 8a to 8i is a first explanatory diagram for setting an additional region to each parallel block illustrated in Figure 7

[0035] Figures 9a to 9i is a second explanatory diagram for setting an additional region to each parallel block illustrated in Figure 7

[0036] Figure 10 is an explanatory diagram for applying an additional region generated in an embodiment of the present application to an encoding / decoding process of other regions;

[0037] Figures 11 to 12 is a flowchart for explaining an encoding / decoding method of a division unit to which an embodiment of the present application is applied;

[0038] Figures 13a to 13g is an explanatory diagram for explaining a region referable by a specific division unit;

[0039] Figures 14a to 14e is a flowchart for explaining a reference possibility of an additional region in a division unit to which an embodiment of the present application is applied;

[0040] Figure 15 is an explanatory diagram for illustrating a block included in a division unit of a current image and a block included in a division unit of another image;

[0041] Figure 16 is a hardware configuration diagram for illustrating an image encoding / decoding apparatus to which an embodiment of the present application is applied;

[0042] Figure 17 is an explanatory diagram for illustrating an intra prediction mode to which an embodiment of the present application is applied;

[0043] Figure 18 is a first explanatory diagram for illustrating a reference pixel configuration used in an intra prediction to which an embodiment of the present application is applied;

[0044] Figures 19a to 19c ​​Fig. 2 is a second example diagram illustrating a reference pixel configuration to which an embodiment of the present application is applied;

[0045] Figure 20 Fig. 3 is a third example diagram illustrating a reference pixel configuration to which an embodiment of the present application is applied;

[0046] Figure 21 Fig. 4 is a fourth example diagram illustrating a reference pixel configuration to which an embodiment of the present application is applied;

[0047] Figures 22a to 22b Fig. 5 is an example diagram illustrating a method of filling a reference pixel to a predetermined position in a reference candidate block which is not available;

[0048] Figures 23a to 23c Fig. 6 is an example diagram illustrating a method of performing interpolation in a reference pixel configured according to an embodiment of the present application on the basis of a fractional pixel unit;

[0049] Figures 24a to 24b Fig. 7 is a first example diagram for explaining an adaptive reference pixel filtering method to which an embodiment of the present application is applied;

[0050] Figure 25 Fig. 8 is a second example diagram for explaining an adaptive reference pixel filtering method to which an embodiment of the present application is applied;

[0051] Figures 26a to 26b Fig. 9 is an example diagram illustrating a case where one reference pixel level is used in reference pixel filtering according to an embodiment of the present application;

[0052] Figure 27 Fig. 10 is an example diagram illustrating a case where a plurality of reference pixel levels are used in reference pixel filtering according to an embodiment of the present application;

[0053] Figure 28 Fig. 11 is a block diagram for explaining an encoding / decoding method of an intra prediction mode to which an embodiment of the present application is applied;

[0054] Figure 29 Fig. 12 is a first example diagram for explaining a bitstream configuration of an intra prediction based on a reference pixel configuration;

[0055] Figure 30 Fig. 13 is a second example diagram for explaining a bitstream configuration of an intra prediction based on a reference pixel configuration;

[0056] Figure 31 Fig. 14 is a third example diagram for explaining a bitstream configuration of an intra prediction based on a reference pixel configuration;

[0057] Figure 32is a flowchart illustrating a video decoding method supporting multiple reference pixel levels according to an embodiment of the present application. DETAILED DESCRIPTION

[0058] The present application is capable of various modifications and alternative forms, and specific embodiments are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that the drawings and detailed description thereto are not intended to limit the application to the particular form disclosed, but on the contrary, the intention is to cover all modifications, equivalents and alternatives falling within the spirit and scope of the present application as defined by the appended claims. Like numbers refer to like elements throughout the description of the figures.

[0059] In the course of describing various elements, terms such as 1st, 2nd, A, B, etc. can be used. However, the elements are not limited by the terms. The terms are used to distinguish one element from another. For example, the 1st element can be named the 2nd element, and the 2nd element can be named the 1st element without departing from the scope of the claims of the present application. The term "and / or" includes a combination of the associated recited items or one of the associated recited items.

[0060] When it is recited that an element is "connected" or "contacted" to another element, it should be understood that the element can be directly connected or contacted to the other element, or there can be another element between the two. In contrast, when it is recited that an element is "directly connected" or "directly contacted" to another element, it should be understood that there is no other element between the two.

[0061] The terms used in the present application are used to describe specific embodiments only and are not intended to limit the present application. Singular forms are intended to include plural forms unless the context clearly indicates otherwise. In the present application, the terms "include" or "have" are used to indicate that the features, numbers, steps, actions, elements, components or combinations thereof described in the specification are present, and should not be understood as precluding the presence or addition of one or more other features, numbers, steps, actions, elements, components or combinations thereof.

[0062] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Terms such as terms generally used in a dictionary are interpreted as having meanings identical to those in the context of related technology, and should not be interpreted as overly idealized or exaggerated meanings unless otherwise defined in the present application.

[0063] Generally, a video can be composed of a series of still images (Still Image), which can be divided in units of a group of pictures (GOP), and each still image can be referred to as a picture or a frame. As a higher-level concept, there can be units such as a group of pictures (GOP) and a sequence (Sequence), and each picture can be divided into specific regions such as a slice, a tile, a block, etc. In addition, a group of pictures (GOP) can include units such as an I picture, a P picture, a B picture, etc. An I picture can refer to a picture that is encoded / decoded by itself without using a reference picture, and a P picture and a B picture can refer to pictures that are encoded / decoded by performing processes such as motion estimation and motion compensation using a reference picture. Generally, a P picture can use an I picture and a P picture as a reference picture, and a B picture can use an I picture and a P picture as a reference picture, but the definition can be changed according to the setting of encoding / decoding.

[0064] Among them, a picture that is referred to as a reference in the encoding / decoding process is referred to as a reference picture, and a block or a pixel that is referred to as a reference is referred to as a reference block, a reference pixel. In addition, reference data (Reference Data) can be a coefficient value of a frequency domain (Frequency Domain) in addition to a pixel value of a spatial domain (Spatial Domain), various encoding / decoding information generated and determined in the encoding / decoding process.

[0065] The minimum unit of the image can be a pixel, and the number of bits used to represent one pixel is referred to as bit depth. Generally, the bit depth can be 8 bits, and other bit depths can be supported according to encoding settings. With respect to bit depth, at least one bit depth can be supported according to a color space. Further, at least one color space can be composed according to a color format of the image. According to the color format, one or more images having a certain size or one or more images having different sizes can be composed. For example, in the case of YCbCr 4:2:0, one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example) can be composed, and the composition ratio of the chrominance components and the luminance component can be 1:2 horizontally and vertically. As another example, in the case of 4:4:4, the composition ratio can be the same horizontally and vertically. In the case of composition by one or more color spaces as described above, partitioning in each color space can be performed.

[0066] In the present application, a part of the color space (Y in this example) of a part of the color format (YCbCr in this example) will be described as a reference, and the same or similar application can be made in other color spaces (Cb, Cr in this example) based on the color format (depending on the setting of the specific color space). However, a part of the difference can also be retained in each color space (independent of the setting of the specific color space). That is, depending on the setting of each color space can refer to the property of being proportional or dependent on the composition ratio of each component (for example, determined according to 4:2:0, 4:2:2, 4:4:4, etc.), and independent of the setting of each color space can refer to the property of being independent of the composition ratio of each component or only applicable to the corresponding color space. In the present application, depending on the encoder / decoder, a part of the composition can have the property of independence or the property of dependence.

[0067] Setting information or syntax elements required in an image encoding process can be determined at a unit level such as a video, a sequence, a picture, a slice, a tile, a block, etc., and can be included in a bitstream and transmitted to a decoder as a unit such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile header, a block header, etc., and can be parsed at the same level in the decoder and used in an image decoding process after decoding the setting information transmitted from the encoder. In addition, related information can be transmitted to the bitstream and used after parsing in the form of, for example, supplemental enhancement information (SEI) or metadata. Each parameter set has an inherent number value, and a lower parameter set can include a number value of an upper parameter set that needs to be referred to. For example, a lower parameter set can refer to information of an upper parameter set having the same number value from one or more upper parameter sets. In the above-described examples of various units, when one unit contains one or more other units, the corresponding unit can be referred to as an upper unit and the contained unit can be referred to as a lower unit.

[0068] Regarding the setting information generated at the above-described units, setting-related contents independent of each unit can be included, and setting-related contents dependent on a previous, subsequent, or upper unit, etc. can be included. Among them, the dependent setting refers to flag information (for example, a 1-bit flag, 1 indicating compliance and 0 indicating non-compliance) for indicating whether to comply with the setting of the previous, subsequent, or upper unit, and can be understood as indicating the setting information of the corresponding unit. Although the setting information will be described in the present application mainly with respect to examples of independent setting, examples of adding or replacing contents using the setting information of the previous, subsequent, or upper unit dependent on the current unit can also be included.

[0069] Next, preferred embodiments to which the present application is applied will be described in detail with reference to the accompanying drawings.

[0070] Figure 1 is a conceptual diagram illustrating an image encoding and decoding system to which an embodiment of the present application is applied.

[0071] Referring to Figure 1The image encoding device 105 and the decoding device 100 can be user terminals such as personal computers (PCs), laptops, personal digital assistants (PDAs), portable multimedia players (PMPs), portable game consoles (PSPs), wireless communication terminals, smartphones, and televisions, or server terminals such as application servers and business servers. They can include various devices such as communication modems for communicating with various devices or wired and wireless communication networks, memory 120 and 125 for storing various applications and data for performing inter-frame or intra-frame prediction in order to encode or decode images, and processors 110 and 115 for performing calculations and control by executing applications. Furthermore, the image encoded into a bitstream by the image encoding device 105 can be transmitted to the image decoding device 100 via wired wireless communication networks such as the Internet, short-range wireless communication networks, wireless local area networks, wireless broadband networks, and mobile communication networks, or via various communication interfaces such as cables and Universal Serial Bus (USB). The image is then decoded and reconstructed in the image decoding device 100 before being played back. Additionally, the image encoded into a bitstream by the image encoding device 105 can also be transmitted from the image encoding device 105 to the image decoding device 100 via a computer-readable storage medium.

[0072] Figure 2 This is a block diagram illustrating an image encoding apparatus to which one embodiment of the present invention is applied.

[0073] The image decoding device 20 applicable to this embodiment is as follows: Figure 2 As shown, it can include a prediction unit 200, a subtraction unit 205, a transformation unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse transformation unit 225, an addition unit 230, a filtering unit 235, an encoded image buffer 240, and an entropy coding unit 245.

[0074] The prediction unit 200 can include an intra prediction unit for performing intra prediction and an inter prediction unit for performing inter prediction. Intra prediction can generate a prediction block by performing spatial prediction using pixels of a block adjacent to the current block, and inter prediction can generate a prediction block by searching for an area that best matches the current block from a reference image and performing motion compensation. The specific information related to each prediction method (e.g., intra prediction mode, motion vector, reference image, etc.) can be determined after determining which of intra prediction or inter prediction is applied to the corresponding unit (coding unit or prediction unit). At this time, the processing unit for performing prediction and the processing unit for determining the prediction method and the specific content can differ according to the encoding / decoding setting. For example, the prediction method and the prediction mode can be determined at the prediction unit, and the execution of prediction can be performed at the transform unit.

[0075] The intra prediction unit can employ directional prediction modes such as horizontal, vertical modes, etc. used according to the prediction direction and non-directional prediction modes such as average (DC), planar (Planar) using the average, interpolation, etc. of reference pixels. Through the directional and non-directional modes, an intra prediction mode candidate group can be constructed, and one of various options such as 35 prediction modes (directional 33 + non-directional 2) or 67 prediction modes (directional 65 + non-directional 2), 131 prediction modes (directional 129 + non-directional 2), etc. can be used as a candidate group.

[0076] The intra prediction unit can include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel construction unit can construct reference pixels for performing intra prediction using pixels included in the adjacent block and adjacent to the current block with the current block as the center. According to the encoding setting, the reference pixels can be constructed using one of the most adjacent reference pixel rows, or other adjacent reference pixel rows, or a plurality of reference pixel rows. When a part of the reference pixels is not available, the reference pixels can be generated using the available reference pixels, and when all are not available, the reference pixels can be generated using a pre-set value (e.g., a middle value of the pixel value range that can be expressed by the bit depth, etc.).

[0077] The reference pixel filter of the intra prediction section can perform filtering on the reference pixel for the purpose of reducing distortion remaining through the encoding process. At this time, the filter used can be a low-pass filter such as a 3-tap filter [1 / 4, 1 / 2, 1 / 4], a 5-tap filter [2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16], etc. Whether or not to apply filtering and the filtering type can be determined according to the encoding information (e.g., the size, shape, prediction mode, etc. of the block).

[0078] The reference pixel interpolator of the intra prediction section can generate a pixel of a decimal unit through a linear interpolation process of the reference pixel according to the prediction mode, and can determine the interpolation filter to be applied according to the encoding information. At this time, the interpolation filter used can include a 4-tap Cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc. Generally, the process of performing low-pass filtering and the process of performing interpolation are independent of each other, but filtering can also be performed after integrating the filters applied in the two processes into one.

[0079] The prediction mode determiner of the intra prediction section can select the best prediction mode from the prediction mode candidate group while taking into account the encoding cost, and the prediction block generator can generate a prediction block using the corresponding prediction mode. In the prediction mode encoder, the above-mentioned best prediction mode can be encoded based on the prediction value. At this time, the prediction information can be adaptively encoded according to whether the prediction value is appropriate or not.

[0080] The above-mentioned prediction value is referred to as the most probable mode (MPM) in the intra prediction section, and a part of all the modes included in the prediction mode candidate group can be selected to constitute the most probable mode (MPM) candidate group. In the most probable mode (MPM) candidate group, a prediction mode (e.g., a DC, a Planar, a vertical, a horizontal, a diagonal mode, etc.) set in advance or a prediction mode of a spatially adjacent block (e.g., a left, an upper, a left upper, a right upper, a left lower block, etc.) can be included. In addition, the most probable mode (MPM) candidate group can be constituted using a mode (a difference of +1, -1, etc. in a directional mode) derived from a mode included in the most probable mode (MPM) candidate group in advance.

[0081] In the prediction mode used for constructing the most probable mode (MPM) candidate group, there can be a priority order. The order in which the prediction modes included in the most probable mode (MPM) candidate group can be determined according to the above-mentioned priority order, and the construction of the most probable mode (MPM) candidate group can be completed when the number of the most probable mode (MPM) candidate group is filled according to the above-mentioned priority order (determined according to the number of prediction mode candidate groups). At this time, the priority order can be determined in the order of the prediction mode of the spatially adjacent block, the prediction mode set in advance, the mode derived from the prediction mode included in the most probable mode (MPM) candidate group earlier, and other variations can also be made.

[0082] For example, in the spatially adjacent blocks, the left-top-left-bottom-right-top block and the like can be included in the candidate group in the order of left-top-left-bottom-right-top, and in the prediction mode set in advance, the mean (DC) -planar (Planar) -vertical -horizontal mode and the like can be included in the candidate group in the order of mean (DC) -planar (Planar) -vertical -horizontal mode, and the mode obtained by performing addition operations such as +1, -1 and the like on the modes included in advance can be included in the candidate group, so that the candidate group is constructed using a total of 6 modes. Alternatively, it can also be included in the candidate group in one priority order of left-top-mean (DC) -planar (Planar) -left-bottom-right-top - (left +1) - (left -1) - (top +1) and the like, so that the candidate group is constructed using a total of 7 modes.

[0083] In the above-mentioned candidate group construction, an effectiveness check can be performed, so that it is included in the candidate group only when it is effective, and it jumps to the next candidate when it is not effective. For example, it can be ineffective when the adjacent block is located outside the image or is included in a different partition unit from the current block, or the encoding mode of the corresponding block is inter prediction, and in addition, it can also be ineffective in the case of non-reference described later in the present application.

[0084] In the above-mentioned selection, the spatially adjacent block can be composed of one block, or can be composed of multiple blocks (sub-blocks). Therefore, in the order such as (left-top) in the above-mentioned candidate group construction, the order can be to jump to the top block after performing an effectiveness check on a certain position (for example, the lowermost block of the left block) in the left block, or the order can be to jump to the top block after performing an effectiveness check on multiple positions (for example, one or more sub-blocks below the uppermost block of the left block), and it can be determined according to the encoding setting.

[0085] In the inter-picture prediction section, it can be classified into a translational motion model and a non-translational motion model according to a motion prediction method. In the translational motion model, prediction is performed while only considering parallel translation, while in the non-translational motion model, prediction is performed while considering motions such as rotation, perspective, zoom in / out, etc. in addition to parallel translation. In the translational motion model, one motion vector will be needed under the assumption of uni-directional prediction, while in the non-translational motion model, more than one motion vector will be needed. In the non-translational motion model, each motion vector can be information applicable to a predetermined position of a current block such as a top-left vertex, a top-right vertex, etc. of the current block, and through the corresponding motion vector, the position of a region in the current block for which prediction is needed can be obtained in pixel units or sub-block units. The inter-picture prediction section can individually apply some of the processes described below while applying the others commonly.

[0086] The inter-picture prediction section can include a reference picture construction section, a motion prediction section, a motion compensation section, a motion information determination section, and a motion information encoding section. The reference picture construction section can include previously or subsequently coded pictures in reference picture lists (L0, L1) with the current picture as the center. From the reference pictures included in the above-mentioned reference picture lists, a prediction block can be obtained, and according to the coding settings, the current picture can also be used to construct a reference picture and included in at least one of the positions in the reference picture lists.

[0087] In the inter-picture prediction section, the reference picture construction section can include a reference picture interpolation section, and can perform an interpolation process for fractional pixel units according to the interpolation accuracy. For example, an 8-tap (8-tap) interpolation filter based on discrete cosine transform (DCT) can be applied to the luminance component, and a 4-tap (4-tap) interpolation filter based on discrete cosine transform (DCT) can be applied to the color difference component.

[0088] In the inter-picture prediction section, the motion prediction section is used to perform a process of exploring a block having a high correlation with the current block through reference pictures, and various methods such as a full search block matching algorithm (FBMA, Full search-based Block Matching Algorithm), a three step search algorithm (TSS, Three Step Search), etc. can be used, while the motion compensation section is used to perform a process of obtaining a prediction block through the motion prediction process.

[0089] In the inter-picture prediction section, the motion information determination section can perform a process for selecting the optimum motion information of the current block, which can be coded by a motion information coding mode such as a skip mode, a merge mode, a competition mode, and the like. The above modes can be combined with supported modes according to a motion model, and examples thereof can include a skip mode (moving), a skip mode (non-moving), a merge mode (moving), a merge mode (non-moving), a competition mode (moving), and a competition mode (non-moving). According to a coding configuration, a part of the above modes can be included in a candidate group.

[0090] The above motion information coding mode can obtain a prediction value of the motion information (motion vector, reference picture, prediction direction, and the like) of the current block from at least one candidate block, and can generate optimum candidate selection information when two or more candidate blocks are supported. The skip mode (no residual signal) and the merge mode (with residual signal) can use the above prediction value directly as the motion information of the current block, and the competition mode can generate difference value information between the motion information of the current block and the above prediction value.

[0091] The candidate group for the prediction value of the motion information of the current block is adaptive according to the motion information coding mode and can have various configurations. The motion information of a block spatially adjacent to the current block (for example, a left, an upper, a top-left, a top-right, a bottom-left block, and the like) can be included in the candidate group, and the motion information of a block temporally adjacent to the current block (for example, a left, a right, an upper, a lower, a top-left, a top-right, a bottom-left, a bottom-right block, and the like, including a block corresponding to or corresponding to the current block in other images) can also be included in the candidate group, and in addition, mixed motion information of a spatial candidate and a temporal candidate (for example, information obtained by averaging, a central value, and the like, from the motion information of a block spatially adjacent to the current block and the motion information of a block temporally adjacent to the current block, and motion information that can be obtained in units of the current block or a sub-block of the current block) can also be included in the candidate group.

[0092] In the configuration of the candidate group of the prediction value of the motion information, there can be a priority order. The order included in the configuration of the prediction value candidate group can be determined according to the above priority order, and the configuration of the candidate group can be completed when the number of the candidate group is filled according to the above priority order (determined according to the motion information coding mode). At this time, the priority order can be determined in the order of the motion information of a block spatially adjacent to the current block, the motion information of a block temporally adjacent to the current block, and mixed motion information of a spatial candidate and a temporal candidate, and other variations can also be performed.

[0093] For example, in spatially adjacent blocks, the blocks can be included in the candidate group in the order of left-top-right-top-left-bottom-left-top, and in temporally adjacent blocks, the blocks can be included in the candidate group in the order of bottom-middle-right-bottom.

[0094] In the above candidate group formation, an effectiveness check can be performed, so that the blocks are included in the candidate group only when they are effective and skipped to the next candidate when they are not effective. For example, the blocks can be not effective when the adjacent blocks are located outside the image or are included in a different partition unit from the current block or the encoding mode of the corresponding blocks is intra prediction, and in addition, can be not effective in the case of non-reference described later in the present application.

[0095] In the above candidate group formation, the blocks adjacent in space or time can be formed of one block or a plurality of blocks (sub-blocks). Therefore, in the order such as (left-top) in the above spatial candidate group formation, the order can be to jump to the top block after performing an effectiveness check on a certain position (for example, the lowermost block of the left block) in the left block, or the order can be to jump to the top block after performing an effectiveness check on a plurality of positions (for example, one or more sub-blocks below the uppermost block of the left block in the predetermined order from the block <2, 2>). In addition, in the order such as (middle-right) in the temporal candidate group formation, the order can be to jump to the right block after performing an effectiveness check on a certain position (for example, <2, 2> when the central block is partitioned into 4 x 4 regions) in the middle block, or the order can be to jump to the bottom block after performing an effectiveness check on a plurality of positions (for example, one or more sub-blocks such as <3, 3>, <2, 3>, and the like on the predetermined order from the position block <2, 2>), and can be determined according to the encoding setting.

[0096] The subtraction operation section 205 generates a residual block by performing a subtraction operation on the current block and the prediction block. That is, the subtraction operation section 205 generates a residual signal in the form of a block, that is, a residual block, by calculating the difference between the pixel value of each pixel of the current block to be encoded and the prediction pixel value of each pixel of the prediction block generated by the prediction section.

[0097] The transform unit 210 transforms each pixel value of the residual block into a frequency coefficient by transforming the residual block into a frequency region. Among them, the transform unit 210 can transform the residual signal into a frequency signal using various transform techniques for transforming a pixel signal of a spatial axis into a frequency axis such as Hadamard transform, DCT based transform, DST based transform, KLT based transform, and the like, and the residual signal transformed into a frequency region will become a frequency coefficient. When transforming, it can be transformed by a 1-dimensional transform matrix. Each transform matrix can be adaptively used in horizontal and vertical units. For example, when the prediction mode in intra prediction is horizontal, a transform matrix based on DCT can be used in the vertical direction and a transform matrix based on DST can be used in the horizontal direction. And when the prediction mode is vertical, a transform matrix based on DCT can be used in the horizontal direction and a transform matrix based on DST can be used in the vertical direction.

[0098] The quantization unit 215 quantizes the residual block containing the frequency coefficient transformed into a frequency region by the transform unit 210. Among them, the quantization unit 215 can quantize the transformed residual block using a quantization technique such as dead zone uniform threshold quantization, quantization weighted matrix, or an improved quantization technique thereof. At this time, one or more quantization techniques can be selected as candidates and can be determined according to the encoding mode, prediction mode information, and the like.

[0099] The entropy encoding unit 245 generates a quantization coefficient sequence by scanning the generated quantization frequency coefficient sequence using various scanning methods, and outputs it after encoding using an entropy encoding technique or the like. As a scanning mode, one of various modes such as zigzag, diagonal, raster, and the like can be set. In addition, it can generate and output encoding data containing encoding information transferred from each constituent unit to a bitstream.

[0100] The inverse quantization unit 220 inverse quantizes the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 generates a residual block containing a frequency coefficient by inverse quantizing the quantization frequency coefficient sequence.

[0101] The inverse transform unit 225 performs inverse transform on the residual block that has been inverse quantized by the inverse quantization unit 220. That is, the inverse transform unit 225 generates a residual block including pixel values, i.e., a reconstructed residual block, by performing inverse transform on the frequency coefficients of the residual block that has been inverse quantized. Here, the inverse transform unit 225 can perform inverse transform by inversely using the transform method used in the transform unit 210.

[0102] The addition unit 230 can reconstruct the current block by performing addition on the predicted block predicted in the prediction unit 200 and the residual block reconstructed by the inverse transform unit 225. The reconstructed current block is stored in the coded picture buffer 240 as a reference picture (or a reference block), and can be used as a reference picture when encoding a next block or other blocks, other pictures, of the current block.

[0103] The filter unit 235 can include one or more post-processing filter processes such as a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), etc. The deblocking filter can remove block distortion that occurs on a boundary between blocks in a reconstructed picture. The adaptive loop filter (ALF) can perform filtering based on a value obtained by comparing a reconstructed picture with an original picture after filtering the blocks by the deblocking filter. The sample adaptive offset (SAO) can reconstruct a difference in offset between an original picture and a residual block to which the deblocking filter has been applied, in pixel units. The post-processing filter as described above can be applied to a reconstructed picture or block.

[0104] The deblocking filter in the filter unit can be applied based on pixels included in several columns or rows included in two blocks with reference to a block boundary. As the blocks, it is preferable to apply to a boundary of an encoding block, a predicted block, a transform block, and can be limited to blocks having a minimum size (e.g., 8x8) or more.

[0105] As to whether to apply filtering, it is possible to determine whether to apply filtering and a filtering strength in consideration of a block boundary characteristic, and to determine one selected from among strong filtering, medium filtering, weak filtering, etc. Further, when the block boundary belongs to a boundary of a partition unit, it is possible to determine whether to apply filtering on the boundary of the partition unit according to a loop filter application flag, and to determine whether to apply filtering according to various conditions to be described later in the present application.

[0106] The sample adaptive offset (SAO) in the filter section can be applied based on a difference value between the reconstructed image and the original image. As the offset type, edge offset and band offset can be supported, and one of the above offsets can be selected according to the characteristics of the image to perform filtering. In addition, the above offset-related information can be encoded in a block unit, and can be encoded by a prediction value related thereto. At this time, the related information can be adaptively encoded according to a suitable case and an unsuitable case of the prediction value. The prediction value can be offset information of a neighboring block (for example, a left block, an upper block, a left upper block, a right upper block, etc.), and selection information related to which block's offset information is acquired can be generated.

[0107] An effectiveness check can be performed in the above candidate composition, so that only when it is effective, it is included in the candidate group, and when it is ineffective, it jumps to the next candidate. For example, a neighboring block can be located outside the image or included in a different partition unit from the current block or in a case where it cannot be referenced as described later in the present invention, it can be ineffective.

[0108] The coded picture buffer 240 can store a block or an image reconstructed by the filter section 235. The reconstructed block or image stored in the coded picture buffer 240 can be provided to the prediction section 200 for performing intra prediction or inter prediction.

[0109] Figure 3 is a configuration diagram illustrating an image decoding apparatus to which an embodiment of the present invention is applied.

[0110] Referring to Figure 3 , the image decoding apparatus 30 can include an entropy decoding section 305, a prediction section 310, an inverse quantization section 315, an inverse transform section 320, an adder / subtracter 325, a filter 330, and a decoded picture buffer 335.

[0111] In addition, the prediction section 310 can further include an intra prediction module and an inter prediction module.

[0112] First, when an image bitstream transferred from the image encoding apparatus 20 is received, it can be transferred to the entropy decoding section 305.

[0113] The entropy decoding section 305 can decode the decoded data including the quantized coefficients and the decoded information transferred from each configuration section by decoding the bitstream.

[0114] The prediction unit 310 can generate a prediction block based on the data passed from the entropy decoding unit 305. At this time, a reference picture list using a default configuration technique can also be constructed based on the decoded reference pictures stored in the decoded picture buffer 335.

[0115] The intra prediction unit can include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction block generation unit, a prediction mode decoding unit, and the inter prediction unit can include a reference picture construction unit, a motion compensation unit, and a motion information decoding unit, some of which can perform the same process as the encoder, and the other of which can perform a process induced in the reverse direction.

[0116] The inverse quantization unit 315 can inverse quantize the quantized transform coefficients supplied from the bitstream and decoded by the entropy decoding unit 305.

[0117] The inverse transform unit 320 can generate a residual block by applying an inverse discrete cosine transform (DCT), an inverse integer transform, or an inverse transform technique similar thereto to the transform coefficients.

[0118] At this time, the inverse quantization unit 315 and the inverse transform unit 320 inversely perform the processes performed in the transform unit 210 and the quantization unit 215 of the image encoding apparatus 20 described above, and can be implemented by various methods. For example, the same processes and inverse transforms shared with the transform unit 210 and the quantization unit 215 can be used, and the transform and quantization processes can be inversely performed using information related to the transform and quantization processes of the image encoding apparatus 20 (e.g., transform size, transform shape, quantization type, etc.).

[0119] The residual block that has undergone the inverse quantization and inverse transform processes can generate a reconstructed image block by being added to the prediction block derived in the prediction unit 310. The addition operation described above can be performed by the adder-subtracter 325.

[0120] The filter 330 can apply a deblocking filter for removing blocking artifacts to the reconstructed image block as necessary, and can additionally use other loop filters before and after the decoding process to improve video quality.

[0121] The reconstructed and filtered image block can be stored in the decoded picture buffer 335.

[0122] Although not shown, the image decoding apparatus 30 can further include a division unit, and at this time, the division unit can include a picture division unit and a block division unit. The division unit is a unit that divides a picture into blocks, and the picture division unit and the block division unit can be implemented by various methods. Figure 2The same or corresponding configurations as those illustrated in the image decoding apparatus are easily understood by a person skilled in the art, and thus detailed descriptions thereof will be omitted here.

[0123] Figures 4a to 4d is a conceptual diagram for explaining a projection format to which an embodiment of the present application is applied.

[0124] Figure 4a An Equi-Rectangular Projection (ERP) format in which a 360-degree image is projected to a 2-dimensional plane is illustrated. Figure 4b A Cube Map Projection (CMP) format in which a 360-degree image is projected to a cube is illustrated. Figure 4c An OctaHedron Projection (OHP) format in which a 360-degree image is projected to an octahedron is illustrated. Figure 4d An IcoSahedral Projection (ISP) format in which a 360-degree image is projected to a polyhedron is illustrated. However, this is not limited thereto, and various projection formats can be used. For example, a Truncated Square Pyramid Projection (TSP), a Segmented Sphere Projection (SSP), or the like can be used. Figures 4a to 4d The left side is a 3-dimensional model, and the right side is an example transformed into a 2-dimensional space through a projection process. The 2-dimensional projection graph can be composed of one or more faces, and each face can have a circular shape, a triangular shape, a quadrangular shape, or the like.

[0125] As illustrated in Figures 4a to 4d , a projection format can be composed of one face (e.g., an Equi-Rectangular Projection (ERP)) or a plurality of faces (e.g., a Cube Map Projection (CMP), an OctaHedron Projection (OHP), an IcoSahedral Projection (ISP), or the like). In addition, each face can be classified into shapes such as a quadrangular shape and a triangular shape. The above classification can be an example of a type, a characteristic, or the like of an image in the present application that can be applied when different encoding / decoding settings are set according to a projection format. For example, the type of an image can be a 360-degree image, and the characteristic of an image can be one of the above classifications (e.g., each projection format, a projection format of one face or a plurality of faces, a projection format in which a face is a quadrangular shape or is not a quadrangular shape, or the like).

[0126] A 2-dimensional plane coordinate system {e.g., (i, j)} can be defined on each surface of a 2-dimensional projection image, and the characteristics of the coordinate system can vary depending on the projection format, the position of each surface, etc. A 2-dimensional plane coordinate system can be included in the case of equirectangular projection (ERP), and other projection formats can include multiple 2-dimensional plane coordinate systems depending on the number of surfaces. At this time, the coordinate system can be expressed as (k, i, j), where k can be index information of each surface.

[0127] In the present application, for the convenience of explanation, a case where the surface shape is a quadrangle will be described as the center, and the number of surfaces projected to 2 dimensions can be one {e.g., equirectangular projection (ERP), i.e., a case where the image is identical to one surface} to two or more {e.g., cube map projection (CMP), etc.}.

[0128] Figures 5a to 5c is a conceptual diagram for explaining the layout of surfaces to which an embodiment of the present application is applied.

[0129] In a projection format in which a 3-dimensional image is projected to 2 dimensions, the layout of the surfaces needs to be determined. At this time, the surface layout can be configured in a manner that maintains the continuity of the image in 3-dimensional space, or in a manner that maximizes the closeness of the interval between the surfaces even if the continuity of the image between some adjacent surfaces is broken. In addition, when the surfaces are configured, some surfaces can be configured after being rotated by a certain angle (0, 90, 180, 270 degrees, etc.).

[0130] Referring to Figure 5a , an example of the layout of surfaces related to the cube map projection (CMP) format can be confirmed, and when configured in a manner that maintains the continuity of the image in 3-dimensional space, a 4 x 3 layout in which four surfaces are configured in the horizontal direction and then one surface is configured above and below each of the four surfaces as shown in the left image can be used. In addition, a 3 x 2 layout in which the surfaces are seamlessly configured on a 2-dimensional plane even if the continuity of the image between some adjacent surfaces is broken as shown in the right image can also be used.

[0131] Referring to Figure 5b , the layout of surfaces related to the octahedral projection (OHP) format can be confirmed, and when configured in a manner that maintains the continuity of the image in 3-dimensional space as shown in the upper image. At this time, when the surfaces are seamlessly configured on the projected 2-dimensional plane even if the continuity is broken in some part, the surfaces can also be configured in the manner shown in the lower image.

[0132] Referring to Figure 5c, it is possible to confirm the surface configuration related to the icosahedron projection (ISP) format, and it is possible to configure the surfaces in a manner that maintains the continuity of the image in the 3-dimensional space as shown in the upper end image, and it is also possible to configure the surfaces in a manner that the surfaces are seamlessly close to each other as shown in the lower end image.

[0133] At this time, the process of configuring the surfaces without gaps can be referred to as frame packing, and the phenomenon that the image continuity is damaged can be minimized by configuring the surfaces after rotating the surfaces. Next, the process of changing the surface configuration to another surface configuration as described above will be referred to as surface reconfiguration.

[0134] Next, the term continuity can be explained as the continuity of the scene that is visible to the naked eye in the 3-dimensional space, or the continuity of the actual image or scene in the 2-dimensional projection space. Having continuity can also mean that the correlation between regions is high. In general, the correlation between regions in a 2-dimensional image can be high or low, but in a 360-degree image, there can be regions that have no continuity even though they are adjacent in space. In addition, according to the surface configuration or reconfiguration as described above, there can be regions that have continuity although they are not adjacent in space.

[0135] The surface reconfiguration can be performed for the purpose of improving the encoding performance. For example, it is possible to configure the surfaces having image continuity adjacent to each other by performing the surface reconfiguration.

[0136] At this time, the surface reconfiguration does not necessarily mean that the surfaces are newly configured after being configured, but can also be understood as a process of setting a specific surface configuration from the beginning. (It can be performed in region-wise packing in the 360-degree image encoding / decoding process)

[0137] In addition, the surface configuration or reconfiguration can include the rotation of the surfaces in addition to the change in the position of each surface (in this example, the simple movement of the surfaces, for example, movement from the upper end of the left side to the lower end of the left side or the lower end of the right side of the image). Among them, the rotation of the surfaces can include 0 degrees without surface rotation, 45 degrees to the right, 90 degrees to the left, etc., and can represent the rotation angle by selecting the divided interval after dividing 360 degrees (equally or unequally) into k (or 2k) intervals.

[0138] The encoder can generate the surface configuration information and / or the surface reconfiguration information from the input image, and the decoder can receive the information from the encoder and decode the information to perform the surface configuration (or reconfiguration) as described above.

[0139] Next, the surfaces referred to below without separate description are premised on a 3x2 layout in Figure 5a and at this time, the numbers of the respective surfaces can be 0 to 5 in a raster scan order from the left upper end.

[0140] Next, in a description premised on continuity between the surfaces in Figure 5a unless otherwise specified, it can be assumed that the surfaces Nos. 0 to 2 have continuity with each other, the surfaces Nos. 3 to 5 have continuity with each other, and the surfaces Nos. 0 and 3, 1 and 4, and 2 and 5 do not have continuity. Whether the surfaces have continuity or not can be determined by the characteristics, type, format, etc. of the image.

[0141] In the encoding / decoding process of the 360-degree image, the encoding apparatus can acquire an input image, pre-process the acquired image, encode the pre-processed image, and transmit the encoded bitstream to the decoding apparatus. The pre-processing can include image stiching, projection of a 3-dimensional image to a 2-dimensional plane, surface configuration, and reconfiguration (or referred to as region-wise packing), etc. In addition, the decoding apparatus can receive the bitstream, decode the received bitstream, perform post-processing (image rendering, etc.) on the decoded image, and generate an output image.

[0142] At this time, the bitstream can include information (supplemental enhancement information (SEI) message or metadata, etc.) generated in the pre-processing process and information (image encoding data) generated in the encoding process, and then transmit the information.

[0143] Figures 6a to 6b is an example diagram for describing a division part to which an embodiment of the present application is applied.

[0144] Figure 2 or Figure 3The image encoding / decoding apparatus in the present application can further include a division section, and the division section can include an image division section and a block division section. The image division section can divide an image into at least one processing unit {e.g., color space (YCbCr, RGB, XYZ, etc.), sub-image, slice, parallel block, basic coding unit (or maximum coding unit), etc.}, and the block division section can divide the basic coding unit into at least one processing unit (e.g., coding, prediction, transform, quantization, entropy, loop filtering unit, etc.).

[0145] The basic coding unit can be obtained by dividing an image in a certain length interval along the horizontal direction or the vertical direction, and can also be a unit applicable to, for example, a sub-image, a parallel block, a slice, a surface, etc. That is, the above-mentioned unit can be configured in an integer multiple of the basic coding unit, but is not limited thereto.

[0146] For example, different basic coding units can be applied in some of the division units (in this example, parallel blocks, sub-images, etc.), and the corresponding division units can adopt independent basic coding unit sizes. That is, the basic coding unit of the above-mentioned division unit can be set to be the same as or different from the basic coding unit of the image unit and the basic coding unit of other division units.

[0147] In the present application, for the convenience of explanation, the basic coding unit and other processing units (coding, prediction, transform, etc.) other than the basic coding unit are referred to as blocks.

[0148] The size or shape of the above-mentioned block can be an N×N square shape (2n×2n, 256×256, 128×128, 64×64, 32×32, 16×16, 8×8, 4×4, etc., n is an integer between 2 and 8) or an M×N rectangular shape (2m×2n) expressed by an exponential power of 2 (2 n ) in the horizontal or vertical length. For example, an input image of an 8k ultra-high definition (UHD) level with a very high resolution can be divided into a size of 256×256, an input image of a 1080p high definition (HD) level can be divided into a size of 128×128, and an input image of a wide video graphics array (WVGA) level can be divided into a size of 16×16.

[0149] An image can be divided into at least one slice. A slice can be composed of a combination of at least one block that is continuous in the scanning order. Each slice can be divided into at least one slice segment, and each slice segment can be divided into a basic coding unit.

[0150] The image can be divided into at least one sub-image or parallel block. The sub-image or parallel block can be divided in a quadrangle (rectangle or square) form, and can be divided into a basic coding unit. The sub-image is similar to the parallel block in terms of being divided in the same form (quadrangle). However, the sub-image is different from the parallel block in terms of being able to be distinguished from the parallel block in terms of being independently coded / decoded. That is, the parallel block receives setting information for performing coding / decoding from a superior unit (e.g., an image, etc.), whereas the sub-image is able to directly acquire at least one setting information for performing coding / decoding from header information of each sub-image. That is, the parallel block and the sub-image are different in that the parallel block is a unit obtained by simply dividing an image, and is not a unit for transmitting data (e.g., a basic unit of a video coding layer (VCL)).

[0151] In addition, the parallel block can be a division unit supported in terms of parallel processing, and the sub-image can be a division unit supported in terms of independent coding / decoding. In detail, the sub-image can not only be set in units of sub-images, but also determine whether to code / decode, and can be constituted and displayed with a region of interest as a center, and settings related thereto can be determined in units of a sequence, an image, etc.

[0152] In the above-described example, it is also possible to change to a manner in which coding / decoding is set in a superior unit of a sub-image and independent coding / decoding is set in a parallel block unit. In the present application, for convenience of explanation, it will be assumed that the parallel block is set independently and in dependence on a superior unit.

[0153] The division information generated when an image is divided in a quadrangle form can be in various forms.

[0154] Referring to Figure 6a, it is possible to obtain a quadrangular division unit by dividing the image once along the horizontal line b7 and the vertical line (at this time, bl and b3, b2 and b4 will divide to constitute one division line). For example, it is possible to generate the number information of the quadrangle based on the horizontal and vertical directions, respectively. At this time, when the quadrangle is equally divided, it is possible to confirm the horizontal and vertical lengths of the divided quadrangle by applying the number of horizontal and vertical lines to the horizontal and vertical lengths of the image, respectively, and when the quadrangle is not equally divided, it is possible to additionally generate information indicating the horizontal and vertical lengths of the quadrangle. At this time, the horizontal and vertical lengths can be expressed in one pixel unit or in a pixel unit. In the case of expressing in a plurality of pixel units, for example, when the size of the basic coding unit is M x N and the size of the quadrangle is 8M x 4N, it is possible to express the horizontal and vertical lengths as 8 and 4, respectively (in this example, the basic coding unit of the corresponding division unit is M x N), or as 16 and 8 (in this example, the basic coding unit of the corresponding division unit is M / 2 x N / 2).

[0155] Further, referring to Figure 6b , it is possible to confirm the case where a quadrangular division unit is obtained by independently dividing the image differently. Figure 6a For example, it is possible to generate the number information of the quadrangle within the image, the start position information of each quadrangle in the horizontal and vertical directions (the positions indicated by reference numerals z0 to z5 can be expressed in x, y coordinates within the image), and the horizontal and vertical length information of each quadrangle. At this time, the start position can be expressed in one pixel unit or a plurality of pixel units, and the horizontal and vertical lengths can also be expressed in one pixel unit or a plurality of pixel units.

[0156] Figure 6a It can be an example of division information of a parallel block or a sub-image, and Figure 6b It can be an example of division information of a sub-image, but is not limited thereto. Next, for the convenience of explanation, it will be assumed that a quadrangular division unit is explained for a parallel block, but the explanation related to the parallel block can be equally or similarly applied to the sub-image (in addition, it can also be applied to a face). That is, in the present invention, only the difference in terminology is distinguished, and actually the explanation related to the parallel block can be used as the definition of the sub-image, and the explanation related to the sub-image can also be used as the definition of the parallel block.

[0157] Some of the division units mentioned in the above do not necessarily have to be included, and all or some of them can be selectively included according to the coding / decoding setting, and other additional units (for example, a face) can also be supported in addition thereto.

[0158] Further, the coding unit can be partitioned into various sizes by a block partitioning section. At this time, the coding unit can be composed of a plurality of coding blocks according to the color format (e.g., one luminance coding block and two color difference coding blocks, etc.). For the convenience of explanation, it will be assumed that one color component unit is explained. The coding block can take a variable size such as MxM (e.g., M is 4, 8, 16, 32, 64, 128, etc.). Further, according to the partitioning method (e.g., tree structure partitioning, i.e., Quadtree (QT) partitioning, Binary Tree (BT) partitioning, Ternary Tree (TT) partitioning, etc.), the coding block can be partitioned into a variable size of MxN (e.g., M and N are 4, 8, 16, 32, 64, 128, etc.). At this time, the coding block can be a unit on which intra prediction, inter prediction, transform, quantization, entropy coding, etc. are based.

[0159] Although the case where the same size and shape of a plurality of sub-blocks are obtained by the partitioning method (symmetry) is assumed in the present application and is explained as an example, it can also be applied to a case including asymmetric sub-blocks (e.g., the horizontal ratio between partitioned blocks (vertical symmetry) is 1:3 or 3:1 or the vertical ratio between partitioned blocks (horizontal symmetry) is 1:3 or 3:1, etc. in the case of BT partitioning, the horizontal ratio between partitioned blocks (vertical symmetry) is 1:2:1 or the vertical ratio between partitioned blocks (horizontal symmetry) is 1:2:1, etc. in the case of TT partitioning).

[0160] The partitioning (MxN) of the coding block can take a recursive tree structure. At this time, whether to partition or not can be indicated by a partitioning flag. For example, when the partitioning flag of the coding block of the partitioning depth (Depth) k is 0, the coding of the coding block is performed on the coding block of the partitioning depth k, and when the partitioning flag of the coding block of the partitioning depth k is 1, the coding of the coding block is performed in 4 sub-coding blocks (QT partitioning), or 2 sub-coding blocks (BT partitioning), or 3 sub-coding blocks (TT partitioning) of the partitioning depth k+1 according to the partitioning method.

[0161] The above-described sub-coding block is re-set to the coding block k+1 and can be again partitioned into a sub-coding block k+2 by the above-described process, and in the case of QT partitioning, a partitioning flag (e.g., for indicating whether to partition or not) can be supported.

[0162] In the binary tree partitioning, a partition flag and a partition direction flag (horizontal or vertical) can be supported. When more than one partition ratio is supported in the binary tree partitioning (e.g., an additional partition ratio other than the horizontal or vertical ratio of 1:1 is supported, i.e., asymmetric partitioning), a partition ratio flag (e.g., one ratio is selected from a horizontal or vertical ratio candidate group <1:1, 1:2, 2:1, 1:3, 3:1>) can be supported, or other forms of flags (e.g., whether symmetric partitioning is performed or not, when 1 is set, it is symmetric partitioning without additional information, when 0 is set, it is asymmetric partitioning requiring additional information related to the ratio) can be supported.

[0163] In the ternary tree partitioning, a partition flag and a partition direction flag can be supported. When more than one partition ratio is supported in the ternary tree partitioning, the same additional partition information as in the binary tree partitioning described above will be required.

[0164] The above-described example is the partition information generated when only one tree partitioning method is valid, and when multiple tree partitioning methods are valid, the partition information can be constructed as described below.

[0165] For example, when multiple tree partitioning methods are supported, in the case where a partitioning priority order is set in advance, partition information corresponding to the priority order can be first constructed. At this time, when the partition flag corresponding to the priority order is true, additional partition information related to the corresponding partitioning method can be further included, and when it is false (no partitioning is performed), the partition information (partition flag, partition direction flag, etc.) of the partitioning method corresponding to the next order can be constructed.

[0166] Alternatively, when multiple tree partitioning methods are supported, selection information related to the partitioning method can be additionally generated, and the partition information related to the selected partitioning method can be constructed.

[0167] Some of the above-described partition flags can be omitted according to the results of the earlier performed upper or previous partitioning.

[0168] The block partitioning can be performed from the maximum coding block to the minimum coding block. Alternatively, it can also be performed from the minimum partitioning depth 0 to the maximum partitioning depth. That is, the partitioning can be recursively performed until the block size reaches the minimum coding block size or the partitioning depth reaches the maximum partitioning depth. At this time, the partitioning can be performed according to the encoding / decoding setting (e.g., image <slice, parallel block> type , coding mode <Intra / Inter>, chrominance component <y cb cr>The maximum coding block size, the minimum coding block size, and the maximum split depth are adaptively set.

[0169] For example, when the maximum coding block is 128x128, the quad-tree split can be performed in the range of 32x32 to 128x128, the binary-tree split can be performed in the range of 16x16 to 64x64 and the maximum split depth is 3, and the ternary-tree split can be performed in the range of 8x8 to 32x32 and the maximum split depth is 3. Alternatively, the quad-tree split can be performed in the range of 8x8 to 128x128, and the binary-tree split and the ternary-tree split can be performed in the range of 4x4 to 128x128 and the maximum split depth is 3. The former can be a setting on an I picture type (e.g., slice), and the latter can be a setting on a P or B picture type.

[0170] As explained in the above example, the maximum coding block size, the minimum coding block size, the maximum split depth, and the like can be commonly or independently set according to the split mode and the encoding / decoding setting as described above.

[0171] When a plurality of split modes are supported, the split is performed within the block support range of each split mode, and when the block support ranges of the split modes overlap, the priority order information of the split modes can be included. For example, the quad-tree split can be performed before the binary-tree split.

[0172] Alternatively, in the case where the split support ranges overlap, split selection information can be generated. For example, selection information related to the split mode that needs to be performed among the binary-tree split and the ternary-tree split can be generated.

[0173] Further, when a plurality of split modes are supported, whether or not to perform the split performed later can be determined according to the result of the split performed earlier. For example, when the result of the split (in this example, the quad-tree split) performed earlier indicates that the split is performed, the split (in this example, the binary-tree split or the ternary-tree split) performed later can not be performed, and the split is continued after the sub-coding block split by the split performed earlier is again set as a coding block.

[0174] Alternatively, when the result of the earlier performed partitioning indicates that the partitioning is not performed, the partitioning can be performed according to the result of the later performed partitioning. At this time, when the result of the later performed partitioning (in the present example, the binary tree partitioning or the ternary tree partitioning) indicates that the partitioning is performed, the partitioning can be continued after the partitioned sub-coding blocks are again set as coding blocks, and when the result of the later performed partitioning indicates that the partitioning is not performed, the partitioning can not be performed again. At this time, when the result of the later performed partitioning indicates that the partitioning is performed and the plurality of partitioning modes are still supported when the partitioned sub-coding blocks are again set as coding blocks (for example, when the block support ranges of the respective partitioning modes overlap), the earlier performed partitioning can not be performed and only the later performed partitioning can be performed. That is, when the plurality of partitioning modes are supported, if the result of the earlier performed partitioning indicates that the partitioning is not performed, the earlier performed partitioning can not be performed again.

[0175] For example, when the MxN coding block can perform the quad tree partitioning and the binary tree partitioning, the quad tree partitioning flag can be first checked, and when the flag is 1, the coding block is partitioned into 4 sub-coding blocks of (M»1)x(N»1) size, and next the partitioning (the quad tree partitioning or the binary tree partitioning) can be performed after the sub-coding blocks are again set as coding blocks. When the flag is 0, the binary tree partitioning flag can be checked, and when the corresponding flag is 1, the coding block is partitioned into 2 sub-coding blocks of (M»1)xN or Mx(N»1) size, and next the partitioning (the binary tree partitioning) can be performed after the sub-coding blocks are again set as coding blocks. When the flag is 0, the partitioning process is ended and the coding is performed.

[0176] The above-described example illustrates the case where the plurality of partitioning modes are performed, but is not limited thereto, and a combination of the plurality of partitioning modes can be supported. For example, the quad tree / binary tree / ternary tree / quad tree+binary tree / quad tree+binary tree+ternary tree, etc. can be used. At this time, information related to whether the additional partitioning mode is supported can be implicitly determined or explicitly included in a unit such as a sequence, a picture, a sub-picture, a slice, a parallel block, etc.

[0177] In the above-described example, information related to the partitioning such as the size information of the coding block, the support range of the coding block, the maximum partitioning depth, etc. can be included in a unit such as a sequence, a picture, a sub-picture, a slice, a parallel block, etc. or implicitly determined. In other words, the supportable block range can be determined according to the size of the maximum coding block, the supported block range, the maximum partitioning depth, etc.

[0178] The coding block obtained by performing the division by the above-described process can be set to the maximum size of the intra prediction or the inter prediction. That is, in order to perform the intra prediction or the inter prediction, the coding block after the block division can be the division start size of the prediction block. For example, when the coding block is 2M x 2N, the size of the prediction block can be the same size or a relatively smaller size of 2M x 2N, M x N. Or, it can be the size of 2M x 2N, 2M x N, M x 2N, M x N. Or, it can be the same size as the coding block, i.e., the size of 2M x 2N. At this time, the same size of the coding block and the prediction block can mean that the division of the prediction block is not performed, and the prediction is performed using the size obtained by the division of the coding block. That is, it means that the division information for the prediction block is not generated. The above-described setting can also be applied to the transform block, and the transform can be performed in the unit of the divided coding block.

[0179] By the above-described coding / decoding setting, various configurations can be implemented. For example, at least one prediction block and at least one transform block can be obtained based on the coding block (after determining the coding block). Or, one prediction block having the same size as the coding block and at least one transform block can be obtained based on the coding block. Or, one prediction block having the same size as the coding block and one transform block can be obtained. In the above-described example of obtaining at least one block, the division information of each block can be generated, and when one block is obtained, the division information of each block will not be generated.

[0180] The blocks of various sizes obtained by the above-described results, in the form of a square or a rectangle, can be blocks used in the intra prediction, the inter prediction, blocks used when transforming and quantizing the residual components, and blocks used in the filtering process.

[0181] The division unit obtained by dividing the image using the image division unit can perform independent coding / decoding or dependent coding / decoding according to the coding / decoding setting.

[0182] The independent coding / decoding can mean that when coding / decoding is performed on a part of the division unit (or region), the data of the other unit cannot be used as a reference. Specifically, the information used or generated in the process of performing the texture coding and the entropy coding on a part of the unit {such as the pixel value or the coding / decoding information (intra prediction related information, inter prediction related information, and entropy coding / decoding related information, etc.)} will not be mutually referenced and independently coded, and the same applies to the process of performing the texture decoding and the entropy decoding on a part of the unit in the decoder, and the analysis information and the reconstructed information of the other unit will not be mutually referenced.

[0183] Further, the dependency of the encoding / decoding can mean that data of other units can be used as a reference when encoding / decoding is performed on a part of the divided units. Specifically, information used or generated in the process of texture encoding and entropy encoding of a part of the units can be encoded in dependency by mutual reference, and also, in the process of texture decoding and entropy decoding of a part of the units in the decoder, the analysis information and the reconstructed information of other units can be mutually referenced.

[0184] Generally, the divided units (e.g., sub-pictures, parallel blocks, slices, etc.) mentioned in the above can employ the independent encoding / decoding setting. That is, the non-referenceable setting can be employed for the purpose of parallelization. Further, the non-referenceable setting can be employed for the purpose of encoding / decoding performance improvement. For example, when a 360-degree image is divided into a plurality of surfaces on a 3-dimensional space and configured on a 2-dimensional space, a case where the correlation (e.g., image continuity) with an adjacent surface can be lowered according to the surface configuration setting can occur. That is, because the necessity of mutual reference is low when there is no correlation between the surfaces, the independent encoding / decoding setting can be employed.

[0185] Further, the referenceable setting between the divided units can be employed for the purpose of encoding / decoding performance improvement. For example, even in the case where a 360-degree image is divided into surface units, a case where the correlation with an adjacent surface can be high according to the surface configuration setting can occur, and in this case, the dependent encoding / decoding setting can be employed.

[0186] Further, in the present application, the independent or dependent encoding / decoding can not only be applied to spatial regions but also be extended to temporal regions. That is, not only the independent or dependent encoding / decoding can be performed on other divided units existing in the same time as the current divided unit, but also the independent or dependent encoding / decoding can be performed on the divided units existing in different times as the current divided unit (in the present example, even if there is a divided unit existing on the same position in the image corresponding to a different time as the current divided unit, it is assumed as other divided units).

[0187] For example, when a bitstream A including data in which a 360-degree image is encoded in a high quality and a bitstream B including data in which it is encoded in a normal quality are simultaneously transmitted, the decoder can analyze and decode the bitstream A transmitted in the high quality in a region corresponding to an area of interest (e.g., an area in which a user's line of sight is focused <Viewport> or an area desired to be displayed, etc.) and analyze and decode the bitstream B transmitted in the normal quality outside the area of interest.

[0188] Specifically, in the case where the video is divided into a plurality of units (e.g., sub-pictures, tiles, slices, surfaces, etc., in this example, it is assumed that the surfaces are data-processed in the same manner as the tiles or sub-pictures), the data (bitstream A) of the divided units included in the region of interest (or the divided units as long as they overlap with the viewport by one pixel) and the data (bitstream B) of the divided units included outside the region of interest can be decoded.

[0189] Alternatively, a bitstream in which the receipt of encoding the entire video is recorded can be transmitted, and the region of interest can be parsed and decoded from the bitstream at the decoder. Specifically, only the data of the divided units included in the region of interest can be decoded.

[0190] In other words, the entire or part of the video can be obtained by generating a bitstream divided into more than one quality at the encoder and decoding only a specific bitstream at the decoder, and the entire or part of the video can be obtained by selectively decoding each bitstream in each video part. In the above example, the case of a 360-degree video is described, but this is a description that can be applied to general videos.

[0191] When encoding / decoding is performed according to the above-described example, because it is not known that the data will be reconstructed at the decoder (in this example, the decoder does not know the position of the region of interest, and is a case of random access with respect to the region of interest), it is necessary to confirm and perform encoding / decoding with respect to the reference setting, etc. in the time region in addition to the spatial region.

[0192] For example, when the decoder determines which type of decoding is performed with respect to a single divided unit, the current divided unit can perform independent encoding in the spatial region and limited dependent encoding in the time region (e.g., only allowing reference to the divided unit at the same position of other time corresponding to the current divided unit and prohibiting reference to other divided units except for this, because generally, there is no limitation in the time region, this is a comparison with unrestricted dependent encoding).

[0193] Or, when the decoder is to determine which type of decoding is to be performed (for example, in this case, the multiple units are decoded as long as any one of the partition units is included in the region of interest) in multiple partition units (the multiple partition units can be obtained by bundling horizontally adjacent partition units or bundling vertically adjacent partition units, or the multiple partition units can be obtained by bundling horizontally and vertically adjacent partition units), the current partition unit can perform independent or dependent decoding in the spatial region and perform limited dependent encoding in the temporal region (for example, in addition to allowing reference to the partition unit at the same position in other time corresponding to the current partition unit, reference to another part of the partition units is allowed).

[0194] In the present application, the surface is a partition unit configured and shaped, which is generally changed according to the projection format and has no independent encoding / decoding setting, although it has different characteristics from other partition units as described above, but can be regarded as a unit obtained in the image partitioning section in terms of being able to divide the image into multiple regions (and adopting a quadrilateral shape or the like).

[0195] As described above, each partition unit can be subjected to independent encoding / decoding for the purpose of parallelization or the like in the spatial region. However, since independent encoding / decoding cannot refer to other partition units, it can cause a problem of degradation of encoding / decoding efficiency. Therefore, as a step before performing encoding / decoding, a partition unit subjected to independent encoding / decoding can be extended by using (or adding) data of adjacent partition units. Among them, since the partition unit to which data of adjacent partition units is added has more data to be referred to, the encoding / decoding efficiency thereof will also be improved. At this time, since the extended partition unit can refer to data in adjacent partition units when encoding / decoding is performed, it can be regarded as dependent encoding / decoding.

[0196] The above information related to the reference setting between the partition units can be included in the bitstream in units of video, sequence, image, sub-image, slice, parallel block, etc. and transmitted to the decoder, and in the decoder, the setting information transmitted from the encoder can be reconstructed by parsing in the same level unit. In addition, the related information can be transmitted to the bitstream in the form of, for example, Supplement Enhancement Information (SEI) or Metadata, and used after parsing. In addition, a definition agreed in advance in the encoder / decoder can be used, so that the encoding / decoding is performed according to the reference setting without transmitting the above information.

[0197] Figure 7 is an example of dividing an image into a plurality of parallel blocks. Figures 8a to 8i is an example of setting an additional area to each parallel block illustrated in Figure 7 Figures 9a to 9i is an example of setting an additional area to each parallel block illustrated in Figure 7

[0198] When an image is divided into two or more division units (or areas) by an image division section and independent encoding / decoding is performed for each division unit, although there is an advantage that parallel processing can be performed, a problem that encoding performance can be degraded due to a decrease in data that each division unit can refer to can occur. In order to solve the problem as described above, processing can be performed by dependency encoding / decoding setting between division units (in the present example, a parallel block particle will be described, and the same or similar setting can be applied to other units).

[0199] Between division units, independent encoding / decoding is usually performed in a non-referable manner. Therefore, a pre-processing or post-processing process for realizing dependent encoding / decoding can be performed. For example, an extension area can be formed on the outer contour of each division unit before encoding / decoding is performed and data of other division units that need to be referred to can be filled in the extension area.

[0200] Although the method as described above has no difference from a manner of performing independent encoding / decoding except that encoding / decoding is performed after each division unit is extended, because the existing division unit acquires data that needs to be referred to from other division units in advance and refers to the data, it can be understood as an example of dependent encoding / decoding.

[0201] In addition, after encoding / decoding is performed, filtering can be applied using a plurality of division unit data with a boundary between division units as a reference. That is, in the case of applying filtering, it belongs to dependency because other division unit data is used, and in the case of not applying filtering, it can belong to independence.

[0202] In the examples described below, a case in which dependent encoding / decoding is performed by performing an encoding / decoding pre-processing process (extension in the present example) will be described as the center. In addition, in the present application, a boundary between the same division units can be referred to as an internal boundary, and an outer contour of an image can be referred to as an external boundary.

[0203] ​​In an embodiment of the present application, an additional area related to the current parallel block can be set. Specifically, the additional area can be set with at least one parallel block (in the present example, including a case where one image is composed of one parallel block, i.e., a case where it is not divided into two or more division units, and more specifically, although the division unit indicates a unit divided into two or more units, it is assumed that it is recognized as one division unit even in a case where it is not divided).

[0204] For example, the additional area can be set in at least one of the up, down, left, right, and the like of the current parallel block. Here, the additional area can be filled with an arbitrary value. In addition, the additional area can be filled with a part of the data in the current parallel block, i.e., it can be filled with the outer pixels of the current parallel block or by copying the pixels within the current parallel block.

[0205] In addition, the additional area can be filled with the image data of other parallel blocks other than the current parallel block. Specifically, the image data in the parallel block adjacent to the current parallel block can be used, i.e., it can be filled by copying the image data in the parallel block adjacent to the current parallel block in a specific direction among the up, down, left, and right.

[0206] At this time, the size (length) of the acquired image data can take the same value in each direction, or it can take an independent value, which can be determined according to the encoding / decoding setting.

[0207] For example, in the case of Figure 6a , it can be extended to all or a part of the boundary directions of b0 to b8. In addition, m or mi can be extended to all of the boundary directions of the division unit or m or mi can be extended according to the boundary direction i (i is each direction index). m or mi can be applied to all of the division units in the image, or it can be set independently for each division unit.

[0208] At this time, the setting information related to the additional region can be generated. At this time, the setting information related to the additional region can be whether the additional region is supported, whether the additional region is supported for each of the divided units, the additional region form on the entire image (for example, determined according to which of the up, down, left, and right directions of the divided unit is expanded, in this example, setting information commonly applied to all of the divided units within the image), the additional region form on each of the divided units (in this example, setting information applied to individual divided units within the image), the additional region size on the entire image (for example, indicating the degree of expansion in the expanded direction after the additional region form is determined, in this example, setting information commonly applied to all of the divided units within the image), the additional region size on each of the divided units (in this example, setting information independently applied to individual divided units within the image), the method of filling the additional region on the entire image, the method of filling the additional region on each of the divided units, and the like.

[0209] The above-described setting related to the additional region can be determined proportionally according to the color space, or can be set independently. The setting information related to the additional region can be generated on the luminance component, and the additional region setting on the color difference component can be determined implicitly according to the color space. Alternatively, the setting information related to the additional region can be generated on the color difference component.

[0210] For example, when the additional region size of the luminance component is m, the additional region size of the color difference component can be determined as m / 2 according to the color format (in this example, 4:2:0). As another example, when the additional region size of the luminance component is m and the color difference component uses independent setting, the size information of the additional region of the color difference component (in this example, n, which can be commonly used or used according to the direction or the expanded region using n1, n2, n3, and the like) can be generated. As another example, the method of filling the additional region of the luminance component can be generated, and the method of filling the additional region of the color difference component can use the method in the luminance component or generate related information.

[0211] The above-described information related to the additional region setting can be included in a bitstream in units of video, sequence, picture, subpicture, slice, and the like, and transmitted, and at the time of decoding, the related information can be parsed and reconstructed from the above-described units. In the embodiments described later, the case where the additional region is supported will be described as an example.

[0212] Referring to Figure 7 It can be confirmed that one picture is divided into each of the parallel blocks labeled 0 to 8. At this time, the additional region is set to each of the parallel blocks illustrated in FIG. 10A. Figure 7 The result of setting the additional region according to an embodiment of the disclosure to each of the parallel blocks illustrated in FIG. 10A is shown in FIG. 10B. Figures 8a to 8i The result of setting the additional region according to an embodiment of the disclosure to each of the parallel blocks illustrated in FIG. 10A is shown in FIG. 10B.

[0213] In Figure 7 and Figure 8a , the 0th parallel block (size T0_W x TO_H) can be expanded by appending the region of E0_R to the right side and the region of E0_D to the lower side. At this time, the appended regions can be acquired from the neighboring parallel blocks. Specifically, the right side expansion region can be acquired from the 1st parallel block, and the lower side expansion region can be acquired from the 3rd parallel block. In addition, the 0th parallel block can set the appended regions using the right lower neighboring parallel block (4th parallel block). That is, the appended regions can be set in the direction of the remaining internal boundaries (or boundaries between the same division units) other than the external boundaries (or image boundaries) of the parallel blocks.

[0214] In Figure 7 and Figure 8e , because the 4th parallel block (size T4_W x T4_H) has no external boundary, it can be expanded by appending regions to the left side, right side, upper side, and lower side. At this time, the left side expansion region can be acquired from the 3rd parallel block, the right side expansion region can be acquired from the 5th parallel block, the upper side expansion region can be acquired from the 1st parallel block, and the lower side expansion region can be acquired from the 7th parallel block. In addition, the 4th expansion region can also set the appended regions to the upper left, lower left, upper right, and lower right. At this time, the upper left expansion region can be acquired from the 0th parallel block, the lower left expansion region can be acquired from the 6th parallel block, the upper right expansion region can be acquired from the 2nd parallel block, and the lower right expansion region can be acquired from the 8th parallel block.

[0215] In FIG. 8, because the L2 block is a block adjacent to the boundary of the parallel block, there is no data that can be referred to from the left side, upper left, and lower left blocks in principle. However, when the appended regions are set to the 2nd parallel block by applying an embodiment of the present application, the L2 block can be encoded / decoded by referring to the appended regions. That is, the L2 block can refer to the data of the blocks located on the left side and upper left as the appended regions (regions that can be acquired from the 1st parallel block), and can refer to the data of the block located on the lower left as the appended region (a region that can be acquired from the 4th parallel block).

[0216] The data included in the additional area by the above-described embodiment can be included in the current parallel block to perform encoding / decoding. In this case, because the data of the additional area is located on the boundary of the parallel block (in the present example, it refers to the parallel block which is updated or expanded because of the additional area), it is also possible that the encoding performance is degraded because there is no data to be referred to in the encoding process. However, because it is only the part which is added in order to provide a reference to the area of the original parallel block boundary, it can be understood as a form of temporary storage for improving the encoding performance. That is, because it can help to improve the quality performance of the final output image and is an area which is finally removed, the degradation of the encoding performance of the corresponding area does not cause any problem. This can be applied to the embodiments described later with similar or the same purpose.

[0217] Further, referring to Figures 9a to 9i , it can be confirmed that the 360-degree image is changed into a 2-dimensional image through a surface configuration (or reconfiguration) process according to a projection format and the 2-dimensional image is divided into each parallel block (it can also be a surface). At this time, because the 2-dimensional image is composed of one surface when the 360-degree image adopts an equirectangular projection, it can be an example in which one surface is divided into parallel blocks. Further, for the convenience of explanation, it will be assumed that the division of the parallel blocks of the 2-dimensional image is the same as the division of the parallel blocks illustrated in Figure 7

[0218] Among them, the divided parallel blocks can be divided into a parallel block composed of only an internal boundary and a parallel block including at least one external boundary, and an additional area can be set for each parallel block in the manner as shown in Figures 8a to 8i . However, the 360-degree image which is changed into a 2-dimensional image can not have continuity on the actual image even if it is adjacent to each other in the 2-dimensional image, and can have continuity on the actual image even if it is not adjacent (see the description of Figures 5a to 5c ). Therefore, even if a part of the boundary of the parallel block is an external boundary, there can be an area having continuity with the external boundary area of the parallel block within the image. In detail, referring to Figure 9b , although the upper end of the first parallel block is an external boundary of the image, because there can be an area having continuity on the actual image within the same image, an additional area can be set at the upper end of the first parallel block. That is, unlike Figures 8a to 8i , in Figures 9a to 9i , an additional area can also be set in all or a part of the direction of the external boundary of the parallel block.

[0219] Referring to Figure 9e ​, the 4th parallel block is a parallel block that contains only internal boundaries among the parallel block boundaries (in this example, the 4th parallel block). Therefore, the additional area of the 4th parallel block can be set in all of the upper, lower, left, right, upper-left, lower-left, upper-right, and lower-right directions. Among them, the left extension area can be image data acquired from the 3rd parallel block, the right extension area can be image data acquired from the 5th parallel block, the upper extension area can be image data acquired from the 1st parallel block, the lower extension area can be image data acquired from the 7th parallel block, the upper-left extension area can be image data acquired from the 0th parallel block, the lower-left extension area can be image data acquired from the 6th parallel block, the upper-right extension area can be image data acquired from the 2nd parallel block, and the lower-right extension area can be image data acquired from the 9th parallel block.

[0220] Referring to Figure 9a , the 0th parallel block is a parallel block that contains at least one external boundary (left, upper direction). Therefore, the 0th parallel block can contain an additional area that extends in the direction of the external boundary (left, upper, upper-left direction) in addition to the right, lower, and lower-right directions that are spatially adjacent. Among them, the right, lower, and lower-right directions that are spatially adjacent can set the additional area using data of the adjacent parallel block, but the additional area in the direction of the external boundary cannot be so. At this time, the additional area in the direction of the external boundary can be set using data that is not spatially adjacent in the image but has continuity in the actual image. For example, when the projection format of the 360-degree image is an equirectangular projection, the left boundary of the image has continuity with the right boundary of the image in the actual image, and the upper boundary of the image has continuity with the lower boundary of the image in the actual image, the left direction boundary of the 0th parallel block has continuity with the right boundary of the 2nd parallel block, and the upper direction boundary of the 0th parallel block has continuity with the lower boundary of the 6th parallel block. Therefore, in the 0th parallel block, the left extension area can be acquired from the 2nd parallel block, the right extension area can be acquired from the 1st parallel block, the upper extension area can be acquired from the 6th parallel block, and the lower extension area can be acquired from the 3rd parallel block. In addition, in the 1st parallel block, the upper-left extension area can be acquired from the 8th parallel block, the lower-left extension area can be acquired from the 5th parallel block, the upper-right extension area can be acquired from the 7th parallel block, and the lower-right extension area can be acquired from the 4th parallel block.

[0221] Because Figure 9a The L0 block in the above-described case is a block located at a parallel block boundary, and thus data referencable from the left, upper-left, lower-left, upper, and upper-right blocks (a case similar to the U0) can not exist. At this time, even if not spatially adjacent in the 2-dimensional image, there can exist blocks having continuity on the actual image within the 2-dimensional image (or picture). Therefore, as described in the above-described premise, when the projection format of the 360-degree image is an equirectangular projection, the left and lower-left blocks of the L0 block can be acquired from the 2nd parallel block, the upper-left block of the L0 block can be acquired from the 8th parallel block, and the upper and upper-right blocks of the L0 block can be acquired from the 6th parallel block, when the left boundary of the picture has continuity with the right boundary of the picture on the actual image, and the upper boundary of the picture has continuity with the lower boundary of the picture on the actual image.

[0222] Table 1 below is a pseudo code for acquiring data corresponding to the additional region from other regions having continuity.

[0223]

Table 1

[0224] i_pos' = overlap(i_pos, minI, maxI)

[0225] overlap(A, B, C)

[0226] {

[0227] if (A < B) output = (A + C - B + 1) % (C - B + 1)

[0228] else if (A > C) output = A % (C - B + 1)

[0229] else output = A

[0230] }

[0231] Referring to the pseudo code in Table 1, the variable i_pos (corresponding to the variable A) of the overlap function is an input pixel position, i_pos' is an output pixel position, minI (corresponding to the variable B) is a minimum value of a pixel position range, maxI (corresponding to the variable C) is a maximum value of the pixel position range, and i is a position component (horizontal, vertical, etc. in the present example). In the present example, minI can be 0, and maxI can be Pic_width (horizontal width of the picture) - 1 or Pic_height (vertical width of the picture) - 1.

[0232] For example, assume that the vertical width range of the picture (general image) is 0 ~ 47 and the picture is arranged as in FIG. 1. Figure 7 The division is performed in the manner shown. When an additional region of m size needs to be set to the lower side of the parallel block 4 and filled with the data of the upper end of the parallel block 7, it is possible to confirm from which position the data needs to be acquired by the above.

[0233] When the vertical length range of the parallel block 4 is 16 to 30 and an additional region of 4 size needs to be set to the lower side, it is possible to fill the data of the positions of 31, 32, 33, 34 corresponding thereto into the additional region of the parallel block 4. At this time, because min and max in the above formula are 0 and 47, respectively, the output values of 31 to 34 will be their own values, i.e., 31 to 34. That is, the data that needs to be filled into the additional region is the data of the positions of 31 to 34.

[0234] Alternatively, assuming that the horizontal length range of an image (360-degree image, equirectangular projection, images having continuity at both ends) is 0 to 95 and the image is divided as shown in FIG. 6, when an additional region of m size needs to be set to the left side of the parallel block 3 and filled with the data of the right side of the parallel block 5, it is possible to confirm from which position the data needs to be acquired by the above. Figure 7

[0235] When the vertical length range of the parallel block 3 is 0 to 31 and an additional region of 4 size needs to be set to the left side, it is possible to fill the data of the positions of -4, -3, -2, -1 corresponding thereto into the additional region of the parallel block 3. Because the above positions do not exist within the horizontal length range of the image, it is possible to confirm from which position the data needs to be acquired by the above known calculation. At this time, because min and max in the above formula are 0 and 95, respectively, the output values of -4 to -1 will be 92 to 95. That is, the data that needs to be filled into the additional region is the data of the positions of 92 to 95.

[0236] Specifically, when the region of m size is data between 360 degrees and 380 degrees (assuming that the range of pixel value positions is 0 degrees to 360 degrees in the present example), it is possible to understand that it is similar to the case of acquiring data from the region between 0 degrees and 20 degrees by adjusting it to the range inside the image. That is, it is possible to acquire it based on the range of pixel value positions between 0 and Pic_width-1.

[0237] In other words, in order to acquire the data of the additional region, it is possible to confirm the position of the data that needs to be acquired through the overlapping process.

[0238] ​The above example illustrates the case of acquiring a surface from a 360-degree image (except where the image boundaries are continuous at both ends), and assumes that spatially adjacent regions within the image are continuous. However, depending on the projection format (e.g., cube map projection), when there are two or more surfaces and each surface has undergone configuration or reconfiguration, there may be cases where even spatially adjacent areas within the image lack continuity. In the cases described above, the location data that are continuous on the actual image can be confirmed using the surface configuration or reconfiguration information, and additional regions can be generated.

[0239] Table 2 below shows the pseudocode for generating additional regions related to the aforementioned specific segmentation unit using the internal data of that specific segmentation unit.

[0240] Table 2

[0241] i_pos'=clip(i_pos,minI,maxI)

[0242] clip(A,B,C)

[0243] {

[0244] if(A <B)output=B

[0245] else if (A>C)output=C

[0246] else output = A

[0247] }

[0248] Since the meanings of the variables in Table 2 are the same as those in Table 1, their detailed explanations will be omitted here. However, in this example, minI can be the left or top coordinate of a specific division unit, while maxI can be the right or bottom coordinate of each unit.

[0249] For example, when the image is according to such Figure 7 The method shown is used for segmentation. The horizontal length of parallel block 2 ranges from 32 to 47 pixels. When an additional region of size m needs to be added to the right of parallel block 2, the data at positions 48, 49, 50, and 51 can be filled using the data at position 47 (corresponding to the interior of parallel block 2) output by the above formula. That is, according to Table 2, the additional region associated with a specific segmentation unit can be generated by copying the outline pixels of the corresponding segmentation unit.

[0250] In other words, in order to obtain data for the additional region, the required location can be confirmed through a clipping process.

[0251] The detailed configuration of the above Table 1 or Table 2 is not fixed but can be changed. For example, the 360-degree image can be changed to apply an overlapping method in consideration of the configuration (or re-configuration) of the surfaces and the coordinate system characteristics between the surfaces.

[0252] Figure 10 is an example diagram in which the additional region generated in an embodiment of the present application is applied in the encoding / decoding process of other regions.

[0253] Further, since the additional region according to an embodiment of the present application is generated using the image data in other regions, it can correspond to repeated image data. Therefore, in order to prevent unnecessary repeated data from existing, the additional region can be removed after the encoding / decoding is performed. However, before the additional region is removed, it can be considered that the additional region is applied to the encoding / decoding and then removed.

[0254] Referring to Figure 10 It can be confirmed that the additional region B of the division unit J is generated using the region A of the division unit I. At this time, before the generated region B is removed, the region B can be applied to the encoding / decoding (specifically, the reconstruction or correction process) of the region A included in the division unit I.

[0255] Specifically, it is assumed that the division unit I and the division unit J are the 0th parallel block and the 1st parallel block in Figure 7 The rightmost part of the division unit I and the left part of the division unit J have continuity of the image with respect to each other. Among them, the additional region B can be used after being used as a reference in the image encoding / decoding of the division unit J and can be used when A is encoded / decoded. In particular, although A and the region B acquire data from the region A when the additional region is generated, a part of values (including quantization errors) that are not the same can be used for reconstruction in the encoding / decoding process. Therefore, when the division unit I is reconstructed, a part corresponding to the region A can be reconstructed using the image data of the reconstructed region A and the image data of the region B. For example, a part of the region C of the division unit I can be replaced by the average or weighted value of the region A and the region B. This is because there are two or more same regions, and thus the reconstructed image (region C, A of the division unit I is replaced by C) can be obtained using the data of the two regions (the process at this time is named Rec_Process in the drawing).

[0256] Further, a part of the region C included in the division unit I can be replaced using the region A and the region B according to which division unit it is closer to. Specifically, because the image data included in a certain range (for example, M pixel intervals) on the left side in the region C is closer to the division unit I, it can be reconstructed using (or copying) the data in the region A, and because the image data included in a certain range (for example, N pixel intervals) on the right side in the region C is closer to the division unit J, it can be reconstructed using (or copying) the data in the region B. This is expressed as a formula as shown in the following Formula 1.

[0257] [Formula 1]

[0258] C(x, y) = A(x, y), (x, y) e M

[0259] B(x, y), (x, y) e N

[0260] Further, a part of the region C included in the division unit I can be replaced using the region A and the region B according to which division unit it is closer to. Specifically, because the image data included in a certain range (for example, M pixel intervals) on the left side in the region C is closer to the division unit I, it can be reconstructed using (or copying) the data in the region A, and because the image data included in a certain range (for example, N pixel intervals) on the right side in the region C is closer to the division unit J, it can be reconstructed using (or copying) the data in the region B. This is expressed as a formula as shown in the following Formula 1.

[0261] As a formula for setting adaptive weight values to the region A and the region B, the following Formula 2 can be derived.

[0262] [Formula 2]

[0263] C(x, y) = A(x, y) x w + B(x, y) x (1 - w)

[0264] w = f(x, y, k)

[0265] Referring to Formula 2, w is a weight value assigned to the pixels of the A region and the B region as (x, y), and at this time, as an average of the weight values of the A region and the B region, the pixels of the A region are multiplied by the weight value w and the pixels of the B region are multiplied by 1 - w. However, in addition to the average of the weight values, different weight values can be assigned to the region A and the region B, respectively.

[0266] After the use of the additional area according to the above-described instructions is completed, the additional area B can be removed in the resizing process of the division unit J and stored in the memory (decoded picture buffer, DPB) (in this example, it is assumed that the process of setting the additional area is resizing <sizing>The above process can be derived by a part of the above-described embodiments (e.g., by the process of confirming the region flag, the following size information confirmation, the following padding method confirmation, etc.), assuming that the process of increasing is performed in the size adjustment process, and conversely, the process of decreasing is performed in the size readjustment process (which can be derived by the reverse process of the above-described process).

[0267] Further, (in particular, immediately after the end of the encoding / decoding of the corresponding image) it is also possible to directly store into the memory without performing the size readjustment, and then perform the output step (in the present example, it is assumed that the display <display>The size resizing process is performed in the step) for the removal. This can be applied to all or a part of the divided units included in the corresponding image.

[0268] The above-mentioned related setting information can be processed implicitly or explicitly according to the encoding / decoding settings, and in the implicit manner (specifically, based on the characteristics, type, format, etc. of the image or according to other encoding / decoding settings <in this case, the settings related to the additional region>), the determination can be made without generating the related syntax element, and in the explicit manner, the setting related to the removal of the additional region can be adjusted by generating the related syntax element, and as a unit related thereto, the video, sequence, picture, sub-picture, slice, tile, etc. can be included.

[0269] Further, the existing encoding method based on the divided unit can include: 1) a step of dividing an image into one or more tiles (or can be collectively referred to as a divided unit) and generating division information; 2) a step of performing encoding in the divided tile unit; 3) a step of performing filtering with information indicating whether or not loop filtering is allowed for the tile boundary; and 4) a step of storing the filtered tile in a memory.

[0270] Further, the existing decoding method based on the divided unit can include: 1) a step of dividing an image into one or more tiles based on tile division information; 2) a step of performing decoding in the divided tile unit; 3) a step of performing filtering with information indicating whether or not loop filtering is allowed for the tile boundary; and 4) a step of storing the filtered tile in a memory.

[0271] Among them, the third step in the encoding / decoding method is a post-processing step of encoding / decoding, and when filtering is performed, it can be dependent encoding / decoding, and when filtering is not performed, it can be independent encoding / decoding.

[0272] The encoding method of the divided unit according to an embodiment of the present application can include: 1) a step of dividing an image into one or more tiles and generating division information; 2) a step of setting an additional region for at least one divided tile unit and filling the additional region with an adjacent tile unit; 3) a step of performing encoding on the tile unit including the additional region; 4) a step of removing the additional region of the tile unit and performing filtering based on information indicating whether or not loop filtering is allowed for the tile boundary; and 5) a step of storing the filtered tile in a memory.

[0273] Further, the decoding method of the partition unit to which the embodiment of the present application is applied can include: 1) a step of partitioning an image into one or more parallel blocks based on parallel block partitioning information; 2) a step of setting an additional area for the partitioned parallel block unit and filling the additional area with decoding information, pre-set information, or other (adjacent) parallel block units that are reconstructed in advance; 3) a step of performing encoding on the parallel block unit including the additional area using the decoding information received from the encoding apparatus; 4) a step of removing the additional area of the parallel block unit and performing filtering based on information indicating whether or not loop filtering is allowed at the parallel block boundary; and 5) a step of storing the filtered parallel block in a memory.

[0274] In the encoding / decoding method of the partition unit to which the embodiment of the present application is applied as described above, the second step can be a pre-encoding / decoding process (dependent encoding / decoding when the additional area is set, or independent encoding / decoding otherwise). Further, the fourth step can be a post-encoding / decoding process (dependent when filtering is performed, or independent otherwise). In this example, the additional area will be used in the encoding / decoding process, and a process of resizing to the original size of the parallel block will be performed before being stored in the memory.

[0275] First, the encoder partitions an image into a plurality of parallel blocks. Based on implicit or explicit setting, an additional area is set for the parallel block unit and relevant data is acquired from adjacent areas. Next, the updated parallel block unit including the original parallel block and the additional area is encoded. After the encoding is completed, the additional area is removed and filtering is performed according to the loop filtering application setting.

[0276] At this time, different filtering settings described above can be used according to the filling method and the removal method of the additional area. For example, in the case of simple removal, the loop filtering application setting described above can be followed, while when removal is performed using the overlapping area, no filtering can be applied or other filtering settings can be followed. That is, because the distortion phenomenon of the parallel block boundary area can be greatly reduced using the overlapping data, filtering can be performed regardless of the loop filtering application of the parallel block unit or a different setting (e.g., a filter with weak filtering strength at the parallel block boundary, etc.) can be applied inside the parallel block while following the loop filtering application described above. After the above process, it is stored in the memory.

[0277] In the decoder, the image is first divided into a plurality of parallel blocks based on the parallel block division information transmitted from the encoder. Next, information related to the additional region is explicitly or implicitly confirmed, and the updated parallel block encoding information transmitted from the encoder is parsed after the additional region is set. Next, decoding is performed in the updated parallel block unit. After decoding is completed, the additional region is removed and filtering is performed according to the same loop filtering application setting as the encoder. Details related thereto have been described in the encoder section, so detailed description thereof will be omitted here. After the above process, storage into the memory is performed.

[0278] Further, a case where the additional region using the division unit in the encoding / decoding process is not removed but directly stored into the memory can also be considered. For example, in the case of a 360-degree image or the like, there can be a problem where the accuracy of prediction decreases in a part of the prediction process (e.g., inter-picture prediction) according to the surface configuration setting or the like (e.g., it is difficult to accurately search at a position where the surface configuration is not continuous when motion search and compensation are performed). Therefore, the additional region can be stored into the memory and used in the prediction process in order to improve the prediction accuracy. When used in inter-picture prediction, the additional region (or the image including the additional region) can be used as a reference image for performing inter-picture prediction.

[0279] In the encoding method when the additional region is stored, it can include: 1) a step of dividing an image into one or more parallel blocks and generating division information; 2) a step of setting an additional region for at least one parallel block unit divided and filling the additional region with a neighboring parallel block unit; 3) a step of performing encoding on the parallel block unit including the additional region; 4) a step of storing the additional region of the parallel block unit (at this time, application of loop filtering can be omitted); and 5) a step of storing the encoded parallel block into the memory.

[0280] In the decoding method when the additional region is stored, it can include: 1) a step of dividing an image into one or more parallel blocks based on parallel block division information; 2) a step of setting an additional region for the divided parallel block unit and filling the additional region with decoding information, pre-set information, or other (neighboring) parallel block units that are reconstructed in advance; 3) a step of performing encoding on the parallel block unit including the additional region with decoding information received from the encoding device; 4) a step of storing the additional region of the parallel block unit (at this time, loop filtering can be omitted); and 5) a step of storing the decoded parallel block into the memory.

[0281] When storing the additional region, the encoder first divides the image into a plurality of parallel blocks. The additional region is set for the parallel blocks according to the implicit or explicit setting, and the relevant data is acquired from the pre-set region. The pre-set region refers to other regions having relevance according to the surface configuration setting of the 360-degree image, and thus can be a region adjacent to the current parallel block or a region not adjacent thereto. Next, encoding is performed in units of the updated parallel blocks. Since the additional region will be stored after decoding is completed, filtering is not performed regardless of the state of the loop filter application setting. This is because the boundaries of each of the updated parallel blocks do not share the actual parallel block boundaries due to the additional region. The above process is stored in the memory.

[0282] When storing the additional region, the decoder first confirms the parallel block division information transmitted from the encoder and divides the image into a plurality of parallel blocks based thereon. Next, information related to the additional region is confirmed, and the encoding information of the updated parallel blocks transmitted from the encoder is parsed after the additional region is set. Next, decoding is performed in units of the updated parallel blocks. After decoding is completed, the additional region is directly stored in the memory without applying loop filtering.

[0283] Next, the encoding / decoding method of the division unit to which one embodiment of the present application is applied will be described with reference to the accompanying drawings.

[0284] Figures 11 to 12 is a flowchart for explaining the encoding / decoding method of the division unit to which one embodiment of the present application is applied. Specifically, as an example of generating an additional region in each division unit and performing encoding / decoding, the encoding method including the additional region is illustrated in Figure 11 , and the decoding method excluding the additional region is illustrated in Figure 12 . Here, the 360-degree image can perform a pre-processing process (stitching, projection lamp) before the steps in Figure 11 , and perform a post-processing process (rendering, etc.) after the steps in Figure 12 .

[0285] First, refer to Figure 11 After the encoder acquires the input image (step A), the input image is divided into two or more division units by the image division section (at this time, setting information related to the division method can be generated, indicated as step B), next, an additional region is generated for the division unit according to the encoding setting or whether or not the additional region is supported (step C), then encoding is performed on the division unit including the additional region and a bitstream is generated (step D). Furthermore, after the bitstream is generated, it can be determined according to the encoding setting whether or not resizing needs to be performed (or whether or not the additional region is deleted, step E), then the encoded data including or having removed the additional region (the image in steps D or E) is stored in the memory (step E).

[0286] Referring to Figure 12 , the decoder divides the image to be decoded into two or more division units with reference to the setting information related to the division obtained by analyzing the received bitstream (step B), next, the size of the additional region is set for each division unit according to the decoding setting acquired from the received bitstream (step C), then image data including the additional region is obtained by decoding the image data included in the bitstream (step D). Next, a reconstructed image is generated by deleting the additional region (step E), then the reconstructed image is output to the display (step F). At this time, it can be determined according to the decoding setting whether or not the additional region is deleted, and the decoded image or image data (data in steps D or E) is stored in the memory. Furthermore, step F can include a process of reducing to a 360-degree image by surface reconfiguration of the reconstructed image.

[0287] Furthermore, according to Figure 11 or Figure 12 , loop filtering (in this example, it is assumed to be a block filter, but other loop filters can also be applied) can be adaptively performed on the division unit boundary according to whether or not the additional region is removed. Furthermore, loop filtering can be adaptively performed according to whether or not the additional region is allowed to be generated.

[0288] When stored in the memory after the additional region is removed, loop filtering can be explicitly applied or not applied according to the flag of whether or not loop filtering is applied on the division unit boundary (in this example, parallel blocks) such as loop_filter_across_enabled_flag (specifically, the initial state).

[0289] Alternatively, the flag of whether or not loop filtering is applied on the division unit boundary can not be supported, but the determination of whether or not filtering is applied and the filtering setting, etc. are implicitly determined in the manner as in the examples described later.

[0290] Further, even in a case where the image continuity exists between the respective divided units, the image continuity on the boundary between the divided units for which the additional region has been generated can disappear after the additional region is generated for each divided unit. If the loop filtering is applied in this case, it can cause an unnecessary increase in the amount of calculation and a decrease in the encoding performance, and thus the loop filtering can be implicitly not applied.

[0291] Further, according to the surface configuration of the 360-degree image, the divided units adjacent on the 2-dimensional space can not have the image continuity therebetween. When the loop filtering is performed on the boundary between the divided units which do not have the image continuity as described above, it can cause a decrease in the image quality. Thus, for the boundary between the divided units which do not have the image continuity, the loop filtering can be implicitly not performed.

[0292] Further, when the weighting values are assigned to the two regions and the part of the region of the current divided unit is replaced in the manner as described in the description of Figure 10 Further, when the weighting values are assigned to the two regions and the part of the region of the current divided unit is replaced in the manner as described in the description of

[0293] Further, it is also possible to determine the application of the loop filtering (in particular, to the additional part of the corresponding boundary) according to the flag indicating whether the loop filtering is applied or not. When the flag is activated, it is possible to apply the filtering according to the loop filtering setting, condition, etc. applied to the inside of the divided unit, or to apply the filtering defined differently in the loop filtering setting, condition, etc. on the boundary of the divided unit (in particular, additionally using the loop filtering setting, condition, etc. different from the case where it is not the boundary of the divided unit).

[0294] In the above-described embodiments, the case where the additional region is stored in the memory after being removed is assumed, but a part of them can also be in the output step other than this (in particular, both the loop filtering section and the post-filtering section <postfilter>the process performed on the additional region.

[0295] The above example is explained assuming that the additional region is supported in each direction of each partition unit, and when it is supported only in a part of the directions according to the setting of the additional region, only a part of the above content can be applied. For example, the original setting can be applied on the boundary on which the additional region is not supported, and the various cases in the above example can be changed to be applied on the boundary on which the additional region is supported. That is, the above application can be adaptively determined in all or a part of the unit boundaries according to the setting of the additional region.

[0296] The above related setting information can be processed implicitly or explicitly according to the encoding / decoding setting, and when the implicit method is used (in particular, based on the characteristics, type, format, etc. of the image or according to other encoding / decoding settings <in this example, the setting related to the additional region>), the determination can be made without generating the related syntax element, and when the explicit method is used, the adjustment can be made by generating the related syntax element, and the unit related thereto can include a video, a sequence, an image, a sub-image, a slice, a parallel block, etc.

[0297] Next, the determination method of the reference possibility of the partition unit and the additional region will be explained in detail. At this time, when it is possible to be referenced, it is dependent encoding / decoding, and when it is not possible to be referenced, it is independent encoding / decoding.

[0298] The additional region to which an embodiment of the present application is applied can be referenced or restrictedly referenced in the encoding / decoding process of the current image or other images. In particular, the additional region removed before being stored in the memory can be referenced or restrictedly referenced in the encoding / decoding process of the current image. In addition, the additional region stored in the memory can be referenced or restrictedly referenced in the encoding / decoding process of the image different in time from the current image.

[0299] In other words, the reference possibility, range, etc. of the above additional region can be determined according to the encoding / decoding setting. Through a part of the above setting, the additional region of the current image will be stored in the memory after being encoded / decoded, which means that it can be referenced or restrictedly referenced by being included in the reference image of other images. This can be applied to all or a part of the partition units included in the corresponding image. In the case of being explained in the example explained later, the current example can be changed to be applied.

[0300] The above-described setting information related to the possibility of reference to the additional area can be processed according to the encoding / decoding settings, either implicitly or explicitly. In the case of implicit processing (specifically, based on the characteristics, type, format, or the like of the video or according to other encoding / decoding settings <in this example, settings related to the additional area>), the determination can be made without generating a related syntax element, while in the case of explicit processing, the setting related to the possibility of reference to the additional area can be adjusted by generating a related syntax element. As a unit related thereto, a video, sequence, picture, subpicture, slice, tile, or the like can be included.

[0301] Generally, a part of the units in the current video (in this example, a divided unit obtained by the picture division section) can refer to the data of the current unit, but cannot refer to the data of other units. Further, a part of the units in the current video can refer to the data of all the units present in other videos. The above-described description is an example related to the general properties of the units obtained by the picture division section, and additional properties related thereto can be defined.

[0302] Further, a flag for indicating whether or not other divided units within the current video can be referred to and whether or not divided units included in other videos can be referred to can be defined.

[0303] As an example, it can be allowed to refer to a divided unit included in other videos and at the same position as the current divided unit, but it can be limited to refer to a divided unit at a different position from the current divided unit. For example, when a plurality of bitstreams in which the same video is encoded in different encoding settings are transmitted and a bitstream for decoding each area (divided unit) in the video (in this example, it is assumed that decoding is performed in tile units) is selectively decided in the decoder, because it is necessary to limit the possibility of reference between each divided unit in the same space and different spaces, it can be performed so as to allow only the same area in different videos to be referred to in encoding / decoding.

[0304] As an example, it can be allowed to refer to or limited to refer to according to the identifier information related to the divided unit. For example, when the identifier information assigned in the divided unit is the same, it is allowed to refer to, while when it is different, it is not allowed to refer to. At this time, the identifier information can be information for indicating that encoding / decoding has been performed (dependently) in an environment in which it can be referred to each other.

[0305] The above-mentioned related setting information can be processed according to the encoding / decoding setting, either implicitly or explicitly. In the implicit manner, the determination can be made without generating the related syntax element, while in the explicit manner, the processing can be made by generating the related syntax element. As the unit related thereto, it can include video, sequence, picture, sub-picture, slice, tile, etc.

[0306] Figures 13a to 13g is an example diagram for explaining the referenceable area for a specific partition unit. In Figures 13a to 13g , the area illustrated with the thick frame line can represent the referenceable area.

[0307] Referring to Figure 13a , various reference arrows for performing inter-picture prediction can be confirmed. At this time, the C0, C1 blocks represent single-direction inter-picture prediction. The C0 block can obtain the RP0 reference block before the current picture and can obtain the RF0 reference block after the current picture. The C2 block represents bi-directional inter-picture prediction and can obtain the RP1, RF1 reference blocks from the image before the current picture or after the current picture. In the drawing, an example of obtaining one reference block from each of the before direction and the after direction is illustrated, but it is also possible to obtain the reference block from only the before direction or the after direction. The C3 block represents no-direction inter-picture prediction and can obtain the RC0 reference block from within the current picture. In the drawing, an example of obtaining one reference block is illustrated, but it is also possible to obtain two or more reference blocks.

[0308] In the following examples, the reference possibility of the pixel value, the prediction mode information in the inter-picture prediction based on the partition unit will be explained as the center, but it can also be understood to include other encoding / decoding information that can be spatially or temporally referenced (e.g., intra-picture prediction mode information, transform and quantization information, loop filtering information, etc.).

[0309] Referring to Figure 13b , the current picture Currnt(t) is partitioned into two or more tiles, and the block C0 in some of the tiles can obtain the reference blocks P0, P1 by performing single-direction inter-picture prediction. The block C1 in some of the tiles can obtain the reference blocks P3, F0 by performing bi-directional inter-picture prediction. That is, this can be understood as an example of allowing reference to the block at other positions included in other images without restrictions such as position-based restrictions and allowing reference only within the same picture.

[0310] Referring to Figure 13c , the image is divided into two or more parallel block units, and a part of the blocks C1 in a part of the parallel blocks can obtain the reference blocks P2, P3 by performing the unidirectional inter-picture prediction. A part of the blocks C0 in a part of the parallel blocks can obtain the reference blocks P0, P1, F0, F1 by performing the bidirectional inter-picture prediction. A part of the blocks C3 in a part of the parallel blocks can obtain the reference block FC0 by performing the no-directional inter-picture prediction.

[0311] That is, Figure 13b and Figure 13c It can be understood as an example of allowing the reference to the blocks in other positions contained in other images without the restrictions such as the position restrictions and the restriction of allowing the reference only within the same image.

[0312] Referring to Figure 13d , the image is divided into two or more parallel block units, and the blocks C0 in a part of the parallel blocks can obtain the reference block P0 by performing the forward inter-picture prediction, but cannot obtain the reference blocks P1, P2, P3 contained in a part of the parallel blocks. The blocks C4 in a part of the parallel blocks can obtain the reference blocks F0, F1 by performing the backward inter-picture prediction, but cannot obtain the reference blocks F2, F3. A part of the blocks C3 in a part of the parallel blocks can obtain the reference block FC0 by performing the no-directional inter-picture prediction, but cannot obtain the reference block FC1.

[0313] That is, in Figure 13d , the reference can be allowed or restricted according to the division of the image (t-1, t, t+1 in the example) and the encoding / decoding setting of the division unit of the image. Specifically, the reference can be allowed only to the blocks contained in the parallel blocks having the same identifier information as the current parallel block.

[0314] Referring to Figure 13e , the image is divided into two or more parallel block units, and the blocks C0 in a part of the parallel blocks can obtain the reference blocks P0, F0 by performing the bidirectional inter-picture prediction, but cannot obtain the reference blocks P1, P2, P3, F1, F2, F3. That is, Figure 13e It can be an example of allowing the reference only to the parallel blocks in the same position as the parallel block containing the current block.

[0315] Referring to Figure 13f , the image is divided into two or more parallel block units, and the blocks C0 in a part of the parallel blocks can obtain the reference blocks P1, F2 by performing the bidirectional inter-picture prediction, but cannot obtain the reference blocks P0, P2, P3, F0, F1, F3. Figure 13f An example can be one in which information indicating a referenceable parallel block of a current partition unit is included in a bitstream and the referenceable parallel block is confirmed by the information.

[0316] Referring to Figure 13g , the image is partitioned into two or more parallel blocks, and some of the blocks C0 in some of the parallel blocks can obtain reference blocks P0, P3, P5 by performing unidirectional inter-picture prediction, but cannot obtain a reference block P4. Some of the blocks C1 in some of the parallel blocks can obtain reference blocks P1, F0, F2 by performing bidirectional inter-picture prediction, but cannot obtain reference blocks P2, F1.

[0317] Figure 13g is an example in which reference is allowed or restricted according to partitioning of an image (in this example, t-3, t-2, t-1, t, t+1, t+2, t+3), encoding / decoding settings of a partition unit of an image (in this example, it is assumed that determination is made according to identifier information of a partition unit, identifier information of an image unit, whether or not the same region of a partition unit, whether or not a similar region of a partition unit, bitstream information of a partition unit, and the like), and the like. Among them, a referenceable parallel block can have the same or similar position as a current block within an image, and can have the same identifier information as a current block (specifically, on an image unit or a partition unit), and can be the same as a bitstream from which a current parallel block is acquired.

[0318] Figures 14a to 14e is a flowchart for describing the reference possibility of an additional region in a partition unit to which an embodiment of the present application is applied. In Figures 14a to 14e , the region illustrated in a thick frame line represents a referenceable region, and the region illustrated in a dotted line represents an additional region of a partition unit.

[0319] In an embodiment of the present application, the reference possibility of some of the images (other images located before or after in time) can be restricted or allowed. Further, the reference possibility of the entire extended partition unit including an additional region can be restricted or allowed. Further, the reference possibility of only the initial partition unit excluding an additional region can be restricted or allowed. Further, the reference possibility of a boundary between an additional region and an initial partition unit can be restricted or allowed.

[0320] Referring to Figure 14a A portion of the blocks C0 in a portion of the parallel blocks can obtain reference blocks P0, P1 by performing uni-directional inter-picture prediction. A portion of the blocks C2 in a portion of the parallel blocks can obtain reference blocks P2, P3, F0, F1 by performing bi-directional inter-picture prediction. A portion of the blocks C1 in a portion of the parallel blocks can obtain reference block FC0 by performing no-directional inter-picture prediction. Among them, the blocks C0 can obtain reference blocks P0, P1, P2, P3, F0 from the initial parallel block region (original parallel block except for the added region) of a portion of the reference pictures t-1, t+1, while the blocks C2 can obtain reference block F1 from the parallel block region containing the added region of the reference picture t+1 while obtaining reference blocks P2, P3 from the initial parallel block region of the reference picture t-1. At this time, as shown in the reference block F1, a reference block containing the added region and the boundary between the initial parallel block region can be obtained.

[0321] Referring to Figure 14b A portion of the blocks C0, C1, C3 in a portion of the parallel blocks can obtain reference blocks P0, P1, P2 / F0, F2 / F1, F3, F4 by performing uni-directional inter-picture prediction. A portion of the blocks C2 in a portion of the parallel blocks can obtain reference blocks FC0, FC1, FC2 by performing no-directional inter-picture prediction.

[0322] A portion of the blocks C0, C1, C3 can obtain reference blocks P0, F0, F3 from the initial parallel block region of a portion of the reference pictures (t-1, t+1 in this example), can obtain reference blocks P1, x, F4 from the updated parallel block region boundary, and can obtain reference blocks P2, F2, F1 from the outside of the updated parallel block region boundary.

[0323] A portion of the blocks C2 can obtain reference block FC1 from the initial parallel block region of a portion of the reference pictures (t in this example), can obtain reference block FC3 from the updated parallel block region boundary, and can obtain reference block FC0 from the outside of the updated parallel block region boundary.

[0324] Among them, a portion of the blocks C0 can be blocks located in the initial parallel block region, a portion of the blocks C1 can be blocks located in the updated parallel block region boundary, and a portion of the blocks C3 can be blocks located outside the updated parallel block boundary.

[0325] Referring to Figure 14c , the image is divided into two or more parallel block units, an additional area is set for a part of the parallel blocks in a part of the image, no additional area is set for a part of the parallel blocks in a part of the image, and no additional area is set in a part of the image. A part of the blocks C0, C1 in a part of the parallel blocks can obtain reference blocks P2, F1, F2, F3 by performing unidirectional inter-picture prediction, but cannot obtain reference blocks P0, P1, P3, F0. A part of the blocks C2 in a part of the parallel blocks can obtain reference blocks FC1, FC2 by performing non-directional inter-picture prediction, but cannot obtain reference block FC0.

[0326] The part of the block C2 cannot obtain reference block FC0 from the initial parallel block area of the part of the reference image (t in this example), but can obtain reference block FC1 from the updated parallel block area (in the method of filling the part of the additional area, FC0 and FC1 can be the same area, although FC0 cannot be referenced in the initial unit parallel block division, it can be referenced by moving the corresponding area to the current parallel block by the additional area).

[0327] The part of the block C2 can obtain reference block FC2 from the part of the parallel block area of the part of the reference image (t in this example) (although the data in the other parallel blocks of the current image cannot be referenced by default, it is allowed to be referenced if it is set to be referenceable by the identifier information and the like in the above embodiment).

[0328] Referring to Figure 14d , the image is divided into two or more parallel block units and the additional area is set. A part of the blocks C0 in a part of the parallel blocks can obtain reference blocks P0, F0, F1, F3 by performing bidirectional inter-picture prediction, but cannot obtain reference blocks P1, P2, P3, F2.

[0329] The part of the block C0 can obtain reference block P0 from the initial parallel block area (0th parallel block) of the part of the reference image t-1, but cannot obtain reference block P3 from the boundary of the extended parallel block area, nor can it obtain reference block P2 from the outside of the boundary of the extended parallel block area (i.e. the additional area).

[0330] The part of the block C0 can obtain reference block F0 from the initial parallel block area (0th parallel block) of the part of the reference image t+1, can also obtain reference block F1 from the boundary of the extended parallel block area, and can also obtain reference block F3 from the outside of the boundary of the extended parallel block area.

[0331] Referring to Figure 14e , the image is divided into two or more parallel block units and an additional area having at least one size and shape is set. In some of the parallel block units, block C0 can obtain reference blocks P0, P3, P5, F0 by performing unidirectional inter-picture prediction, but cannot obtain reference block P2 located on the boundary between the additional area and the original parallel block. In some of the parallel block units, block C1 can obtain reference blocks P1, F2, F3 by performing bidirectional inter-picture prediction, but cannot obtain reference blocks P4, F1, F5.

[0332] As shown in the above example, pixel values can be the object of reference, and the reference of other encoding / decoding information can be limited.

[0333] As an example, when the prediction unit looks for a set of intra-picture prediction mode candidates that need to be used in intra-picture prediction from spatially adjacent blocks, it can be confirmed whether the partition unit containing the current block can reference the partition unit containing the adjacent block by the method shown in Figures 13a to 14e .

[0334] As an example, when the prediction unit looks for a set of motion information candidates that need to be used in inter-picture prediction from temporally and spatially adjacent blocks, it can be confirmed whether the partition unit containing the current block can reference the partition unit containing the block that is spatially adjacent within the current picture or temporally adjacent to the current picture by the method shown in Figures 13a to 14e .

[0335] As an example, when the loop filter unit looks for loop filter related setting information from adjacent blocks, it can be confirmed whether the partition unit containing the current block can reference the partition unit containing the adjacent block by the method shown in Figures 13a to 14e .

[0336] Figure 15 is an example diagram illustrating blocks contained in a partition unit of a current image and blocks contained in a partition unit of another image.

[0337] Referring to Figure 15 , in this example, the spatially adjacent reference candidate blocks can be the left, top-left, bottom-left, top, top-right blocks centered on the current block. In addition, the temporal reference candidate blocks can be the left, top-left, bottom-left, top, top-right, right, bottom-right, bottom, center blocks of the collocated block located at the same or corresponding position in the image (Different picture) adjacent in time to the current image (Current picture). In Figure 15 , the thick outer frame line represents the boundary line of the partition unit.

[0338] When the current block is M, the spatially adjacent blocks G, H, I, L, and Q can all be referenced.

[0339] When the current block is G, some of the spatially adjacent blocks A, B, C, F, and K can be referenced, while the remaining blocks can be restricted from reference. Whether or not reference is possible can be determined based on the reference-related settings between the partitioning units UC, ULC, and LC of the spatially adjacent blocks and the partitioning unit containing the current block.

[0340] When the current block is S, some blocks among the surrounding blocks s, r, m, w, n, x, t, o, and y that are at the same position as the current block in temporally adjacent images can be referenced, while the remaining blocks can be restricted from reference. Whether or not they can be referenced can be determined based on the reference correlation settings between the segmentation units RD, DRD, and DD of the surrounding blocks at the same position as the current block in temporally adjacent images and the unit containing the current block.

[0341] Depending on the current block position, when there are reference-restricted candidates, the block can be filled using candidates whose priority order is next in the candidate group, or it can be replaced by other candidates adjacent to the reference-restricted candidate.

[0342] For example, when the current block in the prediction within the image is G, and the reference of the left-up fast block is limited and the most likely pattern (MPM) candidate group is formed in the order of PDAEU, since A is unreferenceable, it is possible to form a candidate group by performing a validity check according to the remaining EU order, or to use B or F, which are spatially adjacent to A, to replace A.

[0343] Furthermore, when the current block in the inter-screen prediction is S, the temporally adjacent sitting block is restricted in reference, and the temporal candidate of the skipped pattern candidate group is y, since y is unreferenceable, a candidate group can be formed by performing validity checks in the order of spatially adjacent candidates or a mixture of idle and temporal candidates, or by using t, x, s, which are spatially adjacent to y, to replace y.

[0344] Figure 16 This is a hardware configuration diagram illustrating an image encoding / decoding apparatus to which one embodiment of the present invention is applied.

[0345] See Figure 16 The image encoding / decoding apparatus 200 to which an embodiment of the present application is applied can include at least one processor 210 and a memory 220 storing instructions for instructing the at least one processor 210 to perform at least one step.

[0346] The at least one processor 210 can be a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor for performing a method to which an embodiment of the present application is applied. The memory 120 and the storage 260 can be respectively constituted by at least one of a volatile storage medium and a non-volatile storage medium. For example, the memory 220 can be constituted by at least one of a read only memory (ROM) and a random access memory (RAM).

[0347] In addition, the image encoding / decoding apparatus 200 can further include a transceiver 230 for performing communication through a wireless communication network. In addition, the image encoding / decoding apparatus 200 can further include an input interface device 240, an output interface device 250, a storage device 260, etc. The respective constituent elements included in the image encoding / decoding apparatus 200 can be connected to each other through a bus 270 to perform communication with each other.

[0348] The at least one step can include a step of dividing an encoded image included in a received bitstream into at least one division unit by referring to a syntax element acquired from the bitstream, a step of setting an additional area to the at least one division unit, and a step of decoding the encoded image on the basis of the division unit after the additional area is set.

[0349] The step of decoding the encoded image can include a step of determining a reference block related to a current block to be decoded in the encoded image on the basis of information included in the bitstream for instructing possibility of reference.

[0350] The reference block can be a block included in a position overlapping the additional area set on a division unit including the reference block.

[0351] Figure 17 is an example diagram illustrating an intra prediction mode to which an embodiment of the present application is applied.

[0352] Referring to Figure 17 , it is possible to confirm that there are 35 prediction modes in total, and the 35 prediction modes can be classified into 33 directional modes and 2 non-directional modes (DC, Planar). At this time, the directional modes can be identified by a gradient (e.g., dy / dx) or angle information. The above example can refer to a prediction mode candidate group related to a luma component or a chroma component. Alternatively, the chroma component can support a part of the prediction modes (e.g., DC, Planar, vertical, horizontal, diagonal mode, etc.). In addition, after the prediction mode of the luma mode is determined, the corresponding mode can be included in the prediction mode of the chroma component or a mode derived from the corresponding mode can be included in the prediction mode.

[0353] In addition, the reconstructed block in other color spaces that have been completed encoding / decoding can be applied to the prediction of the current block using the correlation between the color spaces, and the supported prediction modes can be included. For example, the chroma component can generate a prediction block of the current block using the reconstructed block of the luma component corresponding to the current block.

[0354] According to the encoding / decoding setting, the prediction mode candidate group can be adaptively determined. The number of candidate groups can be increased for the purpose of improving prediction accuracy, and the number of candidate groups can be reduced for the purpose of reducing the bit amount in the prediction mode.

[0355] For example, one of the candidate groups such as an A candidate group (67, 65 directional modes and 2 non-directional modes), a B candidate group (35, 33 directional modes and 2 non-directional modes), a C candidate group (19, 17 directional modes and 2 non-directional modes), etc. can be used. In the present invention, unless otherwise explicitly stated, it is assumed that intra prediction is performed using a pre-set one prediction mode candidate group (A candidate group).

[0356] Figure 18 is a first example diagram illustrating the composition of reference pixels used in intra prediction to which an embodiment of the present invention is applied.

[0357] The intra prediction method in the image decoding according to an embodiment of the present application can include a reference pixel constructing step, a prediction block generating step of generating a prediction block using one or more prediction modes with reference to the constructed reference pixels, a step of determining an optimal prediction mode, and a step of encoding the determined prediction mode. In addition, the image decoding apparatus can include a reference pixel constructing unit, a prediction block generating unit, a prediction mode determining unit, and a prediction mode encoding unit for performing the reference pixel constructing step, the prediction block generating step, the prediction mode determining step, and the prediction mode encoding step. The above-described processes can omit a part of them or add other processes, and can also be changed to other orders different from the above-described order.

[0358] In addition, the intra prediction method in the image decoding according to an embodiment of the present application can generate a prediction block of a current block according to a prediction mode obtained from a syntax element received from an image encoding apparatus after constructing reference pixels.

[0359] The size and shape (M x N) of the current block for which the intra prediction is performed can be obtained from the block partition unit, and can be 4 x 4 to 256 x 256. The intra prediction is generally performed in a prediction block unit, but can also be performed in a coding block (or coding unit), a transform block (or transform unit), etc. unit according to the setting of the block partition unit. After confirming the block information, the reference pixel constructing unit can construct reference pixels used in the prediction of the current block. At this time, the reference pixels can be temporarily stored in a memory (for example, an array <array>, 1-dimensional, 2-dimensional array, etc.) are managed, generated and removed in each intra-frame prediction process of the block, and the size of the temporary memory can be determined according to the configuration of the reference pixels.

[0360] The reference pixels can be pixels included in the neighboring blocks (which can be referred to as reference blocks) located at the left, top, top-left, top-right, and bottom-left of the current block, but are not limited thereto, and other configured block candidate groups can also be used in the prediction of the current block. Among them, the neighboring blocks located at the left, top, top-left, top-right, and bottom-left can be blocks selected when encoding / decoding is performed using a raster or zigzag scan, and neighboring blocks included in other positions (for example, right, bottom, bottom-right blocks, etc.) can also be used as reference pixels when the scan order is changed.

[0361] In addition, the reference block can be a block corresponding to the current block in a color space different from the color space including the current block. Among them, when the Y / Cb / Cr format is used, the color space can refer to one of Y, Cb, and Cr. In addition, the block corresponding to the current block can refer to a block having the same position coordinates as the current block or having position coordinates corresponding to the current block according to the color component configuration ratio.

[0362] In addition, for the convenience of description, the reference block of the above-mentioned pre-set position (left, top, top-left, top-right, bottom-left) is described as a block, but can also be composed of multiple sub-blocks according to block partitioning.

[0363] In other words, the neighboring block of the current block can be a reference pixel position for intra-frame prediction of the current block, and other color space regions corresponding to the current block can be additionally added as reference pixel positions according to the prediction mode. In addition to the above examples, the defined reference pixel positions can be determined according to the prediction mode, method, etc. For example, when a prediction block is generated by a block matching method, the reference pixel position can be a region that has already been encoded / decoded before the current block of the current image or a region included in the search range (for example, including the left or right or top-left or top-right of the current block, etc.) of the region that has already been encoded / decoded.

[0364] Referring to Figure 18 The reference pixels used in the intra-frame prediction of the current block (with a size of MxN) can be composed of pixels adjacent to the current block (Ref_L, Ref_T, Ref_TL, Ref_TR, Ref_BL) at the left, top, top-left, top-right, and bottom-left. At this time, Figure 18 The content marked as P(x, y) in the above formula can refer to the pixel coordinates. Figure 18 The content marked as P(x, y) in the above formula can refer to the pixel coordinates.

[0365] In addition, pixels adjacent to the current block can be classified into at least one reference pixel level. For example, they can be classified into pixels ref_0 that are closest to the current block {pixels with a pixel value difference of 1 from the boundary pixels of the current block, p(-1, -1) to p(2M - 1, -1), p(-1, 0) to p(-1, 2N - 1)}. Then, the adjacent pixels {with a pixel value difference of 2 from the boundary pixels of the current block, p(-2, -2) to p(2M, -2), p(-2, -1) to p(-2, 2N)} are ref_1, and then the further adjacent pixels {with a pixel value difference of 3 from the boundary pixels of the current block, p(-3, -3) to p(2M + 1, -3), p(-3, -2) to p(-3, 2N + 1)} are ref_2, and so on. That is, the reference pixels can be classified into multiple reference pixel levels according to the pixel distance from the boundary pixels of the current block.

[0366] In addition, different reference pixel levels can be set for each adjacent block at this time. For example, when using the block adjacent to the upper end of the current block as a reference block, the reference pixels at the ref_0 level can be used, and when using the block adjacent to the upper right end as a reference block, the reference pixels at the ref_1 level can be used.

[0367] Among them, the set of reference pixels usually referred to when performing intra-picture prediction is included in the blocks adjacent to the lower left, left, upper left, upper end, and upper right of the current block, and is the pixels belonging to the ref_0 level (the pixels closest to the boundary pixels). Unless otherwise specified in the following content, the above-mentioned pixels are used as the premise. However, it is also possible to use the set of reference pixels included in some of the adjacent blocks mentioned above, and it is also possible to use the pixels included in two or more levels as the set of reference pixels. Among them, the set of reference pixels or levels can be implicitly determined (pre-set in the encoding / decoding device) or explicitly determined (receiving information for determination from the encoding device).

[0368] Here, the case of supporting up to 3 reference pixel levels will be used as a premise for explanation, but larger values can also be used. The number of reference pixel levels and the number of sets of reference pixels (or what can also be called the reference pixel candidate group) based on the positions of the adjacent blocks that can be referred to can be set differently according to the size, shape, prediction mode, video type <I / P / B, the video at this time is an image, slice, parallel block, etc.>, color component, etc., and the relevant information can be included in units such as sequence, image, slice, parallel block, etc.

[0369] In this invention, the description will assume that lower index values ​​(incrementing by 1 from 0) are assigned starting from the reference pixel level most adjacent to the current block, but it is not limited to this. Furthermore, the reference pixel composition information described later can be generated under the index settings described above (such as assigning shorter bits to the smaller index when selecting from multiple reference pixel sets, etc.).

[0370] Furthermore, when there are two or more supported reference pixel levels, the weighted average value can be applied to each reference pixel contained in the two or more reference pixel levels.

[0371] For example, it is possible to utilize through the location located Figure 18 A prediction block is generated from reference pixels obtained by weighting the pixels in the ref_0 and ref_1 levels. At this time, depending on the prediction mode (e.g., the directionality of the prediction mode), the pixels for which the weighted sum is applied in each reference pixel level can be either integer or fractional pixels. Furthermore, a prediction block can be obtained by assigning weights (e.g., 7:1, 3:1, 2:1, 1:1, etc.) to the prediction block obtained using reference pixels in the 1st reference pixel level and the prediction block obtained using reference pixels in the 2nd reference pixel level, respectively. At this time, a higher weight can be assigned to the prediction block of the reference pixel level more adjacent to the current block.

[0372] Assuming that explicit information related to the composition of reference pixels is generated, it is possible to generate adaptive reference pixel composition indication information (adaptive_intra_ref_sample_enabled_flag in this example) on units such as video, sequence, image, strip, parallel block, etc.

[0373] When the above indication information represents an adaptive reference pixel configuration that allows for adaptive configuration (adaptive_intra_ref_sample_enabled_flag = 1 in this example), adaptive reference pixel configuration information (adaptive_intra_ref_sample_flag in this example) can be generated on units such as images, stripes, parallel blocks, and blocks.

[0374] When the above composition information represents an adaptive reference pixel composition (adaptive_intra_ref_sample_flag = 1 in this example), reference pixel composition related information (such as selection information related to reference pixel level and set, which is intra_ref_idx in this example) can be generated on units such as images, strips, parallel blocks, and blocks.

[0375] At this point, when adaptive reference pixel construction is not allowed or is not adaptive reference pixel construction, reference pixels can be constructed according to predefined settings. For example, the most adjacent pixels in adjacent blocks are usually used to construct reference pixels, but this is not limited to this. Several other cases are also allowed (for example, selecting ref_0 and ref_1 as reference pixel levels and using ref_0 and ref_1 to generate predicted pixel values ​​through weighted summation, i.e., the default case).

[0376] Furthermore, reference pixel composition information (such as selection information related to reference pixel level or set) can be constructed after excluding pre-set information (such as the case where the reference pixel level is pre-set to ref_0) (e.g., ref_1, ref_2, ref_3, etc.), but is not limited to this.

[0377] The above examples illustrate a portion of the situation related to the composition of reference pixels. However, the in-frame prediction settings can be determined by combining various encoding / decoding information. This encoding / decoding information can include, for example, image type, color components, the size and shape of the current block, prediction mode {prediction mode type (directional, non-directional), prediction mode direction (vertical, horizontal, diagonal 1, diagonal 2, etc.)}, and the in-frame prediction settings (in this example, the reference pixel composition settings) can be determined based on the encoding / decoding information of adjacent blocks and the combination of the encoding / decoding information of the current block and adjacent blocks.

[0378] Figures 19a to 19c This is a second example illustration illustrating the configuration of reference pixels to which one embodiment of the present invention is applied.

[0379] See Figure 19a It can be used only Figure 18 The case where the reference pixel layer ref_0 constitutes the reference pixel is confirmed. After constructing the reference pixel using the reference pixel layer ref_0 as the object and utilizing the pixels contained in adjacent blocks (e.g., bottom left, left side, top left, top side, top right), subsequent intra-frame prediction (such as reference pixel generation, reference pixel filtering, reference pixel interpolation, prediction block generation, post-processing filtering, etc.) can be performed, and a portion of the intra-frame prediction process can be adaptively performed according to the reference pixel composition. In this example, the case of using a pre-set reference pixel layer, i.e., not generating setting information related to the reference pixel layer and performing intra-frame prediction using a non-directional mode, will be explained.

[0380] See Figure 19b A case where a reference pixel is constituted using two of the supported reference pixel levels at the same time can be confirmed. That is, intra prediction can be performed after a reference pixel is constituted using pixels included in the level ref_0 and the level ref_1 (or a weighted value average of pixels included in both levels). In this example, a case where a plurality of reference pixel levels are used in advance, that is, a case where setting information related to the reference pixel level is not generated and intra prediction is performed using a part of the directional prediction mode (a direction from the upper right to the lower left in the drawing or the reverse direction thereof) will be described.

[0381] Referring to Figure 19c A case where a reference pixel is constituted using only one of the three supported reference pixel levels can be confirmed. In this example, a case where setting information related to the reference pixel level used is generated because there are a plurality of reference pixel level candidates and intra prediction is performed using a part of the directional prediction mode (a direction from the upper left to the lower right in the drawing) will be described.

[0382] Figure 20 FIG. 3 is a third example diagram illustrating the constitution of a reference pixel to which an embodiment of the present application is applied.

[0383] Figure 20 The reference numeral a in FIG. 1 is a block having a size of 64x64 or more, the reference numeral b is a block having a size of 16x16 or more to less than 64x64, and the reference numeral c is a block having a size of less than 16x16.

[0384] When the block of the reference numeral a is a current block in which intra prediction needs to be performed, intra prediction can be performed using one reference pixel level ref_0 that is most adjacent.

[0385] Further, when the block of the reference numeral b is a current block in which intra prediction needs to be performed, intra prediction can be performed using two reference pixel levels ref_0 and ref_1 that can be supported.

[0386] Further, when the block of the reference numeral c is a current block in which intra prediction needs to be performed, intra prediction can be performed using three reference pixel levels ref_0, ref_1, and ref_2 that can be supported.

[0387] As described in the descriptions of the reference numerals a to c, the number of reference pixel levels that can be supported can be differently set according to the size of a current block in which intra prediction needs to be performed. In Figure 20 In this case, the larger the size of the current block, the more likely the size of the neighboring block is small, which can be a result of performing partitioning based on other image characteristics, and thus, in order to prevent prediction from being performed using a pixel having a large distance in pixel value from the current block, it is assumed that the number of reference pixel levels supported is smaller as the size of the block is larger, but other variations including the opposite case are also allowed.

[0388] Figure 21 FIG. 4 is a fourth exemplary diagram illustrating a configuration of reference pixels according to an embodiment of the present application.

[0389] Referring to Figure 21 It can be confirmed that the current block in which intra prediction is performed is in a rectangular shape. If the current block is in a rectangular shape that is horizontally and vertically asymmetric, the number of supported reference pixel levels adjacent to the longer horizontal side boundary surface in the current block can be set to be larger, and the number of supported reference pixel levels adjacent to the shorter vertical side boundary surface in the current block can be set to be smaller. In the drawing, it can be confirmed that the number of reference pixel levels adjacent to the horizontal boundary surface of the current block is set to 2, and the number of reference pixel levels adjacent to the vertical boundary surface of the current block is set to 1. The pixel adjacent to the shorter vertical side boundary surface in the current block can have a problem of decreased accuracy due to a generally relatively large distance from the pixel included in the current block because of the larger horizontal length. Thus, the number of supported reference pixel levels adjacent to the shorter vertical side boundary surface is set to be smaller, but a setting opposite thereto can also be used.

[0390] Further, the reference pixel levels to be used in prediction can be differently set according to the type of the intra prediction mode or the position of the neighboring block adjacent to the current block. For example, a directional mode using pixels included in a block adjacent to the upper end, the upper right end of the current block as reference pixels can use two or more reference pixel levels, and a directional mode using pixels included in a block adjacent to the left end, the lower left end of the current block as reference pixels can use only one most adjacent reference pixel level.

[0391] Further, when the prediction blocks generated through the respective reference pixel levels in the plurality of reference pixel levels are the same or similar to each other, the generation of the setting information of the reference pixel levels can cause a problem of unnecessarily generated data.

[0392] For example, when the distribution characteristics of the pixels constituting each reference pixel level are similar or the same to each other, it is possible to generate similar or the same prediction blocks regardless of which reference pixel level is used, and thus there is no need to generate data for the selected reference pixel level. At this time, the distribution characteristics of the pixels constituting the reference pixel level can be determined by comparing the average or dispersion value of the pixels with a predetermined threshold value.

[0393] That is, when the reference pixel levels based on the finally determined intra prediction mode are the same or similar to each other, the reference pixel level can be selected using a predetermined method (for example, selecting the most adjacent reference pixel level).

[0394] At this time, the decoder can receive the intra prediction information (or intra prediction mode information) from the encoding apparatus and determine whether or not information for selecting the reference pixel level needs to be received based on the received information.

[0395] The above-described various examples are described for the case where the reference pixels are constituted using a plurality of reference pixel levels, but are not limited thereto, and various modified examples can be used, and can be used in combination with other additional configurations.

[0396] The reference pixel constituting section of the intra prediction can include, for example, a reference pixel generating section, a reference pixel interpolation section, a reference pixel filtering section, etc., and can include all or some of the above-described configurations. Among them, a block including pixels that can be used as reference pixels can be referred to as a reference candidate block. In addition, the reference candidate block can be usually a neighboring block adjacent to the current block.

[0397] The reference pixel constituting section can determine whether or not to use the pixels included in the reference candidate block as reference pixels, based on the reference pixel availability set for the reference candidate block.

[0398] With regard to the above-described reference pixel availability, it can be determined that it cannot be used when at least one of the following conditions is satisfied. For example, when the reference candidate block satisfies at least one of the conditions such as being located outside the image boundary, not being included in the same division unit (for example, a slice, a parallel block, etc.) as the current block, not having completed encoding / decoding, having a limitation on its use in the encoding / decoding setting, etc., it can be determined that the pixels included in the corresponding reference candidate block cannot be used for reference. At this time, if none of the above-described conditions is satisfied, it can be determined that it can be used.

[0399] Further, the use of the reference pixels can be limited according to the encoding / decoding setting. For example, when a flag (e.g., constrained_intra_pred_flag) for limiting the reference of the reference candidate block is activated, it can be limited that the pixels included in the corresponding reference candidate block cannot be used as the reference pixels. In order to effectively perform the encoding / decoding even when the error occurs due to various external factors including the communication environment, the above flag can be applied when the reference candidate block is a block reconstructed by referring to the image different in time from the current block.

[0400] Among them, in the case where the flag for limiting the reference is activated (e.g., when constrained_intra_pred_flag = 0 in the I picture type or the P or B picture type), all the pixels in the reference candidate block can be used as the reference pixels. Further, in the case where the flag for limiting the reference is activated (e.g., when constrained_intra_pred_flag = 1 in the P or B picture type), it can be determined whether the reference is possible according to whether the reference candidate block is encoded by the intra prediction or the inter prediction. That is, when the reference candidate block is encoded by the intra prediction, the corresponding reference candidate block can be referred to regardless of whether the above flag is activated, and when the reference candidate block is encoded by the inter prediction, it can be determined whether the corresponding reference candidate block is referable according to whether the above flag is activated.

[0401] Further, the reconstructed block located at the position corresponding to the current block in the other color space can be used as the reference candidate block. At this time, it can be determined whether the reference is possible according to the encoding mode of the reference candidate block. For example, when the current block belongs to a part of the color difference components (Cb, Cr), it can be determined whether the reference is possible according to the encoding mode of the block (= reference candidate block) located at the position corresponding to the current block on the luminance component (Y) and having completed the encoding / decoding. This can be an example corresponding to the case where the encoding mode is determined independently according to the color space.

[0402] The flag for limiting the reference can be a setting applied to a part of the image type (e.g., P or B slice / parallel block type, etc.).

[0403] By referring to the pixel availability, the reference candidate block can be classified into a case of full availability, a case of partial availability, and a case of no availability. In the case of no availability, the reference pixels on the unavailable candidate block positions can be filled or generated.

[0404] In the case of availability of the reference candidate block, the pixels on the predetermined positions of the current block (or the pixels adjacent to the current block) can be stored into the reference pixel memory of the current block. At this time, the pixel data on the corresponding block positions can be directly copied or stored into the reference pixel memory through a reference pixel filtering process or the like.

[0405] In the case of no availability of the reference candidate block, the pixels obtained through the reference pixel generation process can be included into the reference pixel memory of the current block.

[0406] In other words, the reference pixels can be constructed in the state of availability of the reference pixel candidate block, and the reference pixels can be generated in the state of no availability of the reference pixel candidate block.

[0407] The method of filling the reference pixels into the predetermined positions in the unavailable reference candidate block is as follows. First, the reference pixels can be generated using arbitrary pixel values. Among them, the arbitrary pixel values are specific pixel values included in the pixel value range, and can be the minimum value, the maximum value, the median value, or a value derived from the above values, which can be used in the pixel value adjustment process based on the bit depth or the pixel value adjustment process based on the image pixel value range information. Among them, the method of generating the reference pixels using the arbitrary pixel values can also be applicable in the case of no availability of all the reference candidate blocks.

[0408] Next, the reference pixels can be generated using the pixels included in the blocks adjacent to the unavailable reference candidate block. Specifically, the pixels included in the adjacent blocks can be filled into the predetermined positions in the unavailable reference candidate block by the methods of extrapolation, interpolation, or copying. At this time, the method of performing the copying or the extrapolation or the like can be the clockwise direction or the counterclockwise direction, and can be determined according to the encoding / decoding setting. For example, the reference pixel generation direction within the block can follow a predetermined direction or a direction adaptively determined according to the position of the unavailable block.

[0409] Figures 22a to 22b is an exemplary diagram illustrating the method of filling the reference pixels into the predetermined positions in the unavailable reference candidate block.

[0410] Referring to Figure 22a A method of padding pixels included in the unavailable reference candidate block in the reference pixels constituted by one reference pixel level can be confirmed. In Figure 22a when the neighboring block adjacent to the upper end of the current block is the unavailable reference candidate block, the reference pixels included in the neighboring block adjacent to the upper end (indicated as <1>) can be generated by performing the extrapolation in the clockwise direction or the linear extrapolation on the reference pixels included in the neighboring block adjacent to the upper end of the current block.

[0411] Further, in Figure 22a when the neighboring block adjacent to the left side of the current block is the unavailable reference candidate block, the reference pixels included in the neighboring block adjacent to the left side (indicated as <2>) can be generated by performing the extrapolation in the counterclockwise direction or the linear extrapolation on the reference pixels included in the neighboring block adjacent to the upper left end of the current block (corresponding to the available block). At this time, by performing the extrapolation in the clockwise direction or the linear extrapolation, the reference pixels included in the neighboring block adjacent to the lower left end of the current block can be utilized.

[0412] Further, in Figure 22a a part of the reference pixels included in the neighboring block adjacent to the upper side of the current block (indicated as <3>) can be generated by performing the interpolation or the linear interpolation on the available reference pixels on both sides. That is, even in a case where a part of the reference pixels included in the neighboring block is unavailable but not all of them are unavailable, the setting can be made, and in this case, the unavailable reference pixels can be padded using the neighboring pixels of the unavailable reference pixels.

[0413] Referring to Figure 22b a method of padding the unavailable reference pixels when a part of the reference pixels constituted by a plurality of reference pixel levels is unavailable can be confirmed. Referring to Figure 22b when the neighboring block adjacent to the upper right end of the current block is the unavailable reference candidate block, the pixels included in the 3 reference pixel levels included in the corresponding neighboring block (indicated as <1>) can be generated in the clockwise direction using the pixels included in the neighboring block adjacent to the upper end of the current block (corresponding to the available block).

[0414] Further, in Figure 22b when the neighboring block adjacent to the left side of the current block is the unavailable reference candidate block and the neighboring blocks adjacent to the upper left end or the lower left end of the current block are the available reference candidate blocks, the reference pixels of the unavailable reference candidate block can be generated by padding the reference pixels of the available reference candidate blocks in the clockwise direction, the counterclockwise direction, or both directions.

[0415] At this time, the unusable reference samples in each reference sample level can be generated using the samples in the same reference sample level, but this does not exclude the method of using the samples in different reference sample levels. For example, in the case of FIG. 6, the reference samples (indicated as <3>) included in the three reference sample levels in the neighboring block adjacent to the upper end of the current block are assumed to be unusable reference samples. At this time, the samples included in the reference sample level ref_0 most adjacent to the current block and the reference sample level ref_2 farthest from the current block can be generated using the usable reference samples included in the same reference sample level. In addition, the samples included in the reference sample level ref_1 at a distance of one sample from the current block can be generated not only using the samples included in the same reference sample level ref_1 but also using the samples included in different reference sample levels ref_0 and ref_2. At this time, the unusable reference samples can be filled in by twice linear interpolation or the like using the usable reference samples on both sides. Figure 22b

[0416] The above example is an example of generating reference samples when a plurality of reference sample levels are constituted by reference samples and a part of the reference candidate blocks is unusable. Alternatively, adaptive reference sample constitution (adaptive_intra_ref_sample_flag = 0 in the present example) can be set according to the encoding / decoding setting (for example, the case where at least one reference candidate block is unusable or all of the reference candidate blocks are unusable, etc.). That is, the reference samples can be constituted according to the predetermined setting without generating any additional information.

[0417] The reference sample interpolation section can generate reference samples in decimal units by linear interpolation of the reference samples. In the present application, the description is assumed to be a part of the process of the reference sample constitution section, but the constitution can also be adopted to be included in the prediction block generation section, and it can be understood that the process is performed before the generation of the prediction block.

[0418] In addition, the description is assumed to be an independent process separate from the reference sample filtering section described later, but the process can also be adopted to be integrated into one process. This is also a constitution provided to solve the case where the reference samples are distorted due to the increase in the number of times of applying the filter to the reference samples when a plurality of filters are applied by the reference sample interpolation section and the reference sample filtering section.

[0419] ​The reference pixel interpolation process is not performed in some prediction modes (e.g., horizontal, vertical, some diagonal modes <e.g., Diagonal down right, Diagonal down left, Diagonal upright, etc. 45 degree angle modes>, non-directional mode, color mode, color copy mode, etc., i.e., modes in which fractional interpolation is not needed when generating a prediction block), and can be performed only in other prediction modes (modes in which fractional interpolation is needed when generating a prediction block) other than the same.

[0420] The interpolation precision (e.g., 1, 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, etc. pixel unit) can be determined according to the prediction mode (or directionality of the prediction mode). For example, a 45-degree angle prediction mode does not need an interpolation process, but a 22.5-degree or 67.5-degree angle prediction mode needs a 1 / 2 pixel unit difference. As described above, at least one interpolation precision and a maximum interpolation precision can be determined according to the prediction mode.

[0421] For reference pixel interpolation, only one interpolation filter (e.g., a 2-tap linear interpolation filter) can be used, or a filter selected from a plurality of interpolation filter candidates (e.g., a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc.) can be used according to encoder / decoder settings. At this time, the interpolation filter can be distinguished according to differences in the number of filter taps (i.e., the number of pixels to which the filter is applied), filter coefficients, etc.

[0422] Interpolation can be performed in stages in order from lower precision to higher precision (e.g., 1 / 2 → 1 / 4 → 1 / 8), or can be performed at once. The former case refers to performing interpolation based on integer unit pixels and fractional unit pixels (pixels for which interpolation has already been completed using lower precision than the pixels to be interpolated), and the latter case refers to performing interpolation based on integer unit pixels.

[0423] When one of a plurality of filter candidates is used, filter selection information can be explicitly generated or mode-determined, or can be determined according to encoder / decoder settings (e.g., interpolation precision, size, shape, prediction mode, etc. of a block). At this time, the unit of explicit generation can be a video, a sequence, an image, a slice, a parallel block, a block, etc.

[0424] For example, an 8-tap Kalman filter can be applied to the reference pixels of an integer unit when an interpolation accuracy of 1 / 4 or more (1 / 2, 1 / 4) is used, a 4-tap Gaussian filter can be applied to the reference pixels of an integer unit and the interpolated reference pixels of a unit of 1 / 4 or more when an interpolation accuracy of less than 1 / 4 and 1 / 16 or more (1 / 8, 1 / 16) is used, and a 2-tap linear filter can be applied to the reference pixels of an integer unit and the interpolated reference pixels of a unit of 1 / 16 or more when an interpolation accuracy of less than 1 / 16 (1 / 32, 1 / 64) is used.

[0425] Alternatively, an 8-tap Kalman filter can be applied to a block of 64x64 or more, a 6-tap Wiener filter can be applied to a block of less than 64x64 and 16x16 or more, and a 4-tap Gaussian filter can be applied to a block of less than 16x16.

[0426] Alternatively, a 4-tap cubic filter can be applied to a prediction mode having an angle difference of less than 22.5 degrees in a vertical or horizontal mode, and a 4-tap Gaussian filter can be applied to a prediction mode having an angle difference of 22.5 degrees or more.

[0427] Further, a plurality of filter candidate groups can be composed of a 4-tap cubic filter, a 6-tap Wiener filter, and an 8-tap Kalman filter in some encoding / decoding settings, and can be composed of a 2-tap linear filter and a 6-tap Wiener filter in some encoding / decoding settings.

[0428] Figures 23a to 23c is an example diagram illustrating a method of performing interpolation in a reference pixel composed according to an embodiment of the present application on the basis of a fractional pixel unit.

[0429] Referring to Figure 23a A method of interpolating a pixel of a fractional unit in a case where one reference pixel level (ref_i) is used as a reference pixel can be confirmed. Specifically, interpolation can be performed by applying a filter (filter function is marked as int_func_1D) to pixels adjacent to an interpolation target pixel (marked with x) in a case where a filter is used for an integer unit pixel in the present example. Here, since one reference pixel level is used as a reference pixel, interpolation can be performed using adjacent pixels included in the same reference pixel level as the interpolation target pixel x.

[0430] Referring to Figure 23b This allows for the verification of methods to obtain interpolated pixels in fractional units when using more than two reference pixel levels (ref_i, ref_j, ref_k) as reference pixels. Figure 23b In this context, when performing reference pixel interpolation at the reference pixel level ref_j, it is possible to append interpolation of the target pixels using other reference pixel levels ref_k and ref_i with fractional units. Specifically, by interpolating adjacent pixels a... k ~h k a j ~h j a i ~h i Perform filtering (interpolation process, function int_func_1D) to obtain the interpolated pixel (x) respectively. j The position of the pixel x) and the position of the pixel x contained in other reference pixel levels that corresponds to the interpolated pixel (the corresponding position in each reference pixel level according to the direction of the prediction mode). k x i And the obtained first-order interpolated pixel x k x j x i Append filtering is performed (this can be a non-interpolation process such as weighted averaging of [1, 2, 1] / 4, [1, 6, 1] / 8, etc.) to finally obtain the final interpolated pixel x at the reference pixel level ref_j. In this example, it is assumed that the pixel x at other reference pixel levels corresponding to the interpolated pixel is... k x i The case where fractional pixels can be obtained through interpolation is explained.

[0431] In the above examples, the case where first-order interpolated pixels can be obtained by filtering at each reference pixel level, and the final difference pixel can be obtained by performing additional filtering on the first-order interpolated pixels, was explained. However, it is also possible to obtain the final difference pixel by filtering adjacent pixels a at multiple reference pixel levels. k ~h k a j ~h j a i ~h i The final interpolated pixels are obtained in one step by filtering.

[0432] exist Figure 23b Among the 3 reference pixel levels supported in the middle, the level actually used as a reference pixel can be ref_j. That is, in order to interpolate one reference pixel level composed of reference pixels, other reference pixel levels included in a candidate group (for example, not composed of reference pixels means that the pixels in the corresponding reference pixel level are not applicable to prediction, but in the above case, the corresponding pixels are referenced when performing interpolation, so to speak, they can also belong to the use case) can be used.

[0433] Referring to Figure 23c The case where both of the supported 2 reference pixel levels are used as reference pixels is illustrated. The final interpolation pixel x can be obtained by using the pixels adjacent to the decimal unit position where interpolation is to be performed in each of the supported reference pixel levels (d i , d j , e i , e j in this example) as input pixels and performing filtering on the adjacent pixels. Figure 23b At this time, the method of obtaining the final interpolation pixel x by performing additional filtering on the 1st interpolation pixel after obtaining the 1st interpolation pixel on each reference pixel level as shown in

[0434] The above example is not limited to the reference pixel interpolation process, but can also be understood as a process combined with other processes of intra prediction (for example, a reference pixel filtering process, a prediction block generation process, etc.).

[0435] Figures 24a to 24b is a 1st example diagram for explaining an adaptive reference pixel filtering method to which an embodiment of the present application is applied.

[0436] Generally, the main purpose of the reference pixel filtering unit can be to perform smoothing by using a low-pass filter {Low-pass Filter, for example, 3-tap, 5-tap filter such as [1, 2, 1] / 4, [2, 3, 6, 3, 2] / 16, etc.}, but other types of filters (for example, a high-pass filter, etc.) can also be used according to the filter application purpose {for example, sharpening, etc.}. In the present application, the case where distortion generated in the encoding / decoding process is reduced by performing filtering with the purpose of smoothing will be explained as the center.

[0437] The reference pixel filtering can be determined according to the encoding / decoding setting. However, since the batch filtering can not be applied, the local characteristics of the image can not be reflected, and the filtering based on the local characteristics of the image can be more advantageous for the improvement of the encoding performance. The characteristics of the image can be determined according to the image type, the color component, the quantization parameter, the encoding / decoding information of the current block (e.g., the size, the shape, the partition information, the prediction mode, etc. of the current block), the encoding / decoding information of the neighboring block, and the combination of the encoding / decoding information of the current block and the neighboring block. In addition, the characteristics of the reference pixel (e.g., the dispersion, the standard deviation, the flat region, the discontinuous region, etc. of the reference pixel region) can be determined.

[0438] Referring to Figure 24a When the classification (class 0) according to a part of the encoding / decoding setting (e.g., the block size range A, the prediction mode B, the color component C, etc.) is applied, the filtering can not be applied, and when the classification (class 1) according to a part of the encoding / decoding setting (e.g., the prediction mode A of the current block, the prediction mode B of the neighboring block, etc.) is applied, the filtering can be applied.

[0439] Referring to Figure 24b When the classification (class 0) according to a part of the encoding / decoding setting (e.g., the size A of the current block, the size B of the neighboring block, the prediction mode C of the current block, etc.) is applied, the filtering can not be applied, when the classification (class 1) according to a part of the encoding / decoding setting (e.g., the size A of the current block, the shape B of the current block, the size C of the neighboring block, etc.) is applied, the filtering can be performed using the filter A, and when the classification (class 2) according to a part of the encoding / decoding setting (e.g., the parent block A of the current block, the parent block B of the neighboring block, etc.) is applied, the filtering can be performed using the filter B.

[0440] Accordingly, the application of the filtering, the type of the filter, the encoding of the filter information (explicit / implicit), the number of times of the filtering, etc. can be determined according to the size, the prediction mode, the color component, etc. of the current block and the neighboring block, and the type of the filter can be classified according to the number of taps, the difference in the filter coefficient, etc. At this time, when the number of times of the filtering is two or more, the same filter can be applied multiple times or different filters can be applied.

[0441] The above example can be a case where the reference pixel filtering is pre-set according to the characteristics of the image. That is, it can be a case where the filter-related information is determined implicitly. However, when the judgment of the characteristics of the image as described above is not accurate, it can adversely affect the encoding efficiency, and thus this part must be considered.

[0442] To prevent the above-mentioned situation from occurring, explicit setting can be performed on the reference pixel filter. For example, information related to whether or not the filter is applied can be generated. At this time, when there is only one filter, filter selection information can not be generated, and when there are multiple filter candidate groups, filter selection information can be generated.

[0443] The above-mentioned example illustrates the implicit setting and the explicit setting related to the reference pixel filter, and a hybrid approach can be adopted in which the explicit setting is determined in some cases and the implicit setting is determined in other cases. The implicit meaning is that the information related to the reference pixel filter (such as information on whether or not the filter is applied, filter type information) can be derived from the decoder.

[0444] Figure 25 FIG. 2 is a second example diagram for explaining an adaptive reference pixel filtering method to which an embodiment of the present application is applied.

[0445] Referring to Figure 25 The category can be classified by the image characteristics that can be confirmed by using the encoding / decoding information, and the reference pixel filter can be adaptively performed according to the classified category.

[0446] For example, the filter is applied when classified as category 0, and filter A is used when classified as category 1. Category 0 and category 1 can be an example of the implicit reference pixel filter.

[0447] In addition, when classified as category 2, the filter can not be applied or filter A can be applied, and at this time, the generated information can be information on whether or not the filter is applied, but filter selection information can not be generated.

[0448] In addition, when classified as category 3, filter A or filter B can be applied, and at this time, the generated information can be filter selection information, and the application of the filter can be an example of unconditionally performed. That is, when classified as category 3, it can be understood as a case where the filter must be performed but the filter type needs to be selected.

[0449] In addition, when classified as category 4, the filter can not be applied or filter A or filter B can be applied, and at this time, the generated information can be information on whether or not the filter is applied and filter selection information.

[0450] In other words, the explicit or implicit process can be determined according to the category, and when performed by the explicit process, each reference pixel filter-related candidate group setting can be adaptively constituted.

[0451] Regarding the above-mentioned category, the following example can be considered.

[0452] First, for a block of 64x64 or more, one of <filter off>, <filter on-filter A>, <filter on+filter B>, and <filter on+filter C> can be implicitly determined according to the prediction mode of the current block. At this time, the candidate added in consideration of the distribution characteristics of the reference pixels can be <filter on+filter C>. That is, filter A, filter B, or filter C can be applied in the case of filter on.

[0453] Further, for a block of less than 64x64 and 16x16 or more, one of <filter off>, <filter on+filter A>, and <filter on+filter B> can be implicitly determined according to the prediction mode of the current block.

[0454] Further, for a block of less than 16x16, one of <filter off>, <filter on+filter A>, and <filter on+filter B> can be selected according to the prediction mode of the current block. At this time, the mode can be determined as <filter off> in some prediction modes, one of <filter off> and <filter on+filter A> can be explicitly selected in some prediction modes, and one of <filter off> and <filter on+filter B> can be explicitly selected in some prediction modes.

[0455] As an example of the configuration related to a plurality of reference pixel filters, when the reference pixels acquired in each filter (including the case of filter off in the present example) are the same or similar, the generation of reference pixel filter information (for example, reference pixel filter permission information, reference pixel filter information, and the like) can result in the generation of unnecessary duplicate information. For example, when the distribution characteristics of the reference pixels acquired in each filter (for example, the values acquired by averaging, dispersion, and the like of the respective reference pixels, and the threshold values <threshold>When the characteristics judged by the comparison are the same or similar, the reference pixel filter-related information can be omitted. When the reference pixel filter-related information is omitted, filtering can be applied in a predetermined method (e.g., filter off). The decoder can judge whether the reference pixel filter-related information needs to be received in the same manner as the encoder after receiving the intra prediction information, and can determine whether to receive the reference pixel filter-related information based on the judgment.

[0456] In the case of assuming generation of explicit information related to reference pixel filtering, indication information (adaptive_ref_filter_enabled_flag in this example) that allows adaptive reference pixel filtering can be generated on a unit of video, sequence, picture, slice, parallel block, and the like.

[0457] When the above indication information represents that adaptive reference pixel filtering is allowed (adaptive_ref_filter_enabled_flag = 1 in this example), adaptive reference pixel filtering allowance information (adaptive_ref_filter_flag in this example) can be generated on a unit of picture, slice, parallel block, block, and the like.

[0458] When the above allowance information represents that adaptive reference pixel filtering is allowed (adaptive_ref_filter_flag = 1 in this example), reference pixel filter-related information (e.g., reference pixel filter selection information, and the like, ref_filter_idx in this example) can be generated on a unit of picture, slice, parallel block, block, and the like.

[0459] At this time, in the case where adaptive reference pixel filtering is not allowed or cannot be applied, a filtering operation can be performed on the reference pixel in accordance with a predetermined setting (as described above, the filter application or not, the filter type, and the like are determined in advance from the video encoding / decoding information, and the like).

[0460] Figures 26a to 26b is an explanatory diagram illustrating a case where one reference pixel level is used in reference pixel filtering in an embodiment of the present application.

[0461] Referring to Figure 26a It can be confirmed that interpolation is performed by applying filtering (referred to as a smt_func_1 function) to an object pixel d and pixels a, b, c, e, f, g adjacent to the object pixel d included in the reference pixel level ref_i.

[0462] Figure 26a Generally, it can be a case where the successive filtering is applied, but it can also be a case where the multiple filtering is applied. For example, the 2nd filtering can be applied to the reference pixels (a*, b*, c*, d*, etc. in this case) obtained by applying the 1st filtering.

[0463] Referring to Figure 26b , the filtered pixel e* (referred to as a function smt_func_2) can be obtained by performing the linear interpolation proportional to the distance (for example, the distance z from a) on the pixels located at both sides with the object pixel e as the center. Among them, the pixels located at both sides can be the pixels located at both ends of the neighboring pixels within the blocks consisting of the upper block, the left block, the upper block + upper right block, the left block + lower left block, the upper left block + upper block + upper right block, the upper left block + left block + lower left block, the upper left block + left block + upper block + lower left block + upper right block of the current block. Figure 26b It can be the reference pixel filtering performed according to the reference pixel distribution characteristics.

[0464] In Figures 26a to 26b , it is illustrated a case where the pixels in the same reference pixel level as the filtered object reference pixel are used for the reference pixel filtering. At this time, the type of the filter used in the reference pixel filtering can be the same or different according to the reference pixel level.

[0465] In addition, in the case of using a plurality of reference pixel levels, not only the pixels in the same reference pixel level but also the pixels in different reference pixel levels can be used when performing the reference pixel filtering in some of the reference pixel levels.

[0466] Figure 27 is an explanatory diagram illustrating a case where a plurality of reference pixel levels are used in the reference pixel filtering to which one embodiment of the present application is applied.

[0467] Referring to Figure 27 , first, the filtering can be performed using the pixels included in the same reference pixel level on the reference pixel levels ref_k and ref_i, respectively. That is, the filtered pixel d k * can be obtained by performing the filtering (defined as a function smt_func_1D) on the object pixel d k and the neighboring pixels a k to g k on the reference pixel level ref_k, and the filtered pixel d i * can be obtained by performing the filtering (defined as a function smt_func_1D) on the object pixel d i and the neighboring pixels a i to g i on the reference pixel level ref_i.

[0468] Further, when performing the reference pixel filtering on the reference pixel level ref_j, not only the same reference pixel level ref_j but also the pixels included in the other reference pixel levels ref_i, ref_k which are spatially adjacent to the reference pixel level ref_j can be used. Specifically, the interpolated pixel d j can be obtained by applying filtering (defined as a function smt_func_2D) to the spatially adjacent pixels c k , d k , e k , c j , e j , c i , d i , e i around the object pixel d j * (i.e., it can be a filter having a 3x3 square mask). However, it is not limited to the 3x3 square shape, and a filter having a mask such as a 5x2 rectangular shape (b k , c k , d k , e k , f k , b j , c j , e j , f j ), a 3x3 diamond shape (d k , c j , e j , d i ), a 5x3 cross shape (d k , b j , c j , e j , f j , d i , etc. can be used around the object pixel.

[0469] where the reference pixel level is composed of pixels included in the neighboring blocks adjacent to the current block and close to the boundary of the current block as shown in the above Figures 18 to 2 2, etc. In consideration of the above aspect, the filtering using the pixels included in the same reference pixel level in the reference pixel level ref_k and the reference pixel level ref_i, a 1-dimensional mask shape filter using the pixels horizontally or vertically adjacent to the interpolation object pixel can be applied. However, the interpolation pixel of the reference pixel d j in the reference pixel level ref_j can be obtained by applying a 2-dimensional mask shape filter using all the pixels spatially up / down / left / right adjacent.

[0470] Further, in each reference pixel level, 2nd reference pixel filtering can be applied to the reference pixels to which 1st reference pixel filtering has been applied. For example, 1st reference pixel filtering can be performed using the reference pixels included in each reference pixel level ref_k, ref_j, ref_i, and next, in the reference pixel levels (referred to as ref_k*, ref_j*, ref_i*) to which 1st reference pixel filtering has been applied, reference pixel filtering can be performed not only using the respective reference pixel levels but also using the reference pixels in the other reference pixel levels.

[0471] The prediction block generating section can generate a prediction block according to at least one intra prediction mode (which can be simply referred to as a prediction mode) and can use the reference pixels based on the prediction mode. At this time, the prediction block can be generated by extrapolating or interpolating or DC copying the reference pixels according to the prediction mode. Among them, extrapolation can be applied to the directional mode in the intra prediction mode, and the rest can be applied to the non-directional mode.

[0472] Further, when copying the reference pixels, one or more prediction pixels can be generated by copying one reference pixel to a plurality of pixels in the prediction block, and one or more prediction pixels can also be generated by copying one or more reference pixels, and the number of copied reference pixels can be equal to or less than the number of copied prediction pixels.

[0473] Further, the prediction block is usually generated for the prediction of one intra prediction mode, but a final prediction block can also be generated by applying a weighted value or the like to a plurality of prediction blocks obtained after obtaining the plurality of prediction blocks. Among them, the plurality of prediction blocks can refer to the prediction blocks obtained according to the reference pixel levels.

[0474] In the prediction mode determination section of the encoding apparatus, a process for selecting the best mode from a plurality of prediction mode candidates is performed. Generally, the best mode in terms of encoding cost can be determined using a rate-distortion technique that predicts distortion {e.g., distortion between a current block and a reconstructed block (Distortion), Sum of Absolute Difference (SAD), Sum of Square Difference (SSD), etc.} of a block and a bit amount generated in the prediction mode. A prediction block generated based on the prediction mode determined by the above process can be transmitted to the subtraction operation section and the addition operation section (at this time, since the decoding apparatus can acquire information indicating the best prediction mode from the encoding apparatus, the process of selecting the best prediction mode can be omitted).

[0475] The prediction mode encoding section of the encoding apparatus can encode the best intra prediction mode selected by the prediction mode determination section. At this time, the index information indicating the best prediction mode can be directly encoded, or prediction information related to the prediction mode (e.g., a difference value between the predicted prediction mode index and the prediction mode index of the current block) can be encoded after the best prediction mode is predicted by a prediction mode obtainable from other blocks in the periphery, etc. Among them, the former case can be applied to the chrominance component, and the latter case can be applied to the luminance component.

[0476] When the best prediction mode of the current block is predicted and encoded, the prediction value (or prediction information) of the prediction mode can be referred to as the most probable mode (MPM). At this time, the most probable mode (MPM) refers to the prediction mode with the highest probability of becoming the best prediction mode of the current block, and can be composed of a prediction mode of a block spatially adjacent (e.g., a left, an upper, a left upper, a right upper, a left lower block, etc.) or a prediction mode of a prediction mode (e.g., a mean (DC), a planar, a vertical, a horizontal, a diagonal mode, etc.) set in advance. Among them, the diagonal mode refers to a diagonal up right, a diagonal down right, a diagonal down left, and can be a mode corresponding to the 2nd, 18th, 34th mode in the table of FIG. 1. Figure 17

[0477] ​Furthermore, patterns derived from the prediction patterns contained in the most probable patterns (MPMs), i.e., the most probable pattern (MPM) candidate group, can be added to the most probable pattern (MPM) candidate group. In directional patterns, prediction patterns with an index interval equal to a preset value between them and the prediction patterns contained in the most probable pattern (MPM) candidate group can be added to the most probable pattern (MPM) candidate group. For example, a pattern included in the most probable pattern (MPM) candidate group is... Figure 17 In the case of pattern 10, the derived pattern can be equivalent to patterns 9, 11, 8, 12, etc.

[0478] The above example is equivalent to the case where the most likely mode (MPM) candidate group consists of multiple modes. The composition of the most likely mode (MPM) candidate group (e.g., the number of predicted modes included in the most likely mode (MPM) and the priority order of its composition) is determined according to the encoding / decoding settings (e.g., predicted mode candidate group, image type, block size, block shape, etc.) and can include at least one mode composition.

[0479] It is possible to set the priority order of prediction patterns included in the Most Probable Pattern (MPM) candidate group. It is possible to determine the order of prediction patterns included in the MPM candidate group according to the set priority order, and to complete the formation of the MPM candidate group when the number of added prediction patterns reaches a preset number. The priority order can be set as prediction patterns of blocks spatially adjacent to the current block to be predicted, preset prediction patterns, and patterns derived from earlier prediction patterns included in the MPM candidate group, but is not limited to this.

[0480] Specifically, in spatially adjacent blocks, priority can be set in the order of left-top-bottom-top-right-top-left blocks. In pre-defined prediction patterns, priority can be set in the order of mean (DC)-planar-vertical-horizontal pattern. Next, the index values ​​of prediction patterns included in the most likely pattern (MPM) candidate group can be... Figure 17 The predicted patterns obtained by performing addition operations (+1, -1, etc., integer values) on the predicted pattern number are included in the most probable pattern (MPM) candidate group. As one example above, the priority order can be set in the following order: left-top-mean (DC)-planar-bottom-top-right-top-left-(spatial adjacent block pattern)+1-(spatial adjacent block pattern)-1-horizontal-vertical-diagonal.

[0481] In the above example, the priority order of the most probable mode (MPM) candidate set is fixed, but the priority order can be adaptively determined according to the shape, size, etc. of the block.

[0482] When encoding the prediction mode of the current block using the most probable mode (MPM), information (e.g., most_probable_mode_flag) related to whether the prediction mode coincides with the most probable mode (MPM) can be generated.

[0483] When coinciding with the most probable mode (MPM) (e.g., most_probable_mode_flag = 1), most probable mode (MPM) index information (e.g., mpm_idx) can be additionally generated according to the composition of the most probable mode (MPM). For example, when the most probable mode (MPM) is composed of one prediction mode, additional most probable mode (MPM) index information can not be generated, and when composed of multiple prediction modes, index information corresponding to the prediction mode of the current block in the most probable mode (MPM) candidate set can be generated.

[0484] When not coinciding with the most probable mode (MPM) (e.g., most_probable_mode_flag = 0), non-most probable mode (non-MPM) index information (e.g., non_mpm_idx) corresponding to the prediction mode of the current block in the remaining prediction mode candidate set excluding the most probable mode (MPM) candidate set from the supported intra prediction modes can be generated, which can be an example of the case where the non-most probable mode (non-MPM) is composed of one set.

[0485] When the non-most probable mode (non-MPM) candidate set is composed of multiple sets, information related to which set the prediction mode of the current block is included in can be generated. For example, when the non-most probable mode (non-MPM) is composed of two sets of A and B, and the prediction mode of the current block coincides with the prediction mode of the A set (e.g., non_mpm_A_flag = 1), index information corresponding to the prediction mode of the current block can be generated in the candidate set of the A set, and when not coinciding (e.g., non_mpm_A_flag = 0), index information corresponding to the prediction mode of the current block can be generated in the remaining prediction mode candidate set (or the candidate set of the B set). As in the above example, the non-most probable mode (non-MPM) can be composed of multiple sets, and the number of sets can be specified according to the prediction mode candidate set. For example, it can be one when the prediction mode candidate set is 35 or less, and two in other cases.

[0486] At this time, the specific A group can be composed of modes that are determined to have a high probability of being consistent with the prediction mode of the current block after the most probable mode (MPM) candidate group. For example, the next prediction mode that is not included in the most probable mode (MPM) candidate group can be included in the A group or the directional mode with a certain interval can be included in the A group.

[0487] As in the above example, when the non-MPM is composed of a plurality of groups, the effect of reducing the mode coding bit amount can be achieved in the case where the number of prediction modes is large and the prediction mode of the current block is not consistent with the MPM.

[0488] In encoding (or decoding) the prediction mode of the current block using the MPM, a binarization table that is suitable for each prediction mode candidate group (e.g., the MPM candidate group, the non-MPM candidate group, etc.) can be generated individually, and different binarization methods can be applied individually according to each candidate group.

[0489] In the above example, the terms such as the MPM candidate group, the non-MPM candidate group, etc. are only a part of the terms used in the present application and are not limited thereto. Specifically, only the information indicating which category belongs to when the current intra prediction mode is classified into a plurality of categories and the mode information within the corresponding category, terms such as the 1st MPM candidate group and the 2nd MPM candidate group can be used instead.

[0490] Figure 28 is a block diagram for explaining an intra prediction mode encoding / decoding method to which an embodiment of the present application is applied.

[0491] Referring to Figure 28 , first acquires the mpm_flag (S10), next confirms whether or not it is consistent with the 1st most probable mode (MPM) (indicated by the mpm_flag) (S11), and when it is consistent, acquires the most probable mode (MPM) index information (mpm_idx) (S12). When it is not consistent with the most probable mode (MPM), acquires the rem_mpm_flag (S13), next confirms whether or not it is consistent with the 2nd most probable mode (MPM) (indicated by the rem_mpm_flag) (S14), and when it is consistent, acquires the 2nd most probable mode (MPM) index information (rem_mpm_idx) (S16). When it is not consistent with the 2nd most probable mode (MPM), acquires the index information (rem_mode_idx) of the candidate group constituted by the remaining prediction modes (S15). In the present example, a case where the index information generated depending on whether or not it is consistent with the 2nd most probable mode (MPM) is expressed using the same syntax element is explained, but other mode coding settings (for example, binarization) can also be applied, and different index information described above can also be set to be processed.

[0492] In the image decoding method according to an embodiment of the present application, the intra prediction can be constituted in the following manner. The intra prediction of the prediction section can include a prediction mode decoding step, a reference pixel constituting step, and a prediction block generating step. Further, the image decoding apparatus can include a prediction mode decoding section, a reference pixel constituting section, and a prediction block generating section for executing the prediction mode decoding step, the reference pixel constituting step, and the prediction block generating step. The above-described processes can omit a part thereof or add other processes, and can also be changed to other order different from the above-described order.

[0493] Since the reference pixel constituting section and the prediction block generating section of the image decoding apparatus can function in the same manner as those of the image encoding apparatus, detailed description thereof will be omitted here, and the prediction mode decoding section can use the manner used in the prediction mode encoding section in reverse.

[0494] Next, the intra prediction based on the reference pixel constitution of the decoding apparatus will be explained in detail. Figures 29 to 31 Various embodiments of the intra prediction based on the reference pixel constitution of the decoding apparatus will be explained. Among them, the explanation related to the reference pixel level support and the reference pixel filtering method explained in the above description in conjunction with the drawings should be interpreted as being applicable to the decoding apparatus as well, and detailed description thereof will be omitted in order to prevent redundant explanation.

[0495] Figure 29 is a first example diagram for explaining the bitstream constitution of the intra prediction based on the reference pixel constitution.

[0496] In the first example in FIG. 1, the case where multiple reference pixel levels are supported, at least one of the supported reference pixel levels is used as a reference pixel, multiple candidate groups related to reference pixel filtering are supported and one filter is selected therefrom is taken as a premise. Figure 29 After the pixel candidate groups are constituted using multiple reference pixel levels in the encoder (in the state where the reference pixel generation process has been completed in the present example), the reference pixel is constituted using at least one reference pixel level, and then reference pixel filtering and reference pixel interpolation are applied. At this time, multiple candidate groups related to reference pixel filtering are supported.

[0497] Next, a process for selecting the best mode from the prediction mode candidate groups is executed, and after the best prediction mode is determined, a prediction block based on the corresponding mode is generated and passed to the subtraction operation section, and then the encoding process of the intra prediction related information is executed. In the present example, the case where the reference pixel level and the reference pixel filtering are determined implicitly from the encoding information is taken as a premise.

[0498] In the decoder, the intra prediction related information (e.g. prediction mode, etc.) is reconstructed and after the prediction block based on the reconstructed prediction mode is generated and passed to the subtraction operation section. At this time, the case where the reference pixel level and the reference pixel filtering for generating the prediction block are determined implicitly is taken as a premise.

[0499] Referring to FIG. 2, a bitstream (S20) can be constituted using one intra prediction mode (intra_mode). At this time, the reference pixel level ref_idx and the reference pixel filtering category ref_filter_idx supported (or used) in the current block can be determined implicitly from the intra prediction mode (determined as Category A, B, respectively, S21-S22). At this time, the encoding / decoding information (e.g. image type, color component, block size and shape, etc.) can be additionally considered.

[0500] Figure 29 is a second example diagram for explaining the bitstream constitution of the intra prediction based on the reference pixel.

[0501] Figure 30 In the second example in FIG. 2, the case where multiple reference pixel levels are supported and one of the supported multiple reference pixel levels is used as a reference pixel is taken as a premise. In addition, the case where multiple candidate groups related to reference pixel filtering are supported and one filter is selected therefrom is taken as a premise. The difference from

[0502] In the second example in FIG. 2, the case where multiple reference pixel levels are supported and one of the supported multiple reference pixel levels is used as a reference pixel is taken as a premise. In addition, the case where multiple candidate groups related to reference pixel filtering are supported and one filter is selected therefrom is taken as a premise. The difference from Figure 30 Figure 29 is that the information related to the selection is explicitly generated by the encoding device.​​

[0503] After determining the support of multiple reference pixel levels in the encoder, a reference pixel is constructed using one of the reference pixel levels, and then reference pixel filtering and reference pixel interpolation are applied. At this time, multiple filtering methods related to reference pixel filtering are supported.

[0504] When performing the process of deciding the best prediction mode of the current block in the encoder, the process of selecting the best reference pixel level in each prediction mode and the process of selecting the best reference pixel filtering can also be added. After deciding the best prediction mode of the current block and the reference pixel level and the reference pixel filtering, the prediction block generated on this basis is passed to the subtraction operation unit and the encoding process of the intra prediction related information is performed.

[0505] In the decoder, the intra prediction related information (such as prediction mode and reference pixel level, reference pixel filtering information, etc.) is reconstructed and after generating the prediction block using the reconstructed information, it is passed to the subtraction operation unit. At this time, the reference pixel level and the reference pixel filtering used for the prediction block generation comply with the settings determined according to the information transmitted from the encoder.

[0506] Referring to Figure 30 , the decoder confirms the best prediction mode of the current block through the intra prediction mode information (intra_mode) included in the bitstream (S30), and confirms whether multiple reference pixel levels are supported (multi_ref_flag) (S31). When multiple reference pixel levels are supported, the reference pixel level selection information (ref_idx) is confirmed (S32), so that the reference pixel level that can be used in the intra prediction is determined. When multiple reference pixel levels are not supported, the process of obtaining the reference pixel level selection information (ref_idx) (S32) can be omitted.

[0507] Next, whether the support of adaptive reference pixel filtering is confirmed (adap_ref_smooth_flag) (S33), and when the support of adaptive reference pixel filtering is confirmed, the filtering method of the reference pixel is determined through the reference pixel filter information (ref_filter_idx) (S34).

[0508] Figure 31 is a third example diagram for explaining the bitstream construction of the intra prediction based on the reference pixel construction.

[0509] In the third example diagram in Figure 31 , the support of multiple reference pixel levels and the use of one of the multiple reference pixel levels are assumed. In addition, the support of multiple candidate groups related to reference pixel filtering and the selection of one filter from among them are assumed. As in the first example diagram, the support of multiple reference pixel levels and the use of one of the multiple reference pixel levels are assumed. Figure 30 The difference is that the selection information is adaptively generated.

[0510] After the reference pixels are constructed using one of the multiple reference pixel levels supported by the encoder, reference pixel filtering and reference pixel interpolation are applied. At this time, multiple filters related to reference pixel filtering are supported.

[0511] When performing the process for selecting the best mode from the multiple prediction mode candidates, the process for selecting the best reference pixel level in each prediction mode and the process for selecting the best reference pixel filtering can also be considered. After the best prediction mode and the reference pixel level and reference pixel filtering are determined, the prediction block is generated based thereon and passed to the subtraction operation unit, and then the encoding process for the intra prediction related information is performed.

[0512] At this time, the repeatability of the generated prediction block is confirmed, and when it is the same or similar to the prediction block obtained using other reference pixel levels, the selection information related to the best reference pixel level is omitted and the preset reference pixel level is used. At this time, the preset reference pixel level can be the level most adjacent to the current block.

[0513] For example, the repeatability or not can be judged based on the difference value (distortion value) between the prediction block generated by ref_0 in FIG. 8 and the prediction block generated by ref_1. Figure 19c When the above difference value is less than a preset threshold value, it is determined that the prediction block has repeatability, otherwise it is determined that the prediction block has no repeatability. At this time, the above threshold value can be adaptively determined according to the quantization parameter and the like.

[0514] In addition, the best reference pixel filtering information also confirms the repeatability of the prediction block, and when it is the same or similar to the prediction block obtained by applying other reference pixel filtering, the reference pixel filtering information is omitted and the preset reference pixel filtering is applied.

[0515] For example, the repeatability or not of the prediction block obtained by filter A (3-tap filter in this example) and the prediction block obtained by filter B (5-tap filter in this example) is judged based on the difference value between them. At this time, the difference value can also be compared with a preset threshold value, and when it is smaller, it is determined that the prediction block has repeatability. When the prediction block has repeatability, the prediction block can be generated by the preset reference pixel filtering method. Among them, the preset reference pixel filtering can be a filtering method with less number of taps or lower complexity, including the case of omitting the filtering usage.

[0516] The intra prediction related information (e.g. prediction mode and reference pixel level, reference pixel filter information, etc.) is reconstructed in the decoder and passed to the subtraction operation section after generating the prediction block. At this time, the reference pixel level information and the reference pixel filter used to generate the prediction block comply with the settings determined from the information transmitted from the encoder, and the decoder can comply with the pre-set method when the redundancy exists after directly confirming the redundancy (without passing through the syntax element).

[0517] Referring to Figure 31 , the decoder first confirms the intra prediction mode information (intra_mode) of the current block (S40), and confirms whether the support of multiple reference pixel levels (multi_ref_flag) (S41). When supporting multiple reference pixel levels, the redundancy check of the prediction block based on the supported multiple reference pixel levels is performed (represented by the ref_check process, S42), and when the redundancy check result is that the prediction block has no redundancy (redund_ref = 0, S43), the selection information of the reference pixel level (ref_idx) is referred from the bit stream (S44) and the optimal reference pixel level is determined.

[0518] Next, the support of adaptive reference pixel filtering (adap_ref_smooth_flag) is confirmed (S45), and when supporting adaptive reference pixel filtering, the redundancy check of the prediction block of the supported multiple reference pixel filtering methods (represented by the ref_check process, S46). When there is no redundancy of the prediction block (redund_ref = 0, S47), the selection information of the reference pixel filtering method (ref_filter_idx) is referred from the bit stream (S48) and the optimal reference pixel filtering method is determined.

[0519] At this time, the redund_ref in the drawing is a value for indicating the redundancy check result, and when it is 0, it means that there is no redundancy.

[0520] In addition, the decoder can perform the intra prediction using the pre-set reference pixel level and the pre-set reference pixel filtering method when the prediction block has redundancy.

[0521] Figure 32 is a flowchart illustrating an image decoding method supporting multiple reference pixel levels according to an embodiment of the present application.

[0522] Referring to Figure 32 The image decoding method supporting multiple reference pixel levels can include a step of confirming whether multiple reference pixel levels are supported through a bitstream (S100), a step of determining a reference pixel level to be used in a current block by referring to syntax information included in the bitstream when the multiple reference pixel levels are supported (S110), a step of constituting a reference pixel using a pixel included in the determined reference pixel level (S120), and a step of performing an intra prediction of the current block using the constituted reference pixel (S130).

[0523] The method or apparatus can further include a step of confirming whether an adaptive reference pixel filtering method is supported through the bitstream after the step of confirming whether the multiple reference pixel levels are supported.

[0524] The method or apparatus can further include a step of constituting a reference pixel using a preset reference pixel level when the multiple reference pixel levels are not supported after the step of confirming whether the multiple reference pixel levels are supported.

[0525] The method according to the present application can be realized in the form of program instructions executable by various computing means and recorded in a computer-readable medium. The computer-readable medium can include program instructions, data files, data structures, etc., alone or in combination. The program instructions recorded in the computer-readable medium can be program instructions specially designed for the present application or program instructions commonly available to computer software practitioners.

[0526] Examples of the computer-readable medium can include specially configured hardware devices for storing and executing program instructions, such as read-only memory (ROM), random access memory (RAM), flash memory, etc. Examples of the program instructions include not only machine codes generated by a compiler, but also high-level language codes executable in a computer using an interpreter, etc. The hardware devices can be constituted by at least one software module for performing actions according to the present application, and vice versa.

[0527] In addition, all or a part of the above-described method or apparatus can be combined or separated.

[0528] The above-described preferred embodiments according to the present application are described in conjunction with the accompanying drawings, but those skilled in the art should understand that various modifications and changes can be made to the present application within the scope of the ideas and regions of the present application described in the appended claims.< / threshold> < / array> < / postfilter> < / display> < / sizing> < / y>

Claims

1. A method for decoding a video signal using an image decoding device, the method comprising: Obtain the coefficients of the current block of the current image from the bitstream; Perform inverse quantization on the coefficients of the current block; as well as The current block is decoded on a sub-block basis by performing an inverse transform on the inverse-quantized coefficients. Specifically, based on the segmentation information obtained from the bitstream, the sub-block is obtained by segmenting the current block. The number of sub-blocks is 2. The segmentation information includes a segmentation ratio flag, which indicates one of a plurality of segmentation ratios among the sub-blocks. The multiple segmentation ratios mentioned include 1:

3. Specifically, the current image is segmented into multiple sub-images based on sub-image segmentation information. The sub-image segmentation information includes quantity information, position information, and size information. The quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the top-left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. Both the location information and the size information are represented in the largest coded unit. The plurality of sub-images includes a first sub-image and a second sub-image. A portion of the right boundary of the first sub-image is adjacent to the left boundary of the second sub-image, and The upper boundary of the first sub-image and the upper boundary of the second sub-image are discontinuous, while the lower boundary of the first sub-image and the lower boundary of the second sub-image are continuous.

2. The method of claim 1, wherein one of independence decoding or dependency decoding is applied to the sub-image. in, The independence decoding scheme refers to the following: when decoding the sub-image in the current image, it is not allowed to refer to other sub-images other than co-position sub-images in sub-images belonging to different images that are not the current image. The dependency decoding scheme refers to a decoding scheme in which, when decoding the sub-image in the current image, reference is allowed to other sub-images in different images.

3. The method of claim 2, wherein the first flag indicates whether the independence decoding or the dependency decoding is used for the sub-image. The first flag is obtained from the sequence parameters set in the bit stream. The first flag having a first value indicates that the independence decoding is applied to the sub-image, and The first flag having a second value indicates that the dependency decoding is applied to the sub-image.

4. The method according to claim 3, wherein, A second flag is obtained from the bitstream, which indicates whether loop filtering across the boundaries between the sub-images is applicable.

5. The method according to claim 4, wherein, The maximum coding unit refers to the basic coding unit with the largest size among the predefined coding units in the image decoding device.

6. The method according to claim 5, wherein, The sub-image is quadrilateral in shape.

7. A method for encoding a video signal using an image encoding device, the method comprising: The coefficients of the current block in the current image are obtained on a sub-block basis by performing a transformation on the residual sample of the current block. Quantization is performed on the coefficients of the current block; and The quantization coefficients of the current block are encoded into a bit stream. The sub-blocks are obtained by splitting the current block. The segmentation information based on the segmentation of the current block is encoded into the bit stream. The number of sub-blocks is 2. The segmentation information includes a segmentation ratio flag, which indicates one of a plurality of segmentation ratios among the sub-blocks. The multiple segmentation ratios mentioned include 1:

3. Specifically, the current image is segmented into multiple sub-images based on sub-image segmentation information. The sub-image segmentation information includes quantity information, position information, and size information. The quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the top-left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. Both the location information and the size information are represented in the largest coded unit. The plurality of sub-images includes a first sub-image and a second sub-image. A portion of the right boundary of the first sub-image is adjacent to the left boundary of the second sub-image, and The upper boundary of the first sub-image and the upper boundary of the second sub-image are discontinuous, while the lower boundary of the first sub-image and the lower boundary of the second sub-image are continuous.

8. A method for transmitting a bit stream generated by an encoding method, the encoding method comprising: The coefficients of the current block in the current image are obtained on a sub-block basis by performing a transformation on the residual sample of the current block. Quantization is performed on the coefficients of the current block; and The quantization coefficients of the current block are encoded into a bit stream. The sub-blocks are obtained by splitting the current block. The segmentation information based on the segmentation of the current block is encoded into the bit stream. The number of sub-blocks is 2. The segmentation information includes a segmentation ratio flag, which indicates one of a plurality of segmentation ratios among the sub-blocks. The multiple segmentation ratios mentioned include 1:

3. Specifically, the current image is segmented into multiple sub-images based on sub-image segmentation information. The sub-image segmentation information includes quantity information, position information, and size information. The quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the top-left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. Both the location information and the size information are represented in the largest coded unit. The plurality of sub-images includes a first sub-image and a second sub-image. A portion of the right boundary of the first sub-image is adjacent to the left boundary of the second sub-image, and The upper boundary of the first sub-image and the upper boundary of the second sub-image are discontinuous, while the lower boundary of the first sub-image and the lower boundary of the second sub-image are continuous.

Citation Information

Patent Citations

  • Image decoding method and image decoding device

    JP6074743B2

  • Efficient scalable coding concept

    US20150304667A1