Image decoding method and apparatus using a segmentation unit including an additional area

By setting additional areas in the image segmentation unit and referring to the image data of these areas, the problem of difficulty in using adjacent image data in the prior art is solved, and more efficient image compression and more accurate in-screen prediction are achieved.

CN116248864BActive Publication Date: 2025-06-10INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310305068.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-16
Filing Date
2018-07-03
Publication Date
2025-06-10
Estimated Expiration
2038-07-03

AI Technical Summary

Technical Problem

The existing image encoding technology is difficult to effectively utilize adjacent image data as a reference, resulting in inefficient encoding. The traditional in-screen prediction method relies on the most adjacent pixels and may not be suitable for different types of images.

Method used

By setting additional areas in the segmentation unit of the image and referring the image data in these additional areas during the encoding process, the encoding efficiency is improved. In addition, methods that support multiple reference pixel levels allow adaptive selection of the best reference pixel filtering method according to different image characteristics.

Benefits of technology

It improves image compression efficiency and in-screen prediction accuracy, can dynamically adjust reference pixel filtering according to image characteristics, and improves the overall image encoding/decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116248864B_ABST
    Figure CN116248864B_ABST
Patent Text Reader

Abstract

The present invention discloses a video decoding method and apparatus using a segmentation unit including an additional area. The video decoding method using a segmentation unit including an additional area includes: a step of dividing an encoded video included in the bitstream into at least one segmentation unit by referring to a syntax element obtained from the received bitstream; a step of setting an additional area for the at least one segmentation unit; and a step of decoding the encoded video based on the segmentation unit after setting the additional area. Thereby, the video encoding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application filed on July 3, 2018 with the application number 2018800450230 and the invention title "Image Decoding Method and Apparatus Using a Division Unit Including an Additional Region". Technical Field

[0002] The present invention relates to an image decoding method and apparatus using a division unit including an additional region, and more particularly to a technique for improving encoding efficiency by setting additional regions on the upper side, lower side, left side, and right side of a division unit such as a parallel block (tile) in an image and simultaneously referring to the image data in the additional regions during encoding. Background Art

[0003] In recent years, the demand for multimedia data such as videos on the Internet has been increasing rapidly. However, the current development speed of the channel bandwidth still requires a method capable of effectively compressing the rapidly increasing amount of multimedia data. For this reason, the Moving Picture Expert Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) and the Video Coding Expert Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) are working on developing a more efficient video compression standard through continuous collaborative research.

[0004] In addition, when performing independent encoding on an image, independent encoding is usually performed on individual division units including parallel blocks (tiles), so there is a problem that the image data of other division units adjacent in time and space cannot be referred to.

[0005] Therefore, a solution is needed that can refer to adjacent image data while maintaining the existing parallel processing based on independent encoding.

[0006] In addition, in intra prediction based on existing image encoding / decoding methods, the reference pixels are formed using the pixels closest to the current block, and depending on the type of image, the method of forming reference pixels using the closest pixels may not be advisable.

[0007] Therefore, a method is needed that can improve the intra prediction efficiency by adopting a different method of forming reference pixels from the existing method. Summary of the Invention

[0008] Technical Problem

[0009] In order to solve the above-mentioned existing problems, an object of the present invention is to provide an image decoding apparatus and method using a segmentation unit including an additional region.

[0010] In order to solve the above-mentioned existing problems, another object of the present invention is to provide an image encoding apparatus and method using a segmentation unit including an additional region.

[0011] In order to solve the above-mentioned existing problems, an object of the present invention is to provide an image decoding method supporting multiple reference pixel levels.

[0012] In order to solve the above-mentioned existing problems, another object of the present invention is to provide an image decoding apparatus supporting multiple reference pixel levels.

[0013] Technical solution

[0014] In order to achieve the above object, one aspect of the present invention provides an image decoding method using a segmentation unit including an additional region.

[0015] Among them, the image decoding method using a segmentation unit including an additional region may include: a step of dividing the encoded image included in the bitstream into at least one segmentation unit by referring to a syntax element obtained from the received bitstream; a step of setting an additional region for the at least one segmentation unit; and a step of decoding the encoded image based on the segmentation unit after setting the additional region.

[0016] Among them, the step of decoding the encoded image may include: a step of determining a reference block related to a current block to be decoded in the encoded image according to information included in the bitstream for indicating whether reference is possible or not.

[0017] Among them, the reference block may be a block included at a position overlapping with an additional region set in the segmentation unit to which the reference block belongs.

[0018] In order to achieve the above object, another aspect of the present invention provides an image decoding method supporting multiple reference pixel levels.

[0019] An image decoding method supporting multiple reference pixel levels can include: a step of confirming whether multiple reference pixel levels are supported through a bitstream; when multiple reference pixel levels are supported, a step of determining the reference pixel level to be used in the current block by referring to the syntax information included in the bitstream; a step of forming reference pixels using the pixels included in the determined reference pixel level; and a step of performing intra prediction of the current block using the formed reference pixels.

[0020] Among them, after the step of confirming whether multiple reference pixel levels are supported, it can further include: a step of confirming whether an adaptive reference pixel filtering method is supported through a bitstream.

[0021] Among them, after the step of confirming whether multiple reference pixel levels are supported, it can further include: when multiple reference pixel levels are not supported, a step of forming reference pixels using a preset reference pixel level.

[0022] Technical effects

[0023] When adopting the image decoding method and apparatus using a segmentation unit including an additional area applicable to the present invention as described above, since there is more image data that can be used as a reference, the image compression efficiency can be improved.

[0024] When adopting the image decoding method and apparatus supporting multiple reference pixel levels applicable to the present invention as described above, since multiple reference pixels can be used, the accuracy of intra prediction can be improved.

[0025] In addition, in the present invention, since adaptive reference pixel filtering is supported, the best reference pixel filtering can be performed according to the characteristics of the image.

[0026] In addition, the compression efficiency during image encoding / decoding can also be improved. Brief description of the drawings

[0027] Figure 1 It is a conceptual diagram illustrating an image encoding and decoding system according to an embodiment applicable to the present invention;

[0028] Figure 2 It is a block diagram illustrating an image encoding apparatus according to an embodiment applicable to the present invention;

[0029] Figure 3 It is a configuration diagram illustrating an image decoding apparatus according to an embodiment applicable to the present invention;

[0030] Figures 4a to 4d It is a conceptual diagram for explaining a projection format according to an embodiment applicable to the present invention;

[0031] Figures 5a to 5c is a conceptual diagram for explaining the surface configuration to which one embodiment of the present invention is applicable;

[0032] Figures 6a to 6b is an illustrative diagram for explaining a dividing part to which one embodiment of the present invention is applicable;

[0033] Figure 7 is an illustrative diagram for dividing an image into a plurality of parallel blocks;

[0034] Figures 8a to 8i is for Figure 7 the first illustrative diagram for setting additional areas for each of the parallel blocks illustrated in;

[0035] Figures 9a to 9i is for Figure 7 the second illustrative diagram for setting additional areas for each of the parallel blocks illustrated in;

[0036] Figure 10 is an illustrative diagram for applying the additional areas generated in one embodiment to which the present invention is applicable during the encoding / decoding process of other areas;

[0037] Figures 11 to 12 is a flowchart for explaining the encoding / decoding method of a dividing unit to which one embodiment of the present invention is applicable;

[0038] Figures 13a to 13g is an illustrative diagram for explaining the areas that can be referred to by a specific dividing unit;

[0039] Figures 14a to 14e is a flowchart for explaining the referability of additional areas in a dividing unit to which one embodiment of the present invention is applicable;

[0040] Figure 15 is an illustrative diagram for illustrating the blocks of a dividing unit included in the current image and the blocks of a dividing unit included in other images;

[0041] Figure 16 is a hardware configuration diagram for illustrating an image encoding / decoding device to which one embodiment of the present invention is applicable;

[0042] Figure 17 is an illustrative diagram for illustrating an intra prediction mode to which one embodiment of the present invention is applicable;

[0043] Figure 18 is the first illustrative diagram for illustrating the reference pixel configuration used in intra prediction to which one embodiment of the present invention is applicable;

[0044] Figures 19a to 19cIt is the second illustrative diagram showing the reference pixel composition applicable to one embodiment of the present invention;

[0045] Figure 20 It is the third illustrative diagram showing the reference pixel composition applicable to one embodiment of the present invention;

[0046] Figure 21 It is the fourth illustrative diagram showing the reference pixel composition applicable to one embodiment of the present invention;

[0047] Figures 22a to 22b It is the illustrative diagram showing the method of filling reference pixels at a preset position in an unusable reference candidate block;

[0048] Figures 23a to 23c It is the illustrative diagram showing the method of performing interpolation based on sub-pixel units in the reference pixels configured according to one embodiment of the present invention;

[0049] Figures 24a to 24b It is the first illustrative diagram for explaining the adaptive reference pixel filtering method applicable to one embodiment of the present invention;

[0050] Figure 25 It is the second illustrative diagram for explaining the adaptive reference pixel filtering method applicable to one embodiment of the present invention;

[0051] Figures 26a to 26b It is the illustrative diagram showing the case of using one reference pixel level in reference pixel filtering applicable to one embodiment of the present invention;

[0052] Figure 27 It is the illustrative diagram showing the case of using multiple reference pixel levels in reference pixel filtering applicable to one embodiment of the present invention;

[0053] Figure 28 It is the block diagram for explaining the encoding / decoding method of the in-picture prediction mode applicable to one embodiment of the present invention;

[0054] Figure 29 It is the first illustrative diagram for explaining the bitstream composition of in-picture prediction based on the reference pixel composition;

[0055] Figure 30 It is the second illustrative diagram for explaining the bitstream composition of in-picture prediction based on the reference pixel composition;

[0056] Figure 31 It is the third illustrative diagram for explaining the bitstream composition of in-picture prediction based on the reference pixel composition;

[0057] Figure 32It is a flowchart illustrating an image decoding method that supports multiple reference pixel levels for one embodiment of the present invention. Detailed implementation

[0058] The present invention can be subject to various changes and has multiple different embodiments. Next, specific embodiments thereof will be illustrated and described in detail. However, the following content is not intended to limit the present invention to a specific embodiment, but should be understood to include all changes, equivalents, and substitutes within the spirit and technical scope of the present invention. In the process of explaining each drawing, similar reference symbols are used for similar components.

[0059] In the process of explaining different components, terms such as first, second, A, B, etc. can be used, but the above components are not limited by the above terms. The above terms are only used to distinguish one component from other components. For example, without departing from the scope of the claims of the present invention, the first component can also be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes combinations of multiple related recited items or one of the multiple related recited items.

[0060] When it is described that a certain component is "connected" or "contacted" with other components, it should be understood that not only can it be directly connected or contacted with the above other components, but there can also be other components between the two. On the contrary, when it is described that a certain component is "directly connected" or "directly contacted" with other components, it should be understood that there are no other components between the two.

[0061] The terms used in this application are only for explaining specific embodiments and are not intended to limit the present invention. Unless there is a clear contrary meaning in the context, the singular form of a statement also includes the plural meaning. In this application, terms such as "including" or "having" are only used to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should not be understood to preclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof in advance.

[0062] Unless otherwise defined, the meanings of all terms used herein, including technical or scientific terms, are the same as those commonly understood by a person of ordinary skill in the technical field to which the present invention belongs. Terms that are commonly used, such as those defined in a dictionary, should be interpreted as having a meaning consistent with their meaning in the context of the related art. In this application, unless otherwise clearly defined, they should not be interpreted as overly idealized or exaggerated meanings.

[0063] Generally, an image can be composed of a series of still images. The above still images can be divided in units of a group of pictures (GOP). Each still image can be called a picture or a frame. As its superordinate concepts, there can be units such as a group of pictures (GOP) and a sequence. Each picture can also be divided into specific regions such as stripes, parallel blocks, and blocks. In addition, a group of pictures (GOP) can include units such as I pictures, P pictures, and B pictures. An I picture can refer to an image that is encoded / decoded on its own without using a reference image, while P pictures and B pictures can refer to images that are encoded / decoded by performing processes such as motion estimation and motion compensation using reference images. Generally, a P picture can use I pictures and P pictures as reference images, and a B picture can use I pictures and P pictures as reference images. However, the above definitions can also be changed according to the encoding / decoding settings.

[0064] Among them, the image used as a reference during the encoding / decoding process is called a reference picture, and the block or pixel used as a reference is called a reference block or a reference pixel. In addition, reference data can be coefficient values in the frequency domain, in addition to pixel values in the spatial domain, and various encoding / decoding information generated and determined during the encoding / decoding process.

[0065] The smallest unit that makes up an image can be a pixel, and the number of bits used to represent a pixel is called the bit depth. Usually, the bit depth can be 8 bits, and other bit depths can be supported according to the encoding settings. Regarding the bit depth, at least one bit depth can be supported according to the color space. In addition, it can be composed of at least one color space according to the color format of the image. According to the color format, it can be composed of one or more images of a certain size or one or more images of different sizes. For example, in the case of YCbCr 4:2:0, it can be composed of one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example). At this time, the composition ratio of the chrominance component and the luminance component can be 1:2 horizontally and vertically. As another example, in the case of 4:4:4, it can have the same composition ratio horizontally and vertically. In the case of being composed of more than one color space as described above, the image can be segmented on each color space.

[0066] In the present invention, a part of the color space (Y in this example) of a part of the color format (YCbCr in this example) will be used as a reference for description, and it can be applied in the same or similar way in other color spaces (Cb and Cr in this example) based on the color format (depending on the settings of the specific color space). However, some differences can also be retained in each color space (independent of the settings of the specific color space). That is, depending on the settings of each color space can refer to the property that is proportional or dependent on the composition ratio of each component (for example, determined according to 4:2:0, 4:2:2, 4:4:4, etc.), and independent of the settings of each color space can refer to the property that is independent of the composition ratio of each component and is only applicable to the corresponding color space independently. In the present invention, depending on the encoder / decoder, some compositions can have the property of independence or dependence.

[0067] The setting information or syntax elements required during the video encoding process can be determined at the unit levels such as video, sequence, picture, slice, tile, block, etc., and can be included in the bitstream and transmitted to the decoder in units such as Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice Header, Tile Header, Block Header, etc. In the decoder, they can be parsed at the same-level units and used during the video decoding process after decoding the setting information transmitted from the encoder. In addition, relevant information can also be transmitted to the bitstream in forms such as Supplement Enhancement Information (SEI) or Metadata and used after parsing. Each parameter set has an inherent number value, and the lower-level parameter set can include the number values of the upper-level parameter sets that need to be referenced. For example, the lower-level parameter set can reference the information of the upper-level parameter sets with consistent number values from one or more upper-level parameter sets. In the above-mentioned examples of various units, when a certain unit contains one or more other units, the corresponding unit can be called the upper-level unit and the included unit can be called the lower-level unit.

[0068] Regarding the setting information generated at the above units, it can include the content related to the independent settings in each unit, and can also include the content related to the settings that depend on the previous, subsequent, or upper-level units, etc. Among them, the dependent setting refers to the flag information (for example, a 1-bit flag, where 1 means follow and 0 means not follow) used to indicate whether to follow the settings of the previous, subsequent, or upper-level units, and can be understood as representing the setting information of the corresponding unit. Although in the present invention, the description of the setting information will be centered on the examples related to the independent settings, it can also include examples of adding or replacing using the setting information of the previous, subsequent units, or upper-level units that depend on the current unit.

[0069] Next, the preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0070] Figure 1 It is a conceptual diagram illustrating the video encoding and decoding system of the embodiments applicable to the present invention.

[0071] Refer to Figure 1, the video encoding device 105 and the decoding device 100 can be user terminals such as a personal computer (PC), a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a portable game station (PSP), a wireless communication terminal, a smart phone, a television (TV), etc., or server terminals such as an application server and a service server. They can include various devices such as a communication device like a communication modem for communicating with various devices or wired / wireless communication networks, memories 120, 125 for storing various application programs and data for performing inter-frame or intra-frame prediction for video encoding or decoding, processors 110, 115 for performing operations and controls by executing application programs, etc. In addition, the video encoded into a bitstream by the video encoding device 105 can be transmitted to the video decoding device 100 through wired / wireless communication networks such as the Internet, a short-range wireless communication network, a wireless local area network, a wireless broadband network, a mobile communication network, etc., or through various communication interfaces such as a cable, a universal serial bus (USB), etc., and decoded in the video decoding device 100 to reconstruct the video and then played. In addition, the video encoded into a bitstream by the video encoding device 105 can also be transferred from the video encoding device 105 to the video decoding device 100 through a computer-readable storage medium.

[0072] Figure 2 It is a block diagram illustrating a video encoding device to which one embodiment of the present invention is applied.

[0073] The video decoding device 20 applicable to this embodiment, as Figure 2 shown, can include a prediction unit 200, a subtraction operation unit 205, a transformation unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse transformation unit 225, an addition operation unit 230, a filtering unit 235, an encoded image buffer 240, and an entropy encoding unit 245.

[0074] The prediction unit 200 can include an intra-prediction unit for performing intra-picture prediction and an inter-picture prediction unit for performing inter-picture prediction. Intra-picture prediction can generate a prediction block by performing spatial prediction using the pixels of the blocks adjacent to the current block, while inter-picture prediction can generate a prediction block by searching for the region that best matches the current block in the reference picture and performing motion compensation. After determining which of intra-picture prediction or inter-picture prediction is applicable to the corresponding unit (coding unit or prediction unit), the specific information related to each prediction method (such as intra-picture prediction mode, motion vector, reference picture, etc.) can be determined. At this time, the processing unit that performs the prediction and the processing unit that determines the prediction method and the specific content can be different according to the encoding / decoding settings. For example, the prediction method and prediction mode, etc. can be determined in terms of the prediction unit, while the execution of the prediction can be performed in terms of the transform unit.

[0075] The intra-prediction unit can adopt directional prediction modes such as horizontal and vertical modes used according to the prediction direction, and non-directional prediction modes such as mean (DC), planar, etc. that use methods such as the average and interpolation of reference pixels. Through the directional and non-directional modes, a candidate group of intra-picture prediction modes can be formed, and one of the various selections such as 35 prediction modes (33 directional + 2 non-directional), 67 prediction modes (65 directional + 2 non-directional), 131 prediction modes (129 directional + 2 non-directional), etc. can be used as the candidate group.

[0076] The intra-prediction unit can include a reference pixel formation unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel formation unit can form the reference pixels for performing intra-picture prediction by using the pixels included in the adjacent blocks and adjacent to the current block with the current block as the center. According to the encoding settings, the reference pixels can be formed by using the closest adjacent reference pixel row, or by using other adjacent reference pixel rows in addition to this, or by using multiple reference pixel rows. When some of the reference pixels are unavailable, the available reference pixels can be used to generate the reference pixels, and when all are unavailable, the reference pixels can be generated by using a pre-set value (such as the middle value of the pixel value range that can be represented by the bit depth, etc.).

[0077] The reference pixel filter section of the intra prediction section can perform filtering on reference pixels for the purpose of reducing distortion remaining through the encoding process. At this time, the filter used can be a low-pass filter such as a 3-tap filter [1 / 4, 1 / 2, 1 / 4], a 5-tap filter [2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16], etc. The applicability and type of filtering can be determined according to encoding information (such as the size, shape, prediction mode, etc. of the block).

[0078] The reference pixel interpolation section of the intra prediction section can generate pixels with fractional units through a linear interpolation process of reference pixels according to the prediction mode, and can determine the applicable interpolation filter according to encoding information. At this time, the interpolation filters used can include a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc. Generally, the process of performing low-pass filtering and the process of performing interpolation are independent of each other, but it is also possible to perform the filtering process after integrating the filters applicable in the two processes into one.

[0079] The prediction mode determination section of the intra prediction section can select the best prediction mode from a candidate group of prediction modes while taking into account the encoding cost, and the prediction block generation section can generate a prediction block using the corresponding prediction mode. In the prediction mode encoding section, the best prediction mode can be encoded based on the prediction value. At this time, the prediction information can be adaptively encoded according to whether the prediction value is appropriate or not.

[0080] In the intra prediction section, the above prediction value is called the Most Probable Mode (MPM), and a part of the modes included in the candidate group of prediction modes can be selected to form a candidate group of the Most Probable Mode (MPM). In the candidate group of the Most Probable Mode (MPM), it can include preset prediction modes (such as mean (DC), planar, vertical, horizontal, diagonal modes, etc.) or the prediction modes of spatially adjacent blocks (such as the left, upper, upper left, upper right, lower left blocks, etc.). In addition, the candidate group of the Most Probable Mode (MPM) can be formed using modes derived from the modes pre-included in the candidate group of the Most Probable Mode (MPM) (such as differences of +1, -1, etc. in the directional mode).

[0081] Among the prediction patterns used to form the candidate group of the most probable mode (MPM), a priority order can exist. The order included in the candidate group of the most probable mode (MPM) can be determined according to the above priority order, and the formation of the candidate group of the most probable mode (MPM) can be completed when the number of the candidate group of the most probable mode (MPM) is filled according to the above priority order (determined according to the number of the candidate group of prediction patterns). At this time, the priority order can be determined in the order of the prediction patterns of spatially adjacent blocks, preset prediction patterns, and patterns derived from the prediction patterns earlier included in the candidate group of the most probable mode (MPM), and other variations can also be made.

[0082] For example, in spatially adjacent blocks, they can be included in the candidate group in the order of left - upper - lower - left - upper - right - upper - left blocks, etc. In the preset prediction patterns, they can be included in the candidate group in the order of mean (DC) - planar - vertical - horizontal patterns, etc. Then, the patterns obtained by performing addition operations such as +1 and -1 on the pre - included patterns are included in the candidate group, so that the candidate group is formed by a total of 6 patterns. Or, they can also be included in the candidate group in a priority order such as left - upper - mean (DC) - planar (Plana) - lower - left - upper - right - (left + 1) - (left - 1) - (upper + 1), etc., so that the candidate group is formed by a total of 7 patterns.

[0083] In the above candidate group formation, a validity check can be performed, so that only when it is valid, it is included in the candidate group, and when it is invalid, it jumps to the next candidate. For example, when the adjacent block is outside the image or is included in a segmentation unit different from the current block or the coding mode of the corresponding block is inter - picture prediction, it can be invalid. In addition, it can also be invalid in the case of non - reference described later in the present invention.

[0084] In the above candidates, the spatially adjacent blocks can be composed of one block or multiple blocks (sub - blocks). Therefore, in the order such as (left - upper) in the above candidate group formation, the order can be to jump to the upper block after performing the validity check on a certain position in the left block (for example, the lowermost block in the left block), or the order can be to jump to the upper block after performing the validity check on multiple positions (for example, one or more sub - blocks located below the uppermost block in the left block), and it can also be determined according to the coding setting.

[0085] In the inter-picture prediction unit, it can be divided into a translational motion model and a non-translational motion model according to the motion prediction method. In the translational motion model, prediction is performed while only considering translational motion, while in the non-translational motion model, prediction is performed while considering motions such as rotation, distance, and zoom in / out in addition to translational motion. On the premise of assuming unidirectional prediction, one motion vector is required in the translational motion model, while more than one motion vector is required in the non-translational motion model. In the non-translational motion model, each motion vector can be information applicable to a preset position of the current block, such as the upper left vertex or the upper right vertex of the current block. Through the corresponding motion vector, the position of the area to be predicted in the current block can be obtained in pixel units or sub-block units. The inter-picture prediction unit can apply some of the processes described later commonly while applying some other processes individually according to the above motion models.

[0086] The inter-picture prediction unit can include a reference image construction unit, a motion prediction unit, a motion compensation unit, a motion information determination unit, and a motion information encoding unit. The reference image construction unit can include the previously or subsequently encoded images in the reference image list (L0, L1) centered on the current image. From the reference images included in the above reference image list, a predicted block can be obtained, and according to the encoding setting, the current video can also be used to construct a reference image and include it in at least one position in the reference image list.

[0087] In the inter-picture prediction unit, the reference image construction unit can include a reference image interpolation unit, and an interpolation process for sub-pixel units can be performed according to the interpolation accuracy. For example, an 8-tap discrete cosine transform (DCT)-based interpolation filter can be applied to the luminance component, and a 4-tap discrete cosine transform (DCT)-based interpolation filter can be applied to the chrominance component.

[0088] In the inter-picture prediction unit, the motion prediction unit is used to perform the process of exploring a block with a higher correlation with the current block through the reference image, and various methods such as the full search block matching algorithm (FBMA) and the three-step search algorithm (TSS) can be used, while the motion compensation unit is used to perform the process of obtaining the predicted block through the motion prediction process.

[0089] In the inter-picture prediction unit, the motion information determination unit can execute a process for selecting the best motion information for the current block, and the motion information can be encoded by motion information encoding modes such as Skip Mode, Merge Mode, and Competition Mode. The above modes can adopt a configuration that combines the supported modes according to the motion model, and its examples can include Skip Mode (moving), Skip Mode (non-moving), Merge Mode (moving), Merge Mode (non-moving), Competition Mode (moving), and Competition Mode (non-moving). According to the symbolization setting, a part of the above modes can be included in the candidate group.

[0090] The above motion information encoding modes can obtain the predicted values of the motion information (motion vector, reference image, prediction direction, etc.) of the current block from at least one candidate block, and can generate the best candidate selection information when supporting two or more candidate blocks. Skip Mode (no residual signal) and Merge Mode (with residual signal) can directly use the above predicted values as the motion information of the current block, while Competition Mode can generate the difference value information between the motion information of the current block and the above predicted values.

[0091] The candidate group for the predicted value of the motion information of the current block is adaptive according to the motion information encoding mode and can adopt various configurations. The motion information of the blocks spatially adjacent to the current block (such as the left, upper, upper left, upper right, lower left blocks, etc.) can be included in the candidate group, and the motion information of the blocks temporally adjacent to the current block (such as the left, right, upper, lower, upper left, upper right, lower left, lower right blocks, etc. including the block <center> corresponding to or corresponding to the current block in other images) can also be included in the candidate group. In addition, the mixed motion information of spatial candidates and temporal candidates (such as the information obtained by averaging, taking the median, etc. of the motion information of the spatially adjacent blocks and the motion information of the temporally adjacent blocks, and the motion information that can be obtained in units of the current block or sub-blocks of the current block) can also be included in the candidate group.

[0092] In the composition of the candidate group of the predicted value of the motion information, there can be a priority order. The order included in the composition of the predicted value candidate group can be determined according to the above priority order, and the composition of the candidate group can be completed when the number of the candidate group (determined according to the motion information encoding mode) is filled in according to the above priority order. At this time, the priority order can be determined in the order of the motion information of the blocks spatially adjacent to the current block, the motion information of the blocks temporally adjacent to the current block, and the mixed motion information of spatial candidates and temporal candidates, and other deformations can also be performed.

[0093] For example, in spatially adjacent blocks, they can be included in the candidate group in the order of left - upper - upper - right - lower - left - upper blocks, etc., while in temporally adjacent blocks, they can be included in the candidate group in the order of lower - right - middle - right - lower blocks, etc.

[0094] In the above - mentioned candidate group formation, a validity check can be performed, so that it is included in the candidate group only when it is valid and jumps to the next candidate when it is invalid. For example, when an adjacent block is outside the image, or is included in a segmentation unit different from the current block, or the coding mode of the corresponding block is intra - picture prediction, it can be invalid. In addition, it can also be invalid in the case of non - reference described later in the present invention.

[0095] In the above - mentioned candidates, adjacent blocks in space or time can be composed of one block or multiple blocks (sub - blocks). Therefore, in the order such as (left - upper) in the above - mentioned spatial candidate group formation, the order can be to jump to the upper block after performing a validity check on a certain position in the left block (for example, the lowermost block in the left block), or the order can be to jump to the upper block after performing validity checks on multiple positions (for example, one or more sub - blocks located below starting from the uppermost block in the left block). In addition, in the order such as (middle - right) in the temporal candidate group formation, the order can be to jump to the right block after performing a validity check on a certain position in the middle block (for example, <2, 2> when the central block is divided into 4×4 regions), or the order can be to jump to the lower block after performing validity checks on multiple positions (for example, one or more sub - blocks such as <3, 3>, <2, 3>, etc. in a preset order starting from the preset position block <2, 2>), and it can also be determined according to the coding settings.

[0096] The subtraction operation unit 205 generates a residual block by performing subtraction on the current block and the predicted block. That is, the subtraction operation unit 205 calculates the difference between the pixel values of each pixel of the current block to be encoded and the predicted pixel values of each pixel of the predicted block generated by the prediction unit, and generates a residual signal in the form of a block, that is, a residual block.

[0097] The transform unit 210 transforms each pixel value of the residual block into a frequency coefficient by transforming the residual block into the frequency domain. Among them, the transform unit 210 can use various transform techniques for transforming a pixel signal on the spatial axis into a frequency signal, such as Hadamard Transform, DCT Based Transform, DST Based Transform, KLT Based Transform, etc., to transform the residual signal into a frequency signal, and the residual signal transformed into the frequency domain will become the frequency coefficient. The transformation can be performed through a one-dimensional transformation matrix during transformation. Each transformation matrix can be adaptively used in horizontal and vertical units. For example, when the prediction mode in intra-frame prediction is horizontal, a DCT-based transformation matrix can be used in the vertical direction and a DST-based transformation matrix can be used in the horizontal direction. When the prediction mode is vertical, a DCT-based transformation matrix can be used in the horizontal direction and a DST-based transformation matrix can be used in the vertical direction.

[0098] The quantization unit 215 quantizes the residual block containing the frequency coefficients transformed into the frequency domain by the transform unit 210. Among them, the quantization unit 215 can use quantization techniques such as Dead Zone Uniform Threshold Quantization, Quantization Weighted Matrix or improved versions thereof to quantize the transformed residual block. At this time, one or more quantization techniques can be selected as candidates and determined according to coding mode, prediction mode information, etc.

[0099] The entropy coding unit 245 generates a quantized coefficient sequence by scanning the generated quantized frequency coefficient sequence using various scanning methods, and outputs it after encoding using entropy coding techniques, etc. As the scanning mode, it can be set to one of various modes such as zigzag, diagonal, raster, etc. In addition, encoded data containing the encoding information transmitted from each component can be generated and output to the bitstream.

[0100] The inverse quantization unit 220 performs inverse quantization on the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 generates a residual block containing frequency coefficients by performing inverse quantization on the quantized frequency coefficient sequence.

[0101] The inverse transformation unit 225 performs an inverse transformation on the residual block that has been inverse quantized by the inverse quantization unit 220. That is, the inverse transformation unit 225 generates a residual block containing pixel values, i.e., a reconstructed residual block, by performing an inverse transformation on the frequency coefficients of the inverse quantized residual block. Among them, the inverse transformation unit 225 can perform the inverse transformation by reversely using the transformation method used in the transformation unit 210.

[0102] The addition operation unit 230 can reconstruct the current block by adding the predicted block predicted by the prediction unit 200 and the residual block reconstructed by the inverse transformation unit 225. The reconstructed current block will be stored in the encoded image buffer 240 as a reference image (or reference block), so that it can be used as a reference image when encoding the next block or subsequent other blocks or other images of the current block.

[0103] The filtering unit 235 can include one or more post-processing filtering processes such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF). The deblocking filter can eliminate block distortion that appears on the boundary between blocks in the reconstructed image. The adaptive loop filter (ALF) can perform filtering based on the value obtained by comparing the reconstructed image with the original image after filtering the block with the deblocking filter. The sample adaptive offset (SAO) can reconstruct the offset difference between the residual block to which the deblocking filter has been applied and the original image in terms of pixels. The post-processing filters described above can be applied to the reconstructed image or block.

[0104] The deblocking filter in the filtering unit can be applied based on the pixels included in several columns or rows included in two blocks with the block boundary as a reference. As the above-mentioned block, it is suitable to be applied to the boundaries of the encoded block, predicted block, and transformed block, and can be limited to blocks with a preset minimum size (e.g., 8×8) or more.

[0105] Regarding whether to apply filtering, it is possible to determine whether to apply filtering and the filtering intensity in consideration of the block boundary characteristics, and it can be determined as one of the selected options such as strong filtering, medium filtering, and weak filtering. In addition, when the above-mentioned block boundary belongs to the boundary of the segmentation unit, it is determined whether to apply it based on the loop filter application flag at the boundary of the segmentation unit, and it can also be determined whether to apply it according to various situations described later in the present invention.

[0106] Sample Adaptive Offset (SAO) in the filtering unit can be applied based on the difference value between the reconstructed image and the original image. As offset types, it can support, for example, Edge Offset and Band Offset, and can select one of the above offsets to perform filtering according to the characteristics of the image. In addition, the above offset-related information can be encoded in block units and can also be encoded through related prediction values. At this time, the related information can be adaptively encoded according to whether the prediction value is appropriate or not. The prediction value can be the offset information of adjacent blocks (such as the left, upper, upper left, upper right blocks, etc.), and selection information related to obtaining the offset information of which block can be generated.

[0107] In the above candidate group composition, validity checks can be performed, so that it is included in the candidate group only when it is valid and jumps to the next candidate when it is invalid. For example, adjacent blocks can be outside the image or included in a segmentation unit different from the current block or can be invalid in cases where it cannot be referred to as described later in the present invention.

[0108] The encoded image buffer 240 can store the blocks or images reconstructed by the filtering unit 235. The reconstructed blocks or images stored in the encoded image buffer 240 can be provided to the prediction unit 200 for performing intra-picture prediction or inter-picture prediction.

[0109] Figure 3 It is a block diagram illustrating an image decoding apparatus to which one embodiment of the present invention is applied.

[0110] Refer to Figure 3 , the image decoding apparatus 30 can include an entropy decoding unit 305, a prediction unit 310, an inverse quantization unit 315, an inverse transform unit 320, an addition / subtraction arithmetic unit 325, a filter 330, and a decoded image buffer 335.

[0111] In addition, the prediction unit 310 can further include an intra-picture prediction module and an inter-picture prediction module.

[0112] First, when receiving the image bitstream transmitted from the image encoding apparatus 20, it can be transmitted to the entropy decoding unit 305.

[0113] The entropy decoding unit 305 can decode the decoded data including the quantized coefficients and the decoding information transmitted from each component by decoding the bitstream.

[0114] The prediction unit 310 may generate a prediction block based on the data transmitted from the entropy decoding unit 305. At this time, the reference picture list using a default composition technique may be constructed based on the decoded reference picture stored in the picture buffer 335.

[0115] The intra-frame prediction unit can include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction block generation unit, and a prediction mode decoding unit, and the inter-frame prediction unit can include a reference image construction unit, a motion compensation unit, and a motion information decoding unit, one part of which can perform the same process as the encoder, and the other part can perform a reverse induction process.

[0116] The inverse quantization unit 315 can inversely quantize the quantized transform coefficient provided from the bit stream and decoded by the entropy decoding unit 305 .

[0117] The inverse transform unit 320 may generate a residual block by applying an inverse discrete cosine transform (DCT), an inverse integer transform, or an inverse transform technique having a similar concept to the transform coefficient.

[0118] At this time, the inverse quantization unit 315 and the inverse transform unit 320 will inversely perform the processes performed in the transform unit 210 and the quantization unit 215 of the image encoding device 20 described above, and can be implemented by various methods. For example, the same process and inverse transform shared by the transform unit 210 and the quantization unit 215 can be used, and the transform and quantization process can be inversely performed using information related to the transform and quantization process of the image encoding device 20 (such as transform size, transform shape, quantization type, etc.).

[0119] The residual block after the inverse quantization and inverse transformation process can be added to the prediction block derived in the prediction unit 310 to generate a reconstructed image block. The addition operation can be performed by the addition and subtraction operator 325.

[0120] For the reconstructed image blocks, the filter 330 can apply a deblocking filter for eliminating blocking as needed, and can additionally use other loop filters before and after the above decoding process in order to improve video quality.

[0121] The reconstructed and filtered image blocks can be stored in the decoded image buffer 335 .

[0122] Although not shown in the figure, the video decoding device 30 can further include a segmentation unit, and the segmentation unit can include an image segmentation unit and a block segmentation unit. Figure 2In the image decoding device illustrated in the [description], the same or corresponding components are easily understandable to those skilled in the art, so detailed descriptions thereof will be omitted herein.

[0123] Figures 4a to 4d It is a conceptual diagram for explaining the projection format applicable to one embodiment of the present invention.

[0124] Figure 4a The equi-rectangular projection (ERP) format for projecting a 360-degree image onto a two-dimensional plane is illustrated. Figure 4b The cube map projection (CMP) format for projecting a 360-degree image onto a cube is illustrated. Figure 4c The octahedron projection (OHP) format for projecting a 360-degree image onto an octahedron is illustrated. Figure 4d The icosahedron projection (ISP) format for projecting a 360-degree image onto a polyhedron is illustrated. However, it is not limited thereto, and various projection formats can be used. For example, the truncated square pyramid projection (TSP), segmented sphere projection (SSP), etc. can also be used. Figures 4a to 4d On the left side is a 3D model, and on the right side is an example transformed into a two-dimensional space through the projection process. The graph of the two-dimensional projection can be composed of one or more faces, and each face can adopt shapes such as a circle, triangle, quadrilateral, etc.

[0125] As Figures 4a to 4d shown, the projection format can be composed of one face (such as the equi-rectangular projection (ERP)) or multiple faces (such as the cube map projection (CMP), octahedron projection (OHP), icosahedron projection (ISP), etc.). In addition, each face can be classified into forms such as quadrilateral and triangle. The above classification can be an example of the type, characteristics, etc. of the image in the present invention applicable when setting different encoding / decoding settings according to the projection format. For example, the type of the image can be a 360-degree image, and the characteristic of the image can be one of the above classifications (such as each projection format, the projection format of one face or multiple faces, the projection format with a quadrilateral or non-quadrilateral face, etc.).

[0126] A two-dimensional plane coordinate system {e.g., (i, j)} can be defined on each surface of a two-dimensional projected image, and the characteristics of the coordinate system can vary according to the projection format, the position of each surface, etc. In cases such as equirectangular projection (ERP), a two-dimensional plane coordinate system can be included, while other projection formats can include multiple two-dimensional plane coordinate systems according to the number of surfaces. At this time, the coordinate system can be represented as (k, i, j), where k can be the index information of each surface.

[0127] For the convenience of description in the present invention, the case where the surface form is a quadrilateral will be mainly described, and the number of surfaces projected onto two dimensions can be one (e.g., equirectangular projection, that is, the case where the image is equivalent to one surface) to two or more (e.g., cube map projection, etc.).

[0128] Figures 5a to 5c It is a conceptual diagram for explaining the surface configuration applicable to one embodiment of the present invention.

[0129] In the projection format of projecting a three-dimensional image onto two dimensions, it is necessary to determine the configuration of its surface. At this time, the surface configuration can be configured in a way that maintains the image continuity in three-dimensional space, or can be configured in a way that maximally tightens the interval between surfaces even if the image continuity between some adjacent surfaces is broken. In addition, when configuring the surface, some surfaces can be configured after rotating a certain angle (0, 90, 180, 270 degrees, etc.).

[0130] Refer to Figure 5a , an example of the surface layout related to the cube map projection (CMP) format can be confirmed. When configured in a way that maintains the image continuity in three-dimensional space, a 4×3 layout that configures four surfaces horizontally and then one surface each above and below as shown in the left image can be used. In addition, a 3×2 layout that seamlessly configures the surfaces on the two-dimensional plane even if the image continuity between some adjacent surfaces is broken as shown in the right image can also be used.

[0131] Refer to Figure 5b , the surface configuration related to the octahedron projection (OHP) format can be confirmed. When configured in a way that maintains the image continuity in three-dimensional space, it is as shown in the upper image. At this time, when the surfaces are seamlessly configured on the projected two-dimensional plane even if some image continuity is broken, it can also be configured in the way shown in the lower image.

[0132] Refer to Figure 5c, it is possible to confirm the surface configuration related to the icosahedral projection (ISP) format. At this time, it can be configured in a way that maintains the continuity of the image in the three-dimensional space as shown in the upper image, or it can also be configured in a way that fits tightly between the surfaces without gaps as shown in the lower image.

[0133] At this time, the process of tightly fitting the configured surfaces can be called frame packing, and the phenomenon of image continuity being damaged can be minimized by rotating the surfaces and then configuring them. Next, the process of changing the surface configuration to another surface configuration as described above will be called surface reconfiguration.

[0134] Next, the term continuity can be interpreted as the continuity of the scene visible to the naked eye in the three-dimensional space, or the continuity of the actual image or scene in the two-dimensional projection space. Having continuity can also be expressed as a high correlation between regions. Generally, the correlation between regions on a two-dimensional image may be high or low, but in a 360-degree image, there may be regions that are adjacent in space but have no continuity at all. In addition, according to the surface configuration or reconfiguration as described above, there may be regions that have continuity but are not adjacent in space.

[0135] Surface reconfiguration can be performed for the purpose of improving coding performance. For example, by performing surface reconfiguration, surfaces with image continuity can be configured adjacent to each other.

[0136] At this time, surface reconfiguration does not necessarily mean that it must be reconstituted after configuring the surfaces. It can also be understood as a process of setting a specific surface configuration from the beginning. (It can be performed in the region-wise packing during the 360-degree image encoding / decoding process)

[0137] In addition, surface configuration or reconfiguration can include the rotation of the surface in addition to the change in the position of each surface (in this example, the simple movement of the surface, such as moving from the upper left end of the image to the lower left end or the lower right end). Among them, the rotation of the surface can include 0 degrees without surface rotation, 45 degrees to the right, 90 degrees to the left, etc., and the rotation angle can be represented by selecting the divided interval after dividing the 360 degrees (equally or unequally) into k (or 2k) intervals.

[0138] The encoder / decoder is capable of performing surface configuration (or reconfiguration) in accordance with preset surface configuration information (such as the shape of the surface, the number of surfaces, the position of the surface, the rotation angle of the surface, etc.) and / or surface reconfiguration information (information for indicating the position or movement angle, movement direction, etc. of each surface). In addition, the encoder can generate surface configuration information and / or surface reconfiguration information based on the input image, and the decoder can receive the above-mentioned information from the encoder and perform decoding, so as to perform surface configuration (or reconfiguration).

[0139] Next, unless otherwise specified, the surfaces referred to without separate explanation are based on the 3×2 configuration (layout) in Figure 5a , and at this time, the numbers representing each surface can be 0 to 5 respectively in the raster scan order starting from the upper left end.

[0140] Next, when explaining on the premise of the continuity between the surfaces in Figure 5a , unless otherwise specified, it can be assumed that there is continuity between the surfaces numbered 0 to 2, and there is continuity between the surfaces numbered 3 to 5, while there is no continuity between the surfaces numbered 0 and 3, 1 and 4, 2 and 5. Whether there is continuity between the above surfaces can be determined by settings such as the characteristics, type, format, etc. of the image.

[0141] In the encoding / decoding process of 360-degree images, the encoding device can acquire the input image, perform preprocessing on the acquired image, encode the preprocessed image, and transmit the encoded bitstream to the decoding device. Among them, the preprocessing can include image stitching, projecting a 3D image onto a 2D plane, surface configuration and reconfiguration (or what is called region-wise packing), etc. In addition, the decoding device can receive the bitstream, decode the received bitstream, perform postprocessing (such as image rendering) on the decoded image, and generate the output image.

[0142] At this time, the bitstream can include the information generated in the preprocessing process (such as supplementary enhancement information (SEI) messages or metadata, etc.) and the information generated in the encoding process (image encoding data) and then transmit.

[0143] Figures 6a to 6b It is an illustrative diagram for explaining the segmentation unit applicable to one embodiment of the present invention.

[0144] Figure 2 Or Figure 3The image encoding / decoding device therein may further include a segmentation unit, and the segmentation unit may include an image segmentation unit and a block segmentation unit. The image segmentation unit may segment an image into at least one processing unit (such as a color space (YCbCr, RGB, XYZ, etc.), a sub-image, a strip, a parallel block, a basic coding unit (or a maximum coding unit), etc.), and the block segmentation unit may segment a basic coding unit into at least one processing unit (such as a coding, prediction, transform, quantization, entropy, loop filtering unit, etc.).

[0145] The basic coding unit may be obtained by segmenting an image at a certain length interval along the horizontal or vertical direction, and this may also be a unit applicable to units such as sub-images, parallel blocks, strips, surfaces, etc. That is, the above units may be constituted by an integer multiple of the basic coding unit, but it is not limited thereto.

[0146] For example, in some segmentation units (in this example, parallel blocks, sub-images, etc.), different basic coding units may be applicable, and the corresponding segmentation units may adopt independent basic coding unit sizes. That is, the basic coding unit of the above segmentation unit may be set to be the same as or different from the basic coding unit of the image unit and the basic coding unit of other segmentation units.

[0147] In the present invention, for the convenience of description, the basic coding unit and other processing units (coding, prediction, transform, etc.) other than this are referred to as blocks (Block).

[0148] The size or shape of the above block may be an N×N square shape (2n×2n, 256×256, 128×128, 64×64, 32×32, 16×16, 8×8, 4×4, etc., where n is an integer between 2 and 8) or an M×N rectangular shape (2m×2n) represented by an exponential power of 2 (2 n ) in the horizontal or vertical length. For example, an input image of the 8k ultra-high definition (UHD) level with extremely high resolution may be segmented into a size of 256×256, an input image of the 1080p high definition (HD) level may be segmented into a size of 128×128, and an input image of the wide video graphics array (WVGA) level may be segmented into a size of 16×16.

[0149] An image may be segmented into at least one strip. A strip may be composed of a combination of at least one block consecutive in the scanning order. Each strip may be segmented into at least one strip segment, and each strip segment may be segmented into basic coding units.

[0150] An image can be divided into at least one sub-image or parallel block. The sub-image or parallel block can adopt a quadrilateral (rectangle or square) division form and can be divided into basic coding units. The sub-image is similar to the parallel block in terms of adopting the same division form (quadrilateral). However, the sub-image is different from the parallel block and can be distinguished from the parallel block in terms of adopting a single independent coding / decoding setting. That is, the parallel block receives the setting information for performing coding / decoding from a superior unit (such as an image, etc.), while the sub-image can directly obtain at least one setting information for performing coding / decoding from the header information of each sub-image. That is, different from the sub-image, the parallel block is only a unit obtained by dividing the video, and is not a unit for transmitting data {such as the basic unit of the video coding layer (VCL)}.

[0151] In addition, the parallel block can be a division unit supported from the perspective of parallel processing, while the sub-image can be a division unit supported from the perspective of independent coding / decoding. Specifically, the sub-image can not only perform coding / decoding setting in units of sub-images, but also determine whether to perform coding / decoding, and can form and display the corresponding sub-image centered on the region of interest, and the related settings can be determined in units such as sequences and images.

[0152] In the above example, it can also be changed to a method of performing coding / decoding setting at the superior unit of the sub-image and performing independent coding / decoding setting in units of parallel blocks. For the convenience of description in the present invention, it will be assumed that the parallel block can be set independently or depend on the superior unit for setting as an example for description.

[0153] The segmentation information generated when the image is segmented in a quadrilateral form can adopt various forms.

[0154] Refer to Figure 6a, a quadrilateral-shaped segmentation unit can be obtained by dividing the image once along the horizontal line b7 and the vertical line (at this time, b1 and b3, b2 and b4 will divide to form a dividing line). For example, the quantity information of quadrilaterals can be generated respectively based on the horizontal and vertical directions. At this time, when the quadrilaterals are evenly divided, the horizontal length and vertical length of the image can be divided by the number of horizontal lines and vertical lines respectively to confirm the horizontal and vertical lengths of the divided quadrilaterals. When the quadrilaterals are not evenly divided, information for indicating the horizontal and vertical lengths of the quadrilaterals can be additionally generated. At this time, the horizontal and vertical lengths can be represented in one-pixel unit or in pixel units. In the case of representing in multiple pixel units, for example, when the size of the basic coding unit is M×N and the size of the quadrilateral is 8M×4N, the horizontal and vertical lengths can be represented as 8 and 4 respectively (when the basic coding unit of the corresponding segmentation unit in this example is M×N), or represented as 16 and 8 (when the basic coding unit of the corresponding segmentation unit in this example is M / 2×N / 2).

[0155] In addition, referring to Figure 6b , it is possible to confirm the case where a quadrilateral-shaped segmentation unit is obtained by independently dividing the image, which is different from Figure 6a . For example, the quantity information of quadrilaterals in the image, the starting position information of each quadrilateral in the horizontal and vertical directions (the positions indicated by the reference numerals z0 to z5 in the drawings, which can be represented by the x, y coordinates in the image), and the horizontal and vertical length information of each quadrilateral can be generated. At this time, the starting position can be represented in one-pixel unit or multiple pixel units, and the horizontal and vertical lengths can also be represented in one-pixel unit or multiple pixel units.

[0156] Figure 6a can be an example of the segmentation information of a parallel block or a sub-image, and Figure 6b can be an example of the segmentation information of a sub-image, but is not limited thereto. Next, for the convenience of explanation, it will be assumed that the segmentation unit in the shape of a quadrilateral is used to explain the parallel block, but the explanations related to the parallel block can be applied to the sub-image in the same or similar manner (in addition, it can also be applied to the surface). That is, in the present invention, only a distinction is made by the difference in terms. In fact, the explanations related to the parallel block can be used as the definition of the sub-image, and the explanations related to the sub-image can also be used as the definition of the parallel block.

[0157] Among the segmentation units mentioned above, some are not required to be included. All or a part of them can be selectively included according to the coding / decoding settings, and other additional units (such as the surface) can also be supported.

[0158] In addition, the coding units (or blocks) can be divided into various sizes by block division. At this time, the coding units can be composed of multiple coding blocks according to the color format (for example, one luminance coding block and two chrominance coding blocks, etc.). For the convenience of explanation, it will be assumed that one color component unit is used for explanation. The coding blocks can have variable sizes such as M×M (for example, M is 4, 8, 16, 32, 64, 128, etc.). In addition, according to the division method (for example, tree structure division, i.e., quadtree <QuadTree, QT> division, binary tree <Binary Tree, BT> division, ternary tree <Ternary Tree, TT> division, etc.), the coding blocks can be divided into variable sizes of M×N (for example, M and N are 4, 8, 16, 32, 64, 128, etc.). At this time, the coding blocks can be the units that are the basis for in-picture prediction, inter-picture prediction, transformation, quantization, entropy coding, etc.

[0159] Although in the present invention, the case where a plurality of sub-blocks (symmetric) of the same size and shape are obtained by the division method is assumed as an example for explanation, it can also be applied to the case including asymmetric sub-blocks (for example, the horizontal ratio <vertical is the same> between the divided blocks is 1:3 or 3:1 or the vertical ratio <horizontal is the same> is 1:3 or 3:1, etc. in binary tree division, and the horizontal ratio <vertical is the same> between the divided blocks is 1:2:1 or the vertical ratio <horizontal is the same> is 1:2:1, etc. in ternary tree division).

[0160] The division (M×N) of the coding blocks can adopt a recursive tree structure. At this time, whether to divide can be indicated by a division flag. For example, when the division flag of the coding block with a division depth (Depth) of k is 0, the coding of the coding block is performed on the coding block with a division depth of k, and when the division flag of the coding block with a division depth of k is 1, the coding of the coding block will be performed on 4 sub-coding blocks (quadtree division) or 2 sub-coding blocks (binary tree division) or 3 sub-coding blocks (ternary tree division) with a division depth of k + 1 according to the division method.

[0161] The above sub-coding blocks will be reset as coding block k + 1 and can be divided into sub-coding blocks k + 2 again through the above process, and the division flag (for example, used to indicate whether to divide) can be supported in quadtree division.

[0162] In binary tree segmentation, it is possible to support a segmentation flag and a segmentation direction flag (horizontal or vertical). When more than one segmentation ratio is supported in binary tree segmentation (for example, supporting an additional segmentation ratio other than a horizontal or vertical ratio of 1:1, i.e., asymmetric segmentation), it is also possible to support a segmentation ratio flag (for example, selecting one ratio from a candidate group of horizontal or vertical ratios <1:1, 1:2, 2:1, 1:3, 3:1>), or support other forms of flags (for example, whether it is symmetric segmentation. When it is 1, it is symmetric segmentation without additional information, and when it is 0, it is asymmetric segmentation and requires additional information related to the ratio).

[0163] In ternary tree segmentation, it is possible to support a segmentation flag and a segmentation direction flag. When more than one segmentation ratio is supported in ternary tree segmentation, the same additional segmentation information as in the above binary tree segmentation will be required.

[0164] The above example is the segmentation information generated when only one type of tree-like segmentation method is effective. When multiple tree-like segmentation methods are effective, the following described segmentation information can be constructed.

[0165] For example, when multiple tree-like segmentations are supported, in the presence of a preset segmentation priority order, it is possible to first construct segmentation information corresponding to the priority order. At this time, when the segmentation flag corresponding to the priority order is true, it can further include additional segmentation information related to the corresponding segmentation method, and when it is false (no segmentation is performed), it can be composed of the segmentation information (segmentation flag, segmentation direction flag, etc.) of the segmentation method corresponding to the next order.

[0166] Alternatively, when multiple tree-like segmentations are supported, it is possible to additionally generate selection information related to the segmentation method, and it is composed of the segmentation information related to the selected segmentation method.

[0167] Some of the above segmentation flags can be omitted according to the results of earlier executed upper-level or previous segmentations.

[0168] Block segmentation can be performed from the largest coding block to the smallest coding block. Alternatively, it can also be performed from the smallest segmentation depth of 0 to the largest segmentation depth. That is, it is possible to recursively perform segmentation before the block size reaches the smallest coding block size or the segmentation depth reaches the largest segmentation depth. At this time, it is possible to be based on the coding / decoding settings (for example, video <strip, parallel block> type , Coding mode <Intra / Inter>, color difference component <y cb cr>etc.), adaptively set the size of the largest coding block, the size of the smallest coding block, and the maximum splitting depth.

[0169] For example, when the largest coding block is 128×128, quadtree splitting can be performed in the range of 32×32 to 128×128, binary tree splitting can be performed in the range of 16×16 to 64×64 and within the range of a maximum splitting depth of 3, and ternary tree splitting can be performed in the range of 8×8 to 32×32 and within the range of a maximum splitting depth of 3. Or, quadtree splitting can be performed in the range of 8×8 to 128×128, while binary tree splitting and ternary tree splitting can be performed in the range of 4×4 to 128×128 and with a maximum splitting depth of 3. The former can be a setting for the I picture type (such as a slice), and the latter case can be a setting for the P or B picture type.

[0170] As described in the above example, splitting settings such as the largest coding block size, the smallest coding block size, and the maximum splitting depth can adopt common or independent settings according to the splitting method and the coding / decoding settings as described above.

[0171] When multiple splitting methods are supported, splitting will be performed within the block support range of each splitting method. When the block support ranges of each splitting method overlap, priority order information of the splitting method can be included. For example, quadtree splitting can be performed prior to binary tree splitting.

[0172] Or, in the case where the splitting support ranges overlap, splitting selection information can be generated. For example, selection information related to the splitting method that needs to be performed between binary tree splitting and ternary tree splitting can be generated.

[0173] In addition, when multiple splitting methods are supported, it is possible to determine whether to perform the later splitting based on the result of the earlier splitting. For example, when the result of the earlier splitting (quadtree splitting in this example) indicates that splitting is to be performed, it is possible not to perform the later splitting (binary tree splitting or ternary tree splitting in this example), but to continue splitting after re-setting the sub-coding blocks split by the earlier splitting as coding blocks.

[0174] Alternatively, when the result of an earlier performed segmentation indicates not to perform segmentation, segmentation can be performed according to the result of a later performed segmentation. At this time, when the result of the later performed segmentation (in this example, binary tree segmentation or ternary tree segmentation) indicates to perform segmentation, segmentation can be continued after re-setting the segmented sub-coding blocks as coding blocks, and when the result of the later performed segmentation indicates not to perform segmentation, segmentation can be stopped. At this time, when the result of the later performed segmentation indicates to perform segmentation and multiple segmentation methods are still supported when re-setting the segmented sub-coding blocks as coding blocks (for example, when the block support ranges of the respective segmentation methods overlap), the earlier performed segmentation can be not performed and only the later performed segmentation can be performed. That is, when multiple segmentation methods are supported, if the result of the earlier performed segmentation indicates not to perform segmentation, the earlier performed segmentation can be stopped from being performed again.

[0175] For example, when an M×N coding block can perform quadtree segmentation and binary tree segmentation, the quadtree segmentation flag can be first confirmed, and when the above-mentioned segmentation flag is 1, it can be segmented into 4 sub-coding blocks of size (M>>1)×(N>>1), and then segmentation (quadtree segmentation or binary tree segmentation) can be performed after re-setting the above-mentioned sub-coding blocks as coding blocks. When the above-mentioned segmentation flag is 0, the binary tree segmentation flag can be confirmed, and when the corresponding flag is 1, it can be segmented into 2 sub-coding blocks of size (M>>1)×N or M×(N>>1), and then segmentation (binary tree segmentation) can be performed after re-setting the above-mentioned sub-coding blocks as coding blocks. When the above-mentioned segmentation flag is 0, the segmentation process will end and coding will be performed.

[0176] The above example illustrates the case of performing multiple segmentation methods, but it is not limited thereto, and combinations of multiple segmentation methods can also be supported. For example, segmentation methods such as quadtree / binary tree / ternary tree / quadtree + binary tree / quadtree + binary tree + ternary tree can be used. At this time, information related to whether additional segmentation methods are supported can be implicitly determined or explicitly included in units such as sequences, images, sub-images, stripes, parallel blocks, etc.

[0177] In the above example, information related to segmentation such as the size information of the coding block, the support range of the coding block, the maximum segmentation depth, etc. can be included in units such as sequences, images, sub-images, stripes, parallel blocks, etc. or implicitly determined. In other words, the range of admissible blocks can be determined based on the size of the maximum coding block, the range of supported blocks, the maximum segmentation depth, etc.

[0178] The coded blocks obtained by performing segmentation through the above process can be set to the maximum size for intra-picture prediction or inter-picture prediction. That is, in order to perform intra-picture prediction or inter-picture prediction, the coded blocks after block segmentation can be the starting size of the segmentation of the prediction blocks. For example, when the coded block is 2M×2N, the size of the prediction block can be the same as it or relatively smaller sizes of 2M×2N, M×N. Or, it can be the sizes of 2M×2N, 2M×N, M×2N, M×N. Or, it can be the size of 2M×2N which is the same as the coded block. At this time, the same size of the coded block and the prediction block can mean not performing the segmentation of the prediction block, but using the size obtained by the segmentation of the coded block to perform prediction. That is, it means not generating the segmentation information for the prediction block. The above-mentioned setting can also be applied to transform blocks, and the transformation can be performed in units of the segmented coded blocks.

[0179] Through the above-mentioned encoding / decoding settings, various configurations can be achieved. For example, (after determining the coded block), at least one prediction block and at least one transform block can be obtained based on the coded block. Or, one prediction block with the same size as the coded block can be obtained and at least one transform block can be obtained based on the coded block. Or, one prediction block with the same size as the coded block and one transform block can be obtained. When obtaining at least one block in the above examples, the segmentation information of each block can be generated, and when obtaining one block, the segmentation information of each block will not be generated.

[0180] The blocks of various sizes in the form of squares or rectangles obtained through the above results can be the blocks used in intra-picture prediction and inter-picture prediction, can also be the blocks used for transforming and quantizing the residual components, and can also be the blocks used in the filtering process.

[0181] The segmentation units obtained by segmenting an image using an image segmentation unit can perform independent encoding / decoding or dependent encoding / decoding according to the encoding / decoding settings.

[0182] Independent encoding / decoding can mean that when performing encoding / decoding on a part of the segmentation units (or regions), the data of other units cannot be used as a reference. Specifically, the information used or generated during the texture encoding and entropy encoding processes of a part of the units {such as pixel values or encoding / decoding information (intra-picture prediction-related information, inter-picture prediction-related information, and entropy encoding / decoding-related information, etc.)} will be encoded independently without mutual reference, and similarly, during the texture decoding and entropy decoding processes of a part of the units in the decoder, the parsing information and reconstruction information of other units will not be mutually referenced.

[0183] In addition, dependent encoding / decoding can refer to using the data of other units as a reference when performing encoding / decoding on a part of the segmentation units. Specifically, the information used or generated during the texture encoding and entropy encoding of a part of the units can be encoded dependently by mutual reference, and similarly, during the texture decoding and entropy decoding of a part of the units in the decoder, the parsed information and reconstructed information of other units can be mutually referred to.

[0184] Generally, the segmentation units (such as sub-images, parallel blocks, strips, etc.) mentioned above can adopt independent encoding / decoding settings. That is, a non-reference setting can be adopted for the purpose of parallelization. In addition, a non-reference setting can be adopted for the purpose of improving encoding / decoding performance. For example, when a 360-degree image is segmented into multiple surfaces in a three-dimensional space and arranged in a two-dimensional space, the correlation (such as image continuity) with adjacent surfaces may decrease according to the surface arrangement setting. That is, since the need for mutual reference is relatively low when there is no correlation between the surfaces, an independent encoding / decoding setting can be adopted.

[0185] In addition, a referenceable setting between segmentation units can be adopted for the purpose of improving encoding / decoding performance. For example, even when a 360-degree image is segmented into surface units, there may be a high correlation with adjacent surfaces according to the surface arrangement setting, and in this case, a dependent encoding / decoding setting can be adopted.

[0186] In addition, in the present invention, independent or dependent encoding / decoding can be applied not only to spatial regions but also extended to temporal regions. That is, independent or dependent encoding / decoding can be performed not only on other segmentation units that exist at the same time as the current segmentation unit, but also on segmentation units that exist at different times from the current segmentation unit (in this example, even if there is a segmentation unit at the same position in an image corresponding to a different time from the current segmentation unit, it is assumed to be another segmentation unit).

[0187] For example, when simultaneously transmitting bitstream A containing data encoding a 360-degree image with high quality and bitstream B containing data encoded with normal quality, the decoder can parse and decode bitstream A transmitted with high quality in the area corresponding to the area of interest (such as the area where the user's line of sight is focused <viewport> or the area to be displayed, etc.), and parse and decode bitstream B transmitted with normal quality outside the area of interest.

[0188] Specifically, in the case where an image is segmented into multiple units (such as sub-images, parallel blocks, stripes, surfaces, etc., and in this example, it is assumed that the surface is processed in the same way as parallel blocks or sub-images), it is possible to decode the data (bitstream A) of the segmented units included in the region of interest (or segmented units that overlap with the viewport by at least one pixel) and the data (bitstream B) of the segmented units outside the region of interest.

[0189] Alternatively, it is possible to transmit the bitstream containing the receipt for encoding the entire image, and on the decoder side, it is possible to parse and decode the region of interest from the bitstream. Specifically, it is possible to decode only the data of the segmented units included in the region of interest.

[0190] In other words, it is possible to obtain the entire or a part of the image by generating a bitstream divided into more than one image quality level in the encoder and decoding only a specific bitstream in the decoder, or by selectively decoding each bitstream in each image part. In the above example, the case of a 360-degree image is used as an example, but this is an explanation that can be applied to general images.

[0191] When performing encoding / decoding according to the above example, since it is not possible to know which data will be reconstructed in the decoder (in this example, the decoder does not know the position of the region of interest and is in a situation of random access based on the region of interest), it is necessary to confirm and perform encoding / decoding on the reference settings in the temporal region in addition to the spatial region.

[0192] For example, when the decoder determines which type of decoding to perform with a single segmented unit, the current segmented unit can perform independent encoding in the spatial region and limited dependent encoding in the temporal region (for example, only allowing reference to the segmented units at the same position in other times corresponding to the current segmented unit and prohibiting reference to other segmented units, because generally there are no restrictions in the temporal region, so this is a comparison with unrestricted dependent encoding).

[0193] Alternatively, when the decoder determines which type of decoding to perform in multiple segmentation units (multiple segmentation units can be obtained by bundling horizontally adjacent segmentation units or vertically adjacent segmentation units, and can also be obtained by bundling both horizontally and vertically adjacent segmentation units), that is, in this case, as long as any one of the segmentation units is included in the region of interest, decoding is performed on multiple units, the current segmentation unit can perform independent or dependent decoding in the spatial region and perform limited dependent encoding in the temporal region (for example, in addition to allowing reference to the segmentation units at the same position at other times corresponding to the current segmentation unit, it also allows reference to a part of other segmentation units).

[0194] In the present invention, a surface is a segmentation unit whose configuration and form usually change according to the projection format and has no independent encoding / decoding setting. Although it has different characteristics from other segmentation units described above, in terms of being able to divide an image into multiple regions (and having a quadrilateral form, etc.), it can also be regarded as a unit obtained in the image segmentation section.

[0195] As described above, in the spatial region, independent encoding / decoding can be performed on each segmentation unit for the purpose of parallelization, etc. However, since independent encoding / decoding cannot refer to other segmentation units, it will cause a problem of decreased encoding / decoding efficiency. Therefore, as a step before performing encoding / decoding, the segmentation unit performing independent encoding / decoding can be extended by using (or adding) the data of adjacent segmentation units. Among them, since the segmentation unit with the data of adjacent segmentation units added has more data for reference, its encoding / decoding efficiency will also be improved. At this time, since the extended segmentation unit can refer to the data of adjacent segmentation units during encoding / decoding, it can be regarded as dependent encoding / decoding.

[0196] The information related to the reference setting between the above-mentioned segmentation units can be included in the bitstream in units such as video, sequence, image, sub-image, strip, parallel block, etc. and transmitted to the decoder, and in the decoder, the setting information transmitted from the encoder can be reconstructed by parsing at the same level unit. In addition, the relevant information can also be transmitted to the bitstream in the form of supplementary enhancement information (SEI) or metadata, etc. and used after parsing. In addition, the definitions pre-agreed in the encoder / decoder can also be used to perform encoding / decoding according to the reference setting without transmitting the above information.

[0197] Figure 7 It is an illustrative diagram that divides an image into multiple parallel blocks. Figures 8a to 8i It is for Figure 7 the first illustrative diagram that sets additional regions for each of the parallel blocks illustrated in Figures 9a to 9i It is for Figure 7 the second illustrative diagram that sets additional regions for each of the parallel blocks illustrated in

[0198] When an image is divided into two or more segmentation units (or regions) by an image segmentation unit and independent encoding / decoding is performed on each segmentation unit, although there are advantages such as being able to perform parallel processing, etc., at the same time, there may be a problem that the encoding performance deteriorates due to a reduction in the data that each segmentation unit can refer to. To solve the above-mentioned problem, it is possible to perform processing through the encoding / decoding setting of the dependency between segmentation units (in this example, it will be described in terms of parallel block particles, and the same or similar settings can also be applied in other units).

[0199] Between segmentation units, independent encoding / decoding is usually performed in a non-referable manner. Therefore, it is possible to perform pre-processing or post-processing procedures for implementing dependent encoding / decoding. For example, it is possible to form an extended region on the outer contour of each segmentation unit before performing encoding / decoding and fill the data of other segmentation units that need to be referred to in the extended region.

[0200] Although the method described above has no other differences from the method of performing independent encoding / decoding except that encoding / decoding is performed after expanding each segmentation unit, since the existing segmentation units will obtain and refer to the data that needs to be referred to from other segmentation units in advance, it can be understood as an example of dependent encoding / decoding.

[0201] In addition, after performing encoding / decoding, it is possible to apply filtering using the data of multiple segmentation units based on the boundaries between the segmentation units. That is, when applying filtering, it belongs to the dependent situation because the data of other segmentation units is used, and when not applying filtering, it can belong to the independent situation.

[0202] In the examples described later, the case of performing dependent encoding / decoding by performing pre-processing of encoding / decoding (expansion in this example) will be mainly described. In addition, in the present invention, the boundary between the same segmentation units can be called an internal boundary, and the outer contour of the image can be called an external boundary.

[0203] In an embodiment where the present invention is applicable, an additional area related to the current parallel block can be set. Specifically, at least one parallel block (in this example, including the case where one image is composed of one parallel block, that is, including the case where it is not divided into two or more divided units. To be precise, although the divided unit means being divided into two or more units, it is assumed that it is also recognized as one divided unit when not divided) can be used as a reference to set the additional area.

[0204] For example, an additional area can be set in at least one of the directions such as above / below / left / right of the current parallel block. Among them, the additional area can be filled with any value. In addition, the additional area can be filled with a part of the data in the current parallel block, that is, it can be filled by using the outer contour pixels of the current parallel block or by copying the pixels within the current parallel block.

[0205] In addition, the additional area can be filled with the image data of other parallel blocks outside the current parallel block. Specifically, the image data in the parallel block adjacent to the current parallel block can be used, that is, it can be filled by copying the image data in the parallel block adjacent to the current parallel block in a specific direction among above / below / left / right.

[0206] At this time, the size (length) of the acquired image data can adopt the same value in each direction or can also adopt independent values, which can be determined according to the encoding / decoding settings.

[0207] For example, in Figure 6a it can be extended in all or some of the boundary directions of b0 to b8. In addition, it can be extended by m in all boundary directions of the divided unit or extended by m according to the boundary direction i (i is the index of each direction). m or mi can be applied to all divided units in the image or can also be set independently for each divided unit.

[0208] At this time, setting information related to the additional area can be generated. At this time, the setting information related to the additional area can be whether the additional area is supported, whether each segmentation unit supports the additional area, the shape of the additional area on the overall image (for example, determined according to which direction among up / down / left / right of the segmentation unit, and in this example, it is the setting information commonly applicable to all segmentation units within the image), the shape of the additional area on each segmentation unit (in this example, it is the setting information applicable to individual segmentation units within the image), the size of the additional area on the overall image (for example, after determining the shape of the additional area, it represents the degree of expansion in the expansion direction, and in this example, it is the setting information commonly applicable to all segmentation units within the image), the size of the additional area on each segmentation unit (in this example, it is the setting information independently applicable to individual segmentation units within the image), the method of filling the additional area on the overall image, the method of filling the additional area on each segmentation unit, etc.

[0209] The above settings related to the additional area can be determined proportionally according to the color space, or independent settings can also be adopted. Setting information related to the additional area can be generated on the luminance component, and the additional area setting on the chrominance component can be implicitly determined according to the color space. Or, setting information related to the additional area can also be generated on the chrominance component.

[0210] For example, when the size of the additional area of the luminance component is m, the size of the additional area of the chrominance component can be determined as m / 2 according to the color format (in this example, 4:2:0). As another example, when the size of the additional area of the luminance component is m and the chrominance component adopts independent settings, size information of the additional area of the chrominance component can be generated (in this example, n, and n can be commonly used or n1, n2, n3, etc. can be used according to the direction or the expansion area). As another example, a method of filling the additional area of the luminance component can be generated, and the method of filling the additional area of the chrominance component can use the method in the luminance component or generate relevant information.

[0211] The above information related to the additional area setting can be included in the bitstream in units such as video, sequence, image, sub-image, strip, etc. and transmitted, and during decoding, the relevant information can be parsed and reconstructed from the above units. In the embodiments described later, the case of supporting the additional area will be assumed for illustration.

[0212] Refer to Figure 7 , it can be confirmed that an image is segmented into parallel blocks labeled 0 to 8. At this time, the result of setting the additional area of one embodiment of the present invention for each of the parallel blocks illustrated in Figure 7 is as shown in Figures 8a to 8i .

[0213] In Figure 7 and Figure 8a In, the No. 0 parallel block (with a size of T0_W × TO_H) can be expanded by appending an area of E0_R to the right and an area of E0_D to the bottom. At this time, the appended areas can be obtained from adjacent parallel blocks. Specifically, the right-side expanded area can be obtained from the No. 1 parallel block, and the bottom-side expanded area can be obtained from the No. 3 parallel block. In addition, the No. 0 parallel block can set the appended areas by using the adjacent parallel block at the lower right (the No. 4 parallel block). That is, the appended areas can be set in the direction of the remaining internal boundaries (or the boundaries between the same segmentation units) except for the outer boundaries of the parallel blocks (or the image boundaries).

[0214] In Figure 7 and Figure 8e In, since the No. 4 parallel block (with a size of T4_W × T4_H) has no outer boundary, it can be expanded by appending areas to the left, right, top, and bottom. At this time, the left-side expanded area can be obtained from the No. 3 parallel block, the right-side expanded area can be obtained from the No. 5 parallel block, the top-side expanded area can be obtained from the No. 1 parallel block, and the bottom-side expanded area can be obtained from the No. 7 parallel block. In addition, the No. 4 expanded area can also set appended areas to the upper left, lower left, upper right, and lower right. At this time, the upper left expanded area can be obtained from the No. 0 parallel block, the lower left expanded area can be obtained from the No. 6 parallel block, the upper right expanded area can be obtained from the No. 2 parallel block, and the lower right expanded area can be obtained from the No. 8 parallel block.

[0215] In FIG. 8, since the L2 block is a block adjacent to the boundary of the parallel block, in principle, there is no data that can be referenced from the left, upper left, and lower left blocks. However, when an appended area is set for the No. 2 parallel block by applying an embodiment of the present invention, the L2 block can perform encoding / decoding by referring to the appended area. That is, the L2 block can refer to the data of the blocks located on the left and upper left as the appended area (which can be the area obtained from the No. 1 parallel block), and can refer to the data of the block located on the lower left as the appended area (which can be the area obtained from the No. 4 parallel block).

[0216] Data included in the additional area through the above-described embodiments can be included in the current parallel block for encoding / decoding. In this case, since the data in the additional area is located at the boundary of the parallel block (in this example, it refers to the parallel block updated or extended due to the additional area), there may also be a decrease in encoding performance during the encoding process due to the lack of reference data. However, since this is only an additional part provided to the boundary area of the original parallel block for reference, it can be understood as a form of temporary memory for improving encoding performance. That is, since it can help improve the image quality performance of the finally output image and is an area that will ultimately be removed, the decrease in encoding performance in the corresponding area will not cause any problems. This can be applied to the embodiments described later for similar or the same purpose.

[0217] In addition, referring to Figures 9a to 9i , it can be confirmed that the 360-degree image is changed into a 2D image through a surface configuration (or reconfiguration) process according to the projection format and the 2D image is divided into respective parallel blocks (which can also be surfaces). At this time, since the 2D image consists of one surface when the 360-degree image adopts the equirectangular projection, it can be an example of dividing one surface into parallel blocks. In addition, for the convenience of explanation, it is assumed that the parallel block division of the 2D image is the same as the parallel block division illustrated in Figure 7 .

[0218] Among them, the divided parallel blocks can be divided into parallel blocks composed only of internal boundaries and parallel blocks containing at least one external boundary, and additional areas can be set for each parallel block in the manner shown in Figures 8a to 8i . However, even if the 360-degree image transformed into a 2D image is adjacent to each other in the 2D image, there may be no continuity in the actual image, and even if they are not adjacent, there may be continuity in the actual image (refer to the explanation of Figures 5a to 5c ). Therefore, even if a part of the boundary of the parallel block is an external boundary, there may be areas in the image that are continuous with the external boundary area of the parallel block. Specifically, referring to Figure 9b , although the upper end of the first parallel block is the external boundary of the image, since there may be areas in the same image that are continuous in the actual image, an additional area can be set at the upper end of the first parallel block. That is, different from Figures 8a to 8i , in Figures 9a to 9i , additional areas can also be set for all or part of the external boundary direction of the parallel block.

[0219] Referring to Figure 9e , the 4th parallel block is a parallel block whose parallel block boundary only contains internal boundaries (in this example, it is the 4th parallel block). Therefore, the additional area of the 4th parallel block can be set in all directions of the upper side, lower side, left side, right side, as well as the upper left, lower left, upper right, and lower right directions. Among them, the left extended area can be the image data obtained from the 3rd parallel block, the right extended area can be the image data obtained from the 5th parallel block, the upper extended area can be the image data obtained from the 1st parallel block, the lower extended area can be the image data obtained from the 7th parallel block, the upper left extended area can be the image data obtained from the 0th parallel block, the lower left extended area can be the image data obtained from the 6th parallel block, the upper right extended area can be the image data obtained from the 2nd parallel block, and the lower right extended area can be the image data obtained from the 9th parallel block.

[0220] Refer to Figure 9a , the 0th parallel block is a parallel block that contains at least one external boundary (left side, upper side direction). Therefore, in addition to the right side, lower side, and lower right directions that are spatially adjacent, the 0th parallel block can also contain an additional area that extends in the direction of the external boundary (left side, upper side, upper left direction). Among them, the additional area in the spatially adjacent right side, lower side, and lower right directions can be set using the data of the adjacent parallel blocks, but this is not possible for the additional area in the external boundary direction. At this time, for the additional area in the external boundary direction, it can be set using data that is not adjacent in the space within the image but has continuity in the actual image. For example, when the projection format of the 360-degree image is equirectangular projection, the left boundary of the image and the right boundary of the image have continuity in the actual image, and the upper boundary of the image and the lower boundary of the image have continuity in the actual image, the left boundary direction of the 0th parallel block and the right boundary of the 2nd parallel block have continuity, and the upper boundary direction of the 0th parallel block and the lower boundary of the 6th parallel block have continuity. Therefore, in the 0th parallel block, the left extended area can be obtained from the 2nd parallel block, the right extended area can be obtained from the 1st parallel block, the upper extended area can be obtained from the 6th parallel block, and the lower extended area can be obtained from the 3rd parallel block. In addition, in the 1st parallel block, the upper left extended area can be obtained from the 8th parallel block, the lower left extended area can be obtained from the 5th parallel block, the upper right extended area can be obtained from the 7th parallel block, and the lower right extended area can be obtained from the 4th parallel block.

[0221] Because Figure 9a The L0 block in it is a block located at the boundary of parallel blocks. Therefore, data that can be referenced from the left, upper left, lower left, upper, and upper right blocks (similar to the case of U0) may not exist. At this time, even if they are not spatially adjacent in the 2D image, there can still be blocks that are continuous in the actual image within the 2D image (or picture). Therefore, as mentioned in the above premise, when the projection format of the 360-degree image is equirectangular projection, the left boundary of the image and the right boundary of the image are continuous in the actual image, and the upper boundary of the image and the lower boundary of the image are continuous in the actual image, the left and lower left blocks of the L0 block can be obtained from the 2nd parallel block, the upper left block of the L0 block can be obtained from the 8th parallel block, and the upper and upper right blocks of the L0 block can be obtained from the 6th parallel block.

[0222] The following Table 1 is the pseudo code for obtaining data corresponding to the additional area from other continuous areas.

[0223]

Table 1

[0224] i_pos' = overlap(i_pos, minI, maxI)

[0225] overlap(A, B, C)

[0226] {

[0227] if (A < B) output = (A + C - B + 1) % (C - B + 1)

[0228] else if (A > C) output = A % (C - B + 1)

[0229] else output = A

[0230] }

[0231] Referring to the pseudo code in Table 1, the variable i_pos (corresponding to the variable A) of the overlap function is the input pixel position, i_pos' is the output pixel position, minI (corresponding to the variable B) is the minimum value of the pixel position range, maxI (corresponding to the variable C) is the maximum value of the pixel position range, and i is the position component (in this example, horizontal, vertical, etc.). In this example, minI can be 0, and maxI can be Pic_width (the horizontal width of the image) - 1 or Pic_height (the vertical width of the image) - 1.

[0232] For example, assuming that the vertical width range of the image (general picture) is 0 to 47 and the image is as Figure 7 It is segmented in the manner shown. When it is necessary to set an additional area of size m below the parallel block 4 and fill the additional area with the upper end data of the parallel block 7, it is possible to confirm from which position to obtain the data through the above method.

[0233] When the vertical length range of the parallel block 4 is 16 - 30 and it is necessary to set an additional area of size 4 below it, the data at positions 31, 32, 33, and 34 corresponding to it can be filled into the additional area of the parallel block 4. At this time, since min and max in the above formula are 0 and 47 respectively, the output values of 31 - 34 will be their own values, that is, 31 - 34. That is, the data to be filled into the additional area is the data at positions 31 - 34.

[0234] Or, assume that the horizontal length range of an image (360 - degree image, equirectangular projection, with continuity at both ends of the image) is 0 - 95 and the image is segmented in the manner shown. Figure 7 When it is necessary to set an additional area of size m to the left of the parallel block 3 and fill the additional area with the right - hand side data of the parallel block 5, it is possible to confirm from which position to obtain the data through the above method.

[0235] When the vertical length range of the parallel block 3 is 0 - 31 and it is necessary to set an additional area of size 4 to the left, the data at positions - 4, - 3, - 2, and - 1 corresponding to it can be filled into the additional area of the parallel block 3. Since the above positions do not exist within the horizontal length range of the image, it is necessary to calculate through the above well - known method from which position to obtain the data. At this time, since min and max in the above formula are 0 and 95 respectively, the output values of - 4 to - 1 will be 92 - 95. That is, the data to be filled into the additional area is the data at positions 92 - 95.

[0236] Specifically, when the area of size m is data between 360 degrees and 380 degrees (assuming the range of pixel value positions is 0 degrees to 360 degrees in this example), by adjusting it to the internal range of the image, it can be understood as similar to the case of obtaining data from the area between 0 degrees and 20 degrees. That is, it can be obtained based on the pixel value position range between 0 and Pic_width - 1.

[0237] In other words, in order to obtain the data of the additional area, it is possible to confirm the position of the data to be obtained through an overlapping process.

[0238] The above example illustrates the case of obtaining a surface from a 360-degree image, (except for the case where the image boundaries are continuous at both ends), and on the premise that adjacent regions in the image space are continuous with each other. However, depending on different projection formats (such as cube map projection, etc.), in the case of including two or more surfaces and each surface undergoing a configuration or reconfiguration process, there can be a situation where there is no continuity even though they are adjacent in the image space. In the above-mentioned case, it is possible to confirm and generate additional regions for the position data that is continuous on the actual image through the surface configuration or reconfiguration information.

[0239] The following Table 2 is pseudo-code for generating an additional region related to the above specific segmentation unit using the internal data of the specific segmentation unit.

[0240]

Table 2

[0241] i_pos' = clip(i_pos, minI, maxI)

[0242] clip(A, B, C)

[0243] {

[0244] if (A < B) output = B

[0245] else if (A > C) output = C

[0246] else output = A

[0247] }

[0248] Since the meanings of the variables in Table 2 are the same as those in Table 1, their detailed descriptions will be omitted here. However, in this example, minI can be the left or upper coordinate of the specific segmentation unit, and maxI can be the right or lower coordinate of each unit.

[0249] For example, when the image is segmented as shown in Figure 7 and the horizontal length range of parallel block 2 is 32 - 47, and an additional region of size m needs to be set to the right of parallel block 2, the data at positions 48, 49, 50, and 51 can be filled with the data at position 47 (corresponding to the inside of parallel block 2) output by the above formula. That is, according to Table 2, the additional region related to the specific segmentation unit can be generated by copying the outer contour pixels of the corresponding segmentation unit.

[0250] In other words, in order to obtain the data of the additional region, it is possible to confirm the position to be obtained through a clipping process.

[0251] The detailed composition of Table 1 or Table 2 above is not fixed but can be changed. For example, the 360-degree image can change the applicable overlapping method in consideration of the surface configuration (or reconfiguration) and the coordinate system characteristics between surfaces.

[0252] Figure 10 It is an illustrative diagram for applying the additional region generated in one embodiment of the present invention during the encoding / decoding process of other regions.

[0253] In addition, since the additional region of one embodiment of the present invention is generated using the image data in other regions, it can be equivalent to duplicate image data. Therefore, in order to prevent the existence of unnecessary duplicate data, the additional region can be removed after encoding / decoding. However, before removing the additional region, it can be considered to remove it after applying the additional region to the encoding / decoding.

[0254] Refer to Figure 10 , it can be confirmed that the additional region B of the segmentation unit J is generated using the region A of the segmentation unit I. At this time, before removing the generated region B, the region B can be applied to the encoding / decoding (specifically, the reconstruction or correction process) of the region A included in the segmentation unit I.

[0255] Specifically, assuming that the segmentation unit I and the segmentation unit J are respectively Figure 7 the 0th parallel block and the 1st parallel block in, the rightmost part of the segmentation unit I and the left part of the segmentation unit J have image continuity with each other. Among them, after the additional region B is used as a reference in the image encoding / decoding of the segmentation unit J, it can also be used when encoding / decoding A. In particular, although A and the region B obtained data from the region A when generating the additional region, different values (including quantization errors) may be used for reconstruction during the encoding / decoding process. Therefore, when reconstructing the segmentation unit I, the part corresponding to the region A can be reconstructed using the image data of the reconstructed region A and the image data of the region B. For example, a part of the region C of the segmentation unit I can be replaced by the average or weighted value of the region A and the region B. This is because there are data of two or more identical regions, so the image (region C, where A in the segmentation unit I is replaced by C) can be reconstructed using the data of the two regions (the process at this time is named Rec_Process in the attached figure).

[0256] In addition, a part of region C included in segmentation unit I can be replaced by using regions A and B according to which segmentation unit it is closer to. Specifically, since the image data within a certain range (e.g., M pixel intervals) on the left side in region C is closer to segmentation unit I, the data in region A can be used (or copied) for reconstruction. Also, since the image data within a certain range (e.g., N pixel intervals) on the right side in region C is closer to segmentation unit J, the data in region B can be used (or copied) for reconstruction. The result of expressing this as a formula is shown in Formula 1 below.

[0257]

Formula 1

[0258] C(x,y) = A(x,y), (x,y) ∈ M

[0259] B(x,y), (x,y) ∈ N

[0260] In addition, a part of region C included in segmentation unit I can be replaced by assigning weight values to the image data in regions A and B respectively according to which segmentation unit it is closer to. That is, for the image data in region C that is closer to segmentation unit I, a higher weight value can be assigned to the image data in region A, and for the image data that is closer to segmentation unit J, a higher weight value can be assigned to the image data in region B. That is, the weight value can be set based on the distance difference between the horizontal width of region C and the x coordinate of the pixel value to be corrected.

[0261] As a formula for setting adaptive weight values for regions A and B, Formula 2 as described below can be derived.

[0262]

Formula 2

[0263] C(x,y) = A(x,y) × w + B(x,y) × (1 - w)

[0264] w = f(x,y,k)

[0265] Referring to Formula 2, w is the weight value assigned to the pixels of regions A and B as (x,y). At this time, as the average of the weight values for regions A and B, the pixels in region A are multiplied by the weight value w and the pixels in region B are multiplied by 1 - w. However, in addition to the average of the weight values, different weight values can also be assigned to regions A and B respectively.

[0266] After the use of the additional area is completed according to the description as above, the additional area B can be removed during the process of re-sizing (Resing) the size of the segmentation unit J and stored in the memory (decoded picture buffer, DPB). (In this example, it is assumed that the process of setting the additional area is resizing.) <sizing>, The above process can be deduced from a part of the above embodiments (for example, through processes such as additional region flag confirmation, subsequent size information confirmation, subsequent filling method confirmation, etc.). Assume that the process of increasing size is executed during the size adjustment process, while on the contrary, the process of decreasing size is executed during the size readjustment process (which can be deduced from the reverse process of the above process).

[0267] In addition, (specifically, immediately after the encoding / decoding of the corresponding image is completed), it is also possible to directly store it in the memory without performing size readjustment, and then in the output step (assumed to be display in this example) <display>In step ), it is removed by performing a size adjustment process. This can be applied to all or part of the segmentation units included in the corresponding image.

[0268] The above-mentioned relevant setting information can be processed implicitly or explicitly according to the encoding / decoding settings. When using the implicit method (specifically, based on the characteristics, type, format, etc. of the image or according to other encoding / decoding settings <in this example, the settings related to the additional area>), it can be determined without generating relevant syntax elements. When using the explicit method, the settings related to the removal of the additional area can be adjusted by generating relevant syntax elements. The units related to this can include video, sequence, image, sub-image, strip, parallel block, etc.

[0269] In addition, the existing encoding method based on segmentation units can include: 1) a step of dividing an image into one or more parallel blocks (or can be collectively referred to as segmentation units) and generating segmentation information; 2) a step of performing encoding according to the divided parallel block units; 3) a step of performing filtering with information indicating whether loop filtering at the parallel block boundary is allowed; and 4) a step of storing the filtered parallel blocks in a memory.

[0270] In addition, the existing decoding method based on segmentation units can include: 1) a step of dividing an image into one or more parallel blocks based on the parallel block segmentation information; 2) a step of performing decoding according to the divided parallel block units; 3) a step of performing filtering with information indicating whether loop filtering at the parallel block boundary is allowed; and 4) a step of storing the filtered parallel blocks in a memory.

[0271] Among them, the third step in the encoding / decoding method is a post-processing step of encoding / decoding. When filtering is performed, it can be dependent encoding / decoding, and when filtering is not performed, it can be independent encoding / decoding.

[0272] The encoding method of the segmentation unit applicable to one embodiment of the present invention can include: 1) a step of dividing an image into one or more parallel blocks and generating segmentation information; 2) a step of setting an additional area for at least one of the divided parallel block units and filling the additional area with adjacent parallel block units; 3) a step of performing encoding on the parallel block units including the additional area; 4) a step of removing the additional area of the parallel block units and performing filtering based on the information indicating whether loop filtering at the parallel block boundary is allowed; and 5) a step of storing the filtered parallel blocks in a memory.

[0273] In addition, the decoding method of the segmentation unit applicable to an embodiment of the present invention may include: 1) a step of dividing an image into one or more parallel blocks based on parallel block division information; 2) a step of setting an additional area for the divided parallel block unit and filling the additional area with decoding information, preset information, or other (adjacent) parallel block units that have been previously reconstructed; 3) a step of performing encoding on the parallel block unit including the additional area using the decoding information received from the encoding device; 4) a step of removing the additional area of the parallel block unit and performing filtering based on information indicating whether loop filtering is allowed at the parallel block boundary; and 5) a step of storing the filtered parallel block in a memory.

[0274] In the encoding / decoding method of the segmentation unit applicable to an embodiment of the present invention as described above, the second step may be a pre-processing process for encoding / decoding (dependency encoding / decoding when setting the additional area, otherwise independent encoding / decoding). In addition, the fourth step may be a post-processing process for encoding / decoding (dependency when performing filtering, otherwise independent). In this example, the additional area will be used during the encoding / decoding process, and a process of adjusting the size to the initial size of the parallel block will be performed before being stored in the memory.

[0275] First, the encoder divides the image into multiple parallel blocks. According to an implicit or explicit setting, an additional area is set for the parallel block unit and relevant data is obtained from adjacent areas. Next, encoding is performed on the updated parallel block unit including the original parallel block and the additional area. After encoding is completed, the additional area is removed and filtering is performed according to the loop filtering application setting.

[0276] At this time, different above-mentioned filtering settings can be used according to the filling method and removal method of the additional area. For example, in the case of simple removal, the above-mentioned loop filtering application setting can be followed, while when removing using an overlapping area, filtering may not be applied or other filtering settings may be followed. That is, because a large amount of overlapping data can be used to significantly reduce the distortion phenomenon in the parallel block boundary area, filtering may not be performed regardless of whether loop filtering is applied to the parallel block boundary unit, or different settings can be applied while complying with the above-mentioned filtering application or not (for example, applying a filter with a weaker filtering intensity at the parallel block boundary) for the filtering setting inside the parallel block. After the above process, it is stored in the memory.

[0277] In the decoder, first, the image is segmented into a plurality of parallel blocks according to the parallel block segmentation information transmitted from the encoder. Next, the information related to the additional area is confirmed explicitly or implicitly, and after setting the additional area, the encoded information of the updated parallel blocks transmitted from the encoder is analyzed. Next, decoding is performed in units of the updated parallel blocks. After decoding is completed, the additional area is removed and filtering is performed according to the same loop filter application setting as the encoder. The detailed information related thereto has been described in the encoder section, so the detailed description thereof will be omitted here. After the above process, it is stored in the memory.

[0278] In addition, it is also possible to consider the case where the additional area in the segmentation unit is not removed but directly stored in the memory during the encoding / decoding process. For example, in the case of 360-degree images, etc., there may be a problem of a decrease in prediction accuracy in part of the prediction process (such as inter-picture prediction) according to surface configuration settings, etc. (for example, it is difficult to accurately find at positions where the surface configuration is discontinuous during motion search and compensation). Therefore, the additional area can be stored in the memory and used during the prediction process to improve prediction accuracy. When used in inter-picture prediction, the additional area (or the image including the additional area) can be used as a reference image for performing inter-picture prediction.

[0279] The encoding method for storing the additional area may include: 1) a step of segmenting the image into one or more parallel blocks and generating segmentation information; 2) a step of setting an additional area for at least one of the segmented parallel block units and filling the additional area with adjacent parallel block units; 3) a step of encoding the parallel block unit including the additional area; 4) a step of storing the additional area of the parallel block unit (at this time, the application of loop filtering can be omitted); and 5) a step of storing the encoded parallel block in the memory.

[0280] The decoding method for storing the additional area may include: 1) a step of segmenting the image into one or more parallel blocks based on the parallel block segmentation information; 2) a step of setting an additional area for the segmented parallel block units and filling the additional area with decoding information, pre-set information, or other (adjacent) parallel block units that have been pre-reconstructed; 3) a step of encoding the parallel block unit including the additional area using the decoding information received from the encoding device; 4) a step of storing the additional area of the parallel block unit (at this time, loop filtering can be omitted); and 5) a step of storing the decoded parallel block in the memory.

[0281] When storing the additional area, the encoder first divides the image into multiple parallel blocks. According to an implicit or explicit setting, an additional area is set for the parallel blocks and relevant data is obtained from a preset area. The preset area refers to other areas that are relevant according to the surface configuration of the 360-degree image, so it can be an area adjacent to or not adjacent to the current parallel block. Next, encoding is performed in units of the updated parallel blocks. Since the additional area will be stored after decoding is completed, filtering is not performed regardless of the state of the loop filter application setting. This is because the boundaries of each updated parallel block will not share the actual parallel block boundaries due to the additional area. After the above process, it is stored in the memory.

[0282] When storing the additional area, the decoder first confirms the parallel block division information transmitted from the encoder and divides the image into multiple parallel blocks based on this. Next, information related to the additional area is confirmed, and after setting the additional area, the encoded information of the updated parallel blocks transmitted from the encoder is parsed. Next, decoding is performed in units of the updated parallel blocks. After decoding is completed, it is directly stored in the memory without applying loop filtering to the additional area.

[0283] Next, the encoding / decoding method of the division unit applicable to one embodiment of the present invention as described above will be described with reference to the accompanying drawings.

[0284] Figures 11 to 12 It is a flowchart for explaining the encoding / decoding method of the division unit applicable to one embodiment of the present invention. Specifically, as an example of generating an additional area and performing encoding / decoding in each division unit, the encoding method including the additional area is illustrated in Figure 11 and the decoding method for removing the additional area is illustrated in Figure 12 Among them, the 360-degree image can perform a preprocessing process (stitching, projection, etc.) before the step in Figure 11 and perform a postprocessing process (rendering, etc.) after the step in Figure 12

[0285] First, refer to Figure 11 , after the encoder obtains the input image (step A), the input image is segmented into two or more segmentation units by the image segmentation unit (at this time, setting information related to the segmentation method can be generated, denoted as step B). Next, an additional area is generated according to the encoding setting or whether additional areas are supported as segmentation units (step C). Then, encoding is performed on the segmentation units including the additional area to generate a bitstream (step D). In addition, after the bitstream is generated, it is possible to determine whether size readjustment is required according to the encoding setting (or whether the additional area is deleted, step E). Then, the encoded data including or without the additional area (the image in step D or E) is stored in the memory (step E).

[0286] See Figure 12 , the decoder refers to the segmentation-related setting information obtained by parsing the received bitstream and segments the image to be decoded into two or more segmentation units (step B). Next, the size of the additional area is set for each segmentation unit according to the decoding setting obtained from the received bitstream (step C). Then, the image data including the additional area is obtained by decoding the image data included in the bitstream (step D). Next, a reconstructed image is generated by deleting the additional area (step E). Then, the reconstructed image is output to the display (step F). At this time, it is possible to determine whether to delete the additional area according to the decoding setting, and the decoded image or image data (the data in step D or E) is stored in the memory. In addition, step F can include a process of reducing the reconstructed image to a 360-degree image by surface reconfiguration.

[0287] In addition, according to Figure 11 or Figure 12 whether the additional area is removed, loop filtering can be adaptively performed at the segmentation unit boundary (assuming a block filter in this example, and other loop filters can also be applied). In addition, loop filtering can be adaptively performed according to whether additional areas are allowed to be generated.

[0288] When storing in the memory after removing the additional area, it is possible to explicitly apply or not apply loop filtering according to the loop filtering applicability flag at the segmentation unit boundary such as loop_filter_across_enabled_flag (specifically, the initial state) (which is a parallel block in this example).

[0289] Alternatively, it is possible not to support the loop filtering applicability flag at the segmentation unit boundary, but implicitly determine the applicability of filtering and filtering settings in the manner of the examples described later.

[0290] In addition, even when there is video continuity between individual segmentation units, after generating additional regions for the individual segmentation units, the video continuity at the boundaries between the segmentation units for which the additional regions have been generated may be lost. If loop filtering is applied in such a case, it will result in an unnecessary increase in the amount of calculation and a degradation in coding performance. Therefore, it is possible to implicitly not apply loop filtering.

[0291] In addition, according to the surface configuration of the 360-degree video, segmentation units that are adjacent in the two-dimensional space may not have video continuity with each other. When loop filtering is performed on the boundaries between such segmentation units without video continuity, it may result in a degradation in image quality. Therefore, for the boundaries between segmentation units without video continuity, it is possible to implicitly not perform loop filtering.

[0292] In addition, when weighting values are assigned to two regions in the manner described in the description as in Figure 10 and a part of the current segmentation unit is replaced, since the boundaries of the individual segmentation units are internal boundaries in the additional regions, loop filtering can be applied. However, since it is possible to additionally reduce the coding error by, for example, weighting value summation of a part of the current region included in other regions, it may not be necessary to perform loop filtering. Therefore, in the case as described above, it is possible to implicitly not perform loop filtering.

[0293] In addition, it is also possible to determine whether to apply loop filtering (specifically, additionally for the corresponding boundaries) based on a flag indicating whether loop filtering is applicable. When the above flag is activated, it is possible to apply filtering according to the loop filtering settings, conditions, etc. applicable within the segmentation unit, or to apply filtering with different definitions of loop filtering settings, conditions, etc. (specifically, additionally using different loop filtering settings, conditions, etc. from the case when it is not the boundary of the segmentation unit) at the boundary of the segmentation unit.

[0294] In the above embodiment, it is assumed that the situation is stored in the memory after removing the additional regions, but a part of it can also be in other output steps (specifically, it can belong to both the loop filtering unit and, for example, the post-filtering unit <postfilter>Processes executed on etc.

[0295] The above example is described assuming that the additional area is supported in all directions of each segmentation unit. When it is supported only in some directions according to the setting of the additional area, only a part of the above content can be applied. For example, the original setting can be applied at the boundary where the additional area is not supported, and various situations in the above example can be changed and applied at the boundary where the additional area is supported. That is, the above application can be adaptively determined in all or part of the unit boundaries according to the setting of the additional area.

[0296] The above relevant setting information can be processed implicitly or explicitly according to the encoding / decoding setting. When using the implicit method (specifically based on the characteristics, type, format, etc. of the image or according to other encoding / decoding settings <in this example, the setting related to the additional area>), it can be determined without generating relevant syntax elements. When using the explicit method, it can be adjusted by generating relevant syntax elements. The units related to this can include video, sequence, image, sub-image, strip, parallel block, etc.

[0297] Next, a detailed description will be given of the method for determining whether the segmentation unit and the additional area are referable. At this time, when it is referable, it belongs to dependent encoding / decoding, and when it is not referable, it belongs to independent encoding / decoding.

[0298] The additional area applying one embodiment of the present invention can be referred to or restrictedly referred to during the encoding / decoding process of the current image or other images. Specifically, the additional area removed before being stored in the memory can be referred to or restrictedly referred to during the encoding / decoding process of the current image. In addition, the additional area after being stored in the memory can be referred to or restrictedly referred to during the encoding / decoding process of an image that is temporally different from the current image in addition to the current image.

[0299] In other words, the referability and range, etc. of the above additional area can be determined according to the encoding / decoding setting. Through the above part of the setting, the additional area of the current image will be stored in the memory after being encoded / decoded, which means that it can be referred to or restrictedly referred to by being included in the reference image of other images. This can be applied to all or part of the segmentation units included in the corresponding image. The situations described in the examples in the subsequent description can also be changed and applied to the current example.

[0300] The above setting information related to the reference possibility of the additional area can be processed implicitly or explicitly according to the encoding / decoding settings. When using the implicit method (specifically, based on the characteristics, type, format, etc. of the video or according to other encoding / decoding settings <in this example, the settings related to the additional area>), it can be determined without generating the relevant syntax elements. When using the explicit method, the settings related to the reference possibility of the additional area can be adjusted by generating the relevant syntax elements. The units related to this can include video, sequence, image, sub-image, slice, parallel block, etc.

[0301] Generally, a part of the units in the current video (assumed to be the segmentation units obtained by the image segmentation unit in this example) can refer to the data of the current unit, but cannot refer to the data of other units. In addition, a part of the units in the current video can refer to the data of all units existing in other videos. The above description is an example related to the general nature of the units obtained by the image segmentation unit, and the additional properties related to it can also be defined.

[0302] In addition, a flag can be defined to indicate whether it is possible to refer to other segmentation units within the current video and whether it is possible to refer to the segmentation units included in other videos.

[0303] As an example, it can be allowed to refer to the segmentation units at the same position as the current segmentation unit included in other videos, but the reference to the segmentation units at different positions from the current segmentation unit is restricted. For example, when transmitting multiple bitstreams obtained by encoding the same video in environments with different encoding settings and selectively determining the bitstream used for decoding each region (segmentation unit) in the video (assumed to be decoded in parallel block units) in the decoder, since it is necessary to limit the reference possibility between each segmentation unit in the same space and different spaces, the encoding / decoding can be performed in a way that only allows reference to the same region in different videos.

[0304] As an example, the reference can be allowed or restricted according to the identifier information related to the segmentation unit. For example, when the identifier information assigned to the segmentation units is the same, the reference is allowed, and when it is different, the reference is not possible. At this time, the identifier information can be the information indicating that the encoding / decoding has been performed in an environment where mutual reference is possible (dependently).

[0305] The above relevant setting information can be processed implicitly or explicitly according to the encoding / decoding settings. When using the implicit method, it can be determined without generating relevant syntax elements, and when using the explicit method, it can be processed by generating relevant syntax elements. The units related to it can include video, sequence, image, sub-image, strip, parallel block, etc.

[0306] Figures 13a to 13g It is an illustrative diagram for explaining the area that can be referred to for a specific segmentation unit. In Figures 13a to 13g The area outlined by the thick frame line can represent the area that can be referred to.

[0307] Refer to Figure 13a , and various reference arrows used for performing inter-picture prediction can be confirmed. At this time, the C0 and C1 blocks represent unidirectional inter-picture prediction. The C0 block can obtain the RP0 reference block before the current image and can obtain the RF0 reference block after the current image. The C2 block represents bidirectional inter-picture prediction and can obtain the RP1 and RF1 reference blocks from the video before or after the current image. In the accompanying drawings, examples of obtaining one reference block from each of the previous direction and the subsequent direction are illustrated, but it is also possible to obtain reference blocks only from the previous direction or the subsequent direction. The C3 block represents non-directional inter-picture prediction and can obtain the RC0 reference block from within the current image. In the accompanying drawings, an example of obtaining one reference block is illustrated, but it is also possible to obtain two or more reference blocks.

[0308] In the examples described later, the description will be centered on the reference possibility of pixel values and prediction mode information in inter-picture prediction based on the segmentation unit, but it can also be understood to include other encoding / decoding information that can be referred to spatially or temporally (such as intra-picture prediction mode information, transform and quantization information, loop filter information, etc.).

[0309] Refer to Figure 13b , the current image Currnt(t) is divided into two or more parallel blocks, and the block C0 in a part of the parallel blocks can obtain the reference blocks P0 and P1 by performing unidirectional inter-picture prediction. The block C1 in a part of the parallel blocks can obtain the reference blocks P3 and F0 by performing bidirectional inter-picture prediction. That is, this can be understood as an example where reference to blocks at other positions included in other videos is allowed without restrictions such as position restrictions and only allowing reference within the same image.

[0310] Refer to Figure 13c , the image is divided into two or more parallel block units, and a part of the blocks C1 in a part of the parallel blocks can obtain reference blocks P2 and P3 by performing unidirectional inter-picture prediction. A part of the blocks C0 in a part of the parallel blocks can obtain reference blocks P0, P1, F0, and F1 by performing bidirectional inter-picture prediction. A part of the blocks C3 in a part of the parallel blocks can obtain reference block FC0 by performing non-directional inter-picture prediction.

[0311] That is, Figure 13b and Figure 13c It can be understood as an example that allows reference to blocks at other positions included in other images without restrictions such as position restrictions and only allowing reference within the same image.

[0312] Refer to Figure 13d , the current image is divided into two or more parallel block units, and the block C0 in a part of the parallel blocks can obtain the reference block P0 by performing forward inter-picture prediction, but cannot obtain the reference blocks P1, P2, and P3 included in a part of the parallel blocks. The block C4 in a part of the parallel blocks can obtain the reference blocks F0 and F1 by performing backward inter-picture prediction, but cannot obtain the reference blocks F2 and F3. A part of the blocks C3 in a part of the parallel blocks can obtain the reference block FC0 by performing non-directional inter-picture prediction, but cannot obtain the reference block FC1.

[0313] That is, in Figure 13d , it is possible to allow or restrict reference according to the division of the image (in this example, t-1, t, t+1) and the encoding / decoding settings of the image division unit. Specifically, it is possible to only allow reference to the blocks included in the parallel blocks having the same identifier information as the current parallel block.

[0314] Refer to Figure 13e , the image is divided into two or more parallel block units, and a part of the blocks C0 in a part of the parallel blocks can obtain the reference blocks P0 and F0 by performing bidirectional inter-picture prediction, but cannot obtain the reference blocks P1, P2, P3, F1, F2, and F3. That is, Figure 13e It can be an example that only allows reference to the parallel blocks at the same position as the parallel block containing the current block.

[0315] Refer to Figure 13f , the image is divided into two or more parallel block units, and a part of the blocks C0 in a part of the parallel blocks can obtain the reference blocks P1 and F2 by performing bidirectional inter-picture prediction, but cannot obtain the reference blocks P0, P2, P3, F0, F1, and F3. Figure 13f An example is that information indicating parallel blocks that can be referred to for the current segmentation unit is included in the bitstream, and the parallel blocks that can be referred to are confirmed based on the above information.

[0316] Refer to Figure 13g , the image is segmented into two or more parallel blocks, and a part of the blocks C0 in a part of the parallel blocks can obtain reference blocks P0, P3, P5 by performing unidirectional inter-picture prediction, but cannot obtain the reference block P4. A part of the blocks C1 in a part of the parallel blocks can obtain reference blocks P1, F0, F2 by performing bidirectional inter-picture prediction, but cannot obtain the reference blocks P2, F1.

[0317] Figure 13g An example is that it is possible to allow or restrict reference according to whether the image (t-3, t-2, t-1, t, t+1, t+2, t+3 in this example) is segmented, the encoding / decoding setting of the segmentation unit of the image (it is assumed in this example that it is determined according to the identifier information of the segmentation unit, the identifier information of the image unit, whether the same area of the segmentation unit, whether the similar area of the segmentation unit, the bitstream information of the segmentation unit, etc.). Among them, the position of the parallel blocks that can be referred to in the video can be the same as or similar to the current block and can have the same identifier information as the current block (specifically, on the image unit or segmentation unit), and can be the same as the bitstream for obtaining the current parallel block.

[0318] Figures 14a to 14e It is a flowchart for explaining the reference possibility of the additional area in the segmentation unit to which one embodiment of the present invention is applied. In Figures 14a to 14e , the area shown by the thick frame line represents the area that can be referred to, and the area shown by the dotted line represents the additional area of the segmentation unit.

[0319] In one embodiment of the present invention, it is possible to restrict or allow the reference possibility of a part of the image (other images located before or after in time). In addition, it is possible to restrict or allow the reference possibility of the entire extended segmentation unit including the additional area. In addition, it is possible to restrict or allow only the reference possibility of the initial segmentation unit excluding the additional area. In addition, it is possible to restrict or allow the reference possibility of the boundary between the additional area and the initial segmentation unit.

[0320] Refer to Figure 14a , a part of blocks C0 in a part of parallel blocks can obtain reference blocks P0, P1 by performing unidirectional inter-picture prediction. A part of blocks C2 in a part of parallel blocks can obtain reference blocks P2, P3, F0, F1 by performing bidirectional inter-picture prediction. A part of blocks C1 in a part of parallel blocks can obtain reference block FC0 by performing non-directional inter-picture prediction. Among them, block C0 can obtain reference blocks P0, P1, P2, P3, F0 from the initial parallel block area (the original parallel block except the additional area) of a part of reference images t-1, t+1, while block C2 can obtain reference block F1 from the parallel block area including the additional area of reference image t+1 while obtaining reference blocks P2, P3 from the initial parallel block area of reference image t-1. At this time, as shown by reference block F1, a reference block including the boundary between the additional area and the initial parallel block area can be obtained.

[0321] See Figure 14b , a part of blocks C0, C1, C3 in a part of parallel blocks can obtain reference blocks P0, P1, P2 / F0, F2 / F1, F3, F4 by performing unidirectional inter-picture prediction. A part of blocks C2 in a part of parallel blocks can obtain reference blocks FC0, FC1, FC2 by performing non-directional inter-picture prediction.

[0322] A part of blocks C0, C1, C3 can obtain reference blocks P0, F0, F3 from the initial parallel block area of a part of reference images (in this example, t-1, t+1), can also obtain reference blocks P1, x, F4 from the boundary of the updated parallel block area, and can also obtain reference blocks P2, F2, F1 from the outside of the boundary of the updated parallel block area.

[0323] A part of blocks C2 can obtain reference block FC1 from the initial parallel block area of a part of reference images (in this example, t), can also obtain reference block FC3 from the boundary of the updated parallel block area, and can also obtain reference block FC0 from the outside of the boundary of the updated parallel block area.

[0324] Among them, a part of blocks C0 can be blocks located in the initial parallel block area, a part of blocks C1 can be blocks located on the boundary of the updated parallel block area, and a part of blocks C3 can be blocks located outside the boundary of the updated parallel block.

[0325] See Figure 14c , the image is divided into two or more parallel block units. An additional area is set for a part of the parallel blocks in a part of the image, no additional area is set for a part of the parallel blocks in a part of the image, and no additional area is set in a part of the image. Some of the blocks C0, C1 in a part of the parallel blocks can obtain reference blocks P2, F1, F2, F3 by performing unidirectional inter-picture prediction, but cannot obtain reference blocks P0, P1, P3, F0. Some of the blocks C2 in a part of the parallel blocks can obtain reference blocks FC1, FC2 by performing bidirectional inter-picture prediction, but cannot obtain reference block FC0.

[0326] Some of the blocks C2 cannot obtain the reference block FC0 from the initial parallel block area of a part of the reference image (in this example, t), but can obtain the reference block FC1 from the updated parallel block area (in the method of filling a part of the additional area, FC0 and FC1 can be the same area. Although FC0 cannot be referenced in the parallel block division of the initial unit, it can be referenced when the corresponding area is moved to the current parallel block through the additional area).

[0327] Some of the blocks C2 can obtain the reference block FC2 from a part of the parallel block area of a part of the reference image (in this example, t) (although the data in other parallel blocks of the current image cannot be referenced by default, it is allowed to reference when it is set to be referenceable through the identifier information in the above embodiments, etc.).

[0328] Refer to Figure 14d , the image is divided into two or more parallel block units and an additional area is set. Some of the blocks C0 in a part of the parallel blocks can obtain reference blocks P0, F0, F1, F3 by performing bidirectional inter-picture prediction, but cannot obtain reference blocks P1, P2, P3, F2.

[0329] Some of the blocks C0 can obtain the reference block P0 from the initial parallel block area (parallel block No. 0) of a part of the reference image t-1, but cannot obtain the reference block P3 from the boundary of the extended parallel block area, nor can it obtain the reference block P2 from the outside of the boundary of the extended parallel block area (i.e., the additional area).

[0330] Some of the blocks C0 can obtain the reference block F0 from the initial parallel block area (parallel block No. 0) of a part of the reference image t+1, can also obtain the reference block F1 from the boundary of the extended parallel block area, and can also obtain the reference block F3 from the outside of the boundary of the extended parallel block area.

[0331] Refer to Figure 14e , the image is divided into two or more parallel block units and an additional area having at least one size and shape is set. In some of the parallel blocks, block C0 can obtain reference blocks P0, P3, P5, F0 by performing uni-directional inter-picture prediction, but cannot obtain reference block P2 located on the boundary between the additional area and the original parallel block. In some of the parallel blocks, block C1 can obtain reference blocks P1, F2, F3 by performing bi-directional inter-picture prediction, but cannot obtain reference blocks P4, F1, F5.

[0332] As shown in the above example, pixel values can be objects for reference, and the reference to other coding / decoding information can be restricted.

[0333] As an example, when the prediction unit searches for a group of in-picture prediction mode candidates to be used in in-picture prediction from spatially adjacent blocks, it can confirm whether the segmentation unit containing the current block can refer to the segmentation unit containing the adjacent block by the method as Figures 13a to 14e shown.

[0334] As an example, when the prediction unit searches for a group of motion information candidates to be used in inter-picture prediction from temporally and spatially adjacent blocks, it can confirm whether the segmentation unit containing the current block can refer to the segmentation unit containing a block that is spatially adjacent within the current image or temporally adjacent to the current image by the method as Figures 13a to 14e shown.

[0335] As an example, when the loop filter unit searches for loop filter related setting information from adjacent blocks, it can confirm whether the segmentation unit containing the current block can refer to the segmentation unit containing the adjacent block by the method as Figures 13a to 14e shown.

[0336] Figure 15 is an illustrative diagram showing the blocks of the segmentation unit included in the current video and the blocks of the segmentation unit included in other videos.

[0337] Refer to Figure 15 , in this example, the spatially adjacent reference candidate blocks can be the left, upper left, lower left, upper, and upper right blocks centered on the current block. In addition, the temporal reference candidate blocks can be the left, upper left, lower left, upper, upper right, right, lower right, lower, and central blocks of the block (Collocated block) located at the same or corresponding position as the current block in the image (Different picture) that is temporally adjacent to the current image (Currentpicture). In Figure 15 , the thick outer frame line represents the boundary line of the segmentation unit.

[0338] When the current block is M, the spatially adjacent blocks G, H, I, L, and Q can all be referenced.

[0339] When the current block is G, some of the spatially adjacent blocks A, B, C, F, and K can be referenced while the remaining blocks can restrict the reference. Whether it can be referenced or not can be determined according to the reference-related settings between the segmentation units UC, ULC, LC included in the spatially adjacent blocks and the segmentation unit including the current block.

[0340] When the current block is S, some of the surrounding blocks s, r, m, w, n, x, t, o, and y at the same position as the current block in the temporally adjacent images can be referenced while the remaining blocks can restrict the reference. Whether it can be referenced or not can be determined according to the reference-related settings between the segmentation units RD, DRD, DD included in the surrounding blocks at the same position as the current block in the temporally adjacent images and the unit including the current block.

[0341] According to the position of the current block, when there is a candidate with restricted reference, it is possible to fill it with the candidate with the next order in the priority order in the candidate group composition, or replace it with other candidates adjacent to the candidate with restricted reference.

[0342] For example, when the current block in the in-picture prediction is G, the reference of the upper-left block is restricted, and the most probable mode (MPM) candidate group composition adopts the order of P-D-A-E-U. Since A cannot be referenced, the candidate group can be formed by performing a validity check in the remaining order of E-U, or B or F adjacent to A in space can be used to replace A.

[0343] In addition, when the current block in the inter-picture prediction is S, the reference of the adjacent lower block in time is restricted, and the temporal candidate of the skip mode candidate group is y. Since y cannot be referenced, the candidate group can be formed by performing a validity check in the order of a mixture of spatially adjacent candidates or idle candidates and temporal candidates, or t, x, s adjacent to y in space can be used to replace y.

[0344] Figure 16 It is a hardware configuration diagram illustrating an image encoding / decoding device to which one embodiment of the present invention is applied.

[0345] See Figure 16 , the video encoding / decoding device 200 applicable to an embodiment of the present invention may include: at least one processor 210; and a memory 220 storing instructions for instructing the at least one processor 210 to execute at least one step.

[0346] Among them, the at least one processor 210 may refer to a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor for executing the method of the embodiment applicable to the present invention. The memory 120 and the storage device 260 may be respectively composed of at least one of a volatile storage medium and a non-volatile storage medium. For example, the memory 220 may be composed of at least one of a read only memory (ROM) and a random access memory (RAM).

[0347] In addition, the video encoding / decoding device 200 may further include: a transceiver 230 for performing communication through a wireless communication network. In addition, the video encoding / decoding device 200 may further include: an input interface device 240, an output interface device 250, a storage device 260, etc. Each component included in the video encoding / decoding device 200 may be connected through a bus 270 to communicate with each other.

[0348] Among them, the at least one step may include: a step of dividing the encoded video included in the bitstream into at least one division unit by referring to syntax elements obtained from the received bitstream; a step of setting an additional area for the at least one division unit; and a step of decoding the encoded video based on the division unit after setting the additional area.

[0349] Among them, the step of decoding the encoded video may include: a step of determining a reference block related to a current block to be decoded in the encoded video according to information indicating whether reference is possible included in the bitstream.

[0350] Among them, the reference block may be a block included at a position overlapping with the additional area set on the division unit including the reference block.

[0351] Figure 17 It is an exemplary diagram illustrating an intra prediction mode of an embodiment applicable to the present invention.

[0352] Referring to Figure 17 , it can be confirmed that there are a total of 35 prediction patterns, and the 35 prediction patterns can be divided into 33 directional patterns and 2 non-directional patterns (mean (DC), planar). At this time, the directional patterns can be identified by the inclination (e.g., dy / dx) or angle information. The above examples can refer to a candidate group of prediction patterns related to the luminance component or the color difference component. Alternatively, the color difference component can support some prediction patterns (e.g., mean (DC), planar, vertical, horizontal, diagonal patterns, etc.). In addition, after the prediction pattern of the luminance mode is determined, the corresponding pattern can be included in the prediction pattern of the color difference component or the pattern derived from the corresponding pattern can be included in the prediction pattern.

[0353] In addition, the correlation between color spaces can be utilized to apply the reconstructed blocks in other color spaces that have been encoded / decoded to the prediction of the current block, and the supported prediction patterns can be included. For example, the color difference component can generate a prediction block for the current block by using the reconstructed block of the luminance component corresponding to the current block.

[0354] According to the encoding / decoding settings, the candidate group of prediction patterns can be adaptively determined. The number of candidate groups can be increased for the purpose of improving the prediction accuracy, or the number of candidate groups can be reduced for the purpose of reducing the number of bits in the prediction pattern.

[0355] For example, one of the candidate groups such as candidate group A (67, 65 directional patterns and 2 non-directional patterns), candidate group B (35, 33 directional patterns and 2 non-directional patterns), candidate group C (19, 17 directional patterns and 2 non-directional patterns), etc. can be used. In the present invention, unless otherwise clearly stated, it is assumed that the in-picture prediction is performed using a pre-set candidate group of prediction patterns (candidate group A).

[0356] Figure 18 FIG. 1 is a first illustrative diagram showing the composition of reference pixels used in the in-picture prediction to which one embodiment of the present invention is applied.

[0357] The intra prediction method in video decoding applicable to an embodiment of the present invention may include: a reference pixel formation step, a prediction block generation step of referring to the formed reference pixels and using one or more prediction modes, a step of determining the best prediction mode, and a step of encoding the determined prediction mode. In addition, the video decoding apparatus may include a reference pixel formation unit, a prediction block generation unit, a prediction mode determination unit, and a prediction mode encoding unit for performing the reference pixel formation step, the prediction block generation step, the prediction mode determination step, and the prediction mode encoding step. The above-described process may omit a part thereof or add other processes, and may also be changed to another order different from the order described above.

[0358] In addition, the intra prediction method in video decoding applicable to an embodiment of the present invention may generate a prediction block of a current block according to a prediction mode obtained from syntax elements received from a video encoding apparatus after forming reference pixels.

[0359] The size and shape (M×N) of the current block for performing intra prediction may be obtained from a block division unit, and sizes from 4×4 to 256×256 may be adopted. Intra prediction is usually performed in units of prediction blocks, but may also be performed in units of coding blocks (or coding units), transform blocks (or transform units), etc. according to the setting of the block division unit. After confirming the block information, a reference pixel formation unit may form reference pixels used in the prediction of the current block. At this time, the reference pixels may be stored in a temporary memory (e.g., an array <array>, managed for 1D, 2D arrays, etc., generated and removed during the prediction process within each frame of the block, and the size of the temporary memory can be determined according to the composition of the reference pixels.

[0360] The reference pixels can be pixels included in adjacent blocks (which can be called reference blocks) located on the left, upper, upper-left, upper-right, and lower-left sides centered on the current block, but are not limited thereto. In the prediction of the current block, other sets of block candidates with different compositions can also be used. Among them, the adjacent blocks on the left, upper, upper-left, upper-right, and lower-left sides can be the blocks selected when performing encoding / decoding using raster or zigzag scanning. When the scanning order is changed, adjacent blocks at other positions (such as the right, lower, and lower-right blocks, etc.) can also be used as reference pixels.

[0361] In addition, the reference block can be a block corresponding to the current block in a color space different from the color space containing the current block. Among them, when in the Y / Cb / Cr format, the color space can refer to one of Y, Cb, and Cr. In addition, the block corresponding to the current block can refer to a block having the same position coordinates as the current block or having position coordinates corresponding to the current block according to the color component composition ratio.

[0362] In addition, for the convenience of explanation, it is assumed that the reference block at the above-mentioned preset positions (left, upper, upper-left, upper-right, lower-left) is composed of one block, but according to block division, it can also be composed of multiple sub-blocks.

[0363] In other words, the adjacent area of the current block can be the reference pixel position for the intra-frame prediction of the current block, and the area corresponding to the current block in other color spaces can be additionally used as the reference pixel position according to the prediction mode. In addition to the above examples, the defined reference pixel positions can also be determined according to the prediction mode, method, etc. For example, when generating a prediction block by methods such as block matching, the reference pixel position can be the area that has been encoded / decoded before the current block in the current image or the area included within the exploration range of the area that has been encoded / decoded (such as including the left or right or upper-left or upper-right of the current block, etc.).

[0364] See Figure 18 , in the intra-frame prediction of the current block (with a size of M×N), the reference pixels used can be composed of pixels adjacent to the current block on the left, upper, upper-left, upper-right, and lower-left sides ( Figure 18 Ref_L, Ref_T, Ref_TL, Ref_TR, Ref_BL in Figure 18 The content marked in the form of P(x,y) can refer to pixel coordinates.

[0365] In addition, pixels adjacent to the current block can be classified into at least one reference pixel level. For example, they can be classified into pixels ref_0 that are closest to the current block {pixels with a pixel value difference of 1 from the boundary pixels of the current block, p(-1, -1) to p(2M - 1, -1), p(-1, 0) to p(-1, 2N - 1)}. Then, the adjacent pixels {with a pixel value difference of 2 from the boundary pixels of the current block, p(-2, -2) to p(2M, -2), p(-2, -1) to p(-2, 2N)} are ref_1, and then the further adjacent pixels {with a pixel value difference of 3 from the boundary pixels of the current block, p(-3, -3) to p(2M + 1, -3), p(-3, -2) to p(-3, 2N + 1)} are ref_2, and so on. That is, the reference pixels can be classified into multiple reference pixel levels according to the pixel distance from the boundary pixels of the current block.

[0366] In addition, different reference pixel levels can be set for each adjacent block at this time. For example, when using the block adjacent to the upper end of the current block as a reference block, the reference pixels at the ref_0 level can be used, and when using the block adjacent to the upper right end as a reference block, the reference pixels at the ref_1 level can be used.

[0367] Among them, the set of reference pixels usually referred to when performing intra-picture prediction is included in the blocks adjacent to the lower left, left, upper left, upper end, and upper right end of the current block, and is the pixels belonging to the ref_0 level (the pixels closest to the boundary pixels). Unless otherwise specified in the following content, the above-mentioned pixels are used as the premise. However, it is also possible to use the set of reference pixels included in some of the adjacent blocks mentioned above, and it is also possible to use the pixels included in two or more levels as the set of reference pixels. Among them, the set of reference pixels or levels can be implicitly determined (pre-set in the encoding / decoding device), or can be explicitly determined (receiving information for determination from the encoding device).

[0368] Here, the case of supporting at most 3 reference pixel levels will be used as a premise for explanation, but larger values can also be used. The number of reference pixel levels and the number of sets of reference pixels (or what can also be called the reference pixel candidate group) based on the positions of the adjacent blocks that can be referred to can be set differently according to the size, shape, prediction mode, video type <I / P / B, the video at this time is an image, slice, parallel block, etc.>, color component, etc., and the relevant information can be included in units such as sequence, image, slice, parallel block, etc.

[0369] In the present invention, it is described on the premise that lower index values (starting from 0 and incrementing by 1 gradually) are assigned starting from the reference pixel level closest to the current block, but it is not limited thereto. In addition, the reference pixel composition-related information described later can be generated under the above-mentioned index setting (such as binary conversion that assigns shorter bits to smaller indexes when selecting one from multiple reference pixel sets).

[0370] In addition, when there are two or more supported reference pixel levels, weighted value averaging or the like can be applied to each reference pixel included in the two or more reference pixel levels.

[0371] For example, it is possible to generate a predicted block using a reference pixel obtained by combining the weighted values of the pixels in the ref_0-th level and the ref_1-th level located at Figure 18 . At this time, depending on the prediction mode (such as the directionality of the prediction mode), the pixels to which the weighted value combination is applied in each reference pixel level can be either integer unit pixels or fractional unit pixels. In addition, one predicted block can be obtained by assigning weighted values (such as 7:1, 3:1, 2:1, 1:1, etc.) to the predicted block obtained using the reference pixels in the first reference pixel level and the predicted block obtained using the reference pixels in the second reference pixel level. At this time, a higher weighted value can be assigned to the predicted block of the reference pixel level closer to the current block.

[0372] In the case of assuming the generation of explicit information related to the reference pixel composition, it is possible to generate indication information (in this example, adaptive_intra_ref_sample_enabled_flag) that allows an adaptive reference pixel composition in units such as video, sequence, image, stripe, parallel block, etc.

[0373] When the above indication information represents an allowable adaptive reference pixel composition (in this example, adaptive_intra_ref_sample_enabled_flag = 1), it is possible to generate adaptive reference pixel composition information (in this example, adaptive_intra_ref_sample_flag) in units such as image, stripe, parallel block, block, etc.

[0374] When the above composition information represents an adaptive reference pixel composition (in this example, adaptive_intra_ref_sample_flag = 1), it is possible to generate reference pixel composition-related information (such as selection information related to the reference pixel level and set, etc., in this example, intra_ref_idx) in units such as image, stripe, parallel block, block, etc.

[0375] At this time, when an adaptive reference pixel composition is not allowed or the reference pixel composition is not adaptive, the reference pixels can be formed according to a preset setting. For example, usually, the pixels closest to the current block in adjacent blocks are used to form the reference pixels, but it is not limited to this, and various situations are also allowed (for example, the case where ref_0 and ref_1 are selected as the reference pixel levels and the predicted pixel value is generated by weighted combination of ref_0 and ref_1, that is, the default situation).

[0376] In addition, the information related to the reference pixel composition (such as the selection information related to the reference pixel level or set, etc.) can be formed (such as ref_1, ref_2, ref_3, etc.) after excluding the preset information (for example, the case where the reference pixel level is preset as ref_0), but it is not limited to this either.

[0377] Through the above examples, some situations related to the reference pixel composition have been described, but the in-picture prediction setting can be determined by combining with various coding / decoding information. At this time, the coding / decoding information can include, for example, video type, color component, size and shape of the current block, prediction mode {type of prediction mode (directional, non-directional), direction of prediction mode (vertical, horizontal, diagonal 1, diagonal 2, etc.)}, etc., and the in-picture prediction setting (in this example, the reference pixel composition setting) can be determined according to the coding / decoding information of adjacent blocks and the combination of the coding / decoding information of the current block and adjacent blocks.

[0378] Figures 19a to 19c It is the second illustrative diagram showing the reference pixel composition to which one embodiment of the present invention is applied.

[0379] Refer to Figure 19a and the case where the reference pixels are formed only using the ref_0-th reference pixel level in Figure 18 can be confirmed. After using the pixels included in adjacent blocks (such as lower left, left, upper left, upper, upper right) to form the reference pixels with the ref_0-th reference pixel level as the object, subsequent in-picture prediction can be performed (such as reference pixel generation, reference pixel filtering, reference pixel interpolation, predicted block generation, post-processing filtering, etc., and some in-picture prediction processes can be adaptively performed according to the reference pixel composition). In this example, the case of using a preset single reference pixel level, that is, the case where the setting information related to the reference pixel level is not generated and the in-picture prediction is performed using the non-directional mode, will be described.

[0380] Refer to Figure 19b , it is possible to confirm the case where a reference pixel is formed by simultaneously using two supported reference pixel levels. That is, it is possible to perform intra-frame prediction after forming a reference pixel by using the pixels included in level ref_0 and level ref_1 (or the weighted average value of the pixels included in the two levels). In this example, an example of performing intra-frame prediction using a plurality of preset reference pixel levels, that is, without generating setting information related to the reference pixel levels and using a part of the directional prediction modes (from the upper right side to the lower left side in the attached drawing or the opposite direction) will be described.

[0381] Refer to Figure 19c , it is possible to confirm the case where a reference pixel is formed by using only one of the three supported reference pixel levels. In this example, an example of performing intra-frame prediction by generating setting information related to the used reference pixel level due to the existence of a plurality of reference pixel level candidates and using a part of the directional prediction modes (from the upper left side to the lower right side in the attached drawing) will be described.

[0382] Figure 20 FIG. 3 is a third illustrative diagram showing the reference pixel formation to which one embodiment of the present invention is applied.

[0383] Figure 20 In the attached drawing, reference numeral a is a block with a size of 64×64 or more, reference numeral b is a block with a size of 16×16 or more and less than 64×64, and reference numeral c is a block with a size less than 16×16.

[0384] When the block of reference numeral a is used as the current block for which intra-frame prediction needs to be performed, it is possible to perform intra-frame prediction by using one closest reference pixel level ref_0.

[0385] In addition, when the block of reference numeral b is used as the current block for which intra-frame prediction needs to be performed, it is possible to perform intra-frame prediction by using two supported reference pixel levels ref_0 and ref_1.

[0386] In addition, when the block of reference numeral c is used as the current block for which intra-frame prediction needs to be performed, it is possible to perform intra-frame prediction by using three supported reference pixel levels ref_0, ref_1, and ref_2.

[0387] As described in the explanations of reference numerals a to c in the attached drawing, it is possible to set the number of supported reference pixel levels differently according to the size of the current block for which intra-frame prediction needs to be performed. In Figure 20 Among them, the larger the size of the current block, the higher the possibility that the size of the adjacent block is smaller. This may be the result of performing segmentation according to other image characteristics. Therefore, in order to prevent prediction from being performed using pixels with a large pixel value distance from the current block, it is assumed that when the size of the block is larger, the number of reference pixel levels supported is smaller. However, other variations including the opposite case are also allowed.

[0388] Figure 21 It is the fourth exemplary diagram illustrating the reference pixel composition to which one embodiment of the present invention is applied.

[0389] Refer to Figure 21 , it is possible to confirm the case where the current block for performing in-picture prediction is in a rectangular shape. If the current block is a horizontally and vertically asymmetric rectangular shape, the number of supported reference pixel levels adjacent to the longer horizontal side interface in the current block can be set to be larger, and the number of supported reference pixel levels adjacent to the shorter vertical side interface in the current block can be set to be smaller. In the drawings, it is possible to confirm the case where two reference pixel levels adjacent to the horizontal side interface of the current block are set and one reference pixel level adjacent to the vertical side interface of the current block is set. Pixels adjacent to the shorter vertical side interface in the current block may have a problem of decreased accuracy because the distance from the pixels included in the current block is generally relatively far (because of the larger horizontal length). Therefore, the number of supported reference pixel levels adjacent to the shorter vertical side interface is set to be smaller, but the opposite setting can also be adopted.

[0390] In addition, the reference pixel levels to be used in prediction can be set differently according to the type of in-picture prediction mode or the position of the adjacent block adjacent to the current block. For example, the directional mode that uses the pixels included in the block adjacent to the upper end and the upper right end of the current block as reference pixels can use two or more reference pixel levels, while the directional mode that uses the pixels included in the block adjacent to the left end and the lower left end of the current block as reference pixels can use only the closest one reference pixel level.

[0391] In addition, when the prediction blocks generated through each reference pixel level among multiple reference pixel levels are the same or similar to each other, the setting information of the reference pixel level may cause a problem of additional generation of unnecessary data.

[0392] For example, when the distribution characteristics of the pixels constituting each reference pixel level are similar or the same, similar or identical prediction blocks may be generated regardless of which reference pixel level is used. Therefore, there is no need to generate data for selecting a reference pixel level. At this time, the distribution characteristics of the pixels constituting the reference pixel level can be determined by comparing the average value or dispersion value of the pixels with a preset threshold value.

[0393] That is, when the reference pixel levels based on the finally determined intra-picture prediction mode are the same or similar to each other, a reference pixel level can be selected using a preset method (for example, selecting the nearest reference pixel level).

[0394] At this time, the decoder can receive intra-picture prediction information (or intra-picture prediction mode information) from the encoding device and determine whether to receive information for selecting a reference pixel level based on the received information.

[0395] The cases of forming reference pixels using multiple reference pixel levels have been described through the above various examples. However, it is not limited thereto, and various modified examples can also be adopted, and it can also be used in combination with other additional configurations.

[0396] The reference pixel forming unit for intra-picture prediction can include, for example, a reference pixel generation unit, a reference pixel interpolation unit, a reference pixel filtering unit, etc., and can include all or a part of the above configurations. Among them, a block containing pixels that can be used as reference pixels can be called a reference candidate block. In addition, the reference candidate block can usually be an adjacent block adjacent to the current block.

[0397] The reference pixel forming unit can determine whether to use the pixels included in the reference candidate block as reference pixels according to the availability of the reference candidate block set for the reference pixels.

[0398] Regarding the above-mentioned availability of reference pixels, it can be determined as unusable when at least one of the following conditions is met. For example, when the reference candidate block meets at least one of the conditions such as being located outside the image boundary, not being included in the same segmentation unit as the current block (for example, a strip, a parallel block, etc.), not having completed encoding / decoding, and being restricted in its use in the encoding / decoding settings, it can be determined that the pixels included in the corresponding reference candidate block cannot be used as references. At this time, if none of the above conditions are met, it can be determined as usable.

[0399] In addition, the use of reference pixels can be restricted according to encoding / decoding settings. For example, when a flag (such as constrained_intra_pred_flag) used to restrict the reference to a reference candidate block is activated, it can be restricted so that the pixels included in the corresponding reference candidate block cannot be used as reference pixels. In order to effectively perform encoding / decoding even when errors occur due to various external factors including the communication environment, the above flag can be applied when the reference candidate block is a block reconstructed by referring to an image that is temporally different from the current block.

[0400] Among them, when the flag used to restrict the reference is activated (for example, when constrained_intra_pred_flag = 0 in the I picture type or P or B picture type), all pixels in the reference candidate block can be used as reference pixels. In addition, when the flag used to restrict the reference is activated (for example, when constrained_intra_pred_flag = 1 in the P or B picture type), it can be determined whether it can be referenced according to whether the reference candidate block is encoded by intra prediction or inter prediction. That is, when the reference candidate block is encoded by intra prediction, the corresponding reference candidate block can be referenced regardless of whether the above flag is activated, and when the reference candidate block is encoded by inter prediction, it can be determined whether the corresponding reference candidate block can be referenced according to whether the above flag is activated.

[0401] In addition, a reconstructed block located at the position corresponding to the current block in another color space can be used as a reference candidate block. At this time, it can be determined whether it can be referenced according to the encoding mode of the reference candidate block. For example, when the current block belongs to a part of the color difference components (Cb, Cr), it can be determined whether it can be referenced according to the encoding mode of the block (=reference candidate block) located at the position corresponding to the current block in the luminance component (Y) and that has been encoded / decoded. This can be an example corresponding to the case of independently determining the encoding mode according to the color space.

[0402] The flag used to restrict the reference can be a setting applicable to a part of the video types (such as P or B slice / parallel block types, etc.).

[0403] By referring to the usability of reference pixels, reference candidate blocks can be classified into cases where they can be fully used, cases where they can be partially used, and cases where they cannot be used at all. In cases other than the case where they can be fully used, it is possible to fill or generate reference pixels at positions of candidate blocks that cannot be used.

[0404] In the case where a reference candidate block can be used, pixels at a preset position of the current block (or pixels adjacent to the current block) can be stored in the reference pixel memory of the current block. At this time, the pixel data at the corresponding block position can be directly copied or stored in the reference pixel memory through processes such as reference pixel filtering.

[0405] In the case where a reference candidate block cannot be used, pixels obtained through the reference pixel generation process can be included in the reference pixel memory of the current block.

[0406] In other words, reference pixels can be formed in a state where a reference pixel candidate block can be used, and reference pixels can be generated in a state where a reference pixel candidate block cannot be used.

[0407] The method of filling reference pixels at a preset position in a reference candidate block that cannot be used is as follows. First, reference pixels can be generated using any pixel value. Here, any pixel value is a specific pixel value included in the pixel value range, and can be the minimum value, maximum value, median value, or a value derived from the above values used in the pixel value adjustment process based on the bit depth or the pixel value range information of the image. Among them, the method of generating reference pixels using any pixel value can also be applied in the case where all reference candidate blocks cannot be used.

[0408] Next, reference pixels can be generated using the pixels included in the blocks adjacent to the reference candidate block that cannot be used. Specifically, the pixels included in the adjacent blocks can be filled into the preset positions in the reference candidate block that cannot be used by means of extrapolation, interpolation, or copying. At this time, the method of performing copying or extrapolation, etc. can be in the clockwise direction or the counterclockwise direction, and can be determined according to the encoding / decoding settings. For example, the direction of generating reference pixels within a block can follow a preset direction or a direction adaptively determined according to the position of the block that cannot be used.

[0409] Figures 22a to 22b It is an exemplary diagram illustrating the method of filling reference pixels at a preset position in a reference candidate block that cannot be used.

[0410] See Figure 22a , a method for filling pixels included in an unusable reference candidate block in a reference pixel composed of a reference pixel level can be confirmed. In Figure 22a , when an adjacent block adjacent to the upper right end of the current block is an unusable reference candidate block, the reference pixels (denoted as <1>) included in the adjacent block adjacent to the upper right end can be generated by performing clockwise extrapolation or linear extrapolation on the reference pixels included in the adjacent block adjacent to the upper end of the current block.

[0411] In addition, in Figure 22a , when an adjacent block adjacent to the left side of the current block is an unusable reference candidate block, the reference pixels (denoted as <2>) included in the adjacent block adjacent to the left side can be generated by performing counterclockwise extrapolation or linear extrapolation on the reference pixels included in the adjacent block adjacent to the upper left end of the current block (corresponding to a usable block). At this time, by performing clockwise extrapolation or linear extrapolation, the reference pixels included in the adjacent block adjacent to the lower left end of the current block can be utilized.

[0412] In addition, in Figure 22a , a part of the reference pixels (denoted as <3>) included in the adjacent block adjacent to the upper side of the current block can be generated by performing interpolation or linear interpolation on the usable reference pixels on both sides. That is, setting can also be performed when a part of the reference pixels included in the adjacent block is unusable rather than all being unusable. In this case, the adjacent pixels of the unusable reference pixels can be used to fill the unusable reference pixels.

[0413] Refer to Figure 22b , a method for filling unusable reference pixels when a part of the reference pixels in a reference pixel composed of multiple reference pixel levels are unusable can be confirmed. Refer to Figure 22b , when an adjacent block adjacent to the upper right end of the current block is an unusable reference candidate block, the pixels (denoted as <1>) included in the three reference pixel levels included in the corresponding adjacent block can be generated clockwise using the pixels included in the adjacent block adjacent to the upper end of the current block (corresponding to a usable block).

[0414] In addition, in Figure 22b , when an adjacent block adjacent to the left side of the current block is an unusable reference candidate block and the adjacent block adjacent to the upper left end or lower left end of the current block is a usable reference candidate block, the reference pixels of the unusable reference candidate block can be generated by filling the reference pixels of the usable reference candidate block clockwise, counterclockwise, or bidirectionally.

[0415] At this time, the unusable reference pixels in each reference pixel level can be generated using the pixels in the same reference pixel level, but the method of using the pixels in different reference pixel levels is not excluded. For example, in Figure 22b , taking as a premise that the reference pixels (denoted as <3>) in three reference pixel levels included in the adjacent block adjacent to the upper end of the current block are unusable reference pixels. At this time, the pixels included in the reference pixel level ref_0 closest to the current block and the reference pixel level ref_2 farthest from the current block can be generated using the usable reference pixels included in the same reference pixel level. In addition, the pixels included in the reference pixel level ref_1 at a distance of 1 pixel from the current block can be generated not only using the pixels included in the same reference pixel level ref_1, but also using the pixels included in different reference pixel levels ref_0 and ref_2. At this time, the unusable reference pixels can be filled with the usable reference pixels on both sides by methods such as bilinear interpolation.

[0416] The above example is an example of generating reference pixels when multiple reference pixel levels are composed of reference pixels and some reference candidate blocks are unusable. Or, it can be set not to allow adaptive reference pixel composition (in this example, adaptive_intra_ref_sample_flag = 0) according to the encoding / decoding setting (such as the case where at least one reference candidate block is unavailable or all reference candidate blocks are unavailable, etc.). That is, the reference pixels can be composed according to the preset setting without generating any additional information.

[0417] The reference pixel interpolation unit can generate reference pixels in fractional units through linear interpolation of reference pixels. In the present invention, it is assumed to be a part of the process of the reference pixel composition unit for explanation, but it can also adopt a composition included in the prediction block generation unit, and it can also be understood as a process executed before generating the prediction block.

[0418] In addition, it is assumed to be an independent process separate from the reference pixel filtering described later, but it can also adopt a process integrated into one process. This is also a configuration provided to solve the distortion of reference pixels caused by an increase in the number of filter applications to reference pixels when applying various filters through the reference pixel interpolation unit and the reference pixel filtering unit.

[0419] The reference pixel interpolation process is not performed in some prediction modes (e.g., horizontal, vertical, some diagonal modes <such as diagonal down right, diagonal down left, diagonal upright, etc., which are 45-degree angle modes>, non-directional modes, color modes, color copy modes, etc., i.e., modes that do not require interpolation in fractional units when generating the prediction block), and can be performed only in other prediction modes (modes that require interpolation in fractional units when generating the prediction block).

[0420] The interpolation accuracy (e.g., 1, 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, etc. pixel units) can be determined according to the prediction mode (or the directionality of the prediction mode). For example, the 45-degree angle prediction mode does not require the interpolation process, but the 22.5-degree or 67.5-degree angle prediction modes require a 1 / 2 pixel unit difference. As described above, at least one interpolation accuracy and the maximum interpolation accuracy amount can be determined according to the prediction mode.

[0421] For reference pixel interpolation, only one pre-set interpolation filter (e.g., 2-tap linear interpolation filter) can be used, or a filter selected from a plurality of interpolation filter candidate groups (e.g., 4-tap cubic filter, 4-tap Gaussian filter, 6-tap Wiener filter, 8-tap Kalman filter, etc.) according to the encoder / decoder settings can be used. At this time, the interpolation filters can be distinguished according to differences such as the number of filter taps (i.e., the number of pixels to which the filter is applied) and filter coefficients.

[0422] The interpolation can be performed in stages in the order from lower accuracy to higher accuracy (e.g., 1 / 2 → 1 / 4 → 1 / 8), or can be performed at once. The former case means that the interpolation is performed based on integer unit pixels and fractional unit pixels (pixels that have been interpolated with a lower accuracy than the pixels currently to be interpolated), and the latter case means that the interpolation is performed based on integer unit pixels.

[0423] When using one of the multiple filter candidate groups, the filter selection information for generation or mode determination can be explicitly indicated, or can be determined according to the encoder / decoder settings (e.g., interpolation accuracy, size, shape, prediction mode, etc. of the block). At this time, the unit for explicit generation can be video, sequence, image, slice, parallel block, block, etc.

[0424] For example, when an interpolation accuracy of 1 / 4 or more (1 / 2, 1 / 4) is adopted, an 8-tap Kalman filter can be applied to a reference pixel in integer units; when an interpolation accuracy less than 1 / 4 and 1 / 16 or more (1 / 8, 1 / 16) is adopted, a 4-tap Gaussian filter can be applied to a reference pixel in integer units and an interpolated reference pixel in units of 1 / 4 or more; and when an interpolation accuracy less than 1 / 16 (1 / 32, 1 / 64) is adopted, a 2-tap linear filter can be applied to a reference pixel in integer units and an interpolated reference pixel in units of 1 / 16 or more.

[0425] Alternatively, an 8-tap Kalman filter can be applied to a block of 64×64 or more, a 6-tap Wiener filter can be applied to a block less than 64×64 and 16×16 or more, and a 4-tap Gaussian filter can be applied to a block less than 16×16.

[0426] Alternatively, a 4-tap cubic filter can be applied to a prediction mode with an angular difference of less than 22.5 degrees based on a vertical or horizontal mode, and a 4-tap Gaussian filter can be applied to a prediction mode with an angular difference of 22.5 degrees or more.

[0427] In addition, multiple groups of filter candidates can be composed of a 4-tap cubic filter, a 6-tap Wiener filter, and an 8-tap Kalman filter in some encoding / decoding settings, and can be composed of a 2-tap linear filter and a 6-tap Wiener filter in some encoding / decoding settings.

[0428] Figures 23a to 23c It is an illustrative diagram showing a method of performing interpolation based on fractional pixel units in a reference pixel configured according to an embodiment of the present invention.

[0429] See Figure 23a , a method of interpolating pixels in fractional units can be confirmed in the case of supporting the use of one reference pixel level (ref_i) as a reference pixel. Specifically, interpolation can be performed by applying filtering (marking the filtering function as int_func_1D) to pixels adjacent to the pixel to be interpolated (marked with an x). Among them, since one reference pixel level is used as a reference pixel, interpolation can be performed using adjacent pixels included in the same reference pixel level as the pixel x to be interpolated.

[0430] See Figure 23b , a method for obtaining interpolated pixels in fractional units can be confirmed when supporting the use of more than two reference pixel levels (ref_i, ref_j, ref_k) as reference pixels. Figure 23b When the reference pixel interpolation process is performed on the reference pixel level ref_j, the interpolation of the interpolation target pixel in the decimal unit can be performed by using other reference pixel levels ref_k and ref_i. k ~h k 、a j ~h j 、a i ~h i Execute filtering (interpolation process, function int_func_1D) to obtain interpolation object pixels (x j and pixels x at positions corresponding to the interpolation target pixel (corresponding positions at each reference pixel level according to the direction of the prediction mode) contained in other reference pixel levels k 、x i And the obtained 1st interpolation pixel x k 、x j 、x i Perform additional filtering (which may be filtering corresponding to weighted averages such as [1, 2, 1] / 4, [1, 6, 1] / 8, etc., which is not an interpolation process) to finally obtain the final interpolation pixel x on the reference pixel level ref_j. In this example, it is assumed that the pixel x on the other reference pixel level corresponding to the interpolation target pixel k 、x i The case of fractional unit pixels that can be obtained through the interpolation process is explained.

[0431] In the above example, a case where a first interpolation pixel can be obtained by filtering at each reference pixel level and a final difference pixel can be obtained by performing additional filtering on the first interpolation pixel as an object is described. However, it is also possible to obtain a final difference pixel by filtering adjacent pixels a at multiple reference pixel levels. k ~h k 、a j ~h j 、a i ~h i The final interpolated pixel is obtained in one go by filtering.

[0432] exist Figure 23b Among the three reference pixel levels supported, the level actually used as a reference pixel can be ref_j. That is, in order to interpolate a reference pixel level composed of reference pixels, other reference pixel levels included in the candidate group (for example, not being composed of reference pixels means that the pixels in the corresponding reference pixel level are not applicable for prediction, but in the above case, the corresponding pixels will be referred to when performing interpolation, so accurately speaking, they can also belong to the usage situation) can be utilized.

[0433] See Figure 23c , the case of using all two supported reference pixel levels as reference pixels is illustrated. The final interpolated pixel x can be obtained by constructing the input pixels using the pixels (in this example, d i , d j , e i , e j ) adjacent to the fractional unit positions where interpolation is to be performed in each supported reference pixel level and performing filtering on the adjacent pixels. At this time, the method of obtaining the final interpolated pixel x by obtaining the first interpolated pixels on each reference pixel level and then performing additional filtering on the first interpolated pixels as shown in Figure 23b can also be adopted.

[0434] The above example is not limited to the reference pixel interpolation process, and can also be understood as a process combined with other processes of intra prediction (such as the reference pixel filtering process, the prediction block generation process, etc.).

[0435] Figures 24a to 24b is the first illustrative diagram for explaining the adaptive reference pixel filtering method applicable to one embodiment of the present invention.

[0436] Generally, the main purpose of the reference pixel filter can be to perform smoothing by using a low-pass filter {Low-passFilter, such as [1, 2, 1] / 4, [2, 3, 6, 3, 2] / 16, etc., 3-tap, 5-tap filters, etc.}, but other types of filters (such as high-pass filters, etc.) can also be used according to the filter application purpose {such as sharpening, etc.}. In the present invention, the case of reducing the distortion generated in the encoding / decoding process by performing filtering for the purpose of smoothing will be mainly described.

[0437] Reference pixel filtering can be determined whether to execute according to the encoding / decoding settings. However, since applying filtering in batches may cause problems in not being able to reflect the local characteristics of the image, performing filtering based on the local characteristics of the image will be more beneficial to improving the encoding performance. Among them, the characteristics of the image can be judged not only according to the image type, color component, quantization parameter, encoding / decoding information of the current block (such as the size, shape, segmentation information, prediction mode, etc. of the current block), but also according to the encoding / decoding information of adjacent blocks and the combination of the encoding / decoding information of the current block and adjacent blocks. In addition, it can also be judged according to the reference pixel distribution characteristics (such as the dispersion, standard deviation, flat area, discontinuous area, etc. of the reference pixel area).

[0438] Refer to Figure 24a , when belonging to a classification (category 0) according to a part of the encoding / decoding settings (such as block size range A, prediction mode B, color component C, etc.), filtering may not be applied, and when belonging to a classification (category 1) according to a part of the encoding / decoding settings (such as the prediction mode A of the current block, the prediction mode B of the pre-set adjacent block, etc.), filtering can be applied.

[0439] Refer to Figure 24b , when belonging to a classification (category 0) according to a part of the encoding / decoding settings (such as the size A of the current block, the size B of the adjacent block, the prediction mode C of the current block, etc.), filtering may not be applied, when belonging to a classification (category 1) according to a part of the encoding / decoding settings (such as the size A of the current block, the shape B of the current block, the size C of the adjacent block, etc.), filtering can be performed using filter A, and when belonging to a classification (category 2) according to a part of the encoding / decoding (such as the parent block A of the current block, the parent block B of the adjacent block, etc.), filtering can be performed using filter B.

[0440] Therefore, it is possible to determine whether to apply filtering, the type of filter, whether to encode the filter information (explicit / implicit), the number of filtering times, etc. according to the size, prediction mode, color component, etc. of the current block and adjacent blocks, and the type of filter can be classified according to the differences in the number of taps, filter coefficients, etc. At this time, when the number of filtering times is 2 or more, the same filter can be applied multiple times or different filters can be applied separately.

[0441] The above example can be a case where the reference pixel filtering is preset according to the characteristics of the image. That is, it can be a case where the filter-related information is implicitly determined. However, when the judgment of the image characteristics as described above is inaccurate, it may instead have an adverse effect on the encoding efficiency, so this part must be considered.

[0442] To prevent the above-described situation from occurring, explicit setting of reference pixel filtering can be performed. For example, information related to whether filtering is applicable can be generated. At this time, when there is only one filter, filter selection information may not be generated, and when there are multiple filter candidate groups, filter selection information can be generated.

[0443] The above examples illustrate the implicit setting and explicit setting related to reference pixel filtering, and a hybrid method can be adopted where determination is made through explicit setting in some cases and through implicit setting in other cases. Here, the meaning of implicit is that information related to the reference pixel filter (such as information on whether filtering is applicable, filter type information) can be derived from the decoder.

[0444] Figure 25 This is the second illustrative diagram for explaining the adaptive reference pixel filtering method applicable to one embodiment of the present invention.

[0445] Refer to Figure 25 , and the categories can be classified by the image characteristics that can be confirmed by using the encoding / decoding information, and the reference pixel filtering can be adaptively performed according to the classified categories.

[0446] For example, filtering is applicable when classified as category 0, and filter A is used when classified as category 1. Category 0 and category 1 can be an example of implicit reference pixel filtering.

[0447] In addition, when classified as category 2, filtering may not be applicable or filter A may be applicable. At this time, the generated information can be information related to whether filtering is applicable, but filter selection information will not be generated.

[0448] In addition, when classified as category 3, filter A or filter B may be applicable. At this time, the generated information can be filter selection information, and the application of the filter can be an example of unconditional execution. That is, when classified as category 3, it can be understood as a situation where filtering must be performed but the filter type needs to be selected.

[0449] In addition, when classified as category 4, filtering may not be applicable, filter A may be applicable, or filter B may be applicable. At this time, the generated information can be information related to whether filtering is applicable and filter selection information.

[0450] In other words, it is possible to determine explicit or implicit processing according to the category, and when executed through explicit processing, the candidate group settings related to each reference pixel filter can be adaptively configured.

[0451] Regarding the above categories, the following examples can be considered.

[0452] First, for a block larger than 64×64, one of <filter off>, <filter on - filter A>, <filter on + filter B>, <filter on + filter C> can be implicitly determined according to the prediction mode of the current block. At this time, the candidate added considering the reference pixel distribution characteristics can be <filter on + filter C>. That is, filter A, filter B, or filter C can be applied when the filter is on.

[0453] In addition, for a block smaller than 64×64 and larger than 16×16, one of <filter off>, <filter on + filter A>, <filter on + filter B> can be implicitly determined according to the prediction mode of the current block.

[0454] In addition, for a block smaller than 16×16, one of <filter off>, <filter on + filter A>, <filter on + filter B> can be selected according to the prediction mode of the current block. At this time, it can be mode - determined as <filter off> in some prediction modes, and one of <filter off>, <filter on + filter A> can be explicitly selected in some prediction modes, and one of <filter off>, <filter on + filter B> can be explicitly selected in some prediction modes.

[0455] As an example of the settings related to multiple reference pixel filters, when the reference pixels obtained in each filtering (including the case of filter off in this example) are the same or similar, generating reference pixel filter information (such as reference pixel filtering tolerance information, reference pixel filter information, etc.) may lead to the generation of unnecessary duplicate information. For example, when the reference pixel distribution characteristics obtained in each filtering (such as the values obtained by averaging, dispersing, etc. of each reference pixel and the threshold <threshold>When the characteristics (judged by comparison) are the same or similar, information related to the reference pixel filter can be omitted. When the information related to the reference pixel filter is omitted, filtering can be applied in a preset method (e.g., filter off). The decoder can determine whether it is necessary to receive the information related to the reference pixel filter in the same way as the encoder after receiving the in-picture prediction information, and can determine whether to receive the information related to the reference pixel filter based on the above determination.

[0456] Assuming that explicit information related to reference pixel filtering is generated, it is possible to generate indication information (adaptive_ref_filter_enabled_flag in this example) that allows adaptive reference pixel filtering in units such as video, sequence, image, slice, parallel block, etc.

[0457] When the above indication information represents allowing adaptive reference pixel filtering (adaptive_ref_filter_enabled_flag = 1 in this example), it is possible to generate adaptive reference pixel filtering allowance information (adaptive_ref_filter_flag in this example) in units such as image, slice, parallel block, block, etc.

[0458] When the above allowance information represents adaptive reference pixel filtering (adaptive_ref_filter_flag = 1 in this example), it is possible to generate information related to the reference pixel filter (e.g., reference pixel filter selection information, etc., ref_filter_idx in this example) in units such as image, slice, parallel block, block, etc.

[0459] At this time, in the case where adaptive reference pixel filtering is not allowed or cannot be applied, the filtering operation can be performed on the reference pixel according to a preset setting (as described above, the applicability of filtering, the type of filtering, etc. are determined in advance according to video coding / decoding information, etc.).

[0460] Figures 26a to 26b It is an exemplary diagram showing a case where one embodiment of the present invention is applied and one reference pixel level is used in reference pixel filtering.

[0461] See Figure 26a , it can be confirmed that interpolation is performed by applying filtering (referred to as the smt_func_1 function) to the target pixel d and the adjacent pixels a, b, c, e, f, g among the pixels included in the reference pixel level ref_i.

[0462] Figure 26a Examples that can typically be applied to sequential filtering can also be applied to multiple filtering. For example, it is possible to apply two - stage filtering to reference pixels (such as a*, b*, c*, d*, etc.) obtained by applying one - stage filtering.

[0463] Refer to Figure 26b , and it is possible to obtain the filtered pixel e* (referred to as the smt_func_2 function) by performing linear interpolation proportional to the distance (such as the distance z from a) on the pixels located on both sides with the target pixel e as the center. Among them, the pixels located on both sides can be the pixels at the two - side ends of adjacent pixels within the block composed of the upper - side block, left - side block, upper - side block + upper - right block, left - side block + lower - left block, upper - left block + upper - side block + upper - right block, upper - left block + left - side block + lower - left block, and upper - left block + left - side block + upper - side block + lower - left block + upper - right block of the current block. Figure 26b It can be reference - pixel filtering performed according to the reference - pixel distribution characteristics.

[0464] In Figures 26a to 26b , what is illustrated is the case of using pixels in the same reference - pixel level as the reference pixel to be filtered for reference - pixel filtering. At this time, the type of filter used in reference - pixel filtering can be the same or different according to the reference - pixel level.

[0465] In addition, in the case of using multiple reference - pixel levels, when performing reference - pixel filtering in some reference - pixel levels, it is possible to use not only pixels in the same reference - pixel level but also pixels in different reference - pixel levels.

[0466] Figure 27 is an exemplary diagram illustrating the case of using multiple reference - pixel levels in reference - pixel filtering in one embodiment of the present invention.

[0467] Refer to Figure 27 , first, it is possible to perform filtering on the pixels included in the same reference - pixel level on the reference - pixel levels ref_k and ref_i respectively. That is, it is possible to obtain the filtered pixel d k and adjacent pixels a k to g k by performing filtering (defined as the function smt_func_1D) on the reference - pixel level ref_k, and it is also possible to obtain the filtered pixel d k * by performing filtering (defined as the function smt_func_1D) on the reference - pixel level ref_i on the target pixel d i and adjacent pixels a i to g i . i *

[0468] In addition, when performing reference pixel filtering at the reference pixel level ref_j, not only can the same reference pixel level ref_j be used, but also pixels included in other reference pixel levels that are spatially adjacent to the reference pixel level ref_j, namely ref_i and ref_k, can be used. Specifically, it is possible to obtain the interpolated pixel d j by applying filtering (defined as the function smt_func_2D) to the spatially adjacent pixels c k , d k , e k , c j , e j , c i , d i , e i (i.e., it can be a filter with a 3×3 square mask). However, it is not limited to the 3×3 square form, and it is also possible to use a filter with a mask in the form of a 5×2 rectangle (b j *, c k , d k , e k , f k , b k , c j , e j , f j , f j ), 3×3 rhombus form (d k , c j , e j , d i ), 5×3 cross form (d k , b j , c j , e j , f j , d i ), etc.

[0469] Among them, the reference pixel level, as shown in the above Figures 18 to 2 2, etc., is composed of pixels included in adjacent blocks adjacent to the current block and close to the boundary of the current block. Considering the above aspects, for filtering using pixels included in the same reference pixel level in the reference pixel level ref_k and the reference pixel level ref_i, a filter with a 1D mask form using pixels horizontally or vertically adjacent to the interpolation target pixel can be applied. However, for the interpolated pixel of the reference pixel d j in the reference pixel level ref_j, it can be obtained by applying a filter with a 2D mask form using all pixels adjacent in the spatial up / down / left / right directions.

[0470] In addition, in each reference pixel level, the reference pixel filtering can be applied twice to the reference pixels for which the reference pixel filtering has been applied once. For example, the reference pixel filtering can be performed once using the reference pixels included in each reference pixel level ref_k, ref_j, ref_i. Next, in the reference pixel levels (referred to as ref_k*, ref_j*, ref_i*) where the reference pixel filtering has been performed once, the reference pixel filtering can be performed not only using the respective reference pixel levels but also using the reference pixels in other reference pixel levels.

[0471] The prediction block generation unit can generate a prediction block based on at least one intra-picture prediction mode (which can be simply referred to as the prediction mode), and can use reference pixels based on the above prediction mode. At this time, the prediction block can be generated by extrapolating, interpolating, or averaging (DC) copying the reference pixels according to the prediction mode. Among them, extrapolation can be applied to the directional mode in the intra-picture prediction mode, and the rest can be applied to the non-directional mode.

[0472] In addition, when copying the reference pixels, one or more prediction pixels can be generated by copying one reference pixel to multiple pixels within the prediction block, or one or more prediction pixels can be generated by copying more than one reference pixel, and the number of copied reference pixels can be equal to or less than the number of predicted pixels to be copied.

[0473] In addition, usually, one prediction block is generated for the prediction of one intra-picture prediction mode, but after obtaining multiple prediction blocks, the final prediction block can also be generated by applying weighted values and the like to the obtained multiple prediction blocks. Among them, the multiple prediction blocks can refer to the prediction blocks obtained according to the reference pixel levels.

[0474] In the prediction mode determination unit of the encoding device, a process for selecting the best mode from multiple prediction mode candidate groups is performed. Generally, the rate-distortion (Rate-Distortion) technique that can predict using the distortion of the block (such as the distortion (Distortion) between the current block and the reconstructed block, the sum of absolute differences (SAD, Sum of Absolute Dirrefence), the sum of square differences (SSD, Sum ofSquare Difference), etc.) and the amount of bits generated in the prediction mode can be used to determine the best mode in terms of encoding cost. The prediction block generated based on the prediction mode determined through the above process can be transmitted to the subtraction operation unit and the addition operation unit (at this time, since the decoding device can obtain the information for indicating the best prediction mode from the encoding device, the process of selecting the best prediction mode can be omitted).

[0475] The prediction mode encoding unit of the encoding device can encode the best intra prediction mode selected by the prediction mode determination unit. At this time, the index information for indicating the best prediction mode can be directly encoded, or the prediction information related to the prediction mode (such as the difference value between the predicted prediction mode index and the prediction mode index of the current block) can be encoded after predicting the best prediction mode using the prediction mode that can be obtained from other surrounding blocks, etc. Among them, the former case can be applied to the chrominance component, and the latter case can be applied to the luminance component.

[0476] When predicting and encoding the best prediction mode of the current block, the predicted value (or prediction information) of the prediction mode can be referred to as the most probable mode (MPM). At this time, the most probable mode (MPM) refers to the prediction mode with the highest possibility of becoming the best prediction mode of the current block, and can be composed of pre-set prediction modes (such as mean (DC), planar (Planar), vertical, horizontal, diagonal mode, etc.) or the prediction modes of spatially adjacent blocks (such as the left, upper, upper left, upper right, lower left blocks, etc.). Among them, the diagonal mode is diagonal up right, diagonal down right, diagonal down left, and can be the mode corresponding to the 2nd, 18th, and 34th modes in Figure 17 No.

[0477] In addition, a pattern derived from a prediction pattern included in a set of prediction patterns consisting of the most probable mode (MPM), i.e., a candidate group of the most probable mode (MPM), can be added to the candidate group of the most probable mode (MPM). In the case of the directional mode, a prediction pattern whose index interval from the prediction pattern included in the candidate group of the most probable mode (MPM) is equal to a preset value can be added to the candidate group of the most probable mode (MPM). For example, when the pattern included in the candidate group of the most probable mode (MPM) is Figure 17 the 10th pattern, the derived pattern can correspond to the 9th, 11th, 8th, 12th patterns, etc.

[0478] The above example can correspond to the case where the candidate group of the most probable mode (MPM) consists of multiple patterns. The composition of the candidate group of the most probable mode (MPM) (e.g., the number of prediction patterns included in the most probable mode (MPM), the composition priority order) is determined according to the encoding / decoding settings (e.g., prediction pattern candidate group, video type, block size, block shape, etc.) and can include at least one pattern composition.

[0479] The priority order of the prediction patterns included in the candidate group of the most probable mode (MPM) can be set. The order of the prediction patterns included in the candidate group of the most probable mode (MPM) can be determined according to the set priority order, and the composition of the candidate group of the most probable mode (MPM) can be completed when the added prediction patterns reach a preset number. Among them, the priority order can be set as the order of the prediction patterns of the blocks spatially adjacent to the current block to be predicted, preset prediction patterns, and patterns derived from the prediction patterns earlier included in the candidate group of the most probable mode (MPM), but it is not limited thereto.

[0480] Specifically, in the spatially adjacent blocks, the priority order can be set in the order of left - upper - lower - left - upper - right - upper - left blocks. In the preset prediction patterns, the priority order can be set in the order of mean (DC) - planar - vertical - horizontal patterns. Next, prediction patterns obtained by performing addition operations such as +1, -1, etc. (integer values) on the index value ( Figure 17 which is the prediction pattern number in this case) of the prediction patterns included in the candidate group of the most probable mode (MPM) can be included in the candidate group of the most probable mode (MPM). As one of the above examples, the priority order can be set in the order of left - upper - mean (DC) - planar - lower - left - upper - right - upper - left - (spatially adjacent block pattern) +1 - (spatially adjacent block pattern) -1 - horizontal - vertical - diagonal, etc.

[0481] In the above example, the case where the priority order of the candidate group of the most probable mode (MPM) is fixed is described. However, the above priority order can also be adaptively determined according to the shape, size, etc. of the block.

[0482] When encoding the prediction mode of the current block using the most probable mode (MPM), information related to whether the prediction mode matches the most probable mode (MPM) (e.g., most_probable_mode_flag) can be generated.

[0483] When it matches the most probable mode (MPM) (e.g., most_probable_mode_flag = 1), additional most probable mode (MPM) index information (e.g., mpm_idx) can be generated according to the composition of the most probable mode (MPM). For example, when the most probable mode (MPM) consists of one prediction mode, no additional most probable mode (MPM) index information needs to be generated, while when it consists of multiple prediction modes, index information corresponding to the prediction mode of the current block in the candidate group of the most probable mode (MPM) can be generated.

[0484] When it does not match the most probable mode (MPM) (e.g., most_probable_mode_flag = 0), non-most probable mode (non-MPM) index information (e.g., non_mpm_idx) corresponding to the prediction mode of the current block can be generated from the remaining prediction mode candidate group after excluding the candidate group of the most probable mode (MPM) from the supported intra-picture prediction modes (referred to as the non-most probable mode (non-MPM) candidate group). This can be an example of the case where the non-most probable mode (non-MPM) is formed as a single group.

[0485] When the non-most probable mode (non-MPM) candidate group consists of multiple groups, information related to which group the prediction mode of the current block belongs to can be generated. For example, when the non-most probable mode (non-MPM) consists of two groups, A and B, and the prediction mode of the current block matches the prediction mode of group A (e.g., non_mpm_A_flag = 1), index information corresponding to the prediction mode of the current block can be generated in the candidate group of group A, while when it does not match (e.g., non_mpm_A_flag = 0), index information corresponding to the prediction mode of the current block can be generated in the remaining prediction mode candidate group (or the candidate group of group B). As shown in the above example, the non-most probable mode (non-MPM) can consist of multiple groups, and the number of groups can be specified according to the prediction mode candidate group. For example, when the number of prediction mode candidates is 35 or less, it can be 1, while in other cases it can be 2.

[0486] At this time, a specific group A can be composed of patterns that are determined to have a relatively high probability of being consistent with the predicted pattern of the current block after the most probable mode (MPM) candidate group. For example, it is possible to include the next predicted pattern not included in the most probable mode (MPM) candidate group or an oriented pattern with a certain interval in group A.

[0487] As in the above example, when the non-most probable mode (non-MPM) is composed of multiple groups, it is possible to achieve the effect of reducing the number of pattern coding bits when the number of predicted patterns is large and the predicted pattern of the current block is inconsistent with the most probable mode (MPM).

[0488] When encoding (or decoding the predicted pattern) of the predicted pattern of the current block using the most probable mode (MPM), it is possible to individually generate a binarization table applicable to each predicted pattern candidate group (such as the most probable mode (MPM) candidate group, non-most probable mode (non-MPM) candidate group, etc.), and different binarization methods can be individually applied according to each candidate group.

[0489] In the above example, terms such as the most probable mode (MPM) candidate group and non-most probable mode (non-MPM) candidate group are only part of the terms used in the present invention and are not limited thereby. Specifically, it is only information indicating which category a current in-picture prediction mode belongs to when classifying the in-picture prediction mode into multiple categories and the mode information within the corresponding category. Terms such as the first most probable mode (MPM) candidate group and the second most probable mode (MPM) candidate group can also be used for substitution.

[0490] Figure 28 It is a block diagram for explaining an in-picture prediction mode encoding / decoding method to which one embodiment of the present invention is applied.

[0491] Refer to Figure 28 , first, obtain mpm_flag (S10). Next, confirm whether it matches the 1st most probable mode (MPM) (indicated by mpm_flag) (S11), and when it matches, confirm the most probable mode (MPM) index information (mpm_idx) (S12). When it does not match the most probable mode (MPM), obtain rem_mpm_flag (S13). Next, confirm whether it matches the 2nd most probable mode (MPM) (indicated by rem_mpm_flag) (S14), and when it matches, confirm the 2nd most probable mode (MPM) index information (rem_mpm_idx) (S16). When it does not match the 2nd most probable mode (MPM), confirm the index information (rem_mode_idx) of the candidate group composed of the remaining prediction modes (S15). In this example, the case where the index information generated according to whether it matches the 2nd most probable mode (MPM) is represented by the same syntax element is described, but other mode coding settings (such as binary method) can also be applied, and different above-mentioned index information can also be set for processing.

[0492] In the video decoding method applying one embodiment of the present invention, the intra prediction within a picture can be configured in the following manner. The intra prediction of the prediction unit can include a prediction mode decoding step, a reference pixel composition step, and a prediction block generation step. In addition, the video decoding apparatus can include a prediction mode decoding unit, a reference pixel composition unit, and a prediction block generation unit for executing the prediction mode decoding step, the reference pixel composition step, and the prediction block generation step. Some of the above processes can be omitted or other processes can be added, and the order can also be changed to other orders different from the order described above.

[0493] Since the reference pixel composition unit and the prediction block generation unit of the video decoding apparatus can play the same role as those in the video encoding apparatus, the detailed description related thereto will be omitted here, and the prediction mode decoding unit can use the method used in the prediction mode encoding unit in reverse.

[0494] Next, it will be combined with Figures 29 to 31 Multiple embodiments of intra prediction based on the reference pixel composition of the decoding apparatus will be described. Among them, the descriptions related to reference pixel level support and reference pixel filtering method described in combination with the accompanying drawings above should be interpreted as being equally applicable in the decoding apparatus, and in order to prevent repeated description, the detailed description related thereto will be omitted.

[0495] Figure 29 It is the first illustrative diagram for explaining the bitstream composition of intra prediction based on reference pixel composition.

[0496] In Figure 29 the first exemplary diagram, it is premised that multiple reference pixel levels are supported, at least one of the supported reference pixel levels is used as a reference pixel, multiple candidate groups related to reference pixel filtering are supported, and one filter is selected therefrom.

[0497] After forming pixel candidate groups using multiple reference pixel levels in the encoder (in this example, in a state where the reference pixel generation process has been completed), a reference pixel is formed using at least one reference pixel level, and then reference pixel filtering and reference pixel interpolation are applied. At this time, multiple candidate groups related to reference pixel filtering are supported.

[0498] Next, a process for selecting the best mode from the prediction mode candidate group is performed. After determining the best prediction mode, a prediction block based on the corresponding mode is generated and transmitted to the subtraction operation unit, and then an encoding process for the information related to intra-picture prediction is performed. In this example, it is premised that the reference pixel level and reference pixel filtering are implicitly determined according to the encoding information.

[0499] In the decoder, the information related to intra-picture prediction (such as the prediction mode, etc.) is reconstructed, and after generating a prediction block based on the reconstructed prediction mode, it is transmitted to the subtraction operation unit. At this time, it is premised that the reference pixel level and reference pixel filtering used for generating the prediction block are implicitly determined.

[0500] Refer to Figure 29 , a bitstream (S20) can be formed using one intra-picture prediction mode (intra_mode). At this time, the reference pixel level ref_idx and the reference pixel filtering category ref_filter_idx supported (or used) in the current block can be implicitly determined according to the intra-picture prediction mode (determined as categories (Category) A and B respectively, S21 to S22). At this time, the encoding / decoding information (such as the video type, color component, block size, and shape, etc.) can be additionally considered.

[0501] Figure 30 is the second exemplary diagram for explaining the formation of the bitstream of intra-picture prediction based on reference pixels.

[0502] In Figure 30 the second exemplary diagram, it is premised that multiple reference pixel levels are supported and one of the supported multiple reference pixel levels is used as a reference pixel. In addition, it is premised that multiple candidate groups related to reference pixel filtering are supported and one filter is selected therefrom. The difference from Figure 29 is that the information related to generation and selection is explicitly generated by the encoding device.

[0503] After determining that multiple reference pixel levels are supported in the encoder, a reference pixel is formed using one reference pixel level, and then reference pixel filtering and reference pixel interpolation are applied. At this time, multiple filtering methods related to reference pixel filtering are supported.

[0504] When performing the process of determining the best prediction mode of the current block in the encoder, it is also possible to additionally consider the process of selecting the best reference pixel level and the process of selecting the best reference pixel filtering in each prediction mode. After determining the best prediction mode, reference pixel level, and reference pixel filtering of the current block, the prediction block generated based on this is passed to the subtraction operation unit and the encoding process of the intra-prediction related information in the picture is performed.

[0505] In the decoder, the intra-prediction related information (such as prediction mode, reference pixel level, reference pixel filtering information, etc.) is reconstructed, and after generating a prediction block using the reconstructed information, it is passed to the subtraction operation unit. At this time, the reference pixel level and reference pixel filtering used for generating the prediction block comply with the settings determined according to the information transmitted from the encoder.

[0506] See Figure 30 , the decoder confirms the best prediction mode of the current block through the intra-prediction mode information (intra_mode) included in the bitstream (S30), and confirms whether multiple reference pixel levels are supported (multi_ref_flag) (S31). When multiple reference pixel levels are supported, the reference pixel level selection information (ref_idx) is confirmed (S32) to determine the reference pixel level that can be used in intra-prediction. When multiple reference pixel levels are not supported, the process of obtaining the reference pixel level selection information (ref_idx) (S32) can be omitted.

[0507] Next, it is confirmed whether adaptive reference pixel filtering is supported (adap_ref_smooth_flag) (S33), and when adaptive reference pixel filtering is supported, the filtering method of the reference pixel is determined through the reference pixel filter information (ref_filter_idx) (S34).

[0508] Figure 31 It is the 3rd illustrative diagram for explaining the bitstream composition of intra-prediction based on reference pixels.

[0509] In Figure 31 In the 3rd illustrative diagram, it is premised that multiple reference pixel levels are supported and one reference pixel level among the multiple reference pixel levels is used. In addition, it is premised that multiple candidate groups related to reference pixel filtering are supported and one filter is selected from them. And Figure 30 The difference is that the selection information is generated adaptively.

[0510] After constructing the reference pixels using one of the multiple reference pixel levels supported by the encoder, reference pixel filtering and reference pixel interpolation are applied. At this time, various filtering related to reference pixel filtering is supported.

[0511] When performing the process of selecting the best mode from multiple candidate groups of prediction modes, it is also possible to additionally consider the process of selecting the best reference pixel level and the process of selecting the best reference pixel filtering in each prediction mode. After determining the best prediction mode, reference pixel level, and reference pixel filtering, a prediction block is generated based on this and passed to the subtraction operation unit, and then the encoding process of the information related to in-picture prediction is performed.

[0512] At this time, the repeatability of the generated prediction block is confirmed. When it is the same as or similar to the prediction block obtained using other reference pixel levels, the selection information related to the best reference pixel level is omitted and a preset reference pixel level is used. At this time, the preset reference pixel level can be the level closest to the current block.

[0513] For example, it is possible to Figure 19c judge the repeatability based on the difference value (distortion value) between the prediction block generated by ref_0 in

[0514] and the prediction block generated by ref_1. When the above difference value is less than a preset threshold value, it is determined that the prediction block has repeatability; otherwise, it is determined that the prediction block does not have repeatability. At this time, the above threshold value can be determined adaptively according to quantization parameters, etc.

[0515] In addition, the best reference pixel filtering information also confirms the repeatability of the prediction block. When it is the same as or similar to the prediction block obtained using other reference pixel filtering, the reference pixel filtering information is omitted and a preset reference pixel filtering is applied.

[0516] In the decoder, the in-picture prediction related information (such as prediction mode, reference pixel level, reference pixel filtering information, etc.) is reconstructed and passed to the subtraction operation unit after generating a prediction block therefrom. At this time, the reference pixel level information and reference pixel filtering for generating the prediction block comply with the settings determined according to the information transmitted from the encoder, and the decoder can directly confirm whether there is repetition (without passing syntax elements) and then comply with the preset method when there is repetition.

[0517] See Figure 31 , the decoder first confirms the in-picture prediction mode information (intra_mode) of the current block (S40), and confirms the support for multiple reference pixel levels (multi_ref_flag) (S41). When multiple reference pixel levels are supported, a repetition check of the prediction block based on the supported multiple reference pixel levels is performed (represented by the ref_check process, S42). When the repetition confirmation result is that there is no repetition in the prediction block (redund_ref = 0, S43), the selection information (ref_idx) of the reference pixel level is referred to from the bitstream (S44) and the optimal reference pixel level is determined.

[0518] Next, the support for adaptive reference pixel filtering (adap_ref_smooth_flag) is confirmed (S45). When adaptive reference pixel filtering is supported, a repetition check of the prediction block for the supported multiple reference pixel filtering methods is performed (represented by the ref_check process, S46). When there is no repetition in the prediction block (redund_ref = 0, S47), the selection information (ref_filter_idx) of the reference pixel filtering method is referred to from the bitstream (S48) and the optimal reference pixel filtering method is determined.

[0519] At this time, redund_ref in the figure is a value used to indicate the repetition confirmation result, and when it is 0, it means there is no repetition.

[0520] In addition, the decoder can perform in-picture prediction using the preset reference pixel level and the preset reference pixel filtering method when the prediction block has repetition.

[0521] Figure 32 is a flowchart illustrating an image decoding method supporting multiple reference pixel levels according to an embodiment of the present invention.

[0522] See Figure 32 , An image decoding method supporting multiple reference pixel levels may include: a step (S100) of confirming whether multiple reference pixel levels are supported through a bitstream; when multiple reference pixel levels are supported, a step (S110) of determining the reference pixel level to be used in the current block by referring to the syntax information included in the above bitstream; a step (S120) of forming reference pixels using the pixels included in the determined reference pixel level; and a step (S130) of performing intra prediction on the above current block using the formed reference pixels.

[0523] Wherein, after the step (S100) of confirming whether multiple reference pixel levels are supported, it may further include: a step of confirming whether an adaptive reference pixel filtering method is supported through a bitstream.

[0524] Wherein, after the step (S100) of confirming whether multiple reference pixel levels are supported, it may further include: when multiple reference pixel levels are not supported, a step of forming reference pixels using a preset reference pixel level.

[0525] The method applicable to the present invention can be implemented in the form of program instructions executable by various computing means and recorded in a computer-readable medium. The computer-readable medium can include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded in the computer-readable medium can be program instructions specially designed for the present invention or program instructions well-known and available to computer software practitioners.

[0526] Examples of the computer-readable medium can include special hardware devices such as read-only memory (ROM), random access memory (RAM), flash memory, etc. for storing and executing program instructions. Examples of program instructions not only include machine code generated by a compiler, but also include high-level language code that can be executed in a computer using an interpreter, etc. The above hardware device can be composed of at least one software module for performing the actions applicable to the present invention, and vice versa.

[0527] In addition, the above method or device can combine or separate all or part of its composition or functions.

[0528] The above content has been described in combination with preferred embodiments applicable to the present invention, but those skilled in the relevant technical field should be able to understand that the present invention can be variously modified and changed without departing from the spirit and scope of the present invention described in the appended claims.< / threshold> < / array> < / postfilter> < / display> < / sizing> < / y>

Claims

1. A method for decoding an image using segmentation units, comprising: segmenting the received bitstream into at least one segmentation unit by referring to segmentation information obtained from the received bitstream; deriving an intra prediction mode of a current block in the segmentation unit; deriving a reference sample layer of the current block from multiple reference sample layers; generating a reference sample based on the derived reference sample layer; and and generating a prediction sample of the current block based on the reference sample and the intra prediction mode, wherein the segmentation information includes width information and height information, the width information indicates the width of the segmentation unit in units of a basic coding unit, and the height information indicates the height of the segmentation unit in units of the basic coding unit, wherein the multiple reference sample layers are included in reconstructed adjacent blocks of the current block, and wherein each of the reference sample layers of the current block has a different distance from the boundary of the current block.

2. The method according to claim 1, wherein when the intra prediction mode is a non - directional mode, the prediction sample is generated only by using the reference sample layer having the shortest distance from the boundary of the current block.

3. The method according to claim 2, wherein the non - directional mode represents a planar mode.

4. The method according to claim 1, wherein the reference sample layer is derived based on color components.

5. The method according to claim 1, wherein the number of the multiple reference sample layers is equal to or greater than 3.

6. A method for encoding an image using segmentation units, comprising: segmenting the image into at least one segmentation unit; determining an intra prediction mode of a current block in the segmentation unit; determining a reference sample layer of the current block from multiple reference sample layers; generating a reference sample based on the determined reference sample layer; and generating a prediction sample of the current block based on the reference sample and the intra prediction mode, wherein segmentation information is encoded into a bitstream based on the segmentation of the image, wherein the segmentation information includes width information and height information, the width information indicates the width of the segmentation unit in units of a basic coding unit, and the height information indicates the height of the segmentation unit in units of the basic coding unit, wherein the multiple reference sample layers are included in reconstructed adjacent blocks of the current block, and wherein each of the reference sample layers of the current block has a different distance from the boundary of the current block.

7. The method according to claim 6, wherein when the intra prediction mode is a non - directional mode, the prediction sample is generated only by using the reference sample layer having the shortest distance from the boundary of the current block.

8. The method according to claim 7, wherein the non - directional mode represents a planar mode.

9. The method according to claim 6, wherein the reference sample layer is determined based on color components.

10. The method according to claim 6, wherein The number of the multiple reference sample layers is equal to or greater than 3.

11. A method for transmitting a bitstream generated by an encoding method, the encoding method comprising: dividing an image into at least one division unit; determining an intra prediction mode of a current block in the division unit; determining a reference sample layer of the current block from a plurality of reference sample layers; generating a reference sample based on the determined reference sample layer; and generating a prediction sample of the current block based on the reference sample and the intra prediction mode, wherein division information is encoded into the bitstream based on the division of the image, wherein the division information includes width information and height information, the width information indicates the width of the division unit in units of a basic coding unit, and the height information indicates the height of the division unit in units of the basic coding unit, wherein the plurality of reference sample layers are included in reconstructed adjacent blocks of the current block, and wherein each of the reference sample layers of the current block has a different distance from the boundary of the current block.

Citation Information

Patent Citations

  • Video encoding method, video decoding method, video encoding device, video decoding device, and programs for same

    CN103069803A

  • Method and apparatus of image encoding / decoding using reference pixel composition in intra prediction

    KR1020160143586A