Video decoding method and device using division units including additional regions

By introducing additional areas and multiple reference pixel levels in image decoding, the problem of low image coding efficiency in the existing technology is solved, and more efficient image compression and accurate intra-frame prediction are achieved.

CN116320400BActive Publication Date: 2025-10-10INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310305085.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-16
Filing Date
2018-07-03
Publication Date
2025-10-10
Estimated Expiration
2038-07-03

AI Technical Summary

Technical Problem

In existing image coding methods, independently coded parallel blocks cannot effectively utilize temporally and spatially adjacent image data as references, resulting in low coding efficiency. In addition, intra-frame prediction relies on the nearest neighbor pixel reference, which may not be applicable.

Method used

A segmented unit image decoding method that includes an additional area is adopted. Syntax elements are obtained through the bitstream, the additional area is set, and blocks within the additional area are referenced during decoding. Multiple reference pixel levels and adaptive reference pixel filtering are supported to improve image compression efficiency and intra-frame prediction accuracy.

Benefits of technology

It improves image compression efficiency and intra-frame prediction accuracy, improves compression efficiency during image encoding/decoding, and supports adaptive reference pixel filtering to adapt to different image characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320400B_ABST
    Figure CN116320400B_ABST
Patent Text Reader

Abstract

Disclosed is an image decoding method and apparatus using a partition unit including an additional area. The image decoding method using a partition unit including an additional area includes the steps of partitioning an encoded image included in a received bitstream into at least one partition unit by referring to syntax elements acquired from the bitstream; setting an additional area for the at least one partition unit; and decoding the encoded image based on the partition unit after the additional area is set. Thus, the image coding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of application number 2018800450230 filed on July 3, 2018, and the invention name is "Image decoding method and device using segmentation units including additional areas." Technical Field

[0002] The present invention relates to a method and apparatus for decoding an image using a segmentation unit including an additional region, and more particularly to a technique for improving coding efficiency by setting additional regions on the top, bottom, left, and right sides of a segmentation unit such as a tile within an image and simultaneously referencing image data in the additional regions during encoding. Background Art

[0003] In recent years, the demand for multimedia data, such as video, on the internet has been rapidly increasing. However, the current rate of increase in channel bandwidth still requires a method to effectively compress the rapidly increasing amount of multimedia data. To this end, the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / ISE and the Video Coding Experts Group (VCEG) of the International Telecommunication Union's Telecommunication Standardization Sector (ITU-T) are working tirelessly to develop more efficient video compression standards.

[0004] Furthermore, when independent coding is performed on an image, independent coding is usually performed on individual division units including tiles. Therefore, there is a problem in that image data of other division units that are temporally and spatially adjacent cannot be used as a reference.

[0005] Therefore, a solution is needed that can maintain the existing parallel processing based on independence coding while also using adjacent image data as a reference.

[0006] Furthermore, in the intra-frame prediction based on the existing image encoding / decoding method, the pixels closest to the current block are used as reference pixels. However, depending on the type of the image, the method of using the closest pixels as reference pixels may not be desirable.

[0007] Therefore, a method is needed to improve the efficiency of intra-frame prediction by adopting a reference pixel composition method different from the existing method. Summary of the Invention

[0008] Technical issues

[0009] To solve the existing problems as described above, the present application aims to provide an image decoding apparatus and method using a partition unit including an additional area.

[0010] To solve the existing problems as described above, another object of the present application is to provide an image encoding apparatus and method using a partition unit including an additional area.

[0011] To solve the existing problems as described above, the present application aims to provide an image decoding method supporting multiple reference pixel levels.

[0012] To solve the existing problems as described above, another object of the present application is to provide an image decoding apparatus supporting multiple reference pixel levels.

[0013] Technical Solution

[0014] To achieve the above objects, in one aspect of the present application, there is provided an image decoding method using a partition unit including an additional area.

[0015] The image decoding method using a partition unit including an additional area can include the steps of partitioning an encoded image included in a received bitstream into at least one partition unit by referring to a syntax element acquired from the bitstream, setting an additional area for the at least one partition unit, and decoding the encoded image based on the partition unit after the additional area is set.

[0016] The step of decoding the encoded image can include the step of determining a reference block related to a current block to be decoded in the encoded image, according to information indicating reference possibility included in the bitstream.

[0017] The reference block can be a block included in a position overlapping the additional area set in a partition unit to which the reference block belongs.

[0018] To achieve the above objects, in another aspect of the present application, there is provided an image decoding method supporting multiple reference pixel levels.

[0019] An image decoding method that supports multiple reference pixel levels may include: a step of confirming whether multiple reference pixel levels are supported through a bitstream; when multiple reference pixel levels are supported, a step of determining the reference pixel level to be used in the current block by referring to syntax information contained in the above-mentioned bitstream; a step of constructing reference pixels using pixels contained in the determined reference pixel level; and a step of performing intra-frame prediction of the above-mentioned current block using the constructed reference pixels.

[0020] After the above step of confirming whether multiple reference pixel levels are supported, the method may further include: confirming whether an adaptive reference pixel filtering method is supported through a bitstream.

[0021] After the step of confirming whether multiple reference pixel levels are supported, the method may further include: when multiple reference pixel levels are not supported, constructing reference pixels using a preset reference pixel level.

[0022] Technical Effects

[0023] When the video decoding method and apparatus using the division unit including the additional region according to the present invention are used as described above, the video compression efficiency can be improved because more video data can be used as a reference.

[0024] When the image decoding method and apparatus supporting multiple reference pixel levels according to the present invention are used, the accuracy of intra-frame prediction can be improved because multiple reference pixels can be utilized.

[0025] Furthermore, in the present invention, since adaptive reference pixel filtering is supported, optimal reference pixel filtering can be performed according to the characteristics of the image.

[0026] In addition, it can also improve the compression efficiency during image encoding / decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a conceptual diagram illustrating a video encoding and decoding system according to an embodiment of the present invention.

[0028] Figure 2 A block diagram illustrating a video encoding device according to an embodiment of the present invention;

[0029] Figure 3 This is a diagram illustrating the structure of a video decoding device according to one embodiment of the present invention;

[0030] Figures 4a to 4d This is a conceptual diagram for explaining a projection format applicable to one embodiment of the present invention;

[0031] Figures 5a to 5c This is a conceptual diagram for explaining a surface configuration according to one embodiment of the present invention.

[0032] Figures 6a to 6b This is an illustrative diagram for explaining a dividing portion to which one embodiment of the present invention is applicable;

[0033] Figure 7 This is an example diagram of dividing an image into multiple parallel blocks;

[0034] Figures 8a to 8i Yes Figure 7 FIG. 1 is a first example diagram of setting additional areas for each parallel block shown in FIG.

[0035] Figures 9a to 9i Yes Figure 7 FIG2 is a second example diagram showing the setting of additional areas for each parallel block shown in FIG2;

[0036] Figure 10 This is an illustrative diagram showing how an additional region generated in an embodiment of the present invention is used in the encoding / decoding process of other regions.

[0037] Figures 11 to 12 is a flowchart for explaining a method for encoding / decoding a segmentation unit according to an embodiment of the present invention;

[0038] Figures 13a to 13g This is an illustrative diagram for explaining an area that can be referenced by a specific segmentation unit;

[0039] Figures 14a to 14e This is a flowchart for explaining the possibility of referring to an additional area in a division unit to which one embodiment of the present invention is applied;

[0040] Figure 15 This is an example diagram illustrating blocks included in the division unit of the current image and blocks included in the division units of other images;

[0041] Figure 16 This is a diagram illustrating the hardware configuration of a video encoding / decoding device according to one embodiment of the present invention;

[0042] Figure 17 This is an exemplary diagram illustrating an intra-frame prediction mode to which one embodiment of the present invention is applied;

[0043] Figure 18 This is a first exemplary diagram illustrating a reference pixel structure used in intra-frame prediction according to an embodiment of the present invention;

[0044] Figures 19a to 19cA second exemplary diagram illustrating a reference pixel configuration according to an embodiment of the present invention;

[0045] Figure 20 A third exemplary diagram illustrating a reference pixel configuration according to an embodiment of the present invention;

[0046] Figure 21 A fourth exemplary diagram illustrating a reference pixel configuration according to an embodiment of the present invention;

[0047] Figures 22a to 22b is an exemplary diagram illustrating a method of filling reference pixels at predetermined positions in an unusable reference candidate block;

[0048] Figures 23a to 23c is an exemplary diagram illustrating a method of performing interpolation on a fractional pixel basis in reference pixels configured according to an embodiment of the present invention;

[0049] Figures 24a to 24b is a first exemplary diagram for illustrating an adaptive reference pixel filtering method according to an embodiment of the present invention;

[0050] Figure 25 is a second exemplary diagram for illustrating an adaptive reference pixel filtering method according to an embodiment of the present invention;

[0051] Figures 26a to 26b is an exemplary diagram illustrating a case where one reference pixel level is used in reference pixel filtering according to one embodiment of the present invention;

[0052] Figure 27 is an exemplary diagram illustrating the use of multiple reference pixel levels in reference pixel filtering according to one embodiment of the present invention;

[0053] Figure 28 This is a block diagram for explaining an encoding / decoding method of an intra-frame prediction mode to which one embodiment of the present invention is applied;

[0054] Figure 29 This is a first exemplary diagram for explaining the bitstream structure of intra-frame prediction based on reference pixel structure;

[0055] Figure 30 This is a second exemplary diagram for explaining the bit stream structure of intra-frame prediction based on reference pixel structure;

[0056] Figure 31 This is a third exemplary diagram for explaining the bit stream structure of intra-frame prediction based on reference pixel structure;

[0057] Figure 32The flowchart illustrates a method for image decoding supporting multiple reference pixel levels according to an embodiment of the present invention. DETAILED DESCRIPTION

[0058] The present invention is susceptible to various modifications and embodiments. Specific embodiments will be illustrated and described in detail below. However, the following is not intended to limit the present invention to any specific embodiment. Instead, the present invention should be understood to encompass all modifications, equivalents, and even substitutes within the spirit and technical scope of the present invention. Similar reference numerals are used throughout the drawings to denote similar components.

[0059] In the process of describing different constituent elements, terms such as 1st, 2nd, A, B, etc. can be used, but the above constituent elements are not limited by the above terms. The above terms are only used to distinguish one constituent element from other constituent elements. For example, without departing from the scope of the claims of the present invention, the first constituent element can also be named as the second constituent element, and similarly, the second constituent element can also be named as the first constituent element. The term "and / or" includes a combination of multiple related recorded items or a single item from multiple related recorded items.

[0060] When a component is described as being “connected to” or “in contact with” another component, it should be understood that it can not only be directly connected to or in contact with the other component, but also that other components can exist between the two components. Conversely, when a component is described as being “directly connected to” or “directly in contact with” another component, it should be understood that no other components exist between the two components.

[0061] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit the present invention. Unless the context clearly indicates otherwise, singular statements also include plural meanings. In this application, terms such as "including" or "having" are only used to indicate the existence of features, numbers, steps, actions, constituent elements, parts, or combinations thereof described in the specification, and should not be understood as excluding in advance the possibility that one or more other features, numbers, steps, actions, constituent elements, parts, or combinations thereof may exist or be added.

[0062] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meanings as those commonly understood by persons having ordinary knowledge of the technical field to which this invention belongs. Commonly used terms, such as those defined in dictionaries, should be interpreted as having the same meanings as those in the context of the relevant technology and should not be interpreted as idealized or exaggerated unless otherwise clearly defined in this application.

[0063] Typically, an image consists of a series of still images, which can be divided into groups of pictures (GOPs). Each still image is called a picture or frame. Superordinate concepts include groups of pictures (GOPs) and sequences, and each image can be divided into specific regions such as slices, tiles, and blocks. Furthermore, a GOP can include units such as I-pictures, P-pictures, and B-pictures. I-pictures are images that are encoded / decoded independently without using a reference image, while P-pictures and B-pictures are images that are encoded / decoded using reference images to perform processes such as motion estimation and motion compensation. Typically, P-pictures use I-pictures and P-pictures as reference images, while B-pictures use I-pictures and P-pictures as reference images. However, these definitions can vary depending on the encoding / decoding settings.

[0064] The image used as a reference during the encoding / decoding process is called a reference image, and the block or pixel used as a reference is called a reference block or reference pixel. Furthermore, reference data can include not only spatial domain pixel values ​​but also frequency domain coefficient values, as well as various encoding / decoding information generated and determined during the encoding / decoding process.

[0065] The minimum unit of the image can be a pixel, and the number of bits used to represent one pixel is referred to as bit depth. Generally, the bit depth can be 8 bits, and other bit depths can be supported according to encoding settings. With respect to bit depth, at least one bit depth can be supported according to a color space. Further, at least one color space can be composed according to a color format of the image. According to the color format, one or more images having a certain size or one or more images having different sizes can be composed. For example, in the case of YCbCr 4:2:0, one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example) can be composed, and the composition ratio of the chrominance components and the luminance component can be 1:2 horizontally and vertically. As another example, in the case of 4:4:4, the composition ratio can be the same horizontally and vertically. In the case of composition by one or more color spaces as described above, partitioning in each color space can be performed.

[0066] In the present application, a part of the color space (Y in this example) of a part of the color format (YCbCr in this example) will be described as a reference, and the same or similar application can be made in other color spaces (Cb, Cr in this example) based on the color format (depending on the setting of the specific color space). However, a part of the difference can also be retained in each color space (independent of the setting of the specific color space). That is, depending on the setting of each color space can refer to the property of being proportional or dependent on the composition ratio of each component (for example, determined according to 4:2:0, 4:2:2, 4:4:4, etc.), and independent of the setting of each color space can refer to the property of being independent of the composition ratio of each component or only applicable to the corresponding color space. In the present application, depending on the encoder / decoder, a part of the composition can have the property of independence or the property of dependence.

[0067] The configuration information or syntax elements required during the image coding process can be determined at the unit level, such as video, sequence, picture, slice, tile, or block. These can be included in the bitstream and transmitted to the decoder in units such as the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice Header, Tile Header, and Block Header. The decoder can parse the configuration information transmitted from the encoder at the same level and use it in the image decoding process after decoding it. Furthermore, relevant information can be transmitted to the bitstream in the form of supplemental enhancement information (SEI) or metadata and used after parsing. Each parameter set has its own number value, and a lower-level parameter set can include the number value of a higher-level parameter set for reference. For example, a lower-level parameter set can reference information of an upper-level parameter set with a consistent number value from more than one upper-level parameter set. In the example of multiple units as described above, when a unit contains more than one other unit, the corresponding unit can be called a higher-level unit and the contained units can be called lower-level units.

[0068] The configuration information generated for the above-mentioned units can include configuration information related to independent settings in each unit, as well as configuration information related to settings that are dependent on previous, subsequent, or superior units. The dependent configuration refers to flag information (e.g., a 1-bit flag, 1 for compliance and 0 for non-compliance) used to indicate whether the configuration information of the previous, subsequent, or superior units is followed, and can be understood as representing the configuration information of the corresponding unit. Although the present invention will focus on examples related to independent configurations regarding configuration information, examples can also include those that are supplemented or replaced by configuration information that is dependent on previous, subsequent, or superior units of the current unit.

[0069] Next, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0070] Figure 1 This is a conceptual diagram illustrating a video encoding and decoding system according to an embodiment of the present invention.

[0071] See Figure 1The image encoding device 105 and the decoding device 100 can be user terminals such as personal computers (PCs), notebook computers, personal digital assistants (PDAs), portable multimedia players (PMPs), portable game stations (PSPs, PlayStation Portables), wireless communication terminals (Wireless Communication Terminals), smart phones (SmartPhone), televisions (TVs), or server terminals such as application servers and business servers. They can include various devices such as communication devices such as communication modems for communicating with various devices or wired and wireless communication networks, memories (memory) 120, 125 for storing various applications and data for performing inter-frame or intra-frame prediction for encoding or decoding images, and processors (processors) 110, 115 for performing calculations and controls by executing applications. Furthermore, the video encoded into a bitstream by the video encoding device 105 can be transmitted to the video decoding device 100 via a wired or wireless communication network (network), such as the Internet, a short-range wireless communication network, a wireless local area network, a wireless broadband network, or a mobile communication network, or via various communication interfaces, such as a cable or a universal serial bus (USB), and then decoded and reconstructed into a video in the video decoding device 100 for playback. Furthermore, the video encoded into a bitstream by the video encoding device 105 can also be transferred from the video encoding device 105 to the video decoding device 100 via a computer-readable storage medium.

[0072] Figure 2 This is a block diagram illustrating an image encoding device according to one embodiment of the present invention.

[0073] The video decoding device 20 of this embodiment is suitable for Figure 2 As shown, it can include a prediction unit 200, a subtraction unit 205, a transformation unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse transformation unit 225, an addition unit 230, a filtering unit 235, a coded image buffer 240 and an entropy coding unit 245.

[0074] The prediction unit 200 can include an intra prediction unit for performing intra prediction and an inter prediction unit for performing inter prediction. Intra prediction can generate a prediction block by performing spatial prediction using pixels of a block adjacent to the current block, and inter prediction can generate a prediction block by searching for an area that best matches the current block from a reference image and performing motion compensation. The specific information related to each prediction method (e.g., intra prediction mode, motion vector, reference image, etc.) can be determined after determining which of intra prediction or inter prediction is applied to the corresponding unit (coding unit or prediction unit). At this time, the processing unit for performing prediction and the processing unit for determining the prediction method and the specific content can differ according to the encoding / decoding setting. For example, the prediction method and the prediction mode can be determined at the prediction unit, and the execution of prediction can be performed at the transform unit.

[0075] The intra prediction unit can employ directional prediction modes such as horizontal, vertical modes, etc. used according to the prediction direction and non-directional prediction modes such as average (DC), planar (Planar) using the average, interpolation, etc. of reference pixels. Through the directional and non-directional modes, an intra prediction mode candidate group can be constructed, and one of various options such as 35 prediction modes (directional 33 + non-directional 2) or 67 prediction modes (directional 65 + non-directional 2), 131 prediction modes (directional 129 + non-directional 2), etc. can be used as a candidate group.

[0076] The intra prediction unit can include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel construction unit can construct reference pixels for performing intra prediction using pixels included in the adjacent block and adjacent to the current block with the current block as the center. According to the encoding setting, the reference pixels can be constructed using one of the most adjacent reference pixel rows, or other adjacent reference pixel rows, or a plurality of reference pixel rows. When a part of the reference pixels is not available, the reference pixels can be generated using the available reference pixels, and when all are not available, the reference pixels can be generated using a pre-set value (e.g., a middle value of the pixel value range that can be expressed by the bit depth, etc.).

[0077] The reference pixel filter unit of the intra-frame prediction unit can filter the reference pixels to reduce distortion remaining after the encoding process. The filter used can be a low-pass filter such as a 3-tap filter (1 / 4, 1 / 2, 1 / 4) or a 5-tap filter (2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16). The applicability and type of filtering can be determined based on encoding information (such as block size, shape, and prediction mode).

[0078] The reference pixel interpolation unit of the intra-frame prediction unit can generate decimal pixels through a linear interpolation process of reference pixels according to the prediction mode, and can determine the applicable interpolation filter according to the encoding information. At this time, the interpolation filter used can include a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc. Usually, the process of performing low-pass filtering and the process of performing interpolation are independent of each other, but it is also possible to integrate the filters applicable to the two processes into one and then perform the filtering process.

[0079] The prediction mode determination unit of the intra-frame prediction unit selects the optimal prediction mode from a candidate set of prediction modes, taking into account coding costs. The prediction block generation unit generates a prediction block using the corresponding prediction mode. The prediction mode encoding unit encodes the optimal prediction mode based on the predicted value. This allows adaptive encoding of prediction information based on whether the predicted value is appropriate or inappropriate.

[0080] In the intra-frame prediction unit, the prediction value is referred to as the Most Probable Mode (MPM). A subset of modes can be selected from all modes included in the prediction mode candidate set to form the MPM candidate set. The MPM candidate set can include pre-set prediction modes (e.g., mean (DC), planar, vertical, horizontal, and diagonal modes) or prediction modes for spatially adjacent blocks (e.g., left, top, upper left, upper right, and lower left blocks). Furthermore, the MPM candidate set can be formed using modes derived from the modes pre-included in the MPM candidate set (e.g., differences of +1 or -1 in directional modes).

[0081] Prediction modes used to form a most probable mode (MPM) candidate set may have a priority order. The order of inclusion in the most probable mode (MPM) candidate set may be determined based on the priority order, and the formation of the most probable mode (MPM) candidate set may be completed when the number of most probable mode (MPM) candidate sets (determined based on the number of prediction mode candidate sets) is filled according to the priority order. In this case, the priority order may be determined based on the order of prediction modes of spatially adjacent blocks, a pre-set prediction mode, or a mode derived from a prediction mode included earlier in the most probable mode (MPM) candidate set, and other variations are also possible.

[0082] For example, spatially adjacent blocks can be included in the candidate group in the order of left-upper-left-lower-right-upper block, etc., and pre-set prediction modes can be included in the candidate group in the order of mean (DC)-plane (Planar)-vertical-horizontal mode, etc., and modes obtained by adding +1, -1, etc. to the pre-included modes can be included in the candidate group, thereby forming a candidate group with a total of 6 modes. Alternatively, a candidate group can be formed with a total of 7 modes by including them in a priority order such as left-upper-mean (DC)-plane (Plana)-lower-left-upper-right-(left+1)-(left-1)-(top+1).

[0083] In the candidate group formation described above, a validity check can be performed, so that only valid candidates are included in the candidate group, and if invalid, the next candidate is skipped. For example, it can be invalidated when the adjacent block is located outside the image, included in a different partition unit than the current block, or the coding mode of the corresponding block is inter-picture prediction. In addition, it can also be invalidated in the case of non-reference conditions described later in this invention.

[0084] In the above post-selection, spatially adjacent blocks can be composed of a single block or multiple blocks (sub-blocks). Therefore, in the order (left-top) in the above candidate group composition, the order can be to jump to the top block after performing a validity check on a certain position in the left block (e.g., the bottom block of the left block), or the order can be to jump to the top block after performing a validity check on multiple positions (e.g., starting from the top block of the left block and one or more sub-blocks located below it), and the order can also be determined according to the encoding settings.

[0085] The inter-frame prediction unit can be divided into a moving motion model and a non-moving motion model based on the motion prediction method. In the moving motion model, prediction is performed while only parallel motion is considered. In the non-moving motion model, prediction is performed while considering rotation, perspective, zoom in / out, and other motions in addition to parallel motion. Assuming unidirectional prediction, the moving motion model requires one motion vector, while the non-moving motion model requires more than one motion vector. In the non-moving motion model, each motion vector can be information applicable to a predetermined position in the current block, such as the upper left vertex or upper right vertex of the current block. Using the corresponding motion vector, the position of the area to be predicted in the current block can be obtained in units of pixels or sub-blocks. The inter-frame prediction unit can apply some of the processes described below in common according to the above motion models while also applying other processes individually.

[0086] The inter-frame prediction unit can include a reference picture construction unit, a motion prediction unit, a motion compensation unit, a motion information determination unit, and a motion information encoding unit. The reference picture construction unit can include previously or subsequently encoded pictures in a reference picture list (L0, L1) centered around the current picture. Prediction blocks can be obtained from reference pictures included in the reference picture list. Furthermore, depending on encoding settings, a reference picture can be constructed using the current picture and included in at least one position in the reference picture list.

[0087] In the inter-picture prediction unit, the reference image constructing unit can include a reference image interpolating unit, and can perform an interpolation process for fractional pixels according to the interpolation accuracy. For example, an 8-tap discrete cosine transform (DCT)-based interpolation filter can be applied to the luma component, and a 4-tap discrete cosine transform (DCT)-based interpolation filter can be applied to the chroma component.

[0088] In the inter-picture prediction unit, the motion prediction unit is used to perform the process of exploring blocks with high correlation with the current block through the reference image. It can use various methods such as the full search-based block matching algorithm (FBMA) and the three-step search algorithm (TSS). The motion compensation unit is used to perform the process of obtaining the predicted block through the motion prediction process.

[0089] In the inter-picture prediction unit, the motion information determination unit can perform a process for selecting optimal motion information for the current block. The motion information can be encoded using motion information coding modes such as Skip Mode, Merge Mode, and Competition Mode. These modes can be configured by combining supported modes based on the motion model. Examples include Skip Mode (Move), Skip Mode (Non-Move), Merge Mode (Move), Merge Mode (Non-Move), Competition Mode (Move), and Competition Mode (Non-Move). Depending on the symbolization settings, some of these modes can be included in the candidate set.

[0090] The motion information coding modes described above can obtain a predicted value for the current block's motion information (motion vector, reference image, prediction direction, etc.) from at least one candidate block, and can generate optimal candidate selection information when supporting two or more candidate blocks. Skip mode (without a residual signal) and merge mode (with a residual signal) can directly use the predicted value as the motion information for the current block, while competitive mode generates the difference between the current block's motion information and the predicted value.

[0091] The candidate set of motion information prediction values ​​for the current block is adaptive depending on the motion information coding mode and can adopt various configurations. The candidate set includes motion information for spatially adjacent blocks (e.g., blocks to the left, above, upper left, upper right, and lower left) and temporally adjacent blocks (e.g., blocks to the left, right, above, below, upper left, upper right, lower left, and lower right, including the center block in another image corresponding to or associated with the current block). Furthermore, the candidate set includes mixed motion information of spatial and temporal candidates (e.g., information obtained by averaging or medianing the motion information of spatially and temporally adjacent blocks, or motion information obtained for the current block or a subblock of the current block).

[0092] A priority order may exist in the formation of motion information predictor candidate groups. The order of inclusion in the formation of the predictor candidate groups may be determined based on the priority order, and the formation of the candidate groups may be completed when the number of candidate groups (determined by the motion information coding mode) is filled according to the priority order. In this case, the priority order may be determined based on the order of motion information of spatially adjacent blocks, motion information of temporally adjacent blocks, or mixed motion information of spatial and temporal candidates, and other variations are also possible.

[0093] For example, in spatially adjacent blocks, the blocks can be included in the candidate group in the order of left-top-right-top-left-bottom-left-top, and in temporally adjacent blocks, the blocks can be included in the candidate group in the order of bottom-middle-right-bottom.

[0094] In the above candidate group formation, an effectiveness check can be performed, so that the blocks are included in the candidate group only when they are effective and skipped to the next candidate when they are not effective. For example, the blocks can be not effective when the adjacent blocks are located outside the image or are included in a different partition unit from the current block or the encoding mode of the corresponding blocks is intra prediction, and in addition, can be not effective in the case of non-reference described later in the present application.

[0095] In the above candidate group formation, the blocks adjacent in space or time can be formed of one block or a plurality of blocks (sub-blocks). Therefore, in the order such as (left-top) in the above spatial candidate group formation, the order can be to jump to the top block after performing an effectiveness check on a certain position (for example, the lowermost block of the left block) in the left block, or the order can be to jump to the top block after performing an effectiveness check on a plurality of positions (for example, one or more sub-blocks below the uppermost block of the left block in the predetermined order from the block <2, 2>). In addition, in the order such as (middle-right) in the temporal candidate group formation, the order can be to jump to the right block after performing an effectiveness check on a certain position (for example, <2, 2> when the central block is partitioned into 4 x 4 regions) in the middle block, or the order can be to jump to the bottom block after performing an effectiveness check on a plurality of positions (for example, one or more sub-blocks such as <3, 3>, <2, 3>, and the like on the predetermined order from the position block <2, 2>), and can be determined according to the encoding setting.

[0096] The subtraction operation section 205 generates a residual block by performing a subtraction operation on the current block and the prediction block. That is, the subtraction operation section 205 generates a residual signal in the form of a block, that is, a residual block, by calculating the difference between the pixel value of each pixel of the current block to be encoded and the prediction pixel value of each pixel of the prediction block generated by the prediction section.

[0097] The transform unit 210 transforms each pixel value of the residual block into a frequency coefficient by transforming the residual block into a frequency region. Among them, the transform unit 210 can transform the residual signal into a frequency signal using various transform techniques for transforming a pixel signal of a spatial axis into a frequency axis such as Hadamard transform, DCT based transform, DST based transform, KLT based transform, and the like, and the residual signal transformed into a frequency region will become a frequency coefficient. When transforming, it can be transformed by a 1-dimensional transform matrix. Each transform matrix can be adaptively used in horizontal and vertical units. For example, when the prediction mode in intra prediction is horizontal, a transform matrix based on DCT can be used in the vertical direction and a transform matrix based on DST can be used in the horizontal direction. And when the prediction mode is vertical, a transform matrix based on DCT can be used in the horizontal direction and a transform matrix based on DST can be used in the vertical direction.

[0098] The quantization unit 215 quantizes the residual block containing the frequency coefficient transformed into a frequency region by the transform unit 210. Among them, the quantization unit 215 can quantize the transformed residual block using a quantization technique such as dead zone uniform threshold quantization, quantization weighted matrix, or an improved quantization technique thereof. At this time, one or more quantization techniques can be selected as candidates and can be determined according to the encoding mode, prediction mode information, and the like.

[0099] The entropy encoding unit 245 generates a quantization coefficient sequence by scanning the generated quantization frequency coefficient sequence using various scanning methods, and outputs it after encoding using an entropy encoding technique or the like. As a scanning mode, one of various modes such as zigzag, diagonal, raster, and the like can be set. In addition, it can generate and output encoding data containing encoding information transferred from each constituent unit to a bitstream.

[0100] The inverse quantization unit 220 inverse quantizes the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 generates a residual block containing a frequency coefficient by inverse quantizing the quantization frequency coefficient sequence.

[0101] The inverse transform unit 225 inversely transforms the residual block inversely quantized by the inverse quantization unit 220. Specifically, the inverse transform unit 225 inversely transforms the frequency coefficients of the inversely quantized residual block to generate a residual block containing pixel values, i.e., a reconstructed residual block. The inverse transform unit 225 can perform the inverse transform by inversely applying the transform method used by the transform unit 210.

[0102] The addition unit 230 can reconstruct the current block by adding the prediction block predicted by the prediction unit 200 and the residual block reconstructed by the inverse transformation unit 225. The reconstructed current block is stored as a reference image (or reference block) in the coded image buffer 240, and can be used as a reference image when encoding the next block of the current block or other subsequent blocks or other images.

[0103] The filtering unit 235 can include one or more post-processing filtering processes such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can eliminate block distortion that occurs at the boundaries between blocks in the reconstructed image. The adaptive loop filter (ALF) can perform filtering based on the value of comparing the reconstructed image with the original image after filtering the block through the deblocking filter. The sample adaptive offset (SAO) can reconstruct the offset difference between the residual block to which the deblocking filter has been applied and the original image in units of pixels. The post-processing filters described above can be applied to the reconstructed image or block.

[0104] The deblocking filter within the filter unit can be applied based on pixels contained in a number of columns or rows between two blocks, based on a block boundary. The above-mentioned block boundary is preferably applied to the boundaries of coding blocks, prediction blocks, and transform blocks, and can be limited to blocks of a predetermined minimum size (e.g., 8×8) or larger.

[0105] Regarding the applicability of filtering, the applicability and strength of filtering can be determined by taking into account the characteristics of block boundaries. The filtering can be selected from among strong, medium, or weak filtering. Furthermore, when the block boundary is a segmentation unit boundary, the applicability of the loop filter is determined based on a loop filter applicability flag at the segmentation unit boundary. The applicability can also be determined based on various conditions described later in this disclosure.

[0106] The sample adaptive offset (SAO) in the filter section can be applied based on a difference value between the reconstructed image and the original image. As the offset type, edge offset and band offset can be supported, and one of the above offsets can be selected according to the characteristics of the image to perform filtering. In addition, the above offset-related information can be encoded in a block unit, and can be encoded by a prediction value related thereto. At this time, the related information can be adaptively encoded according to a suitable case and an unsuitable case of the prediction value. The prediction value can be offset information of a neighboring block (for example, a left block, an upper block, a left upper block, a right upper block, etc.), and selection information related to which block's offset information is acquired can be generated.

[0107] An effectiveness check can be performed in the above candidate composition, so that only when it is effective, it is included in the candidate group, and when it is ineffective, it jumps to the next candidate. For example, a neighboring block can be located outside the image or included in a different partition unit from the current block or in a case where it cannot be referenced as described later in the present invention, it can be ineffective.

[0108] The coded picture buffer 240 can store a block or an image reconstructed by the filter section 235. The reconstructed block or image stored in the coded picture buffer 240 can be provided to the prediction section 200 for performing intra prediction or inter prediction.

[0109] Figure 3 is a configuration diagram illustrating an image decoding apparatus to which an embodiment of the present invention is applied.

[0110] Referring to Figure 3 , the image decoding apparatus 30 can include an entropy decoding section 305, a prediction section 310, an inverse quantization section 315, an inverse transform section 320, an adder / subtracter 325, a filter 330, and a decoded picture buffer 335.

[0111] In addition, the prediction section 310 can further include an intra prediction module and an inter prediction module.

[0112] First, when an image bitstream transferred from the image encoding apparatus 20 is received, it can be transferred to the entropy decoding section 305.

[0113] The entropy decoding section 305 can decode the decoded data including the quantized coefficients and the decoded information transferred from each configuration section by decoding the bitstream.

[0114] The prediction unit 310 can generate a prediction block based on the data transmitted from the entropy decoding unit 305. At this time, the prediction unit 310 can also construct a reference picture list using a default construction technique based on the decoded reference picture stored in the picture buffer 335.

[0115] The intra-frame prediction unit can include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction block generation unit, and a prediction mode decoding unit. The inter-frame prediction unit can include a reference image construction unit, a motion compensation unit, and a motion information decoding unit. One part of the unit can perform the same process as the encoder, while the other part can perform the reverse induction process.

[0116] The inverse quantization unit 315 can inversely quantize the quantized transform coefficients supplied from the bit stream and decoded by the entropy decoding unit 305 .

[0117] The inverse transform unit 320 may generate a residual block by applying an inverse discrete cosine transform (DCT), an inverse integer transform, or a similar inverse transform technique to the transform coefficients.

[0118] At this time, the inverse quantization unit 315 and the inverse transform unit 320 reverse the processes performed by the transform unit 210 and the quantization unit 215 of the video encoding device 20 described above, and can be implemented using various methods. For example, the same processes and inverse transforms shared by the transform unit 210 and the quantization unit 215 can be used, or the transform and quantization processes can be reversed using information related to the transform and quantization processes of the video encoding device 20 (e.g., transform size, transform shape, quantization type, etc.).

[0119] The residual block after the inverse quantization and inverse transformation process can be added to the prediction block derived in the prediction unit 310 to generate a reconstructed image block. The addition operation can be performed by the addition and subtraction operator 325.

[0120] For the reconstructed image blocks, the filter 330 can apply a deblocking filter for eliminating blocking as needed, and can also use other loop filters before and after the decoding process to improve video quality.

[0121] The reconstructed and filtered image blocks can be stored in the decoded image buffer 335 .

[0122] Although not shown in the figure, the video decoding device 30 can further include a segmentation unit, in which case the segmentation unit can include an image segmentation unit and a block segmentation unit. Figure 2The same or corresponding structures in the image decoding device shown in the figure can be easily understood by general technicians, so their detailed description will be omitted here.

[0123] Figures 4a to 4d This is a conceptual diagram for explaining a projection format to which one embodiment of the present invention is applicable.

[0124] Figure 4a The equirectangular projection (ERP) format for projecting a 360-degree image onto a two-dimensional plane is illustrated. Figure 4b The CubeMap Projection (CMP) format for projecting a 360-degree image onto a cube is illustrated. Figure 4c The Octahedron Projection (OHP) format for projecting a 360-degree image onto an octahedron is illustrated. Figure 4d The icosahedral projection (ISP) format for projecting a 360-degree image onto a polyhedron is illustrated. However, the present invention is not limited to this format, and various projection formats can be used. For example, the truncated square pyramid projection (TSP) and the segmented spherical projection (SSP) can also be used. Figures 4a to 4d The left side is a 3D model, and the right side is an example of a 2D model transformed into a 2D space through projection. A 2D projected image can be composed of one or more surfaces, and each surface can be a circle, triangle, quadrilateral, or other shape.

[0125] like Figures 4a to 4d As shown, the projection format can be composed of one surface (e.g., equirectangular projection (ERP)) or multiple surfaces (e.g., cubic projection (CMP), octahedral projection (OHP), icosahedral projection (ISP), etc.). In addition, each surface can be classified into shapes such as quadrilateral and triangle. The above classification can be an example of the type, characteristics, etc. of the image in the present invention that can be applied when different encoding / decoding settings are set according to the projection format. For example, the type of image can be a 360-degree image, and the characteristics of the image can be one of the above classifications (e.g., each projection format, a projection format of one surface or multiple surfaces, a projection format in which the surface is quadrilateral or not quadrilateral, etc.).

[0126] A two-dimensional coordinate system {e.g., (i, j)} can be defined for each surface of a two-dimensional projection image. The characteristics of the coordinate system may vary depending on the projection format, the location of each surface, and so on. For example, an equirectangular projection (ERP) may include a single two-dimensional coordinate system, while other projection formats may include multiple two-dimensional coordinate systems depending on the number of surfaces. In this case, the coordinate system can be represented by (k, i, j), where k can be the index information of each surface.

[0127] For the convenience of explanation, the present invention will focus on the case where the surface shape is a quadrilateral, and the number of surfaces projected into 2D can be 1 (for example, equirectangular projection, that is, the image is equivalent to one surface) to more than 2 (for example, cube projection, etc.).

[0128] Figures 5a to 5c This is a conceptual diagram for explaining a surface configuration to which one embodiment of the present invention is applied.

[0129] When projecting a 3D image onto a 2D projection format, the surface configuration must be determined. This configuration can be done to maintain image continuity in 3D space, or to minimize the distance between adjacent surfaces even if image continuity is compromised. Furthermore, when configuring surfaces, some surfaces can be rotated at a certain angle (e.g., 0, 90, 180, 270 degrees).

[0130] See Figure 5a , it is possible to confirm examples of surface layouts related to the cubic projection (CMP) format. When arranging surfaces to maintain image continuity in three-dimensional space, a 4×3 layout can be used, as shown in the image on the left, where four surfaces are arranged horizontally and then one surface is arranged above and below. Furthermore, a 3×2 layout can be used, as shown in the image on the right, where surfaces are arranged seamlessly on a two-dimensional plane even if the image continuity between adjacent surfaces is partially lost.

[0131] See Figure 5b , it is possible to check the surface configuration related to the Octahedron Projection (OHP) format. When the configuration is performed in a way that maintains the continuity of the image in 3D space, it is shown in the upper image. At this time, even if the continuity is partially destroyed, the surface can be seamlessly configured on the projected 2D plane, as shown in the lower image.

[0132] See Figure 5c, it is possible to confirm the surface configuration related to the icosahedral projection (ISP) format. At this time, it can be configured in a manner that maintains image continuity in three-dimensional space as shown in the upper image, or it can be configured in a manner that fits tightly between surfaces without gaps as shown in the lower image.

[0133] At this time, the process of seamlessly arranging the surface can be called frame packing, and the phenomenon of image continuity being disrupted can be minimized by rotating the surface before arranging it. Next, the process of changing the surface configuration to another surface configuration as described above will be called surface reconfiguration.

[0134] Next, the term "continuity" can be interpreted as the continuity of scenes visible to the naked eye in three-dimensional space, or the continuity of actual images or scenes in two-dimensional projection space. Continuity can also be expressed as high correlation between regions. Generally, correlation between regions in a two-dimensional image can be high or low, but in a 360-degree image, there may be areas that lack any continuity even if they are spatially adjacent. Furthermore, due to the surface configuration or reconfiguration described above, there may be areas that are not spatially adjacent but nonetheless have continuity.

[0135] Surface reconfiguration can be performed to improve encoding performance. For example, surfaces having image continuity can be arranged adjacent to each other by performing surface reconfiguration.

[0136] In this context, surface reconfiguration does not necessarily mean reconstructing the surface after configuration. It can also be understood as the process of setting a specific surface configuration from the beginning. (This can be performed during region-wise packing during the 360-degree image encoding / decoding process.)

[0137] Furthermore, surface configuration or reconfiguration can include not only positional changes of the respective surfaces (in this example, simple movement of the surfaces, such as movement from the upper left end of the image to the lower left end or the lower right end), but also surface rotation. Surface rotation can include 0 degrees (no surface rotation), 45 degrees to the right, 90 degrees to the left, and so on. Furthermore, after dividing 360 degrees (equally or unequally) into k (or 2k) intervals, the divided intervals can be selected to represent the rotation angle.

[0138] The encoder / decoder can perform surface configuration (or reconfiguration) according to pre-set surface configuration information (surface form, number of surfaces, surface position, surface rotation angle, etc.) and / or surface reconfiguration information (information indicating the position or movement angle, movement direction, etc. of each surface). In addition, the encoder can generate surface configuration information and / or surface reconfiguration information based on the input image, and the decoder can receive and decode the above information from the encoder to perform surface configuration (or reconfiguration).

[0139] In the following, surfaces are referred to as Figure 5a The 3×2 layout in FIG is a precondition, and in this case, the numbers representing the respective surfaces can be 0 to 5 in a grid-like scanning order starting from the upper left end.

[0140] Next in Figure 5a When describing the continuity between surfaces in the image, unless otherwise specified, it is assumed that surfaces 0 through 2 are continuous with each other, surfaces 3 through 5 are continuous with each other, and that surfaces 0 and 3, 1 and 4, and 2 and 5 are not continuous with each other. Whether or not these surfaces are continuous can be determined by settings such as the characteristics, type, and format of the image.

[0141] During the encoding / decoding process of a 360-degree image, an encoding device can obtain an input image, pre-process the obtained image, encode the pre-processed image, and transmit the encoded bitstream to a decoding device. Pre-processing can include image stitching, projection of a 3D image onto a 2D plane, surface configuration and reconfiguration (also known as region-wise packing), etc. Furthermore, a decoding device can receive the bitstream, decode the received bitstream, perform post-processing (such as image rendering) on ​​the decoded image, and generate an output image.

[0142] At this time, the bitstream can include information generated during pre-processing (supplemental enhancement information (SEI) messages or metadata, etc.) and information generated during the encoding process (video encoding data) and then be transmitted.

[0143] Figures 6a to 6b This is an illustrative diagram for explaining a dividing portion to which one embodiment of the present invention is applied.

[0144] Figure 2 or Figure 3The image encoding / decoding device may further include a segmentation unit, which may include an image segmentation unit and a block segmentation unit. The image segmentation unit may segment the image into at least one processing unit {e.g., a color space (YCbCr, RGB, XYZ, etc.), a sub-image, a slice, a parallel block, a basic coding unit (or a maximum coding unit), etc.}, while the block segmentation unit may segment the basic coding unit into at least one processing unit (e.g., a coding, prediction, transform, quantization, entropy, or loop filtering unit, etc.).

[0145] The basic coding unit can be obtained by dividing an image horizontally or vertically at intervals of a certain length. This unit can also be a sub-image, tile, slice, or surface. In other words, the above units can be composed of integer multiples of the basic coding unit, but are not limited to this.

[0146] For example, different basic coding units can be applied to some partition units (in this example, parallel blocks, sub-pictures, etc.), and each partition unit can have an independent basic coding unit size. In other words, the basic coding unit of the above partition unit can be set to be the same as or different from the basic coding unit of the picture unit and the basic coding unit of other partition units.

[0147] For the convenience of explanation in the present invention, the basic coding unit and other processing units (coding, prediction, transformation, etc.) are referred to as blocks.

[0148] The size or shape of the block can be a power of 2 in terms of horizontal or vertical length (2 n ) in an N×N square format (2n×2n, 256×256, 128×128, 64×64, 32×32, 16×16, 8×8, 4×4, etc., where n is an integer between 2 and 8), or an M×N rectangular format (2m×2n). For example, it is possible to segment an extremely high-resolution 8k ultra-high-definition (UHD) input image into 256×256 size, a 1080p high-definition (HD) input image into 128×128 size, and a wide video graphics array (WVGA) input image into 16×16 size.

[0149] An image can be divided into at least one slice. A slice can be composed of a combination of at least one block that is continuous in scanning order. Each slice can be divided into at least one slice segment, and each slice segment can be divided into basic coding units.

[0150] The image can be divided into at least one sub-image or parallel block. The sub-image or parallel block can be divided in a quadrangle (rectangle or square) form, and can be divided into a basic coding unit. The sub-image is similar to the parallel block in terms of being divided in the same form (quadrangle). However, the sub-image is different from the parallel block in terms of being able to be distinguished from the parallel block in terms of being independently coded / decoded. That is, the parallel block receives setting information for performing coding / decoding from a superior unit (e.g., an image, etc.), whereas the sub-image is able to directly acquire at least one setting information for performing coding / decoding from header information of each sub-image. That is, the parallel block and the sub-image are different in that the parallel block is a unit obtained by simply dividing an image, and is not a unit for transmitting data (e.g., a basic unit of a video coding layer (VCL)).

[0151] In addition, the parallel block can be a division unit supported in terms of parallel processing, and the sub-image can be a division unit supported in terms of independent coding / decoding. In detail, the sub-image can not only be set in units of sub-images, but also determine whether to code / decode, and can be constituted and displayed with a region of interest as a center, and settings related thereto can be determined in units of a sequence, an image, etc.

[0152] In the above-described example, it is also possible to change to a manner in which coding / decoding is set in a superior unit of a sub-image and independent coding / decoding is set in a parallel block unit. In the present application, for convenience of explanation, it will be assumed that the parallel block is set independently and in dependence on a superior unit.

[0153] The division information generated when an image is divided in a quadrangle form can be in various forms.

[0154] Referring to Figure 6a, a quadrilateral-shaped segmentation unit can be obtained by segmenting the image along the horizontal line b7 and the vertical line (in this case, b1 and b3, b2 and b4 will form a segmentation line) at one time. For example, the number of quadrilaterals can be generated based on the horizontal and vertical directions. At this time, when the quadrilateral is evenly divided, the horizontal and vertical lengths of the image can be divided by the number of horizontal lines and vertical lines respectively to confirm the horizontal and vertical lengths of the segmented quadrilateral, and when the quadrilateral is not evenly divided, information indicating the horizontal and vertical lengths of the quadrilateral can be additionally generated. At this time, the horizontal and vertical lengths can be expressed in one pixel unit or in pixel units. In the case of expressing in multiple pixel units, for example, when the size of the basic coding unit is M×N and the size of the quadrilateral is 8M×4N, the horizontal and vertical lengths can be expressed as 8 and 4 respectively (in this example, when the basic coding unit of the corresponding segmentation unit is M×N), or as 16 and 8 (in this example, when the basic coding unit of the corresponding segmentation unit is M / 2×N / 2).

[0155] In addition, see Figure 6b , able to Figure 6a The situation of obtaining quadrilateral-shaped segmentation units by independently segmenting the image is confirmed. For example, information on the number of quadrilaterals in the image, information on the starting position of each quadrilateral in the horizontal and vertical directions (the positions indicated by the numbers z0 to z5 in the figure can be expressed as x and y coordinates in the image), and information on the horizontal and vertical lengths of each quadrilateral can be generated. In this case, the starting position can be expressed in units of one pixel or multiple pixels, and the horizontal and vertical lengths can also be expressed in units of one pixel or multiple pixels.

[0156] Figure 6a can be instances of segmentation information for parallel blocks or sub-images, and Figure 6b This can be an example of sub-image segmentation information, but is not limited to this. For ease of explanation, the following description of parallel blocks assumes quadrilateral segmentation units, but the description related to parallel blocks can be applied identically or similarly to sub-images (and also to surfaces). That is, in the present invention, the distinction is made solely through terminology; in reality, the description related to parallel blocks can be used as the definition of sub-images, and the description related to sub-images can also be used as the definition of parallel blocks.

[0157] It is not necessary to include part of the division units mentioned above, and all or part of them can be selectively included according to encoding / decoding settings. Other additional units (such as surfaces) can also be supported.

[0158] In addition, the block partitioning part can be used to divide the coding units (or blocks) into various sizes. In this case, the coding unit can be composed of multiple coding blocks according to the color format (for example, one luminance coding block and two color difference coding blocks, etc.). For the convenience of explanation, it will be assumed that one color component unit is used for explanation. The coding block can adopt a variable size such as M×M (for example, M is 4, 8, 16, 32, 64, 128, etc.). In addition, according to the partitioning method (for example, tree structure partitioning, i.e., quadtree partitioning), the coding unit can be divided into multiple coding blocks ...<QuadTree,QT> Partition, binary tree<Binary Tree,BT> Split, ternary tree<Ternary Tree,TT> The coding block can be divided into M×N (for example, M and N are 4, 8, 16, 32, 64, 128, etc.) variable sizes. In this case, the coding block can be the basic unit for intra-picture prediction, inter-picture prediction, transform, quantization, entropy coding, etc.

[0159] Although the present invention assumes that multiple sub-blocks of the same size and shape are obtained through segmentation (symmetrical) as an example, it can also be applied to situations involving asymmetric sub-blocks (for example, when binary tree segmentation is performed, the horizontal ratio <vertical same> between the segmented blocks is 1:3 or 3:1 or the vertical ratio <horizontal same> is 1:3 or 3:1, etc., and when ternary tree segmentation is performed, the horizontal ratio <vertical same> between the segmented blocks is 1:2:1 or the vertical ratio <horizontal same> is 1:2:1, etc.).

[0160] The partitioning (M×N) of the coding block can adopt a recursive tree structure. In this case, whether the coding block is partitioned or not can be indicated by a partition flag. For example, when the partition flag of the coding block with a partition depth (Depth) of k is 0, the coding of the coding block is performed on the coding block with a partition depth of k, and when the partition flag of the coding block with a partition depth of k is 1, the coding of the coding block will be performed in 4 sub-coding blocks (quadtree partitioning) or 2 sub-coding blocks (binary tree partitioning) or 3 sub-coding blocks (ternary tree partitioning) with a partition depth of k+1 according to the partitioning method.

[0161] The above sub-coding block will be reset to coding block k+1 and can be split again into sub-coding block k+2 through the above process. The quadtree partitioning can support a split flag (for example, to indicate whether to split or not).

[0162] In binary tree partitioning, a split flag and a split direction flag (horizontal or vertical) can be supported. When more than one split ratio is supported in binary tree partitioning (for example, supporting additional split ratios other than 1:1 for horizontal or vertical splitting, i.e., asymmetric splitting), a split ratio flag can also be supported (for example, selecting a ratio from a candidate group of horizontal or vertical ratios <1:1, 1:2, 2:1, 1:3, 3:1>), or other flags can be supported (for example, whether the split is symmetrical or not, when 1, it is a symmetrical split that does not include additional information, when 0, it is an asymmetric split that requires additional information related to the ratio).

[0163] In the ternary tree partitioning, a split flag and a split direction flag can be supported. When more than one split ratio is supported in the ternary tree partitioning, additional split information similar to the above-mentioned binary tree partitioning will be required.

[0164] The above example is segmentation information generated when only one tree-structured segmentation method is valid. However, when multiple tree-structured segmentation methods are valid, the following segmentation information can be generated.

[0165] For example, when multiple tree-like partitioning is supported, if a pre-set partitioning priority order exists, the partitioning information corresponding to the priority order can be first constructed. In this case, if the partition flag corresponding to the priority order is true, additional partitioning information related to the corresponding partitioning method can be further included, while if it is false (partitioning is not performed), the partitioning information (partition flag, partition direction flag, etc.) corresponding to the partitioning method of the next order can be constructed.

[0166] Alternatively, when multiple types of tree-like partitioning are supported, selection information related to the partitioning method can be additionally generated and composed of partitioning information related to the selected partitioning method.

[0167] Some of the above-mentioned segmentation flags can be omitted based on the results of the earlier executed upper or previous segmentation.

[0168] Block segmentation can be performed from the largest coding block to the smallest coding block. Alternatively, it can be performed from the smallest segmentation depth 0 to the largest segmentation depth. That is, segmentation can be performed recursively until the block size reaches the minimum coding block size or the segmentation depth reaches the maximum segmentation depth. In this case, it can be performed according to the encoding / decoding settings (such as image <slice, parallel block> type). , Coding mode <Intra / Inter>, Color difference component <y cb cr>etc.), and adaptively set the size of the maximum coding block, the size of the minimum coding block, and the maximum segmentation depth.

[0169] For example, when the maximum coding block is 128×128, quadtree partitioning can be performed in the range of 32×32 to 128×128, binary tree partitioning can be performed in the range of 16×16 to 64×64 with a maximum partition depth of 3, and ternary tree partitioning can be performed in the range of 8×8 to 32×32 with a maximum partition depth of 3. Alternatively, quadtree partitioning can be performed in the range of 8×8 to 128×128, while binary tree partitioning and ternary tree partitioning can be performed in the range of 4×4 to 128×128 with a maximum partition depth of 3. The former can be set for an I picture type (e.g., a slice), and the latter can be set for a P or B picture type.

[0170] As described in the above examples, the partitioning settings such as the maximum coding block size, the minimum coding block size, and the maximum partitioning depth can be universally or independently set according to the partitioning method and the encoding / decoding settings described above.

[0171] When multiple partitioning methods are supported, partitioning will be performed within the block support range of each partitioning method. When the block support ranges of the partitioning methods overlap, the priority information of the partitioning methods can be included. For example, quadtree partitioning can be performed before binary tree partitioning.

[0172] Alternatively, when the split support ranges overlap, split selection information can be generated. For example, selection information related to the splitting method to be performed in binary tree splitting and ternary tree splitting can be generated.

[0173] Furthermore, when multiple partitioning schemes are supported, the execution of a later partitioning scheme can be determined based on the results of an earlier partitioning scheme. For example, if the result of an earlier partitioning scheme (in this example, quadtree partitioning) indicates that a partitioning scheme should be executed, a later partitioning scheme (in this example, binary or ternary tree partitioning) can be omitted and the subsequent partitioning scheme can be performed after the sub-coding blocks generated by the earlier partitioning scheme are set as coding blocks again.

[0174] Alternatively, when the result of an earlier segmentation indicates that segmentation is not to be performed, segmentation can be performed based on the result of a later segmentation. In this case, when the result of the later segmentation (in this example, binary tree segmentation or ternary tree segmentation) indicates that segmentation is to be performed, segmentation can be continued after the segmented sub-coding blocks are set as coding blocks again, and when the result of the later segmentation indicates that segmentation is not to be performed, segmentation can be skipped. In this case, when the result of the later segmentation indicates that segmentation is to be performed and multiple segmentation methods are still supported when the segmented sub-coding blocks are set as coding blocks again (for example, when the block support ranges of the different segmentation methods overlap), the earlier segmentation can be skipped and only the later segmentation can be performed. That is, when multiple segmentation methods are supported, if the result of the earlier segmentation indicates that segmentation is not to be performed, the earlier segmentation can be skipped.

[0175] For example, when an M×N coding block can be subjected to quadtree splitting and binary tree splitting, the quadtree splitting flag can be first confirmed, and when the splitting flag is 1, it is split into 4 sub-coding blocks of (M>>1)×(N>>1) size, and then the splitting (quadtree splitting or binary tree splitting) can be performed after the sub-coding blocks are set as coding blocks again. When the splitting flag is 0, the binary tree splitting flag can be confirmed, and when the corresponding flag is 1, it is split into 2 sub-coding blocks of (M>>1)×N or M×(N>>1) size, and then the splitting (binary tree splitting) can be performed after the sub-coding blocks are set as coding blocks again. When the splitting flag is 0, the splitting process is terminated and encoding is performed.

[0176] The above examples illustrate the use of multiple partitioning schemes, but the present invention is not limited thereto and supports combinations of multiple partitioning schemes. For example, partitioning schemes such as quadtree, binary tree, ternary tree, quadtree + binary tree, quadtree + binary tree + ternary tree can be used. In this case, information regarding support for additional partitioning schemes can be implicitly determined or explicitly included in units such as sequences, pictures, sub-pictures, slices, and parallel blocks.

[0177] In the above examples, information related to segmentation, such as the size of the coding block, the supported range of coding blocks, and the maximum segmentation depth, can be included in units such as sequences, pictures, sub-pictures, slices, and parallel blocks, or can be implicitly determined. In other words, the range of allowable blocks can be determined based on the maximum coding block size, the range of supported blocks, the maximum segmentation depth, and the like.

[0178] The coding block obtained by performing the division by the above-described process can be set to the maximum size of the intra prediction or the inter prediction. That is, in order to perform the intra prediction or the inter prediction, the coding block after the block division can be the division start size of the prediction block. For example, when the coding block is 2M x 2N, the size of the prediction block can be the same size or a relatively smaller size of 2M x 2N, M x N. Or, it can be the size of 2M x 2N, 2M x N, M x 2N, M x N. Or, it can be the same size as the coding block, i.e., the size of 2M x 2N. At this time, the same size of the coding block and the prediction block can mean that the division of the prediction block is not performed, and the prediction is performed using the size obtained by the division of the coding block. That is, it means that the division information for the prediction block is not generated. The above-described setting can also be applied to the transform block, and the transform can be performed in the unit of the divided coding block.

[0179] By the above-described coding / decoding setting, various configurations can be implemented. For example, at least one prediction block and at least one transform block can be obtained based on the coding block (after determining the coding block). Or, one prediction block having the same size as the coding block and at least one transform block can be obtained based on the coding block. Or, one prediction block having the same size as the coding block and one transform block can be obtained. In the above-described example of obtaining at least one block, the division information of each block can be generated, and when one block is obtained, the division information of each block will not be generated.

[0180] The blocks of various sizes obtained by the above-described results, in the form of a square or a rectangle, can be blocks used in the intra prediction, the inter prediction, blocks used when transforming and quantizing the residual components, and blocks used in the filtering process.

[0181] The division unit obtained by dividing the image by using the image division unit can perform independent coding / decoding or dependent coding / decoding according to the coding / decoding setting.

[0182] The independent coding / decoding can mean that when coding / decoding is performed on a part of the division unit (or region), the data of the other unit cannot be used as a reference. Specifically, the information used or generated in the process of performing the texture coding and the entropy coding on a part of the unit {for example, pixel values or coding / decoding information (intra prediction-related information, inter prediction-related information, and entropy coding / decoding-related information, etc.)} will not be mutually referenced and independently coded, and the same applies to the process of performing the texture decoding and the entropy decoding on a part of the unit in the decoder, and the analysis information and the reconstructed information of the other unit will not be mutually referenced.

[0183] Furthermore, dependent encoding / decoding can mean that when encoding / decoding a portion of a segmented unit, data from other units can be used as a reference. Specifically, information used or generated during texture encoding and entropy encoding of a portion of the unit can be encoded in a dependent manner by referencing each other. Similarly, during texture decoding and entropy decoding of a portion of the unit in the decoder, parsed information and reconstruction information from other units can be cross-referenced.

[0184] Generally, the division units mentioned above (such as sub-images, parallel blocks, strips, etc.) can adopt independent encoding / decoding settings. That is, non-reference settings can be adopted for the purpose of parallelization. In addition, non-reference settings can be adopted for the purpose of improving encoding / decoding performance. For example, when a 360-degree image is divided into multiple surfaces in a 3D space and configured in a 2D space, the correlation with adjacent surfaces (such as image continuity) may decrease depending on the surface configuration setting. That is, because the need for mutual reference is low when there is no correlation between surfaces, independent encoding / decoding settings can be adopted.

[0185] Furthermore, it is possible to employ settings that allow for reference between segmentation units to improve encoding / decoding performance. For example, even when segmenting a 360-degree image into surface units, there may be a high correlation with adjacent surfaces depending on the surface configuration. In such cases, it is possible to employ dependent encoding / decoding settings.

[0186] Furthermore, in the present invention, independent or dependent encoding / decoding can be applied not only to spatial regions but also to temporal regions. Specifically, independent or dependent encoding / decoding can be performed not only on other segments that exist at the same time as the current segment, but also on segments that exist at different times than the current segment (in this example, even if a segment exists at the same position within an image at a different time than the current segment, it is assumed to be another segment).

[0187] For example, when a bit stream A containing data encoding a 360-degree image into higher image quality and a bit stream B containing data encoded into normal image quality are transmitted simultaneously, the decoder can parse and decode the bit stream A transmitted with higher image quality in an area corresponding to the area of ​​interest (such as the area where the user's line of sight is focused <viewport> or the area desired to be displayed, etc.), and parse and decode the bit stream B transmitted with normal image quality outside the area of ​​interest.

[0188] Specifically, in the case where the video is divided into a plurality of units (e.g., sub-pictures, tiles, slices, surfaces, etc., in this example, it is assumed that the surfaces are data-processed in the same manner as the tiles or sub-pictures), the data (bitstream A) of the divided units included in the region of interest (or the divided units as long as they overlap with the viewport by one pixel) and the data (bitstream B) of the divided units included outside the region of interest can be decoded.

[0189] Alternatively, a bitstream in which the receipt of encoding the entire video is recorded can be transmitted, and the region of interest can be parsed and decoded from the bitstream at the decoder. Specifically, only the data of the divided units included in the region of interest can be decoded.

[0190] In other words, the entire or part of the video can be obtained by generating a bitstream divided into more than one quality at the encoder and decoding only a specific bitstream at the decoder, and the entire or part of the video can be obtained by selectively decoding each bitstream in each video part. In the above example, the case of a 360-degree video is described, but this is a description that can be applied to general videos.

[0191] When encoding / decoding is performed according to the above-described example, because it is not known that the data will be reconstructed at the decoder (in this example, the decoder does not know the position of the region of interest, and is a case of random access with respect to the region of interest), it is necessary to confirm and perform encoding / decoding with respect to the reference setting, etc. in the time region in addition to the spatial region.

[0192] For example, when the decoder determines which type of decoding is performed with respect to a single divided unit, the current divided unit can perform independent encoding in the spatial region and limited dependent encoding in the time region (e.g., only allowing reference to the divided unit at the same position of other time corresponding to the current divided unit and prohibiting reference to other divided units except for this, because generally, there is no limitation in the time region, this is a comparison with unrestricted dependent encoding).

[0193] Alternatively, when the decoder determines which type of decoding to perform based on multiple segmentation units (the multiple segmentation units can be obtained by bundling horizontally adjacent segmentation units or bundling vertically adjacent segmentation units, or by bundling horizontally and vertically adjacent segmentation units), for example, in this case, multiple units are decoded as long as any one of the segmentation units is included in the area of ​​interest), the current segmentation unit can perform independence or dependency decoding in the spatial region and perform limited dependency encoding in the temporal region (for example, in addition to allowing reference to the segmentation units at the same position at other times corresponding to the current segmentation unit, reference to a portion of the segmentation units other than the current segmentation unit is also allowed).

[0194] In the present invention, a surface is a segmentation unit whose configuration and shape usually change according to the projection format and has no independent encoding / decoding settings. Although it has different characteristics from the other segmentation units described above, it can also be regarded as a unit obtained in the image segmentation unit in terms of being able to divide the image into multiple areas (and using a quadrilateral shape, etc.).

[0195] As described above, independent encoding / decoding can be performed on each segmentation unit in the spatial region for the purpose of parallelization, etc. However, since independent encoding / decoding cannot refer to other segmentation units, it will lead to the problem of reduced encoding / decoding efficiency. Therefore, as a step before performing encoding / decoding, the segmentation unit for performing independent encoding / decoding can be expanded by utilizing (or appending) the data of adjacent segmentation units. Among them, since the segmentation unit to which the data in the adjacent segmentation unit is appended has more data to refer to, its encoding / decoding efficiency will also be improved. At this time, since the expanded segmentation unit can refer to the data in the adjacent segmentation unit when encoding / decoding, it can be regarded as dependent encoding / decoding.

[0196] The information related to the reference settings between the above-mentioned division units can be included in the bitstream and transmitted to the decoder in units such as video, sequence, image, sub-image, slice, parallel block, etc., and the setting information transmitted from the encoder can be reconstructed in the decoder by parsing the units at the same level. In addition, the relevant information can be transmitted to the bitstream in the form of supplementary enhancement information (SEI) or metadata and used after parsing. In addition, it is also possible to use the definition agreed in advance in the encoder / decoder to perform encoding / decoding according to the reference settings without transmitting the above-mentioned information.

[0197] Figure 7 This is an example diagram of dividing an image into multiple parallel blocks. Figures 8a to 8i Yes Figure 7 FIG. 1 is a diagram showing a first example of setting additional areas for each parallel block shown in FIG. Figures 9a to 9i Yes Figure 7 FIG2 is a second example diagram showing the setting of additional areas for each parallel block shown in FIG2.

[0198] While the image segmentation unit divides an image into two or more segments (or regions) and performs independent encoding / decoding on each segment, while offering advantages such as parallel processing, it can also lead to reduced encoding performance due to the reduced reference data available to each segment. To address this issue, it is possible to implement encoding / decoding settings that prioritize dependencies between segments (in this example, parallel block particles are used; similar or similar settings can also be applied to other units).

[0199] Independent encoding / decoding is typically performed between segmented units without reference to each other. Therefore, pre-processing or post-processing can be performed to implement dependent encoding / decoding. For example, before encoding / decoding, an extension region can be formed on the outline of each segmented unit and filled with data from other segmented units that require reference.

[0200] Although the method described above is no different from the method of performing independent encoding / decoding except that encoding / decoding is performed after each segmentation unit is expanded, because the existing segmentation units will obtain the data required for reference from other segmentation units in advance and refer to them, it can be understood as an example of dependent encoding / decoding.

[0201] Furthermore, after encoding / decoding, filtering can be applied using multiple segmentation unit data based on the boundaries between the segmentation units. That is, when filtering is applied, it is dependent on the use of other segmentation unit data, but when filtering is not applied, it can be independent.

[0202] The examples described below will focus on the case where dependency encoding / decoding is performed by executing the encoding / decoding pre-processing (in this example, extension). Furthermore, in the present invention, the boundaries between identical segmentation units can be referred to as internal boundaries, and the image outline can be referred to as external boundaries.

[0203] In one embodiment of the present invention, an additional region associated with the current parallel tile can be set. Specifically, the additional region can be set based on at least one parallel tile (in this example, including a case where one image is composed of one parallel tile, i.e., including a case where the image is not divided into two or more division units. To be precise, although a division unit means a unit divided into two or more, it is assumed that even if the image is not divided, it is still recognized as a single division unit).

[0204] For example, an additional region can be set in at least one of the following directions: up, down, left, or right. The additional region can be filled with any value. Furthermore, the additional region can be filled with a portion of the data in the current parallel tile, i.e., by using pixels outside the current parallel tile or by copying pixels within the current parallel tile.

[0205] Furthermore, the additional area can be filled with image data from parallel blocks other than the current parallel block. Specifically, the additional area can be filled with image data from parallel blocks adjacent to the current parallel block, that is, by copying image data from parallel blocks adjacent to the current parallel block in a specific direction of up / down / left / right.

[0206] At this time, the size (length) of the acquired image data may be the same value in each direction or may be independent values, and this may be determined according to the encoding / decoding settings.

[0207] For example, in Figure 6a In the example, the expansion can be made in all or part of the boundaries of b0 to b8. In addition, the expansion can be made in all the boundaries of the division unit or in accordance with the boundary direction. i (i is the index of each direction.) m or mi can be applied to all segmentation units in the image, or can be set independently for each segmentation unit.

[0208] At this time, setting information related to the additional area can be generated. At this time, the setting information related to the additional area can be whether the additional area is supported, whether the additional area of ​​each segmentation unit is supported, the form of the additional area on the overall image (for example, it is determined based on which direction of expansion among the up / down / left / right of the segmentation unit, in this example, it is setting information that is commonly applicable to all segmentation units in the image), the form of the additional area on each segmentation unit (in this example, it is setting information that is applicable to individual segmentation units in the image), the size of the additional area on the overall image (for example, after the form of the additional area is determined, it indicates the degree of expansion in the direction of expansion, in this example, it is setting information that is commonly applicable to all segmentation units in the image), the size of the additional area on each segmentation unit (in this example, it is setting information that is independently applicable to individual segmentation units in the image), the method of filling the additional area on the overall image, the method of filling the additional area on each segmentation unit, etc.

[0209] The above-mentioned settings related to the additional area can be determined proportionally to the color space or independently. The setting information related to the additional area can be generated for the luminance component, while the setting for the additional area for the chrominance component can be implicitly determined based on the color space. Alternatively, the setting information related to the additional area can also be generated for the chrominance component.

[0210] For example, when the additional area size for the luma component is m, the additional area size for the chroma component can be determined as m / 2 based on the color format (4:2:0 in this example). As another example, when the additional area size for the luma component is m and the chroma components are set to be independent, size information for the additional areas of the chroma components can be generated (in this example, n; n can be used in common or n1, n2, n3, etc. can be used depending on the direction or expansion area). As another example, a method for filling the additional area for the luma component can be generated, while the method for filling the additional area for the chroma component can use the method used for the luma component or generate related information.

[0211] The information related to the additional region settings can be included in the bitstream and transmitted in units such as video, sequence, picture, sub-picture, and slice. During decoding, the relevant information can be parsed and reconstructed from these units. The following embodiments will assume that additional regions are supported.

[0212] See Figure 7 , it can be confirmed that an image is divided into parallel blocks marked from 0 to 8. Figure 7 The result of setting the additional regions of each parallel block shown in FIG. 1 is as follows: Figures 8a to 8i shown.

[0213] exist Figure 7 as well as Figure 8a , parallel block No. 0 (with a size of T0_W×TO_H) can be expanded by adding an area of ​​E0_R to the right and an area of ​​E0_D to the bottom. At this time, the additional area can be obtained from the adjacent parallel block. Specifically, the right extension area can be obtained from parallel block No. 1, and the bottom extension area can be obtained from parallel block No. 3. In addition, parallel block No. 0 can set an additional area using the parallel block adjacent to the lower right side (parallel block No. 4). That is, the additional area can be set in the direction of the remaining internal boundaries (or boundaries between the same segmentation units) except the external boundaries (or image boundaries) of the parallel block.

[0214] exist Figure 7 as well as Figure 8e In the example, because parallel tile No. 4 (size T4_W×T4_H) has no external boundaries, it can be expanded by adding regions to the left, right, top, and bottom. In this case, the left extension region can be obtained from parallel tile No. 3, the right extension region can be obtained from parallel tile No. 5, the top extension region can be obtained from parallel tile No. 1, and the bottom extension region can be obtained from parallel tile No. 7. In addition, for extension region No. 4, additional regions can be set to the upper left, lower left, upper right, and lower right. In this case, the upper left extension region can be obtained from parallel tile No. 0, the lower left extension region can be obtained from parallel tile No. 6, the upper right extension region can be obtained from parallel tile No. 2, and the lower right extension region can be obtained from parallel tile No. 8.

[0215] In Figure 8 , because the L2 tile is adjacent to the parallel tile boundary, in principle, there is no data that can be referenced from the left, upper-left, or lower-left tiles. However, by applying an embodiment of the present invention and setting an additional region for the second parallel tile, encoding / decoding of the L2 tile can be performed by referencing the additional region. That is, the L2 tile can reference data from the left and upper-left tiles as an additional region (which can be the region obtained from the first parallel tile) and can reference data from the lower-left tile as an additional region (which can be the region obtained from the fourth parallel tile).

[0216] The data included in the additional area through the above-mentioned embodiment can be included in the current parallel block for encoding / decoding. In this case, because the data in the additional area is located on the boundary of the parallel block (in this example, it refers to the parallel block that is updated or expanded due to the additional area), the encoding performance may also be reduced during the encoding process due to the lack of data for reference. However, because this is only a part added to provide a reference to the original parallel block boundary area, it can be understood as a temporary storage form for improving encoding performance. That is, because it can help improve the image quality performance of the final output image and is the area that is ultimately removed, the reduction in encoding performance of the corresponding area does not cause any problems. This can be applied to the embodiments described later for similar or identical purposes.

[0217] In addition, see Figures 9a to 9i , it can be confirmed that the 360-degree image is changed into a 2D image through the surface configuration (or reconfiguration) process according to the projection format, and the 2D image is divided into parallel blocks (which can also be surfaces). In this case, because the 2D image is composed of a single surface when the 360-degree image adopts equirectangular projection, this can be an example of dividing a single surface into parallel blocks. In addition, for the convenience of explanation, the parallel block division of the 2D image is compared with Figure 7 The parallel block division shown in FIG is the same as the premise.

[0218] The divided parallel blocks can be divided into parallel blocks consisting only of internal boundaries and parallel blocks containing at least one external boundary, which can be divided into parallel blocks consisting only of internal boundaries and at least one external boundary. Figures 8a to 8i The method shown sets additional areas for each parallel block. However, the 360-degree images converted into 2D images may not have continuity in the actual image even if they are adjacent to each other in the 2D image, and may have continuity in the actual image even if they are not adjacent (see the Figures 5a to 5c Therefore, even if part of the boundary of the parallel block is an external boundary, there can be an area in the image that is continuous with the external boundary area of ​​the parallel block. Figure 9b Although the top of the first parallel block is the outer boundary of the image, since there can be areas with actual image continuity within the same image, an additional area can be set at the top of the first parallel block. Figures 8a to 8i Different, in Figures 9a to 9i It is also possible to set additional areas in all or part of the outer boundary direction of the parallel block.

[0219] See Figure 9e , parallel block No. 4 is a parallel block that only includes the internal boundary in the parallel block boundary (parallel block No. 4 in this example). Therefore, the additional area of ​​parallel block No. 4 can be set in the upper, lower, left, and right directions, as well as the upper left, lower left, upper right, and lower right directions. Among them, the left extension area can be the image data obtained from parallel block No. 3, the right extension area can be the image data obtained from parallel block No. 5, the upper extension area can be the image data obtained from parallel block No. 1, the lower extension area can be the image data obtained from parallel block No. 7, the upper left extension area can be the image data obtained from parallel block No. 0, the lower left extension area can be the image data obtained from parallel block No. 6, the upper right extension area can be the image data obtained from parallel block No. 2, and the lower right extension area can be the image data obtained from parallel block No. 9.

[0220] See Figure 9a , the parallel block No. 0 is a parallel block that contains at least one external boundary (left side, upper direction). Therefore, in addition to the spatially adjacent right side, bottom side, and lower right direction, the parallel block No. 0 can also include an additional area extending in the external boundary direction (left side, upper side, upper left direction). Among them, the additional area in the spatially adjacent right side, bottom side, and lower right direction can be set using the data of the adjacent parallel blocks, but the additional area in the external boundary direction cannot be set in this way. At this time, the additional area in the external boundary direction can be set using data that is not spatially consistent within the image but has continuity in the actual image. For example, when the projection format of a 360-degree image is equirectangular, the left and right edges of the image are continuous in the actual image, and the top and bottom edges of the image are continuous in the actual image, the left edge of parallel block 0 is continuous with the right edge of parallel block 2, and the top edge of parallel block 0 is continuous with the bottom edge of parallel block 6. Therefore, in parallel block 0, the left extension area can be obtained from parallel block 2, the right extension area can be obtained from parallel block 1, the top extension area can be obtained from parallel block 6, and the bottom extension area can be obtained from parallel block 3. Furthermore, in parallel block 1, the top left extension area can be obtained from parallel block 8, the bottom left extension area can be obtained from parallel block 5, the top right extension area can be obtained from parallel block 7, and the bottom right extension area can be obtained from parallel block 4.

[0221] because Figure 9a The L0 block in is located at the parallel block boundary, so data that can be referenced from the left, upper left, lower left, upper, or upper right blocks (similar to the case with U0) may not exist. In this case, even if they are not spatially adjacent in the 2D image, blocks that are actually continuous can still exist within the 2D image (or image). Therefore, as described above, when the projection format of the 360-degree image is equirectangular, the left and right image boundaries are actually continuous, and the upper and lower image boundaries are actually continuous, the left and lower left blocks of the L0 block can be obtained from parallel block No. 2, the upper left block of the L0 block can be obtained from parallel block No. 8, and the upper and upper right blocks of the L0 block can be obtained from parallel block No. 6.

[0222] Table 1 below is a pseudo code for acquiring data corresponding to the additional area from other areas having continuity.

[0223]

Table 1

[0224] i_pos'=overlap(i_pos,minI,maxI)

[0225] overlap(A,B,C)

[0226] {

[0227] if(A <B)output=(A+C-B+1)%(C-B+1)

[0228] else if(A>C)output=A%(C-B+1)

[0229] else output=A

[0230] }

[0231] Referring to the pseudocode in Table 1, the variable i_pos (corresponding to variable A) of the overlap function is the input pixel position, i_pos' is the output pixel position, minI (corresponding to variable B) is the minimum value of the pixel position range, maxI (corresponding to variable C) is the maximum value of the pixel position range, and i is the position component (horizontal, vertical, etc. in this example). In this example, minI can be 0 and maxI can be Pic_width (horizontal width of the image)-1 or Pic_height (vertical width of the image)-1.

[0232] For example, assuming that the vertical width of an image (general image) ranges from 0 to 47 and the image is as follows Figure 7 The division is performed in the manner shown. When an additional region of m size needs to be set to the lower side of the parallel block 4 and filled with the data of the upper end of the parallel block 7, it is possible to confirm from which position the data needs to be acquired by the above.

[0233] When the vertical length range of the parallel block 4 is 16 to 30 and an additional region of 4 size needs to be set to the lower side, it is possible to fill the data of the positions of 31, 32, 33, 34 corresponding thereto into the additional region of the parallel block 4. At this time, because min and max in the above formula are 0 and 47, respectively, the output values of 31 to 34 will be their own values, i.e., 31 to 34. That is, the data that needs to be filled into the additional region is the data of the positions of 31 to 34.

[0234] Alternatively, assuming that the horizontal length range of an image (360-degree image, equirectangular projection, images having continuity at both ends) is 0 to 95 and the image is divided as shown in FIG. 6, when an additional region of m size needs to be set to the left side of the parallel block 3 and filled with the data of the right side of the parallel block 5, it is possible to confirm from which position the data needs to be acquired by the above. Figure 7

[0235] When the vertical length range of the parallel block 3 is 0 to 31 and an additional region of 4 size needs to be set to the left side, it is possible to fill the data of the positions of -4, -3, -2, -1 corresponding thereto into the additional region of the parallel block 3. Because the above positions do not exist within the horizontal length range of the image, it is possible to confirm from which position the data needs to be acquired by the above known calculation. At this time, because min and max in the above formula are 0 and 95, respectively, the output values of -4 to -1 will be 92 to 95. That is, the data that needs to be filled into the additional region is the data of the positions of 92 to 95.

[0236] Specifically, when the region of m size is data between 360 degrees and 380 degrees (assuming that the range of pixel value positions is 0 degrees to 360 degrees in the present example), it is possible to understand that it is similar to the case of acquiring data from the region between 0 degrees and 20 degrees by adjusting it to the range inside the image. That is, it is possible to acquire it based on the range of pixel value positions between 0 and Pic_width-1.

[0237] In other words, in order to acquire the data of the additional region, it is possible to confirm the position of the data that needs to be acquired through the overlapping process.

[0238] ​The above example describes the acquisition of a single surface from a 360-degree image, assuming that spatially adjacent regions within the image are continuous (except when the image boundaries are continuous at both ends). However, depending on the projection format (e.g., cube map projection), when two or more surfaces are included and each surface undergoes a configuration or reconfiguration process, there may be instances where spatially adjacent regions within the image lack continuity. In such cases, the surface configuration or reconfiguration information can be used to confirm location data that is continuous in the actual image and generate additional regions.

[0239] Table 2 below is a pseudo code for generating an additional area related to a specific division unit using the internal data of the specific division unit.

[0240]

Table 2

[0241] i_pos'=clip(i_pos,minI,maxI)

[0242] clip(A,B,C)

[0243] {

[0244] if(A <B)output=B

[0245] else if(A>C)output=C

[0246] else output=A

[0247] }

[0248] Since the meanings of the variables in Table 2 are the same as those in Table 1, their detailed descriptions will be omitted. However, in this example, minI can be the left or top coordinate of a specific segmentation unit, and maxI can be the right or bottom coordinate of each unit.

[0249] For example, when the image is Figure 7 If the segmentation is performed in the manner shown, and the horizontal length of parallel block 2 ranges from 32 to 47, and an additional region of size m is required to be set to the right of parallel block 2, the data at positions 48, 49, 50, and 51 corresponding to these positions can be filled with the data at position 47 (corresponding to the interior of parallel block 2) output by the above formula. In other words, according to Table 2, the additional region associated with a specific segmentation unit can be generated by copying the outer pixels of the corresponding segmentation unit.

[0250] In other words, in order to obtain data in the additional area, the position to be obtained can be confirmed through a clipping process.

[0251] The detailed structure of Table 1 or Table 2 is not fixed and can be changed. For example, the superposition method of the 360-degree image can be changed by taking into account the layout (or re-layout) of the surface and the coordinate system characteristics between the surfaces.

[0252] Figure 10 This is a diagram illustrating an example of applying an additional region generated in an embodiment of the present invention to the encoding / decoding process of another region.

[0253] Furthermore, because the additional regions to which one embodiment of the present invention is applied are generated using image data from other regions, they can represent duplicate image data. Therefore, to prevent unnecessary duplicate data, the additional regions can be removed after encoding / decoding. However, before removing the additional regions, it is possible to consider removing them after encoding / decoding them.

[0254] See Figure 10 , it can be confirmed that the additional region B of the division unit J is generated using the region A of the division unit I. In this case, before the generated region B is removed, the region B can be applied to the encoding / decoding (specifically, the reconstruction or correction process) of the region A included in the division unit I.

[0255] Specifically, assuming that the segmentation unit I and segmentation unit J are Figure 7 Under the premise of parallel block No. 0 and parallel block No. 1 in , the rightmost part of segmentation unit I and the left part of segmentation unit J have image continuity with each other. Among them, the additional area B can be used in encoding / decoding of A after being used as a reference in the image encoding / decoding of segmentation unit J. In particular, although A and area B obtain data from area A when generating the additional area, they may be reconstructed using some different values ​​(including quantization errors) during the encoding / decoding process. Therefore, when reconstructing segmentation unit I, the part corresponding to area A can be reconstructed using the reconstructed image data of area A and the image data of area B. For example, a part of area C of segmentation unit I can be replaced by the average or weighted value of area A and area B. This is because there are data of two or more identical areas, so the reconstructed image (area C, A of segmentation unit I is replaced by C) can be obtained using the data of two areas (the process at this time is named Rec_Process in the accompanying figure).

[0256] Further, a part of the region C included in the division unit I can be replaced using the region A and the region B according to which division unit it is closer to. Specifically, because the image data included in a certain range (for example, M pixel intervals) on the left side in the region C is closer to the division unit I, it can be reconstructed using (or copying) the data in the region A, and because the image data included in a certain range (for example, N pixel intervals) on the right side in the region C is closer to the division unit J, it can be reconstructed using (or copying) the data in the region B. This is expressed as a formula as shown in the following Formula 1.

[0257] [Formula 1]

[0258] C(x, y) = A(x, y), (x, y) e M

[0259] B(x, y), (x, y) e N

[0260] Further, a part of the region C included in the division unit I can be replaced using the region A and the region B according to which division unit it is closer to. Specifically, because the image data included in a certain range (for example, M pixel intervals) on the left side in the region C is closer to the division unit I, it can be reconstructed using (or copying) the data in the region A, and because the image data included in a certain range (for example, N pixel intervals) on the right side in the region C is closer to the division unit J, it can be reconstructed using (or copying) the data in the region B. This is expressed as a formula as shown in the following Formula 1.

[0261] As a formula for setting adaptive weight values to the region A and the region B, the following Formula 2 can be derived.

[0262] [Formula 2]

[0263] C(x, y) = A(x, y) x w + B(x, y) x (1 - w)

[0264] w = f(x, y, k)

[0265] Referring to Formula 2, w is a weight value assigned to the pixels of the A region and the B region as (x, y), and at this time, as an average of the weight values of the A region and the B region, the pixels of the A region are multiplied by the weight value w and the pixels of the B region are multiplied by 1 - w. However, in addition to the average of the weight values, different weight values can be assigned to the region A and the region B, respectively.

[0266] After the additional area is used as described above, the additional area B can be removed during the resizing process of the segmentation unit J and stored in the memory (decoded picture buffer, DPB) (in this example, it is assumed that the process of setting the additional area is resizing). <sizing>The above process can be derived through some of the above embodiments <for example, through the process of additional area mark confirmation, subsequent size information confirmation, subsequent filling method confirmation, etc.>, assuming that the process of enlarging is performed during the resizing process, and conversely, the process of shrinking is performed during the resizing process <which can be derived through the reverse process of the above process>).

[0267] Furthermore, (specifically immediately after the encoding / decoding of the corresponding image is completed) it is also possible to store the image directly in the memory without resizing and then in the output step (assuming in this example that the image is displayed) <display>This can be applied to all or part of the segmented units included in the corresponding image.

[0268] The above-mentioned relevant setting information can be processed implicitly or explicitly according to the encoding / decoding settings. When the implicit method is adopted (specifically, it is based on the characteristics, type, format, etc. of the image or according to other encoding / decoding settings <in this example, the settings related to the additional area>), it can be determined without generating relevant syntax elements. When the explicit method is adopted, the settings related to the removal of the additional area can be adjusted by generating relevant syntax elements. The units related thereto can include videos, sequences, images, sub-images, slices, parallel blocks, etc.

[0269] In addition, the existing encoding method based on partitioning units may include: 1) a step of partitioning an image into one or more parallel blocks (or collectively referred to as partitioning units) and generating partitioning information; 2) a step of performing encoding according to the partitioned parallel block units; 3) a step of performing filtering with information indicating whether loop filtering is allowed at the parallel block boundary; and, 4) a step of storing the filtered parallel blocks in a memory.

[0270] In addition, the existing decoding method based on segmentation units can include: 1) a step of segmenting an image into one or more parallel blocks based on parallel block segmentation information; 2) a step of performing decoding according to the segmented parallel block units; 3) a step of performing filtering using information indicating whether loop filtering is allowed at the parallel block boundary; and, 4) a step of storing the filtered parallel blocks in a memory.

[0271] The third step in the encoding / decoding method is a post-processing step of encoding / decoding, which can be dependent encoding / decoding when filtering is performed, and can be independent encoding / decoding when filtering is not performed.

[0272] A method for encoding a segmentation unit according to an embodiment of the present invention may include: 1) dividing an image into one or more parallel blocks and generating segmentation information; 2) setting an additional region for at least one of the divided parallel block units and filling the additional region with adjacent parallel block units; 3) encoding the parallel block unit including the additional region; 4) removing the additional region of the parallel block unit and filtering based on information indicating whether loop filtering of the parallel block boundary is permitted; and 5) storing the filtered parallel block in a memory.

[0273] In addition, a decoding method for a segmentation unit applicable to one embodiment of the present invention may include: 1) a step of segmenting an image into one or more parallel blocks based on parallel block segmentation information; 2) a step of setting an additional area for the segmented parallel block units and filling the additional area using decoded information, pre-set information, or pre-reconstructed other (adjacent) parallel block units; 3) a step of encoding the parallel block unit including the additional area using decoded information received from an encoding device; 4) a step of removing the additional area of ​​the parallel block unit and performing filtering based on information indicating whether loop filtering of the parallel block boundary is allowed; and 5) a step of storing the filtered parallel block in a memory.

[0274] In the above-described encoding / decoding method for a segmented unit according to one embodiment of the present invention, the second step can be a pre-encoding / decoding process (dependent encoding / decoding when an additional region is set, otherwise independent encoding / decoding). Furthermore, the fourth step can be a post-encoding / decoding process (dependent when filtering is performed, otherwise independent). In this example, the additional region is used during the encoding / decoding process, and the size is adjusted to the initial size of the parallel block before being stored in the memory.

[0275] First, the encoder divides the image into multiple parallel blocks. Based on implicit or explicit settings, additional regions are set for each parallel block and relevant data is retrieved from adjacent regions. Next, encoding is performed on the updated parallel block units that include the original parallel block and the additional regions. After encoding is complete, the additional regions are removed and filtering is performed according to the applicable loop filtering settings.

[0276] At this time, different filtering settings can be used depending on the filling method and removal method of the additional area. For example, when simply removing data, the aforementioned loop filter application settings can be followed, while when removing data using overlapping areas, filtering can be omitted or different filtering settings can be followed. In other words, because the overlapping data can be used to significantly reduce distortion in the parallel block boundary area, it is possible to not perform filtering regardless of whether loop filtering is applicable to the parallel block convenience unit, or to apply different filtering settings within the parallel block while complying with the aforementioned filtering application (for example, applying a filter with weaker filtering strength at the parallel block boundary, etc.). After the above process, the data is stored in the memory.

[0277] In the decoder, the image is first divided into a plurality of parallel blocks based on the parallel block division information transmitted from the encoder. Next, information related to the additional region is explicitly or implicitly confirmed, and the updated parallel block encoding information transmitted from the encoder is parsed after the additional region is set. Next, decoding is performed in the updated parallel block unit. After decoding is completed, the additional region is removed and filtering is performed according to the same loop filtering application setting as the encoder. Details related thereto have been described in the encoder section, so detailed description thereof will be omitted here. After the above process, storage into the memory is performed.

[0278] Further, a case where the additional region using the division unit in the encoding / decoding process is not removed but directly stored into the memory can also be considered. For example, in the case of a 360-degree image or the like, there can be a problem where the accuracy of prediction decreases in a part of the prediction process (e.g., inter-picture prediction) according to the surface configuration setting or the like (e.g., it is difficult to accurately search at a position where the surface configuration is not continuous when motion search and compensation are performed). Therefore, the additional region can be stored into the memory and used in the prediction process in order to improve the prediction accuracy. When used in inter-picture prediction, the additional region (or the image including the additional region) can be used as a reference image for performing inter-picture prediction.

[0279] In the encoding method when the additional region is stored, it can include: 1) a step of dividing the image into one or more parallel blocks and generating division information; 2) a step of setting an additional region for at least one of the divided parallel block units and filling the additional region with adjacent parallel block units; 3) a step of performing encoding on the parallel block unit including the additional region; 4) a step of storing the additional region of the parallel block unit (at this time, application of loop filtering can be omitted); and 5) a step of storing the encoded parallel block into the memory.

[0280] In the decoding method when the additional region is stored, it can include: 1) a step of dividing the image into one or more parallel blocks based on parallel block division information; 2) a step of setting an additional region for the divided parallel block unit and filling the additional region with decoding information, pre-set information, or other (adjacent) parallel block units that are previously reconstructed; 3) a step of performing encoding on the parallel block unit including the additional region using the decoding information received from the encoding device; 4) a step of storing the additional region of the parallel block unit (at this time, loop filtering can be omitted); and 5) a step of storing the decoded parallel block into the memory.

[0281] When storing additional regions, the encoder first divides the image into multiple parallel blocks. Based on implicit or explicit settings, additional regions are set for the parallel blocks and relevant data is retrieved from pre-set areas. These pre-set areas are other areas related to the surface configuration of the 360-degree image and can be adjacent or non-adjacent to the current parallel block. Encoding is then performed on the updated parallel block units. Because the additional regions are stored after decoding, filtering is not performed regardless of the state of the loop filtering application settings. This is because the boundaries of the updated parallel blocks do not share the actual parallel block boundaries due to the additional regions. After this process, the data is stored in memory.

[0282] When storing the additional region, the decoder first verifies the parallel tile division information transmitted from the encoder and uses this information to divide the image into multiple parallel tiles. Next, it verifies information related to the additional region and, after setting the additional region, parses the updated parallel tile encoding information transmitted from the encoder. Decoding is then performed on the updated parallel tile units. After decoding, the additional region is directly stored in memory without loop filtering.

[0283] Next, the encoding / decoding method of the segmentation unit according to one embodiment of the present invention will be described with reference to the accompanying drawings.

[0284] Figures 11 to 12 This is a flowchart for explaining a method for encoding / decoding a segmentation unit according to an embodiment of the present invention. Specifically, as an example of generating an additional region in each segmentation unit and performing encoding / decoding, in Figure 11 The encoding method including the additional area is illustrated in Figure 12 The decoding method for removing the additional area is illustrated in FIG. Among them, the 360-degree image can be Figure 11 The previous steps perform pre-processing (stitching, projection light) and Figure 12 The following steps perform post-processing (rendering, etc.).

[0285] See first Figure 11 After the encoder receives the input image (step A), the image segmentation unit segments the input image into two or more segmentation units (at this time, setting information related to the segmentation method can be generated, represented as step B). Next, an additional region is generated for the segmentation unit based on the encoding settings or whether the additional region is supported (step C). Then, the segmentation unit containing the additional region is encoded and a bitstream is generated (step D). In addition, after the bitstream is generated, it can be determined whether to resize (or delete the additional region, step E) based on the encoding settings. Then, the encoded data (the image in step D or E) containing or without the additional region is stored in the memory (step E).

[0286] See Figure 12 , the decoder divides the image to be decoded into two or more division units with reference to the setting information related to division obtained by parsing the received bit stream (step B), then sets the size of the additional area for each division unit according to the decoding setting obtained from the received bit stream (step C), and then obtains the image data including the additional area by decoding the image data contained in the bit stream (step D). Next, a reconstructed image is generated by deleting the additional area (step E), and then the reconstructed image is output to the display (step F). At this time, whether to delete the additional area can be determined according to the decoding setting, and the decoded image or image data (data in step D or E) can be stored in the memory. In addition, step F can include a process of reducing the reconstructed image into a 360-degree image by surface reconfiguration.

[0287] In addition, according to Figure 11 or Figure 12 In order to determine whether to remove the additional area in the image, loop filtering can be adaptively performed on the boundary of the segmentation unit (in this example, a block filter is assumed, but other loop filters can also be applied). In addition, loop filtering can be adaptively performed according to whether the generation of the additional area is allowed.

[0288] When the data is stored in the memory after the additional area is removed, loop filtering can be explicitly applied or not based on the loop filtering applicable or not flag on the division unit boundary (specifically, the initial state) such as loop_filter_across_enabled_flag (in this example, the parallel block).

[0289] Alternatively, the flag indicating whether loop filtering is applicable or not on the division unit boundary may not be supported, and instead the application of filtering and the filter settings may be implicitly determined in a manner as described in the examples below.

[0290] Furthermore, even when image continuity exists between the segments, after additional regions are generated for each segment, image continuity may disappear at the boundaries between the segments for which additional regions have been generated. Applying loop filtering in this situation would unnecessarily increase computational complexity and degrade encoding performance, so loop filtering can be implicitly disabled.

[0291] Furthermore, depending on the surface configuration of the 360-degree image, adjacent segments in two-dimensional space may lack image continuity. Performing loop filtering on the boundaries between such segments without image continuity can result in image quality degradation. Therefore, loop filtering can be implicitly omitted for boundaries between segments without image continuity.

[0292] In addition, when following Figure 10 When weighted values ​​are assigned to two regions and a portion of the current segmentation unit is replaced as described in the description, loop filtering can be applied because the boundaries of the respective segmentation units are internal boundaries within the additional region. However, loop filtering may not be required because coding errors can be reduced by additionally weighting the portion of the current region included in other regions. Therefore, loop filtering can be implicitly omitted in such cases.

[0293] Furthermore, the application of loop filtering (specifically, additionally at the corresponding boundary) can be determined based on a flag indicating whether loop filtering is applicable. When this flag is activated, filtering can be applied based on the loop filtering settings and conditions applied within the partition, or filtering with different definitions of loop filtering settings and conditions can be applied at the partition boundary (specifically, using different loop filtering settings and conditions than those not applied at the partition boundary).

[0294] In the above embodiment, it is assumed that the additional area is removed and then stored in the memory, but part of it can also be in other output steps (specifically, it can belong to both the loop filter unit and the post filter unit). <postfilter>etc.)

[0295] The above examples assume that the additional area supports all directions within each segmented unit. However, if the additional area is only supported in some directions, only a portion of the above information can be applied. For example, the original settings can be applied to boundaries where the additional area is not supported, while the various conditions described in the above examples can be modified and applied to boundaries where the additional area is supported. In other words, the above application can be adaptively determined for all or part of the unit boundaries based on the additional area settings.

[0296] The above-mentioned relevant setting information can be processed implicitly or explicitly according to the encoding / decoding settings. When the implicit method is adopted (specifically, it is based on the characteristics, type, format, etc. of the image or according to other encoding / decoding settings <in this example, the settings related to the additional area>), it can be determined without generating relevant syntax elements, and when the explicit method is adopted, it can be adjusted by generating relevant syntax elements. The units related thereto can include video, sequence, image, sub-image, slice, parallel block, etc.

[0297] Next, the method for determining whether the division unit and the additional region can be referenced will be described in detail. At this time, when reference is possible, it is dependent encoding / decoding, and when it is not possible to refer to it, it is independent encoding / decoding.

[0298] An additional region to which one embodiment of the present invention applies can be referenced or limited in reference during the encoding / decoding process of the current image or other images. Specifically, an additional region removed before storage in memory can be referenced or limited in reference during the encoding / decoding process of the current image. Furthermore, an additional region after storage in memory can be referenced or limited in reference during the encoding / decoding process of images different in time, in addition to the current image.

[0299] In other words, the reference possibility and range of the additional region can be determined based on the encoding / decoding settings. Through some of the above settings, the additional region of the current image will be stored in memory after encoding / decoding, which means that it can be referenced or limited by including it in the reference image of other images. This can be applied to all or some of the segmented units included in the corresponding image. The situations described in the examples described later can also be modified and applied to the current example.

[0300] The above-mentioned setting information related to the reference possibility of the additional area can be processed implicitly or explicitly according to the encoding / decoding settings. When the implicit method is adopted (specifically, it is based on the characteristics, type, format, etc. of the image or according to other encoding / decoding settings <in this example, the settings related to the additional area>), it can be determined without generating relevant syntax elements. When the explicit method is adopted, the settings related to the reference possibility of the additional area can be adjusted by generating relevant syntax elements. The units related thereto can include videos, sequences, images, sub-images, slices, parallel blocks, etc.

[0301] Typically, some units in the current image (in this example, the segmented units obtained by the image segmentation unit) can reference data for the current unit, but cannot reference data for other units. Furthermore, some units in the current image can reference data for all units in other images. The above description is an example of general properties of units obtained by the image segmentation unit, and additional properties related to these properties can also be defined.

[0302] Furthermore, it is possible to define a flag for indicating whether other segmentation units in the current image can be referenced and whether segmentation units included in other images can be referenced.

[0303] As an example, it is possible to allow reference to a segmentation unit contained in another image and located at the same position as the current segmentation unit, but restrict reference to a segmentation unit located at a different position from the current segmentation unit. For example, when transmitting multiple bitstreams encoding the same image under different encoding settings and selectively determining the bitstream used to decode each region (segmentation unit) in the image (assuming decoding in parallel block units in this example), it is necessary to restrict the possibility of reference between each segmentation unit in the same space and in different spaces. Therefore, encoding / decoding can be performed in a manner that only allows reference to the same region in different images.

[0304] As an example, reference can be permitted or restricted based on the identifier information associated with the segmentation unit. For example, reference is permitted when the identifier information assigned to the segmentation units is the same, but is not permitted when the identifier information is different. In this case, the identifier information can be information indicating that encoding / decoding was performed in an environment where mutual reference is possible (dependently).

[0305] The above-mentioned related setting information can be processed implicitly or explicitly according to the encoding / decoding settings. When the implicit method is adopted, it can be determined without generating related syntax elements, and when the explicit method is adopted, it can be processed by generating related syntax elements. The related units can include video, sequence, image, sub-image, slice, parallel block, etc.

[0306] Figures 13a to 13g This is an example diagram for explaining the area that can be referenced by a specific segmentation unit. Figures 13a to 13g In the figure, the area shown with a thick frame line can represent a reference area.

[0307] See Figure 13a , it is possible to confirm various reference arrows used to perform inter-frame prediction. At this time, blocks C0 and C1 represent unidirectional inter-frame prediction. Block C0 can obtain the RP0 reference block before the current image, and can obtain the RF0 reference block after the current image. Block C2 represents bidirectional inter-frame prediction, and can obtain RP1 and RF1 reference blocks from images before or after the current image. The accompanying drawings illustrate examples of obtaining one reference block each from the previous direction and the subsequent direction, but it is also possible to obtain reference blocks only from the previous direction or the subsequent direction. Block C3 represents non-directional inter-frame prediction, and can obtain the RC0 reference block from the current image. The accompanying drawings illustrate examples of obtaining one reference block, but it is also possible to obtain more than two reference blocks.

[0308] In the examples described later, the description will focus on the possibility of reference of pixel values ​​and prediction mode information in inter-frame prediction based on partition units, but it can also be understood to include other encoding / decoding information that can be referenced spatially or temporally (such as intra-frame prediction mode information, transformation and quantization information, loop filter information, etc.).

[0309] See Figure 13b , the current image Currnt(t) is divided into two or more parallel blocks, and block C0 in some of the parallel blocks can obtain reference blocks P0 and P1 by performing unidirectional inter-picture prediction. Block C1 in some of the parallel blocks can obtain reference blocks P3 and F0 by performing bidirectional inter-picture prediction. In other words, this can be understood as an example of allowing reference to blocks at other locations contained in other images without restrictions such as positional restrictions or limiting reference to only within the same image.

[0310] See Figure 13c The image is divided into two or more parallel block units, and a portion of block C1 in a portion of the parallel blocks can obtain reference blocks P2 and P3 by performing unidirectional inter-picture prediction. A portion of block C0 in a portion of the parallel blocks can obtain reference blocks P0, P1, F0, and F1 by performing bidirectional inter-picture prediction. A portion of block C3 in a portion of the parallel blocks can obtain reference block FC0 by performing non-directional inter-picture prediction.

[0311] Right now, Figure 13b as well as Figure 13c This can be understood as an example of allowing reference to blocks at other locations contained in other images without restrictions such as positional restrictions and allowing reference only within the same image.

[0312] See Figure 13d The current image is divided into two or more parallel block units, and block C0 in some of the parallel blocks can obtain reference block P0 by performing forward inter-picture prediction, but cannot obtain reference blocks P1, P2, and P3 included in some of the parallel blocks. Block C4 in some of the parallel blocks can obtain reference blocks F0 and F1 by performing backward inter-picture prediction, but cannot obtain reference blocks F2 and F3. Block C3 in some of the parallel blocks can obtain reference block FC0 by performing non-directional inter-picture prediction, but cannot obtain reference block FC1.

[0313] That is, in Figure 13d In the example, it is possible to set whether to allow or restrict reference based on whether the image (in this example, t-1, t, and t+1) is split or not, and the encoding / decoding of the image division unit. Specifically, it is possible to only allow reference to tiles included in parallel tiles that have the same identifier information as the current parallel tile.

[0314] See Figure 13e , the image is divided into two or more parallel block units, and a portion of the parallel blocks C0 can obtain reference blocks P0 and F0 by performing bidirectional inter-frame prediction, but cannot obtain reference blocks P1, P2, P3, F1, F2, and F3. That is, Figure 13e It may be an example that only references to parallel tiles that are co-located with the parallel tile containing the current tile are allowed.

[0315] See Figure 13f , the image is divided into two or more parallel block units, and a part of the parallel blocks C0 can obtain reference blocks P1 and F2 by performing bidirectional inter-frame prediction, but reference blocks P0, P2, P3, F0, F1, and F3 cannot be obtained. Figure 13f This may be an example where the bitstream includes information indicating a parallel tile that can be referenced by the current partition unit, and the referenced parallel tile is confirmed using the information.

[0316] See Figure 13g , the image is divided into two or more parallel blocks, and a portion of block C0 in some of the parallel blocks can obtain reference blocks P0, P3, and P5 by performing unidirectional inter-picture prediction, but cannot obtain reference block P4. A portion of block C1 in some of the parallel blocks can obtain reference blocks P1, F0, and F2 by performing bidirectional inter-picture prediction, but cannot obtain reference blocks P2 and F1.

[0317] Figure 13g This example allows or restricts reference based on whether the image (in this example, t-3, t-2, t-1, t, t+1, t+2, and t+3) is partitioned, encoding / decoding settings of the image partition units (in this example, it is assumed that the settings are determined based on the identifier information of the partition units, the identifier information of the image units, whether the partition units are the same area, whether the partition units are similar areas, and bitstream information of the partition units). The referenced parallel block can be located in the same or similar position within the image as the current block, can have the same identifier information as the current block (specifically, in image units or partition units), and can be the same bitstream as the current parallel block.

[0318] Figures 14a to 14e This is a flowchart for explaining the possibility of referring to additional areas in a division unit to which one embodiment of the present invention is applied. Figures 14a to 14e In FIG. 1 , the area indicated by the thick frame line represents a reference area, and the area indicated by the dotted line represents an additional area of ​​the division unit.

[0319] In one embodiment of the present invention, the possibility of referencing a portion of an image (another image that is temporally preceding or following) can be restricted or permitted. Furthermore, the possibility of referencing the entire expanded segmented unit that includes the additional region can be restricted or permitted. Furthermore, the possibility of referencing only the initial segmented unit excluding the additional region can be restricted or permitted. Furthermore, the possibility of referencing the boundary between the additional region and the initial segmented unit can be restricted or permitted.

[0320] See Figure 14a Part of the parallel blocks, C0, can obtain reference blocks P0 and P1 by performing unidirectional inter-frame prediction. Part of the parallel blocks, C2, can obtain reference blocks P2, P3, F0, and F1 by performing bidirectional inter-frame prediction. Part of the parallel blocks, C1, can obtain reference block FC0 by performing non-directional inter-frame prediction. Block C0 can obtain reference blocks P0, P1, P2, P3, and F0 from the initial parallel block regions (the original parallel blocks excluding the additional region) of a portion of reference images t-1 and t+1, while block C2 can obtain reference blocks P2 and P3 from the initial parallel block region of reference image t-1 and reference block F1 from the parallel block region of reference image t+1 that includes the additional region. In this case, as shown by reference block F1, a reference block can be obtained that includes the boundary between the additional region and the initial parallel block region.

[0321] See Figure 14b Part of the parallel blocks C0, C1, and C3 can obtain reference blocks P0, P1, P2 / F0, F2 / F1, F3, and F4 by performing unidirectional inter-picture prediction. Part of the parallel blocks C2 can obtain reference blocks FC0, FC1, and FC2 by performing non-directional inter-picture prediction.

[0322] A portion of blocks C0, C1, and C3 can obtain reference blocks P0, F0, and F3 from the initial parallel block area of ​​a portion of the reference image (t-1, t+1 in this example), and can also obtain reference blocks P1, x, and F4 from the updated parallel block area boundary, and can also obtain reference blocks P2, F2, and F1 from outside the updated parallel block area boundary.

[0323] A portion of block C2 can obtain reference block FC1 from an initial parallel block area of ​​a portion of the reference image (t in this example), can also obtain reference block FC3 from an updated parallel block area boundary, and can also obtain reference block FC0 from outside the updated parallel block area boundary.

[0324] Part of the blocks C0 may be blocks located in the initial parallel block region, part of the blocks C1 may be blocks located at the boundary of the updated parallel block region, and part of the blocks C3 may be blocks located outside the updated parallel block boundary.

[0325] See Figure 14c The image is divided into two or more parallel block units. Additional regions are set for some parallel blocks in some images, not set for some parallel blocks in some images, and not set for some images. For some blocks C0 and C1 in some parallel blocks, unidirectional inter-picture prediction can be used to obtain reference blocks P2, F1, F2, and F3, but reference blocks P0, P1, P3, and F0 cannot be obtained. For some block C2 in some parallel blocks, non-directional inter-picture prediction can be used to obtain reference blocks FC1 and FC2, but reference block FC0 cannot be obtained.

[0326] A part of the block C2 cannot obtain the reference block FC0 from the initial parallel block area of ​​a part of the reference image (t in this example), but can obtain the reference block FC1 from the updated parallel block area (in the method of filling a part of the additional area, FC0 and FC1 can be the same area. Although FC0 cannot be referenced in the parallel block division of the initial unit, it can be referenced when the corresponding area is moved to the current parallel block through the additional area).

[0327] A portion of block C2 is able to obtain a reference block FC2 from a portion of a parallel block area of ​​a portion of a reference image (t in this example) (although it is not possible to reference data in other parallel blocks of the current image by default, it is assumed that reference is allowed when it is set as referenceable through the identifier information in the above embodiment, etc.).

[0328] See Figure 14d The image is divided into two or more parallel block units and additional regions are defined. A portion of the parallel blocks, C0, can obtain reference blocks P0, F0, F1, and F3 through bidirectional inter-frame prediction, but cannot obtain reference blocks P1, P2, P3, and F2.

[0329] A portion of block C0 can obtain reference block P0 from an initial parallel block area (parallel block No. 0) of a portion of reference image t-1, but cannot obtain reference block P3 from the boundary of the extended parallel block area, nor can it obtain reference block P2 from outside the boundary of the extended parallel block area (i.e., the additional area).

[0330] A part of the block C0 can obtain the reference block F0 from the initial parallel block area (parallel block No. 0) of a part of the reference image t+1, can also obtain the reference block F1 from the extended parallel block area boundary, and can also obtain the reference block F3 from outside the extended parallel block area boundary.

[0331] See Figure 14e The image is divided into two or more parallel tile units and an additional region having at least one size and shape is defined. Block C0 in some parallel tiles can obtain reference blocks P0, P3, P5, and F0 by performing unidirectional inter-frame prediction, but cannot obtain reference block P2 located at the boundary between the additional region and the original parallel tile. Block C1 in some parallel tiles can obtain reference blocks P1, F2, and F3 by performing bidirectional inter-frame prediction, but cannot obtain reference blocks P4, F1, and F5.

[0332] As shown in the above example, pixel values ​​can be used as reference objects, and reference to other encoding / decoding information can be restricted.

[0333] As an example, when the prediction unit searches for a candidate set of intra-picture prediction modes to be used in intra-picture prediction from spatially adjacent blocks, it can be performed as follows: Figures 13a to 14e The method shown confirms whether a partition unit including a current block can refer to a partition unit including an adjacent block.

[0334] As an example, when the prediction unit searches for a motion information candidate group to be used in inter prediction for a picture from temporally and spatially adjacent blocks, it can be done as follows: Figures 13a to 14e The method shown determines whether a partition unit including a current block can refer to a partition unit including a block that is spatially adjacent in a current image or temporally adjacent to the current image.

[0335] As an example, when the loop filter unit searches for loop filter related setting information from adjacent blocks, it can be done as follows: Figures 13a to 14e The method shown confirms whether a partition unit including a current block can refer to a partition unit including an adjacent block.

[0336] Figure 15 This is an example diagram illustrating blocks included in the division unit of the current image and blocks included in the division units of other images.

[0337] See Figure 15 In this example, the spatially adjacent reference candidate blocks can be the left, upper left, lower left, upper side, and upper right blocks centered on the current block. In addition, the temporal reference candidate blocks can be the left, upper left, lower left, upper side, upper right, right side, lower right, lower side, and central blocks of the collocated block located at the same or corresponding position as the current block in the image (Different picture) that is temporally adjacent to the current image (Current picture). Figure 15 The thick outer frame line indicates the boundary line of the division unit.

[0338] When the current block is M, the spatially adjacent blocks G, H, I, L, and Q can all be used for reference.

[0339] When the current block is G, some of the spatially adjacent blocks A, B, C, F, and K can be referenced, while the remaining blocks can be limited in reference. Whether reference is permitted or not can be determined based on the reference correlation settings between the division units UC, ULC, and LC containing the spatially adjacent blocks and the division unit containing the current block.

[0340] When the current block is S, some of the neighboring blocks s, r, m, w, n, x, t, o, and y located at the same position as the current block in temporally adjacent images can be referenced, while the remaining blocks can be limited in reference. Whether reference is permitted is determined by the reference correlation settings between the division units RD, DRD, and DD containing the neighboring blocks located at the same position as the current block in temporally adjacent images and the unit containing the current block.

[0341] According to the position of the current block, when a candidate with limited reference exists, it can be filled with a candidate with the next priority in the candidate group, or replaced with another candidate adjacent to the candidate with limited reference.

[0342] For example, when the current block in the intra-picture prediction is G, the reference of the left-up fast is limited, and the most likely mode (MPM) candidate group is constructed in the order of PDAEU, because A is not a reference, the candidate group can be constructed by performing a validity check according to the remaining EU order, or A can be replaced by B or F that is spatially adjacent to A.

[0343] In addition, when the current block in inter-frame prediction is S, the reference of the temporally adjacent next block is limited and the temporal candidate of the skip mode candidate group is composed of y, since y is not a reference, the candidate group can be composed by performing validity checks in the order of spatially adjacent candidates or mixed candidates of idle candidates and temporal candidates, or using t, x, s that are spatially adjacent to y instead of y.

[0344] Figure 16 This is a diagram illustrating the hardware configuration of a video encoding / decoding device to which one embodiment of the present invention is applied.

[0345] See Figure 16 An image encoding / decoding device 200 applicable to one embodiment of the present invention may include: at least one processor 210; and a memory 220 storing instructions for instructing the at least one processor 210 to perform at least one step.

[0346] The at least one processor 210 may be a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor for executing methods applicable to embodiments of the present invention. The memory 120 and the storage device 260 may each be composed of at least one of a volatile storage medium and a non-volatile storage medium. For example, the memory 220 may be composed of at least one of a read-only memory (ROM) and a random access memory (RAM).

[0347] The video encoding / decoding device 200 can also include a transceiver 230 for communicating via a wireless communication network. Furthermore, the video encoding / decoding device 200 can include an input interface 240, an output interface 250, a storage device 260, and the like. The various components included in the video encoding / decoding device 200 can be connected via a bus 270 to enable communication with each other.

[0348] Among them, at least one step may include: a step of dividing the encoded image contained in the above-mentioned bitstream into at least one segmentation unit by referring to the syntax elements obtained from the received bitstream; a step of setting an additional area for the above-mentioned at least one segmentation unit; and a step of decoding the above-mentioned encoded image based on the segmentation unit after setting the additional area.

[0349] The step of decoding the coded image may include determining a reference block related to a current block to be decoded in the coded image based on information indicating the possibility of reference contained in the bit stream.

[0350] The reference block may be a block included in a position overlapping with an additional area set on a division unit including the reference block.

[0351] Figure 17 This is an exemplary diagram illustrating an intra-frame prediction mode to which one embodiment of the present invention is applied.

[0352] See Figure 17 , it can be confirmed that there are a total of 35 prediction modes, and the 35 prediction modes can be divided into 33 directional modes and 2 non-directional modes (mean (DC), plane (Planar)). At this time, the directional mode can be identified by inclination (such as dy / dx) or angle information. The above example can refer to a prediction mode candidate group related to the brightness component or the color difference component. Alternatively, the color difference component can support a part of the prediction modes (such as mean (DC), plane (Planar), vertical, horizontal, diagonal mode, etc.). In addition, after the prediction mode of the brightness mode is determined, the corresponding mode can be included in the prediction mode of the color difference component or the mode derived from the corresponding mode can be included in the prediction mode.

[0353] Furthermore, the correlation between color spaces can be exploited to apply reconstructed blocks in other color spaces that have already been encoded / decoded to the prediction of the current block, including supported prediction modes. For example, the chrominance component can be used to generate the prediction block of the current block using the reconstructed block of the luma component corresponding to the current block.

[0354] The prediction mode candidate groups can be adaptively determined according to the encoding / decoding settings. The number of candidate groups can be increased to improve prediction accuracy or decreased to reduce the bit size in the prediction mode.

[0355] For example, one of the candidate groups such as A candidate group (67, 65 directional modes and 2 non-directional modes), B candidate group (35, 33 directional modes and 2 non-directional modes), and C candidate group (19, 17 directional modes and 2 non-directional modes) can be used. Unless otherwise specified, the present invention assumes that intra-frame prediction is performed using a pre-set prediction mode candidate group (A candidate group).

[0356] Figure 18 This is a first exemplary diagram illustrating a reference pixel structure used in intra-frame prediction according to an embodiment of the present invention.

[0357] A method for intra-frame prediction in video decoding according to an embodiment of the present invention may include: a reference pixel configuration step; a prediction block generation step using one or more prediction modes with reference to the configured reference pixels; a step of determining an optimal prediction mode; and a step of encoding the determined prediction mode. Furthermore, a video decoding device may include a reference pixel configuration unit, a prediction block generation unit, a prediction mode determination unit, and a prediction mode encoding unit for performing the reference pixel configuration step, the prediction block generation step, the prediction mode determination step, and the prediction mode encoding step. The above-described process may omit portions thereof, or additional processes may be added, and the order may be modified to differ from the order described above.

[0358] Furthermore, the intra-frame prediction method in video decoding according to one embodiment of the present invention can generate a prediction block of a current block according to a prediction mode obtained through a syntax element received from a video encoding apparatus after constructing reference pixels.

[0359] The size and shape (M×N) of the current block for performing intra-frame prediction can be obtained from the block partitioning unit, and can be in a size of 4×4 to 256×256. Intra-frame prediction is usually performed in prediction block units, but can also be performed in units such as coding blocks (or coding units), transform blocks (or transform units) according to the settings of the block partitioning unit. After the block information is confirmed, the reference pixels used in the prediction of the current block can be constructed by the reference pixel construction unit. At this time, the reference pixels can be stored in a temporary memory (such as an array <array>, 1-dimensional, 2-dimensional arrays, etc.), are managed, generated and removed during each intra-frame prediction process of the block, and the size of the temporary memory can be determined according to the composition of the reference pixels.

[0360] Reference pixels may be pixels included in adjacent blocks (referred to as reference blocks) located to the left, above, upper left, upper right, or lower left of the current block, but are not limited thereto. Block candidate groups of other configurations may also be used in the prediction of the current block. The adjacent blocks located to the left, above, upper left, upper right, or lower left may be blocks selected when encoding / decoding is performed using a grid or zigzag scan, and adjacent blocks located elsewhere (e.g., right, lower, or lower right blocks) may also be used as reference pixels when the scan order is changed.

[0361] Furthermore, the reference block may be a block corresponding to the current block in a color space different from the color space containing the current block. When using a Y / Cb / Cr format, the color space may refer to one of Y, Cb, and Cr. Furthermore, the block corresponding to the current block may be a block having the same position coordinates as the current block, or a block having position coordinates corresponding to the current block based on a color component composition ratio.

[0362] In addition, for the convenience of explanation, the reference block at the above-mentioned pre-set position (left side, upper side, upper left, upper right, lower left) will be described as consisting of one block, but it can also be composed of multiple sub-blocks according to block segmentation.

[0363] In other words, the adjacent area of ​​the current block can be the reference pixel position for intra-frame prediction of the current block, and the area corresponding to the current block in other color spaces can be additionally used as a reference pixel position according to the prediction mode. In addition to the above examples, the defined reference pixel position can also be determined according to the prediction mode, method, etc. For example, when generating a prediction block by a method such as block matching, the reference pixel position can be an area that has been encoded / decoded before the current block of the current image or an area included in the exploration range of the area that has been encoded / decoded (for example, including the left or right side or upper left or upper right of the current block, etc.) as the reference pixel position.

[0364] See Figure 18 , the reference pixels used in the intra-frame prediction of the current block (size is M×N) can be the pixels adjacent to the current block on the left, above, upper left, upper right, and lower left ( Figure 18 In this case, Figure 18 The content in the form of P(x,y) can refer to pixel coordinates.

[0365] Furthermore, pixels adjacent to the current block can be classified into at least one reference pixel level, such as ref_0, which is the pixel closest to the current block (pixels whose pixel values ​​differ by 1 from the boundary pixels of the current block, p(-1,-1) to p(2M-1,-1), p(-1,0) to p(-1,2N-1)), followed by ref_1, which is the pixel whose pixel values ​​differ by 2 from the boundary pixels of the current block, p(-2,-2) to p(2M,-2), p(-2,-1) to p(-2,2N), and ref_2, which is the pixel whose pixel values ​​differ by 3 from the boundary pixels of the current block, p(-3,-3) to p(2M+1,-3), p(-3,-2) to p(-3,2N+1), and so on. In other words, reference pixels can be classified into multiple reference pixel levels based on the distance of pixels adjacent to the boundary pixels of the current block.

[0366] Furthermore, different reference pixel levels can be set for adjacent blocks. For example, when the block adjacent to the upper edge of the current block is used as a reference block, reference pixels on the ref_0 level can be used, while when the block adjacent to the upper right edge is used as a reference block, reference pixels on the ref_1 level can be used.

[0367] Among them, the reference pixel set that is usually referenced when performing intra-frame prediction is included in the blocks adjacent to the lower left, left, upper left, upper end, and upper right end of the current block, and is a pixel belonging to the ref_0 level (the pixel most adjacent to the boundary pixel). Unless otherwise specified in the following content, the above-mentioned pixels are taken as a premise. However, it is also possible to use a reference pixel set included in a part of the adjacent blocks mentioned in the above content, and it is also possible to use pixels included in two or more levels as a reference pixel set. Among them, the reference pixel set or level can be implicitly determined (pre-set in the encoding / decoding device) or explicitly determined (information for determination is received from the encoding device).

[0368] Here, the explanation will be based on the premise of supporting a maximum of 3 reference pixel levels, but a larger value can also be used. The number of reference pixel levels and the number of reference pixel sets based on the referenced adjacent block positions (or can also be called reference pixel candidate groups) can <I / P / B,此时的影像为图像、条带、并行区块等>be set differently according to the size, shape, prediction mode, image type, color component, etc. of the block, and the relevant information can be included in units such as sequences, images, strips, parallel blocks, etc.

[0369] ​In the present invention, the description is based on the assumption that lower index values ​​(starting from 0 and gradually increasing by 1) are assigned starting from the reference pixel level most adjacent to the current block, but the present invention is not limited to this. In addition, the reference pixel composition-related information described later can be generated under the above-mentioned index setting (such as binarization that assigns shorter bits to smaller indices when selecting one from multiple reference pixel sets).

[0370] In addition, when two or more reference pixel levels are supported, weighted averaging or the like can be applied to the reference pixels included in the two or more reference pixel levels.

[0371] For example, it is possible to use Figure 18 A prediction block is generated using the reference pixels obtained by summing the weighted values ​​of the pixels in the ref_0th layer and the ref_1th layer. At this time, depending on the prediction mode (e.g., the directionality of the prediction mode), the pixels to which the weighted values ​​are applied in each reference pixel layer can be either integer-unit pixels or fractional-unit pixels. Furthermore, a prediction block can be obtained by assigning weighted values ​​(e.g., 7:1, 3:1, 2:1, 1:1, etc.) to the prediction block obtained using the reference pixels in the first reference pixel layer and the prediction block obtained using the reference pixels in the second reference pixel layer. At this time, a higher weighted value can be assigned to the prediction block in the reference pixel layer that is more adjacent to the current block.

[0372] Assuming that explicit information on reference pixel composition is generated, indication information (in this example, adaptive_intra_ref_sample_enabled_flag) for enabling adaptive reference pixel composition can be generated in units such as video, sequence, picture, slice, and tile.

[0373] When the indication information indicates that adaptive reference pixel composition is allowed (adaptive_intra_ref_sample_enabled_flag=1 in this example), adaptive reference pixel composition information (adaptive_intra_ref_sample_flag in this example) can be generated in units such as picture, slice, tile, or block.

[0374] When the above-mentioned composition information represents adaptive reference pixel composition (adaptive_intra_ref_sample_flag=1 in this example), reference pixel composition-related information (such as selection information related to reference pixel levels and sets, etc., in this example, intra_ref_idx) can be generated on units such as images, strips, parallel blocks, and blocks.

[0375] In this case, when adaptive reference pixel composition is not permitted or is not adaptive, reference pixels can be composed according to a predetermined setting. For example, the most adjacent pixels in adjacent blocks are typically used to form reference pixels, but this is not limited to this, and various situations are also allowed (for example, selecting ref_0 and ref_1 as reference pixel levels and using ref_0 and ref_1 to generate predicted pixel values ​​by weighted summation, i.e., a default situation).

[0376] In addition, information related to reference pixel composition (such as selection information related to the reference pixel level or set, etc.) can be constructed (such as ref_1, ref_2, ref_3, etc.) after excluding pre-set information (such as the case where the reference pixel level is pre-set to ref_0), but is also not limited to this.

[0377] The above examples illustrate some aspects related to reference pixel composition, but the intra-frame prediction setting can be determined by combining various encoding / decoding information. The encoding / decoding information can include, for example, image type, color component, size and shape of the current block, prediction mode {prediction mode type (directional, non-directional), prediction mode direction (vertical, horizontal, diagonal 1, diagonal 2, etc.)}, etc. Furthermore, the intra-frame prediction setting (in this example, the reference pixel composition setting) can be determined based on encoding / decoding information of neighboring blocks and a combination of encoding / decoding information of the current block and neighboring blocks.

[0378] Figures 19a to 19c This is a second exemplary diagram illustrating a reference pixel configuration according to an embodiment of the present invention.

[0379] See Figure 19a , can be used only to Figure 18 The situation in which the reference pixel level ref_0 in the image is used to constitute the reference pixels is confirmed. After the reference pixel level ref_0 is used as the object and the pixels contained in the adjacent blocks (e.g., the lower left, left, upper left, upper side, and upper right) are used to constitute the reference pixels, subsequent intra-frame prediction (such as reference pixel generation, reference pixel filtering, reference pixel interpolation, prediction block generation, post-processing filtering, etc.) can be performed. Part of the intra-frame prediction process can be adaptively performed based on the reference pixel composition. In this example, the case of using a pre-set reference pixel level, that is, an example of performing intra-frame prediction using a non-directional mode without generating setting information related to the reference pixel level, will be described.

[0380] See Figure 19b , it is possible to confirm the case where reference pixels are constructed using both of the supported reference pixel hierarchies. That is, it is possible to perform intra-frame prediction after constructing reference pixels using pixels included in hierarchies ref_0 and ref_1 (or using the weighted average of the pixels included in the two hierarchies). In this example, an example will be described where multiple pre-set reference pixel hierarchies are used, that is, where setting information related to the reference pixel hierarchies is not generated, and intra-frame prediction is performed using a portion of directional prediction modes (from the upper right to the lower left in the figure, or in the opposite direction).

[0381] See Figure 19c , it is possible to confirm the case where reference pixels are constructed using only one of the three supported reference pixel hierarchies. In this example, because there are multiple reference pixel hierarchy candidates, setting information related to the reference pixel hierarchy used is generated and an example of performing intra-frame prediction using a portion of directional prediction modes (from the upper left to the lower right in the figure) will be described.

[0382] Figure 20 This is a third exemplary diagram illustrating a reference pixel configuration according to an embodiment of the present invention.

[0383] Figure 20 The figure number a in the figure is a block with a size of 64×64 or larger, the figure number b is a block with a size of 16×16 or larger but less than 64×64, and the figure number c is a block with a size of less than 16×16.

[0384] When the block numbered a in the figure is used as the current block for which intra prediction needs to be performed, the intra prediction can be performed using the most adjacent reference pixel level ref_0.

[0385] In addition, when the block numbered b in the figure is used as the current block for which intra prediction needs to be performed, the intra prediction can be performed using the two supported reference pixel levels ref_0 and ref_1.

[0386] In addition, when the block numbered c in the figure is used as the current block for which intra prediction needs to be performed, intra prediction can be performed using the three supported reference pixel levels ref_0, ref_1, and ref_2.

[0387] As described in the description of Figures a to c, the number of supported reference pixel levels can be set differently according to the size of the current block for which intra prediction is required. Figure 20 In this case, the larger the size of the current block, the more likely the size of the neighboring block is small, which can be a result of performing partitioning based on other image characteristics, and thus, in order to prevent prediction from being performed using a pixel having a large distance in pixel value from the current block, it is assumed that the number of reference pixel levels supported is smaller as the size of the block is larger, but other variations including the opposite case are also allowed.

[0388] Figure 21 FIG. 4 is a fourth exemplary diagram illustrating a configuration of reference pixels according to an embodiment of the present application.

[0389] Referring to Figure 21 It can be confirmed that the current block in which intra prediction is performed is in a rectangular shape. If the current block is in a rectangular shape that is horizontally and vertically asymmetric, the number of supported reference pixel levels adjacent to the longer horizontal side boundary surface in the current block can be set to be larger, and the number of supported reference pixel levels adjacent to the shorter vertical side boundary surface in the current block can be set to be smaller. In the drawing, it can be confirmed that the number of reference pixel levels adjacent to the horizontal boundary surface of the current block is set to 2, and the number of reference pixel levels adjacent to the vertical boundary surface of the current block is set to 1. The pixel adjacent to the shorter vertical side boundary surface in the current block can have a problem of decreased accuracy due to a generally relatively large distance from the pixel included in the current block because of the larger horizontal length. Thus, the number of supported reference pixel levels adjacent to the shorter vertical side boundary surface is set to be smaller, but a setting opposite thereto can also be used.

[0390] Further, the reference pixel levels to be used in prediction can be differently set according to the type of the intra prediction mode or the position of the neighboring block adjacent to the current block. For example, a directional mode using pixels included in a block adjacent to the upper end, the upper right end of the current block as reference pixels can use two or more reference pixel levels, and a directional mode using pixels included in a block adjacent to the left end, the lower left end of the current block as reference pixels can use only one most adjacent reference pixel level.

[0391] Further, when the prediction blocks generated through the respective reference pixel levels in the plurality of reference pixel levels are the same or similar to each other, the generation of the setting information of the reference pixel levels can cause a problem of unnecessarily generated data.

[0392] For example, when the distribution characteristics of the pixels constituting each reference pixel level are similar or identical, similar or identical prediction blocks can be generated regardless of which reference pixel level is used, and thus, there is no need to generate data for selecting a reference pixel level. In this case, the distribution characteristics of the pixels constituting the reference pixel level can be determined by comparing the average value or dispersion value of the pixels with a pre-set threshold.

[0393] That is, when the reference pixel levels based on the finally determined intra prediction mode are the same or similar to each other, the reference pixel level can be selected using a pre-set method (eg, selecting the most adjacent reference pixel level).

[0394] In this case, the decoder can receive intra-frame prediction information (or intra-frame prediction mode information) from the encoding apparatus and determine whether to receive information for selecting a reference pixel level based on the received information.

[0395] The above-mentioned various examples illustrate the case where reference pixels are constructed using a plurality of reference pixel levels, but the present invention is not limited thereto. Various modified examples can be adopted, and the invention can be used in combination with other additional configurations.

[0396] The reference pixel generation unit for intra-frame prediction can include, for example, a reference pixel generation unit, a reference pixel interpolation unit, and a reference pixel filtering unit, and can include all or part of the aforementioned components. Blocks containing pixels that can serve as reference pixels can be referred to as reference candidate blocks. Furthermore, reference candidate blocks can typically be adjacent blocks to the current block.

[0397] The reference pixel constructing unit can determine whether to use pixels included in the reference candidate block as reference pixels based on the reference pixel availability set for the reference candidate block.

[0398] Regarding the availability of reference pixels, it can be determined that the reference pixels are unavailable if at least one of the following conditions is met. For example, if a reference candidate block satisfies at least one of the following conditions: being outside the image boundary, not included in the same segmentation unit (e.g., slice, tile, etc.) as the current block, not yet fully encoded or decoded, or restrictions on its use are set in the encoding or decoding settings, it can be determined that the pixels included in the corresponding reference candidate block cannot be referenced. In this case, if any of the above conditions are not met, the reference pixel can be determined to be available for use.

[0399] Furthermore, the use of reference pixels can be restricted based on encoding / decoding settings. For example, when a flag (e.g., constrained_intra_pred_flag) for restricting the reference of a reference candidate block is activated, pixels included in the corresponding reference candidate block can be restricted from being used as reference pixels. To enable efficient encoding / decoding even when errors are caused by various external factors, including the communication environment, the above flag can be applied when the reference candidate block is reconstructed by referencing an image that is temporally different from the current block.

[0400] When a flag for restricting reference is activated (e.g., when constrained_intra_pred_flag = 0 in an I picture type or a P or B picture type), all pixels in the reference candidate block can be used as reference pixels. Furthermore, when a flag for restricting reference is activated (e.g., when constrained_intra_pred_flag = 1 in a P or B picture type), whether the reference candidate block can be used as a reference can be determined based on whether it is encoded using intra prediction or inter prediction. That is, when the reference candidate block is encoded using intra prediction, the reference candidate block can be used as a reference regardless of whether the flag is activated. However, when the reference candidate block is encoded using inter prediction, whether the flag is activated determines whether the reference candidate block can be used as a reference.

[0401] Furthermore, a reconstructed block located at a position corresponding to the current block in another color space can be used as a reference candidate block. In this case, whether or not the reference candidate block can be used as a reference can be determined based on the coding mode of the reference candidate block. For example, when the current block is part of the chrominance components (Cb, Cr), whether or not the reference candidate block can be used as a reference can be determined based on the coding mode of a block (=reference candidate block) located at a position corresponding to the current block in the luminance component (Y) and already encoded / decoded. This can be an example corresponding to the case where the coding mode is determined independently of the color space.

[0402] The flag used to restrict the reference may be a setting applicable to a certain image type (eg, P or B slice / panel block type, etc.).

[0403] The reference pixel usage probability can be used to classify the reference candidate blocks into fully usable, partially usable, and completely unusable. In other cases except the fully usable case, reference pixels at the unusable candidate block positions can be filled or generated.

[0404] When a reference candidate block is available, pixels at a predetermined position in the current block (or pixels adjacent to the current block) can be stored in the reference pixel memory of the current block. The pixel data at the corresponding block position can be directly copied or stored in the reference pixel memory through a process such as reference pixel filtering.

[0405] In the case where the reference candidate block is not available, the pixels obtained through the reference pixel generation process can be included in the reference pixel memory of the current block.

[0406] In other words, reference pixels can be constructed in a state where the reference pixel candidate block can be used, and reference pixels can be generated in a state where the reference pixel candidate block cannot be used.

[0407] The method for filling reference pixels at pre-set locations in unusable reference candidate blocks is as follows. First, reference pixels can be generated using arbitrary pixel values. The arbitrary pixel value is a specific pixel value within a pixel value range and can be the minimum, maximum, or median value of pixel values ​​used in a pixel value adjustment process based on bit depth or pixel value adjustment process based on image pixel value range information, or a value derived from these values. The method of generating reference pixels using arbitrary pixel values ​​is also applicable even when all reference candidate blocks are unusable.

[0408] Next, reference pixels can be generated using pixels contained in blocks adjacent to the unusable reference candidate block. Specifically, pixels contained in adjacent blocks can be filled into pre-set positions in the unusable reference candidate block through extrapolation, interpolation, or copying. The copying or extrapolation method can be clockwise or counterclockwise, and can be determined based on encoding / decoding settings. For example, the direction of reference pixel generation within the block can follow a pre-set direction or a direction adaptively determined based on the position of the unusable block.

[0409] Figures 22a to 22b This is an exemplary diagram illustrating a method of filling reference pixels at predetermined positions in an unusable reference candidate block.

[0410] See Figure 22a , a method for filling pixels included in an unusable reference candidate block in reference pixels composed of a reference pixel hierarchy can be confirmed. Figure 22a In the example, when the neighboring block adjacent to the upper right end of the current block is an unusable reference candidate block, the reference pixels (expressed as <1> ) can be generated by performing clockwise extrapolation or linear extrapolation on the reference pixels contained in the adjacent block adjacent to the upper end of the current block.

[0411] In addition, Figure 22a In the example, when the neighboring block on the left side of the current block is an unusable reference candidate block, the reference pixels (expressed as <2> ) can be generated by counterclockwise extrapolation or linear extrapolation of reference pixels included in a neighboring block adjacent to the upper left end of the current block (corresponding to an available block). In this case, by performing clockwise extrapolation or linear extrapolation, reference pixels included in a neighboring block adjacent to the lower left end of the current block can be used.

[0412] In addition, Figure 22a In the example, a portion of the reference pixels contained in the adjacent block above the current block (expressed as <3> ) can be generated by interpolating or linearly interpolating the usable reference pixels on both sides. In other words, it is possible to set a setting where some, but not all, of the reference pixels included in the adjacent blocks are unusable. In this case, the unusable reference pixels can be filled with the pixels adjacent to the unusable reference pixels.

[0413] See Figure 22b , a method for filling unusable reference pixels when some of the reference pixels in a reference pixel hierarchy composed of multiple reference pixel levels are unusable can be confirmed. Figure 22b , when the neighboring block adjacent to the upper right end of the current block is an unusable reference candidate block, the pixels contained in the three reference pixel levels of the corresponding neighboring block (expressed as <1> ) can be generated in a clockwise direction using the pixels contained in the adjacent block (corresponding to the usable block) adjacent to the upper end of the current block.

[0414] In addition, Figure 22b In the method, when the neighboring block adjacent to the left side of the current block is an unusable reference candidate block and the neighboring block adjacent to the upper left end or the lower left end of the current block is a usable reference candidate block, the reference pixels of the unusable reference candidate block can be generated by filling the reference pixels of the usable reference candidate block in a clockwise direction, a counterclockwise direction or both directions.

[0415] At this time, unusable reference pixels in each reference pixel level can be generated using pixels in the same reference pixel level, but this does not exclude the use of pixels in different reference pixel levels. Figure 22b In the example, the reference pixels in the three reference pixel levels contained in the adjacent block adjacent to the upper end of the current block (represented as <3> ) is an unusable reference pixel. At this time, the pixels contained in the reference pixel level ref_0 that is most adjacent to the current block and the reference pixel level ref_2 that is farthest away can be generated using the reference pixels contained in the same reference pixel level and can be used. In addition, the pixels contained in the reference pixel level ref_1 that is 1 pixel away from the current block can be generated not only using the pixels contained in the same reference pixel level ref_1, but also using the pixels contained in different reference pixel levels ref_0 and ref_2. At this time, the unusable reference pixels can be filled in by using usable reference pixels on both sides through methods such as quadratic linear interpolation.

[0416] The above example is an example of generating reference pixels when multiple reference pixel levels are composed of reference pixels and some reference candidate blocks are unavailable. Alternatively, it is possible to configure adaptive reference pixel composition (in this example, adaptive_intra_ref_sample_flag = 0) according to encoding / decoding settings (for example, when at least one reference candidate block is unavailable or all reference candidate blocks are unavailable). In other words, reference pixels can be composed according to pre-defined settings without generating any additional information.

[0417] The reference pixel interpolation unit can generate fractional reference pixels by linear interpolation of reference pixels. In the present invention, the process is described as part of the reference pixel generation unit, but it can also be used as a configuration included in the prediction block generation unit, and it can also be understood as a process performed before generating a prediction block.

[0418] Furthermore, although this is assumed to be an independent process separate from the reference pixel filtering section described later, it is also possible to adopt a process that is integrated into one process. This is also a configuration provided to solve the problem of reference pixel distortion caused by an increase in the number of filters applied to the reference pixels when multiple filters are applied by the reference pixel interpolation section and the reference pixel filtering section.

[0419] The reference pixel interpolation process is not performed in some prediction modes (for example, horizontal, vertical, some diagonal modes (such as diagonal down right, diagonal down left, diagonal upright, etc.) at a 45-degree angle, non-directional mode, color mode, color copy mode, etc., that is, modes that do not require decimal unit interpolation when generating a prediction block), and can only be performed in other prediction modes (modes that require decimal unit interpolation when generating a prediction block).

[0420] The interpolation accuracy (e.g., 1, 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, etc. pixel units) can be determined based on the prediction mode (or the directionality of the prediction mode). For example, a prediction mode at a 45-degree angle does not require interpolation, but a prediction mode at a 22.5-degree or 67.5-degree angle requires a 1 / 2 pixel difference. As described above, at least one interpolation accuracy and a maximum interpolation accuracy limit can be determined based on the prediction mode.

[0421] For reference pixel interpolation, a single pre-set interpolation filter (e.g., a 2-tap linear interpolation filter) can be used, or a filter selected from a plurality of interpolation filter candidate groups (e.g., a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc.) according to the encoder / decoder settings can be used. In this case, the interpolation filters can be differentiated based on the number of filter taps (i.e., the number of pixels to which the filter is applied) and the filter coefficients.

[0422] Interpolation can be performed in stages from lower precision to higher precision (e.g., 1 / 2 → 1 / 4 → 1 / 8), or all at once. In the former case, interpolation is performed based on integer units of pixels and fractional units of pixels (pixels that have already been interpolated using a lower precision than the current pixel to be interpolated), while in the latter case, interpolation is performed based on integer units of pixels.

[0423] When using one of multiple filter candidate groups, filter selection information can be explicitly generated or determined in mode, or can be determined based on encoding / decoding settings (such as interpolation accuracy, block size, shape, prediction mode, etc.). In this case, the unit for explicit generation can be a video, sequence, picture, slice, parallel block, block, etc.

[0424] For example, when an interpolation accuracy of 1 / 4 or more (1 / 2, 1 / 4) is used, an 8-tap Kalman filter can be applied to integer-unit reference pixels; when an interpolation accuracy of less than 1 / 4 and more than 1 / 16 (1 / 8, 1 / 16) is used, a 4-tap Gaussian filter can be applied to integer-unit reference pixels and interpolated reference pixels of more than 1 / 4; and when an interpolation accuracy of less than 1 / 16 (1 / 32, 1 / 64) is used, a 2-tap linear filter can be applied to integer-unit reference pixels and interpolated reference pixels of more than 1 / 16.

[0425] Alternatively, an 8-tap Kalman filter can be applied to blocks larger than 64×64, a 6-tap Wiener filter can be applied to blocks smaller than 64×64 and larger than 16×16, and a 4-tap Gaussian filter can be applied to blocks smaller than 16×16.

[0426] Alternatively, a 4-tap cubic filter may be applied to prediction modes with an angular difference of less than 22.5 degrees based on the vertical or horizontal mode, while a 4-tap Gaussian filter may be applied to prediction modes with an angular difference of 22.5 degrees or more.

[0427] In addition, multiple filter candidate groups can be composed of a 4-tap (4-tap) cubic filter, a 6-tap (6-tap) Wiener filter, and an 8-tap (8-tap) Kalman filter in some encoding / decoding settings, and can be composed of a 2-tap (2-tap) linear filter and a 6-tap (6-tap) Wiener filter in some encoding / decoding settings.

[0428] Figures 23a to 23c This is an exemplary diagram illustrating a method of performing interpolation on a fractional pixel basis in reference pixels configured according to an embodiment to which the present invention is applied.

[0429] See Figure 23a , a method for interpolating fractional pixels can be confirmed when using a single reference pixel level (ref_i) as reference pixels. Specifically, interpolation can be performed by applying a filter (the filter function is labeled int_func_1D) to pixels adjacent to the interpolation target pixel (labeled with x) (assuming that a filter is applied to integer pixels in this example). Since a single reference pixel level is used as reference pixels, interpolation can be performed using adjacent pixels on the same reference pixel level as the interpolation target pixel x.

[0430] See Figure 23b , a method for obtaining interpolated pixels in fractional units can be confirmed when two or more reference pixel levels (ref_i, ref_j, ref_k) are supported as reference pixels. Figure 23b When the reference pixel interpolation process is performed on the reference pixel level ref_j, the interpolation of the interpolation target pixel of the decimal unit can be performed by additionally using the other reference pixel levels ref_k and ref_i. Specifically, by interpolating the adjacent pixel a k ~h k 、a j ~h j 、a i ~h i Execute filtering (interpolation process, function int_func_1D) to obtain interpolation object pixels (x j and the pixel x at the position corresponding to the interpolation target pixel contained in other reference pixel levels (the corresponding position on each reference pixel level according to the direction of the prediction mode). k 、x i And the obtained 1st interpolation pixel x k 、x j 、x i Perform additional filtering (which may be filtering corresponding to weighted averages such as [1, 2, 1] / 4, [1, 6, 1] / 8, etc., which is not an interpolation process) to finally obtain the final interpolated pixel x on the reference pixel level ref_j. In this example, it is assumed that the pixel x on the other reference pixel level corresponding to the interpolation target pixel is k 、x i The case of fractional unit pixels that can be obtained through the interpolation process is described.

[0431] In the above example, a case is described in which a single interpolation pixel is obtained by filtering at each reference pixel level and a final difference pixel is obtained by performing additional filtering on the single interpolation pixel. However, it is also possible to obtain a final difference pixel by filtering adjacent pixels a at multiple reference pixel levels. k ~h k 、a j ~h j 、a i ~h i The final interpolated pixel is obtained in one go by filtering.

[0432] exist Figure 23b Of the three reference pixel levels supported in [ ], the level actually used as reference pixels can be ref_j. That is, in order to interpolate a reference pixel level composed of reference pixels, other reference pixel levels included in the candidate group can be used (for example, not composed of reference pixels means that the pixels in the corresponding reference pixel level are not suitable for prediction, but in the above case, the corresponding pixels are referenced when performing interpolation, so it can be used to be precise).

[0433] See Figure 23c , illustrates the case where both supported reference pixel levels are used as reference pixels. i d j 、e i 、e j ) constitutes the input pixel and performs filtering on the adjacent pixels to obtain the final interpolation pixel x. At this time, it is also possible to use Figure 23b The method shown is to obtain a first interpolation pixel at each reference pixel level and then perform additional filtering on the first interpolation pixel to obtain a final interpolation pixel x.

[0434] The above example is not limited to the reference pixel interpolation process, and can also be understood as a process combined with other processes of intra-frame prediction (such as a reference pixel filtering process, a prediction block generation process, etc.).

[0435] Figures 24a to 24b This is a first exemplary diagram for explaining an adaptive reference pixel filtering method according to an embodiment of the present invention.

[0436] Generally, the main purpose of the reference pixel filter unit can be to perform smoothing by using a low-pass filter (e.g., a 3-tap or 5-tap filter such as [1, 2, 1] / 4 or [2, 3, 6, 3, 2] / 16). However, other types of filters (e.g., a high-pass filter) can also be used depending on the purpose of the filter (e.g., sharpening). In this invention, the description will focus on the case of reducing distortion generated during the encoding / decoding process by performing filtering for the purpose of smoothing.

[0437] Reference pixel filtering can be determined based on the encoding / decoding settings. However, whether or not to apply filtering in batches may result in the inability to reflect the local characteristics of the image, because performing filtering based on the local characteristics of the image will be more conducive to improving encoding performance. Among them, the characteristics of the image can be judged not only based on the image type, color components, quantization parameters, encoding / decoding information of the current block (such as the size, shape, segmentation information, prediction mode, etc. of the current block), but also based on the encoding / decoding information of adjacent blocks and the combination of the encoding / decoding information of the current block and adjacent blocks. In addition, it can also be judged based on the reference pixel distribution characteristics (such as the dispersion of the reference pixel area, standard deviation, flat area, discontinuous area, etc.).

[0438] See Figure 24a When the image belongs to a category (category 0) based on a part of the encoding / decoding settings (for example, block size range A, prediction mode B, color component C, etc.), filtering may not be applied, and when the image belongs to a category (category 1) based on a part of the encoding / decoding settings (for example, prediction mode A of the current block, prediction mode B of the pre-set adjacent block, etc.), filtering may be applied.

[0439] See Figure 24b When the classification belongs to a part of the encoding / decoding settings (for example, the size A of the current block, the size B of the adjacent block, the prediction mode C of the current block, etc.) (category 0), filtering may not be applied; when the classification belongs to a part of the encoding / decoding settings (for example, the size A of the current block, the shape B of the current block, the size C of the adjacent block, etc.) (category 1), filtering can be performed using filter A; and when the classification belongs to a part of the encoding / decoding (for example, the parent block A of the current block, the parent block B of the adjacent block, etc.) (category 2), filtering can be performed using filter B.

[0440] Therefore, the application of filtering, the filter type, whether filter information is encoded (explicit / implicit), the number of filtering times, etc. can be determined based on the size of the current block and adjacent blocks, the prediction mode, the color components, etc., and the filter type can be classified based on the number of taps, the filter coefficients, etc. In this case, when the number of filtering times is two or more, the same filter can be applied multiple times or different filters can be applied respectively.

[0441] The above example could be a case where reference pixel filtering is pre-set based on image characteristics. In other words, filter-related information could be implicitly determined. However, inaccurate judgment of image characteristics, as described above, could negatively impact coding efficiency, so this aspect must be considered.

[0442] To prevent this situation, explicit settings can be made for reference pixel filtering. For example, information can be generated regarding whether filtering is applicable. In this case, filter selection information may not be generated when there is only one filter, but it can be generated when there are multiple candidate filter sets.

[0443] The above examples illustrate implicit and explicit settings related to reference pixel filtering. A hybrid approach can be employed, where explicit settings are used in some cases and implicit settings are used in other cases. Implicit means that information related to the reference pixel filter (e.g., whether filtering is applicable and filter type information) can be inferred from the decoder.

[0444] Figure 25 This is a second exemplary diagram for explaining an adaptive reference pixel filtering method according to an embodiment of the present invention.

[0445] See Figure 25 , it is possible to classify categories by using image characteristics confirmed by encoding / decoding information and adaptively perform reference pixel filtering according to the classified categories.

[0446] For example, filtering is applied when classified into category 0 and filter A is used when classified into category 1. Category 0 and category 1 can be an example of implicit reference pixel filtering.

[0447] Furthermore, when the classification is into category 2, filtering may not be applied or filter A may be applied. In this case, the information generated may be information on whether filtering is applied or not, but filter selection information is not generated.

[0448] Furthermore, when classified as Category 3, either Filter A or Filter B can be applied. In this case, the generated information can be filter selection information, and the application of the filter can be unconditional. In other words, when classified as Category 3, it can be understood that filtering must be performed, but the filter type needs to be selected.

[0449] Furthermore, when the classification is into category 4, filtering may not be applied, filter A may be applied, or filter B may be applied. In this case, the information generated may be information on whether filtering is applied or not and filter selection information.

[0450] In other words, it is possible to determine whether to perform explicit or implicit processing according to the category, and when performing explicit processing, each reference pixel filter-related candidate group setting can be adaptively constructed.

[0451] Regarding the above categories, the following examples can be considered.

[0452] First, for blocks larger than 64×64, one of <Filter Off>, <Filter On - Filter A>, <Filter On + Filter B>, or <Filter On + Filter C> can be implicitly determined based on the prediction mode of the current block. In this case, taking into account the distribution characteristics of reference pixels, an additional candidate can be <Filter On + Filter C>. That is, when the filter is on, filter A, filter B, or filter C can be applied.

[0453] In addition, for blocks smaller than 64×64 and larger than 16×16, one of <filter off>, <filter on + filter A>, and <filter on + filter B> can be implicitly determined according to the prediction mode of the current block.

[0454] Furthermore, for blocks smaller than 16×16, one of <Filter Off>, <Filter On + Filter A>, or <Filter On + Filter B> can be selected based on the prediction mode of the current block. In this case, the mode can be set to <Filter Off> in some prediction modes, while the selection of either <Filter Off> or <Filter On + Filter A> can be explicitly indicated in some other prediction modes, and the selection of either <Filter Off> or <Filter On + Filter B> can be explicitly indicated in some other prediction modes.

[0455] As an example of settings related to multiple reference pixel filters, when the reference pixels obtained in each filter (including the case where the filter is turned off in this example) are the same or similar, generating reference pixel filter information (such as reference pixel filter permission information, reference pixel filter information, etc.) may result in the generation of unnecessary duplicate information. For example, when the reference pixel distribution characteristics obtained in each filter (such as the value obtained by averaging, dispersion, etc. of each reference pixel and the threshold value) are different, <threshold>When the characteristics determined by comparison are the same or similar, the reference pixel filter related information can be omitted. When the reference pixel filter related information is omitted, filtering can be applied using a pre-set method (e.g., filtering off). After receiving the intra-picture prediction information, the decoder can determine whether to receive the reference pixel filter related information in the same manner as the encoder, and can determine whether to accept the reference pixel filter related information based on the above determination.

[0456] Assuming that explicit information related to reference pixel filtering is generated, indication information (adaptive_ref_filter_enabled_flag in this example) for enabling adaptive reference pixel filtering can be generated in units such as video, sequence, picture, slice, or tile.

[0457] When the above indication information indicates that adaptive reference pixel filtering is allowed (adaptive_ref_filter_enabled_flag=1 in this example), adaptive reference pixel filtering enable information (adaptive_ref_filter_flag in this example) can be generated in units such as picture, slice, tile, or block.

[0458] When the above permission information represents adaptive reference pixel filtering (adaptive_ref_filter_flag=1 in this example), reference pixel filtering related information (such as reference pixel filter selection information, etc., ref_filter_idx in this example) can be generated on units such as images, strips, parallel blocks, and blocks.

[0459] At this time, when adaptive reference pixel filtering is not allowed or cannot be applied, filtering can be performed on the reference pixel according to predetermined settings (as mentioned above, whether filtering is applicable and the type of filtering are determined in advance based on image encoding / decoding information, etc.).

[0460] Figures 26a to 26b This is an exemplary diagram illustrating a case where one reference pixel level is used in reference pixel filtering according to one embodiment of the present invention.

[0461] See Figure 26a It can be confirmed that interpolation is performed by applying filtering (referred to as smt_func_1 function) to the object pixel d and the pixels a, b, c, e, f, and g adjacent to the object pixel d among the pixels included in the reference pixel level ref_i.

[0462] Figure 26a Typically, sequential filtering is applied, but multiple filtering steps may be applied. For example, two filtering steps may be applied to reference pixels (a*, b*, c*, d*, etc. in this example) obtained by applying one filtering step.

[0463] See Figure 26b , a pixel e* that has been filtered (referred to as a smt_func_2 function) can be obtained by performing linear interpolation proportional to the distance (e.g., the distance z from a) on the pixels located on both sides with the target pixel e as the center. The pixels located on both sides can be pixels located at the ends of adjacent pixels in a block consisting of the upper block, left block, upper block + upper right block, left block + lower left block, upper left block + upper block + upper right block, upper left block + left block + lower left block, and upper left block + left block + upper block + lower left block + upper right block of the current block. Figure 26b It can be reference pixel filtering performed according to reference pixel distribution characteristics.

[0464] exist Figures 26a to 26b 2 illustrates a case where pixels in the same reference pixel hierarchy as the reference pixel to be filtered are used for reference pixel filtering. In this case, the type of filter used in reference pixel filtering can be the same or different depending on the reference pixel hierarchy.

[0465] Furthermore, when multiple reference pixel levels are used, when reference pixel filtering is performed in a portion of the reference pixel levels, not only pixels in the same reference pixel level but also pixels in different reference pixel levels can be used.

[0466] Figure 27 This is an exemplary diagram illustrating the use of multiple reference pixel levels in reference pixel filtering according to one embodiment of the present invention.

[0467] See Figure 27 , first, filtering can be performed on the reference pixel levels ref_k and ref_i using pixels contained in the same reference pixel level. That is, filtering can be performed on the reference pixel level ref_k using pixels contained in the same reference pixel level. k and adjacent pixel a k to g k Perform filtering (defined as function smt_func_1D) to obtain the filtered pixel d k *, it is also possible to obtain the object pixel d by i and adjacent pixel a i to g i Perform filtering (defined as function smt_func_1D) to obtain the filtered pixel d i *.

[0468] In addition, when performing reference pixel filtering on the reference pixel level ref_j, not only the same reference pixel level ref_j but also pixels included in other reference pixel levels ref_i and ref_k that are spatially adjacent to the reference pixel level ref_j can be used. Specifically, it is possible to filter the reference pixel by taking the target pixel d j The center is the spatially adjacent pixel c k d k 、e k 、c j 、e j 、c i d i 、e i (ie, it can be a filter with a 3×3 square mask) Apply filtering (defined as function smt_func_2D) to obtain the interpolated pixel d j *. However, it is not limited to a 3×3 square form, and a 5×2 rectangular form (b k 、c k d k 、e k 、f k 、b j 、c j 、e j 、f j ), 3×3 diamond shape (d k 、c j 、e j d i )、5×3 cross shape (d k 、b j 、c j 、e j 、f j d i ) and other masked filters.

[0469] Among them, the reference pixel level is as above Figure 18 to Figure 2 2, etc., is composed of pixels included in adjacent blocks adjacent to the current block and close to the boundary of the current block. Considering the above-mentioned aspects, in the reference pixel level ref_k and the reference pixel level ref_i, filtering using pixels included in the same reference pixel level can be applied using a one-dimensional mask-like filter using pixels horizontally or vertically adjacent to the interpolation target pixel. However, in the reference pixel level ref_j, the reference pixel d j The interpolated pixel can be obtained by applying a filter in the form of a two-dimensional mask using all the pixels that are spatially adjacent to each other above, below, left, and right.

[0470] Furthermore, in each reference pixel level, reference pixel filtering can be applied twice to reference pixels that have already been filtered once. For example, reference pixel filtering can be performed once using reference pixels included in each reference pixel level ref_k, ref_j, and ref_i. Subsequently, in the reference pixel levels (referred to as ref_k*, ref_j*, and ref_i*) that have already been filtered once, reference pixel filtering can be performed using not only the reference pixel level itself but also reference pixels in other reference pixel levels.

[0471] The prediction block generator can generate a prediction block based on at least one intra-frame prediction mode (simply referred to as a prediction mode) and can use reference pixels based on the prediction mode. In this case, the prediction block can be generated by performing extrapolation, interpolation, or DC copying of the reference pixels according to the prediction mode. Extrapolation can be applied to directional intra-frame prediction modes, while the remaining non-directional intra-frame prediction modes can be applied.

[0472] In addition, when copying reference pixels, more than one predicted pixel can be generated by copying one reference pixel to multiple pixels within the prediction block, and more than one predicted pixel can be generated by copying more than one reference pixel, and the number of copied reference pixels can be equal to or less than the number of copied predicted pixels.

[0473] Furthermore, a prediction block is typically generated for a prediction in an intra-frame prediction mode. However, it is also possible to obtain multiple prediction blocks and then generate a final prediction block by applying a weighted sum to the obtained multiple prediction blocks. The multiple prediction blocks may refer to prediction blocks obtained based on a reference pixel hierarchy.

[0474] The prediction mode determination unit of the encoding device performs a process for selecting the optimal mode from a group of multiple prediction mode candidates. Typically, the optimal mode in terms of coding cost can be determined using rate-distortion techniques, which predict block distortion (e.g., distortion, Sum of Absolute Difference (SAD), Sum of Square Difference (SSD) between the current block and the reconstructed block) and the amount of bits generated by the prediction mode. The prediction block generated based on the prediction mode determined by this process is transmitted to the subtraction unit and the addition unit. (In this case, the decoding device can obtain information indicating the optimal prediction mode from the encoding device, so the process of selecting the optimal prediction mode can be omitted.)

[0475] The prediction mode encoding unit of the encoding device can encode the optimal intra-frame prediction mode selected by the prediction mode determination unit. In this case, index information indicating the optimal prediction mode can be directly encoded, or prediction information related to the prediction mode (e.g., the difference between the predicted prediction mode index and the prediction mode index of the current block) can be encoded after the optimal prediction mode is predicted using prediction modes available from other neighboring blocks. The former case is applicable to chrominance components, while the latter case is applicable to luma components.

[0476] When predicting and encoding the best prediction mode for the current block, the prediction value (or prediction information) of the prediction mode can be called the most likely mode (MPM). At this time, the most likely mode (MPM) refers to the prediction mode with the highest probability of becoming the best prediction mode for the current block, which can be composed of a pre-set prediction mode (such as mean (DC), plane (Planar), vertical, horizontal, diagonal mode, etc.) or a prediction mode of a spatially adjacent block (such as the left, upper, upper left, upper right, lower left block, etc.). Among them, the diagonal mode is pointing to the right (Diagonal up right), right (Diagonal down right), left (Diagonal down left), and can be the same as Figure 17 The patterns corresponding to patterns No. 2, No. 18, and No. 34 in .

[0477] In addition, it is also possible to add a mode derived from a prediction mode included in a most likely mode (MPM) candidate group, which is a set of prediction modes consisting of most likely modes (MPMs), to the most likely mode (MPM) candidate group. In a directional mode, a prediction mode whose index interval with a prediction mode included in the most likely mode (MPM) candidate group is equal to a preset value can be added to the most likely mode (MPM) candidate group. For example, when the mode included in the most likely mode (MPM) candidate group is Figure 17 In the case of pattern No. 10, the derived pattern can be equivalent to pattern No. 9, No. 11, No. 8, No. 12, etc.

[0478] The above example is equivalent to the situation where the most likely mode (MPM) candidate group is composed of multiple modes. The composition of the most likely mode (MPM) candidate group (for example, the number of prediction modes included in the most likely mode (MPM), the priority of the composition) is determined according to the encoding / decoding settings (for example, the prediction mode candidate group, image type, block size, block shape, etc.) and can include at least one mode composition.

[0479] The priority order of the prediction modes included in the most probable mode (MPM) candidate group can be set. The order of the prediction modes included in the most probable mode (MPM) candidate group can be determined based on the set priority order, and the formation of the most probable mode (MPM) candidate group can be completed when the added prediction modes reach a predetermined number. The priority order can be set to the order of prediction modes of blocks spatially adjacent to the current block to be predicted, predetermined prediction modes, and modes derived from prediction modes included earlier in the most probable mode (MPM) candidate group, but is not limited to this.

[0480] Specifically, the priority order can be set in the order of left-upper-left-lower-right-upper-left blocks in spatially adjacent blocks, and the priority order can be set in the order of mean (DC)-planar-vertical-horizontal mode in the pre-set prediction mode. Then, the index value ( Figure 17 The prediction mode obtained by adding +1, -1, etc. (integer value) to the prediction mode number is included in the most probable mode (MPM) candidate group. As one of the above examples, the priority order can be set in the order of left side - top side - mean (DC) - planar - bottom left - top right - top left - (spatial neighboring block mode) + 1 - (spatial neighboring block mode) - 1 - horizontal - vertical - diagonal.

[0481] In the above example, a case where the priority order of the most probable mode (MPM) candidate group is fixed is described. However, the priority order can also be adaptively determined according to the shape and size of the block.

[0482] When encoding the prediction mode of the current block using the most probable mode (MPM), information (eg, most_probable_mode_flag) regarding whether the prediction mode is consistent with the most probable mode (MPM) can be generated.

[0483] When the most probable mode (MPM) is consistent (e.g., most_probable_mode_flag = 1), most probable mode (MPM) index information (e.g., mpm_idx) can be additionally generated based on the composition of the most probable mode (MPM). For example, when the most probable mode (MPM) consists of a single prediction mode, additional most probable mode (MPM) index information may not be generated, while when the most probable mode (MPM) consists of multiple prediction modes, index information corresponding to the prediction mode of the current block in the most probable mode (MPM) candidate group can be generated.

[0484] When it is inconsistent with the most probable mode (MPM) (for example, most_probable_mode_flag=0), it is possible to generate non-most probable mode (non-MPM) index information (for example, non_mpm_idx) corresponding to the prediction mode of the current block in the remaining prediction mode candidate group (referred to as the non-most probable mode (non-MPM) candidate group) after excluding the most probable mode (MPM) candidate group from the supported intra-picture prediction modes. This can be an example of a case where the non-most probable mode (non-MPM) is constituted into a group.

[0485] When the non-most probable mode (non-MPM) candidate group is composed of a plurality of groups, information related to which group the prediction mode of the current block is included in can be generated. For example, when the non-most probable mode (non-MPM) is composed of two groups A and B and the prediction mode of the current block is consistent with the prediction mode of group A (for example, non_mpm_A_flag=1), index information corresponding to the prediction mode of the current block can be generated in the candidate group of group A, and when they are inconsistent (for example, non_mpm_A_flag=0), index information corresponding to the prediction mode of the current block can be generated in the remaining prediction mode candidate groups (or the candidate group of group B). As shown in the above example, the non-most probable mode (non-MPM) can be composed of a plurality of groups, and the number of groups can be specified according to the prediction mode candidate groups. For example, when the number of prediction mode candidate groups is 35 or less, it can be 1, and in other cases, it can be 2.

[0486] At this time, the specific A group can be composed of modes that are determined to have a high probability of being consistent with the prediction mode of the current block after the most probable mode (MPM) candidate group. For example, the next prediction mode that is not included in the most probable mode (MPM) candidate group can be included in the A group or the directional mode with a certain interval can be included in the A group.

[0487] As in the above example, when the non-MPM is composed of a plurality of groups, the effect of reducing the mode coding bit amount can be achieved in the case where the number of prediction modes is large and the prediction mode of the current block is not consistent with the MPM.

[0488] In encoding (or decoding) the prediction mode of the current block using the MPM, a binarization table that is suitable for each prediction mode candidate group (e.g., the MPM candidate group, the non-MPM candidate group, etc.) can be generated individually, and different binarization methods can be applied individually according to each candidate group.

[0489] In the above example, the terms such as the MPM candidate group, the non-MPM candidate group, etc. are only a part of the terms used in the present application and are not limited thereto. Specifically, only the information indicating which category belongs to when the current intra prediction mode is classified into a plurality of categories and the mode information within the corresponding category, terms such as the 1st MPM candidate group and the 2nd MPM candidate group can be used instead.

[0490] Figure 28 is a block diagram for explaining an intra prediction mode encoding / decoding method to which an embodiment of the present application is applied.

[0491] Referring to Figure 28 First, mpm_flag is obtained (S10). Next, whether it is consistent with the first most probable mode (MPM) (indicated by mpm_flag) is checked (S11). If it is consistent, the most probable mode (MPM) index information (mpm_idx) is checked (S12). If it is inconsistent with the most probable mode (MPM), rem_mpm_flag is obtained (S13). Next, whether it is consistent with the second most probable mode (MPM) (indicated by rem_mpm_flag) is checked (S14). If it is consistent, the second most probable mode (MPM) index information (rem_mpm_idx) is checked (S16). If it is inconsistent with the second most probable mode (MPM), the index information (rem_mode_idx) of the candidate group consisting of the remaining prediction modes is checked (S15). In this example, the case where the index information generated based on the consistency of the two most likely modes (MPMs) is expressed using the same syntax elements is described, but other mode coding settings (such as binarization method) can also be applied, and different index information can also be set for processing.

[0492] In a video decoding method according to an embodiment of the present invention, intra-frame prediction can be configured as follows. The intra-frame prediction performed by the prediction unit can include a prediction mode decoding step, a reference pixel configuration step, and a prediction block generation step. Furthermore, a video decoding device can include a prediction mode decoding unit, a reference pixel configuration unit, and a prediction block generation unit for performing the prediction mode decoding step, the reference pixel configuration step, and the prediction block generation step. The above-described process may omit portions thereof or add other steps, and may also be performed in a different order than described above.

[0493] Because the reference pixel construction unit and prediction block generation unit of the image decoding device can play the same role as those in the image encoding device, detailed descriptions thereof will be omitted here, and the prediction mode decoding unit can reversely use the method used in the prediction mode encoding unit.

[0494] Next, we will combine Figures 29 to 31 Various embodiments of intra-frame prediction based on reference pixel composition in a decoding device are described. The descriptions of reference pixel hierarchy support and reference pixel filtering methods described above in conjunction with the accompanying drawings should be interpreted as being equally applicable to the decoding device. To avoid duplication, the detailed descriptions thereof will be omitted.

[0495] Figure 29 This is a first exemplary diagram for explaining the bit stream structure of intra-frame prediction based on reference pixel structure.

[0496] exist Figure 29 In the first example diagram in , the premise is to support multiple reference pixel levels, use at least one reference pixel level among the supported reference pixel levels as a reference pixel, support multiple candidate groups related to reference pixel filtering and select a filter from them.

[0497] After the encoder constructs a pixel candidate group using multiple reference pixel levels (in this example, the reference pixel generation process has been completed), it constructs reference pixels using at least one reference pixel level, and then applies reference pixel filtering and reference pixel interpolation. In this case, multiple candidate groups related to reference pixel filtering are supported.

[0498] Next, a process is performed to select the optimal mode from the candidate prediction mode set. After the optimal prediction mode is determined, a prediction block based on the corresponding mode is generated and passed to the subtraction operation unit. Then, a process is performed to encode information related to intra-frame prediction. In this example, it is assumed that the reference pixel level and reference pixel filtering are implicitly determined based on the encoding information.

[0499] In the decoder, intra-frame prediction-related information (e.g., prediction mode) is reconstructed and a prediction block based on the reconstructed prediction mode is generated, which is then passed to the subtraction unit. At this point, the reference pixel hierarchy and reference pixel filtering used to generate the prediction block are implicitly determined.

[0500] See Figure 29 , a bitstream can be constructed using an intra-frame prediction mode (intra_mode) ( S20 ). At this time, the reference pixel level ref_idx and reference pixel filter category ref_filter_idx supported (or used) in the current block can be implicitly determined based on the intra-frame prediction mode (determined as Category A or B, respectively, S21-S22 ). At this time, additional encoding / decoding information (e.g., image type, color component, block size, and shape) can be considered.

[0501] Figure 30 This is a second exemplary diagram for explaining the bit stream structure of intra-frame prediction based on reference pixel structure.

[0502] exist Figure 30 In the second example diagram in FIG, it is assumed that multiple reference pixel levels are supported and one of the multiple reference pixel levels is used as a reference pixel. In addition, it is assumed that multiple candidate groups related to reference pixel filtering are supported and a filter is selected from them. Figure 29 The difference is that the information related to the selection is explicitly generated by the encoding device.

[0503] After the encoder determines that multiple reference pixel levels are supported, a reference pixel hierarchy is used to construct reference pixels, and then reference pixel filtering and reference pixel interpolation are applied. In this case, multiple filtering methods related to reference pixel filtering are supported.

[0504] When determining the optimal prediction mode for the current block in the encoder, the process of selecting the optimal reference pixel level and the optimal reference pixel filter for each prediction mode can also be additionally considered. After determining the optimal prediction mode, reference pixel level, and reference pixel filter for the current block, the prediction block generated based on these is passed to the subtraction unit to perform the encoding process of the intra-frame prediction related information.

[0505] The decoder reconstructs intra-frame prediction-related information (such as the prediction mode, reference pixel level, and reference pixel filtering information) and uses this information to generate a prediction block, which is then passed to the subtraction unit. The reference pixel level and reference pixel filtering used to generate the prediction block follow the settings determined by the information transmitted from the encoder.

[0506] See Figure 30 The decoder uses the intra-prediction mode information (intra_mode) included in the bitstream to determine the optimal prediction mode for the current block (S30) and confirms whether multiple reference pixel levels are supported (multi_ref_flag) (S31). When multiple reference pixel levels are supported, the decoder confirms the reference pixel level selection information (ref_idx) (S32) to determine the reference pixel level that can be used for intra-prediction. If multiple reference pixel levels are not supported, the process of obtaining the reference pixel level selection information (ref_idx) (S32) can be omitted.

[0507] Next, whether adaptive reference pixel filtering is supported (adap_ref_smooth_flag) is confirmed (S33), and when adaptive reference pixel filtering is supported, a filtering method for reference pixels is determined using reference pixel filter information (ref_filter_idx) (S34).

[0508] Figure 31 This is a third exemplary diagram for explaining the bit stream structure of intra-frame prediction based on reference pixel structure.

[0509] exist Figure 31 In the third example diagram in FIG, it is assumed that multiple reference pixel levels are supported and one of the multiple reference pixel levels is used. In addition, it is assumed that multiple candidate groups related to reference pixel filtering are supported and a filter is selected from them. Figure 30 The difference is that the selection information is adaptively generated.

[0510] After the reference pixels are constructed using one of the multiple reference pixel levels supported by the encoder, reference pixel filtering and reference pixel interpolation are applied. At this time, multiple filters related to reference pixel filtering are supported.

[0511] When performing the process for selecting the best mode from the multiple prediction mode candidates, the process for selecting the best reference pixel level in each prediction mode and the process for selecting the best reference pixel filtering can also be considered. After the best prediction mode and the reference pixel level and reference pixel filtering are determined, the prediction block is generated based thereon and passed to the subtraction operation unit, and then the encoding process for the intra prediction related information is performed.

[0512] At this time, the repeatability of the generated prediction block is confirmed, and when it is the same or similar to the prediction block obtained using other reference pixel levels, the selection information related to the best reference pixel level is omitted and the preset reference pixel level is used. At this time, the preset reference pixel level can be the level most adjacent to the current block.

[0513] For example, the repeatability or not can be judged based on the difference value (distortion value) between the prediction block generated by ref_0 in FIG. 8 and the prediction block generated by ref_1. Figure 19c When the above difference value is less than a preset threshold value, it is determined that the prediction block has repeatability, otherwise it is determined that the prediction block has no repeatability. At this time, the above threshold value can be adaptively determined according to the quantization parameter and the like.

[0514] In addition, the best reference pixel filtering information also confirms the repeatability of the prediction block, and when it is the same or similar to the prediction block obtained by applying other reference pixel filtering, the reference pixel filtering information is omitted and the preset reference pixel filtering is applied.

[0515] For example, the repeatability or not of the prediction block obtained by filter A (3-tap filter in this example) and the prediction block obtained by filter B (5-tap filter in this example) is judged based on the difference value between them. At this time, the difference value can also be compared with a preset threshold value, and when it is smaller, it is determined that the prediction block has repeatability. When the prediction block has repeatability, the prediction block can be generated by the preset reference pixel filtering method. Among them, the preset reference pixel filtering can be a filtering method with less number of taps or lower complexity, including the case of omitting the filtering usage.

[0516] The intra prediction related information (e.g. prediction mode and reference pixel level, reference pixel filter information, etc.) is reconstructed in the decoder and passed to the subtraction operation section after generating the prediction block. At this time, the reference pixel level information and the reference pixel filter used to generate the prediction block comply with the settings determined from the information transmitted from the encoder, and the decoder can comply with the pre-set method when the redundancy exists after directly confirming the redundancy (without passing through the syntax element).

[0517] Referring to Figure 31 , the decoder first confirms the intra prediction mode information (intra_mode) of the current block (S40), and confirms whether the support of multiple reference pixel levels (multi_ref_flag) (S41). When supporting multiple reference pixel levels, the redundancy check of the prediction block based on the supported multiple reference pixel levels is performed (represented by the ref_check process, S42), and when the redundancy check result is that the prediction block has no redundancy (redund_ref = 0, S43), the selection information of the reference pixel level (ref_idx) is referred from the bit stream (S44) and the optimal reference pixel level is determined.

[0518] Next, the support of adaptive reference pixel filtering (adap_ref_smooth_flag) is confirmed (S45), and when supporting adaptive reference pixel filtering, the redundancy check of the prediction block of the supported multiple reference pixel filtering methods (represented by the ref_check process, S46). When there is no redundancy of the prediction block (redund_ref = 0, S47), the selection information of the reference pixel filtering method (ref_filter_idx) is referred from the bit stream (S48) and the optimal reference pixel filtering method is determined.

[0519] At this time, the redund_ref in the drawing is a value for indicating the redundancy check result, and when it is 0, it means that there is no redundancy.

[0520] In addition, the decoder can perform the intra prediction using the pre-set reference pixel level and the pre-set reference pixel filtering method when the prediction block has redundancy.

[0521] Figure 32 is a flowchart illustrating an image decoding method supporting multiple reference pixel levels according to an embodiment of the present application.

[0522] Referring to Figure 32 , an image decoding method that supports multiple reference pixel levels may include: a step of confirming whether multiple reference pixel levels are supported through a bitstream (S100); when multiple reference pixel levels are supported, a step of determining the reference pixel level to be used in the current block by referring to syntax information contained in the above-mentioned bitstream (S110); a step of constructing reference pixels using pixels contained in the determined reference pixel level (S120); and a step of performing intra-frame prediction of the above-mentioned current block using the constructed reference pixels (S130).

[0523] After the step ( S100 ) of confirming whether multiple reference pixel levels are supported, the method may further include: confirming whether an adaptive reference pixel filtering method is supported through a bitstream.

[0524] After the step of confirming whether multiple reference pixel levels are supported ( S100 ), the method may further include: if multiple reference pixel levels are not supported, constructing reference pixels using a preset reference pixel level.

[0525] The methods applicable to the present invention can be implemented in the form of program instructions executable by various computing means and recorded on a computer-readable medium. The computer-readable medium can contain program instructions, data files, data structures, etc., alone or in combination. The program instructions recorded on the computer-readable medium can be program instructions specially designed for the present invention or program instructions commonly known and available to computer software practitioners.

[0526] Examples of computer-readable media include specially configured hardware devices for storing and executing program instructions, such as read-only memory (ROM), random access memory (RAM), and flash memory. Examples of program instructions include not only machine code generated by a compiler, but also high-level language code that can be executed in a computer using an interpreter. The aforementioned hardware devices can be composed of at least one software module for executing actions applicable to the present invention, and vice versa.

[0527] Furthermore, the above-mentioned methods or apparatuses may be combined or separated in whole or in part in terms of their configurations or functions.

[0528] While the foregoing description is based on preferred embodiments of the present invention, skilled practitioners in the relevant technical field should understand that various modifications and changes can be made to the present invention without departing from the scope of the invention and the scope of the scope of the appended claims.< / threshold> < / array> < / postfilter> < / display> < / sizing> < / y>

Claims

1. A method for decoding a video signal using a decoding device, the method comprising: dividing a current image in the video signal into a plurality of sub-images based on the segmentation information; Splitting a current block in the current image into a plurality of sub-blocks; as well as decoding each of said sub-blocks, wherein the current block is divided based on at least one of quadtree division, binary tree division, and ternary tree division, The segmentation information includes quantity information, position information, and size information, wherein the quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the upper left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. wherein both the position information and the size information are expressed in units of maximum coding units, wherein the plurality of sub-images include a first sub-image and a second sub-image, wherein a portion of the right boundary of the first sub-image is adjacent to the left boundary of the second sub-image, wherein an upper boundary of the first sub-image and an upper boundary of the second sub-image are discontinuous, and a lower boundary of the first sub-image and a lower boundary of the second sub-image are continuous, wherein one of independence decoding or dependency decoding is applied to the sub-image, The independence decoding represents a decoding scheme in which, when decoding the sub-image in the current image, no reference is allowed to other sub-images except for the co-located sub-image in the sub-image belonging to a different image than the current image. The dependency decoding represents a decoding scheme in which, when decoding the sub-image in the current image, reference to the other sub-images in the different image is allowed. Wherein, the first flag is obtained from the sequence parameter set in the bit stream, wherein the first flag indicates whether the independence decoding or the dependency decoding is used for the sub-image, wherein the first flag having a first value indicates that the independence decoding applies to the sub-image, and Wherein the first flag having a second value indicates that the dependency decoding applies to the sub-image.

2. The method according to claim 1, further comprising: Whether loop filtering across the boundary between the sub-images is applicable is determined based on a second flag obtained from the bitstream.

3. The method according to claim 2, wherein: The maximum coding unit refers to a basic coding unit having a maximum size among coding units predefined in the decoding apparatus.

4. The method according to claim 3, wherein: The sub-image is quadrilateral in shape.

5. A method for encoding a video signal using an encoding device, comprising: Splitting a current image in a video signal into a plurality of sub-images; Splitting a current block in the current image into a plurality of sub-blocks; as well as encoding each of said sub-blocks, wherein the current block is divided based on at least one of quadtree division, binary tree division, and ternary tree division, The segmentation information is the basis for segmenting the current image. The segmentation information includes quantity information, position information, and size information, wherein the quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the upper left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. wherein both the position information and the size information are encoded in units of maximum coding units, wherein the plurality of sub-images include a first sub-image and a second sub-image, A portion of the right side border of the first sub-image is adjacent to the left side border of the second sub-image, wherein an upper boundary of the first sub-image and an upper boundary of the second sub-image are discontinuous, and a lower boundary of the first sub-image and a lower boundary of the second sub-image are continuous, wherein one of independence coding or dependency coding is applied to the sub-images, The independence coding represents a coding scheme in which, when coding the sub-image in the current image, no reference is allowed to other sub-images except for the co-located sub-image in the sub-image belonging to a different image than the current image. The dependency coding represents a coding scheme in which, when encoding the sub-image in the current image, reference to the other sub-images in the different image is allowed. The first flag is encoded into the sequence parameters set in the bitstream. wherein the first flag is encoded according to whether the independence coding or the dependency coding is used for the sub-image, wherein, in response to the independence coding being applied to the sub-image, the first flag has a first value, and Wherein, in response to the dependency encoding being applied to the sub-image, the first flag has a second value.

6. A method for transmitting a bit stream generated by an encoding method, the encoding method comprising: Splitting a current image in a video signal into a plurality of sub-images; Splitting a current block in the current image into a plurality of sub-blocks; as well as encoding each of said sub-blocks, wherein the current block is divided based on at least one of quadtree division, binary tree division, and ternary tree division, The segmentation information is the basis for segmenting the current image. The segmentation information includes quantity information, position information, and size information, wherein the quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the upper left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. wherein both the position information and the size information are expressed in units of maximum coding units, wherein the plurality of sub-images include a first sub-image and a second sub-image, A portion of the right side border of the first sub-image is adjacent to the left side border of the second sub-image, wherein an upper boundary of the first sub-image and an upper boundary of the second sub-image are discontinuous, and a lower boundary of the first sub-image and a lower boundary of the second sub-image are continuous, wherein one of independence coding or dependency coding is applied to the sub-images, The independence coding represents a coding scheme in which, when coding the sub-image in the current image, no reference is allowed to other sub-images except for the co-located sub-image in the sub-image belonging to a different image than the current image. The dependency coding represents a coding scheme in which, when encoding the sub-image in the current image, reference to the other sub-images in the different image is allowed. Wherein, the first flag is included in the sequence parameters set in the data stream, wherein the first flag is encoded according to whether the independence coding or the dependency coding is used for the sub-image, wherein, in response to the independence coding being applied to the sub-image, the first flag has a first value, and Wherein, in response to the dependency coding being applied to the sub-image, the first flag has a second value.

Citation Information

Patent Citations

  • Image decoding method and image decoding device

    JP6074743B2

  • Efficient scalable coding concept

    US20150304667A1