Image decoding method and apparatus using a partition unit including an additional area

By introducing an additional region and multiple reference pixel levels into the image decoding method, the problem of low image coding efficiency in the existing technology is solved, and more efficient image compression and accurate in-frame prediction are achieved.

CN116248868BActive Publication Date: 2025-11-25INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310305097.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-16
Filing Date
2018-07-03
Publication Date
2025-11-25
Estimated Expiration
2038-07-03

AI Technical Summary

Technical Problem

In existing image coding methods, parallel blocks coded independently cannot effectively utilize temporally and spatially adjacent image data as references, resulting in low coding efficiency. Furthermore, in-frame prediction relies on the nearest neighbor pixel reference, which may not be applicable.

Method used

The image decoding method employs segmented units that include appended regions. It obtains syntax elements through bitstream, sets appended regions, and references blocks within the appended regions during decoding. It supports multiple reference pixel levels and adaptive reference pixel filtering, thereby improving image compression efficiency and in-frame prediction accuracy.

Benefits of technology

It improves image compression efficiency and in-frame prediction accuracy. Through multiple reference pixel levels and adaptive filtering, it adapts to different image characteristics and improves compression efficiency during encoding/decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116248868B_ABST
    Figure CN116248868B_ABST
Patent Text Reader

Abstract

Disclosed is a method and apparatus for decoding an image using a partition unit including an additional area. The method for decoding an image using a partition unit including an additional area includes the steps of partitioning an encoded image included in a received bitstream into at least one partition unit by referring to syntax elements obtained from the bitstream; setting an additional area for the at least one partition unit; and decoding the encoded image based on the partition unit after the additional area is set. Thus, the efficiency of image encoding can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of application number 2018800450230, filed on July 3, 2018, entitled "Image Decoding Method and Apparatus Utilizing Segmentation Units Including Additional Regions". Technical Field

[0002] This invention relates to an image decoding method and apparatus that utilizes segmentation units including additional regions, and more particularly to a technique that improves encoding efficiency by setting additional regions (top, bottom, left, and right) on segmentation units such as parallel tiles within an image and simultaneously referencing image data in the additional regions during encoding. Background Technology

[0003] In recent years, the demand for multimedia data such as video on the Internet has been increasing dramatically. However, the current pace of channel bandwidth development still necessitates an effective method to compress the rapidly increasing volume of multimedia data. To this end, the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / ISE and the Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) are working tirelessly through collaborative research to develop more efficient video compression standards.

[0004] Furthermore, when performing independent encoding on images, individual segmentation units, including parallel tiles, are typically encoded independently, thus creating the problem of not being able to reference image data from other temporally and spatially adjacent segmentation units.

[0005] Therefore, a solution is needed that can maintain existing parallel processing based on independent coding while also using adjacent image data as a reference.

[0006] Furthermore, intra-frame prediction based on existing image encoding / decoding methods uses the nearest neighbor pixel to the current block as a reference pixel. However, depending on the type of image, using the nearest neighbor pixel as a reference pixel may not be a viable approach.

[0007] Therefore, there is a need for a method that can improve in-frame prediction efficiency by using a different reference pixel configuration than existing methods. Summary of the Invention

[0008] Technical issues

[0009] In order to solve the existing problems as described above, the object of the present invention is to provide an image decoding apparatus and method utilizing segmentation units including additional regions.

[0010] In order to solve the existing problems as described above, another object of the present invention is to provide an image encoding apparatus and method utilizing segmentation units including additional regions.

[0011] In order to solve the existing problems as described above, the object of the present invention is to provide an image decoding method that supports multiple reference pixel levels.

[0012] In order to solve the existing problems as described above, another object of the present invention is to provide an image decoding apparatus that supports multiple reference pixel levels.

[0013] Technical solution

[0014] To achieve the objectives described above, one aspect of the present invention provides an image decoding method utilizing segmentation units that include additional regions.

[0015] The image decoding method using segmentation units including appended regions can include: a step of segmenting an encoded image contained in a received bitstream into at least one segmentation unit by referring to a syntax element obtained from the received bitstream; a step of setting an appended region for the at least one segmentation unit; and a step of decoding the encoded image based on the segmentation unit after setting the appended region.

[0016] The aforementioned steps for decoding the encoded image may include: determining a reference block related to the current block in the encoded image that needs to be decoded, based on information contained in the bitstream that indicates the possibility of reference.

[0017] The aforementioned reference block can be a block located at a position that overlaps with an additional region defined in the segmentation unit to which the aforementioned reference block belongs.

[0018] In order to achieve the objectives described above, another aspect of the present invention provides an image decoding method that supports multiple reference pixel levels.

[0019] An image decoding method supporting multiple reference pixel levels may include: a step of confirming whether multiple reference pixel levels are supported via a bitstream; when multiple reference pixel levels are supported, a step of determining the reference pixel levels to be used in the current block by referring to the syntax information contained in the bitstream; a step of constructing reference pixels using the pixels contained in the determined reference pixel levels; and a step of performing in-frame prediction of the current block using the constructed reference pixels.

[0020] The method may further include, after the step of confirming whether multiple reference pixel levels are supported, a step of confirming whether an adaptive reference pixel filtering method is supported via a bitstream.

[0021] The method may further include, after confirming whether multiple reference pixel levels are supported, a step of constructing reference pixels using pre-defined reference pixel levels when multiple reference pixel levels are not supported.

[0022] Technical effect

[0023] When the image decoding method and apparatus of the present invention, which utilizes segmentation units including additional regions as described above, are used, the image compression efficiency can be improved because there is a large amount of image data that can be used as a reference.

[0024] When the image decoding method and apparatus supporting multiple reference pixel levels of the present invention as described above are used, the accuracy of in-frame prediction can be improved because multiple reference pixels can be utilized.

[0025] Furthermore, in this invention, because adaptive reference pixel filtering is supported, optimal reference pixel filtering can be performed based on the characteristics of the image.

[0026] In addition, it can improve the compression efficiency during image encoding / decoding. Attached Figure Description

[0027] Figure 1 This is a conceptual diagram illustrating an image encoding and decoding system to which embodiments of the present invention are applied;

[0028] Figure 2 This is a block diagram illustrating an image encoding apparatus to which one embodiment of the present invention is applied;

[0029] Figure 3 This is a configuration diagram illustrating an image decoding apparatus to which one embodiment of the present invention is applied;

[0030] Figures 4a to 4d This is a conceptual diagram used to illustrate a projection format applicable to one embodiment of the present invention;

[0031] Figures 5a to 5c This is a conceptual diagram illustrating a surface configuration applicable to one embodiment of the present invention;

[0032] Figures 6a to 6b This is an illustrative diagram used to explain the segmentation portion to which one embodiment of the present invention is applied;

[0033] Figure 7 This is an example of dividing an image into multiple parallel blocks;

[0034] Figures 8a to 8i Yes Figure 7 The first example diagram illustrates the setting of additional regions for each parallel block shown in the figure.

[0035] Figures 9a to 9i Yes Figure 7 The second example diagram illustrates the additional regions set for each parallel block.

[0036] Figure 10 This is an illustrative diagram illustrating the application of an additional region generated in one embodiment of the present invention during the encoding / decoding process in other regions;

[0037] Figures 11 to 12 This is a flowchart illustrating an encoding / decoding method for a segmentation unit applicable to one embodiment of the present invention;

[0038] Figures 13a to 13g It is an example diagram used to illustrate the area that can be referenced for a specific unit of division;

[0039] Figures 14a to 14e This is a flowchart illustrating the reference possibilities for adding regions in a segmentation unit to which one embodiment of the present invention is applicable;

[0040] Figure 15 This is an example diagram illustrating blocks contained in the current image segmentation unit and blocks contained in other image segmentation units;

[0041] Figure 16 This is a hardware configuration diagram illustrating an image encoding / decoding apparatus to which one embodiment of the present invention is applied;

[0042] Figure 17 This is an example diagram illustrating an in-screen prediction mode applicable to one embodiment of the present invention;

[0043] Figure 18 This is a first example illustration illustrating the configuration of reference pixels used in in-picture prediction according to one embodiment of the present invention.

[0044] Figures 19a to 19cThis is a second example illustration showing the configuration of reference pixels to which one embodiment of the present invention is applied;

[0045] Figure 20 This is a third example illustration showing the configuration of reference pixels to which one embodiment of the present invention is applied;

[0046] Figure 21 This is a fourth example illustration showing the configuration of reference pixels to which one embodiment of the present invention is applied;

[0047] Figures 22a to 22b This is an illustration of a method for filling reference pixels into pre-defined positions in unusable reference candidate blocks;

[0048] Figures 23a to 23c This is an illustrative diagram illustrating a method of interpolating on a fractional pixel basis in reference pixels constructed according to an embodiment of the present invention.

[0049] Figures 24a to 24b This is a first illustrative diagram used to explain an adaptive reference pixel filtering method applicable to one embodiment of the present invention;

[0050] Figure 25 This is a second illustration used to explain an adaptive reference pixel filtering method applicable to one embodiment of the present invention;

[0051] Figures 26a to 26b This is an illustrative diagram illustrating the use of a reference pixel level in reference pixel filtering according to one embodiment of the present invention;

[0052] Figure 27 This is an illustrative diagram illustrating a case in which multiple reference pixel levels are used in reference pixel filtering according to one embodiment of the present invention;

[0053] Figure 28 This is a block diagram used to illustrate an encoding / decoding method for an in-screen prediction mode applicable to one embodiment of the present invention;

[0054] Figure 29 This is the first example diagram used to illustrate the composition of the bitstream predicted within the frame based on the reference pixels.

[0055] Figure 30 This is the second illustration used to explain the bitstream composition of in-frame prediction based on reference pixels.

[0056] Figure 31 This is the third illustration used to explain the bitstream composition of in-frame prediction based on reference pixels.

[0057] Figure 32This is a flowchart illustrating an image decoding method supporting multiple reference pixel levels applicable to one embodiment of the present invention. Detailed Implementation

[0058] This invention is capable of various modifications and has many different embodiments. Specific embodiments will be illustrated and described in detail below. However, the following content is not intended to limit the invention to specific implementations, but should be understood to include all modifications, equivalents, and substitutions within the scope of the invention's concept and technology. Similar reference numerals are used for similar constituent elements in the description of the various figures.

[0059] In describing different constituent elements, terms such as "first," "second," "A," and "B" may be used, but these constituent elements are not limited by these terms. These terms are merely used to distinguish one constituent element from others. For example, without departing from the scope of the claims of this invention, a first constituent element can also be named a second constituent element, and similarly, a second constituent element can also be named a first constituent element. The term "and / or" includes a combination of multiple related descriptions or one of multiple related descriptions.

[0060] When a constituent element is described as being "connected" or "in contact" with other constituent elements, it should be understood that it can not only be directly connected or in contact with the aforementioned other constituent elements, but also that other constituent elements can exist between the two. Conversely, when a constituent element is described as being "directly connected" or "directly in contact" with other constituent elements, it should be understood that no other constituent elements exist between the two.

[0061] The terminology used in this application is for illustrative purposes only and is not intended to limit the invention. Singular statements also have plural meanings unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are used only to indicate the presence of features, numbers, steps, actions, constituent elements, components, or combinations thereof as described in the specification, and should not be construed as excluding the possibility of one or more other features, numbers, steps, actions, constituent elements, components, or combinations thereof being present or added.

[0062] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms commonly used, such as those defined in dictionaries, should be interpreted as having the same meaning as they have in the context of the relevant art, and should not be interpreted as having an overly idealized or exaggerated meaning in this application unless otherwise explicitly defined.

[0063] Typically, an image can be composed of a series of still images, which can be divided into groups of pictures (GOPs). Each still image can be called a picture or a frame. Higher-level concepts include GOPs and sequences, and each image can be further divided into specific regions such as strips, parallel blocks, or blocks. Furthermore, a GOP can include units such as I-images, P-images, and B-images. An I-image refers to an image that is encoded / decoded independently without using a reference image, while P-images and B-images refer to images encoded / decoded using reference images through processes such as motion estimation and motion compensation. Typically, P-images can use I-images and P-images as reference images, and B-images can use I-images and P-images as reference images; however, these definitions can change depending on the encoding / decoding settings.

[0064] In this context, the image used as a reference during the encoding / decoding process is called the Reference Picture, and the blocks or pixels used as references are called Reference Blocks and Reference Pixels. Furthermore, Reference Data, in addition to pixel values ​​in the spatial domain, can also include coefficient values ​​in the frequency domain, as well as various encoding / decoding information generated and determined during the encoding / decoding process.

[0065] The smallest unit constituting an image can be a pixel, and the number of bits used to represent a pixel is called bit depth. Typically, bit depth can be 8 bits, but other bit depths can be supported depending on the encoding settings. Regarding bit depth, at least one bit depth can be supported based on a color space. Furthermore, the image can be composed of at least one color space depending on its color format. Depending on the color format, the image can be composed of one or more images of a certain size or one or more images of different sizes. For example, in the case of YCbCr 4:2:0, the image can be composed of one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example), where the ratio of the chrominance component to the luminance component can be 1:2 (horizontal to vertical). As another example, in the case of 4:4:4, the image can have the same horizontal and vertical ratio. When the image is composed of one or more color spaces as described above, the image can be segmented in each color space.

[0066] In this invention, a portion of the color space (Y in this example) of a certain color format (YCbCr) will be used as a reference for explanation. The same or similar application is possible in other color spaces based on the color format (Cb, Cr in this example) (depending on the specific color space setting). However, some differences can also be retained in each color space (independent of the specific color space setting). That is, the setting dependent on each color space can refer to properties that are proportional to or dependent on the composition ratio of each component (e.g., determined according to 4:2:0, 4:2:2, 4:4:4, etc.), while the setting independent of each color space can refer to properties that are unrelated to the composition ratio of each component or apply independently only to the corresponding color space. In this invention, depending on the encoder / decoder, a portion of the composition can have independent or dependent properties.

[0067] The configuration information or syntax elements required during image encoding can be determined at the unit level, such as video, sequence, image, slice, parallel block, and block. These can be included in the bitstream and transmitted to the decoder in units such as Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice Header, Tile Header, and Block Header. In the decoder, these are parsed at the same unit level and used in the image decoding process after decoding the configuration information received from the encoder. Furthermore, relevant information can be transmitted to the bitstream in the form of Supplement Enhancement Information (SEI) or metadata and used after parsing. Each parameter set has its own unique numbering value, and lower-level parameter sets can include the numbering values ​​of higher-level parameter sets that need to be referenced. For example, a lower-level parameter set can reference information from one or more higher-level parameter sets that have consistent number values. In the examples of various units described above, when a unit contains one or more other units, the corresponding unit can be called the higher-level unit and the contained unit can be called the lower-level unit.

[0068] The setting information generated on the aforementioned units can include setting-related content that is independent within each unit, as well as setting-related content that depends on previous, subsequent, or superior units. Here, dependent settings refer to flag information (e.g., a 1-bit flag, where 1 indicates compliance and 0 indicates non-compliance) used to indicate whether the settings of previous, subsequent, or superior units are followed; this can be understood as representing the setting information of the corresponding unit. Although this invention will focus on examples related to independent settings, it can also include examples where setting information is added to or replaced using setting information from previous, subsequent, or superior units that depend on the current unit.

[0069] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0070] Figure 1 This is a conceptual diagram illustrating an image encoding and decoding system to which embodiments of the present invention are applied.

[0071] See Figure 1The image encoding device 105 and the decoding device 100 can be user terminals such as personal computers (PCs), laptops, personal digital assistants (PDAs), portable multimedia players (PMPs), portable game consoles (PSPs), wireless communication terminals, smartphones, and televisions, or server terminals such as application servers and business servers. They can include various devices such as communication modems for communicating with various devices or wired and wireless communication networks, memory 120 and 125 for storing various applications and data for performing inter-frame or intra-frame prediction in order to encode or decode images, and processors 110 and 115 for performing calculations and control by executing applications. Furthermore, the image encoded into a bitstream by the image encoding device 105 can be transmitted to the image decoding device 100 via wired wireless communication networks such as the Internet, short-range wireless communication networks, wireless local area networks, wireless broadband networks, and mobile communication networks, or via various communication interfaces such as cables and Universal Serial Bus (USB). The image is then decoded and reconstructed in the image decoding device 100 before being played back. Additionally, the image encoded into a bitstream by the image encoding device 105 can also be transmitted from the image encoding device 105 to the image decoding device 100 via a computer-readable storage medium.

[0072] Figure 2 This is a block diagram illustrating an image encoding apparatus to which one embodiment of the present invention is applied.

[0073] The image decoding device 20 applicable to this embodiment is as follows: Figure 2 As shown, it can include a prediction unit 200, a subtraction unit 205, a transformation unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse transformation unit 225, an addition unit 230, a filtering unit 235, an encoded image buffer 240, and an entropy coding unit 245.

[0074] The prediction unit 200 can include an intra-frame prediction unit for performing intra-frame prediction and an inter-frame prediction unit for performing inter-frame prediction. Intra-frame prediction can generate a prediction block by performing spatial prediction using pixels from blocks adjacent to the current block, while inter-frame prediction can generate a prediction block by finding the region in the reference image that best matches the current block and performing motion compensation. After determining whether to apply intra-frame prediction or inter-frame prediction to the corresponding unit (encoding unit or prediction unit), specific information related to each prediction method (e.g., intra-frame prediction mode, motion vector, reference image, etc.) can be determined. At this time, the processing unit that performs the prediction and the processing unit that determines the prediction method and specific content can differ depending on the encoding / decoding settings. For example, the prediction method and prediction mode can be determined at the prediction unit level, while the prediction can be performed at the transformation unit level.

[0075] The in-frame prediction unit can employ directional prediction modes such as horizontal and vertical modes, which are based on the prediction direction, as well as non-directional prediction modes such as mean (DC) and planar modes, which use methods such as averaging and interpolation of reference pixels. Using both directional and non-directional modes, candidate groups of in-frame prediction modes can be created. A candidate group can be selected from various options, such as 35 prediction modes (33 directional + 2 non-directional), 67 prediction modes (65 directional + 2 non-directional), or 131 prediction modes (129 directional + 2 non-directional).

[0076] The in-frame prediction unit may include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel construction unit can construct reference pixels for performing in-frame prediction, centered on the current block, using pixels contained in and adjacent to the current block. Depending on the encoding settings, the reference pixel can be constructed using the nearest adjacent row of reference pixels, or using any other adjacent row of reference pixels, or using multiple rows of reference pixels. When some of the reference pixels are unavailable, the reference pixel can be generated using the available reference pixels; when all are unavailable, the reference pixel can be generated using a pre-set value (e.g., the median value of a pixel value range represented by bit depth).

[0077] The reference pixel filtering unit of the prediction section within the image can perform filtering on reference pixels to reduce distortion remaining from the encoding process. The filter used can be a low-pass filter such as a 3-tap filter [1 / 4, 1 / 2, 1 / 4] or a 5-tap filter [2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16]. The applicability and type of filtering can be determined based on encoding information (such as block size, shape, prediction mode, etc.).

[0078] The reference pixel interpolation unit of the in-frame prediction section can generate fractional-unit pixels through a linear interpolation process of the reference pixels according to the prediction mode, and can determine the applicable interpolation filter based on the encoding information. The interpolation filters used can include 4-tap Cubic filters, 4-tap Gaussian filters, 6-tap Wiener filters, 8-tap Kalman filters, etc. Typically, the low-pass filtering process and the interpolation process are independent, but it is also possible to combine the applicable filters from both processes into one before performing the filtering process.

[0079] The prediction mode determination unit of the in-frame prediction unit can select the optimal prediction mode from the candidate prediction mode group while taking into account the encoding cost, and the prediction block generation unit can generate prediction blocks using the corresponding prediction mode. In the prediction mode encoding unit, the optimal prediction mode can be encoded based on the prediction value. At this time, the prediction information can be adaptively encoded according to whether the prediction value is suitable or unsuitable.

[0080] In the in-image prediction unit, the aforementioned predicted value is referred to as the Most Probable Mode (MPM). A subset of modes can be selected from all modes included in the prediction mode candidate group to form the Most Probable Mode (MPM) candidate group. The Most Probable Mode (MPM) candidate group can include pre-defined prediction modes (e.g., mean (DC), planar, vertical, horizontal, diagonal modes, etc.) or prediction modes of spatially adjacent blocks (e.g., left, top, upper left, upper right, lower left blocks, etc.). Furthermore, the Most Probable Mode (MPM) candidate group can be constructed using modes derived from those pre-included in the Most Probable Mode (MPM) candidate group (differences such as +1 and -1 in directional modes).

[0081] Among the prediction patterns used to construct the most probable pattern (MPM) candidate group, a priority order can exist. The order in which the most probable pattern (MPM) candidate groups are included can be determined according to this priority order, and the construction of the most probable pattern (MPM) candidate group can be completed when the number of most probable pattern (MPM) candidate groups (determined based on the number of prediction pattern candidate groups) is filled according to this priority order. At this time, the priority order can be determined according to the order of prediction patterns of spatially adjacent blocks, pre-defined prediction patterns, and patterns derived from prediction patterns earlier included in the most probable pattern (MPM) candidate group, or other variations can be made.

[0082] For example, spatially adjacent blocks can be included in the candidate group in the order of left-top-bottom-top-right-top, etc. In a pre-defined prediction pattern, they can be included in the candidate group in the order of mean (DC)-planar-vertical-horizontal pattern, etc. Patterns obtained by adding 1, subtracting 1, etc., to the pre-included patterns are then included in the candidate group, thus using a total of 6 patterns to form a candidate group. Alternatively, they can be included in the candidate group in a priority order of left-top-mean (DC)-planar-bottom-top-right-(left+1)-(left-1)-(top+1), thus using a total of 7 patterns to form a candidate group.

[0083] In the above candidate group configuration, a validity check can be performed, so that only valid blocks are included in the candidate group, and invalid blocks are skipped to the next candidate. For example, a block can be invalidated when the adjacent block is located outside the image, or is contained in a different segmentation unit than the current block, or when the encoding mode of the corresponding block is inter-picture prediction. In addition, it can also be invalidated in cases where reference is not possible as described later in this invention.

[0084] In the aforementioned subsequent selection, spatially adjacent blocks can consist of a single block or multiple blocks (sub-blocks). Therefore, in the candidate group configuration, such as the (left-top) order, the order can be either jumping to the top block after performing a validity check on a certain position in the left block (e.g., the bottommost block of the left block), or jumping to the top block after performing validity checks on multiple positions (e.g., one or more sub-blocks below the topmost block of the left block), or it can be determined according to the encoding settings.

[0085] In the inter-frame prediction unit, motion prediction methods can be categorized into mobile motion models and non-mobile motion models. In the mobile motion model, prediction is performed considering only parallel movement, while in the non-mobile motion model, prediction is performed considering rotation, distance, and zoom (zoom in / out) motion along with parallel movement. Assuming unidirectional prediction, the mobile motion model requires one motion vector, while the non-mobile motion model requires more than one. In the non-mobile motion model, each motion vector can be information applicable to a pre-defined position within the current block, such as the top-left vertex or top-right vertex. Using the corresponding motion vectors, the position of the area to be predicted within the current block can be obtained in pixel units or sub-block units. The inter-frame prediction unit can apply some of the processes described below along with other processes individually, according to the aforementioned motion models.

[0086] The inter-frame prediction unit may include a reference image composition unit, a motion prediction unit, a motion compensation unit, a motion information determination unit, and a motion information encoding unit. The reference image composition unit can include previously or subsequently encoded images in a reference image list (L0, L1) centered on the current image. From the reference images included in the reference image list, prediction blocks can be obtained, and according to the encoding settings, reference images can be composed using the current image and included in at least one position in the reference image list.

[0087] In the inter-frame prediction unit, the reference image composition unit can include a reference image interpolation unit, and can perform an interpolation process for fractional-unit pixels based on the interpolation accuracy. For example, an 8-tap interpolation filter based on Discrete Cosine Transform (DCT) can be applied to the luminance component, and a 4-tap interpolation filter based on Discrete Cosine Transform (DCT) can be applied to the chrominance component.

[0088] In the inter-frame prediction unit, the motion prediction unit is used to perform the process of exploring blocks that are highly correlated with the current block by using reference images. It can use various methods such as the Full search-based Block Matching Algorithm (FBMA) and the Three Step Search (TSS). The motion compensation unit is used to perform the process of obtaining the prediction block through the motion prediction process.

[0089] In the inter-frame prediction unit, the motion information determination unit performs a process to select the best motion information for the current block. The motion information can be encoded using motion information encoding modes such as Skip Mode, Merge Mode, and Competition Mode. These modes can be configured by combining supported modes according to the motion model. Examples include Skip Mode (moving), Skip Mode (non-moving), Merge Mode (moving), Merge Mode (non-moving), Competition Mode (moving), and Competition Mode (non-moving). Depending on the symbolization settings, a subset of these modes can be included in the candidate group.

[0090] The aforementioned motion information encoding mode can obtain the predicted value of the motion information (motion vector, reference image, prediction direction, etc.) of the current block from at least one candidate block, and can generate the best candidate selection information when supporting more than two candidate blocks. The skip mode (without residual signal) and the merge mode (with residual signal) can directly use the above predicted value as the motion information of the current block, while the competition mode can generate the difference value information between the motion information of the current block and the above predicted value.

[0091] Candidate groups for motion information prediction values ​​of the current block are adaptive based on the motion information coding mode and can adopt various configurations. Motion information of blocks spatially adjacent to the current block (e.g., left, top, top left, top right, bottom left blocks, etc.) can be included in the candidate group. Motion information of blocks temporally adjacent to the current block (e.g., left, right, top, bottom, top left, top right, bottom left, bottom right blocks, etc., including blocks in other images corresponding to or related to the current block <center>) can also be included in the candidate group. In addition, mixed motion information of spatial and temporal candidates (e.g., information obtained by averaging, median values, etc., from the motion information of spatially adjacent blocks and the motion information of temporally adjacent blocks, or motion information obtained by the current block or its sub-blocks) can also be included in the candidate group.

[0092] In the construction of candidate groups for motion information prediction values, a priority order can exist. The order in which the candidate groups are included can be determined according to this priority order, and the construction of candidate groups can be completed when the number of candidate groups is filled according to this priority order (determined based on the motion information encoding mode). At this time, the priority order can be determined according to the order of motion information of spatially adjacent blocks, motion information of temporally adjacent blocks, and mixed motion information of spatial and temporal candidates, or other variations can be made.

[0093] For example, spatially adjacent blocks can be included in the candidate group in the order of left-top-top-left-bottom-top-left blocks, while temporally adjacent blocks can be included in the candidate group in the order of bottom-right-middle-right-bottom blocks.

[0094] In the above candidate group configuration, a validity check can be performed, so that only valid blocks are included in the candidate group, and invalid blocks are skipped to the next candidate. For example, a block can be invalidated when the adjacent block is located outside the image, or is contained in a different segmentation unit than the current block, or when the encoding mode of the corresponding block is intra-image prediction. In addition, it can also be invalidated in cases where reference is not possible as described later in this invention.

[0095] In the aforementioned subsequent selection, spatially or temporally adjacent blocks can consist of a single block or multiple blocks (sub-blocks). Therefore, in the spatial candidate group formation described above, such as the (left-top) order, the order can be either jumping to the top block after performing a validity check on a certain position in the left block (e.g., the bottommost block of the left block), or jumping to the top block after performing validity checks on multiple positions (e.g., one or more sub-blocks below the topmost block of the left block). Furthermore, in the (middle-right) order of the temporal candidate group composition, the order can be to jump to the right block after performing a validity check on a certain position in the middle block (e.g., <2,2> when the middle block is divided into 4×4 regions), or the order can be to jump to the next block after performing a validity check on multiple positions (e.g., one or more sub-blocks such as <3,3>, <2,3>, etc., starting from the pre-defined position block <2,2> and located in the pre-defined order), which can also be determined according to the encoding settings.

[0096] The subtraction operation unit 205 generates a residual block by performing a subtraction operation between the current block and the predicted block. That is, the subtraction operation unit 205 generates a residual signal in the form of a block, i.e., a residual block, by calculating the difference between the pixel values ​​of each pixel in the current block to be encoded and the predicted pixel values ​​of each pixel in the predicted block generated by the prediction unit.

[0097] The transformation unit 210 transforms each pixel value of the residual block into frequency coefficients by transforming the residual block into a frequency region. Specifically, the transformation unit 210 can utilize various transformation techniques for transforming spatial axis pixel signals into frequency axes, such as the Hadamard Transform, Discrete Cosine Transform (DCT-based Transform), Discrete Sine Transform (DST-based Transform), and Caronan-Louis Transform (KLT-based Transform), to transform the residual signal into a frequency signal. The residual signal transformed into a frequency region becomes the frequency coefficients. The transformation can be performed using a 1D transformation matrix. Each transformation matrix can be adaptively used in both horizontal and vertical units. For example, when the prediction mode in in-frame prediction is horizontal, a DCT-based transformation matrix can be used vertically, and a DST-based transformation matrix can be used horizontally. When the prediction mode is vertical, a transformation matrix based on Discrete Cosine Transform (DCT) can be used in the horizontal direction and a transformation matrix based on Discrete Sine Transform (DST) can be used in the vertical direction.

[0098] The quantization unit 215 quantizes the residual block containing frequency coefficients that has been transformed into a frequency region by the transformation unit 210. The quantization unit 215 can quantize the transformed residual block using quantization techniques such as dead zone uniform threshold quantization, quantization weighted matrix, or modified quantization techniques. At this time, one or more quantization techniques can be selected as candidates, and the selection can be based on coding mode, prediction mode information, etc.

[0099] The entropy coding unit 245 generates a quantization coefficient column by scanning the generated quantization frequency coefficient column using various scanning methods, and outputs it after encoding using techniques such as entropy coding. The scanning mode can be set to one of several modes, such as zigzag, diagonal, or raster. Furthermore, it can generate encoded data containing encoding information passed from each component and output it to a bitstream.

[0100] The inverse quantization unit 220 performs inverse quantization on the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 generates a residual block containing frequency coefficients by performing inverse quantization on the quantized frequency coefficient column.

[0101] The inverse transform unit 225 performs an inverse transform on the residual block that has been inversely quantized by the inverse quantization unit 220. That is, the inverse transform unit 225 generates a residual block containing pixel values, i.e., a reconstructed residual block, by performing an inverse transform on the frequency coefficients of the inversely quantized residual block. The inverse transform unit 225 can perform the inverse transform by reversing the transform method used in the transform unit 210.

[0102] The addition unit 230 can reconstruct the current block by performing an addition operation on the prediction block predicted in the prediction unit 200 and the residual block reconstructed by the inverse transform unit 225. The reconstructed current block is stored as a reference image (or reference block) in the encoded image buffer 240, so that it can be used as a reference image when encoding the next block or subsequent blocks or other images of the current block.

[0103] The filtering unit 235 can include one or more post-processing filtering procedures, such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can eliminate block distortion appearing at the boundaries between blocks in the reconstructed image. The adaptive loop filter (ALF) can perform filtering based on the values ​​obtained by comparing the reconstructed image with the original image after filtering the blocks by the deblocking filter. The sample adaptive offset (SAO) can reconstruct the offset difference between the residual blocks that have been treated with the deblocking filter and the original image, pixel by pixel. The post-processing filters described above can be applied to the reconstructed image or blocks.

[0104] The deblocking filter in the filtering section can be applied based on the pixels contained in several columns or rows of two blocks, with the block boundary as the reference. The aforementioned blocks are preferably applicable to the boundaries of coding blocks, prediction blocks, and transform blocks, and can be limited to blocks of a preset minimum size (e.g., 8×8) or larger.

[0105] Regarding the applicability of filtering, its applicability and strength can be determined by considering the characteristics of the block boundaries, and can be selected from strong filtering, medium filtering, weak filtering, etc. Furthermore, when the aforementioned block boundary is the boundary of a segmentation unit, its applicability will be determined at the boundary of the segmentation unit based on the loop filter applicability flag, or it can be determined based on various situations described later in this invention.

[0106] The Sample Adaptive Offset (SAO) in the filtering unit can be applied based on the difference between the reconstructed image and the original image. As offset types, it supports edge offset and band offset, and can select one of these offsets to perform filtering based on the characteristics of the image. Furthermore, the offset-related information can be encoded in block units or using associated predicted values. In this case, the relevant information can be adaptively encoded according to whether the predicted values ​​are suitable or unsuitable. The predicted values ​​can be offset information of adjacent blocks (e.g., left, top, upper left, upper right blocks, etc.), and selection information related to which block's offset information to acquire can be generated.

[0107] The above-described candidate group structure allows for validity checks, ensuring that a valid block is included in the candidate group while an invalid block is moved to the next candidate. For example, adjacent blocks may be located outside the image, contained in a different segmentation unit than the current block, or may be invalid in cases where they are not referenceable, as described later in this invention.

[0108] The encoded image buffer 240 can store the blocks or images reconstructed by the filtering unit 235. The reconstructed blocks or images stored in the encoded image buffer 240 can be provided to the prediction unit 200 for performing intra-frame prediction or inter-frame prediction.

[0109] Figure 3 This is a configuration diagram illustrating an image decoding apparatus to which one embodiment of the present invention is applied.

[0110] See Figure 3 The image decoding device 30 may include an entropy decoding unit 305, a prediction unit 310, an inverse quantization unit 315, an inverse transform unit 320, an addition and subtraction unit 325, a filter 330, and a decoded image buffer 335.

[0111] Furthermore, the prediction unit 310 may further include an in-frame prediction module and an inter-frame prediction module.

[0112] First, when the image bitstream is received from the image encoding device 20, it can be transmitted to the entropy decoding unit 305.

[0113] The entropy decoding unit 305 can decode decoded data containing quantized coefficients and decoding information transmitted from each component unit by decoding the bit stream.

[0114] The prediction unit 310 can generate prediction blocks based on the data transmitted from the entropy decoding unit 305. At this time, it can also construct a list of reference images using the default composition technique based on the decoded reference images stored in the image buffer 335.

[0115] The intra-frame prediction unit may include a reference pixel composition unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction block generation unit, and a prediction mode decoding unit. The inter-frame prediction unit may include a reference image composition unit, a motion compensation unit, and a motion information decoding unit. One part of the unit may perform the same process as the encoder, while the other part may perform the reverse induction process.

[0116] The inverse quantization unit 315 is capable of inverse quantizing the quantized transform coefficients provided from the bitstream and decoded by the entropy decoding unit 305.

[0117] The inverse transform unit 320 can generate residual blocks by applying inverse discrete cosine transform (DCT), inverse integer transform, or similar inverse transform techniques to the transform coefficients.

[0118] At this time, the inverse quantization unit 315 and the inverse transformation unit 320 will reverse the processes executed in the transformation unit 210 and the quantization unit 215 of the image encoding apparatus 20 described above, and this can be achieved through various methods. For example, the same processes and inverse transformations shared with the transformation unit 210 and the quantization unit 215 can be used, or the transformation and quantization processes can be reversed using information related to the transformation and quantization processes of the image encoding apparatus 20 (such as transformation size, transformation shape, quantization type, etc.).

[0119] The residual blocks that have undergone inverse quantization and inverse transform can be added to the prediction blocks derived in the prediction unit 310 to generate reconstructed image blocks. The addition operation described above can be performed by the addition / subtraction operator 325.

[0120] For the reconstructed image blocks, filter 330 can be used as needed to apply a deblocking filter to eliminate blocking phenomena, and can also add other loop filters before and after the above decoding process to improve video quality.

[0121] The reconstructed and filtered image blocks can be stored in the decoded image buffer 335.

[0122] Although not shown in the figure, the image decoding device 30 may also include a segmentation unit, which may include an image segmentation unit and a block segmentation unit. The segmentation unit is related to... Figure 2The same or corresponding components in the image decoding device illustrated herein can be easily understood by those skilled in the art, so detailed descriptions will be omitted here.

[0123] Figures 4a to 4d This is a conceptual diagram used to illustrate a projection format applicable to one embodiment of the present invention.

[0124] Figure 4a An illustration is provided of the equi-rectangular projection (ERP) format, which projects 360-degree images onto a 2D plane. Figure 4b The cube map projection (CMP) format, which projects a 360-degree image onto a cube, is illustrated. Figure 4c An illustration is provided of the octahedral projection (OHP) format, which projects 360-degree images onto an octahedron. Figure 4d The diagram illustrates the icosahedral projection (ISP) format, which projects a 360-degree image onto a polyhedron. However, it is not limited to this; various projection formats can be used. For example, truncated square pyramid projection (TSP) and segmented sphere projection (SSP) can also be used. Figures 4a to 4d The left side shows a 3D model, and the right side shows an instance transformed into 2D space through a projection process. A 2D projected image can be composed of more than one surface, and each surface can take the shape of a circle, triangle, quadrilateral, etc.

[0125] like Figures 4a to 4d As shown, the projection format can consist of one surface (e.g., equidistant cylindrical projection (ERP)) or multiple surfaces (e.g., cubic projection (CMP), octahedral projection (OHP), icosahedral projection (ISP), etc.). Furthermore, each surface can be classified into shapes such as quadrilaterals and triangles. The above classification can be an example of the image type, characteristics, etc., applicable when different encoding / decoding settings are set according to the projection format. For example, the image type can be a 360-degree image, and the image characteristics can be one of the above classifications (e.g., various projection formats, projection formats with one or more surfaces, projection formats with quadrilateral or non-quadrilateral surfaces, etc.).

[0126] A 2D planar coordinate system {e.g., (i, j)} can be defined on each surface of a 2D projected image. The characteristics of the coordinate system can vary depending on the projection format, the position of each surface, etc. In projection formats such as equidistant cylindrical projection (ERP), a single 2D planar coordinate system can be included, while other projection formats can include multiple 2D planar coordinate systems depending on the number of surfaces. In this case, the coordinate system can be represented as (k, i, j), where k can be the index information of each surface.

[0127] For ease of explanation, this invention will focus on the case where the surface shape is quadrilateral. The number of surfaces projected onto a 2D surface can be one (e.g., equirectangular projection, where the image is equivalent to a surface) to two or more (e.g., cube projection).

[0128] Figures 5a to 5c This is a conceptual diagram illustrating a surface configuration applicable to one embodiment of the present invention.

[0129] When projecting a 3D image into a 2D projection format, the arrangement of its surfaces needs to be determined. The surface arrangement can be configured in a way that maintains the continuity of the image in 3D space, or it can be configured to maximize the fit between surfaces even if it disrupts some of the image continuity between adjacent surfaces. Furthermore, when arranging surfaces, some surfaces can be arranged after rotating them by a certain angle (0, 90, 180, 270 degrees, etc.).

[0130] See Figure 5a It can verify instances of surface layouts related to the cube projection (CMP) format. When configuring in a way that maintains image continuity in 3D space, a 4×3 layout can be used, as shown in the left image, which configures four surfaces horizontally and then one surface vertically. Furthermore, a 3×2 layout can be used, as shown in the right image, which seamlessly configures surfaces in a 2D plane even if it disrupts some of the image continuity between adjacent surfaces.

[0131] See Figure 5b It can confirm the surface configuration related to the octahedral projection (OHP) format, as shown in the upper image when configured in a way that maintains the continuity of the image in 3D space. At this time, even if some continuity is disrupted, the surface can be seamlessly configured on the projected 2D plane, as shown in the lower image.

[0132] See Figure 5cIt can confirm the surface configuration related to the icosahedral projection (ISP) format. It can be configured in a way that maintains the continuity of the image in 3D space as shown in the image above, or it can be configured in a way that fits the surfaces seamlessly as shown in the image below.

[0133] At this point, the process of seamlessly fitting the surface to the configuration can be called frame packing, and the disruption of image continuity can be minimized by rotating the surface before configuration. Next, the process of changing the surface configuration to another surface configuration as described above will be called surface reconfiguration.

[0134] Next, the term continuity can be interpreted as the continuity of a scene visible to the naked eye in 3D space, or the continuity of an actual image or scene in 2D projected space. Continuity can also be expressed as a high degree of correlation between regions. Typically, the correlation between regions in a 2D image may be high or low, but in a 360-degree image, there may be regions that are spatially adjacent but have no continuity. Furthermore, according to the surface configurations or reconfigurations described above, there may be regions that are not spatially adjacent but have continuity.

[0135] It is possible to perform surface reconfiguration for the purpose of improving coding performance. For example, it is possible to arrange surfaces with image continuity adjacent to each other by performing surface reconfiguration.

[0136] At this point, surface reconfiguration does not necessarily mean rebuilding the surface after it has been configured; it can also be understood as a process of setting a specific surface configuration from the outset. (This can be performed within region-wise packing during the 360-degree image encoding / decoding process.)

[0137] Furthermore, surface configuration or reconfiguration can include surface rotation in addition to changes in the position of each surface (in this example, simple movement of the surface, such as movement from the upper left to the lower left or lower right of the image). Surface rotation can include 0 degrees (no surface rotation), 45 degrees to the right, 90 degrees to the left, etc., and can be represented by selecting the selected intervals after dividing the 360 ​​degrees (equal or unequal) into k (or 2k) intervals.

[0138] The encoder / decoder can perform surface configuration (or reconfiguration) according to pre-defined surface configuration information (surface shape, number of surfaces, surface position, surface rotation angle, etc.) and / or surface reconfiguration information (information indicating the position or movement angle, movement direction, etc. of each surface). Furthermore, the encoder can generate surface configuration information and / or surface reconfiguration information based on the input image, and the decoder can receive and decode this information from the encoder to perform surface configuration (or reconfiguration).

[0139] Unless otherwise specified, the surfaces referred to below are... Figure 5a The 3×2 layout is a prerequisite, and at this time, it means that the numbering of each surface can be from 0 to 5, starting from the top left and following the grid scanning order.

[0140] Next, in Figure 5a When describing the continuity between surfaces, unless otherwise stated, it can be assumed that surfaces 0 to 2 are continuous with each other, surfaces 3 to 5 are continuous with each other, while surfaces 0 and 3, 1 and 4, and 2 and 5 are not continuous. Whether the aforementioned surfaces are continuous can be determined by the characteristics, type, format, etc. of the image.

[0141] In the encoding / decoding process of 360-degree images, the encoding device can acquire the input image, preprocess the acquired image, encode the preprocessed image, and send the encoded bitstream to the decoding device. Preprocessing can include image stitching, projecting the 3D image onto a 2D plane, surface configuration, and reconfiguration (or region-wise packing). Furthermore, the decoding device can receive the bitstream, decode the received bitstream, perform post-processing (image rendering, etc.) on the decoded image, and generate the output image.

[0142] At this point, the bitstream can include and transmit information generated during preprocessing (such as supplementary enhancement information (SEI) messages or metadata) and information generated during encoding (image encoding data).

[0143] Figures 6a to 6b This is an illustrative diagram used to explain the segmentation portion to which one embodiment of the present invention is applied.

[0144] Figure 2 or Figure 3The image encoding / decoding apparatus may further include a segmentation unit, which may include an image segmentation unit and a block segmentation unit. The image segmentation unit may segment the image into at least one processing unit {e.g., color space (YCbCr, RGB, XYZ, etc.), sub-image, stripe, parallel block, basic coding unit (or maximum coding unit, etc.)}, while the block segmentation unit may segment the basic coding unit into at least one processing unit (e.g., encoding, prediction, transform, quantization, entropy, loop filtering unit, etc.).

[0145] The basic coding unit can be obtained by dividing the image along the horizontal or vertical direction at certain length intervals. This can also be applied to units such as sub-images, parallel blocks, stripes, surfaces, etc. That is, the above units can be constructed in integer multiples of the basic coding unit, but are not limited to this.

[0146] For example, different basic coding units can be applied to some segmentation units (parallel blocks, sub-images, etc. in this example), and the corresponding segmentation units can adopt independent basic coding unit sizes. That is, the basic coding unit of the above-mentioned segmentation units can be set to be the same as or different from the basic coding unit of the image unit and the basic coding unit of other segmentation units.

[0147] For ease of explanation in this invention, the basic coding unit and other processing units (coding, prediction, transformation, etc.) are referred to as blocks.

[0148] The size or shape of the aforementioned block can be a power of 2 in terms of its horizontal or vertical length (2^2). n The format can be represented as an N×N square (2n×2n, 256×256, 128×128, 64×64, 32×32, 16×16, 8×8, 4×4, etc., where n is an integer between 2 and 8) or an M×N rectangle (2m×2n). For example, it can divide a high-resolution 8k Ultra High Definition (UHD) input image into a 256×256 size, a 1080p High Definition (HD) input image into a 128×128 size, and a Wide Image Graphics Array (WVGA) input image into a 16×16 size.

[0149] An image can be segmented into at least one strip. A strip can be composed of a combination of at least one block that is consecutive in the scanning order. Each strip can be segmented into at least one strip segment, and each strip segment can be segmented into basic coding units.

[0150] An image can be segmented into at least one sub-image or parallel block. Sub-images or parallel blocks can adopt a quadrilateral (rectangular or square) segmentation shape and can be segmented into basic coding units. Sub-images are similar to parallel blocks in adopting the same segmentation shape (quadrilateral). However, unlike parallel blocks, sub-images can be distinguished from parallel blocks in that they employ independent encoding / decoding settings. That is, parallel blocks receive setting information for performing encoding / decoding from a higher-level unit (e.g., an image), while sub-images can directly obtain at least one setting information for performing encoding / decoding from the header information of each sub-image. In other words, unlike sub-images, parallel blocks are units obtained simply by segmenting the image, not units for data transmission {e.g., the basic unit of a video coding layer (VCL)}.

[0151] Furthermore, parallel blocks can be segmentation units supported from a parallel processing perspective, while sub-images can be segmentation units supported from an independent encoding / decoding perspective. Specifically, sub-images can not only be encoded / decoded on a sub-image basis, but also their encoding / decoding can be determined. Moreover, corresponding sub-images can be constructed and displayed centered on the region of interest, and related settings can be determined at the sequence, image, or other unit level.

[0152] In the above examples, it is also possible to change the encoding / decoding settings to be performed at the higher level of the sub-image and to perform independent encoding / decoding settings at the parallel block level. In this invention, for the sake of explanation, we will take the case where the parallel blocks can be set independently or dependent on the higher level as an example.

[0153] The segmentation information generated when segmenting an image into quadrilateral shapes can take on multiple forms.

[0154] See Figure 6aThis method can obtain quadrilateral-shaped segmentation units by dividing an image at once along the horizontal line b7 and the vertical line (where b1 and b3, b2 and b4 form a single dividing line). For example, it can generate the number of quadrilaterals based on the horizontal and vertical directions respectively. When the quadrilaterals are equally divided, the horizontal and vertical lengths of the image can be divided by the number of horizontal and vertical lines to confirm the horizontal and vertical lengths of the segmented quadrilaterals. When the quadrilaterals are not equally divided, additional information indicating the horizontal and vertical lengths of the quadrilaterals can be generated. The horizontal and vertical lengths can be represented in one pixel unit or in pixels. When represented in multiple pixel units, for example, when the size of the basic encoding unit is M×N and the size of the quadrilateral is 8M×4N, the horizontal and vertical lengths can be represented as 8 and 4 respectively (in this example, when the basic encoding unit of the corresponding segmentation unit is M×N), or as 16 and 8 (in this example, when the basic encoding unit of the corresponding segmentation unit is M / 2×N / 2).

[0155] In addition, see Figure 6b able to Figure 6a Different methods were used to obtain quadrilateral-shaped segmentation units by independently segmenting the image. For example, it was possible to generate information on the number of quadrilaterals within the image, the starting position information of each quadrilateral in the horizontal and vertical directions (indicated by the positions shown in the attached figures z0 to z5, which can be represented by x, y coordinates within the image), and the horizontal and vertical length information of each quadrilateral. In this case, the starting position can be represented in one or more pixel units, and the horizontal and vertical lengths can also be represented in one or more pixel units.

[0156] Figure 6a It can be an instance of segmentation information from parallel blocks or sub-images, and Figure 6b It can be an instance of segmentation information of a sub-image, but is not limited to this. For ease of explanation, the parallel blocks will be described using quadrilateral-shaped segmentation units, but the descriptions related to parallel blocks can be applied in the same or similar way to sub-images (and also to faces). That is, the distinction in this invention is merely a difference in terminology; in reality, the descriptions related to parallel blocks can be used as definitions of sub-images, and the descriptions related to sub-images can also be used as definitions of parallel blocks.

[0157] The segmentation units mentioned above are not required to be included. It is possible to selectively include all or a portion of them based on the encoding / decoding settings. It is also possible to support other additional units (such as surfaces).

[0158] Furthermore, the encoding unit (or block) can be divided into various sizes through block segmentation. In this case, the encoding unit can consist of multiple encoding blocks depending on the color format (e.g., one luminance encoding block and two chrominance encoding blocks). For ease of explanation, a single color component unit will be assumed. The encoding block can be of variable size, such as M×M (e.g., M is 4, 8, 16, 32, 64, 128, etc.). Furthermore, the size can be determined based on the segmentation method (e.g., tree structure segmentation, i.e., quadtree).<QuadTree,QT> partitioning, binary tree<Binary Tree,BT> partitioning, ternary tree<Ternary Tree,TT> (Segmentation, etc.) can divide the coded block into M×N (e.g., M and N are 4, 8, 16, 32, 64, 128, etc.) variable sizes. In this case, the coded block can serve as the basic unit for intra-frame prediction, inter-frame prediction, transformation, quantization, entropy coding, etc.

[0159] Although this invention is illustrated using the example of obtaining multiple sub-blocks of the same size and shape (symmetrical) through segmentation, it is equally applicable to cases containing asymmetrical sub-blocks (e.g., in binary tree segmentation, the horizontal ratio <vertical similarity> between the segmented blocks is 1:3 or 3:1, or the vertical ratio <horizontal similarity> is 1:3 or 3:1, etc.; in ternary tree segmentation, the horizontal ratio <vertical similarity> between the segmented blocks is 1:2:1, or the vertical ratio <horizontal similarity> is 1:2:1, etc.).

[0160] The partitioning of an M×N coded block can employ a recursive tree structure. Whether or not a partition is performed can be indicated by a partition flag. For example, when the partition flag for a coded block with a partition depth of k is 0, the encoding of the coded block is performed on the coded block with partition depth k. When the partition flag for a coded block with partition depth k is 1, the encoding of the coded block will be performed on four sub-coded blocks (quadtree partitioning), two sub-coded blocks (binary tree partitioning), or three sub-coded blocks (ternary tree partitioning) with partition depth k+1, depending on the partitioning method.

[0161] The aforementioned sub-encoding block will be reset to encoding block k+1 and can be further divided into sub-encoding block k+2 through the above process. In quadtree partitioning, a partitioning flag (e.g., used to indicate whether or not to partition) can be supported.

[0162] In binary tree partitioning, it supports partition flags and partition direction flags (horizontal or vertical). When more than one partition ratio is supported in binary tree partitioning (e.g., supporting additional partition ratios other than 1:1 horizontal or vertical, i.e., asymmetric partitioning), it can also support partition ratio flags (e.g., selecting one of the horizontal or vertical ratio candidate groups <1:1, 1:2, 2:1, 1:3, 3:1>), or support other types of flags (e.g., whether it is a symmetric partition or not; when it is 1, it is a symmetric partition without additional information; when it is 0, it is an asymmetric partition that requires additional information related to the ratio).

[0163] The ternary tree partitioning supports partition flags and partition direction flags. When more than one partition ratio is supported in the ternary tree partitioning, the same additional partitioning information as described above for binary tree partitioning will be required.

[0164] The above example shows the segmentation information generated when only one tree segmentation method is effective. When multiple tree segmentation methods are effective, the segmentation information described below can be formed.

[0165] For example, when multiple tree-like partitions are supported, if a pre-defined partition priority order exists, partition information corresponding to the priority order can be constructed first. At this time, when the partition flag corresponding to the priority order is true, additional partition information related to the corresponding partition method can be further included, while when it is false (partitioning is not performed), the partition information (partition flag, partition direction flag, etc.) corresponding to the partition method of the next order can be constructed.

[0166] Alternatively, when multiple tree-like segmentations are supported, additional selection information related to the segmentation method can be generated, and the information can be composed of segmentation information related to the selected segmentation method.

[0167] Some of the above-mentioned segmentation markers can be omitted based on the earlier segmentation results from the superior or previous levels.

[0168] Block partitioning can proceed from the largest coded block to the smallest coded block. Alternatively, it can proceed from the minimum partition depth of 0 to the maximum partition depth. That is, partitioning can be performed recursively until the block size reaches the minimum coded block size or the partition depth reaches the maximum partition depth. In this case, it can be performed according to the encoding / decoding settings (e.g., image <strip, parallel block> type). Encoding mode <Intra / Inter>, color difference components <y cb cr>(etc.), the size of the maximum coding block, the size of the minimum coding block, and the maximum segmentation depth are adaptively set.

[0169] For example, when the maximum coding block size is 128×128, quadtree segmentation can be performed within the range of 32×32 to 128×128, while binary tree segmentation can be performed within the range of 16×16 to 64×64 with a maximum segmentation depth of 3, and ternary tree segmentation can be performed within the range of 8×8 to 32×32 with a maximum segmentation depth of 3. Alternatively, quadtree segmentation can be performed within the range of 8×8 to 128×128, while binary and ternary tree segmentation can be performed within the range of 4×4 to 128×128 with a maximum segmentation depth of 3. The former can be a setting on I-type imagery (e.g., stripe), while the latter can be a setting on P-type or B-type imagery.

[0170] As illustrated in the examples above, the segmentation settings such as maximum code block size, minimum code block size, and maximum segmentation depth can be set generally or independently depending on the segmentation method and the encoding / decoding settings as described above.

[0171] When multiple partitioning methods are supported, the partitioning will be performed within the block support range of each partitioning method. When the block support ranges of different partitioning methods overlap, priority information of the partitioning methods can be included. For example, quadtree partitioning can be performed before binary tree partitioning.

[0172] Alternatively, when the segmentation support ranges overlap, segmentation selection information can be generated. For example, selection information related to the segmentation methods to be performed in binary tree segmentation and ternary tree segmentation can be generated.

[0173] Furthermore, when multiple segmentation methods are supported, the execution of later segments can be determined based on the results of earlier segments. For example, if the result of an earlier segment (quadtree segmentation in this example) indicates that a segmentation should be performed, the later segmentation (binary tree segmentation or ternary tree segmentation in this example) can be skipped. Instead, the segmentation can continue after the sub-coded blocks segmented by the earlier segmentation are re-defined as coded blocks.

[0174] Alternatively, if the result of an earlier segmentation indicates that no segmentation should be performed, segmentation can be performed based on the result of a later segmentation. In this case, if the result of a later segmentation (in this example, a binary tree segmentation or a ternary tree segmentation) indicates that segmentation should be performed, segmentation can continue after the segmented sub-coded blocks are re-set as coded blocks. Conversely, if the result of a later segmentation indicates that no segmentation should be performed, segmentation can be stopped. In this case, if the result of a later segmentation indicates that segmentation should be performed and multiple segmentation methods are still supported when re-setting the segmented sub-coded blocks as coded blocks (e.g., when the block support ranges of different segmentation methods overlap), only the later segmentation can be performed without performing the earlier segmentation. That is, when multiple segmentation methods are supported, if the result of an earlier segmentation indicates that no segmentation should be performed, the earlier segmentation can be stopped.

[0175] For example, when an M×N coded block can undergo quadtree partitioning and binary tree partitioning, the quadtree partitioning flag is first confirmed. When the flag is 1, it is partitioned into four sub-coded blocks of size (M>>1)×(N>>1). These sub-coded blocks are then re-established as coded blocks before further partitioning (quadtree or binary tree partitioning) is performed. When the flag is 0, the binary tree partitioning flag is confirmed. When the flag is 1, it is partitioned into two sub-coded blocks of size (M>>1)×N or M×(N>>1). These sub-coded blocks are then re-established as coded blocks before further partitioning (binary tree partitioning) is performed. When the flag is 0, the partitioning process ends and encoding begins.

[0176] The above examples illustrate the use of multiple segmentation methods, but are not limited to this; combinations of multiple segmentation methods can also be supported. For example, segmentation methods such as quadtree / binary tree / ternary tree / quadtree+binary tree / quadtree+binary tree+ternary tree can be used. In this case, information related to whether additional segmentation methods are supported can be implicitly determined or explicitly included in units such as sequences, images, sub-images, stripes, and parallel blocks.

[0177] In the above examples, segmentation-related information such as the size of the coded block, the supported range of the coded block, and the maximum segmentation depth can be included in units such as sequences, images, sub-images, stripes, and parallel blocks, or implicitly determined. In other words, the range of permissible blocks can be determined based on the size of the maximum coded block, the range of supported blocks, and the maximum segmentation depth.

[0178] The coded blocks obtained through the segmentation process described above can be set to the maximum size for intra-frame or inter-frame prediction. That is, to perform intra-frame or inter-frame prediction, the coded blocks after segmentation can be the starting size for the prediction blocks. For example, when the coded block is 2M×2N, the prediction block size can be the same as or a smaller 2M×2N or M×N. Alternatively, it can be 2M×2N, 2M×N, M×2N, or M×N. Or, it can be the same size as the coded block, 2M×2N. In this case, the same size for the coded block and the prediction block means that no segmentation of the prediction block is performed, but prediction is performed using the size obtained from the segmentation of the coded block. That is, no segmentation information for the prediction block is generated. The settings described above can also be applied to transform blocks, allowing transformation to be performed in units of segmented coded blocks.

[0179] Various configurations can be achieved through the encoding / decoding settings described above. For example, (after determining the encoded block) at least one prediction block and at least one transform block can be obtained based on the encoded block. Alternatively, a prediction block of the same size as the encoded block can be obtained, and at least one transform block can be obtained based on the encoded block. Or, a prediction block of the same size as the encoded block and a transform block can be obtained. In the above examples, when at least one block is obtained, segmentation information for each block can be generated, while when only one block is obtained, segmentation information for each block will not be generated.

[0180] The square or rectangular blocks of various sizes obtained from the above results can be used for intra-frame prediction, inter-frame prediction, transformation and quantization of residual components, and filtering.

[0181] The segmentation units obtained by segmenting an image using an image segmentation unit can perform independent or dependent encoding / decoding based on the encoding / decoding settings.

[0182] Independent encoding / decoding refers to encoding / decoding a subset of units (or regions) without referencing data from other units. Specifically, the information used or generated during texture and entropy encoding of a subset of units {e.g., pixel values ​​or encoding / decoding information (intra-frame prediction information, inter-frame prediction information, and entropy encoding / decoding information, etc.)} will be encoded independently without cross-referencing. Similarly, during texture and entropy decoding of a subset of units in the decoder, the parsed and reconstructed information from other units will not be cross-referenced.

[0183] Furthermore, dependent encoding / decoding refers to the ability to use data from other units as a reference when encoding / decoding a subset of segments. Specifically, the information used or generated during texture encoding and entropy encoding of a subset of segments can be encoded dependently through mutual reference. Similarly, during texture decoding and entropy decoding of a subset of segments in the decoder, the parsed and reconstructed information from other units can be cross-referenced.

[0184] Typically, the segmentation units mentioned above (e.g., sub-images, parallel blocks, stripes, etc.) can employ independent encoding / decoding settings. That is, a non-referenced setting can be used for parallelization purposes. Furthermore, a non-referenced setting can be used to improve encoding / decoding performance. For example, when a 360-degree image is segmented into multiple surfaces in 3D space and configured in 2D space, the correlation (e.g., image continuity) with adjacent surfaces may decrease depending on the surface configuration. In other words, because the need for mutual reference is low when there is no correlation between surfaces, an independent encoding / decoding setting can be used.

[0185] Furthermore, it is possible to employ referenceable settings between segmentation units for the purpose of improving encoding / decoding performance. For example, even when segmenting a 360-degree image into surface units, there may be cases where the correlation with adjacent surfaces is high depending on the surface configuration settings. In such cases, a dependent encoding / decoding setting can be used.

[0186] Furthermore, in this invention, the independent or dependent encoding / decoding is not only applicable to the spatial region but also extended to the temporal region. That is, it can perform independent or dependent encoding / decoding not only on other segmentation units existing in the same time period as the current segmentation unit, but also on segmentation units existing in different time periods (in this example, even if there are segmentation units at the same location in images corresponding to different time periods as the current segmentation unit, they are assumed to be other segmentation units).

[0187] For example, when simultaneously transmitting bitstream A, which contains data encoded in high quality for 360-degree images, and bitstream B, which contains data encoded in normal quality, the decoder can parse and decode bitstream A, which is transmitted in high quality, in the region corresponding to the area of ​​interest (such as the area where the user's gaze is focused <viewport> or the area that is desired to be displayed), while parsing and decoding bitstream B, which is transmitted in normal quality, outside the area of ​​interest.

[0188] Specifically, when an image is segmented into multiple units (e.g., sub-images, parallel blocks, strips, surfaces, etc., in this example it is assumed that the surface is processed in the same way as the parallel blocks or sub-images), it is possible to decode the data (bit stream A) of the segmented units contained in the region of interest (or segmented units that overlap with the viewport by as many as one pixel) and the data (bit stream B) of the segmented units contained outside the region of interest.

[0189] Alternatively, it can transmit a bitstream containing data encoding the entire image, and the decoder can parse and decode the region of interest from the bitstream. Specifically, it can decode only the data of the segmented units contained within the region of interest.

[0190] In other words, it is possible to obtain a whole or a part of an image by generating bitstreams of more than one quality in the encoder and decoding only specific bitstreams in the decoder, or it is possible to obtain a whole or a part of an image by selectively decoding individual bitstreams in each image segment. The example above uses a 360-degree image, but this explanation is applicable to general images.

[0191] When performing encoding / decoding as described above, since it is impossible to know which data will be reconstructed in the decoder (in this example, the decoder does not know the location of the region of interest and it is based on random access of the region of interest), in addition to the spatial region, it is also necessary to confirm the reference settings in the temporal region and perform encoding / decoding.

[0192] For example, when the decoder determines which type of decoding to perform based on a single segmentation unit, the current segmentation unit can perform independent coding in the spatial region and limited dependency coding in the temporal region (e.g., only allowing reference to segmentation units at the same location at other times corresponding to the current segmentation unit and prohibiting reference to other segmentation units, since there is generally no restriction in the temporal region, so this is in contrast to unrestricted dependency coding).

[0193] Alternatively, when the decoder determines which type of decoding to perform (e.g., in this case, multiple units are decoded as long as any one of the units is contained within the region of interest) is determined by multiple segmentation units (which can be obtained by bundling horizontally adjacent segmentation units or vertically adjacent segmentation units, or by bundling both horizontally and vertically adjacent segmentation units), the current segmentation unit can perform independent or dependent decoding in the spatial region and limited dependent encoding in the temporal region (e.g., it can refer to other segmentation units besides allowing reference to segmentation units at the same location at other times corresponding to the current segmentation unit).

[0194] In this invention, a surface is a segmentation unit whose configuration and shape usually change according to the projection format and which does not have independent encoding / decoding settings. Although it has different characteristics from other segmentation units as described above, it can also be regarded as a unit obtained in the image segmentation unit in terms of being able to divide an image into multiple regions (and adopting a quadrilateral shape, etc.).

[0195] As described above, independent encoding / decoding can be performed on individual segmented units in a spatial region for purposes such as parallelization. However, because independent encoding / decoding cannot reference other segmented units, it leads to a decrease in encoding / decoding efficiency. Therefore, as a step before performing encoding / decoding, the segmented units performing independent encoding / decoding can be expanded by utilizing (or appending) data from adjacent segmented units. Since the segmented unit with appended data from adjacent segmented units has more data to reference, its encoding / decoding efficiency is improved. At this point, because the expanded segmented unit can reference data from adjacent segmented units during encoding / decoding, it can be considered as dependent encoding / decoding.

[0196] The aforementioned information related to the reference settings between segmentation units can be recorded in the bitstream and transmitted to the decoder in units such as video, sequence, image, sub-image, stripe, and parallel block. The decoder can then reconstruct the settings information transmitted from the encoder by parsing at the same level of units. Furthermore, relevant information can be transmitted to the bitstream in the form of Supplement Enhancement Information (SEI) or metadata and used after parsing. Additionally, encoding / decoding can be performed based on the reference settings without transmitting the aforementioned information, utilizing pre-agreed definitions in the encoder / decoder.

[0197] Figure 7 This is an example of dividing an image into multiple parallel blocks. Figures 8a to 8i Yes Figure 7 The first example diagram illustrates the additional regions set for each parallel block. Figures 9a to 9i Yes Figure 7 The second example diagram illustrates the additional regions set for each parallel block.

[0198] When an image is segmented into two or more segments (or regions) by an image segmentation unit and each segment is encoded / decoded independently, while there are advantages such as the ability to perform parallel processing, the encoding performance may degrade due to the reduced amount of data available for reference by each segment. To solve the problem described above, processing can be performed using encoding / decoding settings that are dependent on the segmentation units (in this example, parallel block particles will be used as an example; the same or similar settings can be applied to other units).

[0199] Encoding / decoding is typically performed independently between segmented units in a non-referenceable manner. Therefore, pre-processing or post-processing procedures can be performed to implement dependent encoding / decoding. For example, it is possible to form extended regions on the outline of each segmented unit before performing encoding / decoding and fill these extended regions with data from other segmented units that require reference.

[0200] Although the method described above is no different from performing independent encoding / decoding except that the encoding / decoding is performed after each segmentation unit is expanded, it can be understood as an instance of dependent encoding / decoding because the existing segmentation units obtain the necessary reference data from other segmentation units in advance and refer to it.

[0201] Furthermore, after encoding / decoding, filtering can be applied using multiple segmented unit data based on the boundaries between segmented units. That is, when applying filtering, the data becomes dependent due to the use of other segmented unit data, while when not applying filtering, it becomes independent.

[0202] In the examples described below, the focus will be on cases where dependent encoding / decoding is performed by executing a pre-processing step (an extension in this example). Furthermore, in this invention, the boundary between identical segmentation units can be referred to as the inner boundary, and the outline of the image can be referred to as the outer boundary.

[0203] In one embodiment of the present invention, an additional region associated with the current parallel block can be set. Specifically, the additional region can be set based on at least one parallel block (in this example, this includes cases where an image is composed of one parallel block, i.e., cases where it is not divided into more than two segmentation units; more precisely, although a segmentation unit means being divided into more than two units, it is assumed that it is also identified as a segmentation unit even if it is not segmented).

[0204] For example, an append region can be set in at least one of the directions such as up / down / left / right of the current parallel block. The append region can be filled with arbitrary values. Furthermore, the append region can be filled with a portion of the data in the current parallel block; that is, it can be filled using the outline pixels of the current parallel block or by copying pixels within the current parallel block.

[0205] Furthermore, the additional region can be filled using image data from other parallel blocks besides the current parallel block. Specifically, it can be filled using image data from parallel blocks adjacent to the current parallel block; that is, it can be filled by copying image data from parallel blocks adjacent to the current parallel block in a specific direction (up / down / left / right).

[0206] At this point, the size (length) of the acquired image data can use the same value in all directions or use independent values, which can be determined according to the encoding / decoding settings.

[0207] For example, in Figure 6a It can extend in all or part of the boundary directions of the b0 to b8 boundary. Furthermore, it can extend m in all boundary directions of the segmented unit or extend m according to the boundary direction. i (i is the index for each direction). m or mi can be applied to all segmentation units in the image, or the independence of each segmentation unit can be set.

[0208] At this point, setting information related to the additional region can be generated. This setting information can include whether the additional region is supported, whether the additional region is supported for each segmentation unit, the shape of the additional region on the overall image (e.g., determined based on which direction it expands into (up / down / left / right) of the segmentation unit; in this example, it is setting information that applies to all segmentation units within the image), the shape of the additional region on each segmentation unit (in this example, it is setting information that applies to individual segmentation units within the image), the size of the additional region on the overall image (e.g., indicating the degree of expansion in the direction of expansion after the shape of the additional region is determined; in this example, it is setting information that applies to all segmentation units within the image), the size of the additional region on each segmentation unit (in this example, it is setting information that applies independently to individual segmentation units within the image), the method of filling the additional region on the overall image, and the method of filling the additional region on each segmentation unit, etc.

[0209] The aforementioned settings related to the additional area can be determined proportionally according to the color space, or they can be set independently. Additional area-related setting information can be generated for the luminance component, while the additional area setting for the chromatic difference component can be implicitly determined according to the color space. Alternatively, additional area-related setting information can also be generated for the chromatic difference component.

[0210] For example, when the size of the additional region for the luminance component is m, the size of the additional region for the chrominance component can be determined as m / 2 based on the color format (4:2:0 in this example). As another example, when the size of the additional region for the luminance component is m and the chrominance component is set independently, the size information of the additional region for the chrominance component can be generated (n in this example; n can be used collectively or n1, n2, n3, etc., depending on the direction or extended region). As yet another example, a method for filling the additional region for the luminance component can be generated, while the method for filling the additional region for the chrominance component can use the method from the luminance component or generate related information.

[0211] The information related to the aforementioned append region settings can be recorded in the bitstream and transmitted in units such as video, sequence, image, sub-image, and stripe. During decoding, the relevant information can be parsed and reconstructed from these units. In the embodiments described below, the case of supporting append regions will be used as an example.

[0212] See Figure 7 This confirms that an image has been segmented into parallel blocks labeled 0 to 8. At this point, for... Figure 7 The results of setting additional regions for each parallel block as illustrated in the figure, according to one embodiment of the present invention, are as follows: Figures 8a to 8i As shown.

[0213] exist Figure 7 as well as Figure 8a In this context, parallel block 0 (of size T0_W × TO_H) can be expanded by appending an area of ​​E0_R to the right and an area of ​​E0_D to the bottom. The appended areas can be obtained from adjacent parallel blocks. Specifically, the right-side expansion area can be obtained from parallel block 1, and the bottom expansion area can be obtained from parallel block 3. Furthermore, parallel block 0 can utilize the adjacent parallel block to its lower right (parallel block 4) to define the appended area. That is, the appended area can be defined in the direction of the remaining internal boundaries (or boundaries between the same segmentation units) excluding the outer boundary (or image boundary) of the parallel block.

[0214] exist Figure 7 as well as Figure 8e In this context, because parallel block number 4 (size T4_W×T4_H) has no external boundary, it can be expanded by adding regions to the left, right, top, and bottom. In this case, the left expansion region can be obtained from parallel block number 3, the right expansion region from parallel block number 5, the top expansion region from parallel block number 1, and the bottom expansion region from parallel block number 7. Furthermore, parallel block number 4 can also add regions to the upper left, lower left, upper right, and lower right. In this case, the upper left expansion region can be obtained from parallel block number 0, the lower left expansion region from parallel block number 6, the upper right expansion region from parallel block number 2, and the lower right expansion region from parallel block number 8.

[0215] In Figure 8, since the L2 block is adjacent to the boundary of the parallel block, it theoretically has no data that can be referenced from the left, upper left, and lower left blocks. However, when an additional region is set for the second parallel block using an embodiment of the present invention, the L2 block can be encoded / decoded by referencing the additional region. That is, the L2 block can reference the data of the left and upper left blocks as the additional region (which can be the region obtained from the first parallel block), and can reference the data of the lower left block as the additional region (which can be the region obtained from the fourth parallel block).

[0216] The data included in the appended region in the above embodiments can be included in the current parallel block for encoding / decoding. In this case, because the data in the appended region is located on the boundary of the parallel block (in this example, the parallel block that is updated or expanded due to the appended region), the encoding performance may also degrade due to the lack of reference data during the encoding process. However, since this is only an appended portion to provide a reference to the boundary region of the original parallel block, it can be understood as a form of temporary memory used to improve encoding performance. That is, since it can help improve the image quality performance of the final output image and is the region that is eventually removed, the degradation of the encoding performance of the corresponding region does not cause any problems. This can be applied to the embodiments described below for similar or the same purpose.

[0217] In addition, see Figures 9a to 9i It can be confirmed that, according to the projection format, a 360-degree image is transformed into a 2D image through a surface configuration (or reconfiguration) process, and the 2D image is divided into parallel blocks (or surfaces). In this case, since the 2D image is composed of a single surface when the 360-degree image is projected using an equirectangular projection, it can be an example of dividing a surface into parallel blocks. Furthermore, for ease of explanation, the parallel block division of the 2D image will be compared with... Figure 7 The parallel block division shown in the diagram is the same as the premise.

[0218] The resulting parallel blocks can be divided into parallel blocks consisting only of internal boundaries and parallel blocks containing at least one external boundary, and can be configured according to... Figures 8a to 8i The method shown sets additional regions for each parallel block. However, even if the 360-degree images transformed into 2D images are adjacent in the 2D image, they may not have continuity in the actual image, while even if they are not adjacent, they may have continuity in the actual image (see [reference to...]). Figures 5a to 5c (See the explanation). Therefore, even if part of the boundary of a parallel block is an outer boundary, there can still be regions within the image that are continuous with the outer boundary region of the parallel block. For details, see [link to relevant documentation]. Figure 9b Although the upper end of the first parallel block is the outer boundary of the image, because there can be regions with continuity in the actual image within the same image, an additional region can be set at the upper end of the first parallel block. That is, with Figures 8a to 8i The difference lies in Figures 9a to 9i It also allows setting additional regions in all or part of the outer boundary direction of parallel blocks.

[0219] See Figure 9e Parallel Block No. 4 is a parallel block whose boundary only includes its internal boundary (parallel block No. 4 in this example). Therefore, the additional area of ​​Parallel Block No. 4 can be set in the top, bottom, left, right, and top-left, bottom-left, top-right, and bottom-right directions. Specifically, the left extension area can be image data obtained from Parallel Block No. 3, the right extension area can be image data obtained from Parallel Block No. 5, the top extension area can be image data obtained from Parallel Block No. 1, the bottom extension area can be image data obtained from Parallel Block No. 7, the top-left extension area can be image data obtained from Parallel Block No. 0, the bottom-left extension area can be image data obtained from Parallel Block No. 6, the top-right extension area can be image data obtained from Parallel Block No. 2, and the bottom-right extension area can be image data obtained from Parallel Block No. 9.

[0220] See Figure 9a Parallel block 0 is a parallel block that contains at least one outer boundary (left or top direction). Therefore, in addition to spatially adjacent regions in the right, bottom, and bottom-right directions, parallel block 0 can also contain additional regions extending towards the outer boundary (left, top, and top-left directions). While the spatially adjacent regions in the right, bottom, and bottom-right directions can utilize data from adjacent parallel blocks to define the additional regions, this is not possible for additional regions in the outer boundary directions. In this case, the additional regions in the outer boundary directions can be defined using data that is spatially disjoint within the image but has continuity in the actual image. For example, when the projection format of the 360-degree image is equirectangular projection, and the left and right boundaries of the image are continuous in the actual image, and the top and bottom boundaries of the image are also continuous in the actual image, the left boundary of parallel block 0 is continuous with the right boundary of parallel block 2, and the top boundary of parallel block 0 is continuous with the bottom boundary of parallel block 6. Therefore, in parallel block 0, the left extended region can be obtained from parallel block 2, the right extended region can be obtained from parallel block 1, the top extended region can be obtained from parallel block 6, and the bottom extended region can be obtained from parallel block 3. Furthermore, in parallel block 1, the upper left extended region can be obtained from parallel block 8, the lower left extended region can be obtained from parallel block 5, the upper right extended region can be obtained from parallel block 7, and the lower right extended region can be obtained from parallel block 4.

[0221] because Figure 9a The L0 block in the image is located at the boundary of the parallel blocks, so data that can be referenced from the left, upper left, lower left, upper, and upper right blocks (similar to the case of U0) may not exist. In this case, even if they are not spatially adjacent in the 2D image, there can still be blocks that are continuous in the actual image. Therefore, as mentioned above, when the projection format of the 360-degree image is equirectangular projection, the left and right boundaries of the image are continuous in the actual image, and the upper and lower boundaries of the image are continuous in the actual image, the left and lower left blocks of the L0 block can be obtained from the second parallel block, the upper left block of the L0 block can be obtained from the eighth parallel block, and the upper and upper right blocks of the L0 block can be obtained from the sixth parallel block.

[0222] Table 1 below is the pseudocode for retrieving data corresponding to the appended region from other regions that have continuity.

[0223] Table 1

[0224] i_pos'=overlap(i_pos,minI,maxI)

[0225] overlap(A,B,C)

[0226] {

[0227] if(A <B)output=(A+C-B+1)%(C-B+1)

[0228] else if(A>C)output=A%(C-B+1)

[0229] else output = A

[0230] }

[0231] Referring to the pseudocode in Table 1, the variable i_pos (corresponding to variable A) in the overlap function is the input pixel position, i_pos' is the output pixel position, minI (corresponding to variable B) is the minimum value of the pixel position range, maxI (corresponding to variable C) is the maximum value of the pixel position range, and i is the position component (horizontal, vertical, etc. in this example). In this example, minI can be 0, and maxI can be Pic_width (horizontal width of the image) - 1 or Pic_height (vertical width of the image) - 1.

[0232] For example, suppose the vertical width of the image (general image) ranges from 0 to 47 and the image is arranged according to... Figure 7 The data is divided as shown. When it is necessary to set an additional region of size m below parallel block 4 and fill the additional region with data from the upper end of parallel block 7, the location from which data needs to be obtained can be determined as described above.

[0233] When the vertical length of parallel block 4 ranges from 16 to 30 and an append region of size 4 needs to be set downwards, the data at positions 31, 32, 33, and 34 can be filled into the append region of parallel block 4. In this case, since min and max in the above formula are 0 and 47 respectively, the output values ​​for positions 31 to 34 will be their own values, i.e., 31 to 34. That is, the data to be filled into the append region are the data at positions 31 to 34.

[0234] Alternatively, assume the horizontal length of the image (360-degree image, equirectangular projection, with continuity at both ends) ranges from 0 to 95 and the image is arranged according to... Figure 7 The partitioning is performed as shown. When it is necessary to set an append region of size m to the left of parallel block 3 and fill the append region with data from the right side of parallel block 5, the above method can be used to determine from which position the data needs to be obtained.

[0235] When the vertical length of parallel block 3 ranges from 0 to 31 and an additional region of size 4 needs to be set to the left, the data at positions -4, -3, -2, and -1 can be filled into the additional region of parallel block 3. Since these positions do not exist within the horizontal length range of the image, the data needs to be obtained from which position using the aforementioned known calculation. In this case, because min and max in the above formula are 0 and 95 respectively, the output values ​​for -4 to -1 will be 92 to 95. That is, the data to be filled into the additional region is the data at positions 92 to 95.

[0236] Specifically, when the aforementioned region of size m contains data between 360 and 380 degrees (assuming the pixel value position range is 0 to 360 degrees in this example), by adjusting it to the range within the image, it can be understood as similar to acquiring data from a region between 0 and 20 degrees. That is, it can be acquired based on the pixel value position range from 0 to Pic_width-1.

[0237] In other words, in order to obtain data for the additional region, the location of the data to be obtained can be confirmed through an overlapping process.

[0238] The above example illustrates the case of acquiring a surface from a 360-degree image (except where the image boundaries are continuous at both ends), and assumes that spatially adjacent regions within the image are continuous. However, depending on the projection format (e.g., cube map projection), when there are two or more surfaces and each surface has undergone configuration or reconfiguration, there may be cases where even spatially adjacent areas within the image lack continuity. In the cases described above, the location data that are continuous on the actual image can be confirmed using the surface configuration or reconfiguration information, and additional regions can be generated.

[0239] Table 2 below shows the pseudocode for generating additional regions related to the aforementioned specific segmentation unit using the internal data of that specific segmentation unit.

[0240] Table 2

[0241] i_pos'=clip(i_pos,minI,maxI)

[0242] clip(A,B,C)

[0243] {

[0244] if(A <B)output=B

[0245] else if (A>C)output=C

[0246] else output = A

[0247] }

[0248] Since the meanings of the variables in Table 2 are the same as those in Table 1, their detailed explanations will be omitted here. However, in this example, minI can be the left or top coordinate of a specific division unit, while maxI can be the right or bottom coordinate of each unit.

[0249] For example, when the image is according to such Figure 7 The method shown is used for segmentation. The horizontal length of parallel block 2 ranges from 32 to 47 pixels. When an additional region of size m needs to be added to the right of parallel block 2, the data at positions 48, 49, 50, and 51 can be filled using the data at position 47 (corresponding to the interior of parallel block 2) output by the above formula. That is, according to Table 2, the additional region associated with a specific segmentation unit can be generated by copying the outline pixels of the corresponding segmentation unit.

[0250] In other words, in order to obtain data for the additional region, the required location can be confirmed through a clipping process.

[0251] The detailed structure of Table 1 or Table 2 above is not fixed and can be changed. For example, the applicable overlay method for 360-degree images can be changed to take into account the configuration (or reconfiguration) of the surfaces and the characteristics of the coordinate system between the surfaces.

[0252] Figure 10 This is an example diagram illustrating the application of an additional region generated in one embodiment of the present invention during the encoding / decoding process in other regions.

[0253] Furthermore, since the additional region applicable to one embodiment of the present invention is generated using image data from other regions, it can be equivalent to duplicated image data. Therefore, in order to prevent the existence of unnecessary duplicate data, the additional region can be removed after encoding / decoding. However, before removing the additional region, it is advisable to consider removing it after applying the additional region to the encoding / decoding process.

[0254] See Figure 10 It can be confirmed that region A of segmentation unit I generates the additional region B of segmentation unit J. At this point, before removing the generated region B, region B can be applied to the encoding / decoding (specifically, the reconstruction or correction process) of region A contained in segmentation unit I.

[0255] Specifically, assuming that the segmentation unit I and the segmentation unit J are respectively Figure 7 Given parallel blocks 0 and 1, the rightmost part of segmentation unit I and the leftmost part of segmentation unit J have image continuity. The appended region B, after being used as a reference in the image encoding / decoding of segmentation unit J, can also be used when encoding / decoding A. Specifically, although A and region B obtained data from region A when generating the appended region, reconstruction may utilize some different values ​​(including quantization errors) during the encoding / decoding process. Therefore, when reconstructing segmentation unit I, the part corresponding to region A can be reconstructed using the reconstructed image data of region A and the image data of region B. For example, a portion of region C in segmentation unit I can be replaced by the average or weighted sum of the values ​​of regions A and B. This is because data from two or more identical regions exists, thus the reconstructed image (region C, where A in segmentation unit I is replaced by C) can be obtained using the data from both regions (this process is named Rec_Process in the attached figure).

[0256] Furthermore, a portion of region C contained within segmentation unit I can be replaced using regions A and B depending on which segmentation unit it is closer to. Specifically, because the image data within a certain range (e.g., M pixel intervals) on the left side of region C is closer to segmentation unit I, data from region A can be used for reconstruction. Similarly, because the image data within a certain range (e.g., N pixel intervals) on the right side of region C is closer to segmentation unit J, data from region B can be used for reconstruction. The result, expressed as a formula, is shown in Equation 1 below.

[0257] [Formula 1]

[0258] C(x,y)=A(x,y),(x,y)∈M

[0259] B(x,y),(x,y)∈N

[0260] Furthermore, a portion of region C contained within segmentation unit I can be replaced by assigning weighted values ​​to image data in regions A and B respectively, based on which segmentation unit it is closer to. That is, for image data in region C that is closer to segmentation unit I, a higher weighted value can be assigned to image data in region A, while for image data in region C that is closer to segmentation unit J, a higher weighted value can be assigned to image data in region B. In other words, the weighted value can be set based on the distance difference between the horizontal width of region C and the x-coordinate of the pixel value to be corrected.

[0261] As a formula for setting adaptive weighting values ​​for regions A and B, Formula 2 as described below can be derived.

[0262] [Formula 2]

[0263] C(x,y)=A(x,y)x w+B(x,y)x(1-w)

[0264] w = f(x, y, k)

[0265] Referring to Formula 2, w is the weighting value assigned to the pixels in regions A and B as (x, y). In this case, as the average of the weighting values ​​for regions A and B, the pixels in region A are multiplied by the weighting value w, and the pixels in region B are multiplied by 1-w. However, in addition to the average weighting value, different weighting values ​​can also be assigned to regions A and B respectively.

[0266] After using the appended region as described above, the appended region B can be removed during the resizing process of the segmentation unit J and stored in memory (decoded picture buffer, DPB). (In this example, it is assumed that the process of setting the appended region is resizing.) <sizing>The above process can be derived through some of the above embodiments (e.g., by adding region markers for confirmation, then confirming the size information, then confirming the filling method, etc.). It is assumed that the enlargement process is performed during the size adjustment process, while the shrinking process is performed during the size readjustment process (which can be derived through the reverse process of the above).

[0267] Furthermore, (specifically, immediately after the encoding / decoding of the corresponding image is completed) it can be directly stored in memory without resizing, and then in the output step (assumed to be display in this example). <display>The removal is performed by resizing the image during the step. This can be applied to all or part of the segmented units contained in the corresponding image.

[0268] The aforementioned settings can be processed implicitly or explicitly based on the encoding / decoding settings. When using the implicit method (specifically based on the characteristics, type, format, etc. of the image, or based on other encoding / decoding settings <in this example, settings related to the appended area>), the settings can be determined without generating relevant syntax elements. When using the explicit method, the settings related to the removal of the appended area can be adjusted by generating relevant syntax elements. The related units can include video, sequence, image, sub-image, strip, parallel block, etc.

[0269] Furthermore, existing segmentation-based encoding methods can include: 1) the steps of segmenting an image into one or more parallel blocks (or collectively referred to as segmentation units) and generating segmentation information; 2) the steps of performing encoding according to the segmented parallel block units; 3) the steps of performing filtering with information indicating whether loop filtering is permissible at the boundaries of the parallel blocks; and 4) the steps of storing the filtered parallel blocks in memory.

[0270] Furthermore, existing segmentation-based decoding methods can include: 1) dividing an image into one or more parallel blocks based on parallel block segmentation information; 2) performing decoding according to the segmented parallel block units; 3) performing filtering with information indicating whether loop filtering is permissible at the boundaries of the parallel blocks; and 4) storing the filtered parallel blocks in memory.

[0271] The third step in the encoding / decoding method is a post-processing step of encoding / decoding. When filtering is performed, it can be a dependent encoding / decoding, while when filtering is not performed, it can be an independent encoding / decoding.

[0272] An encoding method for segmentation units according to an embodiment of the present invention may include: 1) dividing an image into one or more parallel blocks and generating segmentation information; 2) setting an append region for at least one segmented parallel block unit and filling the append region with adjacent parallel block units; 3) encoding the parallel block unit containing the append region; 4) removing the append region of the parallel block unit and performing filtering based on information indicating whether loop filtering is permissible at the boundaries of the parallel blocks; and 5) storing the filtered parallel blocks in a memory.

[0273] Furthermore, the decoding method for the segmentation unit applicable to one embodiment of the present invention can include: 1) dividing an image into one or more parallel blocks based on parallel block segmentation information; 2) setting an additional region for the segmented parallel block unit and filling the additional region using decoding information, pre-set information, or other (adjacent) parallel block units that have been pre-reconstructed; 3) encoding the parallel block unit containing the additional region using decoding information received from the encoding device; 4) removing the additional region of the parallel block unit and performing filtering based on information indicating whether loop filtering is permissible at the boundary of the parallel block; and 5) storing the filtered parallel block in a memory.

[0274] In the encoding / decoding method for segmented units applicable to one embodiment of the present invention as described above, the second step can be a pre-encoding / decoding process (dependent encoding / decoding when an appended region is set, otherwise independent encoding / decoding). Furthermore, the fourth step can be a post-encoding / decoding process (dependent when filtering is performed, otherwise independent). In this example, an appended region is used during the encoding / decoding process, and a process of resizing to the initial size of the parallel block is performed before it is stored in memory.

[0275] First, the encoder divides the image into multiple parallel blocks. Depending on implicit or explicit settings, additional regions are defined for each parallel block unit, and relevant data is obtained from adjacent regions. Next, encoding is performed on updated parallel block units containing the original parallel blocks and the additional regions. After encoding is complete, the additional regions are removed, and filtering is performed according to the applicable loop filtering settings.

[0276] At this point, different filtering settings can be applied depending on the filling and removal methods of the appended regions. For example, when performing simple removal, the loop filtering settings described above can be followed, while when removing using overlapping regions, filtering can be omitted or other filtering settings can be followed. That is, because the distortion in the parallel block boundary region can be significantly reduced by utilizing overlapping data, different filtering settings can be applied regardless of whether loop filtering is applied to the parallel block unit or not (e.g., applying a weaker filter at the parallel block boundary). The data is then stored in memory after the above process.

[0277] In the decoder, the image is first segmented into multiple parallel blocks based on the parallel block segmentation information received from the encoder. Next, information related to the appended regions is explicitly or implicitly confirmed, and the updated encoding information of the parallel blocks received from the encoder is parsed after setting the appended regions. Decoding is then performed on a per-updated parallel block basis. After decoding, the appended regions are removed, and filtering is performed using the same loop filtering settings as the encoder. Detailed information related to this has already been described in the encoder section, so it will be omitted here. The image is then stored in memory after the above process.

[0278] Furthermore, it is also possible to consider cases where additional regions added during the encoding / decoding process using segmented units are not removed and are directly stored in memory. For example, in cases such as 360-degree images, the accuracy of prediction may decrease in certain prediction processes (e.g., inter-frame prediction) depending on surface configuration settings (e.g., difficulty in accurately locating locations with discontinuous surface configurations during motion detection and compensation). Therefore, additional regions can be stored in memory and used during the prediction process to improve prediction accuracy. When used in inter-frame prediction, the additional regions (or images containing the additional regions) can be used as reference images for performing inter-frame prediction.

[0279] The encoding method for storing the appended region can include: 1) dividing the image into one or more parallel blocks and generating segmentation information; 2) setting an appended region for at least one segmented parallel block unit and filling the appended region with adjacent parallel block units; 3) encoding the parallel block unit containing the appended region; 4) storing the appended region of the parallel block unit (in which case the application of loop filtering can be omitted); and 5) storing the encoded parallel block in memory.

[0280] The decoding method for storing the additional region may include: 1) dividing the image into one or more parallel blocks based on parallel block segmentation information; 2) setting an additional region for the segmented parallel block unit and filling the additional region with decoding information, pre-set information, or other (adjacent) parallel block units that have been pre-reconstructed; 3) encoding the parallel block unit containing the additional region using decoding information received from the encoding device; 4) storing the additional region of the parallel block unit (in which case loop filtering may be omitted); and 5) storing the decoded parallel block in memory.

[0281] When storing the additional regions, the encoder first divides the image into multiple parallel blocks. Depending on whether it's implicit or explicit, additional regions are defined for the parallel blocks, and relevant data is obtained from pre-defined regions. These pre-defined regions refer to other regions that are relevant based on the surface configuration of the 360-degree image; therefore, they can be regions adjacent to or non-adjacent to the current parallel block. Encoding is then performed on updated parallel block units. Because the additional regions are stored after decoding, filtering is not performed regardless of the loop filtering settings. This is because the boundaries of the updated parallel blocks are not shared with the actual parallel block boundaries due to the additional regions. The data is then stored in memory after the above process.

[0282] When storing the appended regions, the decoder first verifies the parallel block segmentation information received from the encoder and divides the image into multiple parallel blocks based on this. Next, it verifies the information related to the appended regions and, after setting the appended regions, parses the updated encoding information of the parallel blocks received from the encoder. Decoding is then performed on units of updated parallel blocks. After decoding is complete, the appended regions are directly stored in memory without applying loop filtering.

[0283] Next, the encoding / decoding method for segmentation units applicable to one embodiment of the present invention, as described above, will be explained with reference to the accompanying drawings.

[0284] Figures 11 to 12 This is a flowchart illustrating an encoding / decoding method for segmentation units to which one embodiment of the present invention applies. Specifically, as an example of generating appended regions and performing encoding / decoding in each segmentation unit, in Figure 11 The encoding method involving appended regions is illustrated in the figure. Figure 12 The decoding method for removing appended regions is illustrated in the image. Specifically, 360-degree images can... Figure 11 The previous steps involved preprocessing (slicing, projection), and... Figure 12 The subsequent steps involve post-processing (rendering, etc.).

[0285] First refer to Figure 11 After the encoder acquires the input image (step A), the image segmentation unit divides the input image into two or more segmentation units (at this time, setting information related to the segmentation method can be generated, referred to as step B). Next, based on the encoding settings or whether additional regions are supported, additional regions are generated for the segmentation units (step C). Then, the segmentation units containing the additional regions are encoded and a bitstream is generated (step D). Furthermore, after generating the bitstream, it can be determined whether to perform a size readjustment (or whether to delete the additional regions, step E) based on the encoding settings. Then, the encoded data containing or with the additional regions removed (the image in step D or E) is stored in the memory (step E).

[0286] See Figure 12 The decoder, referring to segmentation-related setting information obtained by parsing the received bitstream, segments the image to be decoded into two or more segmentation units (step B). Next, based on the decoding settings obtained from the received bitstream, it sets the size of the appended region for each segmentation unit (step C). Then, it decodes the image data contained in the bitstream to obtain image data containing the appended regions (step D). Next, it generates a reconstructed image by deleting the appended regions (step E), and then outputs the reconstructed image to the display (step F). At this time, it is possible to determine whether to delete the appended regions based on the decoding settings, and store the decoded image or image data (data in step D or E) in the memory. In addition, step F may include a process of reducing the reconstructed image to a 360-degree image by surface reconfiguration.

[0287] In addition, according to Figure 11 or Figure 12 Whether or not additional regions are removed, loop filtering can be adaptively performed at the boundary of the segmentation unit (in this example, a block filter is assumed, but other loop filters can also be applied). Furthermore, loop filtering can be adaptively performed based on whether the generation of additional regions is permitted.

[0288] When the appended region is removed and stored in memory, the loop filter can be explicitly applied or not applied based on a flag indicating whether loop filtering is applicable on the boundary of the segmentation unit (specifically, the initial state) such as loop_filter_across_enabled_flag (in this example, for parallel blocks).

[0289] Alternatively, it could not support a flag indicating whether loop filtering is applicable at the boundary of the segmentation unit, but instead implicitly determine whether filtering is applicable and the filtering settings, as in the examples described later.

[0290] Furthermore, even when there is image continuity between individual segmentation units, the image continuity at the boundaries between segmentation units with newly generated regions may disappear after additional regions are generated for each segmentation unit. Applying loop filtering in this case would lead to an unnecessary increase in computation and a decrease in coding performance; therefore, loop filtering can be implicitly avoided.

[0291] Furthermore, based on the surface configuration of the 360-degree image, adjacent segmentation units in 2D space may lack image continuity with each other. Performing loop filtering on the boundaries between segmentation units lacking image continuity, as described above, could lead to a degradation in image quality. Therefore, loop filtering can be implicitly omitted for boundaries between segmentation units lacking image continuity.

[0292] Furthermore, when following as Figure 10 The method described above, where weights are assigned to two regions and a portion of the current segmentation unit is replaced, allows for loop filtering because the boundaries of each segmentation unit are internal boundaries in the appended region. However, since coding errors can be further reduced by weighting a portion of the current region contained in other regions, loop filtering may not be necessary. Therefore, in the case described above, loop filtering can be implicitly omitted.

[0293] Furthermore, the applicability of loop filtering can be determined (specifically, additionally at the corresponding boundary) based on a flag used to indicate whether loop filtering is applicable. When the aforementioned flag is activated, filtering can be applied based on loop filtering settings and conditions applicable within the segmentation unit, or filtering can be applied at the boundary of the segmentation unit (specifically, using loop filtering settings and conditions different from those used at the boundary of the segmentation unit).

[0294] The above embodiments assume that the additional region is stored in memory after removal, but some of this can also occur in other output steps (specifically, it can belong to either the loop filter section or the post-filter section). <postfilter>The process of execution on (etc.).

[0295] The above examples illustrate the scenario where the additional region is supported in all directions of each segmentation unit. However, when the additional region is supported only in a portion of the directions according to its settings, only a portion of the above can be applied. For example, the original settings can be applied to boundaries where the additional region is not supported, while the application of the above examples can be varied to boundaries where the additional region is supported. In other words, the application of the above methods can be adaptively determined in all or part of the unit boundaries based on the settings of the additional region.

[0296] The aforementioned settings can be processed implicitly or explicitly based on the encoding / decoding settings. When using the implicit method (specifically based on the characteristics, type, format, etc. of the image, or based on other encoding / decoding settings <in this example, settings related to the appended area>), it can be determined without generating relevant syntax elements. When using the explicit method, it can be adjusted by generating relevant syntax elements. The related units can include video, sequence, image, sub-image, strip, parallel block, etc.

[0297] Next, we will explain in detail the methods for determining whether the segmentation unit and the appended region are referable. Here, referentiality is considered dependent encoding / decoding, while non-referability is considered independent encoding / decoding.

[0298] An additional region applicable to one embodiment of the present invention can be referenced or restricted during the encoding / decoding process of the current image or other images. Specifically, an additional region removed before being stored in memory can be referenced or restricted during the encoding / decoding process of the current image. Furthermore, an additional region added after being stored in memory can be referenced or restricted during the encoding / decoding process of images at different times, in addition to the current image.

[0299] In other words, the reference probability and range of the aforementioned additional region can be determined based on the encoding / decoding settings. Through the aforementioned settings, the additional region of the current image will be stored in memory after encoding / decoding, meaning it can be referenced or limited to reference by being included in a reference image of other images. This can be applied to all or part of the segmentation units contained in the corresponding image. The cases described in the examples below can also be modified to apply to the current examples.

[0300] The aforementioned settings related to the reference probability of the additional region can be processed implicitly or explicitly based on the encoding / decoding settings. When using the implicit method (specifically based on the characteristics, type, format, etc. of the image, or based on other encoding / decoding settings <in this example, settings related to the additional region>), it can be determined without generating relevant syntax elements. When using the explicit method, the settings related to the reference probability of the additional region can be adjusted by generating relevant syntax elements. The related units can include video, sequence, image, sub-image, strip, parallel block, etc.

[0301] Typically, a subset of units in the current image (assumed in this example as segmentation units obtained by the image segmentation unit) can reference the data of the current unit, but cannot reference the data of other units. Furthermore, a subset of units in the current image can reference the data of all units existing in other images. The above description is an example related to the general properties of units obtained by the image segmentation unit; additional properties related to this can also be defined.

[0302] Furthermore, it is possible to define flags that indicate whether other segmentation units within the current image can be referenced, and whether segmentation units contained in other images can be referenced.

[0303] As an example, it is permissible to reference segmentation units contained in other images that are at the same location as the current segmentation unit, but reference to segmentation units at different locations from the current segmentation unit is restricted. For instance, when multiple bitstreams encoding the same image under different encoding settings are transmitted, and the decoder selectively determines the bitstream used to decode various regions (segmentation units) in the image (assuming decoding is performed in parallel block units in this example), because it is necessary to restrict the possibility of reference between segmentation units in the same and different spaces, it is possible to perform encoding / decoding in a manner that only allows reference to the same region in different images.

[0304] As an example, referencing can be allowed or restricted based on the identifier information associated with the segmentation unit. For instance, referencing is allowed when the identifier information assigned to the segmentation unit is the same, but not when they are different. In this case, the identifier information can be information used to indicate that encoding / decoding was performed (dependently) in an environment where mutual reference is possible.

[0305] The aforementioned settings can be processed implicitly or explicitly based on the encoding / decoding settings. When using the implicit method, the settings can be determined without generating relevant syntax elements, while when using the explicit method, the settings can be processed by generating relevant syntax elements. The related units can include video, sequence, image, sub-image, strip, parallel block, etc.

[0306] Figures 13a to 13g It is an illustrative diagram used to explain the reference area of ​​a specific unit of division. Figures 13a to 13g In the diagram, the area indicated by the thick outline represents a reference area.

[0307] See Figure 13a This allows for the identification of various reference arrows used to perform inter-frame prediction. In this case, blocks C0 and C1 represent unidirectional inter-frame prediction. Block C0 can obtain the RP0 reference block before the current image and the RF0 reference block after the current image. Block C2 represents bidirectional inter-frame prediction, obtaining the RP1 and RF1 reference blocks from images before or after the current image. An example of obtaining one reference block each in the preceding and following directions is illustrated in the accompanying figure, but it is also possible to obtain reference blocks only in the preceding or following directions. Block C3 represents non-directional inter-frame prediction, obtaining the RC0 reference block within the current image. An example of obtaining one reference block is illustrated in the accompanying figure, but it is also possible to obtain two or more reference blocks.

[0308] In the examples described below, the focus will be on the reference probability of pixel values ​​and prediction mode information in inter-frame prediction based on segmentation units. However, it can also be understood that it includes other encoded / decoded information that can be referenced spatially or temporally (such as intra-frame prediction mode information, transform and quantization information, loop filtering information, etc.).

[0309] See Figure 13b The current image Currnt(t) is divided into two or more parallel blocks, and a portion of these parallel blocks, block C0, can obtain reference blocks P0 and P1 by performing unidirectional inter-image prediction. A portion of these parallel blocks, block C1, can obtain reference blocks P3 and F0 by performing bidirectional inter-image prediction. That is, this can be understood as an instance where referencing blocks at other locations within other images is permitted without restrictions such as location limitations or limitations allowing reference only within the same image.

[0310] See Figure 13c The image is segmented into two or more parallel block units, and a portion of the parallel blocks, C1, can obtain reference blocks P2 and P3 by performing unidirectional inter-frame prediction. A portion of the parallel blocks, C0, can obtain reference blocks P0, P1, F0, and F1 by performing bidirectional inter-frame prediction. A portion of the parallel blocks, C3, can obtain reference block FC0 by performing non-directional inter-frame prediction.

[0311] Right now, Figure 13b as well as Figure 13c This can be understood as an instance that allows referencing of blocks in other locations contained in other images without restrictions such as location or restrictions that only allow referencing within the same image.

[0312] See Figure 13d The current image is segmented into two or more parallel block units. Block C0 within a subset of these parallel blocks can obtain reference block P0 through forward-facing inter-frame prediction, but cannot obtain reference blocks P1, P2, and P3 contained within the subset of parallel blocks. Block C4 within a subset of these parallel blocks can obtain reference blocks F0 and F1 through backward-facing inter-frame prediction, but cannot obtain reference blocks F2 and F3. A portion of block C3 within a subset of these parallel blocks can obtain reference block FC0 through unfacing inter-frame prediction, but cannot obtain reference block FC1.

[0313] That is, in Figure 13d The system can allow or restrict referencing based on whether the image (t-1, t, t+1 in this example) is segmented or not, and the encoding / decoding of the image segmentation unit. Specifically, it can allow referencing only of blocks contained in a parallel block that have the same identifier information as the current parallel block.

[0314] See Figure 13e The image is divided into two or more parallel block units, and a portion of these parallel blocks, C0, can obtain reference blocks P0 and F0 through bidirectional inter-image prediction, but cannot obtain reference blocks P1, P2, P3, F1, F2, and F3. That is, Figure 13e It can be an instance that only allows reference to parallel blocks that are in the same position as the parallel block containing the current block.

[0315] See Figure 13f The image is divided into two or more parallel block units, and a portion of the parallel blocks C0 can obtain reference blocks P1 and F2 by performing bidirectional inter-image prediction, but cannot obtain reference blocks P0, P2, P3, F0, F1, and F3. Figure 13f This can be an instance in which the bitstream contains information indicating the parallel blocks that can be referenced in the current segmentation unit, and the referenced parallel blocks are confirmed using the aforementioned information.

[0316] See Figure 13g The image is segmented into two or more parallel blocks. A portion of these parallel blocks, block C0, can obtain reference blocks P0, P3, and P5 through unidirectional inter-picture prediction, but cannot obtain reference block P4. Similarly, a portion of these parallel blocks, block C1, can obtain reference blocks P1, F0, and F2 through bidirectional inter-picture prediction, but cannot obtain reference blocks P2 and F1.

[0317] Figure 13g This is an instance that allows or restricts referencing based on whether the image (t-3, t-2, t-1, t, t+1, t+2, t+3 in this example) is segmented, the encoding / decoding settings of the image segmentation units (in this example, it is assumed to be determined based on the identifier information of the segmentation unit, the identifier information of the image unit, whether the segmentation unit is the same region, whether the segmentation unit is similar to the region, the bitstream information of the segmentation unit, etc.), etc. Specifically, the referenceable parallel block can be located in the same or similar position as the current block within the image, can have the same identifier information as the current block (specifically, in the image unit or segmentation unit), and can be the same as the bitstream used to acquire the current parallel block.

[0318] Figures 14a to 14e This is a flowchart illustrating the possibilities of adding a region in a segmentation unit to which one embodiment of the present invention applies. Figures 14a to 14e In the diagram, the area indicated by the thick outline represents the reference area, while the area indicated by the dashed line represents the additional area of ​​the dividing unit.

[0319] In one embodiment of the invention, the reference possibilities for a portion of the image (other images located before or after in time) can be restricted or permitted. Furthermore, the reference possibilities for the entire expanded segmentation unit, including the appended region, can be restricted or permitted. Additionally, the reference possibilities for only the initial segmentation unit excluding the appended region can be restricted or permitted. Furthermore, the reference possibilities for the boundary between the appended region and the initial segmentation unit can be restricted or permitted.

[0320] See Figure 14a A subset of parallel blocks, block C0, can obtain reference blocks P0 and P1 by performing unidirectional inter-frame prediction. A subset of parallel blocks, block C2, can obtain reference blocks P2, P3, F0, and F1 by performing bidirectional inter-frame prediction. A subset of parallel blocks, block C1, can obtain reference block FC0 by performing non-directional inter-frame prediction. Block C0 can obtain reference blocks P0, P1, P2, P3, and F0 from the initial parallel block regions (excluding the appended regions) of a subset of reference images t-1 and t+1. Block C2 can obtain reference blocks P2 and P3 from the initial parallel block region of reference image t-1 while simultaneously obtaining reference block F1 from the parallel block region of reference image t+1 containing the appended region. At this point, as shown by reference block F1, a reference block containing the boundary between the appended region and the initial parallel block region can be obtained.

[0321] See Figure 14b A subset of parallel blocks, C0, C1, and C3, can obtain reference blocks P0, P1, P2 / F0, F2 / F1, F3, and F4 by performing unidirectional inter-picture prediction. A subset of parallel blocks, C2, can obtain reference blocks FC0, FC1, and FC2 by performing non-directional inter-picture prediction.

[0322] Some blocks C0, C1, and C3 can obtain reference blocks P0, F0, and F3 from the initial parallel block regions of a portion of the reference image (t-1 and t+1 in this example), and can also obtain reference blocks P1, x, and F4 from the updated parallel block region boundaries, and can also obtain reference blocks P2, F2, and F1 from outside the updated parallel block region boundaries.

[0323] A portion of block C2 can obtain reference block FC1 from the initial parallel block region of a portion of the reference image (t in this example), and can also obtain reference block FC3 from the updated parallel block region boundary, and can also obtain reference block FC0 from outside the updated parallel block region boundary.

[0324] Among them, some blocks C0 can be blocks located in the initial parallel block region, some blocks C1 can be blocks located at the boundary of the updated parallel block region, and some blocks C3 can be blocks located outside the boundary of the updated parallel block region.

[0325] See Figure 14c The image is divided into two or more parallel block units. In one part of the image, additional regions are set for some parallel blocks; in another part, no additional regions are set for some parallel blocks; and in yet another part, no additional regions are set. Within one set of parallel blocks, blocks C0 and C1 can obtain reference blocks P2, F1, F2, and F3 through unidirectional inter-frame prediction, but cannot obtain reference blocks P0, P1, P3, and F0. Similarly, within another set of parallel blocks, block C2 can obtain reference blocks FC1 and FC2 through non-directional inter-frame prediction, but cannot obtain reference block FC0.

[0326] A portion of block C2 cannot obtain reference block FC0 from the initial parallel block region of a portion of the reference image (t in this example), but can obtain reference block FC1 from the updated parallel block region (FC0 and FC1 can be the same region in the method of filling a portion of the appended region; although FC0 cannot be referenced in the initial unit of parallel block segmentation, it can be referenced when the corresponding region is moved to the current parallel block by appending the region).

[0327] A portion of block C2 can obtain reference block FC2 from a portion of the parallel block region of a portion of the reference image (t in this example). (Although by default, data in other parallel blocks of the current image cannot be referenced, it is assumed that reference is allowed when it is set to be referable by means of identifier information, etc., in the above embodiments.)

[0328] See Figure 14d The image is divided into two or more parallel block units and additional regions are defined. A portion of the parallel blocks, block C0, can obtain reference blocks P0, F0, F1, and F3 by performing bidirectional inter-image prediction, but cannot obtain reference blocks P1, P2, P3, and F2.

[0329] A portion of block C0 can obtain reference block P0 from the initial parallel block region (parallel block 0) of a portion of reference image t-1, but cannot obtain reference block P3 from the boundary of the extended parallel block region, nor can it obtain reference block P2 from the outer boundary of the extended parallel block region (i.e., the appended region).

[0330] A portion of block C0 can obtain reference block F0 from the initial parallel block region (parallel block 0) of a portion of reference image t+1, and can also obtain reference block F1 from the boundary of the extended parallel block region, and can also obtain reference block F3 from outside the boundary of the extended parallel block region.

[0331] See Figure 14e The image is segmented into two or more parallel block units, and an additional region with at least one size and shape is defined. In a subset of parallel blocks, block C0 can obtain reference blocks P0, P3, P5, and F0 by performing unidirectional inter-picture prediction, but cannot obtain reference block P2 located on the boundary between the additional region and the original parallel block. In a subset of parallel blocks, block C1 can obtain reference blocks P1, F2, and F3 by performing bidirectional inter-picture prediction, but cannot obtain reference blocks P4, F1, and F5.

[0332] As shown in the examples above, pixel values ​​can be used as references and can restrict the reference to other encoded / decoded information.

[0333] As an example, when the prediction unit searches for candidate groups of intra-frame prediction modes to be used in intra-frame prediction from spatially adjacent blocks, it can do so by means of... Figures 13a to 14e The method shown confirms whether a partition unit containing the current block can reference a partition unit containing adjacent blocks.

[0334] As an example, when the prediction unit searches for candidate groups of motion information needed for inter-frame prediction from temporally and spatially adjacent blocks, it can do so by means of... Figures 13a to 14e The method shown confirms whether the segmentation unit containing the current block can reference segmentation units containing blocks that are spatially adjacent or temporally adjacent to the current image.

[0335] As an example, when the loop filter unit searches for loop filter-related setting information from adjacent blocks, it can do so through methods such as... Figures 13a to 14e The method shown confirms whether a partition unit containing the current block can reference a partition unit containing adjacent blocks.

[0336] Figure 15 This is an example diagram illustrating blocks contained in the current image segmentation unit and blocks contained in other image segmentation units.

[0337] See Figure 15 In this example, spatially adjacent reference candidate blocks can be the blocks to the left, upper left, lower left, upper, or upper right of the current block. Furthermore, temporally adjacent reference candidate blocks can be the blocks to the left, upper left, lower left, upper, upper right, right, lower right, lower, or central of a block located in the same or corresponding position as the current block within a different picture that is temporally adjacent to the current picture. Figure 15 In the text, the thick outer frame line represents the boundary line of the dividing unit.

[0338] When the current block is M, the spatially adjacent blocks G, H, I, L, and Q can all be referenced.

[0339] When the current block is G, some of the spatially adjacent blocks A, B, C, F, and K can be referenced, while the remaining blocks can be restricted from reference. Whether or not reference is possible can be determined based on the reference-related settings between the partitioning units UC, ULC, and LC of the spatially adjacent blocks and the partitioning unit containing the current block.

[0340] When the current block is S, some blocks among the surrounding blocks s, r, m, w, n, x, t, o, and y that are at the same position as the current block in temporally adjacent images can be referenced, while the remaining blocks can be restricted from reference. Whether or not they can be referenced can be determined based on the reference correlation settings between the segmentation units RD, DRD, and DD of the surrounding blocks at the same position as the current block in temporally adjacent images and the unit containing the current block.

[0341] Depending on the current block position, when there are reference-restricted candidates, the block can be filled using candidates whose priority order is next in the candidate group, or it can be replaced by other candidates adjacent to the reference-restricted candidate.

[0342] For example, when the current block in the prediction within the image is G, and the reference of the left-up fast block is limited and the most likely pattern (MPM) candidate group is formed in the order of PDAEU, since A is unreferenceable, it is possible to form a candidate group by performing a validity check according to the remaining EU order, or to use B or F, which are spatially adjacent to A, to replace A.

[0343] Furthermore, when the current block in the inter-screen prediction is S, the temporally adjacent sitting block is restricted in reference, and the temporal candidate of the skipped pattern candidate group is y, since y is unreferenceable, a candidate group can be formed by performing validity checks in the order of spatially adjacent candidates or a mixture of idle and temporal candidates, or by using t, x, s, which are spatially adjacent to y, to replace y.

[0344] Figure 16 This is a hardware configuration diagram illustrating an image encoding / decoding apparatus to which one embodiment of the present invention is applied.

[0345] See Figure 16 An image encoding / decoding apparatus 200 according to one embodiment of the present invention may include: at least one processor 210; and a memory 220 storing instructions for instructing the at least one processor 210 to perform at least one step.

[0346] At least one processor 210 may refer to a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor for performing the methods applicable to embodiments of the present invention. The memory 120 and storage device 260 may be constituted by at least one of volatile storage media and non-volatile storage media, respectively. For example, the memory 220 may be constituted by at least one of read-only memory (ROM) and random access memory (RAM).

[0347] Furthermore, the image encoding / decoding apparatus 200 may also include a transceiver 230 for performing communication via a wireless communication network. Additionally, the image encoding / decoding apparatus 200 may also include an input interface device 240, an output interface device 250, a storage device 260, etc. The various components included in the image encoding / decoding apparatus 200 can be connected to each other via a bus 270 to perform communication.

[0348] At least one step may include: dividing the encoded image contained in the received bitstream into at least one segmentation unit by referring to the syntax elements obtained from the received bitstream; setting an append region for the at least one segmentation unit; and decoding the encoded image based on the segmentation unit after setting the append region.

[0349] The aforementioned step of decoding the encoded image may include: determining a reference block related to the current block in the encoded image that needs to be decoded, based on information contained in the bitstream that indicates the possibility of reference.

[0350] The aforementioned reference block can be a block contained at a location that overlaps with an additional region defined on a segmentation unit containing the aforementioned reference block.

[0351] Figure 17 This is an example diagram illustrating an in-screen prediction mode applicable to one embodiment of the present invention.

[0352] See Figure 17 A total of 35 prediction modes can be identified, which can be divided into 33 directional modes and 2 non-directional modes (mean (DC) and planar)). In this case, the directional mode can be identified by tilt (e.g., dy / dx) or angle information. The above example can refer to a candidate group of prediction modes related to the luminance component or the chrominance component. Alternatively, the chrominance component can support a subset of prediction modes (e.g., mean (DC), planar, vertical, horizontal, diagonal modes, etc.). Furthermore, after the prediction mode for the luminance mode is determined, the corresponding mode can be included in the prediction mode for the chrominance component, or a mode derived from the corresponding mode can be included in the prediction mode.

[0353] Furthermore, it can leverage the correlation between color spaces to apply reconstructed blocks from other color spaces that have already been encoded / decoded to the prediction of the current block, and can include supported prediction modes. For example, the chromatic aberration component can generate the prediction block for the current block using the reconstructed block of the luminance component corresponding to the current block.

[0354] Based on the encoding / decoding settings, candidate groups for prediction modes can be adaptively determined. The number of candidate groups can be increased to improve prediction accuracy, or decreased to reduce the number of bits in the prediction mode.

[0355] For example, one of the candidate groups such as Candidate Group A (67, 65 directional patterns and 2 non-directional patterns), Candidate Group B (35, 33 directional patterns and 2 non-directional patterns), and Candidate Group C (19, 17 directional patterns and 2 non-directional patterns) can be used. In this invention, unless otherwise explicitly stated, it is assumed that in-screen prediction is performed using a pre-defined prediction pattern candidate group (Candidate Group A).

[0356] Figure 18 This is a first example illustration illustrating the configuration of reference pixels used in in-picture prediction according to one embodiment of the present invention.

[0357] An in-frame prediction method for image decoding according to one embodiment of the present invention may include: a reference pixel construction step, a prediction block generation step that refers to the constructed reference pixels and utilizes one or more prediction modes, a step of determining an optimal prediction mode, and a step of encoding the determined prediction mode. Furthermore, the image decoding apparatus may include a reference pixel construction unit, a prediction block generation unit, a prediction mode determination unit, and a prediction mode encoding unit for performing the reference pixel construction step, the prediction block generation step, the prediction mode determination step, and the prediction mode encoding step. The process described above may omit some steps or add other steps, and may also be changed to a different order than that described above.

[0358] Furthermore, the in-frame prediction method in image decoding according to one embodiment of the present invention can generate a prediction block of the current block based on the prediction mode obtained by the syntax element received from the image encoding device after the reference pixel is formed.

[0359] The size and shape (M×N) of the current block for in-frame prediction can be obtained from the block segmentation unit, and the size can be from 4×4 to 256×256. In-frame prediction is usually performed in prediction block units, but it can also be performed in units such as coded blocks (or coded units) or transform blocks (or transform units) depending on the settings of the block segmentation unit. After confirming the block information, the reference pixel composition unit can construct reference pixels to be used in the prediction of the current block. At this time, the reference pixels can be stored in temporary memory (e.g., an array). <array>The temporary storage is managed by 1D, 2D arrays, etc., and is generated and removed during the prediction process within each frame of the block. The size of the temporary storage can be determined based on the composition of the reference pixels.

[0360] Reference pixels can be pixels contained in adjacent blocks (which can be called reference blocks) centered on the current block and located to the left, top, upper left, upper right, or lower left. However, they are not limited to these; other candidate block groups can also be used in the prediction of the current block. The adjacent blocks located to the left, top, upper left, upper right, or lower left can be blocks selected when performing encoding / decoding using a grid or zigzag scan. Adjacent blocks in other positions (e.g., right, bottom, lower right blocks) can also be used as reference pixels when the scan order is changed.

[0361] Furthermore, a reference block can be a block that corresponds to the current block in a color space different from the color space containing the current block. When using Y / Cb / Cr particle format, the color space can refer to one of Y, Cb, or Cr. Additionally, a block corresponding to the current block can be a block that has the same positional coordinates as the current block or a block with positional coordinates corresponding to the current block based on the proportion of its color components.

[0362] Furthermore, for ease of explanation, the explanation will be based on the premise that the reference blocks at the aforementioned pre-defined positions (left side, top side, upper left, upper right, lower left) are composed of a single block. However, depending on the block division, they can also be composed of multiple sub-blocks.

[0363] In other words, the adjacent regions of the current block can be used as reference pixel positions for prediction within the current block's image, and regions in other color spaces corresponding to the current block can be additionally used as reference pixel positions depending on the prediction mode. Besides the examples above, the defined reference pixel positions can also be determined based on the prediction mode, method, etc. For example, when generating prediction blocks using methods such as block matching, the reference pixel positions can be regions that have already been encoded / decoded before the current block of the current image, or regions included within the exploration range of already encoded / decoded regions (e.g., including the left, right, upper left, or upper right of the current block).

[0364] See Figure 18 The reference pixels used in the prediction within the current block (size M×N) can be derived from the pixels adjacent to the current block on the left, top, top left, top right, and bottom left. Figure 18 This is composed of Ref_L, Ref_T, Ref_TL, Ref_TR, and Ref_BL. At this point, Figure 18 The content marked as P(x,y) can refer to pixel coordinates.

[0365] In addition, pixels adjacent to the current block can be classified into at least one reference pixel level. For example, they can be classified into pixels ref_0 that are closest to the current block {pixels with a pixel value difference of 1 from the boundary pixels of the current block, p(-1, -1) to p(2M - 1, -1), p(-1, 0) to p(-1, 2N - 1)}. Then, the adjacent pixels {with a pixel value difference of 2 from the boundary pixels of the current block, p(-2, -2) to p(2M, -2), p(-2, -1) to p(-2, 2N)} are ref_1, and the subsequent adjacent pixels {with a pixel value difference of 3 from the boundary pixels of the current block, p(-3, -3) to p(2M + 1, -3), p(-3, -2) to p(-3, 2N + 1)} are ref_2, and so on. That is, the reference pixels can be classified into multiple reference pixel levels according to the pixel distance from the boundary pixels of the current block.

[0366] In addition, different reference pixel levels can be set for each adjacent block at this time. For example, when using the block adjacent to the upper end of the current block as a reference block, the reference pixels at the ref_0 level can be used, and when using the block adjacent to the upper right end as a reference block, the reference pixels at the ref_1 level can be used.

[0367] Among them, the set of reference pixels usually referred to when performing intra-picture prediction is included in the blocks adjacent to the lower left, left, upper left, upper end, and upper right end of the current block, and is the pixels belonging to the ref_0 level (the pixels closest to the boundary pixels). Unless otherwise specified in the following content, the above-mentioned pixels are used as the premise. However, it is also possible to use the set of reference pixels included in some of the adjacent blocks mentioned above, and it is also possible to use the pixels included in two or more levels as the set of reference pixels. Among them, the set of reference pixels or levels can be determined implicitly (pre-set in the encoding / decoding device) or explicitly (receiving information for determination from the encoding device).

[0368]

[0369] ​In this invention, the description will assume that lower index values ​​(incrementing by 1 from 0) are assigned starting from the reference pixel level most adjacent to the current block, but it is not limited to this. Furthermore, the reference pixel composition information described later can be generated under the index settings described above (such as assigning shorter bits to the smaller index when selecting from multiple reference pixel sets, etc.).

[0370] Furthermore, when there are two or more supported reference pixel levels, the weighted average value can be applied to each reference pixel contained in the two or more reference pixel levels.

[0371] For example, it is possible to utilize through the location located Figure 18 A prediction block is generated from reference pixels obtained by weighting the pixels in the ref_0 and ref_1 levels. At this time, depending on the prediction mode (e.g., the directionality of the prediction mode), the pixels for which the weighted sum is applied in each reference pixel level can be either integer or fractional pixels. Furthermore, a prediction block can be obtained by assigning weights (e.g., 7:1, 3:1, 2:1, 1:1, etc.) to the prediction block obtained using reference pixels in the 1st reference pixel level and the prediction block obtained using reference pixels in the 2nd reference pixel level, respectively. At this time, a higher weight can be assigned to the prediction block of the reference pixel level more adjacent to the current block.

[0372] Assuming that explicit information related to the composition of reference pixels is generated, it is possible to generate adaptive reference pixel composition indication information (adaptive_intra_ref_sample_enabled_flag in this example) on units such as video, sequence, image, strip, parallel block, etc.

[0373] When the above indication information represents an adaptive reference pixel configuration that allows for adaptive configuration (adaptive_intra_ref_sample_enabled_flag = 1 in this example), adaptive reference pixel configuration information (adaptive_intra_ref_sample_flag in this example) can be generated on units such as images, stripes, parallel blocks, and blocks.

[0374] When the above composition information represents an adaptive reference pixel composition (adaptive_intra_ref_sample_flag = 1 in this example), reference pixel composition related information (such as selection information related to reference pixel level and set, which is intra_ref_idx in this example) can be generated on units such as images, strips, parallel blocks, and blocks.

[0375] At this point, when adaptive reference pixel construction is not allowed or is not adaptive reference pixel construction, reference pixels can be constructed according to predefined settings. For example, the most adjacent pixels in adjacent blocks are usually used to construct reference pixels, but this is not limited to this. Several other cases are also allowed (for example, selecting ref_0 and ref_1 as reference pixel levels and using ref_0 and ref_1 to generate predicted pixel values ​​through weighted summation, i.e., the default case).

[0376] Furthermore, reference pixel composition information (such as selection information related to reference pixel level or set) can be constructed after excluding pre-set information (such as the case where the reference pixel level is pre-set to ref_0) (e.g., ref_1, ref_2, ref_3, etc.), but is not limited to this.

[0377] The above examples illustrate a portion of the situation related to the composition of reference pixels. However, the in-frame prediction settings can be determined by combining various encoding / decoding information. This encoding / decoding information can include, for example, image type, color components, the size and shape of the current block, prediction mode {prediction mode type (directional, non-directional), prediction mode direction (vertical, horizontal, diagonal 1, diagonal 2, etc.)}, and the in-frame prediction settings (in this example, the reference pixel composition settings) can be determined based on the encoding / decoding information of adjacent blocks and the combination of the encoding / decoding information of the current block and adjacent blocks.

[0378] Figures 19a to 19c This is a second example illustration illustrating the configuration of reference pixels to which one embodiment of the present invention is applied.

[0379] See Figure 19a It can be used only Figure 18 The case where the reference pixel layer ref_0 constitutes the reference pixel is confirmed. After constructing the reference pixel using the reference pixel layer ref_0 as the object and utilizing the pixels contained in adjacent blocks (e.g., bottom left, left side, top left, top side, top right), subsequent intra-frame prediction (such as reference pixel generation, reference pixel filtering, reference pixel interpolation, prediction block generation, post-processing filtering, etc.) can be performed, and a portion of the intra-frame prediction process can be adaptively performed according to the reference pixel composition. In this example, the case of using a pre-set reference pixel layer, i.e., not generating setting information related to the reference pixel layer and performing intra-frame prediction using a non-directional mode, will be explained.

[0380] See Figure 19b This allows us to verify the case where two supported reference pixel levels are used simultaneously to construct a reference pixel. That is, it enables in-frame prediction after constructing a reference pixel using pixels contained in level ref_0 and level ref_1 (or using the weighted average of the pixels contained in the two levels). In this example, we will illustrate the case where multiple pre-defined reference pixel levels are used, i.e., no setting information related to the reference pixel levels is generated, and in-frame prediction is performed using a portion of the directional prediction mode (from the upper right to the lower left in the attached diagram, or the opposite direction).

[0381] See Figure 19c This allows us to verify the case where only one of the three supported reference pixel levels is used to construct a reference pixel. In this example, we will explain an instance where, because multiple reference pixel level candidates exist, setting information related to the reference pixel level used is generated, and in-frame prediction is performed using a portion of the directional prediction mode (from the upper left to the lower right in the attached figure).

[0382] Figure 20 This is a third example illustration illustrating the configuration of reference pixels to which one embodiment of the present invention is applied.

[0383] Figure 20 In the attached diagram, 'a' represents blocks with a size of 64×64 or larger, 'b' represents blocks with a size of 16×16 or larger but less than 64×64, and 'c' represents blocks with a size of less than 16×16.

[0384] When the block with reference number a in the attached figure is used as the current block for which intra-frame prediction needs to be performed, intra-frame prediction can be performed using the nearest reference pixel level ref_0.

[0385] Furthermore, when the block in Figure b is used as the current block for which in-frame prediction needs to be performed, in-frame prediction can be performed using the two supported reference pixel levels ref_0 and ref_1.

[0386] Furthermore, when the block in Figure c is used as the current block for which in-frame prediction needs to be performed, in-frame prediction can be performed using the three supported reference pixel levels ref_0, ref_1, and ref_2.

[0387] As described in the descriptions of Figures a through c, the number of supported reference pixel levels can be set differently depending on the size of the current block to be predicted within the image. Figure 20 In this context, the larger the current block size, the higher the probability that the adjacent blocks are smaller. This may be a result of segmentation based on other image characteristics. Therefore, in order to prevent prediction from being performed using pixels with large pixel value distances from the current block, it is assumed that the number of reference pixel levels supported is smaller when the block size is larger, but other variations including the opposite are also allowed.

[0388] Figure 21 This is a fourth example illustration illustrating the configuration of reference pixels to which one embodiment of the present invention is applied.

[0389] See Figure 21 This allows us to confirm if the current block being predicted within the frame is rectangular. If the current block is asymmetrical in both horizontal and vertical dimensions, we can set a larger number of reference pixel layers adjacent to the longer horizontal side boundary of the current block, and a smaller number of reference pixel layers adjacent to the shorter vertical side boundary of the current block. In the attached diagram, we can confirm the case where we set two reference pixel layers adjacent to the horizontal boundary of the current block and one reference pixel layer adjacent to the vertical boundary of the current block. Pixels adjacent to the shorter vertical side boundary of the current block may cause accuracy degradation because the distance between them and the pixels contained in the current block is usually relatively large (due to the larger horizontal length). Therefore, we set a smaller number of reference pixel layers adjacent to the shorter vertical side boundary, but the opposite setting is also possible.

[0390] Furthermore, the reference pixel level to be used in the prediction can be set differently depending on the type of prediction mode within the image or the position of adjacent blocks adjacent to the current block. For example, a orientation mode that uses pixels contained in blocks adjacent to the top and top right edges of the current block as reference pixels can use more than two reference pixel levels, while an orientation mode that uses pixels contained in blocks adjacent to the left and bottom left edges of the current block as reference pixels can use only the nearest reference pixel level.

[0391] Furthermore, when the predicted blocks generated by each reference pixel level are the same or similar to each other in multiple reference pixel levels, the setting information for generating the reference pixel level may lead to the problem of generating unnecessary additional data.

[0392] For example, when the distribution characteristics of the pixels constituting each reference pixel level are similar or identical, similar or identical prediction blocks may be generated regardless of which reference pixel level is used. Therefore, it is not necessary to generate data for the selected reference pixel level. In this case, the distribution characteristics of the pixels constituting the reference pixel level can be determined by comparing the average or dispersion value of the pixels with a pre-set threshold value.

[0393] That is, when the reference pixel levels are the same or similar when the final in-frame prediction mode is used as a reference, the reference pixel level can be selected using a pre-set method (e.g., selecting the nearest reference pixel level).

[0394] At this point, the decoder can receive in-frame prediction information (or in-frame prediction mode information) from the encoding device and determine whether it is necessary to receive information for selecting the reference pixel level based on the received information.

[0395] The above examples illustrate the use of multiple reference pixel levels to construct reference pixels, but this is not the only approach. Various variations can be used, and it can be combined with other additional structures.

[0396] The reference pixel composition unit for in-frame prediction can include a reference pixel generation unit, a reference pixel interpolation unit, a reference pixel filtering unit, etc., and can include all or part of the above-mentioned components. A block containing pixels that can be used as reference pixels can be called a reference candidate block. Furthermore, a reference candidate block can typically be a neighboring block adjacent to the current block.

[0397] The reference pixel composition unit can determine whether to use the pixels contained in the reference candidate block as reference pixels based on the reference pixel availability set for the reference candidate block.

[0398] Regarding the possibility of using the aforementioned reference pixels, it can be determined that they are unusable if at least one of the following conditions is met. For example, if a reference candidate block meets at least one of the following conditions: it is located outside the image boundary, it is not contained in the same segmentation unit as the current block (e.g., stripe, parallel block, etc.), encoding / decoding has not yet been completed, or its use is restricted in the encoding / decoding settings, it can be determined that the pixels contained in the corresponding reference candidate block cannot be referenced. In this case, if none of the above conditions are met, it can be determined that they are usable.

[0399] Furthermore, the use of reference pixels can be restricted based on encoding / decoding settings. For example, when a flag used to restrict references for a reference candidate block (e.g., constrained_intra_pred_flag) is activated, it can be restricted that pixels contained in the corresponding reference candidate block cannot be used as reference pixels. In order to perform encoding / decoding effectively even when errors occur due to various external factors, including the communication environment, the aforementioned flag can be applied when the reference candidate block is a block reconstructed by referencing an image that is temporally different from the current block.

[0400] Specifically, when the flag used to restrict reference is activated (e.g., when constrained_intra_pred_flag = 0 in I-image type or P or B-image type), all pixels in the reference candidate block can be used as reference pixels. Furthermore, when the flag used to restrict reference is activated (e.g., when constrained_intra_pred_flag = 1 in P or B-image type), whether the reference candidate block is referential can be determined based on whether it is encoded via intra-frame prediction or inter-frame prediction. That is, when the reference candidate block is encoded via intra-frame prediction, the corresponding reference candidate block can be referenced regardless of whether the aforementioned flag is activated; while when the reference candidate block is encoded via inter-frame prediction, the referentiality of the corresponding reference candidate block can be determined based on whether the aforementioned flag is activated.

[0401] Furthermore, reconstructed blocks located at positions corresponding to the current block in other color spaces can be used as reference candidate blocks. In this case, their referenceability can be determined based on the encoding mode of the reference candidate block. For example, when the current block belongs to a subset of color difference components (Cb, Cr), its referenceability can be determined based on the encoding mode of a block (=reference candidate block) located at the corresponding position in the luminance component (Y) that has already been encoded / decoded at that position. This could be an example of determining the encoding mode independently based on the color space.

[0402] The flag used to restrict the reference can be a setting applicable to a certain image type (such as P or B strip / parallel block type, etc.).

[0403] Based on the usability of reference pixels, reference candidate blocks can be classified into three cases: fully usable, partially usable, and completely unusable. In cases other than fully usable, reference pixels can be filled or generated at the locations of unusable candidate blocks.

[0404] When a reference candidate block is available, pixels at a pre-defined position in the current block (or pixels adjacent to the current block) can be stored in the reference pixel memory of the current block. In this case, pixel data at the corresponding block position can be directly copied or stored in the reference pixel memory through processes such as reference pixel filtering.

[0405] When a reference candidate block is not available, pixels obtained through the reference pixel generation process can be included in the reference pixel memory of the current block.

[0406] In other words, a reference pixel can be constructed when a reference pixel candidate block is available, and a reference pixel can be generated when a reference pixel candidate block is not available.

[0407] The method for filling reference pixels into pre-defined positions in unusable reference candidate blocks is as follows. First, reference pixels can be generated using arbitrary pixel values. Here, arbitrary pixel value refers to a specific pixel value within a pixel value range, which can be the minimum, maximum, median, or a value derived from the above values ​​used in a pixel value adjustment process based on bit depth or based on the pixel value range information of the image. This method of generating reference pixels using arbitrary pixel values ​​is also applicable when all reference candidate blocks are unusable.

[0408] Next, reference pixels can be generated using pixels contained in blocks adjacent to the unusable reference candidate block. Specifically, pixels contained in adjacent blocks can be used to fill pre-defined positions in the unusable reference candidate block through extrapolation, interpolation, or duplication. The method of performing duplication or extrapolation can be clockwise or counterclockwise, determined according to encoding / decoding settings. For example, the direction of reference pixel generation within a block can follow a pre-defined direction or be adaptively determined based on the position of the unusable block.

[0409] Figures 22a to 22b This is an example diagram illustrating a method of filling reference pixels into pre-defined positions in unusable reference candidate blocks.

[0410] See Figure 22a This allows for the verification of a method for filling unusable reference candidate blocks within a reference pixel hierarchy composed of reference pixels. Figure 22a In the context of the current block, when the adjacent block to the upper right corner is an unusable reference candidate block, the reference pixels contained in the adjacent block to the upper right corner (represented as...) <1> It can be generated by interpolating or linearly interpolating the reference pixels contained in the adjacent blocks that are adjacent to the top of the current block in a clockwise direction.

[0411] In addition, Figure 22a In the context of a current block, when the left-adjacent block to the left is an unusable reference candidate block, the reference pixels contained in the left-adjacent block (represented as...) <2> This can be generated by extrapolating or linearly extrapolating the reference pixels contained in the adjacent block (corresponding to the usable block) adjacent to the upper left of the current block. At this time, by performing extrapolation or linear extrapolation in the clockwise direction, the reference pixels contained in the adjacent block adjacent to the lower left of the current block can be utilized.

[0412] In addition, Figure 22a In the context, a portion of the reference pixels contained in the adjacent block above the current block (represented as...). <3> It can be generated by interpolating or linearly interpolating the usable reference pixels on both sides. That is, it can be set even when some of the reference pixels contained in the adjacent block are unusable, rather than all of them. In this case, the unusable reference pixels can be filled with the adjacent pixels of the unusable reference pixels.

[0413] See Figure 22b This allows for the verification of a method to fill in unusable reference pixels when a portion of the reference pixels in a reference pixel hierarchy composed of multiple reference pixel levels are unusable. See also... Figure 22b When the adjacent block to the upper right of the current block is an unusable reference candidate block, the pixels contained in the three reference pixel levels of the corresponding adjacent block (represented as...) <1> It can generate in a clockwise direction using the pixels contained in the adjacent blocks (corresponding to usable blocks) that are above the current block.

[0414] In addition, Figure 22b In the current block, when the adjacent block to the left of the current block is an unusable reference candidate block and the adjacent block to the upper left or lower left of the current block is a usable reference candidate block, the reference pixels of the unusable reference candidate block can be generated by filling the reference pixels of the usable reference candidate block along the clockwise direction, the counterclockwise direction, or both directions.

[0415] At this point, unusable reference pixels in each reference pixel level can be generated using pixels from the same reference pixel level, but this does not preclude the use of pixels from different reference pixel levels. For example, in Figure 22b In the middle, the reference pixels (represented as) in the three reference pixel levels contained in the adjacent block above the current block are selected. <3> Assuming unusable reference pixels are used as a premise, pixels contained in the nearest and furthest reference pixel levels (ref_0 and ref_2) can be generated using usable reference pixels within the same reference pixel level. Furthermore, pixels contained in reference pixel level ref_1, which is 1 pixel away from the current block, can be generated not only using pixels within the same reference pixel level ref_1, but also using pixels from different reference pixel levels ref_0 and ref_2. In this case, usable reference pixels can be used on both sides to fill the unusable reference pixels using methods such as quadratic linear interpolation.

[0416] The above example demonstrates how reference pixels are generated when multiple reference pixel levels consist of reference pixels and some reference candidate blocks are unavailable. Alternatively, the reference pixel configuration can be set to not allow adaptation (adaptive_intra_ref_sample_flag = 0 in this example) based on encoding / decoding settings (e.g., at least one reference candidate block is unavailable or all reference candidate blocks are unavailable). That is, reference pixels can be constructed according to predefined settings without generating any additional information.

[0417] The reference pixel interpolation unit can generate reference pixels with fractional units by linear interpolation of reference pixels. In this invention, the process is described as part of the reference pixel construction unit, but it can also be incorporated into the prediction block generation unit, and can be understood as a process performed before the prediction block is generated.

[0418] Furthermore, while it is assumed to be an independent process separate from the reference pixel filtering section described later, it is also possible to adopt a process that integrates them into one. This configuration is also provided to address the issue of reference pixel distortion caused by the increase in the number of filters applied to the reference pixel when multiple filters are applied through the reference pixel interpolation unit and the reference pixel filtering unit.

[0419] The reference pixel interpolation process is not executed in some prediction modes (such as horizontal, vertical, some diagonal modes <such as diagonal down right, diagonal down left, diagonal upright, etc. at a 45-degree angle>, non-directional modes, color modes, color replication modes, etc., i.e. modes that do not require decimal units of interpolation when generating prediction blocks), but can only be executed in other prediction modes (modes that require decimal units of interpolation when generating prediction blocks).

[0420] Interpolation accuracy (e.g., 1, 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64 pixel units) can be determined based on the prediction mode (or the directionality of the prediction mode). For example, a prediction mode at a 45-degree angle does not require interpolation, but prediction modes at 22.5 degrees or 67.5 degrees require a difference of 1 / 2 pixel unit. As described above, at least one interpolation accuracy and a maximum interpolation accuracy can be determined based on the prediction mode.

[0421] For reference pixel interpolation, a single pre-defined interpolation filter (e.g., a 2-tap linear interpolation filter) can be used, or a filter selected from multiple interpolation filter candidate groups (e.g., a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc.) can be used, selected according to the encoder / decoder settings. In this case, the interpolation filter can be distinguished based on differences in the number of filter taps (i.e., the number of pixels applicable to the filtering) and filter coefficients.

[0422] Interpolation can be performed in stages, progressing from lower precision to higher precision (e.g., 1 / 2 → 1 / 4 → 1 / 8), or it can be performed all at once. The former refers to interpolation based on integer units of pixels and fractional units of pixels (pixels that have already been interpolated using a lower precision of the pixel to be interpolated), while the latter refers to interpolation based on integer units of pixels.

[0423] When using one of multiple filter candidate groups, the filter selection information can be explicitly generated or determined by the mode, or it can be determined based on encoding / decoding settings (such as interpolation accuracy, block size, shape, prediction mode, etc.). In this case, the unit of explicit generation can be video, sequence, image, strip, parallel block, block, etc.

[0424] For example, an 8-tap Kalman filter can be applied to reference pixels of integer units when using interpolation accuracy greater than 1 / 4 (1 / 2, 1 / 4), a 4-tap Gaussian filter can be applied to reference pixels of integer units and interpolated reference pixels of greater than 1 / 4 units when using interpolation accuracy less than 1 / 4 and greater than 1 / 16 (1 / 8, 1 / 16), and a 2-tap linear filter can be applied to reference pixels of integer units and interpolated reference pixels of greater than 1 / 16 units when using interpolation accuracy less than 1 / 16 (1 / 32, 1 / 64).

[0425] Alternatively, an 8-tap Kalman filter can be applied to blocks larger than 64×64, a 6-tap Wiener filter can be applied to blocks smaller than 64×64 and larger than 16×16, and a 4-tap Gaussian filter can be applied to blocks smaller than 16×16.

[0426] Alternatively, a 4-tap cubic filter can be applied to prediction modes with an angular difference of less than 22.5 degrees based on vertical or horizontal patterns, while a 4-tap Gaussian filter can be applied to prediction modes with an angular difference of more than 22.5 degrees.

[0427] Furthermore, multiple filter candidate groups can be composed of 4-tap cubic filters, 6-tap Wiener filters, and 8-tap Kalman filters in some encoding / decoding settings, and 2-tap linear filters and 6-tap Wiener filters in others.

[0428] Figures 23a to 23c This is an illustrative diagram illustrating a method of performing interpolation on a fractional pixel basis in reference pixels constructed according to an embodiment of the present invention.

[0429] See Figure 23a This allows us to confirm a method for interpolating pixels in fractional units when a reference pixel level (ref_i) is used as the reference pixel. Specifically, interpolation can be performed by applying a filter (labeled int_func_1D) to pixels adjacent to the interpolation target pixel (marked with x) (assuming, in this example, a filter is applied to integer pixel units). Since a reference pixel level is used as the reference pixel, interpolation can be performed using adjacent pixels contained in the same reference pixel level as the interpolation target pixel x.

[0430] See Figure 23b This allows for the verification of methods to obtain interpolated pixels in fractional units when using more than two reference pixel levels (ref_i, ref_j, ref_k) as reference pixels. Figure 23b In this context, when performing reference pixel interpolation at the reference pixel level ref_j, it is possible to append interpolation of the target pixels using other reference pixel levels ref_k and ref_i with fractional units. Specifically, by interpolating adjacent pixels a... k ~h k a j ~h j a i ~h i Perform filtering (interpolation process, function int_func_1D) to obtain the interpolated pixel (x) respectively. j The position of the pixel x) and the position of the pixel x contained in other reference pixel levels that corresponds to the interpolated pixel (the corresponding position in each reference pixel level according to the direction of the prediction mode). k x i And the obtained first-order interpolated pixel x k x j x i Append filtering is performed (this can be a non-interpolation process such as weighted averaging of [1, 2, 1] / 4, [1, 6, 1] / 8, etc.) to finally obtain the final interpolated pixel x at the reference pixel level ref_j. In this example, it is assumed that the pixel x at other reference pixel levels corresponding to the interpolated pixel is... k x i The case where fractional pixels can be obtained through interpolation is explained.

[0431] In the above examples, the case where first-order interpolated pixels can be obtained by filtering at each reference pixel level, and the final difference pixel can be obtained by performing additional filtering on the first-order interpolated pixels, was explained. However, it is also possible to obtain the final difference pixel by filtering adjacent pixels a at multiple reference pixel levels. k ~h k a j ~h j a i ~h i The final interpolated pixels are obtained in one step by filtering.

[0432] exist Figure 23b Of the three reference pixel levels supported in the code, the level actually used as a reference pixel can be ref_j. That is, in order to interpolate a reference pixel level composed of reference pixels, other reference pixel levels contained in the candidate group (for example, if it is not composed of reference pixels, it means that the pixels in the corresponding reference pixel level are not suitable for prediction, but in the above case, the corresponding pixels will be referenced when performing interpolation, so it can also be considered as a case of use) can be used.

[0433] See Figure 23c The diagram illustrates the case where both supported reference pixel levels are used as reference pixels. This allows for the use of pixels adjacent to the fractional unit positions (d in this example) of the pixels to be interpolated in each of the supported reference pixel levels. i d j e i e j The input pixels are constructed, and filtering is performed on adjacent pixels to obtain the final interpolated pixel x. At this time, it is also possible to use methods such as... Figure 23b The method shown is to obtain the first interpolated pixel at each reference pixel level, and then perform additional filtering on the first interpolated pixel to obtain the final interpolated pixel x.

[0434] The above examples are not limited to the reference pixel interpolation process, but can also be understood as a process combined with other in-image prediction processes (such as the reference pixel filtering process, the prediction block generation process, etc.).

[0435] Figures 24a to 24b This is a first example diagram used to illustrate an adaptive reference pixel filtering method applicable to one embodiment of the present invention.

[0436] Typically, the main purpose of the reference pixel filtering section is to perform smoothing by using a low-pass filter (e.g., a 3-tap or 5-tap filter such as [1,2,1] / 4, [2,3,6,3,2] / 16, etc.). However, other types of filters (e.g., high-pass filters) can also be used depending on the intended purpose of the filter (e.g., sharpening). In this invention, the focus will be on reducing distortion generated during encoding / decoding by performing filtering for the purpose of smoothing.

[0437] Reference pixel filtering can be performed based on encoding / decoding settings. However, batch filtering might fail to capture the local characteristics of the image, so applying filtering based on these local characteristics is more beneficial for improving encoding performance. These image characteristics can be determined not only by image type, color components, quantization parameters, and the encoding / decoding information of the current block (e.g., block size, shape, segmentation information, prediction mode), but also by the encoding / decoding information of adjacent blocks and combinations thereof. Furthermore, it can be determined based on the distribution characteristics of reference pixels (e.g., the dispersion, standard deviation, flatness, and discontinuity of the reference pixel region).

[0438] See Figure 24a When a block belongs to a category (Category 0) based on a portion of encoding / decoding settings (e.g., block size range A, prediction mode B, color composition C, etc.), filtering is not applicable. However, when a block belongs to a category (Category 1) based on a portion of encoding / decoding settings (e.g., prediction mode A of the current block, prediction mode B of pre-defined adjacent blocks, etc.), filtering is applicable.

[0439] See Figure 24b When a block belongs to a category based on a portion of encoding / decoding settings (e.g., the size of the current block A, the size of the adjacent block B, the prediction mode of the current block C, etc.) (Category 0), filtering is not required. When a block belongs to a category based on a portion of encoding / decoding settings (e.g., the size of the current block A, the shape of the current block B, the size of the adjacent block C, etc.) (Category 1), filtering can be performed using filter A. When a block belongs to a category based on a portion of encoding / decoding settings (e.g., the parent block A of the current block, the parent block B of the adjacent block, etc.) (Category 2), filtering can be performed using filter B.

[0440] Therefore, the applicability of filtering, the type of filter, the encoding of filter information (explicit / implicit), and the number of filtering iterations can be determined based on the size of the current block and adjacent blocks, the prediction mode, and color components. The filter type can be classified based on differences in the number of taps and filter coefficients. In this case, when the number of filtering iterations is two or more, the same filter can be applied multiple times, or different filters can be applied separately.

[0441] The above examples illustrate cases where the reference pixel filtering is pre-set based on the characteristics of the image. That is, the filter-related information can be implicitly determined. However, if the determination of image characteristics as described above is inaccurate, it may negatively impact coding efficiency; therefore, this aspect must be considered.

[0442] To prevent the situation described above, the filtering of the reference pixels can be explicitly set. For example, information related to whether filtering is applicable can be generated. In this case, filter selection information can be not generated when there is only one filter, but filter selection information can be generated when multiple filter candidate groups exist.

[0443] The above examples illustrate the implicit and explicit settings related to reference pixel filtering, demonstrating a hybrid approach where explicit settings are used in some cases and implicit settings in others. Implicit settings refer to information related to the reference pixel filter (such as whether filtering is applicable or not, and filter type information) that can be derived from the decoder.

[0444] Figure 25 This is a second illustration used to explain an adaptive reference pixel filtering method applicable to one embodiment of the present invention.

[0445] See Figure 25 It can classify categories by utilizing image characteristics confirmed by encoding / decoding information and adaptively perform reference pixel filtering according to the classified category.

[0446] For example, a filter is applied when the image is classified as category 0, while filter A is used when the image is classified as category 1. Category 0 and category 1 can be instances of the implicit reference pixel filter.

[0447] Furthermore, when classified as Category 2, it is possible to either not apply filtering or apply filter A. In this case, the generated information may be related to whether or not filtering is applicable, but filter selection information will not be generated.

[0448] Furthermore, when classified as Category 3, either Filter A or Filter B can be applied. In this case, the generated information can be filter selection information, and the application of the filter can be an instance of unconditional execution. That is, when classified as Category 3, it can be understood as a case where filtering must be performed but the filter type needs to be selected.

[0449] Furthermore, when classified as Category 4, it is possible to not apply filtering, apply filter A, or apply filter B. In this case, the generated information may be related to whether filtering is applicable and filter selection information.

[0450] In other words, it is possible to determine explicit or implicit processing based on the category, and when explicit processing is performed, it is possible to adaptively construct the candidate group settings related to each reference pixel filter.

[0451] Regarding the above categories, the following examples can be considered.

[0452] First, for blocks larger than 64×64, one of the following can be implicitly determined based on the prediction mode of the current block: <Filter Off>, <Filter On - Filter A>, <Filter On + Filter B>, and <Filter On + Filter C>. In this case, considering the characteristics of the reference pixel distribution, the additional candidate can be <Filter On + Filter C>. That is, when the filter is on, filter A, filter B, or filter C can be applied.

[0453] In addition, for blocks smaller than 64×64 and larger than 16×16, one of <Filter Off>, <Filter On + Filter A>, and <Filter On + Filter B> can be implicitly determined based on the prediction mode of the current block.

[0454] Furthermore, for blocks smaller than 16×16, the system can select one of <Filter Off>, <Filter On + Filter A>, or <Filter On + Filter B> based on the current block's prediction mode. In this case, the mode can be set to <Filter Off> in some prediction modes, while <Filter Off> or <Filter On + Filter A> can be explicitly selected in others, and <Filter Off> or <Filter On + Filter B> can be explicitly selected in still others.

[0455] As an example of settings related to multiple reference pixel filters, when the reference pixels obtained in various filters (including the case where filtering is off in this example) are the same or similar, generating reference pixel filter information (e.g., reference pixel filter tolerance information, reference pixel filter information, etc.) may lead to the generation of unnecessary duplicate information. For example, when the distribution characteristics of the reference pixels obtained in various filters (e.g., values ​​obtained by averaging, dispersing, etc., of the various reference pixels and the threshold) are similar, the generation of reference pixel filter information may be affected. <threshold>When the characteristics determined by comparison are the same or similar, the reference pixel filter information can be omitted. When the reference pixel filter information is omitted, filtering can be applied in a pre-set method (e.g., filtering off). After receiving in-frame prediction information, the decoder can determine whether to receive the reference pixel filter information in the same way as the encoder, and can determine whether to accept the reference pixel filter information based on the above determination.

[0456] Assuming the generation of explicit information related to reference pixel filtering, it is possible to generate adaptive reference pixel filtering indication information (adaptive_ref_filter_enabled_flag in this example) on units such as video, sequence, image, strip, and parallel block.

[0457] When the above indication information represents an allowable adaptive reference pixel filtering (adaptive_ref_filter_enabled_flag = 1 in this example), adaptive reference pixel filtering allowance information (adaptive_ref_filter_flag in this example) can be generated on units such as images, strips, parallel blocks, and blocks.

[0458] When the above-mentioned permissive information represents adaptive reference pixel filtering (adaptive_ref_filter_flag=1 in this example), reference pixel filtering related information (such as reference pixel filter selection information, ref_filter_idx in this example) can be generated on units such as images, strips, parallel blocks, and blocks.

[0459] At this time, when adaptive reference pixel filtering is not allowed or cannot be applied, the filtering action can be performed on the reference pixel according to the pre-defined settings (as mentioned above, the applicability of filtering and the type of filtering are determined in advance based on image encoding / decoding information, etc.).

[0460] Figures 26a to 26b This is an illustrative diagram illustrating the use of a reference pixel level in reference pixel filtering according to one embodiment of the present invention.

[0461] See Figure 26a It can be confirmed that interpolation is performed by applying a filter (called the smt_func_1 function) to the object pixel d in the pixels contained in the reference pixel level ref_i and the pixels a, b, c, e, f, g adjacent to the object pixel d.

[0462] Figure 26a Typically, sequential filtering is applicable, but multiple filtering can also be applied. For example, two filtering steps can be applied to reference pixels (a*, b*, c*, d*, etc. in this example) obtained by applying one filtering step.

[0463] See Figure 26b This function can obtain a filtered pixel e* by performing linear interpolation (called the smt_func_2 function) on the pixels located on both sides of the object pixel e, with the object pixel e as the center, in proportion to the distance (e.g., the distance z between it and a). The pixels located on both sides are those located at the ends of adjacent pixels within a block consisting of the upper block, left block, upper block + upper right block, left block + lower left block, upper left block + upper block + upper right block, upper left block + left block + lower left block, and upper left block + left block + upper side block + lower left block + upper right block of the current block. Figure 26b It can be reference pixel filtering performed based on the characteristics of the reference pixel distribution.

[0464] exist Figures 26a to 26b The diagram illustrates the case where pixels from the same reference pixel level as the target reference pixel are used for reference pixel filtering. In this case, the type of filter used in reference pixel filtering can vary depending on whether the reference pixel level is the same or different.

[0465] Furthermore, when using multiple reference pixel levels, when performing reference pixel filtering in a subset of reference pixel levels, it is possible to use pixels not only from the same reference pixel level but also pixels from different reference pixel levels.

[0466] Figure 27 This is an illustrative diagram illustrating a case in which multiple reference pixel levels are used in reference pixel filtering, according to one embodiment of the present invention.

[0467] See Figure 27 Firstly, filtering can be performed on the reference pixel levels ref_k and ref_i using pixels contained in the same reference pixel level. That is, filtering can be performed on the object pixel d at the reference pixel level ref_k. k and adjacent pixel a k to g k Perform filtering (defined as function smt_func_1D) to obtain the filtered pixel d. k * It is also possible to modify the object pixel d at the reference pixel level ref_i. i and adjacent pixel a i to g i Perform filtering (defined as function smt_func_1D) to obtain the filtered pixel d. i *

[0468] Furthermore, when performing reference pixel filtering on the reference pixel level ref_j, it is possible not only to use the same reference pixel level ref_j, but also to use pixels contained in other spatially adjacent reference pixel levels, namely ref_i and ref_k. Specifically, it is possible to use the object pixel d j Centered on the spatially adjacent pixels c k d k e k c j e j c i d i e i (That is, a filter that can be a 3×3 square mask) can be applied to the filter (defined as the function smt_func_2D) to obtain the interpolated pixel d. j * However, it is not limited to a 3×3 square shape; it can also use a shape such as a 5×2 rectangle centered on the object pixel (b k c k d k e k f k b j c j e j f j ), 3×3 rhombus shape (d k c j e j d i ), 5×3 cross shape (d k b j c j e j f j d i ) and other masking filters.

[0469] The reference pixel level is as described above. Figures 18 to 2 As shown in Figure 2, it consists of pixels contained in adjacent blocks that are close to the boundary of the current block. Considering the aspects described above, filtering in reference pixel levels ref_k and ref_i using pixels contained in the same reference pixel level allows for the application of filters utilizing a 1D mask shape of pixels horizontally or vertically adjacent to the interpolation target pixel. However, in reference pixel level ref_j, the reference pixel d... j Interpolated pixels can be obtained by applying a filter that utilizes a 2D mask shape of all spatially adjacent pixels (top / bottom / left / right).

[0470] Furthermore, within each reference pixel level, reference pixel filtering can be applied twice to reference pixels that have already undergone one round of filtering. For example, one round of filtering can be performed using the reference pixels contained in each reference pixel level (ref_k, ref_j, ref_i). Subsequently, in the reference pixel levels that have already undergone one round of filtering (referred to as ref_k*, ref_j*, ref_i*), not only can the respective reference pixel levels be used, but reference pixels from other reference pixel levels can also be used to perform reference pixel filtering.

[0471] The prediction block generation unit can generate prediction blocks based on at least one in-frame prediction mode (which can be simply referred to as a prediction mode), and can use reference pixels based on the aforementioned prediction mode. At this time, prediction blocks can be generated by performing extrapolation, interpolation, or DC copying on the reference pixels according to the prediction mode. Extrapolation can be applied to directional modes within the in-frame prediction modes, while the remaining modes can be applied to non-directional modes.

[0472] Furthermore, when copying a reference pixel, it is possible to generate more than one predicted pixel by copying a reference pixel to multiple pixels within the prediction block, or by copying more than one reference pixel. The number of reference pixels copied can be equal to or less than the number of predicted pixels copied.

[0473] Furthermore, a prediction block is typically generated for a single prediction mode within an image. However, it is also possible to generate a final prediction block by applying a weighted summation method to multiple prediction blocks after obtaining them. Here, multiple prediction blocks can refer to prediction blocks obtained based on a reference pixel level.

[0474] In the prediction mode determination unit of the encoding apparatus, a process for selecting the optimal mode from multiple prediction mode candidate groups is performed. Typically, the optimal mode in terms of encoding cost can be determined using rate-distortion techniques, which predict the block distortion (e.g., distortion between the current block and the reconstructed block, sum of absolute errors (SAD), sum of square differences (SSD), etc.)) and the number of bits generated under the prediction mode. The prediction block generated based on the prediction mode determined by the above process can be transmitted to the subtraction and addition units (at this time, since the decoding apparatus can obtain information indicating the optimal prediction mode from the encoding apparatus, the process of selecting the optimal prediction mode can be omitted).

[0475] The prediction mode encoding unit of the encoding device can encode the optimal in-frame prediction mode selected by the prediction mode determination unit. At this time, it can directly encode the index information indicating the optimal prediction mode, or it can encode prediction information related to the prediction mode (e.g., the difference between the predicted prediction mode index and the prediction mode index of the current block) after predicting the optimal prediction mode using prediction modes obtainable from surrounding blocks. The former case applies to chromatic aberration components, while the latter applies to luminance components.

[0476] When predicting and encoding the best prediction pattern for the current block, the predicted value (or prediction information) of the prediction pattern can be called the Most Probable Pattern (MPM). Here, the Most Probable Pattern (MPM) refers to the prediction pattern with the highest probability of becoming the best prediction pattern for the current block. It can be composed of pre-defined prediction patterns (such as mean (DC), planar, vertical, horizontal, diagonal patterns, etc.) or prediction patterns of spatially adjacent blocks (such as left, top, upper left, upper right, lower left blocks, etc.). Among them, the diagonal pattern refers to diagonally up to the right, diagonally down to the right, or diagonally down to the left, and can be... Figure 17 The patterns corresponding to patterns No. 2, No. 18, and No. 34 in the data.

[0477] Furthermore, patterns derived from the prediction patterns contained in the most probable patterns (MPMs), i.e., the most probable pattern (MPM) candidate group, can be added to the most probable pattern (MPM) candidate group. In directional patterns, prediction patterns with an index interval equal to a preset value between them and the prediction patterns contained in the most probable pattern (MPM) candidate group can be added to the most probable pattern (MPM) candidate group. For example, a pattern included in the most probable pattern (MPM) candidate group is... Figure 17 In the case of pattern 10, the derived pattern can be equivalent to patterns 9, 11, 8, 12, etc.

[0478] The above example is equivalent to the case where the most likely mode (MPM) candidate group consists of multiple modes. The composition of the most likely mode (MPM) candidate group (e.g., the number of predicted modes included in the most likely mode (MPM) and the priority order of its composition) is determined according to the encoding / decoding settings (e.g., predicted mode candidate group, image type, block size, block shape, etc.) and can include at least one mode composition.

[0479] It is possible to set the priority order of prediction patterns included in the Most Probable Pattern (MPM) candidate group. It is possible to determine the order of prediction patterns included in the MPM candidate group according to the set priority order, and to complete the formation of the MPM candidate group when the number of added prediction patterns reaches a preset number. The priority order can be set as prediction patterns of blocks spatially adjacent to the current block to be predicted, preset prediction patterns, and patterns derived from earlier prediction patterns included in the MPM candidate group, but is not limited to this.

[0480] Specifically, in spatially adjacent blocks, priority can be set in the order of left-top-bottom-top-right-top-left blocks. In pre-defined prediction patterns, priority can be set in the order of mean (DC)-planar-vertical-horizontal pattern. Next, the index values ​​of prediction patterns included in the most likely pattern (MPM) candidate group can be... Figure 17 The predicted patterns obtained by performing addition operations (+1, -1, etc., integer values) on the predicted pattern number are included in the most probable pattern (MPM) candidate group. As one example above, the priority order can be set in the following order: left-top-mean (DC)-planar-bottom-top-right-top-left-(spatial adjacent block pattern)+1-(spatial adjacent block pattern)-1-horizontal-vertical-diagonal.

[0481] In the above example, the priority order of the most likely pattern (MPM) candidate groups was fixed, but the above priority order can also be adaptively determined according to the shape, size, etc. of the blocks.

[0482] When encoding the prediction mode of the current block using the most probable mode (MPM), information related to whether the prediction mode matches the most probable mode (MPM) can be generated (e.g., most_probable_mode_flag).

[0483] When the most probable mode (MPM) matches (e.g., most_probable_mode_flag = 1), the most probable mode (MPM) index information (e.g., mpm_idx) can be appended based on the composition of the most probable mode (MPM). For example, when the most probable mode (MPM) consists of a single prediction mode, no additional most probable mode (MPM) index information can be generated, while when it consists of multiple prediction modes, index information corresponding to the prediction mode of the current block in the most probable mode (MPM) candidate group can be generated.

[0484] When the most probable mode (MPM) is inconsistent (e.g., most_probable_mode_flag=0), it is possible to generate non-most probable mode (non-MPM) index information (e.g., non_mpm_idx) in the remaining prediction mode candidate group (called the non-most probable mode (non-MPM) candidate group) after excluding the most probable mode (MPM) candidate group from the supported in-frame prediction modes, which can be an instance of forming a group of non-most probable modes (non-MPM).

[0485] When the non-most likely pattern (non-MPM) candidate group consists of multiple groups, information related to which group the current block's predicted pattern belongs to can be generated. For example, when the non-most likely pattern (non-MPM) consists of two groups, A and B, and the current block's predicted pattern matches the predicted pattern of group A (e.g., non_mpm_A_flag = 1), index information corresponding to the current block's predicted pattern can be generated in the candidate groups of group A. When they do not match (e.g., non_mpm_A_flag = 0), index information corresponding to the current block's predicted pattern can be generated in the remaining prediction pattern candidate groups (or the candidate groups of group B). As shown in the above example, the non-most likely pattern (non-MPM) can consist of multiple groups, and the number of groups can be specified according to the prediction pattern candidate groups. For example, it can be 1 when there are fewer than 35 prediction pattern candidate groups, and 2 in other cases.

[0486] At this point, a specific group A can be composed of patterns that are determined to have a high probability of being consistent with the predicted pattern of the current block after the most probable pattern (MPM) candidate group. For example, subsequent predicted patterns that are not included in the most probable pattern (MPM) candidate group can be included in group A, or directional patterns with a certain interval can be included in group A.

[0487] As in the example above, when the non-most likely pattern (non-MPM) consists of multiple groups, it can reduce the number of pattern coding bits even when the number of predicted patterns is large and the predicted pattern of the current block is inconsistent with the most likely pattern (MPM).

[0488] When encoding (or decoding) the prediction mode of the current block using the most probable mode (MPM), it is possible to generate binary tables individually for each prediction mode candidate group (e.g., most probable mode (MPM) candidate group, non-most probable mode (non-MPM) candidate group, etc.), and to apply different binary methods individually for each candidate group.

[0489] In the above examples, terms such as Most Probable Pattern (MPM) candidate group and Non-Most Probable Pattern (non-MPM) candidate group are only a portion of the terms used in this invention and are not intended to limit the scope of the invention. Specifically, these terms are used to indicate which category a predicted pattern in the current frame belongs to when classifying it into multiple categories, as well as the pattern information within that category. Terms such as First Most Probable Pattern (MPM) candidate group and Second Most Probable Pattern (MPM) candidate group can also be used instead.

[0490] Figure 28 This is a block diagram used to illustrate an encoding / decoding method for an in-screen prediction mode applicable to one embodiment of the present invention.

[0491] See Figure 28 First, obtain the `mpm_flag` (S10). Next, confirm whether it matches the first most likely pattern (MPM) (indicated by `mpm_flag`) (S11). If they match, confirm the index information (`mpm_idx`) of the most likely pattern (MPM) (S12). If they do not match the most likely pattern (MPM), obtain the `rem_mpm_flag` (S13). Next, confirm whether it matches the second most likely pattern (MPM) (indicated by `rem_mpm_flag`) (S14). If they match, confirm the index information (`rem_mpm_idx`) of the second most likely pattern (MPM) (S16). If they do not match the second most likely pattern (MPM), confirm the index information (`rem_mode_idx`) of the candidate group composed of the remaining predicted patterns (S15). This example illustrates the case where the index information generated based on the consistency of the two most likely patterns (MPMs) is represented using the same syntax elements. However, other pattern encoding settings (such as binary encoding) can also be applied, and different index information can be set for processing.

[0492] In an image decoding method according to one embodiment of the present invention, intra-frame prediction can be configured as follows: The intra-frame prediction of the prediction unit can include a prediction mode decoding step, a reference pixel construction step, and a prediction block generation step. Furthermore, the image decoding apparatus can include a prediction mode decoding unit, a reference pixel construction unit, and a prediction block generation unit for performing the prediction mode decoding step, the reference pixel construction step, and the prediction block generation step. The process described above can omit some steps or add other steps, and can also be changed to a different order than that described above.

[0493] Since the reference pixel composition unit and the prediction block generation unit of the image decoding device can perform the same function as the composition unit in the image encoding device, detailed descriptions related to them will be omitted here. The prediction mode decoding unit can be used in reverse in the manner used in the prediction mode encoding unit.

[0494] Next, we will combine Figures 29 to 31 Various embodiments of in-frame prediction based on reference pixels of the decoding device will be described. The descriptions of reference pixel level support and reference pixel filtering methods given above in conjunction with the accompanying drawings should be interpreted as equally applicable to the decoding device; however, detailed descriptions related to these aspects will be omitted to avoid redundancy.

[0495] Figure 29 This is the first example diagram used to illustrate the bitstream composition predicted within a frame based on reference pixels.

[0496] exist Figure 29 The first example in the figure assumes the following conditions: supporting multiple reference pixel levels, using at least one of the supported reference pixel levels as a reference pixel, supporting multiple candidate groups related to reference pixel filtering, and selecting a filter from them.

[0497] After constructing a pixel candidate group using multiple reference pixel levels in the encoder (in this example, the reference pixel generation process has been completed), a reference pixel is constructed using at least one reference pixel level. Next, reference pixel filtering and reference pixel interpolation are applied. At this point, multiple candidate groups related to reference pixel filtering are supported.

[0498] Next, a process is executed to select the best prediction mode from the candidate prediction mode group. After the best prediction mode is determined, a prediction block based on the corresponding mode is generated and passed to the subtraction unit. Then, the encoding process for prediction-related information within the image is performed. In this example, the reference pixel level and reference pixel filtering are implicitly determined based on the encoded information.

[0499] In the decoder, prediction-related information (such as prediction patterns) within the image is reconstructed, and after generating prediction blocks based on the reconstructed prediction patterns, it is passed to the subtraction unit. At this time, the reference pixel level and reference pixel filtering used to generate the prediction blocks are implicitly determined.

[0500] See Figure 29 A bitstream can be constructed using an intra-mode prediction method (S20). At this time, the reference pixel level ref_idx and reference pixel filter category ref_filter_idx supported (or used) in the current block can be implicitly determined based on the intra-mode prediction method (determined as Category A and Category B, S21-S22, respectively). Encoding / decoding information (such as image type, color components, block size, and shape) can then be additionally considered.

[0501] Figure 30 This is the second illustration used to explain the bitstream composition predicted within a frame based on reference pixels.

[0502] exist Figure 30 The second example diagram assumes the case where multiple reference pixel levels are supported, and one of these multiple supported reference pixel levels is used as the reference pixel. Furthermore, it assumes the case where multiple candidate groups related to reference pixel filtering are supported, and a filter is selected from these candidate groups. Figure 29 The difference lies in the fact that the information related to the selection is explicitly generated by the encoding device.

[0503] After determining that multiple reference pixel levels are supported in the encoder, a reference pixel is constructed using one reference pixel level. Next, reference pixel filtering and reference pixel interpolation are applied. At this point, multiple filtering methods related to reference pixel filtering are supported.

[0504] When the encoder performs the process of determining the best prediction mode for the current block, it can also consider the process of selecting the best reference pixel level among the various prediction modes and the process of selecting the best reference pixel filter. After determining the best prediction mode, reference pixel level, and reference pixel filter for the current block, the prediction block generated based on this is passed to the subtraction unit to perform the encoding process of prediction-related information within the image.

[0505] In the decoder, prediction-related information within the image (such as prediction mode, reference pixel level, reference pixel filtering information, etc.) is reconstructed, and after generating prediction blocks using the reconstructed information, it is passed to the subtraction unit. At this time, the reference pixel level and reference pixel filtering used for generating the prediction blocks follow settings determined based on information transmitted from the encoder.

[0506] See Figure 30 The decoder confirms the best prediction mode for the current block using the intra-mode prediction information (intra_mode) contained in the bitstream (S30), and confirms whether multiple reference pixel levels are supported (multi_ref_flag) (S31). When multiple reference pixel levels are supported, the reference pixel level selection information (ref_idx) is confirmed (S32), thereby determining the reference pixel levels that can be used for intra-frame prediction. When multiple reference pixel levels are not supported, the process of obtaining the reference pixel level selection information (ref_idx) (S32) can be omitted.

[0507] Next, the support for adaptive reference pixel filtering (adap_ref_smooth_flag) is confirmed (S33). If adaptive reference pixel filtering is supported, the filtering method for the reference pixel is determined by the reference pixel filter information (ref_filter_idx) (S34).

[0508] Figure 31 This is the third illustration used to explain the bitstream composition predicted within a frame based on reference pixels.

[0509] exist Figure 31 In the third example diagram, the case of supporting multiple reference pixel levels and using one of the multiple reference pixel levels is assumed. Furthermore, the case of supporting multiple candidate groups related to reference pixel filtering and selecting one filter from them is assumed. Figure 30 The difference lies in the fact that it adaptively generates selection information.

[0510] After constructing a reference pixel using one of the multiple reference pixel levels supported by the encoder, reference pixel filtering and reference pixel interpolation are applied. At this point, various filters related to reference pixel filtering are supported.

[0511] When performing the process of selecting the best mode from multiple prediction mode candidate groups, it is also possible to additionally consider the process of selecting the best reference pixel level in each prediction mode and the process of selecting the best reference pixel filter. After determining the best prediction mode, reference pixel level, and reference pixel filter, a prediction block is generated based on this and passed to the subtraction unit, and then the encoding process of prediction-related information within the image is performed.

[0512] At this point, the repeatability of the generated prediction blocks is checked. If the predicted blocks are the same as or similar to those obtained using other reference pixel levels, the selected information related to the best reference pixel level is omitted, and a pre-set reference pixel level is used. In this case, the pre-set reference pixel level can be the level most adjacent to the current block.

[0513] For example, it is possible to pass Figure 19c The reproducibility of a prediction block is determined based on the difference (distortion value) between the prediction block generated by ref_0 and the prediction block generated by ref_1. If the difference value is less than a pre-set threshold, the prediction block is considered reproducible; otherwise, it is considered non-reproducible. This threshold can be adaptively determined based on quantization parameters, etc.

[0514] In addition, the best reference pixel filtering information also confirms the repeatability of the predicted block. When the predicted block is the same as or similar to the predicted block obtained by applying other reference pixel filters, the reference pixel filtering information is omitted and a pre-set reference pixel filter is applied.

[0515] For example, the repetition of a prediction block is determined based on the difference between the prediction block obtained through filter A (a 3-tap filter in this example) and the prediction block obtained through filter B9 (a 5-tap filter in this example). Similarly, the difference value can be compared with a pre-set threshold, and if it is smaller, the prediction block is considered repetitive. When a prediction block is repetitive, a prediction block can be generated using a pre-set reference pixel filtering method. This pre-set reference pixel filtering can be a filtering method with fewer taps or lower complexity, including cases where filtering is omitted.

[0516] In the decoder, prediction-related information within the image (such as prediction mode, reference pixel level, reference pixel filtering information, etc.) is reconstructed and used to generate a prediction block, which is then passed to the subtraction unit. At this time, the reference pixel level information and reference pixel filtering used to generate the prediction block follow settings determined based on information transmitted from the encoder. The decoder can directly confirm whether there is repetition (without using syntax elements) and, if there is repetition, follow a pre-defined method.

[0517] See Figure 31 The decoder first confirms the intra-mode prediction information (intra_mode) of the current block (S40) and confirms the support for multiple reference pixel levels (multi_ref_flag) (S41). When multiple reference pixel levels are supported, a repeatability check is performed on the prediction blocks based on the supported multiple reference pixel levels (represented by the ref_check process, S42). When the repeatability confirmation result is that the prediction blocks have no repeatability (redund_ref = 0, S43), the selection information of the reference pixel level (ref_idx) is referenced from the bitstream (S44) and the optimal reference pixel level is determined.

[0518] Next, the support for adaptive reference pixel filtering (adap_ref_smooth_flag) is confirmed (S45). When adaptive reference pixel filtering is supported, the repetition of prediction blocks for multiple supported reference pixel filtering methods is checked (represented by the ref_check procedure, S46). When there is no repetition of prediction blocks (redund_ref = 0, S47), the selection information of the reference pixel filtering method (ref_filter_idx) from the bitstream is referenced (S48), and the optimal reference pixel filtering method is determined.

[0519] In this case, redund_ref in the attached diagram is a value used to indicate the repeatability verification result; a value of 0 indicates no repeatability.

[0520] Furthermore, the decoder can perform in-frame prediction using a pre-defined reference pixel level and a pre-defined reference pixel filtering method when the predicted blocks are repetitive.

[0521] Figure 32 This is a flowchart illustrating an image decoding method supporting multiple reference pixel levels applicable to one embodiment of the present invention.

[0522] See Figure 32 The image decoding method supporting multiple reference pixel levels includes: a step of confirming whether multiple reference pixel levels are supported by a bitstream (S100); when multiple reference pixel levels are supported, a step of determining the reference pixel level to be used in the current block by referring to the syntax information contained in the bitstream (S110); a step of constructing a reference pixel using the pixels contained in the determined reference pixel level (S120); and a step of performing in-frame prediction of the current block using the constructed reference pixel (S130).

[0523] The step of confirming whether multiple reference pixel levels are supported (S100) may further include a step of confirming whether an adaptive reference pixel filtering method is supported by a bit stream.

[0524] The step of confirming whether multiple reference pixel levels are supported (S100) may further include: when multiple reference pixel levels are not supported, a step of constructing reference pixels using pre-set reference pixel levels.

[0525] The method applicable to this invention can be implemented in the form of program instructions executable by various computing means and recorded on a computer-readable medium. The computer-readable medium can contain program instructions, data files, data structures, etc., individually or in combination. The program instructions recorded on the computer-readable medium can be program instructions specifically designed for this invention or program instructions known to those skilled in the art of computer software.

[0526] Examples of computer-readable media can include hardware devices with special configurations for storing and executing program instructions, such as read-only memory (ROM), random access memory (RAM), and flash memory. Examples of program instructions include not only machine code generated by a compiler, but also high-level language code that can be executed in a computer using an interpreter. The aforementioned hardware device can be constructed from at least one software module for performing the actions applicable to this invention, and vice versa.

[0527] Furthermore, the aforementioned methods or apparatus can combine or separate all or part of their components or functions.

[0528] The preferred embodiments of the present invention have been described above in conjunction with applicable embodiments. However, those skilled in the art will understand that various modifications and alterations can be made to the present invention without departing from the spirit and scope of the invention as set forth in the appended claims.< / threshold> < / array> < / postfilter> < / display> < / sizing> < / y>

Claims

1. A method for decoding a video signal using an image decoding device, the method comprising: Obtain the coefficients of the current block of the current image from the bitstream; Perform inverse quantization on the coefficients of the current block; as well as The current block is decoded on a sub-block basis by performing an inverse transform on the inverse-quantized coefficients. Specifically, based on the segmentation information obtained from the bitstream, the sub-block is obtained by segmenting the current block. The number of sub-blocks is 2. The segmentation information includes a segmentation ratio flag and a segmentation flag. The segmentation ratio flag indicates one of multiple segmentation ratios among the sub-blocks, and the segmentation flag indicates whether to segment the current block into the sub-block. The multiple division ratios mentioned include 1:

3. Specifically, the current image is segmented into multiple sub-images based on sub-image segmentation information. The sub-image segmentation information includes quantity information, position information, and size information. The quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the top-left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. Both the location information and the size information are represented in the largest coded unit. The plurality of sub-images includes a first sub-image and a second sub-image. A portion of the right boundary of the first sub-image is adjacent to the left boundary of the second sub-image, and The upper boundary of the first sub-image and the upper boundary of the second sub-image are discontinuous, while the lower boundary of the first sub-image and the lower boundary of the second sub-image are continuous.

2. The method of claim 1, wherein one of independence decoding or dependency decoding is applied to the sub-image. in, The independence decoding scheme refers to the following: when decoding the sub-image in the current image, it is not allowed to refer to other sub-images other than co-position sub-images in sub-images belonging to different images that are not the current image. The dependency decoding scheme refers to a decoding scheme in which, when decoding the sub-image in the current image, reference is allowed to other sub-images in different images.

3. The method of claim 2, wherein the first flag indicates whether the independence decoding or the dependency decoding is used for the sub-image. The first flag is obtained from the sequence parameters set in the bit stream. The first flag having a first value indicates that the independence decoding is applied to the sub-image, and The first flag having a second value indicates that the dependency decoding is applied to the sub-image.

4. The method according to claim 3, wherein, A second flag is obtained from the bitstream, which indicates whether loop filtering across the boundaries between the sub-images is applicable.

5. The method according to claim 4, wherein, The maximum coding unit refers to the basic coding unit with the largest size among the predefined coding units in the image decoding device.

6. The method according to claim 5, wherein, The sub-image is quadrilateral in shape.

7. A method for encoding a video signal using an image encoding device, the method comprising: The coefficients of the current block in the current image are obtained on a sub-block basis by performing a transformation on the residual sample of the current block. Quantization is performed on the coefficients of the current block; and The quantization coefficients of the current block are encoded into a bit stream. The sub-blocks are obtained by splitting the current block. The segmentation information based on the segmentation of the current block is encoded into the bit stream. The number of sub-blocks is 2. The segmentation information includes a segmentation ratio flag and a segmentation flag. The segmentation ratio flag indicates one of multiple segmentation ratios among the sub-blocks, and the segmentation flag indicates whether to segment the current block into the sub-block. The multiple division ratios mentioned include 1:

3. The current image is segmented into multiple sub-images based on sub-image segmentation information. The sub-image segmentation information includes quantity information, position information, and size information. The quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the top-left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. Both the location information and the size information are represented in the largest coded unit. The plurality of sub-images includes a first sub-image and a second sub-image. A portion of the right boundary of the first sub-image is adjacent to the left boundary of the second sub-image, and The upper boundary of the first sub-image and the upper boundary of the second sub-image are discontinuous, while the lower boundary of the first sub-image and the lower boundary of the second sub-image are continuous.

8. A method for transmitting a bit stream generated by an encoding method, the encoding method comprising: The coefficients of the current block in the current image are obtained on a sub-block basis by performing a transformation on the residual sample of the current block. Quantization is performed on the coefficients of the current block; and The quantization coefficients of the current block are encoded into a bit stream. The sub-blocks are obtained by splitting the current block. The segmentation information based on the segmentation of the current block is encoded into the bit stream. The number of sub-blocks is 2. The segmentation information includes a segmentation ratio flag and a segmentation flag. The segmentation ratio flag indicates one of multiple segmentation ratios among the sub-blocks, and the segmentation flag indicates whether to segment the current block into the sub-block. The multiple division ratios mentioned include 1:

3. The current image is segmented into multiple sub-images based on sub-image segmentation information. The sub-image segmentation information includes quantity information, position information, and size information. The quantity information indicates the number of sub-images belonging to the current image, the position information indicates the position of the top-left largest coding unit (LCU) within the sub-image, and the size information indicates the width and height of the sub-image. Both the location information and the size information are represented in the largest coded unit. The plurality of sub-images includes a first sub-image and a second sub-image. A portion of the right boundary of the first sub-image is adjacent to the left boundary of the second sub-image, and The upper boundary of the first sub-image and the upper boundary of the second sub-image are discontinuous, while the lower boundary of the first sub-image and the lower boundary of the second sub-image are continuous.

Citation Information

Patent Citations

  • Image decoding method and image decoding device

    JP6074743B2

  • Efficient scalable coding concept

    US20150304667A1