Image encoding / decoding method and apparatus

By adjusting and dividing the size of the basic coding blocks, the problem of low image coding/decoding efficiency in existing technologies is solved, and efficient coding of images of different sizes is achieved.

CN116506646BActive Publication Date: 2026-04-03INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-07-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle images whose size is not an integer multiple of the basic coding block size, resulting in low encoding/decoding efficiency.

Method used

By adjusting the size of the basic coding block and the size information of the boundary block, the boundary block is divided into at least one coding block and encoded to adapt to the encoding/decoding of images of different sizes.

Benefits of technology

It achieves unified encoding for images whose size is not an integer multiple of the basic coding block, thus improving encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116506646B_ABST
    Figure CN116506646B_ABST
Patent Text Reader

Abstract

This invention discloses an image coding method, comprising: a step of setting the size information and block segmentation information of a boundary block located at the boundary of the image and smaller than the size of the basic coding block, based on the size information of the image and the size information of the basic coding block; a step of segmenting the boundary block into at least one coding block based on the size information of the basic coding block, the size information of the boundary block, and the block segmentation information; and a step of encoding the at least one segmented coding block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application filed on July 12, 2018, with application number 2018800475878 and invention title "Image Encoding / Decoding Method and Apparatus". Technical Field

[0002] This invention relates to a method for encoding / decoding images whose size is not an integer multiple of the size of a basic coding block, and more particularly to a method and apparatus for encoding / decoding images whose size is not an integer multiple of the size of a basic coding block by magnifying the image or adjusting the coding block. Background Technology

[0003] With the widespread adoption of the internet and portable devices, as well as the development of communication technologies, the use of multimedia data is increasing rapidly. Therefore, in order to perform diverse services or tasks through image prediction in various systems, there is an urgent need to improve the performance and efficiency of image processing systems; however, research and development results that can meet these needs are few and far between.

[0004] As mentioned above, in existing image encoding and decoding methods and apparatuses, there is a need to improve the performance of image processing, especially image encoding or image decoding. Summary of the Invention

[0005] Technical problems to be solved

[0006] In order to solve the existing problems as described above, the object of the present invention is to provide a method for encoding / decoding images whose size is not an integer multiple of the basic coding block size.

[0007] In order to solve the existing problems as described above, the object of the present invention is to provide an apparatus for encoding / decoding images whose size is not an integer multiple of the basic coding block size.

[0008] Technical solution

[0009] To achieve the objectives described above, an image coding method according to one embodiment of the present invention may include: a step of setting the size information and block segmentation information of a boundary block located at the boundary of the image and smaller than the size of the basic coding block, based on the size information of the image and the size information of the basic coding block; a step of segmenting the boundary block into at least one coding block based on the size information of the basic coding block, the size information of the boundary block, and the block segmentation information; and a step of encoding the at least one segmented coding block.

[0010] It may also include: a step of enlarging the size of the image based on the size information of the boundary blocks and the size information of the basic coding blocks; and a step of adjusting the size information of the boundary blocks based on the size of the enlarged image.

[0011] The step of dividing a boundary block into at least one coded block may include: adjusting the size of a basic coded block according to block segmentation information; and dividing the block into at least one coded block based on the size information of the basic coded block, the size information of the boundary block, and the adjusted size information of the basic coded block.

[0012] Technical effect

[0013] In this invention, when the size of the image is not an integer multiple of the size of the basic coding block, encoding uniformity can be provided by enlarging the image to an integer multiple.

[0014] In this invention, when the size of the image is not an integer multiple of the size of the basic coding block, more efficient coding can be provided by adjusting the size of the basic coding block. Attached Figure Description

[0015] Figure 1 This is a conceptual diagram illustrating an image encoding and decoding system to which embodiments of the present invention are applied;

[0016] Figure 2 This is a block diagram illustrating an image encoding apparatus to which one embodiment of the present invention is applied;

[0017] Figure 3 This is a block diagram illustrating an image decoding apparatus to which one embodiment of the present invention is applied;

[0018] Figure 4 This is a schematic diagram illustrating the block shape based on a tree structure according to one embodiment of the present invention;

[0019] Figure 5 This is a schematic diagram illustrating various block configurations applicable to one embodiment of the present invention;

[0020] Figure 6 This is a schematic diagram illustrating a block used to explain the block division applicable to one embodiment of the present invention;

[0021] Figure 7 This is a sequence diagram illustrating an image encoding method applicable to one embodiment of the present invention;

[0022] Figure 8a as well as Figure 8bThis is a schematic diagram illustrating an image divided into basic coded blocks of various sizes to which one embodiment of the present invention is applied;

[0023] Figure 9 This is a first illustrative diagram used to explain a basic coding block adjustment method applicable to one embodiment of the present invention;

[0024] Figure 10 This is a second illustration used to explain a basic coding block adjustment method applicable to one embodiment of the present invention;

[0025] Figure 11a as well as Figure 11b This is a schematic diagram illustrating a boundary block used to explain a basic coding block adjustment method applicable to one embodiment of the present invention;

[0026] Figure 12 This is a sequence diagram illustrating a method for performing image encoding by adjusting basic coding blocks according to one embodiment of the present invention;

[0027] Figures 13a to 13b This is a first illustrative diagram used to explain an image size adjustment method applicable to one embodiment of the present invention;

[0028] Figure 14 This is a second illustration used to explain an image size adjustment method applicable to one embodiment of the present invention;

[0029] Figure 15 This is a sequence diagram illustrating a method for performing image encoding by adjusting the image size, applicable to one embodiment of the present invention;

[0030] Figures 16a to 16d This is an illustrative diagram used to explain an adaptive image data processing method applicable to one embodiment of the present invention;

[0031] Figures 17a to 17f This is a first illustration for explaining an adaptive image data processing method based on multiple segmentation methods applicable to one embodiment of the present invention;

[0032] Figures 18a to 18c This is a second illustration used to explain an adaptive image data processing method based on multiple segmentation methods applicable to one embodiment of the present invention;

[0033] Figure 19 This is a sequence diagram illustrating a method for performing image encoding through adaptive image data processing according to one embodiment of the present invention. Detailed Implementation

[0034] This invention is capable of various modifications and has many different embodiments. Specific embodiments will be illustrated and described in detail below. However, the following content is not intended to limit the invention to a particular implementation, but should be understood to include all modifications, equivalents, and substitutions within the scope of the invention's concept and technology.

[0035] In describing different constituent elements, terms such as "first," "second," "A," and "B" may be used, but these constituent elements are not limited by these terms. These terms are merely used to distinguish one constituent element from others. For example, without departing from the scope of the claims of this invention, a first constituent element can also be named a second constituent element, and similarly, a second constituent element can also be named a first constituent element. The term "and / or" includes a combination of multiple related descriptions or one of multiple related descriptions.

[0036] When a constituent element is described as being "connected" or "in contact" with other constituent elements, it should be understood that it can not only be directly connected or in contact with the aforementioned other constituent elements, but also that other constituent elements can exist between the two. Conversely, when a constituent element is described as being "directly connected" or "directly in contact" with other constituent elements, it should be understood that no other constituent elements exist between the two.

[0037] The terminology used in this application is for illustrative purposes only and is not intended to limit the invention. Singular statements also have plural meanings unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are used only to indicate the presence of features, numbers, steps, actions, constituent elements, components, or combinations thereof as described in the specification, and should not be construed as excluding the possibility of one or more other features, numbers, steps, actions, constituent elements, components, or combinations thereof being present or added.

[0038] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms commonly used, such as those defined in dictionaries, should be interpreted as having the same meaning as they have in the context of the relevant art, and should not be interpreted as having an overly idealized or exaggerated meaning in this application unless otherwise explicitly defined.

[0039] Typically, an image can be composed of a series of still images, which can be divided into groups of pictures (GOPs). Each still image can be called a picture or a frame. Higher-level concepts include GOPs and sequences, and each image can be further divided into specific regions such as strips, parallel blocks, or blocks. Furthermore, a GOP can include units such as I-images, P-images, and B-images. An I-image refers to an image that is encoded / decoded independently without using a reference image, while P-images and B-images refer to images encoded / decoded using reference images through processes such as motion estimation and motion compensation. Typically, P-images can use I-images and P-images as reference images, and B-images can use I-images and P-images as reference images; however, these definitions can change depending on the encoding / decoding settings.

[0040] In this context, the image used as a reference during the encoding / decoding process is called the Reference Picture, and the blocks or pixels used as references are called Reference Blocks and Reference Pixels. Furthermore, Reference Data, in addition to pixel values ​​in the spatial domain, can also include coefficient values ​​in the frequency domain, as well as various encoding / decoding information generated and determined during the encoding / decoding process.

[0041] The smallest unit constituting an image can be a pixel, and the number of bits used to represent a pixel is called the bit depth. Typically, the bit depth can be 8 bits, but other bit depths can be supported depending on the encoding settings. Regarding bit depth, at least one bit depth can be supported based on a color space. Furthermore, the image can be composed of at least one color space depending on its color format. Depending on the color format, the image can be composed of one or more images of a certain size or one or more images of different sizes. For example, in the case of YCbCr 4:2:0, the image can be composed of one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example), where the ratio of the chrominance component to the luminance component can be 1:2 (horizontal to vertical). As another example, in the case of 4:4:4, the image can have the same horizontal and vertical ratio. When the image is composed of one or more color spaces as described above, the image can be segmented in each color space.

[0042] In this invention, a portion of the color space (Y in this example) of a certain color format (YCbCr) will be used as a reference for explanation. The same or similar application is possible in other color spaces based on the color format (Cb, Cr in this example) (depending on the specific color space setting). However, some differences can also be retained in each color space (independent of the specific color space setting). That is, the setting dependent on each color space can refer to properties that are proportional to or dependent on the composition ratio of each component (e.g., determined according to 4:2:0, 4:2:2, or 4:4:4, etc.), while the setting independent of each color space can refer to properties that are unrelated to the composition ratio of each component or apply independently only to the corresponding color space. In this invention, depending on the encoder / decoder, a portion of the composition can have independent or dependent properties.

[0043] The configuration information or syntax elements required during video encoding can be determined at the unit level, such as video, sequence, image, slice, parallel block, or block. These elements can be included in the bitstream and transmitted to the decoder in units such as Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice Header, Tile Header, or Block Header. In the decoder, they can be parsed at the same unit level and used in the video decoding process after decoding the configuration information transmitted from the encoder. Each parameter set has its own unique numbering value, and lower-level parameter sets can include the numbering values ​​of higher-level parameter sets that need to be referenced. For example, a lower-level parameter set can reference information from higher-level parameter sets with consistent numbering values ​​from one or more higher-level parameter sets. In the various examples of units described above, when a unit contains one or more other units, the corresponding unit can be called the superior unit and the contained unit can be called the subordinate unit.

[0044] The setting information generated on the aforementioned units can include setting-related content that is independent within each unit, as well as setting-related content that depends on previous, subsequent, or superior units. Here, dependent settings refer to flag information (e.g., a 1-bit flag, where 1 indicates compliance and 0 indicates non-compliance) used to indicate whether the settings of previous, subsequent, or superior units are followed; this can be understood as representing the setting information of the corresponding unit. Although this invention will focus on examples related to independent settings, it can also include examples where setting information is added to or replaced using setting information from previous, subsequent, or superior units that depend on the current unit.

[0045] Video encoding / decoding is typically performed based on the input size, but sometimes it may be performed by resizing. For example, in scalability video coding, which supports spatial, temporal, and scalability, there may be instances where the overall resolution is adjusted by enlarging or reducing the image, or where a portion of the image is enlarged or reduced. Related details can be toggled by assigning selection information to units such as the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), and Slice Header, as described above. In the cases described above, the hierarchy between these units can be set as follows: Video Parameter Set (VPS) > Sequence Parameter Set (SPS) > Picture Parameter Set (PPS) > Slice Header, etc.

[0046] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0047] Figure 1 This is a conceptual diagram illustrating an image encoding and decoding system to which embodiments of the present invention are applied.

[0048] See Figure 1 The image encoding device 105 and the decoding device 100 can be user terminals such as personal computers (PCs), laptops, personal digital assistants (PDAs), portable multimedia players (PMPs), portable game consoles (PSPs), wireless communication terminals, smartphones, or televisions, or server terminals such as application servers and business servers. They can include various devices such as communication modems for communicating with various devices or wired and wireless communication networks, memory 120, 125 for storing various applications and data for performing inter-frame or intra-frame prediction in order to encode or decode images, or processors 110, 115 for performing calculations and control by executing applications.

[0049] Furthermore, the image encoded into a bitstream by the image encoding device 105 can be transmitted to the image decoding device 100 via wired wireless communication networks such as the Internet, short-range wireless communication networks, wireless local area networks, wireless broadband networks, or mobile communication networks, or via various communication interfaces such as cables or Universal Serial Bus (USB). The image is then decoded and reconstructed in the image decoding device 100 before being played back. Additionally, the image encoded into a bitstream by the image encoding device 105 can also be transmitted from the image encoding device 105 to the image decoding device 100 via a computer-readable storage medium.

[0050] The image encoding device and image decoding device described above can be separate and independent devices, or they can be manufactured as a single image encoding / decoding device depending on the specific implementation. In the above cases, a portion of the components of the image encoding device and a portion of the components of the image decoding device are actually the same technical elements, and can at least include the same structure or at least perform the same function.

[0051] Therefore, in the following detailed explanation of the technical elements and their working principles, repeated descriptions of the corresponding technical elements will be omitted. Furthermore, the image decoding apparatus is a computing device that applies the image decoding method executed in the image encoding apparatus to the decoding; therefore, the following explanation will focus on the image encoding apparatus.

[0052] The computing device may include: a memory for storing programs or software modules for implementing image encoding methods and / or image decoding methods; and a processor for executing the programs via connection to the memory. The image encoding device may be referred to as an encoder, and the image decoding device may be referred to as a decoder.

[0053] Figure 2 This is a block diagram illustrating an image encoding apparatus to which one embodiment of the present invention is applied.

[0054] See Figure 2 The image encoding apparatus 20 may include: a prediction unit 200, a subtraction unit 205, a transformation unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse transformation unit 225, an addition unit 230, a filtering unit 235, an encoded image buffer 240, and an entropy encoding unit 245.

[0055] The prediction unit 200 can be implemented using a software module, namely a prediction module, and can generate prediction blocks by performing intra-frame prediction or inter-frame prediction on the blocks to be encoded. The prediction unit 200 can generate prediction blocks by predicting the current block in the image that needs to be encoded. In other words, the prediction unit 200 can generate prediction blocks that include predicted pixel values ​​of each pixel generated by performing intra-frame or inter-frame prediction on the pixel values ​​of each pixel in the current block to be encoded in the image. Furthermore, the prediction unit 200 can also transmit information necessary for generating prediction blocks, such as information related to prediction modes (e.g., intra-frame prediction mode or inter-frame prediction mode), to the encoding unit, thereby enabling the encoding unit to encode the information related to the prediction mode. At this time, the processing unit performing the prediction and the processing unit determining the prediction method and specific content can differ depending on the encoding / decoding settings. For example, prediction methods and prediction models can be determined in terms of prediction units, while the execution of predictions can be carried out in terms of transformation units.

[0056] The in-frame prediction unit can employ directional prediction modes such as horizontal and vertical modes, which are based on the prediction direction, as well as non-directional prediction modes such as mean (DC) and planar modes, which use methods such as averaging and interpolation of reference pixels. Using directional and non-directional modes, candidate groups of in-frame prediction modes can be constructed. A candidate group can be selected from various options, such as 35 prediction modes (33 directional + 2 non-directional), 67 prediction modes (65 directional + 2 non-directional), or 131 prediction modes (129 directional + 2 non-directional).

[0057] The in-frame prediction unit may include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel construction unit can construct reference pixels for performing in-frame prediction, centered on the current block, using pixels contained in and adjacent to the current block. Depending on the encoding settings, the reference pixel can be constructed using the nearest adjacent row of reference pixels, or using any other adjacent row of reference pixels, or using multiple rows of reference pixels. When some of the reference pixels are unavailable, the reference pixel can be generated using the available reference pixels; when all are unavailable, the reference pixel can be generated using a pre-set value (e.g., the median value of a pixel value range represented by bit depth).

[0058] The reference pixel filtering unit of the prediction section within the image can perform filtering on reference pixels to reduce distortion remaining from the encoding process. The filter used can be a low-pass filter such as a 3-tap filter [1 / 4, 1 / 2, 1 / 4] or a 5-tap filter [2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16]. The applicability and type of filtering can be determined based on encoding information (such as block size, shape, prediction mode, etc.).

[0059] The reference pixel interpolation unit of the in-frame prediction section can generate fractional-unit pixels through a linear interpolation process of the reference pixels according to the prediction mode, and can determine the applicable interpolation filter based on the encoding information. The interpolation filters used can include 4-tap Cubic filters, 4-tap Gaussian filters, 6-tap Wiener filters, 8-tap Kalman filters, etc. Typically, the low-pass filtering process and the interpolation process are independent, but it is also possible to combine the applicable filters from both processes into one before performing the filtering process.

[0060] The prediction mode determination unit of the in-frame prediction unit can select at least one optimal prediction mode from the candidate prediction mode group while taking into account the encoding cost, and the prediction block generation unit can generate prediction blocks using the corresponding prediction mode. In the prediction mode encoding unit, the optimal prediction mode can be encoded based on the prediction value. At this time, the prediction information can be adaptively encoded according to whether the prediction value is suitable or unsuitable.

[0061] In the in-image prediction unit, the aforementioned predicted value is referred to as the Most Probable Mode (MPM). A subset of modes can be selected from all modes included in the prediction mode candidate group to form the Most Probable Mode (MPM) candidate group. The Most Probable Mode (MPM) candidate group can include pre-defined prediction modes (e.g., mean (DC), planar, vertical, horizontal, diagonal modes, etc.) or prediction modes of spatially adjacent blocks (e.g., left, top, upper left, upper right, lower left blocks, etc.). Furthermore, the Most Probable Mode (MPM) candidate group can be constructed using modes derived from those pre-included in the Most Probable Mode (MPM) candidate group (differences such as +1 and -1 in directional modes).

[0062] Among the prediction patterns used to construct the most probable pattern (MPM) candidate group, a priority order can exist. The order in which the most probable pattern (MPM) candidate groups are included can be determined according to this priority order, and the construction of the most probable pattern (MPM) candidate group can be completed when the number of most probable pattern (MPM) candidate groups (determined based on the number of prediction pattern candidate groups) is filled according to this priority order. At this time, the priority order can be determined according to the order of prediction patterns of spatially adjacent blocks, pre-defined prediction patterns, and patterns derived from prediction patterns earlier included in the most probable pattern (MPM) candidate group, or other variations can be made.

[0063] For example, spatially adjacent blocks can be included in the candidate group in the order of left-top-bottom-top-right-top, etc. In a pre-defined prediction pattern, they can be included in the candidate group in the order of mean (DC)-planar-vertical-horizontal pattern, etc. Patterns obtained by adding 1, subtracting 1, etc., to the pre-included patterns are then included in the candidate group, thus using a total of 6 patterns to form a candidate group. Alternatively, they can be included in the candidate group in a priority order of left-top-mean (DC)-planar-bottom-top-right-(left+1)-(left-1)-(top+1), thus using a total of 7 patterns to form a candidate group.

[0064] In the inter-frame prediction unit, motion prediction methods can be categorized into mobile motion models and non-mobile motion models. In the mobile motion model, prediction is performed considering only parallel movement, while in the non-mobile motion model, prediction is performed considering rotation, distance, and zoom (zoom in / out) motion along with parallel movement. Assuming unidirectional prediction, the mobile motion model requires one motion vector, while the non-mobile motion model requires more than one. In the non-mobile motion model, each motion vector can be information applicable to a pre-defined position within the current block, such as the top-left vertex or top-right vertex. Using the corresponding motion vectors, the position of the area to be predicted within the current block can be obtained in pixel units or sub-block units. The inter-frame prediction unit can apply some of the processes described below along with other processes individually, according to the aforementioned motion models.

[0065] The inter-frame prediction unit may include a reference image composition unit, a motion prediction unit, a motion compensation unit, a motion information determination unit, and a motion information encoding unit. The reference image composition unit can include previously or subsequently encoded images in a reference image list (L0, L1) centered on the current image. From the reference images included in the reference image list, prediction blocks can be obtained, and according to the encoding settings, reference images can be composed using the current image and included in at least one position in the reference image list.

[0066] In the inter-frame prediction unit, the reference image composition unit can include a reference image interpolation unit, and can perform an interpolation process for fractional-unit pixels based on the interpolation accuracy. For example, an 8-tap interpolation filter based on Discrete Cosine Transform (DCT) can be applied to the luminance component, and a 4-tap interpolation filter based on Discrete Cosine Transform (DCT) can be applied to the chrominance component.

[0067] In the inter-frame prediction unit, the motion prediction unit is used to perform the process of exploring blocks that are highly relevant to the current block by using reference images. It can use various methods such as the Full Search-based Block Matching Algorithm (FBMA) and the Three Step Search (TSS). The motion compensation unit is used to perform the process of obtaining the prediction block through the motion prediction process.

[0068] In the inter-frame prediction unit, the motion information determination unit performs a process to select the best motion information for the current block. The motion information can be encoded using motion information encoding modes such as Skip Mode, Merge Mode, and Competition Mode. These modes can be configured by combining supported modes according to the motion model. Examples include Skip Mode (moving), Skip Mode (non-moving), Merge Mode (moving), Merge Mode (non-moving), Competition Mode (moving), and Competition Mode (non-moving). Depending on the symbolization settings, a subset of these modes can be included in the candidate group.

[0069] The aforementioned motion information encoding mode can obtain the predicted value of the motion information (motion vector, reference image, prediction direction, etc.) of the current block from at least one candidate block, and can generate the best candidate selection information when supporting more than two candidate blocks. The skip mode (without residual signal) and the merge mode (with residual signal) can directly use the above predicted value as the motion information of the current block, while the competition mode can generate the difference value information between the motion information of the current block and the above predicted value.

[0070] Candidate groups for motion information prediction values ​​of the current block are adaptive based on the motion information coding mode and can adopt various configurations. Motion information of blocks spatially adjacent to the current block (e.g., left, top, top left, top right, bottom left blocks, etc.) can be included in the candidate group. Motion information of blocks temporally adjacent to the current block (e.g., left, right, top, bottom, top left, top right, bottom left, bottom right blocks, etc., including blocks in other images corresponding to or related to the current block <center>) can also be included in the candidate group. In addition, mixed motion information of spatial and temporal candidates (e.g., information obtained by averaging, median values, etc., from the motion information of spatially adjacent blocks and the motion information of temporally adjacent blocks, or motion information obtained by the current block or its sub-blocks) can also be included in the candidate group.

[0071] In the construction of candidate groups for motion information prediction values, a priority order can exist. The order in which the candidate groups are included can be determined according to this priority order, and the construction of candidate groups can be completed when the number of candidate groups is filled according to this priority order (determined based on the motion information encoding mode). At this time, the priority order can be determined according to the order of motion information of spatially adjacent blocks, motion information of temporally adjacent blocks, and mixed motion information of spatial and temporal candidates, or other variations can be made.

[0072] For example, spatially adjacent blocks can be included in the candidate group in the order of left-top-top-left-bottom-top-left blocks, while temporally adjacent blocks can be included in the candidate group in the order of bottom-right-middle-right-bottom blocks.

[0073] The subtraction unit 205 can generate a residual block by performing a subtraction operation between the current block and the predicted block. In other words, the subtraction unit 205 can generate a residual signal, i.e., a residual block, in block form by calculating the difference between the pixel values ​​of each pixel in the current block to be encoded and the predicted pixel values ​​of each pixel in the predicted block generated by the prediction unit. Furthermore, the subtraction unit 205 can also generate residual blocks in units other than the block units obtained by the block segmentation unit described later.

[0074] The transformation unit 210 transforms the pixel values ​​of the residual block into frequency coefficients by transforming the residual block into a frequency region. Specifically, the transformation unit 210 can utilize various transformation techniques for transforming spatial axis pixel signals into frequency axis signals, such as the Hadamard Transform, Discrete Cosine Transform (DCT-based Transform), Discrete Sine Transform (DST-based Transform), and Caronan-Louis Transform (KLT-based Transform), to transform the residual signal into a frequency signal. The residual signal transformed into a frequency region becomes the frequency coefficients. The transformation can be performed using a 1D transformation matrix. Each transformation matrix can be adaptively used in both horizontal and vertical units. For example, when the prediction mode in in-frame prediction is horizontal, a DCT-based transformation matrix can be used in the vertical direction and a DST-based transformation matrix can be used in the horizontal direction. When the prediction mode is vertical, a transformation matrix based on Discrete Cosine Transform (DCT) can be used in the horizontal direction and a transformation matrix based on Discrete Sine Transform (DST) can be used in the vertical direction.

[0075] The quantization unit 215 quantizes the residual block containing frequency coefficients that has been transformed into a frequency region by the transformation unit 210. The quantization unit 215 can quantize the transformed residual block using quantization techniques such as dead zone uniform threshold quantization, quantization weighted matrix, or modified quantization techniques. At this time, one or more quantization techniques can be selected as candidates, and the selection can be based on coding mode, prediction mode information, etc.

[0076] The inverse quantization unit 220 performs inverse quantization on the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 generates a residual block containing frequency coefficients by performing inverse quantization on the quantized frequency coefficient column.

[0077] The inverse transform unit 225 performs an inverse transform on the residual block that has been inversely quantized by the inverse quantization unit 220. That is, the inverse transform unit 225 generates a residual block containing pixel values, i.e., a reconstructed residual block, by performing an inverse transform on the frequency coefficients of the inversely quantized residual block. The inverse transform unit 225 can perform the inverse transform by reversing the transform method used in the transform unit 210.

[0078] The addition unit 230 can reconstruct the current block by performing an addition operation on the prediction block predicted in the prediction unit 200 and the residual block reconstructed by the inverse transform unit 225. The reconstructed current block is stored as a reference image (or reference block) in the encoded image buffer 240, so that it can be used as a reference image when encoding the next block or subsequent blocks or other images of the current block.

[0079] The filtering unit 235 can include one or more post-processing filtering procedures, such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can eliminate block distortion appearing at the boundaries between blocks in the reconstructed image. The adaptive loop filter (ALF) can perform filtering based on the values ​​obtained by comparing the reconstructed image with the original image after filtering the blocks by the deblocking filter. The sample adaptive offset (SAO) can reconstruct the offset difference between the residual blocks that have been treated with the deblocking filter and the original image, pixel by pixel. The post-processing filters described above can be applied to the reconstructed image or blocks.

[0080] The encoded image buffer 240 can store the blocks or images reconstructed by the filtering unit 235. The reconstructed blocks or images stored in the encoded image buffer 240 can be provided to the prediction unit 200 for performing intra-frame prediction or inter-frame prediction.

[0081] The entropy coding unit 245 generates a quantization coefficient column by scanning the generated quantization frequency coefficient column using various scanning methods, and outputs it after encoding using techniques such as entropy coding. The scanning mode can be set to one of several modes, such as zigzag, diagonal, or raster. Furthermore, it can generate encoded data containing encoding information passed from each component and output it to a bitstream.

[0082] Figure 3 This is a block diagram illustrating an image decoding apparatus to which one embodiment of the present invention is applied.

[0083] See Figure 3 The image decoding device 30 may include an entropy decoding unit 305, a prediction unit 310, an inverse quantization unit 315, an inverse transform unit 320, an addition and subtraction unit 325, a filter 330, and a decoded image buffer 335.

[0084] Furthermore, the prediction unit 310 may further include an in-frame prediction module and an inter-frame prediction module.

[0085] First, when receiving the image bitstream transmitted from the image encoding device 20, it can be transmitted to the entropy decoding unit 305.

[0086] The entropy decoding unit 305 can decode decoded data containing quantized coefficients and decoding information transmitted from each component unit by decoding the bit stream.

[0087] The prediction unit 310 can generate prediction blocks based on the data transmitted from the entropy decoding unit 305. At this time, it can also construct a list of reference images using the default composition technique based on the decoded reference images stored in the image buffer 335.

[0088] The intra-frame prediction unit may include a reference pixel composition unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction block generation unit, and a prediction mode decoding unit. The inter-frame prediction unit may include a reference image composition unit, a motion compensation unit, and a motion information decoding unit. One part of the unit may perform the same process as the encoder, while the other part may perform the reverse induction process.

[0089] The inverse quantization unit 315 is capable of inverse quantizing the quantized transform coefficients provided from the bitstream and decoded by the entropy decoding unit 305.

[0090] The inverse transform unit 320 can generate residual blocks by applying inverse discrete cosine transform (DCT), inverse integer transform, or similar inverse transform techniques to the transform coefficients.

[0091] At this time, the inverse quantization unit 315 and the inverse transformation unit 320 will reverse the processes executed in the transformation unit 210 and the quantization unit 215 of the image encoding apparatus 20 described above, and this can be achieved through various methods. For example, the same processes and inverse transformations shared with the transformation unit 210 and the quantization unit 215 can be used, or the transformation and quantization processes can be reversed using information related to the transformation and quantization processes of the image encoding apparatus 20 (such as transformation size, transformation shape, quantization type, etc.).

[0092] The residual blocks that have undergone inverse quantization and inverse transform can be added to the prediction blocks derived in the prediction unit 310 to generate reconstructed image blocks. The addition operation described above can be performed by the addition / subtraction operator 325.

[0093] For the reconstructed image blocks, filter 330 can be used as needed to apply a deblocking filter to eliminate blocking phenomena, and can also add other loop filters before and after the above decoding process to improve video quality.

[0094] The reconstructed and filtered image blocks can be stored in the decoded image buffer 335.

[0095] Although not illustrated, the image encoding / decoding device may also include an image segmentation unit and a block segmentation unit.

[0096] The image segmentation unit can segment (or divide) an image into at least one processing unit, such as color space (YCbCr, RGB or XYZ, etc.), parallel blocks, stripes, basic coding units (or maximum coding units), while the block segmentation unit can segment the basic coding units into at least one processing unit (such as coding, prediction, transform, quantization, entropy and loop filtering units, etc.).

[0097] The basic coding unit can be obtained by dividing the image along the horizontal and vertical directions at certain intervals. Based on this, segmentation can be performed on parallel blocks, stripes, etc. (that is, segmentation units such as parallel blocks or stripes are composed of integer multiples of the basic coding block, but special cases may occur on segmentation units located at the image boundary), but it is not limited to this.

[0098] For example, the image can be segmented into basic coding units after being divided into them, or the image can be segmented into basic coding units after being divided into them. In this invention, the former case will be described assuming that the division of each unit and the order of segmentation are as described, but it is not limited to this, and the latter case can also be used depending on the encoding / decoding settings. In the latter case, it can be modified to a case where the size of the basic coding unit is adaptive according to the segmentation unit (parallel block) (i.e., different sizes of basic coding blocks can be applied to each unit).

[0099] In this invention, the following examples will be described with the case where the image is divided into basic coding units as the basic setting (i.e., the image is not divided into parallel blocks or stripes, or the image is a single parallel block or stripe). However, in the case where each segmentation unit is first divided as described above, and then the obtained units are divided into basic coding units (i.e., each segmentation unit is not an integer multiple of the basic coding unit), the settings proposed in each example can be the same or applied after modification.

[0100] In the aforementioned segmentation units, stripes can be composed of combinations of at least one block that is consecutive in the scanning sequence, while parallel blocks can be composed of combinations of rectangular shapes of spatially adjacent blocks. Furthermore, other additional segmentation units can be supported and constructed according to their definitions. In particular, parallel blocks can segment the image using one or more horizontal and vertical lines in a checkerboard shape. Alternatively, they can be composed of rectangles of different sizes that are not segmented using the uniform horizontal lines and digital displays described above.

[0101] Furthermore, it is possible to divide the code into encoding units (or blocks) of various sizes through block segmentation. In this case, the encoding unit can be composed of multiple encoding blocks according to the color format (e.g., one luminance encoding block and two chrominance encoding blocks), and then the size of the block is determined according to the color format. For ease of explanation, the following explanation will be based on a block based on a single color component (luminance component).

[0102] The following explanation uses a single color component as an example. However, it's important to understand that the ratio can be proportionally adjusted based on the length proportions of the color format (e.g., in YCbCr 4:2:0, the horizontal and vertical length ratios of the luminance and chromatic aberration components are 2:1) and applied to other color components. Furthermore, block segmentation can be performed based on other color components (e.g., in CbCr, depending on the block segmentation results of Y). However, it's crucial to understand that block segmentation can be performed independently for each color component. Additionally, a single block segmentation setting can be used (considering its relationship to the length proportions), but it's also important to understand that independent block segmentation settings can be used for each color component.

[0103] The encoded block can be of a variable size, such as M×M (e.g., M is 4, 8, 16, 32, 64, 128, etc.). Alternatively, the encoded block can also be configured based on the segmentation method (e.g., tree-based segmentation, quadtree segmentation).<Quad Tree,QT> binary tree<Binary Tree,BT> (etc.) Variable sizes and shapes such as M×M (e.g., M is 4, 8, 16, 32, 64, 128, etc.) are adopted. In this case, the coded block can be the basic unit for intra-frame prediction, inter-frame prediction, transformation, quantization, and entropy coding, while the block is assumed to be a unit that can be obtained after the various units are determined.

[0104] The block segmentation unit can be set according to the various components of the image encoding device and the image decoding device, and the size and shape of the blocks can be determined through the above process. At this time, the set blocks can be defined differently according to the components, such as defining prediction blocks in the prediction unit, transformation blocks in the transformation unit, and quantization blocks in the quantization unit. However, it is not limited to this; block units based on other components can also be defined. The size and shape of the blocks can be defined according to the horizontal and vertical lengths of the blocks.

[0105] In the block segmentation section, blocks can be retrieved within the maximum and minimum value ranges of each block. For example, when the block shape supports a square and the maximum value of the block is set to 256×256 and the minimum value is set to 8×8, 2 blocks can be retrieved. m ×2 m Blocks of size (where m is an integer from 3 to 8 in this example, such as 8×8, 16×16, 32×32, 64×64, 128×128, 256×256), blocks of size 2m×2m (where m is an integer from 4 to 128 in this example), or blocks of size m×m (where m is an integer from 8 to 256 in this example).

[0106] Alternatively, when the block shape supports squares and rectangles and its extent is the same as in the examples above, it is possible to obtain 2. m ×2 n The block size (in this example, m and n are integers from 3 to 8, assuming a maximum horizontal-to-vertical ratio of 2:1, including 8×8, 8×16, 16×8, 16×16, 16×32, 32×16, 32×32, 32×64, 64×32, 64×64, 64×128, 128×64, 128×128, 128×256, 256×128, and 256×256). However, depending on the encoding / decoding settings, there may be no limit to the horizontal-to-vertical ratio or a maximum limit on the ratio (e.g., 1:2, 1:3, 1:7, etc.). Alternatively, it may be possible to obtain blocks of size 2m×2n (in this example, m and n are integers from 4 to 128), or blocks of size m×n (in this example, m and n are integers from 8 to 256).

[0107] Blocks that can be obtained through the block segmentation unit can be determined based on encoding / decoding settings (such as block type, segmentation method, and segmentation settings). For example, as an encoded block, 2 blocks can be obtained. m ×2 n The size of the block can be used as a prediction block to obtain a block of size 2m×2n or m×n, and as a transformation block to obtain a block of size 2. m ×2 nBlock size. In other words, it is possible to generate information such as block size and range (e.g., exponent, multiple information, etc.) based on the above encoding / decoding settings.

[0108] For example, the maximum and minimum block sizes can be determined based on the block type. Furthermore, the block range information for some blocks can be explicitly generated, while for others it can be implicitly determined. For instance, the relevant information for encoded blocks and transformed blocks can be explicitly generated, while the relevant information for predicted blocks can be implicitly processed.

[0109] In an explicit manner, at least one range information can be generated. For example, range-related information for an encoded block can be generated based on the maximum and minimum values, or based on the difference between the maximum value and a preset minimum value (e.g., 8). In other words, range-related information can be generated based on the above settings, or based on the exponential difference between the maximum and minimum values, but is not limited to these. Furthermore, multiple range information related to the horizontal and vertical lengths of a rectangular block can be generated.

[0110] By default, range information can be obtained based on encoding / decoding settings (such as block type, segmentation method, and segmentation settings). For example, the predicted block can obtain the maximum and minimum value information from the parent unit, i.e., the encoded block (such as the maximum size M×N and minimum size m×n of the encoded block), based on the candidate groups (M×N and m / 2×n / 2 in this example) that can be obtained from the segmentation settings of the predicted block (such as quadtree segmentation + segmentation depth 0).

[0111] The size and shape of the initial (or starting) block of the block segmentation unit can be determined from the higher-level unit. In other words, the coded block can use the basic coded block as its initial block, and the prediction block can use the coded block as its initial block. Furthermore, the transform block can use either the coded block or the prediction block as its initial block, which can be determined according to the encoding / decoding settings.

[0112] For example, when the coding mode is intra, the prediction block can be a higher-level unit than the transform block, while when the coding mode is inter, the prediction block can be a unit independent of the transform block. The initial unit of segmentation, i.e., the initial block, can be divided into smaller blocks, and after determining the optimal size and shape of the block segmentation, these blocks can be determined as the initial blocks of the lower-level units. The initial unit of segmentation, i.e., the initial block, can be considered as the initial block of the higher-level unit. The higher-level unit can be a coding block, and the lower-level unit can be a prediction block or a transform block, but is not limited to these. After determining the initial block in the manner shown in the example above, the segmentation process of finding the optimal size and shape of the block can be performed in conjunction with the higher-level unit.

[0113] In other words, the block segmentation unit can divide a basic coding unit (or the largest coding unit) into at least one coding unit (or a lower-level coding unit). Alternatively, a coding unit can be divided into at least one prediction unit, or it can be divided into at least one transform unit. A coding unit can be divided into at least one coding block, and a coding block can be divided into at least one prediction block, or it can be divided into at least one coding block. A prediction unit can be divided into at least one prediction block, and a transform unit can be divided into at least one transform block.

[0114] The above examples illustrate different segmentation settings based on block type. Furthermore, some blocks can undergo a segmentation process by combining them with other blocks. For instance, when combining a coded block and a transform block into a single unit, a segmentation process can be performed to obtain the optimal block size and shape, which can be the optimal size and shape for both the coded and transform blocks. Alternatively, coded blocks and transform blocks can be combined into a single unit, as can prediction blocks and transform blocks, or even a combination of coded blocks, prediction blocks, and transform blocks. Combining other blocks is also possible.

[0115] When searching for blocks of optimal size and shape as described above, associated pattern information (such as segmentation information) can be generated. All this information, along with information generated on the constituent parts to which the block belongs (such as prediction-related information and transformation-related information), can be included in the bitstream and transmitted to the decoder, where it can be parsed at the same level and used during the image decoding process.

[0116] The segmentation method will be explained next. For the sake of convenience, it will be assumed that the initial block is square. However, the same or similar method can be applied when the initial block is rectangular, so it is not limited to this.

[0117] Figure 4 This is a schematic diagram illustrating the block shape based on a tree structure applicable to one embodiment of the present invention.

[0118] The segmentation pattern can be determined based on the type of tree structure used during segmentation. See also Figure 4 The partitioning pattern can include a 2N×2N partition without partitioning, two 2N×N partitions based on horizontal partitioning of a binary tree, two N×2N partitions based on vertical partitioning of a binary tree, and four N×N partitions based on quadtree partitioning, but it is not limited to these.

[0119] Specifically, when performing a quadtree-based partition, the candidate blocks that can be obtained are 4a and 4d, while when performing a binary tree-based partition, the candidate blocks that can be obtained are 4a, 4b, and 4c.

[0120] When performing quadtree-based partitioning, a partitioning flag can be supported. This flag can indicate whether or not to partition. In other words, when the quadtree partitioning flag is 0, partitioning can be skipped and a block as shown in 4a can be obtained; when it is 1, partitioning can be performed and a block as shown in 4d can be obtained.

[0121] When performing binary tree-based partitioning, multiple partitioning flags can be supported. One of these flags can be a partitioning-to-partition flag, and another can be a partitioning-direction flag. In other words, when the partitioning-to-partition flag in the binary tree partitioning is 0, partitioning can be skipped and a block as shown in 4a can be obtained; when it is 1, a block as shown in 4b or 4d can be obtained based on the partitioning-direction flag.

[0122] Figure 5 This is a schematic diagram illustrating various block configurations applicable to one embodiment of the present invention.

[0123] See Figure 5 A 4N×4N block can be divided into various forms according to the partitioning settings and methods, and can also be partitioned into... Figure 5 Other forms not illustrated in the diagram.

[0124] As one embodiment, asymmetric partitioning based on a tree structure can be allowed. For example, when performing binary tree partitioning, a 4N×4N block can allow symmetric blocks such as 4N×2N (5b) and 2N×4N (5c), and can also allow asymmetric blocks such as 4N×3N / 4N×N (5d), 4N×N / 4N×3N (5e), ​​3N×4N / N×4N (5f), and N×4N / 3N×4N (5g). When the flag allowing asymmetric partitioning in the partitioning flags is explicitly or implicitly disabled according to the partitioning settings, the candidate block can be 5b or 5c depending on its partitioning direction flag. When the flag allowing asymmetric partitioning is enabled, the candidate block can be 5b, 5d, and 5e or 5c, 5f, and 5g depending on its partitioning direction flag. Among them, the illustrations in 5d to 5g for asymmetric division show the cases where the length ratio of the left and right sides or the top and bottom is 1:3 or 3:1, but cases of 1:2, 1:4, 2:3, 2:4 and 3:4 can also be used, so it is not limited to these.

[0125] As a segmentation flag for binary tree partitioning, in addition to generating a segmentation whether flag and a segmentation direction flag, a segmentation shape flag can also be generated. The segmentation shape flag can indicate symmetry or asymmetry. When the segmentation shape flag indicates asymmetry, a flag indicating the segmentation ratio can also be generated. The flag indicating the segmentation ratio can be represented using a pre-assigned index based on pre-defined candidate groups. For example, when supporting candidate groups with a segmentation ratio of 1:3 or 3:1, a 1-bit flag can be used to select the segmentation ratio.

[0126] Furthermore, in addition to generating a splitting flag and a splitting direction flag, a flag indicating the splitting ratio can also be generated as a splitting flag for binary tree splitting. In this case, the candidates related to the splitting ratio can include candidates with a symmetry ratio of 1:1.

[0127] In this invention, the case of distinguishing binary tree segments by using additional segmentation morphology markers will be used as an example. Unless otherwise explicitly stated, binary tree segmentation may refer to symmetric binary tree segmentation.

[0128] As another embodiment, tree-based segmentation can also allow for additional tree-based segmentations. For example, it can allow segmentations such as ternary trees, quad-type trees, and octa trees, thereby obtaining n segmentation blocks (n is an integer). In other words, 3 segmentation blocks can be obtained when segmenting with a ternary tree, 4 segmentation blocks can be obtained when segmenting with a quad-type tree, and 8 segmentation blocks can be obtained when segmenting with an octa tree.

[0129] The blocks supported for ternary tree partitioning can be 4N×2N / 4N×N_2(5h), 4N×N / 4N×2N / 4N×N(5i), 4N×N_2 / 4N×2N(5j), 2N×4N / N×4N_2(5k), N×4N / 2N×4N / N×4N(51) and N×4N_2 / 2N×4N(5m), the blocks supported for quadtree partitioning can be 2N×2N(5n), 4N×N(5o) and N×4N(5p), and the blocks supported for octree partitioning can be N×N(5q).

[0130] Whether tree-based segmentation is supported can be determined implicitly or explicitly based on encoding / decoding settings. Furthermore, it can be used independently based on encoding / decoding settings, or in combination with binary, ternary, and quadtree segmentations.

[0131] For example, when performing binary tree partitioning, blocks such as 5b or 5c can be supported depending on the partitioning direction. When binary tree partitioning and ternary tree partitioning are used in combination, blocks such as 5b, 5c, 5i, and 5l can be supported, assuming that their usage partially overlaps. When the flag allowing additional partitioning beyond the existing tree structure is explicitly or implicitly disabled according to the encoding / decoding settings, the available candidate blocks can be 5b or 5c. When enabled, the available candidate blocks can be 5b, 5i (or 5b, 5h, 5i, 5j) or 5c, 5l (or 5c, 5k, 5l, 5m) depending on their partitioning direction.

[0132] The diagram illustrates the cases where the length ratio of the left / middle / right or top / middle / bottom is 2:1:1, 1:2:1, or 1:1:2 when dividing a ternary tree. However, other ratios can be used depending on the encoding settings, so it is not limited to these.

[0133] When performing ternary tree-based partitioning, multiple partitioning flags can be supported. Among these flags, one can be a partitioning-whether-partitioning flag, another can be a partitioning-direction flag, and a partitioning-ratio flag can also be included.

[0134] Similar to the segmentation markers used in binary tree segmentation, this invention will assume, for example, that the segmentation ratio marker is omitted because it includes a candidate with a ratio of 1:2:1 supported by the segmentation direction.

[0135] In this invention, adaptive encoding / decoding settings can be applied according to the segmentation method.

[0136] As one embodiment, the segmentation method can be determined based on the type of block. For example, coded blocks and transform blocks can be segmented using quadtrees, while prediction blocks can be segmented using quadtrees and binary trees (or ternary trees, etc.).

[0137] As another embodiment, the partitioning method can be determined based on the size of the block. For example, a quadtree partitioning can be performed within a certain range (e.g., a×b to c×d) between the maximum and minimum values ​​of the block, while a binary tree (or ternary tree, etc.) partitioning can be performed within another range (e.g., e×f to g×h). Here, "a certain range" can refer to a larger size than "other certain ranges," but is not limited to this. The range information based on the partitioning method can be explicitly generated or implicitly determined, and the ranges can overlap.

[0138] As another embodiment, the segmentation method can be determined based on the shape of the block (or the block before segmentation). For example, when the block is square, quadtree segmentation and binary tree segmentation (or ternary tree segmentation, etc.) can be performed. Alternatively, when the block is rectangular, binary tree segmentation (or ternary tree segmentation, etc.) can be performed.

[0139] Tree-based segmentation can be performed recursively. For example, when the segmentation flag of a coded block with a segmentation depth of k is 0, the encoding of the coded block is performed on the coded block with a segmentation depth of k. When the segmentation flag of a coded block with a segmentation depth of k is 1, the encoding of the coded block will be performed on four sub-coded blocks (in quadtree segmentation), two sub-coded blocks (in binary tree segmentation), or three sub-coded blocks (in ternary tree segmentation) with a segmentation depth of k+1, depending on the segmentation method. The sub-coded block will be reset to coded block k+1 and can be further segmented into sub-coded block k+2 through the above process. The hierarchical segmentation method described above can be determined based on segmentation settings such as segmentation range and allowable segmentation depth.

[0140] At this point, the structure of the bitstream used to express segmentation information can be selected from more than one scanning method. For example, the bitstream of segmentation information can be constructed based on the segmentation depth order, or it can be constructed based on whether or not segmentation has occurred. For example, in the case of using the segmentation depth order as the basis, a method can be used that obtains the segmentation information of the current level depth based on the initial block and then obtains the segmentation information of the next level depth. In the case of using whether or not segmentation has occurred as the basis, a method can be used that prioritizes obtaining the additional segmentation information in the block segmented based on the initial block. Other additional scanning methods can also be considered. In this invention, the case of constructing the bitstream of segmentation information based on whether or not segmentation has occurred will be assumed.

[0141] As mentioned above, it can support one type of tree segmentation or multiple types of tree segmentation depending on the encoding / decoding settings.

[0142] When multiple tree-structured partitions are available, partitioning can be performed according to a pre-defined priority order. Specifically, partitioning is performed in priority order: if a previous partition has already been performed (i.e., the block is divided into two or more blocks), the partitioning ends at the corresponding partition, and the sub-blocks are re-encoded before proceeding to the next partition depth. If a previous partition has not been performed, the next partition can be performed. If not all tree-structured partitions are performed, the partitioning process within the corresponding block ends.

[0143] In the above example, when the segmentation flag corresponding to the priority order is true (segmentation o), it can further include additional segmentation information (segmentation direction flag, etc.) related to the corresponding segmentation method, while when it is false (segmentation x), it can be composed of segmentation information (segmentation flag, segmentation direction flag, etc.) of the segmentation method corresponding to the next order.

[0144] When performing a segmentation using at least one segmentation method in the above examples, the segmentation can be re-executed in the next sub-block (or the next segmentation depth) according to the above priority order. Alternatively, a segmentation corresponding to a priority order that was not segmented in the previous segmentation results can be excluded from the segmentation candidate group of the next sub-block.

[0145] For example, when supporting quadtree splitting as well as binary tree splitting, with quadtree splitting having a higher priority and performing quadtree splitting at a split depth of k, it is possible to start splitting from the quadtree again on a sub-block with a split depth of k+1 obtained by splitting using a quadtree.

[0146] When a binary tree split is performed instead of a quadtree split at a split depth k, the next sub-block at a split depth of k+1 can be configured to either allow the split to restart from the quadtree or exclude the quadtree from the candidate group and only allow binary tree splits. This can be determined based on the encoding / decoding settings; in this invention, the latter case will be used as an example.

[0147] Alternatively, when multiple tree-like partitions can be performed, additional selection information related to the partitioning method can be generated, and partitioning can be performed according to the selected partitioning method. In this example, partitioning information related to the selected partitioning method can be further included.

[0148] Figure 6 This is a schematic diagram illustrating a block used to explain the block division applicable to one embodiment of the present invention.

[0149] Next, we will explain the size and shape of blocks that can be obtained by using more than one segmentation method starting from a basic coded block. However, this is only an example and is not a limitation. Furthermore, the segmentation settings such as segmentation type, segmentation information, and the order in which the segmentation information is composed, which are explained later, are also only examples and are not limited to this.

[0150] See Figure 6 Thick solid lines represent basic coded blocks, while thick dashed lines represent quadtree partition boundaries, and double solid lines represent symmetric binary tree partition boundaries. Additionally, solid lines represent ternary tree partition boundaries, while thin dashed lines represent asymmetric ternary tree partition boundaries.

[0151] For ease of explanation, it will be assumed that the upper left block, upper right block, lower left block, and lower right block (e.g., 64×64 respectively) are divided into blocks based on the basic encoded block (e.g., 128×128).

[0152] Furthermore, assuming that a state of four sub-blocks (top left block, top right block, bottom left block, and bottom right block) is obtained by performing a splitting operation on the initial block (e.g., 128×128), the splitting depth has increased from 0 to 1. The splitting information related to quadtree splitting is the maximum block size of 128×128, the minimum block size of 8×8, and the maximum splitting depth of 4. The splitting settings related to quadtree splitting are applied to each block.

[0153] Furthermore, when each sub-block supports multiple tree-like segmentation methods, the size and shape of the obtainable blocks can be determined through multiple block segmentation settings. In this example, it is assumed that the maximum block size for binary tree segmentation and quadtree segmentation is 64×64, the minimum block size is 4 on one side, and the maximum segmentation depth is 4.

[0154] Based on the above assumptions, the top left block, top right block, bottom left block, and bottom right block will be explained separately below.

[0155] 1. Top left block (A1 to A6)

[0156] The top-left block 4M×4N illustrates the case where a single tree-like partitioning method is supported. The size and shape of the obtainable blocks can be determined by setting the maximum block size, minimum block size, and partitioning depth.

[0157] The top-left block represents the case where the block format is one that can be obtained during the splitting process. The splitting information required for the splitting action can be a splitting or no-split flag, and the obtainable candidates can be 4M×4N and 2M×2N. Specifically, when the splitting or no-split flag is 0, splitting can be skipped, while when it is 1, splitting can be performed.

[0158] See Figure 6 The top left block 4M×4N can be split into quadtrees at the current depth of 1 when the splitting information is 1. At this time, the depth will increase from 1 to 2 and blocks A0 to A3, A4, A5 and A6 (each 2M×2N) will be obtained.

[0159] When supporting the multi-tree splitting method described in the following examples, the information generated when only quadtree splitting is supported is the same as in this example, consisting of splitting / not splitting flags, so the detailed descriptions related to it will be omitted.

[0160] 2. Top right block (A7 to A11)

[0161] The upper right block 4M×4N illustrates the segmentation situation that supports multiple tree-like methods (quadtree segmentation and binary tree segmentation, with a segmentation priority order, quadtree segmentation → binary tree segmentation). The size and shape of the obtainable blocks can be determined by setting multiple block segmentation.

[0162] The upper right block is a block that can be obtained during the splitting process. There are multiple possible block types. The splitting information required when performing the splitting action can be a splitting status flag, a splitting type flag, a splitting form flag, and a splitting direction flag. The obtainable candidates can be 4M×4N, 4M×2N, 2M×4N, 4M×N / 4M×3N, 4M×3N / 4M×N, M×4N / 3M×4N, and 3M×4N / M×4N.

[0163] When an object block is included in the scope where both quadtree partitioning and binary tree partitioning are executable, partitioning information can be constructed based on the cases where both quadtree partitioning and binary tree partitioning are executable, and the cases where only one type of tree partitioning can be executed.

[0164] In the above explanation, it should be clarified that "executable range" refers to the case where the object block (64×63) can be divided into quadtree partitioning (128×128~8×8) and binary tree partitioning {64×64~(64>>m)×(64>>n), where m+n is 4} according to the block partitioning settings. "Executable case" refers to the case where, under the support of multiple partitioning methods, quadtree partitioning and binary tree partitioning can be executed according to the partitioning order and rules, or the case where only one type of tree partitioning can be executed.

[0165] (1) Cases where both quadtree partitioning and binary tree partitioning are feasible.

[0166] [Table 1]

[0167]

[0168]

[0169] In Table 1, 'a' can refer to a flag indicating whether a quadtree has been split or not, and 'b' can refer to a flag indicating whether a binary tree has been split or not. Furthermore, 'c' can refer to a flag indicating the direction of the split, 'd' can refer to a flag indicating the form of the split, and 'e' can refer to a flag indicating the proportion of the split.

[0170] Referring to Table 1, when 'a' in the splitting information is 1, it can mean performing a quadtree split (QT), while when it is 0, it can also include 'b'. When 'b' in the splitting information is 0, it can mean no split is performed in the corresponding block (No Split), while when it is 1, it can mean performing a binary tree split (BT) and can also include 'c' and 'd'. When 'c' in the splitting information is 0, it can mean a horizontal split (hor), while when it is 1, it can mean a vertical split (ver). When 'd' in the splitting information is 0, it can mean performing a symmetric binary tree (SBT) split, while when it is 1, it can mean performing an asymmetric binary tree (ABT) split and can also include 'e'. When 'e' in the splitting information is 0, it can mean performing a split with a 1:4 left-right or top-bottom block ratio, while when it is 1, it can mean performing a split with the opposite ratio (3:4).

[0171] (2) Cases where only binary tree partitioning can be performed

[0172] [Table 2]

[0173] b c d e No Split 0 Horizontal Symmetric Binary Tree Partitioning (SBT hor) 1 0 0 Horizontal Asymmetric Binary Tree Partition (ABT hor) 1 / 4 1 0 1 0 Horizontal Asymmetric Binary Tree Partitioning (ABT hor) 3 / 4 1 0 1 1 Vertically Symmetric Binary Tree Partition (SBT ver) 1 1 0 Vertical Asymmetric Binary Tree Partition (ABT ver) 1 / 4 1 1 1 0 Vertical Asymmetric Binary Tree Partition (ABT ver) 3 / 4 1 1 1 1

[0174] In Table 2, b to e can refer to the same symbols as b to e in Table 1.

[0175] Referring to Table 2, when only binary tree splitting is possible, everything is the same except for the flag (a) used to indicate whether a quadtree split is possible. In other words, when only binary tree splitting is possible, the splitting information can start from b.

[0176] See Figure 6 Block A7 in the upper right block is a case where a quadtree split could have been performed in the block before the split, but a binary tree split was performed instead. Therefore, the split information can be generated according to Table 1.

[0177] Conversely, blocks A8 to A11 are cases where a binary tree split was performed instead of a quadtree split in the blocks before the split (i.e., the upper right block). Therefore, a quadtree split cannot be performed again, and the split information can be generated according to Table 2.

[0178] 3. Bottom left block (A12 to A15)

[0179] The diagram in the lower left block 4M×4N illustrates the segmentation scenario that supports multiple tree-like methods (quadtree segmentation, binary tree segmentation, and ternary tree segmentation, with a segmentation priority order: quadtree segmentation → binary tree segmentation or ternary tree segmentation, and the segmentation method can be selected: binary tree segmentation or ternary tree segmentation). The size and shape of the obtainable blocks can be determined through multiple block segmentation settings.

[0180] The lower left block is a block that can be obtained during the splitting process. The size and shape of the block can be varied. The splitting information required when performing the splitting action can be a splitting flag, a splitting type flag, and a splitting direction flag. The available candidates can be 4M×4N, 4M×2N, 2M×4N, 4M×N / 4M×2N / 4M×N, and M×4N / 2M×4N / M×4N.

[0181] When an object block is included in the scope where quadtree partitioning, binary tree partitioning, and ternary tree partitioning are all executable, partitioning information can be constructed based on the cases where quadtree partitioning, binary tree partitioning, and ternary tree partitioning are all executable, as well as cases where none of the above are applicable.

[0182] (1) Cases where quadtree partitioning and binary / tritree partitioning can both be performed.

[0183] [Table 3]

[0184] a b c d Quadtree partitioning (QT) 1 No Split 0 0 Horizontal binary tree partitioning (BT hor) 0 1 0 0 Horizontal ternary tree partitioning (TT hor) 0 1 0 1 Vertical binary tree partitioning (BT ver) 0 1 1 0 Vertical ternary tree partitioning (TT ver) 0 1 1 1

[0185] In Table 3, 'a' can refer to a flag indicating whether a quadtree has been split or not, while 'b' can refer to a flag indicating whether a binary or ternary tree has been split or not. Furthermore, 'c' can refer to a flag indicating the direction of the split, and 'd' can refer to a flag indicating the type of split.

[0186] Referring to Table 3, when 'a' in the splitting information is 1, it can indicate a quadtree split (QT), while 'b' can also be included when it is 0. When 'b' in the splitting information is 0, it can indicate that no further splitting is performed in the corresponding block (No Split), while 'b' can indicate a binary tree split (BT) or a ternary tree split (TT), and can also include 'c' and 'd'. When 'c' in the splitting information is 0, it can indicate a horizontal split (hor), while 'c' can indicate a vertical split (ver). When 'd' in the splitting information is 0, it can indicate a binary tree split (BT), while 'd' can indicate a ternary tree split (TT).

[0187] (2) Cases where only binary tree partitioning / ternary tree partitioning can be performed.

[0188] [Table 4]

[0189] b c d No Split 0 Horizontal binary tree partitioning (BT hor) 1 0 0 Horizontal ternary tree partitioning (TT hor) 1 0 1 Vertical binary tree partitioning (BT ver) 1 1 0 Vertical ternary tree partitioning (TT ver) 1 1 1

[0190] In Table 4, b to d can refer to the same symbols as b to d in Table 3.

[0191] Referring to Table 4, when only binary tree splitting / ternary tree splitting can be performed, everything is the same except for the flag (a) used to indicate whether quadtree splitting is performed. In other words, when only binary tree splitting / ternary tree splitting can be performed, the splitting information can start from b.

[0192] See Figure 6 Blocks A12 and A15 in the lower left block are cases where quadtree splitting could have been performed in the lower left block before the split, but binary / ternary tree splitting was performed instead. Therefore, the splitting information can be generated according to Table 3.

[0193] Conversely, blocks A13 and A14 were partitioned using a ternary tree instead of a quadtree in the lower left block before the partitioning, and therefore the partitioning information can be generated according to Table 4.

[0194] 4. Bottom right block (A16 to A20)

[0195] The bottom right block (4M×4N) illustrates a segmentation scenario that supports multiple tree-like structures (similar to the bottom left block, but in this example, asymmetric binary tree segmentation is supported). This allows the size and shape of the obtainable blocks to be determined through multiple block segmentation settings.

[0196] The lower right block is a block that can be obtained during the splitting process. There are multiple possible block types. The splitting information required when performing the splitting action can be a splitting status flag, a splitting type flag, a splitting form flag, and a splitting direction flag. The available candidates can be 4M×4N, 4M×2N, 2M×4N, 4M×N / 4M×3N, 4M×3N / 4M×N, M×4N / 3M×4N, 3M×4N / M×4N, 4M×N / 4M×2N / 4M×N, and M×4N / 2M×4N / M×4N.

[0197] When an object block is included in the scope where quadtree partitioning, binary tree partitioning, and ternary tree partitioning are all executable, partitioning information can be constructed based on the cases where quadtree partitioning, binary tree partitioning, and ternary tree partitioning are all executable, as well as cases where none of the above are applicable.

[0198] (1) Cases where quadtree partitioning and binary / tritree partitioning can both be performed.

[0199] [Table 5]

[0200] a b c d e f Quadtree partitioning (QT) 1 No Split 0 0 Horizontal ternary tree partitioning (TT hor) 0 1 0 0 Horizontal Symmetric Binary Tree Partitioning (SBT hor) 0 1 0 1 0 Horizontal Asymmetric Binary Tree Partition (ABT hor) 1 / 4 0 1 0 1 1 0 Horizontal Asymmetric Binary Tree Partitioning (ABT hor) 3 / 4 0 1 0 1 1 1 Vertical ternary tree partitioning (TT ver) 0 1 1 0 Vertically Symmetric Binary Tree Partition (SBT ver) 0 1 1 1 0 Vertical Asymmetric Binary Tree Partition (ABT ver) 1 / 4 0 1 1 1 1 0 Vertical Asymmetric Binary Tree Partition (ABT ver) 3 / 4 0 1 1 1 1 1

[0201] In Table 5, 'a' can be a marker indicating whether a quadtree has been partitioned, while 'b' can be a marker indicating whether a binary or ternary tree has been partitioned. Furthermore, 'c' can be a marker indicating the direction of the partition, 'd' can be a marker indicating the type of partition, 'e' can be a marker indicating the form of the partition, and 'f' can be a marker indicating the proportion of the partition.

[0202] Referring to Table 5, when 'a' in the splitting information is 1, it can indicate a quadtree split (QT), while 'b' can also be included when it is 0. When 'b' in the splitting information is 0, it can indicate no splitting in the corresponding block (No Split), while 'b' can indicate a binary tree split (BT) or a ternary tree split (TT), and can also include 'c' and 'd'. When 'c' in the splitting information is 0, it can indicate a horizontal split (hor), while 'c' can indicate a vertical split (ver). When 'd' in the splitting information is 0, it can indicate a ternary tree split (TT), while 'd' can indicate a binary tree split (BT), and can also include 'e'. When 'e' in the splitting information is 0, it can indicate a symmetric binary tree split (S BT), while 'e' can indicate an asymmetric binary tree split (ABT), and can also include 'f'. When f in the segmentation information is 0, it can mean that the segmentation is performed with a left-right or top-bottom block ratio of 1:4, while when it is 1, it can mean that the segmentation is performed with the opposite ratio (3:4).

[0203] (2) Cases where only binary tree partitioning / ternary tree partitioning can be performed.

[0204] [Table 6]

[0205] b C d e f No Split 0 Horizontal ternary tree partitioning (TT hor) 1 0 0 Horizontal Symmetric Binary Tree Partitioning (SBT hor) 1 0 1 0 Horizontal Asymmetric Binary Tree Partition (ABT hor) 1 / 4 1 0 1 1 0 Horizontal Asymmetric Binary Tree Partitioning (ABT hor) 3 / 4 1 0 1 1 1 Vertical ternary tree partitioning (TT ver) 1 1 0 Vertically Symmetric Binary Tree Partition (SBT ver) 1 1 1 0 Vertical Asymmetric Binary Tree Partition (ABT ver) 1 / 4 1 1 1 1 0 Vertical Asymmetric Binary Tree Partition (ABT ver) 3 / 4 1 1 1 1 1

[0206] In Table 6, b to f can refer to the same symbols as b to f in Table 5.

[0207] Referring to Table 6, when only binary tree splitting / ternary tree splitting can be performed, everything is the same except for the flag (a) used to indicate whether quadtree splitting is performed. In other words, when only binary tree splitting / ternary tree splitting can be performed, the splitting information can start from b.

[0208] See Figure 6 Block A20 in the lower right block is a case where a quadtree split could have been performed in the block before the split, i.e., the lower right block, but instead a binary tree split or ternary tree split was performed. Therefore, the split information can be generated according to Table 5.

[0209] Conversely, blocks A16 to A19 are blocks that were split before the partitioning, i.e., the lower right block, where a binary tree partitioning was performed instead of a quadtree partitioning. Therefore, the partitioning information can be generated according to Table 6.

[0210] Figure 7 This is a sequence diagram illustrating an image encoding method applicable to one embodiment of the present invention.

[0211] See Figure 7 The image encoding apparatus according to one embodiment of the present invention can set the size information of the image and the size information of the basic encoding blocks (S710), and can also divide the image into basic encoding block units (S720). In addition, the image encoding apparatus can encode the basic encoding blocks according to the block scanning order (S730), and then send the encoded bit stream to the image decoding apparatus.

[0212] Furthermore, the image decoding apparatus according to one embodiment of the present invention can operate in a manner corresponding to the image encoding apparatus. In other words, the image decoding apparatus can reconstruct the image size information and the size information of the basic coded blocks based on the bitstream received from the image encoding apparatus, thereby dividing the image into basic coded block units. Furthermore, the image decoding apparatus can perform decoding on the basic coded blocks according to the block scanning order.

[0213] The basic encoding block size information can refer to the size of the largest block to be encoded, and the block scanning order can be a grid scan, but other scanning orders can also be used, so it is not limited to this.

[0214] As described above, the size information can be represented using powers of 2. The size of the largest block can be directly represented using powers of 2 or using the difference between the maximum block size and a pre-defined power (e.g., 8 or 16). For example, when the basic coded block size is between 8×8 and 128×128, the size information can be generated using the difference between the maximum block size and a power of 8. When the maximum block size is 64×64, the size information can be generated using log2(64)-log2(8)=3. This example describes the size information of the basic coded block, but these descriptions are also applicable when explicitly generating maximum and minimum size information based on various tree structure types in block partitioning settings.

[0215] However, when the image size is not an integer multiple of the basic coded block, encoding / decoding cannot be performed in the encoder / decoder as described above, and therefore various transformations may be required. Detailed methods related to this will be explained later.

[0216] Furthermore, the examples described below will assume that the image consists of a single stripe or parallel block. However, even when divided into multiple stripes or parallel blocks, the block located on the right or bottom boundary relative to the image may not be the size of the basic coding block. This can occur when the image is first divided into blocks and then further divided into parallel blocks or stripes (i.e., the block size at the boundary may not be an integer multiple).

[0217] Alternatively, other situations may occur when the image is first segmented into parallel blocks, stripes, etc., and then further divided into blocks. That is, because the sizes of the parallel blocks, stripes, etc., are initially set, there may be cases where the individual parallel blocks or stripes are not integer multiples of the basic coding block size. This means that, in addition to the right or bottom boundaries relative to the image, there may also be blocks in the middle of the image that are not the size of the basic coding block. Therefore, while this approach can be applied in the examples described later, it must be applied with the consideration that blocks may not be integer multiples of the parallel blocks and stripes.

[0218] Figure 8a as well as Figure 8b This is a schematic diagram illustrating an image divided into basic coded blocks of various sizes to which one embodiment of the present invention is applied.

[0219] Figure 8a This is a schematic diagram illustrating the division of an 832×576 image using basic coded blocks of size 64×64. (See also...) Figure 8a The horizontal length of the image (video) is an integer multiple of the horizontal length of the basic coded block, i.e., 13 times, while the vertical length of the image is an integer multiple of the vertical length of the basic coded block, i.e., 9 times. Therefore, the image can be divided into 117 basic coded blocks for encoding / decoding.

[0220] Figure 8b This is a schematic diagram illustrating the division of an 832×576 image using basic coded blocks of size 128×128. (See also...) Figure 8b The horizontal length of the image is a non-integer multiple of the horizontal length of the basic coding block, i.e., 6.5 times, while the vertical length of the image is a non-integer multiple of the vertical length of the basic coding block, i.e., 4.5 times.

[0221] In this invention, the outer boundary of an image can refer to the left, right, top, bottom, upper left, lower left, upper right, and lower right sides of the image, etc. However, under the assumption that block segmentation is usually performed based on the upper coordinate of the image, the description will focus on the right, bottom, and lower right boundaries.

[0222] See Figure 8bWhen dividing the region located at the outer boundary of the image (right, bottom, and lower right) into basic coding blocks, although the coordinates of the upper left side of the corresponding block are inside the image, the coordinates of the lower right side are outside the image. Moreover, depending on the position of the corresponding block (right boundary, bottom boundary, or lower right boundary), the coordinates of the upper right side may be outside the image (P_R), the coordinates of the lower left side may be outside the image (P_D), or both coordinates may be outside the image (P_DR).

[0223] At this point, the 24 basic coding blocks of size 128×128 located at the outer boundary of the image cannot be encoded / decoded, but the regions located at the outer boundary of the image (P_R, P_D, and P_DR) have not been divided into basic coding blocks, so other processing may be required.

[0224] To address the situation described above, various methods can be employed. For example, methods such as enlarging (or padding) the image before encoding / decoding, and methods that perform encoding / decoding according to specific cases of basic coded blocks, can be used. The method of enlarging the image before encoding / decoding can be an example of maintaining the uniformity of the image encoding / decoding structure, while the method of performing encoding / decoding according to specific cases of basic coded blocks can be an example of improving image encoding / decoding performance.

[0225] Next, as solutions to the problem of the image size not being an integer multiple of the basic coding block size as described above, the following will be explained in detail: 1. Basic coding block adjustment method, 2. Image size adjustment method, 3. Adaptive image data processing method, and 4. Adaptive image data processing method based on multiple segmentation methods.

[0226] 1. Basic Coding Block Adjustment Method

[0227] The basic coding block adjustment method is a method of borrowing at least one of the blocks supported by the block segmentation unit when the image size is not an integer multiple of the basic coding block size.

[0228] Figure 9 This is a first illustration for explaining a basic coding block adjustment method applicable to one embodiment of the present invention.

[0229] See Figure 9When the image size is 1856×928, the basic coding block size is 128×128, and single tree-like segmentation methods such as quadtrees are supported, with a minimum block size of 8×8, the basic coding block can borrow 64×64, 32×32, 16×16, and 8×8 blocks through the block segmentation unit. Specifically, when no segmentation is performed in the 128×128 block with a segmentation depth of 0, a 128×128 block can still be borrowed; however, since the image size is not an integer multiple of 128×128, it can be excluded from the candidate group.

[0230] Because the horizontal length of the image is 1856 and the horizontal length of the basic coding block is 128, a length of 64 can be left at the right boundary of the image. Since the size of the block located at the right boundary of the image can be 64×128, the basic coding block can borrow a 64×64 block from the candidate blocks supported by the block segmentation unit. The borrowed block can be the largest block in the candidate blocks based on 64×128 or the smallest block after segmentation. In the case of borrowing a 64×64 block, 64×128 can be composed of two 64×64 blocks.

[0231] Furthermore, since the vertical length of the image is 928 and the vertical length of the basic coding block is 128, a length of 32 can be left at the lower boundary of the image. At this time, since the size of the block located at the lower boundary of the image can be 128×32, the basic coding block can borrow a 32×32 block from the candidate blocks supported by the block segmentation unit. The borrowed block can be the largest block in the candidate blocks based on 128×32 or the smallest block after segmentation. In the case of borrowing a 32×32 block, 128×32 can be composed of four 32×32 blocks.

[0232] Furthermore, since the size of the block located on the lower right boundary of the image is 64×32, the basic coded block can borrow a 32×32 block from the candidate blocks supported by the block segmentation unit. The borrowed block can be the largest block among the candidate blocks based on 64×32 or the smallest block after segmentation. In the case of borrowing a 32×32 block, 64×32 can be composed of two 32×32 blocks.

[0233] In other words, a basic coding block (128×128) can be composed of two adjusted basic coding blocks (64×64) on the right edge of the image, four adjusted basic coding blocks (32×32) on the bottom edge of the image, and two adjusted basic coding blocks (32×32) on the bottom right edge of the image.

[0234] Figure 10 This is a second illustration used to explain a basic coding block adjustment method applicable to one embodiment of the present invention.

[0235] See Figure 10 In terms of image size and basic coding block size Figure 9 When the basic coding block supports multiple tree-like partitioning methods such as quadtree partitioning and binary tree partitioning, and the minimum block size is 8×8, the basic coding block can borrow square blocks such as 64×64, 32×32, 16×16, and 8×8, as well as rectangular blocks such as 128×64, 64×128, 64×32, 32×64, 32×16, 16×32, 16×8, and 8×16 through the block partitioning part. It is also possible to borrow blocks with a vertical ratio exceeding 2 through quadtree partitioning and binary tree partitioning, but for the sake of explanation, these are excluded from the candidate group, but this is not a limitation.

[0236] Because the horizontal length of the image is 1856 and the horizontal length of the basic coding block is 128, a length of 64 can be left at the right boundary of the image. Since the size of the block located at the right boundary of the image can be 64×128, the basic coding block can borrow a 64×128 block from the candidate blocks supported by the block segmentation unit. The borrowed block can be the largest block in the candidate blocks based on 64×128 or the smallest block after segmentation. In the case of borrowing a 64×128 block, 64×128 can be composed of one 64×128 block.

[0237] Furthermore, since the vertical length of the image is 928 and the vertical length of the basic coding block is 128, a length of 32 can be left at the lower boundary of the image. At this time, since the size of the block located at the lower boundary of the image can be 128×32, the basic coding block can borrow a 64×32 block from the candidate blocks supported by the block segmentation unit. The borrowed block can be the largest block in the candidate blocks based on 128×32 or the smallest block after segmentation. In the case of borrowing a 64×32 block, 128×32 can be composed of two 64×32 blocks.

[0238] Furthermore, since the size of the block located on the lower right boundary of the image is 64×32, the basic coded block can borrow a 64×32 block from the candidate blocks supported by the block segmentation unit. The borrowed block can be either the largest block in the candidate blocks based on 64×32 or the smallest block after segmentation. In the case of borrowing a 64×32 block, 64×32 can be composed of a single 64×32 block.

[0239] In other words, a basic coding block (128×128) can be composed of one adjusted basic coding block (64×128) at the right edge of the image, two adjusted basic coding blocks (64×32) at the bottom edge of the image, and two adjusted basic coding blocks (64×32) at the bottom right edge of the image.

[0240] See Figure 10 The borrowed block can be determined based on the remaining horizontal or vertical length while maintaining the horizontal or vertical length of the block (performing the first process). In other words, the right side of the image can be determined from blocks of size 64×128, 64×64, and 64×32 represented by the remaining horizontal length (64), while the bottom side of the image can be determined from blocks of size 64×32, 32×32, and 16×32 represented by the remaining vertical length (32), while the borrowed block (64×32) can be determined from blocks of size 64×32, 32×32, and 16×32 represented by the remaining vertical length (32).

[0241] However, when the block segmentation unit does not support a block corresponding to the remaining horizontal or vertical length, a borrowed block can be determined by executing the process more than twice. Detailed information related to this will be provided later. Figure 10 Please provide an explanation.

[0242] Figure 11a as well as Figure 11b This is a schematic diagram illustrating a boundary block used to explain a basic coding block adjustment method applicable to one embodiment of the present invention.

[0243] See Figure 11a Assume the size of the basic coding block is 128×128, and because the image size is not an integer multiple of the basic coding block size, there is a 128×80 block at the lower boundary of the image. In this case, the 128×80 block can be called a boundary block.

[0244] The basic coded block can borrow one of the blocks of size 128×64, 64×64 and 32×64 that are approximately represented by 128×80 from the candidate blocks supported by the block segmentation unit, namely a block of size 128×64, which can leave a block of size 128×16.

[0245] At this point, the basic coded block can borrow one of the blocks of size 32×16, 16×16, and 8×16 that are approximately represented by 128×16 from the candidate blocks supported by the block segmentation unit, namely a block of size 32×16.

[0246] See Figure 11b Assume the size of the basic coding block is 128×128, and because the image size is not an integer multiple of the basic coding block size, there is a 56×128 block on the right edge of the image.

[0247] The basic coded block can borrow one of the blocks of size 32×64, 32×32 and 32×16 that are approximately represented by 56×128 from the candidate blocks supported by the block segmentation unit, namely a block of size 32×64, which can leave a block of size 24×128.

[0248] At this point, the basic coded block can borrow one of the blocks of size 16×32, 16×16, and 16×8 that are approximately represented by 24×128 from the candidate blocks supported by the block segmentation unit, namely a block of size 16×32. At this point, there can be a block of size 8×128 remaining.

[0249] The basic coded block can ultimately borrow one of the blocks that is approximately represented by an 8×128 block, i.e., an 8×16 block, from the candidate blocks supported by the block segmentation unit.

[0250] In other words, in such Figure 11a In the case shown, a coding block (128×128) can be composed of one basic coding block (128×64) after the first adjustment and four basic coding blocks (32×16) after the second adjustment, while in the case shown... Figure 11b In the case shown, it can be composed of 2 basic coding blocks (32×64) after the first adjustment, 4 basic coding blocks (16×32) after the second adjustment, and 8 basic coding blocks (8×16) after the third adjustment.

[0251] Figure 12 This is a sequence diagram illustrating a method for performing image encoding by adjusting basic coding blocks according to one embodiment of the present invention.

[0252] See Figure 12An image encoding apparatus for adjusting basic coding blocks according to one embodiment of the present invention can set image size information, basic coding block size information, and block segmentation information (S1210), and then divide the image into basic coding block units redefined based on the basic coding blocks or block segmentation information according to the block positions (S1220). Furthermore, the image encoding apparatus can encode the basic coding blocks according to the redefined block scanning order (S1230), and then send the encoded bitstream to an image decoding apparatus.

[0253] Furthermore, the image decoding apparatus to which one embodiment of the present invention is applicable can operate in a manner corresponding to the image encoding apparatus. In other words, the image decoding apparatus can reconstruct the image size information, the basic coding block size information, and the block segmentation information based on the bitstream received from the image encoding apparatus, thereby dividing the image into basic coding block units that are redefined based on the basic coding blocks or the block segmentation information.

[0254] In addition, the image decoding device can perform decoding on the basic coded blocks according to the reset block scanning order.

[0255] Block segmentation information refers to segmentation information related to blocks that can be obtained between the maximum and minimum block sizes. Specifically, it can refer to explicit generation information (e.g., based on color components) related to block segmentation settings that affect the acquisition of blocks of various sizes and shapes through the block segmentation unit. <y cb cr>Image type It can include information such as the supported tree structure, the maximum and minimum size of different tree structures, the maximum segmentation depth, and block shape constraints (horizontal / vertical ratio information, etc.), as well as implicit information such as segmentation order and segmentation scale that are predetermined in the encoder / decoder.

[0256] Furthermore, when the image is not an integer multiple of the basic coding block size, since it is impossible to divide the image into basic coding block sizes on the right or lower boundary, the basic coding blocks of the supported block size can be reset according to the block segmentation settings supported by the block segmentation unit before further segmentation. In other words, a basic coding block located on the image boundary can be composed of at least one reset basic coding block and divided accordingly.

[0257] Furthermore, because the boundaries of an image can be composed of more than one basic coding block, certain scanning sequences, such as z-scan, can be applied to the basic coding blocks located at the image boundaries. For example, in Figure 11a The scanning sequence can first scan the 128×64 block, and then scan the four 32×16 blocks located on the bottom side from left to right.

[0258] 2. Image size adjustment method

[0259] Image resizing is a method of enlarging the size of an image when the image size is not an integer multiple of the basic coding block size.

[0260] Figures 13a to 13b This is a first example diagram used to illustrate an image size adjustment method applicable to one embodiment of the present invention.

[0261] See Figure 13a When the image size is 832×576 and the basic coding block size is 128×128, a length of 64 can be left at the right and bottom edges of the image. In the case described above, the image size can be adjusted by increasing the length by 64 in both the horizontal and vertical directions.

[0262] See Figure 13b The enlarged image can be 896×640 pixels in size, and can be divided into 35 basic coded blocks of size 128×128.

[0263] Next, the image resizing method will be explained in more detail. When performing image magnification, the image size can be increased to the minimum of integer multiples of the original image's basic coded block size. See [link / reference]. Figure 11a Because the horizontal length of the image is 832, although lengths such as 896 and 1024 exist that are larger than the horizontal length of the image and are integer multiples of the basic coding block size, 896 is the smallest among these lengths, so the horizontal length of the image can be enlarged to 896. Similarly, because the vertical length of the image is 576, although lengths such as 640 and 768 exist that are larger than the vertical length of the image and are integer multiples of the basic coding block size, 640 is the smallest among these lengths, so the vertical length of the image can be enlarged to 896. However, the image enlargement standard can also use other settings and is not limited to these.

[0264] The magnified area can be filled with pre-set pixel values ​​or with pixel values ​​located on the boundary of the magnified area.

[0265] As an example, when filling the magnified area with a preset pixel value, the preset pixel value can be a value that is included in the range of pixel values ​​determined according to the bit depth or a value that is included in the range of pixel values ​​constituting the actual image.

[0266] The range of pixel values ​​determined by the bit depth can be from 0 to a maximum value of 2. n -1 (n is the number of bits). For example, when using the middle value within the pixel value range determined according to the bit depth as the preset pixel value and the number of bits is 10, the preset pixel value can be 512 in the range of 0 to 1023. Furthermore, the middle value within the pixel value range constituting the actual image can be used as the preset pixel value. For example, when the pixel value range constituting the actual image is from the minimum value x to the maximum value y, the preset pixel value can be the middle value in the range of x to y. The method of selecting the preset pixel value as described above is only one example and is not limited thereto.

[0267] As another embodiment, when filling the enlarged area using pixel values ​​located on the region boundaries, pixels located on the right and bottom boundaries of the image can be used. In other words, an area enlarged from the right boundary of the image can be filled by copying pixels located on the right boundary of the image horizontally, while an area enlarged from the bottom boundary of the image can be filled by copying pixels located on the bottom boundary of the image vertically. An area enlarged from the bottom right boundary of the image can be filled by copying pixels located on the bottom right boundary of the image diagonally.

[0268] Furthermore, in some images, such as 360-degree images, data for magnified regions can be obtained from regions that are continuous (or correlated) in 3D space {e.g., a sphere}. For example, when filling the right-hand magnified region of an object image in a projection format such as equirectangular projection, the magnified region can be filled using pixel values ​​located on the left side of the image. This example belongs to the case where the image contains data that is actually correlated, although located at the image boundary (i.e., within 2D space), due to the setting of configuring a 3D image in 2D space.

[0269] In the data filling method, except for the case of 360-degree images, all of them can be classified as cases where the image is magnified by filling the magnified area with unnecessary data before encoding / decoding.

[0270] Figure 14 This is a second illustration used to explain an image size adjustment method applicable to one embodiment of the present invention.

[0271] The magnified area of ​​an image can be determined based on the size of the image and the size of the basic coding block. The proportion of the magnified area from the basic coding block located on the image boundary can be less than or greater than the proportion of the area containing the actual image in the basic coding block.

[0272] See Figure 14 When the size of the basic coding block is 64×64 and the image size is not an integer multiple of the basic coding block size, the image can be magnified by 10 horizontally from the right boundary and by 52 vertically from the bottom boundary. In other words, the area magnified from the right boundary can be smaller than the area of ​​the original image contained in the corresponding basic coding block, while the area magnified from the bottom boundary can be larger than the area of ​​the original image contained in the corresponding basic coding block.

[0273] That is, the magnified area of ​​the basic coding block including the right boundary of the image is relatively small, so its encoding / decoding performance degradation can be relatively small. However, the magnified area of ​​the basic coding block including the bottom boundary of the image is relatively large, so its encoding / decoding performance degradation can be relatively large.

[0274] To solve the problem described above, one can use a method that borrows one of the blocks supported by the block segmentation unit, or a method that minimizes the magnified data. The method of borrowing one of the blocks supported by the block segmentation unit has already been described above, so the method of minimizing the magnified data will be described next.

[0275] As an example, it is assumed that the basic encoded block supports a single tree-like segmentation such as a quadtree, with a minimum block size of 8×8. The block segmentation unit can utilize square blocks of sizes 64×64, 32×32, 16×16, and 8×8. Furthermore, it is assumed that the boundary portion of the image can be enlarged according to one of the aforementioned blocks based on the encoding / decoding settings, thereby allowing the basic encoded block to be (re)set or (readjusted). The encoding / decoding settings can be the block with the smallest size difference among the remaining blocks, or blocks composed of the fewest possible elements, but are not limited to these.

[0276] See Figure 14 A length of 54 can be left at the right boundary of the image, so the size of the block located at the right boundary can be 54×64. The block located at the right boundary can be enlarged by 10 from the right boundary to utilize the 64×64 block with the smallest size difference among the candidate blocks supported by the block segmentation unit. Therefore, the enlarged region (64×64) can be composed of one adjusted basic coded block (64×64). However, since the size of the adjusted coded block is the same as the size of the original basic coded block, it is similar to enlarging the image size to an integer multiple of the basic coded block size.

[0277] A length of 12 is allowed at the bottom edge of the image, so the block size on the right edge can be 64×12. The block on the bottom edge can be enlarged by 4 from the bottom edge to utilize the 16×16 block with the smallest size difference among the candidate blocks supported by the block segmentation unit. Therefore, the enlarged area (64×16) can be composed of four adjusted basic coded blocks (16×16).

[0278] As another embodiment, it is assumed that the basic coding block supports multiple tree-like partitioning such as quadtree partitioning and binary tree partitioning, and that the block partitioning part can not only borrow square blocks of sizes such as 64×64, 32×32, 16×16 and 8×8, but also rectangular blocks such as 64×32, 32×64, 64×16, 16×64, 64×8, 8×64, 32×16, 16×32, 32×8, 8×32, 16×8 and 8×16.

[0279] A length of 54 can be left at the right edge of the image, so the size of the block located at the right edge can be 54×64. At this time, the block located at the right edge can be enlarged by 10 at the right edge in order to borrow the 64×64 block with the smallest difference among the candidate blocks (blocks of size 64×64, 64×32, 64×16 and 64×8) with a horizontal length of 64 that has the smallest difference from its own horizontal length.

[0280] A length of 12 can be left at the bottom edge of the image, so the block size on the right edge can be 64×12. In this case, the block on the bottom edge can be enlarged by 4 in length at the bottom edge to utilize the smallest difference among candidate blocks (64×16, 32×16, 16×16, and 8×16 blocks) with a vertical length of 16 that has a relatively small difference from its own. Therefore, the enlarged area (64×16) can be composed of an adjusted basic coded block (64×16).

[0281] In other words, an image can be enlarged by borrowing the candidate block with the smallest difference based on the remaining length, according to the difference between the image size and an integer multiple of the basic coding block. The basic coding block can be (re)set or (read)adjusted according to the size of the candidate block with a small size difference.

[0282] The image size adjustment method is similar to the basic coding block adjustment method in that it can adjust the basic coding block, but there are differences in that the image size can be enlarged to avoid adjusting the basic coding block, and even if the basic coding block is adjusted, only one adjustment needs to be performed.

[0283] Furthermore, the basic coded block adjustment method only needs to perform encoding / decoding on the original image data and does not require generating information related to the basic coded blocks of the image (such as the size or segmentation information of the readjusted blocks). In other words, it does not require performing encoding / decoding processes on unnecessary data. However, the basic coded block adjustment method constructs variable basic coded blocks on the image boundaries, which can lead to the problem of disrupting the uniformity of the encoding / decoding structure due to special processing.

[0284] In contrast, image resizing methods can maintain the uniformity of the encoding / decoding structure by enlarging the image to an integer multiple of the basic coding block, but they suffer from a degradation in encoding performance due to the need to encode / decode unnecessary data generated during the enlargement process.

[0285] Figure 15 This is a sequence diagram illustrating a method for performing image encoding by adjusting the image size, applicable to one embodiment of the present invention.

[0286] See Figure 15 An image encoding apparatus for adjusting image size according to one embodiment of the present invention can set image size information, basic encoding block size information, and block segmentation information (S1510), and then enlarge the image size according to the basic encoding blocks redefined based on an integer multiple of the basic encoding block size or the block segmentation information (S1520). Furthermore, the image encoding apparatus can divide the enlarged image according to block positions into basic encoding blocks or redefined basic encoding block units (S1530), then encode the basic encoding blocks according to the block scanning order (S1540), and then send the encoded bitstream to an image decoding apparatus.

[0287] Furthermore, the image decoding apparatus according to one embodiment of the present invention can operate in a manner corresponding to the image encoding apparatus. In other words, the image decoding apparatus can reconstruct the image size information, basic coding block size information, and block segmentation information based on the bitstream received from the image encoding apparatus, and then enlarge the image size according to basic coding blocks redefined based on integer multiples of the basic coding block size or the block segmentation information. Furthermore, the image decoding apparatus can divide the enlarged image according to block positions into basic coding blocks or redefined basic coding block units, and then perform decoding on the basic coding blocks according to the block scanning order.

[0288] The image magnification can be performed in a direction that is an integer multiple of the basic coded block, such as to the right or bottom of the image, or it can be performed according to the block size supported by the block segmentation setting supported by the block segmentation unit.

[0289] Furthermore, the scanning order can be reset according to the number of basic coded blocks added, as described above. However, even when the basic coded blocks are reset, they still belong to the same block, so their scanning order can remain unchanged.

[0290] 3. Adaptive Image Data Processing Methods

[0291] Adaptive image data processing is a method that enlarges the image size when the image size is not an integer multiple of the basic coding block size, and adaptively processes the image encoding / decoding data of the enlarged area based on the original image area. In other words, it is a method that performs explicit processing on image encoding / decoding data based on the original image area and implicit processing on image encoding / decoding data based on the enlarged area.

[0292] Figures 16a to 16d This is an illustrative diagram used to explain an adaptive image data processing method applicable to one embodiment of the present invention.

[0293] See Figure 16a Assuming the image size is 392×272 and the basic coding block size is 64×64, there is room for a horizontal length of 8 on the right edge and a vertical length of 16 on the bottom edge. Therefore, the image can be enlarged horizontally by 56 and vertically by 48. The enlarged image size can be 448×320, which is an integer multiple of the basic coding block size.

[0294] Next, we will explain the explicit and implicit processing methods for the segmentation information generated in order to obtain the optimal segmentation form from the basic coded blocks according to the segmentation settings.

[0295] For clarity, the processes described below will be marked with parentheses containing sequence numbers within the blocks. The number following the "-" in the parentheses indicates the block position where the corresponding process is executed. Specifically, when four blocks are obtained through quadtree partitioning, the numbers representing the block positions can be indices (1, 2, 3, and 4) assigned according to the z-scan order (top left, top right, bottom left, and bottom right). When two blocks are obtained through binary tree partitioning, the indices can be assigned according to either top-bottom or left-right order (1 and 2). For example, when four blocks are obtained through quadtree partitioning, (2-1) could refer to the second process, specifically the execution of the top-left block out of the four blocks.

[0296] Referring to Figure 16, the actual image can be the diagonal portion on the upper left side of the block (8×16 size), while the rest of the 64×64 block, excluding the diagonal portion, is an enlarged area, which can refer to the area filled when using expanded pixels. In this case, the basic coded block can support single tree-like segmentation such as a quadtree. Assuming the minimum block size is 8×8, it is assumed that blocks of 64×64, 32×32, 16×16, and 8×8 sizes can be obtained through the block segmentation unit.

[0297] The actual image can be composed of two 8×8 blocks from the blocks supported by the block segmentation unit, and the associated block segmentation information can be adaptively constructed through the process described below.

[0298] The first process is used to obtain a candidate block for segmentation (1) of size 64×64. The obtainable blocks can be 64×64 or 32×32. Since segmentation must be performed to obtain blocks containing only actual image data, 64×64 blocks containing padding data can be excluded from the candidates. Therefore, the candidates can only be 32×32 blocks. Segmentation information can be represented using 1 or 0 depending on whether segmentation is performed, but since there is only one 32×32 block as a candidate, segmentation must be performed. Therefore, the segmentation information can be processed implicitly rather than explicitly based on set values.

[0299] The second process is used to obtain candidate segments for the top-left (2-1), top-right (2-2), bottom-left (2-3), and bottom-right (2-4) blocks, each with a size of 32×32. Within each block, the obtainable blocks can be either 32×32 or 16×16. For the same reason as in the first process, the top-left block (2-1) can only be a 16×16 block as a candidate. The top-right (2-2), bottom-left (2-3), and bottom-right (2-4) blocks, since they only include filling areas and do not require additional segmentation, can only be 32×32 blocks as candidates. That is, because each block has only one candidate, the segmentation information can be implicitly processed.

[0300] The third process is used to obtain candidate segments for the top-left (3-1), top-right (3-2), bottom-left (3-3), and bottom-right (3-4) blocks, each with a size of 16×16. Within each block, the obtainable blocks can be either 16×16 or 8×8. For the same reason as the top-left (2-1) blocks in the first and second processes, the top-left (3-1) block can only be an 8×8 block as a candidate. Similarly, for the same reason as the top-right (3-2), bottom-left (3-3), and bottom-right (3-4) blocks in the second process, the top-right (2-2), bottom-left (2-3), and bottom-right (2-4) blocks can only be 16×16 blocks as candidates. That is, since each block has only one candidate, the segmentation information can be implicitly processed.

[0301] The fourth process is used to obtain four 8×8 blocks by segmenting the upper left block (3-1) from the third process. The block size containing the actual image data is 8×16, which is larger than the 8×8 block obtained by segmenting the upper left block (3-1), but since it can be composed of the upper left and lower left 8×8 blocks, no additional segmentation is required.

[0302] The segmentation information generated by processes 1 to 4 as described above can be represented as (8×8): 1-1-1-0-0-0-0-0-0, which, in sequence, are information related to the top-left block in process 1, the top-left block in process 2, the top-right block in process 3, the bottom-left block in process 3, the bottom-right block in process 3, the top-right block in process 2, the bottom-left block in process 2, and the bottom-right block in process 2. The 1s and 0s in the segmentation information can be syntax elements used to indicate whether a quadtree segmentation has occurred.

[0303] However, the segmentation information generated through processes 1 to 4 as described above can all be implicitly processed. That is, there is no explicit processing when the encoded / decoded data is adaptively processed based on the actual image data, so no segmentation information exists. This allows for the confirmation of the optimal segmentation pattern even without the generation and reconstruction of segmentation information.

[0304] See Figure 16c The actual image can be the diagonal portion on the left side of the block (64×16 size), while the rest of the 64×64 block, excluding the diagonal portion, is an enlarged area, referring to the area filled using expanded pixels. In this case, the basic coded block can support single tree-like segmentation such as a binary tree. Assuming the minimum block has a side length of 8 and a maximum segmentation depth of 3, it is assumed that blocks of 64×64, 64×32, 32×64, 64×16, 16×64, 32×32, 64×8, 8×64, 32×16, and 16×32 sizes can be obtained through the block segmentation unit. The block segmentation information can be adaptively constructed through the process described below.

[0305] The first process is used to obtain candidates for segmentation of a 64×64 block (1). The segmentation depth can increase from 0 to 1, and the obtainable blocks can be 64×64, 64×32, or 32×64 in size. Since segmentation must be performed to obtain blocks containing only actual image data, 64×64 blocks containing padding data can be excluded from the candidates. Furthermore, since 32×64 blocks require additional segmentation, such as horizontal segmentation, in subsequent steps to obtain blocks containing only actual image data, they can be excluded from the candidates. In other words, 32×64 blocks can be excluded from the candidates because they would increase the number of segmentations. That is, the candidate can only be a 64×32 block, and since there is only one candidate for a block, the segmentation information can be implicitly processed.

[0306] The second process is used to obtain candidate segments for each of the 64×32 upper block (2-1) and lower block (2-2). The segmentation depth can increase from 1 to 2. Within each block, the obtainable blocks can be 64×32, 64×16, and 32×32 in size. For the same reason as in the first process, the upper block (2-1) can exclude 64×32 and 32×32 blocks from the candidates, therefore only a 64×16 block can be selected as a candidate. The lower block (2-2), since it only includes the filling region and does not require additional segmentation, can only select a 64×32 block as a candidate. That is, because each block has only one candidate, the segmentation information can be implicitly processed.

[0307] The third process is used to obtain candidate segments for each of the 64×16 upper block (3-1) and lower block (3-2). The segmentation depth can be increased from 2 to 3. The blocks that can be obtained in each block can be 64×16, 64×8, and 32×16 in size. Since the upper block (3-1) has already obtained a block that only includes actual image data, one of the 64×16, 64×8, and 32×16 blocks can be determined as the optimal segmentation shape. However, since the lower block (3-2) only includes the filled area and does not need to perform additional segmentation, only the 64×16 block can be selected as a candidate.

[0308] The fourth process is used to obtain candidates for the split of the upper block (3-1) in the third process. However, since the optimal split shape has been determined in the upper block (3-1) in the third process and the length of one side and the split depth have reached the minimum block condition, no additional split is required.

[0309] The segmentation information generated by the first to fourth processes described above can be represented as follows: (64×16): 1-0-1-0-0-0-0; (64×8): 1-0-1-0-1-0-0-0; and (32×16): 1-0-1-0-1-1-0-0. The segmentation information, in sequence, relates to the upper block in the first process, the upper block in the second process, the lower block in the third process, and the lower block in the second process.

[0310] Specifically, the segmentation information can include two syntax elements related to whether or not the segmentation is performed in the first process and the segmentation direction, two syntax elements related to whether or not the upper block is performed in the second process and the segmentation direction, two syntax elements related to whether or not the upper block is performed in the third process and the segmentation direction, one syntax element related to whether or not the lower block is performed in the third process, and one syntax element related to whether or not the lower block is performed in the second process. When the syntax element related to whether or not the segmentation is performed is 0, it can represent no segmentation, while when it is 1, it can represent the segmentation is performed. When the syntax element related to the segmentation direction is performed is 0, it can represent the horizontal direction, while when it is 1, it can represent the vertical direction.

[0311] However, in the segmentation information generated by processes 1 to 4 as described above, the upper block in process 3 can be explicitly processed, while the rest can be implicitly processed. That is, the segmentation information in which the encoded / decoded data is adaptively processed based on the actual image data can be represented as (64×16):0 when a 64×16 block is determined as the optimal segmentation form, as (64×8):1-0 when a 64×8 block is determined as the optimal segmentation form, and as (32×16):1-1 when a 32×16 block is determined as the optimal segmentation form.

[0312] In other words, the optimal segmentation form of the basic coded block can be confirmed by the segmentation information generated in the upper block of the third process of explicit processing, excluding the case of implicit processing. Although the above describes the case where a portion of the segmentation information is implicitly processed, according to the block segmentation setting, all segmentation information can be implicitly processed when 64×16 meets the minimum block condition.

[0313] See Figure 16d The actual image can be the diagonal section at the top of the block (8×64 size), while the rest of the 64×64 block, excluding the diagonal section, is an enlarged area, which can refer to the area filled using expanded pixels. In this case, the basic coded block can support multiple tree-like segmentation methods such as quadtree segmentation and binary tree segmentation. In the quadtree segmentation setting, the maximum block size can be the same as the basic coded block size (64×64), and the minimum block size can be 16×16. In the binary tree segmentation setting, the maximum block size can be 32×32, the minimum block's side length can be 4, and the maximum segmentation depth can be 3. Furthermore, when the occurrence ranges of various tree-like partitions overlap, they can be executed according to the priority order of the tree-like partitions (e.g., quadtree partitioning takes precedence over binary tree partitioning). It is assumed that the block partitioning unit can obtain blocks of sizes 64×64, 32×32, 32×16, 16×32, 32×8, 8×32, 16×16, 32×4, 4×32, 16×8, and 8×16. The block partitioning information can be adaptively constructed through the process described below.

[0314] The first process is used to obtain a candidate block (1) of size 64×64 for segmentation. The obtainable blocks can be 64×64 or 32×32. Since segmentation must be performed to obtain blocks containing only actual image data, 64×64 blocks containing padding data can be excluded from the candidates. In the existing case, segmentation information such as whether quadtree segmentation is performed, whether binary tree segmentation is performed, and the direction of binary tree segmentation can be generated. As an example, since only one candidate (32×32 block) of quadtree segmentation is supported, the segmentation information can be implicitly processed.

[0315] The second process is used to obtain candidate partitions for the top-left (2-1), top-right (2-2), bottom-left (2-3), and bottom-right (2-4) blocks of size 32×32. Specifically, when performing quadtree partitioning, the obtainable blocks in each partition can be 32×32 or 16×16. When performing binary tree partitioning, the partitioning depth increases from 0 to 1, and the obtainable blocks in each partition can be 32×16 or 16×32.

[0316] In the upper left block (2-1) and lower left block (2-3), the 32×32 block in the candidate blocks must be segmented and thus eliminated from the candidate group. Furthermore, the 32×16 and 16×16 blocks, in order to obtain the block containing actual image data, require additional segmentation (such as vertical segmentation) in subsequent steps, increasing the number of segmentations. Therefore, they can be eliminated from the candidates. The block containing actual image data could be 8×64, but because quadtree segmentation was performed in step 1, 8×32 can be considered the size of the block containing actual image data. Therefore, each block's candidate can only include one 16×32 block. Although the 16×32 block also requires additional segmentation in subsequent steps, it requires fewer segmentations to reach the block containing actual image data compared to other candidate blocks, thus making it the best candidate.

[0317] The segmentation information used to obtain the best candidate (a 16×32 block) should not be performed in the quadtree segmentation with higher priority in the multi-tree segmentation, but rather in the binary tree segmentation with lower priority, with the segmentation direction being vertical. In this case, the segmentation information in the binary tree segmentation includes whether to segment and the segmentation direction information. However, when quadtree and binary tree segments overlap, because the quadtree segmentation with higher priority includes 32×32 blocks as a candidate, the binary tree segmentation may not require support for 32×32 block candidates. In other words, when quadtree and binary tree segments overlap, the segmentation information can omit the segmentation status information of the binary tree segmentation and only include the segmentation direction information. That is, because there is only one best candidate, the segmentation information can be implicitly processed.

[0318] Since the upper right block (2-2) and the lower right block (2-4) only include filling areas and do not require additional segmentation, only a 32×32 block can be selected as a candidate. That is, because there is only one candidate, the segmentation information can be implicitly processed.

[0319] The third process is used to obtain candidate segments for each of the following: the left block (3-1-1) and right block (3-1-2) of size 16×32 obtained from the top-left block (2-1), and the left block (3-2-1) and right block (3-2-2) of size 16×32 obtained from the bottom-left block (2-3). At this point, the segmentation depth can be increased from 1 to 2, and the blocks that can be obtained in the segment can be 16×32, 16×16, and 8×32.

[0320] For the same reason as the upper left block (2-1) in the second process, the 16×32 and 16×16 blocks can be excluded from the candidate group, and only the 8×32 block is selected as the candidate. In the current case, the segmentation information includes whether a binary tree segmentation is performed and the segmentation direction information, and a vertical 2-segmentation must be performed; therefore, the segmentation information can be implicitly processed. The 8×32 block obtained through the process described above can be a block containing only the actual image data.

[0321] Since the right-hand block (23) only includes the filling area and does not require additional segmentation, only a 16×32 block can be selected as a candidate. That is, since there is only one candidate, the segmentation information can be implicitly processed.

[0322] The left block (3-2-1) can perform implicit processing of segmentation information through the same process as the left block (3-1-1), and the obtained 8×32 block can be a block that only includes actual image data.

[0323] Since the right-hand block (3-2-2) only includes the filling area and does not require additional segmentation, only a 16×32 block can be selected as a candidate. That is, because there is only one candidate, the segmentation information can be implicitly processed.

[0324] The fourth process is used to obtain candidate segments for each of the 8×32 left block (4-1-1) and right block (4-1-2) obtained from the left block (3-1-1), and the 8×32 left block (4-2-1) and right block (4-2-2) obtained from the left block (3-2-1). The segmentation depth can be increased from 2 to 3. The blocks obtainable within a block can be 8×32, 8×16, or 4×32 in size. Since the left blocks (4-1-1) and (4-2-1) have already been obtained as blocks containing only actual image data, one of the 8×32, 8×16, or 4×32 blocks can be determined as the optimal segmentation form. However, since the right blocks (4-1-2) and (4-2-2) only contain filled areas and do not require additional segmentation, only 8×32 blocks can be selected as candidates. The segmentation information of the left block (4-1-1) and the left block (4-2-1) can be explicitly processed according to the determined optimal segmentation form, while the right block (4-1-2) and the right block (4-2-2) can be implicitly processed because there is only one candidate.

[0325] The segmentation information generated by the first to fourth processes described above can be represented as follows: (8×32) when an 8×32 block is determined as the optimal segmentation form: 1-0-1-1-1-0-0-0-0-0-1-1-1-0-0-0-0; (8×16) when an 8×16 block is determined as the optimal segmentation form: 1-0-1-1-1-1-0-0-0-0-0-1-1-1-0-0-0-0; and (4×32) when a 4×32 block is determined as the optimal segmentation form: 1-0-1-1-1-1-1-0-0-0-0-1-1-1-0-0-0-0. The segmentation information, in sequence, can be related to the following: the top left block (2-1) in the first process, the left block (3-1-1) in the second process, the left block (4-1-1) in the third process, the right block (4-1-2) in the fourth process, the right block (3-1-2) in the third process, the top right block (2-2) in the second process, the bottom left block (2-3) in the second process, the left block (3-2-1) in the third process, the left block (4-2-1) in the fourth process, the right block (4-2-2) in the fourth process, the right block (3-2-2) in the third process, and the bottom right block (2-4) in the second process.

[0326] Specifically, the segmentation information can include one syntax element related to whether or not a segment is performed in the first process; two syntax elements related to whether or not the upper left block (2-1) is segmented and the segmentation direction in the second process; two syntax elements related to whether or not the left block (3-1-1) is segmented and the segmentation direction in the third process; one syntax element related to whether or not the left block (4-1-1) is segmented in the fourth process; one syntax element related to whether or not the left block (4-1-2) is segmented in the fourth process; one syntax element related to whether or not the right block (3-1-2) is segmented in the third process; one syntax element related to whether or not the upper right block (2-2) is segmented in the second process; and one syntax element related to whether or not the lower right block (2-3) is segmented in the second process. The syntax elements are: 2 related to the splitting direction, 1 related to the splitting of the left block (3-2-1) in the third process, 1 related to the splitting of the left block (4-2-1) in the fourth process, 1 related to the splitting of the left block (4-2-2) in the fourth process, 1 related to the splitting of the right block (3-2-2) in the third process, and 1 related to the splitting of the lower right block (2-4) in the second process. When the syntax element related to splitting is 0, it means that no splitting is performed, and when it is 1, it means that splitting is performed. When the syntax element related to splitting direction is 0, it means that the horizontal direction is performed, and when it is 1, it means that the vertical direction is performed.

[0327] However, in the segmentation information generated by processes 1 to 4 as described above, the left-hand blocks (4-1-1 and 4-2-1) in process 4 can be explicitly processed, while the rest can be implicitly processed. That is, the segmentation information in which the encoded / decoded data is adaptively processed based on the actual image data can be represented as (8×32): 0-0 when an 8×32 block is determined as the optimal segmentation form, as (8×16): 1-0-0 when an 8×16 block is determined as the optimal segmentation form, and as (4×32): 1-1-1 when a 4×32 block is determined as the optimal segmentation form.

[0328] In other words, the optimal segmentation form of the basic coded block can be confirmed by the segmentation information generated from the left block (4-1-1 and 4-2-1) in the fourth process of explicit processing, excluding the case of implicit processing.

[0329] Next, we will explain other applicable scenarios for the adaptive image data processing method.

[0330] 4-1. Adaptive Image Data Processing Method Based on Multiple Segmentation Methods 1

[0331] Figures 17a to 17f This is a first example diagram used to illustrate an adaptive image data processing method based on multiple segmentation methods applicable to one embodiment of the present invention.

[0332] To reiterate, quadtree partitioning and binary tree partitioning are non-directional tree-like partitioning, meaning the resulting sub-blocks are partitioned without a specific direction. Binary tree partitioning, on the other hand, is a directional tree-like partitioning, meaning the resulting sub-blocks are partitioned with a specific direction, such as horizontal or vertical. Therefore, when performing binary tree partitioning, information related to the partitioning direction can be appended to the partitioning information.

[0333] Multiple tree-based partitioning can refer to supporting both quadtree partitioning and binary tree partitioning, but it is not limited to this because it can also support other partitioning methods.

[0334] The block (64×64) supports multiple tree-like partitioning methods. In quadtree partitioning, the largest block can be the same as the basic coded block (64×64), and the smallest block can be 16×16. In binary tree partitioning, the largest block can be 32×32, the smallest block can have a side length of 4, and the maximum partitioning depth can be 3. Furthermore, tree partitioning can have a priority order; quadtree partitioning can be performed first within the overlapping area of ​​quadtree and binary tree partitioning, but this is not a limitation.

[0335] See Figure 17a The actual image data in the block can be a diagonal region (32×16). First, a quadtree segmentation can be performed on block (1) to obtain four 32×32 blocks. Furthermore, a binary tree segmentation can be performed on the upper left block (2) of the four 32×32 blocks to obtain two 32×16 blocks. At this point, the upper block (3) of the two 32×16 blocks can contain only the actual image data, and the segmentation information can be implicitly processed as described above. In other words, the segmentation information can be adaptively processed.

[0336] See Figure 17b The actual image data in the block can be a diagonal region (8×8). First, quadtree segmentation can be performed on block (1) to obtain four 32×32 blocks. In addition, quadtree segmentation can be performed again on the upper left block (2) of the four 32×32 blocks to obtain four 16×16 blocks. Vertical binary tree segmentation can be performed on the upper left block (3) of the four 16×16 blocks to obtain two 8×16 blocks. In addition, horizontal binary tree segmentation can be performed on the left side block (4) of the two 8×16 blocks to obtain two 8×8 blocks. At this time, the upper block (5) of the two 8×8 blocks can only include actual image data, and the segmentation information can be implicitly processed as described above. In other words, the segmentation information can be adaptively processed.

[0337] See Figure 17c The actual image data in the block can be compared with... Figure 17b The same diagonal region (8×8). First, quadtree segmentation can be performed on block (1), thereby obtaining four 32×32 blocks. In addition, quadtree segmentation can be performed again on the top left block (2) of the four 32×32 blocks, thereby obtaining four 16×16 blocks. Quadtree segmentation can be performed again on the top left block (3) of the four 16×16 blocks, thereby obtaining four 8×8 blocks. At this time, the top left block (4) of the four 8×8 blocks can only include actual image data, and the segmentation information can be implicitly processed as described above. In other words, the segmentation information can be adaptively processed.

[0338] See Figure 17b as well as Figure 17c The actual image data areas of the two blocks can be the same, but in Figure 17b In order to obtain block (5), two binary tree splits need to be performed, while Figure 17c This can be obtained by performing a single quadtree partition. Therefore, when considering the number of partitions... Figure 17c This approach can be more effective in situations like those described above.

[0339] However, because the smallest block size of the quadtree partition is 16×16, therefore in Figure 17c The final quadtree partition in the block can be restricted. In other words, a quadtree partition cannot be performed in the corresponding block; only a binary tree partition can be performed.

[0340] To address the issues described above, the block segmentation settings on the basic coded blocks contained within the image boundary can be special settings that differ from the block segmentation settings within the image. For example, for basic coded blocks located on the lower boundary of the image, quadtree segmentation can be allowed when the horizontal and vertical lengths of the region containing the actual image data do not exceed half the horizontal and vertical lengths of the block before segmentation.

[0341] When applying the special block partitioning settings described above, because Figure 17c The size of the region containing the actual image data in the block is 8×8, while the size of the block before segmentation is 16×16. This means the size of the region containing the actual image data does not exceed half the size of the block before segmentation, thus allowing for a third quadtree segmentation. In other words, Figure 17c The situation shown could be a case where the smallest block size of the quadtree partitioning in a portion of the partitioning settings is changed to 8×8.

[0342] See Figure 17d The actual image data in the block can be a diagonal region (48×8). First, quadtree segmentation can be performed on block (1) to obtain four 32×32 blocks. Since the upper left block (2-1) and upper right block (2-2) of the four 32×32 blocks contain actual image data, additional segmentation can be performed on the upper left block (2-1) and upper right block (2-2) separately.

[0343] Regarding the additional segmentation of the two blocks, a horizontal binary tree segmentation can be performed on the upper left block (2-1), thereby obtaining two 32×16 blocks. A horizontal binary tree segmentation can then be performed again on the upper block (3-1) of the two 32×16 blocks, thereby obtaining two 32×8 blocks. At this point, the upper block (4-1) of the 32×8 blocks can contain only the actual image data.

[0344] A horizontal binary tree segmentation can be performed on the upper right block (2-2), thereby obtaining two 32×16 blocks. A vertical binary tree segmentation can be performed on the upper block (3-2) of the two 32×16 blocks, thereby obtaining two 16×16 blocks. Furthermore, a horizontal binary tree segmentation can be performed again on the left block (4-2) of the two 16×16 blocks, thereby obtaining two 16×8 blocks. At this time, the upper block (5) of the 16×8 blocks can include only the actual image data. In other words, Figure 17d In order to obtain blocks that contain only actual image data, the number of segmentations may increase, which may further increase the complexity.

[0345] See Figure 17e The actual image data in the block can be compared with... Figure 17d The same diagonal region (48×8). A horizontal binary tree segmentation can be performed on block (1), thereby obtaining two 64×32 blocks. A horizontal binary tree segmentation can be performed again on the upper block (2) of the two 64×32 blocks, thereby obtaining two 64×16 blocks. A horizontal binary tree segmentation can be performed again on the upper block (3) of the two 64×16 blocks, thereby obtaining two 64×8 blocks. When obtaining blocks containing only actual image data through the process described above, compared with... Figure 17d It requires fewer segmentation steps and is therefore more efficient.

[0346] However, according to the block partitioning settings, only quadtree partitioning is supported in the initial block partitioning step. Therefore, binary tree partitioning cannot be performed in the initial block and only quadtree partitioning is allowed.

[0347] To address the problems described above, the block segmentation settings on the basic coded blocks contained within the image boundary can be special settings different from the block segmentation settings within the image. For example, for a basic coded block located on the lower boundary of the image, binary tree segmentation can be allowed when the horizontal and vertical lengths of the region containing the actual image data exceed half the horizontal and vertical lengths of the block before segmentation. Specifically, horizontal binary tree segmentation is allowed when the horizontal length exceeds half, and vertical binary tree segmentation is allowed when the vertical length exceeds half. However, the segmentation settings described above can be applied to tree-like segmentation requiring direction, such as ternary tree segmentation and asymmetric binary tree segmentation, and are therefore not limited to this.

[0348] When applying the special block partitioning settings described above, because Figure 17e The size of the region containing the actual image data in the segmented block is 48×8, while the size of the block before segmentation is 64×64. This means the size of the region containing the actual image data exceeds half the size of the block before segmentation, thus allowing for binary tree segmentation. In other words, Figure 17e The situation shown could be a case where the maximum block size of the binary tree partition in a certain partitioning setting is changed to 64×64.

[0349] Specifically, when the horizontal and vertical lengths of the region including the actual image data exceed half of the horizontal and vertical lengths before block segmentation, the number of segmentations required for quadtree segmentation can be less than the number of segmentations required for binary tree segmentation.

[0350] The block segmentation settings described above allow quadtree segmentation when both the horizontal and vertical lengths of the region containing the actual image data on the lower right boundary of the image exceed half the horizontal and vertical lengths of the block before segmentation, and when both horizontal and vertical lengths are less than half the horizontal and vertical lengths of the block before segmentation. Binary tree segmentation in either the horizontal or vertical direction is allowed when only one of the horizontal or vertical lengths exceeds half the length of either the horizontal or vertical length of the block before segmentation. When the block segmentation settings do not allow segmentation as described above, special processing (partial segmentation setting modification) can be applied to the corresponding block settings. Therefore, block segmentation can be flexibly performed based on the modified segmentation settings.

[0351] The block segmentation setting allows segmentation as described below on boundaries other than a portion of the original boundary. The block segmentation setting allows horizontal binary tree segmentation on the lower boundary of the image, regardless of the vertical length of the region containing actual image data, and vertical binary tree segmentation on the right boundary of the image, regardless of the horizontal length of the region containing actual image data.

[0352] In other words, quadtree segmentation or binary tree segmentation can be performed on the lower right boundary of the image. Furthermore, binary tree segmentation in a segmentation direction parallel to the boundary, and quadtree segmentation, can be performed on the lower and right boundaries of the image, but binary tree segmentation in a segmentation direction perpendicular to the boundary cannot be performed. The description in this invention assumes that the pre-segmentation block includes both actual image data and padding data; this method is prohibited when the pre-segmentation block only includes actual image data.

[0353] Given a maximum block size of 32×32 in a binary tree partition, the possible block sizes based on the minimum block size (one side length of 4 and maximum partition depth of 3) are 32×32, 32×16, 16×32, 32×8, 8×32, 16×16, 32×4, 4×32, 16×8, and 8×16. However, as... Figure 17e As shown, when the maximum block size of the binary tree partition changes from 32×32 to 64×64, the obtainable block sizes can be 64×64, 64×32, 32×64, 64×16, 16×64, 32×32, 64×8, 8×64, 32×16, and 16×32. That is, some of the blocks obtainable under the existing block partitioning settings will not be obtainable under the changed block partitioning settings.

[0354] In addition to the maximum block size, other block segmentation settings such as the side length of the minimum block or the maximum segmentation depth can also be changed, thereby performing block segmentation in the basic coded blocks of the image boundary.

[0355] See Figure 17f The actual image data in the block can be compared with... Figure 17d as well as Figure 17e The same diagonal area (48×8). When applying the block segmentation settings described above, a horizontal binary tree segmentation can be performed on block (1), thereby obtaining two blocks of size 64×32. A horizontal binary tree segmentation can be performed again on the upper block (2) of the two 64×32 blocks, thereby obtaining two blocks of size 64×16. A horizontal binary tree segmentation can be performed again on the upper block (3) of the two 64×16 blocks, thereby obtaining two blocks of size 64×8. Furthermore, a vertical binary tree segmentation can be performed on the upper block (4) of the two 64×8 blocks, thereby obtaining two blocks of size 32×8. The left block of the two 32×8 blocks can contain only the actual image data, and the right block (5) of the two 32×8 blocks can be split vertically again to obtain two 16×8 blocks. The left block (6) of the two 16×8 blocks can contain only the actual image data.

[0356] However, under the special case of block partitioning settings as described above, partitions smaller than the minimum size used to encode residual coefficients are still not allowed. For example, when the minimum size used to encode residual coefficients is 4×4, blocks smaller than 4×4 are not allowed. However, it is allowed when the horizontal and vertical lengths are greater than or equal to the horizontal and vertical lengths of the minimum size used to encode each residual coefficient, and even when one of the horizontal and vertical lengths is smaller than its corresponding horizontal and vertical length but the other is long enough that the product of the horizontal and vertical lengths is greater than the product of the horizontal and vertical lengths of the minimum size used to encode the residual coefficients. Here, a block can be a block that supports a flag (Coded Block Flag) to confirm the presence or absence of residual coefficients in the block.

[0357] 4-2. Adaptive Image Data Processing Method Based on Multiple Segmentation Methods 2

[0358] Figures 18a to 18c This is a second illustration used to explain an adaptive image data processing method based on multiple segmentation methods applicable to one embodiment of the present invention.

[0359] This design assumes that blocks can support multiple tree-like partitioning methods, such as binary tree partitioning and ternary tree partitioning. The maximum block size is 64×64, the minimum block size is 4 units on one side, and the maximum partitioning depth is 4. There is no priority order associated with tree partitioning, but selection information related to which type of tree partitioning to perform can be generated.

[0360] That is, the segmentation information can generate segmentation-or-not information, tree selection information, and segmentation direction information, which can be generated sequentially in the order listed. Specifically, a segmentation-or-not information value of 0 indicates no segmentation, while a value of 1 indicates segmentation. A tree selection information value of 0 indicates binary tree segmentation, while a value of 1 indicates ternary tree segmentation. Furthermore, a segmentation direction value of 0 indicates a horizontal direction, while a value of 1 indicates a vertical direction.

[0361] Furthermore, it is assumed that the special block segmentation settings described in the above-described adaptive image data processing method 1 based on multiple segmentation methods can also be applied.

[0362] See Figures 18a to 18c The basic coded block size is 64×64, and the actual image data can be a diagonal region (64×16). Because the block can undergo binary tree and ternary tree segmentation, the block obtained through a single segmentation can be 64×64, 64×32, 32×64, 64×16 / 64×32 / 64×16, or 16×64 / 32×64 / 16×64.

[0363] See Figure 18a A horizontal binary tree segmentation can be performed on block (1) to obtain two 64×32 blocks. At this time, the segmentation depth can be increased from 0 to 1 and 1-0-0 can be generated as segmentation information, but implicit processing is also possible. A horizontal binary tree segmentation can be performed again on the upper block (2) of the two 64×32 blocks to obtain two 64×16 blocks. At this time, the segmentation depth can be increased from 1 to 2 and 1-0-0 can be generated as segmentation information, but implicit processing is also possible. Since the upper block (3) of the two 64×16 blocks only includes actual image data, the optimal segmentation form can be determined by adding segmentation, but in this example, etc. Figure 18b as well as Figure 18c In this case, it is assumed that no additional partitioning is performed on the region.

[0364] The segmentation information generated during the segmentation process described above can be represented as (64×16): 1-0-0-1-0-0-0-0-0. However, as segmentation information, the upper block (3) of the two 64×16 blocks can be explicitly processed because it only includes actual image data, while the remaining blocks can be implicitly processed because they only include filled areas. That is, when the encoded / decoded data is adaptively processed based on the actual image data, the segmentation information during the explicit processing can be 0. Here, 0 can refer to the segmentation information of the third generation of segmentation information.

[0365] See Figure 18b The system can perform horizontal ternary tree segmentation on block (1), thereby obtaining blocks of sizes 64×16, 64×32, and 64×16. The segmentation depth can be increased from 0 to 1, and 1-1-0 can be generated as segmentation information, but implicit processing is also possible. Because the upper block (2) of size 64×16 can contain only actual image data, additional segmentation is not required. Therefore, the optimal segmentation shape can be determined directly or through additional segmentation.

[0366] The segmentation information generated during the segmentation process described above can be represented as (64×16): 1-1-0-0-0-0. However, as segmentation information, the upper block (2) of size 64×16 can be explicitly processed because it only includes actual image data, while the rest can be implicitly processed because it only includes filled areas. That is, when the encoded / decoded data is adaptively processed based on the actual image data, the segmentation information during the explicit processing can be 0. Here, 0 can refer to the segmentation information of the second generated segmentation information.

[0367] exist Figure 18a as well as Figure 18b In the cases shown, the segmentation information can all be 0. However, as the number of segments increases during the encoding / decoding process, even when the data is processed implicitly, the number of steps required to reach a block containing only the actual image data can still differ. Therefore, when there is more than one method to reach a block containing only the actual image data from the initial block or basic encoded block, it can be set to implicitly use the method that achieves this with fewer segments. In other words, when Figure 18a as well as Figure 18b When all options are available, it can be set to default to using a relatively small number of splits. Figure 18b The ternary tree partitioning method is shown.

[0368] In one embodiment, vertical ternary tree segmentation can be performed when the horizontal length of the region including the actual image data on the right boundary of the image is consistent with 1 / 4 or 3 / 4 of the horizontal length of the block before segmentation, and vertical binary tree segmentation can be performed when neither is consistent.

[0369] Furthermore, horizontal ternary tree segmentation can be performed when the vertical length of the region including the actual image data on the lower boundary of the image is consistent with 1 / 4 or 3 / 4 of the vertical length of the block before segmentation, while horizontal binary tree segmentation can be performed when neither is consistent.

[0370] The above only applies to some boundary conditions, while the segmentation on other boundaries can be described as follows.

[0371] A horizontal or vertical ternary tree segmentation can be performed when the horizontal and vertical lengths of the region containing the actual image data on the lower right boundary of the image are consistent with at least one-quarter or three-quarters of the horizontal and vertical lengths of the corresponding pre-segmentation block; and a horizontal or vertical binary tree segmentation can be performed when neither is consistent.

[0372] Regarding the direction of binary tree segmentation, segmentation can be performed in the direction perpendicular to the corresponding length when either the horizontal or vertical length of the region containing the actual image data does not exceed 1 / 2; segmentation can be performed in either the horizontal or vertical direction when both exceed or do not exceed 1 / 2. For example, segmentation can be performed in the vertical direction when only the horizontal length of the region containing the actual image data does not exceed 1 / 2.

[0373] In addition, it can confirm whether the horizontal length of the region including the actual image data is consistent with 1 / 4 or 3 / 4 of the horizontal length of the block before segmentation, or whether the vertical length of the region including the actual image data is consistent with 1 / 4 or 3 / 4 of the vertical length of the block before segmentation, and perform ternary tree segmentation if either is consistent, otherwise perform binary tree segmentation.

[0374] In other words, binary tree segmentation or ternary tree segmentation can be performed on the lower right boundary of the image. Furthermore, binary tree segmentation or ternary tree segmentation with the segmentation direction parallel to the boundary can be performed on the lower and right boundaries of the image, but binary tree segmentation or ternary tree segmentation with the segmentation direction perpendicular to the boundary cannot be performed.

[0375] The embodiment of the present invention described above can be applied when the pre-segmentation block includes both actual image data and padding data, and can be prohibited when the pre-segmentation block only includes actual image data. Furthermore, for ease of explanation, the example given is a ternary tree segmentation performed in the above ratio (1:2:1), but other ratios can also be applied depending on the encoding settings.

[0376] Next, please refer to Figure 18c Other embodiments applicable to various segmentation methods are described below.

[0377] This design assumes that the block can support multiple tree-like partitioning methods, such as symmetric binary tree partitioning and asymmetric binary tree partitioning. The maximum block size is 64×64, the minimum block size is 4 units on one side, and the maximum partitioning depth is 4. Furthermore, there is no priority order related to tree partitioning, but information related to which type of tree partitioning to perform can be generated. In asymmetric binary tree partitioning, not only can partitioning direction information be added, but partitioning ratio information can also be added. This can be set to a pre-defined partitioning ratio such as 1:3 or 3:1, but other added partitioning ratios can also be applied.

[0378] Therefore, the segmentation information can generate segmentation status information, segmentation direction information, and tree selection information. When the tree selection information indicates asymmetric binary tree segmentation, segmentation ratio information can also be further generated. Furthermore, the segmentation information can be generated in the order described above. Specifically, a segmentation status of 0 indicates no segmentation, while a status of 1 indicates segmentation. A segmentation direction of 0 indicates a horizontal direction, while a status of 1 indicates a vertical direction. A tree selection information of 0 indicates a symmetric binary tree segmentation, while a status of 1 indicates an asymmetric binary tree segmentation. Additionally, when the asymmetric binary tree segmentation ratio information is 0, it indicates a wider ratio on the top or left side, while a status of 1 indicates a wider ratio on the bottom or right side.

[0379] The basic coded block size is 64×64, and the actual image data can be a diagonal region (64×16). Because the block can undergo symmetric and asymmetric binary tree segmentation, the obtainable block sizes can be 64×64, 64×32, 32×64, 64×48 / 64×16, 64×16 / 64×48, 48×64 / 16×64, and 16×64 / 48×64.

[0380] See Figure 18c The system can perform horizontal asymmetric binary tree segmentation on block (1), thereby obtaining blocks of sizes 64×16 and 64×48. At this time, the segmentation depth can be increased from 0 to 1 and 1-0-1-1 can be generated as segmentation information. However, since only one candidate can be supported, the segmentation information can be implicitly processed. In addition, since the 64×16 block (2) can only include actual image data, it is not necessary to perform additional segmentation. Therefore, the optimal segmentation form can be determined directly, or the optimal segmentation form can be determined by additional segmentation based on the block segmentation settings when additional segmentation is performed.

[0381] The segmentation information generated during the segmentation process described above can be represented as (64×16): 1-0-1-1-0-0. However, as segmentation information, since the 64×16 block (2) only includes actual image data, it can be explicitly processed, while since the rest only includes filling areas, it can be implicitly processed. That is, when the encoded / decoded data is adaptively processed based on the actual image data, the segmentation information during the explicit processing can be 0. Here, 0 can refer to whether the 64×16 block (2) is segmented or not.

[0382] Therefore, in Figure 18a as well as Figure 18c In the process of obtaining a block (64×16) containing only actual image data from the initial block or basic coded block, the segmentation information generated can all be 0. However, the information represented by 0 can be different from each other.

[0383] Furthermore, in order to progress from an initial block or basic coded block to a block that contains only actual image data, Figure 18a as well as Figure 18c Both can be used, thus allowing the segmentation method to be determined implicitly based on fewer segmentations. In other words, it is possible to determine the segmentation method using a relatively small number of segmentations. Figure 18c The asymmetric binary tree partitioning method is illustrated in the figure.

[0384] In one embodiment, vertical asymmetric binary tree segmentation can be performed when the horizontal length of the region including the actual image data on the right boundary of the image is consistent with 1 / 4 or 3 / 4 of the horizontal length of the block before segmentation; vertical symmetric binary tree segmentation can be performed when it is consistent with 1 / 2; and vertical symmetric or asymmetric binary tree segmentation can be performed when none of the three are consistent.

[0385] In this case, when all three are inconsistent, if the horizontal length of the region including the actual image data is less than 1 / 4 of the horizontal length of the block before segmentation, asymmetric binary tree segmentation in the vertical direction with a wider right side can be performed; if it exceeds 1 / 4 but is less than 1 / 2, symmetric binary tree segmentation in the vertical direction can be performed. Furthermore, if it exceeds 1 / 2 but is less than 3 / 4, asymmetric binary tree segmentation in the vertical direction with a wider left side can be performed; and if it exceeds 3 / 4, symmetric binary tree segmentation in the vertical direction can be performed.

[0386] Furthermore, horizontal asymmetric binary tree segmentation can be performed when the vertical length of the region including the actual image data on the lower boundary of the image is consistent with 1 / 4 or 3 / 4 of the vertical length of the block before segmentation; horizontal symmetric binary tree segmentation can be performed when it is consistent with 1 / 2; and horizontal symmetric or asymmetric binary tree segmentation can be performed when none of the three are consistent.

[0387] In the case where all three are inconsistent, when the vertical length of the region including the actual image data is less than 1 / 4 of the vertical length of the block before segmentation, asymmetric binary tree segmentation in the horizontal direction with a wider lower side can be performed; when it exceeds 1 / 4 but is less than 1 / 2, symmetric binary tree segmentation in the horizontal direction can be performed. Furthermore, when it exceeds 1 / 2 but is less than 3 / 4, asymmetric binary tree segmentation in the horizontal direction with a wider upper side can be performed; when it exceeds 3 / 4, symmetric binary tree segmentation in the horizontal direction can be performed.

[0388] The above only applies to some boundary conditions, while the segmentation on other boundaries can be described as follows.

[0389] Asymmetric binary tree segmentation in the horizontal or vertical direction can be performed when the horizontal and vertical lengths of the region including the actual image data on the lower right boundary of the image are consistent with at least one-quarter or three-quarters of the horizontal and vertical lengths of the corresponding pre-segmentation block; symmetric binary tree segmentation in the horizontal or vertical direction can be performed when they are consistent with one-half; and symmetric or asymmetric binary tree segmentation in the horizontal or vertical direction can be performed when none of the three are consistent. Explanations related to all inconsistent cases can be derived from the aforementioned right or lower boundary, so detailed explanations related to them will be omitted here.

[0390] In other words, symmetric or asymmetric binary tree segmentation can be performed on the lower right boundary of the image. Furthermore, symmetric and asymmetric binary tree segmentation with the segmentation direction parallel to the boundary can be performed on the lower and right boundaries of the image, but symmetric or asymmetric binary tree segmentation with the segmentation direction perpendicular to the boundary cannot be performed.

[0391] The embodiment of the present invention described above can be applied when the pre-segmentation block includes both actual image data and padding data, and can be prohibited when the pre-segmentation block only includes actual image data. Furthermore, for ease of explanation, the example is given in asymmetric binary tree segmentation assuming a 2-segmentation is performed according to the above ratio (1:3 or 3:1), but other ratios can also be applied depending on the encoding settings.

[0392] 4-3. Adaptive Image Data Processing Method Based on Multiple Segmentation Methods 3

[0393] The blocks can support multiple tree-like segmentation methods, such as quadtree segmentation, binary tree segmentation, and ternary tree segmentation. Other aspects are the same as those described in the adaptive image data processing methods 1 and 2 based on multiple segmentation methods. Furthermore, it is assumed that the special block segmentation settings described in the adaptive image data processing methods 1 and 2 based on multiple segmentation methods can also be applied.

[0394] In one embodiment, ternary tree segmentation in the horizontal or vertical direction can be performed when the horizontal and vertical lengths of the region including actual image data on the lower right boundary of the image are consistent with at least one-quarter or three-quarters of the horizontal and vertical lengths of the corresponding pre-segmentation block. Where neither is consistent, the horizontal and vertical lengths of the region including actual image data can be compared with the horizontal and vertical lengths of the pre-segmentation block, and quadtree segmentation can be performed when both exceed or do not exceed half; and binary tree segmentation in the horizontal or vertical direction can be performed when only one exceeds half.

[0395] A horizontal ternary tree segmentation can be performed when the horizontal and vertical lengths of the region containing the actual image data on the lower boundary of the image are consistent with 1 / 4 or 3 / 4 of the vertical length of the block before segmentation; otherwise, a horizontal binary tree segmentation can be performed.

[0396] Furthermore, vertical ternary tree segmentation can be performed when the horizontal and vertical lengths of the region containing the actual image data on the right boundary of the image are consistent with 1 / 4 or 3 / 4 of the horizontal length of the block before segmentation, while vertical binary tree segmentation can be performed when they are inconsistent.

[0397] In other words, quadtree, binary tree, and ternary tree segmentation can be performed horizontally or vertically on the lower boundary of the image, while binary tree and ternary tree segmentation with the segmentation direction parallel to the boundary can be performed on the lower and right boundaries of the image. However, although quadtree segmentation can be performed on the lower and right boundaries of the image depending on the situation, binary tree and ternary tree segmentation with the segmentation direction perpendicular to the boundary cannot be performed.

[0398] In the description of this invention, an adaptive image data processing method based on various segmentation methods has been described in conjunction with several embodiments (4-1, 4-2, and 4-3). However, this method can be modified according to the block segmentation settings and is therefore not limited thereto. In other words, a block can be determined as a block that can be obtained through block segmentation settings, thereby enabling adaptive processing of block segmentation information at the boundaries of the image.

[0399] Figure 19 This is a sequence diagram illustrating a method for performing image encoding through adaptive image data processing according to one embodiment of the present invention.

[0400] See Figure 19 An image encoding apparatus for adaptively processing image data according to one embodiment of the present invention can set image size information, basic encoding block size information, and block segmentation information (S1910), and then enlarge the image size to an integer multiple of the basic encoding block size (S1920). Furthermore, the image encoding apparatus can divide the enlarged image into basic encoding block units (S1930), and then encode the basic encoding blocks according to the block scanning order, block position information, and (re)set block segmentation information (S1940), and then send the encoded bitstream to an image decoding apparatus.

[0401] Furthermore, the image decoding apparatus according to one embodiment of the present invention can operate in a manner corresponding to the image encoding apparatus. In other words, the image decoding apparatus can reconstruct the image size information, the basic coded block size information, and the block segmentation information based on the bitstream received from the image encoding apparatus, thereby enlarging the image size to an integer multiple of the basic coded block size. Furthermore, the image decoding apparatus can divide the enlarged image into basic coded block units, and then perform decoding on the basic coded blocks according to the block scanning order, block position information, and (re)set block segmentation information.

[0402] Specifically, image magnification can be performed in a direction that is an integer multiple of the basic coded blocks, moving towards the right or bottom of the image. Encoding within the basic coded blocks located on the image boundaries can be adaptively performed based on the block position and block segmentation settings. These block segmentation settings can be directly used based on the encoding settings or reset by modifying a portion of them.

[0403] In one embodiment, the segmentation information of basic coded blocks located within the image can be explicitly processed, while the information of basic coded blocks located on the image boundary can be explicitly or implicitly processed based on the actual image data. Specifically, when settings such as the size and shape of the blocks obtainable during segmentation are changed and reset, this can be reflected during processing.

[0404] The preferred embodiments of the present invention have been described above in conjunction with applicable embodiments. However, those skilled in the art will understand that various modifications and alterations can be made to the present invention without departing from the spirit and scope of the invention as set forth in the following claims. < / y>

Claims

1. An image decoding method, comprising: Segment the first coded block in the current image to determine the second coded block; as well as Decode the second encoded block. The first coded block is segmented based on a predefined segmentation type in the image decoding device. The segmentation type includes at least one of binary tree segmentation or ternary tree segmentation. Wherein, the binary tree partitioning represents a partitioning type that divides a coded block into two coded blocks, and the binary tree partitioning includes horizontal binary tree partitioning and vertical binary tree partitioning. The ternary tree partitioning refers to a partitioning type that divides a coded block into three coded blocks, and the ternary tree partitioning includes horizontal ternary tree partitioning and vertical ternary tree partitioning. Specifically, binary tree partitioning and ternary tree partitioning are applied only when quadtree partitioning is no longer performed. Specifically, when the boundary of the current image is the bottom boundary, only the horizontal binary tree segmentation and the vertical binary tree segmentation are allowed for the binary tree segmentation, while the vertical ternary tree segmentation is not allowed. Specifically, when the boundary of the current image is the right boundary, only the vertical binary tree segmentation of the horizontal binary tree segmentation and the vertical binary tree segmentation is allowed for binary tree segmentation, while the horizontal ternary tree segmentation is not allowed.

2. The image decoding method according to claim 1, in, The first coded block is segmented based on the segmentation information, and The segmentation information includes at least one of a first flag indicating whether the first coded block is segmented, a second flag indicating the segmentation direction, or a third flag indicating either the binary tree segmentation or the ternary tree segmentation.

3. The image decoding method according to claim 2, in, When the value of the second flag is the first value, the first coded block is divided horizontally, and Specifically, when the value of the second flag is the second value, the first coded block is divided in the vertical direction.

4. The image decoding method according to claim 3, in, When the value of the third flag is the first value, perform the binary tree split, and Specifically, when the value of the third flag is the second value, the ternary tree split is performed.

5. An image coding method, comprising: Segment the first coded block in the current image to determine the second coded block; as well as Encode the second coded block, The first coded block is segmented based on a segmentation type predefined in the image encoding device. The segmentation type includes at least one of binary tree segmentation or ternary tree segmentation. Wherein, the binary tree partitioning represents a partitioning type that divides a coded block into two coded blocks, and the binary tree partitioning includes horizontal binary tree partitioning and vertical binary tree partitioning. The ternary tree partitioning refers to a partitioning type that divides a coded block into three coded blocks, and the ternary tree partitioning includes horizontal ternary tree partitioning and vertical ternary tree partitioning. Specifically, binary tree partitioning and ternary tree partitioning are applied only when quadtree partitioning is no longer performed. Specifically, when the boundary of the current image is the bottom boundary, only the horizontal binary tree segmentation and the vertical binary tree segmentation are allowed for the binary tree segmentation, while the vertical ternary tree segmentation is not allowed. Specifically, when the boundary of the current image is the right boundary, only the vertical binary tree segmentation of the horizontal binary tree segmentation and the vertical binary tree segmentation is allowed for binary tree segmentation, while the horizontal ternary tree segmentation is not allowed.

6. A method for transmitting a bit stream, characterized in that, The bitstream is generated by performing an encoding method, and the method for transmitting the bitstream includes transmitting the bitstream generated based on a second coded block. The encoding method includes: Segment the first coded block in the current image to determine the second coded block; and Encode the second coded block, The first coded block is segmented based on a segmentation type predefined in the image encoding device. The segmentation type includes at least one of binary tree segmentation or ternary tree segmentation. Wherein, the binary tree partitioning represents a partitioning type that divides a coded block into two coded blocks, and the binary tree partitioning includes horizontal binary tree partitioning and vertical binary tree partitioning. The ternary tree partitioning refers to a partitioning type that divides a coded block into three coded blocks, and the ternary tree partitioning includes horizontal ternary tree partitioning and vertical ternary tree partitioning. Specifically, binary tree partitioning and ternary tree partitioning are applied only when quadtree partitioning is no longer performed. Specifically, when the boundary of the current image is the bottom boundary, only the horizontal binary tree segmentation and the vertical binary tree segmentation are allowed for the binary tree segmentation, while the vertical ternary tree segmentation is not allowed. Specifically, when the boundary of the current image is the right boundary, only the vertical binary tree segmentation of the horizontal binary tree segmentation and the vertical binary tree segmentation is allowed for binary tree segmentation, while the horizontal ternary tree segmentation is not allowed.

Citation Information

Patent Citations

  • Method for encoding a coding unit at a picture boundary

    EP3059708A1