Image encoding device, method, and program
The image coding device addresses excessive coding processing in geometric partitioning by applying block-by-block prediction with a subset of geometric partitioning modes, enhancing computational efficiency without compromising image quality.
Patent Information
- Application Number
- JP2022101437
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2042-06-23
AI Technical Summary
The existing geometric partitioning methods in image encoding require excessive coding processing due to large geometric partition mode sets, leading to inefficiencies.
An image coding device that selectively applies block-by-block prediction using a subset of predefined geometric partitioning modes based on block information, reducing the number of candidates for geometric partitioning and motion information to minimize coding processing.
This approach significantly reduces coding processing while maintaining image quality by narrowing down the geometric partitioning mode set, thus optimizing computational efficiency.
Smart Images

Figure 0007802618000001 
Figure 0007802618000002 
Figure 0007802618000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image encoding device, method, and program that apply geometric partitioning. [Background technology]
[0002] Non-Patent Document 1 discloses a technique called geometric partitioning, in which an encoding device divides an encoding block in a predefined manner and performs predictive encoding for each region using motion information generated by a merger. Here, the motion information refers to a motion vector and an index that identifies a reference frame. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11),JVET-T2002. Summary of the Invention [Problem to be solved by the invention]
[0004] The geometric division of Non-Patent Document 1 has a problem in that the amount of coding processing is large.
[0005] That is, in Non-Patent Document 1, one geometric partition mode to be applied to a current coding block is determined from a geometric partition mode set. Here, the geometric partition mode is information configured from a partition boundary (a line dividing the current coding block into two regions) and two pieces of motion information used for predictive coding. The geometric partition mode set is a set whose elements are geometric partition modes. Since a geometric partition mode that minimizes the rate-distortion cost (cost based on a cost function J=D+λR, etc., depending on the bit rate R and distortion D) is selected from the geometric partition mode set, a problem arises in that if the geometric partition mode set has a large number of elements, an extremely large amount of coding processing is required.
[0006] In view of the above-mentioned problems with the conventional technology, an object of the present invention is to provide an image coding device, method, and program that can reduce the amount of coding processing while maintaining quality. [Means for solving the problem]
[0007] To achieve the above object, the present invention provides an image coding device that encodes video by applying block-by-block prediction. When handling a block to be coded that uses inter prediction and geometric partitioning as prediction modes, the device first extracts information about the block to be coded, and, based on the information about the block to be coded, selects only a subset of a geometric partitioning mode set predefined in the geometric partitioning as an evaluation target for use in coding the block to be coded using inter prediction and geometric partitioning. The image coding device also provides a method and program corresponding to the image coding device. The device secondly determines the subset of a block to be coded that uses only a subset of a geometric partitioning mode set predefined in the geometric partitioning as an evaluation target for use in coding the block to be coded using inter prediction and geometric partitioning. The device further provides a method and program corresponding to the image coding device. [Effects of the Invention]
[0008] According to the first feature, by selecting only a partial set from the entire geometric partitioning mode set predefined in the geometric partition based on the information on the block to be coded, as an evaluation target to be used for coding after applying inter prediction and geometric partitioning to the block to be coded, it is possible to reduce the amount of coding processing while maintaining quality. According to the second feature, by using only a portion of the combinations of geometric partitioning mode candidates and motion information candidates that constitute the entire geometric partitioning mode set, with respect to combinations of angles of partitioning line segments and distances from block centers that are predefined as the entire geometric partitioning mode candidates, it is possible to reduce the amount of coding processing while maintaining quality. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram illustrating an image processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of functional blocks of an image encoding device according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating an example of functional blocks of an inter prediction unit according to an embodiment (a reference method based on magnitude information). [Figure 4] 1 is an explanatory example of geometric division. [Figure 5] FIG. 10 is an explanatory diagram showing that a division boundary in geometric division is given by a combination of an angle and a distance. [Figure 6] FIG. 10 is a flow diagram of a speed-up method using the size of an encoding block according to an embodiment. [Figure 7] FIG. 10 illustrates an implementation example of thinning out a geometric partition mode set. [Figure 8] FIG. 10 illustrates an implementation example of thinning out a geometric partition mode set. [Figure 9] FIG. 10 is a diagram illustrating an example of functional blocks of an inter prediction unit according to an embodiment of Modification 1 with respect to the base method. [Figure 10] FIG. 10 is a flowchart of a speed-up method according to the first modified embodiment. [Figure 11]FIG. 10 is a flowchart of a speed-up method according to the second modified embodiment. [Figure 12] FIG. 10 is a flowchart of a speed-up method according to the third modified embodiment. [Figure 13] 13A and 13B are diagrams showing examples of determining the perspective of an angle pattern in the vertical and horizontal directions in Modification 3. FIG. [Figure 14] FIG. 10 is a flowchart of a speed-up method according to the fourth modified embodiment. [Figure 15] 13A and 13B are diagrams illustrating examples of adjacent blocks and the like referred to when deriving motion information, as an explanatory example of the selection process in Modification 4. FIG. [Figure 16] FIG. 13 is a flowchart of a speed-up method according to the fifth modified embodiment. [Figure 17] FIG. 13 is a flowchart of a speed-up method according to the sixth modification. [Figure 18] FIG. 13 is a flowchart of a speed-up method according to an embodiment of Modification 7. [Figure 19] FIG. 1 is a diagram illustrating a hardware configuration of a typical computer. DETAILED DESCRIPTION OF THE INVENTION
[0010] Various embodiments of the present invention will be described below with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, etc., and various variations, including combinations with other existing components, are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0011] (First embodiment: Method for reducing the amount of coding processing for geometric division using the size of coding blocks) An image processing system 10 according to a first embodiment of the present invention will now be described with reference to Figures 1 to 8. Figure 1 is a diagram showing an image processing system 10 according to this embodiment.
[0012] (Image Processing System 10) As shown in FIG. 1, an image processing system 10 includes an image encoding device 100 and an image decoding device 200. The image encoding device 100 is configured to generate encoded data by encoding an input image signal. The image decoding device 200 is configured to generate an output image signal by decoding the encoded data. Here, the encoded data may be transmitted from the image encoding device 100 to the image decoding device 200 via a transmission path. Alternatively, the encoded data may be generated in the image encoding device 100 and stored in a storage medium, and then provided to the image decoding device 200 by reading the encoded data from the storage medium.
[0013] (Image encoding device 100) An image encoding device 100 according to this embodiment will be described below with reference to Fig. 2. Fig. 2 is a diagram showing an example of functional blocks of the image encoding device 100 according to this embodiment. The image encoding device 100 includes an inter prediction unit 111, an intra prediction unit 112, a subtractor 121, an adder 122, a transform and quantization unit 131, an inverse transform and inverse quantization unit 132, an encoding unit 140, an in-loop filtering processing unit 150, a frame buffer 160, and a block division unit 170.
[0014] The inter prediction unit 111 is configured to generate predicted pixels corresponding to a signal to be coded by inter prediction (inter-frame prediction). Specifically, the inter prediction unit 111 compares a frame to be coded (hereinafter referred to as a target frame) with a reference frame stored in the frame buffer 160 (the reference frame is a frame obtained by reconstructing an already coded frame (which may be a portion of a frame) through processing performed by the inverse transform unit / inverse quantization unit 132 and subsequent steps), thereby identifying a reference block included in the reference frame and determining a corresponding motion vector (MV). Furthermore, the inter prediction unit 111 generates predicted pixels for each block to be coded (hereinafter referred to as a target block) based on the reference block and the motion vector. The inter prediction unit 111 outputs the predicted pixels to a subtractor 121 and an adder 122.
[0015] The intra prediction unit 112 is configured to generate predicted pixels corresponding to a current signal by intra prediction (intra-frame prediction). Specifically, the intra prediction unit 112 identifies a reference block included in a current frame and generates predicted pixels for each current block based on the reference block. The intra prediction unit 112 outputs the predicted pixels to the subtractor 121 and the adder 122. The reference block is, for example, a block adjacent to the current block.
[0016] The subtractor 121 is configured to subtract predicted pixels (predicted pixels predicted by the inter prediction unit 111 or the intra prediction unit 112) from the input image signal to generate a predicted residual signal, and output the predicted residual signal to the transformation and quantization unit 131.
[0017] The adder 122 is configured to add predicted pixels (predicted pixels predicted by the inter prediction unit 111 or the intra prediction unit 112) to the predicted residual signal output from the inverse transform / inverse quantization unit 132 to generate a pre-filtered decoded signal, and output the pre-filtered decoded signal to the intra prediction unit 112 and the in-loop filter processing unit 150.
[0018] The transform / quantization unit 131 is configured to perform transform processing and quantization processing on the prediction residual signal output from the subtractor 121 to generate quantized level values, and output the quantized level values to the inverse transform / inverse quantization unit 132 and the encoding unit 140. The transform processing is processing to convert the prediction residual signal into a frequency component signal. In such transform processing, a basis pattern (transform matrix) corresponding to a Discrete Cosine Transform (DCT) may be used, or a basis pattern (transform matrix) corresponding to a Discrete Sine Transform (DST) may be used. The transform processing may not be performed.
[0019] The inverse transform / inverse quantization unit 132 is configured to perform inverse quantization processing and inverse transform processing on the quantized level values to generate a prediction residual signal and output it to the addition unit 122 .
[0020] The encoding unit 140 is configured to encode the quantized level values output from the transform / quantization unit 131 and output encoded data. The encoding is entropy encoding, which assigns codewords of different lengths based on the occurrence probability of the quantized level values. The encoding unit 140 is also configured to encode control data used in the decoding process in the image decoding device 200 and output the encoded data. Here, the control data refers to additional information used by the image decoding device 200 to identify the block partition mode, prediction mode, motion information, etc. Although not shown in the figure to simplify the drawing, these are output from each functional block of the encoding device 100. For example, the block partition mode is output from the block partitioning unit 170, the intra prediction mode is output from the intra prediction unit 112, and the inter prediction mode and motion information are output from the inter prediction unit 111. A control unit (not shown) identifies a single control data from multiple control data candidates output from each functional block. As an identification method, control data that minimizes the rate-distortion cost may be selected.
[0021] The in-loop filtering unit 150 is configured to filter the unfiltered decoded signal, generate a filtered decoded signal (a reconstructed image signal or a reference frame; a reference frame refers to a reconstructed image signal that is used for predictive coding of another frame in the inter prediction unit 111), and output it to the frame buffer 160. The filtering includes deblocking filtering that reduces distortion occurring at boundaries between coding blocks, prediction blocks, or transform blocks, and adaptive loop filtering that switches filters based on filter coefficients and filter selection information transmitted from the image coding device 100, local properties of the image pattern, and the like.
[0022] The frame buffer 160 is configured to accumulate the filtered decoded signal and output it to the inter prediction unit 111 as a reference frame.
[0023] The block division unit 170 is configured to divide the target frame into square or rectangular blocks that do not overlap each other, and output the encoding target signal for each block to the inter predictor 111 and the intra predictor 112. Here, the block division method may be different for the inter predictor 111 and the intra predictor 112, or for the luminance signal and the color difference signal, for example. The block division may also be performed using any of the following methods: quartering by halving the vertical and horizontal lengths, halving by halving the vertical or horizontal lengths, or thirds by halving the vertical or horizontal lengths.
[0024] (Geometric partitioning mode candidate selection method using the size of the target coding block (hereinafter referred to as the reference method)) (Inter prediction unit 111) The inter prediction unit 111 according to this embodiment will be described with reference to Figures 3 to 8. Figure 3 shows an example of functional blocks of the inter prediction unit 111 according to this embodiment. The inter prediction unit 111 in Figure 3 is configured to include a merging unit 1111, a coding block information extraction unit 1112, a geometric partitioning mode candidate selection unit 1113, and a predicted pixel generation unit 1114.
[0025] 3 is configured to perform processing in the case where inter prediction is applied to a coding block, and inter prediction is not applied to the entire coding block, but rather inter prediction is applied to each divided portion after applying geometric partitioning. (As described above, the determination of whether to apply inter prediction to the coding block, or whether to apply a geometric partitioning technique after applying inter prediction, can be made separately by a control unit (not shown) provided in the image coding device 100 using any method.)
[0026] The merge unit 1111 derives a plurality of pieces of motion information for the current block from motion information for adjacent positions (adjacent blocks) in time and space relative to the current coding block, and outputs the motion information.
[0027] By using a merge mode technique, the merge unit 1111 can derive multiple pieces of motion information as candidates for motion information to be applied to the current coding block from motion information already determined for coded adjacent blocks, and output the derived pieces of motion information to the coding block information extraction unit 1112. (Note that each piece of motion information already determined for coded adjacent blocks may include a set of motion vectors and an index specifying a reference frame (hereinafter, a reference index) as a single piece of motion information, or may include two sets of motion vectors and reference indexes. Specifically, when unidirectional prediction using one reference frame is applied to a coding block including spatially and temporally adjacent regions, one set of motion vectors and reference indexes is included, and when bidirectional prediction using two reference frames is applied or when geometric partitioning is applied, two sets of motion vectors and reference indexes are included. Note that when geometric partitioning is applied to spatially and temporally adjacent regions, it is also possible to obtain only a set of motion vectors and reference indexes to be used for predictive coding of the adjacent regions based on block partitioning by geometric partitioning.)
[0028] The coding block information extraction unit 1112 extracts the size of the current coding block and outputs the coding block information to the geometric partitioning mode candidate selection unit 1113. Here, the coding block information can be given as, for example, (1) the length of the short side of the current coding block, (2) the length of the long side, or (3) all or any part (one or two) of the number of pixels of the current coding block.
[0029] The geometric partition mode candidate selection unit 1113 selects multiple geometric partition mode candidates from the geometric partition mode set based on the coding block information, and outputs the selected geometric partition mode candidates. Figure 4 is an explanatory example of geometric partitioning. In this embodiment, an example of a partition boundary is expressed as a combination of angle patterns and distance patterns shown in Figure 5.
[0030] 4 shows an example of a geometric partition (one of multiple partition candidates) applied to a target coding block BL in a frame F to be coded, where the target coding block BL is divided into a first portion PA and a second portion PB by a straight partition boundary DB. Note that a first motion vector MVA for the first portion PA and a second motion vector MVB for the second portion PB are searched for in the downstream predicted pixel generation unit 1114 from among the candidate motion information provided by the merge unit 1111 to optimize cost. These motion vectors MVA and MVB constitute a motion vector MV=(MVA, MVB) for each of the portions PA and PB geometrically divided by the partition boundary DB. Here, the motion vector MVA points to block BLA in the reference frame FA as the reference destination of the first portion PA, and the motion vector MVB points to block BLB in the reference frame FB as the reference destination of the second portion PB.
[0031] Figure 5 shows that the geometric division mode set, which is (all) candidates for the division boundary shown in Figure 4 for one case of the division boundary DB, is given by the combination (φ, ρ) of the angle parameter φ and the distance (offset) parameter ρ.
[0032] The angle parameter φ determines the angle that the line of the division boundary DB makes with the block, and with the direction from the center of the block to the horizontal right as φ=0°=φ1, for example, 20 candidates can be prepared in order of increasing angle, such as φ={φ1,φ2,φ3,...,φ20}, as shown in the figure. Each of the 20 angles φ can be determined so that the possible values of tan φ for each of these angles φ are, for example, as follows: tanφ=0, ±1 / 4, ±1 / 2, ±1, ±2, ±∞
[0033] The distance (offset) parameter ρ determines the displacement of the straight line of the division boundary DB from the center of the block (displacement in the direction given by the angle parameter φ (parallel movement in the direction of φ+90°)), and for each angle φ, four distance parameter candidates can be determined in advance, such as distance parameter ρ={ρ1, ρ2, ρ3, ρ4}, in order of closest to the center. (ρ=ρ1 corresponds to the case where the division boundary DB is closest to the block center and passes through the block center, and ρ=ρ4 corresponds to the case where the division boundary DB is farthest from the block center, and the division boundary DB for each distance parameter ρ is constructed by horizontally moving the straight line from the block center in the direction of φ+90° by the distance parameter ρ as a horizontal movement relative to the straight line in the direction φ.)
[0034] Figure 5 shows the distance parameters ρ = {ρ1, ρ2, ρ3, ρ4}, which are commonly defined as the distance from the center of the block in the direction of the angle φ, for the angle parameters φ = φ3, φ9, φ13, φ19 = 45°, 135°, 225°, and 315°.
[0035] As shown schematically in Figures 4 and 5, when geometric partitioning is applied, the total number of combinations to be searched (total geometric partitioning mode "sets") is an enormous number, for example, 64 * 30 = 1920, which is calculated by multiplying the number of "partition boundary" candidates by the number of "motion information" candidates, as shown below. Note that when one piece of motion information contains two sets of motion vectors and indices that specify a reference frame, a predetermined method is used to determine which of the two sets of motion vectors and indices that specify a reference frame to use. Specifically, this can be determined, for example, by the method described in Non-Patent Document 1. Division boundary: 64 ways (64 combinations are obtained by excluding the same division boundary and the division boundary that divides the block into four, two, or three divisions by the block dividing unit 170 from the 20×4=80 combinations in FIG. 5.) Candidates for movement information: 30 (The merge unit 1111 can obtain six pieces of motion information from the motion information applied during predictive coding in spatiotemporally adjacent coded regions. In geometric division, it is possible to select from these six pieces of motion information to be applied to the two regions after dividing the current coding block. Since the same motion information cannot be selected for the two regions, the number of combinations of motion information for the two regions is 6P2=30.)
[0036] The geometric division mode candidate selection unit 1113 of this embodiment can contribute to reducing the amount of calculation by narrowing down the above-mentioned huge number of candidates in advance using the methods described below (narrowing down all candidates for the geometric division mode).
[0037] Returning to the explanation of Fig. 3, the predicted pixel generator 1114 generates a predicted pixel using the multiple geometric partition mode candidates selected by the geometric partition mode candidate selector 1113. Here, the predicted pixel may be generated by taking a weighted average of the blocks in each of the two reference frames. The weight may be determined according to the partition boundary.
[0038] The predicted pixel generation unit 1114 generates predicted pixels for each of the candidates reduced by the geometric partitioning mode candidate selection unit 1113, and outputs the generated predicted pixels to the inter prediction unit 111. As described above, which one of the reduced geometric partitioning mode candidates is actually used for encoding can be determined by a control unit (not shown) provided in the image encoding device 100 as the one that minimizes the rate-distortion cost.
[0039] The flow of selecting geometric partitioning mode candidates and generating prediction signals executed by the coding block information extraction unit 1112, the geometric partitioning mode candidate selection unit 1113, and the prediction pixel generation unit 1114 of the inter prediction unit 111 of this embodiment will be explained using Figure 6.
[0040] In step S1, it is determined whether geometric partitioning is enabled for the current coding block in the settings of the image coding device 100. If it is enabled, the process proceeds to step S2; if it is disabled, the process ends. (Note that if it is disabled, the current coding block will be coded using a predictive coding method other than the applied geometric partitioning.)
[0041] In step S2, the size of the target coding block is extracted. As described above, the size of the target coding block may be determined by using all or some of the length of the short side, the length of the long side, and the number of pixels.
[0042] In step S3, a plurality of geometric partition mode candidates are selected as candidates narrowed down from the entire geometric partition mode set based on the size of the target coding block.
[0043] In step S4, predicted pixels are generated using the geometric partitioning mode candidates, and the flow of FIG. 6 ends.
[0044] The selection process in step S3 will be described in detail below.
[0045] In this selection process, a predetermined threshold value may be used for the length of the short side or the length of the long side or the area (or the number of pixels) of the target coding block, and specifically, the following methods 1 to 5 may be used.
[0046] <Method 1...Reducing the number of candidates when the long side is determined to be short> For example, geometric partition mode candidates may be selected so that the number of candidates when the length of the long side of the current coding block is equal to or less than a threshold is smaller than the number of candidates when the length is greater than the threshold. Specifically, when the length is equal to or less than the threshold, multiple geometric partition mode candidates may be selected from the geometric partition mode set, and when the length is greater than the threshold, all of the geometric partition mode set may be used as geometric partition mode candidates.
[0047] The reason for this is as follows: When the target coding block is small, even if there are few patterns of division boundaries, it is possible to fit the boundaries of the subject, foreground, and background within the block to a limited division pattern, and it is expected that the amount of coding processing can be reduced while maintaining coding performance. The threshold value can be set to any value, and may be 16 pixels for the length of the long side.
[0048] <Method 2...Reducing the number of candidates when the short side is determined to be long> For example, geometric partition mode candidates may be selected so that the number of candidates when the size of the short side of the current coding block is equal to or greater than a threshold is smaller than the number of candidates when the size is smaller than the threshold. Specifically, when the size is equal to or greater than the threshold, multiple geometric partition mode candidates may be selected from the geometric partition mode set, and when the size is smaller than the threshold, all of the geometric partition mode set may be used as geometric partition mode candidates.
[0049] The reason for this is as follows: When the target coding block is large, the boundaries between the subject and foreground / background within the block are more likely to be composed of multiple straight lines than when the block is small. Furthermore, even if the boundaries within the block are composed of a single type of straight line, it is difficult to accurately match the boundaries within the block at each division boundary in the geometric division mode if the block size is large. Therefore, even if a certain number of similar patterns are removed from the angle and distance patterns that make up the geometric division mode, the prediction accuracy that takes into account the boundaries within the block using geometric division will not be significantly reduced, and it is expected that the amount of coding processing required to select the geometric division mode will be significantly reduced. The threshold value may be set to any value, and may be 32 pixels for the length of the short side.
[0050] <Method 3: Reducing the number of candidates when the area or number of pixels is determined to be small> Furthermore, to select geometric partitioning mode candidates, the area of the current coding block may be thresholded instead of the size of the long or short side described above. If the area is equal to or less than the threshold, geometric partitioning mode candidates may be selected so that the number of candidates is smaller than when the area is greater than the threshold. The area of the current coding block may be determined using, for example, the number of pixels in the current coding block. The threshold may also be set to any value, such as 1024 pixels.
[0051] <Method 4: Reducing the number of candidates when determined to be elongated> In this selection process, a predetermined threshold value may be used for the ratio (aspect ratio) between the length of the short side and the length of the long side of the current coding block. For example, geometric partitioning mode candidates may be selected so that the number of candidates when the result of dividing the length of the long side by the length of the short side is equal to or greater than the threshold value is smaller than the number of candidates when the result is smaller than the threshold value.
[0052] The reason for this is as follows: When the target coding block is elongated, if the division boundary is configured by combining the same angle patterns and distance patterns as when it is square, multiple similar division boundary patterns will exist. (For example, when the target coding block is elongated horizontally, it is expected that angles φ={φ5, φ6, φ7, φ15, φ16, φ17} (angles φ near tan φ=±∞) in Figure 5 will result in almost identical division boundary patterns.) By eliminating similar patterns, it is expected that the amount of coding processing can be reduced while maintaining coding performance. Note that the threshold may be set to any value, such as 2, which is the result of dividing the length of the long side by the length of the short side.
[0053] <Method 5... Thinning out of geometrically divided mode sets> In this selection process, a geometric division mode of a predetermined pattern (a pattern thinned out from a state in which all patterns are covered) for the angle pattern and the distance pattern may be selected from the geometric division mode set.
[0054] <Implementation example of method 5, part 1> For example, as shown in FIG. 7, it is possible to thin out the 20 possible angles φ and use only eight of them, namely {0°, 45°, 90°, 135°, 180°, 225°, 270°, 315°}, and / or thin out the four possible distances ρ by half and use only two of them, namely the second closest and fourth closest division boundaries from the center {ρ2, ρ4}.
[0055] In addition, Figure 7 shows an example in which the angle φ is thinned to eight patterns, and the distance ρ pattern is thinned to two patterns {ρ2, ρ4} for angles φ = 45° and 90°, based on the example in Figure 5.
[0056] The reason for this is as follows. Generally, when capturing video, the camera is often held horizontally or vertically relative to the ground. Therefore, it is believed that many boundaries of objects, such as man-made objects, exist horizontally or vertically in the target block. While it might be possible to limit the angle patterns of the division boundaries to horizontal or vertical, this would result in inconsistencies in matching with non-horizontal or vertical boundaries, potentially reducing prediction accuracy. Therefore, horizontal and vertical angle patterns, as well as intermediate angles such as 45°, 135°, 225°, and 315°, are used. Furthermore, geometric division modes that use horizontal or vertical angle patterns and distance patterns passing through the center are equivalent to performing predictive coding on each block after bisection by the block division unit 170, and therefore are not included in the geometric division mode set to avoid process duplication. (In other words, as mentioned above, the geometric division mode set is defined to avoid duplication.) Therefore, by using the distance patterns second closest to the center and fourth closest to the center, it is possible to reduce the division boundaries without spatial bias. This is expected to reduce the amount of coding processing while maintaining coding performance.
[0057] <Implementation example 2 of Method 5> 8, for example, the division boundaries whose angles are 45°, 135°, 225°, and 315° and whose distance passes through the center and the third closest division boundary, and the division boundaries whose angles are 0°, 90°, 180°, and 270° and whose distances are the second closest and fourth closest from the center may be used. In other words, the following 8+8=16 combinations of set A (4*2=8 combinations) or set B (4*2=8 combinations) may be used as candidates (φ,ρ) to be narrowed down from a total of 64 geometric division mode sets. (φ,ρ)∈A∪B A={45°,135°,225°,315°}×{ρ1,ρ3} B={0°,90°,180°,270°}×{ρ2,ρ4}
[0058] In the example of Figure 8, based on the example of Figure 5, sets A and B are shown for the angle φ, and for the distance ρ, candidates ρ={ρ1, ρ3} for set A when the angle φ=45° and candidates ρ={ρ2, ρ4} for set B when the angle φ=90° are shown.
[0059] The reason for this is as follows. The reason for selecting the angle pattern has been described above. Furthermore, for distance patterns, only when the angles are 45°, 135°, 225°, and 315°, the one that passes through the center and the one third closest to the center can be selected. Here, for combinations where the angles differ by 180°, such as (45°, 225°) and (135°, 315°), distance patterns that pass through the center represent the same division boundary, and therefore only one of the division boundaries is included in the geometric division mode set. (In other words, as described above, the geometric division mode set is defined to avoid overlaps.) Therefore, the selection method shown in Figure 8 can reduce the number of geometric division candidate elements compared to the selection method in Figure 7 while avoiding spatial bias in the division boundary. Therefore, it is expected to further reduce the amount of coding processing while maintaining coding performance.
[0060] Method 5 can be used in combination with any of methods 1 to 4, rather than being used alone. The case where method 5 is used alone will be described as Modified Example 1 below.
[0061] (Variation 1: Geometric Partition Mode Candidate Selection Method for Smaller Configurations) The difference between this modified example and the standard method described above will be described below with reference to Figures 9 and 10. The difference is that this modified example selects geometric partitioning mode candidates without using the size of the target coding block.
[0062] Fig. 9 is a diagram showing the configuration of an inter prediction unit 111 according to Modification 1. The following describes the differences between the configuration of Fig. 9 and that shown in Fig. 3. The inter prediction unit 111 in Fig. 9 is configured not to have (omit) the coded block information extraction unit 1112 shown in Fig. 3.
[0063] As a result of this omission, the geometric partition mode candidate selection unit 1113 selects multiple geometric partition mode candidates from the geometric partition mode set without using target coding block information as in the standard method, and outputs these geometric partition mode candidates.
[0064] 10 shows a flow of selecting geometric partition mode candidates and generating predicted signals in this modified example, which is executed by the geometric partition mode candidate selection unit 1113 and the predicted pixel generation unit 1114. The differences between the flow in FIG. 10 and the flow shown in FIG. 6 will be explained below.
[0065] The flow in Fig. 10 does not have step S2 in Fig. 5, and the processing content of step S13 in Fig. 10, which corresponds to step S3 in Fig. 5, is partially different. Also, as indicated by the same reference numerals, steps S1 and S4 in Fig. 10 are the same as steps S1 and S4 in Fig. 5.
[0066] In step S13, multiple geometric partition mode candidates are selected from the geometric partition mode set regardless of the current coding block. The selection of the geometric partition mode candidates may be performed by selecting geometric partition modes of predetermined patterns for angle patterns and distance patterns, as has already been described in the method for selecting geometric partition mode candidates using the size of the current coding block (reference method).
[0067] That is, in step S13, for example, the method described as Method 5 "thinning out geometric partition mode set" in the reference method can be applied without using information (size information, etc.) of the target coding block.
[0068] The advantages obtained by the above difference are as follows: Even for a target coding block for which a small number of geometric partition mode candidates was not selected in the reference method, the geometric partition mode to be applied to the target coding block is determined from a small number of geometric partition mode candidates, so a further reduction in the amount of coding processing can be expected.
[0069] (Explanation of Modifications 2 to 7) Below, six geometric partition mode candidate selection methods will be described as Modifications 2 to 7, in which the coding block information to be extracted is changed from the standard method. All of the methods have the same functional configuration as the standard method, but differ in the processing content of the coding block information extraction unit 1112. Furthermore, the generation flow differs in the processing content corresponding to steps S2 and S3 in FIG. 6.
[0070] (Variation 2: Geometric division mode candidate selection method using the distance between the target frame and the reference frame) The differences between this modified example and the basic selection method are described below. The coding information block extraction unit 1112 of this modified example extracts the output order (POC: Picture Order Count) of the target frame and outputs it as coding block information. As is well known, the POC is the order in which the target frame is played back, and since frames immediately after reconstruction are not arranged in chronological order, the POC corresponds to information for rearranging them, and can be automatically determined using existing methods during coding.
[0071] Next, differences in the generation flow of this method and this modified example will be described with reference to Fig. 11. Note that steps S1 and S4 in Fig. 11 have the same processing content as steps S1 and S4 in Fig. 6, as indicated by the same reference numerals.
[0072] In step S22, the POC of the target frame is extracted, and in step S23, a plurality of geometric partition mode candidates are selected from the geometric partition mode set based on the POC of the target frame.
[0073] In this selection process, the POC of the target frame and the POC of the reference frame included in the motion information of the geometric partition mode set may be subjected to threshold processing. For example, the absolute value of the difference between the POC of the target frame and the POC of the reference frame of a certain geometric partition mode is calculated, and if the calculation result is equal to or less than a threshold, the geometric partition mode is selected as one of the elements of the geometric partition mode candidates (i.e., it is treated as not to be excluded from the candidates), and if the calculation result is greater than the threshold, it is not included in the geometric partition mode candidates (i.e., it is excluded from the candidates).
[0074] In the example of Figure 4, for the block BL to be coded and the partition boundary DB (this partition boundary DB and corresponding motion vector are, in other words, any one of the entire geometric partition mode set), the POC of the target frame F containing this block BL to be coded (referred to as POC[F]) is extracted, and the POC of the reference frames FA and FB referenced by the motion vectors MV = (MVA, MVB) (referred to as POC[FA] and POC[FB], respectively) are extracted, and the following is done using a threshold value TH. If |POC(FA)-POC(F)|>TH, the geometric division mode that refers to FA is not used (similarly for FB).
[0075] The reason for this is that when the absolute value of the difference in POC between two frames is large, it corresponds to a time series distance, and therefore in geometric segmentation that predicts object boundaries, the prediction of the moving object side is likely to be incorrect, so it is better not to select it as a geometric segmentation candidate.
[0076] (Variation 3: Method for selecting geometric partition mode candidates for a target coding block using geometric partition mode information of neighboring blocks) The differences between this modified example and the reference method are described below: The coding information block extraction unit 1112 extracts information about the geometric partitioning mode applied to neighboring blocks of the current coding block, and outputs it as coding block information.
[0077] Next, differences in the generation flow will be described with reference to Fig. 12. Note that steps S1 and S4 in the flow of Fig. 12 have the same processing contents as steps S1 and S4 in Fig. 6, as indicated by the same reference numerals.
[0078] In step S32, geometric partition mode information applied to neighboring blocks is extracted. These neighboring blocks may be selected from adjacent, coded blocks to which a geometric partition mode has been applied, using a predetermined rule. For example, the blocks adjacent to the left and above the current coding block may be selected. If no such blocks exist, information indicating that no such blocks exist is used as the geometric partition mode information applied to the neighboring blocks.
[0079] In step S33, based on the geometric division mode information applied to the neighboring blocks, multiple geometric division mode candidates are selected from the geometric division mode set using one of the following methods 31 to 33. (Method 33 corresponds to a skip process when Modification 3 cannot be applied.)
[0080] <Method 31> In this selection process, a geometric partitioning mode in which the partition boundary of the neighboring block and the partition boundary of the target coding block are determined to be continuous may be selected as a geometric partitioning mode candidate. The reason for this is as follows: Generally, many of the boundaries between objects and foreground / background in a target frame are continuous. Therefore, by selecting a geometric partitioning mode in which the partition boundary of the neighboring block and the partition boundary of the target coding block are continuous as a geometric partitioning mode candidate, it is expected that the amount of coding processing can be reduced while maintaining coding performance.
[0081] Examples EX1 and EX2 in Fig. 13 show examples in which partition boundaries are determined to be continuous. Example EX1 is an example in which a partition boundary DB for a target coding block BL is selected that is determined to be continuous because its boundary locations (completely) match with those of partition boundaries DBU, DBL applied to two neighboring blocks, an upper adjacent block BLU and a left adjacent block BLL, respectively, and similarly, example EX2 is an example in which a partition boundary DB is selected that is determined to be continuous because its boundary locations match those of adjacent partition boundaries DBU, DBL (including cases in which they do not completely match, the distance between the boundary locations is determined to be close by threshold determination).
[0082] <Method 32> Alternatively, the angle pattern φ of the division boundaries (φ, ρ) of neighboring blocks may be referenced to select a geometric division mode candidate for the current coding block.
[0083] Here, an example will be described in which candidate geometric partition modes for a target coding block are selected assuming that the angle patterns φ are similar. The similarity of angle patterns φ can be determined by, for example, whether |tanφ| is similar. That is, angle patterns φ that differ by ±180° and are opposite in direction can be treated as identical; the closer the difference in φ is to ±90°, the greater the difference in angle patterns can be treated; and the closer the difference in φ is to ±0° or ±180°, the smaller the difference in angle patterns can be treated. (In other words, if two angles φa and φb that are opposite in direction by 180° are treated as identical angles, the absolute value of the difference between the two angles φa and φb will be obtained in the range of 0° to 90°. Therefore, if this absolute value difference is equal to or less than a threshold value of 45°, for example, the angle patterns can be treated as close; otherwise, the angle patterns can be treated as distant. Similarly, the magnitude of |cos(φa-φb)| can be used to determine whether angles φa and φb are close to each other, as a method for automatically determining that angles that differ by 180° are the same.)
[0084] Alternatively, whether the angle patterns φ are similar may be determined by determining which of a predetermined range the angle pattern φ belongs to. For example, if the angle pattern of the neighboring block adjacent to the top is vertical, a geometric division mode combined with a vertical angle pattern may be selected as a geometric division mode candidate. Also, if the angle pattern of the neighboring block adjacent to the left is horizontal, a geometric division mode combined with a horizontal angle pattern may be selected as a geometric division mode candidate.
[0085] Note that here, a vertical angle pattern refers to an angle pattern within the ranges of 45° to 135° and 225° to 315°, as shown in example EX3 of Fig. 13, and a horizontal angle pattern refers to an angle pattern within the ranges of 0° to 45°, 135° to 225°, and 315° to 360°, as shown in example EX4 of Fig. 13. Note that Fig. 13 shows an example of determining close angles in the vertical direction or horizontal direction (horizontal direction), but close angles can also be determined for oblique angles that deviate from the vertical or horizontal direction in the same way.
[0086] <Method 33> Furthermore, if there is no coded block adjacent to the current coding block to which a geometric partitioning mode has been applied, the entire geometric partitioning mode set may be selected as a geometric partitioning mode candidate.
[0087] (Variation 4: Geometric partitioning mode candidate selection method using the positions of adjacent blocks referenced when deriving motion information) The differences between this modified example and the reference method are described below. The coding information block extraction unit 1112 extracts the positions of the neighboring blocks referenced when the motion information of the geometric partition mode set is derived by the merging unit 1111, and outputs the extracted positions as coding block information.
[0088] Next, differences in the generation flow will be described with reference to Fig. 14. Note that steps S1 and S4 in the flow of Fig. 14 have the same processing contents as steps S1 and S4 in Fig. 6, as indicated by the same reference numerals.
[0089] In step S42, the positions of the neighboring blocks referenced when the motion information of the geometric partition mode set was derived by the merge unit 1111 are extracted. In step S43, multiple geometric partition mode candidates are selected from the geometric partition mode set based on the positions of the neighboring blocks referenced when the motion information was derived.
[0090] Next, an example of which geometric partition mode from the geometric partition mode set is selected as a geometric partition mode candidate in this selection process will be described with reference to Fig. 15. First, for each geometric partition mode in the geometric partition mode set as shown in Fig. 15 (in Fig. 15, one geometric partition mode that provides a partition boundary DB is illustrated as the same example as in Fig. 4 above), consider the partitioned target coding block when that geometric partition mode is applied.
[0091] In this case, if the motion information MVA (see FIG. 4) to be applied to the divided left-hand region PA is derived by referring to one of the blocks adjacent to the region PA to which it is applied, such as three adjacent blocks A0, A1, and A2, and if the motion information MVB (see FIG. 4) to be similarly applied to the other region PB also refers to one of the adjacent blocks, such as two adjacent blocks B0 and B1, then this geometric partitioning mode may be selected as a candidate geometric partitioning mode. (Note that in the example of FIG. 15, only adjacent blocks in a spatial range that coincide in time with the current coding block BL are considered, but similarly, adjacent blocks in a temporal range may also be considered, or adjacent blocks in both temporal and spatial ranges may also be considered.)
[0092] As described above, the merge unit 1111 collects, for example, six pieces of motion information actually used for predictive coding in spatially and temporally adjacent regions, and outputs them as candidates (elements of the partition mode set) for motion information to be used for predictive coding of the target coding block. For example, the merge unit 1111 checks whether there is motion information applied to a block including, for example, an adjacent region A0 for the coding block BL, and if there is, sets that motion information as a candidate. Next, the merge unit 1111 checks adjacent regions A1, and so on, until six pieces of motion information are collected. Therefore, the motion information included in the partition mode set may include information that references A0, information that references B0, etc.
[0093] Therefore, in Modification 4, a method can be used in which a geometric partitioning mode in which the motion information to be applied to PA references any of A0, A1, and A2 and the motion information to be applied to PB references any of B1 and B0 may be selected as a geometric partitioning mode candidate. (That is, in Modification 4, as described above, there are comprehensively 6P2=30 possible combinations of motion information for the two parts PA and PB (two parts PA and PB divided by any one partition boundary DB), but by excluding adjacent areas B1, B0, etc. of another part PB from the candidates for part PA in this partition boundary DB, and excluding adjacent areas A0, A1, A12, etc. of another part PA from the candidates for part PB, it is possible to narrow down the motion information candidates corresponding to each of the 64 partition boundary DBs to fewer than 30.)
[0094] The reason for this will be explained. Consider a case where a target coding block BL is divided by a line segment of a division boundary DB as shown in Figure 15. In this case, it can be assumed that each divided area PA, PB has different motion MVA, MVB. Any one of the motion information MV[A0], MV[A1], MV[A2] derived from the adjacent blocks A0, A1, A2 can often be used to accurately predictively code the area PA on the left side of the target coding block BL, which is adjacent to the adjacent blocks A0, A1, A2, but it is unlikely that it can be used to accurately predictively code the area PB on the right side.
[0095] Therefore, by using the positions of the referenced adjacent blocks as described above, it is possible to select a geometric partitioning mode that allows for accurate predictive coding as a geometric partitioning mode candidate, and it is expected that the amount of coding processing can be reduced while maintaining coding efficiency.
[0096] In addition, in the example of Figure 15 above, the selection condition for the geometric division mode is that both areas PA, PB refer to one of the adjacent blocks, but it may also be that one of the areas (only one of the two divided areas PA, PB) refers to one of the adjacent blocks.
[0097] (Modification 5: Geometric Partition Mode Candidate Selection Method Using Prediction Direction of Intra Prediction) The differences between this modified example and the standard method are described below. The coding information block extraction unit 1112 extracts the reference direction of directional prediction of the current coding block in the intra prediction unit 112 and outputs it as coding block information. (Note that the current coding block is coded by the inter prediction unit 111 using the geometric partitioning mode according to this embodiment, but in modified example 5, additional processing is also applied, namely processing by the intra prediction unit 112 (processing that evaluates SAD, SSD, etc. for each reference direction and determines the reference direction that minimizes this), making it possible to narrow down candidates based on the geometric partitioning direction.)
[0098] Next, differences in the generation flow will be described with reference to Fig. 16. Note that steps S1 and S4 in the flow of Fig. 16 have the same processing contents as steps S1 and S4 in Fig. 6, as indicated by the same reference numerals.
[0099] In step S52, a reference direction for directional prediction of the current coding block in the intra prediction unit 112 is extracted. In step S53, a plurality of geometric partition mode candidates are selected from the geometric partition mode set based on the extracted reference direction.
[0100] In this selection method, the reference direction of directional prediction of the current coding block in the intra prediction unit 112 may be referenced, and a geometric partitioning mode using an angle pattern φ whose angle is close to the reference direction may be selected as a geometric partitioning mode candidate. Close angles, for example, mean that the difference between the angles is equal to or less than a threshold. The threshold may be arbitrarily specified, and may be 10°, for example. In this case, as in the case of the third modification example described above, the angle difference may be evaluated by treating opposite directions that are 180° different as the same direction (angle).
[0101] The reason for this is as follows: The reference direction for directional prediction is often determined along the edge or boundary direction of the target coding block. Therefore, by selecting a geometric partitioning mode with an angle pattern close to the reference direction for directional prediction as a geometric partitioning mode candidate, it is possible to select a geometric partitioning mode that is close to the boundary between the foreground and background or the object, which is expected to reduce the amount of coding processing while maintaining coding efficiency.
[0102] (Variation 6: Geometric Partition Mode Candidate Selection Method Using Geometric Partition Mode Information of Reference Block) The differences between this modified example and the standard method are described below: The coding information block extraction unit 1112 extracts geometric partition mode information applied to the reference block of the current coding block, and outputs it as coding block information.
[0103] Next, differences in the generation flow will be described with reference to Fig. 17. Note that steps S1 and S4 in the flow of Fig. 17 have the same processing contents as steps S1 and S4 in Fig. 6, as indicated by the same reference numerals.
[0104] In step S62, geometric partition mode information applied to the reference block is extracted. Here, the reference block is a list of all reference blocks (e.g., six or less) referenced by multiple pieces of motion information (e.g., six types) obtained for the current block by the merge unit 1111. In step S63, multiple geometric partition mode candidates are selected from the geometric partition mode set based on the geometric partition mode information (e.g., six or less types) applied to the reference block.
[0105] In this selection process, the geometric partition mode of the reference block (or the same geometric partition mode) may be selected as a candidate geometric partition mode for the current coding block. (That is, since the number of geometric partition modes applied to the reference block of the current coding block BL is, for example, six or less, the candidates can be narrowed down from the 64 options described in FIG. 5 to six or less. Note that the 30 options of motion information remain as the search range for optimal coding.) Alternatively, a geometric partition mode having an angle pattern similar to that of the reference block may be selected as a candidate geometric partition mode for the current coding block. "Similar angle patterns" refers, for example, to the absolute value of the difference in angle between two angle patterns being equal to or less than a threshold. This threshold may be any numeric value, such as 10°. In this case, as in Modification Example 3, when two angle patterns are opposite in orientation but 180° apart, they are considered to be the same orientation and the angle difference is evaluated.
[0106] The reason for this is as follows: Many reference blocks have pixel values close to those of the target coding block, and therefore the positions of the boundaries between the foreground and background or between objects are likely to be close. Therefore, by selecting the geometric partitioning mode of the reference block as a candidate geometric partitioning mode, it is expected that the amount of coding processing can be reduced while maintaining coding efficiency.
[0107] (Variation 7: Geometric Partition Mode Candidate Selection Method Using Most Probable Mode (MPM) Information) The differences between this modified example and the reference method are described below. The coding information block extraction unit 1112 extracts information about the MPM of the current coding block applied by the intra prediction unit 112, and outputs it as coding block information. (Note that, as with Modification 5, MPM information can be extracted by also applying the processing of the intra prediction unit 112 as additional processing to the current coding block in the geometric partition mode. It may also be possible to check whether the intra prediction unit 112 was applied to an adjacent block for actual coding, and, if so, to construct MPM information based on the mode used in that block.)
[0108] Next, differences in the generation flow will be described with reference to Fig. 18. Note that steps S1 and S4 in the flow of Fig. 18 have the same processing contents as steps S1 and S4 in Fig. 6, as indicated by the same reference numerals.
[0109] In step S72, MPM information (consisting of directional information, as in the case of normal intra prediction) of the current coding block is extracted. In step S73, multiple geometric partition mode candidates are selected from the geometric partition mode set based on the MPM information.
[0110] In this selection process, depending on the MPM information, one or more geometric division modes associated in advance by the MPM information may be selected as geometric division mode candidates. (That is, one or more geometric division modes whose angle φ is determined to be close to one direction indicated by the MPM information may be selected as candidates. In this case, as in Modification 3, in the case of opposite directions that differ by 180 degrees, the angle difference may be evaluated as the same angle.) Also, a geometric division mode with high coding efficiency may be estimated using the MPM information.
[0111] The reason for this is as follows. The MPM information is information on multiple prediction modes that the intra prediction unit 112 estimates are likely to be ultimately applied to the current coding block. Therefore, it is considered that the MPM information often includes information on the boundary and texture direction of the current coding block. Therefore, it is considered that by using the MPM information, it is possible to select and estimate a geometric partitioning mode with high coding efficiency. Therefore, by selecting the selected and estimated geometric partitioning mode as a geometric partitioning mode candidate, it is expected that the amount of coding processing can be reduced while maintaining coding efficiency.
[0112] The above-described base method and each of Modifications 1 to 7 may be combined. For example, the base method and Modification 2 may be combined to select a geometric partitioning mode candidate taking into consideration the size of the current coding block and the distance between the current frame and the reference frame.
[0113] As described above, according to each embodiment of the present invention, rather than using the entire pre-defined geometric partitioning mode set as candidates, a small number of geometric partitioning mode candidates are determined from the entire set, which is expected to reduce the amount of coding processing without degrading coding performance compared to determining the geometric partitioning mode to be applied to the target coding block from the entire geometric partitioning mode set.
[0114] Various supplementary examples, alternative examples, additional examples, etc. will be described below.
[0115] (1) As an example of use of the image encoding device 100 according to an embodiment of the present invention, it can be used to encode video transmitted to a remote location during a remote video conference. Since the encoding process can be accelerated while maintaining video quality, delays during the conference can be reduced, contributing to the realization of a more realistic teleconference. This enables smooth remote communication and eliminates the need for users to travel to a remote location for a conference (travel to a remote location is not necessarily required). Therefore, this use example of the embodiment of the present invention can reduce carbon dioxide emissions by saving the energy resources required for user travel, thereby contributing to Goal 13 of the United Nations' Sustainable Development Goals (SDGs), which states, "Take urgent action to combat climate change and its impacts."
[0116] (2) FIG. 19 is a diagram showing an example of the hardware configuration of a general computer device 60. The image encoding device 100 and the image decoding device 200 of this embodiment can be realized as one or more computer devices 60 having such a configuration. When each device is realized using two or more computer devices 60, information required for processing may be transmitted and received via a network. The computer device 60 includes a dedicated processor 60 that is a processor specialized for video encoding and decoding processes, a CPU (Central Processing Unit) 61 that executes predetermined instructions, a GPU (Graphics Processing Unit) 62 that is a dedicated processor that executes some or all of the CPU 61's execution instructions in place of or in cooperation with the CPU 61, a RAM 63 that serves as a main storage device providing a work area for the CPU 61 (and the dedicated processor 60 and GPU 62), a ROM 64 that serves as an auxiliary storage device, a communication interface 65, a display 66, an input interface 67 that accepts user input via a mouse, keyboard, touch panel, etc., and a bus BS for transmitting and receiving data among them.
[0117] Each functional unit of the image encoding device 100 and the image decoding device 200, and the encoding method and decoding method executed by each device can be realized by all or part of a dedicated processor 60, a CPU 61, and a GPU 62 that reads and executes a predetermined program corresponding to the function of each unit or step from a ROM 64. Note that the dedicated processor 60, the CPU 61, and the GPU 62 are all types of arithmetic devices (processors). Here, when display-related processing is performed, a display 66 also operates in conjunction with the dedicated processors, and when communication-related processing related to data transmission and reception is performed, a communication interface 65 also operates in conjunction with the dedicated processors. [Explanation of symbols]
[0118] 100... image encoding device, 111... inter prediction unit, 1111... merge unit, 1112... encoding block information extraction unit, 1113... geometric partition mode candidate selection unit, 1114... predicted pixel generation unit
Claims
1. In an image coding device that codes video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting information on the encoding target block information including information on the output order of a target frame including the encoding target block and information on the output order of a reference block of the encoding target block; An image encoding device characterized by determining the partial set by excluding, from the combinations of geometric division mode candidates and motion information candidates that constitute the entire geometric division mode set, those that are determined to have a large difference between the output order of reference blocks of the motion information candidates in the first and / or second parts divided by each of the geometric division modes and the output order of the block to be encoded.
2. In an image coding device that codes video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting motion information that has been coded by inter prediction in an adjacent block that is space-time adjacent to the current block as motion information candidate to be applied to the current block; An image encoding device characterized by determining the partial set by excluding, from all of the motion information candidates corresponding to each geometric division mode among the combinations of geometric division mode candidates and motion information candidates that constitute the entire geometric division mode set, motion information that has already been applied to adjacent blocks adjacent in time and space to the first part for the first and second parts divided by each geometric division mode from candidates for motion information to be applied to the second part, and excluding motion information that has already been applied to adjacent blocks adjacent in time and space to the second part from candidates for motion information to be applied to the first part.
3. In an image coding device that codes video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting motion information that has been coded by inter prediction in an adjacent block that is space-time adjacent to the current block as motion information candidate to be applied to the current block; extracting a reference block to which the motion information refers; An image encoding device characterized by determining the partial set by limiting the geometric partitioning mode candidates, among the combinations of geometric partitioning mode candidates and motion information candidates that constitute the entire geometric partitioning mode set, to only those geometric partitioning modes that have already been applied to the reference block.
4. In an image coding device that codes video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting a most probable mode when applying intra prediction to the encoding target block as an additional process; extracting information on the extracted most probable mode into the encoding target block information; An image encoding device, characterized in that the partial set is determined as a set of geometric division modes whose angle is determined to be closest to the direction of the most probable mode.
5. 1. An image coding method for coding video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting information on the encoding target block information including information on the output order of a target frame including the encoding target block and information on the output order of a reference block of the encoding target block; An image encoding method characterized by determining the partial set by excluding, from the combinations of geometric division mode candidates and motion information candidates that constitute the entire geometric division mode set, those that are determined to have a large difference between the output order of the reference blocks of the motion information candidates in the first and / or second parts divided by each of the geometric division modes and the output order of the block to be encoded.
6. 1. An image coding method for coding video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting motion information that has been coded by inter prediction in an adjacent block that is space-time adjacent to the current block as motion information candidate to be applied to the current block; An image encoding method characterized by determining the partial set by excluding, from all of the motion information candidates corresponding to each geometric division mode among the combinations of geometric division mode candidates and motion information candidates that constitute the entire geometric division mode set, motion information that has already been applied to adjacent blocks adjacent in time and space to the first part for a first and second part divided by each geometric division mode from candidates for motion information to be applied to the second part, and excluding motion information that has already been applied to adjacent blocks adjacent in time and space to the second part from candidates for motion information to be applied to the first part.
7. 1. An image coding method for coding video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting motion information that has been coded by inter prediction in an adjacent block that is space-time adjacent to the current block as motion information candidate to be applied to the current block; extracting a reference block to which the motion information refers; An image encoding method characterized by determining the partial set by limiting the geometric partitioning mode candidates, among the combinations of geometric partitioning mode candidates and motion information candidates that constitute the entire geometric partitioning mode set, to only those geometric partitioning modes that have already been applied to the reference block.
8. 1. An image coding method for coding video by applying block-by-block prediction, When handling a block to be coded to which inter prediction and geometric partitioning are applied as prediction modes, coding target block information, which is information about the block to be coded, is extracted, and only a partial set selected from the entire set of geometric partitioning modes predefined in the geometric partitioning based on the coding target block information is used as an evaluation target to be used for coding the block to be coded after applying inter prediction and geometric partitioning to the block to be coded; extracting a most probable mode when applying intra prediction to the encoding target block as an additional process; extracting information on the extracted most probable mode into the encoding target block information; An image coding method, comprising determining the subset as a set of geometric division modes whose angles are determined to be closest to the direction of the most probable mode.
9. 5. A program that causes a computer to function as the image coding device according to claim 1.
Citation Information
Patent Citations
Geometric partition mode with increased efficiency
US20210152825A1
Simplified inter prediction with geometric partitioning
WO2021104433A1
Intra prediction with geometric partition
WO2022106281A1