Method of increasing the temporal resolution of a video sequence
By combining multiple temporal interpolation methods and optimizing weighting coefficients based on object classes, the method addresses the limitations of existing techniques in increasing video sequence temporal resolution, achieving adaptable and high-quality results across diverse video content.
Patent Information
- Application Number
- FR2023014808
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for increasing the temporal resolution of video sequences are limited by their dependence on specific training data domains and struggle with fast-moving objects, varying camera parameters, and different video encoders, making them unsuitable for diverse video content.
A method that combines the results of multiple temporal interpolation methods using a linear combination, where each image is decomposed into segments based on object classes, and weighting coefficients are optimized to minimize the error between the input and reconstructed images, allowing for adaptive optimization across different video sequences.
This approach enables the generation of high-quality interpolated images that are independent of specific video sequence types, effectively improving the detection and identification of objects in diverse scenes by optimizing the combination of interpolation methods based on object classes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for increasing the temporal resolution of a video sequence
[0001] The invention relates to the field of video processing and more specifically relates to a method for increasing the temporal resolution of a video sequence. The invention can be applied in different fields which require a video sequence with high temporal resolution, in particular to analyze details in images, identify events having a high frequency occurrence or more generally to improve the detection and identification of objects in a scene.
[0002] Generating a high temporal resolution video sequence from a low resolution sequence has the advantage of limiting the quantity of data to be stored while allowing, as required, to benefit from an augmented version of the sequence.
[0003] Recent approaches based on convolutional neural networks (CNNs), trained with large amounts of high temporal resolution videos, allow to obtain high quality interpolation results when the test videos are similar to those of the training data.
[0004] However, these methods may fail if the training data differs from the target domain of the useful data. For example, if the intended application is to slow down video sequences of fish moving underwater but the training data does not contain any video sequences of fish, the learning of the temporal interpolation may be imperfect.
[0005] In addition to the domain problem, slow motion generation algorithms have difficulty recovering the trajectories of fast moving objects.
[0006] Other limitations of these methods may also be related to the parameters of the shooting cameras or to the different types of video encoders as well as to other parameters of the scene such as light intensity.
[0007] In practice, it is impossible to train an artificial intelligence model to perform a temporal interpolation function that is precisely adapted to a maximum of possible scenarios in terms of video content as discussed in reference [3].
[0008] In other words, each interpolation method works more or less optimally depending on the content of the sequences, the type of moving objects or the acquisition parameters of the video sequence.
[0009] There is therefore a need to determine a method for automatically increasing the temporal resolution of a video sequence which is suitable for any type of sequence, regardless of content and acquisition parameters.
[0010] The proposed invention is based on the search for an optimal combination of the results provided by different temporal interpolation methods so as to best adapt to each video sequence.
[0011] The images interpolated via the method which is the subject of the invention are obtained by means of a linear combination of the same images interpolated via several state-of-the-art methods, each image being furthermore decomposed from a segmentation so as to associate a different weight with each class of objects contained in an image.
[0012] The weighting coefficients of the linear combination are determined by optimization so as to minimize a distance between an input image and the same image reconstructed via the method.
[0013] Taking into account different weighting coefficients according to the classes of objects contained in the images makes it possible to differentiate the respective impacts of the interpolation methods according to the type of objects and to obtain the most optimal combination of the results provided by these different methods.
[0014] In other words, certain interpolation methods can produce more accurate results for certain classes of objects rather than others, the invention exploits this observation to determine a more optimal version of the interpolated images.
[0015] The subject of the invention is a method, implemented by computer, for increasing the temporal resolution of a video sequence, comprising the steps of, for at least one set of three consecutive initial images of the video sequence: - Selecting several temporal interpolation methods and for each method: i. Apply the method to said set to generate a pair of images at two respective intermediate times between the times corresponding to the first two images of the set and the last two images of the set, ii. Apply the method to the pair of generated images to determine a reconstruction of the second image of said set, - Apply several different segmentation masks to each reconstruction of the second image to decompose it into a sum of components, - Determine a final reconstruction of the second image of said set as equal to a function of the reconstructions of the second image, the function being defined by a linear combination weighted by a set of weighting coefficients, of said components for the set of segmentation masks and the set of interpolation methods, - The set of weighting coefficients being determined by the minimization of a predetermined cost function representative of an error between the final reconstruction of the second image and the second initial image, - Determine a final version of the pair of images at the two intermediate times, each image of the pair being taken equal to said function applied to the respective images generated by all the interpolation methods.
[0016] According to a particular aspect of the invention, the segmentation masks are obtained by applying a regular grid to an image, each mask being defined by at least one box of the grid.
[0017] According to a particular aspect of the invention, the segmentation masks are obtained by applying a segmentation method to at least one of the initial images.
[0018] In an alternative embodiment, the method according to the invention comprises: - Applying the segmentation method to the three consecutive initial images of the video sequence to generate a first set of initial segments for each of the images, - For each object resulting from the segmentation and present in at least two of said images, define a convex hull which encompasses the representations of the object in said images, - Construct a set of final segments including: i. The segments of the first set not belonging to a convex hull, ii. A new set of segments defined by the boundaries of convex hulls and intersections between several convex hulls.
[0019] According to a particular aspect of the invention, the segmentation masks are obtained by applying a segmentation method to at least one of the pairs of images obtained using at least one temporal interpolation method.
[0020] In an alternative embodiment, the method according to the invention comprises: - Applying the segmentation method to one of the pairs of images obtained using at least one temporal interpolation method and to the second initial image, to generate a first set of initial segments for each of the images, - For each object resulting from the segmentation and present in at least two of said images, define a convex hull which encompasses the representations of the object in said images, - Construct a set of final segments including: i. The segments of the first set not belonging to a convex hull, ii. A new set of segments defined by the boundaries of convex hulls and intersections between several convex hulls.
[0021] In an alternative embodiment, the method according to the invention further comprises: - The selection of the temporal interpolation method which makes it possible to minimize said cost function applied between the second image and the reconstruction of the second image, - Application of the segmentation method to the pair of images obtained via the selected method.
[0022] According to a particular aspect of the invention, the segmentation masks are selected so as to minimize the cost function applied between the final reconstruction of the second image and the second image.
[0023] According to a particular aspect of the invention, the cost function is a distance between two images.
[0024] According to a particular aspect of the invention, at least one temporal interpolation method involves the implementation of a machine learning model of a temporal interpolation function, the model being previously trained from training data consisting of video sequences belonging to a given application domain.
[0025] The invention also relates to a computer program comprising code instructions for implementing the method according to the invention, when said program is executed on a computer and a computer-readable recording medium on which the computer program according to the invention is recorded.
[0026] Other characteristics and advantages of the present invention will appear more clearly on reading the description which follows in relation to the following appended drawings.
[0027] [Fig.l] represents a flowchart detailing the steps of implementing a method for increasing the temporal resolution according to an embodiment of the invention,
[0028] [Fig.2] represents an illustrative diagram of a step of the method according to the invention,
[0029] [Fig.3] represents a diagram illustrating a first step of generation of a segmentation mask according to a particular embodiment of the invention,
[0030] [Fig.4a] represents a diagram illustrating a second step of generating a segmentation mask according to a particular embodiment of the invention,
[0031] [Fig.4b] represents a diagram illustrating a third step of generating a segmentation mask according to a particular embodiment of the invention.
[0032] [Fig.l] represents a flowchart of the method of increasing the re- temporal solution according to one embodiment of the invention.
[0033] The method according to the invention receives as input a sequence of images at a given temporal resolution. It aims to determine, for each set of three consecutive images I(t0), I(0), I(t2) of the image sequence taken at three consecutive times t0, tb t2, two new images I(ta) and I(tb) at the respective times ta and tb such that t0 < ta < ti and ti < tb < t2. In other words, it is a question of increasing the temporal resolution of the sequence by adding an additional image between each pair of consecutive images.
[0034] Many temporal interpolation methods can be envisaged, according to the prior art, to achieve this increase in resolution.
[0035] Among the known methods, we can cite those described in references [1], [2] and [4],
[0036] References [1] and [2] describe methods based on machine learning techniques that involve convolutional neural networks. In other words, these methods propose to train artificial intelligence models to learn how to perform this temporal interpolation function.
[0037] As indicated in the preamble, these methods have the disadvantage that the trained model only works optimally for video sequences in the same domain as those used for training.
[0038] Other methods also have limitations for video sequences that include high-speed moving objects.
[0039] Reference [4] describes a method based on depth maps (3D images).
[0040] Reference [5] proposes a method based on a reconstruction of the trajectories of high-speed objects from auxiliary information.
[0041] Reference [6] further describes a method which is applicable to different spatial resolutions.
[0042] Reference [7] relates to an interpolation method based on unsupervised training.
[0043] All the methods described above, in particular those described in references [1], [4], [6] and [7] can be used to generate the images I(ta) and I(tb).
[0044] The method according to the invention thus begins at step 101 by executing several temporal interpolation methods, among the existing methods which include those described above, to generate N pairs of images {îj(ta), îj(tb)} corresponding to N different versions of these two images with N a strictly positive integer and j varying between 1 and N. These N different methods have the aforementioned drawbacks and each provide more or less optimal results depending on the nature of the video sequence. The N methods used can be based on different learning architectures or on identical architectures but data different learning techniques or a combination of these two techniques.
[0045] In step 102, the N interpolation methods are again applied to generate a first reconstruction of the central image îj(ti) at time ti of the triplet of images from each pair of images {îj(ta), îj(tb)}.
[0046] In other words, if the initial video sequence is produced at a regular temporal rate with a period equal to 1, then we have t0=0, ti= 1, t2=2. If we denote ta=t with t between 0 and 1, then we have tb=t+1. By applying the same model used to generate the intermediate images at times ta and tb, that is to say at a time t in the time interval between two consecutive images, then by applying the same model to the two images {îj(ta), îj(tb)} to generate a new image at a time 1-t between times ta and tb, we must obtain an image corresponding to the initial image at time tb. More generally, the temporal rate of sampling of the images can be different from 1, for example equal to a value Tp. In this case, t varies between 0 and Tp, tb=t+Tp and the new image generated from the pair {îj(ta), îj(tb)} corresponds to an instant Tp-t.
[0047] This principle, called cyclic consistency, is presented in reference [7] and is illustrated in [Fig.2] by designating Modj the trained model or more generally the method used to carry out the temporal interpolation function, steps 101 and 102 are translated by the following relations:
[0048] îj(ta)= Modj(I(to), I(tû)
[0049] îj(tb)= Mod / I^û, I(t2))
[0050] ^(0)= Modj(îj(ta), îj(tb))
[0051] This principle therefore makes it possible to generate, for each of the interpolation methods, a first reconstruction of the original image îj(ti).
[0052] In step 103, a segmentation step is then applied to decompose each image îj(ti) into several components.
[0053] According to a first embodiment, the segmentation step 103 consists of a simple regular grid, each segment corresponding to one or more boxes of the grid.
[0054] According to another embodiment, the segmentation 103 is applied by means of a segmentation method such as that described in reference [8]. This segmentation method aims to detect objects in an image, associate a class with each object and decompose the image into a sum of components, each component comprising a class of objects of the image.
[0055] In this way, each image I can be represented as a sum of products between this image and respective segmentation masks Mk. [°056]
[0057] K denotes the number of masks equal to the number of segments obtained via the seg- nutrition.
[0058] Mk denotes a mask associated with the segment k, which is such that Mk=1 for the pixels of the image belonging to the segment and Mk=0 otherwise.
[0059] A constraint is that the set of segments covers the entire image, in other words that there is no part of the image that is not included in a segment.
[0060] To determine the set of segmentation masks Mk, several methods are possible.
[0061] For example, the segmentation masks are obtained by applying a segmentation method to one or more of the initial images I(t0), I(0), I(t2).
[0062] According to another example, the segmentation method is applied to a pair of images {îj (ta), îj (tb)} provided by any one of the interpolation methods j' among the N available methods.
[0063] For example, the interpolation method j' selected is the one for which a distance between the original image 1(0) and the reconstructed image îj (0) obtained in step 102 is the smallest.
[0064] The cost function chosen to calculate the distance between these two images can be a pixel-to-pixel distance or a more elaborate cost function such as that proposed in reference [7].
[0065] In other words, the search for the interpolation method j' which gives the best temporal interpolation result consists of minimizing a cost function L: min 1 j-LV
[0066] In the two examples above, the potential movement of the objects detected by the segmentation between the different images is not taken into account.
[0067] To improve the generation of segmentation masks by taking these movements into account, another embodiment of the segmentation step 103 is proposed and described in an example in FIGS. 3 and 4.
[0068] In this embodiment, the possible movement of the segmented objects between three images denoted Fb F2 and F3 is taken into account.
[0069] In a first variant of this embodiment, the three FB images F2 and F3 are the three original images Fi = I(t0), F2 = 1(0), F3 = I(t2).
[0070] In a second variant of this embodiment, the three images are the images Fi = îj (ta), F2 = 1(0) and F3 = îj (tb), where the image pair {îj (ta), îj (tb)} is provided by any one of the interpolation methods j' among the N available methods as explained above. An advantage of this second variant is that the images {îr (ta), îj (tb)} are closer in time to the central image 1(0) which makes it possible to minimize the impact of excessive movements on the determination of the segmentation masks.
[0071] Whatever the embodiment variant chosen, the segmentation method is first applied to each of the FH images F2 and F3 and the segments corresponding to the same objects on the three images are identified. The same object may of course not be present on all three images but only on two images or just one.
[0072] We note K(Fi), K(F2), K(F3) respectively the number of segments obtained for each of the three images.
[0073] We also note Kmax as the maximum value among these three numbers.
[0074] We then note s;(Fi), s;(F2), s;(F3), the representations, in the three images, of each of the segments s; for i varying from 1 to Kmax. In other words, it is the set of pixels of each image which correspond to the segment Sj.
[0075] As illustrated in an example in [Fig.3] in which the three images Fi, F2, F3 are shown diagrammatically, the same object corresponding to a segment 301 can change position between the three successive images.
[0076] To take this phenomenon into account, a convex envelope 302 is defined (for example using the method proposed in reference [9]) which encompasses the segment 301 according to its three positions in the image.
[0077] Then, when several envelopes have a non-zero intersection, as shown in [Fig.4a] for the case of two envelopes 401,402, the final segments retained are determined from the boundaries between the intersection zones of the different envelopes. In the example of [Fig.4b], five segments 411,412,413,414,415 are thus defined from the two initial envelopes 401,402. A sixth segment 410 is added corresponding to the background of the image which has not been classified.
[0078] Generally speaking, the final set of segments obtained is composed of: - segments s; which do not belong to any envelope, i.e. segments corresponding to objects which only appear in one of the three images, - of a new set of segments whose boundaries are defined by the boundaries of the envelopes and by the intersections between several envelopes as illustrated in [Fig.4b].
[0079] At the end of step 103, we therefore obtain a set of segmentation masks Mk for k varying from 0 to K-1.
[0080] Then, in step 104, a final reconstruction of the image 1(0) is determined as a weighted sum of the N images obtained in step 102 and decomposed using the segmentation masks obtained in step 103.
[0081] In other words, the final reconstruction of the image 1(0) is obtained via the following relation:
[0082] â x vA7 ■ (1) / (fi)
[0083] The coefficients ajk depend on both the interpolation method of index j and the segmentation mask of index k. They are positive or negative and respect the following relation: yN , for each value of k.
[0084] The coefficients ajk are determined in step 105 so as to minimize a distance between the image reconstructed by means of relation (1), called final reconstruction and the original image I(ti) available in the input sequence.
[0085] The distance chosen is a cost function identical to that described above, i.e. a simple pixel-to-pixel distance or the cost function proposed in reference [7].
[0086] The search for coefficients that minimize the cost function is carried out using an optimization algorithm >LV[ \ " / J digital, for example a least squares algorithm or any suitable algorithm.
[0087] Finally, in step 106, the final versions of the images I(ta) and I(tb) are determined from the following relationships:
[0088]
[0089] I {fa) j ^jk^k O j ( fa ) tb) = ï^^^jk^k 0 7 / th) (2) (3)
[0090] The interpolated images obtained in step 106 thus correspond to an optimal combination of the respective versions of these images obtained by the N methods. Taking into account the segmentation masks in the combination makes it possible to give greater weight to the interpolation methods depending on the types of objects for which they are most suitable.
[0091] Thus, the proposed method makes it possible to obtain results independent of the types of video sequence since it automatically combines the particular advantages of each method of the prior art for different types of objects or classes of objects.
[0092] In an alternative embodiment of the invention, the segmentation masks Mk can also be determined, jointly with the coefficients ajk by minimizing the cost function
[0093] The invention may be implemented as a computer program comprising instructions for its execution. The computer program may be recorded on a recording medium readable by a processor.
[0094] Reference to a computer program that, when executed, performs any of the functions described above, is not limited to an application program running on a single host computer. Rather, the terms computer program and software are used herein in a general sense to refers to any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be used to program one or more processors to implement aspects of the techniques described herein. The computing means or resources may in particular be distributed (“Cloud computing”), possibly using peer-to-peer technologies. The software code may be executed on any suitable processor (e.g., a microprocessor) or processor core or a set of processors, whether provided in a single computing device or distributed among several computing devices (e.g., as may be accessible in the device environment). The executable code of each program enabling the programmable device to implement the processes according to the invention may be stored, for example, in the hard disk or in read-only memory.Generally, the program(s) may be loaded into one of the storage means of the device before being executed. The central unit may control and direct the execution of the instructions or portions of software code of the program(s) according to the invention, instructions which are stored in the hard disk or in the read-only memory or in the other aforementioned storage elements.
[0095] The invention can be implemented on a computing device based, for example, on an embedded processor. The processor can be a generic processor, a specific processor, an application-specific integrated circuit (also known as an ASIC for "Application-Specific Integrated Circuit") or an in situ programmable gate network (also known as an FPGA for "Field-Programmable Gate Array"). The computing device can use one or more dedicated electronic circuits or a general-purpose circuit. The technique of the invention can be carried out on a reprogrammable computing machine (a processor or a microcontroller for example) executing a program comprising a sequence of instructions, or on a dedicated computing machine (for example a set of logic gates such as an FPGA or an ASIC, or any other hardware module).
[0096] The invention can be applied to any type of video sequence, for example in the field of the entertainment industry for the production of films or documentaries but also for the sports industry for the automatic analysis of sports content or even in the field of the production of augmented reality content. References
[0097] [1] Huaizu Jiang, Deqing Sun, Varun Jampani, Ming-Hsuan Yang, Erik Learned- Miller, and Jan Kautz. Super slomo: High quality estimation of multiple intermediate frames for video interpolation. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
[0098] [2] Simon Niklaus, Long Mai, and Feng Liu. Video frame interpolation via adaptive separable convolution. IEEE International Conference on Computer Vision (ICCV), 2017.
[0099] [3] Hidetoshi Shimodaira. Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of statistical planning and inference, 90(2):227-244,2000
[0100] [4] Wenbo Bao, Wei-Sheng Lai, Chao Ma, Xiaoyun Zhang, Zhiyong Gao, Ming- Hsuan Yang. Depth-Aware Video Frame Interpolation. IEEE Conférence on Computer Vision and Pattern Récognition, Long Beach, CVPR, 2019
[0101] [5] Avinash Paliwal, Nima Khademi Kalantari. Deep Slow Motion Video Recons truction with Hybrid Imaging System. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) and ICCP, 2020
[0102] [6] Fitsum Reda, Janne Kontkanen, Eric Tabellion, Deqing Sun, Caroline Pantofaru, Brian Curless. ”FILM: Frame Interpolation for Large Motion”. ECCV, 2022
[0103] [7] Fitsum A. Reda, Deqing Sun, Aysegul Dundar, Mohammad Shoeybi, Guilin Liu, Kevin J. Shih, Andrew Tao, Jan Kautz, Bryan Catanzaro, ”Unsupervised Video Interpolation Using Cycle Consistency”. ICCV, 2019.
[0104] [8] Shervin Minaee, Yuri Boykov, Fatih Porikli, Antonio Plaza, Nasser Kehtarnavaz, and Demetri Terzopoulos, "Image Segmentation Using Deep Leaming: A Survey”, arXiv:2001.05566 [cs.CV], 2020.
[0105] [9] Andrew, A. M., ”Another efficient algorithm for convex hulls in two di mensions”, Information Processing Letters, 9 (5): 2161s in two
Claims
1. Claims A computer-implemented method for increasing the temporal resolution of a video sequence, comprising the steps of, for at least one set of three consecutive initial images (I(t0), I(ti), I(t2)) of the video sequence: - Select several time interpolation methods and for each method: i. Apply (101) the method to said set to generate a pair of images (îj(ta), îj(tb)) at two respective intermediate times between the times corresponding to the first two images of the set and to the last two images of the set, ii. Apply (102) the method to the pair of images (îj(ta), îj(tb)) generated to determine a reconstruction (îj (ti)) of the second image of said set, - Apply (103) several different segmentation masks to each reconstruction (îj(ti)) of the second image to decompose it into a sum of components, - Determine (104) a final reconstruction (a A of the w) second image of said set as equal to a function of the reconstructions (îj(ti)) of the second image, the function being defined by a linear combination weighted by a set of weighting coefficients, of said components for the set of segmentation masks and the set of interpolation methods, - The set of weighting coefficients being determined (105) by the minimization of a predetermined cost function representative of an error between the final reconstruction (aj of the second image and the second initial image (I(ti)), - Determine (106) a final version of the image pair ( A / x ) at the two intermediate instants, each image of the couple being taken equal to said function applied to the respective images generated by all the methods of interpolation.
2. A method of increasing the temporal resolution of a video sequence according to claim 1 wherein the segmentation masks are obtained (103) by applying a regular grid to an image, each mask being defined by at least one box of the grid.
3. Method for increasing the temporal resolution of a video sequence according to claim 1 in which the segmentation masks are obtained (103) by applying a segmentation method to at least one of the initial images ( I(t0), 1(0), I(t2) ).
4. A method for increasing the temporal resolution of a video sequence according to claim 3 comprising: - Applying the segmentation method to the three consecutive initial images (I(t0), 1(0), I(t2)) of the video sequence to generate a first set of initial segments for each of the images, - For each object resulting from the segmentation and present on at least two of said images, defining a convex hull which encompasses the representations of the object in said images, - Constructing a set of final segments comprising: i. The segments of the first set not belonging to a convex hull, ii. A new set of segments defined by the boundaries of the convex hulls and intersections between several convex hulls.
5. Method for increasing the temporal resolution of a video sequence according to claim 1 in which the segmentation masks are obtained (103) by applying a segmentation method to at least one of the pairs of images (îj(ta), îj(tb)) obtained using at least one temporal interpolation method.
6. Method for increasing the temporal resolution of a video sequence according to claim 5 comprising: - Applying the segmentation method to one of the pairs of images (îj(ta), îj(tb)) obtained using at least one temporal interpolation method and to the second image I(ti ) initial, to generate a first set of initial segments for each of the images, - For each object resulting from the segmentation and present on at least two of said images, define a convex hull which encompasses the representations of the object in said images, - Construct a set of final segments comprising: i. The segments of the first set not belonging to a convex hull, ii. A new set of segments defined by the boundaries of the convex hulls and the intersections between several convex hulls.
7. Method for increasing the temporal resolution of a video sequence according to any one of claims 5 or 6 further comprising: - The selection of the temporal interpolation method which makes it possible to minimize said cost function applied between the second image and the reconstruction of the second image, - The application of the segmentation method to the pair of images obtained via the selected method.
8. A method of increasing the temporal resolution of a video sequence according to any preceding claim wherein the segmentation masks are selected so as to minimize the cost function applied between the final reconstruction of the second image and the second image.
9. A method of increasing the temporal resolution of a video sequence according to any preceding claim wherein the cost function is a distance between two images.
10. Method for increasing the temporal resolution of a video sequence according to any one of the preceding claims in which at least one temporal interpolation method involves the implementation of a machine learning model of a temporal interpolation function, the model being previously trained from training data consisting of video sequences belonging to a given application domain.
11. A computer program comprising code instructions for implementing the method according to any one of claims 1 to 10, when said program is executed on a computer.
12. A computer-readable recording medium on which the computer program according to claim 11 is recorded.
Citation Information
Patent Citations
Region sizing for macroblocks
US8379720B2