Method and system for motion vector calculation and processing
By calculating and processing motion vector and virtual depth information, generating foreground and background MV fields, and filtering the MV along the motion trajectory, the problems of frame rate conversion and frame interpolation are solved, and a more stable and high-quality image processing effect is achieved.
Patent Information
- Application Number
- CN202411544872.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively process and reduce the void caused by reprojection and extrapolation during frame rate conversion and frame interpolation, resulting in unstable output.
By calculating and processing motion vector (MV) and virtual depth information, the foreground and background MV fields are generated and the MV is filtered along the motion trajectory to reduce the size of the hole and fill the smaller hole area.
It effectively reduces the size and number of holes during frame interpolation and frame rate conversion, improves the stability and quality of output, and reduces the use of computing resources.
Smart Images

Figure CN120075460A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the subject matter disclosed herein relate to the field of three-dimensional (3D) computer graphics and, more particularly, to motion vector calculation, processing, and filtering. Background Art
[0002] Over the years, improvements in computer processing power have made real-time video rendering (e.g., for video games or certain animations) increasingly complex. For example, early video games were characterized by pixelated sprites moving on a fixed background, while contemporary video games are characterized by realistic three-dimensional scenes filled with characters. At the same time, miniaturization of processing components has enabled mobile devices (such as handheld video game devices and smartphones) to effectively support real-time rendering of high frame rate, high resolution video.
[0003] 3D graphics video can be output at various different frame rates and screen resolutions. It may be necessary to convert a video with 3D graphics from one frame rate (and / or resolution) to another. To save computing power while still increasing the frame rate, interpolated frames can be used instead of rendering all frames within the video. Interpolated frames can be effectively generated by using motion vectors (also referred to herein as MVs), which track the position differences of objects between the current frame (CF) and the previous frame (PF). Summary of the Invention
[0004] In one example, a method includes: receiving a plurality of image frames as input; calculating one or more internal motion vector (MV) fields between a previous frame (PF) and a current frame (CF) in the plurality of image frames; generating a foreground MV field and a background MV field of the one or more MV fields; processing the one or more internal MV fields and the foreground MV field and the background MV field to generate one or more depth fields, wherein processing the one or more internal MV fields includes generating virtual depths of the one or more internal MV fields to generate one or more MVD fields; and outputting the one or more MVD fields for image processing.
[0005] It should be understood that the above brief description is provided to introduce some concepts that are further described in the detailed description in a simplified form. It is not meant to identify key or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims that follow the detailed description. Moreover, the claimed subject matter is not limited to embodiments that solve any disadvantages noted above or in any part of this disclosure. Brief Description of the Drawings
[0006] Figure 1 A block diagram of an example computing system is shown.
[0007] Figure 2Shows a block diagram of an example image processing system.
[0008] Figure 3 Shows a flowchart illustrating a method for motion vector calculation, processing, and filtering.
[0009] Figure 4 Shows a flowchart illustrating a method for motion vector calculation.
[0010] Figure 5 Shows a flowchart illustrating a method for foreground and background motion vector detection.
[0011] Figure 6 Shows a flowchart illustrating a method for motion vector post-filtering.
[0012] Figure 7A Shows the first part of a flowchart illustrating a method for motion vector filtering to reduce holes for inpainting.
[0013] Figure 7B Shows the Figure 7A second part of a flowchart of a method for motion vector filtering to reduce holes.
[0014] Figure 8 Shows a first diagram of a motion vector field.
[0015] Figure 9 Shows a second diagram of a motion vector field.
[0016] Figure 10 Shows a third diagram of an unfiltered motion vector field and a corresponding filtered motion vector field.
[0017] Figure 11 Shows for Figure 2 an image processing system use case scenario. Detailed Description
[0018] Systems and methods for calculating, processing, and filtering a motion vector (MV) field for use in intra-frame interpolation, frame rate conversion, hole filling reduction, or other operations are described herein. In Figure 1 a block diagram depicts a computing system equipped with an image processing system. In Figure 2 a block diagram depicts an exemplary image processing system. In Figure 3 a flowchart illustrates a method for MV calculation, processing, and filtering. In Figure 4 a flowchart illustrates a method for MV calculation. In Figure 5 a flowchart illustrates a method for detecting foreground MV and background MV. In Figure 6The method for MV processing is illustrated in the flowchart in. The method for reduced MV filtering for hole filling is illustrated in the flowchart in FIG. 7. In Figure 8 a first example diagram of the MV field is shown. In Figure 9 a second example diagram of the MV field is shown. In Figure 10 a third example diagram of the MV field and the filtered MV field is shown.
[0019] An MV field can be generated that tracks the position differences of objects moving forward in time, such as the position differences between a previous frame (PF) and a current frame (CF). As explained herein, two types of MVs are generated, MV 0 represents the motion from PF to CF, and MV 1 represents the motion from CF to PF. An MV is generated for each pixel (or group of pixels called a block) on the screen, thus forming an MV field or a set of MVs of the pixels on the screen. As used herein, a field is defined as a mapping from a set of pixels in a frame to a set of one or more numbers (e.g., components of a vector or a single number).
[0020] If MVs are used during frame rate conversion, a typical rendering engine only outputs a two-dimensional MV 1 texture. In this way, the texture does not contain depth content and only includes information about the change in relative screen position as observed in the reference frame of the virtual camera. The depth and foreground and background information of the per-pixel MV can inform how to calculate the 2D components of the block-level MV. The block-level MV can represent the average or weighted average of the MVs of a pixel block (e.g., an eight-by-eight pixel block) and can be used for frame interpolation or other image processing tasks to reduce processing requirements. The block-level MV can be converted back to the pixel-level MV to allow other actions or uses based on virtual depth. Regions of a scene with certain relative depth ranges are called foreground (close to the camera), background (far from the camera), and intermediate range (between the foreground and the background).
[0021] As an example, two objects can be located at different distances from a virtual camera or viewpoint. If the two objects move in the same direction at an equal world space distance, the farther object may appear to move a smaller distance in eye space, thus creating a parallax effect where the object farther from the viewpoint appears to move shorter than the object closer to the viewpoint.
[0022] Additionally, each frame can consist of two types of objects: objects with an MV and objects without an MV. Objects characterized by an MV can include moving characters or other objects, the user's view (virtual camera position), and moving parts of the user interface (for a game, this can be a health bar or other similar in-game graphics or statistics). Objects without an MV can include, for example, smoke effects, lighting effects, reflections, full-screen or partial-screen scene transitions (e.g., fades and wipes), and / or particle effects. By separating objects with an MV from objects without an MV, improved image processing can be performed. Traditionally, algorithms may attempt to exclude screen areas characterized by objects without an MV. However, this method is not perfect and may result in the blending of nearby objects during the process of frame rate conversion and / or frame interpolation.
[0023] Rendering images in real time on a computer system typically involves the calculation and ordered blending of multiple virtual layers. Examples of such layers include a bottom black background layer, which can include some or all of the black areas on the screen. Applications and games can be drawn above the black background to make them visible. Note that some layers may be opaque, fully transparent, or partially transparent. The transparency of the pixels within a given layer can be represented by its alpha mask. Above the internal parts of the game / application, a graphical user interface (GUI) for the operating system can be drawn. Traditionally, the parts of the operating system and the information from the game / application can be blended together from bottom to top. The blended image can be directly used for display.
[0024] In addition to frame interpolation, typical methods for frame rate conversion to increase the display frame rate also include extrapolation and reprojection. Reprojection is a general method for accelerating real-time rendering by reusing the rendered pixels of adjacent frames. This involves obtaining one or more previously rendered frames and using newer motion information to extrapolate the previous frames into a prediction of what the normally rendered frames would look like. Although reprojection is useful, it cannot produce a perfect output. For example, if previously hidden content is not covered in the new frame, holes will appear due to the lack of available pixel information in the source image. As an example, when an object with an MV moves between the PF and the CF, the background pixels may not be covered at the location where the object was in the PF but is no longer in the CF. These pixel areas cannot be correctly filled by reprojection alone, and thus there are various repair algorithms to fill such holes, such as copying the pixel values of the background texture to fill the holes, selecting random neighboring pixels to fill the holes, applying a low-pass filtered intermediate image to fill the holes, etc. However, when the uncovered holes are large, the resulting output of the repair algorithms may look unstable.
[0025] Similarly, extrapolation is a method for using existing data (such as MV and MVD data) to estimate future data. In image processing and particularly in video game processing, the MV calculated between the PF and the CF can be used for extrapolation. However, as a result, large holes may appear in the resulting image.
[0026] As discussed herein, frame interpolation refers to a systematic process of generating one or more additional frames from the visual information of a previous frame PF and a current frame CF. The additional frames are referred to as interpolated frames or IFs. The IFs can be sequentially displayed between the PF and the CF, with a uniform spacing between them in some instances and a non-uniform spacing between them in other instances. As an example, frame rate conversion via frame interpolation can be used to generate one IF between a given pair of PF / CF, resulting in an output with a frame rate that is twice that of the original video. The generation of the IFs may be desirable due to its low computational cost, as generating more frames from the game engine itself may require a higher degree of computational complexity and thus more processing power.
[0027] For video games, the depth can be provided by the game engine at the pixel level and used for block-to-pixel MV conversion. However, additional information is usually required. For example, when there is a shadow cast by another object on the object, the depth of the object may be irrelevant to the motion. Generally, the lighting, ray tracing, and other special effects applied in the game make the MV unsuitable for frame interpolation and must be replaced with internally generated MV. Then, such internally generated MV needs to be converted from the block level to the pixel level. The depth provided by the game engine is also not suitable for converting the MV from the block level to the pixel level because its value does not change when the MV changes.
[0028] To accomplish frame rate conversion via one or all of frame interpolation, extrapolation, and reprojection, high-quality MV and depth information can be obtained from the game engine, but this may result in consuming more computational resources and more processing power. As described herein, a flexible architecture for frame interpolation is proposed, which allows the calculation, processing, and filtering of the MV from the game engine to reduce the use of computational resources for image processing. In some instances, the MV and depth information can be obtained from the game engine, referred to as external MV and depth, and / or in other instances, the MV and depth information can be estimated based on the input image, referred to as internal MV and depth. The external MV and depth and the internal MV and depth can be combined and / or used in combination during processing and / or filtering for frame interpolation. The internal virtual depth information can be generated based on the internal block-level MV, allowing the conversion of the MV from the block level to the pixel level.
[0029] Extrapolation and interpolation can be accomplished by calculating and processing MVs in terms of virtual depth, block level, and pixel level. The methods and systems provided herein allow for MV calculation and processing at the block level based on the obtained input images rather than input images from a game engine. The method and system may include detection of foreground MV and background MV, generation of virtual depth information, and conversion from block level to pixel level to allow use in processes such as extrapolation and interpolation. In an instance where extrapolation is performed, only one image may be output. When extrapolation is not performed, two or more images may be output, and then the two or more images may be merged into a single image.
[0030] The added depth information can provide additional flexibility to the image processing system. The depth component of the MV can be used as an input to, for example, a weighting function. The correlation between depth and the MV is generally higher compared to the actual image, even for small regions. By determining the foreground MV and background MV of a region, the virtual depth of a scene can be obtained, for example, in a video game. When the virtual depth is converted to the pixel level, it can allow for conversion of the block level MV to a pixel level MV and allow for more accurate interpolation to be performed by considering the change in depth. As described herein, the MV with added depth information is also referred to as MVD (MV with depth).
[0031] In addition, MV filtering may include filtering the MV along the motion trajectory and changing the MV of the background near the uncovered region. Doing so can divide large uncovered holes generated by the extrapolation and / or reprojection process into smaller hole regions. The smaller hole regions can be further filtered to completely reduce the holes, or a filling algorithm can be utilized to fill the smaller hole regions, which can mitigate any instability that may result from performing a filling algorithm on large holes.
[0032] Aspects of the present disclosure may be described herein with reference to the flowcharts and / or block diagrams of methods, apparatus, and computer program products according to embodiments disclosed herein. To the extent that such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, those skilled in the art will appreciate that each such function and / or operation within such block diagrams, flowcharts, or examples can be implemented, individually and / or collectively, by computer-readable instructions using a variety of hardware, software, firmware, or in fact any combination thereof. The systems described are exemplary in nature and may include additional elements and / or omit elements. The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various systems and configurations disclosed, as well as other features, functions, and / or properties. Accordingly, these methods may be performed by executing instructions stored on a machine-readable storage medium by utilizing one or more logical devices (e.g., processors) in conjunction with one or more additional hardware elements such as storage devices, memories, hardware network interfaces / antennas, switches, actuators, clock circuits, etc. The described methods and associated actions may also be performed in parallel and / or simultaneously in various orders other than the order described in this application. The processors of the logic subsystem may be single-core or multi-core, and the programs executed thereon may be configured for parallel processing or distributed processing. The logic subsystem may optionally include separate components distributed in two or more devices, which may be remotely located and / or configured for coordinated processing. One or more aspects of the logic subsystem may be virtualized and executed by a remotely accessible networked computing device configured in a cloud computing configuration.
[0033] Figure 1 An example of a computer system 100 is schematically depicted, which may include one or more processors 110, volatile and / or non-volatile memory 120 (e.g., random access memory (RAM) and / or one or more hard disk drives (HDD)). The processor 110 may include one or more central processing units (CPUs), GPUs, and / or one or more image processing systems, such as image processing system 115. The computer system may also include one or more displays 130, which may include any number of visual interface technologies. Additionally, example embodiments may include a user interface 140 (e.g., keyboard, computer mouse, touch screen, controller, etc.) to allow a user to provide input to the computer system. In some embodiments, the computer system may be a mobile phone or a tablet. The present disclosure herein is primarily directed to systems that describe the transfer of data within the processor 110. The transfer of data within the processor may include data generated by the CPU / GPU and transferred to the image processing system 115, or vice versa. It should be noted that the image processing system 115 may be a hardware unit external to the CPU / GPU or may be a program operating within the CPU / GPU.
[0034] As used herein, the term "system" or "module" may include a hardware and / or software system that operates to perform one or more functions. For example, a module or system may include a computer processor, a controller, or other logic-based device that performs operations based on instructions stored on a tangible and non-transitory computer-readable storage medium (such as computer memory). Alternatively, a module or system may include a hardwired device that performs operations based on hardwired logic of the device. The various modules or units shown in the figures may represent hardware that operates based on software or hardwired instructions, software that directs the hardware to perform operations, or a combination thereof.
[0035] It should be noted that the techniques discussed herein are applicable not only to games, but also to any animated rendering of 3D models, although the advantages provided by the method may be most apparent in the case of real-time rendering.
[0036] Figure 2 A block diagram of an image processing system 200 is shown. The image processing system 200 may be included as part of a computer system 100 Figure 1 or otherwise communicatively coupled to the computer system. In some instances, the image processing system 200 may be the image processing system 115 of the processor 110. In other instances, the image processing system 200 may be separate from the image processing system 115 of the processor 110.
[0037] The image processing system 200 may include multiple modules to perform various functions and / or store processor-executable instructions in memory. The image processing system 200 may include a luminance generator 202, an MV calculator 204, an MV processor 206, an MV depth combiner 212, a reprojection module 218, an image combiner 220, and a restorer 222, each of which performs or stores executable instructions to perform various tasks for MV calculation, processing, and / or filtering. In some instances, not all of the modules described herein may be included in the image processing system 200.
[0038] The luminance generator 202 may generate luminance images of two consecutive frames (e.g., PF and CF). A luminance image may refer to the luminance component of an image, where the luminance component is the y component, the y-channel component, or Y'CbCr. The luminance generator 202 may obtain the input images as PF and CF and generate a luminance image from the input images. MV calculation and thus the MV calculator 204 may use the luminance of the image as input. If the input image is red-green-blue (RGB), a color space conversion is required to obtain the luminance image, which may be performed by the luminance generator 202.
[0039] As described above, for some applications (such as video games), external MVs can be used for background and object masks as well as special effects (e.g., fog, shadows, etc.). Thus, the luminance image can be generated and / or modified based on whether the pixel belongs to an object, a special effect, or something else. For example, if the current pixel belongs to an object, the output luminance image can be generated based on Equation (1):
[0040] YO (i,j) = YI (i,j) * k 0 + thr 0 (1)
[0041] where YO is the output luminance image, YI is the input luminance, (i, j) is the pixel position, k 0 is the weight of the input luminance value, and thr 0 is the change to the luminance value based on the classification of the pixel as an object.
[0042] If the current pixel belongs to a special effect, the output luminance image can be generated based on Equation (2):
[0043] YO (i,j) = YI (i,j) * k 1 + thr 1 (2)
[0044] where the variables are as described above. Additionally, if the current pixel belongs to another type of pixel other than an object or a special effect, the output luminance image can be generated based on Equation (3):
[0045] YO (i,j) = YI (i,j) * k 2 + thr 2 (3)
[0046] where the variables are as described.
[0047] In some instances, the pixel values of the luminance image can be measured in the range from 0 to r, and typical values of k 0 are 0.5, k 1 is 0.4, k 2 is 0.4, thr 0 is half of r, thr 1 is 0, and thr 2is 0. The k value and the thr value can be used to tune the performance of the system. They can also be included as metadata to allow them to be optimized for a specific game. In some cases, some pixel regions can use internal MVs (e.g., iMV), and some pixel regions can use external MVs (e.g., eMV). In the case of using both iMV and eMV, a hybrid MV (e.g., hMV) may be generated. The modification of the luminance image described above can allow for a more accurate MV for the hMV case.
[0048] The MV calculator 204 can calculate or store executable instructions to calculate the MV and perform MV correction based on the generated luminance image. As will be further described, the MV can be calculated between the PF and the CF for a given pixel. For a pixel moving from the position (i, j) in the PF to the position (u, v) in the CF, the motion at the position (i, j) of the PF is (u - i, v - j) pixels, and the motion at the position (u, v) of the CF is also (u - i, v - j) pixels.
[0049] In some instances, the MV calculator 204 can calculate the MV at the block level instead of the pixel level, as described above. The block can be square, rectangular, or other groups of pixels, such as 8x8 pixels, 4x4 pixels, etc. Each block can be defined by its centered position or pixel. The MV can be calculated for the block based on the center positions in the PF and the CF in a similar manner as described above. The MV calculator 24 can generate an MV field based on the MVs calculated for multiple blocks within the input image.
[0050] In addition, the MV calculator 204 can generate a region foreground MV and a region background MV. In some instances, the input image or the luminance image can be segmented into H*V regions, where H represents the number of regions in the horizontal direction, and V represents the number of regions in the vertical direction. The foreground MV and the background MV can be generated for each region with an overlapping window. In some instances, the ratio between the sizes of the block regions can approximate the aspect ratio of the image.
[0051] The MV processor 206 can include a current frame MV processor 208 and a previous frame MV processor 210. The current frame MV processor 208 can process the MV of the CF, and the previous frame MV processor 210 can process the MV of the PF. Both the current frame MV processor 208 and the previous frame MV processor 210 can output a pixel-level MV and a pixel-level virtual depth by filtering the block-level MV, generating a block-level virtual depth, and decomposing the block-level MV and depth into a pixel-level MV and depth. Having both the current frame MV processor 208 and the previous frame MV processor 210 allows for the separate processing of the MV of phase 0 and the MV of phase 1.
[0052] The MV depth combiner 212 may include a current frame MV depth combiner 214 and a previous frame MV depth combiner 216. The MV depth combiner 212 may combine or store executable instructions to combine the iMV and internal depth with the eMV and external depth based on the marked objects and special effects in the previous frame (e.g., via the previous frame MV depth combiner 216) and / or in the current frame (e.g., via the current frame MV depth combiner 214).
[0053] The architecture of the image processing system 200 may be flexible as it responds to and meets the performance requirements of the coupled computing system and / or game engine. As such, the MV may be obtained from the game engine as the eMV and depth, or the MV may be estimated as the iMV and depth based on the acquired image. In some instances, if the performance requirements of the game engine are high, the image processing system 200 may compensate by using more iMV and depth to reduce the processing requirements of the game engine. Since the image processing system 200 also supports the combination of eMV and iMV, the MV depth combiner 212 may combine these two sets of information according to a mask. The mask may include mask information for objects, shadows, and special effects. For regions whose motion may not be expressible by the eMV, such as objects, shadows, and special effects, iMV calculations may generate iMV for these regions.
[0054] The reprojection module 218 may execute or store executable instructions to perform reprojection of the input image from time t (e.g., the time of the previous frame) to a new time position t + p (or from time t + 1 to time t + p, where p <= 1) based on the calculated MV, depth, and phase values, where the phase value is p when going from time t to t + p (or the phase value is p - 1 when going from time t + 1 to time t + p), assuming a time difference of 1 between adjacent frames. In some instances, the reprojection module 218 may include a current frame reprojection module and a previous frame reprojection module similar to the MV depth combiner 212 and the MV processor 206, where the MV of the calculated and processed PF is reprojected via the previous frame reprojection module, and the MV of the calculated and processed CF is reprojected via the current frame reprojection module. In some instances, the reprojection module 218 may include instructions stored in a memory to filter the obtained MV along the motion trajectory to reduce holes that are not covered during reprojection.
[0055] The image combiner 220 may obtain the reprojected images and merge them together. Similarly, the image combiner 220 may obtain the reprojected images of the previous frame and the current frame from the respective reprojection modules and may merge them together.
[0056] The restorer 222 may include one or more restoration algorithms or store instructions for executing one or more restoration algorithms. The one or more restoration algorithms may be configured to perform hole filling for pixel positions not reprojected by the reprojection module 218. As previously described, when previously hidden content is not covered in a new frame, holes will occur due to the lack of available color information in the source image. As an example, when an object in the foreground with an MV moves between the PF and the CF, background pixels may be left uncovered at the positions where the object was in the PF but is no longer in the CF. These pixel regions cannot be correctly filled only by reprojection, and the restorer 222 may fill the holes via one or more restoration algorithms. In some instances, the processing and filtering of the MV performed by the image processing system 200 may reduce the size of the holes, as will be further described below.
[0057] Now turning to Figure 3 , a flowchart illustrating a method 300 for MV calculation, processing, and filtering is shown. The method 300 may be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, the method 300 may be implemented by one or more processors according to instructions stored in a memory. For example, the instructions may be stored and executed by an image processing system such as the image processing system 200. The image processing system described herein may be part of a computer system (such as the computer system 100) including a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system.
[0058] At 302, the method 300 includes generating a luminance image as an input. In some instances, an image may be obtained and processed to output a luminance image, as described above. The luminance image may include a plurality of pixels at specific positions. As previously described, for some applications (such as video games), external MVs may be used for background and object masks as well as special effects (e.g., fog, shadows, etc.). Thus, the luminance image may be generated and / or modified based on equations (1), (2), and (3) defined above, based on whether the pixel belongs to an object, a special effect, or others. In some instances, the luminance image may include the PF and the CF.
[0059] At 304, the method 300 includes calculating the MV between the PF and the CF. The MV may be calculated for a specific pixel or block at a defined position in one of the PF and the CF. The MV of the pixel from the PF to the CF or from the CF to the PF may be calculated. An MV field including a plurality of MVs between frames may also be generated. In some instances, the MV calculation at 304 may include foreground MV and background MV detection. Additionally, the MV calculation may include calculating the motion and quality of the MV, which includes the calculation of the sum of absolute differences (SAD) and the difference (MVDIFF) between the MV of the current block and the MV of the block pointed to by the current block. It will be described with respect to Figure 4and Figure 5 MV calculation (including quality determination and detection of background MV and foreground MV) is further described. In some examples, MV can also be calculated between the previous frame (PPF) and the PF, such as Figure 5 Further discussion.
[0060] At 306, method 300 includes processing the MV. Figure 6 As further described, processing the MV may include MV post-filtering, generation of virtual depth, and block-level MV decomposition. MV post-filtering may be based on the MV quality determined at 304. MV processing may output a block-level or pixel-level MV with depth information, such as may be used for frame interpolation. In some examples, processing the MV may include filtering along a motion trajectory to reduce the size of uncovered holes between PF and CF.
[0061] At 308, method 300 includes re-projecting the input image according to the MV. In some examples, a block-level MV and a corresponding pixel-level MV may be obtained. Virtual depth information and a phase value of the MV may be obtained and / or determined, and the input image may be rendered based on the depth and phase values. The phase value may refer to a temporal distance between the rendered image and the original input image.
[0062] At 310, method 300 includes merging the multiple rendered images together. In some examples, when interpolating, PF is forward projected from time t to time t+p, and its phase value is p, where p is between 0 and 1. And CF is backward projected from time t+1 to time t+p, and its phase value is p-1. This produces two images that are merged to create the final image. In other examples, such as when extrapolation is used, the MV of CF is 1 The field may be projected forward from time t+1 to time t+1+p, in which case only one image may be generated and merging may be skipped. In some instances, one or more regions in the merged image may include pixels / blocks with no data because the movement of pixels according to the MV is not covered.
[0063] At 312, method 300 includes filling holes in the merged image that are not covered via a inpainting algorithm. In some instances, the inpainting algorithm may be directly applied to the merged image, such as an inpainting algorithm that copies the background texture around the hole, selects random neighboring pixels, and / or uses an intermediate image with low-pass filtering to fill the hole. In other instances, as will be further described below, the MVs may be filtered along the motion trajectory, and the MVs of the background pixels / blocks near the uncovered region may be changed to divide the uncovered hole into smaller holes, as mentioned at 306. In some instances, the inpainting algorithm may then be applied to the smaller holes. In other instances, the MV filtering for reducing holes may be repeated one or more times until the holes are substantially eliminated.
[0064] Method 300 may allow processed and filtered MVs, which may be used for intra-frame interpolation or other image processing including extrapolation and reprojection. The processed and filtered MVs may be generated in a more flexible manner, thus allowing the use of iMVs and eMVs based on GPU / CPU requirements. Compared with eMVs, MVs including internal estimation and / or calculation may allow lower computing and / or processing capabilities. In addition, holes generated during image processing such as reprojection and / or extrapolation may be filed by filtering the MVs.
[0065] Now turning to Figure 4 , a flowchart illustrating a method 400 for calculating MVs is shown. Method 400 may be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, method 400 may be implemented by one or more processors according to instructions stored in a memory. For example, the instructions may be stored and executed by an image processing system such as image processing system 200. The image processing system described herein may be part of a computer system (such as computer system 100) including a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. Method 400 may be incorporated into the Figure 3 method 300, particularly at 302.
[0066] At 402, method 400 includes generating a luminance input. As described above with respect to Figure 2 , luminance images may be generated for two consecutive frames (e.g., PF and CF). The luminance input may be generated and / or modified based on equations (1), (2), and / or (3) based on whether the pixel belongs to an object, a special effect, or others. For example, for a current pixel belonging to an object, an output luminance image may be generated based on equation (1), for a current pixel belonging to a special effect, an output luminance image may be generated based on equation (2), and for a current pixel belonging to a pixel type other than an object or a special effect, an output luminance image may be generated based on equation (3).
[0067] In some cases, some pixel regions may use internal MVs (e.g., iMV), and some pixel regions may use external MVs (e.g., eMV). In cases where both iMV and eMV are used, a hybrid MV (e.g., hMV) may be generated. The modification of the luminance image as described above may allow for a more accurate MV for the hMV case.
[0068] At 404, method 400 includes calculating the MVs between the PF and the CF based on the luminance input image. The multiple MVs calculated between the PF and the CF at different positions may together form an MV field. Calculating the MVs may include calculating the MV for each block of the PF (as shown at 406) and calculating the MV for each block of the CF (as shown at 408). In some instances, as described above, pixel blocks may be defined for the CF and the PF. For example, blocks of 8x8 pixels may be used, although other sizes and shapes of blocks are possible.
[0069] Each block of the luminance image may have a defined position. As an example, for a block (m, n) with a block size of bxb, the upper left position of the block may be [b*m, b*n], and the lower right position of the block may be [b*m + b - 1, b*n + b - 1]. In some instances, the center position of the block may be proportional to (m, n).
[0070] As an example, for a block (m,n) in the PF with an MV of [y mn , x mn , its upper left position moves from the first position [m*b, m*b] in the PF to the second position [m*b + y mn , n*b + x mn in the CF. The MV in the CF has the same polarity as the MV in the PF. If the MV of a block (c,d) in the CF is [y cd , x cd , then its upper left position moves from the position [c*b, d*b] in the CF to the position [c*b - y cd , d*b - x cd in the PF.
[0071] At 410, method 400 includes determining the quality of the calculated MVs. In some instances, the calculation of the MVs may incorporate the determination of various values, including but not limited to the horizontal motion component, the vertical motion component, the sum of absolute differences (SAD) between the matching blocks (as indicated at 412), and the difference (MVDIFF) between the MV of a block and the MV of the block it points to (as indicated at 414).
[0072] For a block in the CF with an MV of [y cd , x cdThe block (c, d), and the SAD of the block (c, d) is: And for the block (m, n) in PF that has MV[y mn , x mn , the SAD of the block (m, n) is: In some instances, a smaller SAD value indicates a good match between blocks, and a larger SAD value indicates a poor match.
[0073] MVDIFF can indicate whether the block (m, n) is in the covered area or the uncovered area. A small MVDIFF indicates that the MV of PF is confirmed by the opposite MV of CF, and thus it may not be in the covered area / uncovered area. A large MVDIFF indicates that the block is in the covered area / uncovered area, such that the current MV is not confirmed by the opposite MV of CF. For MV 0 , a large MVDIFF can indicate that the block is in the covered area. For MV 1 ,, a large MVDIFF can indicate that the block is located in the uncovered area. The MVDIFF of the block (m, n) of PF can be determined by Equation (4):
[0074]
[0075] Where, and iBMV PF0(m,n) ·mvdiff is the MV field of PF for the block (m, n), iBMV PF0(m,n) ·x is its horizontal motion, and iBMV PF0(m,n) ·y is the vertical motion, and iBMV CF1(c′,d′) ·x is the horizontal motion of the MV field of CF for the block (c′, d′), and iBMV CF1(c′,d′) ·y is the vertical motion.
[0076] The MVDIFF of the block (c, d) of CF can be determined by Equation (5):
[0077]
[0078] Where, and iBMV CF1(c,d) ·mvdiff is the MV field of CF for the block (c, d), iBMV CF1(c,d) ·x is its horizontal motion, and iBMV CF1(c,d) ·y is the vertical motion, and iBMV PF0(m′,n′) ·x is the horizontal motion of the MV field of PF for the block (m′, n′), and iBMV PF0(m′,n′) ·y is the vertical motion.
[0079] Briefly turning to Figure 8 , a first diagram 800 of the MV field is shown. The first diagram 800 shows the MVDIFF calculation as described at 414 of method 400. The first line 802 represents the internal block-level phase 0 MV field of the PF (e.g., iBMV PF0 ), and the second line 804 represents the internal block-level phase 1 MV field of the CF (e.g., iBMV CF1 ). A plurality of block-level MVs 850 are represented as arrows between the first line 802 and the second line 804. For example, for the first block 806 (e.g., the first block (m, n)), the MV of the first block 806 is represented by the arrow 808. The first block 806 may be included in the coverage area 814 such that its MV is invalid up to the CF. For the coverage area where the MV is invalid up to the corresponding next frame, the MVDIFF can be large. When the MV is valid for the corresponding next frame, a smaller MVDIFF can be generated.
[0080] To determine the MVDIFF and whether the block is in the coverage area or the uncovered area, an MV projection can be performed to project the MV to hit its next block. The MV projection can be performed from iBMV PF0(m,n) to the second block 812 (e.g., the second block (c′, d′) in iBMV CF1 ). The second block can be determined based on equation (6):
[0081]
[0082] where the variables are as described above. Then the MV from the CF to the PF can be calculated based on the position of the second block. The resulting positions of c′ and d′ and the MV from the CF to the PF can then be used to calculate the MVDIFF, as described with respect to equation (4) above. If the calculated MVDIFF is large, the first block is in the coverage area. The dashed line 810 represents a set of phases where the MV is invalid due to the first block being covered. The determination of the MVDIFF as described herein allows determination of whether the first block is covered or uncovered.
[0083] Returning to Figure 4 , at 416, method 400 includes detecting a foreground MV and a background MV. As will be further described with respect to Figure 5 , the detection of the foreground MV and the background MV can be an algorithm based on four MV fields. In addition to the MV fields between the PF and the CF (e.g., iBMV PF0 and iBMV CF1 ), the block-level MV fields between the PPF and the PF (e.g., iBMV PP0 and iBMV PF1) can also be used to determine the foreground MV and the background MV. It should be understood that the calculation of the block-level MV field between the PPF and the PF can simply be the previously calculated block-level MV field, for example, between the PF and the CF that is one frame delayed from the currently calculated frame. By comparing the SAD value with the MVDIFF value, an occluded region can be found, and if so, the local foreground MV and background MV can be detected. As described herein, the occluded region can be a region where a moving object covers something that was not previously covered or exposes the background that was previously covered.
[0084] The MV field including foreground information and background information calculated as described herein can be processed, as will be further described below, to generate virtual depth. The virtual depth allows the rendered image to be used for frame interpolation.
[0085] Now turning to Figure 5 , a flowchart illustrating a method 500 for detecting the foreground MV and the background MV is shown. The method 500 can be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, the method 500 can be implemented by one or more processors according to instructions stored in a memory. For example, the instructions can be stored and executed by an image processing system such as the image processing system 200. The image processing system described herein can be part of a computer system (such as the computer system 100) including a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. The method 500 can be incorporated into the Figure 4 method 400, particularly at 416.
[0086] At 502, the method 500 includes obtaining an MV for a first block (e.g., block (m, n)). It should be understood that the first block (m, n) representing each block of the input image and the method 500 can be performed sequentially or simultaneously for one or more of the blocks of the input image. As described with respect to Figure 4 , an MV field can be obtained between the PF and the CF and between the PF and the PFF. The MV field between the PF and the CF can include the MV of each block of the PF (e.g., the internal block-level phase 0 MV or iBMV of the PF PF0 ) and the MV of each block of the CF (e.g., the internal block-level phase 1 MV or iBMV of the CF CF1 ). The MV field between the PPF and the PF can include the MV of each block of the PPF (e.g., the internal block-level phase 0 MV or iBMV of the PPF PP0 ) and the MV of each block of the PF (e.g., the internal block-level phase 1 MV or iBMV of the PF PF1)。When calculating the MV for each block for PF, CF, and PPF, the MV of the first block (m, n) can be obtained. Additionally, obtaining the MV of the first block (m, n) can include obtaining the SAD value and MVDIFF value of the block, as they are calculated during the calculation of the MV, as described with respect to Figure 4 as described.
[0087] At 504, method 500 includes determining whether the first block (m, n) is covered or uncovered. If the MVDIFF of iBMV PF0 (e.g., the MV field between PF and CF for (m, n) of PF) is greater than the MVDIFF of iBMV PF1 (e.g., the MV field between PPF and PF for (m, n) of PF), then the first block (m, n) may be a potentially covered block. If the MVDIFF of iBMV PF0 is less than the MVDIFF of iBMV PF1 , then the first block (m, n) may be a potentially uncovered block.
[0088] At 506, method 500 includes determining the potential background MV and occlusion MV. The determination of the potential background MV and occlusion MV is based on whether the first block is covered or uncovered, as determined at 504. As an example, if the first block (m, n) is covered, then the MV field between the PPF and PF of the first block (e.g., iBMV PF1 ) can be set as the background MV, for example as shown in equations (7) and (8):
[0089] MV BG = iBMV PF1(m,n) (7)
[0090] MV occ = iBM PF0(mn) (8)
[0091] where MV BG is the background MV, and MV occ is the occlusion MV.
[0092] At 508, method 500 includes performing MV projection. The MV projection can be performed using the background MV from the first block (m, n) to the CF of the first block (m, n) (e.g., the inner block phase 1 of CF), as indicated at 510. The MV projection may hit the second block (w, v) to obtain the potential foreground. The MV projection can be performed based on equation (9):
[0093] MV FG = iBMV CF1(u,v) (9)
[0094] where MVFG is a foreground MV, and iBMV CF1(u,v) is the MV field between the PF and the CF of the CF of the second block (u, v).
[0095] Then, the foreground MV from the second block (u, v) to the PF (e.g., the internal block phase 0 of the PF) can be utilized to perform MV projection, as indicated at 512. The MV projection may hit the third block (s, t). The MV projection of the third block (s, t) can be based on Equation (10):
[0096]
[0097] where is the hit MV of the foreground, and iBMV PF0(s,t) is the internal block phase 0 MV of the PF of the third block (s, t).
[0098] Then, the foreground MV from the first block (m, n) to the CF (e.g., the internal block phase 1 MV) can be utilized to perform MV projection, as indicated at 514. The MV projection may hit the fourth block (p, q). The MV projection of the fourth block (p, q) can be based on Equation (11):
[0099]
[0100] where is the hit MV of the background, and iBMV CF1(p,q) is the internal block phase 1 MV of the CF of the fourth block (p, q).
[0101] MV projection can be performed for each block in the blocks of the corresponding input image frame. As described herein, MV projection can provide potential foreground MV and background MV via Equations (9), (10), and (11). The foreground MV and background MV can be global and / or regional.
[0102] Briefly turning to Figure 9 , a second figure 900 is shown, which illustrates MV projection for detecting foreground MV and background MV as described at 508, 510, 512, and 514 of method 500. The second figure 900 includes a first line 902 representing the internal block-level phase 0 MV field of the PPF (e.g., iBMV PP0 ), a second line 904 representing the internal block-level phase 1 MV field of the PF (e.g., iBM PF1 ), a third line 906 representing the internal block-level phase 0 MV field of the PF (e.g., iBMV PF0 ), and a fourth line 908 representing the internal block-level phase 1 MV field of the CF (e.g., iBMV CF1The fourth line 908 of ( )). A plurality of block-level MVs are represented as arrows between the first line 902 and the second line 904 and between the third line 906 and the fourth line 908.
[0103] The first block 910 of PF (e.g., the first block (m, n)) may be an overlay block. The internal block phase 1 MV of PF of the first block 910 may be represented by the line 950. As described with respect to method 500, the internal block phase 1 MV of PF of the first block (m, n) (e.g., iBMV PF1(m,n) ) may be set as a potential background MV, such as the MV described in the above equation (7) BG , and when the first block (m, n) is located in the overlay area 914, the phase 0 MV of PF of the first block (m, n) may be an occlusion MV (e.g., MV occ ), as described in the above equation (8). The occlusion MV may be represented by the first arrow 912.
[0104] The MV projection from the first block 910 to the second block 916 may be represented by the first dashed arrow 952 in the second figure 900. The first dashed arrow 952 may be the internal block-level phase 1 MV of the CF of the second block 916, as described with respect to equation (9), which may be a foreground MV (e.g., MV PG ).
[0105] Then, an MV projection from the second block 916 to the third block 920 (e.g., the third block (s, t)) may be performed. The MV generated by the MV projection from the second block 916 to the third block 920 is represented by the second arrow 918 in the second figure 900. When it is PF, the MV represented by the second arrow 918 may be a dc foreground MV (e.g., ), as described in the above equation (10).
[0106] An MV projection may also be performed from the first block 910 to the fourth block 924 (e.g., the fourth block (p, q)). Similar to as described with respect to Figure 8 , the internal block-level phase 0 MV of PF of the first block 910 may be invalid all the way to the fourth block 924. An MV projection may be performed to reach the fourth block 924. The second dashed arrow 922 may represent the MV obtained by the MV projection. The resulting MV may be the internal block-level phase 1 MV of the CF of the fourth block 924, which, as described in the above equation (11), may be a dc background MV (e.g., ).
[0107] Each of the foreground MV and the background MV determined via MV projection as described in method 500 and shown in the second figure 900 may be a potential MV. Some of the potential MVs may be reliable, and some of the potential MVs may be unreliable.
[0108] Return to Figure 5 At 516, method 500 includes determining the reliability of the foreground MV and the background MV. Determining the reliability of the foreground MV and the background MV may include defining a previous global foreground MV and a previous global background MV, as well as a previous regional foreground MV and a previous regional background MV. The previous regional foreground MV and background MV may be related to the region to which the first block (m, n) belongs. Each of the previous foreground MV and background MV including the corresponding global and regional ones may be defined according to horizontal motion and vertical motion. In addition, a δ (e.g., the pixel distance between the positions pointed to by the two MVs) between the foreground MV and the background MV with respect to the horizontal motion and the vertical motion may be determined, such as that described in equation (12) for the foreground MV:
[0109]
[0110] where mvdist FG is the δ of the foreground MV, MV FG ·x is the horizontal component of the foreground MV, and MV FG ·y is the vertical component of the foreground MV. A similar equation may be used for the background MV.
[0111] In some instances, the difference between the foreground MV and the background MV may be determined according to equation (13):
[0112] mvdist FGBG = |MV FG ·x - MV BG ·x| + |MV FG ·y - MV BG ·y| (13)
[0113] where mvdist FGBG is the δ between the foreground MV and the background MV.
[0114] Then the δ from the foreground to the global foreground and the global background, and from the background to the global foreground and the global background may be defined. For example, the MV distance from the foreground to the global foreground may be defined according to the corresponding horizontal and vertical components of the foreground MV and the previous global foreground MV, the MV distance from the foreground to the global background may be defined according to the corresponding horizontal and vertical components of the foreground MV and the previous global background MV, the MV distance from the background to the global foreground may be defined according to the corresponding horizontal and vertical components of the background MV and the previous global foreground MV, and the MV distance from the background to the global background may be defined according to the corresponding horizontal and vertical components of the background MV and the previous global background MV.
[0115] Based on these defined distances, the bGoodBGFG flag can be determined. The bGoodBGFG flag can be true when the following two conditions are met: the difference between the background MV and the global background MV is less than half of the sum of the difference between the background MV and the global foreground MV and the difference between the foreground MVs, and the global foreground MV is less than half of the difference between the foreground MV and the global background MV.
[0116] When bGoodBGFG is true and the SAD of the background MV is minimized, or when 1) the SAD of the background MV is minimized, 2) the maximum value of the MVDIFF of the background MV and the SAD of the background MV is less than the minimum value of thr1 and half of the difference between the foreground MV and the background MV, 3) the difference between the foreground MV and the background MV is greater than thr2, 4) the SAD of the occluded MV is greater than the maximum value of thr3 and the SAD of the background MV, and 5) the MVDIFF of the occluded MV is greater than thr4, the background MV (MV BG ) may be reliable.
[0117] In other words, the background MV can be reliable when one of the following conditions is met:
[0118] Reliable BG 1 :(bGoodBGFG && MV BG ·saad < min(thr0, MV OCC ·sad)
[0119] Reliable BG 2 :
[0120] Reliable FG = Reliable BG 1 || Reliable BG 2
[0121] When 1) the background MV is reliable, 2) the MVDIFF of the occluded MV is greater than thr5, 3) the difference of the foreground MV is less than thr6, half of the SAD of the foreground MV is less than thr7, and 4) the SAD of the foreground MV is less than thr7, the foreground MV may be reliable.
[0122] In other words, the foreground MV may be reliable when the following conditions are met:
[0123]
[0124] At 518, method 500 includes accumulating reliable foreground MVs and background MVs and averaging the foreground MVs and background MVs for each region and the entire frame. For a given region (h, v), if the reliable counts of the foreground MV and the background MV (e.g., fgcnt and bgcnt) are too small, the foreground MV and the background MV can be changed according to equation (14):
[0125]
[0126] Wherein RMV is the area MV of the foreground or background of the horizontal or vertical movement, which depends on the equation, GlbFGMV is the global foreground MV, GlbBGMV is the global background MV, and wfg (h,v) = min(thr FG , RMV FG(h,v) .fgcnt), and wbg (h,v) = min(thr BG , RMV BG(h,v) .bgcnt).
[0127] In this way, reliable foreground MV and background MV can be generated based on the MV fields between CF and PF and between PF and PPF.
[0128] Now turning to Figure 6 , a flowchart illustrating a method 600 for MV processing is shown. Method 600 can be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, method 600 can be implemented by one or more processors according to instructions stored in a memory. For example, the instructions can be stored and executed by an image processing system such as image processing system 200. The image processing system described herein can be part of a computer system (such as computer system 100) that includes a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. Method 600 can be incorporated into Figure 3 method 300 of, particularly at 304.
[0129] At 602, method 600 includes obtaining an MV field. As described with respect to Figure 4 , an MV field can be obtained between PF and CF and between PF and PFF. The MV field between PF and CF can include the MV of each block of PF (e.g., the internal block-level phase 0 MV or iBMV PF0 ) and the MV of each block of CF (e.g., the internal block-level phase 1 MV or iBMV CF1 ). The MV field between PPF and PF can include the MV of each block of PPF (e.g., the internal block-level phase 0 MV or iBMV PP0 ) and the MV of each block of PF (e.g., the internal block-level phase 1 MV or iBMV PF1 ). When calculating the MV of each block for PF, CF, and PPF, the MV of the first block (m, n) can be obtained. Additionally, obtaining the MV of the block (m, n) can include obtaining the SAD value and MVDIFF value of the block, as they are calculated during the calculation of the MV, as described with respect toFigure 4 as described. In some instances, the obtained MV field may include a foreground MV and a background MV (including a global foreground MV and a global background MV, and a regional foreground MV and a regional background MV), which may be detected as described with respect to Figure 5 as described.
[0130] At 604, method 600 includes performing post-filtering of the MV. In some instances, the post-filtering of the MV may include calculating an average value of a specified window, smoothing the MV, and reducing the MV. For example, post-filtering of the MV may be performed on unreliable regions to generate a filtered output MV. In some instances, the post-filtering of the MV may include replacing unreliable MVs with the regional foreground MV and the global foreground MV, and the regional background MV and the global background MV.
[0131] At 606, method 600 includes generating block-level virtual depth. Virtual depth may be required to generate an output image by re-projection and / or intra-frame interpolation. For a given block, for example, the first block (m, n), the depth may be calculated according to Equation (15):
[0132] iBD (m,n) = max(0, depth i + k 3 *(mvdist2fg - mvdist2bg)+ k 4 *iBMV m,n .mvdiff (15) where iBD (m,n ) is the virtual depth of the first block (m, n), depth i is the initial depth value, mvdist2fg is the MV difference between the current MV and the corresponding regional foreground MV, mvdist2bg is the MV difference between the current MV and the regional background MV, and iBMV (m,n) ·mvdiff is the double-checked MVDIFF of the MV of the first block (m, n). The regional foreground MV or background MV used to determine the MV difference may be the bilinear interpolated MV (foreground or background) of the first block (m, n) with four neighboring regional MVs (foreground or background).
[0133] At 608, method 600 includes decomposing the block-level MV field into smaller block-level MV fields. As previously described, the luminance image may include pixels and be divided into a plurality of equally sized blocks. As an example, the luminance image may be segmented into 8x8 pixel blocks. The MVD calculation and processing described with respect to Figures 4 to 5 may be performed at the block level. Optionally, there is an object mask that marks whether each pixel belongs to an object, and only pixels belonging to the object to be processed.
[0134] Decomposing the block-level MVD field into a pixel-level MVD field may include decomposing the block into smaller blocks to improve the accuracy of the pixel MVD. As indicated at 610, for example, the guiding or reference image may be the luminance of the smaller blocks, as indicated at 612. Determine the block-level guiding image for each smaller block, and optionally, if there is an object mask, count the number of object pixels within each smaller block, as indicated at 614. As an example, an 8x8 pixel block may be decomposed into four 4x4 pixel blocks. The guiding image can be used to improve the accuracy of the pixel-level MV. The guiding image can be depth information, luminance information, or some other pixel-level information that identifies different objects in the scene of the image. This guidance can be used to weight the interpolation of the MV between blocks, or can be used in a regression method such as a guiding filter, as will be further described. In some instances, the weight of the block MV that belongs to the same object as the pixel MV is higher than the block MV that does not belong to the same object as the pixel MV.
[0135] At 616, method 600 includes filtering a subset of the smaller blocks to generate a pixel-level MVD field. For each pixel (i, j), multiple MVDs of the window of smaller blocks around that pixel can be obtained, such as the 5×5 blocks around that pixel (these blocks are denoted as nebblks). In some instances, not all blocks within the window (e.g., the 5x5 block window) may be filtered. If there is an object mask, only the blocks that belong to the same object as the current pixel can be filtered. Additionally, only the blocks whose luminance level (e.g., pixel intensity) is within the range of the luminance level of the pixel in question can be filtered.
[0136] In some instances, the filtering may include weight calculations for luminance adjustment and weighting, spatial weighting, and object mask weighting (if present) for a given block. Luminance adjustment and weighting may include calculating the luminance difference between the current pixel and each surrounding block of the window, and then averaging the luminance differences. The average luminance difference can then be adjusted according to Equation (16):
[0137] brtdiff avg =min(thr0, max(thr1, 1 + brtdiff avg )) (16)
[0138] where brtdiff avg is the average luminance difference. The weight of the luminance can then be calculated according to Equation (17):
[0139] w_brt (m,n) =max(0, brtdiff avg - abs(brt_nebblks (m,n) - brt (i,j) )) (17)
[0140] where w_brt (m,n) is the luminance weight of a given block (m, n) of nebblks, brt_nebblks (m,n) is the luminance of the block (m, n) of nebblks, and brt (i,j) is the luminance of the pixel (i, j).
[0141] Based on the luminance weight, the spatial weight, and the object mask weight, the weight of a given block can be determined, where the weight of the given block is its product, as described in Equation (18):
[0142] w (m,n) = w_brt (m,n) * w_spat (m,n) * w_obj (m,n) (18)
[0143] where w (m,n) is the weight of the given block (m, n), w_spat (m,n) is the spatial weight of the given block, and w_obj (m,n) is the weight of the object mask. If there is no object mask, w_obj (m,n) may be equal to a constant value c.
[0144] Then, the weight of the given block can be used to generate an output image. In some instances, for each component of the MVD, the output can be generated according to Equation (19):
[0145]
[0146] where x o (i, j) represents the component of the decomposed pixel-level MVD of the pixel (i, j) (which can be horizontal motion, or vertical motion or depth), and x i(m,n) and is the component of the input block-level MVD of nebblks (m, n).
[0147] Block-level MVD decomposition, including filtering as described herein, can allow for improving the quality of the MVD by filtering based on weights such as luminance differences, spatial, and / or object masks. The methods for block-level MVD decomposition and filtering herein should be understood as merely examples, and other methods can also allow for generating pixel-level MVD. For example, a guided filter finds the best fit between the input data (e.g., block-level MVD) and the guidance data to generate the output data (e.g., pixel-level MVD). The guided filter can generate the pixel-level MVD for each pixel based on the values in a given window around the corresponding pixel. The guided filter can be applied to a guidance image, which can be luminance, panchromatic image, depth, or object mask.
[0148] In addition, in some instances, in addition to the above decomposition options, the results of object segmentation can also be used as guidance. For example, a trainable neural network can be used to identify whether pixels in an image belong to a certain type of object. Each type of object can have a different MVD associated with it, and thus using object identifiers and segmentation allows for MV calculation and processing.
[0149] As described herein for methods 400, 500, and 600, the output of MV calculation and processing can generate filtered MVs with pixel-level virtual depth information (MVD). These MV fields can be used in various image processing methods, including interpolation, extrapolation, and / or reprojection.
[0150] Now turning to Figure 7A , a flowchart of method 700 for reducing holes in an image generated by reprojection (e.g., a triangle projection method) is illustrated. Method 700 can be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, method 700 can be implemented by one or more processors according to instructions stored in a memory. For example, the instructions can be stored and executed by an image processing system such as image processing system 200. The image processing system described herein can be part of a computer system (such as computer system 100) that includes a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. Method 700 can be incorporated into Figure 3 's method 300, specifically at 306 and / or 310.
[0151] At 702, method 700 includes determining the presence of holes in the reprojection (or extrapolation) image. As previously described, in some instances, reprojection and / or extrapolation can result in one or more holes in the output image when previously covered content is not covered between frames. For each pixel in the input image, reprojection (or extrapolation) includes determining a projection position (e.g., position (u, v)) and depth (e.g., d uv ) based on the values of the MV, depth, and phase. The MV, depth, and phase can be calculated or otherwise determined as described above with respect to Figure 4 , Figure 5 and Figure 6 . The projection position can be determined according to equation (20), and the depth can be determined according to equation (21):
[0152] u = i + phase * mv ij .y (20)
[0153] v = j + phase * mv ij ·x (20)
[0154] d uv= D(i, j) + phase * Z(i, j) (21)
[0155] where mv ij .y and mv ij .x are the vertical and horizontal motions of the MV at the position (i, j) of a given pixel respectively, D(i, j) is the input depth of the input (i, j) of the pixel at time t, and Z(i, j) is the change in the depth of the pixel (i, j) from time t to t + 1.
[0156] If the projected image at the projection position (u, v) is invalid or the depth at the projection position (u, v) is greater than the depth d uv , then the depth at the projection position can be replaced with the depth d uv , and the projected image at position t + phase can be replaced with the input image of the given pixel (i, j) at time t. When VALID(u, v) is equal to zero, the projected image at position (u, v) may be invalid, which may be the initial value of the given pixel. Before reprojection, the initial value of VALID for all positions is 0, and when the position (u, v) is projected by the input pixel, then VALID(u, v) is set to 1. Once the depth and the image are replaced, the validity can be confirmed (e.g., VALID(u, v) = 1). For example, the depth can be used to determine which of the one or more pixels being projected are retained, where some pixels may be background pixels and other pixels may be foreground pixels.
[0157] When a background pixel that was first covered by a neighboring foreground (e.g., at time t) is not covered at a second time (e.g., at time t + p), holes will appear in the projected image. The holes can be detected in the MV field from the first time to the second time. When there is relative motion between the foreground and the background, the holes may not be covered. The MV field can be filtered to reduce the size of the holes, as will be described below. Method 700 herein describes the filtering, where the filtering may include finding the foreground MV for each pixel by comparing the depth with the depths of neighboring pixels and filtering the MV field around the uncovered region to generate a filtered output MV field. The range of neighboring pixels can be determined by the MV amplitude and phase of the projected image. The points along the motion trajectory can be filtered.
[0158] At 704, method 700 includes determining the local foreground MV. Determining the local foreground MV includes initializing the foreground MV field and the foreground depth field, as indicated at 706. The initialization may include setting the values to the values of the input MV and the input depth. Determining the local foreground MV may further include determining the foreground MV field and the foreground depth field for each pixel, as indicated at 708.
[0159] As an example, for a given pixel (i, j), the foreground MV field and the depth field may be determined by determining the magnitude of the MV at position (i, j) according to equations (22) and (23), as indicated at 710:
[0160]
[0161] in is the magnitude of the MV at position (i, j), and and are the horizontal and vertical components of the normalized motion vector, respectively.
[0162] For a given pixel (i, j), its processing range can also be determined according to equation (24):
[0163]
[0164] where r ij is the range and k 0 Equals 2.
[0165] Along with the range [-r ij , r ij ], each point of the trajectory may be examined to determine the projected position (u, v), as indicated at 712. For example, for each step t, t is [-r ij , r ij ], the projection position can be determined according to equation (25)
[0166]
[0167] The variables are as described above.
[0168] For each step t, the position The foreground depth of is compared with the input depth of position (i, j). When the input depth of position (i, j) is less than When the foreground depth is The foreground depth of position (i, j) is replaced by the input depth of position (i, j), and the position The foreground MV of position (i, j) is replaced by the MV of position (i, j). In addition, for each step t, the position The input depth of is compared with the foreground depth at position (i, j). When the input depth is less than the foreground depth of (i, j), the foreground depth of (i, j) is calculated using the position The input depth is replaced by , and the foreground MV of (i, j) is replaced by In addition, k can be replaced by, for example, a guided filter, a bilateral filter, etc.1 The size of the specified window smooths the determined foreground MV field and foreground depth field, as indicated at 714. As will be described below, the foreground MV field and foreground depth field determined herein can be filtered to reduce holes.
[0169] It should also be understood that a given pixel as described herein represents each pixel, and the determination of local foreground MV as described herein can be performed for one or more pixels of the input image.
[0170] At 716, method 700 includes defining a hole MV for a given pixel to define the actual size of the hole. In some instances, the hole MV (e.g., hmv ij ) can be the product of the phase and the difference between the MV of the given pixel (i, j) and the foreground MV of the given pixel (i, j). In other instances, the hole MV can be the product of the phase and the MV of the given pixel (i, j), which can reduce the computational requirements as it reduces the computation of the foreground MV.
[0171] At 718, method 700 includes determining a filtering range and step size and determining the amplitude of the hole MV. Similar to the above, the filtering range, as well as the step size and amplitude of the hole MV, can be determined according to equations (26) and (27):
[0172] amp ij = max(abs(hmv ij .x), abs(hmv ij .y)) (26)
[0173]
[0174] where the variables are as described above.
[0175] The filtering range, as well as the step size and amplitude, can be used to define the range of the warping radius and step size based on equations (28) and (29):
[0176] r 弯曲 = min(thr 2 , k 2 *amp ij ) (28)
[0177]
[0178] where r 弯曲 is the range of the warping radius, s ij is the step size, thr 2 is the threshold for limiting the pixel range, taps is the number of sampling points for controlling along the motion trajectory, and k 2 is a parameter for adjusting the pixel range for performing MV filtering. In some instances, k2 The default value can be 0.25. A large k 2 value may indicate more pixels for which the MV will be filtered. A larger taps value can achieve higher precision. In some instances, the number of points for which filtering can be performed can be 2M + 1, where M is the quotient of r 弯曲 and s ij .
[0179] Continuing to refer to Figure 7B , at 720, method 700 includes filtering the MV along the motion trajectory based on the determined filtering range, step size, and amplitude. As indicated at 722, filtering the MV along the motion trajectory can include determining the projected position for each step size t (e.g., the position (u t , v t )) of the MV. In some instances, t can be between -M and N, where M is defined as above. The projected position can be defined according to equation (30):
[0180] u t = round(i + t * s ij * adj y ) (30)
[0181] v t = round(j + t * s ij * adj x ) (30)
[0182] where the variables are as described previously.
[0183] Then, filtering the MV along the motion trajectory includes comparing the MV of a given pixel (i, j) with the MV of the projected position (u t , v t ), as indicated at 724, to determine whether there is an uncovered relationship between the two. The determination of the uncovered relationship can be based on conditional equation (31):
[0184]
[0185] where mv ij ·x and mv ij ·y are the horizontal and vertical motion components of the MV of the given pixel, and mv t ·x and mv t ·y are the horizontal and vertical motion components of the MV of the projected position. The value of the uncovered relationship, uncovered t can be the sum of horizontal component uncovered x and vertical component uncovered y . When the value of the uncovered relationship, uncovered tWhen greater than zero, the MV of the projection pixel can participate in filtering.
[0186] Filtering the MV along the motion trajectory can further include determining the foreground and background relationship between the MV of a given pixel and the MV of the projection position, as indicated at 726. In some instances, the background may be affected by the foreground, but the foreground may not be affected by the background. The MV participating in filtering can be determined by Equation (32):
[0187]
[0188] where mv′ t .x and mv′ t .y are the horizontal motion component and the vertical motion component of the MV participating in filtering, mv ij .depth is the depth of the MV of the projection pixel, and mv ij .depth is the depth of the MV of a given pixel.
[0189] The MV can be accumulated within a specified range to determine the filtered MV, as indicated at 728. The specified range can be [-M, 0] and [0, M]. The accumulation of the MV can be performed according to Equation (33):
[0190]
[0191] where x0 flt and y0 flt are the filtered horizontal motion and vertical motion, the MV is within the range of [-M, 0], and x1 flt and y1 flt are the filtered horizontal motion and vertical motion, and the MV is within the range of [0, M].
[0192] For pixels surrounded by both background pixels and foreground pixels, foreground MV dilation can be performed, as indicated at 730. Foreground MV dilation can reduce the void area between the foreground and the background. In some instances, the repair algorithm can also use depth information to facilitate filling the void with background information rather than foreground information.
[0193] The filtered MVs (x0 flt , y0 flt ) and (x1 flt , y1 flt ) can be merged. According to Equation (34), the MV with the largest difference compared to the MV of a given pixel can be used for the output MV (fmv ij ), as at 732:
[0194]
[0195] where dist0 = |mvij .x - x0 flt |+|mv ij .y - y0 flt |and dist1 = |mv ij .x - x1 flt |+|mv ij .y - y1 flt |,
[0196] and
[0197] The filtered MV can be determined for each pixel of the image in this way, so as to determine the filtered MV field. Along with the output filtered MV field, large void regions can be divided into smaller voids.
[0198] At 734, method 700 includes performing a repair algorithm to fill the smaller voids. Applying the repair algorithm to the smaller voids (instead of the initially generated large voids) can mitigate the unstable results of the repair. In addition, the repair algorithm can be applied to voids surrounded by background pixels.
[0199] In some instances, method 700 can be repeated one or more times to continue reducing the size of the voids, and in some cases substantially eliminate the voids. In either case, whether performing repair or filtering the voids, an output that fills the voids without instability can be generated.
[0200] Now turning to Figure 10 , a third figure 1000 is shown, which depicts the MV field 1002 and the filtered MV field 1004 from time t to t + 1. When the MV field 1002 is filtered to reduce the size of the voids, the filtered MV field 1004 can be the output filtered MV, as described with respect to method 700.
[0201] The MV field 1002 can include a foreground MV 1006 and a background MV 1008. In some instances, as depicted in the third figure 1000, the foreground MV 1006 and the background MV 1008 can move in opposite directions. In other instances, the foreground MV and the background MV can move in the same direction. The MV field 1002 can include a first void 1010 generated by the foreground MV 1006 and the background MV 1008 moving apart from each other from time t to t + 1. The first void 1010 can have a large first size 1012. In some instances, a void can be considered "large" when its size is above a predefined threshold.
[0202] Similar to the MV field 1002, the filtered MV field 1004 may include a foreground MV 1014 and a background MV 1016. The filtered MV field 1004 may include a plurality of second holes 1018. The size of each of the plurality of second holes 1018 may be smaller than the first size 1012 of the first hole 1010. As described with respect to FIG. 7, most of the plurality of second holes 1018 may come from background pixels, however, some of the second holes 1018 may be located at the edge between the foreground MV and the background MV. For example, a third hole 1020 among the plurality of second holes 1018. The third hole 1020 may undergo foreground MV expansion to reduce the size of the hole, as described with respect to FIG. 7.
[0203] Now turning to Figure 11 , a use case scenario of an image processing system is shown. In some instances, the image processing system may be the image processing system 200 described with respect to Figure 2 and thus use similar component numbers. In some instances, various inputs and outputs are demonstrated in the use case scenario that will occur when performing the above-described methods 300, 400, 500, 600, and 700.
[0204] In the use case scenario shown in the figure, the current image frame CF (denoted as I CF ) and the previous image frame PF (denoted as PF ) and their associated masks are input into the luminance generator 202. The mask may be used to mark objects and special effects in the corresponding frames. The luminance generator 202 outputs the adjusted CF luminance (denoted as Y CF ) and the adjusted PF luminance (denoted as (Y PF ). The adjusted luminances of the CF and PF are input into the motion vector calculator 204. Based on one or more methods (such as methods 400 and 500), the motion vector calculator 204 may output the internal block-level phase 0 Mv field of the PF, the foreground and background MV regions, and the internal block-level phase 1 MV field of the CF. The internal block-level phase 0 MV field of the PF and the foreground and background MV fields, as well as the original PF and its mask, may be input into the previous frame MV processor 210. The internal block-level phase 1 MV field of the CF and the foreground and background MV fields, as well as the original CF and its mask, may be input into the current frame MV processor 208.
[0205] The previous frame MV processor 210 and the current frame MV processor 208 may process the input MV fields and foreground / background MV fields according to one or more methods (such as the above-described method 600). The processing may include decomposing the block level into the pixel level and generating virtual depth information. Thus, the previous frame MV processor 210 may output the internal MV field of each pixel of the PF (denoted as iMV PF0) and the internal virtual depth field of PF (denoted as iD PF0 ). Similarly, the current frame MV processor 208 can output the internal MV field of each pixel of CF (denoted as iMV CF1 ) and the internal virtual depth field of CF (denoted as iD CF1 ). The internal MV field and virtual depth field of PF can be input into the previous frame MVD combiner 216 together with the external MV field between PF and CF of each pixel of PF generated by the game engine, the external depth field generated by the game engine for PF, the change of the depth field from PF to CF generated by the game engine, and the mask of PF. The internal MV field and virtual depth field of CF can be input into the current frame MVD combiner 214 together with the external MV field between CF and PF of each pixel of CF generated by the game engine, the external depth field generated by the game engine for CF, the change of the depth field from CF to PF generated by the game engine, and the mask of CF.
[0206] The previous frame MVD combiner 216 can output the depth field of PF (denoted as D PF ), the change of the depth field from PF to CF (denoted as Z PF0 ), and the MV field between PF and CF for each pixel of PF (denoted as MV PF0 ). The current frame MVD combiner 214 can output the depth field of CF (denoted as D CF ), the change of the depth field from CF to PF (denoted as Z CF1 ), and the MV field between PF and CF for each pixel of CF (denoted as MV CF1 ). The output of the previous frame MVD combiner 216, together with the phase of PF (e.g., p projected from time t to time t + p), which is the time distance between the projected or target image and the input image, and the original PF, can be input into the previous frame reprojection module 218a of the reprojection module 218. The output of the current frame MVD combiner 214, together with the phase of CF (e.g., p - 1 projected from time t + 1 to time t + p), and the original CF, can be input into the current frame reprojection module 218b of the reprojection module 218.
[0207] In the case of obtaining extrapolation, only the current MVD combiner 214 is utilized. A large amplitude of the phase value indicates a large time difference between the projected frame and the input frame. When the phase is positive, it indicates that the input frame is projected into the future. When the phase is negative, it indicates that the input frame is projected into the past. When using extrapolation, only the output of the current frame MVD combiner 214 can be input into the reprojection module 218. In this way, only one frame is used for reprojection, and thus the holes in the reprojected image may be more and / or larger than when MV processing is used for interpolation.
[0208] The reprojected image output from the reprojection module 218 can be input into an image combiner 220 for combination. The combined image 220 can then be input into a filler 222 for filling holes by filling. The filler 222 can output a final image.
[0209] In some instances, although not presented in this use case scenario, the MV field output by the MVD combiner can be filtered to reduce the size of holes before filling, as described with respect to method 700.
[0210] The technical effects of the systems and methods provided herein are that the MV and virtual depth can be estimated and / or calculated based on the input image, rather than obtained from a game engine. The estimation / calculation based on the input image allows for a reduction in computational and / or processing power and thus allows for faster rendering. Additionally, a luminance image can be created based on the RGB components and the type of content within the image, which can improve the MV calculation without increasing the requirements on the system. The MV is calculated and processed at the block level and then the block-level MV is decomposed into pixel-level MV using a bilateral filter, weighted average, and / or guided filter, which can allow for a smoother MV that more closely follows the edges of objects in the luminance image and provides virtual depth.
[0211] Furthermore, when previously covered data is not covered due to motion during reprojection, the MV continuation rate along the motion trajectory of the relevant pixels / blocks can allow for a reduction in the size of holes, which in turn can allow for the application of a filling algorithm and the output of a more stable result. This can reduce the processing requirements by allowing the use of a simpler, less demanding filling algorithm. It can further reduce the latency of the entire system.
[0212] As used herein, an element or step recited in the singular and preceded by the word "a" or "an" should be understood as not excluding a plurality of the recited elements or steps, unless expressly stated to the contrary. Additionally, a reference herein to "one embodiment" is not to be construed as excluding the existence of additional embodiments that also incorporate the recited features. Further, unless expressly stated to the contrary, an embodiment comprising one element or multiple elements having a particular property may include other such elements not having that property. The terms "comprising" and "in which" are used as shorthand equivalents of the respective terms "including" and "wherein". Additionally, the terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements or a particular positional order on their objects.
[0213] This written description uses examples to disclose the invention, including the best mode, and also enables one of ordinary skill in the art to practice the invention, including making and using any device or system and performing any incorporated method. The patentable scope of the invention is defined by the claims and may include other examples that occur to one of ordinary skill in the art. If these other examples have structural elements that are not different from the literal language of the claims, or if they contain equivalent structural elements that are not materially different from the literal language of the claims, then they are intended to be within the scope of the claims.
Claims
1. A method comprising: Receiving a plurality of image frames as input; calculating one or more internal motion vector (MV) fields between a previous frame (PF) and a current frame (CF) in the plurality of image frames; generating a foreground MV field and a background MV field of the one or more MV fields; processing the one or more interior MV fields and the foreground MV fields and the background MV fields to generate one or more depth fields, wherein processing the one or more interior MV fields comprises generating a virtual depth of the one or more interior MV fields to generate one or more MVD fields; and The one or more MVD fields are output for image processing.
2. The method of claim 1, wherein the one or more internal MV fields between the PF and the CF include an internal phase 0 MV field of the PF and an internal phase 1 MV field of the CF.
3. The method of claim 1, wherein the one or more intra MV fields are calculated at a block level, wherein each block includes a group of pixels. 4 . The method of claim 3 , wherein processing the one or more internal MV fields comprises decomposing the one or more internal MV fields from a block level to a pixel level. The method of claim 1 , wherein the plurality of image frames are luminance image frames.
6. The method of claim 1, wherein the plurality of image frames are modified according to pixel values, wherein the pixel values indicate whether a given pixel belongs to an object, a special effect, or otherwise.
7. The method of claim 1, wherein the one or more MV fields are generated internally and combined with one or more externally obtained MV fields to generate the one or more MVD fields.
8. The method of claim 1, wherein detecting the foreground MV and the background MV comprises: Determine potential background MV and foreground MV via MV projection; determining the reliability of each of the potential background MV and foreground MV; as well as Accumulate reliable background MV and foreground MV.
9. The method of claim 1, wherein generating a virtual depth comprises determining a virtual depth for each corresponding frame based on an initial depth value, a first MV difference between a current MV field and a region foreground MV field, and a second MV difference between the current MV field and a region background MV field.
10. The method of claim 1, wherein the processing imaging comprises one or more of frame interpolation, extrapolation, and re-projection.