Method and system for motion vector filtering
By generating output images based on motion vectors and depth fields and filtering along the hollow motion trajectory, the problems of hollows in frame rate conversion and frame interpolation are solved, and higher quality image output and lower computing requirements are achieved.
Patent Information
- Application Number
- CN202411470824.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-10-21
- Publication Date
- 2025-05-30
AI Technical Summary
During the process of frame rate conversion and frame interpolation, the prior art is difficult to effectively reduce the voids generated by reprojection and extrapolation, resulting in large voids in the output image, affecting image quality.
By generating an output image based on the motion vector (MV) field and the depth field, the local foreground MV field is determined and filtered along the motion trajectory of the hollow MV field to reduce the size of the hollow.
This method can effectively reduce the size of the hole, improve image quality, reduce the demand for computing resources, and reduce system delay.
Smart Images

Figure CN120075471A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the subject matter of the present disclosure relate to the field of three-dimensional (3D) computer graphics, and more particularly to motion vector calculation, processing, and filtering. Background Art
[0002] Over the years, the improvement of computer processing power has made real-time video rendering (e.g., for video games or certain animations) increasingly complex. For example, early video games were characterized by pixelated sprites moving on a fixed background, while contemporary video games are characterized by realistic three-dimensional scenes filled with characters. At the same time, the miniaturization of processing components has enabled mobile devices (such as handheld video game devices and smartphones) to effectively support the real-time rendering of high frame rate, high resolution videos.
[0003] 3D graphics videos can be output at various different frame rates and screen resolutions. It may be necessary to convert a video with 3D graphics from one frame rate (and / or resolution) to another frame rate (and / or resolution). To save computing power while still increasing the frame rate, interpolated frames can be used instead of rendering all frames within the video. Interpolated frames can be effectively generated by using motion vectors (also referred to as MVs in the present disclosure), which track the position differences of objects between the current frame (CF) and the previous frame (PF). Summary of the Invention
[0004] In one example, a method includes: determining the presence of one or more holes in an output image, where the output image is generated based on one or more motion vector (MV) fields and a depth field; determining a local foreground MV field and a depth field for each pixel of the output image; defining a hole MV field based on the local foreground MV field; filtering the hole MV field along the motion trajectory of the hole MV field to reduce the size of the one or more holes; and outputting a filtered image having one or more smaller holes.
[0005] It should be understood that the above brief description is provided to introduce some concepts that are further described in the detailed description in a simplified form. It is not meant to identify the key or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims that follow the detailed description. Moreover, the claimed subject matter is not limited to embodiments that solve any disadvantages noted above or in any part of this disclosure. Brief Description of the Drawings
[0006] The drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification, which, together with the following detailed description, is used to explain the present disclosure but does not constitute a limitation to the present disclosure. In the drawings:
[0007] Figure 1Shows a block diagram of an example computing system.
[0008] Figure 2 Shows a block diagram of an example image processing system.
[0009] Figure 3 Shows a flowchart illustrating a method for motion vector calculation, processing, and filtering.
[0010] Figure 4 Shows a flowchart illustrating a method for motion vector calculation.
[0011] Figure 5 Shows a flowchart illustrating a method for foreground and background motion vector detection.
[0012] Figure 6 Shows a flowchart illustrating a method for post-filtering of motion vectors.
[0013] Figure 7A Shows the first part of a flowchart illustrating a method for motion vector filtering to reduce holes for inpainting.
[0014] Figure 7B Shows an illustration of a method for motion vector filtering to reduce holes Figure 7A of the second part of the flowchart.
[0015] Figure 8 Shows a first graph of a motion vector field.
[0016] Figure 9 Shows a second graph of a motion vector field.
[0017] Figure 10 Shows a third graph of an unfiltered motion vector field and a corresponding filtered motion vector field.
[0018] Figure 11 Shows for Figure 2 a block diagram of a use case scenario of an image processing system. DETAILED DESCRIPTION
[0019] The present disclosure describes systems and methods for calculating, processing, and filtering a motion vector (MV) field for use in intra-frame interpolation, frame rate conversion, hole filling reduction, or other operations. In Figure 1 a block diagram depicts a computing system equipped with an image processing system. In Figure 2 a block diagram depicts an exemplary image processing system. In Figure 3 a flowchart exemplifies methods for MV calculation, processing, and filtering. In Figure 4 a flowchart exemplifies a method for MV calculation. In Figure 5The flowchart in [ ] illustrates the method for detecting foreground MV and background MV. In Figure 6 the flowchart in [ ] illustrates the method for MV processing. The flowchart in FIG. 7 illustrates the method for reduced MV filtering for hole filling. In Figure 8 a first example diagram of the MV field is shown. In Figure 9 a second example diagram of the MV field is shown. In Figure 10 a third example diagram of the MV field and the filtered MV field is shown.
[0020] An MV field can be generated that tracks the position differences of objects moving forward in time (such as the position differences between a previous frame (PF) and a current frame (CF)). As explained in this disclosure, two types of MVs are generated, where MV0 represents the motion from PF to CF, and MV1 represents the motion from CF to PF. An MV is generated for each pixel on the screen (or a group of pixels called a block), thus forming an MV field or a set of MVs for the pixels on the screen. As used in this disclosure, a field is defined as a mapping from a set of pixels in a frame to a set of one or more numbers (e.g., components of a vector or a single number).
[0021] If MVs are used during frame rate conversion, a typical rendering engine only outputs a two-dimensional MV1 texture. In this way, the texture does not contain depth content and only includes information about the change in relative screen position as observed in the reference frame of a virtual camera. The depth and foreground and background information of per-pixel MVs can inform how to calculate the 2D components of block-level MVs. Block-level MVs can represent the average or weighted average of the MVs of a pixel block (e.g., an eight-by-eight pixel block) and can be used for frame interpolation or other image processing tasks to reduce processing requirements. The block-level MVs can be converted back to pixel-level MVs to allow other actions or uses based on virtual depth. Regions of a scene with certain relative depth ranges are called foreground (close to the camera), background (far from the camera), and middle range (between the foreground and the background).
[0022] As an example, two objects can be located at different distances from a virtual camera or viewpoint. If the two objects move in the same direction at equal world space distances, the farther object may appear to move a smaller distance in eye space, thus creating a parallax effect where the object farther from the viewpoint appears to move shorter than the object closer to the viewpoint.
[0023] Additionally, each frame can consist of two types of objects: objects with an MV and objects without an MV. Objects characterized by an MV can include moving characters or other objects, the user's view (virtual camera position), and moving parts of the user interface (for games, this can be a health bar or other similar in-game graphics or statistics). Objects without an MV can include, for example, smoke effects, lighting effects, reflections, full-screen or partial-screen scene transitions (e.g., fades and wipes), and / or particle effects. By separating objects with an MV from objects without an MV, improved image processing can be performed. Traditionally, algorithms may attempt to exclude screen areas characterized by objects without an MV. However, this method is not perfect and may cause the blending of nearby objects during the process of frame rate conversion and / or frame interpolation.
[0024] Rendering images in real time on a computer system typically involves the calculation and ordered blending of multiple virtual layers. Examples of such layers include a black background layer at the bottom, which can include some or all of the black areas on the screen. Applications and games can be drawn above the black background to make them visible. Note that some layers may be opaque, fully transparent, or partially transparent. The transparency of the pixels within a given layer can be represented by its alpha mask. Above the internal components of the game / application, a graphical user interface (GUI) for the operating system can be drawn. Traditionally, the components of the operating system and the information from the game / application can be blended together from bottom to top. The blended image can be directly used for display.
[0025] In addition to frame interpolation, typical methods for frame rate conversion to increase the display frame rate also include extrapolation and reprojection. Reprojection is a general method for accelerating real-time rendering by reusing the rendered pixels of adjacent frames. This involves obtaining one or more previously rendered frames and using newer motion information to extrapolate the previous frames into a prediction of what the normally rendered frames would look like. Although reprojection is useful, it cannot produce a perfect output. For example, if previously hidden content is not covered in the new frame, holes appear due to the lack of available pixel information in the source image. As an example, when an object with an MV moves between PF and CF, the background pixels may be left uncovered at the location where the object was in PF but is no longer in CF. These pixel areas cannot be correctly filled by reprojection alone, and thus there are various repair algorithms to fill such holes, such as copying the pixel values of the background texture to fill the holes, selecting random neighboring pixels to fill the holes, applying a low-pass filtered intermediate image to fill the holes, etc. However, when the uncovered holes are large, the resulting output of the repair algorithms may look unstable.
[0026] Similarly, extrapolation is a method for using existing data (such as MV and MVD data) to estimate future data. In image processing and particularly in video game processing, the MV calculated between the PF and the CF can be used for extrapolation. However, as a result, large holes may appear in the resulting image.
[0027] As discussed in the present disclosure, intra-frame interpolation refers to a systematic process of generating one or more additional frames from the visual information of a previous frame PF and a current frame CF. The additional frames are referred to as interpolated frames or IFs. The IFs can be sequentially displayed between the PF and the CF, with a uniform spacing between them in some instances and a non-uniform spacing between them in other instances. As an example, frame rate conversion via intra-frame interpolation can be used to generate one IF between a given pair of PF / CF, resulting in an output with a frame rate that is twice that of the original video. The generation of the IFs may be desirable due to its low computational cost, as generating more frames from the game engine itself may require a higher degree of computational complexity and thus more processing power.
[0028] For video games, the depth can be provided by the game engine at the pixel level and used for block-to-pixel MV conversion. However, additional information is usually required. For example, when there is a shadow cast by another object on the object, the depth of the object may be irrelevant to the motion. Generally, the lighting, ray tracing, and other special effects applied in the game make the MV unsuitable for intra-frame interpolation and must be replaced with internally generated MV. Then, such internally generated MV needs to be converted from the block level to the pixel level. The depth provided by the game engine is also not suitable for converting the MV from the block level to the pixel level because its value does not change when the MV changes.
[0029] To complete frame rate conversion via one or all of intra-frame interpolation, extrapolation, and reprojection, high-quality MV and depth information can be obtained from the game engine, but this may result in consuming more computational resources and more processing power. As described in the present disclosure, a flexible architecture for intra-frame interpolation is proposed, which allows the calculation, processing, and filtering of the MV from the game engine to reduce the use of computational resources for image processing. In some instances, the MV and depth information can be obtained from the game engine, referred to as external MV and depth, and / or, in other instances, the MV and depth information can be estimated based on the input image, referred to as internal MV and depth. The external MV and depth and the internal MV and depth can be combined and / or used in combination during processing and / or filtering for intra-frame interpolation. Internal virtual depth information can be generated based on the internal block-level MV, allowing the conversion of the MV from the block level to the pixel level.
[0030] Extrapolation and interpolation can be accomplished by calculating and processing MVs in terms of virtual depth, block level, and pixel level. The methods and systems provided by this disclosure allow for MV calculation and processing at the block level based on the obtained input images rather than the input images from the game engine. The methods and systems may include detection of foreground MVs and background MVs, generation of virtual depth information, and conversion from block level to pixel level to allow use in processes such as extrapolation and interpolation. In instances where extrapolation is performed, only one image may be output. When extrapolation is not performed, two or more images may be output and then merged into a single image.
[0031] The added depth information can provide additional flexibility to the image processing system. The depth component of the MV can be used as an input to, for example, a weighting function. The correlation between depth and the MV is generally higher compared to the actual image, even for small regions. By determining the foreground MV and background MV of a region, the virtual depth of a scene can be obtained, for example, in a video game. When the virtual depth is converted to the pixel level, it can allow the conversion of block-level MVs to pixel-level MVs and allow for more accurate interpolation by considering the change in depth. As mentioned in this disclosure, the MV with added depth information is also referred to as MVD (MV with depth).
[0032] In addition, MV filtering may include filtering the MV along the motion trajectory and changing the MVs of the background near the uncovered region. Doing so can divide large uncovered holes generated by the extrapolation and / or reprojection process into smaller hole regions. The smaller hole regions can be further filtered to completely reduce the holes, or a filling algorithm can be used to fill the smaller hole regions, which can mitigate any instability that may result from performing the filling algorithm on large holes.
[0033] Aspects of the present disclosure may be described in the present disclosure with reference to flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments disclosed herein. To the extent that such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, those skilled in the art will understand that each function and / or operation in such block diagrams, flowcharts, or examples can be implemented, individually and / or jointly, by computer-readable instructions using a variety of hardware, software, firmware, or in fact any combination thereof. The systems described are exemplary in nature and may include additional elements and / or omit elements. The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various systems and configurations disclosed, as well as other features, functions, and / or properties. Accordingly, these methods can be performed by utilizing one or more logical devices (e.g., processors) in combination with one or more additional hardware elements (such as storage devices, memories, hardware network interfaces / antennas, switches, actuators, clock circuits, etc.) to execute instructions stored on a machine-readable storage medium. The described methods and associated actions can also be performed in various orders in parallel and / or simultaneously, in addition to the order described in this application. The processors of the logic subsystem can be single-core or multi-core, and the programs executed thereon can be configured for parallel processing or distributed processing. Optionally, the logic subsystem can include separate components distributed in two or more devices, which can be remotely located and / or configured for coordinated processing. One or more aspects of the logic subsystem can be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration.
[0034] Figure 1 An example of a computer system 100 is schematically depicted, which may include one or more processors 110, volatile and / or non-volatile memory 120 (e.g., random access memory (RAM) and / or one or more hard disk drives (HDD)). The processor 110 may include one or more central processing units (CPUs), GPUs, and / or one or more image processing systems, such as the image processing system 115. The computer system may also include one or more displays 130, which may include any number of visual interface technologies. Additionally, example embodiments may include a user interface 140 (e.g., keyboard, computer mouse, touch screen, controller, etc.) to allow a user to provide input to the computer system. In some embodiments, the computer system may be a mobile phone or a tablet. The present disclosure is primarily concerned with systems that describe the transfer of data within the processor 110. The transfer of data within the processor may include data generated by the CPU / GPU and transferred to the image processing system 115, or vice versa. It should be noted that the image processing system 115 may be a hardware unit external to the CPU / GPU or may be a program operating within the CPU / GPU.
[0035] As used in this disclosure, the terms "system" or "module" may include a hardware and / or software system that operates to perform one or more functions. For example, a module or system may include a computer processor, a controller, or other logic-based device that performs operations based on instructions stored on a tangible and non-transitory computer-readable storage medium, such as a computer memory. Alternatively, a module or system may include a hardwired device that performs operations based on hardwired logic of the device. The various modules or units shown in the figures may represent hardware that operates based on software or hardwired instructions, software that directs the hardware to perform operations, or a combination thereof.
[0036] It should be noted that the techniques discussed in this disclosure are applicable not only to games, but also to any animated rendering of 3D models, although the advantages provided by the method may be most apparent in the case of real-time rendering.
[0037] Figure 2 A block diagram of an image processing system 200 is shown. The image processing system 200 may be included as part of a Figure 1 computer system 100 or otherwise communicatively coupled to the computer system. In some instances, the image processing system 200 may be the image processing system 115 of the processor 110. In other instances, the image processing system 200 may be separate from the image processing system 115 of the processor 110.
[0038] The image processing system 200 may include multiple modules to perform various functions and / or store processor-executable instructions in a memory. The image processing system 200 may include a luminance generator 202, an MV calculator 204, an MV processor 206, an MV depth combiner 212, a reprojection module 218, an image combiner 220, and a restorer 222, each of which performs or stores executable instructions to perform various tasks for MV calculation, processing, and / or filtering. In some instances, not all of the modules described in this disclosure may be included in the image processing system 200.
[0039] The luminance generator 202 may generate luminance images of two consecutive frames (e.g., PF and CF). The luminance image may refer to the luminance component of an image, where the luminance component is the y-component, the y-channel component, or Y'CbCr. The luminance generator 202 may obtain input images as PF and CF and generate a luminance image from the input images. MV calculation and thus the MV calculator 204 may use the luminance of the image as an input. If the input image is red-green-blue (RGB), a color space conversion is required to obtain the luminance image, which may be performed by the luminance generator 202.
[0040] As described above, for some applications such as video games, external MVs can be used for background and object masks as well as special effects (e.g., fog, shadows, etc.). Thereby, a luminance image can be generated and / or modified based on whether a pixel belongs to an object, a special effect, or others. For example, if the current pixel belongs to an object, the output luminance image can be generated based on Equation (1):
[0041] YO (i,j) = YI (i,j) * k 0 + thr 0 (1)
[0042] where YO is the output luminance image, YI is the input luminance, (i, j) is the pixel position, k 0 is the weight of the input luminance value, and thr 0 is the change to the luminance value based on the classification of the pixel as an object.
[0043] If the current pixel belongs to a special effect, the output luminance image can be generated based on Equation (2):
[0044] YO (i,j) = YI (i,j) * k 1 + thr 1 (2)
[0045] where the variables are as described above. Additionally, if the current pixel belongs to another type of pixel other than an object or a special effect, the output luminance image can be generated based on Equation (3):
[0046] YO (i,j) = YI (i,j) * k 2 + thr 2 (3)
[0047] where the variables are as described.
[0048] In some instances, the pixel values of the luminance image can be measured in the range from 0 to r, and typical values of k 0 are 0.5, k 1 is 0.4, k 2 is 0.4, thr 0 is half of r, thr 1 is 0, and thr 2is 0. The k value and the thr value can be used to tune the performance of the system. They can also be included as metadata to allow them to be optimized for a specific game. In some cases, some pixel regions may use internal MVs (e.g., iMV), and some pixel regions may use external MVs (e.g., eMV). In the case of using both iMV and eMV, a hybrid MV (e.g., hMV) may be generated. The modification of the luminance image described above may allow for a more accurate MV for the hMV case.
[0049] The MV calculator 204 can calculate or store executable instructions to calculate the MV and perform MV correction based on the generated luminance image. As will be further described, the MV can be calculated between the PF and the CF for a given pixel. For a pixel moving from position (i, j) in the PF to position (u, v) in the CF, the motion at position (i, j) in the PF is (u - i, v - j) pixels, and the motion at position (u, v) in the CF is also (u - i, v - j) pixels.
[0050] In some instances, the MV calculator 204 can calculate the MV at the block level rather than the pixel level, as described above. The block can be square, rectangular, or other groups of pixels, such as 8x8 pixels, 4x4 pixels, etc. Each block can be defined by the position or pixel at its center. The MV can be calculated for the block based on the center positions in the PF and the CF in a similar manner as described above. The MV calculator 24 can generate an MV field based on the MVs calculated for multiple blocks within the input image.
[0051] In addition, the MV calculator 204 can generate a region foreground MV and a region background MV. In some instances, the input image or the luminance image can be segmented into H*V regions, where H represents the number of regions in the horizontal direction, and V represents the number of regions in the vertical direction. The foreground MV and the background MV can be generated for each region with an overlapping window. In some instances, the ratio between the sizes of the block regions can approximate the aspect ratio of the image.
[0052] The MV processor 206 can include a current frame MV processor 208 and a previous frame MV processor 210. The current frame MV processor 208 can process the MV of the CF, and the previous frame MV processor 210 can process the MV of the PF. Both the current frame MV processor 208 and the previous frame MV processor 210 can output a pixel-level MV and a pixel-level virtual depth by filtering the block-level MV, generating a block-level virtual depth, and decomposing the block-level MV and depth into a pixel-level MV and depth. Having both the current frame MV processor 208 and the previous frame MV processor 210 allows for the separate processing of the MV of phase 0 and the MV of phase 1.
[0053] The MV depth combiner 212 may include a current frame MV depth combiner 214 and a previous frame MV depth combiner 216. The MV depth combiner 212 may combine or store executable instructions to combine the iMV and internal depth with the eMV and external depth based on the marked objects and special effects in the previous frame (e.g., via the previous frame MV depth combiner 216) and / or in the current frame (e.g., via the current frame MV depth combiner 214).
[0054] The architecture of the image processing system 200 may be flexible as it responds to and meets the performance requirements of the coupled computing system and / or game engine. As such, the MV may be obtained from the game engine as the eMV and depth, or the MV may be estimated as the iMV and depth based on the obtained images. In some instances, if the performance requirements of the game engine are high, the image processing system 200 may compensate by using more iMV and depth to reduce the processing requirements of the game engine. Since the image processing system 200 also supports the combination of eMV and iMV, the MV depth combiner 212 may combine these two sets of information according to a mask. The mask may include mask information for objects, shadows, and special effects. For regions whose motion may not be expressed by the eMV, such as objects, shadows, and special effects, iMV calculations may generate iMV for these regions.
[0055] The reprojection module 218 may execute or store executable instructions to perform reprojection of the input image from time t (e.g., the time of the previous frame) to a new time position t + p (or from time t + 1 to time t + p, where p <= 1) based on the calculated MV, depth, and phase values. When going from time t to t + p, the phase value is p (or when going from time t + 1 to time t + p, the phase value is p - 1), assuming the time difference between adjacent frames is 1. In some instances, the reprojection module 218 may include a current frame reprojection module and a previous frame reprojection module similar to the MV depth combiner 212 and the MV processor 206, where the MV of the calculated and processed PF is reprojected via the previous frame reprojection module, and the MV of the calculated and processed CF is reprojected via the current frame reprojection module. In some instances, the reprojection module 218 may include instructions stored in the memory to filter the obtained MV along the motion trajectory to reduce the holes not covered during reprojection.
[0056] The image combiner 220 may obtain the reprojected images and combine them together. Similarly, the image combiner 220 may obtain the reprojected images of the previous frame and the current frame from the respective reprojection modules and may combine them together.
[0057] The restorer 222 may include one or more restoration algorithms or store instructions for executing one or more restoration algorithms. The one or more restoration algorithms may be configured to perform hole filling for pixel positions not re-projected by the re-projection module 218. As previously described, when previously hidden content is not covered in a new frame, holes occur due to the lack of available color information in the source image. As an example, when an object in the foreground with an MV moves between the PF and the CF, background pixels may be left uncovered at the positions where the object was in the PF but is no longer in the CF. These pixel areas cannot be correctly filled by re-projection alone, and the restorer 222 may fill the holes via one or more restoration algorithms. In some instances, the processing and filtering of the MV performed by the image processing system 200 may reduce the size of the holes, as will be further described below.
[0058] Now turning to Figure 3 , a flowchart of an example method 300 for MV calculation, processing, and filtering is shown. The method 300 may be implemented using the systems and components described above with respect to Figures 1 to 2 For example, the method 300 may be implemented by one or more processors according to instructions stored in a memory. For example, the instructions may be stored and executed by an image processing system such as the image processing system 200. The image processing system described in the present disclosure may be part of a computer system (such as the computer system 100) including a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system.
[0059] At 302, the method 300 includes generating a luminance image as an input. In some instances, an image may be obtained and processed to output a luminance image, as described above. The luminance image may include a plurality of pixels at specific positions. As previously described, for some applications (such as video games), external MVs may be used for background and object masks as well as special effects (e.g., fog, shadows, etc.). Thus, based on the equations (1), (2), and (3) defined above, the luminance image may be generated and / or modified based on whether the pixel belongs to an object, a special effect, or others. In some instances, the luminance image may include the PF and the CF.
[0060] At 304, the method 300 includes calculating the MV between the PF and the CF. The MV may be calculated for a specific pixel or block at a defined position in one of the PF and the CF. The MV of the pixel from the PF to the CF or from the CF to the PF may be calculated. An MV field including a plurality of MVs between frames may also be generated. In some instances, the MV calculation at 304 may include foreground MV and background MV detection. Additionally, the MV calculation may include calculating the motion and quality of the MV, which includes the calculation of the sum of absolute differences (SAD) and the difference (MVDIFF) between the MV of the current block and the MV of the block pointed to by the current block. Regarding Figure 4and Figure 5 MV calculation (including quality determination and detection of background MV and foreground MV) is further described. In some examples, MV can also be calculated between the previous frame (PPF) and the PF, such as Figure 5 Further discussion.
[0061] At 306, method 300 includes processing the MV. Figure 6 As further described, processing the MV may include MV post-filtering, generation of virtual depth, and block-level MV decomposition. MV post-filtering may be based on the MV quality determined at 304. MV processing may output a block-level or pixel-level MV with depth information, such as may be used for frame interpolation. In some examples, processing the MV may include filtering along a motion trajectory to reduce the size of uncovered holes between PF and CF.
[0062] At 308, method 300 includes re-projecting the input image according to the MV. In some examples, a block-level MV and a corresponding pixel-level MV may be obtained. Virtual depth information and a phase value of the MV may be obtained and / or determined, and the input image may be rendered based on the depth and phase values. The phase value may refer to a temporal distance between the rendered image and the original input image.
[0063] At 310, method 300 includes merging the multiple rendered images together. In some examples, when interpolating, PF is forward projected from time t to time t+p, and its phase value is p, where p is between 0 and 1. And CF is backward projected from time t+1 to time t+p, and its phase value is p-1. This produces two images that are merged to create the final image. In other examples, such as when extrapolation is used, the MV of CF is 1 The field may be projected forward from time t+1 to time t+1+p, in which case only one image may be generated and merging may be skipped. In some instances, one or more regions in the merged image may include pixels / blocks with no data because the movement of pixels according to the MV is not covered.
[0064] At 312, method 300 includes filling holes in the merged image that are not covered via an inpainting algorithm. In some instances, the inpainting algorithm can be applied directly to the merged image, such as an inpainting algorithm that copies background textures around the holes, selects random neighboring pixels, and / or uses an intermediate image with low-pass filtering to fill the holes. In other instances, as will be further described below, the MVs can be filtered along the motion trajectories, and the MVs of the background pixels / blocks near the uncovered regions can be changed to divide the uncovered holes into smaller holes, as mentioned at 306. In some instances, the inpainting algorithm can then be applied to the smaller holes. In other instances, the MV filtering for reducing holes can be repeated one or more times until the holes are substantially eliminated.
[0065] Method 300 can allow for processed and filtered MVs, which can be used for intra-frame interpolation or other image processing including extrapolation and reprojection. The processed and filtered MVs can be generated in a more flexible manner, allowing the use of iMVs and eMVs based on GPU / CPU requirements. Compared to eMVs, MVs including internal estimation and / or calculation can allow for lower computational and / or processing capabilities. Additionally, holes generated during image processing such as reprojection and / or extrapolation can be filled by filtering the MVs.
[0066] Now turning to Figure 4 , a flowchart of an example method 400 for calculating MVs is shown. Method 400 can be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, method 400 can be implemented by one or more processors according to instructions stored in a memory. For example, the instructions can be stored and executed by an image processing system such as image processing system 200. The image processing system described in the present disclosure can be part of a computer system (such as computer system 100) including a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. Method 400 can be incorporated into Figure 3 method 300, particularly at 302.
[0067] At 402, method 400 includes generating a luminance input. As described above with respect to Figure 2 , luminance images can be generated for two consecutive frames (e.g., PF and CF). The luminance input can be generated and / or modified based on equations (1), (2), and / or (3), based on whether the pixel belongs to an object, special effect, or other. For example, for a current pixel belonging to an object, an output luminance image can be generated based on equation (1), for a current pixel belonging to a special effect, an output luminance image can be generated based on equation (2), and for a current pixel belonging to a pixel type other than an object or special effect, an output luminance image can be generated based on equation (3).
[0068] In some cases, some pixel regions may use internal MVs (e.g., iMV), and some pixel regions may use external MVs (e.g., eMV). In cases where both iMV and eMV are used, a hybrid MV (e.g., hMV) may be generated. The above modification of the luminance image may allow for a more accurate MV for the hMV case.
[0069] At 404, method 400 includes calculating the MV between a PF and a CF based on a luminance input image. The multiple MVs calculated between the PF and the CF at different positions may together form an MV field. Calculating the MV may include calculating the MV for each block of the PF (as shown at 406) and calculating the MV for each block of the CF (as shown at 408). In some instances, as previously described, pixel blocks may be defined for the CF and the PF. For example, blocks of 8x8 pixels may be used, although other sizes and shapes of blocks are possible.
[0070] Each block of the luminance image may have a defined position. As an example, for a block (m, n) with a block size of bxb, the upper left position of the block may be [b*m, b*n], and the lower right position of the block may be [b*m + b - 1, b*n + b - 1]. In some instances, the center position of the block may be proportional to (m, n).
[0071] As an example, for a block (m,n) in the PF with an MV [y mn , x mn , its upper left position moves from the first position [m*b, m*b] in the PF to the second position [m*b + ym n, n*b + x mn in the CF. The MV in the CF has the same polarity as the MV in the PF. If the MV of a block (c,d) in the CF is [y cd , x cd , then its upper left position moves from the position [c*b, d*b] in the CF to the position [c*b - y cd , d*b - x cd in the PF.
[0072] At 410, method 400 includes determining the quality of the calculated MV. In some instances, the calculation of the MV may incorporate the determination of various values, including but not limited to the horizontal motion component, the vertical motion component, the sum of absolute differences (SAD) between the matching blocks (as indicated at 412), and the difference (MVDIFF) between the MV of a block and the MV of the block it points to (as indicated at 414).
[0073] For a block in the CF with an MV [y cd , x cdThe block (c, d), and the SAD of the block (c, d) is: And for the block (m, n) in PF having MV[y mn , x mn , the SAD of the block (m, n) is: In some instances, a smaller SAD value indicates a good match between blocks, and a larger SAD value indicates a poor match.
[0074] MVDIFF can indicate whether the block (m, n) is in a covered area or an uncovered area. A small MVDIFF indicates that the MV of PF is confirmed by the opposite MV of CF, and thus may not be in the covered area / uncovered area. A large MVDIFF indicates that the block is in the covered area / uncovered area, such that the current MV is not confirmed by the opposite MV of CF. For MV 0 , a large MVDIFF can indicate that the block is in the covered area. For MV 1 , a large MVDIFF can indicate that the block is located in the uncovered area. The MVDIFF of the block (m, n) of PF can be determined by Equation (4):
[0075] Where, and iBMV PF0(m,n) .mvdiff is the MV field of PF for the block (m, n), iBMV PF0(m,n) .x is its horizontal motion, and iBMV PF0(m,n) .y is the vertical motion, and iBMV CF1(c′,d′) .x is the horizontal motion of the MV field of CF for the block (c′, d′), and iBMV CF1(c′,d′) .y is the vertical motion.
[0076] The MVDIFF of the block (c, d) of CF can be determined by Equation (5):
[0077]
[0078] Where, and iBMV CF1(c,d) .mvdiff is the MV field of CF for the block (c, d), iBMV CF1(c,d) .x is its horizontal motion, and iBMV CF1(c,d) .y is the vertical motion, and iBMV PF0(m′,n′) .x is the horizontal motion of the MV field of PF for the block (m′, n′), and iBMV PF0(m′,n′) .y is the vertical motion.
[0079] Briefly turning to Figure 8, shows a first diagram 800 of the MV field. The first diagram 800 shows the MVDIFF calculation as described at 414 of method 400. The first line 802 represents the internal block-level phase 0 MV field of the PF (e.g., iBMV PF0 ), and the second line 804 represents the internal block-level phase 1 MV field of the CF (e.g., iBMV CF1 ). A plurality of block-level MVs 850 are represented as arrows between the first line 802 and the second line 804. For example, for the first block 806 (e.g., the first block (m, n)), the MV of the first block 806 is represented by the arrow 808. The first block 806 may be included in the coverage area 814 such that its MV is invalid up to the CF. For a coverage area where the MV is invalid up to the corresponding next frame, the MVDIFF can be large. When the MV is valid for the corresponding next frame, a smaller MVDIFF can be generated.
[0080] To determine the MVDIFF and whether a block is in a covered area or an uncovered area, an MV projection can be performed to project the MV to hit its next block. The MV projection can be performed from the iBMV PF0(m,n) to the second block 812 (e.g., the second block (c′, d′) in the iBMV CF1 ). The second block can be determined based on equation (6):
[0081]
[0082] where the variables are as described above. Then the MV from the CF to the PF can be calculated based on the position of the second block. The resulting positions of c′ and d′ and the MV from the CF to the PF can then be used to calculate the MVDIFF, as described with respect to equation (4) above. If the calculated MVDIFF is large, the first block is in the covered area. The dashed line 810 represents a set of phases where the MV is invalid due to the first block being covered. The determination of the MVDIFF as described in the present disclosure allows determination of whether the first block is covered or uncovered.
[0083] Returning to Figure 4 , at 416, method 400 includes detecting foreground MVs and background MVs. As will be further described with respect to Figure 5 , the detection of foreground MVs and background MVs can be an algorithm based on four MV fields. In addition to the MV fields between the PF and the CF (e.g., iBMV PF0 and iBMV CF1 ), the block-level MV fields between the PPF and the PF (e.g., iBMV PP0 and iBMV PF1) It can also be used to determine the foreground MV and the background MV. It should be understood that the calculation of the block-level MV field between the PPF and the PF can simply be the previously calculated block-level MV field, for example, between the PF and the CF that is one frame delayed from the currently calculated frame. By comparing the SAD value with the MVDIFF value, an occluded area can be found, and if so, the local foreground MV and the background MV can be detected. As described in the present disclosure, the occluded area can be an area where a moving object covers something that was not previously covered or exposes the background that was previously covered.
[0084] The MV field including foreground information and background information calculated as described in the present disclosure can be processed, as will be further described below, to generate virtual depth. The virtual depth allows the rendered image to be used for frame interpolation.
[0085] Now turning to Figure 5 , a flowchart illustrating a method 500 for detecting the foreground MV and the background MV is shown. The method 500 can be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, the method 500 can be implemented by one or more processors according to instructions stored in a memory. For example, the instructions can be stored and executed by an image processing system such as the image processing system 200. The image processing system described in the present disclosure can be a part of a computer system (such as the computer system 100) including a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. The method 500 can be incorporated into Figure 4 the method 400, particularly at 416.
[0086] At 502, the method 500 includes obtaining an MV for a first block (e.g., block (m, n)). It should be understood that the first block (m, n) representing each block of the input image and the method 500 can be executed sequentially or simultaneously for one or more blocks of the input image. As described with respect to Figure 4 , the MV fields can be obtained between the PF and the CF and between the PF and the PFF. The MV field between the PF and the CF can include the MV of each block of the PF (e.g., the internal block-level phase 0 MV or iBMV of the PF PF0 ) and the MV of each block of the CF (e.g., the internal block-level phase 1 MV or iBMV of the CF CF1 ). The MV field between the PPF and the PF can include the MV of each block of the PPF (e.g., the internal block-level phase 0 MV or iBMV of the PPF PP0 ) and the MV of each block of the PF (e.g., the internal block-level phase 1 MV or iBMV of the PF PF1)。When calculating the MV of each block for PF, CF, and PPF, the MV of the first block (m, n) can be obtained. Additionally, obtaining the MV of the first block (m, n) can include obtaining the SAD value and MVDIFF value of the block, as they are calculated during the calculation of the MV, as described with respect to Figure 4 as described.
[0087] At 504, method 500 includes determining whether the first block (m, n) is covered or not covered. If the MVDIFF of the iBMV PF0 (e.g., the MV field between PF and CF for (m, n) of PF) is greater than the MVDIFF of the iBMV PF1 (e.g., the MV field between PPF and PF for (m, n) of PF), then the first block (m, n) may be a potentially covered block. If the MVDIFF of the iBMV PF0 is less than the MVDIFF of the iBMV PF1 , then the first block (m, n) may be a potentially uncovered block.
[0088] At 506, method 500 includes determining the potential background MV and occlusion MV. The determination of the potential background MV and occlusion MV is based on whether the first block is covered or not covered, as determined at 504. As an example, if the first block (m, n) is covered, then the MV field between the PPF and PF of the first block (e.g., iBMV PF1 ) can be set as the background MV, for example as shown in equations (7) and (8):
[0089] MV BG = iBMV PF1(m,n) (7)
[0090] MV occ = iBMV PF0(m,n) (8)
[0091] where MV BG is the background MV, and MV occ is the occlusion MV.
[0092] At 508, method 500 includes performing MV projection. The MV projection can be performed using the background MV from the first block (m, n) to the CF of the first block (m, n) (e.g., the internal block phase 1 of the CF), as indicated at 510. The MV projection may hit the second block (u, v) to obtain the potential foreground. The MV projection can be performed based on equation (9):
[0093] MV FG = iBMV CF1(u,v) (9)
[0094] where MVFG is the foreground MV, and iBMV CF1(u,v) is the MV field between the PF and the CF of the second (u, v) block's CF.
[0095] Then, the foreground MV from the second (u, v) block to the PF (e.g., the internal block phase 0 of the PF) can be utilized to perform MV projection, as indicated at 512. The MV projection may hit the third (s, t) block. The MV projection of the third (s, t) block can be based on Equation (10):
[0096]
[0097] where is the hit MV of the foreground, and IBMV PF0(s,t) is the internal block phase 0 MV of the PF of the third (s, t) block.
[0098] Then, the foreground MV from the first (m, n) block to the CF (e.g., the internal block phase 1 MV of the CF) can be utilized to perform MV projection, as indicated at 514. The MV projection may hit the fourth (p, q) block. The MV projection of the fourth (p, q) block can be based on Equation (11):
[0099]
[0100] where is the hit MV of the background, and iBMV CF1(p,,q) is the internal block phase 1 MV of the CF of the fourth (p, q) block.
[0101] MV projection can be performed for each block in the corresponding input image frame. As described in this disclosure, MV projection can provide potential foreground MV and background MV via Equations (9), (10), and (11). The foreground MV and background MV can be global and / or regional.
[0102] Briefly turning to Figure 9 , a second Figure 900 is shown, which illustrates the MV projection for detecting foreground MV and background MV as described at 508, 510, 512, and 514 of Method 500. The second Figure 900 includes a first line 902 representing the internal block level phase 0 MV field of the PPF (e.g., iBMV PP0 ), a second line 904 representing the internal block level phase 1 MV field of the PF (e.g., iBMV PF1 ), a third line 906 representing the internal block level phase 0 MV field of the PF (e.g., iBMV PF0 ), and a fourth line 908 representing the internal block level phase 1 MV field of the CF (e.g., iBMV CF1) the fourth line 908. A plurality of block-level MVs are represented as arrows between the first line 902 and the second line 904 and between the third line 906 and the fourth line 908.
[0103] The first block 910 of PF (e.g., the first block (m, n)) may be an overlay block. The internal block phase 1 MV of PF of the first block 910 may be represented by line 950. As described with respect to method 500, the internal block phase 1 MV of PF of the first block (m, n) (e.g., iBMV PF1(m,n) ) may be set as a potential background MV, such as the MV described in the above equation (7) BG , and when the first block (m, n) is located in the overlay area 914, the phase 0 MV of PF of the first block (m, n) may be an occlusion MV (e.g., MV occ ), as described in the above equation (8). The occlusion MV may be represented by the first arrow 912.
[0104] The MV projection from the first block 910 to the second block 916 may be represented by the first dashed arrow 952 in the second figure 900. The first dashed arrow 952 may be the internal block-level phase 1 MV of CF of the second block 916, as described with respect to equation (9), which may be a foreground MV (e.g., MV FG ).
[0105] Then, an MV projection from the second block 916 to the third block 920 (e.g., the third block (s, t)) may be performed. The MV generated by the MV projection from the second block 916 to the third block 920 is represented by the second arrow 918 in the second figure 900. When it is PF, the MV represented by the second arrow 918 may be a dc foreground MV (e.g., ), as described in the above equation (10).
[0106] An MV projection may also be performed from the first block 910 to the fourth block 924 (e.g., the fourth block (p, q)). Similar to as described with respect to Figure 8 , the internal block-level phase 0 MV of PF of the first block 910 may be invalid up to the fourth block 924. An MV projection may be performed to reach the fourth block 924. The second dashed arrow 922 may represent the MV obtained by the MV projection. The resulting MV may be the internal block-level phase 1 MV of CF of the fourth block 924, which, as described in the above equation (11), may be a dc background MV (e.g., ).
[0107] Each of the foreground MV and the background MV determined via MV projection as described in method 500 and shown in the second figure 900 may be a potential MV. Some of the potential MVs may be reliable, and some of the potential MVs may be unreliable.
[0108] Return to Figure 5 At 516, method 500 includes determining the reliability of the foreground MV and the background MV. Determining the reliability of the foreground MV and the background MV may include defining a previous global foreground MV and a previous global background MV, as well as a previous regional foreground MV and a previous regional background MV. The previous regional foreground MV and background MV may be related to the region to which the first block (m, n) belongs. Each of the previous foreground MV and background MV including the corresponding global and regional ones may be defined according to horizontal motion and vertical motion. In addition, a δ (e.g., the pixel distance between the positions pointed to by the two MVs) between the foreground MV and the background MV with respect to horizontal motion and vertical motion may be determined, such as that described in equation (12) for the foreground MV:
[0109]
[0110] where mvdist FG is the δ of the foreground MV, MV FG .x is the horizontal component of the foreground MV, and MV FG .y is the vertical component of the foreground MV. Similar equations may be used for the background MV.
[0111] In some instances, the difference between the foreground MV and the background MV may be determined according to equation (13):
[0112] mvdist FGBG = |MV FG .x - MV BG .x| + |MV FG .y - MV BG .y| (13)
[0113] where mvdist FGBG is the δ between the foreground MV and the background MV.
[0114] Then the δ from the foreground to the global foreground and global background and from the background to the global foreground and global background may be defined. For example, the MV distance from the foreground to the global foreground may be defined according to the corresponding horizontal and vertical components of the foreground MV and the previous global foreground MV, the MV distance from the foreground to the global background may be defined according to the corresponding horizontal and vertical components of the foreground MV and the previous global background MV, the MV distance from the background to the global foreground may be defined according to the corresponding horizontal and vertical components of the background MV and the previous global foreground MV, and the MV distance from the background to the global background may be defined according to the corresponding horizontal and vertical components of the background MV and the previous global background MV.
[0115] Based on these defined distances, the bGoodBGFG flag can be determined. The bGoodBGFG flag can be true when the following two conditions are met: the difference between the background MV and the global background MV is less than half of the difference between the background MV and the global foreground MV and the difference between the foreground MVs, and the global foreground MV is less than half of the difference between the foreground MV and the global background MV.
[0116] When bGoodBGFG is true and the SAD of the background MV is minimized, or when 1) the SAD of the background MV is minimized, 2) the maximum value of the MVDIFF of the background MV and the SAD of the background MV is less than the minimum value of thr1 and half of the difference between the foreground MV and the background MV, 3) the difference between the foreground MV and the background MV is greater than thr2, 4) the SAD of the occluded MV is greater than the maximum value of thr3 and the SAD of the background MV, and 5) the MVDIFF of the occluded MV is greater than thr4, the background MV (MV BG ) may be reliable.
[0117] In other words, the background MV can be reliable when one of the following conditions is met:
[0118] Reliable BG 1 :(bGoodBGFG && MV BG .sad < min(thr0, MV occ .sad)
[0119] Reliable BG 2 :
[0120]
[0121] Reliable FG = Reliable BG 1 || Reliable BG 2
[0122] The foreground MV may be reliable when 1) the background MV is reliable, 2) the MVDIFF of the occluded MV is greater than thr5, 3) the difference of the foreground MV is less than (thr6, half of the SAD of the foreground MV is less than thr7, 4) the SAD of the foreground MV is less than thr7.
[0123] In other words, the foreground MV may be reliable when the following conditions are met:
[0124]
[0125] At 518, method 500 includes accumulating reliable foreground MVs and background MVs and averaging the foreground MVs and background MVs for each region and for the entire frame. For a given region (h, v), if the reliable number of foreground MVs and background MVs (e.g., fgcnt and bgcnt) is too small, the foreground MVs and background MVs may be changed according to equation (14):
[0126]
[0127] where RMV is the region MV of the foreground or background for horizontal or vertical motion, depending on the equation, GlbFGMV is the global foreground MV, GlbBGMV is the global background MV, wfg (h,v) = min(thr FG , RMV FG(h,v) .fgcnt), and wbg (h,v) = min(thr BG , RMV BG(h,v) .bgcnt).
[0128] In this way, reliable foreground MVs and background MVs can be generated based on the MV fields between CF and PF and between PF and PPF.
[0129] Now turning to Figure 6 , a flowchart of an example method 600 for MV processing is shown. Method 600 may be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, method 600 may be implemented by one or more processors according to instructions stored in a memory. For example, the instructions may be stored and executed by an image processing system such as image processing system 200. The image processing system described in this disclosure may be part of a computer system (such as computer system 100) that includes a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. Method 600 may be incorporated into Figure 3 method 300, particularly at 304.
[0130] At 602, method 600 includes obtaining an MV field. As described with respect to Figure 4 , an MV field may be obtained between PF and CF and between PF and PFF. The MV field between PF and CF may include the MVs of each block of PF (e.g., the internal block-level phase 0 MV or iBMV PF0 ) and the MVs of each block of CF (e.g., the internal block-level phase 1 MV or iBMV CF1 ). The MV field between PPF and PF may include the MVs of each block of PPF (e.g., the internal block-level phase 0 MV or iBMV PP0) and the MV of each block of PF (e.g., the internal block-level phase 1 MV or iBMV of PF PF1 ). When calculating the MV of each block for PF, CF, and PPF, the MV of the first block (m, n) can be obtained. Additionally, obtaining the MV of this block (m, n) can include obtaining the SAD value and MVDIFF value of this block, as they are calculated during the calculation of the MV, as described with respect to Figure 4 . In some instances, the obtained MV field can include foreground MV and background MV (including global foreground MV and global background MV, as well as regional foreground MV and regional background MV), which can be detected as described with respect to Figure 5 .
[0131] At 604, method 600 includes performing post-filtering of the MV. In some instances, post-filtering of the MV can include calculating the average value of a specified window, smoothing the MV, and reducing the MV. For example, post-filtering of the MV can be performed on unreliable regions to generate a filtered output MV. In some instances, post-filtering of the MV can include using regional foreground MV and global foreground MV, as well as regional background MV and foreground background MV, to replace unreliable MVs.
[0132] At 606, method 600 includes generating block-level virtual depth. Virtual depth may be required to generate the output image through reprojection and / or intra-frame interpolation. For a given block, such as the first block (m, n), the depth can be calculated according to Equation (15):
[0133] iBD (m,n) = max(0, depth i + k 3 *(mvdist2fg - mvdist2bg)+ k 4 *iBMV m,n .mvdiff (15)
[0134] where iBD (m,n) is the virtual depth of the first block (m, n), depth i is the initial depth value, mvdist2fg is the MV difference between the current MV and the corresponding regional foreground MV, mvdist2bg is the MV difference between the current MV and the regional background MV, and iBMV (m,n) .mvdiff is the double-confirmation MVDIFF of the MV of the first block (m, n). The regional foreground MV or background MV used to determine the MV difference can be the bilinear interpolation MV (foreground or background) of the first block (m, n) with four neighboring regional MVs (foreground or background).
[0135] At 608, method 600 includes decomposing a block-level MV field into smaller block-level MV fields. As previously described, a luminance image may include pixels and be divided into a plurality of equally sized blocks. As an example, a luminance image may be segmented into 8x8 pixel blocks. Regarding Figures 4 to 5 the MVD calculations and processing described can be performed at the block level. Optionally, there is an object mask that flags whether each pixel belongs to an object, and only pixels belonging to the object to be processed.
[0136] Decomposing the block-level MVD field into a pixel-level MVD field may include decomposing the blocks into smaller blocks to improve the accuracy of pixel MVDs. As indicated at 610, for example, a guiding or reference image may be the luminance of the smaller blocks, as indicated at 612. Determine the block-level guiding image for each smaller block, and optionally, if there is an object mask, count the number of object pixels within each smaller block, as indicated at 614. As an example, an 8x8 pixel block may be decomposed into four 4x4 pixel blocks. The guiding image can be used to improve the accuracy of pixel-level MVs. The guiding image can be depth information, luminance information, or some other pixel-level information that identifies different objects in the scene of the image. This guidance can be used to weight the interpolation of MVs between blocks, or can be used in regression methods such as a guiding filter, as will be further described. In some instances, the weight of a block MV that is the same object as the pixel MV is higher than that of a block MV that is not the same object as the pixel MV.
[0137] At 616, method 600 includes filtering a subset of the smaller blocks to generate a pixel-level MVD field. For each pixel (i, j), a plurality of MVDs of the window of smaller blocks around the pixel can be obtained, such as the 5×5 blocks around the pixel (denote these blocks with nebblks). In some instances, not all blocks within the window (e.g., the 5x5 block window) may be filtered. If there is an object mask, only the blocks that are the same object as the current pixel can be filtered. Additionally, only blocks whose luminance level (e.g., pixel intensity) is within the range of the luminance level of the pixel in question can be filtered.
[0138] In some instances, the filtering may include weight calculations for luminance adjustment and weighting, spatial weighting, and object mask weighting (if present) for a given block. Luminance adjustment and weighting may include calculating the luminance difference between the current pixel and each surrounding block of the window, and then averaging the luminance differences. Then the average luminance difference can be adjusted according to Equation (16):
[0139] brtdif favg =min(thr0,max(thr1,1+brtdiff avg ) (16)
[0140] where brtdiff avgis the average luminance difference. Then, the weight of luminance can be calculated according to Equation (17):
[0141] w_brt (m,n) =max(0, brtdiff avg - abs(brt_nebblks (m,n) - brt (i,j) )) (17)
[0142] where w_brt (m,n) is the luminance weight of the given block (m, n) of nebblks, brt_nebblks (m,n) is the luminance of the block (m, n) of nebblks, and brt (i,j) is the luminance of the pixel (i, j).
[0143] Based on the luminance weight, spatial weight, and object mask weight, the weight of a given block can be determined, where the weight of the given block is its product, as described in Equation (18):
[0144] w (m,n) =w_brt (m,n) * w_sprat (m,n) * w_obj (m,n) (18) where w (m,n) is the weight of the given block (m, n), w_spat (m,n) is the spatial weight of the given block, and w_obj (m,n) is the weight of the object mask. If there is no object mask, w_obj (m,n) may be equal to a constant value c.
[0145] Then, the weight of the given block allows for the generation of an output image. In some instances, for each component of the MVD, the output can be generated according to Equation (19):
[0146]
[0147] where x o (i, j) represents the component of the decomposed pixel-level MVD of the pixel (i, j) (which can be horizontal motion, vertical motion, or depth), and x i(m,n) and is the component of the input block-level MVD of nebblks (m, n).
[0148] Block-level MVD decomposition, including filtering as described in this disclosure, may allow for improving the quality of the MVD by filtering based on weights such as luminance differences, spatial, and / or object masks. The methods for block-level MVD decomposition and filtering in this disclosure should be understood as merely examples, and other methods may also allow for generating pixel-level MVD. For example, a guided filter finds the best fit between input data (e.g., block-level MVD) and guidance data to generate output data (e.g., pixel-level MVD). The guided filter may generate a pixel-level MVD for each pixel based on values in a given window around the corresponding pixel. The guided filter may be applied to a guidance image, which may be luminance, panchromatic image, depth, or object mask.
[0149] In addition, in some instances, in addition to the above decomposition options, the results of object segmentation may also be used as guidance. For example, a neural network may be trained to identify whether pixels in an image belong to a certain type of object. Each type of object may have a different MVD associated with it, and thus using object identifiers and segmentation may allow for MV calculation and processing.
[0150] The output of MV calculation and processing as described in this disclosure for methods 400, 500, and 600 may generate filtered MVs with pixel-level virtual depth information (MVD). These MV fields may be used in various image processing methods, including interpolation, extrapolation, and / or reprojection.
[0151] Now turning to Figure 7A , a flowchart of method 700 for reducing holes in an image generated by reprojection (e.g., a triangle projection method) is illustrated. Method 700 may be implemented using the systems and components described above with respect to Figures 1 to 2 . For example, method 700 may be implemented by one or more processors according to instructions stored in a memory. For example, the instructions may be stored and executed by an image processing system such as image processing system 200. The image processing system described in this disclosure may be part of a computer system (such as computer system 100) including a game engine (e.g., GPU, CPU, etc.) or otherwise communicatively coupled to the computer system. Method 700 may be incorporated into Figure 3 's method 300, specifically at 306 and / or 310.
[0152] At 702, method 700 includes determining the presence of holes in the reprojection (or extrapolation) image. As previously described, in some instances, reprojection and / or extrapolation may result in one or more holes in the output image when previously covered content is not covered between frames. For each pixel in the input image, reprojection (or extrapolation) includes determining a projection position (e.g., position (u, v)) and depth (e.g., d uv). MV, depth, and phase can be calculated or otherwise determined as described above with respect to Figure 4 , Figure 5 and Figure 6 . The projection position can be determined according to Equation (20), and the depth can be determined according to Equation (21):
[0153] u = i + phase * mv ij .y (20)
[0154] v = j + phase * mv ij .x (20)
[0155] d uv = D(i, j) + phase * Z(i, j) (21)
[0156] where mv ij .y and mv ij .x are the vertical and horizontal motions of the MV at the position (i, j) of a given pixel, respectively, D(i, j) is the input depth of the input (i, j) of the pixel at time t, and Z(i, j) is the change in the depth of the pixel (i, j) from time t to t + 1.
[0157] If the projected image at the projection position (u, v) is invalid or the depth at the projection position (u, v) is greater than the depth d uv , then the depth at the projection position can be replaced with the depth d uv , and the projected image at the position t + phase can be replaced with the input image of the given pixel (i, j) at time t. When VALID(u, v) is equal to zero, the projected image at the position (u, v) may be invalid, which may be the initial value of the given pixel. Before reprojection, the initial value of VALID for all positions is 0. When the position (u, v) is projected from the input pixel, then VALID(u, v) is set to 1. Once the depth and the image are replaced, the validity can be confirmed (e.g., VALID(u, v) = 1). For example, the depth can be used to determine which of the one or more pixels being projected are retained, where some of the pixels may be background pixels and other pixels may be foreground pixels.
[0158] When a background pixel that was covered by an adjacent foreground at a first time (e.g., time t) is not covered at a second time (e.g., time t + p), holes will appear in the projected image. Holes can be detected in the MV field from the first time to the second time. When there is relative motion between the foreground and the background, the holes may not be covered. The MV field can be filtered to reduce the size of the holes, as will be described below. Method 700 in the present disclosure describes filtering, where the filtering may include finding the foreground MV of each pixel by comparing the depth with the depth of neighboring pixels and filtering the MV field around the uncovered region to generate a filtered output MV field. The range of neighboring pixels can be determined by the MV amplitude and phase of the projected image. Points along the motion trajectory can be filtered.
[0159] At 704, method 700 includes determining a local foreground MV. Determining the local foreground MV includes initializing the foreground MV field and the foreground depth field, as indicated at 706. The initialization may include setting values to the values of the input MV and the input depth. Determining the local foreground MV may further include determining the foreground MV field and the foreground depth field of each pixel, as indicated at 708.
[0160] As an example, for a given pixel (i, j), the foreground MV field and the depth field can be determined by determining the amplitude of the MV at position (i, j) according to equations (22) and (23), as indicated at 710:
[0161]
[0162] where is the amplitude of the MV at position (i, j), and and are the horizontal and vertical components of the normalized motion vector, respectively.
[0163] For a given pixel (i, j), its processing range can also be determined according to equation (24):
[0164]
[0165] where r ij is the range and k 0 is equal to 2.
[0166] Along the motion trajectory with a range of [-r ij , r ij , each point of the trajectory can be checked to determine the projection position (u, v), as indicated at 712. For example, for each step t, where t is an integer within the range of [-r ij , r ij , the projection position can be determined according to equation (25)
[0167]
[0168] The variables are as described above.
[0169] For each step t, compare the foreground depth at position with the input depth at position (i, j). When the input depth at position (i, j) is less than the foreground depth at position , then the foreground depth at position is replaced with the input depth at position (i, j), and the foreground MV at position is replaced with the MV at position (i, j). Additionally, for each step t, the input depth at position can be compared with the foreground depth at position (i, j). When the input depth at position is less than the foreground depth at (i, j), then the foreground depth at (i, j) is replaced with the input depth at position , and the foreground MV at (i, j) is replaced with the MV at 1 . Also, the determined foreground MV field and foreground depth field can be smoothed, for example via a guided filter, bilateral filter, etc., with the size of a specified window at k
[0170] It should also be understood that the given pixel described in this disclosure represents each pixel, and the determination of local foreground MV described in this disclosure can be performed for one or more pixels of the input image.
[0171] At 716, method 700 includes defining a hole MV for a given pixel to define the actual size of the hole. In some instances, the hole MV (e.g., hmv ij ) can be the product of the phase and the difference between the MV of the given pixel (i, j) and the foreground MV of the given pixel (i, j). In other instances, the hole MV can be the product of the phase and the MV of the given pixel (i, j), which can reduce the computational requirements as it reduces the calculation of the foreground MV.
[0172] At 718, method 700 includes determining a filtering range and step size and determining the amplitude of the hole MV. Similar to the above, the filtering range, as well as the step size and amplitude of the hole MV, can be determined according to equations (26) and (27):
[0173] amp ij = max(abs(hmv ij .x), abs(hmv ij .y)) (26)
[0174]
[0175] The variables are as described above.
[0176] The filtering range, step size, and amplitude can be used to define the ranges of the distortion radius and step size based on equations (28) and (29):
[0177] r 弯曲 = min(thr 2 , k 2 * amp ij ) (28)
[0178]
[0179] where r 弯曲 is the range of the distortion radius, s ij is the step size, thr 2 is the threshold for limiting the pixel range, taps is the number of sampling points for controlling along the motion trajectory, and k 2 is a parameter for adjusting the pixel range for performing MV filtering. In some instances, the default value of k 2 can be 0.25. A large value of k 2 may indicate more pixels for which the MV will be filtered. A larger taps value can achieve higher precision. In some instances, the number of points that can be filtered can be 2M + 1, where M is the quotient of r 弯曲 and s ij .
[0180] Continuing to refer to Figure 7B , at 720, method 700 includes filtering the MV along the motion trajectory based on the determined filtering range, step size, and amplitude. As indicated at 722, filtering the MV along the motion trajectory can include determining the MV at the projected position (e.g., the position (u t , v t )) for each step size t. In some instances, t can be between -M and M, where M is defined as above. The projected position can be defined according to equation (30):
[0181] u t = round(i + t * s ij * adj y ) (30)
[0182] v t = round(j + t * s ij * adj x ) (30)
[0183] where the variables are as described above.
[0184] Then, filtering the MV along the motion trajectory includes comparing the MV of a given pixel (i, j) with the MV of the projected position (u t , v t ) as indicated at 724, in order to determine whether there is an uncovered relationship between the two. The determination of the uncovered relationship can be based on the conditional equation (31):
[0185]
[0186] where mv ij .x and mv ij .y are the horizontal and vertical motion components of the MV of the given pixel, and mv t .x and mv t .y are the horizontal and vertical motion components of the MV of the projected position. The value of the uncovered relationship, uncovered t can be the sum of the horizontal component being uncovered x and the vertical component being uncovered y . When the value of the uncovered relationship, uncovered t is greater than zero, the MV of the projected pixel can participate in the filtering.
[0187] Filtering the MV along the motion trajectory can further include determining the foreground and background relationship between the MV of the given pixel and the MV of the projected position, as indicated at 726. In some instances, the background may be affected by the foreground, but the foreground may not be affected by the background. The MV participating in the filtering can be determined by the equation (32):
[0188]
[0189] where mv′ t .x and mv′ t .y are the horizontal and vertical motion components of the MV participating in the filtering, mv ij .depth is the depth of the MV of the projected pixel, and mv ij .depth is the depth of the MV of the given pixel.
[0190] The MV can be accumulated within a specified range to determine the filtered MV, as indicated at 728. The specified range can be [-M, 0] and [0, M]. The accumulation of the MV can be performed according to the equation (33):
[0191]
[0192] where x0 flt and y0 flt are the filtered horizontal and vertical motions, the MV is within the range of [-M, 0], and x1 flt and y1 fltYes, the filtered horizontal and vertical motions, where MV is in the range of [0, M].
[0193] For pixels surrounded by both background and foreground pixels, foreground MV dilation can be performed, as indicated at 730. Foreground MV dilation can reduce the hole region between the foreground and the background. In some instances, the inpainting algorithm can also use depth information to facilitate filling the holes with background information rather than foreground information.
[0194] The filtered MV (x0 flt , y0 flt ) and (x1 flt , y1 flt ) can be merged. According to Equation (34), the MV with the largest difference compared to the MV of a given pixel can be used for the output MV (fmv ij ), as at 732:
[0195]
[0196] where dist0 = |mv ij .x - x0 flt | + |mv ij .y - y0 flt | and
[0197] dist1 = |mv ij .x - x1 flt | + mv ij .y - y1 flt , |
[0198]
[0199] and
[0200] The filtered MV can be determined for each pixel of the image in this way, so as to determine the filtered MV field. Along with the output filtered MV field, large hole regions can be divided into smaller holes.
[0201] At 734, method 700 includes performing an inpainting algorithm to fill the smaller holes. Applying the inpainting algorithm to the smaller holes (instead of the initially generated large holes) can mitigate the unstable results of inpainting. In addition, the inpainting algorithm can be applied to holes surrounded by background pixels.
[0202] In some instances, method 700 can be repeated one or more times to continue reducing the size of the holes, and in some cases to substantially eliminate the holes. In either case, whether performing inpainting or filtering the holes, an output that fills the holes without instability can be generated.
[0203] Now turn to Figure 10 , which shows a third figure 1000 that depicts an MV field 1002 and a filtered MV field 1004 from time t to t+1. When the MV field 1002 is filtered to reduce the hole size, the filtered MV field 1004 can be the output filtered MV, as described with respect to method 700.
[0204] The MV field 1002 can include a foreground MV 1006 and a background MV 1008. In some instances, as depicted in the third figure 1000, the foreground MV 1006 and the background MV 1008 can move in opposite directions. In other instances, the foreground MV and the background MV can move in the same direction. The MV field 1002 can include a first hole 1010 generated by the foreground MV 1006 and the background MV 1008 moving apart from each other from time t to t+1. The first hole 1010 can have a large first size 1012. In some instances, a hole can be considered "large" when the size of the hole is higher than a predefined threshold.
[0205] Similar to the MV field 1002, the filtered MV field 1004 can include a foreground MV 1014 and a background MV 1016. The filtered MV field 1004 can include a plurality of second holes 1018. The size of each of the plurality of second holes 1018 can be smaller than the first size 1012 of the first hole 1010. As described with respect to Figure 7, most of the plurality of second holes 1018 can come from background pixels, yet some of the second holes 1018 can be located at the edge between the foreground MV and the background MV. For example, a third hole 1020 among the plurality of second holes 1018. The third hole 1020 can undergo foreground MV expansion, as described with respect to Figure 7, to reduce the hole size.
[0206] Now turn to Figure 11 , which shows a use case scenario of an image processing system. In some instances, the image processing system can be the image processing system 200 described with respect to Figure 2 , and thus uses similar component numbers. In some instances, various inputs and outputs are demonstrated in the use case scenario that will occur when the above methods 300, 400, 500, 600, and 700 are executed.
[0207] In the use case scenario shown in the figure, the current image frame CF (denoted as I CF ) and the previous image frame PF (denoted as PF ) and their associated masks are input into the luminance generator 202. The masks can be used to mark objects and special effects in the corresponding frames. The luminance generator 202 outputs the adjusted CF luminance (labeled as Y CF ) and the adjusted PF luminance (labeled as (Y PF)。The adjusted brightnesses of the CF and PF are input into the motion vector calculator 204. Based on one or more methods (such as methods 400 and 500), the motion vector calculator 204 can output the internal block-level phase 0Mv field of the PF, the region foreground and region background MVs, and the internal block-level phase 1MV field of the CF. The internal block-level phase 0MV field of the PF, the region foreground and region background MV fields, and the original PF and its mask can be input into the previous frame MV processor 210. The internal block-level phase 1MV field of the CF, the region foreground and region background MVs, and the original CF and its mask can be input into the current frame MV processor 208.
[0208] The previous frame MV processor 210 and the current frame MV processor 208 can process the input MV fields and region foreground MV field / background MV field according to one or more methods (such as the above method 600). The processing can include decomposing the block level into the pixel level and generating virtual depth information. Thus, the previous frame MV processor 210 can output the internal MV field (denoted as iMV PF0 ) of each pixel of the PF and the internal virtual depth field (denoted as iD PF0 ) of the PF. Similarly, the current frame MV processor 208 can output the internal MV field (denoted as iMV CF1 ) of each pixel of the CF and the internal virtual depth field (denoted as iD CF1 ) of the CF. The internal MV field and virtual depth field of the PF can be input into the previous frame MVD combiner 216 together with the external MV field between the PF and the CF for each pixel of the PF generated by the game engine, the external depth field generated by the game engine for the PF, the change in the depth field from the PF to the CF generated by the game engine, and the mask of the PF. The internal MV field and virtual depth field of the CF can be input into the current frame MVD combiner 214 together with the external MV field between the CF and the PF for each pixel of the CF generated by the game engine, the external depth field generated by the game engine for the CF, the change in the depth field from the CF to the PF generated by the game engine, and the mask of the CF.
[0209] The previous frame MVD combiner 216 can output the depth field of the PF (denoted as D PF ), the change in the depth field from the PF to the CF (denoted as Z PF0 ), and the MV field between the PF and the CF for each pixel of the PF (denoted as MV PF0 ). The current frame MVD combiner 214 can output the depth field of the CF (denoted as D CF ), the change in the depth field from the CF to the PF (denoted as Z CF1 ), and the MV field between the PF and the CF for each pixel of the CF (denoted as MV CF1)。The output of the previous frame MVD combiner 216, along with the phase of the PF (e.g., p projected from time t to time t + p), which is the temporal distance between the projected or target image and the input image, and the original PF, can be input into the previous frame reprojection module 218a of the reprojection module 218. The output of the current frame MVD combiner 214, along with the phase of the CF (e.g., p - 1 projected from time t + 1 to time t + p), and the original CF can be input into the current frame reprojection module 218b of the reprojection module 218.
[0210] In instances where extrapolation is obtained, only the current MVD combiner 214 is utilized. A large amplitude of the phase value indicates a large temporal difference between the projected frame and the input frame. When the phase is positive, it indicates that the input frame is projected into the future. When the phase is negative, it indicates that the input frame is projected into the past. When using extrapolation, only the output of the current frame MVD combiner 214 can be input into the reprojection module 218. As such, only one frame is used for reprojection, and thus the holes in the reprojected image may be more numerous and / or larger than when MV processing is used for interpolation.
[0211] The reprojected image output from the reprojection module 218 can be input into the image combiner 220 for merging. The merged image 220 can then be input into the inpainter 222 for inpainting to fill the holes. The inpainter 222 can output the final image.
[0212] In some instances, although not presented in this use case scenario, the MV field output by the MVD combiner can be filtered to reduce the size of the holes before inpainting, as described with respect to method 700.
[0213] The technical effects of the systems and methods provided by the present disclosure are that MV and virtual depth can be estimated and / or calculated based on the input image, rather than obtained from a game engine. Estimation / calculation based on the input image allows for a reduction in computational and / or processing power, thus allowing for faster rendering. Additionally, a luminance image can be created based on the RGB components and the type of content within the image, which can improve MV calculation without increasing the requirements on the system. Calculating and processing MV at the block level and then decomposing the block-level MV into pixel-level MV using a bilateral filter, weighted average, and / or guided filter can allow for smoother MV, which more closely adheres to the edges of objects in the luminance image and provides virtual depth.
[0214] Furthermore, when previously covered data is not covered due to motion during reprojection, the MV continuation rate along the motion trajectories of the relevant pixels / blocks can allow for a reduction in the hole size, which in turn can allow for the application of an inpainting algorithm and the output of more stable results. This can reduce the processing requirements by allowing the use of a simpler, less demanding inpainting algorithm. It can further reduce the latency of the entire system.
[0215] As used in this disclosure, an element or step recited in the singular and preceded by the word "a" or "an" should be understood to not exclude a plurality of the elements or steps, unless such exclusion is explicitly stated. In addition, a reference in this invention to "one embodiment" is not to be construed as excluding the existence of additional embodiments that also incorporate the recited features. Further, unless explicitly stated to the contrary, embodiments of an element or elements having a particular property may include other such elements not having that particular property. The terms "comprising" and "in which" are used as shorthand equivalents of the corresponding terms "including" and "wherein". In addition, the terms "first", "second", and "third", etc. are used merely as labels and are not intended to impose numerical requirements or a particular positional order on their objects.
[0216] This written description uses examples to disclose the invention, including the best mode, and also enables any person skilled in the art to practice the invention, including making and using any device or system and performing any incorporated method. The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. If these other examples have structural elements that are not different from the literal language of the claims, or if they include equivalent structural elements that are not materially different from the literal language of the claims, then they are intended to be within the scope of the claims.
Claims
1. A method comprising: determining the presence of one or more holes in an output image, wherein the output image is generated based on one or more motion vector MV fields and a depth field; Determining a local foreground MV field and a depth field for each pixel of the output image; defining a hole MV field based on the local foreground MV field; Filtering the hole MV field along the motion trajectory of the hole MV field to reduce the size of the one or more holes; as well as Outputs a filtered image with one or more smaller holes.
2. The method according to claim 1, further comprising: An inpainting algorithm is performed on the filtered image to fill the one or more smaller holes.
3. The method according to claim 1, wherein: Filtering the hole MV field along the motion trajectory of the hole MV field includes: The MV at the projected position is determined for each step along the motion trajectory.
4. The method according to claim 3, wherein: Filtering the hole MV field further includes: The MV of a given pixel is compared to the MV at the projected position of the pixel in the output image to determine which MV to filter.
5. The method according to claim 4, further comprising: The foreground and background relationships between the MV for the given pixel and the MV at the projected position are determined.
6. The method according to claim 1, wherein: Filtering the hole MV field along the motion trajectory of the hole MV field includes: MVs within a specified range are accumulated to determine a filtered MV field. The method of claim 6 , wherein the filtered MV field includes the one or more smaller holes.
8. The method according to claim 6, further comprising: A foreground MV dilation is performed on the filtered MV field.
9. The method according to claim 1, further comprising: Determine the range and direction of filtering, which includes calculating the foreground MV field and the foreground depth field.
10. The method of claim 1, wherein the hole MV field is one of a pixel-level MV field and a block-level MV field.