A frame processing device, method, and frame processor
Through motion estimation and deformation technology combined with super resolution and anti-aliasing processing, the problem of image aliasing artifacts on mobile devices is solved, achieving efficient image quality improvement.
Patent Information
- Application Number
- CN202111464573.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-07
- Filing Date
- 2021-12-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-03
AI Technical Summary
The prior art is difficult to efficiently process images with aliased artifacts, especially on mobile devices, resulting in a degradation in image quality due to bandwidth and resolution limitations.
The motion data between the current frame and the previous frame is estimated by the motion estimation circuit and deforms the previous frame based on the motion data to align it with the current frame. The deformed frames are then processed using the super-resolution and anti-aliasing engine, generating high-resolution frames and eliminating aliasing artifacts.
Effectively eliminates aliasing artifacts of images, improves the resolution and quality of images, retains information in the original frame, and is suitable for image processing of mobile devices.
Smart Images

Figure CN114596339B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology, and more particularly, to an artificial intelligence (AI) engine for processing images with aliasing artifacts. Background Art
[0002] The description of the background art provided herein is to present the content of the present invention generally. The work of the inventors of the present invention, including the content described in this background art section and those aspects in the specification that do not conform to the prior art at the filing date, should not be explicitly or implicitly regarded as the prior art of the present invention.
[0003] An image or a frame can be displayed on a mobile phone. The frame can include a video frame from a cloud source via the Internet and a game video generated by a processor of the mobile phone (e.g., a Graphics Processing Unit (GPU)). Limited by the Internet bandwidth and the size and resolution of the mobile phone, the video frame and the game frame may have a low resolution and aliasing characteristics. Summary of the Invention
[0004] The present invention provides a frame processing device, method, and frame processor, which can help eliminate the aliasing artifacts of a frame.
[0005] A frame processing device provided by the present invention includes: a motion estimation circuit configured to estimate motion data between a current frame and a previous frame; a warping circuit coupled to the motion estimation circuit and configured to warp the previous frame based on the motion data so that the warped previous frame is aligned with the current frame and determine whether the current frame is consistent with the warped previous frame; and a temporary decision circuit coupled to the warping circuit and configured to generate an output frame, and when the current frame is consistent with the warped previous frame, the output frame includes the current frame and the warped previous frame.
[0006] A frame processor provided by the present invention includes: a super-resolution (SR) and anti-aliasing (AA) engine for receiving a training frame, enhancing the resolution of the training frame, and eliminating the aliasing artifacts of the training frame to generate a first high-resolution frame with aliasing artifacts and a second high-resolution frame without aliasing artifacts; and an attention reference frame generator connected to the SR and AA engines and configured to generate an attention reference frame based on the first high-resolution frame and the second high-resolution frame.
[0007] The method for processing a frame provided by the present invention includes: estimating motion data between a current frame and a previous frame; deforming the previous frame based on the motion data such that the deformed previous frame is aligned with the current frame; and generating an output frame, when the current frame and the deformed previous frame are identical, the output frame includes the current frame and the deformed previous frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 Some embodiments in accordance with the present invention illustrate how an exemplary virtual high - resolution image 110 can be displayed on a low - resolution raster display 100.
[0009] Figure 2 Some embodiments in accordance with the present invention illustrate how an exemplary triangle 110 can be displayed on a low - resolution raster display 100 when MSAA is applied.
[0010] Figure 3 Some embodiments in accordance with the present invention illustrate a functional block diagram of an exemplary device 300 for processing an image or a frame having aliasing artifacts.
[0011] Figure 4 Some embodiments in accordance with the present invention illustrate a functional block diagram of an exemplary frame processor 400 for processing an image or a frame having aliasing artifacts.
[0012] Figure 5 Some embodiments in accordance with the present invention illustrate a functional block diagram of an exemplary frame processor 500 for processing an image or a frame having aliasing artifacts.
[0013] Figure 6 Some embodiments in accordance with the present invention illustrate a flowchart of an exemplary method 600 for processing an image or a frame having aliasing artifacts.
[0014] Figure 7 Some embodiments in accordance with the present invention illustrate a flowchart of an exemplary method 700 for processing an image or a frame having aliasing artifacts. DETAILED DESCRIPTION
[0015] In the specification and claims, certain terms are used to refer to specific components. Those skilled in the art should understand that hardware manufacturers may use different terms to refer to the same component. The specification and claims do not distinguish components by the difference in names, but by the difference in functions of the components. The terms "comprising" and "including" mentioned throughout the specification and claims are open-ended terms and should be interpreted as "including but not limited to". "Substantially" or "about" means within an acceptable error range. Those skilled in the art can solve the technical problems within a certain error range and basically achieve the technical effects. In addition, the term "coupled" or "coupling" herein includes any direct and indirect electrical connection means. Therefore, if it is described in the text that a first device is coupled to a second device, it means that the first device can be directly electrically connected to the second device, or indirectly electrically connected to the second device through other devices or connection means. The following describes the preferred embodiments of the present invention, aiming to illustrate the spirit of the present invention rather than to limit the protection scope of the present invention. The protection scope of the present invention shall be determined by what is defined in the claims.
[0016] The following description is the optimal embodiment expected by the present invention. These descriptions are used to elaborate on the general principles of the present invention and should not be used to limit the present invention. The protection scope of the present invention should be determined based on reference to the claims of the present invention.
[0017] Super-Resolution (SR) technology can reconstruct high-resolution images from low-resolution images, and the low-resolution images can be captured by an image capture device with insufficient number of included sensors. Anti-Aliasing (AA) technology can improve the quality of low-resolution images with aliasing artifacts. However, after SR and AA operations, some information of the image may be lost. For example, when an object (such as a boarding ladder) in the original image (e.g., the current frame in a continuous frame stream) moves horizontally, after removing the aliasing artifacts, some vertical parts of the boarding ladder (such as the handrails) may disappear and not be shown in the processed image. Instead of only processing the original image during SR and AA operations, when an additional image and the original image meet certain requirements, the present invention can further consider at least one additional image (e.g., the previous frame in a continuous frame stream). In one embodiment, the motion data between the additional image and the original image can be determined first, and then the additional image can be warped based on the motion data so that the warped additional image can be aligned with the original image, and when the warped additional image and the original image are consistent, the warped additional image can be further used in the operations of performing SR and AA on the original image. According to some other embodiments of the present invention, the frame with enhanced resolution can be compared with the frame with removed aliasing artifacts to generate an attention reference frame, which includes the key difference information between the two frames. In one embodiment, the attention reference frame can be used to train a Neural Network (NN), and then the trained NN can enhance the resolution of another frame and use the enhanced resolution of the other frame to remove the aliasing artifacts of the other frame, where only the key information (i.e., the key difference information) included in the attention reference frame is focused on to enhance the resolution of the other frame.
[0018] In most digital imaging applications, it is always desirable to have digital images with higher resolution for subsequent image processing and analysis. The higher the resolution of a digital image, the more details the digital image has. The resolution of a digital image can be classified as, for example, pixel resolution, spatial resolution, temporal resolution, and spectral resolution. The spatial resolution may be limited by the image capture device and the image display device. For example, Charge-Coupled Device (CCD) and Complementary Metal-Oxide-Semiconductor (CMOS) are the most widely used image sensors in image capture devices. The size of the sensor and the number of sensors per unit area can determine the spatial resolution of the image captured by the image capture device. An image capture device with a high sensor density can generate high-resolution images, but consumes a lot of power and has a high hardware cost.
[0019] An image capture device with insufficient number of sensors may generate low-resolution images. The low-resolution images thus generated will have aliasing artifacts or jagged edges (referred to as aliasing), which occur whenever non-rectangular shapes are created using pixels located at exact rows and columns. Aliasing occurs when a high-resolution image is represented at a lower resolution. Aliasing may distract the users of a computer (PC) or a mobile device.
[0020] Figure 1 Some embodiments in accordance with the present invention illustrate how an exemplary virtual high-resolution image 110 is displayed on a low-resolution raster display 100. The display 100 may have a plurality of pixels 120 located in rows and columns. The cross “+” represents the sampling point 130 of the pixel 120, and the sampling point 130 is used to determine whether a fragment will be generated for the pixel. For example, when the sampling point 130A is not covered by the image 110 (such as a triangular pixel), even if a part of the pixel 120A is covered by the triangle 110, no fragment will be generated for the pixel 120A with the sampling point 130A; when the sampling point 130B is covered by the triangle 110, even if a part of the pixel 120B is not covered by the triangle 110, a fragment will be generated for the pixel 120B with the sampling point 130B. Therefore, the triangle 110 rendered by the display 100 is shown as having jagged edges.
[0021] Anti-aliasing is a technique that addresses the aliasing problem by oversampling an image at a rate higher than the expected final output rate to smooth the jagged edges of the image. For example, Multisample Anti-Aliasing (MSAA) is one of the Supersampling Anti-Aliasing (SSAA) algorithms proposed to address the aliasing that appears at the edges of triangle 110. It can simulate each pixel of the display as having multiple sub-pixels and determine the color of the pixel based on the number of sub-pixels covered by the target image. Figure 2 Some embodiments according to the present invention illustrate how an exemplary triangle 110 can be displayed on a low-resolution raster display 100 when MSAA is applied. MSAA can simulate each pixel 120 as having 2x2 sub-pixels 220, each sub-pixel having a sub-sampling point 230, and determine the color of pixel 120 based on the number of sub-sampling points 230 covered by triangle 110. For example, when no sub-sampling point 230A is covered by triangle 110, no fragment is generated for pixel 120A with sampling point 230A, and pixel 120A is blank; when triangle 110 covers only one sub-sampling point 230B, pixel 120B with sampling point 230B will have a light color, e.g., one quarter of the color of triangle 110, which can be estimated by the fragment shader; when triangle 110 covers only two sub-sampling points 230C, pixel 120C with sampling point 230C will have a darker color than pixel 120B, e.g., half of the color of triangle 110; when triangle 110 covers up to three sub-sampling points 230D, pixel 120D with sampling point 230D will have a darker color than pixel 120C, e.g., three quarters of the color of triangle 110; when all sub-sampling points 230E are covered by triangle 110, pixel 120E with sampling point 230E will have the same darkest color as Figure 1 the pixel 120B shown. As a result, the MSAA-applied triangle 110 rendered on display 100 has smoother edges compared to the triangle 110 rendered on a Figure 1 display 100 without MSAA applied.
[0022] As Figure 2As shown, each pixel 120 determines its color using a regular grid of 2x2 sub-sampling points 230. In one embodiment, each pixel 120 may also use a regular grid of 1x2 or 2x1 or 4x4 or 8x8 sub-sampling points 230 to determine its color. In another embodiment, each pixel 120 may also use 2x2 sub-sampling points 230, i.e., Rotated Grid SuperSampling (RGSS), and five sub-sampling points 230 (e.g., a plum blossom anti-aliasing), where four of the five sub-sampling points are shared with four other pixels respectively. As the number of sub-pixels increases, the calculation becomes expensive and requires a large amount of memory. MSAA can be executed by an artificial intelligence (AI) processor such as a convolution accelerator and a Graphics Processing Unit (GPU), which is designed to accelerate the creation of an image destined for a display in an image buffer, so as to offload the graphics processing operations from a central processing unit (CPU). A desktop GPU can use real-time mode rendering. A real-time mode GPU requires off-chip main memory (e.g., DRAM) to store a large amount of multi-sampled pixel data, and must access the DRAM to obtain the pixel coordinates of the current fragment from the multi-sampled pixel data for shading each fragment, which consumes a large amount of bandwidth. The present invention proposes a tile-based GPU for a mobile phone to minimize the amount of external memory access required by the GPU during fragment shading. The tile-based GPU moves the image buffer out of off-chip memory and into high-speed on-chip memory (i.e., a tile buffer that requires less power to access). The size of the tile buffer may vary in different GPUs, but the tile buffer can be as small as 16x16 pixels at minimum. To use such a small tile buffer, the tile-based GPU splits the render target into small tiles and renders one tile at a time. After the rendering is completed, the tiles are copied into external memory. Before splitting the render target, the tile-based GPU must store a large amount of geometric data (i.e., data for each vertex change and the intermediate state of the tile) into the main memory, which will compromise a part of the bandwidth savings for the image buffer data.
[0023] Figure 3A functional block diagram of an exemplary device 300 for processing an image or frame with aliasing artifacts is shown in accordance with some embodiments of the present invention. The device 300 may enhance the resolution of the current frame (e.g., by super-resolution techniques), and may eliminate the aliasing artifacts of the current frame with enhanced resolution by processing only the current frame, or the current frame and a warped previous frame aligned with the current frame, so as to preserve as much information as possible contained in the current frame. For example, the device 300 may include a motion estimation circuit 310, a warping circuit 320, and a temporal decision circuit 330.
[0024] The motion estimation circuit 310 may receive a plurality of consecutive images or frames including at least the current frame and a previous frame. For example, the current frame and the previous frame may be a video frame stream, which may be a low-resolution signal obtained from a cloud source through the Internet and have aliasing characteristics. As another example, the current frame and the previous frame may be game frames generated by a processor (e.g., GPU) of a mobile phone. Limited by the size and resolution of the mobile phone, the game frames may also be low-resolution and thus have aliasing characteristics. The motion estimation circuit 310 may estimate the motion data between the current frame and the previous frame. For example, the motion information may include the direction in which the previous frame moves to the current frame and how far (e.g., a plurality of pixels) it needs to move from the previous frame to the current frame. In one embodiment, the motion estimation circuit 310 may be a neural network that can be trained to estimate the motion data between the current frame and the previous frame. In another embodiment, the motion estimation circuit 310 may use methods such as the Sum of Absolute Difference (SAD) method, the Mean Absolute Difference (MAD) method, the Sum of Squared Difference (SSD) method, the Zero-Mean SAD method, the Locally Scaled SAD method, or the Normalized Cross Correlation (NCC) method to estimate the motion data. For example, in a SAD operation, a patch of the previous frame may be extracted and shifted to the right by a value, and the first sum of the absolute differences between the pixels of the shifted patch of the previous frame and the corresponding patch of the current frame may be calculated. The shifted patch of the previous frame may be further shifted to the right by the value, and the second sum of the absolute differences between the pixels of the further shifted patch of the previous frame and the corresponding patch of the current frame may also be calculated. When the first sum is less than the second sum, the motion data may be equal to the value, or when the second sum is less than the first sum, the motion data may be equal to twice the value.
[0025] The warping circuit 320 may be coupled to the motion estimation circuit 310 and warp a previous frame based on motion data such that the warped previous frame is aligned with the current frame. For example, the warping circuit 320 may geometrically align the structure (texture) / shape of the previous frame with the current frame based on the motion data. In one embodiment, when the first sum is less than the second sum, the warping circuit 320 may warp the previous frame to the right based on this value. In another embodiment, when the second sum is less than the first sum, the warping circuit 320 may warp the previous frame to the right by twice this value. For example, the warping circuit 320 may linearly insert pixels along the rows of the previous frame and then along the columns to assign the bilinear function value of the four pixels closest to S in the previous frame to the reference pixel positions in the current frame, and use 16 nearest neighbors and a bicubic waveform in bicubic interpolation to reduce resampling artifacts. In one embodiment, when the shifted previous frame matches the current frame, the warping circuit 320 may warp the previous frame. For example, when the first sum is less than the second sum and less than a sum threshold, the warping circuit 320 may warp the previous frame to the right based on this value. As another example, when the second sum is less than the first sum and less than the sum threshold, the warping circuit 320 may warp the previous frame to the right by twice this value. In another embodiment, when the motion data is less than a motion threshold, the warping circuit 320 may warp the previous frame. For example, the motion threshold may be three times this value, and regardless of whether the third sum of the absolute differences of the pixels between the block of the previous frame shifted three times this value to the right and the corresponding block of the current frame is less than the first sum, the second sum, and the motion threshold, the warping circuit 320 does not warp the previous frame to the right based on three times this value. In another embodiment, the warping circuit 320 may further determine whether the current frame and the warped previous frame are consistent. For example, the warping circuit 320 may determine the consistency information between the current frame and the warped previous frame based on the cross-correlation between the current frame and the warped previous frame. For example, when the cross-correlation exceeds a threshold, the warping circuit 320 may determine that the warped previous frame and the current frame are consistent.
[0026] The provisional decision circuit 330 may be coupled to the warping circuit 320 and configured to generate an output frame. For example, when the current frame and the warped previous frame are consistent, the output frame may include the current frame and the warped previous frame. As another example, when the current frame and the warped previous frame are inconsistent, the output frame may include only the current frame. In some embodiments, the provisional decision circuit 330 may further be coupled to the motion estimation circuit 310, and when the motion data is equal to or exceeds the motion threshold, the output frame may include only the current frame.
[0027] As Figure 3As shown, the device 300 may further include a frame fusion circuit 340. The frame fusion circuit 340 may be coupled to the temporary decision circuit 330 and fuse an output frame including the current frame and the warped previous frame. For example, the frame fusion circuit 340 may concatenate the warped previous frame to the current frame in a channel-wise manner. As another example, the frame fusion circuit 340 may add the warped previous frame to the current frame to generate a single frame. As Figure 3 shown, the device 300 may further include a frame processor 350. The frame processor 350 may be coupled to the frame fusion circuit 340 and process the frame output from the frame fusion circuit 340, which may be the current frame, the current frame concatenated to the warped previous frame, or the single frame. For example, the frame processor 350 may adjust the size of the current frame or enhance the resolution of the current frame, and use the enhanced resolution of the current frame to eliminate aliasing artifacts in the current frame. In one embodiment, the frame fusion circuit 340 may be omitted, and the frame processor 350 may be directly coupled to the temporary decision circuit 330 and process the current frame, or the current frame and the warped previous frame. Since the warped previous frame may also be generated by the warping circuit 320 and output to the frame processor 350 when the warped previous frame is consistent with the current frame, the frame processor 350 may enhance the resolution of the current frame and eliminate aliasing artifacts in the current frame with enhanced resolution by further considering the warped previous frame. In this case, compared with processing the current frame by only considering the current frame, less information will be lost from the processed current frame.
[0028] Figure 4A functional block diagram of an exemplary frame processor 400 for processing an image or a frame with aliasing artifacts is shown in accordance with some embodiments of the present invention. The frame processor 400 may be coupled to the provisional decision circuit 330 or the frame fusion circuit 340 of the device 300. The frame processor 400 may include an attention reference frame generator 430 and an artificial intelligence (AI) neural network (NN) 440 coupled to the attention reference frame generator 430. The attention reference frame generator 430 may generate an attention reference frame based on a first high-resolution frame with aliasing artifacts and a second high-resolution frame with the aliasing artifacts removed. For example, the attention reference frame generator 430 may compare the first frame (i.e., the first high-resolution frame) and the second frame (i.e., the second high-resolution frame) to capture the key information by which the first frame differs from the second frame. The AI NN 440 may remove the aliasing artifacts of another frame (e.g., a low-resolution current frame) based on the attention reference frame. For example, the AI NN 440 may be trained using the attention reference frame, and then the AI NN 440 enhances the resolution of the low-resolution frame and uses the enhanced resolution of the low-resolution frame to remove the aliasing artifacts of the low-resolution frame, wherein only a portion of the low-resolution frame corresponding to the key information included in the attention reference frame is focused on to enhance the resolution of the low-resolution frame.
[0029] Figure 5A functional block diagram of an exemplary frame processor 500 for processing an image or frame having aliasing artifacts is shown in accordance with some embodiments of the present invention. The frame processor 500 may be coupled to the temporal decision circuit 330 or the frame fusion circuit 340 of the device 300. The frame processor 500 may include a resizing or super-resolution (SR) and anti-aliasing (AA) engine 501, a region of interest reference frame generator 430 coupled to the SR and AA engine 501, and an AI NN 440. The SR and AA engine 501 may generate a first high-resolution frame having aliasing artifacts and a second high-resolution frame with the aliasing artifacts removed. In one embodiment, the SR and AA engine 501 may include an SR engine 510 that may enhance the resolution of a frame that may have aliasing artifacts to generate an enhanced-resolution frame, e.g., the first high-resolution frame having aliasing artifacts. For example, the SR engine 510 may be an artificial intelligence (AI) SR engine. In another embodiment, the SR and AA engine 501 may further include an AA engine 520 coupled to the SR engine 510 that may remove the aliasing artifacts of the enhanced-resolution frame to generate an anti-aliased frame, e.g., the second high-resolution frame with the aliasing artifacts removed. For example, the AA engine 520 may be an AI AA engine. In another embodiment, the AA engine 520 may be arranged in front of the SR engine 510. In this case, a frame (a frame having aliasing artifacts) will have its aliasing artifacts removed by the AA engine 520 first, and then two frames (the frame having aliasing artifacts and the frame with the aliasing artifacts removed by the AA engine 520) will have their resolutions enhanced by the SR engine 510, resulting in the first high-resolution frame having aliasing artifacts and the second high-resolution frame with the aliasing artifacts removed. The region of interest reference frame generator 430 may generate a region of interest reference frame based on the enhanced-resolution frame (i.e., the first high-resolution frame having aliasing artifacts) and the anti-aliased frame (i.e., the second high-resolution frame with the aliasing artifacts removed). For example, the region of interest reference frame generator 430 may compare the enhanced-resolution frame and the anti-aliased frame to capture key information that is different in the enhanced-resolution frame from the anti-aliased frame.
[0030] Figure 6 A flowchart of an exemplary method 600 for processing an image or frame having aliasing artifacts is shown in accordance with some embodiments of the present invention. When the frame processors 400 and 500 are removing the aliasing artifacts of a low-resolution frame, the method 600 may provide additional warped previous frames to the frame processors, e.g., the frame processors 400 and 500. In various embodiments, some of the steps shown in the method 600 may be executed simultaneously, executed in a different order than the Figure 6 order shown, replaced by other method steps, or may be omitted. Additional method steps may also be executed as needed. Aspects of the method 600 may be implemented by the device 300 shown or described above based on the figures.
[0031] In step 610, motion data between the current frame and the previous frame can be estimated.
[0032] In step 620, the previous frame can be deformed based on the motion data to align the deformed previous frame with the current frame.
[0033] In step 630, an output frame can be generated. For example, when the current frame and the deformed previous frame are consistent, the output frame can include the current frame and the deformed previous frame. As another example, when the current frame is inconsistent with the deformed previous frame, the output frame can include only the current frame. Then, method 600 can process the output frame. In step 640, an AI model (i.e., a trained AI neural network) can be executed on the input frame (i.e., the output frame of step 630) to generate an AA frame or an AA+SR frame with aliasing artifacts removed.
[0034] Figure 7 A flowchart of an exemplary method 700 for processing an image or frame with aliasing artifacts is shown according to some embodiments of the present invention. In various embodiments, some of the steps shown in method 700 can be executed simultaneously, executed in an order different from the Figure 7 shown order, replaced by other method steps, or can be omitted. Additional method steps can also be executed as needed. Aspects of method 700 can be implemented by frame 300 processors 400 and 500 shown or described above based on the drawings. A trained AI neural network can be obtained through method 700.
[0035] In step 710, a first high-resolution frame with aliasing artifacts and a second high-resolution frame with aliasing artifacts removed are received, for example. The original input for generating the first high-resolution frame with aliasing artifacts and the second high-resolution frame with aliasing artifacts removed can be the training frames provided to the SR and AA engine 501 during the training phase of the AI neural network.
[0036] In step 720, an attention reference frame can be generated based on the first frame (i.e., the first high-resolution frame) and the second frame (i.e., the second high-resolution frame). In one embodiment, the attention reference frame can include key information of the first frame, and the key information is the information that the first frame is different from the second frame.
[0037] In step 730, a low-resolution frame (e.g., the output frame generated by the temporary decision circuit 330 during the training of the AI neural network) and the attention reference frame can be used to train the AI NN.
[0038] In step 740, the parameters of the AI model (AA or AA+SR) can be determined.
[0039] In step 750, an AI model (AA or AA+SR) (i.e., a trained AI neural network) for which parameters are determined or frozen can be obtained (e.g., the AI model used in step 640 of method 600).
[0040] In one embodiment according to the present invention, the motion estimation circuit 310, the warping circuit 320, the provisional decision circuit 330, and the frame fusion circuit 340 may include circuits configured to perform the functions and processes described herein in combination with software or without software. In another embodiment, the motion estimation circuit 310, the warping circuit 320, the provisional decision circuit 330, and the frame fusion circuit 340 may be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a digital enhancement circuit, or similar devices or a combination thereof. In another embodiment according to the present invention, the motion estimation circuit 310, the warping circuit 320, the provisional decision circuit 330, and the frame fusion circuit 340 may be a Central Processing Unit (CPU) configured to execute program instructions to perform the various functions and methods described herein. In various embodiments, the motion estimation circuit 310, the warping circuit 320, the provisional decision circuit 330, and the frame fusion circuit 340 may be different from each other. In some other embodiments, the motion estimation circuit 310, the warping circuit 320, the provisional decision circuit 330, and the frame fusion circuit 340 may be included in a single chip.
[0041] The device 300 and the frame processors 400 and 500 may optionally include other components, such as input and output devices, additional signal processing circuits, etc. Thus, the device 300 and the frame processors 400 and 500 may be capable of performing other additional functions, such as executing application programs, and processing alternative communication protocols.
[0042] Although the present invention is disclosed above in preferred embodiments, it is not intended to limit the scope of the present invention. Any person skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to what is defined by the claims.
Claims
1. A frame processing device, characterized in that, comprising: a motion estimation circuit configured to estimate motion data between a current frame and a previous frame; a warping circuit coupled to the motion estimation circuit and configured to warp the previous frame based on the motion data such that the warped previous frame is aligned with the current frame and to determine whether the current frame is consistent with the warped previous frame; and a provisional decision circuit coupled to the warping circuit and configured to generate an output frame, the output frame including the current frame and the warped previous frame when the current frame and the warped previous frame are consistent; and a frame processor coupled to the provisional decision circuit and configured to process the output frame including the current frame and the warped previous frame to eliminate aliasing artifacts in the output frame.
2. The device according to claim 1, characterized in that, the frame processor comprises: a super-resolution and anti-aliasing engine configured to enhance the resolution of a training frame and eliminate aliasing artifacts in the training frame to generate a first high-resolution frame with aliasing artifacts and a second high-resolution frame with eliminated aliasing artifacts.
3. The device according to claim 2, characterized in that, the frame processor further comprises: a region-of-interest reference frame generator coupled to the super-resolution and anti-aliasing engine and configured to generate a region-of-interest reference frame based on the first high-resolution frame and the second high-resolution frame.
4. The device according to claim 3, characterized in that, the frame processor further comprises: an artificial neural network coupled to the region-of-interest reference frame generator and configured to perform training based on the region-of-interest reference frame to obtain a trained artificial neural network and to eliminate aliasing artifacts in the output frame based on the trained artificial neural network.
5. The device according to claim 1, characterized in that, further comprising: a frame fusion engine coupled to the provisional decision circuit and configured to fuse the current frame and the warped previous frame.
6. The device according to claim 5, characterized in that, the frame fusion engine fuses the current frame and the warped previous frame by connecting the warped previous frame to the current frame in a channel-by-channel manner.
7. The device according to claim 1, characterized in that, the motion estimation circuit uses the sum of absolute differences method to estimate the motion data between the current frame and the previous frame.
8. The device according to claim 7, characterized in that, the warping circuit warps the previous frame based on the motion data and the warped previous frame matches the current frame when warped based on the motion data.
9. The device according to claim 1, characterized in that, when the motion data is less than a motion threshold, the warping circuit warps the previous frame based on the motion data.
10. The device according to claim 1, characterized in that, the motion estimation circuit is a neural network.
11. A method for processing a frame, characterized in that, comprising: estimating motion data between a current frame and a previous frame; Deform the previous frame based on the motion data such that the deformed previous frame is aligned with the current frame; Generate an output frame that includes the current frame and the deformed previous frame when the current frame and the deformed previous frame are identical; and Process the output frame that includes the current frame and the deformed previous frame to eliminate aliasing artifacts in the output frame.
12. The method according to claim 11, wherein, further comprising: Enhance the resolution of the training frame and eliminate the aliasing artifacts in the output frame to generate a first high-resolution frame with aliasing artifacts and a second high-resolution frame with eliminated aliasing artifacts.
13. The method according to claim 12, wherein, further comprising generating an attention reference frame based on the first high-resolution frame and the second high-resolution frame.
14. The method according to claim 13, wherein, further comprising: Performing training based on the attention reference frame to obtain a trained artificial neural network, and eliminating the aliasing artifacts in the output frame based on the trained artificial neural network.
Citation Information
Patent Citations
Image processing device, image processing method, and computer-readable storage device
US20170353666A1
Hierarchical Neural Network Image Registration
US20200327639A1