Real-time video super-resolution using motion vectors

By calculating motion vectors and hardware accelerator to optimize video frame processing, the problem of insufficient video super resolution performance in the prior art is solved, and a high-quality 720p or 1080p video super resolution is achieved in real time, meeting the rendering requirements of 60FPS.

CN120584355APending Publication Date: 2025-09-02MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480006848.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-07
Filing Date
2024-01-31
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing machine learning video super-resolution technology does not perform enough when generating real-time video super-resolution on client devices, and cannot effectively improve the conversion speed of low-resolution video to high-resolution video, resulting in the inability to render 60FPS 480P videos in real time.

Method used

By calculating the motion vector between the input low-resolution frame and the reference low-resolution frame, combined with hardware accelerators such as GPU or NPU, optimize video frame processing, run super-resolution only on keyframes and recalculate the motion vector, and use color lookup tables and block matching methods to process artifacts to achieve video super-resolution based on motion vectors.

Benefits of technology

It realizes the real-time generation of 720p or 1080p video super resolution on client devices, improves visual quality and reduces artifacts, and meets the rendering requirements of 60FPS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120584355A_ABST
    Figure CN120584355A_ABST
Patent Text Reader

Abstract

A low resolution image frame is received. And a previous reference low-resolution image frame and a super-resolution result thereof are reserved. A motion vector is calculated between an input low resolution image frame and a reference low resolution frame. For each pixel in the low resolution image frame, a reference pixel is determined based on the calculated motion vector and a reference low resolution frame. If the reference pixel equals the corresponding input pixel, the reference super-resolution result pixel is rendered for that pixel, otherwise the corresponding pixel from the current input frame is rendered.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Machine learning (ML) is increasingly being used to perform a variety of tasks in various environments where patterns and reasoning can be used instead of explicit programming. Image super-resolution is a computationally intensive technique for upscaling images using ML models to produce higher-fidelity images on client devices. When extending image super-resolution to video, such as upscaling a 360p video to 1080p on a client device, one issue that can arise is that the performance of such ML models is too slow to allow for real-time video super-resolution.

[0002] It is with respect to these and other considerations that the disclosure made herein is presented. Summary of the Invention

[0003] Disclosed are methods and systems for achieving real-time video super-resolution. Video super-resolution is generally a process for generating high-resolution video frames from low-resolution frames. In various embodiments, the disclosed super-resolution process includes an input low-resolution frame, a reference low-resolution frame, and a reference super-resolution result, and outputs a super-resolution frame. Current machine learning video super-resolution techniques attempt to achieve higher performance by reducing model size. However, such techniques may produce lower-quality super-resolution results. Pure machine learning inference cannot render, for example, 480p video at 60 frames per second (FPS), while the disclosed techniques make it possible to obtain super-resolution 720p or 1080p video based on low-resolution frames. In one embodiment, a motion vector is calculated between the input low-resolution frame and a reference low-resolution frame. For each pixel in the input low-resolution frame, a reference pixel is determined based on the calculated motion vector and the reference low-resolution frame. If the reference pixel is equal to the corresponding input pixel, the reference super-resolution result pixel for this pixel is rendered; otherwise, the corresponding pixel from the current input frame is rendered.

[0004] This summary is not intended to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to implementations that solve any or all disadvantages noted in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The detailed description is given with reference to the accompanying drawings. In the drawings, the leftmost digit(s) of a reference number identifies the drawing in which the reference number first appears. The same reference numbers in different drawings indicate similar or identical items.

[0006] Figure 1 is a diagram illustrating the disclosed technology according to one embodiment disclosed herein.

[0007] Figure 2is a diagram illustrating the disclosed technology according to one embodiment disclosed herein.

[0008] Figure 3 is a diagram illustrating aspects of an example system according to one embodiment disclosed herein.

[0009] Figure 4 is a diagram illustrating aspects of an example system according to one embodiment disclosed herein.

[0010] Figure 5 is a flow chart showing aspects of an illustrative routine according to one embodiment disclosed herein.

[0011] Figure 6 is a computer architecture diagram illustrating aspects of an example computer architecture for a computer capable of executing the software components described herein.

[0012] Figure 7 is a data architecture diagram showing an illustrative example of a computer environment. DETAILED DESCRIPTION

[0013] With reference to the accompanying drawings, in which like numerals represent like elements throughout the several figures, various aspects of various techniques for video super-resolution using motion vectors will be described. In the following detailed description, reference is made to the accompanying drawings, which form a part hereof and are shown by way of illustration of specific configurations or examples.

[0014] Image super-resolution is a computationally intensive technique for upscaling images using machine learning models to produce higher-fidelity images. In some embodiments, inference graphics processing units (GPUs) or neural processing units (NPUs) are used to accelerate processing. When extending image super-resolution to video, allowing client devices to upscale from 360p video to 1080p video, a problem that may arise is that the performance of such machine learning models is too slow to allow real-time video super-resolution.

[0015] For example, GPU inference time needs to be below 16 milliseconds in order to render 60 FPS video in real time. ML models can be optimized at the expense of super-resolution inference quality, but there are limits to how far ML optimization can continue before results start to degrade and can only match simple bicubic interpolation, making the cost of continued optimization unfavorable.

[0016] To perform less work per frame and allow for faster processing, work can be reused across space or time, or both. Some video encoders are configured to perform this type of reuse. For example, a video encoder can segment a video into tiles, mark certain frames as key frames, and calculate the position delta between these tiles, i.e., identify where the tile in the current frame is in the key frame and then encode the tile's position delta vector (motion vector) and the absolute pixel difference after the displacement.

[0017] If software video decoders for common video codecs were modified to run super-resolution only on keyframes, and the decoder played back motion vector deltas / tile deltas on top of these super-resolution keyframes, the result would be a large fraction of super-resolution output frames where super-resolution was not applied to every frame. However, this would have to be updated for every video codec, meaning client devices would have to incorporate software video decoding to implement such techniques.

[0018] Alternatively, many modern GPUs provide functionality to recalculate motion vectors between video frames. Many operating systems provide a public application programming interface (API) to use this functionality. For example, DirectX 12 provides hardware encoder functionality in the GPU to calculate motion vectors. Thus, by utilizing this GPU functionality, it is possible to recalculate motion vectors after or after hardware video decoding on any codec, without relying on information from the video codec.

[0019] Using the above-mentioned GPU-based method for calculating motion vectors, the present disclosure provides a method for generating real-time video super-resolution using the following algorithm:

[0020] Algorithm: Motion Vector Super Resolution

[0021] Input: Input low-resolution frame,

[0022] Reference low-resolution frame,

[0023] Reference super-resolution results (140)

[0024] Output: Super-resolution frame

[0025] If there is no reference super-resolution result:

[0026] Render the entire input low-resolution frame as is.

[0027] If a reference super-resolution result exists:

[0028] Calculate the motion vector between (input low-res frame, reference low-res frame) as motion_vector[width, height]

[0029] For each pixel (x,y) in the input low-resolution frame:

[0030] Reference_pixel = read from reference frame (x-motion_vector[x], y-motion-vector[y])

[0031] If reference_pixel~=input_pixel(x,y)

[0032] Renders the reference super-resolution result for this pixel (x-motion_vector[x], y-motion_vector[y]).

[0033] otherwise

[0034] The input pixel (x,y) to render this pixel at.

[0035] If there are no pending super resolve requests:

[0036] By running the full machine learning model, an update request for the reference super-resolution is scheduled for the current input. The determination of the reference pixel ~= input pixel (x, y) can be based on a threshold. The threshold can be a position threshold, such as a distance threshold. When the threshold is reached, the reference pixel is said to match the input pixel, or the reference pixel is equal to or identical to the input pixel. For example, the reference super-resolution result can be an upscaled image of the reference low-resolution frame using the machine learning model.

[0037] Figure 1 An example of the disclosed technology is shown. Input frame 100 is an input low-resolution frame, and reference input 110 is a reference low-resolution frame. Input frame 100 and reference input 110 are inputs for image processing 130. As described above, inference 120 is performed, and a super-resolution frame 140 is output. Image processing 130 is performed using input frame 100, reference input 110, and super-resolution frame 140 to output a motion vector-based super-resolution (SR) frame 135.

[0038] When the blending of super-resolution pixels with input pixels can be performed with acceptable latency and acceptable perceptible visual degradation compared to full super-resolution, the disclosed embodiments enable rendering video frames with a mixture of super-resolution pixels and pixels from the current input. Current machine learning video super-resolution techniques attempt to achieve performance by reducing model size at the expense of lower quality super-resolution results. Pure machine learning inference cannot render 480P video at 60FPS, while the disclosed embodiments make it possible to generate super-resolution 720p or 1080p video.

[0039] In some embodiments, additional improvements can be achieved in visual quality related to motion vector super-resolution when compared to pure inference. These improvements include two sources of visual degradation: color shading and motion vector inaccuracy.

[0040] When the output of motion vector-based super-resolution is scaled, artifacts may begin to appear. One type of artifact is caused by the model learning a color correction technique that changes the color of the reference super-resolution image compared to the reference input image. These color corrections are similar to automatic contrast corrections that enhance video frames by increasing the dynamic range.

[0041] In motion vector-based video super-resolution, since pixels are mixed from the original input and pixels that have been passed through the model, the color difference between these two types of pixels can appear as patches in the final frame. To handle this color difference, in embodiments, the behavior of the model can be quantized and replicated on non-super-resolution pixels. The color correction of the model can be a spatial function of the pixel's position in the image (for example, the pixel can be close to an edge), a side effect of a contrast / sharpen kernel, or a function of the color itself.

[0042] In one embodiment, a color lookup table is learned from the results of pure inference. For each generated reference super-resolution image, each pixel is processed in a compute shader on the GPU. The input image is compared with the output super-resolution image color for the same pixel, the difference is accumulated, and the count is stored in a buffer indexed by a 15-bit key calculated from the color of the input pixel. It should be noted that the key size may vary. The average color difference is calculated for each color bucket. An example implementation is provided below:

[0043] Unsigned integer calculate color key (float r, float g, float b) {

[0044] / / Only 5 bits per channel

[0045] Unsigned integer r_comp = 0x1F&((unsigned integer) rounded down (r*31.0));

[0046] Unsigned integer g_comp = 0x1F&((unsigned integer) rounded down (g*31.0));

[0047] Unsigned integer b_comp = 0x1F&((unsigned integer) rounded down (b*31.0));

[0048] / / 15-digit key

[0049] unsigned integer key = r_comp<<10|g_comp<<5|b_comp;

[0050] return key;

[0051] }

[0052] void colormap()

[0053] {

[0054] …

[0055] unsigned int key = computeColorKey(input_pixels.r, input_pixels.g, input_pixels.b);

[0056] unsigned int current_value;

[0057] interlockedadd(color_map_c[key],1,current_value);

[0058] input_pixel = input_pixel * 255.0;

[0059] input_sr_pixel = input_sr_pixel * 255.0;

[0060] input_r_val = round(input_sr_pixel.r - input_pixel.r);

[0061] int g_val = round(input_sr_pixel.g - input_pixel.g);

[0062] int b_val = round(input_sr_pixel.b - input_pixel.b);

[0063] interlockedadd(color_map_r[key], r_val, current_value);

[0064] interlockedadd(color_map_g[key], g_val, current_value);

[0065] interlockedadd(color_map_b[key], b_val, current_value);

[0066] }

[0067] In an embodiment, four separate buffers are used instead of one buffer indexed by a key so that multiple cores can process the buffers in parallel. During motion vector based super resolution, the same color delta as used when super resolution pixels are not available for frame regions is applied back to the input pixels.

[0068] Another source of artifacts in motion vector-based video super-resolution is due to inaccuracies in the motion vectors. In one example, the motion vector calculation process operates by convolving each 8×8 patch of the frame with a 1 / 4 pixel offset step over the 8×8 area. The best resulting offset is reported as the motion offset for that 8×8 block. The motion vector accounts for the motion of objects in the scene and the camera itself.

[0069] Motion vectors in video coding applications are intended to be used with a pipeline that encodes the delta between the proposed tile offset and the pixel in the actual tile position. They are not expected to match perfectly. Some reasons for not matching perfectly include:

[0070] The motion vectors are 2D translations, while the real scene is undergoing a 3D transformation, such as a rotation of the camera. The 2D translation is a rough approximation of the original transformation for a small 8×8 patch.

[0071] Some operations, such as zooming in on a scene, may not have a consistent offset for every pixel in a tile. Some of these operations may reveal hidden parts of the scene due to parallax effects.

[0072] Entirely new objects can be introduced into the scene, or objects can change color, such as traffic lights. These changes cannot be represented as motion vectors.

[0073] Given that the pixels are not perfectly matched, a determination is made using the given data to select input pixels or super-resolution pixels from a reference frame, which may be referred to as a pure reference frame. In one embodiment, a method may be referred to as block matching, with a threshold adjusted based on the overall quality of the block. This threshold may be referred to as the pixel acceptance threshold. As an output of this step, an image is generated with a value between 0 and 1 indicating:

[0074] 0 – use current input pixel, 1 – use reference super-resolution pixel.

[0075] In an embodiment, the following process may be performed to determine whether to select an input pixel or a super-resolution pixel from a reference frame and to blur the transition between the super-resolution pixel and the input low-resolution pixel:

[0076] Step 1 - Calculate the sum of absolute pixel differences (SAD) between the motion-adjusted reference input and the current input for a motion vector block of 8×8 or 4×4 pixels. This sum can be considered the quality of the motion vector value.

[0077] Step 2: If the SAD value is below the threshold, the super-resolution result of the block is accepted as is.

[0078] Step 3 - If the SAD value is above the threshold, compare each pixel and accept those pixels that are below the individual pixel comparison threshold.

[0079] Step 4 - Apply a Gaussian smoothing operator or filter to blur the transition from the super-resolution pixels to the original input pixels.

[0080] Figure 2A functional diagram illustrating the disclosed technology is shown. In an embodiment, the solid arrows in the diagram run at a high priority, while the dashed arrows run opportunistically at a lower priority / frame rate. Dashed boxes indicate intermediate results. Solid boxes indicate processes. Block 201 shows the calculation of motion vectors using a hardware encoder to generate motion vectors 203 for an 8×8 tile. Block 202 shows color compensation of an input frame 207 to generate a color-compensated input frame 204. Block 217 shows running inference using an ML model based on a reference input 209. In block 218, a color lookup table is calculated using the high-quality super-resolution frame 216 and the reference input 209 to generate a color lookup table 220. Block 221 shows blending the reference super-resolution frame 216 and the color-compensated input frame 204 and outputting a super-resolution frame 219 generated using the motion vectors. Color compensated block matching includes computing block quality 205 for 4x4 tiles using the sum of absolute pixel differences, performing pixel comparison 211 on poor quality tiles, and smoothing the results to avoid patches 213 to output a pixel source 215 for each input pixel.

[0081] In various embodiments, the machine learning model(s) may be run locally on the client. In other embodiments, the machine learning inference may be performed on a server on the network. For example, Figure 3 In the illustrated system, system 300 is shown that implements an ML platform 330. ML platform 330 can be configured to provide output data to various devices 350 and computing device 330 via network 320. User interface 360 ​​can be presented on computing device 330. User interface 360 ​​can be provided in conjunction with application 340, which communicates with ML platform 330 via network 320 using an API. In some embodiments, system 300 can be configured to provide product information to users. In one example, ML platform 330 can implement a machine learning system to perform one or more tasks. ML platform 330 utilizes a machine learning system to perform tasks such as image and handwriting recognition. The machine learning system can be configured to be optimized using the techniques described herein.

[0082] Figure 4 FIG. 1 is a diagram of a computing system architecture according to an embodiment disclosed herein, which shows an overview of the system disclosed herein for implementing a machine learning model. Figure 4As shown, the machine learning system 400 can be configured to perform analysis and perform identification, prediction, or other functions based on various data collected and processed by data analysis components 430 (which may be referred to individually as "data analysis components 430" or collectively as "data analysis components 430"). Data analysis components 430 may include, for example, but are not limited to, physical computing devices such as server computers or other types of hosts, associated hardware components (e.g., memory and mass storage devices), and networking components (e.g., routers, switches, and cables). Data analysis components 430 may also include software such as operating systems, applications, and containers, network services, virtual components such as virtual disks, virtual networks, and virtual machines. Database 450 may include data, such as a database or database shards (i.e., partitions of a database). Feedback can be used to further update various parameters used by the machine learning model 420. Data can be provided to user applications 415 to provide results to various users 410 using user applications 415. In some configurations, the machine learning model 420 can be configured to utilize supervised and / or unsupervised machine learning techniques. The model compression framework based on sparsity-induced regularization optimization as disclosed in this paper can reduce the amount of data that needs to be processed in such systems and applications. When processing iterations over large amounts of data, effective model compression can provide improved latency for many applications that use such techniques, such as image and sound recognition, recommendation systems, and image analysis.

[0083] Now go to Figure 5 , shows an example operating procedure for generating an image according to the present disclosure. The operating procedure can be implemented in a system including one or more computing devices.

[0084] It will be understood by those skilled in the art that the operations of the methods disclosed herein are not necessarily presented in any particular order, and that it is possible and contemplated to perform some or all of the operations in (a plurality of) alternative orders. For ease of description and illustration, the operations have been presented in the order presented. Operations may be added, omitted, performed together, and / or performed simultaneously without departing from the scope of the appended claims.

[0085] It should also be understood that the methods shown can be terminated at any time and do not need to be performed in full. Some or all of the operations of these methods, and / or substantially equivalent operations, can be performed by executing computer-readable instructions, which are included on a computer storage medium, as defined herein. As used in the specification and claims, the term "computer-readable instructions" and variations thereof are used extensively herein to include routines, applications, application modules, program modules, programs, components, data structures, algorithms, and the like. Computer-readable instructions can be implemented on various system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like. Although the example routines described below operate on a computing device, it will be understood that this routine can be executed on any computing system that may include multiple computers working together to perform the operations disclosed herein.

[0086] It should be understood, therefore, that the logical operations described herein are implemented as (1) a sequence of computer-implemented actions or program modules running on a computing system such as those described herein and / or (2) interconnected machine logic circuits or circuit modules within the computing system. Implementation is a matter of choice depending on the performance and other requirements of the computing system. Thus, the logical operations may be implemented in software, in firmware, in dedicated digital logic, or any combination thereof.

[0087] Reference Figure 5 Operation 501 illustrates receiving, by a computing system, a low-resolution image frame as a current input frame, a reference super-resolution result, and a reference low-resolution frame.

[0088] Operation 501 may be followed by operation 503. Operation 503 illustrates calculating a motion vector between the low-resolution image frame and a reference low-resolution frame.

[0089] Operation 503 may be followed by operation 505. Operation 505 illustrates determining, for each pixel in the low-resolution frame, a reference pixel based on the calculated motion vector and the reference low-resolution frame.

[0090] Operation 505 may be followed by operation 507. Operation 507 illustrates that for each pixel in the low-resolution image frame,

[0091] If the reference pixel is equal to the corresponding input pixel, the reference super-resolution result pixel for this pixel is rendered, otherwise the corresponding pixel from the current input frame is rendered.

[0092] Figure 6 An example computer architecture is shown for a computer capable of providing the functionality described herein, for example, configured to implement the above referenced Figures 1 to 6Therefore, Figure 6 The illustrated computer architecture 600 illustrates the architecture of a server computer or another type of computing device suitable for implementing the functionality described herein. The computer architecture 600 can be used to execute the various software components presented herein to implement the disclosed technology.

[0093] Figure 6 The illustrated computer architecture 600 includes a central processing unit 602 ("CPU"), system memory 604 (including random access memory 606 ("RAM") and read-only memory ("ROM") 608), and a system bus 77 that couples the memory 604 to the CPU 602. Firmware, which contains basic routines that facilitate the transfer of information between elements within the computer architecture 600, such as during startup, is stored in the ROM 608. The computer architecture 600 also includes a mass storage device 612 for storing an operating system 614, other data, such as product data 615, or user data 617.

[0094] The mass storage device 612 is connected to the CPU 602 through a mass storage controller (not shown) connected to the bus 77. The mass storage device 612 and its associated computer-readable media provide non-volatile storage for the computer architecture 600. Although the description of computer-readable media contained herein refers to mass storage devices, such as solid-state drives, hard disks, or optical drives, those skilled in the art will understand that computer-readable media can be any available computer-readable storage media or communication media that can be accessed by the computer architecture 600.

[0095] Communication media includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any delivery medium. The term "modulated data signal" means a signal whose characteristics are changed or set in such a way as to encode information in the signal. By way of example and not limitation, communication media include wired media such as a wired network or a direct wired connection, and wireless media such as acoustic, radio frequency, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0096] By way of example, and not limitation, computer-readable storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. For example, computer media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technology, CD-ROM, digital versatile disks ("DVD"), HD-DVD, BLU-RAY, or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and accessed by the computer architecture 600. For purposes of the claims, the phrases "computer storage medium," "computer-readable storage medium," and variations thereof do not, by themselves, include waves, signals, and / or other transient and / or intangible communication media.

[0097] According to various implementations, the computer architecture 600 can operate in a networked environment using logical connections to remote computers through the network 650 and / or another network (not shown). A computing device implementing the computer architecture 600 can connect to the network 650 via a network interface unit 616 connected to the bus 77. It should be understood that the network interface unit 616 can also be used to connect to other types of networks and / or remote computer systems.

[0098] The computer architecture 600 may also include an input / output controller 618 for receiving and processing input from a variety of other devices, including a keyboard, a mouse, or an electronic stylus ( Figure 6 Similarly, the input / output controller 618 can send data to a display screen, printer, or other type of output device ( Figure 6 ) provides an output.

[0099] It should be understood that when loaded into CPU 602 and executed, the software components described herein can transform CPU 602 and the entire computer architecture 600 from a general-purpose computing system into a special-purpose computing system customized to facilitate the functionality presented herein. CPU 602 can be constructed from any number of transistors or other discrete circuit elements, which can individually or collectively assume any number of states. More specifically, CPU 602 can operate as a finite state machine in response to the executable instructions contained in the software modules disclosed herein. These computer-executable instructions can transform CPU 602 by specifying how CPU 602 transitions between states, thereby transforming the transistors or other discrete hardware elements that make up CPU 602.

[0100] Encoding the software modules proposed herein may also transform the physical structure of the computer-readable medium proposed herein. In different implementations of this specification, the specific transformation of the physical structure may depend on various factors. Examples of these factors may include, but are not limited to, the technology used to implement the computer-readable medium, whether the computer-readable medium is characterized as a primary storage device or a secondary storage device, etc. If the computer-readable medium is implemented as a semiconductor-based memory, the software disclosed herein can be encoded on the computer-readable medium by transforming the physical state of the semiconductor memory. For example, the software can transform the state of transistors, capacitors, or other discrete circuit elements that constitute the semiconductor memory. The software can also transform the physical state of such components in order to store data thereon.

[0101] As another example, the computer-readable medium disclosed herein can be implemented using magnetic or optical technology. In such implementations, when software is encoded therein, the software proposed herein can convert the physical state of the magnetic or optical medium. These conversions may include changing the magnetic properties of a given location within the magnetic medium. These conversions may also include changing the physical features or characteristics of a given location within the optical medium to change the optical properties of those locations. Other conversions of physical media are possible without departing from the scope and spirit of this specification, wherein the aforementioned examples are provided only for ease of discussion.

[0102] In view of the foregoing, it should be understood that many types of physical transformations occur within the computer architecture 600 in order to store and execute the software components presented herein. It should also be understood that the computer architecture 600 may include other types of computing devices, including handheld computers, embedded computer systems, personal digital assistants, and other types of computer devices known to those skilled in the art.

[0103] It is also contemplated that computer architecture 600 may not include Figure 6 All components shown may include Figure 6 Other components not explicitly shown in the figure or may utilize Figure 6 For example, and not limitation, the techniques disclosed herein can be used with multiple CPUs for improved performance through parallelization, with graphics processing units ("GPUs") for faster computation, and / or with tensor processing units ("TPUs"). As used herein, the term "processor" encompasses CPUs, GPUs, TPUs, and other types of processors.

[0104] Figure 7 It shows that it is possible to perform the above Figures 1-6An example computing environment for describing the techniques and processes. In various examples, the computing environment includes a host system 702. In various examples, the host system 702 operates on, in communication with, or as part of a network 704.

[0105] The network 704 may be or may include various access networks. For example, one or more client devices 706(1)...706(N) may communicate with the host system 702 via the network 704 and / or other connections. The host system 702 and / or the client devices may include, but are not limited to, any of a variety of devices, including portable or fixed devices, such as a server computer, a smartphone, a mobile phone, a personal digital assistant (PDA), an e-reader device, a laptop computer, a desktop computer, a tablet computer, a portable computer, a game console, a personal media player device, or any other electronic device.

[0106] According to various implementations, the functionality of the host system 702 may be provided by one or more servers that execute as part of or in communication with the network 704. The servers may host various services, virtual machines, portals, and / or other resources. For example, they may host or provide access to one or more portals, websites, and / or other information.

[0107] The host system 702 may include (multiple) processors 708, memory 710. The memory 710 may include an operating system 712, (multiple) applications 714, and / or a file system 716. In addition, the memory 710 may include the above-mentioned Figures 1 to 5 The storage unit 82 is described.

[0108] (Multiple) processors 708 can be a single processing unit or multiple units, each of which can include multiple different processing units. (Multiple) processors can include microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units (CPUs), graphics processing units (GPUs), security processors, etc. Alternatively or additionally, some or all of the techniques described herein can be performed at least in part by one or more hardware logic components. For example, and not limitation, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), state machines, complex programmable logic devices (CPLDs), other logic circuit devices, systems on chip (SoCs), and / or any other device that performs operations based on instructions. Among other functions, (multiple) processors can be configured to obtain and execute computer-readable instructions stored in memory 710.

[0109] Memory 710 may include one or a combination of computer-readable media. As used herein, "computer-readable media" includes computer storage media and communication media.

[0110] Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, phase change memory (PCM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory or other memory technology, compact disc ROM (CD-ROM), digital versatile disk (DVD) or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store information for access by a computing device.

[0111] In contrast, communication media includes computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave. As defined herein, computer storage media does not include communication media.

[0112] The host system 702 can communicate over the network 704 via a network interface 718. The network interface 718 can include various types of network hardware and software for supporting communication between two or more devices. The host system 702 can also include a machine learning model 719.

[0113] Finally, although various techniques have been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended representations is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed subject matter.

[0114] The disclosure set forth herein also covers the subject matter set forth in the following clauses:

[0115] Clause 1: A method of generating an image, the method comprising:

[0116] The computing system receives a low-resolution image frame as a current input frame, a reference super-resolution result, and a reference low-resolution frame;

[0117] Calculating a motion vector between the current input frame and a reference low-resolution frame; and

[0118] For an input pixel in the current input frame:

[0119] determining a position of a reference pixel based on the calculated motion vector;

[0120] determining that the position of the reference pixel is within a distance threshold of the position of the corresponding input pixel; and

[0121] Based on determining that the position of the reference pixel is within the distance threshold, a reference super-resolution result pixel in a reference super-resolution result associated with the input pixel is rendered.

[0122] Clause 2: The method of clause 1, wherein the input pixel is a first input pixel, further comprising:

[0123] For the second input pixel in the current input frame:

[0124] determining a second position of a second reference pixel based on the calculated motion vector;

[0125] determining that a second position of the second reference pixel is not within a distance threshold of a position of a corresponding second input pixel; and

[0126] Based on determining that the second position of the second reference pixel is not within the distance threshold, rendering a corresponding pixel from the current input frame.

[0127] Clause 3: A method according to any one of clauses 1 to 2, wherein the position of the reference pixels is determined as the difference between the x and y values ​​of each pixel and the calculated motion vector.

[0128] Clause 4: The method of any one of clauses 1 to 3, further comprising scheduling a request to update a reference super-resolution for the current input by running a machine learning model.

[0129] Clause 5: The method according to any one of clauses 1 to 4, further comprising:

[0130] Learn a color lookup table from the results of the reference frame;

[0131] Comparing the input image and the output super-resolution image colors; and

[0132] Accumulate the differences and counts for each color bucket to calculate the average color difference for each color bucket.

[0133] Clause 6: The method of any one of clauses 1 to 5, further comprising applying the same color delta to the input pixels when super-resolution pixels are not available for a region of the input frame.

[0134] Clause 7: The method according to any one of clauses 1 to 6, further comprising:

[0135] Evaluate each block's quality using the sum of its absolute pixel differences, and

[0136] Adjust pixel acceptance threshold based on block quality.

[0137] Clause 8: The method according to any one of clauses 1 to 7, further comprising:

[0138] For each n×n block, the quality of the motion vector value is determined by performing the sum of the absolute pixel differences between the reference input and the current input;

[0139] For motion vectors above a threshold, run a pixel-by-pixel comparison, otherwise accept the current result; and

[0140] A Gaussian smoothing filter is applied to the determination of pixel origin to smooth the transition between super-resolution and non-super-resolution regions.

[0141] Clause 9: A computing system comprising:

[0142] one or more processors; and

[0143] A computer-readable storage medium having computer-executable instructions stored thereon that, when executed by a processor, cause a computing system to perform operations comprising:

[0144] receiving a low-resolution image frame, a reference super-resolution result, and a reference low-resolution frame;

[0145] calculating a motion vector between the low-resolution image frame and a reference low-resolution frame; and

[0146] For each pixel in the low-resolution image frame:

[0147] generating reference pixels based on the calculated motion vector and the reference low-resolution frame; and

[0148] When the reference pixel is within the position threshold of the corresponding input pixel, the reference super-resolution result pixel for the current pixel is rendered, otherwise the corresponding pixel from the low-resolution image frame is rendered.

[0149] Clause 10: The system of clause 9, wherein the width and height of the motion vector are calculated.

[0150] Clause 11: The system of any of clauses 9 and 10, wherein the position of the reference pixels is determined as the difference between the x and y values ​​of each pixel and the calculated motion vector.

[0151] Clause 12: The system of any one of clauses 9 to 11, further comprising computer executable instructions stored on the system that, when executed by a processor, cause the computing system to perform operations comprising scheduling a request to update a reference super-resolution for a current input by running a machine learning model.

[0152] Clause 13: The system of any one of clauses 9 to 12, further comprising computer-executable instructions stored on the system that, when executed by a processor, cause the computing system to perform operations comprising:

[0153] Learn a color lookup table from the results of a pure reference frame;

[0154] Comparing the input image and the output super-resolution image colors; and

[0155] Accumulate the differences and counts for each color bucket to calculate the average color difference for each color bucket.

[0156] Clause 14: A system according to any one of clauses 9 to 13, further comprising computer executable instructions stored on the system, which, when executed by a processor, cause the computing system to perform operations comprising applying the same color delta to input pixels when super-resolution pixels are not available for an area of ​​the input frame.

[0157] Clause 15: The system of any one of clauses 9 to 14, further comprising computer-executable instructions stored on the system that, when executed by a processor, cause the computing system to perform operations comprising:

[0158] Evaluate each block's quality using the sum of its absolute pixel differences; and

[0159] Adjust pixel acceptance threshold based on block quality.

[0160] Clause 16: The system of any one of clauses 9 to 15, further comprising computer-executable instructions stored on the system that, when executed by a processor, cause the computing system to perform operations comprising:

[0161] For each n×n block, the quality of the motion vector value is determined by performing the sum of the absolute pixel differences between the reference input and the current input;

[0162] For motion vectors above a threshold, run a pixel-by-pixel comparison, otherwise accept the result; and

[0163] A Gaussian smoothing filter is applied to the determination of pixel origin to smooth the transition between super-resolution and non-super-resolution regions.

[0164] Clause 17: A computer-readable storage medium having computer-executable instructions stored thereon that, when executed by a processor of a computing system, cause the computing system to perform operations comprising:

[0165] receiving a low-resolution image frame as a current input frame, a reference super-resolution result, and a reference low-resolution frame;

[0166] calculating a motion vector between the low-resolution image frame and a reference low-resolution frame; and

[0167] For each pixel in the low-resolution image frame:

[0168] Calculating a reference pixel based on the calculated motion vector and the reference low-resolution frame; and

[0169] When a reference pixel is within a threshold of the corresponding input pixel, the reference super-resolution result pixel for this pixel is rendered, otherwise the corresponding pixel from the current input frame is rendered.

[0170] Clause 18: The computer-readable storage medium of clause 17, further comprising computer-executable instructions stored on the computer-readable storage medium, which, when executed by a processor, cause a computing system to perform operations comprising:

[0171] Learn a color lookup table from the results of a pure reference frame;

[0172] Comparing the input image and the output super-resolution image colors; and

[0173] Accumulate the differences and counts for each color bucket to calculate the average color difference for each color bucket.

[0174] Clause 19: A computer-readable storage medium according to any one of clauses 17 and 18, further comprising computer-executable instructions stored on the computer-readable storage medium, which, when executed by a processor, cause a computing system to perform operations comprising applying the same color delta to input pixels when super-resolution pixels are not available for an area of ​​the input frame.

[0175] Clause 20: The computer-readable storage medium of any one of clauses 17 to 19, further comprising computer-executable instructions stored on the computer-readable storage medium, which, when executed by a processor, cause a computing system to perform operations comprising:

[0176] The quality of each block is evaluated using the sum of the absolute pixel differences of the block;

[0177] Adjust pixel acceptance threshold based on block quality;

[0178] For each n×n block, the quality of the motion vector value is determined by performing the sum of the absolute pixel differences between the reference input and the current input;

[0179] For motion vectors above a threshold, run a pixel-by-pixel comparison, otherwise accept the result as is; and

[0180] A Gaussian smoothing filter is applied to the determination of pixel origin to smooth the transition between super-resolution and non-super-resolution regions.

Claims

1. A method for generating an image, the method comprising: The computing system receives a low-resolution image frame as a current input frame, a reference super-resolution result, and a reference low-resolution frame; Calculating a motion vector between the current input frame and the reference low-resolution frame; as well as For the input pixels in the current input frame: determining a position of a reference pixel based on the calculated motion vector; determining that the position of the reference pixel is within a distance threshold of a position of a corresponding input pixel; as well as Based on determining that the position of the reference pixel is within the distance threshold, a reference super-resolution result pixel in the reference super-resolution result associated with the input pixel is rendered.

2. The method of claim 1 , wherein the input pixel is a first input pixel, further comprising: For the second input pixel in the current input frame: determining a second position of a second reference pixel based on the calculated motion vector; determining that the second position of the second reference pixel is not within the distance threshold of a position of a corresponding second input pixel; as well as Based on determining that the second position of the second reference pixel is not within the distance threshold, rendering a corresponding pixel from the current input frame. 3 . The method of claim 1 , wherein the locations of the reference pixels are determined as the difference between the x and y values ​​of each pixel and the calculated motion vector.

4. The method according to claim 1, further comprising: A request to update the reference super-resolution for the current input is scheduled by running a machine learning model.

5. The method according to claim 1, further comprising: Learn a color lookup table from the results of the reference frame; Compare the colors of the input image and the output super-resolution image; as well as The color bucket-by-color differences and counts are accumulated to calculate the average color difference for each color bucket.

6. The method according to claim 5, further comprising: When super-resolution pixels are not available for a region of the input frame, the same color delta is applied to the input pixels.

7. The method according to claim 1, further comprising: Evaluate per-block quality using the sum of absolute pixel differences for the block, and A pixel acceptance threshold is adjusted based on the block quality.

8. The method according to claim 7, further comprising: For each n×n block, the quality of the motion vector value is determined by performing the sum of the absolute pixel differences between the reference input and the current input; For motion vectors above the threshold, run pixel-by-pixel comparison, otherwise accept the current result; as well as A Gaussian smoothing filter is applied to the determination of pixel origin to smooth the transition between super-resolution and non-super-resolution regions.

9. A computing system comprising: one or more processors; as well as A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by the processor, cause the computing system to perform operations comprising: receiving a low-resolution image frame, a reference super-resolution result, and a reference low-resolution frame; calculating a motion vector between the low-resolution image frame and the reference low-resolution frame; and For each pixel in the low-resolution image frame: generating a reference pixel based on the calculated motion vector and the reference low-resolution frame; and When the reference pixel is within a position threshold of the corresponding input pixel, the reference super-resolution result pixel for the current pixel is rendered; otherwise, the corresponding pixel from the low-resolution image frame is rendered.

10. The computing system of claim 9, wherein the locations of the reference pixels are determined as the difference between the x and y values ​​of each pixel and the calculated motion vector.

11. The computing system of claim 9, further comprising computer-executable instructions stored on the computing system, which, when executed by the processor, cause the computing system to perform operations comprising: A request to update the reference super-resolution for the current input is scheduled by running a machine learning model.

12. The computing system of claim 9, further comprising computer-executable instructions stored on the computing system that, when executed by the processor, cause the computing system to perform operations comprising: Learn a color lookup table from the results of a pure reference frame; Compare the input image and the output super-resolution image colors; as well as Accumulate the differences and counts of each color bucket to calculate the average color difference for each color bucket; and When super-resolution pixels are not available for a region of the input frame, the same color delta is applied to the input pixels.

13. The computing system of claim 9, further comprising computer-executable instructions stored on the computing system that, when executed by the processor, cause the computing system to perform operations comprising: The quality of each block is evaluated using the sum of absolute pixel differences for the block; as well as A pixel acceptance threshold is adjusted based on the block quality.

14. The computing system of claim 13, further comprising computer-executable instructions stored on the computing system, which, when executed by the processor, cause the computing system to perform operations comprising: For each n×n block, the quality of the motion vector value is determined by performing the sum of the absolute pixel differences between the reference input and the current input; For motion vectors above the threshold, run a pixel-by-pixel comparison, otherwise accept the result; as well as A Gaussian smoothing filter is applied to the determination of pixel origin to smooth the transition between super-resolution and non-super-resolution regions.

15. A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor of a computing system, cause the computing system to perform operations comprising: receiving a low-resolution image frame as a current input frame, a reference super-resolution result, and a reference low-resolution frame; Calculating a motion vector between the low-resolution image frame and the reference low-resolution frame; as well as For each pixel in the low-resolution image frame: Calculating a reference pixel based on the calculated motion vector and the reference low-resolution frame; as well as When the reference pixel is within a threshold of the corresponding input pixel, a reference super-resolution result pixel for this pixel is rendered; otherwise, the corresponding pixel from the current input frame is rendered.