Compressed information-based video super-resolution

KR103024489B1Active Publication Date: 2026-09-23GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020237026123
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-26
Filing Date
2021-08-05
Publication Date
2026-09-23
Estimated Expiration
2041-08-05

Smart Images

  • Figure 112023084050848-PCT00036_ABST
    Figure 112023084050848-PCT00036_ABST
Patent Text Reader

Abstract

Exemplary aspects of the present disclosure relate to systems and methods featuring machine learning video super-resolution (VSR) models trained using a bidirectional training approach. In particular, the present disclosure provides a compression information-based (e.g., compression-aware) super-resolution model capable of performing well on real-world videos at different compression levels. Specifically, the exemplary models described herein may include three modules for robustly restoring missing information caused by video compression. First, a bidirectional recirculation module may be used to reduce accumulated warp errors from random intra-frame locations within compressed video frames. Second, a detail-aware flow estimation module may be added to enable the restoration of high-resolution (HR) flows from compressed low-resolution (LR) frames. Finally, a Laplacian enhancement module may add high-frequency information to warped HR frames that have been washed out by video encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This application claims priority and interest to U.S. Provisional Patent Application No. 63 / 179,795 filed on April 26, 2021. By doing so, U.S. Provisional Patent Application No. 63 / 179,795 is incorporated by reference in its entirety.

[0002] The present disclosure generally relates to systems and methods for performing compression-informed video super-resolution. More specifically, the present disclosure relates to systems and methods featuring a machine learning video super-resolution model trained using an interactive training approach. Background Technology

[0003] Super-resolution is a fundamental research problem in computer vision, which has numerous applications. Systems performing super-resolution aim to reconstruct detailed high-resolution image(s) from low-resolution input(s) (LR). When the input is a single image, the reconstruction process typically uses learned image priors to restore high-resolution details of the given image, which can be referred to as single image super-resolution (SISR). When multiple frames of a video are available, certain reconstruction processes can use both image priors and inter-frame information to generate temporally smooth high-resolution results, which can be referred to as video super-resolution (VSR).

[0004] Although significant progress has been made in the field of super-resolution, existing SISR and VSR methods barely consider compression. Specifically, certain previous works used "uncompressed" data to highlight high-quality, low-compression videos. Consequently, these earlier methods tend to generate significant artifacts when operated on overly compressed input videos.

[0005] In particular, most digital videos (e.g., videos existing on the Internet or mobile devices, such as smartphones) are stored and / or streamed with different levels of compression to achieve a selected visual quality level. For example, the popular compression rate (Constant Rate Factor (CRF)) for H.264 encoding is 23, representing a trade-off between quality and file size. Existing techniques designed and optimized for the application of VSR to uncompressed video data do not perform well when applied to videos compressed in that manner.

[0006] One possible solution is to apply a denoising model to remove compression artifacts, followed by the application of one of the latest VSR models. At first glance, this is attractive because high-quality frames are fed to the VSR model, similar to using evaluation data directly. However, experiments have indicated that such a setup will not improve final performance; in fact, it can even worsen it. With pre-processing, it is highly likely that the denoising model in the first stage alters the degradation kernel implicitly used during VSR model training. Therefore, essentially, VSR models are being applied to more demanding data.

[0007] Another possible solution is to train existing, state-of-the-art VSR models on compressed frames. This can introduce additional compression information into the model training. However, experiments have indicated that simply using compressed frames in training yields only moderate improvements. In fact, without specific changes to the designs of the network modules, such training data can even compromise overall performance.

[0008] Therefore, improved systems, methods, model architectures, and training approaches are needed to provide improved VSR for compressed video data.

[0009] A system of one or more computers may be configured to perform specific operations or actions by installing software, firmware, hardware, or a combination thereof on the system, which causes the system to perform actions when in operation. One or more computer programs may be configured to perform specific operations or actions by including instructions that cause the device to perform actions when executed by a data processing device.

[0010] One general aspect includes a computer-implemented method for bidirectionally training a machine learning video super-resolution (VSR) model using compressed video data. The computer-implemented method includes the step of acquiring a set of ground truth video data that may include a plurality of ground truth higher-resolution (HR) video frames and a plurality of lower-resolution (LR) video frames by a computing system comprising one or more computing devices, wherein the plurality of LR video frames correspond to a plurality of ground truth HR video frames, and the plurality of ground truth HR video frames and the plurality of LR video frames are arranged in a time sequence corresponding to the compressed video. The method also includes the step of performing forward time prediction by the computing system to generate a forward-predicted HR video frame for a current position in a time sequence based on one or more video frames associated with one or more previous positions in a time sequence. The method also includes the step of performing backward time prediction by the computing system to generate a backward-predicted HR video frame for a current position in a time sequence based on one or more video frames associated with one or more subsequent positions in a time sequence. The method also includes the step of evaluating a loss function for a machine learning VSR model by a computing system, wherein the loss function compares a real-world data HR video frame with a forward-predicted HR video frame and compares a real-world data HR video frame with a backward-predicted HR video frame. The method also includes the step of modifying one or more values ​​of one or more parameters of a machine learning VSR model based on the loss function by a computing system.Other embodiments of the present aspect include corresponding computer systems, devices and computer programs recorded on one or more computer storage devices, each of which is configured to perform the actions of the methods.

[0011] Another exemplary aspect relates to a computing system comprising one or more non-transient computer-readable media that collectively store machine learning VSR (video super resolution) models and instructions, and one or more processors, wherein the instructions, when executed by one or more processors, cause the computing system to use the machine learning VSR model to super-resolve compressed video.

[0012] A machine learning VSR (video super resolution) model includes a flow estimation part configured to process a previous or subsequent LR video frame and a current LR video frame to generate a lower resolution (LR) flow estimation and a higher resolution (HR) flow estimation; warp the previous or subsequent LR video frame according to the LR flow estimation to generate a predicted LR video frame for the current position within the time sequence; and warp the previous or subsequent HR video frame according to the HR flow estimation to generate an intermediate HR video frame for the current position within the time sequence; a Laplacian enhancement part configured to enhance the intermediate HR video frame; and a frame generation part configured to process the intermediate HR video frame and the current LR video frame to generate a predicted HR video frame for the current position within the time sequence.

[0013] Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium. Brief explanation of the drawing

[0014] Detailed descriptions of embodiments intended for those skilled in the art are provided herein, and the specification refers to the accompanying drawings.

[0015] FIGS. 1a and 1b illustrate graphic diagrams of an exemplary process for bidirectionally training a machine learning video super-resolution model according to exemplary embodiments of the present disclosure.

[0016] FIG. 2 illustrates a graphic diagram of an exemplary architecture of an exemplary machine learning video super-resolution model according to exemplary embodiments of the present disclosure.

[0017] FIG. 3a illustrates a block diagram of an exemplary computing system according to exemplary embodiments of the present disclosure.

[0018] FIG. 3b illustrates a block diagram of an exemplary computing device according to exemplary embodiments of the present disclosure.

[0019] FIG. 3c illustrates a block diagram of an exemplary computing device according to exemplary embodiments of the present disclosure.

[0020] Reference numbers repeated across multiple drawings are intended to identify the same features in various implementations. Specific details for implementing the invention

[0021] Exemplary aspects of the present disclosure relate to systems and methods featuring machine learning video super-resolution (VSR) models trained using a bidirectional training approach. In particular, the present disclosure provides a compression information-based (e.g., compression-aware) super-resolution model capable of performing well on real-world videos at different compression levels. Specifically, the exemplary models described herein may include three modules for robustly restoring missing information caused by video compression. First, a bidirectional recirculation module may be used to reduce accumulated warp errors from random intra-frame locations within compressed video frames. Second, a detail-aware flow estimation module may be added to enable the restoration of high-resolution (HR) flows from compressed low-resolution (LR) frames. Finally, a Laplacian enhancement module may add high-frequency information to warped HR frames that have been washed out by video encoding. Exemplary implementations of the proposed model may be referred to as COMISR (Compression-Informed Video Super-Resolution) in some cases.

[0022] In U.S. Provisional Patent Application No. 63 / 179,795, which is included in part of the present disclosure and forms part thereof, the validity of exemplary implementations of COMISR having three modules is demonstrated through ablation studies. In particular, extensive experiments were conducted on several VSR benchmark datasets using videos compressed with different CRF values. The experiments indicated that the COMISR model achieves significant performance gains for compressed videos (e.g., CRF23); while maintaining competitive performance for uncompressed videos. Additionally, U.S. Provisional Patent Application No. 63 / 179,795 presents evaluation results based on different combinations of modern VSR models and off-the-shelf video denominations. Finally, U.S. Provisional Patent Application No. 63 / 179,795 demonstrates the robustness of the COMISR model for simulating streaming YouTube videos compressed by proprietary encoders.

[0023] Accordingly, one exemplary aspect of the present disclosure relates to a compression-informed model for super-resolve real-world compressed videos for actual applications. Another exemplary aspect includes three new modules in VSR to effectively improve important components of video super-resolution on compressed frames. Finally, extensive experiments were performed on modern VSR models on compressed benchmark datasets.

[0024] The systems and methods of the present disclosure provide a number of technical effects and benefits. As an example, the models described herein can perform improved image processing, such as improved super-resolution imagery (e.g., increasing the resolution of the imagery through image synthesis). For example, by performing bidirectional training of the VSR model, the VSR model can be better equipped / trained to handle temporal artifacts introduced by the compression process.

[0025] Specifically, one common technique used in video compression is to apply different algorithms to compress and encode frames at different locations in a video stream. Typically, a codec randomly selects several reference frames known as intra-frames and compresses them independently without using information from other frames. It then compresses other frames by utilizing consistency and encoding the differences from the intra-frames. Consequently, intra-frames generally require more bits to encode and have fewer compression artifacts than other frames. In video super-resolution, since the locations of intra-frames are not known in advance, the proposed bidirectional approach can be used to enforce forward and backward consistency of LR warped inputs and HR predicted frames to effectively reduce accumulated errors from unknown locations of intra-frames.

[0026] The systems and methods of the present disclosure may be used in a number of applications. In one example, the models described herein may be used to increase the resolution of compressed videos. For example, compressed videos may be transmitted or streamed in a compressed form and then super-resolve at an end device displaying the video. This may provide a technical advantage of conserving network bandwidth and storage space because compressed videos may require fewer computational resources to transmit and / or store. As examples, the compressed videos may be compressed video conferencing video streams, compressed user-generated content videos, and / or any other types of videos.

[0027] Now, with reference to the drawings, exemplary embodiments of the present disclosure will be discussed in more detail. Exemplary Model Training and Inference

[0028] Exemplary COMISR models are designed using cyclic formulations that feed previous information into the current frame, similar to modern video SR methods. Cyclic designs generally involve low memory consumption and can be applied to many inference tasks of videos.

[0029] The exemplary model architecture described herein includes three novel parts—bidirectional recurrent warping, detail-aware flow estimation, and Laplacian enhancement—which can make the model robust to compressed videos. Given LR real-world data frames, the model can apply forward and backward recurrent modules (see Figs. 1a and 1b) to generate HR frame predictions and compute content losses for HR real-world data frames in both directions. The recurrent modules predict flows and generate warped frames for both LR and HR, and can train the network end-to-end using LR and HR real-world data frames.

[0030] Exemplary bidirectional circulation module

[0031] One technique used in video compression is to apply different algorithms to compress and encode frames at different locations in a video stream. Typically, a codec randomly selects several reference frames known as intra-frames and compresses them independently without using information from other frames. It then compresses other frames by utilizing consistency and encoding the differences from the intra-frames. Consequently, intra-frames generally require more bits to encode and have fewer compression artifacts than other frames. In video super-resolution, since the locations of intra-frames are not known in advance, to effectively reduce accumulated errors from unknown locations of intra-frames, the present disclosure proposes a bidirectional recurrent network for enforcing forward and backward consistency of LR warped inputs and HR predicted frames.

[0032] Specifically, a bidirectional recurrent network may include symmetric modules for the forward and reverse directions. In the forward direction, the model first takes LR frames and LR flow using and HR one Both can be estimated. Next, the model can apply different behaviors individually to the LR and HR streams. In the LR stream, the model, warped LR frame In order to obtain Using the previous LR frame It can warp to time t, which will be used in later stages: (1)

[0033] In the HR stream, the model is a warped HR frame In order to obtain Previous predicted frames using It can be warped to time t, and a Laplacian enhancement module follows to generate an accurate HR warped frame. (2) (3)

[0034] Next, the model Applying spatial-depth operations to it to expand its channels while reducing its resolution, and this as the LR input Fusion with, and pass the contiguous frames to the HR frame generator to predict the final HR. can be obtained. The training process is to measure the loss actual measurement data HR It can be compared with.

[0035] Similarly, the model can apply reverse symmetric operations to obtain warped LR frames and predicted HR frames. In this case, the detail-aware flow estimation module can generate a reverse flow from time t to t-1, and warping can be performed by applying the reverse flow to the frame at time t to estimate the frame at time t-1.

[0036] As examples, FIGS. 1a and FIGS. 1b illustrate exemplary VSR models used for forward time prediction and backward time prediction, respectively. In some embodiments of the present disclosure, forward time prediction may be performed during both training and inference, whereas backward time prediction may be performed only during training.

[0037] Specifically, in some implementations, to train a VSR model, a computing system may acquire multiple sets of actual data training data. Training iterations may be performed on batches of training videos, each batch comprising one or more sets of actual data video data.

[0038] In particular, a set of actual data video data may include multiple actual data HR (higher-resolution) video frames and multiple LR (lower-resolution) video frames. Each of the multiple LR video frames corresponds to each of the multiple actual data HR video frames. For example, each LR frame may be a relatively lower resolution version of the corresponding HR frame among the HR frames. In one example, frames of the HR video may be downsampled and / or compressed to generate LR frames. The HR frames may be compressed or may not be compressed.

[0039] Multiple actual data HR video frames and multiple LR video frames can be arranged in a time sequence. As an example, the time sequence may correspond to numbered frames that are aligned in sequence and captured in sequential order by an image capture device.

[0040] Model training can occur across one or more positions in a time sequence. For example, training can occur for all positions in a time sequence.

[0041] Specifically, a VSR model may be used by a computing system to perform forward time prediction to generate a forward-predicted HR video frame for a current position in a time sequence based on one or more video frames associated with one or more previous positions in a time sequence. An example of forward time prediction is illustrated in FIG. 1a. Additionally, according to an aspect of the present disclosure, a VSR model may also be used by a computing system to perform backward time prediction to generate a backward-predicted HR video frame for a current position in a time sequence based on one or more video frames associated with one or more subsequent positions in a time sequence. An example of backward time prediction is illustrated in FIG. 1b.

[0042] In some implementations, forward and backward models are symmetric and share weights. If considered differently, the same model may be used for each of the forward and backward passes, but different (e.g., opposite) orders or sequences may be applied to the frames. For example, the order of the frames may simply be reversed.

[0043] When forward and / or backward time predictions are performed, the computing system can evaluate a loss function for the machine learning VSR model. As an example, the loss function may perform both (1) comparing a real-world HR video frame with a forward-predicted HR video frame generated by forward time prediction and (2) comparing a real-world HR video frame with a backward-predicted HR video frame generated by backward time prediction. The loss function may be evaluated jointly for both (1) and (2) above, or (1) and (2) may be evaluated individually and then summed, or otherwise processed together (e.g., as a batch).

[0044] A computing system can modify one or more values ​​of one or more parameters of a machine learning VSR model based on a loss function. For example, backpropagation of errors can be used to update the values ​​of the parameters of a machine learning VSR model according to the gradient of the loss function.

[0045] In particular, referring specifically to FIG. 1a, in some implementations, performing forward time prediction may include processing a previous HR video frame (14) associated with a previous position in a time sequence, a previous LR video frame (16) associated with a previous position in a time sequence, and a current LR video frame (18) associated with a current position in a time sequence, using a machine learning VSR model (12) by a computing system to generate a forward-predicted HR video frame (20) for a current position in a time sequence.

[0046] Likewise, specifically referring to FIG. 1b, performing reverse time prediction may include processing a subsequent HR video frame (24) associated with a subsequent position in a time sequence, a subsequent LR video frame (26) associated with a subsequent position in a time sequence, and a current LR (28) associated with a current position in a time sequence, in order to generate a reverse-predicted HR video frame (30) for a current position in a time sequence by using a machine learning VSR model (12) by a computing system.

[0047] The previous HR video frame (14) may be a previous predicted HR video frame or a previous actual data HR video frame. Likewise, the subsequent HR video frame (24) may be a subsequent predicted HR video frame or a subsequent actual data HR video frame.

[0048] Exemplary Circular Model Details

[0049] Now, referring to FIG. 2, an exemplary architecture for an exemplary VSR model (200) is illustrated. The model (200) may include a flow estimation part (202), a Laplacian enhancement part (204), and / or a frame generation part (206).

[0050] The flow estimation part (202) may be configured to process a previous or subsequent LR video frame (e.g., a previous LR video frame (16)) and a current LR video frame (18) to generate a lower resolution (LR) flow estimation (210) and a higher resolution (HR) flow estimation (212). The flow estimation part (202) may warp the previous or subsequent LR video frame (e.g., 16) according to the LR flow estimation (210) to generate a predicted LR video frame (214) for the current position in the time sequence. The flow estimation part (202) may warp the previous or subsequent HR video frame (e.g., a previous HR frame (14)) according to the HR flow estimation (212) to generate an intermediate HR video frame (216) for the current position in the time sequence.

[0051] The Laplacian enhancement portion (204) can be configured to enhance the intermediate HR video frame (216).

[0052] The frame generation part (206) may be configured to process the intermediate HR video frame (216) and the current LR video frame (18) (e.g., after enhancement) to generate the predicted HR video frame (20) for the current position in the time sequence.

[0053] Likewise, performing reverse time prediction (not specifically illustrated) by a computing system may include processing a subsequent LR video frame and a current LR video frame to generate an LR reverse flow estimate and an HR reverse flow estimate by using a flow estimation part (202) of a machine learning VSR model by the computing system; warping a subsequent LR video frame according to the LR reverse flow estimate by the computing system to generate a reverse-predicted LR video frame for the current position in the time sequence by the computing system; and warping a subsequent HR video frame according to the HR reverse flow estimate by the computing system to generate a reverse-intermediate HR video frame for the current position in the time sequence.

[0054] Likewise, performing reverse time prediction by a computing system may include applying a Laplacian enhancement filter to a reverse-intermediate HR video frame by the computing system; and, after applying the Laplacian enhancement filter, processing the reverse-intermediate HR video frame and the current LR video frame to generate a reverse-predicted HR video frame for the current position in the time sequence using the frame generation part of a machine learning VSR model by the computing system.

[0055] In some implementations, the loss function may additionally compare (3) a forward-predicted LR video frame (214) for the current position with a current LR video frame (18) associated with the current position in the time sequence; and / or (4) a backward-predicted LR video frame (not specifically shown) for the current position with a current LR video frame associated with the current position in the time sequence.

[0056] In a time sequence, a previous position can be a position immediately preceding it in the time sequence, or a time position that is not directly adjacent. Similarly, a subsequent position in a time sequence can be a position immediately preceding it, or a time position that is not directly adjacent.

[0057] After training, the machine learning VSR model can be used to super-resolve additionally compressed video. For example, using the machine learning VSR model to super-resolve additionally compressed video may involve performing only forward time prediction on the video frames of the additionally compressed video.

[0058] The training techniques described herein may be performed on multiple compressed training videos for each of multiple training iterations. The multiple compressed training videos may be compressed using the same compression algorithm or multiple different compression algorithms. One exemplary compression algorithm is the H.264 codec.

[0059] Exemplary Details - Perception Flow Estimation

[0060] In the proposed iterative model, the model can explicitly estimate both LR and HR flows between neighboring frames and transmit this information in forward and backward directions.

[0061] Figure 2 illustrates the forward direction for illustrative purposes. Operations in the reverse direction are applied similarly. The model first, two adjacent LR frames and Connect them and transmit them through the LR flow estimation network to LR flow Can generate LR flow Instead of directly upsampling, the model can apply several additional deconvolution layers on top of the bilinearly upsampled LR stream. Thus, detailed residual maps can be encouraged to be learned during end-to-end training, and as a result, the model can better preserve high-frequency details in the predicted HR stream.

[0062] Exemplary Laplacian enhancement module

[0063] Laplacian residuals have been widely used in many vision tasks, including image blending, super-resolution, and reconstruction. They are particularly useful for finding fine details in video frames, where these details can be smoothed out during video compression. In some examples of the proposed cyclic VSR model, the warped predicted HR frame retains information learned from previous frames and some details. These details can easily be lost in the upscaling network. Therefore, some exemplary implementations include Laplacian residuals for the predicted HR frame to enhance details.

[0064] The Laplacian-elevated image is a Gaussian kernel blur with a width of σ It can be computed by: (4) Here, are the intermediate results of the predicted HR frames, and α is a weighted factor controlling the residual power. By using the Laplacian, the model can add details back to the warped HR frames. This may be followed by a spatial-depth operation that rearranges blocks of spatial data into depth dimensions and then concatenates them with the LR input frames. The model can pass this to an HR frame generator to produce the final HR prediction.

[0065] Exemplary loss functions

[0066] During training, typically two streams exist: HR and LR frames. Losses can be designed considering the use of both streams. In the case of loss for HR frames, between the final outputs and the HR frames Distance can be computed. represents the actual measurement data frame, and represents a frame generated at time t. For each of the cycle steps, the predicted HR frames can be used to compute the loss. Losses can be optionally combined as follows. (5)

[0067] Each of the warped LR frames from t-1 to t also, for the current LR frame as follows You can receive a penalty for the distance. (6)

[0068] One exemplary total loss may be the sum of HR and LR losses, and (7) Here, β and These are the weights for each loss. Exemplary devices and systems

[0069] FIG. 3a illustrates a block diagram of an exemplary computing system (100) for performing video super-resolution according to exemplary embodiments of the present disclosure. The system (100) includes a user computing device (102), a server computing system (130), and a training computing system (150) coupled to communicate via a network (180).

[0070] The user computing device (102) may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0071] A user computing device (102) includes one or more processors (112) and memory (114). One or more processors (112) may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or a plurality of processors operably connected. Memory (114) may include one or more non-transient computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory (114) may store instructions (118) and data (116) executed by the processor (112) to enable the user computing device (102) to perform operations.

[0072] In some implementations, the user computing device (102) may store or contain one or more machine learning VSR models (120). For example, the machine learning VSR models (120) may be various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models including non-linear models and / or linear models, or may otherwise include them. Neural networks may include feed-forward neural networks, recurrent neural networks (e.g., long short memory recurrent neural networks), convolutional neural networks, or other types of neural networks. Some exemplary machine learning models may leverage attention mechanisms such as self-attention. For example, some exemplary machine learning models may include multi-head self-attention models (e.g., transformer models). Exemplary machine learning VSR models (120) are discussed with reference to FIGS. 1a, FIGS. 1b and FIG. 2.

[0073] In some implementations, one or more machine learning VSR models (120) are received from a server computing system (130) via a network (180), stored in user computing device memory (114), and then used by one or more processors (112) or implemented in other ways. In some implementations, the user computing device (102) may implement multiple parallel instances of a single machine learning VSR model (120) (e.g., to perform parallel video super-resolution over multiple instances of lower resolution videos).

[0074] Additionally or alternatively, one or more machine learning VSR models (140) may be included in a server computing system (130) that communicates with a user computing device (102) according to a client-server relationship, or otherwise stored and implemented by it. For example, machine learning VSR models (140) may be implemented by the server computing system (140) as part of a web service (e.g., a video super-resolution service). Thus, one or more models (120) may be stored and implemented on the user computing device (102), and / or one or more models (140) may be stored and implemented on the server computing system (130).

[0075] The user computing device (102) may also include one or more user input components (122) for receiving user input. For example, the user input component (122) may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may function to implement a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means that enable the user to provide user input.

[0076] A server computing system (130) includes one or more processors (132) and memory (134). One or more processors (132) may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or a plurality of processors operably connected. Memory (134) may include one or more non-transient computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory (134) may store instructions (138) and data (136) executed by the processor (132) to enable the server computing system (130) to perform operations.

[0077] In some implementations, the server computing system (130) includes one or more server computing devices or is otherwise implemented by them. In cases where the server computing system (130) includes multiple server computing devices, these server computing devices may operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0078] As described above, the server computing system (130) may store or otherwise contain one or more machine learning VSR models (140). For example, the models (140) may be various machine learning models or otherwise contain them. Exemplary machine learning models include neural networks or other multilayer nonlinear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some exemplary machine learning models may leverage attention mechanisms such as self-attention. For example, some exemplary machine learning models may include multi-head self-attention models (e.g., transformer models). Exemplary models (140) are discussed with reference to FIGS. 1a, FIGS. 1b, and FIGS. 2.

[0079] A user computing device (102) and / or a server computing system (130) can train models (120 and / or 140) through interaction with a training computing system (150) coupled to communicate via a network (180). The training computing system (150) may be separate from the server computing system (130) or may be part of the server computing system (130).

[0080] The training computing system (150) includes one or more processors (152) and memory (154). The one or more processors (152) may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or a plurality of processors operably connected. The memory (154) may include one or more non-transient computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory (154) may store instructions (158) and data (156) executed by the processor (152) to enable the training computing system (150) to perform operations. In some implementations, the training computing system (150) includes one or more server computing devices or is otherwise implemented by them.

[0081] The training computing system (150) may include a model trainer (160) that trains machine learning models (120 and / or 140) stored in a user computing device (102) and / or a server computing system (130) using various training or learning techniques, such as backpropagation of errors, for example. For example, a loss function may be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on the gradient of the loss function). Various loss functions, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions, may be used. Gradient descent techniques may be used to iteratively update parameters over a number of training iterations.

[0082] In some implementations, performing backpropagation of errors may involve performing truncated backpropagation over time. The model trainer (160) may perform a number of generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the models being trained.

[0083] In particular, the model trainer (160) can train machine learning VSR models (120 and / or 140) based on a set of training data (162). The training data (162) may include, for example, actual data video data. For example, the actual data video data may include videos of both a higher resolution form and a corresponding lower resolution form.

[0084] In some implementations, training data may include REDS and / or Vimeo datasets for training. The REDS dataset contains over 200 video sequences for training, each containing 100 frames with a resolution of 1280 × 720. The Vimeo-90K dataset contains approximately 65k video sequences for training, each containing 7 frames with a resolution of 448 × 256. One major difference between these two datasets is that the REDS dataset has much greater motion between consecutive frames captured from a handheld device. To train and evaluate the COMISR model, frames can first be smoothed by a Gaussian kernel with a width of 1.5 and downsampled by 4x.

[0085] In some implementations, the COMISR model can be evaluated on Vid4 and REDS4 datasets (clip# 000, 011, 015, 020). All test sequences have more than 30 frames.

[0086] In some implementations, the following compression methods may be used. One example follows the most common settings for the H.264 codec at different compression rates (i.e., different CRF values). Recommended CRF values ​​are 18 to 28, and the default is 23 (although the range of values ​​is 0 to 51). In some examples, CRFs of 15, 25, and 35 may be used to evaluate video super-resolution at a wide range of compression rates. In some implementations, the same degradation method is used to generate LR sequences before compression. Finally, such compressed LR sequences are fed to VSR models for inference.

[0087] In some implementations, the following training process may be used. In some implementations, for each input frame, the training process may randomly crop patches (e.g., 128 × 128 patches) from a mini-batch as input. Each mini-batch may contain multiple samples (e.g., 16 samples). α, β and The parameters can be set to 1, 20, and 1, respectively. Model training can be supervised by losses described elsewhere in this invention. The Adam optimizer can be used with β_1=0.9 and β_2=0.999. The learning rate can be set to 5×10^(-5). Video compression can optionally be adopted as an additional data augmentation method for the training pipeline with a 50% probability for input batches.

[0088] In some implementations, if the user has provided consent, training examples may be provided by the user computing device (102). Thus, in these implementations, the model (120) provided to the user computing device (102) may be trained by the training computing system (150) on user-specific data received from the user computing device (102). In some cases, this process may be referred to as personalizing the model.

[0089] The model trainer (160) includes computer logic utilized to provide desired functions. The model trainer (160) may be implemented as hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer (160) includes program files that are stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer (160) includes one or more sets of computer-executable instructions stored on a computer-readable storage medium of a type such as RAM, a hard disk, or optical or magnetic media.

[0090] The network (180) may be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or any combination thereof, and may include any number of wired or wireless links. Generally, communication through the network (180) may be transmitted via any type of wired and / or wireless connection using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XLM), and / or protection methods (e.g., VPN, secure HTTP, SSL).

[0091] FIG. 3a illustrates one exemplary computing system that may be used to implement the present disclosure. Other computing systems may also be used. For example, in some implementations, a user computing device (102) may include a model trainer (160) and a training dataset (162). In these implementations, models (120) may be trained and used locally on the user computing device (102). In some of these implementations, the user computing device (102) may implement the model trainer (160) to personalize the models (120) based on user-specific data.

[0092] FIG. 3b illustrates a block diagram of an exemplary computing device (10) performed according to exemplary embodiments of the present disclosure. The computing device (10) may be a user computing device or a server computing device.

[0093] The computing device (10) includes a plurality of applications (e.g., applications 1 to N). Each application includes its own machine learning library and machine learning model(s). For example, each application may include a machine learning model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0094] As illustrated in FIG. 3b, each application may communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application may communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0095] FIG. 3c illustrates a block diagram of an exemplary computing device (50) performed according to exemplary embodiments of the present disclosure. The computing device (50) may be a user computing device or a server computing device.

[0096] The computing device (50) includes a number of applications (e.g., applications 1 through N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application may communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API across all applications).

[0097] The central intelligence layer includes multiple machine learning models. For example, as illustrated in FIG. 3c, individual machine learning models may be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of the computing device (50).

[0098] The central intelligence layer can communicate with the central device data layer. The central device data layer may be a centralized repository of data for the computing device (50). As illustrated in FIG. 3c, the central device data layer may communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer may communicate with each device component using an API (e.g., a private API). addition Disclosure

[0099] The technology discussed herein refers to servers, databases, software applications, and other computer-based systems, as well as information transmitted to and from such systems and actions taken. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and partitions of tasks and functions between and between components. For example, the processes discussed herein may be implemented using a single device or component or multiple devices or components operating in combination. Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components may operate sequentially or in parallel.

[0100] Although the gist of this document has been described in detail with respect to various specific exemplary embodiments of the present invention, each example is provided by way of description rather than as a limitation of the disclosure. Those skilled in the art can easily make modifications, variations thereof, and equivalents to such embodiments by understanding the foregoing. Accordingly, the disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the gist of this document as will be readily apparent to those skilled in the art. For example, features exemplified or described as part of one embodiment may be used in conjunction with other embodiments to yield another embodiment. Accordingly, the disclosure is intended to cover such modifications, variations, and equivalents.

Claims

Claim 1 A computer-implemented method for bidirectionally training a machine learning video super-resolution (VSR) model using compressed video data, comprising: a step of acquiring a set of ground truth video data including a plurality of ground truth higher-resolution (HR) video frames and a plurality of lower-resolution (LR) video frames by a computing system comprising one or more computing devices — wherein the plurality of LR video frames correspond to the plurality of ground truth HR video frames, and the plurality of ground truth HR video frames and the plurality of LR video frames are arranged in a time sequence corresponding to the compressed video —; for each of one or more positions within the time sequence: a step of performing forward time prediction by the computing system to generate a forward-predicted HR video frame for a current position within the time sequence based on one or more video frames associated with one or more previous positions within the time sequence — wherein the step of performing forward time prediction is performed by the computing system using the machine learning VSR model, the previous HR video frame associated with the previous position within the time sequence, the previous LR video frame associated with the previous position within the time sequence, and the time sequence Includes the step of processing a current LR video frame associated with a current position within to generate a forward-predicted HR video frame for the current position within the time sequence.A step of performing backward time prediction by the computing system to generate a backward-predicted HR video frame for a current position in the time sequence based on one or more video frames associated with one or more subsequent positions in the time sequence — the step of performing backward time prediction includes the step of generating the backward-predicted HR video frame for a current position in the time sequence by processing a subsequent HR video frame associated with a subsequent position in the time sequence, a subsequent LR video frame associated with a subsequent position in the time sequence, and a current LR video frame associated with a current position in the time sequence using the machine learning VSR model by the computing system —; a step of evaluating a loss function for the machine learning VSR model by the computing system — the loss function compares the actual data HR video frame with the forward-predicted HR video frame and compares the actual data HR video frame with the backward-predicted HR video frame —; A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, comprising the step of modifying one or more values ​​of one or more parameters of the machine learning VSR model based on the loss function by the computing system, wherein the previous position in the time sequence includes the position immediately preceding the time sequence, and the subsequent position in the time sequence includes the position immediately following the time sequence. Claim 2 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein, in claim 1, the previous HR video frame includes a previously predicted HR video frame; and the subsequent HR video frame includes a subsequent predicted HR video frame. Claim 3 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein, in claim 1, the preceding HR video frame includes a preceding actual data HR video frame; and the subsequent HR video frame includes a subsequent actual data HR video frame. Claim 4 In claim 1, the step of performing the forward time prediction by the computing system comprises: the step of processing the previous LR video frame and the current LR video frame to generate LR forward flow estimation and HR forward flow estimation by using the flow estimation portion of the machine learning VSR model by the computing system; the step of warping the previous LR video frame according to the LR forward flow estimation to generate a forward-predicted LR video frame for the current position in the time sequence by the computing system; and the step of warping the previous HR video frame according to the HR forward flow estimation to generate a forward-intermediate HR video frame for the current position in the time sequence by the computing system; and the step of performing the backward time prediction by the computing system comprises: the step of processing the subsequent LR video frame and the current LR video frame to generate LR backward flow estimation and HR backward flow estimation by using the flow estimation portion of the machine learning VSR model by the computing system. A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, comprising: a step of warping the subsequent LR video frame according to the LR reverse flow estimation to generate a reverse-predicted LR video frame for the current position in the time sequence by the computing system; and a step of warping the subsequent HR video frame according to the HR reverse flow estimation to generate a reverse-intermediate HR video frame for the current position in the time sequence by the computing system. Claim 5 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein the loss function further compares a forward-predicted LR video frame for the current position with a current LR video frame associated with the current position in the time sequence; and compares a backward-predicted LR video frame for the current position with a current LR video frame associated with the current position in the time sequence. Claim 6 In claim 4, the step of performing the forward time prediction by the computing system further comprises: the step of applying a Laplacian enhancement filter to the forward-intermediate HR video frame by the computing system; and the step of processing the forward-intermediate HR video frame and the current LR video frame to generate a forward-predicted HR video frame for the current position in the time sequence by using the frame generation part of the machine learning VSR model by the computing system after applying the Laplacian enhancement filter; and the step of performing the backward time prediction by the computing system further comprises the step of applying the Laplacian enhancement filter to the backward-intermediate HR video frame by the computing system. A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, comprising the step of processing the reverse-intermediate HR video frame and the current LR video frame to generate a reverse-predicted HR video frame for the current position in the time sequence using the frame generation part of the machine learning VSR model by the computing system after applying the Laplacian enhancement filter. Claim 7 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein, in any one of claims 1 to 6, the compressed video comprises a compressed video conferencing video stream. Claim 8 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein, in any one of claims 1 to 6, the method further comprises the step of using the machine learning VSR model to super-resolve the additionally compressed video by the computing system, and the step of using the machine learning VSR model to super-resolve the additionally compressed video by the computing system comprises the step of performing only forward time prediction on the video frames of the additionally compressed video. Claim 9 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein, in any one of claims 1 to 6, the method described in claim 1 is further included for a plurality of compressed training videos during a plurality of training iterations, and said plurality of compressed training videos are compressed using the same compression algorithm. Claim 10 In claim 9, the compression algorithm comprises an H.264 codec, and the computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data. Claim 11 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein, in any one of claims 1 to 6, the method described in claim 1 is further included for a plurality of training iterations for a plurality of compressed training videos, and the plurality of compressed training videos are compressed using two or more different compression algorithms. Claim 12 A computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data, wherein, in any one of claims 1 to 6, the method further comprises the step of using the machine learning VSR model to super-resolve additionally compressed video by the computing system. Claim 13 In claim 12, the step of using the machine learning VSR model to super-resolve the additionally compressed video by the computing system comprises the step of performing only forward time prediction on the video frames of the additionally compressed video, a computer-implemented method for bidirectionally training a machine learning VSR model using compressed video data. Claim 14 One or more non-transient computer-readable media for collectively storing machine learning VSR models trained according to the method of any one of paragraphs 1 through 6. Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete

Citation Information

Patent Citations

  • Frame-Recurrent Video Super-Resolution

    US20190206026A1