Real-time super-resolution image processing method and system
By using efficient parallel neural network pipeline architecture and global statistical attribute generation weights or biases on small or edge devices, the problem of slow image processing in the prior art is solved, and high-quality real-time super-resolution image processing is achieved.
Patent Information
- Application Number
- CN202510087257.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-26
- Filing Date
- 2022-02-07
- Publication Date
- 2025-05-13
AI Technical Summary
When processing real-world images, it is difficult to achieve high-quality real-time super-resolution image processing, especially on small or edge devices. Traditional neural network technology has large computing capacity and slow speed, and cannot provide high-resolution video in real time.
Generate weights or biases and provide them to super-resolution neural networks to generate high-resolution images by using efficient parallel neural network pipeline architecture and global statistical properties. In addition, the local attributes of the image are included in the training protocol to simulate real-world image capture conditions and improve the quality of HR images.
Real-time super-resolution image processing is implemented on small or edge devices, and the generated high-resolution images are of high quality and sufficient sharpness, which can provide higher efficiency image processing on high-end devices.
Smart Images

Figure CN119991442A_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of the invention patent application with application date of February 7, 2022, application number 202210116333.5, and titled “Method and system for real-time super-resolution image processing”. Technical Field
[0003] The present disclosure relates to a method and system for real-time super-resolution image processing. Background Art
[0004] Super-resolution (SR) imaging is the process of upscaling lower resolution (LR) images to higher resolution (HR) images, and is often performed to render video of a sequence of frames in HR. However, conventional real-time techniques still provide poor quality images when processing real-world, camera-captured images. Moreover, high-quality SR techniques are often based on neural networks, and are therefore computationally intensive and, therefore, relatively slow, such that these high-quality conventional network-based super-resolution techniques are unable to perform super-resolution conversion to provide HR video in real time on small or edge devices. Summary of the invention
[0005] According to one embodiment of the present disclosure, there is provided a computer-implemented method of super-resolution image processing, comprising: obtaining at least one lower resolution (LR) image; generating at least one high resolution (HR) image, comprising inputting image data of the at least one LR image into at least one super-resolution (SR) neural network; separately generating weights or biases or both, comprising using statistical data of the LR image; and providing the weights or biases or both to the SR neural network for generating the at least one HR image.
[0006] According to another embodiment of the present disclosure, a system for image processing is provided, comprising: at least one processor; and at least one memory, communicatively coupled to the at least one processor and storing a video sequence of low-resolution (LR) images, wherein the at least one processor is configured to operate by the following steps: generating a higher-resolution (HR) image, comprising inputting the image data of the LR image into at least one super-resolution (SR) neural network; separately generating weights or biases or both, comprising using LR image statistics, and generating them at intervals of several images along the video sequence; and providing the weights or biases or both to the SR neural network at the intervals for generating the HR image.
[0007] According to another embodiment of the present disclosure, there is provided a method of image processing, comprising: obtaining an initial training image; generating a blurred training image, comprising intentionally blurring one or more of the initial training images; generating a low-resolution training image or a high-resolution training image or both, and comprising using the blurred image; and training a super-resolution neural network to generate high-resolution pixel values, comprising using the low-resolution training image or the high-resolution training image or both.
[0008] According to another embodiment of the present disclosure, at least one non-transitory article is provided, having at least one computer-readable medium, comprising a plurality of instructions, wherein the plurality of instructions, in response to being executed on a computing device, causes the computing device to operate through the following steps: obtaining initial training images; generating blurred training images, including intentionally blurring one or more of the initial training images; generating low-resolution training images or high-resolution training images, or both, and including using the blurred images; and training a super-resolution neural network to generate high-resolution pixel values, including using the low-resolution training images or the high-resolution training images, or both.
[0009] According to another embodiment of the present disclosure, at least one machine-readable medium is provided, comprising a plurality of instructions, wherein in response to being executed on a computing device, the plurality of instructions causes the computing device to execute the above method.
[0010] According to another embodiment of the present disclosure, a device is provided, including means for executing the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The material described herein is illustrated in the accompanying drawings by way of example and not by way of limitation. For simplicity and clarity of illustration, the elements illustrated in the drawings are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. In addition, reference numerals are repeated between the drawings to indicate corresponding or similar elements when considered appropriate. In the drawings:
[0012] Figure 1 is the original LR image enhanced by the traditional method;
[0013] Figure 2 is improved by using a deep back-projection network Figure 1 The same image of
[0014] Figure 3 is an illustrative schematic diagram of the conventional super-resolution process;
[0015] Figure 4 is improved by using real-time super-resolution image processing according to at least one implementation of the present invention Figure 1 The same image of
[0016] Figure 5 is a flow chart of a runtime method of real-time super-resolution image processing according to at least one implementation of the present invention;
[0017] Figure 6 is a flow chart of a training method for real-time super-resolution image processing according to at least one implementation of the present invention;
[0018] Figure 7 is a schematic diagram of a training system for real-time super-resolution image processing according to at least one implementation of the present invention;
[0019] Figure 8 is a detailed flow chart of a training method for real-time super-resolution image processing according to at least one implementation of the present invention;
[0020] Fig. 9 is a schematic diagram of a runtime system for real-time super-resolution image processing according to at least one implementation of the present invention;
[0021] Fig.10 is a detailed flow chart of a runtime method for real-time super-resolution image processing according to at least one implementation of the present invention;
[0022] Fig.11 is a schematic diagram of a statistical neural network for a runtime method of real-time super-resolution image processing according to at least one implementation of the present invention;
[0023] Fig.12 is a schematic diagram of a super-resolution neural network according to at least one implementation of the present invention;
[0024] Fig.13 is an illustrative diagram of an example image processing system;
[0025] Fig.14 is an illustrative diagram of an example system; and
[0026] Fig.15 are illustrative diagrams of example systems, all arranged in accordance with at least some implementations of the present disclosure. DETAILED DESCRIPTION
[0027] One or more implementations are now described with reference to the accompanying drawings. Although specific configurations and arrangements are discussed, it should be understood that this is done for illustration purposes only. Those skilled in the relevant art will recognize that other configurations and arrangements may be used without departing from the spirit and scope of the description. Those skilled in the relevant art will appreciate that the techniques and / or arrangements described herein may also be used in various other systems and applications that are different from those described herein.
[0028] Although the following description sets forth various implementations that may be present in architectures such as system-on-a-chip (SoC) architectures, for example, the implementations of the techniques and / or arrangements described herein are not limited to a particular architecture and / or computing system, but may be implemented by any architecture and / or computing system for similar purposes. For example, the techniques and / or arrangements described herein may be implemented using various architectures such as multiple integrated circuit (IC) chips and / or packages, and / or various computing devices and / or consumer electronic (CE) devices, such as set-top boxes, televisions, smart monitors, smart phones, cameras, laptops, tablet devices, other edge-type devices, such as internet-of-things (IoT) devices including kitchen or laundry appliances, home security systems, and the like. In addition, although the following description may set forth many specific details, such as logic implementations, types and interrelationships of system components, logic partitioning / integration selections, and the like, the claimed subject matter may be implemented without such specific details. In other cases, some material, such as control structures and complete software instruction sequences, may not be shown in detail to avoid obscuring the material disclosed herein.
[0029] The materials disclosed herein may be implemented in hardware, firmware, software, or any combination thereof, unless otherwise stated. The materials disclosed herein may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include a read-only memory (ROM); a random access memory (RAM); a disk storage medium; an optical storage medium; a flash memory device; an electrical, optical, acoustic or other form of propagation signal (e.g., a carrier wave, an infrared signal, a digital signal, etc.), and others. In another form, a non-transient item such as a non-transient computer-readable medium may be used in conjunction with any of the examples or other examples mentioned above, except that it does not include the transient signal itself. It does include those elements, such as RAM, etc., that can temporarily store data in a "transient" manner in addition to the signal itself.
[0030] References in the specification to "one implementation," "an implementation," "an example implementation," etc., indicate that the described implementation may include a particular feature, structure, or characteristic, but not every implementation may include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. In addition, when a particular feature, structure, or characteristic is described in connection with an implementation, it is considered to be within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in connection with other implementations (whether or not explicitly described herein).
[0031] Systems, articles, media, and methods for real-time super-resolution image processing are described herein.
[0032] refer to Figure 1 , conventional non-network super-resolution techniques typically apply interpolation of low-resolution (LR) pixel data to generate high-resolution (HR) pixel data, such as using bicubic interpolation. These techniques typically produce relatively low-quality HR images. For example, the blurred non-network-based image 100 is an original image that is upscaled to a certain image size by simple interpolation to compare with other HR images here. This is obtained by capturing real-world images using a Sony IMX362 sensor for video SR demonstration and comparison, and the image is upscaled 4 times (x4). This same video is used for the comparison below.
[0033] refer to Figure 2 , traditional high-quality SR techniques are more sophisticated and use neural networks to compute HR image data. An example technique is a deep back-projection network (DBPN). Other SR neural network techniques specifically improve the generalization of learning-based SR networks. However, these SR network techniques have a very large computational load due to the large number of network parameters and / or operations per pixel, so that the images generated in real time are still quite inadequate and appear blurry. For example, the DBPN image 200 is less blurry than the non-network-based image 100, but it is still very blurry and inadequate. In fact, many of these network techniques do not work at all on small or edge devices, which do not have enough power capacity and computing capacity to cope with such a heavy load. The experimental results of this comparison are mentioned in Table 1 below.
[0034] Other traditional SR network techniques may provide good quality images in real time, but only on high-end computers. Such techniques require relatively very high power and specialized, large-footprint graphics processing units, such as on desktop computers or servers. For example, such techniques cannot provide high-quality SR images on small mobile or edge devices.
[0035] In general, state-of-the-art SR models (based on deep learning) are very large and computationally intensive. In particular, typical SR neural networks have receptive fields that can be called suspiciously large. As a result, it has been inferred that the neural network is likely to implicitly perform statistical estimation when propagating pixel-related tokens. Typically, only image input is input at the input nodes of the SR NN, where each frame is analyzed. The implicit statistical estimation in the neural network is likely to greatly increase the required size of the neural network, and in turn, also increase the computational load, processing time and power consumption of the neural network. As a result, the resulting sharpness of images obtained in real time is usually insufficient, or impossible to generate in real time. This also leads to uncontrolled randomness in the SR NN, while external, deliberate statistical control can make the process more efficient and perform better, while improving the sharpness or other qualities of the HR image. Note that the terms frame, image and picture are used interchangeably in this article.
[0036] In addition, refer to Figure 3, most of these conventional network techniques rely on a well-known downscaling training protocol, which involves first downscaling the HR image 304 to generate an input downscaled LR image 302 for training. Then, during actual training, the SR neural network parameters are generated by inputting the downscaled LR image 302 into the training network, while the corresponding HR image 304 is set as the output to form a supervised network. It is usually inferred that the real-world input image is sufficiently close to the downscaled training image 302 so that the resulting runtime HR image 306 will have good quality. However, this is incorrect because the downscaled LR training image 302 has different characteristics from the real-world image. Specifically, training the neural network on the corresponding downscaled image 302 ignores the factors of many different image data parameters and conditions, which may introduce various defects to the real-world image, including blur measured by a variable point spread function (PSF), which depends on the camera used (as described in detail below), noise level, tone mapping, etc. In other words, the resulting runtime output HR images 306 generated by using the reduced input images 302 tend to have unnatural, very sharp features, and therefore, conventional CNNs are trained on certain types of overly sharp input images, which causes the CNNs to not generalize well to other lower quality natural (or real-world) images. These reduced images also have artificially low levels of noise due to noise averaging caused by reduction. Thus, these reduced images 302 are not adequately representative of real-world images, and therefore, conventional SR neural networks perform poorly when the network's input is a real-world, relatively low-quality, low-resolution image, rather than an artificial image formed by reducing a high-resolution image. Thus, under this conventional training protocol, the resulting network cannot adapt to changing image conditions and provides low-quality images.
[0037] To address these issues, a super-resolution imaging technique is provided herein that converts LR images into high-quality, relatively and substantially sharper HR images in real time on small or edge devices, and provides much higher efficiency on high-end, larger capacity devices. Here, such small devices may include mobile devices, including tablet devices, smart phones, and other wearable smart devices, while edge devices refer to low-power devices, such as battery-powered devices using about 15 watts or less. These edge devices may include Internet of Things (IoT) devices on home appliances, building systems, vehicles, etc., and they typically provide wireless or wired access to WANs or LANs including the Internet. Otherwise, the disclosed SR methods and systems are not limited to any particular device.
[0038] To achieve these results, the disclosed methods and systems provide two strategies. (1) An efficient neural network (NN) structure operated by an efficient parallel NN pipeline architecture enables real-time application of video SR, at least on low-power small or edge devices. This is achieved by using global statistical properties and user preferences as weights and / or biases of the NN, rather than relying on implicit operation of the NN and adjustment of values to user preferences. (2) The SR NN training protocol improves the quality of the resulting runtime HR images by taking into account local properties of the images, including expected defects in the captured images caused by camera lens changes or camera operation. Any of these strategies can be used alone, or can be arranged to be used together.
[0039] With respect to the disclosed efficient NN and NN architecture, it has been determined that the system can externalize and control image statistics-estimation and user preferences to generate weights and bias values for the processing NN that generates the HR image, rather than implicitly re-estimating the statistics and inputting the user preference values into the input nodes of the processing NN along with the input LR image data. To implement this technique, two parallel processes can be used: one is the pixel processing process of the processing NN, and the other is a control or side process that has a statistics unit to generate (or estimate) statistics based on the LR image, and a weight engine that uses the statistics to generate weights or biases or both, which will be provided to the processing NN. This alone significantly reduces the size of the processing NN in terms of the number of nodes, and in turn reduces the computational load, parameters, and per-pixel operations of the processing NN, thereby enabling real-time SR on small or edge devices.
[0040] In addition, the statistics can be convolved several times in a separate statistics network, or otherwise combined several times, to obtain representative parameter statistics for forming weights and / or biases to be provided to the processing NN to generate global representations, each of which can represent a different calibration of the entire frame or image. Since the video statistics evolve slowly from one frame to another in a video sequence, the operation rate of this global adaptation mechanism is much lower than the frame rate of the processing NN. Thus, this operation only needs to be performed at certain frame intervals rather than at every frame to change the weights and / or bias values.
[0041] The slow evolution (or evolution) of statistics in a video sequence presents several other advantages. Specifically, the visual properties of the output video (e.g., sharpness) can now be adjusted based on user preferences with negligible overhead. In common deep learning systems, this desirable property comes at a significant cost—the neural network must become more powerful to accommodate the additional flexibility. Externalizing this additional adaptability to the control flow eliminates the increased processing load on the NN.
[0042] Additionally, with respect to the statistics themselves, explicitly externalizing adaptability into the control flow, and triggering the control flow much more slowly than the pixel processing flow due to slow evolution, can make it possible to add very large-scale statistics and make them accessible to the processing NN on the pixel processing flow, all without having to repeatedly estimate the statistics within the processing NN. Thus, by adjusting the trigger rate of the control flow, the statistics NN, as well as the weight engine, can be made arbitrarily large (or powerful) without affecting the overall computational cost. In other words, very detailed statistical descriptions are possible, and the control flow can be made more powerful without affecting pixel throughput. If computational requirements increase, the trigger rate can be reduced to compensate. Additionally, this additional capability has the potential to improve the degree of specialization of processing weights, which in turn can (a) improve the quality of the output video, and (b) reduce the load on the processing NN, which is often the bottleneck of the system.
[0043] Combining these two techniques, reducing the processing NN size, and providing control parameters at intervals, can reduce the computational load of SR by as much as two orders of magnitude (see Table 1 below), thereby enabling real-time or near real-time video SR on small or edge devices.
[0044] Turning to the training protocol, in order to improve accuracy and HR image quality, the training protocol involves mimicking the natural image capture process by at least accounting for the naturally occurring blur in the image caused by the camera lens. This involves intentionally blurring the LR image according to the point spread function (PSF), which generalizes better to real-world input videos, thereby greatly improving the sharpness of the HR image. The PSF is a known measure of the performance of an imaging system, and specifically measures how a point of light will look in the image, or in other words, how much the point of light will spread out. Thus, the spread is a measure of how much blur there is at each point given a given lens.
[0045] Thus, in one form, the system modifies the original image by generating an HR training image using a PSF scaled by a stretch factor, and generating a corresponding LR training image using both the lifting ratio and the PSF scaled by the stretch factor. Further details are provided below. Thus, the deep learning model can be trained on inputs similar to real-world images, and therefore performs well on real-world videos. By using these disclosed methods and systems, a relatively sharp, good quality HR image 400 ( Figure 4 ). In fact, such high quality HR images can be produced by the method described in this article, thereby allowing the camera to have a lower quality or lower resolution optical module (with lens, etc.), thereby allowing the size of such optical module to be reduced.
[0046] It will also be understood that the NN structure and hardware architecture disclosed herein can be used for image processing other than super-resolution, and network training can be performed as disclosed herein by taking into account other image conditions in addition to PSF, such as tone mapping, etc.
[0047] refer to Figure 5 , an example process 500 for super-resolution image processing described herein is arranged in accordance with at least some implementations of the present disclosure, and the process is directed to decoupling statistical estimation and / or user preferences from SR neural networks. In the illustrated implementation, the process 500 may include one or more operations, functions, or actions illustrated by one or more of the uniformly numbered operations 502 to 508. As a non-limiting example, the process 500 may be described herein with reference to any of the example image processing systems described herein and in related places.
[0048] Process 500 may include "obtaining at least one lower resolution (LR) image" 502, and according to one example, obtaining a real-world image captured by a camera. This may be a still photo, but the system here is particularly aimed at video sequences of LR images to be converted to HR.
[0049] The process 500 may include "generating at least one high resolution (HR) image, including inputting image data of at least one LR image into at least one super resolution (SR) neural network" 504. The SR neural network may be a processing NN on a pixel processing flow. The neural network is arranged to convert the LR image into a HR image.
[0050] Process 500 may include "separately generating weights or biases or both, including using statistics of the LR image" 506. This operation may be performed rather than having the SR neural network implicitly generate the statistics itself internally (or by treating the statistics as inputs at the input nodes of the SR neural network). Thus, this operation refers to decoupling the control (or configuration) process that performs the statistics estimation and the pixel processing process that operates the processing NN so that the two processes can operate in parallel, so that the generation and use of statistics does not increase the size and computational load of the SR neural network (also referred to as the processing NN). Now, the weights, biases, or both of the processing NN can be based on the statistics, rather than increasing the number of nodes on the processing NN. Decoupling also avoids any significant latency issues. Since the statistics evolve slowly over a sequence of frames, the system can generate the LR image statistics at frame intervals along the video sequence, and then generate the weights or biases or both. According to one example, the interval is 10 frames. In an example, the interval can be uniform, or the interval can be variable. With this arrangement, if the control flow is late in providing new weights or biases for a target frame, the processing NN can switch to using the latest available weights and biases for the target frame without any noticeable degradation in image quality or sharpness. Since the control flow can perform the statistical estimation process in the pixel processing flow at a rate lower than the video frame rate, this can reduce the number of operations per pixel by as much as two orders of magnitude (see Table 1).
[0051] Process 500 may include "Providing weights or biases or both in the SR neural network for generating at least one HR image" 508, wherein the generated weights or biases may then be provided to a processing NN on a pixel processing flow to generate the HR image.
[0052] refer to Figure 6 , an example process 600 for super-resolution image processing described herein is arranged in accordance with at least some implementations of the present disclosure, and the process is directed to generating training images to train a SR neural network. In the illustrated implementation, the process 600 may include one or more operations, functions, or actions illustrated by one or more of the uniformly numbered operations 602 to 608. As a non-limiting example, the process 600 may be described herein with reference to any of the example image processing systems described herein and in related places.
[0053] Process 600 may include "Obtain Initial Training Images" 602. These may be HR training images sized so that decimation of the images achieves desirable smaller HR and LR training image sizes for training the SR processing NN.
[0054] Process 600 may include "generating blurred training images, including intentionally blurring one or more of the initial training images" 604. This involves attempting to simulate at least one real-world image capture condition so that the resulting trained NN generalizes better to real-world images. Here, the process focuses on blur, but the process may additionally modify the initial image data to simulate other camera lens and camera operation effects. Here, this operation includes obtaining one or more point spread function (PSF) values, each representing blur caused by a different lens, and then determining a stretch factor based at least in part on each PSF being used. The stretch factor compensates for decimating the blurred image to a desired HR and LR training image size. The stretch factor can then be used to set the size, shape and / or coefficients of the convolution kernel for traversing the initial training image input in a blurred convolution operation, which will output image data in the form of a blurred image corresponding to the input initial image.
[0055] The SR processing neural network can be trained for a single specific PSF and a specific camera lens. In other alternatives, the initial image is assigned a random PSF from a set or a series of available PSFs, so that the neural network can convert LR images with different blur patterns and thus different PSFs. Other variations are mentioned below.
[0056] Process 600 may include "generating low resolution training images or high resolution training images or both, and including using blurred images" 606. The blurred images will then be decimated to reduce them to the desired HR training image and LR training image sizes.
[0057] Process 600 may include "training a super-resolution neural network to generate high-resolution pixel values using low-resolution training images or high-resolution training images or both" 608. Thus, the LR training images may be input to the SR neural network (or processing NN) to be trained, and the HR training images may be set as supervised network outputs for training. The neural network may be operated during training until sufficient parameters are set for the neural network.
[0058] refer to Figure 7, the image processing system or device 700 may be used to perform training of a super-resolution processing neural network 728, and in particular to prepare images for use as input and supervised output of the processing NN 728. The image processing device 700 intentionally blurs images to better generalize to real-world images captured by non-ideal cameras with imperfections (or variations), at least with possible point spread, and possibly with other variations such as tone mapping, stabilization, etc. The image processing device 700 may receive initial training images 702 that are to be blurred or otherwise modified to better match real-world images, so that the blurred images may be used for training.
[0059] The imaging device 700 has a camera or lens characterization unit 704, which may be a PSF unit here, which provides factors or coefficients for modifying the image 702. In one form, the PSF stretch factors formed by the stretch factor units 706 and 716 are provided to LR and HR main modification (MMOD) units 708 and 718, respectively, which use these factors to modify the image data of the initial training image 702 and may operate the neural network itself. Then, optionally, one or more additional camera characteristic modification units 710 and 720 may be used to further modify the image with factors from other types of camera characteristics, such as tone mapping, stabilization, etc. Thereafter, the decimation units 712 and 722 reduce the size of the blurred image to the desired input HR image resolution and the desired output LR image resolution to form a blurred output HR training image 714 and a blurred input LR training image 724, respectively. The blurred images 714 and 724 may then be provided to the SR NN training unit 726, which uses the images to perform training of the SR processing NN 728. The following process 800 provides details of the operation of the image processing device 700.
[0060] refer to Figure 8 , an example process 800 for super-resolution image processing described herein is arranged in accordance with at least some implementations of the present disclosure, and the process is particularly directed to generating training image data to train an SR neural network. In the illustrated implementation, the process 800 may include one or more operations, functions, or actions illustrated by one or more of the uniformly numbered operations 802 to 834. As a non-limiting example, the process 800 may be described herein with reference to any of the example image processing systems described herein and in related places.
[0061] As previously described, this process 800 includes generating image pairs, which include input images and output images, whose properties are better correlated (or correspond) to real-world, camera-captured images that are likely to be used as input to a learning-based SR model during runtime, and are thus more likely to generalize to real-world videos. These image pairs are generated by digitally simulating the analog transformations of light and camera operation when the camera captures the image. This includes simulating the characteristics of the camera lens or camera capture operation, which result in certain image characteristics and, in turn, affect the raw image data on the digital camera. One of the most notable characteristics is the blur of the image as measured by the point spread function (PSF), and the examples here intentionally blur the image to simulate the blur on a real camera-captured image.
[0062] Process 800 may include "Obtain HR video frame of image data" 802, such as initial image 702. The size of the initial image may depend on predetermined camera lens characteristic factors (e.g., blur stretch factors described below) and the reduction ratio being used, so that subsequent decimation of the image will produce input LR training images and output HR training images of desired sizes to train the SR neural network 728. Thus, for example, when the stretch or scaling factor S is 4, the initial image may be 400x400 pixels to obtain an output HR training image of 100x100 pixels. As explained below, additional reduction ratios R are used to obtain input LR image sizes, and in this example, say R=2, so that after decimation, the input LR training image is 50x50 pixels. It will be understood that many other sizes may be used as desired. Details will be explained below. This may also include training the SR neural network 728 with multiple different image sizes in the same training session.
[0063] The initial images 702 can be various real-world images captured by a camera with different contents of any known visible subject. In some form, known neural network picture collections are specifically used to collect pictures to train image neural networks, and these pictures include as many different visible views of the world as possible that may be experienced by a person or camera. Since these are pictures captured by the camera, the initial images include lens or camera operating characteristics, such as blur from the PSF. However, it has been found that the deliberate image modification performed here in the present system and method to account for such characteristics is so important and influential that naturally occurring image data changes on the initial images, such as blur, can be considered negligible.
[0064] In addition, the intentional camera characteristics taken into account here apply to all channels, because all light is affected by the lens, such as by blur. However, in the present system and method, only the luminance channel is addressed, because the human eye is much more sensitive to luminance than to color. Thus, in this example, the system only uses the luminance (or brightness or brightness) data of the original image. In an alternative form, the color channel can also be blurred.
[0065] The process 800 may include "Setting one or more HR blur factors" 804, specifically a PSF-based stretch factor S, which is used to determine coefficients and their positions to blur the image data of the initial image. According to one form, when a convolution operation is used to perform blurring of the initial image, the stretch factor S sets the size and shape of the convolution kernel (or filter). The convolution may be operated by the main modifier (MMOD) unit 708 or 718 to perform blurring.
[0066] The stretch factor S is used to compensate for the subsequent decimation of the blurred image to achieve the HR and LR training image size. Specifically, decimation can be considered as another form of camera characteristic simulation by sampling the pixel image data by selecting and discarding pixels to simulate the sampling performed by the camera sensor. This is in contrast to the interpolation or averaging of pixel values to change the image size performed by typical lifting or downscaling algorithms. This adds another feature that mimics real-world image capture to increase the generalization range of the SR neural network during runtime.
[0067] When blurring the initial image, the PSF curve or spread alone is usually not large enough to be adequately represented on the decimated image. Too many blurred pixel locations will be discarded during decimation to form the training image. Thus, the PSF spread is first spatially stretched in the pixel region by a stretch factor S to form the convolution kernel prior to convolution so that the blurred, subsequently decimated image will still have a sufficient amount of blurred pixel image data to adequately represent the blur.
[0068] Thus, this operation may include "determine PSF-based HR blur stretch factor S1" 806. Thus, as a preliminary operation to setting up the image blur unit 708 or 718, the system may obtain a specific PSF to be used to calculate the stretch factor S. In one form, the SR NN 728 may be trained for a specific camera with a specific lens arrangement, so that only one specific PSF value needs to be obtained. In this case, the SR model or neural network may be trained to expect images from a specific imaging device.
[0069] In an alternative form, the SR neural network may be more adaptable to several different cameras and may be trained to convert images into HRs for multiple cameras and thus associated with multiple different PSFs. In these cases, the PSFs of several camera lens devices may be represented by using PSFs within a certain PSF range or by using PSFs from a set of specific known PSF values associated with a camera lens or a specific camera. The PSF may also be set to a non-real value, or include a non-real value, to include a wider variety of PSFs. It should also be noted that in addition to the shape and size of the lens or lens combination, the PSF may also vary or remain constant while other camera parameters are changed. According to one example, these parameters may be related to different lighting conditions and may include changes in noise or tone mapping. Although these changes may be controlled to provide a certain number of input images for certain changes, according to another approach, these changes are random so that the PSF can be better detected in random samples.
[0070] In an alternative with multiple PSF representations, a random PSF within the mentioned PSF range or set may be used for each image pair, and in this case the SR model will be trained to various real-world images expected to have different PSFs. Thus, each image pair may be based on a different PSF. As another example, the same initial image may be added to the training set multiple times, so that each blurred image pair from the same initial image is based on a different PSF. Different lighting conditions may also be a variable within the input images.
[0071] In addition to those mentioned above, another way to describe the PSF is that the PSF is a quantitative measure of how each infinitesimal point of light in the scene spreads across the sensor plane. Point spread is often associated with a Gaussian curve, so that the pattern of image pixels that spread (or smear) across the image is typically circular, although typically non-uniform, with the intensity decreasing as the distance from the center or strongest pixel increases. However, the pattern need not always take the form of a perfect circle on the image, and the spread (or smear) may have many different shapes and areas. The actual blur convolution kernel (or filter) used may be rectangular, circular, or other desired shape, for example to match a top view of the actual PSF curve. For the example currently being worked on, assume that the PSF has a characteristic blur covering a 3x3 pixel area, where a point light source in the real world is smeared onto a 3x3 area of the camera sensor. This is a very simplified version for illustration.
[0072] Now, to determine the free parameter or stretch factor S1 for the blurred HR image, the X and Y dimensions of the PSF (or PSF pattern) are multiplied by the stretch factor S1 to generate a pattern of stretched size to serve as the CNN kernel (or filter). The larger the factor S (or S1 here), the more gradual the gradient of the expansion, the closer the values will be to the PSF curve, and therefore the closer the simulation will be to the actual image capture. If S is too high or too low, the resulting image will become too unrealistic and proper generalization will be reduced. Also, as mentioned earlier, S is factored into the decimation scale. Thus, the initial image dimensions (X and Y) will be divided by S to obtain the decimated blurred HR image dimensions, as explained in more detail below.
[0073] To simplify, the correlation between PSF values and stretch factor S can be determined by using an input image of, say, 400x400 pixels, which is very dark except for one bright spot. This spot is smeared using the spatially stretched PSF. Assume the original smear was 3x3 pixels, and that S=4 is being tested. The stretched smear pattern is now 12x12 pixels (3x3 original smear multiplied by the stretch factor S). The bright spot is now considered to be stretched to 12x12. The system is then tested, and S is varied until the results are adequate. Since the PSF defines how an infinitesimal point in the scene spreads out on the imaging plane (sensor), it is defined as the percentage of light energy as a function of radius (or in other words, distance from the center), and so the PSF can be stretched by an amount to scale this radius so that the light or blur reaches farther outward and the smear is "severer."
[0074] Once the blur or stretch factor S1 is set, the process 800 may include "Generate Blur Convolution Kernel" 808. Here, given the point spread function PSF(x, y) and the input image I in , the PSF is stretched to create a candidate convolution kernel:
[0075] h(x,y)=PSF(x / s,y / s) (1)
[0076] Where h() is the kernel, (x,y) is the pixel location on the image, PSF() is the PSF function or curve used to gradient or blur the light, and s is the stretch factor. This will set the maximum intensity according to the PSF, and the intensity level will decrease with each pixel distance from the maximum intensity pixel within the kernel according to the curve of the PSF. The kernel can then be tested to determine generalization for training and whether sufficient sharpness has been achieved. This process is repeated for each PSF being used and each expected S value for that PSF. This is performed separately before processing NN training to determine the correct PSF-stretch factor association, kernel size and kernel coefficients that will be used to blur the initial image and is performed by stretch factor units 706 and 716. Once the stretch factor S1 and the blur convolution kernel size, blur kernel shape and size, and blur kernel coefficient amplitudes are set within the blur kernel by the stretch factor unit 706, the main modifier (MMOD) unit 708 is ready to perform a blur convolution to blur the initial image to form an HR training image (the operation on the LR image is described later below).
[0077] The process 800 may then include "Generate Initial Blurred HR Image" 810. Here, the kernel may be moved on the image with a stride determined during training, such as a stride of 1, and may be moved, for example, in a raster order, so that each pixel position in the image is modified by the blurred convolution, and as previously described may be a single convolutional layer and operated by the MMOD unit 708. The kernel may be placed on the image by placing each current pixel to be modified at the center position of the kernel or at other specific pixel positions within the kernel.
[0078] This operation may then involve "Convolve the data with a filter based on S1" 812, where each coefficient in the kernel is multiplied by the corresponding original image pixel value. Thus, to blur the image by convolving with h:
[0079] I conv (x,y)=(I in *h)(x,y)=∑ m,n I in (nx,my)h(n,m) (2)
[0080] Among them I conv( ) is the convolved pixel of the image data formed on the convolved and now blurred layer or blurred image, and (n,m) is the coefficient position within the kernel. This formula determines the product between the pixel position on the image and the corresponding kernel coefficient. Each product is then subsequently summed to form a new value for the current pixel. This process is repeated for each pixel location (x,y) to account for the blur into the individual pixel values. It will be noted that the convolution formula (2) and kernel treat each pixel location independently and use the initial image data so that the resulting blurred value of one pixel location is not affected by the resulting blurred pixel value of another pixel location.
[0081] It should also be understood that this convolution operation does not change the resolution of the original image. Thus, a 400x400 pixel original image will still be 400x400 pixels on the blurred image just after the convolution operation. In addition, in the event that the complete kernel does not fit on the image, the pixel positions at the outer edges of the original image will have missing kernel-sized areas filled in by known techniques such as mirroring or other techniques.
[0082] Process 800 may include "Modify data according to other camera lens characteristics" 814. To simulate more realistic image capture, the camera characteristic modification unit 710 (and the unit 720 for blurring the LR image as described below) may modify the blurred image data more. Thus, while blurring is the minimum camera characteristic accounting transformation that can be performed, further image data modifications may be performed for other common image transformations, such as with tone mapping, additive sensor noise, stabilization, etc. These are known image modification algorithms. This will further increase the potential for generalization.
[0083] Process 800 may include "decimating the image" 816, which is performed, for example, by decimation unit 712 and digitally simulates the sampling performed by the camera sensor as described above. The blurred image is then decimated by the HR stretch factor S1. According to one example, the blurred image is then decimated according to the following formula:
[0084] I down (x,y)=I conv (x*ds,y*ds) (3)
[0085] Specifically, decimation is the removal of pixels without interpolation, or in other words, the removal of pixels without combining pixel values or generating pixel data combinations (e.g., averaging). In one form, decimation avoids any interpolation or pixel value combination. Decimation retains pixel values at intervals set by the stretch factor S1, and discards any other pixels. This is performed in both the X and Y directions, although different decimation modes can be used.
[0086] Thus, according to one example, when S (or S1) is 4, the decimation will retain only one out of every four pixels in each dimension. According to one example, if the original and blurred images are 400x400 pixels and S=1, then the decimation will generate a HR training image of 100x100 pixels. The resulting image is then "S" times smaller than the original image. The blurring process can be expressed as:
[0087] WxH→à CONV with S→à W / S x H / S (4)
[0088] Where W x H is the width and height of the original image, àCONV is the convolved or blurred image, S is the stretch factor described above, and àW / S x H / S is the resulting decimated image. The resulting image is similar to an image captured with a lens characterized by a particular PSF used to blur the image. The blurred HR training image 714 is then provided by placing it in the image training set to be used to train the SR NN 728.
[0089] Turning now to the generation of blurred LR training images, process 800 may include "setting one or more LR blur factors" 818, and this process is similar to operation 804 described above for blurred HR training images, except that here this operation may also include "obtaining a resolution scaling ratio R" 820. Specifically, in order to prepare image pairs for training a learning-based SR model (or NN 728), two image capture simulations are performed with the same PSF but with different stretch S values. As described above for the blurred HR training images, S1 is a direct extension of the PSF. However, for the blurred LR training images, the stretch factor S2 is based on the PSF and the lifting ratio R as part of the SR model. Specifically, the output of the blur operation is two representations of the same scene at different scales, with a ratio R between the two scales, where the smaller of the two images serves as the input during training, and the larger one serves as the HR "label" or output image to form a supervised network. Thus, the SR NN is being trained to lift the LR image by a factor R.
[0090] Thus, process 800 may include "Determine PSF-Based Blur Stretch Factor S2" 822, where S2 = S L x R, and according to an example form, S L = S1. Otherwise, the PSF to be used is determined, and then a stretch factor S2 is determined based on the PSF, which is the same as or similar to S1. According to one form, the ratio R can be 2, but in other examples it is often 2 to 4, with up to 8 also being used, but there is no limitation here, except that in one example R is a rational number greater than 1.
[0091] Process 800 may include “generate blur convolution kernel” 824, and once S2 is calculated, this operation is the same or similar to the HR blur convolution kernel formation operation 808, except that here it may be performed by the LR stretch factor unit 716.
[0092] Process 800 may include "Generate initial blurred LR image" 826, and this operation includes "Convolve data using S2" 828. These operations run the blurred convolution and apply the kernel to the initial image data, similar to HR operations 810 and 812, except that here these operations may be performed by MMOD 718.
[0093] Process 800 may include "modify data according to other camera lens characteristics" 830, and this may be performed by camera characteristics modifier unit 720, like unit 710 already described above.
[0094] Once the blurred image is generated (and with or without other modifications as described above), process 800 may then include "decimating the image" 832, and may be performed by decimation unit 722. The lifting ratio R is also used to generate an LR image that is smaller than both the initial image and the blurred HR training image. Here, the LR blur process may be represented as:
[0095] WxH→à CONV with S2→à W / (S2) x H / (S2) (5)
[0096] Referring to the example of continuing from the blurred HR training image generation, assume S1 = 4, R = 2, and S2 = (4x2) = 8. Thus, if the initial image is 400x400 pixels, and the blurred HR training image is 100x100 pixels, then the blurred LR training image will be 50x50 when S2 = 8. The result is a blurred LR training image 724, which can be used to train the NN 728, together with the corresponding blurred HR training image 714 as part of an image pair.
[0097] The blurred LR training images may be added to the training set, stored in memory, and matched against the blurred HR training images to form image pairs.
[0098] The process 800 may include "Training the SR processing NN using the blurred LR images as input and the blurred HR images as output" 834, and may be performed by the SR NN training unit 726. The blurred LR and HR training image pairs 714 and 724 may then be used to train the SR NN 728 using the blurred LR training images as input and the blurred HR training images as supervised output. Once sufficiently trained, the SR NN will provide good generalization and may be used during runtime to convert real-world LR images into HR images with very good high-quality images with very good sharpness.
[0099] refer to Fig. 9 , an image processing system or device 900 for performing super-resolution has two parallel processes, both of which receive the same input frame 901: a pixel processing process unit 902 and a control (or side or configuration) process unit 904. Each process unit 902 and 904 operates in parallel so that the pixel processing process unit 902 does not need to pause and wait for a large amount of time (if any) to retrieve data from the control process unit 904. The pixel processing process unit 902 has a processing neural network (NN) 906, which may have been trained by any training method described herein, and the control process 904 has a statistics unit 908 that generates (or estimates) a statistical representation of the image, and a weight engine 910 that uses the statistics. The SR scale R unit 912 provides the weight engine 910 with a scaling ratio between the input and output images, and the user selection control unit 914 provides the user selected value to the weight engine 910. The weight engine uses these values to generate weights and / or biases for the processing NN 906.
[0100] The parallel operation of the two processes 902 and 904 is performed by hardware or firmware depending on the target platform. Thus, according to one example, each process may have its own image signal processor (ISP), graphics accelerator, or graphical processing unit (GPU), including any multiply-accumulate (MAC) circuits, etc. In this case, the control process 904 can be calculated on a separate low-power computing circuit (or (one or more) processors), freeing up the resources of the main processing device (e.g., GPU or HW accelerator). This may even include, for example, executing the control process on a separate remote device such as a server via a wireless connection via the Internet. In another form, the two processes may share the same hardware, for example, by context switching.
[0101] Due to the relatively small size of the NN as described below ( Fig.11), the processing time should be shorter than the time interval between frames, which should also leave enough time for control flow operations.
[0102] Even if the pixel processing flow and the control flow operate in parallel, latency in the output of the control flow is not a serious problem because the statistics evolve slowly over many frames in a video sequence. In other words, image content typically changes relatively slowly from frame to frame (except at scene changes) relative to the speed at which the control flow operates here, and thus, the same image statistics may remain associated with a long string of consecutive frames (whether 10 frames or 20 frames or some other number of frames, depending on the content, just to give a few random examples). Thus, if the output of the control flow arrives too late to be applied to the target frame corresponding to the input LR frames used to form the statistics, then simply applying the output of the control flow to the next available frame within a certain number of frames along the video sequence from the target frame, such as within 10 frames in one example, will not have a significant impact on the system. Operation of the system or device 900 is provided using the following process 1000.
[0103] refer to Fig.10 , an example process 1000 for super-resolution image processing described herein is arranged in accordance with at least some implementations of the present disclosure. In the illustrated implementation, process 1000 may include one or more operations, functions, or actions illustrated by one or more of uniformly numbered operations 1002 to 1032. As a non-limiting example, process 800 may be described herein with reference to any of the example image processing systems described herein and in related places.
[0104] Process 1000 may include "receiving LR video frames" 1002. In one example, one or more cameras or image sensors may provide image data of a captured LR input image 901, and the data is provided to one or more processors as described using device 900. The image data may be obtained directly from the sensor, or may be retrieved from any memory or storage device that holds the image data. This operation may also include any pre-processing of the image data to prepare the data for performing super-resolution, and may include operations other than the generation of statistics for super-resolution described below. It will be appreciated that instead of video frames, the following process may instead be applied to still photographs of individuals.
[0105] The input LR images 901 to be converted to HR can be real-world images captured by one or more cameras. The LR images may include an expected point spread, the size and magnitude of which depends on the camera lens arrangement and camera operation, as described herein. These images can be of any size, especially those image sizes used for training SR NN training images, and according to one example, for 1080p displays.
[0106] As previously described, two separate parallel flows are performed so that the statistical data generation does not become a bottleneck in the neural network processing that generates HR values for the input LR images. Thus, the decoupling of statistical estimation from pixel processing can begin by providing LR images to both the pixel processing flow unit 902 and the control flow unit 904. The pixel processing flow 1004 processes the input pixels to produce a super-resolved (or upscaled) video frame (or output HR image).
[0107] The pixel processing flow operation 1004 may include "input LR image data for the frame into the processing NN" 1006, which may be the SR processing NN 906. This may include retrieving the LR image data from memory or one or more camera sensors and placing the image data in an input buffer of the processing NN, which will then be placed in registers of the ISP, GPU, and more specifically, the MAC. The LR image data may be placed in the buffer and, in turn, placed in the input node, one vector, matrix, or tensor at a time.
[0108] According to one example form, the input to the statistical CNN is an image of size W0×H0×C0, where W is the width, H is the height, and C is the number of channels. Initially, C0 may be three for the three color channels of an RGB image.
[0109] There are no particular restrictions on the arrangement for organizing the input data at the input nodes of the processing NN, except that the input LR image has not been modified by the SR statistical representations (e.g., global weighted values) generated by the control flow specifically for super-resolution. These statistical values are also not used as inputs at the input nodes of the processing NN. Similarly, when user preferences are obtained as described below, the user preferences are not used to modify the image data and are not input to the input nodes of the processing NN. This externalization of data forces the neural network to avoid additional and unnecessary network operations that are directed to adapting the processing to large-scale statistical properties of the signal or user preferences, and are repeated for every pixel in every frame. This thereby produces a small neural network with a relatively small computational load, as well as small power consumption, which greatly increases the speed of HR image generation, thereby enabling real-time (or near real-time) operation on small or edge devices.
[0110] The process 1000 then includes obtaining (1008, 1010) weights and / or biases from the control flow 1014. Therefore, before describing these operations, the parallel control (or configuration) flow 1004 will first be described.
[0111] Process 1000 may include "Execute Control Flow" 1014. The control flow may be considered a configuration flow that configures the processing NN. This refers to the control flow providing parameters, such as weights, in the form of convolution kernels to traverse over the LR images input to the processing NN. This may also refer to bias values provided to the processing NN. Thus, the control flow provides parameters so that the processing NN can be configured and so that the processing NN is adapted to the current statistical properties of the input LR images and / or user preferences.
[0112] Process 1000 may include "generate statistics" 1016. In control flow 904, the input video frame is fed to an example statistics NN 1100 (eg, which may be operated by statistics unit 908). Fig.11 ), the statistics unit 908 may be arranged to output an encoding of full frame statistics. Specifically, the weights and biases used to process pixels are highly specialized for the current input statistics and / or user preferences. Thus, the processing NN can focus only or primarily on local image features and require far fewer resources than traditional CNNs for image and / or video super-resolution. Separately, the statistics unit 910 generates global statistics, and in particular, generates a vector representing a global weighted average of the entire image. This results in an advantageous situation in which the local processing engine (processing NN) is able to access external global information that would otherwise be unavailable, and this results in higher quality processing.
[0113] refer to Fig.11 , the statistics unit 908 can operate a statistical neural network (NN) 1100, which can be a CNN. Thus, the statistics generation 1016 may include a "convolved image value" 1018. The arrangement of the NN 1100 produces a global weighted average vector representation that characterizes the entire frame. The input LR image is propagated through several convolutional layers from layer A (1102) to layer E2 (1114) that are uniformly numbered. According to one approach, each layer performs a convolution. This may be a multiplication with a filter (or kernel) that moves along the image, where at each position, the products are summed to have a single convolution value. The single convolution value replaces the current image value or the last convolution value, thereby maintaining the array size, or the convolution value replaces all element pixel positions within the kernel, thereby reducing the array size of the next layer.
[0114] Each layer may have an activation function, such as a rectified linear unit (ReLU) or a sigmoid. The weights and biases of the activation functions in a statistical CNN may be determined through training. The size of each layer, the size and stride of the kernel, and other specifications of each layer may be set and refined during the training of the system. According to an example form, for an image of 50X50 pixels, the kernel will be 3x3 pixels.
[0115] Then, process 1000 may include "generate weight array" 1020. Specifically, the topology is divided into two branches, including a weight branch with layers D1 (1108) and E1 (1112), and the other branch is a value branch, which continues convolution like the earlier layers A to C (1102 to 1106). The design of the weight layer is the same as the value layer, but is later used as a weight.
[0116] Two parallel branches, each of which is a small feed-forward topology, maintain the same output size as the other branches. At the output of the last layers E1 (1112) and E2 (1114), the end of the value branch E2 can specify an image (or array) V(x, y, c), where x is in the range 1, ..., W final In; y is in the range 1,…,H final and c is in the range 1,…,C final and among them W final , H final and C final It is not necessary to have the same magnitude as the initial W0, H0 and C0. W and H can vary from layer to layer in the statistical CNN 1100 as needed and are determined through experiments and training.
[0117] In addition, although the input LR image can be input to the statistical NN in a single channel (e.g., a channel of brightness values), it is determined (or assumed) that the statistical network implicitly develops multiple channels C final , where each channel appears to encode a different global property of the image, which is used to calibrate the processing NN, and the result is used as one of the outputs of the statistical network and, in turn, as an element on the output statistics vector. One or more channels may be a measure of the noise level, one or more other channels may be a measure of the width of the PSF, another one or more channels may be provided for tone mapping, etc. Thus, according to one possible example, the resulting global weight value vector S(c) may contain the calibration channel C final=8 elements, each element being a measurement accumulated over the entire frame. Many other numbers of channels may be provided instead. Thus, the contents of the statistical vector may be determined during training of the entire system, based on which dynamic characteristics are changed when training the system including the processing NN, where the system is trained by changing the PSF (as described above) and the noise level. The system "learns" that this is global information that is being used to effectively configure the processing NN (incorporated into the weights and / or biases to be provided to the processing NN). Thus, the statistical NN may be arranged by training to implicitly provide a certain number of final channels, each with a specific calibration type. A sufficient number of output nodes may be determined by trial and error, each of which is most likely to be the output of an implicit channel. Here, it has been determined that eight outputs are sufficient.
[0118] At the end of the weight branch at the output of layer E1 (1112), the output image (or array) W(x, y, c) is the same size as the output of the value branch. In this example, if C final =8, the result is eight pairs of arrays of the same size (values and weights).
[0119] Then, the process 1000 includes "generate global weighted value" 1022. Now, the last array (or feature map) from the last value convolution layer E2 (1114) is weighted by the weight array from the last weight layer E1 (1112). Specifically, each pair of two-dimensional value arrays including the last weight layer E1 (1112) array and the last value layer E2 (1114) array is combined to produce a single "global" value that characterizes the entire frame. In one form, this combination refers to a weighted average according to the following formula:
[0120] S(c)=sum xy (V(x,y,c)*W(x,y,c)) / sum xy (W(x,y,c)) (6)
[0121] Wherein, each value V is multiplied by its corresponding weight W. Then, all products are summed over all pixels (or all elements in the final array). This sum is then divided by the sum of all weights in the weight array to calculate the weighted average. When W and V represent the entire image, this is the global weighted average. Note that, in an example form, each channel can be averaged separately, so that the output is an independent average of each channel. These channels together form the output global weighted average vector S(c).
[0122] According to other alternatives, the average can be calculated in other ways, wherein all elements in the array are averaged to obtain a single value representing (or characterizing) the array. For example, in this method, all elements of the value array from layer E2 (1114) are averaged to generate a single representative value. The weight array from the last weight layer E1 (1112) is similarly averaged to generate a single weight from the elements of the last weight array. Then, the single value is multiplied by the single weight to generate a global weighted average of the pair of arrays. In addition, it will be understood that other combinations other than averages, such as modulus, median, etc., or other algorithms, such as interpolation, can be used here, whether between elements in the same array (value or weight array or both), or between corresponding elements from two arrays (values and weights). Many variations are contemplated.
[0123] According to one example, a pair of arrays is generated for each channel C. Thus, when C final =8, the statistical CNN can calculate eight weighted global values for the output global weight value vector. Thus, in this example, the entire frame can be characterized by eight numbers (or other desired numbers as described above), each of which is determined from the entire frame. Since the resulting global vector characterizes the entire video frame or scene, this operation is an encoding of the statistical properties of the frame; it is not a map of the value of each pixel (or a map of groups of pixels). It will be appreciated that other numbers of pairs can be used, depending on which calibrations are desired as described above.
[0124] Once the statistics vector is generated, the process 1000 may include "generating weights and / or biases" 1024, and for example, performed by the weight engine 910. In other words, one role of the weight engine is to provide a convolution kernel of biases and / or weights to be used by the processing NN to process the video frame. In addition to the statistics, the weight engine may also receive (or have) a boost ratio R from the SR scale R unit 912, and receive (or have) a user selection value from the user selection control unit 914.
[0125] When so provided, this operation may include "get user selection data" 1026. The user selection value may represent a desired sharpness value set automatically or manually set by a user, for example, on a menu or interface. The sharpness value may be one from a range of expected sharpness values. The user selection control unit 914 may provide a single value from the range.
[0126] Process 1000 may include "obtain weight kernel" 1028, where initial weights are generated by training the NN using known methods. In one approach, process 1000 may also include "generate combination coefficients" 1030 to account for various parameters and various parameter levels. Specifically, one way to generate a convolution weight kernel is to combine several kernels based on statistics and / or user preferences. For a simplified example, assume that the statistics are encoded as a single current statistic (alpha) α in the range of 0 to 1 and can be one of the statistical value elements from the statistics vector output from the statistics unit. Then, the combined kernel can be calculated as:
[0127] ker process =α·ker1+(1-α)·ker2 (7)
[0128] Where the "·" operation is a dot product, because ker1 and ker2 are two alternative convolution kernels, each of which is a matrix of weights, and each of which is generated by using different kinds of parameters or parameter values in different runs when training the NN. Here, the statistics "select" which option is appropriate for the current input image. Thus, for example, if α is close to zero (e.g., low noise), the system will use kernel 2 more, while if α is high (e.g., high noise) and close to 1, the system will use kernel 1 more (of course, there are combinations in between).
[0129] Likewise, equation (7) is simplified for a single statistical value rather than a full vector of values as output from the statistical unit.
[0130] However, in the current process 1000, as described above, a global weight value vector of statistical data is provided to the weight engine 910. The weight engine 910 may have a large number of convolution kernels ker ij , and can operate a fully connected neural network (FCNN) that controls the combination of kernels. The statistical vector and user preference values are the inputs of the FCNN. The output of the FCNN is the coefficients (α) for the linear combination of convolution kernels.
[0131] In the example formula below, coefficient α = FCNN (statistical vector, user preference vector):
[0132] ker j =∑ i α ij ker ij (8)
[0133] where i is the frame number and the mixed convolution kernel ker jIt is then used as the weight in the jth convolution of the processing NN. The convolution kernel ker ij It can be the trained initial parameters (weights) of the deep learning system and is generated when training the NN, as described above. However, here, many kernels can be combined, each based on a different combination of parameters and parameter values.
[0134] Process 1000 may include "generate bias" 1032. A similar mechanism may also generate bias values. Specifically, the bias is generated by obtaining a previous bias and an expected bias, and then determining a combination coefficient in the same or similar manner as a weight kernel. The initial bias value is also set by training.
[0135] It will be appreciated that the control flow can generate the weight kernels alone, the biases alone, or both.
[0136] Thereafter, returning to the pixel processing flow 1004, the process 1000 may include the operations of "obtaining weights from control flow" 1008 and "obtaining biases from control flow" 1010. The information bandwidth between these two processes is considered to be negligible. As previously described, the slow evolution of statistical properties in the video stream is exploited to perform the estimation process at a rate lower than the video frame rate, thereby reducing the number of operations per pixel. This enables the system to trigger a configuration or control flow at a rate lower than the frame rate (e.g., every 500-1000 milliseconds). Any expected delay in the calculation of the configuration can be tolerated without having much impact on the output video (again due to the slow evolution of statistical data). According to one approach, the control flow can be triggered at intervals of a lower rate (lower than the frame rate), such as every 10 frames.
[0137] The decoupling of configuration generation from the pixel processing flow, and hence from the frame rate, and the slow evolution of the statistics, also make latency in the control flow a very minor issue. Specifically, assume that the system triggers the control flow every 10 frames. Then, the parameters of the processing NN are updated from the control flow at frames 1, 11, 21, ... and so on. If the control flow has a latency of two frames, i.e., the output of the control flow is ready two frames later than the corresponding target frame, then the processing NN in the pixel processing flow is updated anyway, but now the updating is just switched to frames 3, 13, 23, and so on. In other words, the frame-to-frame differences for these 10 frames can be considered negligible for super-resolution, and in most cases the differences in the output video will be negligible. For SR operations and real-world images, this slow evolution of the image content is expected to occur up to some maximum number of frames, e.g. for typical frame rates such as 30fps or 60fps, and this can be determined experimentally and of course can also vary depending on the video content.
[0138] Note that when a scene change occurs along the video while the control flow configuration values are delayed, the pixel processing flow can be operated in several different ways. First, the scene change can be ignored and the first few frames of the new scene (1 to 5 frames according to a random example) can use incorrect weights and / or biases, for example, which are likely not noticed by the user viewing the image. Otherwise, a scene change detection mechanism can be provided that compares consecutive frames and detects a scene change when the comparison fails to meet a criterion, such as the maximum acceptable full image sum of absolute difference (SAD). According to this option, the pixel processing flow can be paused, and in this example only for this situation, to allow the control flow to catch up and provide corresponding weights and biases for the frames of the new scene. According to another option, when a new scene is detected, the control flow does not provide updated weights and / or biases, and the processing NN will continue to operate subsequent frames with the last received weights and biases until the control flow has a chance to process the weights and / or biases for the frames of the new scene. Many other algorithms can be used to detect scene changes.
[0139] refer to Fig.12 , process 1000 may include "generating HR pixel values from LR images" 1012. This involves running the processing NN using weights and / or bias values from the control flow. This may involve placing the weights and / or biases in a buffer that can be accessed by the processing NN for use. The network topology (or architecture) of the processing NN 1200, which may be similar to the processing NN 906 and may be a CNN, may have a lifting layer 1202, layers 1 to N (1204 to 1208), a scrambling layer or unit 1210, and an adder 1212. Since the weights and biases of the processing NN are highly customized according to the current imaging conditions, the processing NN 1200 can have a very low capacity and still perform efficiently. An example processing NN 1200 can be a simple feed-forward topology with two characteristics commonly found in topologies for SR. The residual connection in the form of the lifting layer 1202 uses bicubic interpolation to perform lifting. In addition, CNN layers 1 to 4 are provided and configured through experiments and training. The last layer 4 provides a size of W×H×R 2 , where R is the lifting ratio and W x H is the size of the image (or array). Then, the pixel shuffling layer or operation 1210 operates on the tensor (W×H×R 2 ) are scrambled to produce a single channel of size (R·W)x(R·H). The convolved values are then added to the corresponding values of the lifted image from the lifting unit 1202 to produce the output HR image data.
[0140] It will be appreciated that all operations of the control flow and pixel processing flow (processing NN) can be made differentiable for all parameters. This includes the example statistical NN and example weight engine FCNN described above. Then in one form, the entire system can be trained end-to-end using a common gradient descent algorithm for training deep learning systems.
[0141] In addition, the framework detailed above that decouples the adaptability of the control flow from the pixel processing flow (and processing NN) can be applied in a wide range of real-time video enhancement applications. For example, this framework can be applied to video denoising, local tone mapping, stabilization, etc.
[0142] The disclosed SR method and system are compared with a conventional video SR architecture with a representative SR model. The output video is 1080p at 30 fps with 4x upscaling (i.e., input frame size is 270×320). The numbers given below are for the generated Figure 2 and Figure 4 The results for each method variant are given in .
[0143] Table 1. Comparison between public SR methods and known SR models
[0144]
[0145]
[0146] (*) Quality is measured by a subjective mean opinion score (MOS), which is not possible to measure objective quality when upscaling real-world videos (no “ground truth” is available).
[0147] In addition, the following may be performed in response to instructions provided by one or more computer program products: Figure 5-Figure 6 , Figure 8 and Fig.10Any one or more operations of the process in. Such a program product may include a signal bearing medium providing instructions, which, when executed by, for example, a processor, may provide the functions described herein. A computer program product may be provided in one or more machine-readable media in any form. Thus, for example, a processor including one or more processor cores may engage in one or more operations of the example process here in response to a program code and / or an instruction or an instruction set delivered to the processor by one or more computers or machine-readable media. In general, a machine-readable medium may deliver software in the form of a program code and / or an instruction or an instruction set, which may enable any device and / or system to operate as described herein. A machine or computer-readable medium may be a non-transient article or medium, such as a non-transient computer-readable medium, and may be used in conjunction with any of the examples or other examples mentioned above, except that it does not include the transient signal itself. It does include those elements, such as RAM, etc., that can temporarily store data in a "transient" manner in addition to the signal itself.
[0148] As used in any implementation described herein, the term "module" refers to any combination of software logic and / or firmware logic configured to provide the functionality described herein. Software may be embodied as a software package, code and / or instruction set, and / or firmware that stores instructions executed by programmable circuits. Modules may collectively or individually be embodied as an implementation that is part of a larger system, such as an integrated circuit (IC), a system on a chip (SoC), and the like.
[0149] As used in any implementation described herein, the term "logic unit" refers to any combination of firmware logic and / or hardware logic configured to provide the functionality described herein. As used in any implementation described herein, "hardware" may, for example, include hardwired circuits, programmable circuits, state machine circuits, and / or firmware storing instructions executed by programmable circuits, alone or in any combination. Logic units may be collectively or individually embodied as circuits that form part of a larger system, such as an integrated circuit (IC), a system on a chip (SoC), and the like. For example, logic circuits may be embodied in logic circuits for implementing the systems discussed herein via firmware or hardware. In addition, it will be appreciated by those of ordinary skill in the art that operations performed by hardware and / or firmware may also utilize a portion of software to implement the functionality of a logic unit.
[0150] As used in any implementation described herein, the terms "engine" and / or "component" may refer to a module or to a logical unit, as these terms are described above. Thus, the terms "engine" and / or "component" may refer to any combination of software logic, firmware logic, and / or hardware logic configured to provide the functionality described herein. For example, one of ordinary skill in the art will appreciate that operations performed by hardware and / or firmware may be implemented via software modules, which may be embodied as software packages, codes, and / or instruction sets, and will also appreciate that a logical unit may also utilize a portion of software to implement its functionality.
[0151] refer to Fig.13 , an example image processing system 1300 is arranged in accordance with at least some implementations of the present disclosure. In various implementations, the example image processing system 1300 may have an imaging device 1302 to form or receive captured image data. This may be accomplished in various ways. Thus, in one form, the image processing system 1300 may be a digital camera or other image capture device, and the imaging device 1302 in this case may be camera hardware and camera sensor software, modules or components 1308. In other examples, the image processing system 1300 may have an imaging device 1302 that includes or may be a camera, and the logic module 1304 may be in remote communication with the imaging device 1302 or may be otherwise communicatively coupled to the imaging device 1001 to further process the image data.
[0152] In either case, such technology may include a camera, such as a digital camera system, a dedicated camera device, or an imaging phone, whether a still picture or video camera, or some combination of the two. Thus, in one form, the imaging device 1302 may include camera hardware and optics, including one or more sensors and autofocus, zoom, aperture, ND filter, automatic exposure, flash, and actuator control. These controls may be part of a sensor module or component 1306 for operating the sensor. The sensor component 1206 may be part of the imaging device 1302, or may be part of the logic module 1304, or both. Such a sensor component may be used to generate an image of the viewfinder and take a still picture or video. The imaging device 1302 may also have a lens, an image sensor with an RGB Bayer color filter, an analog amplifier, an A / D converter, other components that convert incident light into a digital signal, etc., and / or a combination of these. The digital signal may also be referred to as raw image data in this article.
[0153] Other forms include camera sensor type imaging devices (e.g., a webcam or webcam sensor or other complementary metal-oxide-semiconductor (CMOS) type image sensors) that do not use a red-green-blue (RGB) depth camera and / or microphone array to locate who is speaking. The camera sensor may also support other types of electronic shutters, such as a global shutter in addition to or instead of a rolling shutter, and many other shutter types, as long as a multi-frame statistics collection window can be used. In other examples, an RGB depth camera and / or microphone array may be used in addition to or in place of a camera sensor. In some examples, the imaging device 1302 may be provided with an eye tracking camera. It will be understood that the device 1300 may not have a camera and that images are retrieved from memory, whether or not transmitted from another device.
[0154] In the illustrated example, the logic module 1304 may include an image intake unit 1310 that pre-processes the raw data or obtains an image from memory (whether an initial training image or an LR image to be converted to HR) so that the image is ready for the SR operation described herein. To this end, the SR unit 1312 may include a training unit 1313 and / or a runtime unit 1323 so that a single device can perform only one or the other (training or runtime) rather than both. The training unit 1313 has a scale R unit 1314, a PSF unit 1316, (one or more) stretch factor units 1317, a convolution unit 1318 that operates a blur convolution operation, (one or more) camera characteristic modifier units 1320, and (one or more) extraction units 1322. The runtime unit 1323 may have a pixel processing flow unit 1324 with a processing NN or CNN 1328, a control flow unit 1326 with a statistical NN 1330, and a weight engine 1332, which may also have a NN, such as a fully connected NN or CNN. A scale R unit 1334 and a user selection control unit 1336 may also be provided by the runtime unit 1323. These units or modules perform the tasks implied by the labels of the units or modules and described above using units or modules with similar or identical labels. These units or modules may perform more tasks than described herein.
[0155] The logic module 1304 may be communicatively coupled to the imaging device 1302 to receive raw image data when provided, but otherwise communicate with the memory repository(s) 1348 to retrieve images. The memory repository(s) 1348 may have a buffer 1350 or other external or internal memory formed by RAM (e.g., DRAM), cache, or many other types of memory.
[0156] The image processing system 1300 may have one or more of the processors 1340, such as an Intel Atom, one or more dedicated image signal processors (ISP) 1342. The processor 1340 may include any graphics accelerator, GPU, etc. The system 1300 may also have one or more displays 1356, an encoder 1352, and an antenna 1354. It will be understood that at least a portion of the units, components, or modules mentioned may be considered to be at least partially formed or on at least one of the processors 1340, such as any NN being at least partially formed by the ISP 1342 or GPU, including the statistical NN 1344 or the processing NN 1346.
[0157] In one example implementation, the image processing system 1300 may have a display 1356, at least one processor 1340 communicatively coupled to the display, at least one memory 1348 communicatively coupled to the processor, the at least one memory 1348 having a buffer 1350 according to one example for storing the initial image, LR image, HR image, statistics, NN parameters, etc., and any data mentioned herein. An encoder 1328 and antenna 1354 may be provided to compress or decompress the image data for transmission to or from other devices that may display or store the image. It will be understood that the image processing system 1300 may also include a decoder (or the encoder 1352 may include a decoder) to receive and decode the image data for processing by the system 1300. Otherwise, the processed image 1358 may be displayed on the display 1356 or stored in the memory 1324. As shown, any of these components may be capable of communicating with each other and / or with some portion of the logic module 1304 and / or the imaging device 1302. Thus, processor 1340 can be communicatively coupled to both imaging device 1302 and logic module 1304 to operate these components. Fig.13 The illustrations may include a specific set of blocks or actions associated with a specific component or module, but these blocks or actions may be associated with components or modules other than the specific components or modules illustrated herein.
[0158] refer to Fig.14, one or more aspects of the image processing system described herein are operated according to the example system 1400 of the present disclosure. It can be understood from the nature of the system components described below that such components may be associated with one or some parts of the above-mentioned image processing system, or used to operate this or these parts. In various implementations, the system 1400 can be a media system, although the system 1400 is not limited to this context. For example, the system 1400 can be included in a digital still camera, a digital video camera, or other mobile devices with camera or video functions. Otherwise, the system 1400 can be any device, whether it has a camera or not, such as a mobile small device or an edge device. The system 1400 can be any one of the following: an imaging phone, a webcam, a personal computer (personal computer, PC), a laptop computer, an ultra-portable laptop computer, a tablet device, a touchpad, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (personaldigital assistant, PDA), a cellular phone, a combined cellular phone / PDA, a television, a smart device (e.g., a smart phone, a smart tablet or a smart TV), a mobile Internet device (mobile internet device, MID), a messaging device, a data communication device, and the like.
[0159] In various implementations, system 1400 includes a platform 1402 coupled to a display 1420. Platform 1402 may receive content from a content device, such as content services device(s) 1430 or content delivery device(s) 1440 or other similar content sources. A navigation controller 1450 including one or more navigation features may be used to interact with, for example, platform 1402 and / or display 1420. Each of these components is described in further detail below.
[0160] In various implementations, the platform 1402 may include any combination of a chipset 1405, a processor 1410, a memory 1412, a storage device 1414, a graphics subsystem 1415, applications 1416, and / or a radio device 1418. The chipset 1405 may provide intercommunication between the processor 1410, the memory 1412, the storage device 1414, the graphics subsystem 1415, the applications 1416, and / or the radio device 1418. For example, the chipset 1405 may include a storage adapter (not depicted) capable of providing intercommunication with the storage device 1414.
[0161] Processor 1410 may be implemented as a complex instruction set computer (CISC) or reduced instruction set computer (RISC) processor; an x86 instruction set compatible processor, a multi-core or any other microprocessor or central processing unit (CPU). In various implementations, processor 1410 may be a (one or more) dual-core processor, (one or more) dual-core mobile processor, and the like.
[0162] The memory 1412 may be implemented as a volatile memory device, such as but not limited to a random access memory (RAM), a dynamic random access memory (DRAM), or a static RAM (SRAM).
[0163] Storage 1414 may be implemented as a non-volatile storage device, such as, but not limited to, a magnetic disk drive, an optical disk drive, a tape drive, an internal storage device, an attached storage device, flash memory, a battery-backed SDRAM (synchronous DRAM), and / or a network accessible storage device. In various implementations, such as when multiple hard disk drives are included, storage 1414 may include techniques to increase storage performance and enhance protection for valuable digital media.
[0164] The graphics subsystem 1415 may perform processing of images such as still or video for display. The graphics subsystem 1415 may be, for example, a graphics processing unit (GPU), an image signal processor (ISP), or a visual processing unit (VPU). An analog or digital interface may be used to communicatively couple the graphics subsystem 1415 and the display 1420. For example, the interface may be any of a high-definition multimedia interface, a display port, wireless HDMI, and / or wireless HD compliant technology. The graphics subsystem 1415 may be integrated into the processor 1410 or the chipset 1405. In some implementations, the graphics subsystem 1415 may be a stand-alone card communicatively coupled to the chipset 1405.
[0165] The graphics and / or video processing techniques described herein may be implemented in various hardware architectures. For example, graphics and / or video functions may be integrated into a chipset. Alternatively, a discrete graphics and / or video processor may be used. As another implementation, graphics and / or video functions may be provided by a general purpose processor including a multi-core processor. In another implementation, these functions may be implemented in a consumer electronic device.
[0166] The radio 1418 may include one or more radios capable of sending and receiving signals using various suitable wireless communication technologies. Such technologies may involve communication across one or more wireless networks. Example wireless networks include, but are not limited to, wireless local area networks (WLANs), wireless personal area networks (WPANs), wireless metropolitan area networks (WMANs), cellular networks, and satellite networks. When communicating across such networks, the radio 818 may operate in accordance with one or more applicable standards in any version.
[0167] In various implementations, the display 1420 may include any television-type monitor or display. The display 1420 may include, for example, a computer display screen, a touch screen display, a video monitor, a television-like device, and / or a television. The display 1420 may be digital and / or analog. In various implementations, the display 1420 may be a holographic display. In addition, the display 1420 may be a transparent surface that can receive a visual projection. Such a projection may convey information, images, and / or objects in various forms. For example, such a projection may be a visual overlay of a mobile augmented reality (MAR) application. Under the control of one or more software applications 1416, the platform 1402 may display a user interface 1422 on the display 1420.
[0168] In various implementations, content services device(s) 1430 may be hosted by any national, international, and / or independent service and thus accessible to platform 1402 via the Internet, for example. Content services device(s) 1430 may be coupled to platform 1402 and / or display 1420. Platform 1402 and / or content services device(s) 1430 may be coupled to network 1460 to communicate (e.g., send and / or receive) media information to and from network 1460. Content delivery device(s) 1440 may also be coupled to platform 1402 and / or display 1420.
[0169] In various implementations, content services device(s) 1430 may include a cable box, a personal computer, a network, a phone, an Internet-enabled device or appliance capable of delivering digital information and / or content, and any other similar device capable of transmitting content unidirectionally or bidirectionally between content providers and platform 1402 and / or display 1420 via network 1460 or directly. It will be appreciated that content may be transmitted unidirectionally and / or bidirectionally to and from any one of the components in system 1400 and content providers via network 1460. Examples of content may include any media information including, for example, video, music, medical and gaming information, and the like.
[0170] (One or more) content service device 1430 can receive content, such as cable television programs, including media information, digital information, and / or other content. Examples of content providers may include any cable or satellite television or radio or Internet content providers. The examples provided are not intended to limit implementations according to the present disclosure in any way.
[0171] In various implementations, platform 1402 may receive control signals from a navigation controller 1450 having one or more navigation features. The navigation features of controller 1450 may be used, for example, to interact with user interface 1422. In an implementation, navigation controller 1450 may be a pointing device that may be a computer hardware component (specifically, a human interface device) that allows a user to input spatial (e.g., continuous and multi-dimensional) data into a computer. Many systems, such as graphical user interfaces (GUIs) and televisions and monitors, allow a user to control and provide data to a computer or television using physical gestures.
[0172] Movements of the navigation features of controller 1450 may be replicated on a display (e.g., display 1420) by movements of a pointer, cursor, focus ring, or other visual indicators displayed on the display. For example, under the control of software applications 1416, the navigation features on navigation controller 1450 may be mapped to virtual navigation features displayed on user interface 1422, for example. In implementations, controller 1450 may not be a separate component, but may be integrated into platform 1402 and / or display 1420. The present disclosure, however, is not limited to the elements or in the context shown or described herein.
[0173] In various implementations, for example, when enabled, drivers (not shown) may include technology that enables a user to instantly turn platform 1402 on and off like a television with the touch of a button after initial boot-up. Program logic may allow platform 1402 to stream content to a media adapter or other content services device(s) 1430 or content delivery device(s) 1440 even when the platform is "off." Additionally, chipset 1405 may include hardware and / or software support for, for example, 8.1 surround sound audio and / or high definition (7.1) surround sound audio. The drivers may include a graphics driver for an integrated graphics platform. In an implementation, the graphics driver may include a high-speed peripheral component interconnect (PCI) graphics card.
[0174] In various implementations, any one or more of the components shown in system 1400 may be integrated. For example, platform 1402 and (one or more) content service device 1430 may be integrated, or platform 1402 and (one or more) content delivery device 1440 may be integrated, or platform 1402, (one or more) content service device 1430, and (one or more) content delivery device 1440 may be integrated. In various implementations, platform 1402 and display 1420 may be an integrated unit. For example, display 1420 and (one or more) content service device 1430 may be integrated, or display 1420 and (one or more) content delivery device 1440 may be integrated. These examples are not intended to limit the present disclosure.
[0175] In various implementations, the system 1400 may be implemented as a wireless system, a wired system, or a combination of the two. When implemented as a wireless system, the system 1400 may include components and interfaces suitable for communicating via a wireless shared medium, such as one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, and the like. Examples of wireless shared media may include some portions of a wireless spectrum, such as an RF spectrum, and the like. When implemented as a wired system, the system 1400 may include components and interfaces suitable for communicating via a wired communication medium, such as an input / output (I / O) adapter, a physical connector that connects an I / O adapter to a corresponding wired communication medium, a network interface card (NIC), a disk controller, a video controller, an audio controller, and the like. Examples of wired communication media may include wires, cables, metal leads, printed circuit boards (PCBs), backplanes, switching structures, semiconductor materials, twisted pairs, coaxial cables, optical fibers, and the like.
[0176] Platform 1402 may establish one or more logical or physical channels to convey information. The information may include media information and control information. Media information may refer to any data representing content intended for a user. Examples of content may include, for example, data from a voice conversation, a video conference, streaming video, an electronic mail ("email") message, a voice mail message, alphanumeric symbols, graphics, images, video, text, and the like. Data from a voice conversation may be, for example, voice information, periods of silence, background noise, comfort noise, tones, and the like. Control information may refer to any data representing a command, instruction, or control word intended for an automated system. For example, control information may be used to route media information through a system, or to instruct a node to process media information in a predetermined manner. However, implementations are not limited to Fig.14 The elements or situations shown or described in.
[0177] refer to Fig.15 , small form factor device 1500 is one example of various physical styles or form factors in which system 1300 or 1400 may be implemented. In this way, device 1500 may be implemented as a small or edge mobile computing device with wireless capabilities. A mobile computing device may refer to, for example, any device having a processing system and a mobile power source or power supply (e.g., one or more batteries).
[0178] As mentioned above, examples of mobile computing devices may include a digital still camera, a digital video camera, a mobile device with a camera or video capability such as an imaging phone, a webcam, a personal computer (PC), a laptop computer, an ultraportable laptop computer, a tablet device, a touch pad, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (PDA), a cellular phone, a combination cellular phone / PDA, a television, a smart device (e.g., a smart phone, a smart tablet, or a smart television), a mobile Internet device (MID), a messaging device, a data communication device, and the like.
[0179] Examples of mobile computing devices may also include computers that are arranged to be worn by a person, such as a wrist computer, finger computer, ring computer, eyeglass computer, belt-clip computer, arm-band computer, shoe computers, clothing computers, and other wearable computers. In various implementations, for example, a mobile computing device may be implemented as a smart phone capable of executing computer applications in addition to voice communications and / or data communications. Although some implementations may be described using a mobile computing device implemented as a smart phone as an example, it may be appreciated that other implementations may also be implemented using other wireless mobile computing devices. Implementations are not limited in this context.
[0180] like Fig.15As shown, device 1500 may include a housing having a front side 1501 and a back side 1502. Device 1500 includes a display 1504, an input / output (I / O) device 1506, and an integrated antenna 1508. Device 1500 may also include a navigation feature 1512. I / O device 1506 may include any appropriate I / O device for inputting information into a mobile computing device. Examples of I / O device 1506 may include an alphanumeric keyboard, a numeric keypad, a touch pad, input keys, buttons, switches, microphones, speakers, voice recognition devices and software, and the like. Information may also be input into device 1500 via microphone 1514, or may be digitized by a voice recognition device. As shown, device 1500 may include a camera 1505 (e.g., including at least one lens, aperture, and imaging sensor) and a flash 1510 integrated into the back side 1502 (or elsewhere) of device 1500. Implementation is not limited to this context.
[0181] The various forms of devices and processes described herein can be implemented using hardware elements, software elements, or a combination of the two. Examples of hardware elements may include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software may include software components, programs, applications, computer programs, applications, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination of these. Determining whether an implementation is implemented using hardware elements and / or software elements can vary according to any number of factors, such as desired computing rates, power levels, heat tolerance, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints.
[0182] One or more aspects of at least one implementation may be implemented by representative instructions stored on a machine-readable medium that represent various logic within a processor, which when read by a machine causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to load into fabrication machines that actually fabricate the logic or processor.
[0183] Although certain features described herein have been described with reference to various implementations, this description is not intended to be interpreted in a limiting sense. Therefore, various modifications and other implementations of the implementations described herein that are obvious to those skilled in the art to which the present disclosure belongs are considered to be within the spirit and scope of the present disclosure.
[0184] The following examples are further implementations.
[0185] According to one or more first implementations of the examples, a computer-implemented method of super-resolution image processing includes obtaining at least one lower resolution (LR) image; generating at least one high resolution (HR) image, including inputting image data of the at least one LR image into at least one super-resolution (SR) neural network; separately generating weights or biases or both, including using statistical data of the LR image; and providing the weights or biases or both to the SR neural network for generating the at least one HR image.
[0186] According to one or more second implementations, regarding the first implementation, the weights or biases are generated at a control process, and the control process is operated in parallel with a pixel processing process that operates the SR neural network.
[0187] In accordance with one or more third implementations, with respect to the first implementation, wherein the weights or biases are generated at a control process, the control process is operated in parallel with a pixel processing process that operates the SR neural network, and wherein the rate at which the weights or biases or both are generated at the control process is slower than the frame rate at the SR neural network.
[0188] In accordance with one or more fourth implementations, regarding any one of the first to third implementations, wherein the method includes obtaining a video sequence of the LR images, and wherein the weights or biases or both are updated along the video sequence at predetermined image intervals rather than being updated with each image in the video sequence.
[0189] In accordance with one or more fifth implementations, regarding any one of the first to fourth implementations, wherein the weights or biases or both are generated for a target image and are arranged to be used on an image close to the target image along a video sequence of images when the weights or biases are generated too late to be used on the target image.
[0190] According to one or more sixth implementations, regarding any one of the first to fifth implementations, the method includes inputting at least one of the LR images into a convolutional neural network (CNN) to generate statistical values.
[0191] In accordance with one or more seventh implementations, regarding any one of the first to fifth implementations, wherein the method includes inputting at least one of the LR images into a convolutional neural network (CNN) to generate statistical values, and wherein the statistical data includes global weighted averages, each global weighted average representing the entire image.
[0192] According to one or more eighth implementations, regarding any one of the first to seventh implementations, wherein the weights or biases are generated by a convolution kernel that generates weights using the statistical data.
[0193] In accordance with one or more ninth implementations, regarding any one of the first to eighth implementations, generating the weights or biases includes inputting user preference values and the statistical data into a weight engine neural network to generate a scaling coefficient to be applied to a convolution kernel of the weight.
[0194] According to one or more tenth implementations, a system for image processing comprises: at least one processor; and at least one memory communicatively coupled to the at least one processor, the memory storing a video sequence of low resolution (LR) images, the at least one processor being configured to operate through the following steps: generating a higher resolution (HR) image, comprising inputting image data of the LR image into at least one super resolution (SR) neural network; generating weights or biases or both, comprising using LR image statistics, and generating them at intervals of several images along the video sequence; and providing the weights or biases or both to the SR neural network at the intervals for generating the HR image.
[0195] In accordance with one or more eleventh implementations, regarding the tenth implementation, wherein the statistical data is generated in a statistical neural network, which receives the image data as input and generates weights from a weight layer to apply to statistical values of a value layer to generate output statistics.
[0196] According to one or more twelfth implementations, regarding any one of the tenth to eleventh implementations, wherein the statistical data comprises a vector of global weighted values, each value representing the entire image.
[0197] In accordance with one or more thirteenth implementations, regarding any one of the tenth to twelfth implementations, wherein the generation of the weights or biases includes inputting the statistical data into a weight engine neural network, and the weight engine neural network generates a scaling coefficient to be applied to a convolution kernel of the weights.
[0198] According to one or more fourteenth implementations, regarding any one of the tenth to thirteenth implementations, when generation of the weights or biases or both is late, the neural network uses the latest available weights or biases or both, rather than waiting for updated weights or biases or both to be generated.
[0199] According to one or more fifteenth implementations, regarding any one of the tenth to fourteenth implementations, wherein the SR neural network is trained by intentionally blurring an input training image.
[0200] In accordance with one or more sixteenth implementations, a method of image processing comprises: obtaining an initial training image; generating a blurred training image, comprising intentionally blurring one or more of the initial training images; generating a low-resolution training image or a high-resolution training image or both, and comprising using the blurred image; and training a super-resolution neural network to generate high-resolution pixel values, comprising using the low-resolution training image or the high-resolution training image or both.
[0201] According to one or more seventeenth implementations, with respect to the sixteenth implementation, the method includes blurring the initial training image, including changing image data of the initial training image to simulate blur at least partially caused by a real camera lens.
[0202] According to one or more eighteenth implementations, regarding the sixteenth or seventeenth implementation, wherein the method includes blurring the initial training image, including applying a point spread function (PSF) related value to image data of the initial training image.
[0203] In accordance with one or more nineteenth implementations, regarding any one of the sixteenth or seventeenth implementations, wherein the method includes blurring the initial training image, including applying a point spread function (PSF) related value to image data of the initial training image, and randomly selecting a PSF related value from among a plurality of PSF related values that can be used for blurring and each of which is associated with a different PSF value to apply to an individual initial training image.
[0204] According to one or more twentieth implementations, regarding any one of the sixteenth to nineteenth implementations, wherein the method includes training the SR neural network to convert the resolution on the input LR image using blur caused by various PSF values.
[0205] According to one or more twenty-first implementations, regarding any one of the sixteenth to twentieth implementations, the method includes decimating the blurred image to generate a low-resolution training image or a high-resolution training image or both of a desired size.
[0206] In accordance with one or more twenty-second implementations, at least one non-transitory item has at least one computer-readable medium comprising a plurality of instructions, wherein in response to being executed on a computing device, the computing device operates by: obtaining initial training images; generating blurred training images, including intentionally blurring one or more of the initial training images; generating low-resolution training images or high-resolution training images, or both, and including using the blurred images; and training a super-resolution neural network to generate high-resolution pixel values, including using the low-resolution training images or the high-resolution training images, or both.
[0207] In accordance with one or more twenty-third implementations, regarding the twenty-second implementation, wherein the instructions cause the computing device to operate through the following steps: generating a stretch factor, which stretches the size of the point spread function (PSF) and is set to compensate for reducing the size of the initial training image.
[0208] In accordance with one or more twenty-fourth implementations, regarding the twenty-second implementation, wherein the instructions cause the computing device to operate through the following steps: generating a stretch factor, wherein the stretch factor stretches the size of the point spread function (PSF) and is set to compensate for reducing the size of the initial training image, and wherein the stretch factor is used to convolve the initial training image to form the blurred image.
[0209] In accordance with one or more twenty-fifth implementations, regarding any one of the twenty-second to twenty-fourth implementations, wherein the instructions cause the computing device to operate through the following steps during runtime: inputting a low-resolution image to the SR neural network, and separately generating weights or biases or both of the SR neural network based on statistics of the low-resolution image.
[0210] In one or more twenty-sixth implementations, a device or system includes a memory and a processor to execute the method according to any of the above implementations.
[0211] In one or more twenty-seventh implementations, at least one machine-readable medium includes a plurality of instructions, which in response to being executed on a computing device cause the computing device to perform the method according to any one of the above implementations.
[0212] In one or more twenty-eighth implementations, a device may include means for executing the method according to any one of the above implementations.
[0213] The above examples may include specific combinations of features. However, the above examples are not limited thereto, and in various implementations, the above examples may include engaging only a subset of such features, engaging such features in a different order, engaging such features in a different combination, and / or engaging additional features other than those explicitly listed. For example, all features described for any example method herein may be implemented for any example apparatus, example system, and / or example article, and vice versa.
Claims
1. At least one non-transitory computer-readable medium comprising instructions for causing at least one processor circuit to perform at least the following operations: accessing an input video having a first resolution; providing an input frame of the input video to a neural network, the neural network being trained to upscale the input frame to a second resolution, the neural network being trained to reduce the presence of one or more types of image defects in the input frame; Obtaining an output frame at the second resolution from the neural network; as well as The output frame is caused to be presented as part of an output video at the second resolution.
2. The at least one non-transitory computer readable medium of claim 1, wherein: The one or more defects include the presence of noise in the input frame.
3. The at least one non-transitory computer readable medium of claim 1, wherein: The one or more defects include the presence of blur in the input frame.
4. The at least one non-transitory computer readable medium of claim 1, wherein: The neural network is a deep learning neural network.
5. The at least one non-transitory computer readable medium of claim 1, wherein: The neural network is a convolutional neural network.
6. The at least one non-transitory computer readable medium of claim 1, wherein: The instructions are configured to cause one or more of the at least one processor circuit to cause the output frame to be presented as part of the output video concurrently with access to the input video.
7. The at least one non-transitory computer readable medium of claim 1, wherein: The instructions are configured to cause one or more of the at least one processor circuits to perform operations such as altering the neural network based on a selection before the neural network upscales the input frame to the second resolution.
8. An apparatus comprising: Interface circuit; machine-readable instructions; as well as at least one processor circuit programmed based on the machine-readable instructions to: causing a neural network to process an input frame of an input video, the input video having a first resolution, the neural network being trained to upscale the input frame to a second resolution, the neural network being trained to reduce the presence of one or more types of image defects in the input frame; Obtaining an output frame at the second resolution from the neural network; as well as The output frame is caused to be presented as part of an output video at the second resolution.
9. The device according to claim 8, wherein: The one or more defects include the presence of noise in the input frame.
10. The device according to claim 8, wherein: The one or more defects include the presence of blur in the input frame.
11. The device according to claim 8, wherein: The neural network is a deep learning neural network.
12. The device according to claim 8, wherein: The neural network is a convolutional neural network.
13. The device according to claim 8, wherein: One or more of the at least one processor circuits performs operations such that the output frame is presented as part of the output video concurrently with access to the input video.
14. The device according to claim 8, wherein: One or more of the at least one processor circuits performs operations such as altering the neural network based on the selection before the neural network upscales the input frame to the second resolution.
15. An apparatus comprising: means for accessing an input video having a first resolution; means for processing an input frame of the input video using a neural network, the neural network being trained to upscale the input frame to a second resolution, the neural network being trained to reduce the presence of one or more types of image defects in the input frame, the neural network being used to produce an output frame at the second resolution; as well as Means for presenting the output frame as part of an output video at the second resolution.
16. The device according to claim 15, wherein: The one or more defects include the presence of noise in the input frame.
17. The device according to claim 15, wherein: The one or more defects include the presence of blur in the input frame.
18. The device according to claim 15, wherein: The neural network is a deep learning neural network.
19. The device according to claim 15, wherein: The neural network is a convolutional neural network.
20. The device according to claim 15, wherein: The means for presenting is for presenting the output frame as part of the output video concurrently with access to the input video.
21. The apparatus according to claim 15, comprising: Means for altering the neural network based on the selection before the neural network upscales the input frame to the second resolution.