Image Processing Method, Image Processing Apparatus, and Non-Transitory Computer-Readable Medium

By introducing feature learning DNN and amplified DNN into the SISR framework, combining binary masks and microstructure masks, the problem of resource waste in the processing of multi-scale factors is solved, and resource saving and computational efficiency are improved.

CN114586055BActive Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180005995.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-30
Filing Date
2021-07-12
Publication Date
2025-05-30
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

Existing super-resolution image recovery (SISR) methods require the deployment and storage of large number of model instances when processing multiple scale factors, resulting in waste of resources, especially on devices with limited storage and computing resources.

Method used

By designing a multi-scale factor SISR framework, using feature learning deep neural networks (DNNs) and amplified DNNs, binary masks and microstructure masks are used to reduce model size and computational overhead. The framework only requires one basic SISR model instance, which can adapt to multiple scale factors.

Benefits of technology

Implementing the deployment storage of SISR model instances that reduce multiple scale factors provides a flexible and universal framework to adapt to various types of basic SISR models and reduce the computational volume through microstructure masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114586055B_ABST
    Figure CN114586055B_ABST
Patent Text Reader

Abstract

Provided is an image processing method and apparatus. The image processing method includes: obtaining an input low-resolution (LR) image including a height, a width, and a number of channels; implementing a feature learning deep neural network (DNN) configured to calculate a feature tensor based on the input LR image; generating a high-resolution (HR) image by an upscaling DNN based on the feature tensor calculated by the feature learning DNN, the high-resolution (HR) image having a higher resolution than the input LR image, wherein a networking structure of the upscaling DNN is different according to different scale factors, and wherein a networking structure of the feature learning DNN has the same structure for each of the different scale factors.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of U.S. Provisional Application No. 63 / 065,608 filed on August 14, 2020 and U.S. Application No. 17 / 363,280 filed on June 30, 2021, the entire contents of both of which are hereby incorporated by reference into this application. Background Art 1. Technical Field

[0004] The present disclosure relates to image processing, and particularly to an image processing method, an image processing apparatus, and a non - transitory computer - readable medium.

[0005] 2. Description of Related Art

[0006] ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) has actively sought potential requirements for the standardization of future video coding technologies. ISO / IEC JPEG has established the JPEG - AI group focusing on AI - based neural image compression using deep neural networks (DNNs). The Chinese AVS standard has also established the AVS - AI special group dedicated to neural image and video compression technologies. The success of AI technologies and DNNs in a wide range of video applications such as semantic classification, object detection / recognition, object tracking, video quality enhancement, etc. has posed a strong demand for the compression of DNN models, and both MPEG and AVS are working on a neural network compression standard (NNR) that compresses DNN models to save both storage and computation.

[0007] Meanwhile, with the increasing popularity of high - resolution (HR) displays such as 4K (3840×2160) and 8K (7680×4320) resolutions, image / video super - resolution (SR) has attracted great attention in the industry to generate matching HR image / video content. SISR aims to generate HR images from corresponding low - resolution (LR) images, and it has a wide range of applications in supervised imaging, medical imaging, immersive experiences, etc. In real - world scenarios, it is necessary for SR systems to magnify LR images with various scale factors customized for different users.

[0008] Due to the latest development of DNNs, single - image super - resolution (SISR) has achieved great success. However, these SISR methods treat each scale factor as a single task and train a single model for each scale factor (e.g., ×2, ×3, or ×4). Therefore, all these model instances need to be stored and deployed, which is too expensive and impractical, especially for cases with limited storage and computational resources like mobile devices.

[0009] Therefore, technical solutions to these problems are needed. SUMMARY OF THE INVENTION

[0010] According to an exemplary embodiment, to solve one or more different technical problems, the present disclosure provides a technical solution to reduce network overhead and server computing overhead, while transmitting immersive videos for updating one or more viewport margins.

[0011] It includes a method and an apparatus. The apparatus includes a memory configured to store computer program code and one or more processors configured to access the computer program code and operate according to the instructions of the computer program code. The computer program includes: an acquisition code configured to cause at least one processor to acquire an input low-resolution (LR) image including height, width, and number of channels; an implementation code configured to cause at least one processor to implement a feature learning deep neural network (DNN) configured to calculate a feature tensor based on the input LR image; a generation code configured to cause at least one processor to generate a high-resolution (HR) image by an upscale DNN based on the feature tensor calculated by the feature learning DNN, where the resolution of the high-resolution (HR) image is higher than that of the input LR image, wherein the network structure of the upscale DNN varies according to different scale factors, and wherein the network structure of the feature learning DNN has the same structure for each of the different scale factors.

[0012] According to an exemplary embodiment, the generation code is also configured to cause at least one processor to perform the following operations during the test phase: generate masked weight coefficients of the feature learning DNN based on a target scale factor; and select a sub-network of the upscale DNN based on the selected weight coefficients for the target scale factor and calculate the feature tensor through inference.

[0013] According to an exemplary embodiment, use the selected weight coefficients to cause the feature tensor to pass through an upscale module to generate an HR image.

[0014] According to an exemplary embodiment, wherein the weight coefficients of at least one of the feature learning DNN and the upscale DNN include a 5-dimensional (5D) tensor of size c 1 ,k 1 ,k 2 ,k 3 ,c 2 The input of at least one layer of the feature learning DNN and the upscale DNN includes a 4-dimensional (4D) tensor A of size h 1 ,w 1 ,d 1 ,c 1 and the output of the layer is of size h 2 ,w 2 ,d 2, c 2 4D tensor B, and c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , c 1 , h 2 , w 2 , d 2 and c 2 are all integers greater than or equal to 1, and h 1 , w 1 , d 1 represent the height, weight, and depth of tensor A, and h 2 , w 2 , d 2 represent the height, weight, and depth of tensor B, and c 1 and c 2 represent the number of input channels and output channels respectively, and k 1 , k 2 , k 3 represent the size of the convolutional kernel and correspond to the height axis, weight axis, and depth axis respectively.

[0015] According to an exemplary embodiment, there is also reshaping code configured to cause at least one processor to reshape a 5D tensor into a 3 - dimensional (3D) tensor and to reshape a 5D tensor into a 2 - dimensional (2D) matrix.

[0016] According to an exemplary embodiment, the size of the 3D tensor is c′ 1 , c′ 2 , k, where c′ 1 ×c′ 2 ×k = c 1 ×c 2 ×k 1 ×k 2 ×k 3 , and the size of the 2D matrix is c′ 1 , c′ 2 , where c′ 1 ×c′ 2 = c 1 ×c 2 ×k 1 ×k 2 ×k 3 .

[0017] According to an exemplary embodiment, in the training phase, there is also: determining code configured to cause at least one processor to determine masked weight coefficients; obtaining code configured to cause at least one processor to obtain updated weight coefficients through a weight filling module based on a learning process of weights of a feature learning DNN and weights of an amplification DNN; and performing code configured to cause at least one processor to perform a micro-structure pruning process based on the updated weight coefficients to obtain a model instance and a mask.

[0018] According to an exemplary embodiment, the learning process includes re-initializing those weight coefficients by setting weight coefficients having zero values to any random initial values and corresponding weights of a previously learned model.

[0019] According to an exemplary embodiment, the micro-structure pruning process includes: calculating a loss for each of a plurality of micro-structure blocks in at least one of a 3D tensor and a 2D matrix; and sorting the micro-structure blocks based on the calculated loss.

[0020] According to an exemplary embodiment, the micro-structure pruning process further includes determining whether to stop the micro-structure pruning process based on whether a distortion loss reaches a threshold.

[0021] The present application also provides an image processing apparatus, including an acquisition module, a calculation module, and a generation module. The acquisition module is configured to obtain an input low-resolution (LR) image including a height, a width, and a number of channels. The calculation module is configured to implement a feature learning deep neural network (DNN), and the feature learning deep neural network is configured to calculate a feature tensor based on the input LR image. The generation module is configured to generate a high-resolution (HR) image through an amplification DNN based on the feature tensor calculated by the learning DNN, and the resolution of the high-resolution (HR) image is higher than that of the input LR image. The network structure of the amplification DNN is different according to different scale factors. The network structure of the feature learning DNN has the same structure for each of the different scale factors.

[0022] According to the technical solution of the present application, the following beneficial effects can be achieved: Compared with existing SISR methods, the deployment storage of model instances for implementing SISR with multiple scale factors is greatly reduced, and a flexible and general framework that can adapt to various types of basic SISR models is provided. In addition, the micro-structure mask can reduce the amount of computation. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0024] Figure 1 is a simplified schematic illustration according to an embodiment.

[0025] Figure 2 is a simplified flowchart according to an embodiment.

[0026] Figure 3 is a simplified block diagram according to an embodiment.

[0027] Figure 4 is a simplified block diagram according to an embodiment.

[0028] Figure 5 is a simplified flowchart according to an embodiment.

[0029] Figure 6 is a simplified block diagram according to an embodiment.

[0030] Figure 7 is a simplified block diagram according to an embodiment.

[0031] Figure 8 is a simplified block diagram according to an embodiment.

[0032] Figure 9 is a simplified block diagram according to an embodiment.

[0033] Figure 10 is a schematic illustration according to an embodiment. Detailed implementation manners

[0034] The proposed features discussed below can be used alone or in any combination. In addition, an embodiment can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0035] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, a first terminal 103 may encode video data at a local location for transmission to another terminal 102 via the network 105. A second terminal 102 may receive the encoded video data of another terminal from the network 105, decode the encoded data, and display the restored video data. Unidirectional data transmission may be common in media service applications and the like.

[0036] Figure 1A second pair of terminals 101 and 104 are shown, which are provided to support two-way transmission of encoded video that may occur, for example, during a video conference. For two-way transmission of data, each of terminals 101 and 104 can encode video data captured at a local location for transmission via network 105 to the other terminal. Each of terminals 101 and 104 can also receive the encoded video data sent by the other terminal, can decode the encoded data, and can display the restored video data on a local display device.

[0037] In Figure 1 , terminals 101, 102, 103, and 104 may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. Network 105 represents any number of networks that transmit encoded video data between terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communication network 105 can exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise stated below, the architecture and topology of network 105 may be unimportant for the operation of the present disclosure.

[0038] Figure 2 A flow 200 is shown in accordance with an exemplary embodiment of the multi-scale factor SISR framework described herein, which is used to learn and deploy only one SISR model instance that supports SISR with a multi-scale factor. In particular, the embodiment learns a set of binary masks, with one binary mask for each target scale factor, to guide the SISR model instance to generate HR images with different scale factors as described below.

[0039] For example, at S201, given an input image I of size (h, w, c) LR , where h, w, and c are the number of height, width, and channels respectively, the entire SISR is a DNN that can be divided into two parts: a feature learning DNN at S202 and an upscaling DNN at S204. The feature learning DNN at S202 is designed to calculate a feature tensor F at S203 based on the LR input image I LR , and the upscaling DNN is designed to generate an HR image I at S205 based on the feature tensor F LR , LR HR ​​According to an exemplary embodiment, the network structure of the upscaling DNN at S204 is related to the scale factor and is different for different scale factors, while the network structure of the feature learning DNN at S202 is independent of the scale factor and is the same for all scale factors. The feature learning DNN is generally much larger (in terms of the number of DNN parameters) than the upscaling DNN. According to an exemplary embodiment, a set of N scale factors s in ascending order that the SISR model may be expected to achieve can be obtained 1 ,…,s N . Figure 3 The overall structure of the SISR DNN is given in Box 300 of FIG. 3. For example, at Box 301, {W j f} represents a set of weight coefficients of the feature learning DNN at S202, where each W j f is the weight coefficient of the j-th layer. Additionally, can represent a set of binary masks corresponding to the scale factor s i , where each has the same shape as W j f , and each entry of is 1 or 0, indicating whether the corresponding weight entry in W j f participates in the inference calculation for generating the SR image of the scale factor s i . Different from the feature learning DNN at S202, the upscaling DNN at S204 includes several lightweight sub-networks such as those at Boxes 302, 303, and 304. According to an exemplary embodiment, such {W j s (s i )} represents a set of weight coefficients of the sub-network of the upscaling DNN at S204 corresponding to the scale factor s i , where each is the weight coefficient of the j-th layer.

[0040] Each weight coefficient W j f or is a general 5-dimensional (5D) tensor of size (c 1 ,k 1 ,k 2 ,k 3 ,c 2 ). The input to the layer is a 4-dimensional (4D) tensor A of size (h 1 ,w 1 ,d 1 ,c 1 ), and the output of the layer is of size (h 2 ,w2 , d 2 , c 2 ) of the 4D tensor B. The size c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , h 2 , w 2 , d 2 is an integer greater than or equal to 1. When the size c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , h 2 , w 2 , d 2 of any one takes the number 1, the corresponding tensor is reduced to a lower dimension. Each term in each tensor is a floating-point number. The parameters h 1 , w 1 and d 1 (h 2 , w 2 and d 2 ) are the height, weight, and depth of the input tensor A (output tensor B). The parameter c 1 (c 2 ) is the number of input (output) channels. The parameters k 1 , k 2 and k 3 are the sizes of the convolutional kernels, corresponding to the height axis, weight axis, and depth axis respectively. The output B is calculated as follows: Based on the input A, the weight W j f or and the mask (if available) (Note that for implementations, the mask can also be associated with it, and all entries of are set to 1) of the convolutional operation ⊙. That is, the implementation generally considers that B is calculated as A convolved with the masked weight

[0041] or

[0042] where · is element-wise multiplication.

[0043] Given the learned weight coefficients {W mentioned abovej f}, and a mask Figure 4 400 of shows an exemplary embodiment of the test phase according to an exemplary embodiment. For example, given an input LR image I LR and a target scale factor s i , the corresponding mask is used to generate masked weight coefficients of the feature learning DNN and the weight coefficients are used to select the corresponding sub-network of the upscaling DNN for the scale factor s i . Then, the masked weight coefficients are used by the feature learning module at block 401 to compute a feature F LR through inference calculation, and this feature F LR is further passed through the upscaling module at block 402 to generate an HR result I using the selected weight coefficients HR .

[0044] According to an exemplary embodiment, at process 500 of Figure 5 , each W j f (so that each mask ) can have its shape changed such that the convolution with the reshaped input and the reshaped W j f yields the same output. For example, such an embodiment employs the following configuration: (1) At S502, a 5D weight tensor is reshaped into a 3D tensor of size (c′ 1 , c′ 2 , k), where

[0045] c′ 1 × c′ 2 × k = c 1 × c 2 × k 1 × k 2 × k 3 (for example, a preferred configuration is c′ 1 = c 1 , c′ 2 = c 2 , k = k 1 × k 2 × k 3 ), - Equation 2

[0046] and (2) at S503, the 5D weight tensor is reshaped into a 2D matrix of size (c′ 1 , c′ 2 ), where

[0047] c′ 1 ×c′ 2 =c 1 ×c 2 ×k 1 ×k 2 ×k 3 (For example, some preferred configurations are c′ 1 =c 1 , c′ 2 =c 2 ×k 1 ×k 2 ×k 3 , or c′ 2 =c 2 , c′ 1 =c 1 ×k 1 ×k 2 ×k 3 ). - Equation 3

[0048] The exemplary embodiments are designed to have the desired mask microstructure to align with the basic GEMM matrix multiplication process of how to implement the convolution operation, so that the inference calculation using the masked weight coefficients can be accelerated. According to the exemplary embodiments, a block microstructure can be used for the mask (for the masked weight coefficients) of each layer in the 3D reshaped weight tensor or the 2D reshaped weight matrix, and specifically, for the case of the reshaped 3D weight tensor, at S504, there is a division that divides it into blocks of size (g i , g o , g k ), and for the case of the reshaped 2D weight matrix, at S505, it is divided into blocks of size (g i , g o ). All the terms in the block of the mask will have the same binary value 1 or 0; that is, according to the exemplary embodiments, the weight coefficients are masked in a block microstructure manner.

[0049] Figure 7 The exemplary embodiment shown as 700 in j f shows the overall workflow of the training phase according to the exemplary embodiments, where the goal is to have a model instance with weights {W and (for i = 1,..., N), and each and is for each scale factor s of interest i . Thus, a progressive multi-stage training framework representing the technical advantages of the problems pointed out herein is disclosed to achieve this goal.

[0050] As Figure 7 shown, the embodiments are described under the assumption of attempting to train a mask for s i , and there is a current model instance with weights {W j f (i - 1)} and a corresponding mask . In addition, for the current scale factor s i , there is a corresponding upsampling DNN with weight coefficients to be learned. In other words, the goal is to obtain the mask and the updated weight coefficients {W j f (i)} as well as the new weight coefficients Therefore, first, the weight coefficients in {W j f (i - 1)} that are masked by are determined. For example, in a preferred embodiment, if the entry in is 1, the corresponding weight in W j f (i - 1) will be determined. In fact, the remaining weight coefficients corresponding to the 0 entries in have a value of 0. Then, there is a learning process of filling the undetermined zero-valued weights of the feature learning DNN and the weights of the upsampling DNN by the weight filling module 701 . This results in a set of updated weight coefficients {W j f′ (i)} and Then, based on {W j f′ (i)}, and the embodiment performs microstructure pruning through the microstructure pruning module 702 to obtain the model instance and the mask {W j f},

[0051] Other exemplary embodiments, for example Figure 8 at, other exemplary embodiments show a process 800, where the detailed workflow of the weight filling module 801 is also shown. For example, given the current weights {W Figure 8 (i - 1)} and the corresponding mask j f in the weight determination and filling module 801, the weight coefficients in {W that are masked by j f (i - 1)} are determined, and for {W the ones masked by jf re-initialize the remaining weight coefficients having zero values in {W^{(i - 1)}} (e.g., by setting them to some random initial values or using the corresponding weights of a previously learned complete model such as a first complete model having weights {W^{(0)}}). This gives the weight coefficients {W^{(i)}} of the feature learning DNN. Additionally, initialize the weights of the magnification DNN (e.g., by setting them to some randomly initialized values or using the corresponding weights of some previously learned complete model such as a single complete model trained for the current scale factor s). j f′ (0)}). This gives the weight coefficients {W^{(i)}} of the feature learning DNN. Additionally, initialize the weights of the magnification DNN (e.g., by setting them to some randomly initialized values or using the corresponding weights of some previously learned complete model such as a single complete model trained for the current scale factor s). j f′ (i)}. Additionally, initialize the weights of the magnification DNN (e.g., by setting them to some randomly initialized values or using the corresponding weights of some previously learned complete model such as a single complete model trained for the current scale factor s). (e.g., by setting them to some randomly initialized values or using the corresponding weights of some previously learned complete model such as a single complete model trained for the current scale factor s). i After that, the training input image I is passed through the feature learning DNN to compute the feature F using {W^{(i)}} in the feature learning module 802. LR is passed through the feature learning DNN to compute the feature F using {W^{(i)}} in the feature learning module 802. j f′ (i)}. LR F LR is additionally passed through the magnification DNN to compute the HR image I using in the magnification module 803. For training purposes, each training input LR image I has a corresponding ground truth HR image for the current scale factor s HR . LR for the current scale factor s i The overall training objective is to minimize the distortion between the ground truth image and the estimated HR image I . The distortion loss HR can be computed in the compute loss module 804 to measure the distortion, e.g., the L norm of the difference between HR and I 1 or the L 2 norm. The gradient of this loss can be computed to update the undetermined weight coefficients in {W^{(i)}} of the feature learning DNN and the weight coefficients of the magnification DNN j f′ (i)} and the weight coefficients of the magnification DNN Typically, multiple epoch iterations will be employed in this backpropagation and weight update module 805 to optimize the loss e.g., until a maximum number of iterations is reached or until the loss converges.

[0052] According to an exemplary embodiment, Figure 9 a flow 900 is shown that further describes the more detailed workflow of the microstructure pruning module 702. For example, given the updated weights {W Figure 7 of the feature learning DNN from the above weight filling module 701 jf′ (i)} and the magnified DNN and the current mask First, calculate the pruning mask in the pruning mask calculation module 901

[0053] According to an exemplary embodiment, determine the masked weight coefficients in {W j f′ (i)} and the remaining undetermined weight coefficients in {W At block 904, calculate the pruning loss L j f′ (b) for each micro-structure block b (3D block of the 3D reshaped weight tensor or 2D block of the 2D reshaped weight matrix) as described above (e.g., the L p of the weights in the block 1 or L 2 norm). Additionally, these micro-structure blocks are sorted in ascending order based on their pruning losses, and these blocks are pruned top-down according to the sorted list (i.e., by setting the corresponding weights in the pruned blocks to 0) until a stopping criterion is reached. For example, given a validation dataset S val , an SISR model with weights {W j f′ (i)} and can produce a distortion loss

[0054]

[0055] As more and more micro-blocks are pruned, this distortion loss will gradually increase. The stopping criterion can be an admissible percentage threshold for the allowable increase in the distortion loss. The stopping criterion can also be a simple percentage of the micro-structure blocks to be pruned (e.g., 50%). A set of binary pruning masks can be generated, where an entry of 1 in the mask means that the corresponding weight in W j f′ (i) is pruned. Then, the other undetermined weights in W j f′ (i) that are masked by are determined to be pruned, and the remaining weights in W j f′ (i) that are not masked by or are updated, and the weights are optimized for the distortion loss on the training data through conventional backpropagation Typically, multiple epoch iterations will be employed in this backpropagation and weight update process at block 905 to optimize the distortion loss, e.g., until a maximum number of iterations is reached or until the loss converges.

[0056] The corresponding mask can be calculated as

[0057]

[0058] That is to say, the untrimmed entries in that are not masked in will be additionally set to 1 because they are masked in j f (i)}. Additionally, the above micro-structure weight pruning process will output the updated weights {W Note that the above micro-structure pruning process can also optionally be applied to to further reduce the model size and inference computation. That is to say, in the computation pruning mask module 901, the weights of the magnified DNN from box 903 can also be reshaped and partitioned into micro-structures, the pruning losses of these micro-structures can be calculated, and the top-ranked micro-structures with small pruning losses can be pruned. Additionally, the implementation can optionally choose to do so as a trade-off between reducing SISR distortion and saving storage and computation.

[0059] Finally, the last updated weights {W j f (N)} are the final output weights {W j f} of the feature learning DNN of the learning model instance according to the exemplary implementation.

[0060] In view of the above implementation, and compared with any previous attempts at SISR methods, the present disclosure provides technical benefits such as significantly reducing the deployment storage to achieve SISR with multiple scale factors, a flexible and general framework that adapts to various types of base SISR models, and the micro-structure mask provides an additional benefit of reduced computation.

[0061] The above technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media, or implemented by one or more specially configured hardware processors. For example, Figure 10 FIG. shows a computer system 1000 suitable for implementing certain implementations of the disclosed subject matter.

[0062] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be assembled, compiled, linked, etc. by mechanisms to create code including instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc. or executed through interpretation, microcode execution, etc.

[0063] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0064] Figure 10 The components of the computer system 1000 shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the present disclosure. Nor should the configuration of the components be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system 1000.

[0065] The computer system 1000 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs by one or more human users through, for example, tactile inputs (such as keystrokes, swipes, data glove movements), audio inputs (such as voice, clapping), visual inputs (such as gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that is not necessarily directly related to conscious human input, such as audio (such as voice, music, ambient sound), images (such as scanned images, photographic images obtained from a still image camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0066] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard 1001, mouse 1002, touchpad 1003, touch screen 1010, joystick 1005, microphone 1006, scanner 1008, and camera device 1007.

[0067] The computer system 1000 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen 1010 or joystick 1005, but there may also be tactile feedback devices that do not serve as input devices), audio output devices (such as speakers 1009, headphones (not depicted)), visual output devices (such as screen 1010, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual output or more than three-dimensional output, such as through stereoscopic graphics output; virtual reality glasses (not depicted), holographic displays, and smell pots (not depicted)), and printers (not depicted).

[0068] The computer system 1000 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 1020 with media such as CD / DVD 1011, thumb drive 1022, removable hard disk drive or solid state drive 1023, conventional magnetic media (such as tapes and floppy disks (not depicted)), devices based on dedicated ROM / ASIC / PLD (such as security dongles (not depicted)), etc.

[0069] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.

[0070] The computer system 1000 may also include an interface 1099 to one or more communication networks 1098. The communication network 1098 can be, for example, wireless, wired, optical. The network 1098 can also be local, wide area, urban, in-vehicle, and industrial, real-time, delay-tolerant, etc. Examples of the network 1098 include: local area networks such as Ethernet, wireless LAN, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), in-vehicle and industrial (including CAN bus), etc. Some networks 1098 typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (1051 and 1051) (such as, for example, the USB port of the computer system 1000); other networks are typically integrated into the core of the computer system 1000 by attaching to a system bus as described below (such as an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smart phone computer system). Using any of these networks 1098, the computer system 1000 can communicate with other entities. Such communication can be unidirectional and only receive (such as broadcast television), unidirectional and only transmit (such as CAN bus to certain CAN bus devices), or bidirectional (such as using a local or wide area digital network to other computer systems). Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0071] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core 1040 of the computer system 1000.

[0072] The core 1040 may include one or more central processing units (CPUs) 1041, a graphics processing unit (GPU) 1042, a graphics adapter 1017, a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) 1043, a hardware accelerator 1044 for certain tasks, etc. These devices, together with a read-only memory (ROM) 1045, a random access memory 1046, an internal mass storage device 1047 such as an internal non-user-accessible hard disk drive, an SSD, etc., may be connected via a system bus 1048. In some computer systems, the system bus 1048 may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus 1051 to the system bus 1048 of the core. The architecture of the peripheral bus includes PCI, USB, etc.

[0073] The CPU 1041, GPU 1042, FPGA 1043, and accelerator 1044 may execute certain instructions that, when combined, may constitute the aforementioned computer code. The computer code may be stored in the ROM 1045 or the RAM 1046. Interim data may also be stored in the RAM 1046, while permanent data may be stored in, for example, the internal mass storage device 1047. Fast storage and retrieval of any memory device in the memory devices may be achieved by using a cache memory that may be closely associated with one or more CPUs 1041, GPUs 1042, mass storage devices 1047, ROM 1045, RAM 1046, etc.

[0074] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specifically designed and constructed for the purposes of the present disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the field of computer software.

[0075] By way of example and not limitation, a computer system 1000 having an architecture and in particular a core 1040 can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software implemented in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the core 1040 having non-transitoriness, such as an on-core mass storage device 1047 or a ROM 1045. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 1040. Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core 1040 and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in a RAM 1046 and modifying such data structures in accordance with processes defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of logic implemented hard-wired or otherwise in circuitry (e.g., an accelerator 1044), which can operate in place of or in conjunction with the software to perform specific processes or specific portions of specific processes described herein. In appropriate cases, a reference to software can encompass logic, or a reference to logic can encompass software. In appropriate cases, a reference to computer-readable media can encompass circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry implementing logic for execution, or both of the above. The present disclosure encompasses any suitable combination of hardware and software.

[0076] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood that those skilled in the art can envision many systems and methods that, although not explicitly shown or described herein, implement the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. An image processing method, characterized in that, the method is executed by at least one processor and includes: obtaining an input low-resolution (LR) image including the number of height, width, and channels; implementing a feature learning deep neural network DNN, the feature learning deep neural network being configured to calculate a feature tensor based on the input low-resolution (LR) image; generating a high-resolution (HR) image by an amplification DNN based on the feature tensor calculated by the feature learning DNN, the resolution of the high-resolution (HR) image being higher than that of the input low-resolution (LR) image, wherein the network structure of the amplification DNN varies according to different scale factors, learning a set of binary masks, there being one binary mask for each target scale factor, to guide the SISR model instance to generate HR images with different scale factors, and wherein the network structure of the feature learning DNN has the same structure for each of the different scale factors.

2. The method according to claim 1, characterized in that, the method further includes in the test phase: generating masked weight coefficients of the feature learning DNN based on a target scale factor; and selecting a sub-network of the amplification DNN for the target scale factor based on the selected weight coefficients and calculating the feature tensor by inference.

3. The method according to claim 2, characterized in that, the method further includes: using the selected weight coefficients to cause the feature tensor to pass through an amplification module to generate the HR image.

4. The method according to claim 1, characterized in that, The weight coefficients of at least one of the feature learning DNN and the amplification DNN include a 5-dimensional (5D) tensor of size c 1 , k 1 , k 2 , k 3 , c 2 of Among them, the input of at least one layer of the feature learning DNN and the amplification DNN includes a 4D tensor A of size h 1 , w 1 , d 1 , c 1 . where the output of the layer is a 4D tensor B of size h 2 , w 2 , d 2 , c 2 , and Among them, c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , c 1 , h 2 , w 2 , d 2 and c 2 are all integers greater than or equal to 1. where h 1 、 w 1 、 d 1 represent the height, weight, and depth of the tensor A, Among them, h 2 , w 2 , d 2 represent the height, weight, and depth of the tensor B, where c 1 and c 2 represent the number of input channels and output channels, respectively, and Among them, k 1 , k 2 , k 3 represent the size of the convolution kernel and correspond to the height axis, width axis, and depth axis respectively.

5. The method according to claim 4, characterized in that, the method further includes: reshaping the 5-dimensional (5D) tensor into a 3-dimensional (3D) tensor; and reshaping the 5-dimensional (5D) tensor into a 2-dimensional (2D) matrix.

6. The method according to claim 5, characterized in that, The size of the 3D tensor is c′ 1 , c′ 2 , k, where c′ 1 ×c′ 2 ×k = c 1 ×c 2 ×k 1 ×k 2 ×k 3 , and Among them, the size of the 2D matrix is c′ 1 , c′ 2 , where c′ 1 ×c′ 2 =c 1 ×c 2 ×k 1 ×k 2 ×k 3 .

7. The method according to claim 6, characterized in that, the method further includes in the training phase: determining masked weight coefficients; obtaining updated weight coefficients through a weight filling module based on the learning process of the weights of the feature learning DNN and the weights of the amplification DNN; and performing a micro-structure pruning process based on the updated weight coefficients to obtain a model instance and a mask.

8. The method according to claim 7, characterized in that, the learning process includes setting the weight coefficients with zero values to any random initial values, or re-initializing the weight coefficients with zero values and not masked using the corresponding weights of the previously learned model.

9. The method according to claim 7, characterized in that, the micro-structure pruning process includes: calculating the loss of each of a plurality of micro-structure blocks in at least one of the 3-dimensional (3D) tensor and the 2-dimensional (2D) matrix; and sorting the micro-structure blocks based on the calculated loss.

10. The method according to claim 9, characterized in that, the micro-structure pruning process further includes: Determine whether to stop the micro-structure pruning process based on whether the distortion loss reaches a threshold value.

11. An image processing apparatus, characterized in that the apparatus includes: at least one memory configured to store computer program code; at least one processor configured to access the computer program code and operate according to what is indicated by the computer program code to execute the method according to any one of claims 1 to 10.

12. A computer device, characterized in that the computer device includes: one or more computer-readable non-transitory storage media configured to store computer program code; and one or more computer processors configured to access the computer program code and execute the method according to any one of claims 1 to 10 according to the indication of the computer program code.

13. A non-transitory computer-readable medium storing a program, characterized in that the program causes a computer to execute a process such as the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image super-resolution method based on dense connection network

    CN106991646A

  • Multi-scale image super-resolution method based on dual-path network

    CN109064405A