Image processing method and system based on fast neural network training

By reparameterizing the convolutional kernel of a convolutional neural network using frequency-aware reparameterization, the problems of excessive iterations and poor image quality in traditional neural network training are solved, resulting in faster training and higher image quality.

CN121569326APending Publication Date: 2026-02-24INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380099605.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-12
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Traditional deep neural networks suffer from learning bias in image processing, resulting in excessive iterations, slow training convergence, and poor output image quality. This imbalance, particularly when processing low-frequency and high-frequency image data, affects artifact detection and removal.

Method used

Frequency-aware reparameterization (FAR) is used to reparameterize the convolutional kernels of the convolutional neural network. By weighting the kernels in the frequency domain, the emphasis on low-frequency image data is reduced, and more high-frequency image data is processed earlier, resulting in a more uniform iterative distribution.

Benefits of technology

It significantly reduces the number of training iterations, accelerates the training process, and improves the quality of image processing, especially in image restoration and video compression, thereby improving image quality and reducing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121569326A_ABST
    Figure CN121569326A_ABST
Patent Text Reader

Abstract

Methods, systems, and devices for image processing include receiving initial convolutional kernel coefficients to train a neural network having at least one convolutional layer; and converting the coefficients into a frequency domain kernel. The method further includes generating, by the processor circuitry, at least one convolution kernel having convolution coefficients generated by modifying the frequency domain kernel in the frequency domain.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Many image processing applications use neural networks to analyze the details and patterns of image data, such as for video compression, image restoration, denoising, super-resolution, and synthetic image generation. It has been found that conventional training of many different types of deep neural networks used for image processing suffers from learning bias, which leads to excessive iterations and consequently, relatively slow convergence during training. Attached Figure Description

[0002] The materials described herein are illustrated in the accompanying drawings by way of example rather than limitation. For simplicity and clarity, the elements illustrated in the drawings are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to others for clarity. Furthermore, reference labels are repeated in the drawings where appropriate to indicate corresponding or similar elements. In the figures:

[0003] Figure 1 This is a schematic diagram of an image processing system based on at least one of the implementations described herein;

[0004] Figure 2 This is a flowchart of a method for fast training of a neural network for image processing according to at least one of the implementations described herein;

[0005] Figure 3 This is a schematic diagram illustrating the operation for rapid training of an image processing neural network according to at least one of the implementations described herein;

[0006] Figure 4 This is a flowchart of a method for image processing using a neural network with fast training, based on at least one of the implementations described herein;

[0007] Figure 5 It is a set of subband frequency component plots for an image, which is generated by the neural network during the experiment and plotted for training iterations of both the conventional (vanilla) implementation and the implementation according to at least one of the implementations in this paper, with mean absolute discrete cosine coefficients.

[0008] Figure 6 The graph shows the experimental results as peak signal-to-noise ratio (PSNR) versus training iterations, and the results are correlated with both the traditional basic implementation and the fast training according to the implementation described in this paper.

[0009] Figure 7This is another set of subband frequency component plots for another image, which is generated by the neural network during the experiment and plotted as training iterations of both the conventional basic implementation and the implementation described in this paper with the mean absolute discrete cosine coefficients.

[0010] Figure 8 This is a schematic flowchart of a system for image compression that uses various compression codecs with overfitting neural networks, and simultaneously employs both a traditional basic implementation and fast training based on at least one of the implementations disclosed herein.

[0011] Figure 9 It is a graph of the relevant images for fast training and raw JPEG image data according to at least one of the implementations disclosed herein, based on a traditional basic implementation method. The graph shows an example result using JPEG compression as a rate-distortion curve, which is plotted as PSNR versus bits per pixel.

[0012] Figures 10A to 10D It is an image that compares the image compression result with block artifacts based on fast training and raw JPEG image data according to at least one of the traditional basic implementation methods disclosed in this paper;

[0013] Figures 10E to 10H It is an image that compares the image compression result with ringing artifacts based on fast training and raw, efficient image file (HEIF) image data according to at least one of the traditional basic implementation methods disclosed in this paper;

[0014] Figures 10I to 10L It is an image that compares the image compression result with chroma subsampling artifacts based on fast training and raw general video coding (VVC) image data according to at least one of the implementations disclosed in this paper, against the traditional basic implementation method;

[0015] Figure 11 This is a graph showing the bit distortion rate versus total number of iterations for both the conventional basic implementation and fast training according to at least one of the implementations disclosed herein for the Challenge of Learning-Oriented Image Compression (CLIC) on a neural network platform.

[0016] Figure 12 This is an explanatory diagram of the example system;

[0017] Figure 13 This is another illustrative diagram of the example system; and

[0018] Figure 14The illustrations depict example devices arranged according to at least some implementations of this disclosure. Detailed Implementation

[0019] One or more implementations will now be described with reference to the accompanying drawings. While specific configurations and arrangements are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be employed without departing from the spirit and scope of the description. It will be apparent to those skilled in the art that the techniques and / or arrangements described herein can also be used in a variety of other systems and applications besides those described herein.

[0020] While the following description illustrates various implementations that can be embodied in architectures such as System-on-a-Chip (SoC), the implementations of the technologies and / or arrangements described herein are not limited to any particular architecture and / or computing system and can be implemented by any architecture and / or computing system for similar purposes. For example, various architectures and / or various computing devices and / or consumer electronics (CE) devices (such as servers, desktop computers, laptop computers, set-top boxes, smartphones, tablet computers, one or more cameras or camera arrays, and virtual reality (VR) or augmented reality (AR) headsets or other display devices, etc.) employing, for example, multiple integrated circuit (IC) chips and / or packages can implement at least a portion of the technologies and / or arrangements described herein. Furthermore, while the following description may set forth many specific details (such as logical implementations, types and interrelationships of system components, logical partitioning / integration choices, etc.), the claimed subject matter can be practiced without such specific details. In other instances, some material, such as control structures and complete sequences of software instructions, may not be shown in detail so as not to obscure the material disclosed herein.

[0021] The materials disclosed herein can be implemented in hardware, firmware, software, or any combination thereof. The materials disclosed herein can also be implemented as instructions stored on a machine-readable medium that can be read and executed by one or more processors. A machine-readable medium can include any medium and / or mechanism for storing or transmitting information in a machine-readable (e.g., computing device) form. For example, a machine-readable medium can include read-only memory (ROM), random access memory (RAM), disk storage media, optical storage media, flash memory devices, electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). In another form, non-transitory articles of art (such as non-transitory computer-readable media) can be used with any of the above examples or other examples, but non-transitory articles of art themselves do not include transient signals. Non-transitory articles of art do include those elements other than the signal itself that can temporarily store data in a “transitory” manner (such as RAM, etc.).

[0022] References to "an implementation," "implementation method," "example implementation," etc., in the specification indicate that the described implementation may include a specific feature, structure, or characteristic, but each implementation may not necessarily include that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an implementation method, it should be assumed that implementing that feature, structure, or characteristic in conjunction with other implementation methods is within the knowledge of someone skilled in the art, whether or not it is explicitly described herein.

[0023] This paper describes methods, devices, apparatuses, computing platforms, media, and articles related to image processing based on fast neural network training.

[0024] It has been found that deep neural networks (DNNs) that process image data for image analysis, modification, or generation exhibit a learning bias that causes low-frequency image data to be processed earlier than high-frequency image data. Specifically, image data is referred to as low-frequency image data when changes in image data (such as chroma or luminance (or brightness or intensity) pixel values) occur slowly (with small variations) across a relatively large number of pixels, while changes in image data values ​​over small pixel distances or regions involve high-frequency image data. It has been found that DNNs are generally biased or tend to prioritize processing low-frequency data before processing high-frequency data in an image. This neural network training effort can be referred to as processing or fitting low-frequency or high-frequency image data. It has been found that the learning bias increases the number of iterations required to determine sufficient weights or other hyperparameter values ​​during training (or in other words, to achieve sufficient convergence), which in turn leads to delays in the training of the neural network. It has also been found that the learning bias degrades the quality of the output images of those neural networks that modify or generate image data, where tests have shown that the learning bias can lead to missed detection of artifacts that should be identified and removed using an image processing pipeline of the neural network.

[0025] To address these issues, this system, method, and apparatus employ Frequency Aware Reparameterization (FAR) technology, which reparameterizes the convolutional kernels of a convolutional neural network. Reparameterization is achieved by performing a weighted sum of multiple frequency-domain kernels, each representing a different frequency. As illustrated in the example below, the frequency-domain kernels are obtained by applying a frequency transformation algorithm, such as Discrete Cosine Transform (DCT), to the initial spatial-domain convolutional kernel coefficients. The frequency-domain kernels can be orthogonal and normalized. Weighting is achieved by applying reparameterized weights to the frequency-domain (or DCT here) kernel to modify the frequency components of the kernel in the frequency domain. The reparameterized weights can be updated during neural network training to reduce the emphasis on low-frequency image data and process more high-frequency image data earlier. This results in a more uniform or homogeneous distribution of training low-frequency and high-frequency image data throughout the training iterations.

[0026] With a better distribution, and because reparameterization occurs directly in the frequency domain, it leads to a significant reduction in the number of iterations required to train the neural network, thereby accelerating training and providing better image quality for image processing pipelines such as image restoration.

[0027] Although examples of image restoration and video compression are provided in this paper, it will be understood that the FAR method disclosed herein can be used with any image and / or video processing CNN, any image and / or video processing CNN can benefit from the efficiency improvement of fast convergence or capturing high-frequency components (including denoising, super-resolution, image synthesis or generation, etc.).

[0028] It will be understood that convolutional kernels, frequency domain kernels, and reparameterized weights can all have values ​​or elements referred to as weights. In this paper, for clarity, spatial domain convolutional kernels will be referred to as having convolution coefficients, DCT kernels will be referred to as having frequency components in the frequency domain, and reparameterized weights will be referred to as reparameterized (RP) weights.

[0029] refer to Figure 1 An example image processing device or system 100 may be provided for the rapid training of an image processing neural network. The image processing system 100 may receive image data 102, which is any data related to an image, including pixel chroma or brightness (or illuminance) values, as well as gradient, depth, transparency, etc. This is not specifically limited.

[0030] The image processing system 100 may also be or have a neural network application 104 to process image data 102 for many different types of applications. This may include (1) image analysis applications, such as image compression, image quality assessment, object segmentation, detection, recognition and 3D modeling; (2) image modification applications, such as image restoration, denoising and / or super-resolution; and (3) image generation applications, such as image synthesis for images with virtual perspectives, and many other image processing applications.

[0031] This example neural network application 104 may have a neural network (NN) 106 being trained and generating output 108, a controller 110 controlling the NN convolutional kernel unit 112, bias unit 114, NN weight unit 116, activation function unit 118, and any other units for loading data or parameters into the NN 106. The NN 106 may have any desired layer arrangement, provided that the NN 106 has at least one convolutional layer 105, regardless of whether the at least one convolutional layer 105 is defined as part of a convolutional block, and, as desired, the NN 106 may have any other type of layer 107 (such as input, rectified linear (ReLU), pooling, fully connected (or dense), normalized, etc.), and any desired channel, batch size, layer size, activation function, hyperparameter layer, and / or arrangement. This arrangement may be at the hardware or firmware level where buffers for one or more image accelerators or image signal processors (ISPs) are loaded, or alternatively at the software level.

[0032] The NN application 104 may also have an NN training unit 120 with a setting unit 121, which receives parameters (or hyperparameter instructions) for training and instructs the controller 110 to operate according to those training parameters and instructions. A loss unit 122 analyzes the output 108 from the NN 106 and determines modified parameters for the next cycle or iteration of training. Particularly for the fast training described herein, the NN training unit 120 may have an initial spatial domain kernel coefficient unit 124, a frequency transformation unit 126, a reparameterized weights (RP weights) unit 128, and a reparameterization unit 130.

[0033] In operations used for fast training, the initial spatial domain kernel coefficient unit (or coefficient unit only) 124 can be initialized with random values ​​as the coefficients of the convolutional kernel. Alternatively, non-frequency-aware initialization can be performed to generate initial convolutional kernel coefficients, which can be used without using the FAR technique disclosed herein. Similarly, the RP weight unit 128 can also be initialized with random values. Subsequently, the frequency transformation unit 126 transforms the convolutional kernel into the frequency domain using a transformation algorithm such as DCT to generate a frequency domain kernel. This can include any orthogonalization and / or normalization of the frequency domain kernel. Each frequency domain kernel, or DCT kernel here, represents a different frequency and has multiple frequency components arranged as elements in a matrix (such as 3x3 in the current example).

[0034] The RP weights can be scalar values ​​modified by the RP weight unit 128 during training iterations, regardless of whether the scalar values ​​are positive or negative integers, but as desired, the RP weights can be other types of values. The RP weights are set to reduce the low-frequency emphasis of the convolutional kernel and perform more high-frequency fitting in earlier iterations, making low-frequency and high-frequency processing more balanced in earlier iterations, and in a form that runs through all iterations, as described in detail elsewhere in this paper.

[0035] RP weights and frequency domain coefficients are provided to reparameterization unit 130, which multiplies or applies one (or at least one) of the multiple RP weights to the corresponding frequency domain kernel to form a weighted frequency domain kernel. The weighted frequency domain kernels are then summed to generate a weighted sum convolutional kernel (or a FAR-only convolutional kernel) with frequency-aware reparameterization (FAR) coefficients. It will be understood that the weighted sum can be described as an inverse DCT operation. The FAR convolutional kernel is then provided from reparameterization unit 130 to NN convolutional kernel unit 112 to train NN 106.

[0036] It will be understood that, when desired, the publicly disclosed FAR operation described herein can be disabled for non-frequency-based training. In this case, the loss unit 122 can directly provide instructions to the NN convolutional kernel unit 112 to modify the convolutional kernel coefficients for the next training iteration.

[0037] It will also be understood that, in addition to or in place of the arrangement described herein with respect to system 100, many different arrangements can be used to perform the fast training disclosed herein.

[0038] refer to Figure 2 According to at least some implementations of this disclosure, an example process 200 for rapid training of a neural network for image processing is arranged. The training process 200 may include one or more operations 202-206, typically evenly numbered. Furthermore, references may be made herein to at least... Figure 1 and Figures 12-14 The system or device is 100, 1200, 1300 or 1400 to describe process 200.

[0039] As a preliminary step or stage, the neural network to be trained can be used in any image processing application with any desired architecture as described above, provided that the neural network has at least one convolutional layer using a convolutional kernel. Furthermore, for training the neural network, preliminary operations may include any initial setup operations that must be performed to set any hyperparameters other than the convolutional kernel coefficient values, including bias, initial neural network weight values, activation functions, etc. Such preliminary operations may also set the convolutional kernel size, stride, padding, dilation, kernel type, and other specifications in addition to specific kernel coefficient values. This can be achieved through custom programming or by feeding data into a known machine or deep learning framework application. In the example used in this paper, the convolutional kernel is 3x3.

[0040] This initial stage may also include selecting a dataset for training, feeding image data into the neural network being trained to generate the network output for supervised or unsupervised training. The output can be used to generate a loss, which is then used to compute updated reparameterized weights, as explained below.

[0041] Process 200 may include "receiving initial convolutional kernel coefficients to train a neural network with at least one convolutional layer" 202. As mentioned, the initial convolutional kernel coefficients may be initial random coefficient values ​​or initial values ​​that can be used without frequency-aware reparameterization. Otherwise, known initialization algorithms (which may also be random values) are typically provided by the deep learning framework application and are usually performed by determining the initial kernel coefficients using a uniform or normal distribution and the number of inputs or outputs. The result is one or more initial convolutional kernels, each with initial coefficients arranged in a matrix (such as 3x3 in the current example).

[0042] Process 200 may include "converting coefficients into frequency domain kernels" 204. This can be performed using a transformation algorithm that decouples the spatial patterns of the convolutional kernel into patterns representing different frequencies. In the example used herein, a Discrete Cosine Transform (DCT) is applied to the initial convolutional kernel. The DCT outputs a set of frequency domain (or DCT) kernels, each representing a different frequency, thus establishing the frequency domain. Each frequency domain kernel may have a matrix of frequency components, each frequency component representing a different frequency subband. In one form, the size of the frequency domain kernel is the same as the size of the spatial input convolutional kernel, in this example, 3x3.

[0043] According to an example, to compute the frequency components of each frequency domain kernel, the following DCT equation (1) represents the kernel of a convolutional layer with M input channels, N output channels, and a kernel size of H×W (for both spatial and frequency domain kernels) as follows: Specifically, the frequency domain kernel D has frequency components D that can be computed using the orthogonal DCT-II equation. ijhw The orthogonal DCT-II equation also normalizes the frequency component values. The frequency component D at subband i,j∈H×W... ijhw Represented as:

[0044]

[0045] Where (i,j) are the spatial coordinates within the frequency domain kernel D (of size H×W as described above), and (h,w) are the frequency coordinates within the frequency domain kernel D, and if k=0, then the constant c k =1, otherwise = Therefore, in the spatial domain, H and W are the spatial lengths of each dimension. In the frequency domain, H and W are the number of frequency components in each dimension.

[0046] In addition, a constant c is introduced. k Make D ijhwOrthogonal. It will be understood that orthogonality with the example DCT equation used in this paper refers to an arrangement where the dot product of the frequency domain kernels for every possible pair is zero. It will also be understood that equation (1) results in the frequency domain component D ijhw The normalization falls within the range of 0 to 1. Therefore, since DCTs are orthogonal, the D kernels are mutually orthogonal, such that the amplitude of each frequency domain kernel is 1. Note that the DCT basis of each frequency domain kernel is also orthogonal.

[0047] The initialization of the frequency domain kernel can be performed once for the first training iteration and, through an example, does not require subsequent updates via transformation equations. For subsequent iterations, the frequency domain kernel components are updated by alternatively applying reparameterized weights to the frequency domain kernel, as described below.

[0048] Subsequently, the frequency domain kernel is used to reparameterize the convolution kernel in the frequency domain. In other words, process 200 may include "generating at least one convolution kernel via a processor circuit system, the at least one convolution kernel having convolution coefficients generated by modifying the frequency domain kernel in the frequency domain" 206. Therefore, in this example, the convolution kernel is generated by weighted sum computation over the H×W subband (or the reparameterization of the convolution kernel is performed by weighted sum computation over the H×W subband), as follows:

[0049]

[0050] in, It is applied to the frequency domain kernel component D ijhw The reparameterization (RP) weights are used to modify or weight frequency components to form a weighted frequency domain (or DCT) kernel. For equation (2), the variable (i,j) indicates the RP weights and, consequently, which DCT kernel D is weighted, while (h,w) indicates the coordinates of the frequency components summed and fixed for equation (2) within the DCT kernel. This is achieved through the following... Figure 3 In the example, the variable (i,j) is the RP weight V 0,0 To RP weight V 2,2 The subscript indicates that the DCT kernel is transformed into a spatial domain kernel, which has a value at the coordinates within the spatial domain kernel indicated by the RP subscript (i,j) and is zero at all other orientations within the spatial domain kernel. Therefore, for example, the RP weight V... 0,0 The weighted DCT kernel has a spatial domain kernel that has a value at orientation 0,0 and is zero at all other orientations. For each different combination of the input channels, output channels, and kernel coefficient positions on the convolutional kernel K, the convolutional kernel coefficients K are provided. mnhwThe weighted sum of frequency domain kernels (and further, the sum of basis functions, where each frequency domain kernel represents a basis function) effectively performs the inverse DCT. The result is a spatial domain coefficient K with reparameterization over m x n channels. mnhw The convolutional kernel K.

[0051] Therefore, equation (2) can be described as involving three general operations or stages: (a) generating RP weights V; (b) applying the RP weights to the frequency domain kernel; and (c) summing the weighted frequency domain kernels to generate the convolution kernel.

[0052] refer to Figure 3 A weighted sum setting 300 is provided to aid in interpreting the weighted sum equation (2) and stages (a) to (c). For example, initially, the input to equation (1) or other frequency transforms can be the initial spatial domain convolution kernel coefficients, such as in a single 3x3 convolution kernel, where the coefficients are randomly selected and have orientations (i,j). Here, a set 301 of nine uniformly numbered 3x3 frequency domain kernels 302 to 318 is generated for the transform equation of the Discrete Cosine Transform (DCT), although it is possible to use frequency domain kernels of different sizes and numbers, since multiple frequency domain kernels have the same size as the convolution kernel. Each frequency domain kernel 302 to 318 represents a different frequency and has different modes of frequency components 320. The frequency domain kernels are shown as DCT kernels from the lowest frequency 302 to the highest frequency 318, as indicated by different shading modes on each frequency domain kernel.

[0053] Regarding the RP weight V, from V 0,0 To V 2,2 Nine RP weights are shown, with one RP weight for each frequency domain kernel 302 to 318. The RP weights are specified using subscript indices from 0,0 to 2,2 to represent different frequency components. In other words, the subscript indices conventionally start from 0. Then, for example, V 1,1 Corresponding to the (1,1) frequency component or sub-band position. The above rule is repeated for each combination of input and output channels. Therefore, each pair of m-th input channels to different n-th output channels has an RP weight V, and the RP weight shown on setting 300 can actually be specified as V. m,n,0,0 .

[0054] In one form, RP weight V 0,0 To RP weight V 2,2 Initialize randomly using either negative or positive integers. (From RP weight V) 0,0 To RP weight V 2,2The initial random values ​​can be calculated as follows: First, the convolutional kernel K is randomly initialized, then the randomly initialized K is projected onto each frequency domain kernel D, and then the initial RP weight V corresponding to a specific frequency domain kernel D is the dot product between the randomly initialized convolutional kernel K and the specific frequency domain kernel D. The randomly initialized convolutional kernel K can be the same kernel used to generate the frequency domain kernels.

[0055] Once the RP weights V are generated, they are applied to the frequency domain kernels. Specifically, the RP weights corresponding to the frequency domain kernels are applied to each frequency component (e.g., V) within that kernel. 0,0 (Applied to individual or all frequency components in kernel 302). The application here is a direct multiplication, as shown in setting 300. Therefore, the RP weight V 0,0 To RP weight V 2,2 It is directly applied to frequency domain kernels 302 to 318, and in particular, frequency component 320 is directly modified in the frequency domain.

[0056] Once the weighted frequency domain kernels 302 to 318 are generated, they can be summed together into a single spatial domain convolutional kernel 322, thus completing the inverse DCT or transform process. This process can be performed by summing the weighted frequency components 320 at the same or corresponding frequency component orientations (or at the same frequency index (i,j) in (H,W)). For example, the weighted frequency component at orientation (0,0) on frequency domain kernel 304 will be added to all other weighted frequency components at orientation (0,0) on the other frequency domain kernels 302 to 318. This operation is repeated for each of the 3x3 positions on frequency domain kernels 302 to 318. The result is a single convolutional kernel 322 with coefficients 324 reparameterized in the frequency domain as mentioned. For one or more convolutional layers being used by the neural network being trained, this reparameterization can be repeated for each or a single convolutional kernel 322 to be used in the filters or groups 330 of the convolutional kernel 322.

[0057] By an alternative form, it will be understood that by determining the transform basis functions, modifying the basis functions using randomly selected RP weights, and summing the weighted basis to form the convolution kernel coefficients, the calculation of equations (1) and (2) for initializing the frequency domain kernel and RP weights can be performed simply.

[0058] Subsequently, the RP weights are updated by the neural network training units. Specifically, the loss function used for training can be used to update the RP weights in each or individual iterations to minimize the loss, which in turn causes the RP weight updates to alter the distribution in low-frequency and high-frequency fits. Therefore, this mechanism has the effect of reducing the emphasis on fitting low-frequency image data in earlier iterations and increasing the generally more uniform distribution of low-frequency and high-frequency image data fits throughout iterations or over time. This mechanism can also be described as increasing the emphasis on or prioritizing high-frequency image data in earlier iterations compared to training without reparameterization.

[0059] A concrete example of such a loss function could be the mean squared error (MSE) between the recovered image and the target image, but it could also be the mean absolute difference (MSD), the structural similarity index (SSIM), or other functions.

[0060] As mentioned above, throughout the training iterations, with a more uniform distribution of processing for low-frequency and high-frequency image data (where reparameterization occurs directly in the frequency domain), training is accelerated by significantly reducing the number of training iterations required to achieve sufficient convergence.

[0061] refer to Figure 4 According to at least some implementations of this disclosure, an example runtime process 400 for image processing using a neural network trained rapidly through frequency-aware reparameterization is arranged, and may include one or more operations 402-408 that are typically evenly numbered. Furthermore, reference may be made herein to at least... Figure 1 and Figures 12-14 The system or device is described as 100, 1200, 1300 or 1400 to describe process 400.

[0062] Process 400 may include "receiving image data" 402, and "receiving image data" 402 may involve receiving image data for at least one image. Therefore, image data can be provided for individual photos or frames of one or more video sequences. As mentioned, there are no particular limitations on the type of convolutional neural network, the neural network architecture, the application or program being executed, etc., as long as the neural network performs some kind of image processing and receives some version of image data as input. A convolutional neural network is any network having at least one convolutional layer that uses at least one convolutional kernel.

[0063] This operation can also include performing any preprocessing before the image data is input into the neural network. Any preprocessing can include any desired preprocessing, such as depigmentation, scaling, enhancement, denoising, decompression, any other artifact removal operations, etc.

[0064] Process 400 may include "inputting a version of the image data into a neural network having at least one convolutional layer" 404. If the neural network uses at least one convolutional layer, this at least one convolutional layer further uses at least one convolutional kernel with coefficients. The size, type, and other specifications of the convolutional kernel are not particularly limited, as long as rapid training can be performed to generate the convolutional kernels described herein. Here, "version of the image data" refers to any modifications that may be applied before the image data is input into the neural network, such as normalization, grayscale changes, transformations, conversions, etc.

[0065] Process 400 may include "applying at least one convolutional kernel to a version of image data at at least one convolutional layer" 406. The version of image data here may or may not be the same as the version of image data input to the neural network. Here, the version of image data may at least refer to input image data values ​​or intermediate neural network values ​​output from layers preceding the current convolutional layer, which has at least one convolutional kernel providing input to the current convolutional layer. In one form, the convolutional kernel can be applied by traversing the surface of the convolutional layer according to kernel parameters to generate the output values ​​of the convolutional layer.

[0066] Process 400 may further include "wherein at least one convolutional kernel has coefficients generated by modifying the frequency components of the frequency domain kernel" 408. Operation 408 refers to the training process 200 described above, which uses reparameterized weights to modify the frequency domain kernel, and, by way of an example, uses the weighted sum equation described above. As described, the reparameterized weights reduce the emphasis on fitting low-frequency image data and perform better fitting of high-frequency image data in earlier iterations. This mechanism throughout the iterations creates a more uniform or homogeneous distribution of high-frequency and low-frequency image data, thereby accelerating the training and generation of at least one convolutional kernel having coefficients generated by frequency-aware reparameterization.

[0067] experiment

[0068] Experiment 1: This section presents a schematic (or simplified) example comparing the frequency domain behavior of the publicly disclosed FAR method and basic (or regular or non-frequency-based) convolutions. For this experiment, a three-layer CNN was trained using two 512 intermediate channels activated by ReLU to recover a 256×256 image via overfitting. The input image was compressed to JPEG quality 15. The network was trained using the Adaptive Moment Estimation (Adam) Optimizer Neural Network Training Algorithm with a learning rate of 1e-5 for 100,000 iterations. Both the publicly disclosed FAR method and the basic convolution have a kernel size of 3×3.

[0069] refer to Figure 5The diagram shows a set of 500 graphs, each targeting a different frequency sub-band. This set of 500 graphs is used to examine how different frequency components in the reconstructed image change over time during training iterations. The image is decomposed using 4×4 DCT, and the mean absolute value frequency component at each sub-band is determined as a function of training iterations.

[0070] Set 500 shows that images recovered by basic convolutions follow a typical spectral bias, where high-frequency components converge much slower than low-frequency components. Network training based on publicly available FAR methods is much less affected by spectral bias. The mean absolute component (MAS) of most high-frequency components converges much faster than the MAS of networks based on basic convolutions.

[0071] refer to Figure 6 Graph 600 illustrates the noise (PSNR) along the iterative tracking and how different frequency-related behaviors can be visually observed in the resulting images. For example, base image 602 and FAR image 604 are generated around iteration 10. Although there is color bias, FAR image 604 still has superior sharpness compared to base image 602. Similarly, at iteration 100000, FAR image 608 outperforms base image 606. Therefore, the images recovered by the proposed FAR method almost always have higher sharpness, especially in the early iterations.

[0072] refer to Figure 7 This shows a set of 700 graphs, each graph targeting a different RP weight V. 0,0 To RP weight V 2,2 Furthermore, this is applied to different frequency subbands. Set 700 is used to examine training dynamics by visualizing RP weight updates in the frequency domain. Therefore, each plot in Set 700 shows the average absolute change of the DCT coefficients of the 3×3 convolution in each iteration. The plots show that the magnitude of the RP weights is updated more uniformly across all subbands, while the basic convolution updates low-frequency subbands more than high-frequency subbands in most iterations. Figures 5-7 The results shown in the paper demonstrate that the proposed FAR method learns information better when low-frequency and high-frequency processing are more evenly distributed along the iterations.

[0073] Experiment 2: In the second experiment, image compression was performed simultaneously with overfit-based image restoration. This involved overfitting the residuals from image compression using a neural network trained with the proposed FAR training method disclosed herein.

[0074] refer to Figure 8A system 800 for image compression used in experiments is shown. The inputs are images 804, 806, and 808 at progressively decreasing downsampling resolutions. System 800 has a neural network architecture 802, which has volume blocks 810, 812, and 814, cascaded layers (C) 816, convolutional layers (Cv) 818, and summation operations 820 for each image 804, 806, and 808, respectively. Each volume block 810, 812, and 814 has two cascaded repeating sets, each repeating set having, in sequence, a first convolutional layer (Cv) 822 or 828, a second instance normalization (N) layer 824 or 830, and a third rectified linear (ReLU) layer 826 or 832. Each volume block 810, 812, and 814 has another convolutional layer (Cv) 834 at the output of the volume block.

[0075] In the operation, three input images are fed into blocks 810, 812, and 814. Then, the output values ​​from the blocks and the output values ​​from the lower-resolution images are scaled up to match the resolution of image 804. The scaled and unscaled block outputs are then cascaded through layer C 816. Next, the cascaded images are combined with the original highest-resolution image 804 at a summation operation 820 to generate the output image 822.

[0076] Each image was compressed using a conventional codec, overfitted, and compared using networks employing the proposed FAR method and basic convolutions, respectively. Evaluations were conducted using compression codecs for JPEG (cjpeg 9e), High Efficiency Image File (HEIF) (High Efficiency Video Coding (HEVC), libheif 1.12), and Universal Video Coding (VVC) (Intra-Frame Mode, Video Test Model (VTM) 19.0), evaluating Kodak, Tecnick, and Learning-Oriented Image Compression Challenge (CLIC) 2020 neural network applications. The default number of channels for JPEG, HEIF, and VVC were 64, 32, and 16, respectively. For Kodak, the default number of channels was halved due to the smaller image size. For JPEG and HEIF, images were provided at quality levels of 15, 40, 65, and 90. For VVC, the quantization parameters (QP) were 37, 32, 27, and 22. The pixel formats evaluated were YUV420 and YUV444.

[0077] The training objective of Experiment 2 was to determine the mean squared error between the restored image and the original image. The network was trained for 200 iterations using the Adam optimizer, with an L1 penalty of 1e-3 and a linearly decaying learning rate starting at 0.05. After training, the FAR RP weights were quantized and compressed using DeepCABAC. The quantization step size was calculated as |V|max / L, where L = 127 for the experiments. The quantized weights were then reloaded to measure the PSNR, the multi-scale structural similarity index metric (MS-SSIM) of the corresponding RD (rate-distortion) curves, and the Björtegaard delta (BD) rate. The results are shown in the table below.

[0078] Table 1: BD Rate of Image Recovery Based on Overfitting. Bold text indicates better results compared to the corresponding codecs, and underlined text indicates worse results. This includes CLIC-M and CLIC-P, representing CLIC Mobile and CLIC Pro respectively.

[0079]

[0080] In all evaluations, the publicly disclosed FAR method demonstrates significantly better results compared to the basic convolution. In many cases, the basic convolution fails to improve RD (underlined in the table), while the proposed FAR method only fails to improve MS-SSIM for Kodak images compressed with VVC (4:4:4). Therefore, aside from cases where small images are compressed by significantly optimized image codecs, the proposed FAR method holds promise for optimized image compression beyond traditional image codecs. For modern codecs such as VVC, the scheme is able to leverage native support for practical image compression.

[0081] refer to Figure 9 Graph 900 shows an example RD curve plotted from the results of Experiment 2, where JPEG was used in the case of both the FAR method and the base method disclosed in this paper, as well as the original image.

[0082] refer to Figures 10A-10H The images illustrate the quality improvement resulting from using the proposed FAR method disclosed herein. Specifically, image 1000 shows the block effect, while image 1002 shows the original image, image 1004 shows the FAR-generated image (JPEG 4:2:0 + FAR: 35.67 dB, 0.330 bpp (bits per pixel)), and image 1006 shows the base image (JPEG 4:2:0 33.05 dB, 0.337 bpp).

[0083] Image 1010 shows ringing artifacts, while image 1012 shows the original image, image 1014 shows the image generated by FAR (HEIF 4:2:0+FAR:34.37dB, 0.0725bpp), and image 1016 shows the base image (HEEF 4:2:033.85dB, 0.0730bpp).

[0084] Image 1020 shows the chroma subsampling artifacts, while image 1022 shows the original image, image 1024 shows the image generated by FAR (VVC 4:2:0+FAR:30.20dB, 0.190bpp), and image 1026 shows the base image (VVB4:2:028.60dB, 0.192bpp).

[0085] As shown in the figure, the publicly available FAR method provides the ability to more easily detect artifacts for removal using traditional codecs, thereby significantly improving visual quality.

[0086] refer to Figure 11 Figure 1100 further illustrates the convergence advantage of the publicly available FAR method. It shows the BD rates of FAR and the basic convolution on CLIC Pro with different total training iterations, where the publicly available FAR method exhibits better convergence than the basic convolution, especially with fewer total training iterations. Although the gap between the publicly available FAR and the basic convolution decreases with increasing training iterations, the publicly available FAR method is more practical because achieving the same BD rate typically requires less computation.

[0087] While the implementation of the example procedures discussed herein may include performing all the operations shown in the order described, this disclosure is not limited in this respect, and in various examples, the implementation of the example procedures herein may include only a subset of the operations shown, operations performed in a different order than that shown, or additional operations.

[0088] Furthermore, any one or more of the operations discussed herein can be performed in response to instructions provided by one or more computer program products. Such program products may include signal-bearing media that provide instructions, when executed by, for example, a processor, to provide the functionality described herein. Computer program products may be provided in any form of one or more machine-readable media. Thus, for example, a processor including one or more graphics processing units or processor cores may perform one or more blocks of the exemplary processes described herein in response to program code and / or instructions or instruction sets transmitted to the processor from one or more machine-readable media. Typically, machine-readable media may deliver software in the form of program code and / or instructions or instruction sets that enable any of the devices and / or systems described herein to implement at least a portion of the operations discussed herein and / or any part of the devices, systems, or any modules or components as discussed herein.

[0089] As used in any implementation described herein, the term "module" refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. Software may be embodied as software packages, code, and / or instruction sets or instructions, and "hardware" as used in any implementation described herein may (alone or in any combination) include, for example, hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, and / or firmware storing instructions executed by the programmable circuitry. Modules may collectively or individually be embodied as parts of a larger system (e.g., an integrated circuit (IC), a system-on-a-chip (SoC), etc.).

[0090] As used in any implementation described herein, the term "logic unit" refers to any combination of firmware logic and / or hardware logic configured to provide the functions described herein. A logic unit may be embodied collectively or individually as a part of a larger system (e.g., an integrated circuit (IC), a system-on-a-chip (SoC), etc.). For example, a logic unit may be embodied in the logic circuitry of the firmware or hardware used to implement the codec system discussed herein. Those skilled in the art will recognize that operations performed by hardware and / or firmware may alternatively be implemented via software, which may be embodied as software packages, code, and / or instruction sets or instructions, and that a logic unit may also utilize a portion of software to implement its functions.

[0091] As used in any implementation described herein, the term "component" can refer to a module or logical unit, as stated above. Therefore, the term "component" can refer to any combination of software logic, firmware logic, and / or hardware logic configured to provide the functionality described herein. For example, those skilled in the art will recognize that operations performed by hardware and / or firmware can alternatively be implemented via a software module, which can be embodied as a software package, code, and / or instruction set, and that logical units can also utilize a portion of software to implement their functionality. Components herein can also refer to processors and other specific hardware devices.

[0092] As used in any implementation herein, the term "circuit" or "circuit system" may (alone or in any combination) include or form, for example, a hardwired circuit system, a programmable circuit system such as a computer processor including one or more separate instruction processing cores, a state machine circuit system, and / or firmware storing instructions executed by the programmable circuit system. A circuit system may include a processor ("processor circuit system") and / or a controller configured to execute one or more instructions to perform one or more operations described herein. Instructions may be embodied, for example, as an application, software, firmware, etc., configured to cause the circuit system to perform any of the foregoing operations. Software may be embodied as software packages, code, instructions, instruction sets, and / or data recorded on a computer-readable storage device. Software may be embodied or implemented as including any number of processes, and processes may further be embodied or implemented hierarchically as including any number of threads, etc. Firmware may be embodied as code, instructions, or instruction sets and / or data hard-coded (e.g., non-volatile) in a memory device. A circuit system can be collectively or individually embodied as a part of a larger system (e.g., an integrated circuit (IC), an application-specific integrated circuit (ASIC), a system-on-a-chip (SoC), a desktop computer, a laptop computer, a tablet computer, a server, a smartphone, etc.). Other implementations can be implemented as software executed by a programmable control device. In this context, the term "circuit" or "circuit system" is intended to include a combination of software and hardware, such as a programmable control device or a processor capable of executing software. As described herein, various implementations can be implemented using hardware elements, software elements, or any combination thereof that form a circuit, circuit system, or processor circuit system. Examples of hardware elements can include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, microchips, chipsets, etc.

[0093] refer to Figure 12According to at least some implementations of this disclosure, an example image processing system 1200 for rapidly training neural networks is provided. (As in...) Figure 12 As shown, system 1200 may include: a processor formed by processor circuitry 1250, which may or may not include one or more image signal processors (ISP) 1252 and other accelerators 1242; and a memory 1254 having one or more neural network buffers 1244.

[0094] Additionally, as shown in the figure, system 1200 may have a logic unit or module 1204 including a preprocessing unit 1208, which can process image data from imaging device 1202 (such as one or more cameras), which is on a device having the module or communicates with module 1204. The logic unit or module 1204 may have an image processing unit 1206 with preprocessing unit 1208, which can preprocess image data from camera 1202 to provide the image to the neural network application 104 as described above. Figure 1 The image processing unit 1206 may also have a neural network (NN) training unit 120. Figure 1 The neural network training unit 120 performs frequency-aware reparameterization of the convolutional kernel by applying reparameterized weights to the frequency domain kernel in the frequency domain, as described above. Other image processing applications 1212 (whether for image transmission, modification, or generation) can use the output of the neural network application.

[0095] The system 1200 may also have an antenna 1256 for transmitting or receiving image data, an encoder 1258 for encoding or decoding image data, and a display 1260 with a screen capable of displaying an image 1262.

[0096] The names of the units here may be the same as or similar to the names of the units on device 100, such that the units here perform the same or similar tasks as the units on device 100. Otherwise, the operation of the units will be described as needed during the detailed description herein.

[0097] In some examples, one or more of the operations of processes 200 and 400 may be implemented via ISP 1252 or accelerator 1242. In other examples, one or more of the operations are implemented via a central processing unit, image processing unit, image processing pipeline, image signal processor 1252, etc., forming the processor circuit system 1250. In some examples, the rapid training of the neural network or one or more of the operations is implemented in hardware as a system-on-a-chip (SoC) or other dedicated hardware or other shared hardware. In some examples, one or more of the rapid training is implemented in hardware via a field-programmable gate array (FPGA).

[0098] The processor circuitry 1250, image signal processor 1252, and accelerator 1242 may include any number and type of image or graphics processing units that can provide the operations discussed herein. Such operations may be implemented via software or hardware, or a combination thereof. For example, image signal processor 1252 may include circuitry dedicated to manipulating and / or analyzing images obtained from memory 1254. Central processing unit 1250 may include any number and type of processing units or modules that can provide control and other high-level functions for system 1200 and / or provide any operations discussed herein.

[0099] Memory 1254 can be any type of memory, such as volatile memory (e.g., static random access memory, dynamic random access memory, etc.) or non-volatile memory (e.g., hard disk drive, flash memory, etc.). In a non-limiting example, memory 1254 can be implemented by a cache memory. In an implementation, one or more or a portion of the fast training is implemented via an execution unit (EU) of the processor circuit system 1250. For example, the EU may include programmable logic or circuit systems, such as one or more logic cores that can provide various programmable logic functions. In an implementation, one or more or a portion of the fast training is implemented via dedicated hardware (such as a fixed-function circuit system, etc.). The fixed-function circuit system may include dedicated logic or circuit systems and may provide a set of fixed-function entry points that can be mapped to dedicated logic for a fixed purpose or function.

[0100] The various components of the systems described herein can be implemented in software, firmware, and / or hardware and / or any combination thereof. For example, the various components of the devices or systems discussed herein may be provided at least in part by hardware such as a computing system-on-a-chip (SoC) found in computing systems (e.g., a smartphone or camera array). Those skilled in the art will recognize that the systems described herein may include additional components not depicted in the corresponding figures. For example, the systems discussed herein may include additional components not depicted for clarity.

[0101] refer to Figure 13 Example system 1300 is arranged according to at least some implementations of this disclosure. In various implementations, system 1300 may be a mobile device system, but system 1300 is not limited to this context. For example, system 1300 or portions thereof may be integrated into servers, personal computers (PCs), notebook computers, ultranotebook computers, tablet computers, touchpads, portable computers, handheld computers, PDAs, personal digital assistants (PDAs), cellular phones, combined cellular phones / PDAs, televisions, smart devices (e.g., smartphones, smart tablets, or smart TVs), mobile internet devices (MIDs), messaging devices, data communication devices, cameras (e.g., fully automatic portable cameras, super zoom cameras, digital single-lens reflex (DSLR) cameras), surveillance cameras, surveillance systems including cameras, etc. In other aspects, at least a portion of system 1300 may be on one or more servers.

[0102] In various implementations, system 1300 includes a platform 1302 coupled to display 1320. Platform 1302 can receive content from a content device such as content serving device 1330 or content delivery device 1340, or from other content sources such as image sensor 1313. For example, platform 1302 can receive image data, as discussed herein, from image sensor 1313 or any other content source. A navigation controller 1350, including one or more navigation features, can be used to interact with, for example, platform 1302 and / or display 1320. Each of these components is described in more detail below.

[0103] In various implementations, platform 1302 may include any combination of chipset 1305, processor 1310, memory 1312, antenna, storage device 1314, graphics subsystem 1315, application 1316, image signal processor (ISP) 1317, and / or radio device 1318. Chipset 1305 can provide communication between processor 1310, memory 1312, storage device 1314, graphics subsystem 1315, application 1316, image signal processor 1317, and / or radio device 1318. For example, chipset 1305 may include a storage device adapter (not depicted) capable of providing communication with storage device 1314.

[0104] The processor 1310 can be implemented as a Complex Instruction Set Computer (CISC) processor or a Reduced Instruction Set Computer (RISC) processor; an x86 instruction set compatible processor, a multi-core processor, or any other microprocessor or central processing unit (CPU). In various implementations, the processor 1310 can be a dual-core processor, a dual-core mobile processor, etc.

[0105] The memory 1312 can be implemented as a volatile memory device, such as, but not limited to, random access memory (RAM), dynamic random access memory (DRAM), or static RAM (SRAM).

[0106] Storage device 1314 can be implemented as a non-volatile storage device, such as, but not limited to, disk drives, optical disc drives, tape drives, internal storage devices, attached storage devices, flash memory, battery-backed SDRAM (synchronous DRAM), and / or network-accessible storage devices. In various implementations, for example, when including multiple hard disk drives, storage device 1314 may include techniques for enhancing storage performance protection for valuable digital media.

[0107] The image signal processor 1317 can be implemented as a dedicated digital signal processor (DSP) for image processing, etc. In some examples, the image signal processor 1317 can be implemented based on a single instruction multiple data (SID) architecture or a multiple instruction multiple data (MIDD) architecture, etc. In some examples, the image signal processor 1317 is characterized as a media processor. As discussed herein, the image signal processor 1317 can be implemented based on a system-on-a-chip (SoC) architecture and / or a multi-core architecture.

[0108] The graphics subsystem 1315 can perform image processing, such as still or video processing, for display. The graphics subsystem 1315 can be, for example, a graphics processing unit (GPU) or a visual processing unit (VPU). An analog or digital interface can be used to communicatively couple the graphics subsystem 1315 and the display 1320. For example, the interface can be any of a high-resolution multimedia interface, a display port, wireless HDMI, and / or wireless HD compatible technologies. The graphics subsystem 1315 can be integrated into the processor 1310 or the chipset 1305. In some implementations, the graphics subsystem 1315 can be a standalone device communicatively coupled to the chipset 1305.

[0109] The graphics and / or video processing techniques described herein can be implemented in various hardware architectures. For example, graphics and / or video functions can be integrated within a chipset. Alternatively, discrete graphics and / or video processors can be used. As yet another implementation, graphics and / or video functions can be provided by a general-purpose processor, including a multi-core processor. In a further implementation, the functions can be implemented in a consumer electronic device.

[0110] Radio device 1318 may include one or more radio devices capable of transmitting and receiving signals using a variety of suitable wireless communication technologies. Such technologies may involve communication across one or more wireless networks. Example wireless networks include (but are not limited to) wireless local area networks (WLANs), wireless personal area networks (WPANs), wireless metropolitan area networks (WMANs), cellular networks, and satellite networks. When communicating across such networks, radio device 1318 may operate according to one or more applicable standards of any version.

[0111] In various implementations, display 1320 may include any type of television monitor or display. Display 1320 may include, for example, a computer screen, a touchscreen display, a video monitor, a television-like device, and / or a television. Display 1320 may be digital and / or analog. In various implementations, display 1320 may be a holographic display. Additionally, display 1320 may be a transparent surface capable of receiving visual projections. Such projections may convey various forms of information, images, and / or objects. For example, such projections may be visual overlays used in mobile augmented reality (MAR) applications. Under the control of one or more software applications 1316, platform 1302 may display user interface 1322 on display 1320.

[0112] In various implementations, content service device 1330 can be hosted by any national, international, and / or independent service, and therefore can be accessed via, for example, the Internet, platform 1302. Content service device 1330 can be coupled to platform 1302 and / or display 1320. Platform 1302 and / or content service device 1330 can be coupled to network 1360 to transmit media information to and from network 1360 (e.g., sending media information to network 1360 and / or receiving media information from network 1360). Content delivery device 1340 can also be coupled to platform 1302 and / or display 1320.

[0113] Image sensor 1313 may include any suitable image sensor capable of providing image data based on a scene. For example, image sensor 1313 may include a semiconductor charge-coupled device (CCD) based sensor, a complementary metal-oxide-semiconductor (CMOS) based sensor, an N-type metal-oxide-semiconductor (NMOS) based sensor, etc. For example, image sensor 1313 may include any device capable of detecting information about the scene to generate image data.

[0114] In various implementations, content service device 1330 may include a cable TV box, personal computer, network, telephone, internet-enabled device or appliance capable of delivering digital information and / or content, and any other similar device capable of transmitting content unidirectionally or bidirectionally between the content provider and platform 1302 and / or display 1320 via network 1360 or directly. It will be appreciated that content can be transmitted unidirectionally and / or bidirectionally to any of the various components in system 1300 and the content provider via network 1360. Examples of content may include any media information, including, for example, video, music, medical and gaming information.

[0115] Content service device 1330 can receive content such as cable television programs, including media information, digital information, and / or other content. Examples of content providers may include any cable or satellite television or radio or internet content provider. The examples provided are not intended to limit the implementation methods according to this disclosure in any way.

[0116] In various implementations, platform 1302 may receive control signals from navigation controller 1350, which has one or more navigation features. For example, the navigation features of navigation controller 1350 may be used to interact with user interface 1322. In various implementations, navigation controller 1350 may be a pointing device, which may be a computer hardware component (specifically, a human-machine interface device) that allows users to input spatial (e.g., continuous and multidimensional) data into a computer. Many systems, such as graphical user interfaces (GUIs) and televisions and monitors, allow users to control and provide data to a computer or television using physical gestures.

[0117] Movement of navigation features on navigation controller 1350 can be replicated on a display by the movement of a pointer, cursor, focus ring, or other visual indicator displayed on a display (e.g., display 1320). For example, under the control of software application 1316, navigation features on navigation controller 1350 can be mapped to, for example, virtual navigation features displayed on user interface 1322. In various implementations, navigation controller 1350 may not be a separate component but may be integrated into platform 1302 and / or display 1320. However, this disclosure is not limited to the elements or context shown or described herein.

[0118] In various implementations, the driver (not shown) may include technology that allows a user to immediately turn the platform 1302 on and off; for example, when enabled, the television can be turned on and off using a touch of a button after initial startup. Even when the platform is “off,” program logic may allow the platform 1302 to stream content to a media adapter or other content service device 1330 or content delivery device 1340. Additionally, for example, the chipset 1305 may include hardware and / or software support for 5.1 surround sound audio and / or high-resolution 7.1 surround sound audio. The driver may include a graphics driver for an integrated graphics platform. In various implementations, the graphics driver may include a Peripheral Component Interconnect (PCI) Fast Graphics Card.

[0119] In various implementations, any one or more of the components shown in system 1300 can be integrated. For example, platform 1302 and content service device 1330 can be integrated, or platform 1302 and content delivery device 1340 can be integrated, or platform 1302, content service device 1330, and content delivery device 1340 can be integrated. In various implementations, platform 1302 and display 1320 can be integrated units. For example, display 1320 and content service device 1330 can be integrated, or display 1320 and content delivery device 1340 can be integrated. These examples are not intended to limit this disclosure.

[0120] In various implementations, System 1300 can be implemented as a wireless system, a wired system, or a combination of both. When implemented as a wireless system, System 1300 may include components and interfaces suitable for communication via a wireless shared medium, such as one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, etc. Examples of a wireless shared medium may include portions of the wireless spectrum, such as the RF spectrum. When implemented as a wired system, System 1300 may include components and interfaces suitable for communication via a wired communication medium, such as input / output (I / O) adapters, physical connectors for connecting I / O adapters to corresponding wired communication media, network interface cards (NICs), disk controllers, video controllers, audio controllers, etc. Examples of a wired communication medium may include wires, cables, metal leads, printed circuit boards (PCBs), backplanes, switch structures, semiconductor materials, twisted pairs, coaxial cables, optical fibers, etc.

[0121] Platform 1302 can establish one or more logical or physical channels to transmit information. Information may include media information and control information. Media information can refer to any data representing content for a user. Examples of content may include data such as that from voice conversations, video conferencing, streaming video, email (“email”) messages, voicemail messages, alphanumeric symbols, graphics, images, video, text, etc. Data from voice conversations may be, for example, voice information, silence periods, background noise, comfort noise, tone, etc. Control information can refer to any data representing commands, instructions, or control words for an automated system. For example, control information may be used to route media information through the system or to instruct nodes to process media information in a predetermined manner. However, the implementation is not limited to... Figure 13 The elements or context shown or described in the text.

[0122] As mentioned above, system 1300 can be embodied in varying physical styles or physical specifications. Figure 14 The illustration depicts an example small form factor device 1400 arranged according to at least some implementations of this disclosure. In some examples, system 1200 or system 1300 may be implemented via device 1400. In other examples, other systems, components, or modules, or portions thereof, discussed herein may be implemented via device 1400. In various implementations, for example, device 1400 may be implemented as a mobile computing device with wireless capabilities. For example, a mobile computing device may refer to any device having a processing system and a mobile power source or power supply (such as one or more batteries).

[0123] Examples of mobile computing devices may include personal computers (PCs), laptops, ultrabooks, tablets, touchpads, portable computers, handheld computers, PDAs, personal digital assistants (PDAs), cellular phones, combined cellular phones / PDAs, smart devices (e.g., smartphones, smart tablets, or smart mobile TVs), mobile internet devices (MIDs), messaging devices, data communication devices, cameras (e.g., fully automatic compact cameras, super zoom cameras, digital single-lens reflex (DSLR) cameras), and the like.

[0124] Examples of mobile computing devices may also include computers deployed as implemented by motor vehicles or robots, or worn by individuals, such as wrist computers, finger computers, ring computers, glasses computers, belt-clipped computers, armband computers, shoe computers, clothing computers, and other wearable computers. In various implementations, for example, a mobile computing device may be implemented as a smartphone capable of performing computer applications and voice and / or data communications. While some implementations can be described by way of example using a mobile computing device implemented as a smartphone, it is understood that other implementations may also be implemented using other wireless mobile computing devices. The implementation is not limited in this context.

[0125] As in Figure 14 As shown, device 1400 may include a housing having a front portion 1401 and a rear portion 1402. Device 1400 includes a display 1404, an input / output (I / O) device 1406, a camera 1421, a camera 1422, and an integrated antenna 1408. In some implementations, device 1400 does not include cameras 1421 and 1422, and device 1400 obtains input image data (e.g., any input image data discussed herein) from another device. Device 1400 may also include navigation features 1412. I / O device 1406 may include any suitable I / O device for inputting information into a mobile computing device. Examples of I / O device 1406 may include an alphanumeric keypad, numeric keypad, touchpad, input keys, buttons, switches, microphone, speaker, voice recognition device, and software, etc. Information may also be input into device 1400 via microphone 1414, or may be digitized by a voice recognition device. As shown in the figure, device 1400 may include camera 1421, camera 1422 and flash 1410 integrated into the rear 1402 (or other location) of device 1400.

[0126] Various implementations can be achieved using hardware components, software components, or a combination of both. Examples of hardware components can include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software can include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application programming interfaces (APIs), instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. The determination that an implementation is carried out using hardware components and / or software components can vary depending on any number of factors, such as desired computational speed, power level, thermal tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.

[0127] One or more aspects of at least one implementation can be implemented by representative instructions stored on a machine-readable medium, which represent various logics within a processor and, when read by a machine, cause the machine to manufacture the logic to perform the techniques described herein. This representation (referred to as an "IP core") can be stored on a tangible machine-readable medium and supplied to various customers or manufacturing facilities for loading into manufacturing machines that actually manufacture the logic or processor.

[0128] While certain features set forth herein have been described with reference to various implementations, this description is not intended to be construed in a limiting sense. Therefore, various modifications to the implementations described herein, as well as other implementations that will be apparent to those skilled in the art to which this disclosure pertains, are considered to be within the spirit and scope of this disclosure.

[0129] The following examples involve additional implementation methods.

[0130] In Example 1, an image processing method includes: receiving initial convolutional kernel coefficients to train a neural network having at least one convolutional layer; converting the coefficients into a frequency domain kernel; and generating at least one convolutional kernel having convolutional coefficients generated by modifying the frequency domain kernel in the frequency domain via a processor circuit system.

[0131] In Example 2, the subject of Example 1 is repeated, where each frequency domain kernel is associated with a different frequency.

[0132] In Example 3, the subject of Example 1 or Example 2 is used, where the frequency components of the frequency domain kernel are modified so that low-frequency and high-frequency image data are used more evenly throughout the training iterations of the neural network compared to training iterations in the spatial domain.

[0133] In Example 4, the subject of any of Examples 1-3 is used, where the Discrete Cosine Transform (DCT) is applied to the initial convolution kernel coefficients to generate the frequency domain kernel.

[0134] In Example 5, the subject of Example 4 is used, where the DCT is an orthogonal DCT-II.

[0135] In Example 6, the subject of any one of Examples 1-5 is used, where the values ​​of the initial convolutional kernel coefficients are randomly selected.

[0136] In Example 7, the subject of any of Examples 1-6 is provided, wherein generating includes generating a weighted frequency domain kernel, and generating a weighted frequency domain kernel includes applying reparameterized weights to the frequency domain kernel.

[0137] In Example 8, the subject of Example 7 is repeated, where at least one reparameterized weight is applied to each frequency component in a separate frequency domain kernel.

[0138] In Example 9, the topic of Example 7 is repeated, where the reparameterized weights are initially determined randomly.

[0139] In Example 10, the topic of Example 7 is discussed, where reparameterized weights are adjusted iteratively during the training of the neural network.

[0140] In Example 11, a computer-implemented system includes: a memory; and a processor circuitry communicatively coupled to the memory, the processor circuitry being arranged to operate by: receiving image data; inputting a version of the image data into a neural network having at least one convolutional layer; and applying at least one convolutional kernel to the version of the image data at the at least one convolutional layer, wherein the at least one convolutional kernel has coefficients generated by modifying the frequency components of a frequency-domain kernel in the frequency domain.

[0141] In Example 12, the subject of Example 11 is discussed, wherein generating includes generating a weighted frequency domain kernel, which includes applying reparameterized weights to individual frequency domain kernels and combining the weighted frequency domain kernels to form a single convolutional kernel.

[0142] In Example 13, the subject of Example 12 is used, wherein the combination includes: summing the frequency components of multiple weighted frequency domain kernels; and summing the frequency components of the same coordinates within multiple weighted frequency domain kernels.

[0143] In Example 14, the subject of Example 12 is repeated, where the frequency domain kernel has a frequency component that is modified by reparameterizing the weights to increase the processing of high-frequency image data in earlier iterations used to train the neural network.

[0144] In Example 15, the subject of Example 14 is repeated, where the reparameterized weights are randomly initialized and modified during the iteration of training the neural network.

[0145] In Example 16, at least one non-transitory machine-readable medium includes a plurality of instructions that, in response to execution on a computing device, cause the computing device to operate in such a way as to: receive initial convolutional kernel coefficients to train a neural network having at least one convolutional layer; convert the coefficients into a frequency domain kernel; and generate at least one convolutional kernel having convolutional coefficients generated by modifying the frequency domain kernel in the frequency domain.

[0146] In Example 17, the subject of Example 16 is repeated, where the frequency components of the frequency domain kernel are modified in the frequency domain.

[0147] In Example 18, the subject of Example 16 or Example 17 is used, where the frequency components of the frequency domain kernel are normalized.

[0148] In Example 19, the subject of any of Examples 16-18 is provided, wherein generation includes generating a weighted frequency domain kernel, and generating a weighted frequency domain kernel includes applying reparameterized weights to a separate frequency domain kernel.

[0149] In Example 20, the subject of Example 19 is discussed, where reparameterized weights are arranged to generate a more uniform temporal distribution of processing high-frequency and low-frequency image data during iterations used to train the neural network.

[0150] In one example, the device, apparatus, or system includes facilities for performing methods according to any of the above implementations.

[0151] In another example, at least one machine-readable medium includes a plurality of instructions that, in response to execution on a computing device, cause the computing device to perform a method according to any of the above implementations.

[0152] It will be recognized that the implementation is not limited to the described implementation, but can be practiced in modified and alternative ways without departing from the scope of the appended claims. For example, the above implementation may include specific combinations of features. However, the above implementation is not limited to this, and in various implementations, the above implementation may include only a subset of such features, different orders of implementation of such features, different combinations of implementation of such features, and / or additional features that differ from those expressly listed. Therefore, the scope of the implementation should be determined with reference to the appended claims and the full scope of the equivalents given by such claims.

Claims

1. An image processing method, comprising: Receive initial convolutional kernel coefficients to train a neural network with at least one convolutional layer; Convert the coefficients into a frequency domain kernel; as well as The processor circuitry generates at least one convolution kernel, which has convolution coefficients generated by modifying the frequency domain kernel in the frequency domain.

2. The method according to claim 1, wherein, Each frequency domain kernel is associated with a different frequency.

3. The method according to claim 1 or 2, wherein, The frequency components of the frequency domain kernel are modified so that, compared to training iterations in the spatial domain, low-frequency and high-frequency image data are used more evenly throughout the training iterations of the neural network.

4. The method according to any one of claims 1-3, wherein, The Discrete Cosine Transform (DCT) is applied to the initial convolution kernel coefficients to generate the frequency domain kernel.

5. The method according to claim 4, wherein, The DCT is an orthogonal DCT-II.

6. The method according to any one of claims 1-5, wherein, The values ​​of the initial convolution kernel coefficients are randomly selected.

7. The method according to claim 1, wherein, The generation includes generating a weighted frequency domain kernel, which includes applying reparameterized weights to the frequency domain kernel.

8. The method according to claim 7, wherein, At least one reparameterized weight is applied to each frequency component in a separate frequency domain kernel.

9. The method according to claim 7, wherein, The reparameterized weights are initially determined randomly.

10. The method according to claim 7, wherein, The reparameterized weights are adjusted iteratively during the training of the neural network.

11. A computer-implemented system, comprising: Memory; as well as A processor circuitry communicatively coupled to the memory, the processor circuitry being arranged to operate in such a manner as follows: Receive image data; Image data is input into a neural network with at least one convolutional layer; as well as At least one convolutional kernel is applied to the image data at the at least one convolutional layer. The at least one convolutional kernel has coefficients generated by modifying the frequency components of the frequency domain kernel in the frequency domain.

12. The system according to claim 11, wherein, The generation includes generating weighted frequency domain kernels, which involves: applying reparameterized weights to individual frequency domain kernels; and combining the weighted frequency domain kernels to form a single convolutional kernel.

13. The system according to claim 12, wherein, The combination includes: summing the frequency components of a plurality of weighted frequency domain kernels; and summing the frequency components at the same coordinates within the plurality of weighted frequency domain kernels.

14. The system according to claim 12, wherein, The frequency domain kernel has frequency components that are modified by reparameterizing weights to increase the processing of high-frequency image data in earlier iterations used to train the neural network.

15. The system according to claim 14, wherein, The reparameterized weights are randomly initialized and modified during the iterations of training the neural network.

16. At least one non-transitory machine-readable medium, comprising a plurality of instructions that, in response to execution on a computing device, cause the computing device to operate in such a way as: Receive initial convolutional kernel coefficients to train a neural network with at least one convolutional layer; Convert the coefficients into a frequency domain kernel; and The processor circuitry generates at least one convolution kernel, which has convolution coefficients generated by modifying the frequency domain kernel in the frequency domain.

17. The medium according to claim 16, wherein, The frequency components of the frequency domain kernel are modified in the frequency domain.

18. The medium according to claim 16, wherein, The frequency components of the frequency domain kernel are normalized.

19. The medium according to claim 16, wherein, The generation includes generating a weighted frequency domain kernel, which includes applying reparameterized weights to a single frequency domain kernel.

20. The medium according to claim 19, wherein, The reparameterized weights are arranged to generate a more uniform temporal distribution for processing high-frequency and low-frequency image data during iterations used to train the neural network.