Image processing method
By training and mixing the parameters of multiple neural network models, a hybrid neural network model is generated, which solves the problem of high hardware costs caused by multiple neural network processors. It achieves the optimization of the effect of a single neural network processor performing multiple image processing tasks, reduces hardware costs, and improves processing efficiency and accuracy.
Patent Information
- Application Number
- CN202410597452.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, using multiple neural network processors for image processing results in high hardware costs, making it impossible to effectively reduce hardware costs.
By training multiple neural network models and mixing their parameters, a hybrid neural network model is generated for use in a single neural network processor to perform multiple image processing tasks. The model parameters are optimized by combining a hybrid loss function to achieve multiple image processing effects.
It reduced hardware costs while optimizing multiple image processing effects, improving the efficiency and accuracy of image processing.
Smart Images

Figure CN120952061A_ABST
Abstract
Description
Technical Field
[0001] The embodiments described in this application relate to an image processing method, and more particularly to an image processing method based on a neural network processor. Background Technology
[0002] Because machine learning technology offers advantages in improving the efficiency and accuracy of image processing, the use of neural network circuits for image processing has become a major trend in technological development in recent years. On the other hand, image processing implementation typically involves the optimization of multiple image functions, not just for a specific image effect. Different image functions require optimization using different image processing circuits, and each image processing circuit requires a neural network processor. Using multiple image processing circuits to achieve the desired overall image effect inevitably results in high hardware costs.
[0003] Therefore, this application aims to develop an image processing method and apparatus that uses a single neural network circuit to process multiple image effects. Summary of the Invention
[0004] Some embodiments of this application relate to an image processing method, which includes: Based on multiple training data, a first neural network model for performing a first image processing is trained to generate multiple first parameters associated with the first neural network model, wherein these first parameters include multiple weight values; based on the training data and these weight values, a second neural network model for performing a second image processing different from the first image processing is trained to generate multiple second parameters associated with the second neural network model; and these first parameters and these second parameters are mixed to generate multiple hybrid parameters for a hybrid neural network model, wherein the hybrid neural network model is used to perform the first image processing and the second image processing on an input image to output an optimized image.
[0005] Some embodiments of this application relate to an image processing method comprising the steps of: (a) sequentially training multiple neural network models to generate a set of hybrid parameters based on multiple sets of model parameters corresponding to these neural network models, wherein each of these neural network models is used to perform one of multiple image processing operations; and (b) adjusting this set of hybrid parameters based on a set of output data of the hybrid neural network model and a hybrid loss function so that the hybrid neural network model performs hybrid image processing on the input image. Step (a) includes: Based on multiple training data and multiple weight values of a set of model parameters of the preceding neural network model in these neural network models, train the subsequent neural network model in these neural network models to generate a set of model parameters of the subsequent neural network model; and mix these sets of model parameters of these neural network models to generate this set of mixed parameters for the hybrid neural network model. Attached Figure Description
[0006] The content of this application can be better understood by referring to the embodiments in the following paragraphs and the accompanying drawings: Figure 1 A schematic diagram illustrating an image processing apparatus according to some embodiments of the present application is provided. Figure 2 A schematic diagram illustrating the operation of various training methods according to embodiments of this application is provided. Figure 3 This is a flowchart of an image processing method according to some embodiments of this application; Figure 4 This is a flowchart of an image processing method according to some embodiments of this application; Figure 5 A schematic diagram illustrating a hybrid loss function according to some embodiments of this application; and Figure 6 This is a flowchart of an image processing method according to some embodiments of this application. Detailed Implementation
[0007] The spirit of this application will be clearly explained below with reference to the accompanying drawings and detailed description. Anyone skilled in the art can make changes and modifications based on the technology taught in this application after understanding the embodiments of this application, without departing from the spirit and scope of this application.
[0008] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the scope of this application. The singular forms such as “one,” “this,” “this,” “this,” and “the” as used in this article also include the plural forms.
[0009] The terms "include", "include", "have", "contain", etc., used in this article are all open-ended terms, meaning they include but are not limited to.
[0010] Unless otherwise specified, the terms used herein generally have their ordinary meaning in the context of the art, the application, and the specific content thereof. Certain terms used to describe this application will be discussed below or elsewhere in this specification to provide additional guidance to those skilled in the art in describing this application.
[0011] Now for reference Figure 1 . Figure 1 A schematic diagram of an image processing apparatus 10 according to some embodiments of this application is shown. Figure 1 As shown, the image processing apparatus 10 of this application includes an image processing circuit 11 with a neural network processor 12, which performs multiple image processing on an input image LR received from an input terminal IN and outputs a corresponding optimized image HR.
[0012] In some embodiments, the image processing circuit 11 may be an integrated circuit of various circuit components responsible for image processing. In some embodiments, the neural network processor 12 may be a graphics processing unit (GPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC).
[0013] In some embodiments, the neural network model in the image processing circuit 11 is used to perform image processing on the input image LR, such as noise suppression, image detail enhancement, sharpening, and contrast enhancement, to produce a refined image HR.
[0014] In some embodiments, the image processing apparatus 10 can be used to perform multiple image processing operations by programming a hybrid neural network model 140 generated by the image processing methods 300, 400, 600 described in the following paragraphs into a neural network processor 12.
[0015] Using the image processing apparatus 10 of this application, multiple image effects can be processed by a single neural network processor 12, thus reducing hardware costs. This is because this application first trains a neural network model capable of performing multiple image processing using deep learning technology, and then implements this neural network model in hardware. The training method will be described in detail below.
[0016] Now for reference Figure 2 . Figure 2 The illustrations depict various embodiments corresponding to this application. Figure 1 This is a schematic diagram illustrating the training method of the neural network model in the neural network processor 12. In some embodiments, Figure 1 The training of the neural network model in the neural network processor 12 involves multiple training data sets 110, a convolutional neural network (CNN) model 120, a generative adversarial neural network (GAN) model 130, and a hybrid neural network model 140. In some embodiments, Figure 1 Training the neural network model in the neural network processor 12 may also involve a hybrid loss function 150. In other embodiments, Figure 1 Training of the neural network model in the neural network processor 12 may also involve a preprocessing procedure 160.
[0017] Now, please refer to them together. Figure 2 and Figure 3 ,in Figure 3 This is a flowchart of an image processing method 300 according to some embodiments of this application. It should be understood that... Figure 3 The preceding, intermediate, and subsequent steps of method 300 illustrated may replace or eliminate some of the steps described below for additional embodiments of the method. The order of the steps / methods may be interchangeable. In the various figures and illustrative embodiments, the same reference numerals are used to denote the same components. Method 300 includes the following references Figure 2 The training method consists of steps 310 to 330.
[0018] In step 310, as Figure 2 As shown, a CNN model 120 is trained to perform a first image processing (e.g., noise suppression) based on multiple training data 110 to generate multiple CNN model parameters associated with the CNN model 120.
[0019] In some embodiments, the multiple training data 110 include multiple non-golden data and multiple golden data. The non-golden data corresponds to the input image LR, while the golden data is the image to be learned from the input image through the training process; that is, the image HR to be output by the image processing device 10 corresponding to the input image LR. Therefore, non-golden data and golden data have a one-to-one correspondence. In other words, when the non-golden data is... {x1,x2…xn}, correspondingly, the golden ratio data is {X1,X2…Xn}.
[0020] Specifically, training data 110 (i.e., both high-quality and low-quality data) is input into CNN model 120. Then, following general deep learning methods, the CNN model algorithm is used to process the training data 110, and after the processing, the CNN model parameters containing {a1, a2…an} are output, where a1 to an are the weights PTW of the CNN model and correspond to the training data x1 to xn, respectively. In some embodiments, the CNN model may also include a bias a0.
[0021] In step 320, as Figure 2 As shown, a GAN model 130 is trained to perform a second image processing (e.g., generating image details) based on training data 110 and weight values PTW, thereby generating multiple GAN model parameters associated with the GAN model 130. Similar to the training of a CNN model, in addition to inputting multiple training data 110 into the GAN model 130, the training of a GNN model also includes inputting the weight values PTW as pre-trained weight values into the GAN model 130.
[0022] In detail, the GAN model 130 includes a generator model and a discriminator model. The step of inputting multiple training data 110 into the GAN model 130 includes inputting non-golden data into the generator model and inputting golden data into the discriminator model. In each iteration of training the GAN model 130, the generator model produces multiple output data corresponding to the non-golden data; the discriminator model further compares the output data with the golden data; when the discriminator model determines that the output data is different from the golden data, it updates the weight values of the current GAN model; when the discriminator model cannot distinguish between the golden data and the output data, the training of the GAN model 130 ends and outputs GAN model parameters containing {b1, b2…bn}, where b1 to bn are the weight values of the GAN model and correspond to the training data x1 to xn respectively. In some embodiments, the GAN model may also include a bias value b0.
[0023] In step 330, as Figure 2 As shown, these first parameters are mixed with these second parameters to generate multiple mixed parameters for the hybrid neural network model. The hybrid neural network model 140, which takes the mixed parameters as input, can output an optimized image HR with the hybrid image processing effects of the CNN model 120 and the GAN model 130. The method for generating the mixed parameters for the hybrid neural network model 140 is explained in detail below.
[0024] For example, the above training produces a set of CNN model parameters {a1, a2…an} and a set of GAN model parameters {b1, b2…bn}. These can be proportionally α... p :(1-α p Each parameter in the CNN model is mixed with the corresponding parameter in the GAN model as follows: {c1…cn}, has cm=α p *am+(1-α p )*bm, m=1,2…n α p α is a constant between 0 and 1, and its value depends on the desired image processing effect. p For example, if a greater emphasis is placed on suppressing noise amplification, an α value greater than 0.5 can be used. p .
[0025] When the CNN and GAN model parameters also include bias values, the above parameter mixing further includes mixing the bias value a0 of the CNN model with the bias value b0 of the GAN model. Corresponding to the above-described proportional mixing method, the mixing parameters also include a mixing bias value c0: c0=α p *a0+(1-αp )*b0.
[0026] In some embodiments, the hybrid neural network model 140 is a CNN model. In other embodiments, the hybrid neural network model 140 is a GAN model.
[0027] While the training method is illustrated above using a hybrid of CNN model 120 and GAN model 130, the training method described in this application is not limited to a hybrid of CNN model 120 and GAN model 130. In some embodiments, GAN model 130 can be replaced by another CNN model, i.e., two CNN models used for different image processing (in addition to noise suppression, CNN models can generally be used to perform image processing such as sharpening, deblurring, and resolution improvement, depending on the algorithm design). In some embodiments, CNN model 120 and / or GAN model 130 can be replaced by a deep neural network (DNN) model or a recurrent neural network (RNN) model. In some embodiments, CNN model 120 and GAN model 130 can be replaced by an unsupervised neural network model, and the training data 110 does not need to contain gold data.
[0028] Now, please refer to them together. Figure 2 and Figure 4 ,in Figure 4 This is a flowchart of an image processing method 400 according to some embodiments of this application. It should be understood that... Figure 4 Additional steps are provided before, during, and after method 400, and some steps described below may be substituted or eliminated for additional embodiments of the method. The order of steps / methods may be interchangeable. In the various figures and illustrative embodiments, the same reference numerals are used to denote the same components. Method 400 includes the following references Figure 2 The training method steps 410 to 440 are the same as steps 310, 320 and 330 in image processing method 300, so they will not be described again here.
[0029] In step 440, as Figure 4 As shown, based on multiple first output data generated by CNN model 120 or multiple second output data generated by GAN model 130, a hybrid neural network model 140 is further trained using a hybrid loss function 150 to optimize the hybrid parameters of the hybrid neural network model 140.
[0030] In some embodiments, when the hybrid neural network model 140 is a CNN model, the first output data received from the CNN model 120 is input into the hybrid neural network model 140, and the hybrid neural network model 140 is further trained using a hybrid loss function 150. In each iteration of training the hybrid neural network model 140, the hybrid neural network model 140 generates multiple optimization data based on the first output data and inputs the optimization data into the hybrid loss function 150 to calculate the magnitude of the hybrid loss function 150, and determines whether to further adjust the weight values Wi of the hybrid neural network model 140 based on this magnitude.
[0031] For example, if the calculated magnitude of the hybrid loss function 150 is greater than a threshold, the weight values Wi are adjusted and the updated weight values Wi are input back into the hybrid neural network model 140 for the next iteration. Conversely, if the calculated magnitude of the hybrid loss function 150 is less than the threshold, the training of the hybrid neural network model 140 is terminated and the current weight values Wi are used as the hybrid parameters of the hybrid neural network model 140.
[0032] When the hybrid neural network model 140 is a GAN model, the second output data received from the GAN model 130 is input into the hybrid neural network model 140, and the hybrid loss function 150 is used to further train the hybrid neural network model 140. Compared with the embodiment where the hybrid neural network model 140 is a CNN model, the training method is the same as the example where the hybrid neural network model 140 is a GAN model, except that the first output data is replaced by the second output data, so it will not be described again here.
[0033] In some embodiments, the hybrid loss function 150 is as follows: Figure 5 As shown. The hybrid loss function 150 is designed to achieve the desired image optimization effect. Specifically, the hybrid loss function 150 may be a noise suppression loss function 510, a sharpening loss function 520, an image edge enhancement loss function 530, or a combination thereof.
[0034] In some embodiments, the hybrid loss function 150 is a linear superposition of multiple loss functions. For example, such as Figure 5 As shown, the hybrid loss function 150 is formed by superimposing the noise suppression loss function 510, the sharpening loss function 520, and the image edge enhancement loss function 530 in a ratio of α:β:γ. In other words, when the values of the noise suppression loss function 510, the sharpening loss function 520, and the image edge enhancement loss function 530 are f(x1,x2…xn), g(x1,x2…xn), and h(x1,x2…xn), respectively, the value of the hybrid loss function 150, BLF(x1,x2…xn), is: BLF(x1,x2…xn)=α*f(x1,x2…xn)+β*g(x1,x2…xn)+γ*h(x1,x2…xn), The values of α, β, and γ depend on the desired image effect. For example, if a sharpening effect is desired, then β should be increased.
[0035] More specifically, in the above embodiments, since the scales of f(x1, x2...xn), g(x1, x2...xn), and h(x1, x2...xn) are different, a balancing action needs to be performed before determining α, β, and γ to make the contributions of the noise suppression loss function 510, the sharpening loss function 520, and the image edge enhancement loss function 530 to the hybrid loss function 150 comparable. In other words, the scales of f(x1, x2...xn), g(x1, x2...xn), and h(x1, x2...xn) are first estimated to determine the initial values of α, β, and γ so that the products of each initial value and the corresponding function value are comparable, that is, the scales of the products α*f(x1, x2...xn), β*g(x1, x2...xn), and γ*h(x1, x2...xn) are of the same order of magnitude. Then, based on the initial values, the magnitudes of α, β, and γ are adjusted.
[0036] Now, please refer to them together. Figure 2 and Figure 6 ,in Figure 6 This is a flowchart of an image processing method 600 according to some embodiments of this application. It should be understood that... Figure 6 Additional steps are provided before, during, and after method 600, and some steps described below may be substituted or eliminated for additional embodiments of the method. The order of steps / methods may be interchangeable. In the various figures and illustrative embodiments, the same reference numerals are used to denote the same components. Method 600 includes the following references Figure 2 The training method steps 610-640 are the same as steps 310, 320, and 330 in image processing method 300, so they will not be described again here.
[0037] In step 610, as Figure 6 As shown, a preprocessing procedure 160 is performed to optimize multiple training data sets 110. Before training the CNN model, the golden data in the training data set 110 is preprocessed using the golden data to optimize the golden data.
[0038] In some embodiments, preprocessing 160 is a denoising process, a sharpening process, or an edge enhancement process. In some embodiments, preprocessing 160 includes a denoising process, a sharpening process, and an edge enhancement process. Preprocessing 160 can be used to replace the aforementioned hybrid loss function 150. In other words, if preprocessing 160 is performed, it is not necessary to further train the hybrid neural network model 140 with the hybrid loss function 150. Using preprocessing 160 and hybrid loss function 150 can achieve the same or similar image processing results and both can reduce hardware costs.
[0039] While all the above embodiments involve the mixing of only two neural network models, it should be understood that the above image processing methods 300, 400, and 600 can also be applied to the mixing of more neural network models performing different image processing. The training method for the GAN neural network model in image processing methods 300, 400, and 600 can be used to train the next neural network model based on the training data 110 and the weight values of the previous neural network model, and then train subsequent third, fourth, etc. neural network models to generate multiple model parameters and multiple output data for the third, fourth, etc. neural network models.
[0040] Correspondingly, after training, similar to image processing methods 300, 400, and 600, these model parameters are further mixed with CNN model parameters and GAN model parameters to generate multiple mixed parameters for hybrid neural network model 140.
[0041] Similar to image processing method 400, the mixing parameters can be further adjusted based on the output data of one of the trained neural network models and the mixing loss function 150. On the other hand, the training order of the neural network models is determined by their convergence difficulty. For example, CNN models converge more easily than GAN models, so CNN models are trained first. As another example, when the preceding and following neural network models are of the same model type, the following neural network model converges more easily, so neural network models of the same type as the preceding neural network model are trained next to the preceding neural network model. For example, three models are now being trained: a first CNN model, a second CNN model, and a GAN model. In some embodiments, the first CNN model can be trained first, then the second CNN model, and finally the GAN model. In some embodiments, the second CNN model can be trained first, then the first CNN model, and finally the GAN model. However, training the first CNN model first, then the GAN model, and finally the second CNN model should be avoided as this slows down convergence.
[0042] Although the above embodiments of this application focus on image processing, it should be understood that by changing the type of training data 110, the algorithm of the trained neural network model, and the design of the hybrid loss function 150, this application can also be applied to the processing of audio, speech, files, etc.
[0043] In summary, this application provides a pipeline for significantly reducing hardware costs by generating a neural network model with multiple image processing functions using deep learning technology and then implementing it in hardware.
[0044] While this application discloses detailed embodiments as described above, it does not exclude other possible implementations. Therefore, the scope of protection of this application should be determined by the appended claims, and not by the limitations of the foregoing embodiments.
[0045] For those skilled in the art, without departing from the spirit and scope of this application, Various modifications and refinements can be made to this application. Based on the foregoing embodiments, all modifications and refinements made to this application are also covered within the protection scope of this application.
[0046] [Symbol Explanation] 10: Image processing device 11: Image processing circuit 12: Neural Network Processor 110: Training Data 120: CNN (Convolutional Neural Network) Model 130: GAN (Generative Adversarial Network) Model 140: Hybrid Neural Network Model 150: Mixed Loss Function 160: Preprocessing 300: Method 310~330: Steps 400: Method 410~440: Steps 510: Noise Suppression Loss Function 520: Sharpening Loss Function 530: Image Edge Enhancement Loss Function 600: Method 610~640: Steps IN: Input terminal OUT: Output terminal PTW: Weight Value Wi: Weight value α: proportionality coefficient β: proportionality coefficient γ: proportionality coefficient
Claims
1. An image processing method, characterized in that, Include: Based on multiple training data, a first neural network model for performing a first image processing is trained to generate multiple first parameters associated with the first neural network model, wherein the multiple first parameters include multiple weight values; Based on the multiple training data and the multiple weight values, a second neural network model is trained to perform a second image processing to generate multiple second parameters associated with the second neural network model, wherein the second image processing is different from the first image processing; as well as The plurality of first parameters are mixed with the plurality of second parameters to generate a plurality of mixed parameters for a hybrid neural network model, wherein the hybrid neural network model is used to perform the first image processing and the second image processing on an input image to output an optimized image.
2. The image processing method as described in claim 1, characterized in that, Also includes: After the first neural network model is trained, multiple first output data are generated through the first neural network model; After the second neural network model is trained, multiple second output data are generated through the second neural network model; as well as Based on the multiple first output data or the multiple second output data, the hybrid neural network model is further trained using a hybrid loss function to optimize the multiple hybrid parameters.
3. The image processing method as described in claim 2, characterized in that, The first neural network model differs from the second neural network model, and the image processing method further includes one of the following: When the hybrid neural network model is of the same model type as the first neural network model, the hybrid neural network model is trained based on the multiple first output data; and When the hybrid neural network model and the second neural network model are of the same model type, the hybrid neural network model is trained based on the multiple second output data.
4. The image processing method as described in claim 2, characterized in that, The hybrid loss function is a linear superposition of multiple loss functions.
5. The image processing method as described in claim 1, characterized in that, Also includes: Before training the first neural network model, a preprocessing procedure is performed to optimize the multiple training data.
6. The image processing method as described in claim 1, characterized in that, The first neural network model is a convolutional neural network model and the second neural network model is a generative adversarial network model.
7. The image processing method as described in claim 6, characterized in that, The multiple training data packets include multiple first data packets and multiple second data packets, and the generative adversarial network model includes a generative model and a discriminative model. The image processing method further includes: In each iteration of training the generative adversarial network model, the generative model generates multiple output data corresponding to the multiple first data, and the discriminative model compares the multiple output data with the multiple second data.
8. The image processing method as described in claim 1, characterized in that, Each of the plurality of second parameters is mixed with a corresponding one of the plurality of first parameters in a certain proportion.
9. An image processing method, characterized in that, Include: Multiple neural network models are trained sequentially to generate a set of hybrid parameters for a hybrid neural network model based on complex array model parameters corresponding to the multiple neural network models, wherein each of the multiple neural network models is used to perform one of multiple image processing operations, and the hybrid neural network model is used to perform the multiple image processing operations based on the set of hybrid parameters, including: Based on multiple training data and multiple weight values of a set of model parameters of a previous neural network model among the multiple neural network models, train a subsequent neural network model among the multiple neural network models to generate a set of model parameters of the subsequent neural network model. The parameters of the multiple neural network models are mixed to produce the mixed parameters. as well as The set of mixed parameters is adjusted based on a set of output data from one of the multiple neural network models and a mixed loss function.
10. The image processing method as described in claim 9, characterized in that, Also includes: The training order of the multiple neural network models is determined by the ease of convergence.