Training Method of Image Denoising Model and Image Denoising Method
By combining the cyclic iteration method of full gradient and stochastic gradient loss value in image denoising model training, the problem of excessive computing resource occupation and slow convergence speed is solved, and a more efficient image denoising effect is achieved.
Patent Information
- Application Number
- CN202211438914.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-16
AI Technical Summary
The existing image denoising model occupies too much computing resources and has poor denoising effect during training. This is mainly due to the excessive computing resources of the full gradient descent algorithm and the slow convergence speed of the stochastic gradient descent algorithm.
The cyclic iteration method is used to combine full gradient and stochastic gradient loss values, and the model parameters are updated by fusion gradient loss values, where the update frequency of the full gradient loss value is lower than that of the stochastic gradient, reducing the use of computing resources and speeding up the convergence speed.
While reducing the use of computing resources, the denoising effect and convergence speed of the image denoising model are improved, so that the model can achieve the optimal solution within the allowable time.
Smart Images

Figure CN115809967B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular, to a method for training an image denoising model and an image denoising method. Background Art
[0002] With the rapid development of computer science and image processing technology, images have been widely used in medical imaging, pattern recognition, and other aspects. However, during the formation and transmission of images, they will inevitably be interfered by noise, and the noise in some images is very serious. The noise in the image often interweaves with the signal, making the details of the image itself, such as the boundary contours, lines, etc., become blurred. Therefore, it is necessary to perform noise reduction processing on the image to facilitate higher-level image analysis and understanding.
[0003] Image denoising is an important part of image processing. Currently, the training of image denoising models usually uses gradient descent algorithms (including full gradient descent algorithm and stochastic gradient descent algorithm) to solve the optimal solution of the model and obtain the model parameters of the model. However, when using the full gradient descent algorithm, the entire sample image set needs to be used each time to update the model parameters, which will result in too long learning time each time and consume a large amount of computing resources. While using the stochastic gradient descent algorithm, a sample image is randomly selected from the sample image set each time to update the model parameters. Although it can reduce the calculation of the gradient and speed up the learning time each time, the randomness of the stochastic gradient descent may cause the update of the model parameters not to be in the correct direction each time, resulting in optimization fluctuations. Therefore, the number of iterations (learning times) will increase, that is, the convergence speed becomes slower, which may cause the image denoising model not to converge to the optimal solution within the allowed time, and thus the denoising effect of the image denoising model is poor when applied. Summary of the Invention
[0004] The main technical problem to be solved by the present application is to provide a method for training an image denoising model and an image denoising method, which can reduce the computing resources occupied during the model training process and improve the denoising effect of the image denoising model when applied.
[0005] To solve the above technical problems, a first aspect of the present application provides a method for training an image denoising model, including: obtaining a sample image set, where the sample image set includes a plurality of sample image pairs, and each sample image pair respectively includes a first feature matrix obtained by performing feature extraction on an original sample image and a second feature matrix obtained by performing feature extraction on a noisy sample image obtained by adding noise to the original sample image; for the image denoising model, calculating a full gradient loss value and a stochastic gradient loss value through the sample image set in a cyclic iteration manner, and using a fused gradient loss value formed by fusing the full gradient loss value and the stochastic gradient loss value to update the model parameters of the image denoising model, where the update frequency of the full gradient loss value is lower than that of the stochastic gradient loss value.
[0006] To solve the above technical problems, a second aspect of the present application provides an image denoising method, including: obtaining an image denoising model trained using a sample image set, where the model parameters of the image denoising model are obtained using the method for training an image denoising model as described above; using the image denoising model to perform denoising processing on a noisy image to obtain a denoised image.
[0007] The beneficial effects of the present application are as follows: Different from the prior art, the present application obtains a sample image set, where the sample image set includes a plurality of sample image pairs, and each sample image pair respectively includes a first feature matrix obtained by performing feature extraction on an original sample image and a second feature matrix obtained by performing feature extraction on a noisy sample image obtained by adding noise to the original sample image. Then, for the image denoising model, a full gradient loss value and a stochastic gradient loss value are calculated through the sample image set in a cyclic iteration manner, and the model parameters of the image denoising model are updated using a fused gradient loss value formed by fusing the full gradient loss value and the stochastic gradient loss value, where the update frequency of the full gradient loss value is lower than that of the stochastic gradient loss value. As described above, during the gradient descent process of model training, a full gradient loss value and a stochastic gradient loss value are calculated through the sample image set, and the update frequency of the full gradient loss value is lower than that of the stochastic gradient loss value, so that the number of times of traversing the entire sample image set to calculate the full gradient loss value can be reduced, thereby reducing the occupation of computing resources. In addition, updating the model parameters using a fused gradient loss value formed by fusing the full gradient loss value and the stochastic gradient loss value can avoid the slow convergence speed caused by only using the stochastic gradient loss value deviating from the correct gradient descent direction, improve the convergence speed of the model, enable the model to converge to the optimal solution within the allowed time, and further improve the denoising effect of the image denoising model during application. Description of the Drawings
[0008] To more clearly illustrate the technical solutions in this application, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings described below are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings. Among them:
[0009] Figure 1 is a schematic flowchart of an embodiment of the image denoising method of this application;
[0010] Figure 2 is a schematic flowchart of an embodiment of the training method of the image noise reduction model of this application;
[0011] Figure 3 is a schematic flowchart of another embodiment of the training method of the image noise reduction model of this application;
[0012] Figure 4 is a schematic flowchart of yet another embodiment of the training method of the image noise reduction model of this application;
[0013] Figure 5 is Figure 3 a schematic flowchart of an embodiment of step S33 in
[0014] Figure 6 is Figure 5 a schematic flowchart of an embodiment of step S332 in
[0015] Figure 7 is Figure 5 a schematic flowchart of an embodiment of step S333 in
[0016] Figure 8 is Figure 5 a schematic flowchart of an embodiment of step S334 in
[0017] Figure 9 is Figure 5 a schematic flowchart of an embodiment of step S335 in
[0018] Figure 10 is Figure 3 a schematic flowchart of an embodiment of step S34 in
[0019] Figure 11 is a schematic block diagram of an embodiment of the image denoising device of this application;
[0020] Figure 12 is a schematic block diagram of an embodiment of the training device of the image noise reduction model of this application;
[0021] Figure 13 is a schematic block diagram of an embodiment of the electronic device of this application;
[0022] Figure 14 It is a structural schematic block diagram of an embodiment of the computer-readable storage medium of the present application. Specific embodiments
[0023] Referring to "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0024] The terms "first" and "second" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0026] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the image denoising method of the present application. Among them, the execution subject of the present application is an electronic device, such as an electronic device with data processing capabilities such as a mobile phone, a computer, a server, etc.
[0027] The method may include the following steps:
[0028] Step S11: Obtain an image denoising model trained using a sample image set, where the model parameters of the image denoising model are trained using the training method of the image denoising model in any embodiment of the present application.
[0029] Among them, the sample image set includes multiple sample image pairs. Each sample image pair respectively includes a first feature matrix obtained by performing feature extraction on the original sample image, and a second feature matrix obtained by performing feature extraction on the noisy sample image obtained by adding noise to the original sample image. The sample image can be a single captured image, or can also be an image frame in video data, which is not limited here. The image denoising model is used to perform denoising processing on the noisy image to obtain a denoised image. In one example, the sample image is obtained through medical imaging, such as a magnetic resonance image obtained by magnetic resonance imaging (MRI).
[0030] Step S12: Use the image denoising model to perform denoising processing on the noisy image to obtain a denoised image.
[0031] Image acquisition devices are vulnerable to factors such as camera shake, moving objects, low light, and noise, resulting in unclean captured image data, that is, there is noise. The goal of image denoising is to recover the original true image as much as possible from the degraded image affected by noise, which is a key step for subsequent image processing (such as image recognition).
[0032] In some embodiments, the image denoising model can be a total variation model (TV). The total variation model is an anisotropic model that relies on the gradient descent flow to smooth the image. It is hoped that the image can be smoothed as much as possible inside the image (the difference between adjacent pixels is small), while not being smoothed as much as possible at the image edge (image contour).
[0033] As above, in this application, an image denoising model trained using a sample image set is obtained, and then the image denoising model is used to perform denoising processing on the noisy image to obtain a denoised image. Among them, the model parameters of the image denoising model are trained using the training method of the image denoising model in any embodiment of this application. Among them, by accelerating the convergence rate of the model output to the optimal solution of the model during the training process, the effect of the image denoising model is better, thereby improving the denoising effect of the image denoising model.
[0034] Currently, during the training process of an image denoising model, only the stochastic gradient loss value or the full gradient loss value is used to update the original variables. However, although the stochastic gradient loss value can reduce the gradient calculation, thereby accelerating the iterative speed of the model output towards the optimal solution of the model, due to the noise variance between the stochastic gradient and the full gradient, as the iteration progresses, to ensure convergence, the learning rate can only be continuously decreased or even approach 0. As a result, the iterative speed towards the optimal solution will become slower and slower, leading to poor performance of the trained image denoising model during application. And during the entire training process, if the full gradient loss value is used to update the model parameters, the full gradient loss value needs to be updated for the entire sample image set each time, resulting in a relatively large amount of computing resources occupied during the entire training process. Based on this, the present application provides a training method for an image denoising model that can reduce the computing resources occupied during the model training process and simultaneously improve the denoising effect of the image denoising model during application.
[0035] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the training method for the image denoising model of the present application. Among them, the execution subject of the present application is an electronic device, such as a mobile phone, a computer, a server, etc., which is an electronic device with data processing capabilities.
[0036] The method may include the following steps:
[0037] Step S21: Obtain a sample image set, where the sample image set includes a plurality of sample image pairs, and each sample image pair respectively includes a first feature matrix obtained by performing feature extraction on the original sample image and a second feature matrix obtained by performing feature extraction on the noisy sample image obtained by adding noise to the original sample image.
[0038] Among them, the sample image set is denoted as ξ I , and the i-th sample image pair is denoted as {ξ i =(a i , b i )}, where ai is the first feature matrix obtained by performing feature extraction on the original sample image, and b i is the second feature matrix obtained by performing feature extraction on the noisy sample image obtained by adding noise to the original sample image. a i all follow the standard normal distribution and are independently and identically distributed. Specifically, the sample image set includes {(a1, b1), (a2, b2),..., (a n , b n )} n sample image pairs, and n is a positive integer.
[0039] Among them, the noisy sample image can be obtained by adding standard Gaussian noise to the original sample image. Specifically, the noise addition method is not limited here.
[0040] Step S22: For the image denoising model, calculate the full gradient loss value and the stochastic gradient loss value through the sample image set in a cyclic iterative manner, and update the model parameters of the image denoising model using the fused gradient loss value formed by the full gradient loss value and the stochastic gradient loss value, where the update frequency of the full gradient loss value is lower than that of the stochastic gradient loss value.
[0041] In this embodiment, the image denoising model may be a total variation model (TV), but is not limited thereto.
[0042] The target model can be abstracted into a kind of target problem as follows:
[0043]
[0044] In the above target problem expression, is a compact convex set with a diameter of D x The regular function term and the regular function term are both convex but may be non-smooth. r2 contains a penalty matrix that characterizes the desired structured sparse pattern in x For a sample image pair {ξ i =(a i , b i )}, use to represent the convex smooth loss function on method x. The penalty matrix can be a non-diagonal matrix.
[0045] The image denoising model of the present application can also be abstracted into a kind of target problem. In the target problem, the regular term r1(x)=0, r2 = τ1||z||1 = τ1||Fx||1, τ1 is a known coefficient, that is, the target problem corresponding to the image denoising model can be obtained as:
[0046]
[0047] Among them, regarding the penalty matrix F ij , when i = j, F ij =1, when j = i + 1, F ij =-1, in other cases F ij =0.
[0048] Since the target problem is relatively complex, it is very difficult to obtain the optimal solution of the problem by means of derivation or the computational cost is extremely high. Therefore, an iterative method is generally adopted to gradually approach the optimal solution of the target problem, and it can be proved that after iterating enough times, the output of the model can approach the optimal solution of the target problem with a given error.
[0049] In some embodiments, the model parameters of the image denoising model can be updated in a nested manner of outer-loop iteration and inner-loop iteration. When performing outer-loop iteration, the full-gradient loss value is updated, and when performing inner-loop iteration, the stochastic-gradient loss value is updated. That is, when the inner loop ends and exits the inner loop to enter the outer loop, the full-gradient loss value is updated. Therefore, the update frequency of the full-gradient loss value is lower than that of the stochastic-gradient loss value. Since a two-layer loop structure is adopted, the exact full-gradient loss value is first calculated in the outer loop as a full traversal of the sample image set. A layer of inner loop is nested in the outer loop, and the stochastic-gradient loss value is updated in the inner loop. Then, the fused-gradient loss value formed by fusing the full-gradient loss value and the stochastic-gradient loss value is used to update the model parameters, which can avoid the deviation of the iteration result from the convergence direction of the model optimal solution due to the existence of noise variance in the stochastic gradient, thereby accelerating the convergence rate of the model output to the model optimal solution and improving the performance of the trained image denoising model during application.
[0050] In the above solution, by obtaining a sample image set, where the sample image set includes multiple sample image pairs, each sample image pair respectively includes a first feature matrix obtained by performing feature extraction on the original sample image and a second feature matrix obtained by performing feature extraction on the noisy sample image obtained by adding noise to the original sample image. Then, for the image denoising model, the full-gradient loss value and the stochastic-gradient loss value are calculated through the sample image set in a cyclic iteration manner, and the model parameters of the image denoising model are updated using the fused-gradient loss value formed by fusing the full-gradient loss value and the stochastic-gradient loss value. The update frequency of the full-gradient loss value is lower than that of the stochastic-gradient loss value. In the above, during the gradient descent process of model training, the full-gradient loss value and the stochastic-gradient loss value are calculated through the sample image set, and the update frequency of the full-gradient loss value is lower than that of the stochastic-gradient loss value, which can reduce the number of times of traversing the entire sample image set to calculate the full-gradient loss value, thereby reducing the occupation of computing resources. In addition, updating the model parameters based on the fused-gradient loss value formed by fusing the full-gradient loss value and the stochastic-gradient loss value can avoid the slow convergence speed caused by the deviation of only using the stochastic-gradient loss value from the correct gradient descent direction, improve the convergence speed of the model, enable the model to converge to the optimal solution within the allowed time, and further improve the denoising effect of the image denoising model during application.
[0051] Please refer to Figures 3 to 4 , Figure 3 which is a schematic flowchart of another embodiment of the training method of the image denoising model of the present application. Figure 4 which is a schematic flowchart of yet another embodiment of the training method of the image denoising model of the present application. Among them, the execution subject of the present application is an electronic device, such as a mobile phone, a computer, a server, etc., which are electronic devices with data processing capabilities.
[0052] The method may include the following steps:
[0053] Step S31: Initialize the outer loop value of the variables in the expression of the image denoising model; wherein the expression is in the form of an augmented Lagrangian function, and the variables in the expression are in matrix form, and include a first original variable corresponding to the model parameters of the image denoising model, a second original variable that satisfies a constraint relationship defined by a predetermined constraint function with the first original variable, and a dual variable related to the constraint relationship.
[0054] In this embodiment, the expression of the image denoising model has a first original variable (denoted as x) corresponding to the model parameters, a second original variable (denoted as z) that satisfies a constraint relationship defined by a predetermined constraint function with the first original variable, and a dual variable (denoted as λ) related to the constraint relationship. The first original variable x, the second original variable z, the dual variable λ, and the constraint function are all in matrix form. In some embodiments, the constraint function may be a penalty matrix, specifically a penalty matrix in a structured sparse pattern, denoted as
[0055] wherein, the variable refers to at least one of the first original variable, the second original variable, and the dual variable, and the original variable refers to at least one of the first original variable and the second original variable.
[0056] wherein, the first original variable and the second original variable satisfy a constraint relationship defined by a predetermined constraint function. For example, the constraint residual between the first original variable and the second original variable is 0, that is, the first original variable and the second original variable satisfy the following constraint relationship: Fx - z = 0.
[0057] In some embodiments, the expression of the image denoising model may include a loss estimation term, and the loss estimation term has the above-mentioned first original variable x, second original variable z, and dual variable λ. In one embodiment, the loss estimation term may be a loss function minus a correction function. The loss function is used to calculate the distance between the product result of the transpose matrix of the first original variable and the first feature matrix and the second feature matrix, and the correction function is used to calculate the inner product of the dual variable and the constraint residual between the first original variable and the second original variable.
[0058] Specifically, the loss estimation term can be seen in the following formula:
[0059] φ(z, x, λ) = l(x) - <λ, Fx - z>,
[0060]
[0061] wherein, φ(z, x, λ) is the loss estimation term; l(x) is the loss function, x T ai The transposed matrix x of the first original variable T and the product result of the first eigenmatrix a i where bi is the second eigenmatrix; <λ, Fx - z> is a correction function for calculating the dual variable λ and the constraint residual between the first original variable x and the second original variable z. is the dual variable related to Fx = z.
[0062] In some embodiments, the expression of the image denoising model may include a penalty term based on the constraint residual of the first original variable and the second original variable. In one implementation example, the penalty term is Fx - z is the constraint residual of the first original variable and the second original variable, γ is the penalty term coefficient, and γ can be a constant.
[0063] In some embodiments, the expression of the image denoising model may include a first regularization function term and a second regularization function term. The first regularization function term takes the first original variable as a variable, and the second regularization function term takes the second original variable as a variable.
[0064] In some embodiments, both the regularization function terms r1 and r2 are continuous but may be non - smooth. Then, for each regularization function term, its corresponding proximal mapping can obtain a closed - form solution, that is, the following proximal mapping function has a closed - form solution for i = 1, 2:
[0065]
[0066] where is the calculation result of the proximal mapping function. Given a variable x, find an optimal point such that is minimized.
[0067] In some specific embodiments, the expression of the image denoising model is as follows:
[0068]
[0069] where L γ (z, x, λ) is the augmented Lagrangian function, r1(x) is the regularization function term of the first original variable x, r2(z) is the regularization function term of the second original variable z, φ(z, x, λ) is the loss estimation term, is the penalty term. Specifically, in the expression of the total - variation image denoising model, the regularization term r1(x) = 0, r2 = τ1||z||1 = τ1||Fx||1, and τ1 is a known coefficient.
[0070] Specifically, initialize the outer - loop values of the variables in the expression of the image denoising model, that is, assign the initial values of the outer loop to the variables.
[0071] Among them, is the initial value of the outer loop of the first original variable, is the initial value of the outer loop of the second original variable, is the initial value of the outer loop of the dual variable. Among them, and The specific values of can be set as needed and are not limited here.
[0072] Step S32: Initialize the inner loop values of the variables using the outer loop values, and calculate the full gradient loss value at the outer loop values based on all the sample images in the sample image set.
[0073] Among them, the sample image set is denoted as ξ I , and the i-th sample image pair is denoted as {ξ i =(a i , b i )}.
[0074] In some embodiments, the number of iterations of the outer loop is denoted as s, s = 0, 1, 2,..., s - 1. The number of iterations of the inner loop is denoted as k, k == 0, 1, 2,..., k - 1. Among them, the initial value of the outer loop value can be used as the initial value of the inner loop to initialize the inner loop values of the variables. Correspondingly, the expression is as follows:
[0075]
[0076] Among them, is the outer loop value of the first original variable when the outer loop iterates for the s-th time;
[0077] is the inner loop value of the first original variable when the outer loop iterates for the s-th time and the inner loop iterates for the 0-th time, that is, when the outer loop iterates for the s-th time, the initial value of the inner loop of the first original variable;
[0078] is the outer loop value of the second original variable when the outer loop iterates for the s-th time;
[0079] is the inner loop value of the second original variable when the outer loop iterates for the s-th time and the inner loop iterates for the 0-th time, that is, when the outer loop iterates for the s-th time, the initial value of the inner loop of the second original variable;
[0080] is the outer loop value of the dual variable when the outer loop iterates for the s-th time;
[0081] is the inner loop value of the dual variable when the outer loop iterates for the s-th time and the inner loop iterates for the 0-th time, that is, when the outer loop iterates for the s-th time, the initial value of the inner loop of the dual variable.
[0082] When the outer loop iterates 0 times and the inner loop iterates for the 0th time, the initialization of the inner loop values of the above three variables by the outer loop values can be expressed as:
[0083]
[0084] Among them, for each variable, the initial value when the inner loop iterates for the 0th time is equal to the initial value when the outer loop iterates for the 0th time.
[0085] Among them, the full gradient loss value can be obtained by taking the derivative of the loss function, that is, the full gradient loss value Among them, is the loss function.
[0086] Step S33: Update the inner loop values in the inner loop iteration mode. Among them, in each inner loop iteration process, based on the first stochastic gradient loss value at the inner loop values calculated from the sample images randomly selected from the sample image set and the second stochastic gradient loss value at the outer loop values, and use the fusion gradient loss value formed by fusing the first stochastic gradient loss value, the second stochastic gradient loss value, and the full gradient loss value to update the inner loop values.
[0087] In some embodiments, the first stochastic gradient and the second stochastic gradient are obtained by taking the derivative of the loss function, and the fusion gradient loss value is further subtracted from the product of the transpose matrix of the constraint function and the inner loop values of the dual variables based on the calculation results of the full gradient loss value, the first stochastic gradient, and the second stochastic gradient.
[0088] In some embodiments, in each inner loop iteration process, the fusion gradient loss value is obtained by adding the first stochastic gradient to the full gradient loss value and subtracting the second stochastic gradient. Specifically, based on the calculation result of adding the first stochastic gradient to the full gradient loss value and subtracting the second stochastic gradient, the product of the transpose matrix of the constraint function and the inner loop values of the dual variables is further subtracted.
[0089] In an example, the calculation expression of the fusion gradient loss value is as follows:
[0090]
[0091] Among them, is the fusion gradient loss value, is the calculation result of the full gradient loss value, the first stochastic gradient, and the second stochastic gradient, F T λ is the product of the transpose matrix of the constraint function and the inner loop values of the dual variables, is the full gradient loss value, is the first stochastic gradient at the inner loop values, The second stochastic gradient at the outer loop value.
[0092] Step S34: Update the outer loop value for the next outer loop iteration based on the inner loop value of the inner loop iteration, and return the set of steps for initializing the inner loop value of the variable using the outer loop value to perform the outer loop iteration.
[0093] After step S34 is executed, if the outer loop iteration has not ended, step S32 is repeatedly executed. If the outer loop iteration has ended, step S35 is executed.
[0094] In some embodiments, for each variable, the average value of the update values of the inner loop values obtained in each inner loop iteration can be calculated as the outer loop value of the first original variable for the next outer loop iteration. Among them, the update value can be the first update value, the second update value, etc. The first update value is updated based on the initial inner loop value using the initial fused gradient loss value, and the second update value is updated based on the first update value using the fused gradient loss value, and so on.
[0095] Step S35: In response to the end of the outer loop iteration, set the model parameters of the image denoising model using the outer loop value of the first original variable.
[0096] Among them, setting the model parameters of the image denoising model using the outer loop value means that the training of the image denoising model is completed.
[0097] In the above solution, an inner and outer two-layer loop structure is adopted. In the outer loop, the exact full gradient loss value at the outer loop value is first calculated, which is a full traversal of the sample image set. An inner loop is nested in the outer loop. In the inner loop, the full gradient loss value and the latest stochastic gradient loss value are used to construct a variance-reduced stochastic gradient loss value (i.e., the fused gradient loss value). At the end of the inner loop, the iteration result of the inner loop is used to calculate the outer loop for the next outer loop iteration, so that the obtained value is used to initialize the initial value of the inner loop in the next outer loop. This is a correction process to avoid the iteration result deviating from the convergence direction of the model optimal solution due to the existence of noise variance in the stochastic gradient, thereby accelerating the convergence rate of the model output to the model optimal solution and improving the performance of the trained image denoising model in application, such as improving the accuracy of image denoising.
[0098] Please refer to Figures 5 to 9 , Figure 5 is Figure 3 a schematic flowchart of an embodiment of step S33 in Figure 6 is Figure 5 a schematic flowchart of an embodiment of step S332 in Figure 7 is Figure 5Flow diagram of an embodiment of step S333 Figure 8 is Figure 5 Flow diagram of an embodiment of step S334 Figure 9 is Figure 5 Flow diagram of an embodiment of step S335
[0099] In this embodiment, the inner loop iteration process adopts a way of alternately updating the original variables and the dual variables. In each inner loop iteration process, the following steps S331 to S336 are sequentially executed:
[0100] Step S331: Minimize the expression with respect to the second original variable to obtain an updated value of the inner loop value of the second original variable.
[0101] In some embodiments, the step of minimizing the expression with respect to the second original variable includes: based on the initial value of the inner loop value of the second original variable, performing a proximal mapping on the second regular function term to obtain an updated value of the inner loop value of the second original variable. Correspondingly, the expression is as follows:
[0102]
[0103] where is the updated value of the inner loop value of the second original variable, are respectively the inner loop value of the first original variable and the inner loop value of the dual variable at the s-th outer loop iteration and the k-th inner loop iteration, which are known.
[0104] Above, the present application solves for z by minimizing the expression L with respect to z λ which is equivalent to calculating the proximal mapping of the second regular function term r2 and can obtain a closed-form solution under the given assumptions. k+1
[0105] Step S332: Based on the initial value of the inner loop value of the first original variable and the initial value of the inner loop value of the dual variable, calculate the first fusion gradient loss value using the first group of randomly selected sample images, and use the first fusion gradient loss value to update the initial value of the inner loop value of the first original variable once to obtain a first updated value of the inner loop value of the first original variable.
[0106] Among them, the corresponding expression for calculating the first fusion gradient loss value is: where is the initial value of the inner loop value of the first original variable, is the initial value of the inner loop value of the dual variable, the updated value of the inner loop value of the second original variable, For the first group of randomly selected sample image pairs, which include b sample images, where b is a positive integer. The calculation formula for the fused gradient loss value can be referred to the above embodiments and will not be elaborated here.
[0107] In some embodiments, step S332 may include steps S3321 to S3323, as follows:
[0108] Step S3321: Use the first fused gradient loss value and a preset first learning rate to adjust the initial value of the inner loop value of the first original variable once to obtain a first adjustment result.
[0109] In some embodiments, using the first fused gradient loss value and a preset first learning rate to adjust the initial value of the inner loop value of the first original variable once, specifically:
[0110]
[0111] Where, is the initial value of the inner loop value of the first original variable, is the first fused gradient loss value, and c is the first learning rate.
[0112] In some embodiments, the first learning rate may be a constant that does not change with the inner loop iteration and the outer loop iteration.
[0113] Step S3322: Perform a proximal mapping on the first regularization function term based on the first adjustment result to obtain a first proximal mapping result.
[0114] Specifically, the result of the proximal mapping of the first regularization function term is
[0115] Step S3323: Determine the first update value of the inner loop value of the first original variable based on the first proximal mapping result.
[0116] In some embodiments, the first proximal mapping result and the outer loop value of the first original variable may be summed to be used as the first update value of the inner loop value of the first original variable.
[0117] In other embodiments, a weighted sum of the first proximal mapping result and the outer loop value of the first original variable may be performed based on a preset weight coefficient to be used as the first update value of the inner loop value of the first original variable, as shown below:
[0118]
[0119] Where, is the first update value of the inner loop value of the first original variable, is the first proximal mapping result of the first original variable, is the outer loop value of the first original variable, θ s is the weight coefficient of the first proximal mapping result of the first original variable in the s-th outer loop iteration, (1 - θ s ) is the weight coefficient of the outer loop value of the first original variable, and the sum of the two weight coefficients is 1.
[0120] In some embodiments, the sum of the weight coefficient of the first proximal mapping result of the first original variable and the weight coefficient of the outer loop value of the first original variable may not be 1, and can be specifically set according to actual needs.
[0121] In some embodiments, the weight coefficient is a constant that does not change with the inner loop iteration and the outer loop iteration, or a variable that changes with the outer loop iteration. The following shows an implementation where the weight coefficient changes with the outer loop iteration, and the formula is as follows:
[0122]
[0123] where, θ s+1 is the weight coefficient of the first proximal mapping result of the first original variable in the (s + 1)-th outer loop iteration, θ s is the weight coefficient of the first proximal mapping result of the first original variable in the s-th outer loop iteration. When θ s+1 changes, (1 - θ s ) also changes accordingly.
[0124] Step S333: Use the initial value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable to perform a first update on the initial value of the inner loop value of the dual variable, so as to obtain a first updated value of the inner loop value of the dual variable.
[0125] In some embodiments, step S333 may include steps S3331 to S3332, as follows:
[0126] Step S3331: Calculate the first constraint residual between the initial value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable.
[0127] where, the first constraint residual is is the initial value of the inner loop value of the first original variable, is the updated value of the inner loop value of the second original variable, and F is the constraint function (such as a penalty matrix).
[0128] Step S3332: Use the preset second learning rate and the first constraint residual to adjust the initial value of the inner loop value of the dual variable, so as to obtain a first updated value of the inner loop value of the dual variable.
[0129] In some embodiments, the second learning rate is determined by the term coefficient γ of the penalty term in the augmented Lagrangian function.
[0130] In some embodiments, the following formula is used to calculate an update value of the inner loop value of the dual variable:
[0131]
[0132] Where, is an update value of the inner loop value of the dual variable, is the initial value of the inner loop value of the dual variable, γ is the second learning rate, is the first constraint residual.
[0133] Step S334: Based on an update value of the inner loop value of the first primal variable and an update value of the inner loop value of the dual variable, calculate a second fusion gradient loss value using a second set of randomly selected sample image pairs, and use the second fusion gradient loss value to perform a secondary update on the initial value of the inner loop value of the first primal variable to obtain a secondary update value of the inner loop value of the first primal variable.
[0134] Where, the corresponding expression for calculating the second fusion gradient loss value is: Where, is an update value of the inner loop value of the first primal variable, is an update value of the inner loop value of the dual variable, An update value of the inner loop value of the second primal variable, is the second set of randomly selected sample image pairs, which includes b sample images, and b is a positive integer. The calculation formula of the fusion gradient loss value can be referred to the above embodiments and will not be elaborated here.
[0135] In some embodiments, step S334 may include steps S3341 to S3343, as follows:
[0136] Step S3341: Use the second fusion gradient loss value and the first learning rate to perform a secondary adjustment on the initial value of the inner loop value of the first primal variable to obtain a secondary adjustment result.
[0137] Step S3342: Perform a secondary proximal mapping on the first regular function term based on the secondary adjustment result to obtain a secondary proximal mapping result.
[0138] Step S3343: Determine a secondary update value of the inner loop value of the first primal variable based on the secondary proximal mapping result.
[0139] In some embodiments, the secondary proximal mapping result and the outer loop value of the first primal variable are summed to be used as the secondary update value of the inner loop value of the first primal variable.
[0140] In some other embodiments, a weighted sum of the quadratic proximal mapping result and the outer loop value of the first original variable is calculated based on the weight coefficient, and used as the quadratic update value of the inner loop value of the first original variable, as follows:
[0141]
[0142] Wherein, is the quadratic update value of the inner loop value of the first original variable, is the quadratic proximal mapping result of the first original variable, is the outer loop value of the first original variable, θ s is the weight coefficient of the quadratic proximal mapping result of the first original variable, (1 - θ s ) is the weight coefficient of the outer loop value of the first original variable, and the sum of the two weight coefficients is 1. Wherein, calculating θ in s and calculating θ in s are the same, and will not be elaborated here.
[0143] Step S335: Quadratically update the initial value of the inner loop value of the dual variable by using the first update value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable, so as to obtain the second update value of the inner loop value of the dual variable.
[0144] In some embodiments, step S335 may include steps S3351 to S3352, as follows:
[0145] Step S3351: Calculate the second constraint residual between the first update value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable.
[0146] Step S3352: Adjust the initial value of the inner loop value of the dual variable by using the second learning rate and the second constraint residual, so as to obtain the second update value of the inner loop value of the dual variable.
[0147] In some embodiments, the second update value of the inner loop value of the dual variable is calculated by using the following formula:
[0148]
[0149] Wherein, is the second update value of the inner loop value of the dual variable, is the initial value of the inner loop value of the dual variable, γ is the second learning rate, is the second constraint residual.
[0150] Step S336: Use the second update value of the inner loop value of the first original variable, the update value of the inner loop value of the second original variable, and the second update value of the inner loop value of the dual variable as the initial values of the inner loop values of the first original variable, the second original variable, and the dual variable for the next inner loop iteration, and return to the step of minimizing the expression with respect to the second original variable.
[0151] Specifically, use the second update value of the k-th inner loop iteration as the initial values for the next inner loop iteration, i.e., the (k + 1)-th inner loop iteration, and then continue the iteration.
[0152] Please refer to Figure 10 , Figure 10 which Figure 3 is a schematic flowchart of an embodiment of step S34 in
[0153] In this embodiment, step S34, i.e., the step of updating the outer loop value of the next outer loop iteration based on the inner loop value of the inner loop iteration, may include steps S341 to S343 as follows:
[0154] Step S341: Calculate the average value of the second update values of the inner loop values of the first original variable obtained in each inner loop iteration as the outer loop value of the first original variable for the next outer loop iteration.
[0155] Step S342: Calculate the weighted sum of the update value of the inner loop value of the second original variable obtained in each inner loop iteration and the outer loop value of the second original variable of the current outer loop iteration, and then calculate the average value as the outer loop value of the second original variable for the next outer loop iteration.
[0156] Step S343: Use the second update value of the inner loop value of the dual variable obtained in the last inner loop iteration as the outer loop value of the dual variable for the next outer loop iteration.
[0157] In some embodiments, the expressions corresponding to the above steps S341 to S343 are as follows:
[0158]
[0159] where K is the number of inner loop iterations, is the outer loop value of the first original variable for the next outer loop iteration, is the outer loop value of the second original variable for the next outer loop iteration, is the outer loop value of the dual variable for the next outer loop iteration.
[0160] In addition, the present application also provides a Stochastic Primal-Dual Proximal Extragradient Descent Method (abbreviated as SPDPEG) for model training, as a comparative example of the training method of the image denoising model of the present application (which can be abbreviated as AVR-SPDPEG). The idea of the SPDPEG method is to randomly select 2 image samples pair and at the k-th iteration to calculate 2 noisy gradient loss values of the loss function l, and then perform extragradient descent along the noisy gradient loss values. The difference between its implementation process and that of AVR-SPDPEG of the present application lies in that after initializing the variables (the loop initial values are x 0 、z 0 、λ 0 ), when the SPDPEG method executes the loop, only one layer of iteration is performed, and the iteration steps are as follows:
[0161] For iteration number k = 0, 1, 1, … do
[0162] Randomly select 2 image samples pair and
[0163] Update the second primal variable:
[0164] First update the first primal variable:
[0165]
[0166] First update the dual variable:
[0167] Second update the second primal variable:
[0168]
[0169] Second update the dual variable:
[0170] End for
[0171] After the loop ends, the image denoising model outputs:
[0172]
[0173] Among them, α k+1 is the weight coefficient of the quantity obtained at the (k + 1)-th iteration. Based on α k+1 , finally, the variable value output by the image denoising model is the uniform average or weighted average of the quantities obtained at each iteration.
[0174] Among them, the stochastic gradient loss value calculated in the SPDPEG method is
[0175] The SPDPEG method model output adopts a strategy of non-uniform averaging (i.e., weighted averaging) of the iteratively obtained quantities, and slightly adjusts the learning rate to obtain an accelerated convergence rate with an expectation of O(1 / t).
[0176] It can be obtained that in the above SPDPEG method, only the stochastic gradient loss value G(z, x, λ; ξ) is used to update the original variables. As is well known, although the stochastic gradient loss value can reduce the gradient calculation and thus accelerate the iterative speed of the model output towards the model optimal solution, due to the noise variance between the stochastic gradient and the full gradient, as the iteration progresses, to ensure convergence, the learning rate can only be continuously reduced or even approach 0, so the iterative speed towards the optimal solution will become slower and slower. After testing, for a generally convex objective function, taking the uniform average of the iteratively obtained quantities, the expected convergence rate of the SPDPEG method is For a strongly convex objective function, the expected convergence rates when taking the uniform average and the weighted average of the iteratively obtained quantities are O(log(t) / t) and O(1 / t), respectively.
[0177] In the above SPDPEG method,
[0178] The comparison between the SPDPEG method and the AVR-SPDPEG method is as follows:
[0179] As above, the SPDPEG method uses a single for loop, while the AVR-SPDPEG method of the present application has a two-layer loop structure in the iterative update. In the outer loop, the exact full gradient loss value at the outer loop value (reference point) is first calculated as a full traversal of the sample image set. A nested inner loop is used in the outer loop. In the inner loop, the full gradient loss value and the latest stochastic gradient loss value are used to construct a variance-reduced stochastic gradient loss value (i.e., the fused gradient loss value). At the end of the inner loop, the iteratively obtained quantities of the inner loop are used to calculate the outer loop of the next outer loop iteration and ), so that the obtained value is used to initialize the initial value of the inner loop in the next outer loop. This is a correction process to avoid the iteratively obtained quantities deviating from the convergence direction of the model optimal solution due to the existence of noise variance in the stochastic gradient, thereby accelerating the convergence rate of the model output towards the model optimal solution, so that the image denoising model has better performance in practical applications.
[0180] Furthermore, when updating the first original variable x, the AVR-SPDPEG method uses the variance-reduced fused gradient loss value and replaces the learning rate c that changes with the number of iterations in SPDPEG with a constant learning rate c k+1 .
[0181] Furthermore, on the basis of adopting the variance-reduced fused gradient loss value, a momentum acceleration technique is further introduced, that is, a momentum weight θ is introduced when updating the first original variable s , and the reference point calculated by using the quantity obtained from the iteration of the previous inner loop is used to correct the update of the first original variable, further accelerating the convergence rate of the output of the image denoising model towards the optimal solution of the model. It is worth mentioning that, in some embodiments, the momentum weight (i.e., the weight coefficient) θ s varies with the number of outer loop iterations s for the case where the loss function is generally convex. The reason for adopting the update rule in the flowchart is to facilitate the convergence proof of the method of the present application. For a strongly convex objective loss function, θ s can take a constant value.
[0182] Furthermore, in the AVR-SPDPEG method, the model output does not need to take the uniform average or weighted average of each iteration result, but directly takes the obtained at the end of the last loop.
[0183] It can be easily seen from the implementation steps of the AVR-SPDPEG method of the present application that this scheme uses the variance-reduced stochastic gradient loss value (i.e., the fused gradient loss value ) to replace the stochastic gradient loss value G(z, x, λ; ξ) in the SPDPEG method to update the original variable, so that the noise variance introduced by the stochastic gradient loss value can be significantly reduced, that is, the noise variance between the stochastic gradient and the full gradient. Therefore, a constant learning rate c can be used to ensure that the convergence of the model output towards the optimal solution of the image denoising model will not decrease due to the decrease of the learning rate, thus accelerating the convergence rate of the model output towards the optimal solution of the model; on the other hand, this proposal introduces a momentum acceleration technique on the basis of adopting the variance reduction technique to further accelerate the convergence. Through experiments, it can be proved that for a generally convex objective function, the expected convergence rate of the AVR-SPDPEG method is O(1 / t 2 ), and for a strongly convex objective function, it can achieve linear convergence and has better performance in actual use.
[0184] It can be understood that the training method disclosed in the present application can also be used for any target model that can abstract the composite convex minimization problem. The gradient descent method provided in the present application can be used for the optimization and solution of the target model in machine learning. The target problems it addresses can be a class of composite convex minimization problems as follows:
[0185]
[0186] For the description of the expressions, please refer to the corresponding positions in the above embodiments, and details will not be repeated here.
[0187] The above target problem expression includes many common models abstracted from statistics and machine learning. For example r1(x) = 0, r2 = τ1||z||1 = τ1||Fx||1, the total variation denoising (TV) model can be obtained. For another example, let r1(x) = λ||x||1 (where λ > 0 is a parameter), r2 = 0, the Lasso model can be obtained. For another example, let r2 = 0, the linear SVM model can be obtained. For yet another example, by adding a non-trivial regularization function term r2(Fx), the above target problem can include more complex structures, such as the fused Lasso model, the fused logistic regression model, and the graph-guided regularized minimization model, etc. In this embodiment, the image denoising model can include but is not limited to: the Lasso model, the linear SVM model, the fused Lasso model, the fused logistic regression model, and the graph-guided regularized minimization model. Correspondingly, the training method of the present application can be applied to the training of any of the above target models. The sample sets of different target models may be different (the sample sets are not limited to using images, texts, audio, etc.), and can be specifically selected according to the application scenarios of the models. Correspondingly, after the target model is trained, it can be applied to the corresponding application scenarios. For example, the fused logistic regression model can be applied to the image classification scenario.
[0188] Please refer to Figure 11 , Figure 11 which is a structural schematic block diagram of an embodiment of the image denoising device of the present application.
[0189] The image denoising device 100 may include an acquisition module 110 and a denoising module 120. Among them, the acquisition module 110 is used to acquire an image denoising model trained using a sample image set, where the model parameters of the image denoising model are obtained using the training method of the image denoising model in any of the above embodiments; the denoising module 120 is used to perform denoising processing on the noisy image using the image denoising model to obtain a denoised image.
[0190] Please refer to Figure 12 , Figure 12 which is a structural schematic block diagram of an embodiment of the training device of the image denoising model of the present application.
[0191] The training device 200 of the image denoising model includes an acquisition module 210 and an update module 220. Among them, the acquisition module 210 is used to acquire a sample image set, where the sample image set includes a plurality of sample image pairs, and each sample image pair respectively includes a first feature matrix obtained by performing feature extraction on the original sample image and a second feature matrix obtained by performing feature extraction on the noisy sample image obtained by adding noise to the original sample image; the update module 220 is used to calculate the full-gradient loss value and the stochastic gradient loss value through the sample image set in a cyclic iterative manner for the image denoising model, and update the model parameters of the image denoising model by using the fused gradient loss value formed by the full-gradient loss value and the stochastic gradient loss value, where the update frequency of the full-gradient loss value is lower than the update frequency of the stochastic gradient loss value.
[0192] In some embodiments, the update module 220 is specifically configured to update the model parameters of the image denoising model in a nested manner of outer loop iteration and inner loop iteration, where the full-gradient loss value is updated during the outer loop iteration, and the stochastic gradient loss value is updated during the inner loop iteration.
[0193] In some embodiments, the update module 220 is specifically configured to initialize the outer loop value of the variable in the expression of the image denoising model; where the expression is in the form of an augmented Lagrangian function, and the variable of the expression is in matrix form, and includes a first original variable corresponding to the model parameters of the image denoising model, a second original variable that satisfies the constraint relationship defined by a predetermined constraint function with the first original variable, and a dual variable related to the constraint relationship; initialize the inner loop value of the variable by using the outer loop value, and calculate the full-gradient loss value at the outer loop value based on all sample image pairs in the sample image set; update the inner loop value in an inner loop iteration manner, where in each inner loop iteration process, calculate the first stochastic gradient loss value at the inner loop value and the second stochastic gradient loss value at the outer loop value based on the sample image pair randomly selected from the sample image set, and update the inner loop value by using the fused gradient loss value formed by the first stochastic gradient loss value, the second stochastic gradient loss value, and the full-gradient loss value; update the outer loop value of the next outer loop iteration based on the inner loop value of the inner loop iteration, and return to the step set of initializing the inner loop value of the variable by using the outer loop value to perform the outer loop iteration; in response to the end of the outer loop iteration, set the model parameters of the image denoising model by using the outer loop value of the first original variable.
[0194] In some embodiments, in each inner loop iteration process, the fused gradient loss value is obtained by adding the first stochastic gradient loss value to the full-gradient loss value and subtracting the second stochastic gradient loss value.
[0195] In some embodiments, the expression includes a loss estimation term, where the loss estimation term is the loss function minus the correction function. The loss function is used to calculate the distance between the product result of the transposed matrix of the first original variable and the first feature matrix and the second feature matrix. The correction function is used to calculate the inner product of the dual variable and the constraint residual between the first original variable and the second original variable. The full gradient loss value, the first stochastic gradient loss value, and the second stochastic gradient loss value are obtained by taking the derivative of the loss function. The fused gradient loss value is further obtained by subtracting the product of the transposed matrix of the constraint function and the inner loop value of the dual variable from the calculation results of the full gradient loss value, the first stochastic gradient, and the second stochastic gradient.
[0196] In some embodiments, during each inner loop iteration, the following steps are sequentially performed: minimizing the expression with respect to the second original variable to obtain an updated value of the inner loop value of the second original variable; based on the initial value of the inner loop value of the first original variable and the initial value of the inner loop value of the dual variable, calculating a first fused gradient loss value using a first set of randomly selected sample images, and using the first fused gradient loss value to update the initial value of the inner loop value of the first original variable once to obtain a first updated value of the inner loop value of the first original variable; using the initial value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable to update the initial value of the inner loop value of the dual variable once to obtain a first updated value of the inner loop value of the dual variable; based on the first updated value of the inner loop value of the first original variable and the first updated value of the inner loop value of the dual variable, calculating a second fused gradient loss value using a second set of randomly selected sample images, and using the second fused gradient loss value to update the initial value of the inner loop value of the first original variable twice to obtain a second updated value of the inner loop value of the first original variable; using the first updated value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable to update the initial value of the inner loop value of the dual variable twice to obtain a second updated value of the inner loop value of the dual variable; taking the second updated value of the inner loop value of the first original variable, the updated value of the inner loop value of the second original variable, and the second updated value of the inner loop value of the dual variable as the initial values of the inner loop values of the first original variable, the second original variable, and the dual variable for the next inner loop iteration, and returning to the step of minimizing the expression with respect to the second original variable.
[0197] In some embodiments, the expression includes a first regularization function term and a second regularization function term. The first regularization function term takes the first original variable as a variable, and the second regularization function term takes the second original variable as a variable. The step of minimizing the expression with respect to the second original variable includes: performing a proximal mapping on the second regularization function term based on the initial value of the inner loop value of the second original variable to obtain an updated value of the inner loop value of the second original variable.
[0198] In some embodiments, based on the initial value of the inner-loop value of the first original variable and the initial value of the inner-loop value of the dual variable, the steps of calculating a first fused gradient loss value using a first group of randomly selected sample images and performing a first update on the initial value of the inner-loop value of the first original variable using the first fused gradient loss value include: adjusting the initial value of the inner-loop value of the first original variable once using the first fused gradient loss value and a preset first learning rate to obtain a first adjustment result; performing a first proximal mapping on the first regularization function term based on the first adjustment result to obtain a first proximal mapping result; determining a first update value of the inner-loop value of the first original variable based on the first proximal mapping result; based on the first update value of the inner-loop value of the first original variable and the first update value of the inner-loop value of the dual variable, the steps of calculating a second fused gradient loss value using a second randomly selected sample images and performing a second update on the initial value of the inner-loop value of the first original variable using the second fused gradient loss value include: adjusting the initial value of the inner-loop value of the first original variable twice using the second fused gradient loss value and the first learning rate to obtain a second adjustment result; performing a second proximal mapping on the first regularization function term based on the second adjustment result to obtain a second proximal mapping result; determining a second update value of the inner-loop value of the first original variable based on the second proximal mapping result.
[0199] In some embodiments, the first learning rate is a constant that does not change with inner-loop iteration and outer-loop iteration.
[0200] In some embodiments, the step of determining a first update value of the inner-loop value of the first original variable based on the first proximal mapping result includes: performing a weighted sum of the first proximal mapping result and the outer-loop value of the first original variable based on a preset weight coefficient as the first update value of the inner-loop value of the first original variable; the step of determining a second update value of the inner-loop value of the first original variable based on the second proximal mapping result includes: performing a weighted sum of the second proximal mapping result and the outer-loop value of the first original variable based on the weight coefficient as the second update value of the inner-loop value of the first original variable.
[0201] In some embodiments, the weight coefficient is a constant that does not change with inner-loop iteration and outer-loop iteration, or is a variable that changes with outer-loop iteration.
[0202] In some embodiments, the step of updating the initial value of the inner-loop value of the dual variable once using the initial value of the inner-loop value of the first original variable and the updated value of the inner-loop value of the second original variable includes: calculating a first constraint residual between the initial value of the inner-loop value of the first original variable and the updated value of the inner-loop value of the second original variable; adjusting the initial value of the inner-loop value of the dual variable using a preset second learning rate and the first constraint residual to obtain a first updated value of the inner-loop value of the dual variable; the step of updating the initial value of the inner-loop value of the dual variable twice using the first updated value of the inner-loop value of the first original variable and the updated value of the inner-loop value of the second original variable includes: calculating a second constraint residual between the first updated value of the inner-loop value of the first original variable and the updated value of the inner-loop value of the second original variable; adjusting the initial value of the inner-loop value of the dual variable using the second learning rate and the second constraint residual to obtain a second updated value of the inner-loop value of the dual variable.
[0203] In some embodiments, the expression includes a penalty term based on the constraint residuals of the first original variable and the second original variable, and the second learning rate is determined by the term coefficient of the penalty term.
[0204] In some embodiments, the step of updating the outer-loop value of the next outer-loop iteration based on the inner-loop value updated by the inner-loop iteration includes: calculating the average value of the second updated values of the inner-loop values of the first original variable obtained in each inner-loop iteration process as the outer-loop value of the first original variable of the next outer-loop iteration; calculating the weighted sum result of the updated value of the inner-loop value of the second original variable obtained in each inner-loop iteration process and the outer-loop value of the second original variable of the current outer-loop iteration, and then calculating the average value as the outer-loop value of the second original variable of the next outer-loop iteration; the second updated value of the inner-loop value of the dual variable obtained in the last inner-loop iteration process is used as the outer-loop value of the dual variable of the next outer-loop iteration.
[0205] For the description of the above steps, please refer to the corresponding positions in the method embodiments, and details are not described herein again.
[0206] Please refer to Figure 13 , Figure 13 which is a structural schematic diagram of an embodiment of an electronic device according to the present application.
[0207] The electronic device 300 includes a memory 310 and a processor 320 that are coupled to each other. The memory 310 is used to store program data, and the processor 320 is used to execute the program data to implement the steps in any of the above method embodiments.
[0208] The electronic device 300 may include, but is not limited to: personal computers (such as desktop computers, laptop computers, tablet computers, handheld computers, etc.), mobile phones, servers, wearable devices, as well as augmented reality (AR), virtual reality (VR) devices, televisions, etc., which are not limited herein.
[0209] Specifically, the processor 320 is used to control itself and the memory 310 to implement the steps in any of the above method embodiments. The processor 320 may also be referred to as a central processing unit (CPU). The processor 320 may be an integrated circuit chip with signal processing capabilities. The processor 320 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 320 may be implemented by multiple integrated circuit chips together.
[0210] Please refer to Figure 14 , Figure 14 which is a structural schematic block diagram of an embodiment of the computer-readable storage medium of the present application.
[0211] The computer-readable storage medium 400 stores program data 410, which, when executed by the processor, is used to implement the steps in any of the above method embodiments.
[0212] The computer-readable storage medium 400 may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store computer programs, or it may be a server storing the computer program. The server may send the stored computer program to other devices for running, or it may also run the stored computer program itself.
[0213] As described above, in the present application, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in the present application, the term "at least one" means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0214] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical, or other forms.
[0215] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0216] In addition, in each embodiment of the present application, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0217] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0218] The above are only the embodiments of this application, and do not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A training method for an image denoising model, characterized in that, The method includes: Obtaining a set of sample images, where the set of sample images includes multiple pairs of sample images, and each pair of sample images respectively includes a first feature matrix obtained by performing feature extraction on an original sample image and a second feature matrix obtained by performing feature extraction on a noisy sample image obtained by adding noise to the original sample image; For the image denoising model, calculating a full gradient loss value and a stochastic gradient loss value through the set of sample images in a cyclic iterative manner, and updating the model parameters of the image denoising model by using a fused gradient loss value formed by fusing the full gradient loss value and the stochastic gradient loss value, where the update frequency of the full gradient loss value is lower than the update frequency of the stochastic gradient loss value.
2. The method according to claim 1, wherein The step of calculating a full gradient loss value and a stochastic gradient loss value through the set of sample images in a cyclic iterative manner for the image denoising model, and updating the model parameters of the image denoising model by using a fused gradient loss value formed by fusing the full gradient loss value and the stochastic gradient loss value includes: Updating the model parameters of the image denoising model in a nested manner of outer loop iteration and inner loop iteration, where the full gradient loss value is updated during the outer loop iteration and the stochastic gradient loss value is updated during the inner loop iteration.
3. The method according to claim 2, characterized in that, The step of updating the model parameters of the image denoising model in a nested manner of outer loop iteration and inner loop iteration includes: Initializing an outer loop value of a variable in the expression of the image denoising model; where the expression is in the form of an augmented Lagrangian function, and the variable of the expression is in matrix form and includes a first original variable corresponding to the model parameters of the image denoising model, a second original variable that satisfies a constraint relationship defined by a predetermined constraint function with the first original variable, and a dual variable related to the constraint relationship; Initializing an inner loop value of the variable by using the outer loop value, and calculating the full gradient loss value at the outer loop value based on all the pairs of sample images in the set of sample images; Updating the inner loop value in an inner loop iterative manner, where in each inner loop iteration process, a first stochastic gradient loss value at the inner loop value and a second stochastic gradient loss value at the outer loop value calculated based on the pair of sample images randomly selected from the set of sample images are used, and the inner loop value is updated by using a fused gradient loss value formed by fusing the first stochastic gradient loss value, the second stochastic gradient loss value, and the full gradient loss value; Updating the outer loop value of the next outer loop iteration based on the inner loop value of the inner loop iteration, and returning to the step of initializing the inner loop value of the variable by using the outer loop value to perform the outer loop iteration; In response to the end of the outer loop iteration, setting the model parameters of the image denoising model by using the outer loop value of the first original variable.
4. The method according to claim 3, wherein In each iteration of the inner loop, the fused gradient loss value is obtained by adding the first stochastic gradient loss value to the full gradient loss value and subtracting the second stochastic gradient loss value; and / or The expression includes a loss estimation term, which is the loss function minus the correction function. The loss function is used to calculate the distance between the product result of the transposed matrix of the first original variable and the first feature matrix and the second feature matrix. The correction function is used to calculate the inner product of the dual variable and the constraint residual between the first original variable and the second original variable. The full gradient loss value, the first stochastic gradient loss value, and the second stochastic gradient loss value are obtained by taking the derivative of the loss function. The fused gradient loss value is further obtained by subtracting the product of the transposed matrix of the constraint function and the inner loop value of the dual variable from the calculation results of the full gradient loss value, the first stochastic gradient, and the second stochastic gradient.
5. The method according to claim 4, characterized in that, In each iteration of the inner loop, the following steps are sequentially executed: Minimize the expression with respect to the second original variable to obtain an updated value of the inner loop value of the second original variable; Based on the initial value of the inner loop value of the first original variable and the initial value of the inner loop value of the dual variable, calculate a first fused gradient loss value using a first set of randomly selected sample images, and use the first fused gradient loss value to perform a first update on the initial value of the inner loop value of the first original variable to obtain a first updated value of the inner loop value of the first original variable; Use the initial value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable to perform a first update on the initial value of the inner loop value of the dual variable to obtain a first updated value of the inner loop value of the dual variable; Based on the first updated value of the inner loop value of the first original variable and the first updated value of the inner loop value of the dual variable, calculate a second fused gradient loss value using a second set of randomly selected sample images, and use the second fused gradient loss value to perform a second update on the initial value of the inner loop value of the first original variable to obtain a second updated value of the inner loop value of the first original variable; Use the first updated value of the inner loop value of the first original variable and the updated value of the inner loop value of the second original variable to perform a second update on the initial value of the inner loop value of the dual variable to obtain a second updated value of the inner loop value of the dual variable; Take the second updated value of the inner loop value of the first original variable, the updated value of the inner loop value of the second original variable, and the second updated value of the inner loop value of the dual variable as the initial values of the inner loop values of the first original variable, the second original variable, and the dual variable for the next iteration of the inner loop, and return to the step of minimizing the expression with respect to the second original variable.
6. The method according to claim 5, wherein Based on the initial value of the inner-loop value of the first original variable and the initial value of the inner-loop value of the dual variable, calculating a first fusion gradient loss value using a first group of randomly selected sample images, and updating the initial value of the inner-loop value of the first original variable once using the first fusion gradient loss value includes: Adjusting the initial value of the inner-loop value of the first original variable once using the first fusion gradient loss value and a preset first learning rate to obtain a first adjustment result; Performing a first proximal mapping on the first regularization function term based on the first adjustment result to obtain a first proximal mapping result; Determining a first updated value of the inner-loop value of the first original variable based on the first proximal mapping result; Based on the first updated value of the inner-loop value of the first original variable and the first updated value of the inner-loop value of the dual variable, calculating a second fusion gradient loss value using a second randomly selected sample image, and updating the initial value of the inner-loop value of the first original variable twice using the second fusion gradient loss value includes: Adjusting the initial value of the inner-loop value of the first original variable twice using the second fusion gradient loss value and the first learning rate to obtain a second adjustment result; Performing a second proximal mapping on the first regularization function term based on the second adjustment result to obtain a second proximal mapping result; Determining a second updated value of the inner-loop value of the first original variable based on the second proximal mapping result.
7. The method according to claim 6, wherein The first learning rate is a constant that does not change with the inner-loop iteration and the outer-loop iteration; and / or The step of determining a first updated value of the inner-loop value of the first original variable based on the first proximal mapping result includes: Performing a weighted sum of the first proximal mapping result and the outer-loop value of the first original variable based on a preset weight coefficient as the first updated value of the inner-loop value of the first original variable; The step of determining a second updated value of the inner-loop value of the first original variable based on the second proximal mapping result includes: Performing a weighted sum of the second proximal mapping result and the outer-loop value of the first original variable based on the weight coefficient as the second updated value of the inner-loop value of the first original variable.
8. The method according to claim 4, characterized in that The step of updating the initial value of the inner-loop value of the dual variable once using the initial value of the inner-loop value of the first original variable and the updated value of the inner-loop value of the second original variable includes: Calculating a first constraint residual between the initial value of the inner-loop value of the first original variable and the updated value of the inner-loop value of the second original variable; Adjusting the initial value of the inner-loop value of the dual variable using a preset second learning rate and the first constraint residual to obtain a first updated value of the inner-loop value of the dual variable; The step of updating the initial value of the inner-loop value of the dual variable twice using the first updated value of the inner-loop value of the first original variable and the updated value of the inner-loop value of the second original variable includes: Calculate a second constraint residual between an updated value of an inner-loop value of the first original variable and an updated value of an inner-loop value of the second original variable; Adjust an initial value of an inner-loop value of the dual variable by using the second learning rate and the second constraint residual to obtain a second updated value of the inner-loop value of the dual variable; wherein the expression includes a penalty term based on a constraint residual of the first original variable and the second original variable, and the second learning rate is determined by a term coefficient of the penalty term.
9. The method according to claim 4, characterized in that The step of updating an outer-loop value of the next outer-loop iteration based on the inner-loop value of the inner-loop iteration includes: Calculate an average value of second updated values of the inner-loop value of the first original variable obtained in each inner-loop iteration process as the outer-loop value of the first original variable of the next outer-loop iteration; Calculate a weighted summation result of the updated value of the inner-loop value of the second original variable obtained in each inner-loop iteration process and the outer-loop value of the second original variable of the current outer-loop iteration, and then calculate an average value as the outer-loop value of the second original variable of the next outer-loop iteration; Use the second updated value of the inner-loop value of the dual variable obtained in the last inner-loop iteration process as the outer-loop value of the dual variable of the next outer-loop iteration.
10. An image denoising method, characterized in that, including: Obtain an image denoising model trained by using a set of sample images, wherein model parameters of the image denoising model are obtained by using the training method of the image noise reduction model according to any one of claims 1-9; Denoise a noisy image by using the image denoising model to obtain a denoised image.
Citation Information
Patent Citations
Malicious software detection method and device based on deep learning
CN110765458A
Training method of two-dimensional cryoelectron microscope image denoising modeland denoising method
CN113962887A