Artificial intelligence model training and image foreground noise reduction method, device and storage medium
By training image segmentation and noise reduction models, the image foreground noise reduction is automatically processed, which solves the problems of low efficiency and poor accuracy in the existing technology, and achieves efficient and accurate image foreground noise reduction, and retains image details and continuity.
Patent Information
- Application Number
- CN202410328429.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-03-21
AI Technical Summary
The existing image noise reduction technology requires manual operation, which has low efficiency, poor accuracy, and is prone to loss of image details and continuity.
Using the artificial intelligence model training method, by obtaining the training data set, using the image segmentation model to segment the image data into foreground and background, the image noise reduction model is used to denoiser the foreground data, and the model parameters are updated through backpropagation of the loss function, and the image segmentation model and image noise reduction model are trained to obtain the image segmentation model and image noise reduction model.
It realizes automated image foreground noise reduction, improves processing efficiency and accuracy, preserves image details and continuity, and reduces data processing volume.
Smart Images

Figure CN118195942B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an artificial intelligence model training and image foreground denoising method, a computer device, and a storage medium. Background Art
[0002] During the image capture process, such as with cameras or microscopic digital imaging devices, noise can easily interfere with the image, creating noise points that degrade image quality. Therefore, there is a need for image noise reduction. However, current image noise reduction technologies often suffer from issues such as manual labor, low efficiency, poor accuracy, and the tendency to lose detail and continuity. Summary of the Invention
[0003] In view of the common technical problems of current image denoising technology, such as the need for manual operation, low efficiency, poor precision, and easy loss of details and continuity, the purpose of the present invention is to provide an artificial intelligence model training and image foreground denoising method, computer device and storage medium.
[0004] In one aspect, an embodiment of the present invention includes an artificial intelligence model training method, the artificial intelligence model training method comprising the following steps:
[0005] Obtaining a training data set; the training data set includes training image data and label data;
[0006] Acquire an image segmentation model; the image segmentation model is used to segment the received image data into foreground data and background data;
[0007] Acquire an image denoising model; the image denoising model is used to acquire the foreground data from the image segmentation model and perform denoising on the foreground data;
[0008] Inputting the training image data into the image segmentation model, and processing the training image data in sequence by the image segmentation model and the image denoising model;
[0009] Obtaining an output result of the image denoising model;
[0010] Determine a loss function value based on the output result and the corresponding label data;
[0011] Based on the loss function value, back-propagation is performed on the model parameters of the image denoising model and / or the image segmentation model.
[0012] Furthermore, obtaining a training data set includes:
[0013] Acquire low-noise image data;
[0014] Using a semantic segmentation and annotation tool, annotating the foreground content in the low-noise image data to obtain the label data corresponding to the low-noise image data;
[0015] Noise is added to the low-noise image data to obtain the training image data.
[0016] Furthermore, the adding noise to the low-noise image data to obtain the training image data includes:
[0017] generating a two-dimensional random number matrix of the same size as the low-noise image data;
[0018] For each pixel in the low-noise image data, obtaining a random number from a corresponding same position in the two-dimensional random number matrix;
[0019] Set the noise density;
[0020] The pixel values of the pixels in the low-noise image data are set according to the following formula:
[0021]
[0022] Where x represents the position of the pixel point, I(x) represents the pixel value after the pixel point is set, P(x) represents the pixel value before the pixel point is set, and N s (x) represents the pixel value corresponding to white, N P (x) represents the pixel value corresponding to black, rand is the random number at the corresponding x position in the two-dimensional random number matrix, and p is the noise density.
[0023] Furthermore, the adding noise to the low-noise image data to obtain the training image data includes:
[0024] generating a multidimensional random number matrix of the same size as the low-noise image data, wherein the random numbers in the multidimensional random number matrix obey a Gaussian distribution;
[0025] performing mean normalization processing on the low-noise image data;
[0026] For each pixel in the low-noise image data, obtaining a random number from a corresponding same position in the two-dimensional random number matrix;
[0027] The pixel values of the pixels in the low-noise image data are set according to the following formula:
[0028]
[0029] Wherein, x represents the position of the pixel point, I(x) represents the pixel value after the pixel point is set, P(x) represents the pixel value before the pixel point is set, z is the random number corresponding to the x position in the multidimensional random number matrix, N(z) represents the random noise of the Gaussian distribution satisfied by the random number, and σ is the standard deviation of the Gaussian distribution.
[0030] Furthermore, the image segmentation model is used to:
[0031] Performing a first processing process on the received image data multiple times in succession; wherein the first processing process includes continuous double-layer convolution and maximum pooling processing, and the size of the feature map of each first processing process gradually decreases and the number of channels gradually increases;
[0032] After the processing result of the last first processing process is subjected to double-layer convolution, the second processing process is performed multiple times in succession; wherein the second processing process includes continuous upsampling processing and double-layer convolution; the feature map size of each second processing process gradually increases and the number of channels gradually decreases; the result of the upsampling processing in any second processing process is spliced with the result of the double-layer convolution in the first processing process of the corresponding feature map size and number of channels, and used as the processing result of the second processing process;
[0033] The processing result of the last second processing process is obtained, thereby obtaining the foreground data and the background data corresponding to the image data.
[0034] Furthermore, the image denoising model is used to:
[0035] performing convolution processing, multiple third processing steps, and convolution processing on the foreground data output by the image segmentation model to obtain a residual image; wherein the third processing step includes continuous convolution processing and batch normalization processing;
[0036] A pixel subtraction operation is performed on the foreground data and the residual image using a difference method to obtain an output result of the image denoising model.
[0037] Furthermore, performing back propagation update on the model parameters of the image denoising model and / or the image segmentation model according to the loss function value includes:
[0038] When the loss function value is greater than or equal to a loss function threshold, obtaining an update amplitude of model parameters between the image denoising model and the image segmentation model;
[0039] When the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model are both greater than or equal to the amplitude threshold, backpropagation updating is performed on the model parameters of the image denoising model and the image segmentation model according to the loss function value;
[0040] When at least one of the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model is less than an amplitude threshold, performing backpropagation update on the model parameters of the model corresponding to the larger model parameter update amplitude in the image denoising model and the image segmentation model according to the loss function value;
[0041] When the loss function value is less than the loss function threshold, back propagation updating of the model parameters of the image denoising model and the image segmentation model is stopped.
[0042] On the other hand, an embodiment of the present invention further includes a method for image foreground noise reduction, the method comprising the following steps:
[0043] Obtain image data to be processed;
[0044] Obtain an image segmentation model and an image denoising model; the image segmentation model and the image denoising model are trained using the artificial intelligence model training method of the embodiment;
[0045] Inputting the image data to be processed into the image segmentation model, and processing the image data in sequence by the image segmentation model and the image denoising model;
[0046] Obtain an output result of the image denoising model.
[0047] On the other hand, an embodiment of the present invention also includes a computer device comprising a memory and a processor, the memory being used to store at least one program, and the processor being used to load at least one program to execute an artificial intelligence model training method and / or an image foreground denoising method in an embodiment.
[0048] On the other hand, an embodiment of the present invention also includes a storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute an artificial intelligence model training method and / or an image foreground denoising method in the embodiment.
[0049] The beneficial effects of the present invention are: the artificial intelligence model training method in the embodiment can train an image segmentation model and an image denoising model, and the image segmentation model and the image denoising model together have the performance of foreground-background segmentation and foreground denoising of image data, and can denoise the foreground part of the image, while the background part does not need to be denoised. Since the foreground part of the image is more affected by noise, the amount of data processing required for denoising can be reduced while maintaining a good denoising effect. On the other hand, by using the image segmentation model and the image denoising model, manual operations can be reduced and processing efficiency and accuracy can be improved. The image foreground denoising method in the embodiment can use the trained image segmentation model and image denoising model to automatically denoise the image data to be processed, has good processing efficiency and accuracy, and can retain the details and continuity of the image data to be processed. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the steps of the artificial intelligence model training method in an embodiment;
[0051] Figure 2 A schematic diagram of an artificial intelligence model training method in an embodiment;
[0052] Figure 3 Schematic diagram of the structure of the image segmentation model in the embodiment;
[0053] Figure 4 Schematic diagram of the structure of the image denoising model in the embodiment;
[0054] Figure 5 Schematic diagram of the steps of the image foreground noise reduction method in the embodiment. DETAILED DESCRIPTION
[0055] In this embodiment, refer to Figure 1 , the artificial intelligence model training method includes the following steps:
[0056] S1. Obtain training dataset;
[0057] S2. Obtain an image segmentation model;
[0058] S3. Obtain an image denoising model;
[0059] S4. Input the training image data into the image segmentation model, and process it in sequence by the image segmentation model and the image denoising model;
[0060] S5. Obtain the output result of the image denoising model;
[0061] S6. Determine the loss function value based on the output result and the corresponding label data;
[0062] S7. Based on the loss function value, backpropagate and update the model parameters of the image denoising model and / or the image segmentation model.
[0063] The principles of steps S1-S7 are as follows Figure 2 shown.
[0064] In step S1 , the training data set includes training image data and label data. Specifically, the training data set includes a plurality of training image data, and each training image data has corresponding label data.
[0065] In step S2, a U-Net network can be established as an image segmentation model. The image segmentation model can receive image data and segment the received image data into foreground data and background data. The image segmentation model can receive a training image data, process the training image data, and segment it into foreground data and background data.
[0066] In step S3, a convolutional neural network can be established as an image denoising model. The image denoising model can receive the foreground data output by the image segmentation model and perform denoising on the foreground data.
[0067] In step S4, the training image data obtained in step S1 can be input into the image segmentation model, where it is processed sequentially by the image segmentation model and the image denoising model. Specifically, the image segmentation model segments the training image data into foreground data and background data, and then passes the foreground data to the image denoising model, which then denoises the foreground data, thereby obtaining the output result in step S5.
[0068] In step S6, a loss function value is calculated based on the output of the image denoising model obtained in step S5 and the label data corresponding to the training image data input to the image segmentation model in step S4. Specifically, the L2 distance between the output and the label data can be calculated as the loss function value.
[0069] In step S7, based on the loss function value calculated in step S6, it is determined whether to perform backpropagation update on the model parameters of the image denoising model and / or image segmentation model, or if it is determined to perform backpropagation update, the update amplitude of the model parameters is determined based on the size of the loss function value.
[0070] In this embodiment, by executing steps S1-S7, an image segmentation model and an image denoising model can be trained. The image segmentation model and the image denoising model together have the performance of performing foreground and background segmentation and foreground denoising on image data. The foreground part of the image can be denoised, while the background part does not need to be denoised. Since the foreground part of the image is more affected by noise, the amount of data processing required for denoising can be reduced while maintaining a good denoising effect. On the other hand, by using the image segmentation model and the image denoising model, manual operations can be reduced and processing efficiency and accuracy can be improved.
[0071] In this embodiment, when executing step S1, that is, the step of obtaining a training data set, the following steps may be specifically performed:
[0072] S101. Obtain low-noise image data;
[0073] S102. Use a semantic segmentation and annotation tool to annotate the foreground content in the low-noise image data to obtain label data corresponding to the low-noise image data;
[0074] S103. Add noise to the low-noise image data to obtain training image data.
[0075] When executing step S101, a visible light imaging device can be used to collect image data under conditions such as good lighting. The collected image data is less affected by noise than images collected under other conditions. Therefore, the image data obtained by executing step S101 is low-noise image data.
[0076] In step S102 , a semantic segmentation and annotation tool is used to annotate the foreground content in each low-noise image data. The obtained label data represents the content of the foreground part in the corresponding low-noise image data.
[0077] In step S103, at least one type of noise may be added to the noisy image data to obtain training image data.
[0078] In this embodiment, when executing step S103, that is, adding noise to the low-noise image data to obtain training image data, the following steps may be specifically performed:
[0079] S10301A. Generate a two-dimensional random number matrix of the same size as the low-noise image data;
[0080] S10302A. For each pixel in the low-noise image data, obtain a random number from the corresponding same position in the two-dimensional random number matrix;
[0081] S10303A. Set noise density;
[0082] S10304A. Set the pixel values of the pixels in the low-noise image data according to the following formula:
[0083]
[0084] Where x represents the position of the pixel point, I(x) represents the pixel value after the pixel point is set, P(x) represents the pixel value before the pixel point is set, and N S (x) represents the pixel value corresponding to white, N P (x) represents the pixel value corresponding to black, rand is the random number at the corresponding x position in the two-dimensional random number matrix, and p is the noise density. rand is uniformly distributed in the range [0, 1).
[0085] Steps S10301A-S10304A are the first execution method of step S103.
[0086] In step S10304A, the pixel value corresponding to white is 255, and the pixel value corresponding to black is 0. By executing steps S10301A-S10304A, some pixels in the low-noise image data can be randomly selected according to the noise density p (which means the ratio of the number of noise points per unit area in the low-noise image data to the total number of pixels) and set to white pixels or black pixels, while the pixel values of other pixels remain unchanged.
[0087] In this embodiment, when executing step S103, that is, adding noise to the low-noise image data to obtain training image data, the following steps may be specifically performed:
[0088] S10301B generates a multidimensional random number matrix of the same size as the low-noise image data;
[0089] S10302B. Perform mean normalization on low-noise image data;
[0090] S10303B. For each pixel in the low-noise image data, a random number is obtained from the corresponding same position in the two-dimensional random number matrix;
[0091] S10304B. Set the pixel values of the pixels in the low-noise image data according to the following formula:
[0092]
[0093] Wherein, x represents the position of the pixel point, I(x) represents the pixel value after the pixel point is set, P(x) represents the pixel value before the pixel point is set, z is the random number corresponding to the x position in the multidimensional random number matrix, N(z) represents the random noise of the Gaussian distribution satisfied by the random number, and σ is the standard deviation of the Gaussian distribution.
[0094] Steps S10301B-S10304B are a second execution method of step S103.
[0095] In step S10301B, the random numbers in the multidimensional random number matrix obey a Gaussian distribution with a mean of 0 and a standard deviation of σ, where the standard deviation σ can be set to a fixed value or adjusted according to the fitting of the hyperparameters of the image segmentation model and the image denoising model.
[0096] In steps S10301B-S10304B, after the original low-noise image data is mean-normalized, the noise matrix is superimposed on the low-noise image data, the elements in the multidimensional random number matrix are cropped, and the elements greater than 1 in the multidimensional random number matrix are set to 1, and the elements less than 0 are set to 0. Finally, the superimposed multidimensional random number matrix is inversely mean-normalized to obtain an image with random noise, i.e., the training image data.
[0097] Steps S10301A-S10304A and steps S10301B-S10304B can respectively add a type of noise to the original low-noise image data, so that the low-noise image data less affected by noise becomes training image data with a known type of noise. Steps S10301A-S10304A or steps S10301B-S10304B or steps for adding other types of noise can be executed separately to add one type of noise to the low-noise image data; steps S10301A-S10304A or steps S10301B-S10304B or steps for adding other types of noise can also be executed jointly. For example, on the basis of first executing steps S10301A-S10304A, the obtained training image data is used as the "low-noise image data" in steps S10301B-S10304B to execute steps S10301B-S10304B, thereby adding multiple types of noise to the low-noise image data, which is conducive to training the image denoising model's denoising ability for image data mixed with multiple noises.
[0098] Steps S10301A-S10304A and steps S10301B-S10304B are two types of image noise that may be generated by visible light imaging devices during the image collection process. Based on the content, application scenario and requirements of the original image, you can select the appropriate noise type to add to the image for constructing a dataset to better simulate the actual image noise and effectively improve the training effect of the model.
[0099] In this embodiment, the image segmentation model is based on the U-Net semantic segmentation network, and its structure is as follows: Figure 3 As shown. Figure 3The input of the image segmentation model can receive a set of noisy image data. It first passes through two layers of convolutional neural networks and a maximum pooling layer to extract a feature layer. This step is repeated four times, each time gradually reducing the size of the feature map and increasing the number of channels to extract higher-level features. Finally, it passes through two more layers of convolutional neural networks to achieve feature extraction at five scales. The purpose of repeatedly performing convolution and pooling operations is to capture contextual information at different scales, enabling the model to understand the global semantic information of the image. After feature extraction, an upsampling operation is performed to align the feature map size with the corresponding layer. Then, skip connections are used to splice and fuse the features of the corresponding layer. The model then passes through two layers of convolutional neural networks to reconstruct features using low-level features and contextual information. This step is repeated four times, each time gradually increasing the size of the feature map and reducing the number of channels to achieve more refined feature reconstruction. Finally, another layer of convolutional neural network is used to separate the foreground and background of the noisy image.
[0100] based on Figure 3 In the structure shown, when the image segmentation model processes the training image data in step S4, the following steps may be performed:
[0101] S401. Performing a first processing step on the received image data multiple times in succession; wherein the first processing step includes a continuous double-layer convolution and maximum pooling process, wherein the feature map size of each first processing step gradually decreases and the number of channels gradually increases;
[0102] S402. After the result of the last first processing step is subjected to double-layer convolution, a second processing step is performed in succession multiple times. The second processing step includes successive upsampling and double-layer convolution. The feature map size of each second processing step gradually increases, and the number of channels gradually decreases. The upsampling result of any second processing step is concatenated with the result of the double-layer convolution of the first processing step with the corresponding feature map size and number of channels, and the result is used as the processing result of the current second processing step.
[0103] S403. Obtain the processing result of the last second processing step, thereby obtaining foreground data and background data corresponding to the image data.
[0104] Reference Figure 3 In step S401, along Figure 3 From top to bottom on the left, the feature map size of each first processing process (that is, double-layer convolution + maximum pooling layer) gradually decreases, and the number of channels gradually increases; in step S402, along Figure 3From bottom to top on the left, the feature map size of each second process (i.e., upsampling + double-layer convolution) gradually increases, and the number of channels gradually decreases. In this way, except for one second process, each of the other second processes has a corresponding first process, and the feature map size and number of channels used by the first and second processes are equal.
[0105] In this embodiment, the image denoising model is based on a convolutional neural network, and its structure is as follows: Figure 4 As shown. Figure 4 The image segmentation model processes the image to obtain the separated foreground and background of the noisy image. The foreground part is taken and cropped and filled. The foreground data of the noisy image is input into the image denoising model constructed using a denoising convolutional neural network. In the image denoising model, a layer of convolutional neural network is first used to extract image features and change the number of channels in the feature map. Then, a layer of convolutional neural network and a repeated batch normalization module are used to extract features. The batch normalization layer can help the gradient propagate faster, accelerate the training process and improve accuracy. Repeating the module multiple times can extract multi-scale features. The last layer of convolutional neural network receives the data from the previous layer. The number of output channels of this convolutional layer is consistent with the number of input channels of the denoising convolutional neural network. The output result is the residual image of the noisy foreground. The difference method is used to perform a pixel-by-pixel subtraction operation between the noisy foreground image and the output residual image to obtain the denoised foreground image.
[0106] based on Figure 4 In the structure shown, when the image denoising model processes the training image data in step S4, the following steps may be performed:
[0107] S404. Performing convolution processing, multiple third processing steps, and convolution processing on the foreground data output by the image segmentation model to obtain a residual image; wherein the third processing step includes continuous convolution processing and batch normalization processing;
[0108] S405. Use the difference method to perform pixel subtraction operation on the foreground data and the residual image to obtain the output result of the image denoising model.
[0109] In this embodiment, based on Figure 3 The image segmentation model shown and Figure 4 The image denoising model shown in FIG. 1 performs steps S4-S7 as follows:
[0110] In step S4, the training image data in the training data set obtained in process S1 are shuffled and batched, and input into the image segmentation model and image denoising model for forward propagation layer-by-layer calculation; in step S5, the output result of the image denoising model is the prediction result; in step S6, the loss function value is calculated by comparing the prediction result with the label data corresponding to the same training image data in the training data set; in step S7, the gradient of each parameter in the image segmentation model and the image denoising model is calculated by back propagation based on the loss function value, and the gradient obtained by back propagation is combined with the gradient descent algorithm to update the model parameters.
[0111] If the loss function value obtained after executing steps S4-S7 once does not converge (for example, it is not less than a threshold, or the cumulative number of executions of steps S4-S7 does not reach a preset value), after executing steps S4-S7 once, jump back to step S4, select another training image data in the training data set and input it into the image segmentation model, and then execute steps S4-S7 again until the loss function value obtained does not converge. By repeating steps S4-S7, the parameters of the image segmentation model and the image denoising model are iteratively updated, the loss value is minimized, and the optimal model is saved, completing the training of the image segmentation model and the image denoising model.
[0112] In this embodiment, when executing step S7, that is, performing backpropagation update on the model parameters of the image denoising model and / or the image segmentation model according to the loss function value, the following steps may be specifically performed:
[0113] S701. When the loss function value is greater than or equal to the loss function threshold, obtain the model parameter update amplitude between the image denoising model and the image segmentation model;
[0114] S702. When the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model are both greater than or equal to the amplitude threshold, the model parameters of the image denoising model and the image segmentation model are back-propagated and updated according to the loss function value;
[0115] S703. When at least one of the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model is less than the amplitude threshold, backpropagation update is performed on the model parameters of the image denoising model and the image segmentation model corresponding to the larger model parameter update amplitude according to the loss function value;
[0116] S704. When the loss function value is less than the loss function threshold, stop backpropagation updating of the model parameters of the image denoising model and the image segmentation model.
[0117] First, determine whether the loss function value calculated in step S6 is greater than the loss function threshold. If the loss function value is greater than or equal to the loss function threshold, then it indicates that the training of the image denoising model and / or image segmentation model needs to continue, and therefore execute steps S701-S703. If the loss function value is less than the loss function threshold, then it indicates that the training of the image denoising model and / or image segmentation model is complete, and execute step S704 to stop backpropagation updating of the model parameters of the image denoising model and image segmentation model, and do not execute the next step S4-S7.
[0118] In step S701, the model parameter update amplitudes of the previous image denoising model and the previous image segmentation model are obtained. The model parameter update amplitude of the previous image denoising model refers to the difference (which can be an absolute value) between the model parameters before and after the image denoising model update when step S7 of steps S4-S7 was last executed. The same applies to the model parameter update amplitude of the previous image segmentation model.
[0119] If the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model are both greater than or equal to the amplitude threshold, then step S702 is executed to backpropagate and update the model parameters of the image denoising model and the image segmentation model according to the loss function value.
[0120] If at least one of the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model is less than the amplitude threshold, then step S703 is executed. Based on the loss function value, the model parameters of the image denoising model and the image segmentation model corresponding to the larger model parameter update amplitude are backpropagated and updated, while the model parameters of the model corresponding to the larger model parameter update amplitude remain unchanged. For example, if the model parameter update amplitude of the image denoising model is smaller than the model parameter update amplitude of the image segmentation model, then when step S703 is executed this time, only the model parameters of the image segmentation model are backpropagated and updated based on the loss function value, while the model parameters of the image segmentation model remain unchanged.
[0121] In this embodiment, the principle of executing steps S701-S704 is that: when the loss function value is greater than or equal to the loss function threshold, the model parameters of at least one of the image denoising model and the image segmentation model will be back-propagated and updated; and when the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model are both greater than or equal to the amplitude threshold, it indicates that in the previous round of training, the model parameters of the image denoising model and the image segmentation model have been greatly adjusted, and the training completion of the image denoising model and the image segmentation model is low. Therefore, this round of training has a negative impact on the image denoising model and the image segmentation model. The model parameters of the image denoising model and the image segmentation model are all back-propagated and updated, which is conducive to speeding up the training progress; when at least one of the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model is less than the amplitude threshold, it indicates that in the previous round of training, the adjustment amplitude of the model parameters of at least one of the image denoising model and the image segmentation model is small, and the training completion degree of the image denoising model and the image segmentation model is high. Therefore, in this round of training, only the model parameters of the model with the larger adjustment amplitude are back-propagated and updated, which is conducive to reducing the amount of training data processing and realizing the synchronous training of the image denoising model and the image segmentation model.
[0122] The optimal model produced during the training process can be used to perform neural network inference tasks. Specifically, a set of preprocessed image data is input into the prediction model. Using the trained optimal model parameters, the neural network output is calculated through forward propagation to obtain the predicted result for the input sample, and the inference process is completed.
[0123] The above reasoning task is actually an image foreground denoising method performed using the trained image segmentation model and image denoising model. Figure 5 , the image foreground denoising method includes the following steps:
[0124] P1. Get the image data to be processed;
[0125] P2. Obtain image segmentation model and image denoising model;
[0126] P3. Input the image data to be processed into the image segmentation model, and the image segmentation model and image denoising model process it in sequence;
[0127] P4. Get the output of the image denoising model.
[0128] In this embodiment, the image segmentation model and image denoising model used in step P2 have been trained using an artificial intelligence model training method.
[0129] The image segmentation model and image denoising model trained in steps S1-S7 are capable of performing foreground segmentation and denoising on the image data being processed, thereby automatically denoising the image data being processed with high processing efficiency and accuracy. The output result obtained in step P4 is image data with the same foreground content as the image data being processed in step P1, but with lower noise.
[0130] A computer program that executes the artificial intelligence model training method and / or image foreground denoising method in this embodiment can be written and written into a computer device or storage medium. When the computer program is read out and run, the artificial intelligence model training method and / or image foreground denoising method in this embodiment is executed, thereby achieving the same technical effect as the artificial intelligence model training method and / or image foreground denoising method in the embodiment.
[0131] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. In addition, the descriptions of up, down, left, right, etc. used in this disclosure are only relative to the relative positional relationships of the components of the present disclosure in the accompanying drawings. The singular forms of "a", "" and "the" used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as those generally understood by those skilled in the art. The terms used in the specification of this embodiment are only for describing specific embodiments and are not intended to limit the invention. The term "and / or" used in this embodiment includes any combination of one or more related listed items.
[0132] It should be understood that, although the present disclosure may adopt the term first, second, third etc. to describe various elements, these elements should not be limited to these terms.These terms are only used to distinguish the elements of the same type from each other.For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element.The use of any and all examples or exemplary language ("for example", "such as" etc.) provided by the present embodiment is only intended to better illustrate embodiments of the present invention, and unless otherwise required, the scope of the present invention will not be limited.
[0133] It should be appreciated that embodiments of the present invention can be implemented or practiced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The methods can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and figures described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed application-specific integrated circuit for this purpose.
[0134] In addition, the operations of the processes described in this embodiment may be performed in any suitable order, unless otherwise indicated in this embodiment or otherwise clearly contradicted by the context. The processes described in this embodiment (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. A computer program includes multiple instructions that can be executed by one or more processors.
[0135] Furthermore, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, RAM, ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted over a wired or wireless network. When such media includes instructions or programs that implement the above steps in conjunction with a microprocessor or other data processor, the invention of this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0136] The computer program can be applied to input data to perform the functions of the present embodiment, thereby converting the input data to generate output data that is stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.
[0137] The above are merely preferred embodiments of the present invention. The present invention is not limited to the aforementioned embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods may be made.
Claims
1. An artificial intelligence model training method, characterized in that: The artificial intelligence model training method includes: Obtaining a training data set; the training data set includes training image data and label data; Acquire an image segmentation model; the image segmentation model is used to segment the received image data into foreground data and background data; Obtaining an image denoising model; the image denoising model is used to obtain the foreground data from the image segmentation model and perform denoising on the foreground data; the image denoising model is used to: sequentially perform convolution processing, multiple third processing steps, and convolution processing on the foreground data output by the image segmentation model to obtain a residual image; wherein the third processing step includes continuous convolution processing and batch normalization processing; performing a pixel subtraction operation on the foreground data and the residual image using a difference image method to obtain an output result of the image denoising model; Inputting the training image data into the image segmentation model, and processing the training image data in sequence by the image segmentation model and the image denoising model; Obtaining an output result of the image denoising model; Determine a loss function value based on the output result and the corresponding label data; Performing backpropagation update on model parameters of the image denoising model and / or the image segmentation model according to the loss function value; The back-propagation updating of the model parameters of the image denoising model and / or the image segmentation model according to the loss function value includes: When the loss function value is greater than or equal to a loss function threshold, obtaining an update amplitude of model parameters between the image denoising model and the image segmentation model; When the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model are both greater than or equal to the amplitude threshold, backpropagation updating is performed on the model parameters of the image denoising model and the image segmentation model according to the loss function value; When at least one of the model parameter update amplitude of the image denoising model and the model parameter update amplitude of the image segmentation model is less than an amplitude threshold, performing backpropagation update on the model parameters of the model corresponding to the larger model parameter update amplitude in the image denoising model and the image segmentation model according to the loss function value; When the loss function value is less than the loss function threshold, back propagation updating of the model parameters of the image denoising model and the image segmentation model is stopped.
2. The artificial intelligence model training method according to claim 1, characterized in that: The obtaining of the training data set includes: Acquire low-noise image data; Using a semantic segmentation and annotation tool, annotating the foreground content in the low-noise image data to obtain the label data corresponding to the low-noise image data; Noise is added to the low-noise image data to obtain the training image data.
3. The artificial intelligence model training method according to claim 2, characterized in that: The adding noise to the low-noise image data to obtain the training image data includes: generating a two-dimensional random number matrix of the same size as the low-noise image data; For each pixel in the low-noise image data, obtaining a random number from a corresponding same position in the two-dimensional random number matrix; Set the noise density; The pixel values of the pixels in the low-noise image data are set according to the following formula: Where x represents the position of the pixel point, I(x) represents the pixel value after the pixel point is set, P(x) represents the pixel value before the pixel point is set, and N S (x) represents the pixel value corresponding to white, N P (x) represents the pixel value corresponding to black, rand is the random number at the corresponding x position in the two-dimensional random number matrix, and p is the noise density.
4. The artificial intelligence model training method according to claim 2, characterized in that: The adding noise to the low-noise image data to obtain the training image data includes: generating a multidimensional random number matrix of the same size as the low-noise image data, wherein the random numbers in the multidimensional random number matrix obey a Gaussian distribution; performing mean normalization processing on the low-noise image data; For each pixel point in the low-noise image data, obtaining a random number from a corresponding same position in the multidimensional random number matrix; The pixel values of the pixels in the low-noise image data are set according to the following formula: Wherein, x represents the position of the pixel point, I(x) represents the pixel value after the pixel point is set, P(x) represents the pixel value before the pixel point is set, z is the random number corresponding to the x position in the multidimensional random number matrix, N(z) represents the random noise of the Gaussian distribution satisfied by the random number, and σ is the standard deviation of the Gaussian distribution.
5. The artificial intelligence model training method according to claim 1, characterized in that: The image segmentation model is used to: Performing a first processing process on the received image data multiple times in succession; wherein the first processing process includes continuous double-layer convolution and maximum pooling processing, and the size of the feature map of each first processing process gradually decreases and the number of channels gradually increases; After the processing result of the last first processing process is subjected to double-layer convolution, the second processing process is performed multiple times in succession; wherein the second processing process includes continuous upsampling processing and double-layer convolution; the feature map size of each second processing process gradually increases and the number of channels gradually decreases; the result of the upsampling processing in any second processing process is spliced with the result of the double-layer convolution in the first processing process of the corresponding feature map size and number of channels, and used as the processing result of the second processing process; The processing result of the last second processing process is obtained, thereby obtaining the foreground data and the background data corresponding to the image data.
6. A method for image foreground noise reduction, characterized in that: The image foreground noise reduction method comprises: Obtain image data to be processed; Obtain an image segmentation model and an image denoising model; the image segmentation model and the image denoising model are trained by the artificial intelligence model training method according to any one of claims 1 to 5; Inputting the image data to be processed into the image segmentation model, and processing the image data in sequence by the image segmentation model and the image denoising model; Obtain an output result of the image denoising model.
7. A computer device, characterized in that: It includes a memory and a processor, the memory is used to store at least one program, and the processor is used to load at least one program to execute the artificial intelligence model training method described in any one of claims 1 to 5 and / or the image foreground denoising method described in claim 6.
8. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program, when executed by the processor, is used to execute the artificial intelligence model training method described in any one of claims 1-5 and / or the image foreground denoising method described in claim 6.
Citation Information
Patent Citations
Image processing method and electronic equipment
CN110781899A
Video semantic noise reduction method and device and electronic equipment
CN116188320A