Image processing method, device, equipment, storage medium and chip
Patent Information
- Application Number
- CN202511317798.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-09-15
AI Technical Summary
但是当前由低分辨率生成的高分辨率图像通常伴随着大量的噪声
[0055] This disclosure acquires image data and first noisy image data, processes the image data and first noisy image data using a pre-trained target image enhancement model, and obtains an enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on hyperparameters. The first noisy image data is determined by hyperparameters. By enhancing the image through the target enhancement model, the resolution of the image can be enhanced while suppressing the noise in the image.
Smart Images

Figure CN120852220B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to an image processing method, apparatus, device, storage medium, and chip. Background Technology
[0002] To improve image quality, related fields typically require high-resolution output images. However, current high-resolution images generated from low-resolution images often come with significant noise. Furthermore, traditional algorithms for noise suppression often degrade image quality.
[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0004] To overcome the problems existing in related technologies, this disclosure provides an image processing method, apparatus, device, storage medium, and chip.
[0005] According to a first aspect of the present disclosure, an image processing method is provided, comprising:
[0006] Acquire image data and first noisy image data;
[0007] The image data and the first noisy image data are processed by a pre-trained target image enhancement model to obtain the enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on the hyperparameters. The first noisy image data is determined by the hyperparameters.
[0008] In one embodiment of this disclosure, the training method for the target image enhancement model includes:
[0009] A first training dataset comprising multiple first training samples and second noisy image data are obtained; each first training sample includes at least a corresponding first low-quality image and a first high-quality image; the second noisy image data is determined by hyperparameters.
[0010] The first training dataset and the second noisy image data are respectively input into the pre-trained first image enhancement model and the target image enhancement model;
[0011] The total loss function value is determined by the first high-quality image, hyperparameters, the training results of the pre-trained first image augmentation model, and the training results of the target image augmentation model.
[0012] The target image enhancement model is trained based on the total loss function value. When the total loss function value converges, the pre-trained target image enhancement model is obtained.
[0013] In one embodiment of this disclosure, the total loss function value is determined using a first high-quality image, hyperparameters, the training result of a pre-trained first image enhancement model, and the training result of a target image enhancement model, including:
[0014] Determine the difference between the training result of the pre-trained first image augmentation model and the training result of the target image augmentation model;
[0015] The first loss function value corresponding to the target image enhancement model is determined based on the training results of the target image enhancement model and the first high-quality image.
[0016] The total loss function value is determined by the hyperparameters, the difference value, and the first loss function value.
[0017] In one embodiment of this disclosure, the total loss function value is determined by hyperparameters, difference values, and a first loss function value, including:
[0018] The weights corresponding to the difference values and the first loss function values are determined by using hyperparameters.
[0019] The total loss function value is determined by multiplying the difference value by its corresponding weight and by multiplying the first loss function value by its corresponding weight.
[0020] In one embodiment of this disclosure, the training process of the first image enhancement model includes:
[0021] Obtain a second training dataset that includes multiple second training samples, wherein the second training samples include at least the corresponding second low-quality image and the second high-quality image;
[0022] Input the second training dataset into the first image enhancement model to be trained, and obtain the training result corresponding to the first image enhancement model to be trained.
[0023] The value of the second loss function is determined based on the training result of the first image enhancement model to be trained and the second high-quality image.
[0024] The first image augmentation model to be trained is trained based on the second loss function value. When the second loss function value converges, the pre-trained first image augmentation model is obtained.
[0025] In one embodiment of this disclosure, the image noise parameters corresponding to the first high-quality image and the second high-quality image are less than a first preset threshold, and the corresponding image detail parameters are greater than a second preset threshold.
[0026] In one embodiment of this disclosure, the image noise parameters corresponding to the enhanced image are adjusted by adjusting hyperparameters.
[0027] According to a second aspect of the present disclosure, an image processing apparatus is provided, comprising:
[0028] In one embodiment of this disclosure, a first acquisition module is used to acquire image data and first noisy image data;
[0029] The processing module is used to process the image data and the first noisy image data through a pre-trained target image enhancement model to obtain the enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on the hyperparameters. The first noisy image data is determined by the hyperparameters.
[0030] In one embodiment of this disclosure, the apparatus further includes:
[0031] The second acquisition module is used to acquire a first training dataset including multiple first training samples and second noisy image data; each first training sample includes at least a corresponding first low-quality image and a first high-quality image; the second noisy image data is determined by hyperparameters;
[0032] The input module is used to input the first training dataset and the second noisy image data into the pre-trained first image enhancement model and the target image enhancement model, respectively;
[0033] The first determining module is used to determine the total loss function value through the first high-quality image, hyperparameters, the training result of the pre-trained first image enhancement model, and the training result of the target image enhancement model;
[0034] The first training module is used to train the target image enhancement model based on the total loss function value. In response to the convergence of the total loss function value, a pre-trained target image enhancement model is obtained.
[0035] In one embodiment of this disclosure, the determining module includes:
[0036] The first determining unit is used to determine the difference between the training result corresponding to the pre-trained first image enhancement model and the training result corresponding to the target image enhancement model.
[0037] The second determining unit is used to determine the first loss function value corresponding to the target image enhancement model based on the training result of the target image enhancement model and the first high-quality image;
[0038] The third determining unit is used to determine the total loss function value through hyperparameters, difference values, and the first loss function value.
[0039] In one embodiment of this disclosure, the third determining unit includes:
[0040] The first determining subunit is used to determine the weights corresponding to the difference value and the first loss function value through hyperparameters;
[0041] The second determining subunit is used to determine the total loss function value by multiplying the difference value by the corresponding weight and the first loss function value by the corresponding weight.
[0042] In one embodiment of this disclosure, the apparatus further includes:
[0043] The third acquisition module is used to input the second training dataset into the first image enhancement model to be trained, and obtain the training result corresponding to the first image enhancement model to be trained.
[0044] The second determining module is used to determine the value of the second loss function based on the training result of the first image enhancement model to be trained and the second high-quality image.
[0045] The second training module is used to train the first image augmentation model to be trained based on the second loss function value. In response to the convergence of the second loss function value, the pre-trained first image augmentation model is obtained.
[0046] In one embodiment of this disclosure, the image noise parameters corresponding to the first high-quality image and the second high-quality image are less than a first preset threshold, and the corresponding image detail parameters are greater than a second preset threshold.
[0047] In one embodiment of this disclosure, the apparatus further includes:
[0048] The adjustment module is used to adjust the image noise parameters corresponding to the enhanced image by adjusting hyperparameters.
[0049] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0050] processor;
[0051] Memory used to store processor-executable instructions;
[0052] The processor is configured to implement any of the image processing methods described in the first aspect above.
[0053] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, which, when the instructions in the storage medium are executed by a processor of a terminal, enables the terminal to perform any of the image processing methods described in the first aspect.
[0054] According to a fifth aspect of the present disclosure, a computer program product is provided, a chip including a processing circuit for performing the steps of the image processing method as described in the first aspect. The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0055] This disclosure acquires image data and first noisy image data, processes the image data and first noisy image data using a pre-trained target image enhancement model, and obtains an enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on hyperparameters. The first noisy image data is determined by hyperparameters. By enhancing the image through the target enhancement model, the resolution of the image can be enhanced while suppressing the noise in the image.
[0056] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0057] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0058] Figure 1 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 1 .
[0059] Figure 2 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 2 .
[0060] Figure 3 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 3 .
[0061] Figure 4 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 4 .
[0062] Figure 5 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 5 .
[0063] Figure 6 This is a block diagram of an image processing apparatus according to an exemplary embodiment of the present disclosure.
[0064] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0065] Exemplary embodiments of this disclosure will be described in detail herein, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.
[0066] The embodiments described below, which are examples of some of the embodiments of this disclosure, do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0067] The specific implementation methods of the embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0068] Figure 1 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 1 .
[0069] like Figure 1 As shown, it includes the following steps.
[0070] S101, acquire image data and first noisy image data.
[0071] In some embodiments, the image data includes images whose resolution and / or noise parameters are difficult to meet user requirements. This disclosure does not limit the type of image; for example, images in this disclosure may include image frames from video, grayscale images, indexed images, depth images, thermal imaging images, medical images, synthetic images, and panoramic images.
[0072] In some embodiments, the first noisy image data includes an image used to add noise to an enhanced image obtained by enhancing image data. In embodiments of this disclosure, the first noisy image data can be determined by hyperparameters. For example, the first noisy image data may be a Gaussian noise image with a mean hyperparameter.
[0073] It should be noted that after the image data is enhanced by the target image enhancement model, the resulting enhanced image may have insufficient high-frequency noise. Insufficient high-frequency noise may lead to distortion of the enhanced image. To avoid distortion of the enhanced image, this disclosure allows the image data and the first noise image data to be input into the target image enhancement model together.
[0074] This disclosure allows for the acquisition of image data through various methods, such as camera capture and online resource download. This disclosure does not limit the methods for acquiring image data.
[0075] S102, the image data and the first noisy image data are processed by the pre-trained target image enhancement model to obtain the enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on the hyperparameters. The first noisy image data is determined by the hyperparameters.
[0076] In some embodiments, the pre-trained target image augmentation model can be an intelligent model based on an AISR model. AISR models are primarily based on a fusion architecture of Generative Adversarial Networks (GANs) and Diffusion Models. For example, the architecture of the pre-trained target image augmentation model may include a generator and a discriminator. The generator can use a U-Net structure as its foundation, extracting multi-scale features through an encoder-decoder design. The encoder downsamples to extract abstract features, and the decoder upsamples to recover details. The discriminator can use a PatchGAN structure to distinguish between real and fake local regions of the image, improving the local realism of the generated image.
[0077] For example, the target image enhancement model in this disclosure can also be augmented with an Adversarial Diffusion Compression (ADC) framework to improve the inference speed of the model.
[0078] In some embodiments, the total loss function value may be the loss function value used to train the target image enhancement model during the training process.
[0079] In some embodiments, the loss function value can be adjusted by tuning hyperparameters, thereby adjusting the training of the target image enhancement model.
[0080] In some embodiments, the image noise parameters corresponding to the enhanced image can be adjusted by adjusting hyperparameters. Since the first noisy image data can be adjusted by adjusting hyperparameters, the noise parameters in the resulting enhanced image will also be different when different first noisy image data are input to the target image enhancement model.
[0081] In this embodiment of the disclosure, image data and first noisy image data are acquired, and the image data and first noisy image data are processed by a pre-trained target image enhancement model to obtain an enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on hyperparameters. The first noisy image data is determined by hyperparameters. By enhancing the image through the target enhancement model, the resolution of the image can be enhanced while suppressing the noise in the image.
[0082] Figure 2 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 2 .
[0083] Figure 2 Steps S205 and S206 correspond to steps S101 and S102, and will not be repeated here. Figure 2 As shown, in Figure 1 In addition to the implementation process shown, the following steps are also included:
[0084] S201, acquire a first training dataset including multiple first training samples and second noisy image data; each first training sample includes at least a corresponding first low-quality image and a first high-quality image; the second noisy image data is determined by hyperparameters.
[0085] In some embodiments, the second noise image data can be the same as the first noise image data, and the method for obtaining the second noise image data can also be the same as the method for obtaining the first noise image data, which will not be described in detail here.
[0086] In some embodiments, the first low-quality image and the first high-quality image may be two images with different resolutions but the same content.
[0087] In some embodiments, the first high-quality image can be one where the image noise parameter is less than a first preset threshold, and the corresponding image detail parameter is greater than a second preset threshold. In this embodiment, the first and second preset thresholds can be user-defined and are not limited here. For example, an image from a high-quality game can be acquired as the first high-quality image, and then a first low-quality image corresponding to the first high-quality image can be generated based on bicubic interpolation. Alternatively, an image from a low-quality game can be acquired as the first low-quality image, and then a first high-quality image corresponding to the first low-quality image can be generated based on an image restoration model using win transformers (winIR). It should be noted that the image restoration model can transform a low-quality image into a corresponding high-quality image through shallow feature extraction, deep feature extraction, and a high-quality reconstruction module.
[0088] S202, the first training dataset and the second noisy image data are respectively input into the pre-trained first image enhancement model and the target image enhancement model.
[0089] In some embodiments, the target image enhancement model can be a model to be trained. The pre-trained first image enhancement model can be a first image enhancement model that has already been trained. Inputting the first training dataset and the second noisy image data into the pre-trained first image enhancement model and the target image enhancement model respectively can include inputting the same first training dataset and the second noisy image data into the pre-trained first image enhancement model and the target image enhancement model, so that the pre-trained first image enhancement model and the target image enhancement model process the same data respectively to obtain the data processing results.
[0090] It should be noted that no noisy images were added during the training of the first image enhancement model. Therefore, the images processed by the first image enhancement model lack noise, which may cause distortions such as "skin smoothing" after processing natural images.
[0091] S203, determine the total loss function value using the first high-quality image, hyperparameters, the training results of the pre-trained first image enhancement model, and the training results of the target image enhancement model.
[0092] In some embodiments, the total loss function value can be determined by comparing the processing result of the first high-quality image with that of the target image enhancement model and the training result of the target image enhancement model. To ensure that the total loss function value exerts a good constraint on the target image enhancement model in terms of both image enhancement and image preservation, hyperparameters can be used to weight the comparison results obtained above. The method for determining the total loss function value is not limited in the embodiments of this disclosure.
[0093] S204. The target image enhancement model is trained based on the total loss function value. In response to the convergence of the total loss function value, the pre-trained target image enhancement model is obtained.
[0094] In some embodiments, training the target image augmentation model based on the total loss function value may include fixing the parameters of the pre-trained first image augmentation model during the training process and adjusting the parameters of the target image augmentation model only based on the total loss function value.
[0095] In some embodiments, convergence of the total loss function may include the difference between the total loss function values obtained from multiple consecutive training iterations falling within a preset range. This disclosure does not limit the preset range.
[0096] In this embodiment of the disclosure, the first training dataset and the second noisy image data are respectively input into the pre-trained first image enhancement model and the target image enhancement model. The total loss function value is determined by the pre-trained first image enhancement model, the target image enhancement model, hyperparameters, and the first high-quality image used for comparison. This allows the total loss function value to constrain the target image enhancement model in terms of image enhancement and image distortion during the training process of the target image enhancement model, making the trained target image enhancement model more able to meet user needs.
[0097] Figure 3 This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 3 .
[0098] Figure 3 Steps S301, S302, and S306 to S307 correspond to steps S201, S202, and S205 to S206, and will not be repeated here. Figure 3 As shown, in Figure 2 In addition to the implementation process shown, the following steps are also included:
[0099] S303, determine the difference between the training result of the pre-trained first image enhancement model and the training result of the target image enhancement model.
[0100] In some embodiments, the training result corresponding to the first image enhancement model can be a training result represented as a vector. Similarly, the training result corresponding to the target image enhancement model can also be a training result represented as a vector.
[0101] In some embodiments, the difference value may include the difference value determined based on the vector corresponding to the training result of the first image enhancement model and the training result of the target image enhancement model. It should be noted that when both the training result of the first image enhancement model and the training result of the target image enhancement model are images, the difference value can be a value used to represent the aforementioned image difference.
[0102] S304, determine the first loss function value corresponding to the target image enhancement model based on the training result of the target image enhancement model and the first high-quality image.
[0103] In some embodiments, the first loss function value may include a numerical value that reflects the difference between the training result of the target image enhancement model and the first high-quality image.
[0104] In some embodiments, the value of the first loss function can be determined based on the first loss function. An exemplary first loss function may include:
[0105]
[0106] Where L can be the value of the loss function. It can represent the training result corresponding to the target image enhancement model. It can represent the highest quality image.
[0107] S305 determines the total loss function value by using hyperparameters, difference values, and the first loss function value.
[0108] In some embodiments, the total loss function value is determined by hyperparameters, difference values, and the first loss function value, including:
[0109] The weights corresponding to the difference values and the first loss function values are determined by using hyperparameters.
[0110] The total loss function value is determined by multiplying the difference value by its corresponding weight and by multiplying the first loss function value by its corresponding weight.
[0111] It should be noted that the hyperparameters are positive real numbers not greater than 1. It should also be noted that the hyperparameters can be used to determine the first noisy image data and the second noisy image data.
[0112] For example, the distribution of Gaussian noise in the first and second noisy image data can be determined based on hyperparameters. The probability of the Gaussian noise distribution can be determined based on the following formula:
[0113]
[0114] in, However, the probability is a Gaussian noise distribution. For random probability, For hyperparameters, It can be the standard deviation, and the standard deviation can be 1.
[0115] For example, hyperparameters can be the weights corresponding to the difference values, and the difference between 1 and the hyperparameters can be the weights corresponding to the first loss function value. The product of the hyperparameter and its corresponding weight can be summed with the product of the first loss function value and its corresponding weight.
[0116] For example, the total loss function value can be determined based on the following formula:
[0117]
[0118] This is the total loss function value. It can be a hyperparameter. L1 is the difference value, and L1 is the value of the first loss function.
[0119] Figure 4This is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure. Figure 4 .
[0120] like Figure 4 As shown, it includes the following steps:
[0121] S401, Obtain a second training dataset including multiple second training samples, wherein the second training samples include at least the corresponding second low-quality image and the second high-quality image.
[0122] In some embodiments, the second high-quality image and the first high-quality image can be images of the same type. For example, an image from a low-quality game can be captured as the second low-quality image. A second high-quality image corresponding to the second low-quality image is generated based on the SwinIR method. Here, SwinIR is an image restoration method based on the Swin Transformer architecture, rather than a traditional image generation method. Its core objective is to reconstruct high-quality images from low-quality images (such as low-resolution, noisy, or compressed images), belonging to the field of image restoration in low-level vision tasks.
[0123] In some embodiments, the image noise parameters corresponding to the first high-quality image and the second high-quality image are less than a first preset threshold, and the corresponding image detail parameters are greater than a second preset threshold.
[0124] For example, both the first high-quality image and the second high-quality image can be images with high detail parameters and low noise parameters. Both the first preset threshold and the second preset threshold can be defined by the user and are not limited in this disclosure.
[0125] S402, input the second training dataset into the first image enhancement model to be trained, and obtain the training result corresponding to the first image enhancement model to be trained.
[0126] S403, determine the value of the second loss function based on the training result of the first image enhancement model to be trained and the second high-quality image.
[0127] S404, Train the first image augmentation model to be trained based on the second loss function value, and obtain the pre-trained first image augmentation model in response to the convergence of the second loss function value.
[0128] In some embodiments, the methods for training the first image enhancement model and the target image enhancement model, as well as the loss functions used, are the same and will not be described again here.
[0129] In some embodiments, the training method of the first image enhancement model in this disclosure can be applied to the above embodiments.
[0130] In this embodiment of the disclosure, a second training sample is generated, and the second training sample is input into the first image enhancement model to be trained to obtain the training result. Based on the training result, a second loss function value is determined, and the first image enhancement model to be trained is trained based on the second loss function value to obtain the trained first image enhancement model. The first image enhancement model is used as a comparison item for training the target image enhancement model, which can avoid excessive noise in the image enhanced by the trained target image enhancement model.
[0131] Figure 5 Steps S501 and S502 correspond to steps S101 and S102, and will not be repeated here. Figure 5 As shown, in Figure 1 In addition to the implementation process shown, the following steps are also included:
[0132] S503 adjusts the image noise parameters corresponding to the enhanced image by adjusting the hyperparameters.
[0133] In some embodiments, as can be seen from the method for determining the noisy image described in the above embodiments, the generated first noisy image data can be adjusted by adjusting the hyperparameters. Furthermore, since the total loss function corresponding to the target image enhancement model is generated based on the hyperparameters, adjusting the hyperparameters will inevitably also adjust the noise parameters in the enhanced image generated by the target image enhancement model.
[0134] To provide a detailed description of this disclosure, an exemplary embodiment is also provided.
[0135] Acquire game image data, which can include both high-quality and low-quality images. If the game image data is high-quality, bicubic interpolation can be used to generate a low-quality image corresponding to the high-quality image. If the game image data is low-quality, the SwinIR method can be used to generate a high-quality image corresponding to the low-quality image. Pair the low-quality and high-quality images together, creating at least 2000 such image pairs as a training set. Use the Lightweight Image Super-Resolution Network (RLFN) as the base network, inputting the training set into the base network to obtain the training results. Determine the first loss function value based on the training results, and train the base network based on the first loss function value to obtain the first image enhancement model. Reuse the RLFN network, adding one channel as a modulation term. Reuse the trained first image enhancement model, fixing its parameters. Create a Gaussian noise image with mean hyperparameters. The Gaussian noise image and the training set are input into the trained first image augmentation model and the trained target image augmentation model, respectively. The difference between the training results of the pre-trained first image augmentation model and the training results of the target image augmentation model is determined. Based on the training results of the target image augmentation model and the first high-quality image, the first loss function value of the target image augmentation model is determined. The total loss function value is determined using hyperparameters, the difference value, and the first loss function value. The target image augmentation model is then trained based on the total loss function value to obtain the trained target image augmentation model.
[0136] Figure 6 This is a block diagram of an image processing apparatus according to an exemplary embodiment of the present disclosure. (Refer to...) Figure 6 The image processing apparatus 600 includes:
[0137] The first acquisition module 610 is used to acquire image data and first noisy image data;
[0138] The processing module 620 is used to process the image data and the first noisy image data through a pre-trained target image enhancement model to obtain the enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on the hyperparameters. The first noisy image data is determined by the hyperparameters.
[0139] In one embodiment of this disclosure, the apparatus further includes:
[0140] The second acquisition module is used to acquire a first training dataset including multiple first training samples and second noisy image data; each first training sample includes at least a corresponding first low-quality image and a first high-quality image; the second noisy image data is determined by hyperparameters;
[0141] The input module is used to input the first training dataset and the second noisy image data into the pre-trained first image enhancement model and the target image enhancement model, respectively;
[0142] The first determining module is used to determine the total loss function value through the first high-quality image, hyperparameters, the training result of the pre-trained first image enhancement model, and the training result of the target image enhancement model;
[0143] The first training module is used to train the target image enhancement model based on the total loss function value. In response to the convergence of the total loss function value, a pre-trained target image enhancement model is obtained.
[0144] In one embodiment of this disclosure, the determining module includes:
[0145] The first determining unit is used to determine the difference between the training result corresponding to the pre-trained first image enhancement model and the training result corresponding to the target image enhancement model.
[0146] The second determining unit is used to determine the first loss function value corresponding to the target image enhancement model based on the training result of the target image enhancement model and the first high-quality image;
[0147] The third determining unit is used to determine the total loss function value through hyperparameters, difference values, and the first loss function value.
[0148] In one embodiment of this disclosure, the third determining unit includes:
[0149] The first determining subunit is used to determine the weights corresponding to the difference value and the first loss function value through hyperparameters;
[0150] The second determining subunit is used to determine the total loss function value by multiplying the difference value by the corresponding weight and the first loss function value by the corresponding weight.
[0151] In one embodiment of this disclosure, the apparatus further includes:
[0152] The third acquisition module is used to input the second training dataset into the first image enhancement model to be trained, and obtain the training result corresponding to the first image enhancement model to be trained.
[0153] The second determining module is used to determine the value of the second loss function based on the training result of the first image enhancement model to be trained and the second high-quality image.
[0154] The second training module is used to train the first image augmentation model to be trained based on the second loss function value. In response to the convergence of the second loss function value, the pre-trained first image augmentation model is obtained.
[0155] In one embodiment of this disclosure, the image noise parameters corresponding to the first high-quality image and the second high-quality image are less than a first preset threshold, and the corresponding image detail parameters are greater than a second preset threshold.
[0156] In one embodiment of this disclosure, the apparatus further includes:
[0157] The adjustment module is used to adjust the image noise parameters corresponding to the enhanced image by adjusting hyperparameters.
[0158] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. For example, device 700 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.
[0159] Reference Figure 7 The device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0160] Processing component 702 typically controls the overall operation of device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.
[0161] Memory 704 is configured to store various types of data to support the operation of device 700. Examples of this data include instructions for any application or method operating on device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0162] Power supply assembly 706 provides power to various components of device 700. Power supply assembly 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 700.
[0163] Multimedia component 708 includes a screen that provides an output interface between device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0164] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.
[0165] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0166] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of device 700. For example, sensor assembly 714 may detect the on / off state of device 700, the relative positioning of components such as the display and keypad of device 700, changes in the position of device 700 or a component of device 700, the presence or absence of user contact with device 700, the orientation or acceleration / deceleration of device 700, and temperature changes of device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0167] Communication component 716 is configured to facilitate wired or wireless communication between device 700 and other devices. Device 700 can access wireless networks based on communication standards such as Wi-Fi, 3G, 4G, 5G, other communication standards, or combinations thereof. In some embodiments of this disclosure, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of this disclosure, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0168] In some embodiments of this disclosure, the apparatus 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0169] In some embodiments of this disclosure, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions that can be executed by a processor 720 of the device 700 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0170] In some embodiments of this disclosure, a non-transitory computer-readable storage medium enables a terminal to perform an image processing method when instructions in the storage medium are executed by a terminal's processor.
[0171] In some embodiments of this disclosure, a computer program product is also provided, including a computer program / instructions that, when executed by a processor, implement an image processing method.
[0172] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.
[0173] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.
[0174] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0175] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, include: Acquire image data and first noisy image data; The image data and the first noisy image data are processed by a pre-trained target image enhancement model to obtain an enhanced image corresponding to the image data. The target image enhancement model is trained based on the total loss function value, which is determined at least based on hyperparameters. The first noisy image data is determined by the hyperparameters. Training methods for target image enhancement models include: A first training dataset comprising multiple first training samples and second noisy image data are obtained; each first training sample includes at least a corresponding first low-quality image and a first high-quality image; the second noisy image data is determined by the hyperparameters, and the first low-quality image is generated from the first high-quality image based on bicubic interpolation; The first training dataset and the second noisy image data are respectively input into the pre-trained first image enhancement model and the target image enhancement model; The total loss function value is determined by the first high-quality image, the hyperparameters, the training results of the pre-trained first image enhancement model, and the training results of the target image enhancement model. The target image enhancement model is trained based on the total loss function value, and the pre-trained target image enhancement model is obtained in response to the convergence of the total loss function value. The step of determining the total loss function value using the first high-quality image, the hyperparameters, the training results of the pre-trained first image augmentation model, and the training results of the target image augmentation model includes: Determine the difference between the training result of the pre-trained first image enhancement model and the training result of the target image enhancement model; The first loss function value corresponding to the target image enhancement model is determined based on the training result of the target image enhancement model and the first high-quality image. The total loss function value is determined by the hyperparameters, the difference value, and the first loss function value.
2. The method according to claim 1, characterized in that, Determining the total loss function value using the hyperparameters, the difference value, and the first loss function value includes: The weights corresponding to the difference value and the first loss function value are determined by the hyperparameters. The total loss function value is determined by multiplying the difference value by its corresponding weight and by multiplying the first loss function value by its corresponding weight.
3. The method according to claim 1, characterized in that, The training process of the first image enhancement model includes: Obtain a second training dataset comprising multiple second training samples, wherein the second training samples include at least corresponding second low-quality images and second high-quality images; Input the second training dataset into the first image enhancement model to be trained to obtain the training result corresponding to the first image enhancement model to be trained. The second loss function value is determined based on the training result of the first image enhancement model to be trained and the second high-quality image. The first image enhancement model to be trained is trained based on the second loss function value, and the pre-trained first image enhancement model is obtained in response to the convergence of the second loss function value.
4. The method according to claim 3, characterized in that, The image noise parameters corresponding to the first high-quality image and the second high-quality image are less than the first preset threshold, and the corresponding image detail parameters are greater than the second preset threshold.
5. The method according to claim 1, characterized in that, The method further includes: The image noise parameters corresponding to the enhanced image are adjusted by adjusting the hyperparameters.
6. An image processing apparatus, characterized in that, include: The first acquisition module is used to acquire image data and first noisy image data; The processing module is used to process the image data and the first noisy image data using a pre-trained target image enhancement model to obtain an enhanced image corresponding to the image data. The target image enhancement model is trained based on a total loss function value, which is determined at least based on hyperparameters. The first noisy image data is determined by the hyperparameters. The training method of the target image enhancement model includes: A first training dataset comprising multiple first training samples and second noisy image data are obtained; each first training sample includes at least a corresponding first low-quality image and a first high-quality image; the second noisy image data is determined by the hyperparameters, and the first low-quality image is generated from the first high-quality image based on bicubic interpolation; The input module is used to input the first training dataset and the second noisy image data into the pre-trained first image enhancement model and the target image enhancement model, respectively. The first determining module is used to determine the total loss function value based on the first high-quality image, the hyperparameters, the training result of the pre-trained first image enhancement model, and the training result of the target image enhancement model. The first training module is used to train the target image enhancement model based on the total loss function value, and to obtain the pre-trained target image enhancement model in response to the convergence of the total loss function value. The determining module includes: The first determining unit is used to determine the difference value between the training result corresponding to the pre-trained first image enhancement model and the training result corresponding to the target image enhancement model; The second determining unit is used to determine the first loss function value corresponding to the target image enhancement model based on the training result corresponding to the target image enhancement model and the first high-quality image; The third determining unit is used to determine the total loss function value through the hyperparameters, the difference value, and the first loss function value.
7. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the image processing method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of a terminal, enable the terminal to perform the steps of an image processing method according to any one of claims 1 to 5.
9. A chip comprising a processing circuit for performing the steps of the image processing method as claimed in any one of claims 1-5.
Citation Information
Patent Citations
Loss function construction method for image enhancement model, storage medium and device
CN117274127A
Joint denoising training method combining noise-free image and noise image
CN119831886A