A model training method, an image processing method, an apparatus, a medium, and an equipment
By using a capability transfer training method for the model to be transferred, sample images and label images are generated, solving the problem of high training difficulty in image enhancement processing and achieving high-precision image enhancement processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2023-09-11
- Publication Date
- 2026-04-24
AI Technical Summary
Existing neural network models face significant challenges and high costs in image enhancement processing due to the wide variety of image types and the poor regularity of their content distribution.
The target image is processed by a pre-trained transfer processing model to obtain a first label image. The target image is then perturbed and mirrored to generate a sample image and a second label image, which are used to train the image processing model and transfer the capabilities of the transfer processing model to the image processing model.
It simplifies the training process of image processing models, reduces the difficulty of collecting sample images and model fitting, and improves the processing performance and generalization ability of image processing models.
Smart Images

Figure CN117217294B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to deep learning technology, and more particularly to a model training method, image processing method, apparatus, medium, and device. Background Technology
[0002] Image enhancement processing is used to improve image quality and reduce noise in images. It is widely used in various application fields.
[0003] With the continuous development of deep learning technology, image enhancement processing through neural network models has become a research hotspot in the field of computer vision. Currently, when using neural network models for image enhancement, the training process of these models is difficult and costly due to the wide variety of image types and the poor regularity of image content distribution. Summary of the Invention
[0004] This disclosure provides a model training method, image processing method, apparatus, medium, and device to simplify the training process of image processing models and reduce the training difficulty of image processing models.
[0005] In a first aspect, embodiments of this disclosure provide a method for training an image processing model, comprising:
[0006] Acquire the target image, process the target image based on the pre-trained transfer processing model, and obtain the first label image;
[0007] The target image is perturbed to obtain a perturbed image;
[0008] The perturbation image and the first label image are mirrored to obtain a sample image corresponding to the perturbation image and a second label image corresponding to the first label image;
[0009] The image processing model to be trained is trained based on the sample image and the second label image to obtain the trained image processing model.
[0010] Secondly, embodiments of this disclosure also provide an image processing method, including:
[0011] Obtain the image to be processed;
[0012] The image to be processed is processed based on the image processing model to obtain the processed image, wherein the image processing model is trained based on the image processing model training method provided in the embodiments of this disclosure.
[0013] Thirdly, embodiments of this disclosure also provide a training apparatus for an image processing model, comprising:
[0014] The first label image acquisition module is used to acquire a target image and process the target image based on a pre-trained transfer processing model to obtain a first label image.
[0015] An image perturbation processing module is used to perturb the target image to obtain a perturbed image;
[0016] The image morphology processing module is used to perform mirror morphology processing on the perturbation image and the first label image to obtain a sample image corresponding to the perturbation image and a second label image corresponding to the first label image.
[0017] The image training module is used to train the image processing model to be trained based on the sample image and the second label image, so as to obtain the trained image processing model.
[0018] Fourthly, embodiments of this disclosure also provide an image processing apparatus, including:
[0019] The image acquisition module is used to acquire the image to be processed.
[0020] An image processing module is used to process the image to be processed based on an image processing model to obtain a processed image, wherein the image processing model is trained based on the image processing model training method provided in the embodiments of this disclosure.
[0021] In this embodiment, a target image is processed using a pre-trained transfer processing model to obtain a first label image; perturbation and mirror morphology processing are applied to obtain a sample image and a second label image, and the image processing model is then trained. The image processing capabilities of the transfer processing model are transferred to the image processing model, resulting in a trained image processing model capable of performing high-precision enhancement processing on non-target images. This transfer-based training method reduces the difficulty of sample image acquisition and model fitting, simplifying the training process of the image processing model. Attached Figure Description
[0022] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0023] Figure 1 This is a schematic flowchart illustrating a training method for an image processing model provided in an embodiment of the present disclosure.
[0024] Figure 2 This is a flowchart of a training method for an image processing model provided in an embodiment of this disclosure;
[0025] Figure 3 This is a schematic diagram illustrating the process of determining distillation characteristic loss terms provided in the disclosed embodiments;
[0026] Figure 4 This is a schematic flowchart of an image processing method provided in a disclosed embodiment;
[0027] Figure 5 A schematic diagram of the structure of a training device for an image processing model provided in an embodiment of this disclosure;
[0028] Figure 6 This is a schematic diagram of the structure of an image processing apparatus provided in the disclosed embodiments;
[0029] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0032] The term "comprising" and its variations as used herein are open-ended inclusion, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0036] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0037] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0038] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0039] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0040] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0041] The transfer processing model is an image processing algorithm for target image enhancement. It can enhance low-quality target images that are blurry, noisy, or contain artifacts, improving the clarity of facial regions in the image. In some embodiments, the transfer processing model can be a model for enhancing images of a specific scene, where the facial image is a scene-specific image. For example, the transfer processing model can be a facial processing model that performs high-precision image enhancement on facial images, where the target image can be a facial image. Enhancement processing of a single scene-specific image has advantages over image enhancement processing of all types of scene images, mainly because the enhancement processing of a single scene-specific image has certain standardization. Taking a facial image as an example, this standardization is reflected in the fact that the facial image can be aligned using alignment techniques, allowing the facial image to undergo radial transformation so that the facial key points fall into fixed positions. The aligned facial image is then cropped to obtain a standard facial image. This facial image is processed by the facial processing model to obtain a high-definition facial image, which is then pasted back into the original image through inverse affine transformation to obtain the processed facial image.
[0042] Based on the above facial image enhancement process, it is known that because the image content distribution of a single specific scene image is standardized and the data distribution is uniform, the distribution that the transfer processing model needs to fit is relatively simple, resulting in good image enhancement effects. However, natural images include multiple types of scene images, and the image content distribution of natural images is much wider, making it much more difficult for the image processing model to fit the data. To address the above technical problems, this disclosure provides an embodiment that transfers the processing capabilities of the transfer processing model to an image processing model that enhances natural images, thereby simplifying the training process of the image processing model while improving its processing performance.
[0043] Figure 1 This is a flowchart illustrating a training method for an image processing model provided in an embodiment of this disclosure. This embodiment is applicable to situations where an image processing model is trained by transferring processing capabilities from a processing model to be transferred. This method can be executed by an image processing model training device, which can be implemented in software and / or hardware, optionally through an electronic device such as a mobile terminal, PC, or server. Figure 1 As shown, the method includes:
[0044] S110. Obtain the target image and process the target image based on the pre-trained transfer processing model to obtain the first label image.
[0045] S120. The target image is perturbed to obtain a perturbed image.
[0046] S130. Perform mirror image processing on the perturbation image and the first label image to obtain a sample image corresponding to the perturbation image and a second label image corresponding to the first label image.
[0047] S140. The image processing model to be trained is trained based on the sample image and the second label image to obtain the trained image processing model.
[0048] The target image can be an image of a specific scene, and correspondingly, the transfer processing model is a model used to enhance images of that specific scene. The target image can be obtained by acquiring the original image of the specific scene, preprocessing the original image to obtain the target image, where preprocessing may include, but is not limited to, image cropping and image alignment. Optionally, the target image can be an image obtained under different shooting conditions within a specific scene to improve the diversity of the target image.
[0049] Taking a facial image as an example, a portrait image is acquired, facial regions within the portrait image are identified, and cropping and alignment processing is performed on the portrait image based on the recognition results to obtain the target image. The portrait image can be imported from an external source, stored locally, or acquired in real-time via an image acquisition device connected to it. The portrait image may include one or more facial regions; cropping and alignment processing is performed on each facial region to obtain one or more facial images. Optionally, the facial images may include facial regions located in different lighting environments, increasing the environmental complexity of the facial images. Optionally, the facial images may include facial images of people of different genders and ages, increasing the diversity of the facial images. Optionally, the facial images may include facial images acquired from different intersection angles, further increasing the diversity of the facial images.
[0050] A pre-trained transfer processing model is invoked, which has the function of enhancing the target image. The target image is enhanced by the transfer processing model to obtain a first label image, which is an enhanced image of the target image. The image quality of the first label image is higher than that of the target image.
[0051] The face processing is performed using the transfer processing model, and the resulting enhanced image is used as the first label image of the target image, simplifying the label image determination process. The target image and the first label image are stored in correspondence to form a sample set for training the image processing model, thus transferring the processing capabilities of the transfer processing model to the image processing model.
[0052] Since natural images (including images of any type of scene) do not possess the normalized distribution of target images, to improve the processing capability and generalization of image processing models for natural images, the aforementioned target images are processed to increase image randomness. This processing of the target images includes perturbation and morphological processing, with the processed target images serving as perturbed images. Perturbation reduces the image quality of the target images, while morphological processing increases the morphological diversity of the target images, reducing the uniformity of content across multiple target images.
[0053] Optionally, the target image is perturbed, including: blurring the contour information of the target image. The contour information of the target image is extracted, for example, through a contour extraction model, specifically by inputting the target image into the contour extraction model to obtain the contour information; alternatively, edge recognition can be performed on the target image, and the identified edges can be used as the contour information; or high-frequency information in the target image can be identified and used as the contour information. Blurring the contour information of the target image reduces the contour sharpness; for example, the blurring can be implemented using a Gaussian blur algorithm, which is not limited here.
[0054] The contour information of the target image is identified, resulting in a contour image and a non-contour image (i.e., a high-frequency information image and a low-frequency information image). The contour image of the target image is blurred to obtain a blurred contour image. The blurred contour image and the non-contour image are then fused to obtain the processed image, which can be used as a perturbation image. The image processing model is trained using the perturbation image to enable the model to learn contour sharpening capabilities.
[0055] Based on the above embodiments, the perturbation processing of the target image further includes one or more of the following: adding white noise to the target image; performing random blurring processing on the target image; and compressing the target image. Specifically, random noise addition processing can be performed on the target image using a noise addition algorithm, where different levels of noise addition processing can be applied to different target images to add different levels of white noise to different target images; or, white noise can be added to at least a local region of the target image, where the size and number of local regions can be randomly determined to improve the randomness of white noise addition.
[0056] Random blurring of a target image can be applied to the entire target image to reduce its overall image quality.
[0057] During the compression process of the target image, multiple different compression algorithms can be preset, and one compression algorithm can be randomly selected from the multiple compression algorithms to compress the target image and reduce the image quality.
[0058] In some embodiments, the target image can be processed using one or more perturbation methods to obtain a perturbed image. For example, the target image can be sequentially blurred by contour information and white noise can be added to obtain a perturbed image; alternatively, the target image can be sequentially blurred by contour information and compressed to obtain a perturbed image; alternatively, the target image can be sequentially added to white noise and compressed to obtain a perturbed image; or alternatively, the target image can be sequentially blurred by contour information, added to white noise, and compressed to obtain a perturbed image. By employing at least one perturbation method to process the target image, the randomness and complexity of the perturbed image are improved.
[0059] Based on the correspondence between the target image and the first label image, a correspondence is established between the perturbated image (obtained through perturbation processing of the target image) and the first label image. The corresponding target image and the first label image are then mirrored to obtain a set of sample images and a second label image, forming a sample dataset. Specifically, the mirroring of the target image and the first label image involves performing morphological processing on each image separately, with the morphological processing method for the target image being the same as that for the first label image. This mirroring process ensures that the sample images and the second label image maintain morphological consistency.
[0060] The mirror morphology processing includes one or more of mirror deformation and mirror cropping. Mirror cropping involves cropping the same region from both the target image and the first label image; this means the target image and the first label image have the same size. Mirror deformation can be performed on both the target image and the first label image in the same way, including but not limited to distortion, rotation, and flipping.
[0061] In some embodiments, by employing different combinations of perturbation and mirror morphology processing, a target image and a first label image are processed to obtain multiple groups of sample images and second label images, thereby increasing the amount of sample data and reducing the number of target images required. Simultaneously, in this embodiment, by using a capability transfer training method for the transfer processing model, it is unnecessary to collect non-target images with different image content, reducing the difficulty of sample image collection.
[0062] By perturbating the target image to reduce its quality, the image processing model learns image enhancement capabilities during training. Morphological processing breaks down the strong prior information of the target image's structure and increases the morphological randomness of the sample images, thereby enabling the trained image processing model to have robustness and generalization in image processing.
[0063] Multiple sets of sample images and second-label images are obtained through the above method to form a sample dataset. The image processing model to be trained is iteratively trained using the sample dataset to transfer the image processing capabilities of the model to be transferred to the image processing model, thus obtaining a trained image processing model.
[0064] For example, see Figure 2 , Figure 2 This is a flowchart illustrating a training method for an image processing model provided in an embodiment of this disclosure. Wherein, Figure 2 Taking a face image as the target image and the face processing model as the example, the processing capabilities of the face processing model are transferred to train an image processing model that can enhance images in any type of scene.
[0065] The technical solution provided in this embodiment processes the target image using a pre-trained transfer processing model to obtain a first label image; perturbation and mirror morphology processing are then applied to obtain a sample image and a second label image, which are then used to train the image processing model. The image processing capabilities of the transfer processing model are transferred to the image processing model, resulting in a trained image processing model capable of performing high-precision enhancement processing on non-target images. This transfer-based training method reduces the difficulty of sample image acquisition and model fitting, simplifying the training process of the image processing model.
[0066] In some embodiments, the training process of the image processing model may include: iteratively executing the following training process to obtain a trained image processing model when the training termination condition is met: inputting the sample image into the image processing model to be trained to obtain a predicted image; determining a loss function based on the predicted image and the second label image; and backpropagating the loss function to the image processing model to adjust the model parameters of the image processing model. The loss function includes at least one loss term, which includes one or more of the following: an image feature loss term, an image quality loss term, a first high-frequency feature loss term, a second high-frequency feature loss term, an image similarity loss term, and a difference loss function.
[0067] In some embodiments, the training process of the image processing model includes two training phases. In the first training phase, the image processing model is iteratively trained. Upon meeting a first termination condition, an intermediate image processing model is obtained, and the process proceeds to the second training phase. In the second training phase, the intermediate image processing model is iteratively trained. Upon meeting a second termination condition, a trained image processing model is obtained. The first and second termination conditions can be different. For example, the first termination condition could be that the number of training iterations reaches a first threshold, the training accuracy reaches a first accuracy threshold, or convergence is achieved; the second termination condition could be that the training accuracy reaches a second accuracy threshold or convergence is achieved.
[0068] The training process of the image processing model may include: in a first training phase, inputting the sample image into the image processing model to be trained to obtain a first predicted image; determining a first loss function based on the first predicted image and a second label image; and adjusting the model parameters of the image processing model based on the first loss function and a first learning rate to obtain an intermediate image processing model; in a second training phase, inputting the sample image into the intermediate image processing model to obtain a second predicted image; determining a second loss function based on the second predicted image and the second label image; and adjusting the model parameters of the image processing model based on the second loss function and a second learning rate to obtain a trained image processing model, wherein the second learning rate is less than the first learning rate. In this embodiment, training the image processing model with different learning rates in the two training phases is beneficial to improving the processing performance of the image processing model.
[0069] In some embodiments, the first loss function used in the first training phase and the second loss function used in the second training phase are different. The first loss function includes a first number of loss terms; the second loss function includes a second number of loss terms; and the loss terms included in the first and second loss functions may partially overlap.
[0070] Optionally, the loss term includes one or more of the following: image feature loss term, image quality loss term, first high-frequency feature loss term, second high-frequency feature loss term, image similarity loss term, and difference loss function.
[0071] The image feature loss term represents the loss between the image features of the predicted image and the image features of the second label image. It can be understood that in the first training phase, the predicted image is the first predicted image, and in the second training phase, the predicted image is the second predicted image. Optionally, the image feature loss term is determined by: extracting label features of the second label image based on the feature extraction module; extracting prediction features of the first or second predicted image based on the feature extraction module; and determining the image feature loss term based on the label features and the prediction features. The feature extraction model can be a neural network model with feature extraction capabilities, such as the VGG model. Predicted features are obtained by inputting the first or second predicted image into the feature extraction model. Label features are obtained by inputting the second label image into the feature extraction model. The loss between the label features and the prediction features is determined as the image feature loss term, where the loss between the label features and the prediction features includes, but is not limited to, L1 loss, L2 loss, and cross-entropy loss. For example, the image feature loss term can be represented by the following formula: Losslpips=L1loss(vgg(f(x)),vgg(y)), where y is the second label image, vgg(y) is the label feature, f(x) is the first or second predicted image, vgg(f(x)) is the predicted feature, and the image feature loss term is the L1 loss of the label feature and the predicted feature.
[0072] The image quality loss term represents the loss of image quality data in the predicted image. Optionally, the image quality loss term is determined by: determining the first image quality data of the first predicted image or the second predicted image, and generating the image quality loss term based on the first image quality data. In any training phase, the predicted image output by the image processing model is evaluated using an image evaluation model to obtain the first image quality data. The larger the first image quality data, the higher the image quality, indicating a smaller loss. Correspondingly, the image quality loss term is negatively correlated with the first image quality data. For example, the image quality loss term can be expressed as: lossiiqa = -Miqa(f(x)), where Miqa(f(x)) represents the first image quality data obtained through the image evaluation model.
[0073] Optionally, the image quality loss term is determined by: determining first image quality data of the first predicted image or the second predicted image, and second image quality data of the second labeled image; and generating the image quality loss term based on the first image quality data and the second image quality data. The predicted image and the second labeled image are evaluated separately using an image evaluation model, and the image quality loss term is formed based on the difference between the first image quality data and the second image quality data. For example, the image quality loss term can be obtained through any one of the L1 loss, L2 loss, cross-entropy loss, etc., between the first image quality data and the second image quality data.
[0074] In some embodiments, since the target image is cropped from the original portrait image, before image quality evaluation, the predicted image and the second label image are restored to the portrait image. Specifically, the cropped facial region in the portrait image is replaced by the predicted image (either the first or second predicted image) to obtain the first portrait image, and the cropped facial region is replaced by the second label image to obtain the second portrait image. A quality evaluation is performed on the first portrait image to obtain first image quality data, and an image quality loss term is determined based on this first image quality data. Alternatively, a quality evaluation is performed on both the first and second portrait images to obtain first image quality data and second image quality data, and an image quality loss term is determined based on both the first and second image quality data.
[0075] The first high-frequency feature loss term characterizes the loss between the high-frequency features of the predicted image and the high-frequency features of the second label image. Optionally, the method for determining the first high-frequency feature loss term includes: extracting first high-frequency data from the first predicted image or the second predicted image, and extracting second high-frequency data from the second label image; determining the first high-frequency feature loss term based on the first high-frequency data and the second high-frequency data. In this embodiment, the method for determining the high-frequency data in the image is not limited; for example, it can be determined by an edge recognition algorithm or by a high-frequency feature extraction model. The first high-frequency feature loss term is determined by any one of the L1 loss, L2 loss, cross-entropy loss, etc., between the first high-frequency data and the second high-frequency data.
[0076] For example, edge recognition is performed on the predicted image to obtain an edge recognition result, which includes gray values representing edge information. A high-frequency feature threshold is compared with the gray values of pixels. A high-frequency mask for the predicted image is formed based on pixels with gray values greater than the high-frequency feature threshold. First high-frequency data is determined based on the high-frequency mask and the predicted image. Similarly, second high-frequency data for the second label image can be determined using the above method. A first high-frequency feature loss term is determined based on the L1 loss between the first and second high-frequency data.
[0077] The first high-frequency feature loss term can be represented as follows: lossedge = L1loss(Edge) x *f(x),Edge y *y), where Edge x =sobel(f(x))>threshold,Edge y =sobel(y)>threshold; Edge x To predict the high-frequency mask of an image, Edge y is the high-frequency mask of the second label image, Threshold is the high-frequency feature threshold, sobel(f(x)) is the edge recognition result of the predicted image, and sobel(y) is the edge recognition result of the second label image.
[0078] The second high-frequency feature loss term characterizes the loss between the high-frequency features of the predicted image and the high-frequency features of the second label image from another dimension. The method for determining the second high-frequency feature loss term includes: extracting first high-frequency data from the first or second predicted image, and extracting second high-frequency data from the second label image; performing discrimination processing on the first and second high-frequency data respectively based on a preset discriminator; and determining the second high-frequency feature loss term based on the discrimination results of the first and second high-frequency data. The method for determining the first and second high-frequency data will not be elaborated here.
[0079] A discriminator is pre-trained to judge high-frequency data and output the discrimination results, namely the discrimination results of the first high-frequency data and the discrimination results of the second high-frequency data. The second high-frequency feature loss term is determined by the discrimination results of the first high-frequency data and the second high-frequency data. For example, the second high-frequency feature loss term can be expressed as: LEdgeAdv=E[logD(sobel(y))]+E[log(1-D(sobel(f(x))))], where D(sobel(y)) is the discrimination result of the second high-frequency data, D(sobel(f(x)))) is the discrimination result of the first high-frequency data, and E represents the expectation function.
[0080] The image similarity loss term characterizes the similarity loss between the predicted image and the second labeled image. Optionally, the pixel mean and variance of the first or second predicted image, the pixel mean and variance of the second labeled image, and the covariance of the predicted image and the second labeled image are determined. The image similarity loss term is determined using the pixel mean and variance of the first or second predicted image, the pixel mean and variance of the second labeled image, and the covariance. For example, the image similarity loss term can be expressed by the following formula:
[0081]
[0082] Among them, u x To predict the pixel mean of an image, u y σ is the pixel mean of the second labeled image. xy To predict the covariance of the image and the second-label image, To predict the variance of the image, Let c be the variance of the second labeled image, and c1 and c2 be hyperparameters.
[0083] The difference loss function characterizes the loss between the predicted image and the second labeled image. For example, it can be the L2 loss between the predicted image and the second labeled image. For example, it can be expressed as lossl2 = ||f(x) - y||2.
[0084] Based on the above embodiments, in the iterative training of the first training stage, a training method of knowledge distillation can be interspersed. That is, when a sample image is determined to have a loss term in the above manner, the image processing model is also subjected to knowledge distillation through the transfer processing model to determine the distilled feature loss term. The loss function is obtained based on the combination of the distilled feature loss term and one or more of the above.
[0085] The method for determining the distillation feature loss term includes: inputting the sample image into the transfer processing model to obtain the processed result image and the first feature data in the prediction processing layer of the transfer processing model during processing; obtaining the second feature data in the prediction processing layer during the processing of the sample image by the image processing model; and determining the distillation feature loss function based on either the first or second predicted image, the processed result image, the first feature data, and the second feature data. For example, see [link to example]. Figure 3 , Figure 3 This is a schematic diagram illustrating the process of determining the characteristic loss term of distillation.
[0086] It is understood that the image processing model has the same network structure as the model to be transferred, including multiple downsampled network blocks and multiple upsampled network blocks. Each network block may include multiple network layer groups, and each network layer group may include normalization layers, activation function layers and convolutional layers. The network block may also include jump-connected convolutional layers, which may be connected to the input layer and the last network layer group of the network block.
[0087] Sample images are input into an image processing model and a transfer processing model, respectively. The image processing model outputs a predicted image, and the transfer processing model outputs a processed image. During the processing of the sample images by the image processing model and the transfer processing model, feature data from the corresponding network layers of the image processing model and the transfer processing model are extracted, namely, the first feature data in the prediction processing layer of the transfer processing model and the second feature data in the prediction processing layer of the image processing model. The preset processing layer can be at least a local upsampled network block, such as... Figure 3 As shown.
[0088] The distillation feature loss term is determined by the loss between the second feature data and the first feature data, and the loss between the predicted data and the processed image. For example, the distillation feature loss term can be characterized as follows: lossdistill=L1(f(x),y')+∑L1(f i (f(x)),f face (x)), where y' is the processed image, f i (f(x) is the second feature data, f) face (x) is the first feature data.
[0089] In one optional embodiment, during the first training phase of the image processing model, the first loss function includes an image feature loss term, an image quality loss term, a first high-frequency feature loss term, a second high-frequency feature loss term, a difference loss function, and a distillation feature loss term. The sum or weighted sum of these loss terms is used to determine the first loss function. During the first training phase, the second loss function includes an image feature loss term, an image quality loss term, a first high-frequency feature loss term, a second high-frequency feature loss term, and an image similarity loss term. The sum or weighted sum of these loss terms is used to determine the second loss function.
[0090] The technical solution of this disclosure improves the training accuracy of the image processing model by setting two training phases and sampling different loss functions and learning rates in the two training phases. In each training phase, a loss function formed by multiple loss terms is used to train the image processing model from multiple dimensions, thereby improving the image processing performance of the model.
[0091] Figure 4 This is a schematic flowchart of an image processing method provided in a disclosed embodiment. This disclosure is applicable to situations where image enhancement processing is performed on images of arbitrary content. The method can be executed by an image processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 4 As shown, the method includes:
[0092] S210. Obtain the image to be processed.
[0093] S220. The image to be processed is processed based on the image processing model to obtain the processed image.
[0094] In this embodiment, the image to be processed can be an image from any type of scene. Optionally, the image to be processed can be a facial image or a non-facial image. Correspondingly, the image processing model is a model that can enhance images from any scene, and can enhance the image to be processed to obtain the processed image, i.e., the enhanced image of the image to be processed.
[0095] The image processing model in this embodiment is trained using the training method of any of the image processing models in the above embodiments. It can process any image, whether it is a target image or a non-target image, to improve image quality. Simultaneously, during the enhancement process of the image to be processed, noise reduction and other processing can be performed to remove interfering data, thereby reducing the amount of image data and achieving image data compression.
[0096] Figure 5 This is a schematic diagram of the structure of a training device for an image processing model provided in an embodiment of this disclosure, as shown below. Figure 5 As shown, the device includes: a first label image acquisition module 310, an image perturbation processing module 320, an image morphology processing module 330, and an image training module 340.
[0097] The first label image acquisition module 310 is used to acquire a target image and process the target image based on a pre-trained transfer processing model to obtain a first label image.
[0098] The image perturbation processing module 320 is used to perturb the target image to obtain a perturbed image.
[0099] The image morphology processing module 330 is used to perform mirror morphology processing on the perturbed image and the first label image to obtain a sample image corresponding to the perturbed image and a second label image corresponding to the first label image.
[0100] The image training module 340 is used to train the image processing model to be trained based on the sample image and the second label image, so as to obtain the trained image processing model.
[0101] The technical solution provided in this disclosure processes a target image using a pre-trained transfer processing model to obtain a first label image; perturbation and mirror morphology processing are then applied to obtain a sample image and a second label image, which are then used to train the image processing model. The image processing capabilities of the transfer processing model are transferred to the image processing model, resulting in a trained image processing model capable of performing high-precision enhancement processing on non-target images. This transfer-based training method reduces the difficulty of sample image acquisition and model fitting, simplifying the training process of the image processing model.
[0102] Based on the above embodiments, optionally, the image perturbation processing 320 is used to: blur the contour information of the target image.
[0103] Based on the above embodiments, optionally, the image perturbation processing 320 is also used to perform one or more of the following:
[0104] Add white noise to the target image;
[0105] The target image is subjected to random blurring.
[0106] The target image is compressed.
[0107] Based on the above embodiments, optionally, the mirror morphology processing includes one or more of mirror deformation and mirror clipping.
[0108] Based on the above embodiments, optionally, the image training module 340 is used for:
[0109] In the first training phase, the sample image is input into the image processing model to be trained to obtain a first predicted image. A first loss function is determined based on the first predicted image and the second label image. The model parameters of the image processing model are adjusted based on the first loss function and the first learning rate to obtain an intermediate image processing model.
[0110] In the second training phase, the sample image is input into the intermediate image processing model to obtain a second predicted image. A second loss function is determined based on the second predicted image and the second label image. The model parameters of the image processing model are adjusted based on the second loss function and the second learning rate to obtain a trained image processing model, wherein the second learning rate is less than the first learning rate.
[0111] Optionally, the first loss function includes a first number of loss terms; the second loss function includes a second number of loss terms.
[0112] The loss term includes one or more of the following: image feature loss term, image quality loss term, first high-frequency feature loss term, second high-frequency feature loss term, image similarity loss term, and difference loss function.
[0113] Optionally, the image training module 340 is further configured to: extract label features of the second labeled image based on the feature extraction module, extract prediction features of the first predicted image or the second predicted image based on the feature extraction module, and determine the image feature loss term based on the label features and the prediction features.
[0114] Optionally, the image training module 340 is further configured to: determine first image quality data of the first predicted image or the second predicted image, and generate the image quality loss term based on the first image quality data; or, determine first image quality data of the first predicted image or the second predicted image, and second image quality data of the second labeled image, and generate the image quality loss term based on the first image quality data and the second image quality data.
[0115] Optionally, the image training module 340 is further configured to: extract first high-frequency data from the first predicted image or the second predicted image, extract second high-frequency data from the second labeled image, and determine the first high-frequency feature loss term based on the first high-frequency data and the second high-frequency data.
[0116] Optionally, the image training module 340 is further configured to: extract first high-frequency data from the first predicted image or the second predicted image, and extract second high-frequency data from the second labeled image; perform discrimination processing on the first high-frequency data and the second high-frequency data based on a preset discriminator, and determine the second high-frequency feature loss term based on the discrimination results of the first high-frequency data and the second high-frequency data.
[0117] Optionally, the loss term may further include a distillation characteristic loss term;
[0118] Image training module 340 is also used for:
[0119] The sample image is input into the transfer processing model to obtain the processed result image and the first feature data in the prediction processing layer of the transfer processing model during the processing; the second feature data in the prediction processing layer during the processing of the sample image by the image processing model is obtained; the distillation feature loss function is determined based on either the first prediction image or the second prediction image, the processed result image, the first feature data, and the second feature data.
[0120] The image processing model training apparatus provided in this disclosure can execute the image processing model training method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0121] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0122] Figure 6 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure, as shown below. Figure 6 As shown, the device includes:
[0123] Image acquisition module 410 is used to acquire the image to be processed;
[0124] The image processing module 420 is used to process the image to be processed based on an image processing model to obtain a processed image. The image processing model is trained using the training method provided in the above embodiments.
[0125] The image processing apparatus provided in this disclosure can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0126] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0127] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0128] like Figure 7As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0129] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0130] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0131] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0132] The electronic device provided in this embodiment belongs to the same inventive concept as the training method of the image processing model or the image processing method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0133] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the training method or image processing method of the image processing model provided in the above embodiments.
[0134] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0135] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0136] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0137] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0138] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a target image; process the target image based on a pre-trained transfer processing model to obtain a first label image; perturb the target image to obtain a perturbed image; perform mirror image processing on the perturbed image and the first label image to obtain a sample image corresponding to the perturbed image and a second label image corresponding to the first label image; and train the image processing model to be trained based on the sample image and the second label image to obtain a trained image processing model.
[0139] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire an image to be processed, wherein the image to be processed includes a target image and a non-target image; process the image to be processed based on an image processing model to obtain a processed image, wherein the image processing model is trained based on the image processing model training method provided in the embodiments of this disclosure.
[0140] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0142] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0145] [In the detailed implementation section, after the entire text ends, please repeat all the content that you wish to protect in the form of claims in the following form:]
[0146] According to one or more embodiments of this disclosure, [Example 1] provides a method for training an image processing model, including:
[0147] The target image is acquired, and the target image is processed based on a pre-trained transfer processing model to obtain a first label image;
[0148] The target image is perturbed to obtain a perturbed image;
[0149] The perturbation image and the first label image are mirrored to obtain a sample image corresponding to the perturbation image and a second label image corresponding to the first label image;
[0150] The image processing model to be trained is trained based on the sample image and the second label image to obtain the trained image processing model.
[0151] According to one or more embodiments of this disclosure, [Example Two] provides a training method for the image processing model of Example One, further comprising:
[0152] The perturbation processing of the target image includes: blurring the contour information of the target image.
[0153] According to one or more embodiments of this disclosure, Example 3 provides a training method for the image processing model of Example 1, further comprising:
[0154] The perturbation processing of the target image further includes one or more of the following: adding white noise to the target image; performing random blurring processing on the target image; and performing compression processing on the target image.
[0155] According to one or more embodiments of this disclosure, Example 4 provides a training method for the image processing model of Example 1, further comprising:
[0156] The mirror morphology processing includes one or more of mirror deformation and mirror clipping.
[0157] According to one or more embodiments of this disclosure, Example 5 provides a training method for the image processing model of Example 1, further comprising:
[0158] The training of the image processing model to be trained based on the sample image and the second label image includes: in a first training phase, inputting the sample image into the image processing model to be trained to obtain a first predicted image, determining a first loss function based on the first predicted image and the second label image, and adjusting the model parameters of the image processing model based on the first loss function and the first learning rate to obtain an intermediate image processing model; in a second training phase, inputting the sample image into the intermediate image processing model to obtain a second predicted image, determining a second loss function based on the second predicted image and the second label image, and adjusting the model parameters of the image processing model based on the second loss function and the second learning rate to obtain a trained image processing model, wherein the second learning rate is less than the first learning rate.
[0159] According to one or more embodiments of this disclosure, Example Six provides a method for training the image processing model of Example One, further comprising:
[0160] The first loss function includes a first number of loss terms; the second loss function includes a second number of loss terms; wherein the loss terms include one or more of the following: image feature loss term, image quality loss term, first high-frequency feature loss term, second high-frequency feature loss term, image similarity loss term, and difference loss function.
[0161] According to one or more embodiments of this disclosure, [Example Seven] provides a training method for the image processing model of Example One, further comprising:
[0162] The method for determining the image feature loss term includes: extracting label features of the second labeled image based on the feature extraction module, extracting prediction features of the first predicted image or the second predicted image based on the feature extraction module, and determining the image feature loss term based on the label features and the prediction features.
[0163] According to one or more embodiments of this disclosure, Example 8 provides a training method for the image processing model of Example 1, further comprising:
[0164] The method for determining the image quality loss term includes: determining the first image quality data of the first predicted image or the second predicted image, and generating the image quality loss term based on the first image quality data;
[0165] Alternatively, determine the first image quality data of the first predicted image or the second predicted image, and the second image quality data of the second labeled image, and generate the image quality loss term based on the first image quality data and the second image quality data.
[0166] According to one or more embodiments of this disclosure, [Example Nine] provides a training method for the image processing model of Example One, further comprising:
[0167] The method for determining the first high-frequency feature loss term includes: extracting first high-frequency data from the first predicted image or the second predicted image, extracting second high-frequency data from the second label image; and determining the first high-frequency feature loss term based on the first high-frequency data and the second high-frequency data.
[0168] According to one or more embodiments of this disclosure, Example 10 provides a training method for the image processing model of Example 1, further comprising:
[0169] The method for determining the second high-frequency feature loss term includes: extracting first high-frequency data from the first predicted image or the second predicted image, and extracting second high-frequency data from the second label image; performing discrimination processing on the first high-frequency data and the second high-frequency data based on a preset discriminator, and determining the second high-frequency feature loss term based on the discrimination results of the first high-frequency data and the second high-frequency data.
[0170] According to one or more embodiments of this disclosure, Example 11 provides a training method for the image processing model of Example 1, further comprising:
[0171] The loss term further includes a distillation feature loss term, and the method for determining the distillation feature loss term includes: inputting the sample image into the transfer processing model to obtain the processed result image and the first feature data in the prediction processing layer of the transfer processing model during the processing; obtaining the second feature data in the prediction processing layer during the processing of the sample image by the image processing model; and determining the distillation feature loss function based on either the first prediction image or the second prediction image, the processed result image, the first feature data, and the second feature data.
[0172] According to one or more embodiments of this disclosure, [Example Twelve] provides an image processing method, including:
[0173] Obtain the image to be processed;
[0174] The image to be processed is processed based on the image processing model to obtain the processed image, wherein the image processing model is trained based on the training method of the image processing model provided in the above example.
[0175] According to one or more embodiments of this disclosure, [Example Thirteen] provides a training apparatus for an image processing model, comprising:
[0176] The first label image acquisition module is used to acquire a target image and process the target image based on a pre-trained transfer processing model to obtain a first label image.
[0177] An image perturbation processing module is used to perturb the target image to obtain a perturbed image;
[0178] The image morphology processing module is used to perform mirror morphology processing on the perturbation image and the first label image to obtain a sample image corresponding to the perturbation image and a second label image corresponding to the first label image.
[0179] The image training module is used to train the image processing model to be trained based on the sample image and the second label image, so as to obtain the trained image processing model.
[0180] According to one or more embodiments of this disclosure, [Example Fourteen] provides an image processing apparatus, including:
[0181] The image acquisition module is used to acquire the image to be processed.
[0182] An image processing module is used to process the image to be processed based on an image processing model to obtain a processed image, wherein the image processing model is trained based on the training method of the image processing model provided in the example above.
[0183] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0184] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0185] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A training method for an image processing model, characterized in that, include: Acquire the target image, process the target image based on the pre-trained transfer processing model, and obtain the first label image; The target image is perturbed to obtain a perturbed image; The perturbation image and the first label image are mirrored to obtain a sample image corresponding to the perturbation image and a second label image corresponding to the first label image; In the first training phase, the sample image is input into the image processing model to be trained to obtain a first predicted image. A first loss function is determined based on the first predicted image and the second label image. The model parameters of the image processing model are adjusted based on the first loss function and the first learning rate to obtain an intermediate image processing model. In the second training phase, the sample image is input into the intermediate image processing model to obtain a second predicted image. A second loss function is determined based on the second predicted image and the second label image. The model parameters of the image processing model are adjusted based on the second loss function and the second learning rate to obtain a trained image processing model, wherein the second learning rate is less than the first learning rate.
2. The method according to claim 1, characterized in that, The perturbation processing of the target image includes: blurring the contour information of the target image.
3. The method according to claim 2, characterized in that, The perturbation processing of the target image also includes one or more of the following: Add white noise to the target image; The target image is subjected to random blurring. The target image is compressed.
4. The method according to claim 1, characterized in that, The mirror morphology processing includes one or more of mirror deformation and mirror clipping.
5. The method according to claim 1, characterized in that, The first loss function includes a first number of loss terms; the second loss function includes a second number of loss terms; The loss term includes one or more of the following: image feature loss term, image quality loss term, first high-frequency feature loss term, second high-frequency feature loss term, image similarity loss term, and difference loss function.
6. The method according to claim 5, characterized in that, The method for determining the image feature loss term includes: extracting label features of the second labeled image based on the feature extraction module, extracting prediction features of the first predicted image or the second predicted image based on the feature extraction module, and determining the image feature loss term based on the label features and the prediction features.
7. The method according to claim 5, characterized in that, The method for determining the image quality loss term includes: determining the first image quality data of the first predicted image or the second predicted image, and generating the image quality loss term based on the first image quality data; Alternatively, determine the first image quality data of the first predicted image or the second predicted image, and the second image quality data of the second labeled image, and generate the image quality loss term based on the first image quality data and the second image quality data.
8. The method according to claim 5, characterized in that, The method for determining the first high-frequency feature loss term includes: Extract the first high-frequency data from the first predicted image or the second predicted image, and extract the second high-frequency data from the second label image; The first high-frequency feature loss term is determined based on the first high-frequency data and the second high-frequency data.
9. The method according to claim 5, characterized in that, The method for determining the second high-frequency feature loss term includes: Extract the first high-frequency data from the first predicted image or the second predicted image, and extract the second high-frequency data from the second label image; The first high-frequency data and the second high-frequency data are discriminated based on a preset discriminator, and the second high-frequency feature loss term is determined based on the discrimination results of the first high-frequency data and the second high-frequency data.
10. The method according to claim 5, characterized in that, The loss term also includes a distillation characteristic loss term, and the method for determining the distillation characteristic loss term includes: The sample image is input into the model to be transferred to obtain the processed result image and the first feature data in the prediction processing layer of the model to be transferred during the processing process; Obtain the second feature data in the prediction processing layer during the image processing model's processing of the sample image; The distillation feature loss function is determined based on either the first predicted image or the second predicted image, the processed result image, the first feature data, and the second feature data.
11. An image processing method, characterized in that, include: Obtain the image to be processed; The image to be processed is processed based on the image processing model to obtain the processed image, wherein the image processing model is trained based on the training method of the image processing model as described in any one of claims 1-10.
12. A training device for an image processing model, characterized in that, include: The first label image acquisition module is used to acquire a target image and process the target image based on a pre-trained transfer processing model to obtain a first label image. An image perturbation processing module is used to perturb the target image to obtain a perturbed image; The image morphology processing module is used to perform mirror morphology processing on the perturbation image and the first label image to obtain a sample image corresponding to the perturbation image and a second label image corresponding to the first label image. The image training module is used to train the image processing model to be trained based on the sample image and the second label image, so as to obtain the trained image processing model. The image training module is specifically used for: In the first training phase, the sample image is input into the image processing model to be trained to obtain a first predicted image. A first loss function is determined based on the first predicted image and the second label image. The model parameters of the image processing model are adjusted based on the first loss function and the first learning rate to obtain an intermediate image processing model. In the second training phase, the sample image is input into the intermediate image processing model to obtain a second predicted image. A second loss function is determined based on the second predicted image and the second label image. The model parameters of the image processing model are adjusted based on the second loss function and the second learning rate to obtain a trained image processing model, wherein the second learning rate is less than the first learning rate.
13. An image processing apparatus, characterized in that, include: The image acquisition module is used to acquire the image to be processed. An image processing module is used to process the image to be processed based on an image processing model to obtain a processed image, wherein the image processing model is trained based on the training method of the image processing model as described in any one of claims 1-10.
14. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of any image processing model as claimed in claims 1-10, or the image processing method as claimed in claim 11.
15. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform a training method for any of the image processing models as claimed in claims 1-10, or an image processing method as claimed in claim 11.
Citation Information
Patent Citations
Deep learning training sample optimization method
CN110070548A
Adversarial sample image generation method and device, electronic equipment and storage medium
CN114741701A