Method, device and computer readable storage medium for generating sketch image
By using a feature-learning generative adversarial network model trained with an adversarial loss function of control factors and an image detail loss function, the problem of insufficient image detail and texture features in sketch image generation is solved, thus improving the quality of generated sketch images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
- Filing Date
- 2021-12-16
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for generating sketch images are insufficient in terms of image detail and sketch texture features, resulting in low-quality generated sketch images.
A well-trained image processing model based on an adversarial loss function with control factors and an image detail loss function is adopted. Through a generative adversarial network model with feature learning, the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image is enhanced, thereby improving the quality of the generated sketch image.
By improving the performance of the image processing model, the image details and texture features of the generated sketch images are enhanced, thereby improving the quality of the generated sketch images.
Smart Images

Figure CN116266372B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, device, and computer-readable storage medium for generating sketch images. Background Technology
[0002] With the rapid development of artificial intelligence technology, image recognition has been applied to various industries and fields. The generation or synthesis of sketch images is an important technology in the field of image recognition. The synthesis of sketch images is the process of converting an optical image photograph into a sketch image. It usually utilizes an existing optical image photograph-sketch image database, enabling the computer to learn the complex mapping relationship between the two types of images, thereby building a model to generate a sketch image corresponding to an unknown optical image photograph.
[0003] In related technologies, methods for generating sketch images are generally divided into two categories: traditional sketch image generation methods and deep learning-based sketch image generation methods. Traditional sketch image generation methods are further divided into data-driven sketch image generation methods and model-driven sketch image generation methods. However, sketch images generated using data-driven, model-driven, and deep learning-based methods often lack rich facial details and sketch texture features, resulting in low-quality sketch images. Summary of the Invention
[0004] To address the aforementioned technical problems, embodiments of this application provide a method, apparatus, device, and computer-readable storage medium for generating sketch images, which can enhance the image details and sketch texture features of the generated sketch images and improve the quality of the generated sketch images.
[0005] The technical solution of this application is implemented as follows:
[0006] This application provides a method for generating a sketch image, including:
[0007] The optical image to be processed and the trained image processing model are acquired. The loss function of the trained image processing model is determined based on the adversarial loss function with a control factor and the image detail loss function. The control factor is used to improve the image processing model's ability to learn the difference between the predicted sketch image and the training sketch image.
[0008] The optical image to be processed is input into the trained image processing model to obtain the target sketch image of the optical image to be processed.
[0009] Output the target sketch image.
[0010] This application provides a sketch image generation apparatus, comprising:
[0011] An image acquisition unit is used to acquire an optical image to be processed and a trained image processing model. The loss function of the trained image processing model is determined based on an adversarial loss function with a control factor and an image detail loss function. The control factor is used to improve the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image.
[0012] The image processing unit is used to input the optical image to be processed into the trained image processing model to obtain a target sketch image of the optical image to be processed;
[0013] An image output unit is used to output the target sketch image.
[0014] This application provides a device for generating a sketch image, comprising:
[0015] Memory, used to store executable instructions for generating sketch images;
[0016] The processor is configured to execute the executable sketch image generation instructions stored in the memory to implement the sketch image generation method provided in the embodiments of this application.
[0017] This application provides a computer-readable storage medium storing instructions for generating a sketch image, which, when executed by a processor, implements the sketch image generation method provided in this application.
[0018] The sketch image generation method, apparatus, device, and computer storage medium provided in this application, employing this technical solution, firstly, acquire the optical image to be processed and a trained image processing model. The loss function of the trained image processing model is determined based on an adversarial loss function with control factors and an image detail loss function, used to improve the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image, thereby improving the performance of the image processing model. Then, the optical image to be processed is input into the trained image processing model to obtain the target sketch image of the optical image to be processed. Finally, the target sketch image is output from the trained image processing model. Thus, by using a trained image processing model whose loss function is determined based on an adversarial loss function with control factors and an image detail loss function, the image processing model is ensured to more fully learn the differences between the predicted sketch image and the training sketch image during training. This enhances the image detail and sketch texture features of the generated sketch image when using the trained image processing model for sketch image generation, thereby improving the quality of the generated sketch image. Attached Figure Description
[0019] Figure 1A schematic flowchart illustrating a method for generating a sketch image provided in an embodiment of this application;
[0020] Figure 2 This application provides a schematic flowchart of a method for training an image processing model.
[0021] Figure 3 This application provides a schematic diagram of the structure of a discrimination network module.
[0022] Figure 4 This is a schematic diagram of the structure of a feature extraction network module provided in an embodiment of this application;
[0023] Figure 5 A schematic flowchart illustrating a method for obtaining the discrimination result of a training optical image, provided in an embodiment of this application;
[0024] Figure 6 A flowchart illustrating a method for determining the loss function of an image processing model provided in this application;
[0025] Figure 7 This is a schematic flowchart of a method for obtaining a trained image processing model provided in an embodiment of this application;
[0026] Figure 8 This application provides a schematic flowchart of a method for generating an adversarial loss function.
[0027] Figure 9 This is a schematic diagram of the structure of a generative network module provided in an embodiment of this application;
[0028] Figure 10 This application provides a schematic flowchart of a method for acquiring a target sketch image of an optical image to be processed, as shown in the embodiments of this application.
[0029] Figure 11 This is a schematic diagram of the structure of a feature learning generative adversarial network model provided in an embodiment of this application;
[0030] Figure 12 A schematic diagram illustrating a training method for a generative adversarial network model based on feature learning, provided in an embodiment of this application;
[0031] Figure 13 A schematic diagram of the structure of a sketch image generation device provided in an embodiment of this application;
[0032] Figure 14 This is a schematic diagram of the structure of a sketch image generation device provided in an embodiment of this application. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0034] In the following description, references to “some embodiments / other embodiments” are made, which describe a subset of all possible embodiments. However, it is understood that “some embodiments / other embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0035] Based on different research methods, existing sketch image generation methods can be mainly divided into data-driven sketch image generation methods, model-driven sketch image generation methods, and deep learning-based sketch image generation methods.
[0036] Data-driven sketching image generation methods typically consist of four parts: image segmentation, nearest neighbor selection, weight calculation, and image patch stitching. Usually, similar image patches are first found in the training set, then the weights between different image patches are calculated, and the entire sketch image is stitched together according to these weights. Since the generated sketch image patches are linear combinations of sketch image patches in the training set, data-driven sketching image generation methods can generate relatively complete image detail features.
[0037] However, traditional data-driven sketch image generation methods rely heavily on the training dataset, and the size of the training dataset has varying effects on the quality of the generated sketch images. Furthermore, these methods require traversing the entire dataset to find similar image patches, consuming significant time and limiting their applicability. Therefore, data-driven methods suffer from high time complexity, cumbersome generation processes, and a failure to consider images outside the dataset.
[0038] Model-driven sketch image generation methods primarily work by constructing a mapping function between photographs and sketch images. During the training phase, the function distribution and mapping relationship between photographs and sketch images are learned. Then, the photograph to be generated is input, and the corresponding sketch image is generated through prediction using the offline-learned mapping function.
[0039] However, while traditional model-driven sketch generation methods are unaffected by dataset size, they struggle to find the optimal function distribution and mapping relationship between the photograph and the sketch image. This makes it difficult to determine a nonlinear regression model with strong image detail representation capabilities. Consequently, the generated results often lose some key information (such as glasses, hair clips, etc.) and lack sketch texture features, resulting in overly smooth images. Therefore, sketch images generated by this type of method are prone to image distortion and low sharpness.
[0040] Deep learning-based sketch generation methods learn the non-linear mapping relationship between photographs and sketches using neural networks. Currently, deep learning-based sketch generation methods often use convolutional neural networks and generative adversarial networks as their basic architecture, and then innovate and optimize upon this architecture.
[0041] Deep learning-based sketch image generation methods can solve the problems of blurry and distorted generated sketch images, but they often lose some image detail features, resulting in artifacts and blurry images in the generated sketch images.
[0042] It is evident that, among the relevant technologies, both traditional sketch image generation methods and deep learning-based sketch image generation methods are insufficient in terms of image detail and richness of sketch texture features, resulting in low-quality generated sketch images.
[0043] This application provides a method for generating a sketch image, which can enhance the image details and sketch texture features of the generated sketch image and improve the quality of the generated sketch image.
[0044] The following will describe the method for generating a sketch image provided in the embodiments of this application. This method can be applied to a sketch image generation device, which can be a terminal such as a smartphone, desktop computer, laptop computer, tablet computer, or server. Figure 1 The diagram shown is a flowchart illustrating a method for generating a sketch image according to an embodiment of this application. The method includes the following steps:
[0045] S101. Obtain the optical image to be processed and the trained image processing model.
[0046] Optical images can be color images captured by image acquisition devices, such as cameras or video cameras. They can include facial images, real-time scene images, landscape images, etc. The optical image to be processed is the one that needs to be processed to generate a corresponding sketch image. The optical image to be processed can be a dataset of optical images that has been captured beforehand and stored in a database. In this way, when generating a sketch image based on the optical image, the optical image to be processed can be directly retrieved from the database. The optical image to be processed can also be an image downloaded from the network or an image received from other terminals or servers.
[0047] The image processing model is a feature-learning-based generative adversarial network (GAN) model, which includes multiple deep neural networks. This model can generate sketch images. Before using the image processing model to generate sketch images, it needs to be trained to its optimal state, thus determining the best-trained image processing model. Inputting the optical image to be processed into the trained image processing model yields the corresponding sketch image. A sketch image is an image with a single color except for the background color, which can represent variations in brightness.
[0048] In some embodiments, the loss function of the trained image processing model is determined based on an adversarial loss function with a control factor and an image detail loss function. The control factor is used to improve the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image. During the training process of the image processing model, the loss function of the image processing model is required. The loss function characterizes the difference between the sketch image generated by the image processing model based on the optical image to be processed and the real sketch image corresponding to the optical image to be processed. Judging the value of the loss function can determine whether the image processing model has finished training. For example, the condition for the image processing model to finish training can be that the value of the loss function is less than a preset threshold, and the preset threshold can be any pre-set positive real number.
[0049] It should be noted that the predicted sketch image can be the sketch image generated after the optical image to be processed is input into the image processing model and processed during the image processing model training phase, and the training sketch image can be the real sketch image corresponding to the optical image to be processed.
[0050] Understandably, the control factor constrains the adversarial loss function, enabling the image processing model to learn more deeply about the differences between the output sketch image and the real sketch image during the training phase, as the number of training iterations increases. This ensures that the model is fully trained, improves its processing capabilities, and enhances the quality of the generated sketch image.
[0051] S102. Input the optical image to be processed into the trained image processing model to obtain the target sketch image of the optical image to be processed.
[0052] In this embodiment of the application, the trained image processing model includes a generative network module. When implementing step S102, the optical image to be processed can be input into the trained image processing model, and the generative network module in the trained image processing model processes the optical image to be processed to obtain a target sketch image. The target sketch image can be a sketch image corresponding to the optical image to be processed input into the trained image processing model. It has rich image details and sketch texture features, and has a high similarity to the real sketch image corresponding to the optical image to be processed.
[0053] S103. Output the target sketch image.
[0054] When step S103 is implemented by a terminal device such as a smartphone or desktop computer, the output target sketch image can be the target image displayed on the display device of the terminal device; when the step is implemented by a server, the output target sketch image can be the target sketch image sent to the terminal requesting image processing, and then the terminal requesting image processing, after receiving the target sketch image, displays the target sketch image on its own display device.
[0055] In this embodiment, an optical image to be processed and a trained image processing model are acquired. The loss function of the trained image processing model is determined by an adversarial loss function with control factors and an image detail loss function. This enhances the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image, thereby improving the model's performance. The optical image to be processed is input into the trained image processing model to obtain the target sketch image. The target sketch image is then output from the trained image processing model to obtain the generated sketch image. By using a trained image processing model whose loss function is determined by an adversarial loss function with control factors and an image detail loss function, the model is able to learn the differences between the predicted sketch image and the training sketch image more fully during training. This enhances the image detail and sketch texture features of the generated sketch image, improving its quality.
[0056] like Figure 2The diagram shown is a flowchart illustrating a method for training an image processing model according to an embodiment of this application. In some embodiments, the image processing model includes an adversarial learning module and a feature extraction network module. The adversarial learning module includes a generator network module and a discriminator network module. The generator network module can directly generate a corresponding sketch image from the input optical image. The process of training the image processing model includes the following steps:
[0057] S201. Obtain training data, which includes multiple training optical images and training sketch images corresponding to each training optical image.
[0058] When training an image processing model, training data is required. In this embodiment, the training data may include training optical images, which can be an optical image dataset formed by multiple optical images. The training optical images can be multiple optical images obtained after being captured in advance. These multiple optical images constitute an optical image dataset, which is stored in a database and can be directly retrieved from the database during actual use.
[0059] It should be noted that the training sketch image corresponds to the training optical image. It can be a real sketch image obtained by the user drawing based on the training optical image. The method of obtaining the corresponding real sketch image based on the training optical image is only an example and is not specifically limited.
[0060] In some embodiments, training sketch images can also be stored in a database, stored in a manner corresponding to training optical images. This way, when a training sketch image is needed, it and its corresponding training optical image can be directly retrieved from the database. Alternatively, all training sketch images can be stored in the same location, but separately from the training optical images. Each training sketch image and training optical image can be assigned a corresponding identifier. When training data is needed, the required training optical image and its corresponding training sketch image can be retrieved based on the identifier.
[0061] S202. Input each training optical image into the generation network module to obtain the predicted sketch image corresponding to each training optical image.
[0062] After acquiring the training optical images, the acquired training optical images are sequentially input into the generator network module for processing. The processed images are then output from the generator network module, which yields the predicted sketch images corresponding to the training optical images.
[0063] S203. Input the training sketch image and predicted sketch image corresponding to each training optical image into the discrimination network module to obtain the discrimination result of each training optical image.
[0064] It should be noted that the discriminant network module is used to determine the authenticity of the sketch images generated by the generator network (i.e., the discrimination result), and feeds back the authenticity of the sketch images to the generator network module, so that the network parameters of the generator network module can be adjusted based on the discrimination result, thereby improving the performance of the generator network module.
[0065] Understandably, the discriminative network module can verify local image patches of the input training sketch image and the predicted sketch image, focusing only on the local structure of the image, thus avoiding the information loss caused by directly verifying the complete sketch image, thereby effectively obtaining high-frequency information of the local parts of the sketch image.
[0066] In some embodiments, the discriminative network can use a convolution module to convolve the input training sketch image or predicted sketch image to obtain multiple convolutional image blocks. Each image block consists of multiple pixels, and each pixel corresponds to a pixel value, which can be in the range of [0, 255]. The difference between all adjacent pixel values in each image block is calculated, and the obtained differences are summed to obtain the total pixel difference or average pixel difference for each image block. Image blocks with a total pixel difference or average pixel difference greater than a preset pixel value are considered to have high-frequency information. The preset pixel value can be pre-set and is any integer greater than 0. Of course, the method for determining image blocks with high-frequency information is only illustrative and is not specifically limited here. Then, the image blocks with high-frequency information are verified to obtain the discrimination result of the predicted sketch image's authenticity.
[0067] Understandably, the discriminative network module only needs to learn the high-frequency information of the image and does not need to consider the size constraint of the input sketch image, which reduces the number of parameters in the training process of the generative network to a certain extent and speeds up the calculation.
[0068] In some embodiments, such as Figure 3 The diagram shown is a schematic representation of a discrimination network provided in an embodiment of this application. Figure 3 As can be seen, the discriminant network module 1 includes an input module 11, four convolutional modules (convolutional module 12, convolutional module 13, convolutional module 14, and convolutional module 15), and an output module 16. The input module 11 is used as the input for training or predicting sketch images. The four convolutional modules are used to distinguish between true and false input sketch images. Finally, the output module 16 outputs the result after discrimination. For example, the size of the convolutional kernel in the convolutional module can be 4×4, and the stride of the convolutional kernel can be set to 2. Table 1 shows a model parameter design table for a discriminant network module. The network parameters of the discriminant network model can be designed according to the table below.
[0069] Table 1. A model parameter design table for a discriminant network module.
[0070] Module Name Input Channel Step length Convolution kernel number Output Channel Convolutional Module 12 3 2 64 64 Convolutional Module 13 64 2 128 128 Convolutional Module 14 128 2 256 256 Convolutional Module 15 256 2 512 512 Output module 16 512 1 1 1
[0071] S204. Input the training sketch image and the predicted sketch image corresponding to each training optical image into the feature extraction network module to obtain the first sketch feature of each training sketch image and the second sketch feature of each predicted sketch image.
[0072] It should be noted that the feature extraction network module is used to extract image features from the predicted sketch image and the training sketch image, and to obtain the image features of the predicted sketch image and the training sketch image. The first sketch feature of the training sketch image can be the image feature of the training sketch image, and the second sketch feature of the predicted sketch image can be the image feature of the predicted sketch image.
[0073] In some embodiments, such as Figure 4 The diagram shown is a structural schematic of a feature extraction network module provided in an embodiment of this application. Figure 4 As shown, the feature extraction network module is designed based on VGGNet-16, but removes the last two convolutional modules and three fully connected modules. Feature extraction network module 2 includes an input module 21, three convolutional modules (convolutional module 22, convolutional module 23, and convolutional module 24), and an output module 25. Input module 21 is used to input either the predicted sketch image or the training sketch image; the output module is used to output the extracted image features from the predicted and training sketch images. Convolutional modules 22 and 23 have the same structure, each including two convolutional layers and one fully connected layer, with the connection method being convolutional layer-convolutional layer-pooling layer. Convolutional module 24 includes three sequentially connected convolutional layers. For example, the convolutional modules of the convolutional layers can all be set to 3×3, and the pooling windows of the pooling layers can all be set to 2×2.
[0074] Understandably, feature extraction networks can reduce the number of network parameters by using convolutional layers in multiple convolutional modules in series. Compared to deep neural networks built with only one convolutional module, they have more non-linear transformations and are more suitable for extracting image features.
[0075] S205. Compare the first sketch features of each training sketch image with the second sketch features of each predicted sketch image to obtain the feature comparison results.
[0076] It should be noted that the first sketch feature can be the feature matrix output by the feature extraction network module of the training sketch image, and the second sketch feature can be the feature matrix output by the feature extraction network module of the predicted sketch image. After obtaining the image features of the training sketch image and the predicted sketch image, the method for comparing the first sketch feature and the second sketch feature can be to calculate the error between the feature matrix of the training sketch image and the feature matrix of the predicted sketch image corresponding to the same training optical image.
[0077] S206. Determine the loss function of the image processing model based on the discrimination results and feature comparison results, and use the discrimination results and feature comparison results to perform backpropagation training on the image processing model to obtain the trained image processing model.
[0078] It should be noted that the loss function of an image processing model can characterize the feature error between the output sketch image and the real sketch image, and is used to determine the processing performance of the image processing model. The loss function of the image processing model can be determined based on the discrimination results output by the discriminant network module, and the error between the feature matrix of the training sketch image and the feature matrix of the predicted sketch image.
[0079] Backpropagation can be interpreted as feeding back the outputs of the discriminant network module and the feature extraction network to the image processing model. In some embodiments, the discrimination result of the discriminant network module can be fed back to the image processing model, and the error between the feature matrix of the trained sketch image and the feature matrix of the predicted sketch image can also be fed back to the image processing model, allowing the image processing model to continue training and become a well-trained image processing model, thereby improving the quality of the sketch images generated by the generative network module in the image processing model.
[0080] like Figure 5 The diagram shown is a flowchart of a method for obtaining the discrimination result of a training optical image according to an embodiment of this application. In some embodiments, the training sketch image and the predicted sketch image corresponding to each training optical image are input into the discrimination network module to obtain the discrimination result of each training optical image. That is, S203 can be implemented by the following S2031 to S2036. The steps are described below.
[0081] S2031. Input the training sketch image and the predicted sketch image corresponding to each training optical image into the discrimination network module.
[0082] It should be noted that the training sketch image and the predicted sketch image corresponding to each training optical image can be input into the discriminator network separately, and the input order is not limited. For example, the input order can be to input the training sketch image first and then the predicted sketch image, or to input the predicted sketch image first and then the training sketch image.
[0083] S2032. Divide the training sketch image into multiple image regions, determine the pixel values of each image region in the multiple image regions, and determine the first target image region based on the pixel values of each image region.
[0084] In some embodiments, the sketch image can be segmented by convolving the training sketch image with multiple convolutional kernels of the convolutional module in the discriminant network module to obtain multiple convolutional image blocks. Each image block is an image region, and each image region has multiple pixel values. The first target image region can be an image block with high-frequency information in the training sketch image. The pixel difference between all adjacent pixel values in each image region is calculated, and the calculated pixel differences are summed or averaged. The image block whose sum or average pixel value difference is greater than a preset pixel value is selected, and the image region corresponding to this image block is designated as the first target image region.
[0085] S2033. Obtain the first target feature information of the first target image region in the training sketch image.
[0086] It should be noted that target feature information refers to high-frequency information in the target image region, such as edges, noise, and details in a sketch image. Target feature information can be the pixel values corresponding to multiple pixels in the target image region. For example, it can be determined by calculating the pixel differences between the corresponding pixel values of all adjacent pixels in the target image region, and identifying pixels with pixel differences greater than a preset pixel difference as the first target feature information. The preset pixel difference can be any positive integer greater than zero. In practice, image feature extraction methods, including histogram matrix method, gray-level co-occurrence matrix method, and autoregressive texture feature method, can also be used to obtain the first target feature information in the first target image region.
[0087] S2034. Determine the image region in the predicted sketch image that corresponds to the first target image region in the training sketch image, and define the image region as the second target image region.
[0088] The second target image region can be an image patch with high-frequency information in the predicted sketch image. After determining the first target image region in the training sketch image, the corresponding position in the predicted sketch image can be determined as the second target image region based on the position of the first target image region in the training sketch image.
[0089] In some embodiments, the method for determining the second target image region is similar to the method for determining the training image, but it does not need to determine the image block corresponding to the target image region based on the pixel values of multiple image blocks obtained after convolution. Instead, it only needs to determine the image block corresponding to the first target image region as the second target feature by the arrangement order of the image blocks obtained after convolution.
[0090] S2035. Obtain the second target features of the second target image region.
[0091] The second target feature indicates high-frequency information in the second target image region. The method is the same as that for obtaining the first target feature information of the first target image region. The second target feature in the second target image region is determined based on the pixel difference between the corresponding pixel values of all adjacent pixels in the multiple pixels in the second target image region.
[0092] S2036. Based on the first target feature and the second target feature, obtain the output discrimination result of each training optical image.
[0093] After obtaining high-frequency information from the training and predicted sketch images, the discrimination result of the predicted sketch image can be obtained based on the high-frequency information in the image patches. The discrimination result can represent the similarity between the predicted sketch image and the real sketch image, and can be represented by a probability value [0, 1]. The closer the probability value is to 1, the more similar the predicted sketch image and the training sketch image are; conversely, the closer the probability value is to 0, the greater the difference between the predicted sketch image and the training sketch image.
[0094] like Figure 6 The diagram shown is a flowchart of a method for determining the loss function of an image processing model according to this application. In some embodiments, the loss function of the image processing model is determined based on the discrimination result and the feature comparison result, which can be implemented through steps S301 to S304. The steps are described below.
[0095] S301. Generate an adversarial loss function based on the discrimination results.
[0096] Inputting training and predicted sketch images into the discriminator network module yields multiple probability values. These multiple probability values corresponding to the training sketch images form the probability distribution of the real sketch image, and similarly, the multiple probability values corresponding to the predicted sketch images form the probability distribution of the generated sketch image. However, in practice, the discriminator and generator networks can be mutually antagonistic. The better the discriminator network is trained (for example, when the output probability values are all close to 1 after inputting training sketch images), the performance of the generator network will deteriorate. This can manifest as a severe vanishing gradient in the generator network after the discriminator result is backpropagated, preventing further optimization of the generator network.
[0097] Therefore, it is necessary to balance the distribution differences between the training sketch images and the predicted sketch images output by the discriminant network module to ensure the processing performance of the discriminant network module and the generator network module. The adversarial loss function represents the distribution differences between the training sketch images and the predicted sketch images and can be used to balance the distribution differences between the real sketch images and the generated sketch images. Optimizing this loss function can improve the performance of the generator network module and the discriminant network module.
[0098] S302. Based on the feature comparison results, generate an image detail loss function.
[0099] Feature comparison results can represent the differences between the image features of the predicted sketch image and the image features of the training sketch image. This allows for the assessment of the realism of the predicted sketch image from the perspective of image features, i.e., the sketch image generated by the generative network. Based on the differences between the image features of the predicted sketch image and the training sketch image, an image detail loss function can be determined. This image detail loss function characterizes the image detail error between the training sketch image and the predicted sketch image. The image detail loss function calculates the difference in the image feature space and performs detail error calculations on the image features extracted by the feature extraction network.
[0100] For example, the image detail loss function L_detail can be defined as Equation (1):
[0101]
[0102] Where w and h represent the dimensions of the feature map, s is the real sketch image or the training sketch image, G(p) is the predicted sketch image or the generated sketch face image, and φ(s) and φ(G(p)) represent the feature matrices output by the feature extraction network module of the training sketch image and the predicted sketch image, respectively, and can be: Where n w,k and k w,k Let ||φ(s)-φ(G(p))|| represent the pixel values corresponding to the feature matrices of the training and predicted sketch images, respectively. Thus, in equation (1), ||φ(s)-φ(G(p))|| F It can be derived from the following formula (2):
[0103]
[0104] Where tr represents the trace of the matrix, let [φ(s)-φ(G(p))] T ·[φ(s)-φ(G(p))] is M,
[0105]
[0106] Then in equation (2) It can be derived from the following formula (3):
[0107]
[0108] S303, Obtain the preset first weight value and second weight value.
[0109] It should be noted that the first weight value can be the weight of the adversarial loss function, representing the degree to which the adversarial loss function contributes to the overall loss function. The second weight value can be the weight of the image detail loss function, representing the degree to which the image detail loss function contributes to the overall loss function. The first and second weight values are balancing factors used to balance the initial values of the adversarial loss and facial detail loss. At the start of image processing model training, the first and second weight values can be preset to initial values. During training, these preset initial values can be continuously adjusted based on the value corresponding to the loss function. Both the first and second weight values are arbitrary real numbers greater than 0 and less than 1, and their sum is 1.
[0110] S304. Based on the adversarial loss function, the image detail loss function, the first weight value, and the second weight value, generate the loss function of the image processing model.
[0111] In some embodiments, the adversarial loss function and the image detail loss function can be determined by a weighted sum to define the loss function of the image processing model. Adversarial loss function L D Image detail loss function L d The loss function L, consisting of the first weight α and the second weight β, can be expressed by the following equation (4):
[0112] L=αL D +βL d (4);
[0113] It is understood that the loss function determined in the embodiments of this application can improve the performance of the generative network module, so that the sketch image generated by the generative network module has clear detail features and realistic sketch texture.
[0114] like Figure 7 The diagram shown is a flowchart of a method for obtaining a trained image processing model according to an embodiment of this application. In some embodiments, the image processing model is trained by backpropagation using the discrimination result and feature comparison result to obtain the trained image processing model. This can be achieved through the following steps S401 to S404. Each step is described below.
[0115] S401. Feed back the discrimination result to the discrimination network module and the generation network module, and adjust the network parameters of the discrimination network module and the generation network module.
[0116] It should be noted that the output of the discriminant network module can be fed back to the discriminant network module itself, or to the generator network module, so that the discriminant network module and the generator network module can continue to train and continuously adjust their respective network parameters.
[0117] In some embodiments, assuming θ G With θ D These represent the network parameters of the generator network module and the discriminator network module, respectively. r(s) For the data distribution of realistic sketched human face images, f g(p) To generate the data distribution for the sketched face image, the objective function for the discrimination network module can be expressed as formula (5):
[0118]
[0119] In equation (5) above, the first term represents the real sketch image s input to the discriminant network module. i The objective function value at time t, the second term is the sketch image s generated by the input generator network module. i The objective function value at time ', where D(s) represents the input real sketch face image s. i The probability value obtained by the discriminant network module, G(p), is the input optical image p to be processed. i The generative network module generates a sketched face image, and D(G(p)) represents the probability value obtained by the discriminative network when the generated sketched image G(p) is input. During the training of the image processing model, for the discriminative network module, when the input is real data s... i When the input is generated data G(p), the value of D(s) should be as close to 1 as possible, that is, the larger the value of the first term in equation (5), the better; when the input is generated data G(p), the value of D(G(p)) should be as close to 0 as possible, that is, the larger the value of the second term in equation (5), the better. Therefore, given the generator network module, maximizing V(D,G) yields the optimal discriminator network module.
[0120] The optimization process of equation (5) can be transformed into finding the optimal solution of equation (6):
[0121] ∫[f r(s) logD(s)+f g(s) log(1-D(s))]ds (6);
[0122] The function f in the integral term of the above equation r(s) logD(s)+f g(s) Taking the derivative of log(1-D(s)) with respect to D(s) and setting its value to 0, the expression for the optimal discriminant network is as follows:
[0123]
[0124] Thus, for a training dataset M = {(P} with N (N is a positive integer greater than or equal to 1) training optical images and training sketch images, i ,S i ), i=1,2,3…,N},θ G This can be obtained by optimizing the loss function of the generator network module, i.e.:
[0125]
[0126] It should be noted that during the training of the image processing module, the generation network module and the discriminator network module are trained alternately. For example, the parameters θ of the discriminator network are fixed first. D The network parameters θ of the generator network are trained using equation (8). G Then fix the parameters θ of the generator network. G The parameters θ of the discrimination network are updated and optimized using equation (7). D Until θ G and θ D All parameters have been optimized, and the loss function of the image processing model is in a balanced state, indicating that the training of the image processing model is complete.
[0127] S402. Based on the discrimination results, obtain the distribution difference between the training sketch image and the predicted sketch image, feed the distribution difference between the training sketch image and the predicted sketch image back to the generation network module, and adjust the network parameters of the generation network module and the network parameters of the discrimination network module.
[0128] It should be noted that the discriminant network outputs a probability value for either the training sketch image or the predicted sketch image. This probability value characterizes the degree of difference between the generated sketch image and the real sketch image. Therefore, based on the probability values of all training and predicted sketch images, the distributional differences between the training and predicted sketch images can be determined.
[0129] In some embodiments, the Wasserstein distance is used to measure the difference in probability distributions between the predicted and trained sketch images. Thus, the original adversarial loss function L is constructed based on the difference in probability distributions between the predicted and trained sketch images. D As shown in equation (9):
[0130]
[0131] Among them, f r(s) For the data distribution of real sketch images, f g(s)For the data distribution of the generated sketch image, the first two terms of equation (9) represent all joint distributions combining the data distribution of the generated sketch image with the data distribution of the real sketch image, and the third term is the gradient penalty term. The sample distribution... For the distribution of real sketch image samples f r(s) Distribution of pseudo-sketched face image samples f g(s) Random interpolation sampling on the connection line.
[0132] By continuously optimizing the original adversarial loss function and feeding the optimization results back to the generator and discriminator networks, the network parameters of the generator and discriminator modules are continuously adjusted, so that the function value of the original adversarial loss function reaches its optimum.
[0133] S403. Based on the feature extraction results, obtain the image detail error between the training sketch image and the predicted sketch image, feed the image detail error between the training sketch image and the predicted sketch image back to the generator network module, and adjust the network parameters of the generator network module and the network parameters of the feature extraction network module.
[0134] In some embodiments, after extracting image features from the training sketch image and the predicted sketch image, the image detail error between the two sketch images is determined by the error between the feature matrices of the training sketch image and the predicted sketch image. The image detail error between the training sketch image and the predicted sketch image is characterized by an image detail loss function, such as the image detail loss function L_detail shown in Equation (1) above. By continuously optimizing the image detail loss function L_detail and feeding the results of the optimization process back to the generator network module, the network parameters of the generator network module and the network parameters of the feature extraction module are continuously adjusted so that the image detail loss function L_detail reaches its optimal value.
[0135] S404. When the training termination condition is met, the trained image processing model is obtained.
[0136] In some embodiments, the training termination condition may be that the function value corresponding to the loss function of the image processing model is less than a preset loss value. The preset loss value can be any pre-defined positive real number, such as 2.3, 3.5, etc. The loss function of the image processing model is based on the original adversarial loss function L. D The image processing model's network parameters are obtained from the image detail loss function L_detail. When the loss function value of the image processing model is less than the preset loss value, the training of the image processing model is stopped, and the network parameters of the trained model are obtained.
[0137] like Figure 8The diagram shown is a flowchart of a method for generating an adversarial loss function according to an embodiment of this application. In some embodiments, the adversarial loss function is generated based on the discrimination result. That is, S301 can be implemented by the following S3011 to S3013. The steps are described below.
[0138] S3011. Based on the discrimination results, determine the initial adversarial loss function.
[0139] It is understandable that the initial adversarial loss function can be the original adversarial loss function L. D This is used to stabilize the training process of the discriminant network and the generative network model, and to prevent the gradient vanishing of the generative network from becoming more severe when the discriminant network module is trained better, thereby improving the gradient performance of the generative network module.
[0140] S3012. Obtain the preset decay coefficient and total number of training iterations, and determine the control factor based on the decay coefficient, total number of training iterations, and current number of training iterations.
[0141] It should be noted that the total number of training iterations represents the sum of training iterations for the image processing model and can be preset; the total number of training iterations can be any positive integer. The decay coefficient is a real number greater than 0 and less than 1. Initially, the decay coefficient can be preset to any real number between (0, 1) and can be adjusted during training based on the loss function value of the image processing model. The current training iteration can be the ranking of the image processing model within the total number of training iterations at the current training stage, and its value can be any positive integer less than the total number of training iterations.
[0142] In some embodiments, the control factor determined based on the decay coefficient ω, the total number of training iterations N, and the current number of training iterations n can be (1+ωn) / N, and the control factor is a real number greater than 0 and less than 1.
[0143] Understandably, the control factor ensures that during the image processing model training phase, as the number of iterations increases, the discriminative network module can progressively and deeply learn the differences between the generated sketch image and the real sketch image. This allows the discriminative network module to be fully trained, improving its discrimination ability and thus enhancing the quality of the sketch images generated by the generative network module.
[0144] S3013. Determine the adversarial loss function based on the control factor and the initial adversarial loss function.
[0145] In some embodiments, to address the problem of the model reaching equilibrium prematurely during image processing model training, causing the generator and discriminator network modules to cease optimization, this application embodiment adds a control factor to the adversarial loss function, based on the control factor (1+ωn) / N and the initial adversarial loss function L. D'Determined adversarial loss function L' D It can be shown in the following formula (10):
[0146] L D = (1+ωn) / N·L D '(10);
[0147] In some embodiments, the generating network module includes at least a strided convolutional module, a residual module, and a strided deconvolutional module, such as... Figure 9 The diagram shown is a structural schematic of a generative network module provided in an embodiment of this application. Figure 9 As can be seen, the generator network module 3 includes an input module 31, a convolution module 32, a residual module 33, a deconvolution module 34, and an output module 35.
[0148] In other embodiments, the convolutional module 31 may include three convolutional layers (convolutional layer 311, convolutional layer 312, and convolutional layer 313), the residual module 33 may include nine residual units, the deconvolutional module 34 may include two deconvolutional layers (deconvolutional layer 341 and deconvolutional layer 342), and the output module 35 may include one output layer 351. The size of the convolutional kernel in the convolutional layers can be 4×4. Batch Norm layers can be added to both the convolutional and deconvolutional layers to ensure that the generated network modules are fully trained and to avoid the gradient vanishing problem. In addition, to avoid gradient saturation during training and to make the training process more stable, the ReLU activation function is used inside the convolutional and deconvolutional layers.
[0149] For example, Table 2 shows a parameter design table for a generative network module provided in an embodiment of this application. As can be seen from Table 2, the stride of the first convolutional layer, convolutional layer 321, in the convolutional module is 2. This can be used to reduce the dimensionality of the input optical image, reducing the complexity of generating the sketch image while ensuring that the positional information of the optical image is not lost. The strides of convolutional layers 322 and 323, as well as deconvolutional layers 341 and 342, can all be set to 1 / 2, which can be used to learn the detailed features of the optical image more deeply, thereby improving the quality of the sketch image generated by the generative network module.
[0150] Table 2. Parameter Design Table for a Generative Network Module
[0151] Layer name Input Channel Step length Convolution kernel number Output Channel Convolutional layer 321 3 2 64 64 Convolutional layer 322 64 2 128 128 341 deconvolution layers 128 2 256 256 342 deconvolution layers 256 2 512 512 Output layer 351 512 1 1 1
[0152] like Figure 10 The diagram shown is a flowchart illustrating a method for acquiring a target sketch image of an optical image to be processed, based on an embodiment of this application. Figure 9The structure of the generative network module shown takes the optical image to be processed as input into the trained image processing model and obtains the target sketch image of the optical image to be processed, i.e., S102. This can be achieved by the following steps S1021 to S1024. Each step is explained below.
[0153] S1021. Input the optical image to be processed into the generator network module of the trained image processing model. The generator network module's convolution module with stride performs convolution processing on the optical image to be processed, and obtains the convolution-processed image.
[0154] In some embodiments, the optical image to be processed is divided into three channels, R, G, and B, in the generator network module. The optical image to be processed can be input from the input module 31 of the generator network module into the convolution module 32, and then pass through the convolution layer 321, convolution layer 322, and convolution layer 323 in sequence to perform convolution processing on the optical image to be processed. When passing through the convolution layer 321 with a stride of 2, the dimensionality of the input optical image to be processed can be reduced to obtain multiple optical image blocks after convolution processing. The number of multiple optical image blocks is the same as the number of convolution kernels, and the image size of the multiple optical image blocks is smaller than the image size of the optical image to be processed.
[0155] It is understandable that the optical image block output from the convolutional layer 321 is input into the convolutional layer 322 with a stride of 1 / 2. Convolution on it can effectively expand the height and width of the optical image block, that is, the dimension or size of the optical image block. Convolution in the convolutional layer 323 with the same stride of 1 / 2 can further expand the optical image block, and ensure that the positional information of the optical image is not lost during the downsampling process, such as average pooling or max pooling.
[0156] S1022. Use the residual module to perform identity mapping on the convolutional image to obtain the identity mapping result.
[0157] It should be noted that after the optical image to be processed is convolved by the convolution module, it is directly used as the input of the residual module 33. The residual module 33 may include nine residual units connected in sequence. Each residual unit includes multiple stacked network layers, which may include convolutional layers and batch normalized layers. The identity mapping process means that the input after convolution by the convolution module 32 can be directly mapped to the output of any residual unit, and the output of the previous residual unit can also be directly mapped to all residual units after the current residual unit. By establishing a direct correlation channel between the input and output, the network layers with network parameters in the residual unit can learn the residual between the input and output in a concentrated manner, thereby obtaining the output of the residual module, that is, the result of the identity mapping.
[0158] Understandably, the residual network module 33 represents the training of each residual unit as learning a residual function based on the input, making the generator network module easier to optimize, reducing network degradation problems such as gradient vanishing or gradient exploding due to the increase in the number of network layers, and prompting the generator network module to generate higher quality sketch images.
[0159] S1023. Use a deconvolution module with stride to perform deconvolution processing on the identity mapping result to obtain the deconvolution processed image.
[0160] In some embodiments, the identity mapping result output from the residual module 33 may be an image with a size smaller than the size of the optical image to be processed and different from the multiple image blocks output after the convolution module convolution processing. All image blocks are sequentially input into the deconvolution module 34 for deconvolution processing to achieve upsampling of all image blocks, expand the size of the image blocks, and make the size of the image after deconvolution processing the same as the size of the optical image to be processed.
[0161] S1024. Generate a target sketch image based on the image after deconvolution.
[0162] It should be noted that the optical image obtained after deconvolution processing is output through the output module 34. The output module may include a Tanh function layer to reduce the dimensionality of the optical image after deconvolution processing, converting the RGB three-dimensional color image into a one-dimensional grayscale image, i.e., the target sketch image.
[0163] It is understood that this embodiment achieves effective expansion of image height and width during the optical image upsampling process by generating convolutional, residual, and deconvolutional modules in the network module. Furthermore, as the number of network layers gradually increases, the image information does not lose positional information and continuously adds detailed feature information, thereby improving the image quality of the generated sketch image.
[0164] The implementation process of the embodiments of this application in a practical application scenario will be described below.
[0165] The sketch image generation method in this application embodiment is based on a feature-learning generative adversarial network model. Figure 11 This is a schematic diagram of the structure of a feature learning generative adversarial network model provided in an embodiment of this application. The feature learning generative adversarial network model 4 (image processing model) includes an adversarial learning module 41 (adversarial learning module) and a feature extraction network module 42 (feature extraction network module). The adversarial learning module includes a generator network 411 (generator network module) and a discriminator network 412 (discriminator network module), and the feature extraction network module 42 includes a feature extraction network 421 (feature extraction module).
[0166] In some embodiments, such as Figure 12 The diagram shown is a flowchart illustrating a training method for a generative adversarial network model based on feature learning, provided in an embodiment of this application. The method includes the following steps:
[0167] S501. Acquire an optical facial image (training optical image) and a real sketch image (training sketch image) corresponding to the optical facial image.
[0168] Optical facial images can be optical facial photographs, while real sketch images can be corresponding sketch images drawn by artists based on optical facial photographs. There can be multiple optical facial photographs and real sketch images. All optical facial photographs form an optical facial photograph dataset, and all real sketch images form a real sketch image dataset.
[0169] S502. Input the optical image into the generator network (generator network module) and output the corresponding synthetic sketch image (predicted sketch image).
[0170] When an optical facial photograph is input, a generative network directly generates a corresponding sketched face image (predicted sketch image). Assume the dataset M = {(p i ,s i (i = 1, 2, 3, ..., n) (The training data includes multiple training optical images and the corresponding training sketch images for each training optical image), p i Represents optical facial photograph data, s i This represents the data from sketched facial images (training sketched images). The generative network primarily learns from optical facial photographs. i To sketch human face images i The mapping relationship, namely: s i '=F(p i ).
[0171] S503. Input the real sketch image (training sketch image) and the synthetic sketch image (predicted sketch image) into the discriminant network (discriminant network module), output the discrimination result (discrimination result of the training optical image), generate the original adversarial loss (initial adversarial loss function) based on the discrimination result, and calculate the adversarial loss (adversarial loss function) based on the original adversarial loss and the control factor.
[0172] The discriminative network provides adversarial supervision during the feature learning adversarial network training phase to distinguish between the generated pseudo-sketch image (predicted sketch image) s' and the real sketch image s (training sketch image), and feeds the discrimination results back to the generator network.
[0173] The discriminant network feeds its output discrimination results back to the generator network (feeding back the distribution difference between the training sketch image and the predicted sketch image to the generator network module), which can be done by feeding back the adversarial loss generated based on the discrimination results to the generator network.
[0174] Understandably, to address the issue of prematurely reaching equilibrium during the training of a feature-learning generative adversarial network (GAN) model, causing the generator and discriminator networks to cease optimization, this embodiment adds a control factor to the original robust loss. This control factor ensures that during training, as the number of iterations increases, the discriminator network can progressively and deeply learn the differences between pseudo-sketched face images and real sketched face images. This allows the network model to be fully trained, improving its discriminative ability and enhancing the quality of the synthesized sketched face images by the generator network.
[0175] S504. Input the real sketch image and the synthetic sketch image into the feature extraction network (feature extraction network module) for feature extraction, and calculate the facial detail loss (facial detail loss function) based on the result of feature extraction.
[0176] After extracting special functions from real and synthetic sketch images, feature maps (first sketch features and second sketch features) of real and synthetic sketch images are obtained. Error calculation is performed on the feature maps of the two (to obtain the image detail error between the training sketch image and the predicted sketch image), and the calculation result is fed back to the generator network (to feed back the image detail error between the training sketch image and the predicted sketch image to the generator network module).
[0177] During the feature learning and generative adversarial network training phase, the feature extraction network maps the pseudo-images s' and the real sketch images s to the latent feature space f. s ',f s :F:s→f s ,s'→f s By learning image features in the latent feature space, the detailed information of the synthetic sketch face image (predicted sketch image) is enhanced.
[0178] S505. Generate a composite loss function (loss function of the image processing model) based on adversarial loss, facial detail loss, and balance factors (preset first and second weight values).
[0179] The determination of the loss function is crucial to the synthesis effect of the generative network. To ensure that the synthesized sketch face image has clear facial features and realistic sketch texture, this method proposes a composite loss function: L total =αL D +βL detail L DTo counteract the loss, a measure is used to measure the distributional difference between the generated sketched face image and the real sketched face image. L detail α represents the facial detail loss, used to measure the detail error between the generated sketched face image and the real sketched face image. α and β are balancing factors used to balance the initial values of the adversarial loss and the facial detail loss.
[0180] Understandably, during the training phase of the generative adversarial network (GAN) model, after inputting an optical facial photograph, the generator network directly generates a corresponding sketched face image. A discriminator network and a facial feature extraction network are also introduced. The discriminator network is used to determine the authenticity of the synthesized sketched face image and employs an adversarial training strategy to optimize the learning effect of the generator network and the discriminator network's discrimination ability. The facial feature extraction network extracts features from both the real and synthesized sketched face images and calculates facial detail errors. The calculation results are then fed back online to the generator network to improve the quality of the synthesized sketched face image.
[0181] This application also provides an apparatus for generating a sketch image. Figure 13 This is a schematic diagram of the structure of a sketch image generation device provided in an embodiment of this application, as shown below. Figure 13 As shown, the sketch image generation device 5 includes:
[0182] Image acquisition unit 51 is used to acquire the optical image to be processed and the trained image processing model. The loss function of the trained image processing model is determined based on the adversarial loss function with control factors and the image detail loss function. The control factors are used to improve the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image.
[0183] Image processing unit 52 is used to input the optical image to be processed into a trained image processing model to obtain a target sketch image of the optical image to be processed;
[0184] Image output unit 53 is used to output the target sketch image.
[0185] In some embodiments of this application, the sketch image generation device 5 further includes: an image processing model training unit 54, used to acquire training data, the training data including multiple training optical images and training sketch images corresponding to each training optical image; input each training optical image into a generation network module to obtain a predicted sketch image corresponding to each training optical image; input the training sketch images and predicted sketch images corresponding to each training optical image into a discriminant network module to obtain a discrimination result for each training optical image; input the training sketch images and predicted sketch images corresponding to each training optical image into a feature extraction network module to obtain a first sketch feature for each training sketch image and a second sketch feature for each predicted sketch image; compare the first sketch feature for each training sketch image and the second sketch feature for each predicted sketch image to obtain a feature comparison result; determine the loss function of the image processing model based on the discrimination result and the feature comparison result, and perform backpropagation training on the image processing model using the discrimination result and the feature comparison result to obtain a trained image processing model.
[0186] In some embodiments of this application, the image processing model training unit 54 is further configured to input the training sketch image and the predicted sketch image corresponding to each training optical image into the discriminative network module; divide the training sketch image to obtain multiple image regions, determine the pixel value of each image region in the multiple image regions, and determine a first target image region based on the pixel value of each image region; obtain first target feature information of the first target image region in the training sketch image; determine the image region in the predicted sketch image corresponding to the first target image region in the training sketch image, and determine the image region as a second target image region; obtain a second target feature of the second target image region; and obtain the output discriminative result of each training optical image based on the first target feature and the second target feature.
[0187] In some embodiments of this application, the image processing model training unit 54 is further configured to generate the adversarial loss function based on the discrimination result, the adversarial loss function representing the distribution difference between the training sketch image and the predicted sketch image; generate the image detail loss function based on the feature comparison result, the image detail loss function representing the image detail error between the training sketch image and the predicted sketch image; obtain a preset first weight value and a second weight value; and generate the loss function of the image processing model based on the adversarial loss function, the image detail loss function, the first weight value, and the second weight value.
[0188] In some embodiments of this application, the image processing model training unit 54 is further configured to: feed back the discrimination result to the discrimination network module and the generation network module, and adjust the network parameters of the discrimination network module and the generation network module; based on the discrimination result, obtain the distribution difference between the training sketch image and the predicted sketch image, feed back the distribution difference between the training sketch image and the predicted sketch image to the generation network module, and adjust the network parameters of the generation network module and the discrimination network module; based on the feature extraction result, obtain the image detail error between the training sketch image and the predicted sketch image, feed back the image detail error between the training sketch image and the predicted sketch image to the generation network module, and adjust the network parameters of the generation network module and the network parameters of the feature extraction module; and when the training termination condition is determined to be met, obtain the trained image processing model.
[0189] In some embodiments of this application, the image processing model training unit 54 is further configured to determine an initial adversarial loss function based on the discrimination result; obtain a preset decay coefficient and total number of training iterations; determine a control factor based on the decay coefficient, total number of training iterations, and current number of training iterations, wherein the decay coefficient is a real number greater than 0 and less than 1, and the control factor is a real number greater than 0 and less than 1; and determine the adversarial loss function based on the control factor and the initial adversarial loss function.
[0190] In some embodiments of this application, the image processing model training unit 54 is further configured to input the optical image to be processed into the generative network module in the trained image processing model, and to perform convolution processing on the optical image to be processed by the strided convolution module in the generative network module to obtain a convolutionally processed image; to perform identity mapping processing on the convolutionally processed image using the residual module to obtain an identity mapping result; to perform deconvolution processing on the identity mapping result using the strided deconvolution module to obtain a deconvolutionally processed image, wherein the size of the deconvolutionally processed image is the same as the size of the optical image to be processed; and to generate the target sketch image based on the deconvolutionally processed image.
[0191] This application also provides a device for generating sketch images. Figure 14 This is a schematic diagram of the structure of a sketch image generation device provided in an embodiment of this application, as shown below. Figure 14 As shown, the sketch image generation device 6 includes: a memory 61 for storing executable sketch image generation instructions; and a processor 62 for executing the executable sketch image generation instructions stored in the memory to implement the method provided in the embodiments of this application, for example, implementing the sketch image generation method provided in the embodiments of this application.
[0192] This application provides a computer-readable storage medium storing executable instructions for generating a sketch image, which, when executed by a processor 62, implements the method provided in this application, such as the sketch image generation method provided in this application.
[0193] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0194] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0195] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0196] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0197] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for generating a sketch image, characterized in that, include: The process involves acquiring an optical image to be processed and a trained image processing model. The loss function of the trained image processing model is determined based on an adversarial loss function with a control factor and an image detail loss function. The control factor is used to improve the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image. The trained image processing model is determined at least based on the discrimination result. The discrimination result is the discrimination result of each training optical image obtained by inputting the training sketch image and the predicted sketch image corresponding to each training optical image into the discrimination network module. The optical image to be processed is input into the trained image processing model to obtain the target sketch image of the optical image to be processed. Output the target sketch image; The step of inputting the training sketch image and predicted sketch image corresponding to each training optical image into the discriminant network module to obtain the discrimination result of each training optical image includes: The training sketch image and the predicted sketch image corresponding to each training optical image are input into the discrimination network module; The training sketch image is divided into multiple image regions, and the pixel values of each image region are determined. Based on the pixel values of each image region, a first target image region is determined. The first target image region is an image region in the training sketch image that has high-frequency information, and the high-frequency information includes at least the edge part, noise part, and detail part of the training sketch image. Based on the pixel difference between all adjacent pixels in the first target image region, the first target feature information of the first target image region in the training sketch image is determined. Based on the position of the first target image region in the training sketch image, an image region corresponding to the first target image region in the predicted sketch image is determined, and the image region is determined as the second target image region; Based on the pixel difference between all adjacent pixels in the second target image region, the second target feature of the second target image region is determined. Based on the first target feature and the second target feature, the output discrimination result of each training optical image is obtained.
2. The method according to claim 1, characterized in that, The image processing model includes an adversarial learning module and a feature extraction network module. The adversarial learning module includes a generator network module and a discriminator network module. The method for training the image processing model includes: Acquire training data, which includes multiple training optical images and training sketch images corresponding to each training optical image; Each training optical image is input into the generation network module to obtain the predicted sketch image corresponding to each training optical image; The training sketch image and the predicted sketch image corresponding to each training optical image are input into the discrimination network module to obtain the discrimination result of each training optical image; The training sketch image and the predicted sketch image corresponding to each training optical image are input into the feature extraction network module to obtain the first sketch feature of each training sketch image and the second sketch feature of each predicted sketch image. The first sketch features of each training sketch image and the second sketch features of each predicted sketch image are compared to obtain the feature comparison results. The loss function of the image processing model is determined based on the discrimination result and the feature comparison result, and the image processing model is trained by backpropagation using the discrimination result and the feature comparison result to obtain the trained image processing model.
3. The method according to claim 2, characterized in that, The step of determining the loss function of the image processing model based on the discrimination result and the feature comparison result includes: The adversarial loss function is generated based on the discrimination result, and the adversarial loss function characterizes the distribution difference between the training sketch image and the predicted sketch image; Based on the feature comparison results, the image detail loss function is generated, which characterizes the image detail error between the training sketch image and the predicted sketch image. Obtain the preset first and second weight values; The loss function of the image processing model is generated based on the adversarial loss function, the image detail loss function, the first weight value, and the second weight value.
4. The method according to claim 3, characterized in that, The step of using the discrimination result and the feature comparison result to perform backpropagation training on the image processing model to obtain a trained image processing model includes, The discrimination result is fed back to the discrimination network module and the generation network module, and the network parameters of the discrimination network module and the generation network module are adjusted. Based on the discrimination result, the distribution difference between the training sketch image and the predicted sketch image is obtained, and the distribution difference between the training sketch image and the predicted sketch image is fed back to the generation network module to adjust the network parameters of the generation network module and the network parameters of the discrimination network module. Based on the feature extraction results, the image detail error between the training sketch image and the predicted sketch image is obtained, and the image detail error between the training sketch image and the predicted sketch image is fed back to the generator network module to adjust the network parameters of the generator network module and the network parameters of the feature extraction module. When the training termination condition is met, the trained image processing model is obtained.
5. The method according to claim 3, characterized in that, The step of generating the adversarial loss function based on the discrimination result includes: Based on the discrimination results, the initial adversarial loss function is determined; Obtain a preset decay coefficient and total number of training iterations, and determine a control factor based on the decay coefficient, total number of training iterations, and current number of training iterations. The decay coefficient is a real number greater than 0 and less than 1, and the control factor is a real number greater than 0 and less than 1. The adversarial loss function is determined based on the control factor and the initial adversarial loss function.
6. The method according to any one of claims 2 to 5, characterized in that, The generator network module includes at least a strided convolutional module, a residual module, and a strided deconvolutional module. The step of inputting the optical image to be processed into the trained image processing model to obtain a target sketch image of the optical image to be processed includes: The optical image to be processed is input into the generative network module of the trained image processing model. The convolutional module with stride in the generative network module performs convolution processing on the optical image to be processed to obtain the convolution-processed image. The residual module is used to perform an identity mapping process on the convolutional image to obtain the identity mapping result. The identity mapping result is deconvolved using the deconvolution module with stride to obtain a deconvolutioned image. The size of the deconvolutioned image is the same as the size of the optical image to be processed. The target sketch image is generated based on the image after deconvolution.
7. A device for generating a sketch image, characterized in that, include: An image acquisition unit is used to acquire the optical image to be processed and a trained image processing model. The loss function of the trained image processing model is determined based on an adversarial loss function with a control factor and an image detail loss function. The control factor is used to improve the image processing model's ability to learn the differences between the predicted sketch image and the training sketch image. The trained image processing model is determined at least based on the discrimination result. The discrimination result is the discrimination result of each training optical image obtained by inputting the training sketch image and the predicted sketch image corresponding to each training optical image into the discrimination network module. The image processing unit is used to input the optical image to be processed into the trained image processing model to obtain a target sketch image of the optical image to be processed; An image output unit is used to output the target sketch image; The image processing model training unit is used to input the training sketch image and the predicted sketch image corresponding to each training optical image into the discrimination network module; The training sketch image is divided into multiple image regions, the pixel values of each image region are determined, and the first target image region is determined based on the pixel values of each image region. The first target image region is an image region with high-frequency information in the training sketch image, and the high-frequency information includes at least: the edge part, the noise part, and the detail part of the training sketch image; based on the pixel difference between all adjacent pixels in the first target image region, a first target feature information of the first target image region in the training sketch image is determined; based on the position of the first target image region in the training sketch image, an image region corresponding to the first target image region in the predicted sketch image is determined, and the image region is determined as the second target image region; based on the pixel difference between all adjacent pixels in the second target image region, a second target feature of the second target image region is determined; based on the first target feature and the second target feature, the output discrimination result of each training optical image is obtained.
8. A device for generating a sketch image, characterized in that, include: Memory, used to store executable instructions for generating sketch images; A processor, when executing the executable sketch image generation instructions stored in the memory, implements the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The system stores instructions for generating a sketch image, which, when executed by a processor, implement the method as described in any one of claims 1 to 6.