Artistic portrait generation model for simulating drawing process of real artist
Through an artistic portrait generation model that mimics the painting process of a realistic artist, using sketch convolution and superpixel component attention modules, the challenge of generating high-fidelity portrait sketches is solved, achieving high-quality and detailed portrait sketch generation.
Patent Information
- Application Number
- CN202510029409.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has challenges in generating high-fidelity portrait sketches, especially in dynamic face adaptation, precise overall facial structure and local texture capture.
A artistic portrait generation model that imitates the painting process of a realistic artist is proposed. Through the facial perception module, including sketch convolution and superpixel component attention, it is integrated into the U-Net architecture to simulate the "draft" and "refinement" steps of painting, dynamically perceive and render facial texture details.
The model is able to generate excellent and stable portrait sketches, significantly improving the quality and detail fidelity of generated images, reaching the latest performance levels.
Smart Images

Figure CN120070651A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artistic portrait creation, and particularly to an artistic portrait generation model that mimics the painting process of real artists. Background Art
[0002] Portrait painting aims to generate a gray-scale sketch portrait from a color facial image, and is usually applied to scenarios such as criminal investigation and digital entertainment. Portrait line drawing is a traditional minimalist art style that constructs the facial structure only through sparse lines. In the past decade, generative adversarial networks have achieved great success in portrait painting. Recently, researchers have proposed various techniques, including semantic region adaptive normalization, semantic labels for facial organs, semi-supervised learning, and cyclic structures, to improve the quality of portrait synthesis. Despite significant progress in drawing synthesis, there are still deficiencies in accurately adapting the dynamics of the human face to the overall facial structure and local texture, and there are considerable challenges in generating high-fidelity drawings. There are considerable challenges in generating high-quality images. To address this challenge, referring to the process of real artists painting portraits, as Figure 1 shown, the process of different artists painting portraits when observing different groups. Although there are various artistic styles, most artists adhere to the following two steps during the painting process:
[0003] 1. Sketching: In the initial stage, artists carefully observe the subject, including the shape and proportion of the head. Sketching is the foundation of the drawing process. Artists use basic geometric shapes such as ellipses to draw basic shapes and outlines. On this basis, dividing lines are added to preliminarily determine the positions of the eyes, nose, and mouth. This crucial step ensures that the proportion and composition of the artwork are accurate from the very beginning. This careful planning helps prevent errors during the detailed drawing stage and ensures that the final work is accurate and harmonious.
[0004] 2. Refinement: After completing the sketching step, artists start to pay more attention to capturing details and expressions. They gradually perfect the facial features, and detail the characteristics such as eyes, nose, mouth, and ears. By layering details such as shadows and highlights, the portrait gradually becomes rich in details. This step-by-step enhancement adds vitality to the artwork, increasing depth and dimension.
[0005] Based on the above process of artists painting portraits, an artistic portrait generation model that mimics the painting process of real artists is proposed. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides an artistic portrait generation model that mimics the painting process of real artists, and solves the problems raised in the above background art.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: An artistic portrait generation model that mimics the painting process of real artists, including a facial perception module, which consists of a sketch convolution and a superpixel component attention, and the facial perception module is integrated into the U-Net architecture; the sketch convolution is used to capture the basic contours in the global structure; the superpixel component attention is used to deepen the local texture and depth of the face, and the U-Net architecture is composed of an encoder E and a decoder D.
[0008] Optionally, the encoder adapts the depth map for spatial adaptive normalization and then connects with the decoder for feature integration. In the first four layers of the decoder, four FP blocks and U-Net attention blocks are adopted. The features pass through the draft convolution, which consists of deformable convolutions in four different directions, simulating the steps of sketching in real life. Through the routing mechanism, texture features with different weights are dynamically output. The features are processed by component attention and will undergo T iterations of repeated clustering to further enhance the texture details of the components, mimicking the refinement steps. Residual connections are added in each module to enhance continuity and stability. In addition, to further strengthen the features, cross-attention calculation is performed with the encoder features before inputting to the next layer and then input to the next facial perception module.
[0009] Optionally, the sketch convolution includes deformable convolution, single-direction deformable convolution, and multi-directional dynamic fusion;
[0010] For the deformable convolution: taking the 3x3 convolution as an example, the sampling receptive field R with a dilation rate of 1 can be expressed as:
[0011] R = {(-1, -1), (-1, 0), …, (1, 1)}. (1)
[0012] Given the input X ∈ R N×H×W , for the 3x3 deformable convolution, the feature value y(p0) at each position in the output feature map can be expressed as:
[0013]
[0014] where p0 enumerates the positions in R, w(pn) is the convolution weight corresponding to the point p0, pn is the relative offset within the receptive field R, Δpn is the learnable offset generated by the offset field, usually generated by an additional convolution. Since the learnable offset Δpn is usually a fractional value, the value of x(p0 + pn + Δpn) needs to be obtained through bilinear interpolation;
[0015] (0++Δ) is the corresponding element in the input feature map. Since the learnable offset Δ is usually a fraction, the value of (0++Δ) needs to be obtained through bilinear interpolation:
[0017]
[0018] Among them, g(m,n) = max(0, 1 - |m - n|). Q represents the area where p + p + p can contain point p0 + pn + Δpn;
[0019] The one-way deformable convolution: To ensure that the learnable offset moves in a specific direction, an iterative strategy is adopted to transform the receptive field, ensuring that the deformable convolution maintains a certain degree of continuity and avoiding the over-expansion of the receptive field due to large-range offsets, so that the model can more finely perceive diverse texture details. Four directions are set for deformation: the x-axis, the diagonal of the first and third quadrants, the y-axis, and the diagonal of the second and fourth quadrants;
[0020] Taking the offset learning process along the y-axis as an example, a 3x3 convolution kernel with the center Kc as the core is set. The deformation of the convolution starts from the center and continuously expands and deforms along the y-axis to both sides. In the convolution, each grid position can be represented by its relative position with respect to the convolution kernel center Kc: Kc ± i = (xc ± i, yc ± i), where i = {0, 1, 2, 3, 4} represents the horizontal or vertical distance from the convolution kernel center;
[0021] The initial iteration starts from the center Kc of the convolution kernel. The deformation is an iterative process, where the position offset of each convolution kernel depends on the previously deformed convolution kernel. +1 is obtained by adding a single-step offset Δ\DeltaΔ to Kc. To ensure a certain degree of continuity, |Δ| ≤ 1 is set. Therefore, through the cumulative iterative process, the deformed coordinates of each position along the y-axis can be expressed as:
[0022]
[0023] Similarly, the deformed coordinates of each position along the x-axis can be expressed as:
[0024]
[0025] The deformation along the diagonal of the first and third quadrants is more complex because offsets are added to both the x and y coordinates simultaneously, which may cause the convolution kernel group to deviate spatially from the previous position. To maintain the continuity of the receptive field, fixed offsets are added to the x-axis and y-axis simultaneously at each step. According to the vector operation rules, the single-step offset is decomposed into Δx′ and Δy′ offsets, and the distance of the single-step offset is limited to a maximum receptive field of 1:
[0026]
[0027] Among them and Similarly, the deformations in the diagonal directions of the second and fourth quadrants can be expressed by the following formula:
[0028]
[0029] In addition, the cumulatively added offset is usually a fractional value, which is implemented using bilinear interpolation:
[0030]
[0031] where K represents the fractional position, and K′ enumerates all neighboring spatial positions.
[0032] Optionally, for the multi-directional dynamic fusion: Human facial features include texture details in different directions such as eyebrows, nose bridge, lip edges, and eye sockets. To dynamically capture these different texture details, deformable convolutions in four different directions are used to perceive the texture through dynamic fusion:
[0033] Y(X in ) = F(Concat(Y x (X in ), Y y (X in ), Y xy (X in ), Y yx (X in ))), (9)
[0034] where Y(Xin) is the output feature after dynamic fusion, and the fusion function F uses a 1x1 convolution operation to reduce the feature dimension while retaining important feature information.
[0035] Optionally, for the superpixel component attention: The SLIC algorithm is divided into the following four steps:
[0036] S1: Associate each pixel with the nearest cluster center;
[0037] S2: Calculate the distance between each pixel and the cluster center in five dimensions (including x and y coordinates and RGB pixel values);
[0038] S3: Update the cluster center coordinates within each cluster;
[0039] S4: Iterate T times until the cluster partition remains unchanged. In the initialization stage, the center coordinates of each cluster are randomly generated;
[0040] To handle the non-differentiability of the nearest neighbor operation, differentiable SLIC is used, which calculates the soft association between pixels and the cluster center Q pi t :
[0041]
[0042] Among them, D represents distance calculation: D(x, y) = ||x – y|| 2 Q pi t represents the weight of pixel p relative to the cluster center in the t-th iteration. The updated center point coordinates will use the weighted sum of the pixels in the group:
[0043]
[0044] By assigning different weights, the attention can more flexibly adapt to different structures and patterns in the image, and it has better robustness to various changes in the image (such as illumination changes, occlusion, etc.). When the attention is concentrated on the key parts of the image, it can reduce the processing of the background or irrelevant regions and enhance the characterization of facial features. Therefore, attention is used to replace the soft association distance calculation:
[0045]
[0046] Among them, d represents the number of channels, and the matrix form is as follows:
[0047]
[0048] Among them, the updated center point coordinates S i t The matrix form is as follows:
[0049]
[0050] Among them, represents the normalized weight matrix. After k iterations, the center point matrix of each region will be inversely mapped back to the pixel matrix X' of each region:
[0051]
[0052] Among them, [(Q^k)T]-1 represents the inverse mapping matrix.
[0053] Optionally, the U-Net attention: uses a cross-attention mechanism to enhance the model's ability to perceive diverse texture details. Compared with using skip connections, the attention mechanism introduces more context information during the decoding process, enhancing the model's understanding of the overall image structure. The query (Query, Q) of the cross-attention comes from the encoding layer, while the key (Key, K) and value (Value, V) come from the decoding layer. The features in the decoder are processed through residual connections and passed to the next decoding layer after upsampling.
[0054] Optionally, the loss function: The model will be trained with the following loss function:
[0055] GAN Loss: The adversarial loss is used to measure whether \(\hat{Y}\) is real or fake:
[0056]
[0057] Texture Loss: To make the texture details of the synthesized sketch consistent with the target sketch, gradient constraints are imposed on it. When calculating, the Sobel operator is used to obtain their gradient information, and then the average cosine distance between them is calculated as the optimization objective:
[0058]
[0059] where \(H\) and \(W\) are the width and height, and \(\|\cdot\|\) represents the norm of the vector, represents the gradients of \(Y\) in the \(x\) and \(y\) directions, represents the gradients of \(\hat{Y}\) in the \(x\) and \(y\) directions;
[0060] Reconstruction Loss: Reconstruction constraints are imposed on the synthesized sketch and the target sketch at the pixel - level scale:
[0061]
[0062] Total Loss: All the above losses are combined as the total loss:
[0063]
[0064] where \(\lambda_1\) and \(\lambda_2\) represent weight hyperparameters.
[0065] The present invention provides an artistic portrait generation model that mimics the painting process of real artists, having the following beneficial effects:
[0066] The artistic portrait generation model that mimics the painting process of real artists simulates two key stages of painting: "sketching" and "refinement". "Sketching" is completed through sketch convolution, which can dynamically perceive the texture details of the face from multiple directions, while "refinement" uses super - pixel component attention to render more detailed textures. Different from the traditional way of directly learning to map a face photo to a sketch, the APP model simulates the actual artistic painting process and can generate excellent and stable human sketches. In the artistic sketch generation task, a large number of experiments show that this model has reached the latest performance level in both quantitative and qualitative evaluations. Brief Description of the Drawings
[0067] Figure 1 are the drawing process diagrams of different artists of the present invention;
[0068] Figure 2 is the schematic diagram of the process architecture of the present invention;
[0069] Figure 3 These are the deformation process diagrams of the present invention in four directions;
[0070] Figure 4 These are the qualitative comparison diagrams of the present invention with the best on the FS2K dataset;
[0071] Figure 5 These are the comparison diagrams of different model variants in the ablation experiment of the present invention;
[0072] Figure 6 These are the comparison diagrams of different convolutions of the present invention;
[0073] Figure 7 These are the setting diagrams of different FP modules of the present invention. Detailed implementation manners
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0075] Please refer to Figures 1 to 7 , the present invention provides a technical solution: an artistic portrait generation model that mimics the painting process of a real artist, including a face perception module, the face perception module is composed of a sketch convolution and a superpixel component attention, and the face perception module is integrated into the U-Net architecture; the sketch convolution is used to capture the basic contours in the global structure; the superpixel component attention is used to deepen the local texture and depth of the face, and the U-Net architecture is composed of an encoder E and a decoder D.
[0076] Furthermore, the encoder adapts the depth map for spatial adaptive normalization, and then connects with the decoder for feature integration. In the first four layers of the decoder, four FP blocks and U-Net attention blocks are adopted. The features pass through the draft convolution, and the draft convolution is composed of deformable convolutions in four different directions, simulating the steps of sketching in real life. Through the routing mechanism, texture features with different weights are dynamically output. The features are processed by component attention, and the features will undergo repeated clustering for T times to further enhance the texture details of the components, mimicking the refinement steps. Residual connections are added in each module to enhance continuity and stability. In addition, in order to further strengthen the features, before being input to the next layer, cross-attention calculation is also performed with the encoder features, and then input to the next face perception module.
[0077] Furthermore, the sketch convolution includes deformable convolution, single-direction deformable convolution and multi-direction dynamic fusion;
[0078] For the deformable convolution: taking the 3x3 convolution as an example, the sampling receptive field R with a dilation rate of 1 can be expressed as:
[0079] R = {(-1, -1), (-1, 0), …, (1, 1)}. (1)
[0080] Given the input X ∈ R N×H×W , for 3x3 deformable convolution, the eigenvalue y(p0) at each position in the output feature map can be expressed as:
[0081]
[0082] where p0 enumerates the positions in R, w(pn) is the convolution weight corresponding to the point p0, pn is the relative offset within the receptive field R, Δpn is the learnable offset generated by the offset field, usually generated by an additional convolution. Since the learnable offset Δpn is usually a fractional value, the value of x(p0 + pn + Δpn) needs to be obtained through bilinear interpolation;
[0083] (x(p0 + pn + Δpn)) is the corresponding element in the input feature map. Since the learnable offset Δ is usually fractional, the value of (x(p0 + pn + Δpn)) needs to be obtained through bilinear interpolation:
[0084]
[0085] where g(m, n) = max(0, 1 - |m - n|). Q represents the region where p + p + p can contain the point p0 + pn + Δpn;
[0086] The one - direction deformable convolution: To ensure that the learnable offset moves in a specific direction, an iterative strategy is adopted to transform the receptive field, ensuring that the deformable convolution maintains a certain degree of continuity and avoiding the over - expansion of the receptive field due to large - range offsets, so that the model can more finely perceive diverse texture details. Four directions are set for deformation: the x - axis, the diagonal of the first and third quadrants, the y - axis, and the diagonal of the second and fourth quadrants;
[0087] Taking the offset learning process along the y - axis as an example, a 3x3 convolution kernel with the center Kc as the core is set. The deformation of the convolution starts from the center and continuously expands and deforms along the y - axis to both sides. In the convolution, each grid position can be represented by its relative position with respect to the convolution kernel center Kc: Kc ± i = (xc ± i, yc ± i) where i = {0, 1, 2, 3, 4} represents the horizontal or vertical distance from the convolution kernel center;
[0088] The initial iteration starts from the center \(K_c\) of the convolutional kernel. Deformation is an iterative process, where the position offset of each convolutional kernel depends on the previously deformed convolutional kernel. \(K_{c + 1}\) is obtained by adding a single-step offset \(\Delta\) to \(K_c\). To ensure a certain continuity, it is set that \(|\Delta|\leq1\). Therefore, through the cumulative iterative process, the deformed coordinates at each position along the y-axis can be expressed as:
[0089]
[0090] Similarly, the deformed coordinates at each position along the x-axis can be expressed as:
[0091]
[0092] The deformation along the diagonal of the first and third quadrants is more complex because offsets are added simultaneously in both the x and y coordinates, which may cause the convolutional kernel group to deviate spatially from the previous position. To maintain the continuity of the receptive field, fixed offsets are added along the x-axis and y-axis simultaneously at each step. According to the vector operation rules, the single-step offset is decomposed into \(\Delta x'\) and \(\Delta y'\) offsets, and the distance of the single-step offset is limited to a maximum receptive field of 1:
[0093]
[0094] where and Similarly, the deformation in the direction of the diagonal of the second and fourth quadrants can be expressed by the following formula:
[0095]
[0096] In addition, the cumulatively added offsets are usually fractional values, which are implemented using bilinear interpolation:
[0097]
[0098] where \(K\) represents the fractional position, and \(K'\) enumerates all adjacent spatial positions;
[0099] The multi-directional dynamic fusion: Human facial features include texture details in different directions such as eyebrows, nose bridges, lip edges, and eye sockets. To dynamically capture these different texture details, deformable convolutions in four different directions are used to perceive the texture through dynamic fusion:
[0100] Y(X in ) = F(Concat(Y x (X in ),Y y (X in ),Y xy (X in ),Y yx (Xin )),(9)
[0101] Among them, Y(Xin) is the output feature after dynamic fusion. The fusion function F uses 1x1 convolution operation to reduce the feature dimension while retaining important feature information.
[0102] Furthermore, the superpixel component attention: The SLIC algorithm is divided into the following four steps:
[0103] S1: Associate each pixel with the nearest cluster center;
[0104] S2: Calculate the distance between each pixel and the cluster center in five dimensions (including x and y coordinates and RGB pixel values);
[0105] S3: Update the cluster center coordinates within each cluster;
[0106] S4: Iterate T times until the cluster partition remains unchanged. In the initialization stage, the center coordinates of each cluster are randomly generated;
[0107] To handle the non-differentiability of the nearest neighbor operation, differentiable SLIC is used, which calculates the soft association between pixels and the cluster center Q pi t as follows:
[0108]
[0109] Among them, D represents the distance calculation: D(x,y) = ||x – y|| 2 . Q pi t represents the weight of pixel p relative to the cluster center in the t-th iteration. The updated center point coordinates will use the weighted sum of the n pixels within the group:
[0110]
[0111] By assigning different weights, the attention can more flexibly adapt to different structures and patterns in the image, and it has better robustness to various changes in the image (such as illumination changes, occlusion, etc.). When the attention is focused on the key parts of the image, it can reduce the processing of the background or irrelevant regions and enhance the characterization of facial features. Therefore, attention is used to replace the soft association distance calculation:
[0112]
[0113] Among them, d represents the number of channels, and the matrix form is as follows:
[0114]
[0115] Among them, the updated center point coordinates Si t The matrix form is as follows:
[0116]
[0117] Among them, (Q^t)T represents the normalized weight matrix. After k iterations, the center point matrix of each region will be inversely mapped back to the pixel matrix X′ of each region:
[0118]
[0119] Among them, [(Q^k)T]-1 represents the inverse mapping matrix.
[0120] Furthermore, the U-Net attention: uses a cross-attention mechanism to enhance the model's ability to perceive diverse texture details. Compared with using skip connections, the attention mechanism introduces more context information during the decoding process, enhancing the model's understanding of the overall image structure. The query (Query, Q) of the cross-attention comes from the encoding layer, while the key (Key, K) and value (Value, V) come from the decoding layer. The features in the decoder are processed through residual connections and passed to the next decoding layer after upsampling.
[0121] Furthermore, the loss function: The model will be trained with the following loss function:
[0122] GAN loss: Uses adversarial loss to measure whether Y^ is real or fake:
[0123]
[0124] Texture loss: In order to make the texture details of the synthesized sketch consistent with the target sketch, gradient constraints are imposed on it. During calculation, the Sobel operator is used to obtain their gradient information, and then the average cosine distance between them is calculated as the optimization target:
[0125]
[0126] Among them, H and H are the width and height, and ||·|| represents the norm of the vector, represents the gradient of Y in the x and y directions, represents the gradient of Y^ in the x and y directions;
[0127] Reconstruction loss: Imposes reconstruction constraints on the synthesized sketch and the target sketch at the pixel level scale:
[0128]
[0129] Total loss: Combines all the above losses as the total loss:
[0130]
[0131] Among them, λ1 and λ2 represent weight hyperparameters.
[0132] Embodiment
[0133] Perform experimental tests on this model:
[0134] Experimental settings
[0135] S1: Experimental details
[0136] Use the pytorch1.7.1 version to train the model. All experiments are run on a single NVIDIA RTX3090. The model is trained for 800 epochs with a batch size of 4. Use the Adam optimizer, set β1 to 0.5, set β2 to 0.999, set the learning rate to 2e-4, set both λ1 and λ2 of the loss function to 4, and configure the number of clustering iterations to 1 in each superpixel component attention.
[0137] S2: Dataset
[0138] Use the FS2K dataset for experiments. This dataset is the largest publicly available photo-sketch pair dataset in the FSS field. This high-quality dataset contains 2,104 pairs of photos and sketches, including various real scenes, Internet resources, and compiled data from other datasets. The sketches in FS2K are divided into three different artistic styles. According to the standard settings, the dataset is divided into two parts: the training set and the test set; for each artistic style, the training set consists of 353 pairs of images, a total of 1,059 pairs for model training. Similarly, the test sets for each style consist of 623 pairs, 379 pairs, and 43 pairs of images respectively, totaling 1,045 pairs, and use the corresponding depth maps in HIDA.
[0139] S3: Evaluation metrics
[0140] Select six metrics to evaluate the quality of the generated sketches, including Frechet inception distance (FID), peak signal-to-noise ratio (PSNR), learned perceptual image patch similarity (LPIPS), structural similarity index (SSIM), structure co-occurrence texture (SCOOT), and feature similarity measure (FSIM). When evaluating the generated sketches, multiple metrics are used. Each metric indicates the image quality from a different perspective. Lower FID and LPIPS values indicate that the image is visually closer to the target sketch; on the contrary, higher PSNR, SSIM, SCOOT, and FSIM values indicate higher structural and texture similarity. These metrics will be marked with "↓" indicating the lower the better, and "↑" indicating the higher the better for simplified representation;
[0141] S4: Comparison with existing technologies:
[0142] The technical solution proposed in the present invention was compared with several excellent methods, including Pix2pix, Pix2PixH, and CycleGAN. In addition, it was also compared with some latest existing technologies, such as FSGAN, HIDA, and MDAL. To ensure a fair comparison, each model was retrained using the code released by the original authors, and the same default environment and parameters were adopted to maintain the consistency of the evaluation.
[0143] S41: Quantitative evaluation;
[0144] Table 1: Quantitative comparisons with SOTAs on the FS2K dataset.
[0145]
[0146] Table 1 shows the quantitative performance criteria of each method on the FS2K test dataset. The best results are shown in bold, and the second-best results are underlined. Obviously, the solution proposed in the present invention performs excellently in all evaluation metrics and achieves excellent performance. In particular, the FID score of the present invention is significantly reduced, decreasing by 29.474 points compared with FSGAN and 5.528 points compared with the second-best result (MDAL), indicating that the model has an obvious advantage in perceptual quality. The SCOOT score is approximately 21.1% higher than the second-best result (MDAL), indirectly indicating that the generated sketches have significant advantages in texture details. In addition, the PSNR of the present invention is increased by 13.4%, which means that the technical solution proposed in the present invention retains most of the details of the original image and generates the highest-quality images. In terms of the perceptual similarity metrics - LPIPS (Alex and VGG), SSIM, and FSIM - the values of the present invention are 0.264, 0.378, 0.469, and 0.634 respectively, all of which are the best records, indicating that the sketches proposed in the present invention are very similar to the target sketches in terms of structure and local gradients and can generate more realistic face sketches;
[0147] S42: Qualitative evaluation:
[0148] Figure 4Shows three styles of sketches generated by different methods on the FS2K dataset. As shown in the figure, other methods usually result in facial distortion, blurring, and discontinuous lines. In contrast, the present invention can clearly define the facial edges and render the facial texture, generating satisfactory results. Especially in low-light environments or when the skin color is close to the background, other methods often exhibit severe distortion, blurring, and structural discontinuity. The present invention not only well preserves the shapes of the eyes and mouth but also maintains clear texture. In addition, the present invention ensures a stable facial structure and delicate details. Through the advanced generation technology that simulates human painting, the generated sketches are closer to real art sketches. This success further verifies the effectiveness of the sketch convolution and component attention in the model of the present invention.
[0149] S43: Ablation experiments:
[0150] First, a series of ablation experiments were conducted to study the contributions of different modules in the proposed method. The HIDA model was selected as the baseline, which not only allows for comparison of the work with the latest state-of-the-art in the field but also helps to highlight the potential improvements or advantages of the proposed method in dealing with dynamic perception and texture detail diversity.
[0151] S431: Quantitative evaluation:
[0152] Table 2: Ablation study results on the FS2K dataset. (1) DrC: Draft Convolution. (2) SCA: Superpixel Component Attention. (3) UA: U-Net Attention.
[0153]
[0154]
[0155] Table 2 lists the performance metrics achieved by different variants of the model. The draft convolution (DrC), superpixel component attention (SCA), and U-Net attention (UA) were tested individually, and most performance metrics were observed to improve relative to the baseline model. When further evaluating these dual-combination modules, the performance metrics were enhanced again. In particular, the combination of draft convolution and superpixel component attention significantly improved most performance metrics, reaching a sub-optimal level, highlighting the important role of these two modules in enhancing the model performance. When all three modules were integrated, all performance metrics except SCOOT and SSIM reached the best level, which further verified the effectiveness of each module and especially highlighted the key contributions of draft convolution and component attention.
[0156] S4322: Qualitative Evaluation:
[0157] Figure 5 The sketches generated by different model variants are shown. It can be seen that the generated sketches are blurred around the eyes and the lines of the chin are discontinuous, which may be due to the weak perception of facial lines and components. After adding sketch convolution, the continuity and structure of the lines are enhanced, indicating that sketch convolution has a strong ability to capture the line directions in the sketches. The model enhanced only by component attention generates more realistic eyes, which indirectly proves that this module can improve the details of components and reduce the obvious disharmony in the lines. Finally, the model integrating all three modules generates visually more advantageous facial features;
[0158] S4323: Analysis of Sketch Convolution:
[0159] To further analyze the impact of sketch convolution, it was compared with standard convolution (SC), deformable convolution version 2 (DCv2), and directional deformable convolution (DDC). Figure 6 The visualization results of this comparison are shown. The sketches generated using standard convolution and deformable convolution show chaotic lines in the nose area, while sketch convolution successfully captures the correct texture details in the real image and generates cleaner lines. Although DDC performs similarly in the nose area, it significantly destroys the details around the eyes, and the metrics in Table 3 also indirectly support this. Compared with using a single directional deformable convolution, the mouth lines generated based on sketch convolution are more complete and clean, reflecting the ability of this module to dynamically perceive texture details;
[0160] S4324: Analysis of Facial Perception Module:
[0161] Table 4:Quantitative comparisons of FP Block
[0162]
[0163]
[0164] The effects of deploying the facial perception (FP) module in each layer of the decoder were explored. It was systematically added layer by layer and the performance changes were observed. Table 4 shows that when the FP module is added to the first four layers, the performance metrics gradually improve, and the generated results reach the best level in most metrics. However, when the FP module is added to more layers, the performance decreases. This decrease may be because in high-level decoding, the features have been relatively fully reconstructed, and the feature extraction ability of the FP module becomes redundant and may even affect the reconstruction effect. Figure 7It also shows that adding the FP module to the first four layers can produce the best subjective effect, while adding it to higher layers will cause distortion;
[0165] From the above experiments, inspired by the artist's painting process, the present invention simulates two key stages of painting: "drafting" and "refinement". "Drafting" is completed through sketch convolution, which can dynamically perceive the texture details of the face from multiple directions, while "refinement" uses superpixel component attention to render more detailed textures. Different from the traditional way of directly learning to map a face photo to a sketch, the model simulates the actual art painting process, aiming to generate excellent and stable human sketches. In the challenging art sketch generation task, a large number of experiments show that the model has reached the latest performance level in both quantitative and qualitative evaluations.
[0166] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An artistic portrait generation model that imitates the painting process of real artists, characterized by: The invention comprises a face perception module, which is composed of a sketch convolution and a super-pixel component attention, and is integrated into a U-Net architecture; the sketch convolution is used to capture the basic contour in the global structure; the super-pixel component attention is used to deepen the local texture and depth of the face, and the U-Net architecture is composed of an encoder E and a decoder D.
2. The artistic portrait generation model according to claim 1 that imitates the painting process of real artists, characterized in that: The encoder adapts to the depth map for spatial adaptive normalization, and then connects with the decoder for feature integration. In the decoder, the first four layers use four FP blocks and U-Net attention blocks. The features are passed through draft convolution, which consists of four deformable convolutions in different directions, simulating the steps of drawing a draft in real life. Texture features of different weights are dynamically output through the routing mechanism. The features are processed through component attention. The features will be repeatedly clustered for T iterations to further enhance the texture details of the components, imitating the refinement step. Residual connections will be added to each module to enhance continuity and stability. In addition, in order to further strengthen the features, before entering the next layer, they will be cross-attention calculated with the encoder features and then input into the next facial perception module.
3. The artistic portrait generation model according to claim 1 that imitates the painting process of real artists, characterized in that: The sketch convolution includes deformable convolution, unidirectional deformable convolution and multi-directional dynamic fusion; The deformable convolution: Taking 3x3 convolution as an example, the sampling receptive field R with a dilation rate of 1 can be expressed as: R={(-1,-1),(-1,0),…,(1,1)}. (1) Given an input X∈R N×H×W , for 3x3 deformable convolution, the feature value y(p0) at each position in the output feature map can be expressed as: Where p0 enumerates the position in R, w(pn) is the convolution weight corresponding to point p0, pn is the relative offset within the receptive field R, Δpn is the learnable offset generated by the offset field, usually generated by additional convolution. Since the learnable offset Δpn is usually a fractional value, the value of x(p0+pn+Δpn) needs to be obtained by bilinear interpolation; (0++Δ) is the corresponding element in the input feature map. Since the learnable offset Δ is usually a fraction, bilinear interpolation is required to obtain the value of (0++Δ): Where g(m,n)=max(0,1-|mn|), Q represents the area where p+p+p can contain the point p0+pn+Δpn; The unidirectional deformable convolution: In order to ensure that the learnable offset moves in a specific direction, an iterative strategy is adopted to transform the receptive field, ensuring that the deformable convolution maintains a certain degree of continuity, avoiding excessive expansion of the receptive field due to a large range of offsets, so that the model can perceive diverse texture details more finely, and setting four directions for deformation: the x-axis, the diagonals of the first and third quadrants, the y-axis, and the diagonals of the second and fourth quadrants; Taking the offset learning process along the y-axis as an example, a 3x3 convolution kernel with the center Kc as the core is set, and the deformation of the convolution is set to start from the center and continuously expand and deform on both sides along the y-axis. In the convolution, each grid position can be represented by its relative position relative to the convolution kernel center Kc: Kc±i=(xc±i,yc±i) where i={0,1,2,3,4} represents the horizontal or vertical distance from the convolution kernel center; The initial iteration starts from the center of the convolution kernel Kc. The deformation is an iterative process, in which the position offset of each convolution kernel depends on the previously deformed convolution kernel. +1 is obtained by adding a single-step offset Δ\DeltaΔ to Kc. In order to ensure a certain continuity, |Δ|≤1 is set. Therefore, through the cumulative iterative process, the deformed coordinates of each position along the y-axis can be expressed as: Likewise, the deformed coordinates of each position along the x-axis can be expressed as: The deformation along the diagonal of the first and third quadrants is more complicated because offsets are added to both the x and y coordinates, which may cause the convolution kernel group to deviate from its previous position in space. In order to maintain the continuity of the receptive field, a fixed offset is added along both the x and y axes at each step. According to the vector operation rules, the single-step offset is decomposed into Δx′ and Δy′ offsets, and the distance of the single-step offset is limited to a maximum receptive field of 1: in as well as Similarly, the deformation in the diagonal direction of the second and fourth quadrants can be expressed by the following formula: Additionally, the cumulatively added offset is usually a fractional value, implemented using bilinear interpolation: Among them, K represents the fractional position and K′ enumerates all adjacent spatial positions.
4. The artistic portrait generation model according to claim 1 that imitates the painting process of real artists, characterized in that: The multi-directional dynamic fusion: Human facial features include texture details in different directions such as eyebrows, nose bridge, lip edges and eye sockets. In order to dynamically capture these different texture details, four deformable convolutions in different directions are used to perceive the texture through dynamic fusion: Y(X in )=F(Concat(Y x (X in ),Y y (X in ),Y xy (X in ),Y yx (X in )), (9) Among them, Y(Xin) is the output feature after dynamic fusion, and the fusion function F uses a 1x1 convolution operation to reduce the feature dimension while retaining important feature information.
5. The artistic portrait generation model that imitates the painting process of real artists according to claim 1, characterized in that: The superpixel component attention: SLIC algorithm is divided into the following four steps: S1: Associate each pixel with the nearest cluster center; S2: Calculate the distance between each pixel and the cluster center in five dimensions (including x and y coordinates and RGB pixel values); S3: Update the coordinates of the cluster center within each cluster; S4: Iterate T times until the cluster partition remains unchanged. In the initialization stage, the center coordinates of each cluster are randomly generated; In order to deal with the non-differentiability of the nearest neighbor operation, the differentiable SLIC is used, which separates the pixels from the cluster center Q pi t Soft correlation calculation of: Where D represents the distance calculation: D(x,y)=||x–y|| 2 , Q pi t Represents the weight of pixel p relative to the cluster center in the tth iteration. Updating the center point coordinates will use the weighted sum of the pixels in the group: By assigning different weights, attention can adapt to different structures and patterns in the image more flexibly, and it has better robustness to various changes in the image (such as illumination changes, occlusion, etc.). When attention is focused on the key parts of the image, it can reduce the processing of background or irrelevant areas and enhance the depiction of facial features. Therefore, attention is used to replace the soft correlation distance calculation: Where d represents the number of channels and the matrix form is as follows: Among them, update the center point coordinates S i t The matrix form is as follows: Among them, (Q^t)T represents the normalized weight matrix. After k iterations, the center point matrix of each region will be reversely mapped back to the pixel matrix X′ of each region: Wherein, [(Q^k)T]-1 represents the inverse mapping matrix.
6. The artistic portrait generation model according to claim 1 that imitates the painting process of real artists, characterized in that: The U-Net attention: uses a cross-attention mechanism to enhance the model's ability to perceive diverse texture details. Compared with using skip connections, the attention mechanism introduces more contextual information in the decoding process and enhances the model's understanding of the overall image structure. The cross-attention query (Query, Q) comes from the encoding layer, while the key (Key, K) and value (Value, V) come from the decoding layer. The features in the decoder are processed through residual connections and passed to the next decoding layer after upsampling.
7. The artistic portrait generation model according to claim 1 that imitates the painting process of real artists, characterized in that: The loss function: The model will be trained with the following loss function: GAN loss: Use adversarial loss to measure whether Y^ is real or fake: Texture loss: In order to make the texture details of the synthesized sketch consistent with the target sketch, a gradient constraint is imposed on it. During calculation, the Sobel operator is used to obtain their gradient information, and then the average cosine distance between them is calculated as the optimization target: Where H and H are the width and height, ||·|| represents the norm of the vector, represents the gradient of Y in the x and y directions, represents the gradient of Y^ in the x and y directions; Reconstruction loss: imposes reconstruction constraints on the synthesized sketch and the target sketch at the pixel level: Total loss: Combine all the above losses to get the total loss: Among them, λ1 and λ2 represent weight hyperparameters.