Face sketch image generation method and system based on spatial adaptive attention

Through the spatially adaptive attention method, the problems of structural distortion and texture blur in face sketch generation are solved, and high-quality face sketch image generation is achieved, which improves the clarity and reality of the image.

CN120259140AActive Publication Date: 2025-07-04JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510745318.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The prior art has problems of structural distortion and texture blur in the generation of face sketches, making it difficult to effectively preserve key facial details and improve the clarity and reality of the generated images.

Method used

Using a spatially adaptive attention method, through dual-modal feature decoupling, dynamic feature fusion and multi-scale adversarial optimization mechanism, the semantic content and sketch style features of face images are explicitly separated, and a spatially adaptive attention mechanism is introduced to dynamically align the cross-modal feature distribution, combining multi-layer reflection filling and convolution and upsampling to gradually restore image resolution.

Benefits of technology

It significantly improves the clarity and reality of the generated sketch images, effectively alleviates the edge artifacts and information interference caused by modal differences, enhances the ability to retain key facial details and the global structural consistency of the generated images, and improves the perceived quality of the generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259140A_ABST
    Figure CN120259140A_ABST
Patent Text Reader

Abstract

The invention provides a face sketch image generation method and system based on spatial adaptive attention. The method comprises the following steps: obtaining content features and style features through a face photo image and a face sketch image; respectively inputting the content features and the style features into the content branches and the style branches to respectively obtain depth content features and depth style features; performing adaptive weighted fusion on the depth content features and the depth style features by using a space attention mechanism to obtain adaptive weighted features; obtaining a reconstructed face sketch image by using the self-adaptive weighted features; optimizing the model by using the reconstructed face sketch image to obtain an optimized model; and obtaining a face sketch image by using the optimized model. According to the method, the row vector information and the column vector information are adaptively aggregated on the multi-scale semantic level, so that key facial details such as five sense organ contours and hairline textures are reserved, and the global structure consistency of the generated image is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image style transfer, and particularly to a method and system for generating face sketch images based on spatial adaptive attention. Background Art

[0002] The face sketch image generation technology has shown important application value in the fields of digital entertainment, judicial criminal investigation, and medical diagnosis by converting face photo images into structured sketch images. In the field of digital entertainment, this technology can quickly generate personalized artistic images and lower the user's creation threshold; in judicial criminal investigation, it solves the matching problem of low-quality surveillance images and database photos through cross-modal matching and improves the efficiency of suspect identification; in medical diagnosis, the sketch form can enhance the visual expression of pathological features and assist doctors in accurate analysis. In the prior art, the research methods for face sketch generation are mainly divided into two categories: one is the traditional method, such as the edge detection technology based on filters; the other is the deep learning method, including the method based on convolutional neural network (CNN), the method based on Transformer, and the method based on diffusion model.

[0003] Traditional methods mostly rely on manually designed feature extraction and edge detection algorithms, which have the advantages of simple implementation and high computational efficiency, but have obvious deficiencies in dealing with complex backgrounds, maintaining edge continuity, and retaining detailed information. Summary of the Invention

[0004] In view of the above situation, the main purpose of the present invention is to propose a method and system for generating face sketch images based on spatial adaptive attention to solve the above technical problems.

[0005] The present invention proposes a method for generating face sketch images based on spatial adaptive attention, and the method includes the following steps: Step 1: After inputting the face photo image and the face sketch image into the face sketch image generation network, both are sequentially subjected to cropping and blocking and convolutional processing to obtain initial content features and initial style features respectively, and the initial content features and initial style features are concatenated with the corresponding learnable position embeddings to obtain content features and style features; Step 2: Perform a cascaded max pooling operation and a large kernel convolution extraction operation on the content features to obtain the content features output after large kernel convolution processing; Perform a multi-level average pooling strategy and a small kernel convolution extraction operation on the style features to obtain the style features output after small kernel convolution processing; Spatially normalize the content features output after large-kernel convolution processing and the style features output after small-kernel convolution processing to obtain attention weights; Using the attention weights, spatially align the attention map with the original feature map through bilinear interpolation, and obtain the spatially aligned attention weights; Using the spatially aligned attention weights, dynamically calibrate the output content features and output style features in a weighted manner to obtain depth features; Step 3: Use the spatial attention mechanism to adaptively weight and fuse the depth content features and depth style features to obtain the adaptively weighted features. At the same time, perform a linear projection operation on the depth style features alone to obtain the aligned style features; Using the adaptively weighted features and the aligned style features, process them through the gating interaction mechanism in the row and column dimensions to obtain the fused features; Step 4: Continuously upsample and decode the fused features to obtain the reconstructed face sketch image; Step 5: Use the reconstructed face sketch image, the face photo image, and the face sketch image to construct a loss function; Optimize the face sketch image generation network in combination with the loss function to obtain the trained face sketch image generation network; Input the face photo image and the face sketch image into the trained face sketch image generation network to generate the face sketch painting image.

[0006] The present invention also proposes a face sketch image generation system based on spatial adaptive attention, and the system includes: A feature extraction module for: Input the face photo image and the face sketch image into the face sketch image generation network, and both are sequentially subjected to cropping and block division and convolution processing to respectively obtain the initial content features and the initial style features. Concatenate the initial content features and the initial style features with the corresponding learnable position embeddings to obtain the content features and the style features; A bimodal feature decoupling module for: Perform a cascaded max-pooling operation and a large-kernel convolution extraction operation on the content features to obtain the content features output after large-kernel convolution processing; Adopt a multi-level average pooling strategy and a small-kernel convolution extraction operation on the style features to obtain the style features output after small-kernel convolution processing; Spatially normalize the content features output after large-kernel convolution processing and the style features output after small-kernel convolution processing to obtain attention weights; Using the attention weights, spatially align the attention map with the original feature map through bilinear interpolation, and obtain the spatially aligned attention weights; Using the attention weights after spatial alignment, dynamically calibrate the output content features and output style features in a weighted manner to obtain deep features; A row-column vector dynamic fusion module, configured to: Use the spatial attention mechanism to adaptively weight and fuse the deep content features and deep style features to obtain the adaptively weighted features. At the same time, perform a linear projection operation on the deep style features alone to obtain the aligned style features; Use the adaptively weighted features and the aligned style features, and process them through the gating interaction mechanism in the row and column dimensions to obtain the fused features; An upsampling decoding module, configured to: Perform continuous upsampling decoding processing on the fused features to obtain the reconstructed face sketch image; A training module, configured to: Use the reconstructed face sketch image, face photo image, and face sketch image to construct a loss function; Combine the loss function to optimize the face sketch image generation network to obtain the trained face sketch image generation network; A result output module, configured to: Input the face photo image and the face sketch image into the trained face sketch image generation network to generate a face sketch painting image.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention systematically solves the problems of structural distortion and texture blur existing in traditional cross-modal image generation by integrating the dual-modal feature decoupling, dynamic feature fusion, and multi-scale adversarial optimization mechanisms, and significantly improves the clarity and realism of the generated sketch images; 2. The present invention explicitly separates the semantic content and sketch style features of the face image through a parallel branch structure, and introduces a spatial adaptive attention mechanism to dynamically align the cross-modal feature distributions, effectively alleviating the edge artifacts and information interference caused by the modal differences, and improving the edge retention ability; 3. The present invention adaptively aggregates the row vector and column vector information at the multi-scale semantic level, not only retaining the key facial details such as the facial feature contours and hair textures, but also enhancing the global structural consistency of the generated images; 4. The present invention adopts a combination of multi-layer reflection filling, convolution, and upsampling to gradually restore the image resolution, effectively retaining the edge and texture details, and improving the structure restoration ability and the detail quality of the generated images; 5. The present invention constructs a multi-granularity adversarial supervision mechanism by parallelly deploying a global discriminator and a local discriminator, while strengthening the authenticity of the image details, maintaining the reasonable proportion and consistency of the facial contour structure, and comprehensively improving the perceptual quality of the generated images.

[0008] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a flowchart of a method for generating a face sketch image based on spatial adaptive attention proposed by the present invention; Figure 2 is a schematic diagram of the dynamic fusion of column and row vectors of a method for generating a face sketch image based on spatial adaptive attention proposed by the present invention; Figure 3 is a schematic diagram of the overall framework of a system for generating a face sketch image based on spatial adaptive attention proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0011] These and other aspects of the embodiments of the present invention will be clear from the following description and the drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0012] Please refer to Figure 1 , an embodiment of the present invention proposes a method for generating a face sketch image based on spatial adaptive attention, and the method includes the following steps: Step 1: After inputting the face photo image and the face sketch image into the face sketch image generation network, both are successively subjected to cropping and partitioning and convolutional processing to respectively obtain an initial content feature and an initial style feature, and the initial content feature and the initial style feature are concatenated with corresponding learnable position embeddings to obtain a content feature and a style feature; Further, in this step, the specific steps for processing the face photo image and the face sketch image are as follows: The input face photo image and face sketch image are cropped into images with a size of 256×256 and a channel number of 3; The cropped images are divided into multiple image blocks with a size of 8×8; Each image block is subjected to a convolution operation with a convolution kernel size of 7×7 and then embedded into a 512-dimensional high-dimensional feature space; Finally, a learnable positional embedding is assigned to each image patch, and the positional embedding is a 256-dimensional vector.

[0013] Step 2: Perform a cascaded max pooling operation and a large kernel convolution extraction operation on the content features to obtain the content features output after the large kernel convolution; Perform a multi-level average pooling strategy and a small kernel convolution extraction operation on the style features to obtain the style features output after the small kernel convolution; Perform spatial normalization on the content features output after the large kernel convolution and the style features output after the small kernel convolution to obtain the attention weights; Using the attention weights, spatially align the attention map with the original feature map through bilinear interpolation and obtain the spatially aligned attention weights; Using the spatially aligned attention weights, dynamically calibrate the output content features and output style features in a weighted manner to obtain the depth features; In Step 2, perform a cascaded max pooling operation and a large kernel convolution extraction operation on the content features to obtain the content features output after the large kernel convolution. The specific steps are as follows: Perform a cascaded max pooling operation on the content features to obtain the output content features. The relationship formula for the corresponding process is: ; where represents the output content features, represents the max pooling process with a pooling kernel size of 2×2, represents the max pooling process with a pooling kernel size of 3×3, represents the content features; Use alternately stacked 13×13 large kernel convolutions to process the output content features to obtain the content features output after the large kernel convolution. The relationship formula for the corresponding process is: ; where represents the content features output after the large kernel convolution, represents the ReLU activation function, represents the process of using alternately stacked 13×13 large kernel convolutions for processing; Perform a multi-level average pooling strategy and a small kernel convolution extraction operation on the style features to obtain the style features output after the small kernel convolution. The specific steps are as follows: Perform a multi-level average pooling strategy on the style features to obtain the output style features. The relationship formula for the corresponding process is: ; where Indicates the output style feature, Indicates average pooling with a pooling kernel size of 3×3, Indicates average pooling with a pooling kernel size of 5×5, Indicates the style feature; Use small kernel convolution to replenish the local structure and detailed texture of the output style feature to obtain the style feature output after small kernel convolution processing. The relational formula for the corresponding process is: ; Among them, Indicates the style feature output after small kernel convolution processing, Indicates processing using alternately stacked 3×3 small kernel convolutions; Perform spatial normalization on the content feature output after large kernel convolution processing and the style feature output after small kernel convolution processing to obtain the attention weight. The relational formula for the corresponding process is: ; Among them, Indicates the position The attention weight of the content feature map at the position, Indicates the position The attention weight of the style feature map at the position, Indicates the position The content feature output after large kernel convolution processing at the position, Indicates the position The style feature output after small kernel convolution processing at the position, Indicates a small constant to prevent the denominator from being zero; Use the attention weight to spatially align the attention map with the original feature map through bilinear interpolation and obtain the spatially aligned attention weight. The relational formula for the corresponding process is: ; Among them, Indicates the content attention weight after spatial alignment, Indicates the style attention weight after spatial alignment, Indicates the attention weight of the content feature map, Indicates the attention weight of the style feature map, Indicates processing through a hyperbolic linear interpolation function, Indicates specifying the length and width after upsampling and using the spatially aligned attention weight to dynamically calibrate the output content feature and the output style feature in a weighted manner to obtain the depth feature. The relational formula for the corresponding process is: ; Among them, Represents the depth content feature, Represents the depth style feature, Represents element-wise multiplication; Furthermore, in this step, a collaborative optimization framework for cross-modal feature decoupling is constructed through heterogeneous modality-driven feature extraction and dynamic spatial attention calibration, for separating and aligning the face content feature and the sketch style feature.

[0014] Step 3: Use the spatial attention mechanism to adaptively weight and fuse the depth content feature and the depth style feature to obtain an adaptively weighted feature. At the same time, perform a linear projection operation on the depth style feature alone to obtain an aligned style feature; Use the adaptively weighted feature and the aligned style feature, and process them through the gating interaction mechanism in the row and column dimensions to obtain a fused feature; Please refer to Figure 2 , in Step 3, use the spatial attention mechanism to adaptively weight and fuse the depth content feature and the depth style feature to obtain an adaptively weighted feature. At the same time, perform a linear projection operation on the depth style feature alone to obtain an aligned style feature. The relational expressions for the corresponding process are: ; Among them, Represents the adaptively weighted feature, Represents the aligned style feature, Represents the linear projection operation, Represents the self-attention operation, Represents the channel dimension concatenation operation; Use the adaptively weighted feature and the aligned style feature, and through the processing of the gating interaction mechanism in the row and column dimensions, obtain a fused feature. The specific steps are as follows: Construct multi-norm attention branches for the adaptively weighted feature and the aligned style feature along the row and column vector directions respectively, and obtain the row vector attention score of the adaptively weighted feature, the column vector attention score of the adaptively weighted feature, the row vector attention score of the aligned style feature, and the column vector attention score of the aligned style feature. The relational expressions for the corresponding process are: ; Among them, Represents the total number of rows or columns of the feature, Represents the row vector attention score of the adaptively weighted feature, Represents the column vector attention score of the adaptively weighted feature, Represents the row vector attention score of the aligned style feature, Represents the column vector attention score of the aligned style feature, Indicates that it has undergone L1 norm processing, Indicates that it has undergone L2 norm processing, Indicates L1 norm, Indicates L2 norm, Indicates the maximum norm; Using an exponential-enhanced normalization strategy, non-linearly amplify and probabilistically constrain the row and column vector attention scores of the features to obtain the row vector attention score of the normalized adaptive weighted feature, the column vector attention score of the normalized adaptive weighted feature, the row vector attention score of the normalized aligned style feature, and the column vector attention score of the normalized aligned style feature. The relational expressions for the corresponding processes are: ; Among them, Indicates that it has undergone exponential function processing, Indicates the row vector attention score of the normalized adaptive weighted feature, Indicates the column vector attention score of the normalized adaptive weighted feature, Indicates the row vector attention score of the normalized aligned style feature, Indicates the column vector attention score of the normalized aligned style feature; Using the row and column vector attention scores of the normalized features to perform weighted fusion on the adaptive weighted feature and the aligned style feature to obtain the row vector fusion feature and the column vector fusion feature respectively. The relational expressions for the corresponding processes are: ; Among them, Indicates the row vector fusion feature, Indicates the column vector fusion feature; Fuse the row vector fusion feature and the column vector fusion feature to obtain the fusion feature. The relational expressions for the corresponding processes are: ; Among them, Indicates the fusion feature.

[0015] Step 4: Process the fusion feature through continuous upsampling decoding to obtain the reconstructed face sketch image; In this step, the processing of the fusion feature is as follows: Perform reflection padding on the fusion feature to preserve edge information; Use a convolutional layer to reduce the number of channels from 512 dimensions to 256 dimensions and enhance the non-linear expression through the ReLU activation function; Use the nearest neighbor interpolation method for the first upsampling to expand the spatial size; The feature expression ability is further enhanced through multiple repeated reflection padding and convolution operations, which include three consecutive convolution modules, each of which consists of reflection padding, 3×3 convolution, and ReLU activation function; Reduce the channel dimension from 256 to 128 and perform the second upsampling operation; Repeat the above operation, further reduce the channel dimension from 128 to 64, and perform the third upsampling; Finally, map the channel dimension from 64 to 3 channels of the output image through convolution to obtain the reconstructed face sketch image.

[0016] Step 5: Construct a loss function using the reconstructed face sketch image, face photo image, and face sketch image; Optimize the face sketch image generation network in combination with the loss function to obtain the trained face sketch image generation network; Input the face photo image and face sketch image into the trained face sketch image generation network to generate a face sketch painting image; In step 5, construct a loss function using the reconstructed face sketch image, face photo image, and face sketch image. Among them, the loss function includes a content loss function, a style loss function, a detail loss function, an identity loss function, and an adversarial loss function. Content loss function: ; Among them, represents the content loss function, represents the reconstructed face sketch image, represents the face photo image, represents the feature extraction operation using a deep convolutional neural network; Style loss function: ; Among them, represents the style loss function, represents the face sketch image, represents calculating the channel mean of the feature map, represents calculating the standard deviation of the feature map; Detail loss function: ; Among them, represents the detail loss function, represents the mask of the edge region, represents being processed by a Gaussian filter, represents passing through a Laplacian operator; Identity loss function: ; Among them, and both represent the identity loss function, and are processed by extraction through the th layer of the deep convolutional neural network; Adversarial loss function: ; Among them, represents the adversarial loss function, represents the number of reconstructed face sketch images, represents the th discriminator, represents the face sketch image generated from the th face photo image; Total loss function: ; Among them, represents the total loss function, represents the content loss weight, represents the style loss weight, represents the detail loss weight, and both represent the identity loss weight, represents the adversarial loss weight; Furthermore, in this step, the generative adversarial network framework is adopted, and adversarial training is carried out through the model architectures of the generator and the multi-scale perception discriminator. The multi-scale perception discriminator includes the following steps: The discriminator forms a multi-scale structure perception ability by introducing multiple discriminant sub-networks, and each sub-network acts on face image inputs of different scales respectively; downsampling is performed between each scale by means of average pooling, thereby constructing a discriminant pyramid; Each discriminant sub-network adopts a convolutional neural network based on the PatchGAN structure, which is composed of multiple layers of convolutional layers, non-linear activation layers and normalization layers, including: The initial convolutional layer maps the input image from the input channel to the basic channel dimension, followed by the LeakyReLU activation function; Multiple layers of residual convolutional modules, where each layer doubles the number of channels through convolutional operations, with a maximum of 512 dimensions, and feature regularization is combined with instance normalization (InstanceNorm); The last layer outputs a single-channel feature map for discriminating the authenticity of the image patch; Supports outputting discriminant feature maps of intermediate layers for multi-level feature supervision, thereby enhancing the sensitivity to local texture and structural changes; The network model parameters are initialized with a Gaussian distribution and support GPU-accelerated deployment and multi-scale distributed parallel training; The results of the discriminant sub-networks will be combined and used to provide discriminant feedback signals from different scales to the generator, thereby guiding it to generate face sketch images with a sense of reality in both the overall structure and local texture.

[0017] Furthermore, calculate metrics such as SSIM for the generated face sketch images and real sketches, and finally complete the quality evaluation of the face sketch images.

[0018] Please refer to Figure 3 , this embodiment of the present invention also provides a face sketch image generation system based on spatial adaptive attention, and the system includes: A feature extraction module, which is used for: After inputting the face photo image and the face sketch image into the face sketch image generation network, both are successively subjected to cropping and blocking and convolutional processing to respectively obtain initial content features and initial style features, and the initial content features and initial style features are concatenated with corresponding learnable position embeddings to obtain content features and style features; A dual-modal feature decoupling module, which is used for: Perform a cascaded max pooling operation and a large kernel convolution extraction operation on the content features to obtain the content features output after large kernel convolution processing; Perform a multi-level average pooling strategy and a small kernel convolution extraction operation on the style features to obtain the style features output after small kernel convolution processing; Perform spatial normalization processing on the content features output after large kernel convolution processing and the style features output after small kernel convolution processing to obtain attention weights; Use the attention weights to spatially align the attention map with the original feature map through bilinear interpolation and obtain the spatially aligned attention weights; Use the spatially aligned attention weights to dynamically calibrate the output content features and output style features in a weighted manner to obtain depth features; A row-column vector dynamic fusion module, which is used for: Use the spatial attention mechanism to adaptively weight and fuse the depth content features and depth style features to obtain adaptively weighted features. At the same time, perform a linear projection operation on the depth style features alone to obtain aligned style features; Use the adaptively weighted features and the aligned style features to process through a gating interaction mechanism in the row and column dimensions to obtain fused features; An upsampling decoding module, which is used for: Perform continuous upsampling decoding processing on the fused features to obtain the reconstructed face sketch image; A training module, configured to: Construct a loss function by using the reconstructed face sketch image, face photo image, and face sketch image; Optimize the face sketch image generation network in combination with the loss function to obtain a trained face sketch image generation network; A result output module, configured to: Input the face photo image and the face sketch image into the trained face sketch image generation network to generate a face sketch painting image.

[0019] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0020] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0021] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A method for generating face sketch images based on spatial adaptive attention, characterized in that, The method includes the following steps: Step 1: After inputting the face photo image and the face sketch image into the face sketch image generation network, both are successively subjected to cropping and blocking as well as convolutional processing to respectively obtain the initial content feature and the initial style feature. The initial content feature and the initial style feature are concatenated with the corresponding learnable position embeddings to obtain the content feature and the style feature; Step 2: Perform a cascaded max pooling operation and a large kernel convolution extraction operation on the content feature to obtain the content feature output after large kernel convolution processing; Perform a multi-level average pooling strategy and a small kernel convolution extraction operation on the style feature to obtain the style feature output after small kernel convolution processing; Perform spatial normalization processing on the content feature output after large kernel convolution processing and the style feature output after small kernel convolution processing to obtain the attention weight; Using the attention weight, spatially align the attention map with the original feature map through bilinear interpolation and obtain the spatially aligned attention weight; Using the spatially aligned attention weight, dynamically calibrate the output content feature and the output style feature in a weighted manner to obtain the depth feature; Step 3: Use the spatial attention mechanism to adaptively weight and fuse the depth content feature and the depth style feature to obtain the adaptively weighted feature. At the same time, perform a linear projection operation on the depth style feature alone to obtain the aligned style feature; Using the adaptively weighted feature and the aligned style feature, process through the gating interaction mechanism in the row and column dimensions to obtain the fused feature; Step 4: Perform continuous upsampling decoding processing on the fused feature to obtain the reconstructed face sketch image; Step 5: Construct a loss function using the reconstructed face sketch image, the face photo image, and the face sketch image; Optimize the face sketch image generation network in combination with the loss function to obtain the trained face sketch image generation network; Input the face photo image and the face sketch image into the trained face sketch image generation network to generate a face sketch image.

2. The method for generating a face sketch image based on spatial adaptive attention according to claim 1, wherein In the said Step 2, performing a cascaded max pooling operation and a large kernel convolution extraction operation on the content feature to obtain the content feature output after large kernel convolution processing, the specific steps are: Perform a cascaded max pooling operation on the content feature to obtain the output content feature, and the relational expression existing in the corresponding process is: ; Among them, represents the output content feature, represents the maximum pooling process with a pooling kernel size of 2×2, represents the maximum pooling process with a pooling kernel size of 3×3, represents the content feature; Use alternately stacked 13×13 large kernel convolutions to process the output content feature to obtain the content feature output after large kernel convolution processing, and the relational expression existing in the corresponding process is: ; Among them, represents the content feature output after large kernel convolution processing, represents the ReLU activation function, represents the processing using alternately stacked 13×13 large kernel convolutions.

3. The method for generating a face sketch image based on spatial adaptive attention according to claim 2, wherein In the said Step 2, performing a multi-level average pooling strategy and a small kernel convolution extraction operation on the style feature to obtain the style feature output after small kernel convolution processing, the specific steps are: Perform a multi-level average pooling strategy on the style feature to obtain the output style feature, and the relational expression existing in the corresponding process is: ; Among them, represents the output style feature, represents average pooling processing with a pooling kernel size of 3×3, represents average pooling processing with a pooling kernel size of 5×5, represents the style feature; Use small kernel convolutions to supplement the local structure and detailed texture of the output style feature to obtain the style feature output after small kernel convolution processing, and the relational expression existing in the corresponding process is: ; Among them, represents the style feature output after the small kernel convolution processing, represents the processing using the alternately stacked 3×3 small kernel convolutions.

4. The method for generating a face sketch image based on spatial adaptive attention according to claim 3, wherein In step 2, the content features output after large kernel convolution processing and the style features output after small kernel convolution processing are subjected to spatial normalization processing to obtain attention weights. The relational expression for the corresponding process is as follows: ; Among them, represents the attention weight of the content feature map at the position , represents the attention weight of the style feature map at the position , represents the content feature output after large kernel convolution processing at the position , represents the style feature output after small kernel convolution processing at the position , represents a small constant to prevent the denominator from being zero.

5. The method for generating a face sketch image based on spatial adaptive attention according to claim 4, wherein, In step 2, using the attention weights, the attention map and the original feature map are spatially aligned through bilinear interpolation, and the attention weights after spatial alignment are obtained. The relational expression for the corresponding process is as follows: ; Among them, represents the content attention weight after spatial alignment, represents the style attention weight after spatial alignment, represents the attention weight of the content feature map, represents the attention weight of the style feature map, represents being processed by the hyperbolic linear interpolation function, represents the specified length and width after upsampling.

6. The method for generating a face sketch image based on spatial adaptive attention according to claim 5, wherein, In step 2, using the attention weights after spatial alignment, the output content features and the output style features are dynamically calibrated in a weighted manner to obtain depth features. The relational expression for the corresponding process is as follows: ; Among them, represents the depth content feature, represents the depth style feature, represents element-wise multiplication.

7. The method for generating a face sketch image based on spatial adaptive attention according to claim 6, wherein, In step 3, the depth content features and the depth style features are adaptively weighted and fused using the spatial attention mechanism to obtain adaptively weighted features. At the same time, a linear projection operation is performed on the depth style features alone to obtain the aligned style features. The relational expression for the corresponding process is as follows: ; Among them, represents the adaptive weighted feature, represents the aligned style feature, represents the linear projection operation, represents the self-attention operation, represents the channel dimension concatenation operation.

8. The method for generating a face sketch image based on spatial adaptive attention according to claim 7, wherein In step 3, using the adaptively weighted features and the aligned style features, they are processed through a gating interaction mechanism in the row and column dimensions to obtain fused features. The specific steps are as follows: Multi-norm attention branches are constructed for the adaptively weighted features and the aligned style features along the row and column vector directions respectively to obtain the row vector attention scores of the adaptively weighted features, the column vector attention scores of the adaptively weighted features, the row vector attention scores of the aligned style features, and the column vector attention scores of the aligned style features. The relational expression for the corresponding process is as follows: ; Among them, represents the total number of rows or columns of features, represents the row vector attention score of the adaptive weighted features, represents the column vector attention score of the adaptive weighted features, represents the row vector attention score of the aligned style features, represents the column vector attention score of the aligned style features, represents being processed by the L1 norm, represents being processed by the L2 norm, represents the L1 norm, represents the L2 norm, represents the maximum norm; Using an exponential enhanced normalization strategy, the row and column vector attention scores of the features are non-linearly amplified and probabilistically constrained to obtain the normalized row vector attention scores of the adaptively weighted features, the normalized column vector attention scores of the adaptively weighted features, the normalized row vector attention scores of the aligned style features, and the normalized column vector attention scores of the adaptively weighted features. The relational expression for the corresponding process is as follows: ; Among them, represents being processed by an exponential function, represents the row vector attention score of the normalized adaptive weighted features, represents the column vector attention score of the normalized adaptive weighted features, represents the row vector attention score of the normalized aligned style features, represents the column vector attention score of the normalized aligned style features; Using the normalized row and column vector attention scores of the features, the adaptively weighted features and the aligned style features are weighted and fused to obtain the row vector fused features and the column vector fused features respectively. The relational expression for the corresponding process is as follows: ; Among them, represents the fused feature of row vectors, represents the fused feature of column vectors; The row vector fused features and the column vector fused features are fused to obtain fused features. The relational expression for the corresponding process is as follows: ; Among them, represents the fusion feature.

9. The method for generating a face sketch image based on spatial adaptive attention according to claim 8, wherein, In step 5, a loss function is constructed using the reconstructed face sketch image, the face photo image, and the face sketch image. The loss function includes a content loss function, a style loss function, a detail loss function, an identity loss function, and an adversarial loss function. The expression of the content loss function is as follows: ; Among them, represents the content loss function, represents the reconstructed face sketch image, represents the face photo image, represents the feature extraction operation using a deep convolutional neural network; The expression of the style loss function is as follows: ; Among them, represents the style loss function, represents the face sketch image, represents calculating the channel mean of the feature map, represents calculating the standard deviation of the feature map; The expression of the detail loss function is as follows: ; Among them, represents the detail loss function, represents the mask of the edge region, represents being processed by a Gaussian filter, represents through the Laplacian operator; The expression of the identity loss function is as follows: ; Among them, and both represent the identity loss function, after being extracted and processed by the layer of the deep convolutional neural network; The expression of the adversarial loss function is as follows: ; Among them, represents the adversarial loss function, represents the number of reconstructed face sketch images, represents the th discriminator, represents the face sketch image generated from the th face photo image; The expression of the total loss function is as follows: ; Among them, represents the total loss function, represents the content loss weight, represents the style loss weight, represents the detail loss weight, and both represent the identity loss weight, represents the adversarial loss weight.

10. A face sketch image generation system based on spatial adaptive attention, characterized in that, The system applies the method for generating a face sketch image based on spatial adaptive attention according to any one of claims 1 to 9. The system includes: A feature extraction module for: After inputting the face photo image and the face sketch image into the face sketch image generation network, they both go through cropping and partitioning, as well as convolutional processing in sequence, to obtain the initial content feature and the initial style feature respectively. The initial content feature and the initial style feature are concatenated with the corresponding learnable position embeddings to obtain the content feature and the style feature; The dual-modal feature decoupling module is used for: Performing a cascaded max pooling operation and a large kernel convolution extraction operation on the content feature to obtain the content feature output after large kernel convolution processing; Performing a multi-level average pooling strategy and a small kernel convolution extraction operation on the style feature to obtain the style feature output after small kernel convolution processing; Performing spatial normalization processing on the content feature output after large kernel convolution processing and the style feature output after small kernel convolution processing to obtain the attention weight; Using the attention weight, spatially aligning the attention map with the original feature map through bilinear interpolation, and obtaining the spatially aligned attention weight; Using the spatially aligned attention weight to dynamically calibrate the output content feature and the output style feature in a weighted manner to obtain the depth feature; The row-column vector dynamic fusion module is used for: Using the spatial attention mechanism to adaptively weight and fuse the depth content feature and the depth style feature to obtain the adaptively weighted feature. At the same time, performing a linear projection operation on the depth style feature alone to obtain the aligned style feature; Using the adaptively weighted feature and the aligned style feature, processing through the gating interaction mechanism in the row and column dimensions to obtain the fused feature; The upsampling decoding module is used for: Performing continuous upsampling decoding processing on the fused feature to obtain the reconstructed face sketch image; The training module is used for: Constructing a loss function using the reconstructed face sketch image, the face photo image, and the face sketch image; Combining the loss function to optimize the face sketch image generation network to obtain the trained face sketch image generation network; The result output module is used for: Inputting the face photo image and the face sketch image into the trained face sketch image generation network to generate the face sketch painting image.

Citation Information

Patent Citations

  • System for searching real face based on hand-painted sketch

    CN114299218A

  • Multi-style face sketch generation method guided by depth information

    CN115457160A

  • Face sketch synthesis method and system based on perception loss

    CN119228938A

  • Method for image motion deblurring, apparatus, electronic device and medium therefor

    US20240404025A1