Personalized portrait-oriented multi-style image conversion method and device, and storage medium
By extracting facial identity feature vectors and applying identity consistency constraints in real time, the problems of identity consistency and flexibility of multi-style switching in personalized portrait style transfer in existing technologies are solved, and high-fidelity conversion of facial structures is achieved in complex style transfer.
Patent Information
- Application Number
- CN202511796273.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-02-24
AI Technical Summary
Existing personalized portrait style transfer technology is insufficient in maintaining identity consistency and flexibility in switching between multiple styles. In particular, facial structure distortion and identity feature drift occur during complex style transfer, and there is a lack of on-demand dynamic positioning and activation mechanisms for switching between multiple styles.
By extracting facial identity feature vectors from source portrait images, combining them with style transfer control parameters, a pre-trained style transfer model is dynamically activated, and identity consistency constraints are applied in real time during image generation to ensure that the facial structure of the person in the generated image is consistent with the source portrait.
It achieves high consistency in complex style migration, supports quick switching between independent or combined styles, reduces the risk of identity distortion, and improves the ability to adapt to multiple styles on demand.
Smart Images

Figure CN121563754A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device and storage medium for multi-style image conversion for personalized portraits. Background Technology
[0002] Personalized portrait style transfer technology aims to apply a user-specified artistic style (such as oil painting, watercolor, comics, etc.) to portrait images while preserving the person's identity (ID). This technology is widely used in entertainment applications, digital art creation, and virtual avatar design. Its core challenge lies in balancing the expression of stylistic features with the consistency of identity features during the stylization process—it must fully render the target style while ensuring that the facial structure, features, and other biometric characteristics of the person in the generated image are highly consistent with the original image, avoiding the loss or distortion of identity information.
[0003] Current mainstream personalized portrait style transfer technologies mainly rely on generative adversarial networks (GANs) or style transfer networks (e.g., AdaIN, StyleGAN). Although existing technologies can achieve basic style transfer, they still have limitations in personalized portrait scenarios, such as insufficient identity consistency and low flexibility in switching between multiple styles. Summary of the Invention
[0004] This application provides a method, device, and storage medium for multi-style image conversion of personalized portraits, which can dynamically constrain the style transfer process and ensure a high degree of consistency in portrait identity features.
[0005] On the one hand, this application provides a multi-style image conversion method for personalized portraits, the method comprising the following steps: S101: Extract and output facial identity feature vectors representing the identity of a specific person from the input source portrait image; S102: Receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information; S103: Generate an image based on a preset image generation framework, and in the image generation framework, apply identity consistency constraints in real time based on the facial identity feature vector and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image. S104: Obtain the intermediate image result obtained after processing in step S103; S105: Perform post-processing optimization on the intermediate image results and output the optimized target portrait image.
[0006] On the other hand, this application provides a multi-style image conversion device for personalized portraits, the device comprising: The extraction module is used to extract and output facial identity feature vectors that represent the identity of a specific person from the input source portrait image; The activation module is used to receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information. The fusion module is used to generate images based on a preset image generation framework. Within the image generation framework, identity consistency constraints are applied in real time based on the facial identity feature vector and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image. The acquisition module is used to acquire the intermediate image results obtained after processing by the fusion module. The output module is used to perform post-processing optimization on the intermediate image results and output the optimized target portrait image.
[0007] Thirdly, this application provides an electronic device, the device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the technical solution of the multi-style image conversion method for personalized portraits described above.
[0008] Fourthly, this application provides a storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described multi-style image conversion method for personalized portraits.
[0009] As can be seen from the technical solution provided in this application, on the one hand, it receives style transfer control parameters provided by the user, locates and activates the corresponding pre-trained style transfer model based on these parameters, thus allowing the user to dynamically activate the matching target style model using simple parameters (e.g., style labels, reference images, etc.), supporting rapid switching between independent or combined styles without retraining the entire network, thereby improving the system's ability to adapt to multiple styles on demand. On the other hand, by dynamically injecting an identity constraint mechanism (e.g., loss function guidance or feature space alignment) into the core style transfer processing stage, it forces the facial structure of the generated image to be consistent with the source portrait, significantly reducing the risk of identity distortion in complex style transfer. Even for abstract and highly deformable styles, it can still maintain the identity features of facial features and facial contours. In summary, the technical solution of this application can dynamically constrain the style transfer process, ensuring a high degree of consistency in portrait identity features. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart of a multi-style image conversion method for personalized portraits provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the multi-style image conversion device for personalized portraits provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] In this specification, adjectives such as "first" and "second" are used only to distinguish one element or action from another, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.
[0014] For ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn to actual scale.
[0015] Personalized portrait style transfer technology aims to apply a user-specified artistic style (such as oil painting, watercolor, comics, etc.) to portrait images while preserving the person's identity (ID). This technology has wide applications in entertainment, digital art creation, and virtual avatar design. Its core challenge lies in balancing style expression with identity consistency during the stylization process—fully rendering the target style while ensuring that the facial structure, features, and other biometric characteristics of the person in the generated image are highly consistent with the original image, avoiding the loss or distortion of identity information. Current mainstream methods mainly rely on Generative Adversarial Networks (GANs) or style transfer networks (e.g., AdaIN, StyleGAN). A typical process includes: extracting image content and style features through a pre-trained model; fusing content and style information in the feature space; generating a stylized image through a decoder; some schemes attempt to introduce facial recognition features as auxiliary input to enhance identity preservation capabilities. Although existing technologies can achieve basic style transfer, the following limitations still exist in personalized portrait scenarios: 1) Insufficient identity consistency, which is manifested in the fact that when transferring complex styles (e.g., abstract, heavily distorted styles), identity constraints are only embedded as static inputs and cannot be dynamically applied to specific constraints during the style transfer process, resulting in facial structure distortion and identity feature drift (e.g., misaligned facial features, deformed facial contours); 2) Low flexibility in switching between multiple styles, which is mainly manifested in the fact that independent models or fixed combination models are usually pre-trained for different styles, and there is a lack of mechanisms for dynamically locating and activating the target style model on demand, making it difficult to support users to interactively select or mix multiple styles in real time.
[0016] To address the aforementioned problems in the existing technology, this application proposes a multi-style image conversion method for personalized portraits, the flowchart of which is attached. Figure 1 As shown, the main steps include S101 to S105, which are detailed below: Step S101: Extract and output facial identity feature vectors representing the identity of a specific person from the input source portrait image.
[0017] The fundamental contradiction in portrait style transfer lies in the antagonism between style statistical characteristics and identity topology. Using traditional shallow features (e.g., color histograms, SIFT descriptors) or general deep features (e.g., VGG convolution activation values) fails to decouple "transferable style features" from "immutable identity features." Since identity feature vectors can be optimized in a metric learning space using a dedicated identity recognition network, identical identity feature vectors are tightly clustered in a high-dimensional space, while different identity feature vectors are significantly separated. In other words, identity feature vectors possess the characteristic of resisting style interference, meaning that even if the stylization process introduces texture distortion or geometric deformation, the high-dimensional manifold structure of the identity vectors can still maintain stability. Therefore, in this embodiment, facial identity feature vectors representing a specific person's identity are extracted and output from the input source portrait image. It should be noted that in this embodiment, the facial identity feature vector is a high-dimensional mathematical representation extracted from the portrait face through a deep neural network; its essence is encoding facial biometrics (e.g., geometric relationships of facial features, texture details) into a computable numerical form. The extraction and output of facial identity feature vectors can provide calculable identity benchmark anchors for subsequent dynamic constraints, and establish a mathematically optimizable objective for identity consistency.
[0018] As an embodiment of this application, extracting and outputting facial identity feature vectors representing the identity of a specific person can be achieved through steps S1011 to S1013, as detailed below: Step S1011: Use a pre-trained face detection model to process the source portrait image and locate the face region in the image.
[0019] Specifically, step S1011 can be implemented as follows: After receiving the input source portrait image, scale it according to a preset target size, and then adjust the aspect ratio of the image through an interpolation algorithm to ensure that all input images are of uniform size, thereby completing the normalization of the source portrait image and obtaining a normalized portrait image; perform pixel value offset and scaling operations on each color channel of the normalized portrait image and adjust the pixel distribution range based on the statistical values of the standard dataset to eliminate the influence of illumination differences and obtain a color distribution standardized portrait image; input the color distribution standardized portrait image into a pre-trained face detection model, which analyzes the image content and outputs several candidate regions and their position confidence scores; filter high-confidence regions, i.e., only retain candidate regions with confidence scores exceeding a preset threshold, and when there are multiple candidate regions, select the region with the highest confidence score; convert the relative coordinate parameters output by the face detection model into absolute coordinate values in the image pixel coordinate system, and calculate and determine the final bounding box through the center point coordinates and aspect ratio; finally, extract the face region from the source portrait image based on the final bounding box coordinates.
[0020] In the above implementation of step S1011, adjusting the aspect ratio of the image by interpolation algorithm can be as follows: preset a target size (e.g., fixed width and height values) and calculate the difference between the current aspect ratio of the source portrait image and the target aspect ratio based on the target size; select an interpolation algorithm (e.g., bilinear interpolation or bicubic interpolation) according to the difference between the current aspect ratio and the target aspect ratio, perform a scaling operation on the source portrait image, and adjust the aspect ratio of the image to the target ratio; output the normalized portrait image after aspect ratio adjustment.
[0021] In the implementation of step S1011 above, performing pixel value offset and scaling operations on each color channel of the normalized portrait image and adjusting the pixel distribution range based on standard dataset statistics can be as follows: obtaining the pixel value matrix of each color channel (e.g., RGB channels) of the normalized portrait image; and adjusting the pixel distribution range based on pre-stored standard dataset statistics (e.g., the mean of the ImageNet dataset). μ and standard deviation σ For each color channel, an offset operation is performed on the pixel value, that is, the pixel value is subtracted from the mean offset of the corresponding channel; a scaling operation is performed on the offset pixel value, that is, the pixel value is divided by the standard deviation scaling factor of the corresponding channel, so that the pixel distribution range is adjusted to the standard distribution; the output is a color distribution standardized portrait image, where the pixel value distribution is consistent with the standard dataset.
[0022] In the implementation of step S1011 above, the face detection model is trained using a large-scale face dataset to train a deep learning model (e.g., YOLO, SSD, or MTCNN). The training process includes inputting a face image, labeling the ground truth bounding box, optimizing the model parameters through a loss function, until the model can accurately output the candidate bounding boxes and confidence scores of the face region.
[0023] In the implementation of step S1011 above, converting the relative coordinate parameters output by the face detection model into absolute coordinate values in the image pixel coordinate system, and determining the final bounding box by calculating the center point coordinates and aspect ratio, can be done as follows: receiving the relative coordinate parameters output by the face detection model (e.g., normalized center point coordinates (x_rel, y_rel) and aspect ratio (w_rel, h_rel)); based on the original dimensions (width W and height H) of the source portrait image, converting the relative coordinate parameters into absolute coordinate values, i.e.: absolute center point coordinates (x_abs, y_abs): x_abs = x_rel × W, y_abs = y_rel × H, absolute width w_abs = w_rel × W, and absolute height h_abs = h_rel × H; calculating the upper left corner coordinates (x_min, y_abs) of the bounding box based on the absolute center point coordinates (x_abs, y_abs) and the absolute width and height values (w_abs, h_abs). The coordinates of the lower right corner (x_min) and the lower right corner (x_max, y_max) are calculated as follows: x_min = x_abs - w_abs / 2, y_min = y_abs - h_abs / 2, x_max = x_abs + w_abs / 2, y_max = y_abs + h_abs / 2; based on the calculated coordinate values, the final bounding box is determined, and the face region is extracted from it.
[0024] Step S1012: Input the located face region into the pre-trained identity feature extraction network, wherein the identity feature extraction network is constructed based on a deep learning model.
[0025] The aforementioned identity feature extraction network, built upon a deep learning model, employs a deep convolutional neural network (e.g., ResNet or ArcFace) as its backbone architecture. The network contains multiple convolutional layers, pooling layers, and fully connected layers, and is pre-trained on a large face recognition dataset. The training objective is to minimize the intra-class distance and maximize the inter-class distance between identity feature vectors, enabling the network to extract high-level semantic features (e.g., facial topological relationships). The network's end-point encoding features are fixed-dimensional vectors, and their length is standardized to ensure they can be used for similarity measurement. The core challenge of portrait style transfer lies in the fact that pose differences in the input image (such as side profile and pitch angle) can lead to distortion in identity feature extraction, specifically manifested as spatial deformation interference and inaccurate biometric features. On the one hand, facial landmark detection can establish a biological reference coordinate system (e.g., the distance between the inner canthi of the eyes, the position of the tip of the nose), and affine transformations can remove non-identity-related perspective distortions and eliminate geometric distortions caused by head rotation / tilt. On the other hand, normalized alignment ensures that the identity feature extraction network always operates under uniform spatial parameters, guaranteeing that the extracted high-dimensional vectors encode only essential identity features (e.g., the topological relationships of facial features) without coupling with pose variables. Therefore, before inputting the located face region into the pre-trained identity feature extraction network, landmark detection can be performed on the located face region to obtain the positional information of multiple facial landmarks. Then, based on the positional information of the facial landmarks, normalized alignment processing is performed on the face region to generate a normalized facial image. In this way, the input to the pre-trained identity feature extraction network is a normalized facial image, which significantly enhances the stability of the similarity of the same person in different poses, ensuring that the stylization process maintains the precise consistency of facial feature positions.
[0026] In the above embodiments, the standardization alignment process for the face region based on the location information of facial key points can be as follows: key point detection is performed on the located face region to obtain the location information of multiple facial key points (e.g., the coordinates of the inner canthi of both eyes, the tip of the nose, and the corners of the mouth); based on a predefined standard facial template (e.g., an average face model), affine transformation parameters between the current facial key points and the template key points are calculated, including rotation, scaling, and translation matrices, etc.; the affine transformation parameters are applied to perform geometric transformations (e.g., affine transformation or perspective transformation) on the face region to align the key points to the standard positions; and the aligned and normalized facial image is output, wherein the positions of the facial features are consistent with the standard facial template.
[0027] Step S1013: The identity feature extraction network extracts high-level semantic features of the face region and encodes them into a fixed-dimensional vector as the facial identity feature vector.
[0028] Specifically, step S1013 can be implemented as follows: the identity feature extraction network gradually abstracts high-level semantic information of the face region through cascaded feature extraction modules. That is, the bottom-level module identifies basic visual patterns (e.g., edges, textures), and its output serves as the input of the middle-level module. The middle-level module combines local features (e.g., facial features), and its output serves as the input of the high-level module. The high-level module establishes a global structural representation (e.g., facial topological relationships). Then, the final feature map obtained by the high-level module of the identity feature extraction network is compressed into a single-point vector through spatial pooling, so that the dimension of the single-point vector is consistent with the predefined feature space structure. In the third step, a length standardization operation is performed on the original feature vector, and all dimensional values are scaled proportionally to the unit length range. Finally, a floating-point vector with a fixed dimension is generated as the identity feature carrier, i.e., the facial identity feature vector. The distance relationship of the facial identity feature vector in mathematical space reflects the similarity of facial biometric features.
[0029] In the above implementation of S1013, the predefined feature space refers to the output vector dimension structure (e.g., a 512-dimensional Euclidean space) that is pre-set during network design. This feature space is optimized through the training process so that feature vectors of the same identity are close together and feature vectors of different identities are far apart. The dimension of the single-point vector is consistent with it, ensuring the computability and comparability of features.
[0030] In the above implementation of S1013, the original feature vector refers to the unstandardized initial vector obtained by the spatial pooling operation after the output of the high-level module of the identity feature extraction network. Performing a length standardization operation on the original feature vector, scaling all dimensional values to a unit length range, can be done by: calculating the L2 norm of the original feature vector (i.e., the square root of the sum of the squares of all dimensional values of the vector); dividing each dimensional value of the original feature vector by the L2 norm to obtain the scaled value; and outputting the length-standardized vector with a norm of 1, where all dimensional values are scaled to a unit length range.
[0031] The above embodiment generates a floating-point vector with a fixed dimension as an identity feature carrier, namely a facial identity feature vector. The distance relationship of the facial identity feature vector in mathematical space reflects the similarity of facial biometric features. The fixed dimension refers to the predefined feature space dimension (e.g., 512 dimensions), which is set during network design and is consistent with the predefined feature space structure. The floating-point vector is a mathematical vector composed of real numbers, representing facial identity features. It is obtained by the following steps: the identity feature extraction network extracts features → spatial pooling to obtain the original feature vector → length standardization to generate the vector.
[0032] Step S102: Receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information.
[0033] Using a single model to handle all styles can lead to problems such as style confusion (i.e. the model has difficulty distinguishing between semantically similar styles, resulting in outputs that deviate from user expectations) and catastrophic forgetting. The flexibility of multi-style support depends on the decoupling capability of the model architecture. To map control parameters to style vector space or model index, establish a bridge from discrete style symbols to continuous representation, enable model weights to dynamically adapt to the target style distribution, and avoid parameter interference between multiple styles, this embodiment can receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer parameters to provide style information. The user-provided style transfer control parameters are obtained through one of the following methods: receiving style tags or identifiers selected from a preset style list by the user through a graphical user interface (GUI); receiving style parameter codes contained in a request called by the user through an application programming interface (API); or receiving a reference style image uploaded by the user and performing style feature analysis on the reference style image to parse out the corresponding style transfer control parameters. Specifically, performing style feature analysis on the reference style image to parse out the corresponding style transfer control parameters can be: inputting the reference style image into a pre-trained general style feature extractor for processing; the general style feature extractor outputting a style feature vector or style code representing the main style features of the input image; and mapping the style feature vector or style code onto a predefined multi-style model index to generate style transfer control parameters.
[0034] As one embodiment of this application, locating and activating the corresponding pre-trained style transfer model based on style transfer control parameters can specifically involve: querying a pre-stored multi-style model index based on the style transfer control parameters; determining a specific pre-trained style transfer model matching the style transfer control parameters based on the multi-style model index; and loading the structure and weight parameters of the specific pre-trained style transfer model into computational memory. It should be noted that the pre-trained style transfer model is obtained by training a style transfer model. One way to train a style transfer model can be: constructing a dataset containing images of multiple styles; designing and training a basic generative model to enable it to learn to separate content from style; introducing an identity consistency constraint loss term during training to force the model to retain the identity features of the source portrait when generating images of a specified style; and iteratively optimizing the model parameters until convergence.
[0035] In the above specific implementation method of locating and activating the corresponding pre-trained style transfer model based on style transfer control parameters, the pre-stored multi-style model index is a pre-stored mapping table. Its function is to quickly locate and activate a specific style model. The multi-style model index mainly includes style labels or style codes, corresponding pre-trained style transfer model identifiers (e.g., model file path or unique ID), and optional parameters (e.g., fusion weights), etc.
[0036] In the specific implementation of locating and activating the corresponding pre-trained style transfer model based on the style transfer control parameters described above, determining the specific pre-trained style transfer model that matches the style transfer control parameters according to the multi-style model index can be as follows: receiving style transfer control parameters provided by the user, such as style labels or codes, etc.; querying the pre-stored multi-style model index and matching the control parameters with the entries in the index, including string matching or code mapping, etc.; determining the corresponding specific pre-trained style transfer model based on the matching result (e.g., obtaining the model path); and loading the structure and weight parameters of the model into memory.
[0037] In the above-mentioned specific implementation of locating and activating the corresponding pre-trained style transfer model based on style transfer control parameters, the basic generative model is usually a generative adversarial network (GAN) or variational autoencoder (VAE), which includes an encoder-decoder structure. Designing and training the basic generative model to enable it to learn to separate content from style can be achieved by using a dataset of images with multiple styles, constraining it through loss functions (e.g., content loss, style loss), so that the encoder extracts content features, i.e., the structural information of the image, such as facial contours, and the decoder renders style features such as texture and color.
[0038] It should be noted that, in the embodiments of this application, the binding of the model to a specific style can be either an independent model binding method or a parameterized control method. The former means that each pre-trained style transfer model is trained for one style, and mapping is done directly through a multi-style model index. The latter means that a single model supports multiple styles, and specific style branches are activated by controlling parameters (e.g., style vectors). A multi-style dataset is used during training, but the model is associated with styles through indexes or parameters to ensure activation on demand.
[0039] In multi-style fusion technology, inherent conflicts exist in feature space, such as style non-orthogonality and identity compatibility breaks. The former manifests as, for example, "watercolor" and "ink painting" sharing wet brushstrokes, but conflicting requirements for saturation and white space. The latter manifests as, if style vectors are simply averaged, watercolor blurring at facial edges and ink hardening of internal lines, disrupting identity continuity. To avoid these style conflicts tearing apart the identity structure and to ensure that the blended features remain topologically compatible with the face, style elements can be decoupled and recombined in layers. Based on this, in the above embodiment, the style information provided by the pre-trained style transfer model can be: extracting style feature representations from multiple activated pre-trained style transfer models; and fusing multiple style feature representations according to preset or user-specified fusion weights or rules to generate a mixed style feature as style information. In this way, the freedom of artistic expression is expanded, and the application of mixed style features under constraints avoids style rendering damaging the facial structure. Extracting style feature representations from multiple activated pre-trained style transfer models can be achieved by: activating multiple pre-trained style transfer models based on style transfer control parameters, such as loading the pre-trained style transfer models into memory; for each activated pre-trained style transfer model, inputting style reference data, such as style images or parameters, etc., and extracting the output of its internal style-related layers as style feature representations; and outputting the style feature representations of each model for subsequent fusion.
[0040] Step S103: Generate an image based on a preset image generation framework, and in the image generation framework, apply identity consistency constraints in real time based on facial identity feature vectors and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image.
[0041] On the one hand, facial identity feature vectors and style information belong to heterogeneous spaces and require spatial alignment through specific operators to be understood by the generation framework. On the other hand, the non-affine transformation characteristics of style transfer constitute the fundamental reason for the failure of static constraints. Therefore, in order to dynamically inject identity constraint mechanisms (e.g., loss function guidance or feature space alignment) into the core processing stage of style transfer, and force the facial structure of the generated image to be consistent with the source portrait, thereby reducing the risk of identity distortion in complex style transfer, in this embodiment, image generation can be performed based on a preset image generation framework. In the image generation framework, identity consistency constraints are applied in real time based on facial identity feature vectors and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image. It should be noted that the above-mentioned spatial alignment through specific operators refers to using mathematical transformation operations to map the facial identity feature vector from the identity extraction network and the style information from the style transfer model to the same feature space, thereby eliminating the incompatibility of heterogeneous spaces and enabling the generation framework to uniformly process these two types of features. Specifically, the operators may include linear transformations or attention mechanisms to ensure that the geometric structure of identity and style features is consistent before fusion.
[0042] Considering that a three-level decoupled architecture of encoder-style transfer module-decoder can achieve a balance of contradictions, where the encoder can be used to strip away low-level textures unrelated to identity (e.g., lighting noise) and extract topologically invariant content features (e.g., relative positions of facial features), the style transfer module acts as a decoupling buffer, receiving facial identity feature vectors and style information and performing controlled fusion to avoid identity vectors directly interfering with style parameters, and the decoder reconstructs the image from the decoupled features, ensuring that the output intermediate image satisfies both style intensity and identity topology, the above image generation framework can include an image generator based on generative adversarial networks, which includes an encoder, style transfer module, and decoder. Since the style transfer module receives facial identity feature vectors, style information, and content features from the encoder, and generates decoupled features through controlled fusion, these features contain both stylized textures and retain the identity topology; therefore, the decoupled features are the fused features output by the style transfer module. Specifically, the decoder reconstructs the image from the decoupled features by: receiving the decoupled features output by the style transfer module, which includes fusion information of content, identity, and style; performing spatial resolution restoration on the decoupled features through deconvolutional layers or upsampling layers, gradually expanding the feature map size to the target image size; applying activation functions (e.g., ReLU) and normalization operations on each layer to enhance the nonlinear expression of the features; and generating three-channel image data through the final convolutional layer to output the reconstructed intermediate image result.
[0043] On the other hand, portrait style transfer needs to simultaneously satisfy three conflicting requirements: sufficient style rendering, fidelity of identity structure, and cross-style generalization ability. However, traditional end-to-end generation models suffer from various problems due to feature space coupling. For example, if style rendering is emphasized, the identity structure is excessively distorted (e.g., nose misalignment), or if identity constraints are enforced, stylization is insufficient (e.g., only color transfer), and so on. To solve the above problems, as an embodiment of this application, the above-mentioned real-time application of identity consistency constraints based on facial identity feature vectors and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image can be as follows: the source portrait image is processed by an encoder and content features are extracted; the content features, facial identity feature vectors, and style information provided by the pre-trained style transfer model are input into the style transfer module; in the style transfer module, stylization transformation is performed based on the content features, and identity consistency constraints are applied based on facial identity feature vectors and style information during the transformation process. In the above embodiments, by separating content features through the encoder, the style transfer module can focus on texture and structural transformation, thereby improving the decoupling ability of identity and style. Meanwhile, the decoder reconstructs the image from the normalized constraint features, which also significantly reduces the risk of facial distortion, thus ensuring the high fidelity of the generated results.
[0044] It's important to note that the encoder is able to process source portrait images and extract content features because it's built on a convolutional neural network. Through multiple layers of convolution and pooling operations, it progressively abstracts the semantics of the image. Its lower-level convolutional layers capture local patterns (e.g., edges), while higher-level convolutional layers combine global structures (e.g., facial feature positions), thus separating content from style. The content features extracted by the encoder are high-level semantic feature maps, typically three-dimensional tensors (width × height × number of channels), encoding the facial topological relationships of the source portrait (e.g., relative positions of facial features), decoupled from style texture. As for performing stylistic transformation based on content features and applying identity consistency constraints based on facial identity feature vectors and style information during the transformation process, the specific implementation can be as follows: transfer the statistical characteristics of style information (e.g., mean, variance, etc.) to content features through affine transformation; generate a preliminary stylistic feature map; apply identity consistency constraints in real time during the transformation process, specifically including: calculating the similarity measure between the facial identity feature vector and the preset identity regions (e.g., eyes, nose) in the stylistic feature map; constructing an identity preservation loss function term based on the similarity measure; using the loss function term to guide the direction of stylistic transformation; adjusting the feature map through gradient backpropagation or feature reweighting to force identity feature consistency; and outputting the constrained stylistic features for the decoder to reconstruct the image.
[0045] Furthermore, due to the challenges of maintaining identity features, such as the non-transient nature of style transfer and the cumulative nature of identity drift, the former manifests as the generator needing to perform multiple steps to gradually render the style, for example, from local texture to global structure. The latter manifests as small geometric distortions in early layers being amplified step by step during the decoding stage, such as a shift in the corner of the mouth leading to facial asymmetry. These problems can be solved by intervening in the similarity metric calculation, loss function construction, and weight adjustment guidance in real time at key transformation nodes of the style transfer module (e.g., after affine transformation). Therefore, as an embodiment of this application, the above-mentioned application of identity consistency constraints based on facial identity feature vectors and style information during the transformation process can be as follows: based on style information, locate the identity-related region of the internal intermediate feature map generated by the style transfer module during the transformation process; calculate the similarity metric between the facial identity feature vector and the identity-related region; construct an identity preservation loss function term based on the similarity metric, which also incorporates the style consistency constraints corresponding to the style information; and use the identity preservation loss function term to guide the generation of feature maps or adjust network weights during the forward inference process of the style transfer module or during the generation parameter update process of the image generator to force the satisfaction of identity consistency requirements. In the above embodiments, calculating the similarity measure between the facial identity feature vector and the preset identity-related regions in the intermediate feature map output by the style transfer module essentially compares the semantic distance between the current feature map and the identity vector. This allows for the localization of real-time distortion regions and forces key facial features (e.g., inter-eye distance, nasal bridge curvature, etc.) to closely approximate the source image, enhancing resistance to style interference. Constructing an identity-preserving loss function based on the similarity measure transforms the distortion into a loss gradient, which inversely corrects the style transfer direction in the feature space, for example, by strengthening the weights of the eye region. This weight adjustment guidance, through gradient backpropagation or inference intervention, forces subsequent layers to project onto the identity space, thereby adapting to different style intensities and avoiding overfitting to a single style. Furthermore, it should be noted that in deep learning frameworks, the style transfer module typically consists of multiple network layers (e.g., convolutional layers, normalization layers, etc.), each outputting internal intermediate feature maps. These internal intermediate feature maps are numerical representations of high-level semantics, not the final image. For example, during the transformation process, the style transfer module progressively stylizes the input content features while simultaneously generating intermediate feature maps in real time. Furthermore, the application of identity consistency constraints is dynamic; that is, during the forward inference process of the style transfer module, the stylization direction is corrected in real time by calculating the similarity measure between the facial identity feature vector and specific regions (e.g., identity-related regions such as eyes and nose) in these intermediate feature maps, preventing identity distortion. Therefore, the internal intermediate feature maps generated by the style transfer module during the transformation process are intermediate results within the module, not the final output.
[0046] In the above embodiment where identity consistency constraints are applied based on facial identity feature vectors and style information during the transformation process, the location of identity-related regions in the internal intermediate feature maps generated by the style transfer module during the transformation process based on style information can be achieved by: receiving the internal intermediate feature maps (such as convolutional layer output feature maps) generated by the style transfer module during the transformation process; analyzing the semantic regions related to identity in the feature maps based on style information (e.g., statistical characteristics of style feature vectors), for example, calculating style weight maps through an attention mechanism; locating identity-related regions (e.g., feature regions corresponding to facial features) according to the style weight magnitude, and generating region masks or coordinate markers; and outputting the located identity-related region information for subsequent similarity calculations.
[0047] In the above embodiment where identity consistency constraints are applied based on facial identity feature vectors and style information during the transformation process, the similarity measure between the facial identity feature vector and the identity-related region can be calculated by using cosine similarity or Euclidean distance to calculate the similarity between the facial identity feature vector and the identity-related region.
[0048] In the above embodiment where identity consistency constraints are applied based on facial identity feature vectors and style information during the transformation process, constructing an identity preservation loss function term based on similarity measurement can be achieved by: receiving similarity measurement values and style information; constructing an identity preservation loss function term based on the similarity measurement values, for example, using a distance function such as L2 loss: Lid = 1 − similarity; constructing a style consistency constraint term based on style information, such as a style loss term, to ensure that the generated features statistically match the target style; weightedly fusing the identity preservation loss function term and the style consistency constraint term to generate a joint loss function term; and outputting the joint loss function term for subsequent optimization.
[0049] In the above embodiment where identity consistency constraints are applied based on facial identity feature vectors and style information during the transformation process, the identity preservation loss function term is used to guide the generation of feature maps or adjust network weights during the forward inference process of the style transfer module or during the generation parameter update process of the image generator to enforce the identity consistency requirement. This can be achieved by: receiving a joint loss function term; applying the loss function term to calculate gradients in real time during the forward inference process of the style transfer module to guide the generation of feature maps, for example, by enhancing the response of identity-related features through feature reweighting, or by adjusting network weights through backpropagation using the loss function term during the generation parameter update process of the image generator, for example, by optimizing convolutional layer parameters; and forcing the generated image to meet the identity consistency requirement based on the adjusted features or weights.
[0050] Figure 1The example method may also include guiding the generation of style features for non-facial regions (such as hairstyles and accessories) in a target portrait image based on facial identity feature vectors, so that the stylization results of non-facial regions are visually consistent with facial identity features. Specifically, this can involve: parsing global attributes in the facial identity feature vector, such as skin color, hair texture, etc., and calculating the semantic association strength between these global attributes and non-facial regions (including hairstyles, accessories, etc.), for example, determining the degree of influence of facial features on non-facial regions through vector inner product or attention weights; in the feature transformation stage of the style transfer module, modulating the feature maps corresponding to non-facial regions in real time based on semantic association strength, including: adjusting the mean and variance of the feature distribution of non-facial regions through affine transformation to make them consistent with facial identity features in color and texture, using conditional normalization to integrate facial identity information into the feature generation of non-facial regions to ensure visual consistency, etc.; during training or inference, constructing an identity consistency loss function term, comparing the similarity between non-facial regions in the generated image and the facial features of the source portrait, and dynamically adjusting the style transfer parameters using the loss gradient to force the stylization results of non-facial regions to remain consistent with facial identity features.
[0051] Since generative models based on image generation frameworks have irreversible risks of errors such as local generation failure and adversarial deception, local generation failure manifests as the loss of key identity features due to occlusion (e.g., sunglasses obscuring the eyes), while adversarial deception manifests as the generator potentially forging faces that can be checked by the loss function but are invalid due to biometric features, such as symmetrical fake eyes without iris texture. Introducing an external verification agent can construct an error correction loop. Therefore, in the image generation framework, applying identity consistency constraints in real time based on facial identity feature vectors and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image may also include: using a pre-trained facial keypoint discriminator or identity verification discriminator to identify the authenticity and identity consistency of the intermediate feature maps or real-time generated intermediate image data generated by the image generation framework in step S103; generating a feedback signal based on the identification result; and applying the feedback signal to dynamically adjust the processing parameters of the image generation framework in real time during the execution of step S103 or to adjust the optimization parameters during the post-processing optimization in step S105 to optimize image quality. In the above embodiments, macroscopic distortions can be captured by detecting the integrity of facial geometry (e.g., whether the distance between the eyes conforms to the golden ratio) through a pre-trained facial key point discriminator. Microscopic forgeries can be identified by comparing the biometric features (e.g., nasolabial fold microtexture) of the generated image with those of the source image through a pre-trained identity verification discriminator. Dynamic parameter adjustment can enhance the feature response of weak areas in real time according to the discrimination results, such as increasing the feature extraction intensity of occluded parts. In short, the closed-loop control mechanism of the above generation process can still output a structurally reasonable human image even with poor inputs such as occlusion and low light, which enhances the fault tolerance capability. The collaborative work of multiple discriminators can cover the dual constraints of geometric structure and biometric features, thereby ensuring identity consistency.
[0052] Step S104: Obtain the intermediate image result obtained after processing in step S103.
[0053] The intermediate image result obtained by step S103 can be a feature map that has been transformed and constrained by the style transfer module after being decoded by the decoder.
[0054] Step S105: Perform post-processing optimization on the intermediate image results and output the optimized target portrait image.
[0055] In the embodiments of this application, the processing of intermediate image results includes optimization operations, which may be one or more of the following: performing resolution enhancement processing, performing fine-grained repair processing for facial details, performing color harmonization processing based on the source portrait image or target style, and detecting occluded areas of the face region and performing repair processing, etc.
[0056] Furthermore, since automated image generation faces limitations in algorithmic determinism and subjective intent biases such as sensitivity to style parameters, in order to achieve the effect of users quickly verifying the balance between style and identity through preview, and to enhance system fault tolerance by repairing generation defects (e.g., facial blurring caused by over-stylization), the method in the above embodiments may further include: providing a real-time preview interface to show users the intermediate image result of style transfer or the target portrait image; receiving adjustment instructions input by the user based on the preview result; dynamically modifying the style transfer control parameters, the strength of identity consistency constraints, the generation parameters of the image generation framework, or the post-processing optimization parameters of step S105 according to the adjustment instructions; and re-executing or updating the processing in steps S103, S104, and / or S105 based on the modified parameters. The human-machine collaboration mechanism of the above embodiments maps the intermediate result to a user-perceptible space (e.g., rendering a preview image), eliminating the black box nature of the generation process, while incremental regeneration, i.e., only locally updating modules affected by parameters (e.g., not re-extracting identity features when adjusting style intensity), reduces interaction latency.
[0057] From the above appendix Figure 1 As illustrated in the example of the multi-style image conversion method for personalized portraits, on the one hand, the method receives style transfer control parameters provided by the user, locates and activates the corresponding pre-trained style transfer model based on these parameters. Therefore, users can dynamically activate the matching target style model using simple parameters (e.g., style labels, reference images, etc.), supporting rapid switching between independent or combined styles without retraining the entire network, thus improving the system's dynamic multi-style on-demand adaptation capability. On the other hand, by dynamically injecting identity constraint mechanisms (e.g., loss function guidance or feature space alignment) into the core style transfer processing stage, the facial structure of the generated image is forced to be consistent with the source portrait, significantly reducing the risk of identity distortion in complex style transfers. Even for abstract and highly deformable styles, the identity features of facial features and contours can still be maintained. In summary, the technical solution of this application can dynamically constrain the style transfer process, ensuring a high degree of consistency in portrait identity features.
[0058] Please see the appendix Figure 2 This application provides a multi-style image conversion device for personalized portraits. The device may include an extraction module 201, an activation module 202, a fusion module 203, an acquisition module 204, and an output module 205, as detailed below: Extraction module 201 is used to extract and output facial identity feature vectors representing the identity of a specific person from the input source portrait image; The activation module 202 is used to receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information. The fusion module 203 is used to generate images based on a preset image generation framework, and in the image generation framework, it applies identity consistency constraints in real time based on facial identity feature vectors and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image. The acquisition module 204 is used to acquire the intermediate image result obtained by the fusion module; The output module 205 is used to perform post-processing optimization on the intermediate image results and output the optimized target portrait image.
[0059] From the above appendix Figure 2 As illustrated in the example of the multi-style image conversion device for personalized portraits, on the one hand, it receives style transfer control parameters provided by the user, locates and activates the corresponding pre-trained style transfer model based on these parameters. Therefore, the user can dynamically activate the matching target style model using simple parameters (e.g., style labels, reference images, etc.), supporting rapid switching between independent or combined styles without retraining the entire network, thus improving the system's dynamic multi-style on-demand adaptation capability. On the other hand, by dynamically injecting an identity constraint mechanism (e.g., loss function guidance or feature space alignment) into the core style transfer processing stage, the facial structure of the generated image is forced to be consistent with the source portrait, significantly reducing the risk of identity distortion in complex style transfers. Even for abstract and highly deformable styles, the identity features of facial features and contours can still be maintained. In summary, the technical solution of this application can dynamically constrain the style transfer process, ensuring a high degree of consistency in portrait identity features.
[0060] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 3 As shown, the electronic device 3 in this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for a multi-style image conversion method for personalized portraits. When the processor 30 executes the computer program 32, it implements the steps in the above-described embodiment of the multi-style image conversion method for personalized portraits, for example... Figure 1 The steps S101 to S105 are shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of the extraction module 201, activation module 202, fusion module 203, acquisition module 204, and output module 205 are shown.
[0061] For example, the computer program 32 for a multi-style image conversion method for personalized portraits mainly includes: extracting and outputting facial identity feature vectors representing the specific identity of a person from the input source portrait image; receiving style transfer control parameters provided by the user, locating and activating the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information; generating an image based on a preset image generation framework, and applying identity consistency constraints in real time based on the facial identity feature vector and style information within the image generation framework to force the facial structure features of the person in the generated image to remain consistent with the source portrait image; obtaining the intermediate image result processed by the fusion module; performing post-processing optimization on the intermediate image result, and outputting the optimized target portrait image. The computer program 32 can be divided into one or more modules / units, one or more of which are stored in the memory 31 and executed by the processor 30 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of an extraction module 201, an activation module 202, a fusion module 203, an acquisition module 204, and an output module 205 (a module in the virtual device). The specific functions of each module are as follows: the extraction module 201 is used to extract and output facial identity feature vectors representing the identity of a specific person from the input source portrait image; the activation module 202 is used to receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information; the fusion module 203 is used to generate images based on a preset image generation framework, and in the image generation framework, apply identity consistency constraints in real time based on facial identity feature vectors and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image; the acquisition module 204 is used to acquire the intermediate image result obtained by the fusion module; and the output module 205 is used to perform post-processing optimization on the intermediate image result and output the optimized target portrait image.
[0062] Electronic device 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device may also include input / output devices, network access devices, buses, etc.
[0063] The processor 30 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0064] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard disk or RAM. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 31 can include both internal and external storage units of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0065] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed. That is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above-described device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0066] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0067] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0068] In the embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0069] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0070] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0071] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a storage medium. Based on this understanding, all or part of the processes in the above-described embodiments can also be implemented by a computer program instructing related hardware. The computer program for the multi-style image conversion method for personalized portraits can be stored in a storage medium. When executed by a processor, the computer program can implement the steps of the above-described method embodiments, namely: extracting and outputting facial identity feature vectors representing the identity of a specific person from the input source portrait image; receiving style transfer control parameters provided by the user, locating and activating the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information; generating an image based on a preset image generation framework, and applying identity consistency constraints in real time based on the facial identity feature vectors and style information within the image generation framework to force the facial structure features of the person in the generated image to be consistent with the source portrait image; obtaining the intermediate image result processed by the fusion module; performing post-processing optimization on the intermediate image result, and outputting the optimized target portrait image. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. Storage media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of storage media can be appropriately added to or removed according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, storage media may not include electrical carrier signals and telecommunication signals.
[0072] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application. The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the protection scope of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this invention.
Claims
1. A multi-style image conversion method for personalized portraits, characterized in that, The method includes the following steps: S101: Extract and output facial identity feature vectors representing the identity of a specific person from the input source portrait image; S102: Receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information; S103: Generate an image based on a preset image generation framework, and in the image generation framework, apply identity consistency constraints in real time based on the facial identity feature vector and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image. S104: Obtain the intermediate image result obtained after processing in step S103; S105: Perform post-processing optimization on the intermediate image results and output the optimized target portrait image.
2. The multi-style image conversion method for personalized portraits according to claim 1, characterized in that, The extraction and output of facial identity feature vectors representing a specific person's identity includes: The source portrait image is processed using a pre-trained face detection model to locate the face region in the image; The located facial region is input into a pre-trained identity feature extraction network, which is built based on a deep learning model; The identity feature extraction network extracts high-level semantic features of the face region and encodes them into a fixed-dimensional vector as the facial identity feature vector.
3. The method according to claim 1, characterized in that, The image generation framework includes an image generator built on a generative adversarial network, which includes an encoder, a style transfer module, and a decoder. The step of applying identity consistency constraints in real time based on the facial identity feature vector and style information within the image generation framework to force the facial structure features of the person in the generated image to remain consistent with the source portrait image includes: The encoder is used to process the source portrait image and extract content features; The content features, the style information, and the facial identity feature vector are input into the style transfer module; In the style transfer module, stylization transformation is performed based on the content features, and identity consistency constraints are applied based on the facial identity feature vector and style information during the transformation process.
4. The method according to claim 3, characterized in that, The step of applying identity consistency constraints based on the facial identity feature vector and style information during the transformation process includes: Based on the style information, the identity-related region of the internal intermediate feature map generated by the style transfer module during the transformation process is located; Calculate the similarity measure between the facial identity feature vector and the identity-related region; An identity preservation loss function term based on the similarity metric is constructed, and the identity preservation loss function term also incorporates the style consistency constraint corresponding to the style information; The identity preservation loss function term is used to guide the generation of feature maps or adjust network weights during the forward inference process of the style transfer module or during the generation parameter update process of the image generator, so as to force the identity consistency requirement to be met.
5. The multi-style image conversion method for personalized portraits according to claim 1, characterized in that, The method of applying identity consistency constraints in real time based on the facial identity feature vector and style information in the image generation framework to force the facial structure features of the person in the generated image to be consistent with the source portrait image also includes: A pre-trained facial landmark discriminator or identity verification discriminator is used to verify the authenticity and identity consistency of the intermediate feature map or real-time generated intermediate image data generated by the image generation framework in step S103. A feedback signal is generated based on the identification result; The feedback signal is used to dynamically adjust the processing parameters of the image generation framework in real time during the execution of step S103, or to adjust the optimization parameters during the post-processing optimization process in step S105, so as to optimize the image quality.
6. The method according to claim 1, characterized in that, The style information provided by the pre-trained style transfer model includes: Extract style feature representations from multiple activated pre-trained style transfer models; According to preset or user-specified fusion weights or rules, the multiple style feature representations are fused to generate a hybrid style feature as the style information.
7. The multi-style image conversion method for personalized portraits according to claim 1, characterized in that, The method further includes: Provides a real-time preview interface to show users intermediate image results or target portrait images of style transfer; Receive adjustment instructions from the user based on the preview results in the preview interface; According to the adjustment instructions, the style transfer control parameters, the strength of the identity consistency constraint, the generation parameters of the image generation framework, or the post-processing optimization parameters of step S105 are dynamically modified. Based on the modified parameters, re-execute or update the processes in steps S103, S104, and / or S105.
8. A multi-style image conversion device for personalized portraits, characterized in that, The device includes: The extraction module is used to extract and output facial identity feature vectors that represent the identity of a specific person from the input source portrait image; The activation module is used to receive style transfer control parameters provided by the user, locate and activate the corresponding pre-trained style transfer model based on the style transfer control parameters to provide style information. The fusion module is used to generate images based on a preset image generation framework. Within the image generation framework, identity consistency constraints are applied in real time based on the facial identity feature vector and style information to force the facial structure features of the person in the generated image to be consistent with the source portrait image. The acquisition module is used to acquire the intermediate image results obtained after processing by the fusion module. The output module is used to perform post-processing optimization on the intermediate image results and output the optimized target portrait image.
9. An electronic device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.