Biological feature perception-based portrait image generation method and device, and storage medium
By extracting and nonlinearly fusing age and gender feature vectors, generating images using generative adversarial networks and iteratively optimizing them, the problem of inconsistent control of multi-dimensional biometric features in existing technologies is solved, resulting in natural and harmonious human portrait images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING THUNDERSTONE TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
When controlling multiple biometric features simultaneously, existing technologies struggle to guarantee the coordination of facial features, the accuracy of target attribute expression, and physiological rationality in the generated images, leading to human portrait generation effects that deviate from the true evolutionary laws of the human body.
By extracting initial biometric features representing age and gender, converting them into target feature vectors, and performing nonlinear fusion in a preset feature space, a generative adversarial network is used to generate an image, and iterative optimization is performed to ensure that the generated image meets the target parameters.
It achieves precise and coordinated age and gender conversion, generates natural and harmonious portrait images, eliminates facial feature conflicts, and ensures that the generated results are highly consistent with user needs and conform to human physiological laws.
Smart Images

Figure CN121883646A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device and storage medium for generating human portrait images based on biometric perception. Background Technology
[0002] With the deepening application of artificial intelligence technology in the field of computer vision, biometric-based portrait editing technology has gradually become an important branch of digital image processing. This type of technology aims to automatically adjust biometric attributes such as age and gender of a person's image through algorithms, providing users with personalized image generation services. It has broad application prospects in entertainment, social networking, film and television production, and identity authentication. Currently, mainstream portrait biometric editing schemes mainly include two types: attribute decoupling models and cascaded optimization models. The former separates identity features and biometric features through an encoder-decoder structure and uses conditional vectors to control the generation direction, while the latter processes age and gender features sequentially through a serial network, using the output of the previous stage as the input of the next stage.
[0003] When controlling multiple biometric features simultaneously, the aforementioned existing technologies cannot guarantee the coordination of facial features, the accuracy of target attribute expression, and physiological rationality of the generated image, resulting in the human portrait generation effect under multi-parameter control deviating from the real human evolution law. Summary of the Invention
[0004] This application provides a method, device, and storage medium for generating human portrait images based on biometric perception, which can achieve accurate and coordinated age and gender conversion to generate natural and harmonious human portrait images.
[0005] On the one hand, this application provides a method for generating human portrait images based on biometric perception, the method comprising: Based on the input portrait image, extract initial biometric features representing age and gender; In response to the target age and target gender parameters specified by the user, the initial biometrics are converted into corresponding target age and target gender feature vectors respectively through preset feature transformation operations; In a preset feature space, the target age feature vector and the target gender feature vector are nonlinearly fused to generate a comprehensive biometric vector that includes the target age and gender attributes. Based on the comprehensive biometric feature vector, an initial generated image is generated within a generative adversarial network framework; The initial generated image is iteratively optimized and adjusted based on the target age feature vector and the target gender feature vector to output a personalized portrait image that meets the requirements of the target age parameter and the target gender parameter.
[0006] On the other hand, this application provides a human image generation device based on biometric perception, the device comprising: The extraction module is used to extract initial biometric features representing age and gender based on the input portrait image; The conversion module is used to convert the initial biometrics into corresponding target age feature vectors and target gender feature vectors respectively in response to the target age parameter and target gender parameter specified by the user through preset feature conversion operations; The fusion module is used to nonlinearly fuse the target age feature vector and the target gender feature vector in a preset feature space to generate a comprehensive biometric vector that includes the target age and gender attributes. The generation module is used to generate an initial human portrait image based on the comprehensive biometric feature vector within a generative adversarial network framework. The iterative module is used to iteratively optimize and adjust the initially generated image based on the target age feature vector and the target gender feature vector, and output a personalized portrait image that meets the requirements of the target age parameter and the target gender parameter.
[0007] Thirdly, this application provides an electronic device, the device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described biometric-based human image generation method.
[0008] Fourthly, this application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described biometric-based human image generation method.
[0009] As can be seen from the technical solution provided in this application, on the one hand, by processing age and gender features separately in feature extraction and preset feature transformation operations, the interference of a single attribute adjustment on another feature is avoided. Simultaneously, nonlinear fusion ensures that age-related and gender-related features in the final generated image can coordinate changes according to biological evolutionary laws, eliminating facial feature conflicts. On the other hand, by generating an initial image based on a comprehensive biometric vector and iteratively optimizing it, a closed-loop control link from feature parameters to image output is constructed, ensuring that the requirements of the target age and gender parameters can be accurately mapped to the generated image, significantly improving the fit between the generated result and user needs. Thirdly, the entire processing flow places biometric control under a unified feature space architecture. Through the dual constraints of feature space fusion and iterative optimization, the anatomical structure of the generated image conforms to the actual physiological laws of the human body, effectively avoiding the generation of anti-physiological features and ensuring that the facial structure after age and gender conversion maintains a natural and harmonious biological evolutionary continuity. In summary, the technical solution of this application can achieve accurate and coordinated age and gender conversion, generating natural and harmonious portrait images. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart of the biometric image generation method provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the human image generation device based on biometric perception provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] In this specification, adjectives such as "first" and "second" are used only to distinguish one element or action from another, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.
[0014] For ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn to actual scale.
[0015] With the deepening application of artificial intelligence technology in the field of computer vision, biometric-based portrait editing technology has gradually become an important branch of digital image processing. This type of technology aims to automatically adjust biometric attributes such as age and gender of a person's image through algorithms, providing users with personalized image generation services. It has broad application prospects in entertainment, social networking, film and television production, and identity authentication. Currently, mainstream portrait biometric editing solutions are mainly divided into two categories: 1) Attribute decoupling models, which specifically separate identity features and biometric features through an encoder-decoder structure and use conditional vectors to control the generation direction. While this method can adjust a single attribute such as age or gender, it faces two limitations: First, when adjusting age and gender simultaneously, the coupling between the two features in the latent space often leads to facial feature inconsistencies in the generated image (e.g., an elderly woman's face retaining a male skeletal structure); second, feature adjustment uses unidirectional linear superposition, without considering the physiological constraints of age changes on gender feature expression (e.g., the shape of the mandibular angle increases with age). 1) Changes); 2) Cascaded optimization model, which processes age and gender features sequentially through a serial network and uses the output of the previous stage as the input of the next stage. Although this method can control biometric features step by step, its fragmented operation leads to three core defects: First, the age adjustment process may destroy key gender features (e.g., changes in skin texture weaken gender markers); second, it lacks a cross-modal feature fusion mechanism, and the target age and target gender attributes cannot form a synergistic expression in a unified feature space; finally, the generated result depends on the accumulation of errors in the preceding steps, which is prone to producing artifacts that violate the laws of human anatomy (e.g., abnormal soft tissue distribution in images of juvenile males). The common problem with the above-mentioned existing technologies is that when controlling multiple dimensions of biometric features simultaneously, it is difficult to ensure the coordination of facial features, the accuracy of target attribute expression, and physiological rationality of the generated image. The root cause of this problem is that existing solutions have failed to establish a synergistic conversion mechanism between age and gender features and lack the ability to model the interaction relationship of biometric features, resulting in the image generation effect under multi-parameter control deviating from the real human evolution law.
[0016] To address the aforementioned problems in existing technologies, this application proposes a method for generating human portrait images based on biometric perception, the flowchart of which is attached. Figure 1 As shown, the main steps include S101 to S105, which are detailed below: Step S101: Based on the input portrait image, extract initial biometric features representing age and gender.
[0017] In the field of image processing, human skeletal topology, facial proportions, and inherent facial features are considered identity features, while age features (including skin texture and wrinkle distribution) and gender features (including the angle of the jaw and the prominence of the brow bone) are considered biological features. Identity features are subject to the hard constraint of spatiotemporal invariance (e.g., skull proportions remain unchanged after age 20), while biological features have the soft property of continuous variability, such as the gradual loss of collagen in the skin over the years. The reason for extracting initial biological features representing age and gender here is that separating biological features from identity features is the physical basis for decoupling control. If they are not separated and are forcibly coupled, it will lead to a series of problems, such as age adjustment distorting identity features such as the curvature of the nasal bridge, widening of the forehead and bone deformation during aging treatment, and abnormally increased interocular distance during gender conversion.
[0018] It should be noted that, in order to eliminate the interference of Euclidean space pose bias, photon noise, geometric distortion, and optical attenuation on biometrics, the input portrait image can be preprocessed before extracting the initial biometrics representing age and gender. Specifically, this includes: locating the face bounding box in the image using a face detection model; determining the coordinates of a preset number of facial key points using a facial key point localization model based on the detected face bounding box; aligning the face region to a standard pose using geometric transformation based on the facial key point coordinates; and performing image quality enhancement processing on the face region aligned to the standard pose. Here, image quality enhancement processing includes at least image denoising using a nonlinear filter and applying contrast-limited adaptive histogram equalization.
[0019] As an embodiment of this application, the initial biometric features representing age and gender can be extracted based on the input portrait image through steps S1011 to S1015, as detailed below: Step S1011: Use a face detection model to locate the face bounding box in the portrait image, and use a facial key point localization model to determine the coordinates of a preset number of facial key points based on the detected face bounding box.
[0020] In this embodiment, both the face detection model and the facial landmark localization model can be neural networks based on a deep learning framework. After pre-training, these models can obtain the face bounding box in the face image by inputting the portrait image into the pre-trained face detection model, and determine the coordinates of a preset number of facial landmarks by inputting the face bounding box into the pre-trained facial landmark localization model.
[0021] Step S1012: Apply geometric transformation based on the coordinates of facial key points to align the face region to the standard pose.
[0022] Based on the coordinates of facial key points, existing geometric transformation techniques can be used to align the facial region to a standard pose.
[0023] Step S1013: Extract multi-scale texture feature maps based on the face region aligned to the standard pose.
[0024] Specifically, step S1013 can be implemented by extracting multi-scale texture feature maps using a pre-trained convolutional neural network (e.g., a variant of VGG or ResNet) backbone. This network is pre-trained on a large face dataset (e.g., FFHQ) to learn general face representations. The process of extracting multi-scale texture feature maps includes: inputting an aligned face image into the network; extracting feature maps from different intermediate layers (e.g., shallow, medium, and deep layers) of the pre-trained convolutional neural network, which capture detailed texture, local structure, and global semantic information, respectively, thus forming multi-scale representations; these feature maps at different scales are then upsampled or downsampled to a uniform size and concatenated to form the final multi-scale texture feature map, providing rich input for subsequent age feature extraction.
[0025] Step S1012: Input the multi-scale texture feature map into the pre-trained age feature extraction network to obtain the initial age feature expression in the initial biometrics.
[0026] In this embodiment, the age feature extraction network can be a convolutional neural network structure designed for regression or classification tasks, with multi-scale texture feature maps as input. The network is trained on a large dataset of face images labeled with precise ages (e.g., IMDB-WIKI dataset, MORPH dataset, etc.) to learn how to map age-related information from texture features. The training objective is to minimize the error between the predicted age and the true age (e.g., L1 loss or cross-entropy loss). The output of the final layer of the network or the feature vector of an intermediate layer is used as the initial age feature representation, which encodes the age-related information of the input portrait.
[0027] Step S1013: Calculate the geometric feature set based on the facial key point coordinates, and input the geometric feature set into the pre-trained gender feature extraction network to obtain the initial gender feature expression in the initial biometrics.
[0028] In the above embodiments, the calculation of the geometric feature set based on facial keypoint coordinates can be based on a pre-processed set of facial keypoint coordinates (e.g., 68 or 106), calculating a series of geometric measures reflecting facial shape and skeletal structure. These measures may include distance ratios between different keypoints (e.g., the ratio of eye distance to face width), angles (e.g., mandibular angle), curvature, and the area of the convex hull formed by the keypoints or the aspect ratio of a specific region, etc. These calculated geometric measures are combined into a feature vector, i.e., the geometric feature set. As for the gender feature extraction network, it can be a multilayer perceptron or a lightweight convolutional network, trained on a face dataset containing labeled gender information, learning to distinguish gender-related morphological differences from the geometric feature set. The trained network takes the geometric feature set as input, and its output or the feature vector of a certain intermediate layer is the initial gender feature expression.
[0029] Step S102: In response to the target age parameter and target gender parameter specified by the user, the initial biometric features are converted into the corresponding target age feature vector and target gender feature vector respectively through a preset feature transformation operation.
[0030] Considering the non-linear coupling relationship between age and gender—for example, the rate of mandibular angle resorption is greater in older men than in older women—it generally requires processing through an independent conversion path. Otherwise, the facial incoordination rate of a single-path adjustment model would be high. Therefore, this application can convert the initial biometrics into corresponding target age feature vectors and target gender feature vectors through preset feature conversion operations. Specifically, as an embodiment of this application, in response to the user-specified target age and target gender parameters, the conversion of the initial biometrics into corresponding target age and target gender feature vectors through preset feature conversion operations can be achieved through steps S1021 to S1025, as detailed below: Step S1021: Determine the age adjustment factor based on the age difference and its direction coefficient and amplitude coefficient, wherein the age difference is the difference between the target age parameter specified by the user and the initial age value corresponding to the initial age feature expression.
[0031] Here, the directional coefficient of age difference is used to indicate the evolutionary direction of biological aging and rejuvenation. Specifically, if using Let represent the direction coefficient, then when the age difference is not less than 0, This can represent the aging process of organisms, for example, from 30 to 50 years old. Conversely, when the age difference is less than 0, This can represent the evolution of an organism towards juvenileity, for example, 30 years old → 10 years old. Amplitude coefficient. This is used to reasonably control the adjustment magnitude when determining the age adjustment factor, avoiding extreme adjustments, such as generating infantile characteristics through premature aging. It is used to obtain the directional coefficients. and amplitude coefficient Afterwards, if used k Indicating a phased adjustment factor, the age adjustment factor... It can be represented as .
[0032] Step S1022: Input the initial age feature expression and age adjustment factor into the age feature transformation network for processing, and output the target age feature vector.
[0033] As mentioned earlier, biometric characteristics such as age and gender are inherently decoupled from identity characteristics at the physical level (e.g., changes in skin texture are independent of bone structure). Forced decoupling aligns with the evolutionary laws of human biometric characteristics. Failure to separate core age feature factors and identity preservation factors will disrupt the continuity of facial topology, leading to the annihilation of identity information. On the other hand, the tensor product operation of age features and physiological correlation matrices can encode biomechanical constraints, which is beneficial for establishing anatomical correlations. Therefore, inputting the initial age feature expression and age adjustment factor into the age feature transformation network for processing, and outputting the target age feature vector, can specifically involve: separating the age-related feature components and identity-related feature components in the initial age feature expression that are related to age changes; and dynamically normalizing the age-related feature components based on the age adjustment factor, i.e., dynamically adjusting the scaling factor of the normalization operation. Translation factor Affine transformation is performed on the age-related feature components after dynamic normalization to generate the target age feature vector. The identity-related feature components are directly passed to subsequent nonlinear fusion for identity preservation. In the above embodiment, the identity preservation factor is independent of biometric conversion, ensuring that the output portrait retains the original core identity features. Physiologically constrained age feature components ensure that skin aging / juvenileization conforms to the actual anatomical evolution path, enhancing the naturalness of the conversion. Furthermore, the separate recombination avoids age adjustment distorting gender characteristics (e.g., aging weakens the soft contours of women), effectively eliminating feature conflicts.
[0034] It should be noted that in the above embodiments, the age-related feature components are dynamically normalized based on the age adjustment factor, i.e., the scaling factor of the normalization operation is dynamically adjusted. Translation factor This can be achieved through a Conditional Instance Normalization (CIN) layer. The parameters of this layer (scaling factor) Translation factor The value is not fixed, but is adjusted by a small feedforward network, i.e., a modulation network, based on the input age factor. Dynamically generated. The scaling factor of the normalization operation is dynamically adjusted. Translation factor The process is as follows: Adjust the age factor The input is fed into a modulation network; the output of the modulation network corresponds to a specific scaling factor that meets the current age adjustment requirements. Translation factor These dynamically generated and It is used for instance normalization of age-related feature components, that is, subtracting the mean and dividing by the standard deviation for each channel of the feature component before using... Scaling / Scaling A translation is performed. This process allows the normalization operation to adaptively adjust according to the direction and magnitude of changes in the target age, thereby more accurately guiding features to transform towards the target age.
[0035] Step S1023: Encode the user-specified target gender parameter into a target gender indicator vector.
[0036] In this embodiment, the target gender parameter (e.g., "male" or "female") is encoded as a fixed-dimensional one-hot vector or embedding vector. For example, in a binary classification scenario, a two-dimensional one-hot vector can be used, where [1, 0] represents male and [0, 1] represents female. This target gender indicator vector provides clear directional guidance for subsequent feature transformation.
[0037] Step S1024: Calculate the matching degree between the initial gender feature expression and the target gender indicator vector using the feature transformation function.
[0038] The feature transformation function in the above embodiments can be a learnable similarity calculation function, such as a cosine similarity function or an attention-based matching network. This function takes an initial gender feature representation and a target gender indicator vector as input and calculates a matching score between them. This score quantifies the degree of difference between the gender features of the current image and the target gender. If a simple cosine similarity is used, the matching score is the cosine of the angle between the two vectors in space; if a more complex matching network is used, the network learns a non-linear mapping from feature and indicator vectors to the matching score.
[0039] Step S1025: Transform the initial gender feature expression based on the matching degree between the initial gender feature expression and the target gender indicator vector, and output the target gender feature vector.
[0040] Specifically, step S1025 can be implemented through a gender feature transformation module, including: performing a nonlinear transformation on the initial gender feature expression based on the matching degree between the initial gender feature expression calculated in step S1024 and the target gender indicator vector; if the matching degree is high (indicating that the current feature is close to the target gender), the transformation amplitude is small, mainly fine-tuning; if the matching degree is low (indicating a large difference), a more significant transformation is performed. This transformation operation can be an affine transformation implemented through one or more fully connected layers or convolutional layers, and its parameters can be dynamically adjusted according to the matching degree and the target gender indicator vector; the transformation process aims to minimize the difference between the initial gender feature and the target gender in the feature space, and finally outputs a target gender feature vector that can represent the target gender.
[0041] Step S103: In the preset feature space, the target age feature vector and the target gender feature vector are nonlinearly fused to generate a comprehensive biometric feature vector containing the target age and gender attributes.
[0042] Simple concatenation of different vectors leads to dimensional collapse of the feature space. Therefore, after obtaining the target age feature vector and the target gender feature vector, the target age feature vector and the target gender feature vector can be non-linearly fused in a preset feature space. It should be noted that, in the embodiments of this application, the feature space is a high-dimensional representation space learned internally by the model, in which the attribute features of each portrait, such as age and gender, can be expressed in vector form, and the operations between vectors (e.g., fusion) can correspond to the semantic changes of attributes on the image.
[0043] Furthermore, the elements of the physiological characteristic correlation matrix M between the age and gender dimensions can explicitly define the physiological coupling coefficients of age and gender characteristics (e.g., the decay gradient of brow bone prominence in male aging), thereby accurately modeling cross-dimensional correlations. Establishing an anthropologically accurate mapping based on anatomical markers from a 3D face scan database ensures that the algorithm follows biological evolution rather than data fitting. Therefore, as an embodiment of this application, the nonlinear fusion of the target age feature vector and the target gender feature vector to generate a comprehensive biometric vector containing target age and gender attributes can be achieved through steps S1031 to S1034, as detailed below: Step S1031: Based on the anthropological facial feature database, construct the physiological feature correlation matrix M between the age dimension and the gender dimension.
[0044] Elements of the physiological characteristic correlation matrix M This represents the biological coupling strength between the i-th type of age-related feature and the j-th type of sex-related feature. Specifically, constructing the physiological feature association matrix M between the age dimension and the sex dimension includes: extracting key anatomical markers from a 3D face scan database; calculating the distribution of skeletal structure parameters and soft tissue thickness distribution for a specific age group; measuring the gradient of the sex difference coefficient on the age axis; and establishing a mapping relationship matrix between the age feature set and the sex feature set through canonical correlation analysis.
[0045] Step S1032: Convert the target age feature vectors respectively and target gender feature vector Input gating fusion network, output dynamic gating vector Among them, the target age feature vector It directly participates in subsequent nonlinear fusion.
[0046] In this embodiment, the gated fusion network is a lightweight neural network comprising an age stream processing branch, a gender stream processing branch, and a coupling prediction layer. The age stream processing branch uses a temporal convolutional network to capture the evolution pattern of age features, while the gender stream processing branch uses orientation-sensitive convolutional kernels to capture gender-related geometric features. The coupling prediction layer is responsible for mapping the two types of features to a unified gating space and outputting a dimension-adaptive gating vector. The gated fusion network is trained end-to-end during the overall model training process, and its training data is consistent with the training data of the entire portrait generation model (e.g., large face datasets such as FFHQ and CelebA). The training objective is to minimize the difference between the final generated image and the real image in terms of target age and gender attributes. Through training, the gated fusion network learns how to generate appropriate gating signals based on the input target feature vector to control the information flow in the subsequent fusion process. Furthermore, it should be noted that although the target age feature vector... and target gender feature vector All inputs are into the gating fusion network; however, the output is only the dynamic gating vector. Target age feature vector It directly participates in subsequent nonlinear fusion without generating independent dynamic gating vectors.
[0047] Step S1033: Utilize dynamic gating vectors For the target gender feature vector After modulation, the target age feature vector The modulated target gender feature vector is subjected to feature decoupling and recombination, and feature interaction analysis based on the dual-stream mutual attention module is used to generate cross-enhanced biofeature representations.
[0048] Here, dynamic gating vectors are used. For the target gender feature vector Modulation can be achieved by using a dynamic gate vector. With the target gender feature vector This is achieved by performing channel-by-channel multiplication (or weighting), i.e. , This is the modulated target gender feature vector. This operation aims to dynamically enhance or suppress the salience of different channels in the gender feature vector based on an age-gender physiological correlation model, thereby preparing for subsequent feature fusion and making the expression of gender features more adaptable to the context of the target age. Specifically, this is achieved using dynamic gating vectors. For the target gender feature vector After modulation, the target age feature vector Performing feature decoupling and recombination on the modulated target gender feature vector, and generating cross-enhanced biometric representations based on feature interaction analysis using a two-stream mutual attention module, can be achieved by: separating the features using the feature decoupling module. The core age characteristic factor and identity retention factor in the data were simultaneously separated. The core age feature factor and its associated expression factors are analyzed. The core age feature factor is multiplied by a tensor product with the physiological feature correlation matrix M to generate a physiologically constrained age feature component. This physiologically constrained age feature component is then input into a two-stream mutual attention module for feature interaction analysis with the core age feature factor, generating a cross-enhanced biometric representation. In the above embodiments, the feature decoupling module can be implemented based on an encoder-decoder structure or by introducing a specific decoupling loss function (e.g., correlation minimization loss). It is designed to process the input feature vector (e.g., the target age feature vector) into a two-stream mutual attention module for feature interaction analysis. and target gender feature vector The module is decomposed into components that are strongly correlated with specific attributes (core feature factors) and weakly correlated / uncorrelated (identity preservation factors, secondary expression factors). This module learns decoupled representations during joint training with the portrait generation model. The two-stream mutual attention module is a neural network structure containing two parallel processing streams (age stream and gender stream) and an attention mechanism that computes the mutual attention between the two streams of features. This module, trained on a large face dataset, learns how to model the interaction between age and gender features. Its core is to compute the attention weights between the two streams of features, allowing the feature information of one stream to modulate the feature expression of the other. Core age feature factors mainly refer to features strongly correlated with age changes, such as wrinkle depth, skin laxity, etc.; identity preservation factors refer to features related to individual identity that do not change significantly with age or gender, such as basic facial skeletal structure; core gender feature factors refer to features strongly correlated with gender differences, such as jawline contour, brow bone prominence, etc.; secondary expression factors refer to features related to gender expression but may be affected by other factors (e.g., makeup, hairstyle).
[0049] In the above embodiments, the feature interaction analysis based on the dual-stream mutual attention module can specifically be: calculating the cross-correlation matrix between each channel of the core age feature factor and each channel of the core gender feature factor, and generating a feature association energy map. , where matrix elements The signal represents the biological coupling strength between the i-th age feature channel and the j-th gender feature channel; a modulation signal is generated based on the feature correlation energy map E; and the signal is generated according to the age feature saliency vector. and gender feature saliency vector Highly significant feature channels are selected; using a pre-defined feature channel-anatomical region mapping table, these channels are associated with corresponding facial anatomical regions; for the anatomical regions associated with these highly significant feature channels, region-selective enhancement processing is performed, and the enhanced features are integrated back into the cross-enhanced biometric representation; the modulated signal includes an age feature saliency vector. and gender feature saliency vector The above embodiments enhance key biomarkers by precisely targeting gender-sensitive areas (such as the protrusion of the Adam's apple in men and the distribution of fat in the cheekbones in women). Local enhancement effectively avoids the unnatural feeling caused by excessive modification of irrelevant areas, thereby maintaining facial harmony. Anatomically driven enhancement makes the target gender features conform to human visual cognitive expectations, thereby improving visual recognition.
[0050] In the above embodiments, the technical solution for generating the modulated signal based on the feature correlation energy E can be as follows: Global average pooling is performed on the feature correlation energy map E along the age channel dimension and the gender channel dimension, respectively, to obtain two one-dimensional vectors; these vectors are then converted into age feature saliency vectors with values in the range [0,1] using a small fully connected network or a sigmoid function. and gender feature saliency vector These vectors represent the importance of each feature channel in the age-gender interaction. Based on the saliency vector of age features... and gender feature saliency vector Channels with high significance can be identified by setting an adjustable threshold. For example, channels with significance scores higher than this threshold can be classified as high significance channels.
[0051] The feature channel-anatomical region mapping table is a predefined lookup table constructed before model training by analyzing the model's activation map on the validation set or using interpretable AI techniques to associate specific feature channels with the facial anatomical regions (e.g., corners of the eyes, corners of the mouth, jawline, and forehead) to which they respond most sensitively. This table establishes a semantic relationship between the feature space and the image pixel space. For the anatomical regions associated with highly saliency feature channels, region-selective enhancement processing is performed, and the enhanced features are integrated back into the cross-enhanced biometric representation, which can be achieved through a spatial attention mechanism. Specifically, the facial regions corresponding to highly saliency channels can be found based on the mapping table, and binary masks or soft attention maps of these regions can be generated. At the feature map or image level, enhancement operations are applied to these regions, such as increasing the intensity of the region's features (through multiplication or addition) or using specific filters to enhance the details of the region. The enhanced features are then integrated back into the cross-enhanced biometric representation. The process of integrating the enhanced features back into the cross-enhanced biometric representation can be as follows: Based on the pre-defined mapping relationship between feature channels and facial anatomical regions, a region enhancement mask is generated for the selected highly significant feature channels corresponding to the target anatomical regions; using the region enhancement mask, feature enhancement processing is performed on the corresponding anatomical regions in the cross-enhanced biometric representation to generate enhanced features; the enhanced features are adaptively weighted and fused with the original cross-enhanced biometric representation to generate an updated biometric representation; and the updated biometric representation is output for subsequent image generation.
[0052] Step S1034: Selectively fuse the cross-enhanced biometric representation with the identity preservation factor separated from the initial age feature expression to output a comprehensive biometric vector. .
[0053] Specifically, the cross-enhanced biometric representation is selectively fused with the identity preservation factor separated from the initial age feature expression to output a comprehensive biometric vector. This can be achieved by: concatenating the cross-enhanced biometric representation vector with the identity preservation factor vector along a preset dimension to form a fused input feature; inputting the fused input feature into a gated generative network containing a two-layer fully connected structure, and outputting a gated vector with values between 0 and 1 through a Sigmoid activation function; using the gated vector to perform weighted calculations on the biometric representation, and simultaneously using the complement of the vector to weight the identity preservation factor, implementing selective fusion at the feature channel level; performing L2 norm normalization on the fused feature vector, and implementing a feature protection mechanism when the modulus of the fused vector is less than a preset threshold; wherein, the weighting intensity of the identity preservation factor is set with a lower threshold in the key facial feature channels, and the weighting intensity of the biometric representation is set with an upper threshold in the physiological attribute channels.
[0054] As can be seen from steps S1031 to S1034 of the above embodiments, matrix-driven feature fusion effectively avoids generating anti-physiological features such as juvenile bones paired with elderly skin texture. It forces age and gender features to change synergistically under biological constraints (e.g., the absorption rate of the female mandibular angle is optimized with age), which improves feature synergy. The correlation matrix provides a medically verifiable white-box model of feature interaction, which enhances the interpretability of the technology.
[0055] In the above embodiments, by processing age and gender features separately through feature extraction and preset feature transformation operations, while avoiding interference from single attribute adjustment on the other feature, nonlinear fusion is performed in a unified feature space through feature fusion, establishing a physiological coupling relationship between age and gender features. This allows age-related features (e.g., wrinkle distribution) and gender-related features (e.g., jaw contour) in the final generated image to change in a coordinated manner according to the laws of biological evolution, thereby effectively eliminating facial feature conflicts.
[0056] Step S104: Based on the comprehensive biometric feature vector, generate an initial generated image within the generative adversarial network framework.
[0057] Compared to generative models such as VAEs, the gradient local response characteristics of generative adversarial networks (GANs) can precisely control facial details, such as the direction of a single wrinkle. In tests on the FFHQ dataset, GANs also showed significantly improved Local Feature Fidelity (LIPIS) compared to VAEs. Therefore, an initial generated image can be generated within the generative adversarial network framework based on a comprehensive biometric vector. Specifically, as an embodiment of this application, generating an initial generated image within the generative adversarial network framework based on a comprehensive biometric vector can be achieved through steps S1041 to S1044, as detailed below: Step S1041: The integrated biometric feature vector is mapped into a latent space vector of a generative adversarial network through a multilayer perceptron network.
[0058] In this embodiment, the multilayer perceptron network can be stacked with multiple fully connected layers and activation functions (e.g., ReLU, Tanh). Its function is to nonlinearly map the high-dimensional comprehensive biometric feature vector to the latent space distribution desired by the generative adversarial network (GAN) generator. This network is trained end-to-end along with the entire generative model, using the same training data as the GAN (e.g., FFHQ and other face datasets). Its parameters are optimized using the backpropagation algorithm to ensure that the mapped latent vector can be effectively decoded by the generator.
[0059] Step S1042: Input the latent space vector into the input layer of the pre-trained image generator.
[0060] In this embodiment, the image generator may employ a style-based generator architecture (e.g., a variant of StyleGAN) that is pre-trained on a large, diverse face dataset (e.g., FFHQ) to learn the manifold distribution of natural face images. The image generator may consist of a series of upsampling blocks, each containing operations such as convolution, upsampling, normalization, and activation functions. The input layer receives a latent spatial vector and progressively upsamples and transforms it into a high-resolution face image.
[0061] Step S1043: In the convolutional layers of the image generator, style parameters dynamically generated based on the comprehensive biological feature vector are injected for feature modulation.
[0062] Specifically, dynamic instance normalization is applied between each convolutional layer of the image generator, with normalization parameters (e.g., scaling factors)... Translation factor Modulation feature maps are generated from the integrated biometric feature vectors through a small fully connected network. These modulation feature maps contain data generated by a scaling factor. Controlled target age texture distribution features, by translation factor Maintaining basic identity structure features and gendered geometric expression features formed through channel interactions, etc. The process of injecting style parameters dynamically generated based on the comprehensive biometric feature vector for feature modulation is as follows: through a style mapping network, the injected comprehensive biometric feature vector is used to generate style parameter pairs corresponding to each convolutional layer ( , Before the convolution operation, instance normalization is applied to the input feature map; using the corresponding... and Affine transformations are performed to inject semantic information about the target's age and gender into feature maps of different resolutions, thereby achieving refined feature modulation.
[0063] Step S1044: The image generator outputs the initial generated image.
[0064] As can be seen from step S104 of the above embodiment, by generating an initial generated image based on the comprehensive biometric feature vector and iteratively optimizing it, a closed-loop control link from feature parameters to image output is constructed to ensure that the requirements of the target age parameter and the target gender parameter can be accurately mapped to the generated image: on the one hand, the initial generated image is generated based on the fusion vector to realize the expression of basic attributes; on the other hand, the target attribute matching is strengthened by iterative adjustment, which significantly improves the fit between the generated result and the user's needs.
[0065] Step S105: Iteratively optimize and adjust the initially generated image based on the target age feature vector and the target gender feature vector to output a personalized portrait image that meets the requirements of the target age parameter and the target gender parameter.
[0066] Considering that the initial generated image has an average age deviation due to feature fusion errors, and that this average age deviation can be reduced through closed-loop feedback optimization, conversely, if the initial portrait image is not iteratively optimized, the standard deviation of facial key point offset in gender conversion will be significantly increased. Therefore, after generating the initial generated image, iterative optimization and adjustment can be performed on the initial generated image based on the target age feature vector and the target gender feature vector to output a personalized portrait image that meets the requirements of the target age and target gender parameters. Specifically, as an embodiment of this application, iterative optimization and adjustment of the initial generated image based on the target age feature vector and the target gender feature vector to output a personalized portrait image that meets the requirements of the target age and target gender parameters can be achieved through the following steps S1051 to S1057, which are described in detail below: Step S1051: Input the initially generated image into the feature verification module.
[0067] In this embodiment of the application, the feature verification module includes a pre-trained 3D face reconstruction model branch, which is used to reconstruct 3D geometric information from the initially generated image and use it as an auxiliary input for calculating apparent age features and apparent gender features.
[0068] Step S1052: In the feature verification module, the pre-trained age estimation model and gender classification model are used simultaneously to extract the apparent age features and apparent gender features of the initially generated image, respectively.
[0069] Step S1053: Calculate the difference between the apparent age feature and the target age feature vector, and the difference between the apparent gender feature and the target gender feature vector, respectively.
[0070] Specifically, the difference between apparent age features and target age feature vectors can be calculated using Euclidean distance, while the difference between apparent gender features and target gender feature vectors can be calculated using vector cosine similarity.
[0071] Step S1054: Combine age difference and gender difference to generate overall difference loss value.
[0072] Specifically, the process of fusing age difference and gender difference to generate the overall difference loss value can be as follows: calculate the age difference between the apparent age feature and the target age feature vector, and the gender difference between the apparent gender feature and the target gender feature vector, respectively; standardize the age difference and gender difference to obtain standardized age difference and standardized gender difference; assign corresponding fusion weights to the standardized age difference and standardized gender difference, and perform weighted summation to generate the overall difference loss value; output the overall difference loss value to guide the parameter optimization of the generative model.
[0073] Step S1055: When the overall difference loss value is higher than the preset threshold, update the parameters of the gated fusion network and / or the multilayer perceptron network according to the loss value and through the backpropagation algorithm.
[0074] Step S1056: Re-execute steps S1041 to S1044 using the updated network parameters to perform nonlinear fusion of the target age feature vector and the target gender feature vector, and generate the adjusted image.
[0075] Step S1057: Repeat steps S1051 to S1056 until the overall difference loss value is lower than the preset threshold or the maximum number of iterations is reached.
[0076] On the one hand, considering that negative user feedback can label unknown feature spaces of the model (such as aging patterns of specific ethnic groups), the lack of such a mechanism would limit the system's adaptability. On the other hand, updating the feature extraction, transformation, and fusion modules can maintain subsystem version compatibility, while updating a single network in isolation would disrupt the coupling balance. The method in the above embodiment also includes a result feedback and update mechanism, namely: collecting user feedback on the biometric conformity of the output personalized portrait image; when the conformity of the user feedback is lower than a preset level, based on the input image of the feedback sample, the target parameters specified by the user, and the output image, updating the model parameters of at least one of the age feature extraction network, gender feature extraction network, age feature transformation network, gated fusion network, and feature verification module. Through the above result feedback and update mechanism, not only can it dynamically adapt to long-tail scenarios such as different ethnicities and pathological special cases, enhancing the universality of the technology, but also, through feedback loops, the model can continuously approximate the real biometric distribution patterns, maintaining the system's evolutionary capability. The online update mechanism can also avoid the frequent overall replacement of traditional models, thereby reducing deployment and maintenance costs.
[0077] From the above appendix Figure 1As illustrated by the example of a biometric-based portrait image generation method, on the one hand, by processing age and gender features separately through feature extraction and preset feature transformation operations, interference from a single attribute adjustment on the other feature is avoided. Simultaneously, nonlinear fusion ensures that age-related and gender-related features in the final generated image can coordinate changes according to biological evolutionary laws, eliminating facial feature conflicts. On the other hand, by generating an initial image based on a comprehensive biometric vector and iteratively optimizing it, a closed-loop control link from feature parameters to image output is constructed, ensuring that the requirements of target age and gender parameters can be accurately mapped to the generated image, significantly improving the fit between the generated result and user needs. Thirdly, the entire processing flow places biometric control under a unified feature space architecture. Through the dual constraints of feature space fusion and iterative optimization, the anatomical structure of the generated image conforms to actual human physiological laws, effectively avoiding the generation of anti-physiological features and ensuring that the facial structure after age and gender conversion maintains natural and harmonious biological evolutionary continuity. In summary, the technical solution of this application can achieve accurate and coordinated age and gender conversion, generating natural and harmonious portrait images.
[0078] Please see the appendix Figure 2 This application provides a biometric image generation device, which may include an extraction module 201, a conversion module 202, a fusion module 203, a generation module 204, and an iteration module 205, as detailed below: Extraction module 201 is used to extract initial biometric features representing age and gender based on the input portrait image; The conversion module 202 is used to convert the initial biometric features into the corresponding target age feature vector and target gender feature vector respectively through preset feature conversion operations in response to the target age parameter and target gender parameter specified by the user. The fusion module 203 is used to nonlinearly fuse the target age feature vector and the target gender feature vector in a preset feature space to generate a comprehensive biometric vector containing the target age and gender attributes. The generation module 204 is used to generate an initial generated image based on the comprehensive biometric feature vector within the generative adversarial network framework. The iterative module 205 is used to iteratively optimize and adjust the initially generated image based on the target age feature vector and the target gender feature vector, and output a personalized portrait image that meets the requirements of the target age parameter and the target gender parameter.
[0079] From the above appendix Figure 2As illustrated by the example of a biometric-based portrait image generation device, on the one hand, by processing age and gender features separately through feature extraction and preset feature transformation operations, interference from a single attribute adjustment on the other feature is avoided. Simultaneously, nonlinear fusion ensures that age-related and gender-related features in the final generated image can coordinate changes according to biological evolutionary laws, eliminating facial feature conflicts. On the other hand, by generating an initial image based on a comprehensive biometric vector and iteratively optimizing it, a closed-loop control link from feature parameters to image output is constructed, ensuring that the requirements of target age and gender parameters can be accurately mapped to the generated image, significantly improving the fit between the generated result and user needs. Thirdly, the entire processing flow places biometric control under a unified feature space architecture. Through the dual constraints of feature space fusion and iterative optimization, the anatomical structure of the generated image conforms to actual human physiological laws, effectively avoiding the generation of anti-physiological features and ensuring that the facial structure after age and gender conversion maintains natural and harmonious biological evolutionary continuity. In summary, the technical solution of this application can achieve accurate and coordinated age and gender conversion, generating natural and harmonious portrait images.
[0080] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 3 As shown, the electronic device 3 in this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for a biometric-based portrait image generation method. When the processor 30 executes the computer program 32, it implements the steps described in the above-described biometric-based portrait image generation method embodiment, for example... Figure 1 The steps S101 to S105 are shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of the extraction module 201, conversion module 202, fusion module 203, generation module 204, and iteration module 205 are shown.
[0081] For example, the computer program 32 of the biometric image generation method mainly includes: extracting initial biometric features representing age and gender based on the input portrait image; in response to user-specified target age and target gender parameters, converting the initial biometric features into corresponding target age and target gender feature vectors respectively through preset feature transformation operations; nonlinearly fusing the target age and target gender feature vectors in a preset feature space to generate a comprehensive biometric vector containing target age and gender attributes; generating an initial generated image based on the comprehensive biometric vector within a generative adversarial network framework; iteratively optimizing and adjusting the initial generated image based on the target age and target gender feature vectors to output a personalized portrait image that meets the requirements of the target age and target gender parameters. The computer program 32 can be divided into one or more modules / units, one or more of which are stored in the memory 31 and executed by the processor 30 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 32 in the electronic device 3. For example, computer program 32 can be divided into the functions of extraction module 201, transformation module 202, fusion module 203, generation module 204, and iteration module 205 (modules in the virtual device). The specific functions of each module are as follows: Extraction module 201 is used to extract initial biometric features representing age and gender based on the input portrait image; Transformation module 202 is used to convert the initial biometric features into corresponding target age feature vectors and target gender feature vectors respectively through preset feature transformation operations in response to the target age parameters and target gender parameters specified by the user; Fusion module 203 is used to perform nonlinear fusion of the target age feature vector and the target gender feature vector in a preset feature space to generate a comprehensive biometric feature vector containing target age and gender attributes; Generation module 204 is used to generate an initial generated image based on the comprehensive biometric feature vector under the generative adversarial network framework; Iteration module 205 is used to iteratively optimize and adjust the initial generated image based on the target age feature vector and the target gender feature vector to output a personalized portrait image that meets the requirements of the target age parameters and the target gender parameters.
[0082] Electronic device 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device may also include input / output devices, network access devices, buses, etc.
[0083] The processor 30 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0084] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard disk or RAM. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 31 can include both internal and external storage units of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0085] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed. That is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above-described device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0086] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0087] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0088] In the embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0089] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0090] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0091] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of this application can also be implemented by a computer program instructing related hardware. The computer program for the biometric image generation method can be stored in a storage medium. When executed by a processor, the computer program can implement the steps of the above-described method embodiments, namely, extracting initial biometric features representing age and gender based on the input portrait image; in response to the target age and target gender parameters specified by the user, converting the initial biometric features into corresponding target age feature vectors and target gender feature vectors respectively through preset feature transformation operations; performing nonlinear fusion of the target age feature vector and the target gender feature vector in a preset feature space to generate a comprehensive biometric vector containing target age and gender attributes; generating an initial generated image based on the comprehensive biometric vector under a generative adversarial network framework; iteratively optimizing and adjusting the initial generated image based on the target age feature vector and the target gender feature vector to output a personalized portrait image that meets the requirements of the target age and target gender parameters. Computer programs include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Storage media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of storage media can be appropriately added or removed according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, storage media do not include electrical carrier signals and telecommunication signals.
[0092] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application. The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the protection scope of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this invention.
Claims
1. A method for generating human portrait images based on biometric perception, characterized in that, The method includes: Based on the input portrait image, extract initial biometric features representing age and gender; In response to the target age and target gender parameters specified by the user, the initial biometrics are converted into corresponding target age and target gender feature vectors respectively through preset feature transformation operations; In a preset feature space, the target age feature vector and the target gender feature vector are nonlinearly fused to generate a comprehensive biometric vector that includes the target age and gender attributes. Based on the comprehensive biometric feature vector, an initial generated image is generated within a generative adversarial network framework; The initial generated image is iteratively optimized and adjusted based on the target age feature vector and the target gender feature vector to output a personalized portrait image that meets the requirements of the target age parameter and the target gender parameter.
2. The method of claim 1, wherein, The initial biometric features representing age and gender based on the input portrait image include: The face bounding box in the portrait image is located using a face detection model, and the coordinates of a preset number of facial key points are determined based on the detected face bounding box using a facial key point localization model. Based on the coordinates of the facial key points, a geometric transformation is applied to align the face region to a standard pose; Multi-scale texture feature maps are extracted based on face regions aligned to standard poses. The multi-scale texture feature map is input into a pre-trained age feature extraction network to obtain the initial age feature expression in the initial biometrics; A geometric feature set is calculated based on the facial key point coordinates, and the geometric feature set is input into a pre-trained gender feature extraction network to obtain the initial gender feature expression in the initial biometrics.
3. The method of claim 2, wherein, The step of converting the initial biometrics into corresponding target age feature vectors and target gender feature vectors respectively through a preset feature transformation operation, in response to the target age parameter and target gender parameter specified by the user, includes: An age adjustment factor is determined based on the age difference and its direction and amplitude coefficients, wherein the age difference is the difference between the target age parameter specified by the user and the initial age value corresponding to the initial age feature expression; The initial age feature expression and the age adjustment factor are input into an age feature transformation network for processing, and the target age feature vector is output. Encode the user-specified target gender parameter into a target gender indication vector; The matching degree between the initial gender feature expression and the target gender indicator vector is calculated using a feature transformation function; The initial gender feature expression is transformed based on the matching degree to output the target gender feature vector.
4. The method of claim 3, wherein, The step of inputting the initial age feature representation and the age adjustment factor into an age feature transformation network for processing, and outputting the target age feature vector, includes: In the initial age feature expression, the age-related feature component and the identity-related feature component related to age change are separated. The identity-related feature component is directly passed to the subsequent nonlinear fusion for identity preservation. scaling factor for the dynamic normalization operation of the age-related feature component according to the age adjustment factor and a translation factor ; An affine transformation is performed on the age-related feature components after dynamic normalization to generate the target age feature vector.
5. The method of claim 1, wherein, The step of nonlinearly fusing the target age feature vector and the target gender feature vector to generate a comprehensive biometric vector containing target age and gender attributes includes: Based on an anthropological face feature database, a physiological feature correlation matrix between the age dimension and the gender dimension is constructed M ; The target age feature vectors are respectively and target gender feature vector Input gating fusion network, output dynamic gating vector The target age feature vector Directly participate in subsequent nonlinear fusion; Using the dynamic gating vector For the target gender feature vector After modulation, the target age feature vector The modulated target gender feature vector is decoupled and recombined, and the feature interaction analysis based on the dual-stream mutual attention module generates a cross-enhanced biofeature representation. selectively fusing the cross-enhanced biometric representation with an identity preserving factor isolated from the initial age feature expression, outputting a composite biometric vector .
6. The method of claim 5, wherein, The feature interaction analysis based on the dual-stream mutual attention module includes: Calculate the cross-correlation matrices between each channel of the core age feature factor and each channel of the core gender feature factor obtained in the feature decoupling and recombination step, and generate a feature correlation energy map. The matrix elements of the feature-related energy map E This represents the biological coupling strength between the i-th age characteristic channel and the j-th gender characteristic channel. and These represent the number of channels for age characteristics and the number of channels for gender characteristics, respectively. Based on the aforementioned feature-correlation energy map E, a modulation signal is generated, which includes an age feature saliency vector. and gender feature saliency vector ; Based on the significance vector of the age characteristics and gender feature saliency vector Filter out channels with highly significant features; By using a pre-defined feature channel-anatomical region mapping table, highly significant feature channels are associated with their corresponding anatomical regions. For the anatomical regions associated with the highly significant feature channels, region-selective enhancement processing is performed, and the enhanced features are integrated back into the cross-enhanced biometric representation.
7. The method of claim 3, wherein, After outputting a personalized portrait image that meets the requirements of the target age and target gender parameters, the method further includes: Collect user feedback on the biometric conformity of the output personalized portrait images; When the user feedback is less than the preset level, based on the input image of the feedback sample, the user-specified target age parameters and target gender parameters, and the personalized portrait image, update the model parameters of the age feature extraction network, gender feature extraction network, age feature transformation network, gated fusion network, and at least one of the feature verification modules involved in the iterative optimization adjustment step.
8. A biometric feature perception-based portrait image generation apparatus characterized by comprising: The device includes: The extraction module is used to extract initial biometric features representing age and gender based on the input portrait image; The conversion module is used to convert the initial biometrics into corresponding target age feature vectors and target gender feature vectors respectively in response to the target age parameter and target gender parameter specified by the user through preset feature conversion operations; The fusion module is used to nonlinearly fuse the target age feature vector and the target gender feature vector in a preset feature space to generate a comprehensive biometric vector containing the target age and gender attributes. The generation module is used to generate an initial generated image based on the comprehensive biofeature vector within a generative adversarial network framework. The iterative module is used to iteratively optimize and adjust the initially generated image based on the target age feature vector and the target gender feature vector, and output a personalized portrait image that meets the requirements of the target age parameter and the target gender parameter.
9. An electronic device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized by When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.