Network face replacement method, device and equipment based on dynamic symmetric coding and decoding, and medium
The face replacement method based on a dynamic symmetric encoding and decoding network solves the problem of insufficient quality of face replacement image generation in the existing technology, achieves high-quality and secure image generation, and ensures the coordination and accuracy between the face and the environment.
Patent Information
- Application Number
- CN202510802376.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
AI Technical Summary
In existing technologies, face replacement has artifacts or identity information loss in image generation quality, resulting in insufficient data accuracy and security. Especially in high-resolution scenarios, the face is not coordinated with the surrounding environment, affecting the visual effect and application accuracy.
A face replacement method based on a dynamic symmetric codec network is adopted. By obtaining the preprocessing and feature extraction of the source image and the replacement image, the noise vector is combined to generate a fusion feature, and the preset dynamic symmetric codec network model is used for iterative denoising. The multi-scale discriminator is combined for quality assessment and model optimization to finally generate a high-quality replacement image.
The generation quality of face replacement images is improved, details and clarity are enhanced, the harmonious unity between the face and the environment is ensured, and the accuracy and security of the generated images are improved.
Smart Images

Figure CN120689198A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection technology, and in particular to a face replacement method, device, equipment and medium based on a dynamic symmetric coding and decoding network. Background Art
[0002] Face replacement is the process of automatically replacing a face in a video or image with another face while keeping the expression, movement and posture of the original video unchanged. This requires not only that the face replacement be extremely natural, but also that it be in harmony with the surrounding environment.
[0003] For example, in healthcare, face replacement can be used to protect medical data privacy. For example, in telemedicine and medical research, face replacement can be used to prevent privacy leaks by replacing a patient's face. It can also assist in medical education and surgical simulations, simulating facial features or surgical outcomes for different conditions. In fintech, face replacement can be used as an auxiliary means of identity verification, simulating identity information in different environments to enhance security. It can also be used to generate customized digital avatars that match the cultural background of different users during interactions with virtual banking assistants, improving service intimacy.
[0004] However, current face replacement approaches often suffer from image quality issues, with generated images often exhibiting artifacts or losing identity information, impacting data accuracy and security. In high-resolution scenarios, insufficient fusion of details can cause the face to appear out of sync with its surroundings, reducing visual quality and application accuracy.
[0005] Therefore, the accuracy and security issues of face replacement in the existing technology need to be solved urgently. Summary of the Invention
[0006] The present invention provides an artificial intelligence-based face replacement method, device, computer equipment and medium based on dynamic symmetric coding and decoding networks to solve the accuracy and security problems of face replacement in the prior art.
[0007] In a first aspect, a face replacement method based on a dynamic symmetric encoding and decoding network is provided, comprising:
[0008] Acquire a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image to obtain a feature vector;
[0009] Initializing a noise vector, and combining the noise vector with the feature vector through multi-level injection to obtain a fusion feature;
[0010] Iteratively denoising the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image;
[0011] Using a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, and adjusting the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model;
[0012] The optimized generation model is used to generate a maximum resolution replacement image corresponding to the source image.
[0013] In a second aspect, a face replacement device based on a dynamic symmetric coding and decoding network is provided, comprising:
[0014] a feature extraction module, configured to obtain a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image, and obtain a feature vector;
[0015] A feature fusion module is used to initialize a noise vector and combine the noise vector with the feature vector through multi-level injection to obtain a fused feature;
[0016] A denoising module, configured to iteratively denoise the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image;
[0017] An updating module is configured to perform a quality assessment on the multi-resolution denoised image using a preset multi-scale discriminator, and adjust the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model;
[0018] A generation module is used to generate a maximum resolution replacement image corresponding to the source image using the optimized generation model.
[0019] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the above-mentioned face replacement method based on dynamic symmetric coding and decoding network are implemented.
[0020] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned face replacement method based on dynamic symmetric coding and decoding network are implemented.
[0021] In the above-mentioned solution implemented by the face replacement method, device, computer equipment and storage medium based on the dynamic symmetric codec network, the source image and the replacement image can be obtained, and preprocessed and feature extracted to obtain a feature vector. Then, the noise vector is initialized and combined with the feature vector through multi-level injection to generate a fused feature. The fused feature is iteratively denoised using a preset dynamic symmetric codec network model to generate a multi-resolution denoised image. Finally, the denoised image is quality evaluated using a multi-scale discriminator and the learnable parameters of the codec network model are updated based on the evaluation results, thereby obtaining an optimized generation model. The optimized generation model can improve the quality of the generated image, making it closer to the original image, and can also enhance the details and clarity of the image through iterative denoising and multi-resolution processing. In addition, through continuous quality evaluation and updating of model parameters, the optimized generation model can continuously learn and improve its generation strategy, thereby providing more accurate and higher-quality results when generating the maximum resolution replacement image. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0023] Figure 1 This is a schematic diagram of an application environment of a face replacement method based on a dynamic symmetric coding and decoding network according to an embodiment of the present invention;
[0024] Figure 2 This is a flow chart of a face replacement method based on a dynamic symmetric coding and decoding network in one embodiment of the present invention;
[0025] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S1;
[0026] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S2;
[0027] Figure 5 This is a structural diagram of a face replacement device based on a dynamic symmetric coding and decoding network in one embodiment of the present invention;
[0028] Figure 6 is a structural diagram of a computer device in one embodiment of the present invention;
[0029] Figure 7 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0031] The face replacement method based on dynamic symmetric coding and decoding network provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the client communicates with the server through a network. The server can obtain the source image and the replacement image through the client, pre-process and extract features of the source image and the replacement image to obtain a feature vector; initialize the noise vector, combine the noise vector with the feature vector through multi-level injection to obtain a fusion feature; use a preset dynamic symmetric codec network model to iteratively denoise the fusion feature to generate a multi-resolution denoised image; use a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, adjust the learnable parameters in the dynamic symmetric codec network model according to the quality assessment result, and obtain an optimized generation model; use the optimized generation model to generate the maximum resolution replacement image corresponding to the source image. The client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0032] See also Figure 2 As shown, Figure 2 A flowchart of a face replacement method based on a dynamic symmetric coding and decoding network provided by an embodiment of the present invention includes the following steps:
[0033] S1. Acquire a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image, and obtain a feature vector.
[0034] In this embodiment of the present invention, the feature vector is high-level semantic information extracted from the source image and the replacement image, which is used to guide the generative model to perform face replacement. Specifically, the feature vector includes two parts: identity features and non-identity features. Identity features are used to retain facial identity information in the source image, such as skin color, facial contours, eyes, nose, etc., to ensure that the generated facial image is consistent with the source image in terms of identity. Non-identity features are used to retain non-identity information in the replacement image, such as expression, lighting, background, etc., to ensure that the generated facial image matches the replacement image in these aspects.
[0035] In the present invention, see Figure 3 As shown, the preprocessing and feature extraction of the source image and the replacement image to obtain a feature vector includes:
[0036] S10, cropping the source image and the replacement image to obtain a cropped source image and a cropped replacement image;
[0037] S11, unifying the resolutions of the cropped source image and the cropped replacement image to obtain a standard source image and a standard replacement image;
[0038] S12. Perform multi-layer convolution and pooling on the standard source image and the standard replacement image to obtain a feature vector.
[0039] In this embodiment of the present invention, the cropping process involves performing face detection on the source image and the replacement image using a face detection network, locating the face region, cropping the face portion, and removing irrelevant background information. The face detection network is a multi-task convolutional neural network, a three-tiered cascade network with a three-layer architecture consisting of a candidate network, a refinement network, and an output network. The execution of the multi-task convolutional neural network is divided into three cascade stages: first, the candidate network rapidly scans the input image to generate a large number of candidate regions that may contain faces. This network uses sliding window technology, combined with the feature extraction capabilities of a convolutional neural network, to quickly identify possible face regions in the image and generate preliminary bounding boxes. The refinement network then further refines the candidate regions generated by the candidate network. The refinement network further adjusts the position and size of the bounding boxes to more accurately locate the face region. Finally, the output network further optimizes the position and size of the bounding boxes to ensure accurate cropping of the face region. At the same time, the output network will accurately locate the key points of the face (such as eyes, nose, mouth, etc.). The output of the output network includes the bounding box coordinates and key point coordinates of the face. The multi-task convolutional neural network crops out the face part based on this information and removes irrelevant background information.
[0040] The sliding window technique uses a fixed-size window to slide across the image pixel by pixel or in fixed steps, analyzing and processing the image area within each window. For example, the sliding window starts at the upper left corner of the image and gradually moves rightward and downward, one or several pixels at a time, until the entire image is covered. At each window position, a preset convolutional neural network is used to evaluate whether the image area within the window contains a face. If the preset convolutional neural network determines that the area is a face, the bounding box coordinates of the area are recorded.
[0041] For example, in the fintech sector, sliding window technology is used for customer authentication and anti-fraud detection. A camera captures a customer's facial image. Using sliding window technology, the image is analyzed and processed pixel by pixel or in steps of a certain size within each window. A pre-defined convolutional neural network evaluates whether the image region within the window contains a face. If so, the bounding box coordinates of the region are recorded. The detected face is then compared with a registered customer image in a database to verify the authenticity of the customer's identity. If the detected face does not match the registered image, the system can trigger an alarm to prevent fraud.
[0042] In the examples of the present invention, unifying the resolutions of the cropped source image and the cropped replacement image to obtain a standard source image and a standard replacement image involves first performing a size assessment to determine the original width and height dimensions of the two cropped images by reading the number of pixel rows and columns. A unified strategy is then developed to select a target resolution (e.g., 128×128, 512×512, etc.) based on task requirements, directly scaling the width and height by a fixed ratio. A bilinear interpolation algorithm is then used to perform resampling operations to generate an image that meets the target resolution.
[0043] Specifically, the bilinear interpolation algorithm calculates the new value of a target pixel by taking into account the values and relative positions of the four nearest pixels, using a weighted average. Specifically, bilinear interpolation first linearly interpolates the two nearest pixels horizontally to obtain two intermediate values; these two intermediate values are then linearly interpolated again vertically to obtain the value of the target pixel. Linear interpolation is a mathematical method for smoothing the distance between two known data points, calculating the intermediate values by proportionally assigning weights based on the relative positions of the interpolated point and the two endpoints.
[0044] In the examples of the present invention, the multi-layer convolution and pooling of the standard source image and the standard replacement image to obtain a feature vector refers to a convolutional neural network (such as Swin Transformer) extracting multi-scale features of the image through multi-layer convolution and pooling operations, and then generating a feature vector through a fully connected layer or a global pooling operation. The image is first input into the network, and is processed through a series of convolution layers and nonlinear activation functions to gradually extract high-level features of the image. These features are converted into fixed-length feature vectors after global average pooling or maximum pooling.
[0045] Specifically, the multi-layer convolution operation refers to the stacking of multiple convolution layers in the network, each convolution layer uses a set of learnable convolution kernels to slide on the input image to detect local features in the image. These convolution kernels are element-wise multiplied and summed with the local area of the input image through the convolution operation to generate a new feature map, highlighting specific patterns or features in the image. The pooling operation is used to reduce the spatial dimension of the feature map, reduce the amount of calculation and the number of parameters, while retaining important features. Common pooling methods include maximum pooling and average pooling, where maximum pooling takes the maximum value in the local area and average pooling takes the average value in the local area. The nonlinear activation function is a rectified linear unit (ReLU), which is applied after the convolution operation to introduce a nonlinear activation function to the network, enabling the network to learn more complex feature representations. The nonlinear activation function sets all negative values to zero and retains positive values. This nonlinear transformation helps the network learn complex patterns in the data.
[0046] In the examples of the present invention, the source image and the replacement image are preprocessed and feature extracted to obtain a feature vector. The cropping operation can reduce redundant information interference, unify the resolution to ensure input consistency, and avoid feature offset caused by size differences. The convolutional neural network can capture image semantic information and enhance the model's generalization ability. In the field of financial technology, this process can be applied to identity authentication scenarios, such as preprocessing ID photos uploaded by customers and real-time facial images, extracting feature vectors by cropping the facial area and unifying the image size, and then performing a similarity comparison. This effectively improves the accuracy of identity authentication, while reducing the impact of factors such as lighting and posture to ensure transaction security.
[0047] S2. Initialize a noise vector, and combine the noise vector with the feature vector through multi-level injection to obtain a fusion feature.
[0048] In the examples of the present invention, the time steps are a series of discrete time points used to control the denoising process, which are pre-set and used throughout the model training and inference process. In the context of the diffusion model, these time steps are usually expressed as a sequence of integers from 0 to T, where T represents the total number of steps in the entire diffusion process, and each time step t corresponds to a specific stage in the denoising process. The diffusion model is a generative model based on deep learning, which generates new data samples by gradually adding noise to the data and then learning how to recover the original data from the noisy data. Specifically, in the forward diffusion stage of the diffusion model, noise is gradually added to the data until it is completely converted into noise; while in the reverse denoising stage, the model starts from the noise and gradually removes the noise until clear data is restored.
[0049] In the present invention, see Figure 4As shown, the noise vector is initialized, and the noise vector is combined with the feature vector through multi-level injection to obtain a fusion feature, including:
[0050] S20, determining a time step embedding vector corresponding to the feature vector;
[0051] S21, performing a multi-layer perceptron transformation on the time step embedding vector to generate a scaling factor;
[0052] S22. Perform attention scaling on the scaling factor to obtain an attention weight;
[0053] S23, multiplying the attention weight by the feature vector to obtain an intermediate feature vector;
[0054] S24, initializing a noise vector, and adding the noise vector and the intermediate feature vector vector by vector to obtain an updated feature vector;
[0055] S25: Using the updated feature vector as a feature vector, and returning to the step of determining the time step embedding vector corresponding to the feature vector, until a preset number of iterations is reached, and using the updated feature vector as a fusion feature.
[0056] In the examples of the present invention, determining the time step embedding vector corresponding to the feature vector is to map the corresponding time step of the feature vector sequence having time series characteristics to a high-dimensional space to obtain the time step embedding vector. The method uses a sinusoidal position encoding method to map each time step t to a high-dimensional space to generate a time step embedding vector Embed(t). The time step embedding vector not only contains the absolute position information of time step t, but also captures the relative position relationship of the time step in the entire denoising process through mapping in the high-dimensional space.
[0057] Among them, the sinusoidal position encoding is a technology for converting discrete sequence position information into a continuous vector representation, which is used to introduce the position information of elements in the sequence. By calculating the specific frequencies of the sine and cosine functions to generate a unique vector for each position, the model can use this position information to understand the order relationship of the elements in the sequence. Specifically, for each position p in a sequence, the sinusoidal position encoding generates a vector of dimension d, where each dimension corresponds to a different sine wave frequency. For the i-th dimension, if i is an even number, the sine function is used; if i is an odd number, the cosine function is used. The formula for position encoding can be expressed as:
[0058]
[0059] Among them, p represents the position index in the sequence, i represents the dimension index, and d modelIndicates the dimension size of the model, and PE represents position encoding.
[0060] In the embodiment of the present invention, the multilayer perceptron transformation of the time step embedding vector to generate the scaling factor is to input the time step embedding vector into a multilayer perceptron (MLP). The multilayer perceptron is processed through a series of linear transformations and nonlinear activation functions to output a scalar value. The scalar value is then incremented by 1 to obtain the noise-dependent scaling factor α(t). The calculation formula is as follows:
[0061] α(t)=1+MLP(Embed(t))
[0062] Where α(t) represents the noise-dependent scaling factor at time step t, t represents the time step, Embed(t) represents the embedding vector of the time step, and MLP(Embed(t)) is a multi-layer perceptron.
[0063] Specifically, the multilayer perceptron is a simplified neural network structure that processes input data through a series of linear transformations and nonlinear activation functions to extract higher-level features or make predictions. In the process of processing the time-step embedding vector to output a scalar value, the multilayer perceptron first receives the time-step embedding vector generated by sinusoidal position encoding as input, and then maps the input vector to a new space through at least one layer of linear transformation (such as a fully connected layer). Next, a nonlinear activation function (such as a linear rectified unit function) is applied to introduce nonlinear characteristics. Finally, after processing through all layers, a single scalar value is output, which is an intermediate result used for other calculations (such as the noise-dependent scaling factor α(t) in this example).
[0064] In the present invention, the attention weight is obtained by performing attention scaling on the scaling factor, which means adjusting the attention weight calculation in the standard attention mechanism by introducing a scaling factor α(t) calculated dynamically based on the time step when processing the hth attention head at time step t. The attention weight is calculated using the following formula:
[0065]
[0066] in, is the attention weight matrix of the h-th attention head at time step t, Softmax is a function that converts a real vector into a probability distribution, Q h is the query matrix of the h-th attention head, is the transpose of the key matrix of the h-th attention head, d kis the dimension of the key vector, α(t) represents the noise-dependent scaling factor at time step t, t represents the time step, h represents the index of the attention head, and k represents the index of the element in the key vector.
[0067] In the example of the present invention, the attention weight is multiplied by the feature vector to obtain the intermediate feature vector, which is to multiply the attention weight vector obtained by calculation with the corresponding feature vector according to the element position one by one. Through this weighted operation, the feature parts that the model considers important are highlighted, while relatively unimportant features are suppressed, thereby obtaining the intermediate feature vector containing attention information.
[0068] In this embodiment of the present invention, the initialization of the noise vector and the vector-by-vector addition of the noise vector to the intermediate feature vector to obtain the updated feature vector involves randomly generating a noise vector that conforms to a Gaussian distribution. This noise vector has the same dimensions as the feature vector currently processed by the model. This noise vector is then element-wise added to the intermediate feature vector. This addition operation incorporates the noise information into the feature representation, thereby generating the updated feature vector.
[0069] Specifically, the random generation of a noise vector that conforms to a Gaussian distribution is to generate a random noise vector of the same size as the input image based on a Gaussian (normal) distribution. This process involves determining the probability distribution of each element of the noise vector so that its mean is 0 and its standard deviation is 1, thereby ensuring that the generated noise vector conforms to the characteristics of a standard normal distribution. In practice, by randomly drawing the same number of samples as the number of image channels from a standard normal distribution, these samples are then reshaped into the same dimensions (H, W, C) as the input image, where H and W represent the height and width of the image, respectively, and C represents the number of channels.
[0070] In the example of the present invention, the updated feature vector is used as the feature vector, and the step of determining the time step embedding vector corresponding to the feature vector is returned until the preset number of iterations is reached. The updated feature vector is used as the fusion feature, and the updated feature vector obtained by vector-by-vector addition is used as the new input feature vector, and is sent back to the step of calculating the time step embedding vector to recalculate the new time step embedding vector and the corresponding scaling factor. This process is repeated until the preset number of iterations is reached (such as 1000 times). In each iteration, the attention weights are updated through the attention mechanism using the newly calculated scaling factor, and these weights are applied to the feature vector to obtain weighted features. Subsequently, the randomly generated noise vector is added to the weighted feature vector to generate an updated feature vector. When the iteration reaches the specified number of times, the final updated feature vector will be output as a fusion feature.
[0071] In this example, by determining the time-step embedding vector corresponding to the feature vector and using a multi-layer perceptron transform to generate a scaling factor, the model can adjust the attention mechanism based on the contextual information of the current time step. By multiplying the attention weights obtained through attention scaling with the feature vector, the model can highlight important features and suppress unimportant features, thereby obtaining a more refined intermediate feature vector. Then, by initializing a noise vector and adding it vector-by-vector to the intermediate feature vector, the model not only increases the diversity of the input but also simulates real-world noise, making the model more robust. Finally, by iteratively updating the feature vector until a preset number of iterations is reached, the model can gradually optimize its learned feature representation. The final output fused feature vector will contain richer and more robust information, helping to improve the model's generalization ability on various tasks. For example, in the field of healthcare, it can be used to generate diverse facial samples. By adjusting the scaling factor and noise vector, facial features of different styles and expressions can be generated, assisting in the display of virtual effects and design solutions in fields such as medical beauty.
[0072] S3. Iteratively denoise the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image.
[0073] In the examples of the present invention, the symmetric encoding and decoding network model refers to the dynamic U-Net, which inherits the encoder-decoder structure of the traditional U-Net. The encoder gradually converts the input data into multi-scale feature representations and predicts the related noise; the decoder performs the opposite process, using these feature representations to gradually restore clear output data. In the denoising task, the model dynamically adjusts internal parameters to adapt to inputs of different resolutions, implements layer-by-layer denoising, and ultimately generates multi-resolution denoised images. However, a dynamic adjustment mechanism is introduced to adapt to input images of different resolutions. The dynamic U-Net can adaptively process features of different scales and optimize the fusion effect of multi-scale features through a resolution-aware gating mechanism and noise-dependent attention scaling. This design enables the network to dynamically adjust its behavior according to the current time step and noise level during the denoising process, thereby focusing more on the global structure in the early stages and more on local details in the later stages, achieving dynamic adaptation to the denoising process.
[0074] In the examples of the present invention, the multi-resolution denoised image refers to the denoising results of multiple different resolution versions generated for the same image during the image denoising process. These images of different resolutions usually include multiple levels such as lower resolution (such as 64×64 pixels), medium resolution (such as 128×128 pixels) and higher resolution (such as 256×256 pixels). This multi-resolution denoising image generation method can capture the details and structural information of the image at different scales, thereby more comprehensively evaluating and optimizing the denoising effect. Lower resolution images can provide global image structure information, while higher resolution images can retain more detailed features.
[0075] In an embodiment of the present invention, the iterative denoising of the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image includes:
[0076] Perform feature parsing on the fused features using a preset dynamic symmetric encoding and decoding network model to obtain a feature map and corresponding resolution embedding;
[0077] generating gating weights based on the feature map and the resolution embedding;
[0078] Calculating a weighted sum of the gating weight and the feature map to obtain an output feature map;
[0079] Performing noise prediction on the output feature map using a preset dynamic symmetric encoding and decoding network model;
[0080] The output feature map is subtracted from the noise prediction result, and upsampled using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image.
[0081] In the example of the present invention, the fusion features are analyzed by using a preset dynamic symmetric codec network model to obtain a feature map and a corresponding resolution embedding, which is to input the fusion features into the symmetric codec network, gradually extract multi-scale feature representations through the encoder part, and map these features back to the original space in the decoder part to generate a feature map. In this process, the dynamic symmetric codec network dynamically adjusts its internal parameters according to the resolution of the input image to adapt to inputs of different resolutions. Specifically, the symmetric codec network first gradually extracts features through the downsampling layer of the encoder, and each layer outputs a feature map FI and a corresponding resolution embedding R, where R is calculated from the image size (H, W) by a multi-layer perceptron, and is used to represent the resolution information of the input image.
[0082] In the example of the present invention, generating the gating weight based on the feature map and the resolution embedding means dynamically adjusting the weight gl,l′ of the jump connection between the encoder layer l and the decoder layer l′ using the feature map Fl and the resolution embedding vector R in the dynamic symmetric codec network. The resolution embedding vector R is concatenated with the feature map Fl of the encoder layer l to form a new feature representation. Then, the gating weight gl,l′ is calculated using a learnable weight matrix Wg and a bias bg, as well as a sigmoid activation function σ, as follows:
[0083] g l,l' =σ(W g Concat(F l ,R)+b g )
[0084] Among them, g l,l' represents the gate weight, σ represents the sigmoid activation function, W g Represents a learnable weight matrix, Concat(F l ,R) represents the feature map F of the encoder layer l l and resolution embedding vector R are concatenated on the feature dimension, b g represents the learnable bias vector, F l represents the feature map of the encoder layer l, and R represents the resolution embedding vector.
[0085] In the present embodiment, the calculation of the weighted sum of the gating weight and the feature map to obtain the output feature map refers to using the calculated gating weight gl,l′ to adjust the information flow between the feature map Fl of the encoder layer l and the feature map Fl′ of the decoder layer l′ in the dynamic symmetric encoding and decoding network. The output feature map is calculated using the following formula:
[0086]
[0087] in, It represents the output feature map after dynamic fusion of the resolution-aware gating mechanism, which is the jump connection feature between the encoder layer l and the decoder layer l′. l,l' represents the gate weight, F l Represents the feature map of the encoder layer l, Identity (F l ) is the identity mapping function, directly returning the input feature F l .
[0088] In the example of the present invention, the noise prediction of the output feature map using a preset dynamic symmetric codec network model is to input the output feature map into a pre-built codec network with a symmetric structure, which can dynamically adjust its own parameters or structure according to the characteristics of the input features; in the encoding stage, the model extracts the abstract representation of the feature map through operations such as convolution and downsampling to capture its semantic information; in the decoding stage, the spatial details of the feature map are gradually restored through operations such as deconvolution and upsampling, and multi-scale fusion is achieved by combining the intermediate features of the encoding stage; the model analyzes the feature map at multiple levels of the network, and based on the mapping relationship between the learned noise distribution pattern and the feature representation, predicts the distribution of noise in the feature map (such as noise intensity, type, etc.), and finally outputs multi-level noise prediction results.
[0089] In the example of the present invention, the output feature map is subtracted from the noise prediction result, and the preset dynamic symmetric codec network model is used for upsampling to generate a multi-resolution denoised image. This is achieved by gradually optimizing these fused features through multiple iterative denoising steps in the decoder of the dynamic symmetric codec network, combining upsampling operations and multi-scale information transmission of jump connections, and finally generating denoised image outputs with different resolution levels, wherein deep features retain global structural information and shallow features enhance local details, thereby achieving high-quality generation of multi-resolution images.
[0090] In the examples of the present invention, the introduction of resolution embedding and dynamic gating weights enables the model to adaptively adjust the contribution of features at different scales, effectively solving the problem of detail loss or noise amplification in cross-resolution feature fusion in traditional methods. Secondly, the symmetric encoding and decoding structure is combined with the iterative optimization mechanism to achieve progressive reconstruction from global semantics to local details, significantly improving denoising accuracy and image quality. Finally, the multi-resolution output mechanism can simultaneously meet the image accuracy and computational efficiency requirements of different application scenarios. For example, in the field of financial technology, this technology can be used for document image enhancement and anti-counterfeiting recognition, such as multi-scale denoising and super-resolution reconstruction of blurred ID card photos, which can clearly restore text information (high resolution) while maintaining the recognizability of facial features (medium resolution).
[0091] S4. Use a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, and adjust the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model.
[0092] In the examples of the present invention, the multi-scale discriminator is a neural network structure used for tasks such as evaluating image quality. It simulates the human visual system's perception of multi-level details of images by analyzing images at multiple scales (such as different resolution levels). At each scale, the multi-scale discriminator extracts image features, such as texture and structure, and compares them with the corresponding features of high-quality reference images. It uses a specific loss function (such as adversarial loss) to measure the difference between the input image and the real image in the feature space, thereby quantitatively evaluating the image quality and judging the image's performance in terms of structural rationality, detail integrity, etc. at different scales, providing a basis for subsequent optimization of the image generation model and other operations.
[0093] In an embodiment of the present invention, the method of using a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image and adjusting the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model includes:
[0094] Using a preset multi-scale discriminator to evaluate the multi-resolution denoised images respectively to determine the loss value of each image;
[0095] Weighting the loss values of the multi-resolution denoised image according to a preset weight to calculate a total loss value;
[0096] The learnable parameters in the dynamic symmetric encoding and decoding network are iteratively optimized according to the total loss value, and the iterative optimization is stopped when the total loss value is less than a preset loss value to obtain an optimized generation model.
[0097] In this embodiment of the present invention, the multi-resolution denoised images are evaluated using a preset multi-scale discriminator to determine the loss value of each image. For each denoised image of a specific resolution, a matching multi-scale discriminator is called. For example, for a denoised image with a resolution of 64×64, D64 is used for evaluation; for a denoised image with a resolution of 128×128, D128 is used for evaluation; and for a denoised image with a resolution of 256×256, D256 is used for evaluation. Each discriminator outputs a loss value, such as L64, L128, or L256, which reflects the similarity or quality difference between the generated image and the real image at the corresponding resolution.
[0098] In the example of the present invention, the loss values of the multi-resolution denoised image are weighted according to preset weights to calculate the total loss value. This is to assign corresponding weights to the loss values at different resolutions based on the importance of each resolution to the final image quality. For example, the loss value L64 of the denoised image with a resolution of 64×64, the loss value L128 of the denoised image with a resolution of 128×128, and the loss value L256 of the denoised image with a resolution of 256×256 are assigned weights of 0.2, 0.3, and 0.5, respectively. The total loss value is calculated by weighted summation, and the calculation formula is as follows:
[0099] L total =0.2L 64 +0.3L 128 +0.5L 256
[0100] Among them, L total Represents the total loss value, L 64 Represents the loss value of the denoised image with a resolution of 64×64, L 128 represents the loss value of the denoised image with a resolution of 128×128, L 256 Represents the loss value of the denoised image with a resolution of 256×256.
[0101] In an embodiment of the present invention, the iterative optimization of the learnable parameters in the dynamic symmetric encoding and decoding network according to the total loss value, and stopping the iterative optimization when the total loss value is less than a preset loss value to obtain the optimized generation model, includes:
[0102] Determining the gradient of the total loss value with respect to a learnable parameter in the dynamic symmetric codec network using a back-propagation algorithm;
[0103] Adjust the learnable parameters in the dynamic symmetric encoding and decoding network using a preset optimizer according to the gradient, and iteratively calculate a new total loss value;
[0104] When the new total loss value is less than the preset loss value, the iteration is stopped to obtain the optimized generation model.
[0105] In the example of the present invention, the use of the back propagation algorithm to determine the gradient of the total loss value with respect to the learnable parameters in the dynamic symmetric encoding and decoding network refers to the core mechanism of back propagation, starting from the output layer of the network, and reversely calculating the gradient of the total loss value Ltotal for each learnable parameter in the network (such as the weight matrix Wg and the bias bg) layer by layer to the input layer. Specifically, the back propagation algorithm first calculates the derivative of the total loss value with respect to the activation function of the output layer, and then uses the chain rule to pass these derivatives back to each layer of the network layer by layer, thereby obtaining the degree of contribution of the weights and biases of each layer to the total loss value. These gradient information reflects the potential impact direction and magnitude of adjusting each parameter on reducing the total loss value under the current parameter settings.
[0106] In the example of the present invention, the learnable parameters in the dynamic symmetric encoding and decoding network are adjusted according to the gradient using a preset optimizer, and a new total loss value is iteratively calculated. This is based on these gradient information, with the help of a preset optimizer (such as Adam), and in accordance with the specific update rules of the optimizer, the learnable parameters in the network are adjusted so that the parameters move in a direction that can reduce the total loss value. After the parameter update is completed, the updated model is used again to perform forward propagation calculations on the training data to obtain a new total loss value. This process is repeated continuously, and each iteration calculates a new gradient based on the new parameters, and then updates the parameters and calculates a new total loss value.
[0107] In the examples of the present invention, the iteration is stopped when the new total loss value is less than the preset loss value to obtain the optimized generation model, which means that after updating the learnable parameters of the dynamic symmetric codec network in each iteration and calculating the new total loss value, the new total loss value is checked to see if it is less than the preset loss value. The preset loss value is a preset minimum threshold, which means that the error of the model is small enough to fit the training data well. The dynamic symmetric codec network model that is less than the preset loss value is regarded as an optimized generation model. The model has shown good performance on the training data, can generate high-quality images in tasks such as multi-resolution denoising, meets the image quality requirements in practical applications, and can be used for various subsequent image processing and analysis tasks.
[0108] In this example, the method of evaluating and optimizing a generative model using a preset multi-scale discriminator offers significant advantages. It analyzes images at different resolutions, simulating human vision and comprehensively assessing image structure and detail, avoiding the limitations of a single scale. By weighting the loss values for different resolutions, the importance of each scale can be flexibly adjusted, such as increasing the weight of the high-resolution loss to capture fine details. Iteratively optimizing the learnable parameters allows the model to gradually reduce the difference between the generated and real images, improving adaptability and robustness. For example, in the healthcare field, this method can effectively improve the quality of facial image synthesis. In a medical aesthetic simulation scenario, a doctor first obtains a patient's original facial image and adds simulated noise. A preset multi-scale discriminator is then used to calculate a loss value for the multi-resolution denoised image. The low-resolution layer focuses on the symmetry of the overall facial contour and the proportions of the facial features, while the high-resolution layer focuses on the authenticity of skin texture, pore details, and other details. A total loss is calculated using preset weights, focusing on the accuracy of high-resolution details. The parameters of the dynamic symmetry encoding and decoding network are then iteratively optimized based on the total loss. After multiple rounds of training, the model can generate highly realistic post-operative simulated facial images that preserve the patient's facial features while presenting a natural and aesthetically pleasing cosmetic effect.
[0109] S5. Generate a maximum resolution replacement image corresponding to the source image using the optimized generation model.
[0110] In the example of the present invention, the use of the optimized generation model to generate the maximum resolution replacement image corresponding to the source image refers to applying the optimized dynamic symmetric encoding and decoding network model to the source image, and through complex feature extraction, denoising and detail enhancement processing within the model, finally outputting a high-quality replacement image with the same maximum resolution as the source image.
[0111] In an embodiment of the present invention, the step of generating a maximum resolution replacement image corresponding to the source image using the optimized generation model includes:
[0112] Extracting image fusion features of the source image using the optimized generation model;
[0113] Iteratively denoising the image fusion features using the optimized generative model to obtain an iterative denoised image;
[0114] The iterative denoised image is upsampled to generate a maximum resolution replacement image corresponding to the source image.
[0115] In the example of the present invention, the image fusion features of the source image are extracted by using the optimized generative model, which is to downsample the image through the encoder part, extract multi-scale feature representations layer by layer, and transform the time step embedding vector through a multi-layer perceptron to obtain a scaling factor, and then weight these features through the attention mechanism to obtain an intermediate feature vector. Subsequently, the model initializes the noise vector and adds it to the intermediate feature vector to simulate the noise interference in the real world to generate an updated feature vector. These updated feature vectors are then sent back to the codec network for further processing until the preset number of iterations is reached. Ultimately, the fused feature vector output by the model integrates the multi-scale features of the image, the attention weighted information, and the influence of noise interference.
[0116] In the examples of the present invention, the iterative denoising of the image fusion features using the optimized generative model to obtain an iterative denoising image is to iteratively denoise the image fusion features using the optimized generative model to obtain an iterative denoising image, which means that the model adopts an iterative optimization strategy to gradually reduce the noise components in the image fusion features through repeated denoising processes. In each iteration, the model first predicts the noise distribution in the fusion features, and then subtracts the predicted noise from the features to generate updated denoising features. These updated features are then used as input for the next iteration to continue noise prediction and subtraction. As the number of iterations increases, the noise in the image features is gradually removed, while retaining and enhancing the important details and structural information of the image. Ultimately, after multiple iterations, the iterative denoising image output by the model has less noise interference and more clearly shows the original content and details of the image, thereby improving the visual quality of the image and the performance of subsequent processing tasks.
[0117] In the examples of the present invention, the upsampling of the iterative denoising image to generate the maximum resolution replacement image corresponding to the source image is the upsampling of the iterative denoising image to generate the maximum resolution replacement image corresponding to the source image, which means that in the image processing process, the image feature map obtained through the iterative denoising process is enlarged to the resolution size of the original image through a specific upsampling technology. This process is usually completed by the network structure of the decoder part, which uses the denoised feature map to gradually increase the spatial resolution of the image through a series of upsampling operations (such as transposed convolution, bilinear interpolation, etc.). While upsampling, the network also combines the multi-scale features extracted in the encoder to restore the details and texture information of the image. Ultimately, the generated replacement image achieves the same maximum resolution as the original source image while maintaining the quality of the denoised image, so that it can be used as a high-quality replacement image for the original image.
[0118] In the examples of the present invention, the process of using the optimized generative model to generate the maximum resolution replacement image corresponding to the source image fully utilizes the learning ability and denoising effect of the optimized generative model at multiple resolutions, which can effectively remove the noise in the source image while retaining and enhancing the details and structural information of the image, so that the generated replacement image is clearer, more natural and richer in details in visual effect. For example, in online financial services in the field of financial technology, customers need to authenticate their identities through facial recognition, but the original image may be noisy or blurred due to poor shooting conditions (such as insufficient light, angle deviation, etc.), affecting the accuracy of recognition. At this time, the optimized generative model can be applied to these source images to generate clear, high-quality maximum resolution replacement images through the denoising and detail enhancement capabilities of the model. These replacement images can more accurately reflect the customer's true facial features, thereby improving the accuracy and reliability of the face recognition system.
[0119] It can be seen that in the above scheme, for the target business, the source image and the replacement image are first obtained, and the source image and the replacement image are preprocessed and feature extracted to obtain a feature vector; the noise vector is initialized, and the noise vector is combined with the feature vector through multi-level injection to obtain a fusion feature; the fusion feature is iteratively denoised using a preset dynamic symmetric codec network model to generate a multi-resolution denoised image; the multi-resolution denoised image is quality evaluated using a preset multi-scale discriminator, and the learnable parameters in the dynamic symmetric codec network model are adjusted according to the quality evaluation result to obtain an optimized generation model; the optimized generation model is used to generate a maximum resolution replacement image corresponding to the source image.
[0120] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0121] In one embodiment, a device for replacing a face based on a dynamic symmetric coding network is provided. The device for replacing a face based on a dynamic symmetric coding network corresponds to the method for replacing a face based on a dynamic symmetric coding network in the above embodiment. Figure 5 As shown, the face replacement device based on dynamic symmetric coding and decoding network includes a feature extraction module 101, a feature fusion module 102, a denoising module 103, an updating module 104, and a generating module 105. The functional modules are described in detail as follows:
[0122] A feature extraction module 101 is configured to obtain a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image, and obtain a feature vector;
[0123] A feature fusion module 102 is used to initialize a noise vector and combine the noise vector with the feature vector through multi-level injection to obtain a fused feature;
[0124] A denoising module 103 is configured to iteratively denoise the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image;
[0125] An updating module 104 is configured to perform a quality assessment on the multi-resolution denoised image using a preset multi-scale discriminator, and adjust the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model;
[0126] The generation module 105 is configured to generate a maximum resolution replacement image corresponding to the source image using the optimized generation model.
[0127] In one embodiment, the feature extraction module 101 is specifically configured to:
[0128] cropping the source image and the replacement image to obtain a cropped source image and a cropped replacement image,
[0129] Unifying the resolutions of the cropped source image and the cropped replacement image to obtain a standard source image and a standard replacement image;
[0130] Multi-layer convolution and pooling are performed on the standard source image and the standard replacement image to obtain a feature vector.
[0131] In one embodiment, the feature fusion module 102 is specifically configured to:
[0132] Determining a time step embedding vector corresponding to the feature vector;
[0133] Performing a multi-layer perceptron transformation on the time step embedding vector to generate a scaling factor;
[0134] Performing attention scaling on the scaling factor to obtain an attention weight;
[0135] Multiplying the attention weight by the feature vector to obtain an intermediate feature vector;
[0136] Initializing a noise vector, and adding the noise vector and the intermediate eigenvector vector by vector to obtain an updated eigenvector;
[0137] The updated feature vector is used as a feature vector, and the step of determining the time step embedding vector corresponding to the feature vector is returned to until a preset number of iterations is reached, and the updated feature vector is used as a fusion feature.
[0138] In one embodiment, the denoising module 103 is specifically configured to:
[0139] Perform feature parsing on the fused features using a preset dynamic symmetric encoding and decoding network model to obtain a feature map and corresponding resolution embedding;
[0140] generating gating weights based on the feature map and the resolution embedding;
[0141] Calculating a weighted sum of the gating weight and the feature map to obtain an output feature map;
[0142] Performing noise prediction on the output feature map using a preset dynamic symmetric encoding and decoding network model;
[0143] The output feature map is subtracted from the noise prediction result, and upsampled using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image.
[0144] In one embodiment, the update module 104 is specifically configured to:
[0145] Using a preset multi-scale discriminator to evaluate the multi-resolution denoised images respectively to determine the loss value of each image;
[0146] Weighting the loss values of the multi-resolution denoised image according to a preset weight to calculate a total loss value;
[0147] The learnable parameters in the dynamic symmetric encoding and decoding network are iteratively optimized according to the total loss value, and the iterative optimization is stopped when the total loss value is less than a preset loss value to obtain an optimized generation model.
[0148] In one embodiment, the update module 104 is further configured to:
[0149] Determining the gradient of the total loss value with respect to a learnable parameter in the dynamic symmetric codec network using a back-propagation algorithm;
[0150] Adjust the learnable parameters in the dynamic symmetric encoding and decoding network using a preset optimizer according to the gradient, and iteratively calculate a new total loss value;
[0151] When the new total loss value is less than the preset loss value, the iteration is stopped to obtain the optimized generation model.
[0152] In one embodiment, the generating module 105 is specifically configured to:
[0153] Extracting image fusion features of the source image using the optimized generation model;
[0154] Iteratively denoising the image fusion features using the optimized generative model to obtain an iterative denoised image;
[0155] The iterative denoised image is upsampled to generate a maximum resolution replacement image corresponding to the source image.
[0156] The present invention provides a face replacement device based on a dynamic symmetric coding and decoding network. The device first obtains a source image and a replacement image, performs preprocessing and feature extraction on the source image and the replacement image, and obtains a feature vector; initializes a noise vector, combines the noise vector with the feature vector through multi-level injection, and obtains a fused feature; uses a preset dynamic symmetric coding and decoding network model to iteratively denoise the fused feature to generate a multi-resolution denoised image; uses a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, and adjusts the learnable parameters in the dynamic symmetric coding and decoding network model according to the quality assessment result to obtain an optimized generation model; and uses the optimized generation model to generate a maximum resolution replacement image corresponding to the source image.
[0157] Regarding the specific limitations of the network face replacement device based on dynamic symmetric coding and decoding, please refer to the limitations of the network face replacement method based on dynamic symmetric coding and decoding above, and will not be repeated here. The various modules in the above-mentioned network face replacement device based on dynamic symmetric coding and decoding can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0158] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a network face replacement method based on dynamic symmetric encoding and decoding.
[0159] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a method for face replacement based on dynamic symmetric encoding and decoding.
[0160] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0161] Acquire a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image to obtain a feature vector;
[0162] Initializing a noise vector, and combining the noise vector with the feature vector through multi-level injection to obtain a fusion feature;
[0163] Iteratively denoising the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image;
[0164] Using a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, and adjusting the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model;
[0165] The optimized generation model is used to generate a maximum resolution replacement image corresponding to the source image.
[0166] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0167] Acquire a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image to obtain a feature vector;
[0168] Initializing a noise vector, and combining the noise vector with the feature vector through multi-level injection to obtain a fusion feature;
[0169] Iteratively denoising the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image;
[0170] Using a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, and adjusting the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model;
[0171] The optimized generation model is used to generate a maximum resolution replacement image corresponding to the source image.
[0172] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0173] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0174] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0175] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.
[0176] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A face replacement method based on dynamic symmetric coding and decoding network, characterized in that: include: Acquire a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image to obtain a feature vector; Initializing a noise vector, and combining the noise vector with the feature vector through multi-level injection to obtain a fusion feature; Iteratively denoising the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image; Using a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, and adjusting the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model; The optimized generation model is used to generate a maximum resolution replacement image corresponding to the source image.
2. The face replacement method based on dynamic symmetric coding and decoding network according to claim 1, characterized in that: The preprocessing and feature extraction of the source image and the replacement image to obtain a feature vector includes: Cropping the source image and the replacement image to obtain a cropped source image and a cropped replacement image; Unifying the resolutions of the cropped source image and the cropped replacement image to obtain a standard source image and a standard replacement image; Multi-layer convolution and pooling are performed on the standard source image and the standard replacement image to obtain a feature vector.
3. The face replacement method based on dynamic symmetric coding and decoding network according to claim 1, characterized in that: The initialization noise vector is combined with the feature vector by multi-level injection to obtain a fusion feature, including: Determining a time step embedding vector corresponding to the feature vector; Performing a multi-layer perceptron transformation on the time step embedding vector to generate a scaling factor; Performing attention scaling on the scaling factor to obtain an attention weight; Multiplying the attention weight by the feature vector to obtain an intermediate feature vector; Initializing a noise vector, and performing vector-by-vector addition on the noise vector and the intermediate feature vector to obtain an updated feature vector; The updated feature vector is used as a feature vector, and the step of determining the time step embedding vector corresponding to the feature vector is returned to until a preset number of iterations is reached, and the updated feature vector is used as a fusion feature.
4. The face replacement method based on dynamic symmetric coding and decoding network according to claim 1, characterized in that: The iterative denoising of the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image includes: Perform feature parsing on the fused features using a preset dynamic symmetric encoding and decoding network model to obtain a feature map and corresponding resolution embedding; generating gating weights based on the feature map and the resolution embedding; Calculating a weighted sum of the gating weight and the feature map to obtain an output feature map; Performing noise prediction on the output feature map using a preset dynamic symmetric encoding and decoding network model; The output feature map is subtracted from the noise prediction result, and upsampled using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image.
5. The face replacement method based on dynamic symmetric coding and decoding network according to claim 1, characterized in that: The method of using a preset multi-scale discriminator to perform quality assessment on the multi-resolution denoised image, and adjusting the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model includes: Using a preset multi-scale discriminator to evaluate the multi-resolution denoised images respectively to determine the loss value of each image; Weighting the loss values of the multi-resolution denoised image according to a preset weight to calculate a total loss value; The learnable parameters in the dynamic symmetric encoding and decoding network are iteratively optimized according to the total loss value, and the iterative optimization is stopped when the total loss value is less than a preset loss value to obtain an optimized generation model.
6. The face replacement method based on dynamic symmetric coding and decoding network according to claim 5, characterized in that: The iterative optimization of the learnable parameters in the dynamic symmetric encoding and decoding network according to the total loss value, and stopping the iterative optimization when the total loss value is less than a preset loss value to obtain an optimized generation model, includes: Determining the gradient of the total loss value with respect to a learnable parameter in the dynamic symmetric codec network using a back-propagation algorithm; Adjust the learnable parameters in the dynamic symmetric encoding and decoding network using a preset optimizer according to the gradient, and iteratively calculate a new total loss value; When the new total loss value is less than the preset loss value, the iteration is stopped to obtain the optimized generation model.
7. The face replacement method based on dynamic symmetric coding and decoding network according to claim 1, characterized in that: The step of generating a maximum resolution replacement image corresponding to the source image by using the optimized generation model includes: Extracting image fusion features of the source image using the optimized generation model; Iteratively denoising the image fusion features using the optimized generative model to obtain an iterative denoised image; The iterative denoised image is upsampled to generate a maximum resolution replacement image corresponding to the source image.
8. A face replacement device based on dynamic symmetric coding and decoding network, characterized in that: include: a feature extraction module, configured to obtain a source image and a replacement image, perform preprocessing and feature extraction on the source image and the replacement image, and obtain a feature vector; A feature fusion module is used to initialize a noise vector and combine the noise vector with the feature vector through multi-level injection to obtain a fused feature; A denoising module, configured to iteratively denoise the fused features using a preset dynamic symmetric encoding and decoding network model to generate a multi-resolution denoised image; An updating module is configured to perform a quality assessment on the multi-resolution denoised image using a preset multi-scale discriminator, and adjust the learnable parameters in the dynamic symmetric encoding and decoding network model according to the quality assessment result to obtain an optimized generation model; A generation module is used to generate a maximum resolution replacement image corresponding to the source image using the optimized generation model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the face replacement method based on dynamic symmetric coding and decoding network is implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for face replacement based on a dynamic symmetric coding and decoding network as described in any one of claims 1 to 7 is implemented.