Human face replacement method and device based on artificial intelligence, computer equipment and medium
By combining a hierarchical feature extractor with an adaptive feature alignment layer, the problem of insufficient image quality in existing face replacement techniques is solved, high-fidelity face replacement is achieved, the realism and naturalness of the image are improved, and the robustness of the identity verification system is enhanced.
Patent Information
- Application Number
- CN202510685243.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-16
AI Technical Summary
Existing face replacement technology has bottlenecks in generating image quality and improving resolution. Especially in remote identity verification scenarios in the financial field, the generated replacement images often suffer from local texture loss or global structural dislocation, resulting in an increased misjudgment rate of the identity verification system.
A method based on a hierarchical feature extractor and an adaptive feature alignment layer is adopted to achieve collaborative optimization of multi-scale features through the combination of cross-modal attention operation and diffusion model to generate high-quality face replacement images.
It improves the realism and naturalness of face replacement images, enhances the robustness of the identity verification system, reduces the misjudgment rate, and meets the needs of high-precision fields.
Smart Images

Figure CN120656220A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology and can be applied to fields such as financial technology and digital medicine, and in particular to methods, devices, computer equipment and storage media based on artificial intelligence for face replacement. Background Art
[0002] In traditional face replacement technology, the quality and resolution of generated images have long faced bottlenecks. Existing methods (such as DiffFace and DiffSwap) based on diffusion models have achieved breakthroughs in image generation, but they generally lack the ability to explicitly model multi-scale features during the inverse denoising process. Specifically, such methods often use single-level features or global attention mechanisms for feature fusion, resulting in insufficient coordinated optimization of the global structure and local details of the face. This, in turn, makes the generated face replacement images poorly realistic and natural, resulting in low overall quality.
[0003] This technical flaw is particularly prominent in remote identity verification scenarios in the financial sector. For example, in the online account opening process of insurance institutions, users need to use face replacement technology to generate facial images that simulate different scenarios to test the robustness of the identity verification system. However, due to insufficient multi-scale feature modeling in existing methods, the generated replacement images often suffer from local texture loss (such as blurred eye details) or global structural misalignment (such as distorted facial contours), resulting in an increased false positive rate in the identity verification system. If the simulated image submitted by the customer in a dark environment is misjudged by the system as not being the customer due to distorted details, the account opening process will be directly interrupted, resulting in a decline in customer experience and loss of business efficiency. Such problems not only limit the implementation of face replacement technology in financial security scenarios, but also exacerbate the contradiction between the demand for high-precision face generation and technical capabilities.
[0004] Therefore, there is an urgent need for a face replacement technology that can achieve collaborative optimization of global structure and local details to generate high-quality face replacement images and meet the stringent requirements of face replacement in high-precision fields. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to propose an artificial intelligence-based face replacement method, apparatus, computer device and storage medium to solve the technical problem that the existing diffusion model-based face replacement technology leads to insufficient coordinated optimization of the global structure and local details of the face, thereby resulting in low quality of the generated face replacement image.
[0006] In a first aspect, a face replacement method based on artificial intelligence is provided, comprising:
[0007] Get the input source face image and target face image;
[0008] Preprocessing the source facial image and the target facial image to obtain a first standard image corresponding to the source facial image and a second standard image corresponding to the target facial image;
[0009] Performing feature extraction on the first standard image and the second standard image based on a preset hierarchical feature extractor to obtain corresponding first multi-level feature pyramids and second multi-level feature pyramids;
[0010] Performing a cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on a preset adaptive feature alignment layer to obtain corresponding aligned features;
[0011] Injecting the alignment features into a noise prediction network in a preset diffusion model to obtain an improved target noise prediction network;
[0012] Performing conditional denoising processing based on the target noise prediction network to generate a corresponding intermediate image;
[0013] The intermediate image and the target face image are fused based on a preset fusion strategy to generate a corresponding face replacement image.
[0014] In a second aspect, a face replacement device based on artificial intelligence is provided, comprising:
[0015] A first acquisition module is used to acquire an input source face image and a target face image;
[0016] a preprocessing module, configured to preprocess the source facial image and the target facial image to obtain a first standard image corresponding to the source facial image and a second standard image corresponding to the target facial image;
[0017] an extraction module, configured to perform feature extraction on the first standard image and the second standard image based on a preset hierarchical feature extractor to obtain a corresponding first multi-level feature pyramid and a second multi-level feature pyramid;
[0018] an operation module, configured to perform a cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on a preset adaptive feature alignment layer to obtain corresponding alignment features;
[0019] A processing module, configured to inject the alignment features into a noise prediction network in a preset diffusion model to obtain an improved target noise prediction network;
[0020] A generation module, configured to perform conditional denoising processing based on the target noise prediction network to generate a corresponding intermediate image;
[0021] The fusion module is used to fuse the intermediate image with the target face image based on a preset fusion strategy to generate a corresponding face replacement image.
[0022] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned artificial intelligence-based face replacement method are implemented.
[0023] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned artificial intelligence-based face replacement method are implemented.
[0024] In the scheme implemented by the above-mentioned artificial intelligence-based face replacement method, device, computer equipment and storage medium, the input source face image and target face image are first obtained; then the source face image and the target face image are preprocessed to obtain a first standard image corresponding to the source face image, and a second standard image corresponding to the target face image; then, based on a preset hierarchical feature extractor, feature extraction is performed on the first standard image and the second standard image to obtain corresponding first multi-level feature pyramid and second multi-level feature pyramid; subsequently, based on a preset adaptive feature alignment layer, a cross-modal attention operation is performed on the first multi-level feature pyramid and the second multi-level feature pyramid to obtain corresponding alignment features; and the alignment features are injected into the noise prediction network in the preset diffusion model to obtain an improved target noise prediction network; further, conditional denoising processing is performed based on the target noise prediction network to generate a corresponding intermediate image; finally, the intermediate image and the target face image are fused based on a preset fusion strategy to generate a corresponding face replacement image. This application performs face replacement processing related to the input source face image and target face image through the combined use of a hierarchical feature extractor, an adaptive feature alignment layer, a diffusion model, a target noise prediction network and a fusion strategy. By deeply integrating multi-scale feature extraction and a dynamic alignment mechanism into the reverse process of the diffusion model, high-fidelity face replacement is achieved, overcoming the problem of insufficient collaborative optimization of the global structure and local details of the face in traditional methods, and being able to generate more realistic and natural face replacement results, effectively improving the quality of the generated face replacement image. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0026] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0027] Figure 2 is a flowchart of an embodiment of an artificial intelligence-based face replacement method according to the present application;
[0028] Figure 3 1 is a schematic structural diagram of an embodiment of an artificial intelligence-based face replacement device according to the present application;
[0029] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0031] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0032] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0033] like Figure 1As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0034] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0035] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0036] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0037] It should be noted that the artificial intelligence-based face replacement method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the artificial intelligence-based face replacement device is generally set in the server / terminal device.
[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0039] Continue to refer Figure 2, shows a flow chart of an embodiment of the artificial intelligence-based face replacement method according to the present application. According to different needs, the order of the steps in the flow chart can be changed, and some steps can be omitted. The artificial intelligence-based face replacement method provided in the embodiment of the present application can be applied to any scenario that requires face replacement processing, and the artificial intelligence-based face replacement method can be applied to products in these scenarios, for example, in face replacement processing scenarios in the financial field and the medical field. The artificial intelligence-based face replacement method includes the following steps:
[0040] Step S201: Obtain input source face image and target face image.
[0041] In this embodiment, the artificial intelligence-based face replacement method is run on an electronic device (eg Figure 1 The server / terminal device shown in the figure) can obtain the input source face image and target face image through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or developed in the future. The executive subject of this application is specifically a face replacement processing system, which can be referred to as the system for short. The above-mentioned source face image and target face image are data input by the user according to the actual face replacement needs. The source face image (which can be referred to as the source image for short) contains the face to be replaced. The target face image (which can be referred to as the target image for short) contains the face to be replaced and the background.
[0042] This application proposes an innovative hierarchical feature-aware diffusion framework that achieves high-fidelity face replacement by deeply integrating multi-scale feature extraction and dynamic alignment mechanisms into the reverse process of the diffusion model. The core breakthroughs of this technology are: 1. Volumetric feature modeling: Using 3D convolutional autoencoders (3D-CAE) to construct a hierarchical feature space, for the first time, multi-scale feature decoupling from macro facial contours to micro skin textures is achieved in the diffusion model; 2. Dynamic feature adaptation: Through a cross-level attention mechanism, pixel-level alignment of source identity features and target posture / expression is achieved, resolving the contradiction between identity preservation and posture adaptation in traditional methods; 3. Inverse process enhancement: Introducing feature-aware jump connections in the denoising U-Net, the generation process is simultaneously guided by the high-level semantics of the diffusion model and constrained by local detail features.
[0043] The system architecture of this application includes the following core modules: 1. Preprocessing module: uses 3DDFA for facial detection and key point alignment to generate standardized input; 2. Hierarchical feature extractor: composed of 3D-CAE, processes source face images and target face images, and outputs a multi-level feature pyramid; 3. Adaptive feature alignment layer: receives the output of the hierarchical feature extractor and generates aligned features through cross-attention; 4. Enhanced reverse process (ERP): injects the aligned features into the U-Net of the diffusion model, performs conditional denoising to generate intermediate results; 5. Post-processing module: fuses the intermediate image with the target face image background through Poisson fusion, outputs the final result, and uses it as the face replacement image.
[0044] The face replacement method extracted in this application has potential application value in the insurance and medical fields, and can be used in scenarios such as identity verification, privacy protection, and data enhancement. The following are specific examples in two fields, including scenario descriptions of source and target images.
[0045] 1. Insurance field: customer identity authentication and claims processing. Application scenario: During the insurance claims process, insurance companies need to verify the identity of customers to ensure the authenticity of the claims application. Face replacement technology can be used to generate realistic customer facial images for testing or verifying the robustness of the identity authentication system. Source face image and target face image example: Source face image: Scenario: Customer facial photo obtained from the insurance company's customer database. Features: High resolution, clear, frontal shot, used as a benchmark image for identity authentication. Target face image: Scenario: Simulates customer facial images in different scenarios (such as low light, blur, partial occlusion, etc.). Features: May contain noise, blur or occlusion, used to test the performance of the identity authentication system under complex conditions. The role of face replacement: Replace the customer's face in the source face image with the target face image to generate a realistic synthetic image, which is used to test the adaptability of the identity authentication system to different scenarios. Help insurance companies optimize identity authentication algorithms and improve the accuracy and robustness of the system under complex conditions.
[0046] 2. Medical field: privacy protection and data enhancement. Application scenarios: In medical image analysis, face replacement technology can be used to protect the privacy of patients while generating realistic synthetic data for model training. Examples of source face images and target face images: Source face image: Scenario: CT scan or MRI image of the patient's face. Features: Contains the patient's facial structure information, but may have low resolution due to limitations of medical equipment. Target face image: Scenario: High-resolution face image (such as healthy faces in public datasets). Features: Clear, high-resolution, used to replace the patient's face in medical images to protect privacy. The role of face replacement: Replace the high-resolution face in the target face image with the source face image to generate a synthetic image, which not only retains the anatomical structure information in the medical image but also protects the patient's privacy. The generated synthetic data can be used to train medical image analysis models (such as disease diagnosis models) to improve the generalization ability of the model.
[0047] Step S202 : pre-processing the source facial image and the target facial image to obtain a first standard image corresponding to the source facial image and a second standard image corresponding to the target facial image.
[0048] In this embodiment, the above-mentioned preprocessing of the source facial image and the target facial image to obtain the first standard image corresponding to the source facial image and the second standard image corresponding to the target facial image is specifically implemented. This application will provide further details in subsequent specific embodiments and will not be elaborated on here.
[0049] Step S203 : performing feature extraction on the first standard image and the second standard image based on a preset hierarchical feature extractor to obtain a corresponding first multi-level feature pyramid and a second multi-level feature pyramid.
[0050] In this embodiment, the above-mentioned specific implementation process of performing feature extraction on the first standard image and the second standard image based on the preset hierarchical feature extractor to obtain the corresponding first multi-level feature pyramid and second multi-level feature pyramid will be further described in detail in the subsequent specific embodiments of this application and will not be elaborated on here.
[0051] Step S204: performing a cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on a preset adaptive feature alignment layer to obtain corresponding aligned features.
[0052] In this embodiment, the above-mentioned specific implementation process of performing cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on the preset adaptive feature alignment layer to obtain corresponding alignment features will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0053] Step S205 : injecting the alignment features into the noise prediction network in the preset diffusion model to obtain an improved target noise prediction network.
[0054] In this embodiment, the above-mentioned diffusion model can specifically adopt a U-Net-based model. U-Net is a classic convolutional neural network architecture that is widely used in diffusion models as the core network for noise prediction. The basic structure of U-Net includes a contraction path and a symmetrical expansion path. The contraction path captures contextual information through multiple downsampling operations, while the expansion path combines low-level features and high-level features through upsampling operations to achieve accurate pixel-level segmentation. This U-shaped structural design enables it to efficiently utilize limited labeled samples and execute quickly on modern GPUs.
[0055] In the diffusion model, U-Net is a core component used to predict noise, that is, to predict the added noise from the current noise image and time step t. The input of the original U-Net is the noise image and time step, and the output is the predicted noise.
[0056] The improved noise prediction network (i.e., target noise prediction network) adds alignment features to the original U-Net. The alignment features are generated by the adaptive feature alignment layer and contain the alignment information of the source and target face images.
[0057] Specifically, the improved noise prediction network formula is:
[0058]
[0059] Among them, U-Net(x t ,t) is the output of the original U-Net. AdaIN(A l ,f l ) is the alignment feature A through adaptive instance normalization (AdaIN) l Injected into the lth layer decoding feature f of U-Net l middle. It refers to the fusion of alignment features at all levels.
[0060] In addition, AdaIN's functions include:
[0061]
[0062] Where μ(f) and σ(f) are the mean and standard deviation of feature f. μ(a) and σ(a) are the mean and standard deviation of the aligned feature σ(a). AdaIN transfers the style of U-Net features using the statistical information (mean and standard deviation) of the aligned features, thereby integrating the identity information of the source face image and the pose information of the target face image into the noise prediction process.
[0063] Step S206: Perform conditional denoising processing based on the target noise prediction network to generate a corresponding intermediate image.
[0064] In this embodiment, the above-mentioned conditional denoising process, also referred to as iterative execution of the denoising process, includes the following formula:
[0065]
[0066] Among them, x t is the noise image at step t, α t is the noise scheduling parameter, t is the current denoising step, which is used to control the step size of the denoising process, is the noise standard deviation, and z is the random noise.
[0067] Step S207 : fusing the intermediate image with the target face image based on a preset fusion strategy to generate a corresponding face replacement image.
[0068] In this embodiment, the above-mentioned fusion strategy can specifically adopt a strategy based on illumination-aware Poisson blending technology. Illumination-aware Poisson blending is an advanced image fusion technology that aims to seamlessly blend a specific area of a source image (such as a face) into the background of a target image while maintaining illumination consistency. The specific implementation process of fusing the intermediate image with the target face image based on the preset fusion strategy to generate the corresponding face replacement image will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0069] This application first obtains an input source face image and a target face image; then preprocesses the source face image and the target face image to obtain a first standard image corresponding to the source face image, and a second standard image corresponding to the target face image; then, based on a preset hierarchical feature extractor, feature extraction is performed on the first standard image and the second standard image to obtain corresponding first multi-level feature pyramids and second multi-level feature pyramids; subsequently, based on a preset adaptive feature alignment layer, a cross-modal attention operation is performed on the first multi-level feature pyramid and the second multi-level feature pyramid to obtain corresponding alignment features; and the alignment features are injected into the noise prediction network in the preset diffusion model to obtain an improved target noise prediction network; further, conditional denoising is performed based on the target noise prediction network to generate a corresponding intermediate image; finally, the intermediate image and the target face image are fused based on a preset fusion strategy to generate a corresponding face replacement image. This application performs face replacement processing related to the input source face image and target face image through the combined use of a hierarchical feature extractor, an adaptive feature alignment layer, a diffusion model, a target noise prediction network and a fusion strategy. By deeply integrating multi-scale feature extraction and a dynamic alignment mechanism into the reverse process of the diffusion model, high-fidelity face replacement is achieved, overcoming the problem of insufficient collaborative optimization of the global structure and local details of the face in traditional methods, and being able to generate more realistic and natural face replacement results, effectively improving the quality of the generated face replacement image.
[0070] In some optional implementations, step S202 includes the following steps:
[0071] Call the preset preprocessing module.
[0072] In this embodiment, the pre-processing module is a module having the function of performing face detection and key point alignment using face alignment technology to generate standardized input.
[0073] Get the preset face alignment strategy.
[0074] In this embodiment, the facial alignment strategy is specifically a strategy based on 3DDFA (3D Dense Face Alignment) technology. 3DDFA is a facial alignment technology based on a 3D model that can accurately detect facial key points and adjust posture.
[0075] Based on the preprocessing module, the face alignment strategy is used to perform face detection and key point alignment on the source face image and the target face image to obtain a processed source face image and a processed target face image.
[0076] In this embodiment, the above-mentioned preprocessing module can be used to utilize a facial alignment strategy to perform facial detection and key point alignment on the above-mentioned source face image and target face image, that is, to detect the positions of key points such as facial contours, eyes, nose, and mouth, and adjust the faces of the two images to a unified posture and position to generate a standardized input, thereby obtaining a processed source face image (that is, a standardized first standard image) and a processed target face image (that is, a standardized second standard image).
[0077] The processed source face image is used as the first standard image, and the processed target face image is used as the second standard image.
[0078] In this embodiment, 3DDFA technology is used to perform facial detection and key point alignment on the source and target facial images to generate a standardized input image. This step aims to align facial images of varying poses, expressions, and positions to a uniform standard, providing a good foundation for subsequent feature extraction and processing.
[0079] The present application calls a preset preprocessing module; then obtains a preset facial alignment strategy; then, based on the preprocessing module, uses the facial alignment strategy to perform facial detection and key point alignment on the source facial image and the target facial image to obtain a processed source facial image and a processed target facial image; subsequently, the processed source facial image is used as the first standard image, and the processed target facial image is used as the second standard image. The present application performs facial detection and key point alignment on the source facial image and the target facial image based on the combination of the preprocessing module and the facial alignment strategy, thereby achieving efficient and accurate preprocessing of the source facial image and the target facial image, improving the processing efficiency of image preprocessing, and ensuring the standardization of the obtained first standard image and the second standard image.
[0080] In some optional implementations of this embodiment, step S203 includes the following steps:
[0081] Down-sampling the first standard image and the second standard image is performed to generate corresponding first multi-scale image and second multi-scale image.
[0082] In this embodiment, the first standard image may be downsampled to generate first images at 1 / 2 scale, 1 / 4 scale, and 1 / 8 scale to obtain the first multi-scale image. Furthermore, the second standard image may be downsampled to generate second images at 1 / 2 scale, 1 / 4 scale, and 1 / 8 scale to obtain the second multi-scale image.
[0083] Get the preset stacking strategy.
[0084] In this embodiment, the stacking strategy includes stacking the original image, the 1 / 2 scale, the 1 / 4 scale, and the 1 / 8 scale images in the depth dimension to form a multi-resolution representation corresponding to the original image, that is, a multi-resolution facial block corresponding to the original image.
[0085] Among them, stacked multi-resolution facial patches can capture facial features at different scales, thereby improving the robustness of feature extraction. In addition, by stacking images of different scales, the model can better learn the relationship between facial structures at different scales, which helps to generate more natural replacement results.
[0086] The first multi-scale image and the second multi-scale image are stacked based on the stacking strategy to obtain corresponding first multi-resolution facial blocks and second multi-resolution facial blocks.
[0087] In this embodiment, based on the stacking strategy, the first multi-scale image and the second multi-scale image may be stacked to obtain a first multi-resolution facial block corresponding to the first multi-scale image and a second multi-resolution facial block corresponding to the second multi-scale image.
[0088] Among them, the stacked multi-resolution face block X is the normalized source face image I s and target face image I t The multi-scale representation of the face image is stacked into a four-dimensional tensor, providing richer spatial information for the hierarchical feature extractor. This representation helps the model capture the relationship between facial features at different scales, thereby improving the quality and robustness of face replacement.
[0089] Based on the hierarchical feature extractor, a layer-by-layer convolution operation is performed on the first multi-resolution facial block and the second multi-resolution facial block to obtain a corresponding first multi-level feature pyramid and a second multi-level feature pyramid.
[0090] In this embodiment, the hierarchical feature extractor is a feature extractor composed of a 3D convolutional autoencoder (3D-CAE). The input of the hierarchical feature extractor is a stacked multi-resolution facial block X∈R(H×W×D×3), where D=4 represents four scales (original image, 1 / 2, 1 / 4, 1 / 8 downsampling). H and W refer to the height and width of the image, respectively. The encoding process of 3D-CAE uses a 3D convolution with a kernel size of 3, and extracts features through 5 layers of feature levels (corresponding to spatial resolutions from 64×64 to 4×4) to obtain the source face image I s Multi-level feature pyramid and target face image I t Multi-level feature pyramid Among them, L represents the number of levels of the feature pyramid, L=5, Refers to the feature representation of the source face image at layer l, Refers to the feature representation of the target face image at layer l. The decoder reconstructs the input through a symmetrical structure, and the loss function contains Where E(·) refers to the encoder function, D(·) refers to the decoder function, refers to the adversarial loss, λ adv is the weight of the adversarial loss, and λ adv =0.1. By adopting 3D convolutional autoencoder (3D-CAE) to construct a hierarchical feature space, multi-scale feature decoupling from macro facial contour to micro skin texture can be achieved in the diffusion model.
[0091] The present application generates corresponding first and second multiscale images by downsampling the first and second standard images; then obtains a preset stacking strategy; then stacks the first and second multiscale images based on the stacking strategy to obtain corresponding first and second multiresolution facial blocks; and then performs layer-by-layer convolution operations on the first and second multiresolution facial blocks based on the hierarchical feature extractor to obtain corresponding first and second multilevel feature pyramids. The present application generates corresponding first and second multiscale images by downsampling the first and second standard images; then stacks the first and second multiscale images based on the stacking strategy to obtain corresponding first and second multiresolution facial blocks; and then performs layer-by-layer convolution operations on the first and second multiresolution facial blocks obtained using the hierarchical feature extractor. This allows for efficient and accurate extraction of volumetric features from the first and second standard images, helping to capture the relationship between facial features at different scales, thereby improving the quality and robustness of face replacement.
[0092] In some optional implementations, step S204 includes the following steps:
[0093] The first multi-level feature pyramid and the second multi-level feature pyramid are input into the adaptive feature alignment layer.
[0094] In this embodiment, the adaptive feature alignment (AFA) layer is a technology for cross-modal feature alignment, which aims to align the multi-level features of the source face image and the target face image, thereby providing a consistent feature representation for subsequent image generation tasks (i.e., face replacement).
[0095] In each level of the adaptive feature alignment layer, query-key-value pairs are calculated based on the first multi-level feature pyramid and the second multi-level feature pyramid to obtain corresponding calculation results.
[0096] In this embodiment, the calculation process of the query-key-value pair includes:
[0097]
[0098] in, is the learnable projection matrix, i.e. is a learnable projection matrix used to map features into query, key, and value spaces. Used to convert the source feature Projected into the query space. Key: To target features Projected into the key space. Value: To target features Projected into the value space. In addition, information interaction between source features and target features is achieved through query-key-value pair calculation.
[0099] Generate a corresponding attention weight based on the calculation result.
[0100] In this embodiment, the attention weight can be generated based on the following formula:
[0101]
[0102] Among them, c l Refers to the number of feature channels, which is used to scale similarity to avoid gradient vanishing or exploding. l , w l are the height and width of the features at layer l. Softmax() is used to normalize the similarity into a probability distribution, representing the attention weight of each position in the target feature to each position in the source feature. The goal of the above formula is to generate the attention weight of the target feature to the source feature through Softmax normalization.
[0103] Generate corresponding target features based on the attention weights and the preset residual connection.
[0104] In this embodiment, the generation process of the target feature (i.e., the alignment feature) includes: attention weighting:
[0105] Use the attention weight M l Value V l Weighted: Attention(Q l ,K l ,V l )=Ml V l This step fuses the target feature information into the source feature according to the attention weight. Residual connection: Generation of alignment feature Al:
[0106]
[0107] Among them, the residual connection This is used to preserve the identity information of the source face image, avoiding complete loss of source feature information. This design ensures that the aligned features contain both the pose and expression information of the target image while preserving the identity information of the source image, ensuring the consistency of the generated results.
[0108] The target feature is used as the alignment feature.
[0109] In this embodiment, the adaptive feature alignment layer aligns the multi-level features of the source and target face images using a cross-modal attention mechanism. Attention weights dynamically learn which parts of the target features are most important to which parts of the source features, enabling precise feature alignment. Residual connections ensure that the aligned features retain the identity of the source image, preventing the generated results from being completely biased towards the target image.
[0110] The present application inputs the first multi-level feature pyramid and the second multi-level feature pyramid into the adaptive feature alignment layer, and then in each level of the adaptive feature alignment layer, performs query-key-value pair calculation processing based on the first multi-level feature pyramid and the second multi-level feature pyramid to obtain calculation results, generates attention weights based on the calculation results, and then generates corresponding target features based on the attention weights and preset residual connections and uses them as the required alignment features, so as to achieve efficient and accurate cross-modal attention operations on the first multi-level feature pyramid and the second multi-level feature pyramid, thereby achieving pixel-level alignment of source identity features and target posture / expression through a cross-level attention mechanism, thereby resolving the contradiction between identity preservation and posture adaptation in traditional face replacement methods.
[0111] In some optional implementations, step S207 includes the following steps:
[0112] Based on the target face image, the illumination information of the intermediate image is adjusted to obtain a corresponding designated image.
[0113] In this embodiment, the intermediate image is generated by a denoising process based on a diffusion model. It already has the facial pose and expression of the target image, but may be inconsistent with the background in terms of lighting, color tone, etc. The intermediate image is a preliminary face replacement result. The facial identity information (such as facial features and expression) has been correctly replaced, but the overall lighting and color tone may not match the background of the target image. It is the input for the subsequent lighting-aware Poisson fusion and needs to be further optimized through fusion technology. The target face image refers to the image that contains the original background and the face area to be replaced.
[0114] Illumination information (such as brightness, contrast, and hue) can be extracted by performing illumination estimation on the intermediate image and the target facial image. Specifically, an illumination estimation model (such as an algorithm based on color constancy) or a simple statistical method (such as calculating the average brightness and standard deviation of an image) can be used. The illumination information of the intermediate image is then adjusted to be consistent with the background illumination of the target facial image, thereby obtaining the adjusted designated image.
[0115] A corresponding binary mask is generated based on the specified image.
[0116] In this embodiment, a binary mask can be generated by a face detection algorithm (such as Dlib, MTCNN) to mark the area in the intermediate image that needs to be fused into the target face image (such as the face area).
[0117] Based on the binary mask, the designated image and the target face image are fused in a preset gradient domain to obtain a corresponding fused image.
[0118] In this embodiment, the above-mentioned fusion processing includes: the core idea of Poisson fusion is to perform fusion in the gradient domain to maintain the smoothness and continuity of the image. The specific steps are as follows: 1. Calculate the gradient field of the specified image and the target face image in the mask area. 2. At the mask boundary, force the gradient of the specified image to be consistent with the gradient of the target face image to achieve seamless fusion. 3. Generate a smooth fusion result by solving the Poisson equation. In addition, on the basis of Poisson fusion, illumination perception constraints can also be added. Specifically, in the fusion process, not only the gradient is matched, but also the illumination information (such as brightness and hue) is matched to ensure that the illumination of the fusion area is consistent with the background.
[0119] The fused image is used as the face replacement image.
[0120] In this embodiment, after light-aware Poisson fusion, a final fused image is obtained and used as the face replacement image, wherein the replaced face and the background of the target image are consistent in terms of lighting, color tone, etc.
[0121] This application obtains a designated image by adjusting the illumination information of an intermediate image based on the use of a target facial image, and generates a corresponding binary mask based on the designated image. Then, based on the use of the binary mask, the designated image and the target facial image are fused within a preset gradient domain, and the obtained fused image is used as the desired face replacement image. This application combines illumination information adjustment, binary mask generation, and fusion processing based on a fusion strategy to achieve seamless fusion of the intermediate image into the background of the target facial image while maintaining illumination consistency. This processing method avoids manual parameter adjustment and can automatically generate high-quality replacement results, thereby effectively improving the high quality of the generated face replacement image.
[0122] In some optional implementations of this embodiment, after step S207, the electronic device may further perform the following steps:
[0123] The face replacement image is optimized based on a preset optimization strategy to obtain a corresponding target face replacement image.
[0124] In this embodiment, the above-mentioned specific implementation process of optimizing the face replacement image based on the preset optimization strategy to obtain the corresponding target face replacement image will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0125] Obtain a target output mode corresponding to the target face replacement image.
[0126] In this embodiment, the target output method is not specifically limited and can be determined according to the actual needs of the user, for example, any one of email sending, interface display, message sending, etc. can be used. The user is the designated person who inputs the source face image and the target face image.
[0127] Based on the target output mode, output processing is performed on the target face replacement image.
[0128] In this embodiment, the target face replacement image may be sent to the user according to the selected image output mode to complete the output processing of the target face replacement image.
[0129] After generating a face replacement image, this application automatically optimizes it based on an optimization strategy to improve the quality and naturalness of the generated target face replacement image. Furthermore, by outputting the target face replacement image based on the target output mode obtained for that target face replacement image, the accuracy and quality of the output target face replacement image are effectively ensured, improving the output intelligence of the target face replacement image and ultimately enhancing the user experience.
[0130] In some optional implementations of this embodiment, optimizing the face replacement image based on a preset optimization strategy to obtain a corresponding target face replacement image includes the following steps:
[0131] Perform image sharpening processing on the face replacement image to obtain a corresponding first processed image.
[0132] In this embodiment, the goal of the above-mentioned image sharpening processing is to enhance the details and edges of the image and make facial features clearer. The specific implementation process includes: 1) Gaussian blurring the original image to generate a blurred image; subtracting the original image from the blurred image to obtain high-frequency details (i.e., edge and detail information); superimposing the high-frequency details with the original image at a certain weight (e.g., 1.5 times) to obtain a sharpened image. 2) Using the Laplacian operator (e.g., [[0,1,0], [1,-4,1], [0,1,0]]) to perform a convolution operation on the image to extract edge information; superimposing the edge information with the original image to enhance details and edges.
[0133] De-noising is performed on the first processed image to obtain a corresponding second processed image.
[0134] In this embodiment, the purpose of the denoising process is to remove any noise or artifacts that may remain during the image generation process. The specific implementation process includes: for each pixel in the image, searching for similar pixel blocks in its surrounding neighborhood; calculating weights based on the similarity of the pixel blocks (e.g., Euclidean distance); and performing a weighted average of the similar pixel blocks to obtain the denoised pixel value.
[0135] Performing color correction processing on the second processed image to obtain a corresponding third processed image.
[0136] In this embodiment, the goal of the color correction process is to adjust the color balance of the image to make it appear more natural. The specific implementation process includes: 1) White balance adjustment: Detecting white or gray areas (reference areas) in the image; calculating the RGB channel mean of the reference area to determine the color deviation under the current lighting conditions; adjusting the RGB channel gain of the image to make the RGB mean of the reference area close to neutral (e.g., [255, 255, 255]). 2) Histogram equalization: Calculating the RGB channel histogram of the image; Equalizing the histogram of each channel to expand the dynamic range; Merging the equalized RGB channels to obtain a color-balanced image.
[0137] Perform contrast enhancement processing on the third processed image to obtain a corresponding fourth processed image.
[0138] In this embodiment, the contrast enhancement process aims to increase image contrast and make facial features more distinct. The specific implementation process includes: dividing the image into multiple local regions (e.g., 8×8 blocks); performing histogram equalization on each local region while limiting the magnitude of the contrast enhancement (by clipping the high-frequency portion of the histogram); and using interpolation to smooth the boundaries between the local regions to obtain the final contrast-enhanced image.
[0139] The fourth processed image is used as the target face replacement image.
[0140] This application can automatically and intelligently complete the optimization processing of face replacement images by performing image sharpening, denoising, color correction and contrast enhancement on the face replacement images, thereby significantly improving the quality and naturalness of the generated target face replacement images.
[0141] In some optional implementations, the user information obtained is obtained with the user's consent and complies with relevant laws and policies.
[0142] In addition, any software tools or components not provided by our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0143] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0144] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned face replacement images, the above-mentioned face replacement images can also be stored in a blockchain node.
[0145] The blockchain referred to in this application is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0146] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0147] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0148] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0149] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0150] Further references Figure 3, as a response to the above Figure 2 The present application provides an embodiment of a face replacement device based on artificial intelligence. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0151] like Figure 3 As shown, the artificial intelligence-based face replacement device 300 of this embodiment includes: a first acquisition module 301, a pre-processing module 302, an extraction module 303, an operation module 304, a processing module 305, a generation module 306 and a fusion module 307. Among them:
[0152] A first acquisition module 301 is used to acquire an input source face image and a target face image;
[0153] A preprocessing module 302 is configured to preprocess the source facial image and the target facial image to obtain a first standard image corresponding to the source facial image and a second standard image corresponding to the target facial image;
[0154] An extraction module 303 is configured to perform feature extraction on the first standard image and the second standard image based on a preset hierarchical feature extractor to obtain a corresponding first multi-level feature pyramid and a second multi-level feature pyramid;
[0155] An operation module 304 is configured to perform a cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on a preset adaptive feature alignment layer to obtain corresponding alignment features;
[0156] A processing module 305 is configured to inject the alignment features into a noise prediction network in a preset diffusion model to obtain an improved target noise prediction network;
[0157] A generating module 306 is configured to perform conditional denoising based on the target noise prediction network to generate a corresponding intermediate image;
[0158] The fusion module 307 is configured to fuse the intermediate image with the target face image based on a preset fusion strategy to generate a corresponding face replacement image.
[0159] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based face replacement method in the aforementioned embodiment, and are not repeated here.
[0160] In some optional implementations of this embodiment, the preprocessing module 302 includes:
[0161] Call submodule, used to call the preset preprocessing module;
[0162] A first acquisition submodule is used to acquire a preset face alignment strategy;
[0163] A first processing submodule is configured to perform face detection and key point alignment on the source face image and the target face image using the face alignment strategy based on the preprocessing module to obtain a processed source face image and a processed target face image;
[0164] The first determining submodule is configured to use the processed source face image as the first standard image and the processed target face image as the second standard image.
[0165] In some optional implementations of this embodiment, the extraction module 303 includes:
[0166] a first generating submodule, configured to downsample the first standard image and the second standard image to generate corresponding first multi-scale image and second multi-scale image;
[0167] The second acquisition submodule is used to obtain a preset stacking strategy;
[0168] a second processing submodule, configured to stack the first multi-scale image and the second multi-scale image based on the stacking strategy to obtain corresponding first multi-resolution facial blocks and second multi-resolution facial blocks;
[0169] an operation submodule, configured to perform a layer-by-layer convolution operation on the first multi-resolution facial block and the second multi-resolution facial block based on the hierarchical feature extractor to obtain a corresponding first multi-level feature pyramid and a second multi-level feature pyramid.
[0170] In some optional implementations of this embodiment, the operation module 304 includes:
[0171] An input submodule, configured to input the first multi-level feature pyramid and the second multi-level feature pyramid into the adaptive feature alignment layer;
[0172] a calculation submodule, configured to perform, in each level of the adaptive feature alignment layer, calculation processing of query-key-value pairs based on the first multi-level feature pyramid and the second multi-level feature pyramid to obtain corresponding calculation results;
[0173] A second generating submodule, configured to generate a corresponding attention weight based on the calculation result;
[0174] A third generation submodule is used to generate corresponding target features based on the attention weight and the preset residual connection;
[0175] The second determining submodule is configured to use the target feature as the alignment feature.
[0176] In some optional implementations of this embodiment, the fusion module 307 includes:
[0177] an adjustment submodule, configured to adjust the illumination information of the intermediate image based on the target face image to obtain a corresponding designated image;
[0178] a fourth generating submodule, configured to generate a corresponding binary mask based on the specified image;
[0179] a fusion submodule, configured to fuse the designated image with the target face image in a preset gradient domain based on the binary mask to obtain a corresponding fused image;
[0180] The third determining submodule is configured to use the fused image as the face replacement image.
[0181] In some optional implementations of this embodiment, the artificial intelligence-based face replacement device further includes:
[0182] an optimization module, configured to optimize the face replacement image based on a preset optimization strategy to obtain a corresponding target face replacement image;
[0183] A second acquisition module is used to acquire a target output mode corresponding to the target face replacement image;
[0184] An output module is used to output the target face replacement image based on the target output mode.
[0185] In some optional implementations of this embodiment, the optimization module includes:
[0186] a third processing submodule, configured to perform image sharpening processing on the face replacement image to obtain a corresponding first processed image;
[0187] a fourth processing submodule, configured to perform denoising on the first processed image to obtain a corresponding second processed image;
[0188] a fifth processing submodule, configured to perform color correction processing on the second processed image to obtain a corresponding third processed image;
[0189] a sixth processing submodule, configured to perform contrast enhancement processing on the third processed image to obtain a corresponding fourth processed image;
[0190] The fourth determining submodule is configured to use the fourth processed image as the target face replacement image.
[0191] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0192] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0193] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0194] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for the artificial intelligence-based face replacement method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0195] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions or process data stored in the memory 41, such as computer-readable instructions for executing the artificial intelligence-based face replacement method.
[0196] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0197] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned artificial intelligence-based face replacement method.
[0198] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0199] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A face replacement method based on artificial intelligence, characterized in that: The steps include: Get the input source face image and target face image; Preprocessing the source facial image and the target facial image to obtain a first standard image corresponding to the source facial image and a second standard image corresponding to the target facial image; Performing feature extraction on the first standard image and the second standard image based on a preset hierarchical feature extractor to obtain corresponding first multi-level feature pyramids and second multi-level feature pyramids; Performing a cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on a preset adaptive feature alignment layer to obtain corresponding aligned features; Injecting the alignment features into a noise prediction network in a preset diffusion model to obtain an improved target noise prediction network; Performing conditional denoising processing based on the target noise prediction network to generate a corresponding intermediate image; The intermediate image and the target face image are fused based on a preset fusion strategy to generate a corresponding face replacement image.
2. The artificial intelligence-based face replacement method according to claim 1, characterized in that: The step of preprocessing the source facial image and the target facial image to obtain a first standard image corresponding to the source facial image and a second standard image corresponding to the target facial image specifically includes: Call the preset preprocessing module; Get the preset face alignment strategy; Based on the preprocessing module, using the facial alignment strategy to perform facial detection and key point alignment on the source facial image and the target facial image to obtain a processed source facial image and a processed target facial image; The processed source face image is used as the first standard image, and the processed target face image is used as the second standard image.
3. The artificial intelligence-based face replacement method according to claim 1, characterized in that: The step of extracting features from the first standard image and the second standard image based on a preset hierarchical feature extractor to obtain corresponding first multi-level feature pyramids and second multi-level feature pyramids specifically includes: Downsampling the first standard image and the second standard image to generate corresponding first multi-scale image and second multi-scale image; Get the preset stacking strategy; stacking the first multi-scale image and the second multi-scale image based on the stacking strategy to obtain corresponding first multi-resolution facial blocks and second multi-resolution facial blocks; Based on the hierarchical feature extractor, a layer-by-layer convolution operation is performed on the first multi-resolution facial block and the second multi-resolution facial block to obtain a corresponding first multi-level feature pyramid and a second multi-level feature pyramid.
4. The artificial intelligence-based face replacement method according to claim 1, characterized in that: The step of performing a cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on a preset adaptive feature alignment layer to obtain corresponding alignment features specifically includes: Inputting the first multi-level feature pyramid and the second multi-level feature pyramid into the adaptive feature alignment layer; In each level of the adaptive feature alignment layer, performing query-key-value pair calculation processing based on the first multi-level feature pyramid and the second multi-level feature pyramid to obtain corresponding calculation results; Generate a corresponding attention weight based on the calculation result; Generate corresponding target features based on the attention weight and the preset residual connection; The target feature is used as the alignment feature.
5. The artificial intelligence-based face replacement method according to claim 1, characterized in that: The step of fusing the intermediate image with the target face image based on a preset fusion strategy to generate a corresponding face replacement image specifically includes: Based on the target face image, adjusting the illumination information of the intermediate image to obtain a corresponding designated image; Generating a corresponding binary mask based on the specified image; Based on the binary mask, the designated image and the target face image are fused in a preset gradient domain to obtain a corresponding fused image; The fused image is used as the face replacement image.
6. The artificial intelligence-based face replacement method according to claim 1, characterized in that: After the step of fusing the intermediate image with the target face image based on a preset fusion strategy to generate a corresponding face replacement image, the method further includes: Optimizing the face replacement image based on a preset optimization strategy to obtain a corresponding target face replacement image; Obtaining a target output mode corresponding to the target face replacement image; Based on the target output mode, output processing is performed on the target face replacement image.
7. The artificial intelligence-based face replacement method according to claim 6, characterized in that: The step of optimizing the face replacement image based on a preset optimization strategy to obtain a corresponding target face replacement image specifically includes: Performing image sharpening processing on the face replacement image to obtain a corresponding first processed image; performing denoising on the first processed image to obtain a corresponding second processed image; performing color correction processing on the second processed image to obtain a corresponding third processed image; performing contrast enhancement processing on the third processed image to obtain a corresponding fourth processed image; The fourth processed image is used as the target face replacement image.
8. A face replacement device based on artificial intelligence, characterized in that: include: A first acquisition module is used to acquire an input source face image and a target face image; a preprocessing module, configured to preprocess the source facial image and the target facial image to obtain a first standard image corresponding to the source facial image and a second standard image corresponding to the target facial image; an extraction module, configured to perform feature extraction on the first standard image and the second standard image based on a preset hierarchical feature extractor to obtain a corresponding first multi-level feature pyramid and a second multi-level feature pyramid; an operation module, configured to perform a cross-modal attention operation on the first multi-level feature pyramid and the second multi-level feature pyramid based on a preset adaptive feature alignment layer to obtain corresponding alignment features; A processing module, configured to inject the alignment features into a noise prediction network in a preset diffusion model to obtain an improved target noise prediction network; A generation module, configured to perform conditional denoising processing based on the target noise prediction network to generate a corresponding intermediate image; The fusion module is used to fuse the intermediate image with the target face image based on a preset fusion strategy to generate a corresponding face replacement image.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the artificial intelligence-based face replacement method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based face replacement method according to any one of claims 1 to 7.