Image processing method and apparatus, electronic device, computer-readable storage medium, and computer program product

By employing 3D representation and expression-driven real-time image processing methods, this approach addresses the shortcomings of existing expression transfer technologies in achieving naturalness and realism in real-time scenarios. It enables highly efficient expression transfer, enhancing the user experience in applications such as video calls and virtual reality.

WO2026092771A1PCT designated stage Publication Date: 2026-05-07TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-11-24
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing facial expression transfer technologies cannot achieve highly natural and realistic effects in real-time facial expression transfer scenarios, especially in highly dynamic interactive environments such as video calls and virtual reality applications. Existing methods have significant limitations in terms of facial expression following effect and real-time processing efficiency.

Method used

Employing a real-time image processing method based on 3D representation and expression-driven features, a dynamic and deformable 3D face model is created through 3D reconstruction and expression feature extraction. This captures the detailed geometric structure and subtle facial dynamics of the face, and introduces an expression encoder to achieve expression transfer, ensuring real-time synchronization and naturalness of facial features.

Benefits of technology

It improves the naturalness and realism of expression transfer, enhances the performance and efficiency of real-time processing, meets the needs of modern multimedia applications for high-quality and high-performance expression transfer, and achieves a high degree of naturalness and continuity of facial features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025137170_07052026_PF_FP_ABST
    Figure CN2025137170_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an image processing method and apparatus, and a device. The method comprises: acquiring a first image and a second image, the first image containing a first face of a first object, and the second image containing a second face of a second object; performing three-dimensional reconstruction on the first face to obtain a first face feature of the first face; respectively extracting expression features of the first face and the second face to correspondingly obtain a first expression feature of the first face and a second expression feature of the second face; on the basis of the first expression feature and the second expression feature, performing feature conversion on the first face feature to obtain a second face feature; and decoding the second face feature to obtain a third image.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing methods, apparatuses, electronic devices, computer-readable storage media, and computer program products

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 2024115593770, filed on November 1, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of artificial intelligence technology, and in particular to an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0004] Facial expression transfer is a technology that uses artificial intelligence to accurately map the facial expressions and demeanor features of a source person onto a target person's facial image. It is widely used in fields such as emoji generation, film and television special effects, and virtual anchors.

[0005] However, among related technologies, facial expression transfer technology mainly focuses on two-dimensional image processing and rudimentary three-dimensional modeling techniques. While these methods show promising results in certain application scenarios, they exhibit significant limitations in real-time facial expression transfer and dynamic facial expression capture. Especially in highly dynamic interactive environments, such as video calls and virtual reality applications, these technologies fail to achieve highly natural and realistic effects. Summary of the Invention

[0006] This application provides an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can improve the output quality of a third image obtained through expression transfer.

[0007] The technical solution of this application embodiment is implemented as follows:

[0008] This application provides an image processing method executed by an electronic device. The method includes: acquiring a first image and a second image; the first image includes a first face of a first object, and the second image includes a second face of a second object; performing three-dimensional reconstruction on the first face to obtain a first facial feature of the first face; extracting expression features from the first face and the second face respectively to obtain a first expression feature of the first face and a second expression feature of the second face; performing feature transformation on the first facial feature using the first expression feature and the second expression feature to obtain a second facial feature; and decoding the second facial feature to obtain a third image.

[0009] This application provides an image processing apparatus, comprising: an acquisition module for acquiring a first image and a second image; the first image including a first face of a first object, and the second image including a second face of a second object; a three-dimensional reconstruction module configured to perform three-dimensional reconstruction on the first face to obtain a first facial feature of the first face; an expression feature extraction module configured to extract expression features from the first face and the second face respectively to obtain a first expression feature of the first face and a second expression feature of the second face; a feature conversion module configured to perform feature conversion on the first facial feature using the first expression feature and the second expression feature to obtain a second facial feature; and a decoding module configured to decode the second facial feature to obtain a third image.

[0010] This application provides an electronic device, including: a memory for storing computer-executable instructions; and a processor for executing the computer-executable instructions stored in the memory to implement the image processing method provided in this application.

[0011] This application provides a computer-readable storage medium storing computer-executable instructions for implementing the image processing method provided in this application when executed by a processor.

[0012] This application provides a computer program product including computer-executable instructions stored in a computer-readable storage medium. When the processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, it implements the image processing method provided in this application.

[0013] The embodiments of this application have the following beneficial effects:

[0014] In the process of expression transfer, after obtaining a high-precision 3D facial model of the first face in the first image through 3D reconstruction, and obtaining the first expression feature of the first face and the second expression feature of the second face through expression feature extraction, the 3D facial model of the first face in the first image is transformed using the first and second expression features. That is, the 3D facial model is driven by the first and second expression features, so as to accurately synchronize the expression dynamics of the second face in the second image, ensuring the high naturalness and continuity of facial features during the expression transfer process, thereby improving the output quality of the third image obtained after expression transfer. Attached Figure Description

[0015] Figure 1 is a schematic diagram of the image processing system architecture provided in an embodiment of this application;

[0016] Figure 2 is a schematic diagram of the image processing apparatus provided in an embodiment of this application;

[0017] Figure 3 is a schematic flowchart of an optional image processing method provided in an embodiment of this application;

[0018] Figure 4 is another optional flowchart of the image processing method provided in the embodiments of this application;

[0019] Figure 5 is a schematic diagram illustrating the implementation of the second facial feature provided in an embodiment of this application;

[0020] Figure 6 is a flowchart illustrating the image expression transfer model training method provided in an embodiment of this application;

[0021] Figure 7 is a schematic diagram illustrating the implementation of the first image of the third sample image provided in an embodiment of this application;

[0022] Figure 8 is a schematic diagram illustrating the implementation of the second image of the fourth sample image provided in the embodiments of this application;

[0023] Figure 9 is a schematic diagram illustrating the implementation of the fifth sample image provided in the embodiments of this application;

[0024] Figure 10 is a schematic diagram illustrating the implementation of the loss result provided in an embodiment of this application;

[0025] Figure 11 is another optional flowchart of the image processing method provided in the embodiment of this application;

[0026] Figure 12 is a schematic diagram of the expression-driven process provided in an embodiment of this application;

[0027] Figure 13 is a schematic diagram of the nonlinear transformation of input data by LTW transform with different parameters provided in the embodiments of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments, but it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0030] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0031] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0032] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0033] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0034] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0035] 1) Three-Dimensional (3D) Representation: This refers to 3D representation technology, which involves creating and utilizing 3D models to represent objects, faces, or other shapes in order to accurately capture and display their three-dimensional structure and details in a digital environment. 3D representation technology typically includes multiple aspects such as 3D modeling, scanning, and rendering, and is used in various fields, such as computer graphics, computer vision, virtual reality (VR), and augmented reality (AR). Through 3D representation technology, high-precision representation and interaction of objects or features can be achieved. For example, by creating a 3D model to represent a face, facial structure and details can be accurately captured.

[0036] 2) Expression-Driven: This technique utilizes captured and analyzed facial expression information to control and alter the facial expressions of a 3D facial model. This technology typically involves face recognition and detection, as well as the generation of 3D model animations. First, facial expression information, including facial muscle movements and expression changes, is captured using a camera or other sensors. Then, the captured expression information is mapped onto a 3D facial model, and algorithms control corresponding parts of the model to produce expression changes similar to a real human face. This process is usually performed in real-time to ensure that the 3D facial model's expressions instantly reflect the changes in the target face's facial expressions.

[0037] 3) Generative Adversarial Networks (GANs): These are deep learning models proposed by Ian Goodfellow et al. in 2014. They consist of two neural networks: a generator and a discriminator. The core idea of ​​GANs is to train the generator and discriminator adversarially. The generator learns how to generate realistic data, while the discriminator learns how to distinguish between real data and fake data generated by the generator. GANs are widely used to generate high-quality face images, specifically images of non-existent faces that are visually very similar to real faces.

[0038] One image processing method in related technologies involves acquiring the original facial image, extracting its features, reconstructing the face in a 3D model, and then fusing these images back into the original image to produce a face-swapping effect. However, this method suffers from poor expression tracking. This means that when the original face's expression changes, the generated third image cannot accurately reflect these changes, resulting in an unnatural appearance, especially in dynamic scenes. Another image processing method first utilizes style description text and control conditions to generate latent representations and perform noise reduction using a stable diffusion model, then restores the target style and facial features through face enhancement processing. However, methods based on stable diffusion models are slow and struggle to achieve real-time results in expression transfer requiring real-time processing, limiting their practicality. Yet another image processing method processes the first and second images using 3D reconstruction and fusion techniques to generate natural facial features and shapes. However, this face-swapping is based solely on facial features, missing other detailed facial features, which is a significant drawback in applications requiring consistent facial features across the entire face. Therefore, it is evident that current facial expression transfer technologies primarily focus on 2D image processing and rudimentary 3D modeling. While these methods demonstrate promising performance in certain applications, they exhibit significant limitations in real-time facial expression transfer and dynamic facial expression capture. Particularly in highly dynamic interactive environments, such as video calls and virtual reality applications, these technologies fail to achieve highly natural and realistic results.

[0039] To address at least one of the problems in the aforementioned related technologies, this application proposes a real-time image processing method based on 3D representation and expression-driven techniques. It not only utilizes implicit 3D facial representation to create dynamic and deformable 3D facial models, capturing detailed geometric structures of the face—not just surface textures—but also significantly improves the naturalness of face swapping, making the generated facial images visually indistinguishable from real faces. Furthermore, it introduces an expression encoder that can not only recognize and analyze basic facial expressions but also capture more subtle facial dynamics, such as minute eye movements and slight distortions of the corners of the mouth. This allows for real-time and accurate synchronization of the 3D facial model in terms of expression, precisely reproducing the expression dynamics of the target face while preserving the features of the source face. This precise expression mapping and dynamic synchronization enable this method to demonstrate more realistic face-swapping effects and application potential in real-time expression transfer scenarios, such as video calls, film production, advertising, or virtual reality applications. Therefore, the technical solution of this application not only improves the naturalness and realism of facial expression transfer, but also significantly enhances the performance and efficiency of real-time processing, thereby meeting the urgent need for high-quality and high-performance facial expression transfer technology in modern multimedia applications.

[0040] The following describes exemplary applications of the image processing device (i.e., electronic device) provided in the embodiments of this application. The image processing device provided in the embodiments of this application can be implemented as various types of user terminals capable of image processing, such as laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals. It can also be implemented as a server. The following will describe exemplary applications when the image processing device is implemented as a server.

[0041] Referring to Figure 1, which is a schematic diagram of the architecture of the image processing system 100 provided in an embodiment of this application, in order to support an image editing application, the terminal 400 connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of both.

[0042] Terminal 400 sends an image processing request to server 200, which constitutes the image processing device of this embodiment. Server 200 responds to the image processing request by acquiring a first image and a second image. The first image includes a first face of a first object, and the second image includes a second face of a second object. Then, three-dimensional reconstruction is performed on the first face to obtain first facial features. Expression features are extracted from both the first and second faces to obtain first and second expression features. Next, feature transformation is performed on the first facial features using the first and second expression features to obtain second facial features. Finally, the second facial features are decoded to obtain a third image. The third image is then returned to terminal 400 so that terminal 400 can output the third image or continue further business processing based on the third image.

[0043] In film production and visual effects scenarios, during movie shooting, a well-known actor (the first subject) completes all his scenes, but the director wants to showcase the classic expression of another specific actor (the second subject, possibly an expression capture actor or a deceased actor) in a close-up shot. Image acquisition: The source image (the first image) is a still or 3D scan data of the well-known actor with a neutral expression; the target image (the second image) is a reference video frame of the specific actor making the desired expression. The system performs 3D reconstruction of the well-known actor's face to obtain its unique facial structure (bones, muscles). Then, it extracts facial features (such as upturned corners of the mouth and furrowed brows) from the reference frame of the specific actor and uses these features to drive the 3D model of the well-known actor. In the generated expression transfer image (the third image), the audience sees the face of the well-known actor, but it presents the precise and vivid expression of the specific actor, seamlessly integrating into the plot, greatly reducing the cost and difficulty of using a stand-in to recreate the expression.

[0044] In virtual digital human and live streaming scenarios, the goal is for virtual idols (source objects) to reproduce the rich expressions of real-life anchors (target objects) during live streams, achieving more natural and engaging interactions. Image acquisition: The source image is a preset 3D model of the virtual idol (first facial features); the target image is a real-time video stream of the real-life anchor captured by a camera. The system extracts the second facial expression features of the real-life anchor in real time and immediately maps these features onto the virtual idol's 3D facial features for driving the interaction. During the live stream, the virtual idol (third image) is no longer a stiff, expressionless face, but can make expressions such as smiles, blinks, and surprise identical to those of the real-life anchor in real time. This not only enhances the virtual idol's approachability but also makes its performances and interactions more realistic and believable, greatly improving the user experience. In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal 400 may be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, in-vehicle terminal, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.

[0045] Referring to Figure 2, which is a schematic diagram of the structure of an electronic device 40 provided in an embodiment of this application, the electronic device 40 shown in Figure 2 may be an image processing device, which includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the image processing device are coupled together through a bus system 440. It is understood that the bus system 440 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 440 in Figure 2.

[0046] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0047] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0048] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0049] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0050] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0051] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0052] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, WiFi, and Universal Serial Bus (USB); the presentation module 453 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.); the input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0053] In some embodiments, the apparatus provided in this application can be implemented in software. FIG2 shows an image processing apparatus 455 stored in memory 450, which can be software in the form of programs and plug-ins, including the following software modules: acquisition module 4551, three-dimensional reconstruction module 4552, facial expression feature extraction module 4553, feature conversion module 4554, and decoding module 4555. These modules are logically related, and therefore can be arbitrarily combined or further split according to the functions they implement. The functions of each module will be described below.

[0054] In other embodiments, the memory 450 shown in FIG2 may further include a data processing device, which may be software in the form of programs and plug-ins, including the following software modules: a data acquisition module and a data processing result determination module. These modules are logically linked and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.

[0055] In some other embodiments, the apparatus provided in this application can be implemented in hardware. For example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the image processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0056] In some embodiments, the terminal or server can implement the image processing method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0057] The image processing methods provided in the embodiments of this application can be executed by an electronic device, which can be a server or a terminal. That is, the image processing methods in the embodiments of this application can be executed by a server, by a terminal, or by interaction between a server and a terminal.

[0058] Referring to Figure 3, which is an optional flowchart of an image processing method provided in an embodiment of this application, the method will be described in conjunction with the steps shown in Figure 3. Taking a server as the execution subject of the image processing method as an example, the method includes the following steps S101 to S105:

[0059] Step S101: Obtain the first image and the second image.

[0060] Here, expression transfer technology can also apply one person's facial features (such as smiling, frowning, surprise, etc.) to another person's face. For example, in video game scenarios, a player's character can change their appearance as needed, allowing the player character's facial features to be mixed or replaced with those of other characters in the game to create a new character image. In virtual reality environments, expression transfer technology can be used to map a user's facial expressions onto a virtual character in real time, so that the virtual character can reflect the user's expressions and movements, enhancing immersion.

[0061] In this embodiment, the first image is the original image of the facial identity feature and expression feature to be replaced during the expression transfer process. The first image, as the base facial image, includes non-facial identity features to be retained, such as hairstyle, beard, and glasses. During expression transfer, the non-facial identity features of the first image are retained, while at least one of the facial identity feature and expression feature is replaced. The first image includes the first face of a first object, which refers to the object in the original image during the expression transfer process; the first face can be the face of the first object to be replaced. The second image is a facial image used to replace at least one of the facial identity feature and expression feature of the first image during the expression transfer process. The second image provides new exchange information, such as the shape of expressions, eyes, nose, and mouth. The second image includes the second face of a second object, which can be the object to be replaced during the expression transfer process; the second face can be the face of the second object to be replaced. The first and second images can be acquired using an image acquisition device, or downloaded from the internet or obtained from social media. Continuing the above example, in a virtual reality environment, the first image can be an image including the face of a virtual character, and the second image can be an image including the face of a user.

[0062] Step S102: Perform three-dimensional reconstruction on the first face to obtain the first facial features of the first face.

[0063] In this embodiment of the application, three-dimensional reconstruction is a technique for recovering the three-dimensional structure of an object from two-dimensional images from multiple perspectives. In the application of facial expression transfer, three-dimensional reconstruction of the first face of the first object in the first image refers to extracting information from the two-dimensional image through technical means to create a three-dimensional model of the corresponding face. That is, three-dimensional reconstruction is used to create a high-precision three-dimensional model of the first face. The high-precision three-dimensional model is used to capture the detailed three-dimensional geometric structure of the first face, rather than just surface texture information.

[0064] In some embodiments, after obtaining the first image, the 3D reconstruction process may include face detection and labeling, feature extraction, 3D modeling, and texture mapping. First, the first face of the first object is detected in the first image, and facial key points, such as eyes, nose, and mouth, are labeled. Next, two-dimensional features of the first face are extracted, which may include the contour, texture, and expression of the first face. Then, the 3D structure of the first face is inferred from the 2D image using a deep learning model or other 3D reconstruction techniques, such as Structure from Motion (SfM) or stereo matching. Typically, a set of facial images taken from different angles is needed to calculate depth information. Then, the reconstructed 3D model is mapped using the texture information from the first image, giving the 3D model realistic skin texture and detail, ultimately generating a 3D model containing the facial identity features of the first object, i.e., obtaining the first facial features of the first face.

[0065] Here, the first facial features of the first face can be used for subsequent image processing operations. The first facial features contain the three-dimensional geometric shape and texture information of the first face, which helps to more accurately match and transform facial identity features or expression features during expression transfer.

[0066] Step S103: Extract facial expression features from the first face and the second face respectively to obtain the first facial expression feature of the first face and the second facial expression feature of the second face.

[0067] In this embodiment, facial expression feature extraction refers to identifying and extracting information related to facial expressions from facial images using algorithms. This information may include facial muscle movement, eye opening and closing, mouth shape changes, etc. The first facial expression feature refers to the facial expression information of a first object extracted from the first image, representing the facial expression characteristics of the first object. The second facial expression feature refers to the facial expression information of a second object extracted from the second image, representing the facial expression characteristics of the second object.

[0068] In some embodiments, the process of extracting expression features from a first face may include: analyzing the first face of a first object in a first image to identify the current expression state of the first face, such as happiness, sadness, surprise, etc.; detecting changes in the position and shape of facial key points (e.g., eyes, eyebrows, mouth, etc.) on the first face and extracting expression-related features of the first face, which constitute the first expression features of the first face. Similarly, the process of extracting expression features from a second face may include: analyzing the second face of a second object in a second image to identify the current expression state of the second face, such as happiness, sadness, surprise, etc.; detecting changes in the position and shape of facial key points on the second face and extracting expression-related features of the second face, which constitute the second expression features of the second face.

[0069] Here, the first facial expression features of the first face and the second facial expression features of the second face are obtained through the facial expression feature extraction process, so that the facial expression features in the second image can be preserved and correctly transferred to the first image during the subsequent facial expression transfer process, thereby making the third image after facial expression transfer look more natural and realistic.

[0070] Step S104: The first facial feature is transformed using the first expression feature and the second expression feature to obtain the second facial feature.

[0071] In this embodiment, feature transformation refers to the process of adjusting the three-dimensional facial model of the first face in the first image to match the facial expression features in the second image during expression transfer, using facial expression features from the first and second images. Specifically, it involves mapping and applying the second facial expression features of the second face in the second image to the facial structure of the first face in the first image. The second facial feature refers to the result of feature transformation, meaning that the adjusted three-dimensional facial model of the first face in the first image can display the second facial expression features of the second object in the second image.

[0072] In some embodiments, after obtaining the first expression feature of the first face in the first image, the second expression feature of the second face in the second image, and the first facial feature of the first face, the feature conversion process may be to map the extracted first expression feature in the first image onto the three-dimensional facial model of the first image, deform the three-dimensional facial model according to the first expression feature, and adjust the vertices and texture of the model so that the second expression feature in the second image can be mapped and applied to the three-dimensional facial model to match the expression feature of the second object in the second image.

[0073] In some embodiments, while mapping and applying the second expression features in the second image to the three-dimensional facial model of the first face, the three-dimensional facial model of the first face in the first image can also be adjusted by extracting facial identity features from the first image and the second image to match the facial identity features in the second image. That is, the facial identity features of the second face in the second image are mapped and applied to the facial structure of the first face in the first image.

[0074] In other words, the image processing method in this application embodiment includes three cases. In the first case, in some application scenarios, through the expression transfer process, the facial identity features of the first object in the first image can be completely replaced with the facial identity features of the second object in the second image, and the expression features of the first object in the first image can be completely replaced with the expression features of the second object in the second image. At this time, the three-dimensional facial structure corresponding to the second facial features obtained after feature conversion will look very similar to the second object in the second image, and reflect the expression dynamics of the second object in the second image.

[0075] In the second scenario, in other applications, the expression transfer process can replace only a portion of the facial features of the first object in the first image, such as the mouth, while retaining other facial features. In this case, the resulting 3D facial structure, obtained through feature transformation, will visually resemble only part of the second object in the second image and reflect the dynamic expression of the second object.

[0076] In the third scenario, in some applications, the facial expression transfer process can avoid replacing facial identity features. Instead, it can simply copy the facial expression features of the second object in the second image to the facial structure of the first face in the first image. In this case, the 3D facial structure corresponding to the second facial features obtained after feature transformation retains the other facial identity features of the first object in the first image, but can reflect the facial expression features of the second object in the second image.

[0077] Here, through the feature transformation process, the facial expression dynamics of the second face in the second image can be reproduced by driving the 3D facial model, ensuring the high naturalness and continuity of facial features during the expression transfer process and improving the expression following effect.

[0078] Step S105: Decode the second facial features to obtain the third image.

[0079] In this embodiment, decoding the second facial features refers to the process of converting a three-dimensional facial model, which has been obtained through feature transformation and contains the second facial expression features of the second face in the second image, into a two-dimensional image. The decoding process involves projecting the information of the three-dimensional facial model back into two-dimensional space and may include steps such as texture mapping, lighting processing, background blending, and rendering. The third image refers to the face-swapped image obtained through the expression transfer process. This third image possesses the facial expression features of the second face in the second image, while also retaining some visual features of the first face in the first image. These visual features include, but are not limited to, face shape, moles, and hairstyle.

[0080] In some embodiments, decoding the second facial features to obtain a third image that matches the expression information of the second object and the facial information of the first object can be achieved by the following technical solution: performing background encoding on the target background to obtain background features; and combining the second facial features and the background features to decode the third image.

[0081] As an example, the target background here usually refers to the target image (the second image). More specifically, it is the entire image region after the second face has been removed (or covered by a mask). An encoder (usually a convolutional neural network) is used to analyze this background image to obtain background features. This is not a simple image, but a condensed feature tensor containing rich semantic information. It encodes lighting conditions: the direction, intensity, and color (color temperature) of the light; ambient color: the overall hue and contrast; texture and detail: the blurriness or sharpness of background objects, noise particles, etc.; and spatial context: the spatial position of the face in the scene.

[0082] As an example, the face is combined with environmental information. The second facial feature represents the 3D features of the source face identity with a new expression after being driven, and the background feature represents the features of the target image environment information. These two features are concatenated in the channel dimension or fused through a cross-modal attention mechanism. This step is equivalent to telling the decoder to render the face according to the information of this environment and output a joint feature that integrates identity, expression and environmental context.

[0083] As an example, the final image rendered is in harmony with the environment. The aforementioned joint features are input to a decoder (usually a deconvolutional network or generator). Based on these information-rich joint features, the final pixel image is generated. The generated face not only has the correct identity and accurate expression, but also consistent lighting: the bright, dark, and highlight areas of the face match the background light source; color harmony: the skin tone interacts with the ambient light color temperature to avoid appearing pasted on; natural edges: the boundary between the face and the background is softened based on environmental information to achieve seamless integration; texture matching: if the background is blurry, the generated face will also have a corresponding blurriness; if the background is high-definition, the face will remain sharp.

[0084] Introducing background encoding significantly enhances realism and immersion: the generated face-swapped image is no longer an isolated face, but truly part of the scene. The coordination of lighting and color is crucial for deceiving the human eye, solving the edge blending problem: traditional methods easily produce unnatural hard boundaries or artifacts at facial edges. By allowing the decoder to see background information simultaneously, it can intelligently generate transition pixels, making edge blending more perfect; enhancing the model's versatility: the model learns to adaptively adjust its output according to the environment, thus producing robust and high-quality results when processing target images with various lighting, image quality, and styles; achieving end-to-end optimization: background features and facial features are trained together, enabling the entire system to optimize with the final visual effect in mind, rather than just optimizing the facial region itself.

[0085] The image processing method provided in this application, after obtaining a high-precision three-dimensional facial model of the first face in the first image through a three-dimensional reconstruction process, and obtaining the first expression feature of the first face and the second expression feature of the second face through expression feature extraction, performs feature transformation on the three-dimensional facial model of the first face in the first image through the first expression feature and the second expression feature. That is, the three-dimensional facial model is driven by the first expression feature and the second expression feature to accurately synchronize the expression dynamics of the second face in the second image, ensuring the high naturalness and continuity of facial features during the expression transfer process, thereby improving the output quality of the third image.

[0086] The image processing method in this application embodiment will be described below in conjunction with the interaction between the terminal and the server in the expression transfer system. It should be noted that the image processing method here is implemented through interaction between the terminal and the server, and is essentially the same as the image processing method executed by the server in the above embodiments. The only difference is that this application embodiment also describes the actions performed by the terminal during the execution of the image processing method. Furthermore, some steps can be executed by either the terminal or the server. Therefore, for steps in this embodiment that are the same as those in the above embodiments but have different execution subjects, this embodiment is merely illustrative. In the implementation process, any execution subject can perform the steps, and this application embodiment does not limit this.

[0087] Figure 4 is another optional flowchart of the image processing method provided in the embodiment of this application. As shown in Figure 4, the method includes the following steps S201 to S214:

[0088] Step S201: The terminal receives image processing operations input by the user.

[0089] In this embodiment of the application, an image editing application can run on the terminal. Users can input image processing operations in the client of the image editing application. The image editing application can provide an expression migration function. Users can input image processing operations on the expression migration function page to trigger an expression migration request.

[0090] In some embodiments, when a user inputs an image processing operation, they can simultaneously input a first image and a second image. Upon receiving the first and second images, the terminal will display a confirmation window for the expression migration function. After the terminal detects that the user has clicked the confirmation button, further processing of the first and second images will be performed to achieve the expression migration. Alternatively, in other embodiments, the user can directly input the first and second images on the expression migration function page. Upon receiving the first and second images, the terminal can directly trigger the expression migration function to perform further processing on the first and second images to achieve the expression migration.

[0091] In step S202, the terminal generates an image processing request in response to the image processing operation.

[0092] In this embodiment, user-inputted data can be encapsulated into an image processing request. For example, in the display interface of an image editing application, a first image and a second image are displayed. The user can select or confirm data on this display interface according to actual needs, and then the user-inputted first and second images can be encapsulated into an image processing request.

[0093] In step S203, the terminal sends an image processing request to the server.

[0094] In step S204, the server responds to the image processing request and obtains the first image and the second image.

[0095] In this embodiment, the first image includes a first face of a first object, and the second image includes a second face of a second object. In response to an image processing request, if the image processing request encapsulates both the first and second images, the first and second images can be directly parsed and obtained. For an explanation of the specific meaning and implementation details of obtaining the first and second images, please refer to the description of step S101 above; it will not be repeated here.

[0096] Step S205: The server performs three-dimensional reconstruction of the first face to obtain the first facial features of the first face.

[0097] In this embodiment of the application, the specific meaning and implementation of the step of performing three-dimensional reconstruction of the first face to obtain the first facial features of the first face can be found in the description of step S102 above, and will not be repeated here.

[0098] In step S206, the server extracts facial expression features from the first face and the second face respectively, thereby obtaining the first facial expression feature of the first face and the second facial expression feature of the second face.

[0099] In this embodiment of the application, the specific meaning and implementation of the step of extracting facial expression features from the first face and the second face respectively to obtain the first facial expression feature of the first face and the second facial expression feature of the second face can be found in the description of step S102 above, and will not be repeated here.

[0100] In step S207, the server extracts the driving parameters of the first facial expression features to obtain the first driving parameters of the first face.

[0101] In this embodiment, driving parameter extraction refers to analyzing the first facial expression features of a first object using an algorithm (e.g., a driving encoder) and extracting a set of parameters from these features that can describe the muscle movements and facial expression changes of the first face when making a specific expression; that is, the first driving parameters of the first face. In other words, the first driving parameters are the set of parameters extracted from the first facial expression features. This set of parameters not only describes the muscle movements and facial expression changes of the first face when making a specific expression, but can also be used to control and adjust (i.e., drive) the deformation of the 3D facial model. After obtaining the first facial expression features of the first face in the first image, driving parameters are extracted from these features to obtain the first driving parameters of the first face, so that effective driving of the first facial features can be achieved subsequently based on these first driving parameters.

[0102] In step S208, the server extracts the driving parameters of the second facial expression features to obtain the second driving parameters of the second face.

[0103] In this embodiment, the driving parameter extraction can also be achieved by analyzing the second facial expression features of the second object through an algorithm (e.g., a driving encoder), and extracting a set of parameters from the second facial expression features that can describe the muscle movements and facial expression changes of the second face when making a specific expression; that is, the second driving parameters of the second face. In other words, the second driving parameters refer to the set of parameters extracted from the second facial expression features. This set of parameters can not only describe the muscle movements and facial expression changes of the second face when making a specific expression, but can also be used to control and adjust (i.e., drive) the deformation of the 3D facial model. After obtaining the second facial expression features of the second face in the second image, driving parameters are extracted from these second facial expression features to obtain the second driving parameters of the second face, so that the first facial features can be effectively driven subsequently based on these second driving parameters.

[0104] In step S209, the server performs feature mapping on the first facial feature based on the first driving parameter and the second driving parameter to obtain the second facial feature.

[0105] In this embodiment, feature mapping refers to the process of adjusting the first facial features of a first face using a first driving parameter and a second driving parameter. The feature mapping process typically involves the movement and deformation of vertices of a 3D facial model to simulate a target expression. The second facial feature is the result of the above feature mapping process. The second facial feature is used to characterize the 3D facial model after expression adjustment according to the first and second driving parameters, and this 3D facial model can display the expression features of the second face.

[0106] In some embodiments, referring to Figure 5, Figure 5 illustrates that in step S209, the server performs feature mapping on the first facial feature based on the first driving parameter and the second driving parameter to obtain the second facial feature, which can be achieved through the following steps S2091 to S2092:

[0107] Step S2091: Using the first driving parameters, perform a first feature mapping on the first facial features to obtain the first facial features of the first face in a non-expression state.

[0108] In this embodiment, the first facial feature in a blank expression state refers to the first facial feature of a face without any expression. The first feature mapping refers to the process of reverse-driving the first facial feature using the first driving parameters of the first face; that is, mapping the first driving parameters onto a 3D facial model and adjusting the vertex positions of the 3D facial model to restore it to a blank expression state. The first feature mapping process involves moving the vertices of the 3D facial model of the first face back to their original positions in the blank expression state. After obtaining the first driving parameters and the first facial feature, the first facial feature is reverse-driven using the first driving parameters to obtain the first facial feature in a blank expression state.

[0109] Step S2092: Using the second driving parameters, perform a second feature mapping on the first facial features in the expressionless state to obtain the second facial features.

[0110] In this embodiment, the second feature mapping refers to the process of positively driving the first facial features of the first face in a blank state using the second driving parameters of the second face. Specifically, the second driving parameters are mapped onto a blank 3D facial model, and the vertex positions of the 3D facial model are adjusted to match the expression state of the second face in the second image. The second feature mapping process involves moving the vertices of the blank 3D facial model to their positions corresponding to the expression state of the second face. After obtaining the second driving parameters of the second face and the first facial features in a blank state, the first facial features in a blank state are positively driven using the second driving parameters to obtain the first facial features with the expression state of the second face, i.e., the second facial features are obtained.

[0111] Here, through the reverse driving process of the first facial features by the first driving parameters of the first face and the positive driving process of the first facial features by the second driving parameters of the second face, the first facial features of the first face can be effectively adjusted and updated to the first facial features with the expression state of the second face.

[0112] In summary, by extracting the driving parameters of the first facial expression features of the first face in the first image and the second facial expression features of the second face in the second image, as well as the feature mapping process of the first facial features of the first face, the 3D facial model is updated to reflect the facial expression dynamics of the second face in the second image, ensuring the natural transformation and high realism of the expression during the expression transfer process.

[0113] This technology achieves highly controllable and realistic 3D expression driving through a two-step feature mapping process. Its core technological advantages are: Modularization and Decoupling: Expression driving is explicitly decomposed into two independent stages: "neutral expression generation" and "expression driving." This modular design allows for separate control of the neutral facial model and facial movements, greatly enhancing the system's flexibility and controllability. Data-Driven Universality: This method does not rely on a specific 3D model topology but instead uses feature mapping through driving parameters. This allows it to adapt to different 3D character models; as long as the corresponding driving parameters can be extracted or defined, expression transfer and generation can be achieved, improving the technology's universality and portability. High Fidelity and Naturalness: First, a neutral 3D model consistent with the driving source (first face) is obtained, ensuring the correct geometric basis of expression driving. A second expression feature is then applied, ensuring that the final generated 3D expression (such as a smile or surprise) is not only accurate in movement but also naturally coordinated with the facial structure, avoiding the distortion or falsity that may result from direct mapping and significantly improving the visual realism of the generated expression. In step S210, the server decodes the second facial features to obtain the decoded second facial features.

[0114] In this embodiment of the application, decoding refers to the process of converting the second facial feature from a three-dimensional model format to a two-dimensional image format, and the decoded second facial feature refers to the feature displayed by the decoded three-dimensional model.

[0115] In some embodiments, the decoded second facial feature may be the result of synthesizing the expression features of the second face in the second image and the facial identity features of the first face in the first image. Alternatively, the decoded second facial feature may be the result of synthesizing the expression features of the second face in the second image and some facial identity features of the first face in the first image.

[0116] After obtaining the second facial features, the second facial features are decoded to convert the second facial features from a three-dimensional model format to a two-dimensional image format, that is, to obtain the decoded second facial features.

[0117] Step S211: The server performs background encoding on the target background to obtain background features.

[0118] In this embodiment, the target background can be any background in a video or image. Background encoding refers to the process of analyzing and representing the background portion of a video or image to extract the features of the target background. First, the target background region in the video or image needs to be identified and separated. Once the target background region is identified, its features are extracted. For example, these features can be color distribution, texture information, shape description, etc., that represent the visual characteristics of the target background. The extracted features are then encoded into a compact representation, namely, background features. Background features can be numerical vectors, matrices, or other data structures, and they can effectively describe the statistical characteristics of the target background.

[0119] In step S212, the server performs feature fusion on the decoded second facial features and background features to obtain the third image.

[0120] In this embodiment, the third image refers to an expression-transferred image obtained through the expression transfer process, and this expression-transferred image possesses the expression features of the second face in the second image. Feature fusion refers to the process of combining the decoded second facial features after expression transfer with the background features of the target background.

[0121] Here, after obtaining the decoded second facial features and the background features of the target background, the decoded second facial features and the background features are fused together to ensure seamless fusion of the decoded second facial features with the target background and to guarantee the natural effect of the third image.

[0122] In step S213, the server sends a third image to the terminal.

[0123] Step S214: The terminal outputs the third image.

[0124] In this embodiment, after obtaining a high-precision 3D facial model of the first face in the first image through 3D reconstruction, and obtaining the first expression feature of the first face and the second expression feature of the second face through expression feature extraction, driving parameters are extracted from the first and second expression features to obtain the first driving parameters of the first face and the second driving parameters of the second face. Through the reverse driving process of the first face's features using the first driving parameters, and the forward driving process of the second face's features using the second driving parameters, the 3D facial model of the first face can be effectively adjusted and updated to a 3D facial model with the expression state of the second face. This ensures that the 3D facial model is updated to reflect the dynamic expression of the second face in the second image, guaranteeing the natural transition and high realism of the expression during expression transfer, thereby improving the output quality of the third image.

[0125] In some embodiments, the above-described image processing method can be implemented using an image expression transfer model. Referring to Figure 6, which shows a flowchart of the image expression transfer model training method provided in this embodiment, the image expression transfer model training method can be executed by a model training module. This model training module can be a module in an electronic device used to implement the image processing method, i.e., the image expression transfer model training method can be executed by a terminal or a server. Alternatively, the model training module can be a module in another electronic device different from the electronic device used to implement the image processing method, i.e., the image expression transfer model training method can be executed by another terminal or another server. As shown in Figure 6, the image expression transfer model is trained through the following steps S301 to S306:

[0126] Step S301: Obtain sample data.

[0127] In this embodiment, the sample data includes a first sample image, a second sample image, and a ground truth image. The first sample image and the second sample image serve as input data for the image expression transfer model during training. For the specific meanings and implementation methods of the first and second sample images, please refer to the description of step S101 above. The ground truth image refers to a known real image. It serves as a supervision signal for the model during training, guiding the learning and optimization process of the image expression transfer model.

[0128] In some embodiments, for the second case in step S104 above, during the expression transfer process, only some facial features of the first object in the first image can be replaced, i.e., the facial shape features of the first object in the first image are retained, while replacing the expression features. In this case, the face in the ground truth image during the image expression transfer model training process can possess the facial shape information of the first object in the first image. In addition, the face in the ground truth image can also possess the facial lighting and color information of the second object in the second image.

[0129] Step S302: Perform first data enhancement on the first sample image to obtain the first image of the third sample image.

[0130] In this embodiment, the first sample image includes a first sample face of a first sample object, and the second sample image includes a second sample face of a second sample object. First data enhancement refers to adjusting a certain facial feature of the first sample face in the first sample image, so that the adjustment makes the facial feature consistent with the corresponding facial feature of the second sample face in the second sample image. For example, the facial feature can be facial lighting features and color features; correspondingly, the adjustment can be modifying the brightness, color, or contrast of the first sample face.

[0131] In some embodiments, referring to FIG7, FIG7 illustrates that in step S302, the server performs a first data augmentation on the first sample image to obtain a third sample image, which can be achieved through the following steps S3021 to S3022:

[0132] Step S3021: Extract facial attributes from the first sample face and the second sample face respectively to obtain the first facial attribute set and the second facial attribute set.

[0133] In this embodiment, facial attribute extraction refers to the process of recognizing facial attributes of a first sample face and a second sample face. For example, facial attributes can be facial lighting information or facial color information. The first facial attribute set refers to the data set of facial attributes obtained after extracting facial attributes from the first sample face. The second facial attribute set refers to the data set of facial attributes obtained after extracting facial attributes from the second sample face.

[0134] Step S3022: Based on the first facial attribute set and the second facial attribute set, perform a first image transformation on the first sample image to obtain a third sample image.

[0135] In this embodiment, performing a first image transformation on the first sample image refers to modifying a certain facial attribute of the first sample face in the first sample image to the facial attribute corresponding to the second sample face. The third sample image is the first sample image after the facial attributes have been updated. For example, given a first set of facial attributes and a second set of facial attributes, the facial lighting information and color information in the first set of facial attributes can be modified to the facial lighting information and color information in the second set of facial attributes, thereby updating the first sample image to a third sample image that possesses the facial lighting information and color information in the second set of facial attributes.

[0136] Here, through facial attribute extraction and the first image transformation process, some facial attributes of the first sample face in the first sample image are changed and modified to be consistent with the corresponding facial attributes in the second sample face in the second sample image, so that the learning of some facial attributes of the second sample face in the subsequent image expression transfer model training process can be carried out, thereby allowing some facial attributes of the second sample face to be preserved after expression transfer.

[0137] By introducing facial attribute extraction and attribute-based image transformation, the accuracy and effectiveness of data augmentation are significantly improved. Controllable and semantic data augmentation is achieved: by extracting a set of facial attributes, data augmentation is elevated from simple pixel perturbations to the semantic level. Precise control and transformation of specific attributes such as identity, expression, and lighting can be applied, generating diverse training samples that conform to facial anatomy. The robustness and generalization ability of the model are enhanced: this augmentation method can simulate more real-world facial changes (such as the same expression under different lighting conditions), forcing the model to learn decoupled representations of facial identity and expression attributes, reducing dependence on irrelevant variables, and thus performing more stably when facing complex real-world images. A foundation is laid for high-quality face-swapping effects: through attribute-level operations, high-quality, well-aligned intermediate data is provided for subsequent expression transfer models, ensuring the accuracy of identity transfer and expression preservation, directly contributing to the naturalness and fidelity of the generated images.

[0138] Step S303: Perform second data augmentation on the second sample image to obtain the fourth sample image.

[0139] In this embodiment of the application, the second data enhancement refers to adjusting a certain facial feature of the second sample face in the second sample image, so that the facial feature is consistent with the facial feature corresponding to the first sample face in the first sample image. For example, the facial feature can be a face shape feature (i.e., a facial contour feature). Accordingly, the adjustment process can be modifying the face shape of the second sample face, etc.

[0140] In some embodiments, referring to FIG8, FIG8 illustrates that in step S303, the server performs a second data augmentation on the second sample image to obtain a fourth sample image, which can be achieved through the following steps S3031 to S3032:

[0141] Step S3031: Extract facial key points from the first sample face and the second sample face respectively to obtain the first set of facial key points and the second set of facial key points.

[0142] In this embodiment, facial key point extraction refers to the process of recognizing the coordinates of facial key points in a first sample face and a second sample face. The first set of facial key points refers to the set of coordinates of facial key points obtained after extracting facial key points from the first sample face. For example, the first set of facial key points may include the set of coordinates of the mouth key points, the set of coordinates of the eyes key points, or the set of coordinates of the facial contour key points in the first sample face image. The second set of facial key points refers to the set of coordinates of facial key points obtained after extracting facial key points from the second sample face. For example, the second set of facial key points may include the set of coordinates of the mouth key points, the set of coordinates of the eyes key points, or the set of coordinates of the facial contour key points in the second sample face image.

[0143] Step S3032: Based on the first set of facial key points and the second set of facial key points, perform a second image transformation on the second sample image to obtain a fourth sample image.

[0144] In this embodiment, performing a second image transformation on the second sample image refers to modifying the coordinates of a facial keypoint of the second sample face in the second sample image to the coordinates of the corresponding facial keypoint of the first sample face in the first sample image. For example, the second image transformation can be a linear time warping (LTW) transformation. The fourth sample image is the second sample image with updated facial keypoint coordinates. For example, given a first set of facial keypoints and a second set of facial keypoints, the coordinates of keypoints corresponding to facial features in the second set of facial keypoints can be modified to the coordinates of keypoints corresponding to facial features in the first set of facial keypoints, thereby updating the second sample image to a fourth sample image that possesses the facial features of the first set of facial keypoints.

[0145] Here, through facial key point extraction and the second image transformation process, some facial key points of the second sample face in the second sample image are changed and modified to be consistent with the corresponding facial key points in the first sample face of the first sample image. This is so that the facial features corresponding to some facial key points of the first sample face can be learned during the subsequent training of the image expression transfer model, thereby preserving some facial features of the first sample face after expression transfer.

[0146] By introducing facial keypoints for data augmentation, the robustness of the expression transfer model to pose and expression changes is significantly improved. The main technical effects are: achieving precise geometric alignment; image transformation based on keypoint sets can accurately simulate and correct differences in pose (such as rotation and tilt) and expression shape between different faces. This provides more accurate geometrically aligned training data for subsequent expression transfer models, laying the foundation for high-quality face swapping; enhancing the model's geometric robustness; through keypoint-guided transformations, synthetic training samples containing various exaggerated poses and expressions can be generated. This forces the model to learn to decouple facial identity features from geometric shapes (pose, expression), thus enabling more stable and accurate expression transfer and identity replacement when facing target images with varied poses in the real world; improving the coordination and realism of the generated images; pre-performing geometric alignment at the data level effectively reduces artifacts such as facial feature misalignment and facial distortion that may occur after face swapping. This makes the facial structure of the final generated face-swapped image more coordinated and its integration with the original body more natural, significantly improving visual realism.

[0147] Step S304: Using the image expression transfer model to be trained, the facial expression of the fourth sample image is transferred to the third sample image to obtain the fifth sample image.

[0148] In this embodiment, the image expression transfer model may include a three-dimensional feature encoder, an expression encoder, a driving encoder, and a decoder. The third sample image includes a third sample face, and the fourth sample image includes a fourth sample face.

[0149] In some embodiments, referring to Figure 9, Figure 9 illustrates that in step S304, the server uses the image expression transfer model to be trained to transfer the facial expression of the fourth sample image to the third sample image to obtain the fifth sample image. This can be achieved through the following steps S3041 to S3044:

[0150] Step S3041: The third sample face is reconstructed in three dimensions using a three-dimensional feature encoder to obtain the third facial features of the third sample face.

[0151] In this embodiment, during the training of the image expression transfer model, the 3D reconstruction process can be implemented through the 3D feature encoder in the image expression transfer model. The network structure of the 3D feature encoder is not specifically limited here. For the specific meaning and implementation of the step of performing 3D reconstruction on a third sample face to obtain the third facial features of the third sample face, please refer to the description of step S102 above; it will not be repeated here.

[0152] Step S3042: Using an expression encoder, expression features are extracted from the third sample face and the fourth sample face respectively, resulting in the first sample expression features of the third sample face and the second sample expression features of the fourth sample face.

[0153] In this embodiment, during the training of the image expression transfer model, the expression feature extraction process can be implemented through the expression encoder in the image expression transfer model. The network structure of the expression encoder is not specifically limited here. For the step of extracting expression features from the third and fourth sample faces respectively, resulting in the first sample expression feature of the third sample face and the second sample expression feature of the fourth sample face, the specific meaning and implementation method can be found in the description of step S103 above, and will not be repeated here.

[0154] Step S3043: By driving the encoder, based on the first sample expression features and the second sample expression features, the third facial features are transformed to obtain the fourth facial features.

[0155] In this embodiment, during the training of the image expression transfer model, the feature transformation process can be implemented through the driving encoder in the image expression transfer model. The network structure of the driving encoder is not specifically limited here. For the step of performing feature transformation on the third facial feature based on the first and second sample expression features to obtain the fourth facial feature, the specific meaning and implementation method involved can be found in the description of step S104 above, and will not be repeated here.

[0156] Step S3044: The fourth facial feature is decoded using a decoder to obtain the fifth sample image.

[0157] In this embodiment, during the training of the image expression transfer model, the decoding process can be implemented through the decoder in the image expression transfer model. The network structure of the decoder is not specifically limited here. For the specific meaning and implementation of the step of decoding the fourth facial feature to obtain the fifth sample image, please refer to the description of step S105 above; it will not be repeated here.

[0158] Here, through the three-dimensional feature encoder, expression encoder, driving encoder and decoder in the image expression transfer model, the facial expression of the fourth sample image can be transferred to the third sample image to obtain the fifth sample image, so that the loss of the model can be calculated based on the fifth sample image and the model can be effectively updated based on the loss.

[0159] This technology utilizes a multi-module collaborative neural network architecture to achieve high-fidelity and robust facial expression transfer and face swapping. Its key advantages include: decoupling identity and expression for precise control; effectively separating facial identity information (3D geometry) from expression features using 3D reconstruction and an independent encoder, ensuring a clean combination of the source face's identity and the target face's expression, guaranteeing the integrity of identity preservation and the accuracy of expression transfer; enhancing the geometric consistency and realism of generated images by performing expression-driven features and feature transformation in 3D space, strictly adhering to facial anatomy priors; naturally handling changes in lighting and muscle coordination under different postures, generating face-swapped images with harmonious features, reasonable structure, and realistic texture, significantly outperforming methods that operate directly in 2D pixel space; and enhancing the model's ability to handle complex data. This structure, through explicit 3D representation learning, enhances the model's robustness to non-rigid changes in lighting and angle, enabling the trained model to better generalize to unaligned field data, generating stable and high-quality face-swapped results.

[0160] Step S305: Calculate the loss based on the ground truth image and the fifth sample image to obtain the loss result.

[0161] In some embodiments, referring to Figure 10, Figure 10 illustrates that in step S305, the server performs loss calculation based on the ground truth image and the fifth sample image to obtain the loss result, which can be achieved through the following steps S3051 to S3055:

[0162] Step S3051: Calculate the generative adversarial loss based on the ground truth image and the fifth sample image to obtain the first loss value;

[0163] In this embodiment, the image expression transfer model can be a generative adversarial network (GAN). The calculation of the generative adversarial loss is a process used in the GAN to evaluate the performance of the generator and discriminator. The first loss value includes the generator's generation loss value and the discriminator's discrimination loss value. The discrimination loss value measures the discriminator's ability to distinguish between real samples (i.e., ground truth images) and generated samples (i.e., the fifth sample image). The generation loss value measures the generator's ability to generate realistic samples (i.e., generate realistic fifth sample images).

[0164] In some embodiments, the generative adversarial loss calculation based on the ground truth image and the fifth sample image to obtain a first loss value can be achieved through the following steps: First, perform feature mapping on the fifth sample image to obtain a first discrimination probability that characterizes the fifth sample image as a natural image; perform feature mapping on the ground truth image to obtain a second discrimination probability that characterizes the ground truth image as a natural image; then, perform generative adversarial loss calculation based on the first discrimination probability and the second discrimination probability to obtain the first loss value.

[0165] In this embodiment, the first discrimination probability refers to the probability obtained by inputting the fifth sample image into the discriminator. The discriminator can perform feature mapping on the fifth sample image and output a probability value, namely the first discrimination probability. The first discrimination probability represents the degree to which the discriminator determines that the fifth sample image is a natural image (i.e., a non-fake image or a real image). If the discriminator determines that the fifth sample image is a natural image, it outputs a probability value close to 1. The second discrimination probability refers to the probability obtained by inputting the ground truth image (i.e., the original image without face swapping) into the same discriminator. The second discrimination probability represents the degree to which the discriminator determines that the ground truth image is a natural image. For the ground truth image, the discriminator should output a value close to 1, because the ground truth image is indeed a natural and unaltered image. Generative adversarial loss is calculated based on the first discrimination probability and the second discrimination probability to obtain the first loss value. The specific calculation formula for the generative adversarial loss can be found in formula (3) below.

[0166] By introducing adversarial training of the discriminator, the visual realism and detail quality of the generated face-swapping images are significantly improved. The main technical effects are: Firstly, it drives the generated images to approximate the real data distribution: By having the discriminator learn to distinguish between "sample face-swapping images" and "ground truth images," and using this to calculate the loss, it provides a powerful optimization signal for the generator. This compels the generator to produce face-swapping results that are indistinguishable from real natural images in terms of overall visual effect, texture detail, and lighting, effectively overcoming the limitations of traditional losses (such as pixel loss) that easily lead to blurry results and lack of detail. Secondly, it significantly enhances high-frequency details: Generative adversarial loss is particularly adept at capturing and stimulating high-frequency information in the generated images, such as skin texture, hair details, pores, and subtle changes in gloss. This enables the model to generate facial images with rich texture and fine details, greatly improving the realism of the output results. Thirdly, it improves the overall generalization ability of the model: This adversarial objective based on data distribution learning encourages the generator to learn more fundamental and robust facial synthesis rules, rather than simply memorizing the training set. This helps the model synthesize high-quality, highly realistic face-swapped images even when faced with unfamiliar new identities, expressions, or lighting conditions.

[0167] In some embodiments, the above-mentioned calculation of generative adversarial loss based on the first discrimination probability and the second discrimination probability to obtain the first loss value can be achieved by the following technical solution: performing logarithmic calculation on the second discrimination probability to obtain a first logarithmic calculation result; subtracting the value from the first discrimination probability to obtain a first subtraction result, and performing logarithmic calculation on the first subtraction result to obtain a second logarithmic calculation result; fusing the first logarithmic calculation result and the second logarithmic calculation result to obtain the first loss value.

[0168] As an example, the first logarithmic calculation encourages the discriminator to correctly identify the ground truth image as "real" (the second discrimination probability should be close to 1), while the second logarithmic calculation encourages the discriminator to correctly identify the generated image as "fake" (the first discrimination probability should be close to 0, hence 1 - the first discrimination probability is close to 1). The two terms are fused (usually summed) to obtain a total loss, which is used to train the discriminator to become a more accurate quality inspector. The stronger the discriminator's capabilities, the stronger the improvement signal it provides to the generator, ultimately driving the generator to produce images that are sufficiently realistic to be mistaken for genuine ones.

[0169] A stable and efficient optimization objective is provided for training Generative Adversarial Networks (GANs) through a differentiable and well-defined mathematical formula. Its core effect lies in providing a clear optimization direction for the discriminator. This formula (typically part of the loss in standard GANs or BCE) simultaneously constrains the discriminator's ability to distinguish between ground truth and generated images. It guides the discriminator to maximize the output probability of high values ​​for ground truth images (second discriminant probability) while maximizing the output probability of low values ​​for generated images (first discriminant probability). This bidirectional optimization pressure forces the discriminator to become a powerful quality inspector, thereby continuously improving its ability to distinguish between real and fake images. Indirectly driving generator improvement: The improved discriminator capabilities, through adversarial training mechanisms, in turn provide the generator with clear and powerful gradient signals. The generator's goal is to minimize this first loss value (in generator optimization, this manifests as making the discriminator output a high probability for the generated image). This loss function, based on probability and logarithmic operations, effectively transforms the abstract problem of image realism into an optimizable numerical objective, guiding the generator to continuously improve its output to deceive the discriminator. Ensuring numerical stability during training: Using logarithmic calculations (log) to handle probability values ​​transforms multiplicative relationships into additive relationships and avoids the gradient vanishing problem when calculating gradients, contributing to the stable convergence of the entire model training process.

[0170] Step S3052: Calculate the gaze loss based on the ground truth image and the fifth sample image to obtain the second loss value;

[0171] In this embodiment, the second loss value refers to the gaze loss value obtained after calculating the gaze loss. The second loss value can be used to evaluate the difference between the eye gaze in the fifth sample image generated by the image expression transfer model and the ground truth image, reflecting the performance of the image expression transfer model in gaze features. To ensure that the eye gaze in the fifth sample image is consistent with the eye gaze in the ground truth image, the gaze loss value can be considered when loss calculation is required during model training.

[0172] In some embodiments, the second loss value is obtained by calculating the gaze loss based on the ground truth image and the fifth sample image. This can be achieved through the following steps: First, gaze features are extracted from the fifth sample image to obtain the first gaze features of the fifth sample image; gaze features are extracted from the ground truth image to obtain the second gaze features of the ground truth image; then, gaze loss is calculated based on the first gaze features and the second gaze features to obtain the second loss value.

[0173] In this embodiment, the eye gaze in both the fifth sample image and the ground truth image can be a gaze direction. The difference between eye gazes refers to the difference between the left and right eye gaze directions in the ground truth image and the left and right eye gaze directions in the fifth sample image, respectively. The gaze feature extraction performed on both the ground truth image and the fifth sample image follows the same feature extraction process. This gaze feature extraction can be gaze direction estimation performed on both the fifth sample image and the ground truth image. The first gaze feature of the fifth sample image obtained through gaze direction estimation can also be a gaze direction feature in the fifth sample image, and the second gaze feature of the ground truth image can also be a gaze direction feature in the ground truth image.

[0174] In some embodiments, the second loss value can be determined based on the distance between the first gaze feature and the second gaze feature. The distance between the first gaze feature and the second gaze feature can be the Euclidean distance between the left eye gaze direction in the fifth sample image and the left eye gaze direction in the ground truth image, and the Euclidean distance between the right eye gaze direction in the fifth sample image and the right eye gaze direction in the ground truth image. The sum of the two Euclidean distances is then used as the second loss value.

[0175] Here, by calculating the gaze loss of the first gaze feature of the fifth sample image and the second gaze feature of the ground truth image, the gaze loss value during the model training process can be obtained. This gaze loss value can be used as a feedback signal during the training process of the image expression transfer model, and can be used to adjust the model parameters of the image expression transfer model in the future, thereby gradually optimizing the face-swapping effect of the image expression transfer model in terms of gaze.

[0176] By introducing a specialized gaze loss function, the realism and accuracy of the eye region in face-swapping images are significantly improved. The main technical effects are: accurately preserving key facial features, as gaze direction is a core element of facial expression; even slight deviations can lead to a dull, out-of-focus, or unnatural appearance in the generated image. By independently extracting and comparing the gaze features of the generated image with the ground truth image, this loss function can precisely constrain the direction of the eyeballs and the gaze point, ensuring that the gaze of the face-swapped character remains consistent with the expected ground truth; effectively enhancing the realism of visual interaction, as correct gaze is the foundation for believable interaction between the character and the audience or virtual environment. This mechanism directly optimizes this detail, enabling the generated virtual character or face-swapped character to present vivid, believable, and clearly directional gazes, greatly enhancing the immersion and expressiveness of the image, especially crucial in scenarios emphasizing eye contact, such as dialogue and speeches; and compensating for the supervision blind spots of other loss functions, as traditional pixel loss or general adversarial loss may not effectively capture such subtle, high-level semantic features as gaze. The specialized gaze loss, as a strong semantic constraint, fills this supervision gap and complements other loss functions, jointly driving the model to generate more complete face-swapping results in both overall structure and detail.

[0177] In some embodiments, the above-mentioned calculation of the gaze loss based on the first gaze feature and the second gaze feature to obtain the second loss value can be achieved by the following technical solution: subtracting the first gaze feature from the second gaze feature to obtain a second subtraction result; squaring each element in the second subtraction result to obtain a squared result for each element; summing the squared results of multiple elements and taking the square root to obtain the second loss value.

[0178] As an example, the error of the two gaze feature vectors is calculated in each dimension. Squaring each error amplifies the impact of larger errors, causing the model to focus more on correcting obvious directional errors. The final result is the Euclidean distance between the two feature vectors. This distance, as the loss, directly reflects the overall deviation between the generated image's gaze direction and the real gaze direction. By minimizing this L2 loss, the model is forced to learn how to accurately replicate the correct eye orientation and gaze angle, thus ensuring that the gaze of the person after face swapping is natural and accurate.

[0179] The gaze loss calculation step employs specific mathematical methods to provide accurate and stable gradient signals for model optimization. Its main technical benefits are: Firstly, it achieves precise numerical supervision of gaze. This method, through element-wise subtraction, squaring, summation, and square root operations, essentially calculates the Euclidean distance (L2 norm) between the generated image gaze features and the ground truth gaze features. This calculation transforms the abstract gaze difference into a concrete, differentiable scalar loss value, providing the model with a clear and explicit optimization objective: minimizing the distance between the generated and ground truth gaze in the feature space, thereby ensuring accurate reconstruction of eye orientation. Secondly, it ensures stable convergence during training: the L2 loss function possesses favorable mathematical properties, producing smooth and stable gradients. This allows the model to perform smooth and controllable parameter updates during the optimization process, effectively avoiding gradient explosion or drastic fluctuations that may occur when using certain loss functions. This is conducive to the stable convergence of the entire training process, ultimately resulting in a more reliable expression transfer model. It effectively complements other loss functions. This method is a specialized component of the multi-task loss function, which focuses on solving the specific and critical detail of gaze direction. It complements the adversarial loss responsible for overall realism and the pixel loss responsible for pixel-level fidelity, and together drive the model to generate highly realistic face-swapping images at both the macro and micro levels, especially achieving higher naturalness in subtle details such as eye light and gaze direction.

[0180] Step S3053: Calculate pixel loss based on the ground truth image and the fifth sample image to obtain the third loss value;

[0181] In this embodiment, the third loss value refers to the pixel loss value calculated through pixel loss. The third loss value can be used to evaluate the difference between facial pixels in the fifth sample image generated by the image expression transfer model and the ground truth image. The difference in facial pixels reflects whether the face in the fifth sample image is similar to the face in the ground truth image. To ensure that the facial pixels in the fifth sample image are as consistent as possible with the facial pixels in the ground truth image, the pixel loss value can be considered when loss calculation is required during model training.

[0182] In some embodiments, the pixel loss calculation based on the ground truth image and the fifth sample image to obtain the third loss value can be achieved through the following steps: First, pixel features are extracted from the fifth sample image to obtain the first pixel set of the fifth sample image; pixel features are extracted from the ground truth image to obtain the second pixel set of the ground truth image; then, pixel loss calculation is performed based on the first pixel set and the second pixel set to obtain the third loss value.

[0183] In this embodiment, the pixel feature extraction performed on the ground truth image and the fifth sample image is the same type of feature extraction process. Here, pixel feature extraction can be performed separately on each pixel of the fifth sample image and the ground truth image. The first pixel set includes pixel features corresponding to multiple pixels in the fifth sample image, and the second pixel set includes pixel features corresponding to multiple pixels in the ground truth image.

[0184] The pixel loss calculation step provides a fundamental and crucial constraint for model optimization by directly comparing the generated image with the ground truth image at the pixel level. Its technical effects are mainly reflected in: ensuring the overall fidelity of the image structure; through pixel-by-pixel comparison, this loss function forces the generated face-swapped image to maintain a high degree of consistency with the ground truth image in terms of overall structure, contour, and the positions of major facial features. This provides the most basic geometric correctness guarantee for the face-swapping effect, effectively preventing severe structural distortion, misalignment, or background anomalies in the generated image, ensuring the initial usability of the output results; providing stable and reliable training gradients; pixel loss (such as L1 or L2 loss) is simple to calculate, with smooth and stable gradients, and is less prone to gradient vanishing or exploding problems. In the early stages of training or when adversarial training is unstable, it can provide a clear and robust optimization direction, helping the model quickly converge to a reasonable initial solution, laying a good foundation for subsequent optimization details through adversarial loss; and complementing other loss functions. Pixel loss focuses on low-frequency information and global structure, but can easily lead to overly smooth results lacking detail. It effectively complements generative adversarial loss, which focuses on high-frequency texture and visual realism, and gaze loss, which is responsible for key expressions. The three components work together to drive the model to generate face-swapped images that are rich in detail and have natural expressions while maintaining the correct overall structure.

[0185] In some embodiments, the third loss value can be determined based on the pixel difference between the first pixel set and the second pixel set. The pixel difference between the first pixel set and the second pixel set can be the absolute difference between the pixel features of each pixel in the fifth sample image and the pixel features of each pixel in the ground image. The sum of the absolute differences corresponding to the above multiple pixels is used as the third loss value.

[0186] In some embodiments, the above-mentioned pixel loss calculation based on the first pixel set and the second pixel set to obtain the third loss value can be implemented by the following technical solution: performing the following processing for each pixel position: determining the first difference between the first pixel feature (pixel value) corresponding to the pixel position in the first pixel set and the second pixel feature (pixel value) corresponding to the pixel position in the second pixel set; summing the absolute values ​​of the first differences corresponding to multiple pixel positions to obtain the third loss value.

[0187] Here, by calculating the pixel loss of the first pixel set of the fifth sample image and the second pixel set of the ground truth image, the pixel loss value during the model training process can be obtained. This pixel loss value can be used as a feedback signal during the training process of the image expression transfer model, and can be used to adjust the model parameters of the image expression transfer model in the future, thereby gradually optimizing the face-swapping effect of the image expression transfer model in terms of pixels.

[0188] This calculation process uses the standard L1 loss (absolute value loss). Its core function is to impose precise pixel-level constraints: by calculating the sum of the absolute differences in color values ​​at each pixel between the generated image and the ground truth image, it forces the generated face-swapped image to be as close as possible to the ground truth image at the pixel level. Compared to L2 loss (squared loss), L1 loss is less sensitive to outliers, better preserves edge information, and helps the generated image maintain clearer contours and sharper details, avoiding overly smoothed and blurred results. This loss provides the model with the most basic reconstruction constraints, ensuring the correctness of the generated result in terms of overall structure and color distribution.

[0189] By calculating the L1 loss (sum of absolute values) of all pixels in the image, the technical effects are as follows: It ensures basic alignment of the image structure: through pixel-by-pixel comparison, it forces the generated image to maintain a basic consistency with the ground truth in terms of overall structure and contour, providing the most basic geometric correctness guarantee for the face-swapping effect; it reduces blurring and preserves edges: compared with L2 loss, L1 loss is less sensitive to outliers and has a gentler penalty, helping the generated image retain clearer edges and sharper details, avoiding overly smooth results; it provides stable training gradients: its calculation is simple, the gradient is stable, and it can effectively complement higher-order losses such as adversarial loss, jointly driving the model to synthesize realistic details while maintaining the correct structure.

[0190] Step S3054: Calculate the expression loss based on M sixth sample images and N seventh sample images to obtain the fourth loss value;

[0191] In this embodiment, M and N are both integers greater than 1, and their values ​​can be the same or different. The M sixth sample images refer to M source images containing different objects, and the N seventh sample images refer to N target images containing different objects. Because facial features may be included in the facial expression features extracted during the expression feature extraction process in the image expression transfer model, these facial features may be transferred from the target image to the face-swapped image during actual expression transfer, resulting in severe distortion in the generated face-swapped image and a poorer expression tracking effect. Therefore, to decouple facial features from expression features during the expression feature extraction process in the image expression transfer model, pixel loss values ​​can be considered when calculating loss during model training.

[0192] In some embodiments, the expression loss is calculated based on M sixth sample images and N seventh sample images to obtain a fourth loss value. This can be achieved through the following steps: First, data augmentation is performed on the M sixth sample images and N seventh sample images respectively to obtain M eighth sample images and N ninth sample images. Then, 3D reconstruction is performed on the M fifth sample faces to obtain the third facial features of each of the M fifth sample faces. Next, expression features are extracted from the M fifth sample faces and N sixth sample faces respectively to obtain the third sample expression features of each of the M fifth sample faces and the fourth sample expression features of each of the N sixth sample faces. Then, driving parameters are extracted from the M third sample expression features and N fourth sample expression features respectively to obtain M third driving parameters and N fourth driving parameters. Finally, the expression loss is calculated based on the M third facial features, M third driving parameters, and N fourth driving parameters to obtain the fourth loss value.

[0193] In this embodiment, the M eighth sample images include M fifth sample faces corresponding to M different third sample objects, and the N ninth sample images include N sixth sample faces corresponding to N different fourth sample objects. Data augmentation is performed on the M sixth sample images and the N seventh sample images to obtain the corresponding M eighth sample images and N ninth sample images; 3D reconstruction is performed on the M fifth sample faces to obtain the third facial features of each of the M fifth sample faces; expression features are extracted from the M fifth sample faces and the N sixth sample faces to obtain the corresponding third expression features of each of the M fifth sample faces and the corresponding fourth expression features of each of the N sixth sample faces; driving parameters are extracted from the M third expression features and the N fourth expression features to obtain the corresponding M third driving parameters and N fourth driving parameters. The specific meanings and implementation methods of these steps can be found in the descriptions of steps S302-S304 above, and will not be repeated here.

[0194] In some embodiments, the expression loss is calculated based on M third facial features, M third driving parameters, and N fourth driving parameters to obtain a fourth loss value. This can be achieved through the following steps: First, the M third facial features are mapped using the M third driving parameters and the i-th fourth driving parameter to obtain M three-dimensional first expression driving features; i is a positive integer less than or equal to N. Second, the M third facial features are mapped using the M third driving parameters and N fourth driving parameters to obtain M×N three-dimensional second sample expression driving features. Third, the M three-dimensional first expression driving features and the M×N three-dimensional second sample expression driving features are decoded to obtain M first fifth sample images and M×N second fifth sample images. Fourth, expression features are extracted from the M first fifth sample images and the M×N second fifth sample images to obtain M fifth expression features and M×N sixth expression features. Finally, the expression loss is calculated based on the M fifth expression features and the M×N sixth expression features to obtain the fourth loss value.

[0195] In some embodiments, M third facial features are mapped using M third driving parameters and the i-th fourth driving parameter to obtain M first three-dimensional expression driving features; i is a positive integer less than or equal to N; M third facial features are mapped using M third driving parameters and N fourth driving parameters to obtain M×N second fourth facial features; the M first three-dimensional expression driving features and the M×N second fourth facial features are decoded respectively to obtain M first fifth sample images and M×N second fifth sample images. The specific meanings and implementation methods involved in these steps can be found in the explanation of step S304 above, and will not be repeated here. It should be noted that the M first three-dimensional expression driving features are obtained by extracting driving parameters and mapping third features from any one of the M third sample expression features and N fourth sample expressions, so that all M three-dimensional sample facial models are updated using the same fourth sample expression to reflect the expression dynamics of the sixth sample face in the same ninth sample image. The M×N three-dimensional second sample expression-driven features are obtained by extracting driving parameters and mapping fourth features from M third sample expression features and N fourth sample expressions, so that the M three-dimensional sample facial models are updated through different fourth sample expressions to reflect the expression dynamics of the sixth sample face in different ninth sample images.

[0196] M first three-dimensional facial expression driving features are decoded to obtain M first fifth sample images, and M×N second fourth facial features are decoded to obtain M×N second fifth sample images. That is, the M first fifth sample images are fifth sample images obtained by facial expression transfer based on the same fourth sample expression, and the M×N second fifth sample images are fifth sample images obtained by facial expression transfer based on different fourth sample expressions. Facial expression features are extracted from the M first fifth sample images and the M×N second fifth sample images respectively to obtain M fifth facial expression features and M×N sixth facial expression features. Facial expression loss is calculated based on the M fifth facial expression features and the M×N sixth facial expression features to obtain the fourth loss value. This facial expression loss calculation refers to the process of calculating the cosine distance between the M fifth facial expression features and the M×N sixth facial expression features. That is, the fourth loss value is the cosine distance between the M fifth facial expression features and the M×N sixth facial expression features. The fourth loss value can be used to measure the difference between the facial expression features corresponding to the fifth sample images obtained by facial expression transfer based on different fourth sample expressions and the facial expression features corresponding to the fifth sample images obtained by facial expression transfer based on different fourth sample expressions. The specific formula for calculating the facial expression loss can be found in formula (4) below.

[0197] Here, through the expression loss calculation process, the fourth loss value in the model training process can be obtained. This loss value can be used as a feedback signal in the training process of the image expression transfer model, and can be used to adjust the model parameters of the image expression transfer model in the future, thereby gradually optimizing the decoupling of expression features and facial features in the expression feature extraction process of the image expression transfer model and improving the face swapping effect.

[0198] Step S3055: The first loss value, the second loss value, the third loss value and the fourth loss value are fused to obtain the loss result.

[0199] In this embodiment of the application, loss value fusion refers to the process of weighted summation of the first loss value, the second loss value, the third loss value, and the fourth loss value. The first loss value, the second loss value, the third loss value, and the fourth loss value are weighted using four preset loss weights (i.e., hyperparameters), and then summed to obtain the loss result.

[0200] Here, through generative adversarial loss calculation, gaze loss calculation, pixel loss calculation, and expression loss calculation, we obtain the following corresponding loss values: a first loss value to measure the realism of the fifth sample image generated by the image expression transfer model; a second loss value to measure the face-swapping effect of the image expression transfer model in terms of gaze; a third loss value to measure the face-swapping effect of the image expression transfer model in terms of pixels; and a fourth loss value to measure the degree of decoupling between expression features and facial features during the expression feature extraction process of the image expression transfer model. Based on these loss values, we obtain the loss result of the image expression transfer model. This loss result can be used to guide the next step of image expression transfer model training in the correct direction and improve the model performance.

[0201] By designing a multi-dimensional, multi-task hybrid loss function, the overall performance of the expression transfer model and the quality of the generated images are significantly improved. The main technical effects are: Comprehensive improvement in the visual fidelity of the generated images: By combining generative adversarial loss (ensuring overall visual realism and texture details) and pixel loss (ensuring low-error reconstruction of image structure), the generated face-swapped images are highly consistent with the ground truth in both global perception and local structure, effectively avoiding image blurring or distortion; Precise preservation and transfer of key facial attributes: Specifically, gaze loss and expression loss are introduced to strongly constrain the two crucial dimensions of eye direction and facial expression. This ensures that the gaze of the person after face swapping is natural, the gaze direction is correct, and the target expression is accurately and clearly transferred to the source identity, greatly enhancing the naturalness and expressiveness of the generated results; Enhanced stability and convergence of model training: By fusing loss functions with different supervision signals, more comprehensive and robust gradient guidance is provided for model optimization. This multi-task learning strategy helps the model learn more general and balanced feature representations, avoiding model degradation that may be caused by a single loss, thereby stabilizing the training process and prompting the model to converge to a better performance state.

[0202] Step S306: Based on the loss results, update the model parameters in the image expression transfer model to obtain the trained image expression transfer model.

[0203] In this embodiment, based on the loss result, the model parameters in the image expression transfer model are iteratively updated according to preset iterative conditions to obtain the trained image expression transfer model. The loss result guides the next step of image expression transfer model training in the correct direction. The preset iterative conditions can be a loss threshold, a maximum iteration threshold, and an iteration cutoff time. Based on the derivative of the loss function corresponding to the loss result of the image expression transfer model to be trained, the loss result is backpropagated along the direction of minimum gradient to update the model parameters in the image expression transfer model to be trained, such as the various weight values ​​in the image expression transfer model to be trained. A total loss threshold is preset. When the loss result is less than the preset loss threshold, iterative training stops, i.e., model parameter updates stop. Alternatively, a maximum iteration threshold can be preset. When the number of iterations exceeds the maximum iteration threshold, model parameter updates stop. Alternatively, an iteration cutoff time can be preset. When the iteration time reaches the iteration cutoff time, model parameter updates stop, resulting in the trained image expression transfer model.

[0204] In this embodiment, the facial expression of the image expression transfer model is transferred from the data-enhanced fourth sample image to the data-enhanced third sample image using a 3D feature encoder, expression encoder, driver encoder, and decoder, resulting in a fifth sample image. Generative adversarial loss calculation, gaze loss calculation, pixel loss calculation, and expression loss calculation are used to obtain a first loss value to measure the realism of the fifth sample image generated by the image expression transfer model, a second loss value to measure the face-swapping effect of the image expression transfer model in terms of gaze, a third loss value to measure the face-swapping effect of the image expression transfer model in terms of pixels, and a fourth loss value to measure the degree of decoupling between expression features and facial features during the expression feature extraction process. Based on these loss values, the loss result of the image expression transfer model is obtained. This loss result can be used to guide the next step of image expression transfer model training in the correct direction, improving the model performance.

[0205] This application effectively improves the performance and robustness of the image expression transfer model. The technical effects are mainly reflected in the following aspects: Enhanced model generalization and robustness: By independently augmenting the source and target images, more diverse and challenging training samples (such as different lighting, poses, occlusions, etc.) are generated. This forces the model to learn more essential facial identity and expression features, rather than relying on the randomness of training data, thus significantly improving the model's generalization ability and robustness in complex real-world scenarios. Improved realism and accuracy of face-swapping results: This process clearly defines the task objective of transferring the expression from the target image to the source image. Through end-to-end training and using ground-value images to calculate loss, the quality of the generated image can be directly optimized in multiple dimensions, including pixel-level and feature-level. This ensures that the face-swapped image not only converts identity features to the source identity but also accurately retains and integrates the expression information of the target image, resulting in a final face-swapped result with harmonious facial features, natural expressions, and high realism and fidelity. Ensured stability and effectiveness of the training process: The loss results calculated based on ground-value images provide a clear and reliable optimization direction for model parameter updates. This supervised learning approach helps the model converge stably and gradually reduces artifacts, blurring, or identity-expression confusion in the generated images, ultimately resulting in a high-performance, highly reliable expression transfer model.

[0206] The following describes an exemplary application of the embodiments of this application in a practical application scenario. This application provides a real-time image processing method based on 3D representation and expression-driven processing. A schematic diagram of the expression transfer process of this method is shown in Figure 11. First, a source face 110 (i.e., the aforementioned first face) is obtained from a source video or a first image, and a target face 111 (i.e., the aforementioned second face) is obtained from another input source (such as a real-time video stream or other images). Then, a 3D model is constructed for the source face using a 3D representation encoder 112 (i.e., the aforementioned 3D feature encoder), resulting in a 3D representation 113 of the source face (i.e., the aforementioned first facial features), to accurately reconstruct the geometry and surface details of the source face. Next, the expressions of the source face 110 and the target face 111 are analyzed using an expression encoder 114, and key expression features of the source face 110 and the target face 111 are extracted to obtain source face expression features 115 (i.e., the aforementioned first expression features) and target face expression features 116 (i.e., the aforementioned second expression features). The source face's 3D representation 113 is driven by the source face's expression features 115 and the target face's expression features 116 to match the target face's expression, resulting in an expression-driven 3D representation 117 (i.e., the aforementioned second facial feature). That is, the update of the source face's 3D representation 113 reflects the dynamic expression of the target face 111, ensuring natural transitions and high realism during expression transfer. Then, the expression-driven 3D representation 117 is input into the decoder 118, converting it back to a 2D image format to obtain the decoded second facial feature 119. Finally, the feature fusion unit 120 combines the decoded second facial feature 119 with the background features 121 in the target video or second image to obtain a third image 122, ensuring seamless integration of the swapped face with the background environment. The background feature 121 is obtained by encoding the target background 124 in the target video or second image using the background encoder 123. The output third image 122 will show a face with the target facial expression features, while retaining other visual features of the source face, such as facial shape features.

[0207] Specifically, since facial expression features are two-dimensional representations and cannot directly drive the three-dimensional representation of a face, this embodiment requires a driving encoder to obtain the three-dimensional driving corresponding to the facial expression features. A flowchart of the facial expression driving process in this embodiment is shown in Figure 12. First, the source driving encoder 125 extracts driving parameters from the source face facial expression features 115 to obtain the three-dimensional driving parameters 126 of the source face (i.e., the aforementioned first driving parameters). Using the source face's three-dimensional driving parameters 126, the source face's three-dimensional representation 113 is driven to an expressionless state (i.e., no expression), resulting in the expressionless three-dimensional representation 127 (i.e., the aforementioned first facial feature in an expressionless state). 3) The target driving encoder 128 extracts driving parameters from the target face facial expression features 116 to obtain the target face's three-dimensional driving parameters 129 (i.e., the aforementioned second driving parameters). Using the target face's three-dimensional driving parameters 129, the expressionless three-dimensional representation 127 is driven to the target face's expression, thus obtaining the expression-driven three-dimensional representation 130 (i.e., the aforementioned second facial feature). Furthermore, since one drive encoder is for the reverse process and the other is for the forward process, this embodiment requires two different drive encoders (i.e., source drive encoder 125 and target drive encoder 128).

[0208] In image processing methods, data augmentation design is particularly important because there is no ground truth image after face swapping as a supervision signal during model training. The source face and target face during model training are actually different frames from the same video, so the frame containing the target face can be used as the ground truth image for supervision. However, two problems arise during model training: Problem 1: It is difficult for the model to learn the swapping between different face shapes during training, because the face shapes of people in the same video may be the same. Problem 2: The face after expression transfer only retains the source face and cannot adapt to the lighting changes of the target face, resulting in reduced naturalness.

[0209] To address the above two issues, this application employs two data augmentation methods. 1) Linear Time Warping (LTW) transform: Figure 13 illustrates the nonlinear transformation of input data using LTW transform with different parameters. By extracting facial key points from the source and target faces (i.e., the aforementioned facial key point extraction) and applying LTW transform, the target face is altered, while the face shape in the ground truth image remains that of the source face. Therefore, the face shape provided by the face in the ground truth image during loss supervision is that of the source face, helping the model-generated third image maintain the face shape of the source face, thus ensuring that the face shape of the third image originates from the source face. 2) Color enhancement: To better ensure that the face in the third image maintains the lighting environment of the target face, this application performs color enhancement on the source face, for example, by modifying the brightness, color, or contrast of the source face. The face in the ground truth image still retains the lighting and color of the target face, making the third image generated during model training more likely to generate the lighting environment and color of the source face during loss supervision, thus improving the naturalness of the third image.

[0210] The loss functions in this application mainly include four loss functions: the GAN loss function for generative adversarial networks, the line-of-sight loss function, the L1 loss function, and the recurrence loss function.

[0211] The GAN loss is calculated based on the Generative Adversarial Network (GAN), which consists of two parts: a generator and a discriminator. During model training, the generator and discriminator compete against each other to improve the quality of the generated images and their discriminative ability. The GAN loss function is key to measuring the effectiveness of this adversarial training. The formulas for calculating the GAN loss function are as follows: (1), (2), and (3): L GAN =L D +L G (3)

[0212] Among them, L D L represents the discriminator loss of the GAN, which is maximized during model training. G This represents the generator loss of the GAN. Minimizing this generator loss during model training yields the final GAN ​​loss L. GAN (i.e., the first loss value mentioned above), log represents the logarithmic function, x is the ground truth image (i.e., the real sample), D(x) represents the discrimination probability of the discriminator in the generative adversarial network for the ground truth image, G(z) is the third image generated by the generator in the generative adversarial network, and D(G(z)) represents the discrimination probability corresponding to the third image. This indicates the distribution of the real data p.data The expectation of the real sample x obtained from sampling is used to measure the discriminator's ability to distinguish real samples. Represents the generator noise distribution p z The expectation of the noise z obtained from the sampling is used to measure the generator's ability to generate fake samples.

[0213] The gaze loss function measures the difference between the gaze direction of the swapped face and the original face, helping to improve the gaze accuracy of the third image. The gaze loss function can be expressed as the following formula (4): L eye =‖θ real -θ gen ‖twenty four)

[0214] Where, θ real and θ gen These are the gaze direction vectors of the eye regions in the ground truth image and the third image, respectively. These vectors are obtained using a typical feature extraction network, such as the VGG network. ‖‖2 denotes the L2 norm, L... eye This refers to the loss of sight (i.e., the second loss value mentioned above).

[0215] The L1 loss function is a commonly used loss function to measure the pixel difference between the third image and the ground truth image. It is used to calculate the absolute difference in pixel values ​​between the third image and the ground truth image. The L1 loss function can be expressed as the following formula (5): L L1 =‖x real -x gen ‖1 (5)

[0216] Where, x real x represents the pixel value of the truth image. gen This represents the pixel value of the third image, and |||1 represents the L1 norm. L1 This is the L1 loss (i.e., the third loss value mentioned above).

[0217] The cyclic loss function is used to prevent appearance leakage through facial expression representation. This task is necessary during training because the embodiments of this application only expect the target face to provide facial expressions. However, the face in the ground truth image used in actual training is the same as the target face. Therefore, the facial expression features of the target face may also provide facial representation, resulting in severe distortion of the third image during actual face swapping. This is because the facial representation will leak from the second image to the third image. Therefore, a cyclic loss is needed to decouple the facial expression encoder from the facial expression features and the facial representation.

[0218] In calculating the cyclic loss, embodiments of this application construct positive and negative sample pairs. The positive sample pair consists of expression features obtained from a third image generated using different source and target facial expressions. s→d and and the expression features Z obtained from the same target face d Finally, positive sample pairs were obtained. Negative sample pairs are expression features obtained from third images generated using different source and target facial expressions. s→d and However, the facial features obtained are not from the same target face. Finally, negative sample pairs are obtained. In other words, if the providers of the final facial expression features are consistent, the sample is positive; otherwise, it is negative. This ensures that the facial expression encoder is only related to facial expression features and not to facial representation. The formula for calculating the recurrent loss function is as follows (6):

[0219] Where exp represents the exponential function, log represents the logarithmic function, d represents the cosine distance, ∑ represents the accumulation function, and L COS This is the cyclical loss (i.e., the fourth loss value mentioned above).

[0220] Therefore, the formula for calculating the total loss function during final model training is shown in formula (7): L total =λ GAN L GAN +λ eye L eye +λ L1 L L1 +λ COS L COS (7)

[0221] Where, λ GAN , λ eye , λ L1 , λ COS There are four hyperparameters, which control the weights of the GAN loss, gaze loss, L1 loss, and recurrence loss, respectively. total This represents the total loss (i.e., the loss outcome described above).

[0222] The method provided in this application embodiment can preserve the facial details of the source face to a greater extent and follow the expression well, ultimately obtaining a third image with high fidelity and high naturalness. At the same time, the image processing method based on generative adversarial networks can achieve real-time face swapping, laying the foundation for the large-scale application of this method in highly dynamic interactive environments.

[0223] It is understood that in the embodiments of this application, if the content involving user information, such as the first image, the second image, and the image expression transfer model, involves data related to user information or enterprise information, when the embodiments of this application are applied to specific products or technologies, it is necessary to obtain the user's permission or consent, or to obfuscate this information to eliminate the correspondence between this information and the user; and the collection and processing of related data should strictly comply with the requirements of relevant national laws and regulations when applied in practice, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0224] The following description continues to illustrate the exemplary structure of the image processing apparatus 455 provided in this application embodiment as a software module. In some embodiments, as shown in FIG2, the software module stored in the image processing apparatus 455 in the memory 450 may include: an acquisition module 4551 configured to acquire a first image and a second image; the first image includes a first face of a first object, and the second image includes a second face of a second object; a three-dimensional reconstruction module 4552 configured to perform three-dimensional reconstruction on the first face to obtain a first facial feature of the first face; an expression feature extraction module 4553 configured to extract expression features from the first face and the second face respectively to obtain a first expression feature of the first face and a second expression feature of the second face; a feature conversion module 4554 configured to perform feature conversion on the first facial feature using the first expression feature and the second expression feature to obtain a second facial feature; and a decoding module 4555 configured to decode the second facial feature to obtain a third image.

[0225] In some embodiments, the feature conversion module 4554 is further configured to: extract driving parameters from the first expression features to obtain first driving parameters of the first face; extract driving parameters from the second expression features to obtain second driving parameters of the second face; and perform feature mapping on the first facial features based on the first driving parameters and the second driving parameters to obtain the second facial features.

[0226] In some embodiments, the feature conversion module 4554 is further configured to: perform a first feature mapping on the first facial feature using the first driving parameter to obtain a first facial feature of the first face in an expressionless state; and perform a second feature mapping on the first facial feature in an expressionless state using the second driving parameter to obtain a second facial feature.

[0227] In some embodiments, the decoding module 4555 is further configured to: decode the second facial feature to obtain the decoded second facial feature; perform background encoding on the target background to obtain background features; and perform feature fusion on the decoded second facial feature and the background features to obtain the third image.

[0228] In some embodiments, the image processing method is implemented through an image expression transfer model; the device 455 further includes a model training module, which is configured to: acquire sample data; the sample data includes a first sample image, a second sample image, and a ground truth image; perform a first data augmentation on the first sample image to obtain a third sample image; perform a second data augmentation on the second sample image to obtain a fourth sample image; transfer the facial expression of the fourth sample image to the third sample image using the image expression transfer model to be trained, to obtain a fifth sample image; perform loss calculation based on the ground truth image and the fifth sample image to obtain a loss result; and update the model parameters in the image expression transfer model based on the loss result to obtain a trained image expression transfer model.

[0229] In some embodiments, the first sample image includes a first sample face of a first sample object, and the second sample image includes a second sample face of a second sample object; the model training module is further configured to: extract facial attributes from the first sample face and the second sample face respectively to obtain a first set of facial attributes and a second set of facial attributes; and perform a first image transformation on the first sample image based on the first set of facial attributes and the second set of facial attributes to obtain the third sample image.

[0230] In some embodiments, the model training module is further configured to: extract facial key points from the first sample face and the second sample face respectively to obtain a first set of facial key points and a second set of facial key points; and perform a second image transformation on the second sample image based on the first set of facial key points and the second set of facial key points to obtain the fourth sample image.

[0231] In some embodiments, the image expression transfer model includes a 3D feature encoder, an expression encoder, a driving encoder, and a decoder; the third sample image includes a third sample face, and the fourth sample image includes a fourth sample face; the model training module is further configured to: perform 3D reconstruction on the third sample face using the 3D feature encoder to obtain the third facial features of the third sample face; extract expression features from the third sample face and the fourth sample face using the expression encoder to obtain the first sample expression features of the third sample face and the second sample expression features of the fourth sample face; perform feature transformation on the third facial features based on the first sample expression features and the second sample expression features using the driving encoder to obtain the fourth facial features; and decode the fourth facial features using the decoder to obtain the fifth sample image.

[0232] In some embodiments, the model training module is further configured to: calculate a generative adversarial loss based on the ground truth image and the fifth sample image to obtain a first loss value; calculate a gaze loss based on the ground truth image and the fifth sample image to obtain a second loss value; calculate a pixel loss based on the ground truth image and the fifth sample image to obtain a third loss value; calculate an expression loss based on M sixth sample images and N seventh sample images to obtain a fourth loss value; where M and N are both integers greater than 1; and fuse the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the loss result.

[0233] In some embodiments, the model training module is further configured to: perform feature mapping on the fifth sample image to obtain a first discrimination probability configured to characterize the fifth sample image as a natural image; perform feature mapping on the ground truth image to obtain a second discrimination probability configured to characterize the ground truth image as a natural image; and calculate the generative adversarial loss based on the first discrimination probability and the second discrimination probability to obtain the first loss value.

[0234] In some embodiments, the model training module is further configured to: perform logarithmic calculation on the second discrimination probability to obtain a first logarithmic calculation result; subtract the value one from the first discrimination probability to obtain a first subtraction result, and perform logarithmic calculation on the first subtraction result to obtain a second logarithmic calculation result; and fuse the first logarithmic calculation result and the second logarithmic calculation result to obtain the first loss value.

[0235] In some embodiments, the model training module is further configured to: extract gaze features from the fifth sample image to obtain a first gaze feature of the fifth sample image; extract gaze features from the ground truth image to obtain a second gaze feature of the ground truth image; and calculate gaze loss based on the first gaze feature and the second gaze feature to obtain a second loss value.

[0236] In some embodiments, the model training module is further configured to: subtract the first gaze feature from the second gaze feature to obtain a second subtraction result; square each element in the second subtraction result to obtain a squared result for each element; and sum and take the square root of the squared results of multiple elements to obtain the second loss value.

[0237] In some embodiments, the model training module is further configured to: extract pixel features from the fifth sample image to obtain a first pixel set of the fifth sample image; extract pixel features from the ground truth image to obtain a second pixel set of the ground truth image; and calculate pixel loss based on the first pixel set and the second pixel set to obtain the third loss value.

[0238] In some embodiments, the model training module is further configured to perform the following processing for each pixel location: determine a first difference between a first pixel feature corresponding to the pixel location in the first pixel set and a second pixel feature corresponding to the pixel location in the second pixel set; and sum the absolute values ​​of the first differences corresponding to multiple pixel locations to obtain the third loss value.

[0239] In some embodiments, the model training module is further configured to: perform data augmentation on the M sixth sample images and the N seventh sample images respectively, to obtain M eighth sample images and N ninth sample images; the M sixth sample images include M fifth sample faces corresponding to M third sample objects respectively, and the N ninth sample images include N sixth sample faces corresponding to N fourth sample objects respectively; perform three-dimensional reconstruction on the M fifth sample faces to obtain the third facial features of each of the M fifth sample faces; extract expression features from the M fifth sample faces and the N sixth sample faces respectively, to obtain the third expression features of each of the M fifth sample faces and the fourth expression features of each of the N sixth sample faces respectively; extract driving parameters from the M third expression features and the N fourth expression features respectively, to obtain the M third driving parameters and the N fourth driving parameters; and calculate the expression loss based on the M third facial features, the M third driving parameters, and the N fourth driving parameters to obtain the fourth loss value.

[0240] It should be noted that the description of the apparatus in this application embodiment is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment; therefore, it will not be repeated. For technical details not disclosed in this apparatus embodiment, please refer to the description of the method embodiment of this application for understanding.

[0241] This application provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor will execute the image processing method provided in this application, such as the image processing method shown in FIG3.

[0242] This application provides a computer program product including computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image processing method described in this application.

[0243] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0244] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0245] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0246] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0247] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. An image processing method, the method being performed by an electronic device, the method comprising: Acquire a first image and a second image; the first image includes a first face of a first object, and the second image includes a second face of a second object; The first face is reconstructed in three dimensions to obtain the first facial features of the first face; Expression features are extracted from the first face and the second face respectively, resulting in the first expression feature of the first face and the second expression feature of the second face. The first facial feature is transformed using the first facial feature and the second facial feature to obtain the second facial feature; The second facial features are decoded to obtain a third image that matches the expression information of the second object and the facial information of the first object.

2. The method according to claim 1, wherein, The step of performing feature transformation on the first facial feature using the first expression feature and the second expression feature to obtain the second facial feature includes: The driving parameters of the first facial expression feature are extracted to obtain the first driving parameters of the first face. The driving parameters of the second facial expression feature are extracted to obtain the second driving parameters of the second face. Based on the first driving parameter and the second driving parameter, feature mapping is performed on the first facial feature to obtain the second facial feature.

3. The method according to claim 2, wherein, The step of performing feature mapping on the first facial features based on the first driving parameters and the second driving parameters to obtain the second facial features includes: Using the first driving parameters, the first facial features are mapped to obtain the first facial features of the first face in a blank expression state. The second facial feature is obtained by performing a second feature mapping on the first facial feature in the expressionless state using the second driving parameter.

4. The method according to any one of claims 1 to 3, wherein, Decoding the second facial features to obtain a third image that matches the expression information of the second object and the facial information of the first object includes: The three-dimensional expression-driven features are decoded to obtain the decoded second facial features; The target background is encoded to obtain background features; The decoded second facial features and the background features are fused to obtain the face-swapped image.

5. The method according to any one of claims 1 to 4, wherein, The image processing method is implemented through an image expression transfer model; the image expression transfer model is trained through the following steps: Acquire sample data; the sample data includes a first sample image, a second sample image, and a ground truth image; The first sample image is subjected to a first data augmentation to obtain a third sample image; The second sample image is subjected to a second data augmentation to obtain a fourth sample image; The facial expression in the fourth sample image is transferred to the third sample image using the image expression transfer model to be trained, thus obtaining the fifth sample image. Loss calculation is performed based on the ground truth image and the fifth sample image to obtain the loss result; Based on the loss result, the model parameters in the image expression transfer model are updated to obtain the trained image expression transfer model.

6. The method according to claim 5, wherein, The first sample image includes a first sample face of a first sample object, and the second sample image includes a second sample face of a second sample object; The step of performing a first data augmentation on the first sample image to obtain a third sample image includes: Facial attributes are extracted from the first sample face and the second sample face respectively, resulting in a first set of facial attributes and a second set of facial attributes. Based on the first set of facial attributes and the second set of facial attributes, the first sample image is subjected to a first image transformation to obtain the third sample image.

7. The method according to claim 5 or 6, wherein, The step of performing a second data augmentation on the second sample image to obtain a fourth sample image includes: Facial key points are extracted from the first sample face and the second sample face respectively, resulting in a first set of facial key points and a second set of facial key points. Based on the first set of facial key points and the second set of facial key points, the second sample image is subjected to a second image transformation to obtain the fourth sample image.

8. The method according to claim 7, wherein, The image expression transfer model includes a 3D feature encoder, an expression encoder, a driver encoder, and a decoder; the third sample image includes a third sample face, and the fourth sample image includes a fourth sample face; The process of transferring the facial expression from the fourth sample image to the third sample image using a trained image expression transfer model to obtain the fifth sample image includes: The third facial features of the third sample are obtained by performing three-dimensional reconstruction on the third sample face using the three-dimensional feature encoder. The facial expression encoder extracts facial expression features from the third sample face and the fourth sample face respectively, thereby obtaining the first sample facial expression feature of the third sample face and the second sample facial expression feature of the fourth sample face. The third facial feature is transformed by the drive encoder based on the first sample expression feature and the second sample expression feature to obtain the fourth facial feature; The decoder decodes the fourth facial feature to obtain the fifth sample image.

9. The method according to claim 8, wherein, The loss calculation based on the ground truth image and the fifth sample image, to obtain the loss result, includes: Based on the ground truth image and the fifth sample image, a generative adversarial loss is calculated to obtain a first loss value; Based on the ground truth image and the fifth sample image, the gaze loss is calculated to obtain the second loss value; Pixel loss is calculated based on the ground image and the fifth sample image to obtain a third loss value; The expression loss is calculated based on M sixth sample images and N seventh sample images to obtain the fourth loss value; M and N are both integers greater than 1. The first loss value, the second loss value, the third loss value, and the fourth loss value are fused to obtain the loss result.

10. The method according to claim 9, wherein, The step of calculating the generative adversarial loss based on the ground truth image and the fifth sample image to obtain the first loss value includes: The fifth sample image is subjected to feature mapping to obtain a first discrimination probability that characterizes the fifth sample image as a natural image; The ground truth image is subjected to feature mapping to obtain a second discrimination probability that characterizes the ground truth image as a natural image; Generative adversarial loss is calculated based on the first discrimination probability and the second discrimination probability to obtain the first loss value.

11. The method according to claim 10, wherein, The step of calculating the generative adversarial loss based on the first discrimination probability and the second discrimination probability to obtain the first loss value includes: The second discrimination probability is logarithmically calculated to obtain the first logarithm result; Subtract the value one from the first discrimination probability to obtain the first subtraction result, and perform logarithmic calculation on the first subtraction result to obtain the second logarithmic calculation result; The first logarithmic calculation result and the second logarithmic calculation result are fused to obtain the first loss value.

12. The method according to claim 9, wherein, The step of calculating the gaze loss based on the ground truth image and the fifth sample image to obtain the second loss value includes: The gaze feature is extracted from the fifth sample image to obtain the first gaze feature of the fifth sample image; The gaze feature is extracted from the ground truth image to obtain the second gaze feature of the ground truth image; Based on the first and second line-of-sight features, line-of-sight loss is calculated to obtain the second loss value.

13. The method according to claim 12, wherein, The step of calculating the gaze loss based on the first gaze feature and the second gaze feature to obtain the second loss value includes: Subtract the first gaze feature from the second gaze feature to obtain the second subtraction result; Squaring each element in the second subtraction result yields the squared result of each element. The second loss value is obtained by summing the squares of the multiple elements and then taking the square root.

14. The method according to claim 9, wherein, The step of calculating pixel loss based on the ground truth image and the fifth sample image to obtain a third loss value includes: Pixel features are extracted from the fifth sample image to obtain the first pixel set of the fifth sample image; Pixel features are extracted from the ground truth image to obtain a second set of pixels in the ground truth image; The third loss value is obtained by calculating pixel loss based on the first pixel set and the second pixel set.

15. The method according to claim 14, wherein, The step of calculating the pixel loss based on the first pixel set and the second pixel set to obtain the third loss value includes: For each pixel location, the following process is performed: a first difference is determined between a first pixel feature corresponding to the pixel location in the first pixel set and a second pixel feature corresponding to the pixel location in the second pixel set; The third loss value is obtained by summing the absolute values ​​of the first differences corresponding to multiple pixel positions.

16. The method according to claim 9, wherein, The expression loss is calculated based on M sixth sample images and N seventh sample images to obtain a fourth loss value, including: Data augmentation is performed on the M sixth sample images and the N seventh sample images respectively to obtain M eighth sample images and N ninth sample images; the M eighth sample images include M fifth sample faces corresponding to M third sample objects respectively, and the N ninth sample images include N sixth sample faces corresponding to N fourth sample objects respectively; Three-dimensional reconstruction is performed on the M fifth sample faces to obtain the third facial features of each of the M fifth sample faces; Expression features are extracted from the M fifth sample faces and the N sixth sample faces respectively, resulting in the third sample expression features of the M fifth sample faces and the fourth sample expression features of the N sixth sample faces respectively. Driving parameters are extracted from the M third sample facial expression features and the N fourth sample facial expression features respectively, resulting in M ​​third driving parameters and N fourth driving parameters. The expression loss is calculated based on the M third facial features, the M third driving parameters, and the N fourth driving parameters to obtain the fourth loss value.

17. An image processing apparatus, comprising: The acquisition module acquires a first image and a second image; the first image includes a first face of a first object, and the second image includes a second face of a second object; The 3D reconstruction module is configured to perform 3D reconstruction on the first face to obtain the first facial features of the first face. The facial expression feature extraction module is configured to extract facial expression features from the first face and the second face respectively, thereby obtaining the first facial expression feature of the first face and the second facial expression feature of the second face. The feature conversion module is configured to perform feature conversion on the first facial feature using the first expression feature and the second expression feature to obtain the second facial feature; The decoding module is configured to decode the second facial features to obtain a third image.

18. An electronic device comprising: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the image processing method according to any one of claims 1 to 16.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, which, when executed by a processor, implement the image processing method according to any one of claims 1 to 16.

20. A computer program product comprising computer-executable instructions or a computer program, wherein the image processing method according to any one of claims 1 to 16 is executed by a processor.

Citation Information

Patent Citations

  • Expression migration method and device, electronic equipment and storage medium

    CN115330980A

  • Expression generation method and device, electronic equipment and storage medium

    CN117115320A

  • Face driving method

    CN118505863A