Face replacement method and device, equipment and medium

By extracting and fusing features from three heterogeneous face recognition models in parallel and combining them with facial key point detection to establish spatial correspondence, the problem of insufficient generalization ability of existing face replacement technology in cross-domain scenarios is solved, high-quality face-changing effects are achieved, and the visual authenticity and user experience in film and television special effects production and financial marketing scenarios are improved.

CN120807704APending Publication Date: 2025-10-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510711298.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing face replacement technology lacks generalization capabilities and cannot achieve high-quality face-swapping effects in scenarios with large differences in race, age, and appearance, resulting in distorted face-swapping effects and poor user experience.

Method used

Three heterogeneous face recognition models (such as Arcface, Insightface, and SphereFace) are used in parallel to extract multi-level feature vectors of the original object's facial image. A more inclusive original face feature vector is constructed through feature fusion, and a spatial correspondence is established through the facial key point detection model to ensure that dynamic attribute features such as expression and posture remain consistent during the replacement process.

Benefits of technology

The accuracy and naturalness of face replacement are improved. The generated face-changing image not only retains the basic identity characteristics of the target object, but also fully conveys the personalized attributes of the original object, reducing the cost of manual image retouching, and improving content production efficiency and user immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807704A_ABST
    Figure CN120807704A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence, and relates to a face replacement method, which comprises the following steps: acquiring a first face image of an original object, extracting features by using three face recognition models, and fusing the features to obtain an original face feature vector; synchronously acquiring a second face image of the target object and extracting a feature vector of the second face image; then key point detection is carried out on the two images through a face key point detection model, and a spatial corresponding relation is established; and based on the spatial correspondence, replacing the attribute features in the target face feature vector with the corresponding attribute features of the original face feature vector, and generating a synthetic feature vector. And finally, inputting the synthesized feature vector into a generative model to generate a face changing image of the target object. The invention further provides a device, equipment and a medium. The method can be applied to the business fields of financial science and technology, medical health and the like, and the accuracy and naturalness of face replacement can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and is applied to online processing business scenarios such as financial technology, and in particular to a face replacement method, device, equipment and medium. Background Art

[0002] With the rapid development of computer vision and deep learning technologies, face replacement technology, as an emerging image and video processing technology, is gradually showing great application potential in entertainment, social networking, film and television production, and marketing content creation. Face replacement technology aims to replace the face in an image or video with the face of a target person through algorithms, while preserving the original face's posture, expression, lighting, and other characteristics as much as possible, thereby achieving high-quality face-swapping results.

[0003] In the entertainment and social sectors, face replacement technology provides users with rich, personalized experiences. For example, users can replace their facial features with cartoon or celebrity images, creating engaging video content and enhancing social interactions. This technology also plays a vital role in film, television, and marketing content production. Using face-swapping technology, production teams can quickly generate new videos with different looks for different audiences, reducing production costs and improving content creation efficiency.

[0004] In addition to the above-mentioned fields, this facial technology can also demonstrate unique value in the field of medical health. In the field of orthodontic surgery, doctors can use facial replacement technology to simulate postoperative effects. By integrating the patient's facial structure with a standardized model, they can intuitively display the changes in facial contours after correction, assisting doctor-patient communication and surgical plan design. In dermatology monitoring scenarios, this technology can perform high-precision feature mapping of skin lesions and provide visual data support for the management of chronic skin diseases by tracking facial images of the same patient at different times over a long period of time. These applications place higher demands on the precise restoration of feature details, such as the need to fully preserve key medical diagnostic information such as skin texture, micro-expression dynamics, and three-dimensional facial structure.

[0005] However, despite the significant progress made in face replacement technology, current mainstream techniques still face many challenges. In particular, in terms of generalization ability, existing technologies perform poorly in scenarios with large differences in age, race, and appearance. For example, when replacing the facial features of a young person onto an image of an old person, existing technologies often struggle to accurately preserve age-related features such as wrinkles and skin sagging, resulting in distorted face replacement effects. In addition, differences in facial structure between different races (such as nose height, eye shape, etc.) can also cause noticeable "dissonance" after face replacement, affecting user experience. The root cause of these problems lies in the limited ability of existing technologies to extract and represent facial features, which cannot fully capture and preserve the subtle differences between different ages, races, and appearance characteristics. Therefore, how to improve the generalization ability of face replacement technology to enable high-quality face replacement effects in various complex scenarios has become a pressing technical problem. SUMMARY

[0006] The purpose of the embodiments of the present application is to propose a face replacement method, device, computer equipment and storage medium to solve the problem that existing face replacement technology lacks generalization ability and cannot achieve high-quality face replacement effects in various complex scenarios.

[0007] In a first aspect, a face replacement method is provided, which employs the following technical solution:

[0008] A first face image of an original object is obtained. A first face recognition model, a second face recognition model, and a third face recognition model are used to extract features from the first face image, respectively, to obtain a first original feature vector, a second original feature vector, and a third original feature vector. The first original feature vector, the second original feature vector, and the third original feature vector are fused to obtain an original face feature vector of the first face image. A target face feature vector of a second face image of a target object is obtained. A face key point detection model is used to detect face key points in the first face image and the second face image, and a spatial correspondence between the face key points of the original object and the face key points of the target object is established. Based on the spatial correspondence, the attribute features in the target face feature vector are replaced with the corresponding attribute features of the original face feature vector to obtain a synthetic feature vector. Based on the synthetic feature vector, a generation model is used to generate a face-replaced image of the target object.

[0009] In a second aspect, a face replacement device is provided, which employs the following technical solution:

[0010] A first acquisition module is configured to obtain a first face image of an original object.

[0011] The extraction module is configured to extract features of the first face image by using the first face recognition model, the second face recognition model, and the third face recognition model respectively to obtain a first original feature vector, a second original feature vector, and a third original feature vector.

[0012] The fusion module is configured to fuse the first original feature vector, the second original feature vector, and the third original feature vector to obtain an original face feature vector of the first face image.

[0013] The second acquisition module is configured to acquire a target face feature vector of a second face image of a target object.

[0014] The establishment module is configured to detect face key points of the first face image and the second face image by using a preset face key point detection model, and establish a spatial correspondence between face key points of the original object and face key points of the target object.

[0015] The replacement module is configured to replace attribute features in the target face feature vector with attribute features corresponding to the original face feature vector based on the spatial correspondence to obtain a synthetic feature vector.

[0016] The generation module is configured to generate a face-changing image of the target object based on the synthetic feature vector by using a preset generation model.

[0017] In a third aspect, a computer device is provided, which adopts the technical scheme as follows:

[0018] The first face image of the original object is acquired, and features of the first face image are extracted by using the first face recognition model, the second face recognition model, and the third face recognition model respectively to obtain a first original feature vector, a second original feature vector, and a third original feature vector. The first original feature vector, the second original feature vector, and the third original feature vector are fused to obtain an original face feature vector of the first face image. A target face feature vector of a second face image of a target object is acquired. Face key points of the first face image and the second face image are detected by using a preset face key point detection model, and a spatial correspondence between face key points of the original object and face key points of the target object is established. Attribute features in the target face feature vector are replaced with attribute features corresponding to the original face feature vector based on the spatial correspondence to obtain a synthetic feature vector. A face-changing image of the target object is generated based on the synthetic feature vector by using a preset generation model.

[0019] In a fourth aspect, a computer readable storage medium is provided, which adopts the technical scheme as follows:

[0020] obtaining a first face image of an original object; performing feature extraction on the first face image by using a preset first face recognition model, a second face recognition model and a third face recognition model respectively to obtain a first original feature vector, a second original feature vector and a third original feature vector; fusing the first original feature vector, the second original feature vector and the third original feature vector to obtain an original face feature vector of the first face image; obtaining a target face feature vector of a second face image of a target object; performing face key point detection on the first face image and the second face image by using a preset face key point detection model to establish a spatial correspondence relationship between face key points of the original object and face key points of the target object; replacing attribute features in the target face feature vector with attribute features corresponding to the original face feature vector based on the spatial correspondence relationship to obtain a synthetic feature vector; and generating a face-swapped image of the target object based on the synthetic feature vector by using a preset generation model.

[0021] Compared with the prior art, the embodiments of the present application have the following beneficial effects: by using three heterogeneous face recognition models to extract multi-level feature vectors of the face image of the original object in parallel, using the differential capture capabilities of different models for race, age and other feature dimensions, and constructing a more inclusive original face feature vector through a feature fusion strategy, the limitations of a single model in cross-domain feature extraction are effectively overcome. In the feature replacement stage, the spatial geometric constraint relationship established by the face key point detection model can accurately align the face structures of the original object and the target object, ensure that dynamic attribute features such as expressions and postures remain spatially consistent during the replacement process, and avoid misplacement of facial features or texture distortion due to differences in facial structures. The finally generated face-swapped image not only retains the basic identity features of the target object, but also completely transfers the personalized attributes of the original object, improving the accuracy and naturalness of face replacement. In the financial marketing scenarios (such as customized advertisement production) that require high visual authenticity, such as film and television special effect production and virtual anchor generation, the cost of manual retouching can be significantly reduced, and the content output efficiency and user immersion can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the schemes in the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0023] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0024] Figure 2 Flowchart of an embodiment of the face replacement method according to the present application;

[0025] Figure 3 is a structural schematic diagram of an embodiment of a face replacement device according to the present application;

[0026] Figure 4 is a structural schematic diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application; the use herein of terms such as "comprise" and "have" and any variations thereof are intended to cover a non-exclusive inclusion; the use herein of terms such as "first", "second" and the like are intended to distinguish between similar objects unless the context indicates otherwise.

[0028] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase that "an embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive or alternative embodiments. It is expressly understood that the embodiments described herein are merely examples from a whole class of comparable embodiments which those skilled in the art will readily appreciate. It is further understood that the descriptions of various embodiments do not imply that the compositions are an exhaustive list of possible implementations.

[0029] In order to make the technical personnel in the art better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings.

[0030] As shown in Figure 1 , the system architecture 100 can include a terminal device 101, a network 102 and a server 103, and the terminal device 101 can be a notebook computer 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0031] A user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0032] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing, in addition to the notebook computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an electronic book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer and a desktop computer, etc.

[0033] The server 103 can be a server providing various services, for example, a background server providing support for a page displayed on the terminal device 101.

[0034] It should be noted that the face replacement method provided by the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the face replacement apparatus is generally arranged in the server / terminal device.

[0035] It should be understood that, Figure 1 The number of terminal devices, networks and servers in

[0036] With reference to Figure 2 , a flow chart of one embodiment of the method for service recommendation according to the present application is shown. The face replacement method comprises the following steps:

[0037] Step S201, a first face image of an original object is acquired.

[0038] In the present embodiment, the electronic device (for example, the server / terminal device shown in Figure 1 ) on which the face replacement method runs can acquire the first face image of the original object through a wired connection mode or a wireless connection mode. It should be noted that the wireless connection mode can include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other now known or future developed wireless connection modes.

[0039] The original object refers to an individual whose facial features will be extracted and applied to the face of the target object in face swapping technology. The original object is the source data provider for face swapping operations, and its facial features, such as facial feature shapes, expressions, and poses, will be used to replace the facial features of the target object. For example, in entertainment applications, users can choose to replace their own faces onto cartoon characters or celebrity images.

[0040] The first face image is the facial image data of the original object, which is used for subsequent feature extraction and face swapping operations. This image is usually taken by a camera or obtained from an existing image library and contains complete facial information of the original object. For example, a selfie uploaded by a user can be used as the first face image for extracting its facial features.

[0041] In step S202, the first face image is subjected to feature extraction using a preset first face recognition model, a second face recognition model, and a third face recognition model, respectively, to obtain a first original feature vector, a second original feature vector, and a third original feature vector.

[0042] The first face recognition model is a deep learning model used to extract facial features from the first face image. Taking the Arcface model as an example, it learns deep feature representations of faces through training. This model can capture key information in facial images, such as facial feature shapes and textures, and encode them into feature vectors.

[0043] The second face recognition model is another deep learning model used for feature extraction. Taking the Insightface model as an example, it can also extract features with discriminative power from facial images. Compared with the first face recognition model, the second face recognition model can use different network structures or training strategies to capture different aspects of facial features. For example, the Insightface model has better generalization ability when dealing with cross-racial and cross-age faces.

[0044] The third face recognition model is a deep learning model used to further enrich feature representations. Taking the SphereFace model as an example, it enhances the discriminability of feature vectors by introducing an angular margin loss function. This model can extract complementary feature information to the first and second face recognition models, thereby improving the representation ability after feature fusion. For example, the SphereFace model can more accurately distinguish different individuals with similar facial structures.

[0045] Feature extraction refers to the process of extracting representative feature information from a face image using a face recognition model. These feature information is usually represented in the form of a feature vector, which can reflect the key attributes of the face image. For example, through feature extraction, the facial shape, texture, and other features of the original subject can be extracted from the first face image.

[0046] The first original feature vector is the feature representation obtained by the first face recognition model after feature extraction on the first face image. This vector contains key information of the original subject's face image, such as facial feature shape, pose, etc.

[0047] The second original feature vector is the feature representation obtained by the second face recognition model after feature extraction on the first face image. Compared with the first original feature vector, the second original feature vector may capture different aspects or details of the face image.

[0048] The third original feature vector is the feature representation obtained by the third face recognition model after feature extraction on the first face image. This vector, together with the first and second original feature vectors, constitutes a complete representation of the original subject's facial features.

[0049] Step S203, the first original feature vector, the second original feature vector and the third original feature vector are fused to obtain the original face feature vector of the first face image.

[0050] Fusion refers to the process of combining or weighted averaging the first original feature vector, the second original feature vector and the third original feature vector to obtain a more comprehensive and accurate original face feature vector. The fusion process can fully utilize the advantages of different face recognition models, improve the robustness and discriminability of feature representation. For example, through weighted averaging, the information of the three feature vectors can be integrated to obtain a more representative original face feature vector.

[0051] The original face feature vector is obtained by fusing the first original feature vector, the second original feature vector and the third original feature vector. This vector contains all the key information of the original subject's face image.

[0052] Step S204, obtaining the target face feature vector of the second face image of the target object.

[0053] The target object refers to the individual whose face will be replaced by the facial features of the original object in face replacement technology. The target object is the receiver of the face replacement operation, whose facial features will be modified or replaced. For example, in film and television production, the production team may choose to replace the face of an actor with the face of a target role to achieve rapid role conversion.

[0054] wherein the second face image is a facial image data of the target object, for subsequent feature extraction and face swapping operation. Similar to the first face image, the second face image also contains complete facial information of the target object. For example, the facial image of the target character intercepted from a movie clip can be used as the second face image.

[0055] wherein the target face feature vector is a feature representation obtained by feature extraction on the second face image. For example, the first face recognition model, the second face recognition model, and the third face recognition model are used to extract features from the second face image and fuse them to obtain the target face feature vector. The vector contains key information of the target object's facial image, such as facial feature shape and expression. In the face swapping process, the target face feature vector will be replaced as the object to be replaced, and part of its attribute features will be replaced by the corresponding attribute features in the original face feature vector.

[0056] Step S205, a pre-set face key point detection model is used to detect the face key points of the first face image and the second face image, and a spatial correspondence relationship between the face key points of the original object and the face key points of the target object is established.

[0057] wherein the face key point detection model is a deep learning model used to detect the positions of key points in a face image. This model can identify key points (such as eye corners, nose tips, and mouth corners) in a facial image, providing a basis for subsequent spatial correspondence establishment and feature replacement. For example, through the face key point detection model, the position information of the face key points of the original object and the target object can be determined.

[0058] wherein the face key point detection refers to the process of using the face key point detection model to process the first face image and the second face image to determine the positions of the face key points. This process can extract key geometric information from the face image, providing accurate positioning basis for subsequent feature replacement and image generation. For example, through face key point detection, the alignment points of the faces of the original object and the target object can be determined.

[0059] wherein the spatial correspondence relationship refers to the geometric mapping relationship between the face key points of the original object and the face key points of the target object. This relationship is established through face key point detection and is used to guide subsequent feature replacement operations. For example, in the face swapping process, the spatial correspondence relationship can ensure that the facial features of the original object are accurately mapped to the facial structure of the target object.

[0060] Step S206, based on the spatial correspondence relationship, the attribute features in the target face feature vector are replaced by the corresponding attribute features in the original face feature vector to obtain a synthesized feature vector.

[0061] Among them, the attribute feature in the target face feature vector refers to the part of the target object's face feature vector that represents a specific facial attribute (such as expression, lighting, posture, etc.). In the face swapping process, these attribute features will be replaced by the corresponding attribute features in the original face feature vector. For example, the expression feature in the target face feature vector will be replaced by the expression feature in the original face feature vector to achieve the transfer of expression.

[0062] Among them, the attribute feature corresponding to the original face feature vector refers to the part of the original object's face feature vector that corresponds to the attribute feature in the target face feature vector. These features will be used to replace the corresponding attribute features in the target face feature vector in the face swapping process. For example, the eye shape feature in the original face feature vector will be used to replace the eye shape feature in the target face feature vector.

[0063] Among them, the synthesized feature vector is an intermediate result obtained in the feature replacement process, which contains the identity features (such as facial contours, feature positions, etc.) of the target face feature vector and the attribute features (such as expression, lighting, etc.) of the original face feature vector. This vector is the basic data for generating face-swapped images.

[0064] Step S207, based on the synthesized feature vector, a pre-set generation model is used to generate a face-swapped image of the target object.

[0065] Among them, the generation model is a deep learning model used to generate face-swapped images based on the synthesized feature vector. This model can learn the mapping relationship between facial features and images, and thus generate realistic face-swapped images based on the input synthesized feature vector. For example, generative adversarial networks or variational autoencoders can be used for this purpose.

[0066] Among them, the face-swapped image is an image data generated by the generation model based on the synthesized feature vector. This image retains the identity features (such as facial contours, feature positions, etc.) of the target object, while presenting the facial attribute features (such as expression, lighting, etc.) of the original object. For example, in entertainment applications, the generated face-swapped image may demonstrate the effect of replacing the user's face with a cartoon character or a celebrity's image.

[0067] The embodiments of the present application can overcome the limitations of single model in cross-domain feature extraction by adopting three heterogeneous face recognition models to extract multi-level feature vectors of the original object's face image, utilizing the differential capturing ability of different models for race, age and other feature dimensions, and constructing a more inclusive original face feature vector through feature fusion strategy. In the feature replacement stage, the spatial geometric constraint relationship established by the face key point detection model can accurately align the face structure of the original object and the target object, ensuring that the dynamic attribute features such as expression and posture remain spatial consistency during the replacement process, avoiding the misplacement of facial features or texture distortion caused by the difference in facial structure. The final face replacement image not only retains the basic identity features of the target object, but also completely transfers the personalized attributes of the original object, improving the accuracy and naturalness of face replacement. In the financial marketing scenarios (such as customized advertisement production) that require strict visual authenticity such as film and television special effect production and virtual anchor generation, the cost of manual retouching can be significantly reduced, and the content output efficiency and user immersion can be improved.

[0068] In some optional implementations of the present embodiment, step 203, the first original feature vector, the second original feature vector, and the third original feature vector are fused to obtain the original face feature vector of the first face image, specifically including the following steps:

[0069] Based on the preset feature alignment rule, the first original feature vector, the second original feature vector, and the third original feature vector are reduced to obtain the first feature vector, the second feature vector, and the third feature vector; the first feature vector, the second feature vector, and the third feature vector are averaged and weighted summed to obtain the original face feature vector of the first face image.

[0070] The feature alignment rule refers to a series of standards or methods used when processing original feature vectors extracted from different face recognition models to ensure that these feature vectors are consistent or comparable in dimension, scale, or other related attributes. In the present embodiment, the feature alignment rule is used to reduce the first, second, and third original feature vectors so that subsequent fusion operations can proceed smoothly.

[0071] The reduction refers to a series of transformations or adjustments of the feature vector according to the preset feature alignment rule to meet specific requirements or standards. For example, the shape of the feature vector is reduced, including adjusting the length, width, and dimension of the vector to ensure that they have the same structure, facilitating subsequent average weighted sum operations.

[0072] The first feature vector refers to a feature vector obtained by reducing the first original feature vector. The vector is derived from the result of feature extraction of the first face image by the first face recognition model (such as the Arcface model), and is processed by the reduction rule of feature alignment. The first feature vector represents the feature representation of the original object face image under certain reduction conditions, which is used for subsequent fusion with other feature vectors. For example, the first feature vector after reduction has the same dimension and scale as the second and third feature vectors, facilitating the average weighted sum operation.

[0073] The second feature vector refers to a feature vector obtained by reducing the second original feature vector. The vector is derived from the result of feature extraction of the first face image by the second face recognition model (such as the Insightface model), and is processed by the reduction rule of feature alignment. The second feature vector represents the reduced feature representation of the original object face image under another feature extraction method, which is also used for subsequent feature fusion. For example, the second feature vector may retain some specific features of the original object face image after reduction, such as nose height, eye shape, etc., which can provide complementary information when fused with other feature vectors.

[0074] The third feature vector refers to a feature vector obtained by reducing the third original feature vector. The vector is derived from the result of feature extraction of the first face image by the third face recognition model (such as the SphereFace model), and is processed by the reduction rule of feature alignment. The third feature vector represents the reduced feature representation of the original object face image under another feature extraction method, which is used together with other feature vectors to generate the original face feature vector. For example, the third feature vector may extract unique features of the original object face image through different network structures or training strategies, which can enhance the representation ability of the original face feature vector when fused with other feature vectors.

[0075] In an example, in the financial field, marketing video production often needs to customize the image for different customer groups to enhance the attractiveness of advertising. For example, a bank needs to produce promotional videos for different age groups (young, middle-aged, and old) of customer groups to demonstrate the universality of its financial products. First, the original face image of a young customer can be obtained, and the first, second, and third original feature vectors are extracted by inputting the Arcface, Insightface, and SphereFace models respectively. Based on the feature alignment rules (such as uniforming the dimension to 256, L2 norm normalization), the three original feature vectors are reduced to obtain the first, second, and third feature vectors. This step ensures that the features extracted by different models are consistent in dimension and scale, providing a basis for subsequent fusion. The first, second, and third reduced feature vectors are averaged and weighted (the weight is 1 / 3), and the original face feature vector is obtained. This vector combines the age feature capturing capabilities of different models, such as the sensitivity of the Arcface model to young facial details, the Insightface model's ability to analyze facial structures, and the SphereFace model's robustness to lighting changes. The second face image of a middle-aged or old customer is obtained, and the target face feature vector is extracted. Through a face key point detection model, a spatial correspondence relationship is established, and the age-related attributes (such as wrinkles and skin relaxation) in the target face feature vector are replaced with the corresponding attributes of the original face feature vector to generate a synthetic feature vector. Finally, a generated model is used to generate a marketing video frame after face replacement, showing the natural integration of young customer facial features in middle-aged or old images.

[0076] In an example, in the field of medical monitoring, the technical solutions of the present embodiment can be applied to preoperative effect simulation of orthopedic surgery, assisting doctors and patients in visual communication of surgical plans. For example, a certain orthopedic surgery of a first-class hospital needs to provide personalized surgical plan simulation for rhinoplasty patients to improve the efficiency of doctor-patient decision-making. First, the original face image of the patient (such as a preoperative front view) is collected as the baseline data for feature extraction. Then, three face recognition models, Arcface, Insightface, and SphereFace, are used to extract features from the original image to obtain the first, second, and third original feature vectors. For example, the Arcface model can accurately capture the subtle features of the patient's nose contour, the Insightface model can analyze the three-dimensional structure information of the face, and the SphereFace model can enhance the robustness to light differences. Based on the preset alignment rules (such as uniforming the dimension to 512, PCA dimension reduction and denoising), the three original feature vectors are reduced to eliminate the differences between models. Then, through average weighted summation (weights are 0.4, 0.35, and 0.25, respectively, according to the model's analysis accuracy of the nose area), the original face feature vector that integrates facial details and structural information is generated. Collect medical template images containing ideal nose types (such as standard models in the golden triangle nose type database), and extract their target face feature vectors. Through a face key point detection model (such as dlib or MediaPipe), the nose key points (such as 68 key points including nose tip, nose wing, and nose bridge) of the patient's original image and the template image are located, and an anatomical space mapping is constructed. Replace the attributes related to the nose type (such as the height of the nose bridge, the width of the nose wing, and the angle of the nose tip) in the target feature vector with the corresponding attributes in the patient's original feature vector. For example, the patient's nose skin texture features are retained, and the nose bridge curvature parameters of the template are integrated into the synthesized feature vector. Based on the synthesized feature vector, a generative adversarial network is used to generate postoperative simulation images to show the natural fusion effect of the patient's facial features and the ideal nose type.

[0077] The embodiments of the present application can effectively solve the difference problems existing in the dimensions, scales and representation forms of the feature vectors output by different face recognition models by respectively performing reduction processing on the first, second and third original feature vectors based on the preset feature alignment rules, ensuring that each feature vector has a unified format and comparability, and providing a reliable foundation for subsequent feature fusion. Subsequently, the first, second and third reduced feature vectors are averaged and weighted, which can fully integrate the unique capturing capabilities of different models for face features, forming an original face feature vector with more robustness and generalization. This fusion method combines the advantages of each model, such as the accuracy of the Arcface model in identity recognition and the fineness of the Insightface model in analyzing facial structures, so that the original face feature vector can more comprehensively and accurately represent the various attributes of the face, thereby providing more abundant and accurate feature replacement basis for the target object in the subsequent face changing process, and significantly improving the quality and authenticity of the face changed image.

[0078] In some optional implementations, step 203, the first, second and third original feature vectors are fused to obtain an original face feature vector of the first face image, specifically including the following steps:

[0079] A target domain test set corresponding to the first face image is obtained; the performance of the first, second and third face recognition models is evaluated based on the target domain test set to obtain a first accuracy rate of the first face recognition model, a second accuracy rate of the second face recognition model and a third accuracy rate of the third face recognition model; the first, second and third accuracy rates are normalized to obtain normalized first, second and third accuracy rates; based on the normalized first, second and third accuracy rates, first, second and third weights of the first, second and third face recognition models are obtained; the first, second and third original feature vectors and the first, second and third weights are weighted and summed to obtain an original face feature vector of the first face image.

[0080] The target domain test set is a set of representative image or video data collected and sorted from a target application scenario (such as a complex scenario of cross-race, cross-age, different appearance characteristics, etc.). The test set is used to evaluate the performance of the face recognition model in the target domain to ensure that the model can adapt to and accurately process the data in the domain. For example, in the cross-race face changing scenario, the target domain test set can include face images of different races to evaluate the accuracy and robustness of the model in processing different facial structures.

[0081] The performance evaluation refers to testing the face recognition model on the target domain test set and quantitatively analyzing the performance of the model on the test set. It represents the accuracy and efficiency of the model on a specific task. For example, by calculating the accuracy, recall rate and other indicators of the model on the target domain test set, the performance of the model in different scenarios can be evaluated, and the model parameters or model structure can be adjusted accordingly.

[0082] The first accuracy refers to the recognition accuracy of the first face recognition model on the target domain test set. It is used to measure the performance of the first face recognition model in the target domain and provides a basis for subsequent weight allocation. For example, if the first accuracy is high, it means that the model has good recognition ability in the target domain and can be given a higher weight in the weighted sum.

[0083] The second accuracy refers to the recognition accuracy of the second face recognition model on the target domain test set. It represents the performance of the second face recognition model in the target domain and is also used for subsequent weight allocation.

[0084] The third accuracy refers to the recognition accuracy of the third face recognition model on the target domain test set. It represents the performance of the third face recognition model in the target domain and is also used for subsequent weight allocation.

[0085] The normalization is used to convert data of different dimensions or orders of magnitude to a unified scale or range, which can eliminate the dimensional differences between different model accuracies and make the subsequent weight allocation more reasonable and accurate.

[0086] The first weight refers to the weight value allocated to the first face recognition model according to its normalized accuracy on the target domain test set. It is used to determine the contribution of the first face recognition model in generating the original face feature vector.

[0087] The second weight refers to the weight value allocated to the second face recognition model according to its normalized accuracy on the target domain test set. The second weight represents the relative importance of the second face recognition model in the weighted sum process and is used to determine the contribution of the model in generating the original face feature vector.

[0088] The third weight refers to the weight value allocated to the third face recognition model according to its normalized accuracy on the target domain test set. It is used to determine the contribution of the model in generating the original face feature vector.

[0089] In an example, for a financial anti-fraud scene, a financial user face image dataset containing different races, ages, genders, and facial expressions can be collected to simulate the complex facial feature situations that may be encountered in real business as a target field test set. Using Arcface, Insightface, and SphereFace three face recognition models, performance evaluation is carried out on the target field test set respectively, and the first accuracy of the Arcface model (such as 92%), the second accuracy of the Insightface model (such as 88%), and the third accuracy of the SphereFace model (such as 85%) are obtained. The accuracy of each model is normalized to obtain the normalized accuracy value, and the weights are allocated accordingly, such as the first weight 0.4, the second weight 0.35, and the third weight 0.25, to ensure that the model with better performance contributes more in feature fusion. The first, second, and third original feature vectors extracted by each model are combined with the above weights for weighted summation to generate a more robust original face feature vector.

[0090] The embodiments of the present application can evaluate the performance of each face recognition model in a specific application scenario (such as cross-race, cross-age, etc.) by obtaining the target field test set corresponding to the first face image, ensuring that the evaluation results are highly consistent with the actual needs. Based on the performance evaluation and accuracy normalization of the target field test set, the quantitative comparison of the recognition capabilities of different models is realized, and reasonable weights are allocated accordingly to avoid feature extraction bias caused by the performance limitations of a single model. Finally, by weighted summation of multiple original feature vectors and their corresponding weights, the advantages of each model can be utilized to generate a more comprehensive and accurate original face feature vector. This technical effect not only enhances the applicability of face replacement technology in complex scenarios, but also provides a high-quality feature basis for subsequent face replacement image generation, thereby effectively improving the realism and naturalness of the face replacement effect.

[0091] In some optional implementations, the step of "obtaining a target field test set corresponding to the first face image" specifically includes the following steps:

[0092] Obtaining a plurality of face images similar to the scene of the first face image to obtain an image set; performing identity labeling on each face image of the image set to obtain a labeled image set; dividing the labeled image set into a training set and a test set to obtain the target field test set.

[0093] Among them, the plurality of face images refers to a collection of different individual facial features images similar to the scene of the first face image collected from different sources (such as databases, web crawling, actual shooting, etc.).

[0094] Identity annotation refers to assigning a unique identity label to each face image in the image set to specify the individual that the image belongs to. It is used to support subsequent model training and performance evaluation, ensuring that the model can accurately recognize and distinguish the facial features of different individuals. For example, when constructing the labeled image set, each face image is annotated with an identity label such as "Zhang San" or "Li Si" to clearly specify the class to which each sample belongs during model training.

[0095] The training set refers to a portion of image data divided from the labeled image set for model training. For example, during the training process of a face recognition model, the labeled image set is divided into a training set and a test set, where the training set is used for parameter updating and optimization of the model to improve the recognition accuracy of the model on unknown data.

[0096] The test set refers to another portion of image data divided from the labeled image set for model performance evaluation. For example, after completing the training of the face recognition model, the test set is used to evaluate the performance of the model by comparing the predicted results of the model on the test set with the true labels to calculate the accuracy, recall rate, and other indicators of the model to verify the effectiveness of the model.

[0097] In an example, multiple face images similar to the remote identity verification scenario can be collected, including but not limited to user self-portrait images under different indoor and outdoor lighting conditions, different background environments (such as home, office, public places, etc.). These images need to cover users of different age groups, genders, and races to ensure the diversity and representativeness of the image set. Through data cleaning and preprocessing, low-quality, blurred, or severely occluded images are removed to obtain a high-quality image set. Each face image in the image set is annotated with an identity label. In the financial scenario, manual or automatic identity matching and annotation can be performed based on the identity documents (such as ID cards, passports, etc.) submitted by users during registration to ensure accurate correspondence between each image and the real user identity. After annotation, a labeled image set is formed for subsequent model training and evaluation. The labeled image set is divided into a training set and a test set in proportion (such as 7:3) to obtain a target domain test set. The training set is used for learning and optimizing the model parameters, and the test set is used to evaluate the generalization ability of the model on unknown data. Through cross-validation and other methods, the rationality of the division of the training set and the test set is ensured to avoid evaluation bias caused by data leakage.

[0098] The embodiments of the present application can ensure high consistency of training data and actual application environment through the image set similar to the scene, so that the model can learn more close to the feature representation of the real scene, effectively alleviating the performance decline problem caused by environmental differences. The identity label provides a clear supervision signal for the model, which helps the model to learn more discriminative facial features. By dividing the labeled image set into training set and test set, the performance of the model is effectively evaluated and optimized, avoiding the occurrence of overfitting phenomenon. Finally, the generated target field test set provides a high-quality data basis for subsequent face recognition model training and evaluation, which helps to improve the recognition accuracy and robustness of the model in complex scenes.

[0099] In some optional implementations, the step S204 of obtaining the target face feature vector of the second face image of the target object specifically includes the following steps:

[0100] Obtaining a second face image of a target object; respectively using a first face recognition model, a second face recognition model and a third face recognition model to perform feature extraction on the second face image to obtain a first target feature vector, a second target feature vector and a third target feature vector; and fusing the first target feature vector, the second target feature vector and the third target feature vector to obtain a target face feature vector of the second face image.

[0101] The first target feature vector refers to a vector representation obtained by performing feature extraction on the second face image of the target object using the first face recognition model. It represents the unique facial features and structural information of the second face image from the perspective of the first face recognition model.

[0102] The second target feature vector refers to a vector representation obtained by performing feature extraction on the second face image of the target object using the second face recognition model. It represents the position and attributes of the second face image in the facial feature space defined by the second face recognition model.

[0103] The third target feature vector refers to a vector representation obtained by performing feature extraction on the second face image of the target object using the third face recognition model. It represents the information of the second face image in the facial feature dimension concerned by the third face recognition model.

[0104] In an example, in the process of film and television production, a second face image of a target object (such as a body double) can be collected by a high-definition camera to ensure that the image quality meets the subsequent processing requirements. The second face image is subjected to feature extraction by using a preset first face recognition model, a second face recognition model and a third face recognition model respectively. A first target feature vector (representing the overall structure of the face), a second target feature vector (representing the detailed features of the face) and a third target feature vector (representing the multi-scale fusion features) are obtained. The first target feature vector, the second target feature vector and the third target feature vector extracted are fused. In the fusion process, according to the importance and complementarity of each feature vector, methods such as average weighted summation, weighted average or attention mechanism can be used to effectively integrate the feature information extracted by different models to obtain the target face feature vector of the second face image. The vector integrates the overall structure, detailed features and multi-scale information of the face of the target object, providing a rich feature basis for subsequent face replacement.

[0105] After obtaining the second face image of the target object, the embodiments of the present application can use the first face recognition model, the second face recognition model and the third face recognition model for feature extraction respectively. Based on the unique algorithm design of each model, different levels and dimensions of feature information in the face image can be captured, such as facial contour, local texture and key point distribution, so as to obtain the first target feature vector, the second target feature vector and the third target feature vector. Fusing these feature vectors can comprehensively utilize the advantages of each model and overcome the limitations of a single model, to obtain a more rich and robust target face feature vector. The feature vector not only contains the basic facial features of the target object, but also integrates the unique understanding and representation of features by different models, providing a solid and accurate feature basis for subsequent generation of face replacement images, which helps to improve the naturalness and realism of face replacement effect.

[0106] In some optional implementations, step S206, fusing the first target feature vector, the second target feature vector and the third target feature vector to obtain the target face feature vector of the second face image, specifically includes the following steps:

[0107] Based on the preset feature alignment rule, the first target feature vector, the second target feature vector and the third target feature vector are reduced respectively to obtain the first standardized feature vector, the second standardized feature vector and the third standardized feature vector. The first standardized feature vector, the second standardized feature vector and the third standardized feature vector are subjected to average weighted summation to obtain the target face feature vector of the second face image.

[0108] In an example, in a film and television shooting site, a high-definition camera is used to collect multi-angle and multi-expression second face images of a target object (such as a stand-in actor), to ensure that the images contain rich facial information. A preset first face recognition model (good at extracting overall contour features of a face), a second face recognition model (specializing in extracting texture and detail features of a face), and a third face recognition model (fusing facial expression and posture change features) are respectively used to extract features of the second face image, to obtain a first target feature vector, a second target feature vector, and a third target feature vector. Based on a preset feature alignment rule, the three extracted target feature vectors are reduced. For example, a normalization method is used to adjust the numerical range of each feature vector to a unified interval, to obtain a first standardized feature vector, a second standardized feature vector, and a third standardized feature vector, to eliminate the dimensional differences between features of different models. The three standardized feature vectors are averaged and weighted to obtain a target face feature vector of the second face image. This vector integrates the overall structure, texture details, and expression and posture information of the face of the target object, and provides a comprehensive and consistent feature representation for subsequent face-changing special effects.

[0109] The embodiments of the present application can extract features of the second face image through the first face recognition model, the second face recognition model, and the third face recognition model, to obtain three target feature vectors of different dimensions. These vectors describe the facial features of the target object from different angles, and enrich the diversity of feature information. Subsequently, the feature vectors are reduced based on the preset feature alignment rule, to obtain standardized feature vectors. This step effectively eliminates the dimensional differences and scale inconsistencies between features extracted by different models, and ensures the comparability and compatibility of the features. Finally, the target face feature vector obtained by averaging and weighting the standardized feature vectors not only integrates the advantages of multiple models, but also maintains the consistency and stability of the features.

[0110] In some optional implementations, step S206, the first target feature vector, the second target feature vector, and the third target feature vector are fused to obtain a target face feature vector of the second face image, specifically including the following steps:

[0111] The first target feature vector, the second target feature vector, the third target feature vector, the first weight, the second weight, and the third weight are weighted and summed to obtain a target face feature vector of the second face image.

[0112] In an example, in a film shooting scene, a second face image of a target object (such as a stand-in actor) is collected by using a high-definition camera device, ensuring that the image is clear and the angle is appropriate to capture the facial details of the target object. The second face image is subjected to feature extraction by using a first face recognition model (good at extracting facial contour and structural features), a second face recognition model (good at extracting facial expression and micro-expression features), and a third face recognition model (good at extracting robust features such as facial illumination and posture changes), to obtain a first target feature vector, a second target feature vector, and a third target feature vector. According to the performance, accuracy, and robustness of each face recognition model in previous face swapping tasks, the first weight of the first face recognition model, the second weight of the second face recognition model, and the third weight of the third face recognition model are determined. For example, if the first model performs well in facial contour recognition, it is given a higher weight. The first, second, and third target feature vectors extracted by the first, second, and third face recognition models are weighted and summed according to the first, second, and third weights. The weights are set according to the performance, accuracy, and contribution to face swapping of each model in feature extraction, such as a first face recognition model weight of 0.5 (good at overall structural feature extraction), a second face recognition model weight of 0.3 (good at detailed feature extraction), and a third face recognition model weight of 0.2 (providing auxiliary features). The three target feature vectors and their corresponding weights are weighted and summed to obtain a target face feature vector of the second face image. This vector combines the unique capturing ability of multiple models for facial features of the target object, enhancing the richness and robustness of the features.

[0113] The embodiments of the present application can extract different feature dimensions of the target face image by using the first, second, and third face recognition models to obtain the first, second, and third target feature vectors, which describe the facial characteristics of the target object from multiple angles. By determining the weights of the models, different importance can be given to the features extracted by different models according to their advantages and focuses in feature extraction. Then, the feature vectors and their corresponding weights are weighted and summed, so that the final target face feature vector not only combines the advantages of each model, but also highlights the key features. This processing method enhances the representation ability of the feature vector for the facial information of the target object, provides a more accurate and comprehensive feature basis for subsequent face processing tasks (such as face swapping and expression transfer), and helps to improve the quality and stability of the overall processing effect.

[0114] It should be emphasized that, in order to further ensure the privacy and security of the first face image, the first original feature vector, the second original feature vector, the third original feature vector, the second face image, and the face-swapped image, the first face image, the first original feature vector, the second original feature vector, the third original feature vector, the second face image, and the face-swapped image can also be stored in a node of a blockchain.

[0115] The blockchain referred to in the present application is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. The blockchain is essentially a decentralized database, which is a series of data blocks associated using cryptographic methods, each data block containing information of a batch of network transactions, used to verify the validity (anti-fake) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.

[0116] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Artificial intelligence (AI) is the use of digital computers or computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0117] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0118] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by computer readable instructions instructing related hardware, and the computer readable instructions can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of each method. The storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0119] It should be understood that although each step in the flowchart of the accompanying drawings is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other orders. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or sub-steps or stages of other steps.

[0120] Further referring to Figure 3 , as an implementation of the method shown in the above Figure 2 , the present application provides an embodiment of a face replacement device, which corresponds to the method embodiment shown in Figure 2 , and the device can be applied in various electronic devices.

[0121] As shown in Figure 4 , the face replacement device 400 of the present embodiment comprises a first acquisition module 401, an extraction module 402, a fusion module 403, a second acquisition module 404, an establishment module 405, a replacement module 406 and a generation module 407. Among them:

[0122] The first acquisition module 401 is configured to acquire a first face image of an original subject.

[0123] The extraction module 402 is configured to extract features of the first face image by using a preset first face recognition model, a second face recognition model and a third face recognition model respectively, to obtain a first original feature vector, a second original feature vector and a third original feature vector.

[0124] The fusion module 403 is configured to fuse the first original feature vector, the second original feature vector and the third original feature vector to obtain an original face feature vector of the first face image.

[0125] The second acquisition module 404 is configured to acquire a target face feature vector of a second face image of a target subject.

[0126] The establishment module 405 is configured to detect face key points of the first face image and the second face image by using a preset face key point detection model, and establish a spatial correspondence between the face key points of the original subject and the face key points of the target subject.

[0127] The replacement module 406 is configured to replace attribute features in the target face feature vector with attribute features corresponding to the original face feature vector based on the spatial correspondence, to obtain a synthetic feature vector.

[0128] The generation module 407 is configured to generate a face-changing image of the target subject by using a preset generation model based on the synthetic feature vector.

[0129] In this embodiment, the multi-level feature vectors of the original face image of the original object can be extracted in parallel by adopting three heterogeneous face recognition models, the different models can capture the differences in the dimensions of race and age, and a more inclusive original face feature vector can be constructed by a feature fusion strategy, so as to effectively overcome the limitations of a single model in cross-domain feature extraction. In the feature replacement stage, the spatial geometric constraint relationship established by the face key point detection model can accurately align the face structures of the original object and the target object, ensure that the dynamic attribute features such as expressions and postures remain spatially consistent during the replacement process, and avoid misplacement of facial features or texture distortion caused by differences in facial structures. The final face replacement image not only retains the basic identity features of the target object, but also completely transfers the personalized attributes of the original object, improves the accuracy and naturalness of face replacement, and significantly reduces the cost of manual retouching in the financial marketing scenarios (such as customized advertisement production) that require high visual authenticity, such as special effects production and virtual anchor generation, and improves the content output efficiency and user immersion.

[0130] In an embodiment, the fusion module 403 comprises:

[0131] The reduction submodule is configured to reduce the first original feature vector, the second original feature vector, and the third original feature vector respectively based on a preset feature alignment rule to obtain a first feature vector, a second feature vector, and a third feature vector.

[0132] The average submodule is configured to perform average weighted summation on the first feature vector, the second feature vector, and the third feature vector to obtain an original face feature vector of the first face image.

[0133] The embodiments of the present application can effectively solve the difference problems in the dimensions, scales, and representation forms of the feature vectors output by different face recognition models by performing reduction processing on the first, second, and third original feature vectors based on the preset feature alignment rule, ensuring that each feature vector has a unified format and comparability, and providing a reliable foundation for subsequent feature fusion. Subsequently, performing average weighted summation on the reduced first, second, and third feature vectors can fully integrate the unique capturing capabilities of different models for face features, forming an original face feature vector with more robustness and generalization. This fusion method combines the advantages of each model, such as the accuracy of the Arcface model in identity recognition and the fineness of the Insightface model in face structure analysis, so that the original face feature vector can more comprehensively and accurately represent the multiple attributes of the face, thereby providing more abundant and accurate feature replacement basis for the target object in the subsequent face replacement process, and significantly improving the quality and authenticity of the face replacement image.

[0134] In an embodiment, the fusion module 403 comprises:

[0135] The first obtaining sub-module is configured to obtain a target field test set corresponding to the first face image.

[0136] The evaluation sub-module is configured to perform performance evaluation on the first face recognition model, the second face recognition model and the third face recognition model respectively based on the target field test set, and obtain a first accuracy rate of the first face recognition model, a second accuracy rate of the second face recognition model and a third accuracy rate of the third face recognition model.

[0137] The normalization sub-module is configured to normalize the first accuracy rate, the second accuracy rate and the third accuracy rate, and obtain normalized first accuracy rate, second accuracy rate and third accuracy rate.

[0138] The obtaining sub-module is configured to obtain the first weight of the first face recognition model, the second weight of the second face recognition model and the third weight of the third face recognition model based on the normalized first accuracy rate, the second accuracy rate and the third accuracy rate.

[0139] The weighting sub-module is configured to perform weighted summation on the first original feature vector, the second original feature vector, the third original feature vector, the first weight, the second weight and the third weight, and obtain the original face feature vector of the first face image.

[0140] The embodiments of the present application can obtain the target field test set corresponding to the first face image, evaluate the performance of each face recognition model in a specific application scenario (such as cross-race, cross-age, etc.) and ensure that the evaluation results are highly consistent with the actual needs. Based on the performance evaluation and accuracy rate normalization of the target field test set, the quantitative comparison of the recognition ability of different models is realized, and reasonable weights are allocated accordingly, avoiding the feature extraction deviation caused by the performance limitation of a single model. Finally, through the weighted summation of multiple original feature vectors and their corresponding weights, the advantages of each model can be comprehensively utilized to generate a more comprehensive and accurate original face feature vector. This technical effect not only enhances the applicability of face replacement technology in complex scenarios, but also provides a high-quality feature basis for subsequent face-changing image generation, thereby effectively improving the realism and naturalness of the face-changing effect.

[0141] In an embodiment, the first obtaining sub-module further obtains a plurality of face images similar to the first face image scene to obtain an image set; performs identity labeling on each face image of the image set to obtain a labeled image set; divides the labeled image set into a training set and a test set to obtain the target field test set.

[0142] The embodiment of the application can ensure high consistency of training data and actual application environment through a similar scene image set, enable the model to learn more close-to-reality feature representation, and effectively alleviate the performance decline problem caused by environment difference. The identity label provides an explicit supervision signal for the model, which helps the model to learn more discriminative facial features. By dividing the labeled image set into a training set and a test set, effective evaluation and optimization of the model performance are realized, and the overfitting phenomenon is avoided. Finally, the generated target field test set provides a high-quality data basis for subsequent face recognition model training and evaluation, which helps to improve the recognition accuracy and robustness of the model in complex scenes.

[0143] In an embodiment, the second acquisition module 404 comprises:

[0144] The second acquisition sub-module is configured to acquire the second facial image of the target object.

[0145] The feature extraction sub-module is configured to perform feature extraction on the second facial image by using the first facial recognition model, the second facial recognition model, and the third facial recognition model respectively, to obtain a first target feature vector, a second target feature vector, and a third target feature vector.

[0146] The fusion sub-module is configured to fuse the first target feature vector, the second target feature vector, and the third target feature vector to obtain a target facial feature vector of the second facial image.

[0147] After acquiring the second facial image of the target object, the embodiment of the application performs feature extraction by using the first facial recognition model, the second facial recognition model, and the third facial recognition model respectively. Based on the unique algorithm design of each model, different levels and dimensions of feature information in the facial image can be captured, such as facial contour, local texture, key point distribution, etc., so as to obtain the first target feature vector, the second target feature vector, and the third target feature vector. By fusing these feature vectors, the advantages of each model can be comprehensively utilized, and the limitations of a single model can be overcome, so as to obtain a more rich and robust target facial feature vector. The feature vector not only contains the basic facial features of the target object, but also fuses the unique understanding and representation of features by different models, which provides a solid and accurate feature basis for subsequent face changing image generation, and helps to improve the naturalness and realism of the face changing effect.

[0148] In an embodiment, the fusion submodule is further configured to reduce the first target feature vector, the second target feature vector, and the third target feature vector respectively based on a preset feature alignment rule to obtain a first standardized feature vector, a second standardized feature vector, and a third standardized feature vector; and perform average weighted summation on the first standardized feature vector, the second standardized feature vector, and the third standardized feature vector to obtain the target face feature vector of the second face image.

[0149] The embodiments of the present application can extract features of the second face image through the first face recognition model, the second face recognition model, and the third face recognition model to obtain three target feature vectors of different dimensions, which describe the facial features of the target object from different angles and enrich the diversity of feature information. Then, the feature vectors are reduced based on the preset feature alignment rule to obtain standardized feature vectors, which effectively eliminate the dimensional difference and scale inconsistency between the features extracted by different models and ensure the comparability and compatibility of the features. Finally, the target face feature vector obtained by performing average weighted summation on the standardized feature vectors not only combines the advantages of multiple models but also maintains the consistency and stability of the features.

[0150] In an embodiment, the fusion submodule is further configured to perform weighted summation on the first target feature vector, the second target feature vector, the third target feature vector, the first weight, the second weight, and the third weight to obtain the target face feature vector of the second face image.

[0151] The embodiments of the present application can extract different feature dimensions of the target face image through the first face recognition model, the second face recognition model, and the third face recognition model to obtain the first target feature vector, the second target feature vector, and the third target feature vector, which describe the facial characteristics of the target object from multiple angles. By determining the weights of the models, different importance can be given to the features extracted by different models according to the advantages and emphases of the models in feature extraction. Then, weighted summation is performed on the feature vectors and their corresponding weights, so that the target face feature vector obtained finally not only combines the advantages of the models but also highlights the key features. This processing mode enhances the representation ability of the feature vector for the facial information of the target object, provides a more accurate and comprehensive feature basis for subsequent face processing tasks (such as face changing and expression transfer), and helps to improve the quality and stability of the overall processing effect.

[0152] To solve the above technical problems, the embodiments of the present application also provide a computer device. For details, please refer to Figure 4 , Figure 4 The basic structure block diagram of the computer device of the present embodiment is shown in FIG. 1.

[0153] The computer device 4 includes a memory 61, a processor 62, and a network interface 63, which are communicatively connected by a system bus. It should be noted that the computer device 6 is shown as having a memory 61, a processor 62, and a network interface 63, but it should be understood that not all of the illustrated components need be implemented, and that more or fewer components can alternatively be implemented. As will be understood by those skilled in the art, the computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and the hardware thereof includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, and the like.

[0154] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server, or the like. The computer device can interact with a user through a keyboard, a mouse, a remote controller, a touchpad, a voice control device, or the like.

[0155] The memory 61 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, or the like), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, or the like. In some embodiments, the memory 61 can be an internal storage unit of the computer device 6, such as a hard disk or a memory of the computer device 6. In other embodiments, the memory 61 can also be an external storage device of the computer device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, or the like. Of course, the memory 61 can include both an internal storage unit and an external storage device of the computer device 6. In the present embodiment, the memory 61 is generally used to store an operating system and various application software installed in the computer device 6, such as computer readable instructions of the face replacement method, and the like. In addition, the memory 61 can also be used to temporarily store various data that has been output or will be output.

[0156] The processor 62 may, in some embodiments, be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 62 is generally used to control the overall operation of the computer device 6. In the present embodiment, the processor 62 is used to run computer readable instructions or process data stored in the memory 61, for example, computer readable instructions of a face replacement method.

[0157] The network interface 63 can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0158] The embodiments of the present application can effectively overcome the limitations of a single model in cross-domain feature extraction by using three heterogeneous face recognition models to extract multi-level feature vectors of the face image of the original object in parallel, using the differential capture capabilities of different models for race, age, and other feature dimensions, and constructing a more inclusive original face feature vector through a feature fusion strategy. In the feature replacement stage, the spatial geometric constraint relationship established by the face key point detection model can accurately align the face structures of the original object and the target object, ensuring that the dynamic attribute features such as expressions and postures remain spatially consistent during the replacement process, avoiding misplacement of facial features or texture distortion due to differences in facial structures. The final face-swapped image not only retains the basic identity features of the target object, but also completely transfers the personalized attributes of the original object, improving the accuracy and naturalness of face replacement. In the financial marketing scenarios (such as customized advertisement production) that require strict visual authenticity, such as film and television special effect production and virtual anchor generation, the face-swapped image can significantly reduce the cost of manual retouching, improve the efficiency of content production, and enhance the user's sense of immersion.

[0159] The present application also provides another implementation, namely providing a computer readable storage medium, the computer readable storage medium stores computer readable instructions, the computer readable instructions can be executed by at least one processor to make the at least one processor execute the steps of the face replacement method as described above.

[0160] The embodiments of the present application can overcome the limitations of a single model in cross-domain feature extraction by adopting three heterogeneous face recognition models to extract multi-level feature vectors of the original object's face image in parallel, utilizing the differential capturing capabilities of different models for race, age and other feature dimensions, and constructing a more inclusive original face feature vector through a feature fusion strategy. In the feature replacement stage, the spatial geometric constraint relationship established by the face key point detection model can accurately align the face structures of the original object and the target object, ensuring that dynamic attribute features such as expressions and postures remain spatially consistent during the replacement process, and avoiding misplacement of facial features or texture distortion due to differences in facial structures. The final face replacement image not only retains the basic identity features of the target object, but also completely transfers the personalized attributes of the original object, improving the accuracy and naturalness of face replacement. In the financial marketing scenarios (such as customized advertisement production) that require high visual authenticity, such as film and television special effect production and virtual anchor generation, the cost of manual retouching can be significantly reduced, and the content output efficiency and user immersion can be improved.

[0161] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and the necessary general hardware platform, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the methods of various embodiments of the present application.

[0162] Obviously, the above-described embodiments are only some of the embodiments of the present application, not all the embodiments, and the preferred embodiments of the present application are given in the drawings, but do not limit the patent scope of the present application. The present application can be realized in many different forms, and on the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can modify the technical solutions described in the foregoing embodiments or make equivalent replacements to some technical features. Any equivalent structure made by using the contents of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the scope of the patent protection of the present application.

[0163] The non-company software tools or components appearing in the embodiments of the present application are only examples for introduction, not actual use.

Claims

1. A face replacement method, characterized in that: The steps include: Obtaining a first face image of an original subject; Using a preset first face recognition model, a second face recognition model, and a third face recognition model, respectively, to perform feature extraction on the first face image to obtain a first original feature vector, a second original feature vector, and a third original feature vector; fusing the first original feature vector, the second original feature vector, and the third original feature vector to obtain an original facial feature vector of the first facial image; Obtain a target facial feature vector of a second facial image of a target object; Performing facial key point detection on the first face image and the second face image using a preset facial key point detection model to establish a spatial correspondence between the facial key points of the original object and the facial key points of the target object; Based on the spatial correspondence, the attribute features in the target facial feature vector are replaced with the attribute features corresponding to the original facial feature vector to obtain a synthetic feature vector; Based on the synthesized feature vector, a preset generation model is used to generate a face-swapped image of the target object.

2. The method according to claim 1, characterized in that The step of fusing the first original feature vector, the second original feature vector, and the third original feature vector to obtain the original facial feature vector of the first facial image specifically includes: Based on a preset feature alignment rule, the first original feature vector, the second original feature vector, and the third original feature vector are respectively reduced to obtain a first feature vector, a second feature vector, and a third feature vector; Performing weighted average summation on the first eigenvector, the second eigenvector, and the third eigenvector to obtain an original facial feature vector of the first facial image.

3. The method according to claim 1, characterized in that The step of fusing the first original feature vector, the second original feature vector, and the third original feature vector to obtain the original facial feature vector of the first facial image specifically includes: Obtaining a target domain test set corresponding to the first face image; Based on the target domain test set, respectively perform performance evaluation on the first face recognition model, the second face recognition model, and the third face recognition model to obtain a first accuracy rate of the first face recognition model, a second accuracy rate of the second face recognition model, and a third accuracy rate of the third face recognition model; Normalize the first accuracy rate, the second accuracy rate, and the third accuracy rate to obtain the normalized first accuracy rate, the second accuracy rate, and the third accuracy rate Based on the normalized first accuracy rate, second accuracy rate, and third accuracy rate, obtain a first weight of the first face recognition model, a second weight of the second face recognition model, and a third weight of the third face recognition model; Performing a weighted summation on the first original feature vector, the second original feature vector, the third original feature vector, the first weight, the second weight, and the third weight to obtain an original facial feature vector of the first facial image.

4. The method according to claim 3, characterized in that The step of obtaining a target domain test set corresponding to the first face image specifically includes: Acquire multiple facial images with scenes similar to the first facial image to obtain an image set; Performing identity labeling on each face image in the image set to obtain a labeled image set; The labeled image set is divided into a training set and a test set to obtain a target domain test set.

5. The method according to claim 3, characterized in that The step of obtaining a target facial feature vector of the second facial image of the target object specifically includes: Acquire a second facial image of the target object; respectively using the first face recognition model, the second face recognition model, and the third face recognition model to perform feature extraction on the second face image to obtain a first target feature vector, a second target feature vector, and a third target feature vector; The first target feature vector, the second target feature vector, and the third target feature vector are fused to obtain a target facial feature vector of the second facial image.

6. The method according to claim 5, characterized in that The step of fusing the first target feature vector, the second target feature vector, and the third target feature vector to obtain the target facial feature vector of the second facial image specifically includes: Based on a preset feature alignment rule, the first target feature vector, the second target feature vector, and the third target feature vector are respectively reduced to obtain a first normalized feature vector, a second normalized feature vector, and a third normalized feature vector; Performing an average weighted summation on the first normalized feature vector, the second normalized feature vector, and the third normalized feature vector to obtain a target facial feature vector of the second facial image.

7. The method according to claim 5, characterized in that The step of fusing the first target feature vector, the second target feature vector, and the third target feature vector to obtain the target facial feature vector of the second facial image specifically includes: A weighted sum is performed on the first target feature vector, the second target feature vector, the third target feature vector, the first weight, the second weight, and the third weight to obtain a target facial feature vector of the second facial image.

8. A face replacement device, characterized in that: include: A first acquisition module, configured to acquire a first face image of an original object; an extraction module, configured to perform feature extraction on the first face image using a preset first face recognition model, a second face recognition model, and a third face recognition model, respectively, to obtain a first original feature vector, a second original feature vector, and a third original feature vector; a fusion module, configured to fuse the first original feature vector, the second original feature vector, and the third original feature vector to obtain an original facial feature vector of the first facial image; A second acquisition module is used to obtain a target facial feature vector of a second facial image of a target object; An establishment module is used to perform facial key point detection on the first face image and the second face image using a preset facial key point detection model, and establish a spatial correspondence between the facial key points of the original object and the facial key points of the target object; a replacement module, configured to replace the attribute features in the target facial feature vector with the attribute features corresponding to the original facial feature vector based on the spatial correspondence, to obtain a synthesized feature vector; A generation module is used to generate a face-swapped image of the target object based on the synthesized feature vector using a preset generation model.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the face replacement method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the face replacement method according to any one of claims 1 to 7.