Video face replacement method, device and electronic device
By using multiple face replacement models to replace the original face area in the video frame and selecting the optimal permutation result, the problem of poor replacement effect caused by changes in face angles in different video frames is solved, and the quality of face replacement is significantly improved.
Patent Information
- Application Number
- CN202510387551.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-28
AI Technical Summary
When replacing the original face in a video, the replacement effect is poor due to the changes in the face angle and occlusion situation in different video frames.
At least two different face replacement models are used to process the original face area to be replaced, and the optimal permutation result is selected through ratings to improve the video quality of face replacement.
Through parallel processing and scoring selection of different models, various facial angles of the face are ensured to be optimally replaced, improving the quality of face replacement in videos.
Smart Images

Figure CN119904784B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a video face replacement method, apparatus, and electronic device. Background Art
[0002] Currently, in various actual scenarios such as game entertainment and film and television drama production scenarios, there is a demand for applications that replace the face of a specified object in the original video with a target face. However, in different video frames of the original video, the face of the specified object usually presents different angles and may even be partially occluded. In this case, the effect of replacing the face of the specified object is poor. Summary of the Invention
[0003] In view of this, this application provides a video face replacement method, apparatus, and electronic device to improve the quality of the video obtained by face replacement.
[0004] The technical solutions provided by this application are as follows:
[0005] According to an embodiment of the first aspect of this application, a video face replacement method is provided. The method includes:
[0006] For the original face area to be replaced in the video frame, use at least two different existing face replacement models to perform face replacement processing on the original face area to obtain a face replacement result for the original face area; the face replacement result represents the face area where the original face of the target object in the original face area is replaced with the target face;
[0007] Score the face replacement results output by each face replacement model, and select the optimal face replacement result based on the scores of each face replacement result;
[0008] Replace the original face area to be replaced in the video frame with the optimal face replacement result.
[0009] According to an embodiment of the second aspect of this application, a video face replacement apparatus is provided. The apparatus includes:
[0010] A first replacement unit for, for the original face area to be replaced in the video frame, using at least two different existing face replacement models to perform face replacement processing on the original face area to obtain a face replacement result for the original face area; the face replacement result represents the face area where the original face of the target object in the original face area is replaced with the target face;
[0011] A scoring and selection unit for scoring the face replacement results output by each face replacement model and selecting the optimal face replacement result based on the scores of each face replacement result;
[0012] A second replacement unit, configured to replace the original face area to be replaced in the video frame with the optimal face replacement result.
[0013] According to an embodiment of the third aspect of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in the first aspect is implemented.
[0014] As can be seen from the above technical solutions, in the present application, for the original face area to be replaced in the video frame, at least two different existing face replacement models are used to perform face replacement processing on the original face area to obtain a face replacement result of the original face area. And based on the scoring of each face replacement result, the optimal face replacement result is selected, and the original face area to be replaced in the video frame is replaced with the optimal face replacement result; by performing face replacement on the same original face area to be replaced through different face replacement models, and then selecting the optimal face replacement result, it is ensured that various facial angles of the original face of the target object in the original face area to be replaced can be optimally replaced, improving the quality of the replaced video. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.
[0016] Figure 1 It is a flowchart of the video face replacement method provided by the embodiment of the present application;
[0017] Figure 2 It is a schematic diagram of the overall architecture of the video face replacement method provided by the embodiment of the present application;
[0018] Figure 3 It is a schematic diagram of the overall process of the video face replacement method provided by the embodiment of the present application;
[0019] Figure 4 It is a schematic diagram of the structure of an electronic device provided by the embodiment of the present application;
[0020] Figure 5 It is a structural diagram of a video face replacement device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0022] Please refer toFigure 1 , Figure 1 is a flowchart of the video face replacement method provided by the embodiment of the present application.
[0023] In this embodiment, the video face replacement method can be applied to an electronic device with video processing capabilities, such as mobile devices like smartphones and tablets, or imaging devices like smart cameras. The present application does not limit this.
[0024] As Figure 1 shown, the method may include the following steps:
[0025] Step 101, for the original face area to be replaced in the video frame, use at least two different existing face replacement models to perform face replacement processing on the original face area to obtain the face replacement result of the original face area.
[0026] Among them, the original face area to be replaced is the area of the face image (denoted as the original face) that includes the specified object (denoted as the target object) that needs to be face-replaced; the face replacement result is the face area where the original face of the target object in the original face area is replaced with the target face.
[0027] For example, the original face area to be replaced is the rectangular frame area where the face of the target object is identified by the face detection model. The rectangular frame area includes the original face of the target object and a part of the background area. Then, the face replacement result obtained by performing face replacement processing on the face area to be replaced can be to replace the original face included in the above rectangular frame area with the target face, without changing the background area, to obtain a new rectangular frame area. In this embodiment, the method for determining the original face area to be replaced will be described in detail below and will not be elaborated here.
[0028] In this embodiment, for the original face area to be replaced in the video frame, at least two different face replacement models can be used to perform face replacement processing on the original face area to be replaced, that is, input the face area to be replaced into each face replacement model, and for each face replacement model, obtain a face replacement result.
[0029] Exemplarily, the optimal face angles required by at least two different face replacement models can be different. The optimal face angle can be the face angle at which the model is good at performing face replacement processing. For example, the optimal face angles required by different face replacement models can be different angles such as a frontal face, a profile face, and an occluded face.
[0030] As an example, for each face replacement model, the face replacement effect is optimal at the optimal face angle required by the model. Each face replacement model can be selected from different models that are good at the same angle (such as the front face) and have the best face replacement effect.
[0031] After performing replacement processing on the original face region through at least two different face replacement models, multiple face replacement results of the original face region are obtained.
[0032] In this embodiment, the original face region to be replaced can be determined through the following steps:
[0033] Feature extraction is performed on the reference face image of the target object and each face image region included in the video frame to obtain the reference face image features of the target object and the image features of each face image region; the face image region is determined according to the image size information of the reference face image, and the face image region includes at least the original face region.
[0034] According to the similarity calculation results between the reference face image features of the target object and the image features of each face image region, the face image region corresponding to the image features that meet the specified conditions is determined; wherein, the specified conditions include: the similarity calculation result is the maximum value among the similarity results of each image feature, and this similarity result is greater than the first specified value.
[0035] The original face region included in the face image region corresponding to the image features that meet the specified conditions is determined as the original face region to be replaced.
[0036] First, the method for determining each face image region included in the video frame is introduced below.
[0037] In this embodiment, face image detection can be performed on the video frame to obtain each original face region included in the video frame.
[0038] According to the image size information of the reference face image of the target object that has been obtained and each original face region included in the video frame, the face image region is determined; the face image region refers to the region that includes at least the original face region.
[0039] The original face region to be replaced is determined according to the reference face image of the target object and each face image region.
[0040] Specifically, the video frame can be detected through a trained face image detection model to obtain the rectangular frames for identifying the regions where each original face is located in the video frame, and the position coordinates of 5 key points (such as the right eye pupil, left eye pupil, tip of the nose, right corner of the mouth, left corner of the mouth) in each original face, and the rectangular frame of the region where the original face is located is determined as the original face region.
[0041] After determining each original face region in the video frame, based on the image size information of the reference face image of the target object that has been obtained and the original face regions included in the video frame, a face image region corresponding to each original face region can be determined.
[0042] In this embodiment, a reference face image of the target object can be obtained in advance. The reference face image can be a recent photo of the target object or a face image of the target object directly intercepted from any video frame. The present application does not limit this.
[0043] To facilitate subsequent comparison of face features, based on the image size information of the reference face image of the target object that has been obtained and the original face regions included in the video frame, the specific method for determining the face image region corresponding to the original face region can be: intercept the face image region from the video frame according to the aspect ratio of the obtained reference face image. Each face image region includes at least one original face region, that is, the rectangular frame including the region where the above-mentioned original face is located.
[0044] So far, the introduction of the method for determining the face image region ends.
[0045] After determining the face image region, the same feature extraction method is used to extract features from the obtained reference face image of the target object and each face image region included in the video frame, to obtain the reference face image features of the target object and the image features of each face image region.
[0046] Furthermore, the similarity between the reference face image features of the target object and the image features of each face image region can be calculated to obtain the similarity calculation result.
[0047] If there is a face image region in the video frame whose similarity result is greater than the first specified value, it indicates that there is an original face of the target object in the video frame, that is, there is an original face region to be replaced.
[0048] At this time, the original face region included in the face image region with the largest similarity calculation result can be determined as the original face region to be replaced.
[0049] In this embodiment, the method for calculating similarity can be to calculate the cosine similarity between feature values. For example, the cosine similarity is calculated through the following formula:
[0050]
[0051] Among them, similarity represents cosine similarity, A and B respectively represent the reference face image feature vector of the target object and the image feature vector of the face image region, and A i represents the i-th element in the feature vector A, and B i represents the i-th element in the feature vector B.
[0052] So far, the description of the step of determining the original face region to be replaced is completed.
[0053] Since the solution proposed in the embodiment of the present application is to replace the original face image of the target object included in the video to be face replaced (denoted as the original video), therefore, before performing face replacement processing on the original face region to be replaced in the video frame using at least two different existing face replacement models, it is necessary to first determine whether there is an original face image of the target object in the current video frame.
[0054] As an embodiment, first, the original video can be parsed to obtain multiple video frames corresponding to the original video. For each video frame, perform the above steps of determining the original face region to be replaced.
[0055] After detecting the face images in the current video frame to obtain each original face region included in the video frame, the original face regions can be numbered in a specified order, such as from left to right.
[0056] In this embodiment, each video frame also records an index number value used to indicate whether the video frame includes a face region.
[0057] Specifically, the index number value can be recorded by the following method.
[0058] In the case where there is no original face region to be replaced in the current video frame, the number value corresponding to the original face region to be replaced in the current video frame can be recorded, and this number value is the index number value recorded in the video frame.
[0059] In the case where there is no original face region to be replaced in the current video frame, or there is no original face region in the video frame, a second specified value different from the number value corresponding to any original face region can be recorded as the index number value corresponding to the video frame.
[0060] Before performing face replacement processing on the original face region to be replaced in the video frame using at least two different existing face replacement models, it can be first determined whether the video frame includes an original face region to be replaced according to the index number value recorded in the video frame.
[0061] Specifically, for each video frame, detect the recorded index number value in the video frame. If the recorded index number value is the number value corresponding to any original face region in the video frame, it indicates that there is an original face region to be replaced in the video frame, that is, the original face region corresponding to the recorded index number value. At this time, the steps of performing face replacement processing on the original face region to be replaced in the video frame by using at least two different face replacement models that are currently available can be continued.
[0062] If the recorded index number value is the second specified value, it indicates that there is no original face region to be replaced in the video frame, or there is no original face region in the video frame. At this time, the video frame does not need to be subjected to face replacement processing, and the steps of performing face replacement processing on the original face region to be replaced in the video frame by using at least two different face replacement models that are currently available can be rejected.
[0063] In this embodiment, the second specified value is different from the number value of any original face region. For example, the number values of the original face regions can be 1, 2, 3,...; and the second specified value can be set to a value different from the number value of any original face region, such as -1, etc. The present application does not limit this.
[0064] So far, the description of step 101 is ended, and step 102 is executed below.
[0065] Step 102: Score the face replacement results output by each face replacement model, and select the optimal face replacement result based on the scores of each face replacement result.
[0066] In this embodiment, after obtaining the face replacement results output by each face replacement model through step 101, the face replacement results output by each face replacement model can be scored to select the optimal face replacement result based on the scores of each face replacement result.
[0067] Specifically, for each face replacement result, perform a convolution operation on the gray value of each pixel point in the face replacement result based on the Laplace operator to obtain the convolution operation result corresponding to the face replacement result;
[0068] Take the variance of the convolution operation results corresponding to each face replacement result as the score of each face replacement result; the scores of each face replacement result are used to indicate the image blur degree of each face replacement result, and the image blur degree is inversely correlated with the variance of the convolution operation result;
[0069] Determine the face replacement result with the highest score among each face replacement result as the optimal face replacement result.
[0070] In this embodiment, each face replacement result can be first converted into a grayscale image. Further, based on the Laplace operator, a convolution operation is performed on the grayscale value of each pixel point in the face replacement result to obtain a convolution operation result corresponding to the face replacement result.
[0071] After determining the convolution operation results corresponding to each face replacement result, the variance corresponding to each convolution operation result can be further determined, and this variance is used as the score of each face replacement result.
[0072] The variance corresponding to each convolution operation result can be used to indicate the image blurring degree of each face replacement result. The image blurring degree is inversely correlated with the variance of the convolution operation result, that is, the larger the variance result, the lower the image blurring degree, the clearer the image, the better the display effect, and the higher the score of the face replacement result.
[0073] In this embodiment, the face replacement result with the highest score can be determined as the optimal face replacement result.
[0074] So far, the description of step 102 ends. Next, step 103 is executed.
[0075] Step 103: Replace the original face area to be replaced in the video frame with the optimal face replacement result.
[0076] In this embodiment, after determining the optimal face replacement result corresponding to the video frame, the original face area to be replaced in the video frame can be replaced with the optimal face replacement result, that is, the original face area to be replaced in the video frame is replaced with the optimal face replacement result to obtain a video frame after face replacement. Further, each video frame after face replacement and the video frame in the original video that does not have the original face area to be replaced can be video-encoded to obtain a face replacement video corresponding to the original video.
[0077] As an embodiment, the specific method for replacing the original face area to be replaced in the video frame with the optimal face replacement result may include:
[0078] Based on the trained facial restoration model, image optimization is performed on the optimal face replacement result to obtain an optimally image-processed optimal face replacement result; the optimally image-processed optimal face replacement result is used to replace the original face area to be replaced in the video frame;
[0079] Or, the original face area to be replaced in the video frame is replaced with the optimal face replacement result; based on the trained facial restoration model, image optimization is performed on all face areas included in the video frame.
[0080] In this embodiment, after determining the optimal face replacement result, the optimal face replacement result can be further optimized. For example, the optimal face replacement result can be optimized through a trained face restoration model, and the original face area to be replaced in the video frame is replaced with the optimized optimal face replacement result to obtain the video frame after face replacement.
[0081] In this embodiment, after determining the optimal face replacement result, the original face area to be replaced in the video frame can also be replaced with the optimal face replacement result first to obtain a reference video frame; further, based on the trained face restoration model, image optimization is performed on all face areas included in the reference video frame (including the original face area that does not need to be replaced and the optimal face replacement result) to obtain the video frame after face replacement.
[0082] The above two methods can be selected based on actual needs or server resource occupancy, and the present application does not limit this.
[0083] So far, the description of step 103 ends.
[0084] As an embodiment, before performing face replacement processing on the original face area to be replaced in the video frame by using at least two different existing face replacement models, the method may further include:
[0085] Performing classification prediction on the image features of the video frame and the text features of the text information included in the video frame based on a trained neural network model to obtain a classification prediction result for the video frame;
[0086] If the classification prediction results of all video frames meet the normal video frame conditions, then perform the step of performing face replacement processing on the original face area to be replaced in the video frame by using at least two different existing face replacement models.
[0087] In this embodiment, before performing face replacement on the original video, it can be first detected whether the original video is a video that allows face replacement operations, that is, it is determined whether the video frames of the original video all meet the normal video frame conditions.
[0088] Specifically, for each video frame included in the original video, text recognition can be performed on the video frame to obtain the text information included in the video frame; for the audio file that may be included in the original video, the audio file can also be converted into corresponding text information.
[0089] Further, feature extraction can be performed on the video frame to obtain the image features of the video frame; feature extraction is performed on the text information (including the text information included in the video frame and the text information converted from the audio file) to obtain the text features of the text information.
[0090] The image features and text features are classified and predicted through a trained neural network model to obtain a classification and prediction result for the video frame.
[0091] In this embodiment, inception_v3 can be used to extract the image features of the picture frame. After OCR text recognition and BERT text feature extraction, the image features and text features (ensuring that images, subtitles, etc. in the material are compliant at the same time) are classified and predicted for the picture through the neural network MLP (Multi-Layer Perceptron) to obtain the final classification and prediction result.
[0092] As an embodiment, the classification and prediction result may include a normal type and an abnormal type, and the abnormal type is used to indicate that face replacement for the video frame is not allowed.
[0093] In this embodiment, that the classification and prediction results of the video frames all meet the normal video frame conditions may mean that the classification and prediction results of all video frames in the original video are of the normal type, which indicates that the original video is a video allowing face replacement. At this time, the step of performing face replacement processing on the original face area to be replaced in the video frame by using at least two different existing face replacement models can be continued.
[0094] If the classification and prediction result of any video frame in the original video is of the abnormal type, the subsequent face replacement steps are refused to be executed, and a prompt message can be generated to indicate that the original video is a video not allowing face replacement.
[0095] So far, the description of Figure 1 the video face replacement method is completed.
[0096] In this application, for the original face area to be replaced in the video frame, at least two different existing face replacement models are used to perform face replacement processing on the original face area to obtain a face replacement result for the original face area, and based on the scoring of each face replacement result, the optimal face replacement result is selected, and the optimal face replacement result is used to replace the original face area to be replaced in the video frame; by performing face replacement on the same original face area to be replaced through different face replacement models, and then selecting the optimal face replacement result, it is ensured that various facial angles of the original face of the target object in the original face area to be replaced can be optimally replaced, improving the quality of the replaced video.
[0097] Further, before performing face replacement on the original video, the present application also performs a pre - security compliance check, extracts the image features and text features of each video frame in the original video, classifies the extracted image features and text features through a trained neural network model to obtain a predicted classification result, and only allows face replacement of the original video when the classification prediction results of the video frames all meet the conditions of normal video frames, effectively ensuring the compliance of the input material and the reasonable and compliant use of the technology.
[0098] The following combines Figure 2 and Figure 3 to describe the video face replacement method proposed by the present application.
[0099] Please refer to Figure 2 , Figure 2 which is the overall architecture schematic diagram of the video face replacement method provided by the embodiments of the present application.
[0100] As Figure 2 shown, the overall architecture included in this method consists of the following parts:
[0101] 1. Input part, used to input the original video, the reference face image of the target object (the face image of the target object that needs to perform face replacement and has been obtained), and the target face image (the replaced face image) into the electronic device that executes this method.
[0102] 2. Video parsing part, used to perform video parsing on the original video to obtain an audio file, the frame rate of the original video, and multiple video frames.
[0103] 3. Security detection part, used to extract features from the images and text information included in the video frames, and perform classification prediction on the image features and text features obtained by feature extraction to determine the category of each video frame.
[0104] 4. Index calculation part, for the original video whose classification prediction results of the video frames all meet the conditions of normal video frames, for each video frame, extract the image features of the original face area included in the video frame and the image features of the reference face image of the target object, calculate the face similarity between the image features of the reference face image and the image features of each original face area, and determine the index number value of the original face area that meets the specified conditions. The specified conditions and the index number value have been described in detail above and will not be elaborated here.
[0105] 5. The model processing part is used to replace the original face area to be replaced through at least two face replacement models (such as replacement model A, replacement model B, and replacement model C), and determine the replacement result with the highest score through Laplacian convolution calculation, and perform face restoration on the replacement result with the highest score.
[0106] 6. The video synthesis part is used to perform video encoding on the video frames after face replacement, the video frames that do not need face replacement, and the audio file of the original video according to the frame rate of the original video.
[0107] 7. The output part is used to obtain the result of video synthesis, that is, the face replacement video corresponding to the original video.
[0108] So far, the description of Figure 2 ends.
[0109] Please refer to Figure 3 , Figure 3 , which is the overall process schematic diagram of the video face replacement method provided by the embodiment of this application.
[0110] As Figure 3 shown, the video face replacement method proposed in this embodiment can specifically include the following steps:
[0111] 1. Parse the original video.
[0112] Parse the input original video, and separate information such as each video frame, audio file, and frame rate of the original video.
[0113] 2. Detect whether the original video is compliant material.
[0114] Use the Inception_v3 model to extract the image features of each video frame, perform text recognition on the text information in the video frame through optical character recognition (OCR: Optical Character Recognition), and extract the text features of the text information through the trained language model BERT. Classify and predict the image features and text features through a multi-layer perceptron (MLP: Multi-Layer Perceptron) (to ensure that images, subtitles, etc. in the material are compliant at the same time), and classify the video frame into a normal type and an abnormal type. If the predicted types of all video frames in the original video are normal types, continue to execute the following step 3. If the predicted type of any video frame is an abnormal type, directly end the process and give that the original video is a video that does not allow face replacement.
[0115] 3. Picture alignment.
[0116] Detect the original face regions in each video frame, and number all the faces in the order from left to right, such as 1, 2, 3...
[0117] Crop the face image regions in each video frame according to the aspect ratio of the reference face image to avoid changes in facial features caused by stretching the picture and deforming it. The face image region includes the original face region.
[0118] 4. Similarity calculation.
[0119] Extract the facial features of all face image regions and the reference face image of the target object, and calculate the cosine similarity between the facial features of the face image regions in the video frame and the reference face image.
[0120] 5. Index output.
[0121] Make a judgment based on the similarity calculated in step 4. When the similarity is greater than the first specified value (such as the default 0.6, which can be adjusted according to the actual situation and is usually set in the range of 0.5 - 0.7), then determine that the number value of the original face region included in the face image region with the highest similarity is the index number value corresponding to this video frame. When the similarity is less than the first specified value, it is considered that there is no face to be replaced in this video frame, and at this time, the second specified value, such as -1, is determined as the index number value corresponding to this video frame.
[0122] Through this step, the position of the face to be replaced can be automatically locked, ensuring that no matter how many faces appear in the video and no matter where the face to be replaced appears, it can be accurately replaced.
[0123] 6. Index judgment.
[0124] Before performing face replacement, detect the index number value corresponding to each video frame. If the index number value is -1, it indicates that there is no original face region to be replaced in this video frame, that is, this video frame does not need to be replaced.
[0125] If the index number value is greater than -1, such as 1, 2, 3..., it indicates that there is an original face region to be replaced in this video frame. The original face region to be replaced is the original face region corresponding to this index number value, and at this time, the subsequent replacement steps can be continued.
[0126] 7. Replacement processing.
[0127] Use multiple face replacement models (such as replacement model A, replacement model B, and replacement model C) to perform parallel face replacement processing on the original face region to be replaced, and at the same time obtain the face replacement results of each different model (such as multiple models good at different angles such as frontal face, side face, occlusion, etc.) in the original face region corresponding to the index number value.
[0128] 8. Replacement result scoring.
[0129] Use the opencv Laplacian operator to perform convolution operations on the face replacement results of different models to obtain the convolution operation results. Take the variance of the convolution operation results as the score of the face replacement result, and select the face replacement result corresponding to the highest score as the optimal face replacement result.
[0130] 9. Model repair processing.
[0131] By step 8, the face replacement process is basically completed. However, the facial details of the optimal face replacement result may not be natural enough, and the splicing position may not be smooth enough. You can optimize the face details through a face repair model according to actual needs.
[0132] When repairing the face, you can only optimize the optimal face replacement result to avoid repairing all the original face areas in the overall video frame, so as to reduce the consumption of server resources. You can also repair all the original face areas according to actual needs.
[0133] 10. Merge the video. According to the video frame rate, perform video encoding on the video frames after face replacement of each person, the video frames that do not require face replacement, and the audio file of the original video to obtain the face replacement video corresponding to the original video.
[0134] In this embodiment, you can select some or all of the above steps according to actual needs and resource performance requirements to meet the video face replacement needs.
[0135] So far, the description of Figure 3 ends.
[0136] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device proposed in an embodiment of this application. At the hardware level, this electronic device includes a processor, an internal bus, a network interface, and a computer-readable storage medium. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program instructions from the computer-readable storage medium and runs, forming a terminal interaction device at the logical level. Of course, in addition to the software implementation method, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logic device.
[0137] Please refer to Figure 5 , Figure 5 which is a structural diagram of a video face replacement device proposed in an embodiment of this application. As shown in Figure 5As shown, the device may include a first data replacement unit 501, a scoring and selection unit 502, and a second replacement unit. Specifically, the device includes:
[0138] A first replacement unit 501, configured to perform face replacement processing on the original face area to be replaced in the video frame by using at least two different existing face replacement models, so as to obtain a face replacement result of the original face area; the face replacement result represents that the original face of the target object in the original face area is replaced by the target face area;
[0139] A scoring and selection unit 502, configured to score the face replacement results output by each face replacement model, and select the optimal face replacement result based on the scores of each face replacement result;
[0140] A second replacement unit 503, configured to replace the original face area to be replaced in the video frame with the optimal face replacement result.
[0141] Optionally, the first replacement unit 501 is further configured to:
[0142] Extract features from the reference face image of the target object and each face image area included in the video frame, to obtain the reference face image features of the target object and the image features of each face image area; the face image area is determined according to the image size information of the reference face image, and the face image area includes at least the original face area;
[0143] Determine the face image area corresponding to the image feature that meets the specified condition according to the similarity calculation result between the reference face image features of the target object and the image features of each face image area; wherein, the specified condition includes: the similarity calculation result is the maximum value among the similarity results of each image feature, and the similarity result is greater than the first specified value;
[0144] Determine the original face area included in the face image area corresponding to the image feature that meets the specified condition as the original face area to be replaced;
[0145] And / or, the optimal face angles required by at least two different face replacement models are different;
[0146] And / or, the original face areas in each video frame are numbered in a specified order, and an index number value is also recorded in each video frame; before performing face replacement processing on the original face area to be replaced in the video frame by using at least two different existing face replacement models, the first replacement unit 501 is further configured to:
[0147] For each video frame, detect the recorded index number value in the video frame. If the recorded index number value is the number value corresponding to any original face region in the video frame, continue to execute the step of performing face replacement processing on the original face region to be replaced in the video frame by using at least two different existing face replacement models; the original face region to be replaced refers to the original face region corresponding to the recorded index number value.
[0148] If the recorded index number value is the second specified value, reject the step of performing face replacement processing on the original face region to be replaced in the video frame by using at least two different existing face replacement models; the second specified value is different from the number value corresponding to any original face region, and the second specified value is used to indicate that there is no original face region to be replaced in the video frame, or there is no original face region in the video frame.
[0149] And / or, the scoring and selection unit 502 is specifically configured to:
[0150] For each face replacement result, perform a convolution operation on the gray value of each pixel point in the face replacement result based on the Laplacian operator to obtain the convolution operation result corresponding to the face replacement result.
[0151] Use the variance of the convolution operation results corresponding to each face replacement result as the score of each face replacement result; the scores of each face replacement result are used to indicate the image blurring degree of each face replacement result, and the image blurring degree is inversely correlated with the variance of the convolution operation result.
[0152] Determine the face replacement result with the highest score among each face replacement result as the optimal face replacement result.
[0153] And / or, before performing face replacement processing on the original face region to be replaced in the video frame by using at least two different existing face replacement models, the first replacement unit 501 is further configured to:
[0154] Perform classification prediction on the image features of the video frame and the text features of the text information included in the video frame based on the trained neural network model to obtain the classification prediction result for the video frame.
[0155] If the classification prediction results of all video frames meet the normal video frame conditions, execute the step of performing face replacement processing on the original face region to be replaced in the video frame by using at least two different existing face replacement models.
[0156] And / or, the second replacement unit 503 is specifically configured to:
[0157] Optimize the image of the optimal face replacement result based on the trained face restoration model to obtain the optimized optimal face replacement result; replace the original face area to be replaced in the video frame with the optimized optimal face replacement result;
[0158] Alternatively, replace the original face area to be replaced in the video frame with the optimal face replacement result; optimize the images of all face areas included in the video frame based on the trained face restoration model.
[0159] Thus, the description of the Figure 5 video face replacement device ends.
[0160] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium. A number of computer instructions are stored on the computer-readable storage medium. When the computer instructions are executed, the methods disclosed in the above examples of the present application can be implemented.
[0161] Exemplarily, the above computer-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the computer-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0162] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. A video face replacement method, characterized in that: The method includes: For an original face region to be replaced in a video frame, using at least two different face replacement models currently available, face replacement processing is performed on the original face region to obtain a face replacement result of the original face region; the face replacement result represents a face region in which the original face of a target object in the original face region is replaced with a target face; the at least two different face replacement models require different optimal face angles; Scoring the face replacement results output by each face replacement model, and selecting the optimal face replacement result based on the scores of each face replacement result; The original face region to be replaced in the video frame is replaced with the optimal face replacement result.
2. The method according to claim 1, characterized in that The original face area to be replaced is determined by the following steps: Performing feature extraction on the obtained reference face image of the target object and each face image region included in the video frame to obtain reference face image features of the target object and image features of each face image region; the face image region is determined according to image size information of the reference face image, and the face image region at least includes an original face region; Determining the facial image region corresponding to the image feature that meets a specified condition based on the similarity calculation result between the reference facial image feature of the target object and the image feature of each facial image region; wherein the specified condition includes: the similarity calculation result is the maximum value among the similarity results of each image feature, and the similarity result is greater than a first specified value; The original face region included in the face image region corresponding to the image feature that meets the specified condition is determined as the original face region to be replaced.
3. The method according to claim 2, characterized in that The original face area in each video frame is numbered in a specified order, and an index number value is also recorded in each video frame; before performing face replacement processing on the original face area to be replaced in the video frame using at least two different face replacement models currently available, the method further includes: For each video frame, the recorded index number value in the video frame is detected. If the recorded index number value is the number value corresponding to any original face area in the video frame, the step of performing face replacement processing on the original face area to be replaced in the video frame by using at least two different face replacement models currently available is continued; the original face area to be replaced refers to the original face area corresponding to the recorded index number value; If the recorded index number value is a second specified value, the step of performing face replacement processing on the original face area to be replaced in the video frame using at least two different face replacement models currently available is refused to be executed; the second specified value is different from the number value corresponding to any original face area, and the second specified value is used to indicate that there is no original face area to be replaced in the video frame, or there is no original face area in the video frame.
4. The method according to claim 1, characterized in that: The step of scoring the face replacement results output by each face replacement model and selecting the optimal face replacement result based on the scores of each face replacement result includes: For each face replacement result, a convolution operation is performed on the grayscale value of each pixel in the face replacement result based on the Laplace operator to obtain a convolution operation result corresponding to the face replacement result; The variance of the convolution operation result corresponding to each face replacement result is used as the score of each face replacement result; the score of each face replacement result is used to indicate the image blur degree of each face replacement result, and the image blur degree is inversely correlated with the variance of the convolution operation result; The face replacement result with the highest score among all face replacement results is determined as the optimal face replacement result.
5. The method according to claim 1, characterized in that Before performing face replacement processing on the original face region to be replaced in the video frame using at least two different face replacement models currently available, the method further includes: Based on the trained neural network model, classify and predict the image features of the video frame and the text features of the text information included in the video frame to obtain a classification prediction result for the video frame; If the classification prediction results of all video frames meet the normal video frame conditions, the step of performing face replacement processing on the original face area to be replaced in the video frame using at least two different face replacement models currently available is executed.
6. The method according to claim 1, characterized in that The replacing the original face region to be replaced in the video frame with the optimal face replacement result comprises: Based on the trained face restoration model, the optimal face replacement result is image optimized to obtain the optimal face replacement result after image optimization; the original face area to be replaced in the video frame is replaced with the optimal face replacement result after image optimization; Alternatively, the original face region to be replaced in the video frame is replaced with the optimal face replacement result; and image optimization is performed on all face regions included in the video frame based on the trained face restoration model.
7. A video face replacement device, characterized in that: The device includes: A first replacement unit is used for performing face replacement processing on an original face region to be replaced in a video frame by using at least two different face replacement models currently available, so as to obtain a face replacement result of the original face region; the face replacement result represents a face region in which an original face of a target object in the original face region is replaced with a target face; and the at least two different face replacement models require different optimal face angles; A scoring selection unit, used to score the face replacement results output by each face replacement model, and select the optimal face replacement result based on the scores of each face replacement result; The second replacement unit is used to replace the original face area to be replaced in the video frame with the optimal face replacement result.
8. The device according to claim 7, characterized in that The first replacement unit is also used for: Performing feature extraction on the obtained reference face image of the target object and each face image region included in the video frame to obtain reference face image features of the target object and image features of each face image region; the face image region is determined according to image size information of the reference face image, and the face image region at least includes an original face region; Determining the facial image region corresponding to the image feature that meets a specified condition based on the similarity calculation result between the reference facial image feature of the target object and the image feature of each facial image region; wherein the specified condition includes: the similarity calculation result is the maximum value among the similarity results of each image feature, and the similarity result is greater than a first specified value; Determine the original face region included in the face image region corresponding to the image feature that meets the specified condition as the original face region to be replaced; And / or, the original face area in each video frame is numbered in a specified order, and an index number value is also recorded in each video frame; before performing face replacement processing on the original face area to be replaced in the video frame using at least two different face replacement models currently available, the first replacement unit is further used to: For each video frame, the recorded index number value in the video frame is detected. If the recorded index number value is the number value corresponding to any original face area in the video frame, the step of performing face replacement processing on the original face area to be replaced in the video frame by using at least two different face replacement models currently available is continued; the original face area to be replaced refers to the original face area corresponding to the recorded index number value; If the recorded index number value is a second specified value, the step of performing face replacement processing on the original face region to be replaced in the video frame by using at least two different face replacement models currently available is rejected; the second specified value is different from the number value corresponding to any original face region, and the second specified value is used to indicate that there is no original face region to be replaced in the video frame, or there is no original face region in the video frame; And / or, the scoring selection unit is specifically used for: For each face replacement result, a convolution operation is performed on the grayscale value of each pixel in the face replacement result based on the Laplace operator to obtain a convolution operation result corresponding to the face replacement result; The variance of the convolution operation result corresponding to each face replacement result is used as the score of each face replacement result; the score of each face replacement result is used to indicate the image blur degree of each face replacement result, and the image blur degree is inversely correlated with the variance of the convolution operation result; The face replacement result with the highest score among all face replacement results is determined as the optimal face replacement result; And / or, before performing face replacement processing on the original face region to be replaced in the video frame using at least two different face replacement models currently available, the first replacement unit is further used to: Based on the trained neural network model, classify and predict the image features of the video frame and the text features of the text information included in the video frame to obtain a classification prediction result for the video frame; If the classification prediction results of all video frames meet the normal video frame condition, the step of performing face replacement processing on the original face region to be replaced in the video frame by using at least two different face replacement models currently available is executed; And / or, the second replacement unit is specifically used for: Based on the trained face restoration model, the optimal face replacement result is image optimized to obtain the optimal face replacement result after image optimization; the original face area to be replaced in the video frame is replaced with the optimal face replacement result after image optimization; Alternatively, the original face region to be replaced in the video frame is replaced with the optimal face replacement result; and image optimization is performed on all face regions included in the video frame based on the trained face restoration model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Intelligent face replacement technology suitable for film and television post-production
CN117315089A