An automatic target locking algorithm for video face swapping based on target tracking

By using an automatic locking algorithm based on target tracking in video face exchange, detecting and replacing face areas, the problem of insufficient face exchange accuracy, nature and real-time performance in the prior art is solved, and a more efficient and reliable face exchange effect is achieved.

CN118887714BActive Publication Date: 2025-05-16HANGZHOU XUANYE DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410907105.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-05-16
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

The prior art has problems with insufficient accuracy, nature and real-time performance in video face exchange, especially when dealing with complex backgrounds and dynamic scenes, which are not effective.

Method used

The video face exchange target automatic locking algorithm based on target tracking is used to detect the face area of ​​each frame in the source video and the target video, and the difference in the face area is calculated using a pre-set difference detection model, and the face area in the target video is replaced with its corresponding closest face area.

Benefits of technology

It significantly improves the accuracy, nature, real-time and robustness of face exchange, and can maintain high-precision face detection and matching under complex lighting conditions and backgrounds, ensuring the reliability and efficiency of the face exchange process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887714B_ABST
    Figure CN118887714B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic target locking algorithm for video face swapping based on target tracking, and relates to the field of deep learning technology. It comprises: step 1: detecting the face area of ​​each frame in the source video and the target video, and obtaining the source video face area sequence and the target video face area sequence respectively; step 2: using a pre-set difference detection model, calculating the difference between each face area in the target video face area sequence and each face area in the source video face area sequence, and taking the face area of ​​the source video face area sequence corresponding to the smallest difference as the closest face area; step 3: replacing each face area in the target video with its corresponding closest face area. The present invention significantly improves the accuracy, naturalness, real-time and robustness of face swapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to an automatic locking algorithm for video face swapping targets based on target tracking. Background Art

[0002] Video editing and face swapping technologies are widely used in film and television production, entertainment media, security monitoring and other fields. However, the existing technologies still have significant deficiencies in accuracy, naturalness and real-time performance. In order to better understand the technical background of the present invention, this article will be described in detail in combination with the existing technologies and point out the problems therein.

[0003] Early face swapping technologies mainly relied on template matching methods. These methods usually achieve face swapping by searching for areas in the target image that are similar to the facial features in the source image. The advantage of this method is that it is relatively simple to implement and has a small amount of computation, but its accuracy and robustness are poor, especially when facing illumination changes, posture changes, and partial occlusion. In addition, template matching methods often have difficulty maintaining the naturalness of the replacement effect when dealing with complex backgrounds and dynamic scenes, and are prone to edge discontinuities and image distortion. With the development of computer vision technology, methods based on feature point detection and deformation have gradually become mainstream. These methods detect the key feature points of the face (such as eyes, nose, mouth, etc.), and then use these feature points for geometric deformation and alignment. For example, the active shape model (ASM) and active appearance model (AAM) are common feature point detection methods. By establishing a statistical model of the face, feature points can be accurately located and affine transformation can be performed. This method improves the accuracy and naturalness of face swapping to a certain extent, but the effect is still not ideal when dealing with large-angle posture changes and complex expressions. In addition, these methods have high requirements for computing resources and poor real-time performance.

[0004] In recent years, the rapid development of deep learning technology has brought new breakthroughs in face swapping technology. Deep learning-based methods can achieve high-precision and high-natural face replacement by training convolutional neural networks (CNNs) and generative adversarial networks (GANs). For example, technologies such as FaceSwap and DeepFake use GANs to generate high-quality synthetic face images, greatly improving the effect of face swapping. These methods train models with a large amount of face image data and can achieve natural face replacement in various complex scenarios. However, deep learning methods also have some problems. First, these methods require a large amount of training data and computing resources, and the training process takes a long time. Second, although the generated effect is realistic, discontinuous and unnatural edge transitions may still occur in practical applications. In addition, the robustness of deep learning methods to changes in lighting and posture needs to be improved. Summary of the invention

[0005] In view of this, the present invention provides a video face swap target automatic locking algorithm based on target tracking, which significantly improves the accuracy, naturalness, real-time and robustness of face swapping.

[0006] The technical solution adopted by the present invention is as follows:

[0007] An automatic target locking algorithm for video face swapping based on target tracking, which includes:

[0008] Step 1: Detect the face region of each frame in the source video and the target video, and obtain the source video face region sequence and the target video face region sequence respectively;

[0009] Step 2: Using a preset difference detection model, calculate the difference between each face region in the target video face region sequence and each face region in the source video face region sequence, and take the face region in the source video face region sequence corresponding to the smallest difference as the closest face region;

[0010] Step 3: Replace each face region in the target video with its closest corresponding face region.

[0011] Furthermore, in step 1, before detecting the face area of ​​each frame in the source video and the target video, the video frames of the source video and the target video are first extracted to obtain a source video frame set and a target video frame set, respectively; then, each frame image in the source video frame set and the target video frame set is subjected to image denoising processing and converted into a grayscale image.

[0012] Furthermore, the method for detecting the face area in each frame of the source video and the target video in step 1 includes: using a pre-trained multi-scale multi-resolution model based on a regression tree to slide windows of different sizes in each frame of the source video and the target video, extracting the weighted multi-scale multi-resolution value in the window, if the weighted multi-scale multi-resolution value is greater than a set threshold, it is judged that it belongs to the face area, and the face area of ​​each frame of the image includes all parts covered by the window belonging to the face area.

[0013] Furthermore, the pre-trained regression tree-based multi-scale multi-resolution model is trained through the following process: for each pixel in each sample image in the training set, a label corresponds to the label, which marks whether the pixel belongs to the face area; for each pixel in each sample image in the training set, the pixel itself is taken as the center pixel, a radius R and the number of sampling points P are set to define the neighborhood range, and within the neighborhood range, the difference between its weighted gray value at different scales and different resolutions and the gray value of the center pixel is calculated to obtain a weighted multi-scale multi-resolution value; a weighted multi-scale multi-resolution regression tree is constructed, specifically including: taking each pixel as a node in the regression tree, and at each node, selecting a split feature and a split point so that the weighted mean square error of the weighted multi-scale multi-resolution value inside the child node after the split is minimized, thereby selecting the best split point to recursively construct the tree structure until the preset tree depth is reached.

[0014] Furthermore, the weighted multi-scale multi-resolution value WMSMLBP is calculated using the following formula:

[0015]

[0016] Among them, i c is the gray value of the center pixel; is the grayscale value of the pth sampling point at scale s and radius r; S is the number of scales; w(p,s,r) is the combined weight of distance, scale and radius, calculated using the following formula:

[0017]

[0018] in, is the distance from the pth sampling point to the center pixel at scale s and radius r, σ sr and σ r is the standard deviation of the scale s and radius r, and n is the order of the preset polynomial.

[0019] Furthermore, the weighted mean square error MSMR-WMSE of the weighted multi-scale multi-resolution values ​​is calculated using the following formula:

[0020]

[0021] Where w(i,s,r) is the weight of the i-th sample image in the training set at scale s and radius r, N is the number of samples, and y i is the label of the i-th sample image, is the predicted label.

[0022] Furthermore, in step 2: let the i-th face region in the face region sequence in the target video be T i; The jth face region in the face region sequence in the source video is S j , use the following formula to calculate the difference D H (T i ,S j ):

[0023]

[0024] in, Indicates the calculation of the histogram mean; m is the number of buckets in the histogram; H k (T i ) means calculating the histogram mean of the kth bucket.

[0025] Further, according to the difference calculation result of step 2, the face area T of each face in the target video is determined. i The corresponding closest face area S j ; Use facial key point detection to detect T i and S j Align, the key point set P T and P S Respectively represent T i and S j The key point position of S j to T i The TPS nonlinear transformation matrix is: the area closest to the face S j Based on the nonlinear transformation matrix, affine transformation is performed to the face area T i The position of; is the face area after affine transformation Use the multi-resolution fusion method to create a multi-layer mask M; transform the face area after affine transformation With face area T i Perform multi-resolution fusion; replace the fused face area R to the corresponding position of the target video.

[0026] Furthermore, the Laplace pyramid fusion method is used to transform the face area after affine transformation With face area T i Perform multi-resolution fusion.

[0027] By adopting the above technical scheme, the present invention produces the following beneficial effects: the present invention adopts the weighted multi-scale multi-resolution value (WMSMLBP) and weighted mean square error (MSMR-WMSE) methods to perform high-precision detection and matching of face regions in video frames. WMSMLBP extracts image features at different scales and resolutions and uses weighted processing to ensure the accuracy and robustness of feature extraction. This method can effectively capture the details and overall features of the face region, so that high accuracy can be maintained even under complex lighting conditions and backgrounds during face detection. MSMR-WMSE optimizes the construction of the regression tree by weighted averaging the errors at different scales and resolutions, so that the model can accurately predict the face region in different environments. The comprehensive application of these technologies significantly improves the accuracy of face detection and matching, and ensures the reliability of the subsequent face exchange process. The present invention fully considers real-time performance and computational efficiency in algorithm design. Although the multi-scale multi-resolution weighted processing method and weighted mean square error optimization strategy are more complex in calculation, they can achieve real-time processing while ensuring high accuracy through reasonable algorithm optimization and efficient computing architecture. Specifically, the present invention adopts a pre-trained regression tree model and an efficient feature matching algorithm, so that the face detection and matching process can be completed in a relatively short time, which is suitable for real-time video processing requirements. In addition, the present invention also adopts a multi-resolution fusion technology, which not only retains the detail features but also ensures the smoothness of the overall transition by processing and fusing images layer by layer at different resolution levels. This method not only improves the processing efficiency, but also effectively reduces the computational overhead, so that the present invention has high real-time performance in practical applications. In the design process of the present invention, the robustness and adaptability under various complex scenes are fully considered. Through the multi-scale and multi-resolution feature extraction method, the model can still maintain high-precision face detection and matching under different lighting conditions, different posture angles and partial occlusion. This method captures image features at different scales and resolutions, ensuring that the model fully grasps the details and overall features, thereby improving the adaptability of the system in complex environments. In addition, the weighted processing method of the present invention uses Gaussian distribution and smoothing functions to comprehensively weight distance, scale and radius, making feature extraction more stable and accurate. Especially when dealing with dynamic videos and complex backgrounds, this method can effectively reduce the impact of background noise and dynamic interference, ensuring the robustness of face detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 The present invention is a flowchart of a method for automatically locking a target in a video face swap based on target tracking in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] All features disclosed in this specification, or steps in all methods or processes disclosed, except mutually exclusive features and / or steps, can be combined in any manner.

[0030] Any feature disclosed in this specification (including any additional claims and abstract), unless otherwise stated, may be replaced by other equivalent or alternative features having similar purposes. That is, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.

[0031] Example 1: Reference Figure 1 , an automatic target locking algorithm for video face swapping based on target tracking, which includes:

[0032] Step 1: Detect the face region of each frame in the source video and the target video, and obtain the source video face region sequence and the target video face region sequence respectively;

[0033] The core of this step is to use advanced computer vision technology to accurately locate the face area in the video frame, providing a basis for subsequent face swapping. Specifically, the source video and the target video need to be processed frame by frame first, and each frame goes through preprocessing steps such as color space conversion and scale scaling to meet the needs of the face detection algorithm. Next, a deep learning model is used, such as a convolutional neural network (CNN)-based detector for face detection. This model can learn the characteristics of the face by training a large amount of face data, so that it can accurately identify the face area under different lighting, posture and expression conditions.

[0034] In the specific implementation, a multi-task convolutional neural network (MTCNN) can be used for face detection. MTCNN uses a three-level cascade network structure to refine the face detection results step by step, thereby improving the accuracy and robustness of detection. First, P-Net quickly screens out candidate face regions through simple convolution and pooling operations; then R-Net further refines these candidate regions and removes overlapping detection boxes through non-maximum suppression (NMS); finally, O-Net performs key point detection and bounding box regression on the final face region to accurately locate the face region. Through this step-by-step refinement process, MTCNN is able to provide high-precision face detection results while maintaining high efficiency. The detected face region is usually represented in the form of a bounding box, each of which contains the location information of the face region, such as the coordinates of the upper left corner and the lower right corner. In order to improve the stability and accuracy of the detection, the detection results can be further processed in combination with the target tracking algorithm. Target tracking algorithms such as KCF (Kernelized Correlation Filters), MOSSE (Minimum Output Sum of Squared Error) or deep learning driven trackers such as DeepSort can maintain continuous tracking of the same face area in a video frame sequence by tracking the position changes of the face between adjacent frames, thereby reducing detection errors caused by factors such as instantaneous occlusion and lighting changes.

[0035] In the process of face region detection and tracking, some practical problems need to be dealt with, such as multi-scale detection, face rotation and partial occlusion. Multi-scale detection ensures that faces of different sizes can be accurately detected by applying face detection algorithms in different scale spaces. For face rotation and partial occlusion, data enhancement technology and rotation invariance design can be used to improve the robustness of the detector. For example, face samples with various rotation angles and occlusion conditions can be introduced when training the detection model to enhance the generalization ability of the model. Through the above process, the face region can be accurately detected in each frame of the source video and the target video, and the continuity and stability of these detection results in time are ensured by the target tracking algorithm. Finally, the face region sequence of the source video and the target video is completely extracted, laying a solid foundation for the subsequent difference calculation and face region replacement. This step requires not only efficient computer vision algorithms, but also powerful computing resources and optimized algorithm implementation to meet the needs of real-time video processing. Based on this high-precision face detection and tracking technology, the present invention can achieve accurate, stable and efficient face exchange in complex video environments, thereby showing strong application potential and technical advantages in the fields of video editing, film and television production, etc.

[0036] Step 2: Using a preset difference detection model, calculate the difference between each face region in the target video face region sequence and each face region in the source video face region sequence, and take the face region in the source video face region sequence corresponding to the smallest difference as the closest face region;

[0037] The selection and design of the difference detection model is crucial. An effective difference detection model needs to be able to capture various subtle features of the face, including facial contours, facial features, texture features, etc. To achieve this goal, a feature extraction method based on deep learning is usually used. For example, a convolutional neural network (CNN) is used to extract facial feature vectors, which can effectively represent the high-dimensional features of the face.

[0038] By training a deep learning model, such as FaceNet or VGG-Face, the face image in each frame can be converted into a fixed-length feature vector that preserves the uniqueness and similarity of the face in the feature space.

[0039] Next, the difference between each face region in the target video and the source video is calculated. The purpose of difference calculation is to evaluate the similarity between two feature vectors. Common methods include Euclidean distance, cosine similarity, etc. Euclidean distance measures the difference between two feature vectors by calculating the geometric distance between them in the feature space, while cosine similarity measures the similarity between them by calculating the cosine value of the angle between the two vectors. For the problem of facial feature matching, Euclidean distance is more intuitive and has good results. After calculating the distance between the feature vectors of each pair of face regions, the pair with the smallest distance is selected as the closest face region. This means that in the feature space, the two feature vectors are most similar to each other, so it can be considered that the corresponding two face regions are also visually the most similar. In order to improve the accuracy and robustness of matching, more contextual information can be introduced in the difference calculation. For example, the position and shape of key points such as eyes, nose, and mouth can be matched in combination with facial key point detection technology. This method can further refine the difference calculation so that it not only depends on the global feature vector, but also considers the matching of local features. In addition, a two-way matching strategy can be used to calculate the difference between the target video face area and the source video face area, as well as the difference between the source video face area and the target video face area. Only when both reach the minimum difference is the match considered to be optimal.

[0040] In practical applications, the difference detection model also needs to deal with some complex situations, such as illumination changes, facial expression changes, and partial occlusion. Illumination changes can be alleviated by normalization or illumination invariance feature extraction methods, while facial expression changes can be dealt with by introducing more training samples and enhancing the generalization ability of the model. For partial occlusion, robust feature extraction methods, such as local feature aggregation or occlusion repair technology, can be used to ensure the stability and reliability of the feature vector. Through the above-mentioned difference calculation and matching process, each face area in the target video and the closest face area in the source video can be accurately found. This not only ensures the accuracy of face exchange, but also improves the naturalness and realism of the final video effect. This step plays a key role in the algorithm, connecting the two core links of face detection and face replacement, and ensuring the fluency and consistency of the whole process. Based on the high-precision difference detection model and the fine feature matching method, the present invention can realize efficient and accurate face exchange in various complex scenes, and has significant technical advantages and application prospects.

[0041] Step 3: Replace each face region in the target video with its closest corresponding face region.

[0042] The face regions of the target video and the source video have been accurately detected and matched through the first two steps. Next, preprocessing must be performed to ensure seamless fusion of the two face regions when replacing. This preprocessing process usually includes geometric transformation and color adjustment. The purpose of geometric transformation is to adjust the face region of the source video to the same position, size and angle as the face region in the target video. This can be achieved through affine transformation or perspective transformation, depending on the degree of deformation of the face region. Affine transformation can effectively handle rotation, scaling and translation, while perspective transformation is suitable for more complex deformation situations. After completing geometric alignment, color adjustment is another important step to ensure the natural replacement effect. Since the shooting conditions of the source and target videos may be different, such as lighting, color temperature, etc., this will cause obvious differences in the color and brightness of the face region. To eliminate these differences, color matching algorithms such as histogram matching or color conversion can be used. These algorithms can adjust the color distribution of the face region of the source video to be consistent with the face region of the target video, thereby visually eliminating inconsistencies.

[0043] After geometric transformation and color adjustment of the face region, the next step is to seamlessly blend the face region of the source video into the target video. This step usually requires the use of image fusion technology to ensure a natural transition. Common image fusion techniques include Laplacian pyramid fusion and Poisson fusion. Laplacian pyramid fusion decomposes the multi-level frequency components of the image and fuses them at each frequency level, thereby ensuring the consistency of the fused image at different scales. Poisson fusion solves a Poisson equation to make the gradient of the fusion area consistent with the surrounding area, thereby achieving a seamless fusion effect. In addition, in order to further improve the realism of face replacement, a generative adversarial network (GAN) based on deep learning can be introduced for detail restoration. GAN can learn the overall style and detail features of the target video, thereby generating high-quality face regions consistent with the style of the target video. In this process, the generative network is responsible for generating new details, while the discriminative network evaluates the authenticity of the generated results. The two continuously improve the quality of the generated results through adversarial training. Ultimately, face replacement not only needs to be aligned and fused in space, but also needs to maintain consistency in time. Since the video is a series of continuous frames, the face replacement process must ensure that the face area of ​​each frame has a natural transition in time after replacement without flickering or jumping. This can be achieved through technologies such as optical flow algorithms or temporal convolutional networks (TCN). The optical flow algorithm calculates the pixel motion between adjacent frames to ensure that the motion trajectory of the replaced face area is consistent with the original video. The temporal convolutional network captures the temporal dependencies by modeling the video frame sequence, thereby achieving smooth transition in the temporal dimension.

[0044] Embodiment 2: Step 1 Before detecting the face area of ​​each frame in the source video and the target video, first extract the video frames of the source video and the target video to obtain a source video frame set and a target video frame set, respectively; then perform image denoising on each frame image in the source video frame set and the target video frame set, and convert them into grayscale images.

[0045] Specifically, after extracting the video frame, each frame of the image needs to be denoised. The purpose of image denoising is to eliminate random noise in the image and improve the image quality, thereby providing a clear input image for subsequent face detection. Noise is a common problem in image processing, especially in low-light conditions or videos shot with high ISO sensitivity. Common denoising methods include mean filtering, median filtering, and Gaussian filtering. Mean filtering smoothes the details and noise in the image by averaging the local area of ​​the image; median filtering eliminates sharp noise points by taking the median of the local area, which is particularly suitable for removing salt and pepper noise; Gaussian filtering performs Gaussian smoothing on the image to better retain the edge information of the image while removing noise. Through these denoising techniques, the noise interference in the image can be greatly reduced, ensuring the high accuracy and stability of the face detection algorithm in subsequent processing. After denoising, the image needs to be converted into a grayscale image. A grayscale image means that each pixel in the image contains only one grayscale value, which represents the brightness information of the pixel. Compared with color images, grayscale images reduce the dimension of data, thereby reducing the computational complexity, making subsequent face detection more efficient and accurate. The process of grayscale conversion is usually performed by a weighted summation method to linearly combine the three color channels of the RGB image according to certain weights to generate a single-channel grayscale image. Common weighted combinations are: grayscale value = 0.299*R+0.587*G+0.114*B, where R, G, and B represent the pixel values ​​of the red, green, and blue channels, respectively. This weighted combination reflects the difference in sensitivity of the human eye to different colors, thereby generating a more natural and realistic grayscale image. Through this conversion, the efficiency and accuracy of image processing have been significantly improved, providing simplified input data for subsequent face detection and tracking. These processing steps in Example 2 ensure that each frame image has been optimized before entering the face detection stage, minimizing the impact of noise and other interference factors. Face detection algorithms such as multi-task convolutional neural networks (MTCNN) can more accurately locate face areas when processing these pre-processed images. MTCNN uses its three-level cascade structure to gradually refine the face detection results, thereby improving the accuracy and robustness of detection. The first layer of the network quickly screens out candidate face regions, the second layer of the network further refines these candidate regions and removes overlapping detection frames through non-maximum suppression (NMS), and the third layer of the network performs key point detection and bounding box regression on the final face region to accurately locate the face region.

[0046] Embodiment 3: The method for detecting the face area in each frame of the source video and the target video in step 1 includes: using a pre-trained multi-scale multi-resolution model based on a regression tree to slide windows of different sizes in each frame of the source video and the target video, extracting the weighted multi-scale multi-resolution value in the window, and if the weighted multi-scale multi-resolution value is greater than a set threshold, it is judged that it belongs to the face area, and the face area of ​​each frame of the image includes all parts covered by the window belonging to the face area.

[0047] Specifically, the pre-trained multi-scale and multi-resolution model based on regression trees is the core of the entire detection step. The regression tree model can capture the complex features of the face by learning from a large amount of face image data and make predictions based on this. The multi-scale and multi-resolution model means analyzing the image at different scales and resolutions, which can ensure accurate detection of faces of various sizes and distances. The regression tree can handle non-linear feature relationships and build a series of decision trees through training to predict whether the input image contains a face. The advantage of this model lies in its efficient computing power and strong generalization ability, which can maintain high accuracy in different environments. In actual operation, the multi-scale and multi-resolution analysis of each frame of the image is achieved through sliding windows of different sizes. The sliding window technology means using fixed-size windows to slide step by step on the image to extract local image features through these windows. These feature values ​​are input into the pre-trained regression tree model for prediction. If the weighted multi-scale and multi-resolution value of a window exceeds the set threshold, the window is marked as containing a face area. The calculation of weighted multi-scale multi-resolution values ​​is based on the pixel intensity, edge features and texture information of the local image, which can more accurately reflect the characteristics of the face after weighted processing. The advantage of this detection method is that it can handle complex image scenes. Traditional face detection methods may encounter difficulties when dealing with lighting changes, partial occlusions and complex backgrounds, while the multi-scale multi-resolution model based on regression trees can work stably under different conditions through multi-level feature extraction and comprehensive analysis. For example, in the case of uneven lighting conditions, multi-scale analysis can capture the feature changes under different lighting conditions, thereby ensuring the robustness of face detection. Similarly, for partial occlusion, the regression tree model can identify that the occluded part still belongs to the face area through comprehensive judgment of local features.

[0048] When all windows are scanned and analyzed, the face region of each frame image will be determined. Due to the overlapping parts of the sliding windows, the detected face region is usually the overlapping part of multiple windows, which can more accurately outline the boundary of the face. This step ensures the high accuracy of face detection and provides reliable regional information for subsequent face exchange. Next, by comparing the detection results of each frame in the target video with the detection results of the corresponding frame in the source video, the similarity between each face region is calculated using the aforementioned difference detection model to determine the best match. This process is carried out on the basis of the previous two steps. By comparing the facial features in the source video and the target video, the most similar region is selected for replacement. This multi-scale and multi-resolution detection method based on regression trees is not only superior to traditional methods in detection accuracy, but also has obvious advantages in computational efficiency. The pre-trained model is trained offline, and only fast prediction is required in practical applications, which greatly reduces the computational overhead. In addition, the multi-scale and multi-resolution analysis method enables the system to maintain high accuracy when processing faces of different sizes and distances.

[0049] Embodiment 4: The pre-trained multi-scale multi-resolution model based on regression tree is trained through the following process: for each pixel in each sample image in the training set, a label corresponds to each pixel, and the label marks whether the pixel belongs to the face area; for each pixel in each sample image in the training set, taking itself as the center pixel, a radius R and the number of sampling points P are set to define the neighborhood range, and within the neighborhood range, the difference between its weighted gray value at different scales and different resolutions and the gray value of the center pixel is calculated to obtain a weighted multi-scale multi-resolution value; constructing a weighted multi-scale multi-resolution regression tree, specifically including: taking each pixel as a node in the regression tree, and at each node, selecting a split feature and a split point so that the weighted mean square error of the weighted multi-scale multi-resolution value inside the child node after the split is minimized, thereby selecting the best split point to recursively construct a tree structure until a preset tree depth is reached.

[0050] Specifically, the first step of the training process is to label each sample image in the training set. This means that for each pixel, the system needs to mark whether it belongs to the face area. These labels are the basis of supervised learning. Through a large amount of labeled data, the model can learn the characteristic differences between the face area and the non-face area. In each sample image, each pixel is regarded as an independent sample, which provides a data basis for subsequent feature extraction. In order to extract more features from the image, a radius R and the number of sampling points P are set with each pixel as the center to define its neighborhood range. Within this neighborhood, the weighted grayscale value of each sampling point at different scales and resolutions is calculated. The weighted grayscale value is smoothed and enhanced by a specific weighting function. This method ensures that the features of the image can be fully extracted at different scales and resolutions, thereby capturing more facial details.

[0051] Specifically, the weighted multi-scale multi-resolution values ​​are obtained by calculating the difference between the grayscale value of each sampling point in the neighborhood and the grayscale value of the central pixel. These differences reflect the local pattern of pixel intensity changes and are very effective in distinguishing face areas from background areas. The combination of different scales and resolutions enables the model to capture image features at different levels, from details to the whole, and comprehensively cover possible face features. For example, a smaller scale can capture subtle texture changes, while a larger scale can capture a wider range of structural features. After obtaining these feature values, the next step is to build a weighted multi-scale multi-resolution regression tree. The construction of a regression tree is a recursive process, the purpose of which is to select the best splitting feature and splitting point at each node so that the weighted mean square error of the weighted multi-scale multi-resolution value within the child node is minimized. Specifically, each pixel is regarded as a node in the regression tree. When the node splits, all possible splitting features and splitting points are evaluated to select the splitting point that can minimize the error. This process is achieved by traversing all features and possible splitting points, and finding the best splitting solution by calculating the error difference before and after the split. The process of recursively building the tree structure continues until the preset tree depth is reached. The preset tree depth is to prevent overfitting while keeping the model's complexity moderate so that it has good generalization ability. In this way, the regression tree can refine the division of image features layer by layer, so that the pixel set corresponding to each leaf node is more clustered in the feature space and the error is smaller. Each node of the tree divides the input feature space into several regions through decision rules, each region corresponds to a different distribution of facial features, thereby achieving accurate recognition of facial regions. In practical applications, the pre-trained regression tree model is used for frame-by-frame face detection of video frames. After each frame of the image is processed by the model, an accurate face region can be obtained. Due to the application of sliding window technology, the model can gradually scan the image at different positions and scales to ensure that all potential face regions are fully detected. The sliding window technology combined with multi-scale and multi-resolution analysis greatly improves the coverage and accuracy of detection.

[0052] This method is efficient and accurate in detecting face regions, and can stably identify faces under complex backgrounds and changing lighting conditions. The pre-trained regression tree model is trained offline, so only fast prediction is required in practical applications, which greatly reduces the computational overhead. The multi-scale and multi-resolution analysis method enables the system to maintain high accuracy when processing faces of different sizes and distances. Through this multi-level feature extraction and regression tree construction method, the system can achieve accurate face detection in complex environments. For example, in the case of uneven lighting conditions, multi-scale analysis can capture feature changes under different lighting conditions, thereby ensuring the robustness of face detection. Similarly, for partial occlusion, the regression tree model can identify that the occluded part still belongs to the face region through comprehensive judgment of local features. In addition, this method has good scalability. By adjusting the parameters of the regression tree, such as the depth of the tree and the number of split nodes, the performance of the model can be optimized to adapt to different application scenarios. For example, for high-resolution videos, the depth of the regression tree can be increased to capture more detailed features; for real-time applications, the computational efficiency of the model can be optimized to ensure high-precision face detection at high frame rates. Throughout the process, the pre-trained regression tree model not only improves the accuracy and robustness of face detection, but also ensures the consistency of detection under different lighting, posture and expression conditions. In this way, the face area in each frame of the source video and the target video is accurately detected and extracted, laying a solid foundation for the subsequent difference calculation and face area replacement. The core of the difference calculation and face area replacement process is based on the detected accurate face area. By comparing the face area features in the source video and the target video, the system can calculate the difference between each pair of face areas to find the best match. This process utilizes the aforementioned multi-scale and multi-resolution features to ensure the accuracy and stability of the match. Finally, the automatic target locking algorithm for video face swapping based on target tracking achieves high-quality automatic face swapping through accurate face detection, difference calculation and face area replacement. The whole system demonstrates strong robustness and efficiency in practical applications, can run stably in complex environments, and has broad application prospects and significant technical advantages. This method not only has important application value in the fields of video editing and film and television production, but also provides new ideas and methods for the development of computer vision and image processing technology. Through continuous optimization and expansion, the present invention is expected to play a greater role in more practical application scenarios.

[0053] Embodiment 5: The weighted multi-scale multi-resolution value WMSMLBP is calculated using the following formula:

[0054]

[0055] The core idea of ​​the formula is to accumulate the grayscale differences of pixels at different scales and resolutions by weighted method. c The gray value of all other sampling points is used as the reference. They are weighted according to their differences from the center pixel. This process not only takes into account the difference in grayscale values, but also combines the weights of scale and radius, so that features of different scales and resolutions can be effectively fused. Specifically, the formula obtains a comprehensive feature value by weighted accumulation of the difference between the grayscale value of each sampling point at different scales and radii and the grayscale value of the center pixel. This weighting process is determined by the weight function w(p,s,r), which is calculated based on the distance from the sampling point to the center pixel and the preset scale and radius standard deviation. In this way, the grayscale value that is closer to the center pixel and has a more significant difference at the corresponding scale and radius will have a greater impact on the final WMSMLBP value.

[0056] This multi-scale and multi-resolution weighted processing method has many advantages. First, by analyzing at different scales, the formula can capture subtle changes and local features in the image. For example, details such as edges and textures can be detected at small scales, while larger scales can capture more extensive structural features, such as the outline and overall shape of the face. This approach allows the detection process to focus on details while not ignoring the overall structure, thereby maintaining high accuracy when dealing with complex backgrounds and lighting changes. Second, the weight function w(p,s,r) introduced in the formula further enhances the effectiveness of feature extraction. The weight function takes into account the distance factor and weights the distance through a Gaussian function, so that the closer the sampling point is to the center pixel, the greater its contribution. This design reflects the continuity of facial features in space, that is, the grayscale changes of adjacent pixels have a strong correlation. At the same time, the weight function also combines the standard deviation of scale and radius to ensure that feature extraction at different scales and radii is consistent and robust. Through this comprehensive feature extraction method, the WMSMLBP value can fully express the local and global features of the face area, providing strong support for subsequent face detection. Especially in target tracking video face swapping, this method can stably identify the face area in each frame of the image, regardless of fast motion or lighting changes, and can maintain a high detection accuracy. Furthermore, this multi-scale and multi-resolution feature extraction method also provides a solid foundation for difference calculation. In practical applications, by calculating the WMSMLBP values ​​of the face areas in the source video and the target video, the similarity between them can be quantified, thereby achieving accurate matching and replacement. This feature value-based difference calculation method can effectively distinguish different face areas, ensuring the natural effect and consistency after face swapping.

[0057] Among them, i cis the gray value of the center pixel; is the grayscale value of the pth sampling point at scale s and radius r; S is the number of scales; w(p,s,r) is the combined weight of distance, scale and radius, calculated using the following formula:

[0058]

[0059] The weighting function is designed to weight different sampling points according to their relative positions and the difference in their grayscale values, so that points that are closer to the center pixel and have significant feature differences contribute more to the final WMSMLBP value. Specifically, this weighting function consists of several parts: one is the Gaussian distribution part, which is used to weight the distance; the second is the polynomial weighting part, which is used to combine distance and scale information; and the third is another Gaussian distribution, which is used to process the weight of a specific radius. First, the Gaussian distribution part is used to weight the distance between the sampling point and the center pixel. The Gaussian distribution function has the characteristics of central symmetry, and its weight decreases exponentially with the increase of distance. This means that the closer the sampling point is to the center pixel, the higher its weight is and the greater its influence on the final WMSMLBP value is. This design reflects a basic principle in image processing: spatially adjacent pixels usually have stronger correlation and similarity. Secondly, the polynomial weighting part further combines the information of distance and scale. This part adjusts the weight by the preset polynomial order n, so that the contribution of the sampling point can be finely controlled at different scales and radii. Specifically, the distance factor in the formula The pixels are normalized to a preset radius R and then weighted by a polynomial function. The design of this part enables the model to flexibly adapt to different image resolutions and scales, ensuring that features can be effectively extracted under various conditions. Finally, the second Gaussian distribution part is used to process the weight of a specific radius. This part is similar to the previous Gaussian distribution, but it focuses on the distribution of pixels within a specific radius. By introducing this Gaussian weight, the model can more finely weight the sampling points within different radii, thereby capturing more delicate image features at different scale levels. Through the combined effect of these three parts, the weighting function w(p,s,r) achieves effective weighting of feature points at different scales and resolutions. As a result, the model can accurately capture the features of the face area in multi-scale and multi-resolution image processing, thereby improving the accuracy and robustness of face detection. In the processing of video frames, the face area detection in each frame needs to maintain high accuracy under complex backgrounds and changing lighting conditions. Through the weighting function w(p,s,r), the system can ensure that facial features at different positions, scales and radii are effectively extracted and weighted, thereby achieving accurate recognition of facial areas. This multi-level weighted processing method not only improves the accuracy of feature extraction, but also enhances the robustness of the model in practical applications. For example, when dealing with scenes with uneven lighting, the weighting of the Gaussian distribution can effectively smooth the impact of lighting changes, while the polynomial weighting can further adjust the feature weights at different scales, making the detection results more stable. In addition, for partially occluded facial areas, the introduction of the second Gaussian distribution part enables the model to better handle these complex situations and maintain high detection accuracy.

[0060] in, is the distance from the pth sampling point to the center pixel at scale s and radius r, σ sr and σ r is the standard deviation of the scale s and radius r, and n is the order of the preset polynomial.

[0061] Example 6: The weighted mean square error MSMR-WMSE of the weighted multi-scale multi-resolution values ​​is calculated using the following formula:

[0062]

[0063] Where w(i,s,r) is the weight of the i-th sample image in the training set at scale s and radius r, N is the number of samples, and y i is the label of the i-th sample image, is the predicted label.

[0064] Specifically, the core idea of ​​the formula is to perform a weighted average of the prediction errors of each sample image at different scales and radii. MSMR-WMSE is calculated by weighting and summing the squared errors of each sample and then taking the average. This weighted process ensures that important features at different scales and radii can be fully considered, making the model more adaptable to various image features. Specifically, the weight function w(p,s,r) is introduced in the formula, which combines the weights of the distance, scale and radius of the sampling points to ensure that the points closer to the center pixel and with more significant feature differences have a greater impact on the error calculation. The mean square error (MSE) part involved in the formula is, It is a measure of the predicted value and the true value y i The standard method for calculating the difference between them. By calculating the square of the error, larger prediction errors can be amplified, thereby focusing on reducing these errors during the optimization process. However, traditional mean square error calculations may not be sufficient to capture complex image features in multi-scale and multi-resolution environments. Therefore, in the present invention, the accuracy of error assessment is improved by introducing multi-scale and multi-resolution weight functions. The weight function w(p,s,r) is the key part of the formula, which weights the distance, scale and radius through a Gaussian distribution function and a smoothing function. This design reflects a basic principle in image processing, that is, spatially adjacent pixels usually have stronger correlation and similarity. Specifically, a Gaussian function is used in the formula to weight the distance between the sampling point and the center pixel, so that the closer the sampling point is to the center pixel, the higher its weight. This part of the weight function is as follows in Represents the distance from the sampling point to the center pixel at scale s and radius r, σ sr is the standard deviation of the scale s and radius r.

[0065] In addition, the polynomial weighting part further combines the distance and scale information, and adjusts the weights by the preset polynomial order n, so that the contribution of the sampling points can be finely controlled at different scales and radii. Specifically, the weight function of this part is as follows Here R is the preset radius. This design allows the model to flexibly adapt to different image resolutions and scales, ensuring that features can be effectively extracted under various conditions. Finally, another Gaussian distribution part is used to process the weights of a specific radius. The weight function of this part is as follows σ ris the standard deviation of a specific radius r. By introducing this Gaussian weight, the model can perform more refined weighting on the sampling points within different radii, thereby capturing more delicate image features at different scale levels. Through the combined effect of these parts, MSMR-WMSE can accurately capture the features of the face region in multi-scale and multi-resolution image processing. As a result, the model can achieve efficient and accurate face detection in complex environments. This method is particularly important in the processing of video frames, because the face region detection in each frame image needs to maintain high accuracy under changing background and lighting conditions. Through MSMR-WMSE, the system can ensure that the face features at different positions, scales and radii are effectively extracted and weighted, thereby achieving accurate recognition of the face region. When dealing with scenes with uneven lighting, the weighting of the Gaussian distribution can effectively smooth the impact of lighting changes, while the polynomial weighting can further adjust the feature weights at different scales, making the detection results more stable. In addition, for partially occluded face regions, the design of the weight function enables the model to better handle these complex situations and maintain high detection accuracy. Through this comprehensive weight processing method, the present invention demonstrates strong application potential and technical advantages in the fields of video editing, film and television production, etc.

[0066] Example 7: In step 2: let the i-th face region in the face region sequence in the target video be T i ; The jth face region in the face region sequence in the source video is S j :

[0067]

[0068] in, Indicates the calculation of the histogram mean; n is the number of buckets in the histogram; H k (T i ) means calculating the histogram mean of the kth bucket.

[0069] Specifically, the histogram mean and It is the result of statistical analysis of the grayscale value distribution of the face area in the target and source videos. The grayscale value range of each face area is divided into n buckets, and each bucket records the number of pixels within a certain range. By calculating the mean value H of each bucket k (T i ) and H k (S j ), we can get the histogram features of the image. These features provide a detailed description of the grayscale distribution of the image, which helps to compare the similarities of different images. The similarity metric in the formula is implemented by calculating the product of the bucket means of the two histograms and taking the square root. Specifically, the formula The partial representation is to match each histogram bucket of the face area in the target and source videos, multiply and sum the values ​​of each bucket, and then get an overall similarity measure. This method can effectively capture the similarity of the grayscale distribution of the two images. In order to make the similarity measure more objective and accurate, a normalization factor is introduced in the formula This normalization factor takes into account the product of the two histogram means and the square root of the number of buckets. Through normalization, the influence of the brightness and contrast differences of different images can be eliminated, making the similarity measurement more stable and reliable. Finally, the formula calculates the difference D H (T i ,S j ), provides an effective method to measure the similarity between the face regions in the target video and the source video. The smaller the difference value, the more similar the grayscale distribution of the two regions is, so that face matching can be performed more accurately. This method plays a key role in the algorithm of the present invention, and improves the effect and naturalness of face swapping through accurate face region matching.

[0070] Embodiment 8: Determine each face region T in the target video according to the difference calculation result of step 2 i The corresponding closest face area S j ; Use facial key point detection to detect T i and S j Align, the key point set P T and P S Respectively represent T i and S j The key point position of S j to T i The TPS nonlinear transformation matrix is: the area closest to the face S j Based on the nonlinear transformation matrix, affine transformation is performed to the face area T i The position of; is the face area after affine transformation Use the multi-resolution fusion method to create a multi-layer mask M; transform the face area after affine transformation With face area T i Perform multi-resolution fusion; replace the fused face area R to the corresponding position of the target video.

[0071] Specifically, according to the difference calculation result in step 2, the system determines each face region T in the target video i The corresponding closest face area S j This matching process is based on the previously calculated histogram similarity measure D H (T i ,S j), select a pair of face regions with the smallest difference. After finding the matching region, the next step is to use the facial key point detection technology to i and S j Align. Key point set P T and P S Respectively represent T i and S j The key points of the alignment set P are usually the facial feature points such as eyes, nose, mouth, etc. T and P S , can be calculated from S j to T i The thin plate spline (TPS) nonlinear transformation matrix is ​​a commonly used image deformation method. Through nonlinear mapping of key points, a more natural and realistic face deformation effect can be achieved. After calculating the TPS transformation matrix, the closest face area S j Perform an affine transformation based on this transformation matrix to match the target face area T i The position and shape of the face area obtained after affine transformation It cannot be used directly for replacement, because a simple affine transformation may introduce image discontinuity and abrupt edge transition. In order to overcome these problems, Example 8 uses a multi-resolution fusion method to process the transformed face area. Specifically, the face area after affine transformation Create a multi-layer mask M, which is used to fuse images layer by layer at different resolutions to ensure smooth and natural edge transitions. Multi-resolution fusion methods usually include Laplacian pyramid or Gaussian pyramid technology, which decomposes the image into multiple resolution levels, processes and fuses each level separately, and then combines these levels. For the face area after affine transformation and the target face region T i , a mask M is applied to each resolution level for fusion processing, which not only retains the detail features but also ensures the smoothness of the overall transition. After completing the multi-resolution fusion, the fused face area R obtained has a high degree of naturalness and consistency. Finally, the fused face area R is replaced to the corresponding position of the target video to achieve the final effect of face swap. This process ensures the high quality and naturalness of the replacement effect through fine processing of multiple steps. This method based on key point alignment, nonlinear transformation and multi-resolution fusion not only improves the accuracy of face replacement, but also enhances the robustness of the system in processing complex scenes and different lighting conditions. Through this multi-level processing, the present invention can realize efficient and accurate face swapping in various complex video environments, demonstrating strong application potential and technical advantages in the fields of video editing, film and television production, etc.

[0072] According to the difference calculation results of step 2, determine each face area T in the target video i The corresponding face area S closest to the source video j . Use facial key point detection (such as Dlib's 68-point model) to detect T i and S j Align to ensure facial features are aligned. Key point set P T and P S Respectively represent T i and S j The key point position of S is calculated using the ThinPlateSpline (TPS) algorithm. j to T i The nonlinear transformation matrix T. The TPS transformation can be solved by the following formula:

[0073]

[0074] Where U(r) = r 2 log(r),P i is the control point, a1,a x ,a y and w i is the TPS parameter; (x, y) is the pixel position; solved by the least squares method:

[0075]

[0076] Among them, P i and Q i is a matching point pair, λ is a regularization parameter. j Perform TPS transformation to the target face area T i Location:

[0077]

[0078] is the face area after affine transformation Create a multi-layer mask M to ensure smooth transitions at the edges. Use the multi-resolution fusion method to create the mask:

[0079]

[0080] Among them, M l is the lth layer mask, σ l is the blur parameter of the lth layer. Using the LaplacianPyramid fusion method, the face area after affine transformation is and the target face region T i Perform multi-resolution fusion:

[0081]

[0082] Where L is the number of resolution layers, and They are the transformed image of the lth layer and the target face area. Replace the fused face area R to the corresponding position of the target video:

[0083]

[0084] Among them, T i ′(x,y) is the replaced video frame.

[0085] Example 9: Using the Laplacian pyramid fusion method, the face area after affine transformation With face area T i Perform multi-resolution fusion.

[0086] The present invention is not limited to the above-mentioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.

Claims

1. A video face swap target automatic locking algorithm based on target tracking, characterized in that: It includes: Step 1: Detect the face region of each frame in the source video and the target video, and obtain the source video face region sequence and the target video face region sequence respectively; Step 2: Using a preset difference detection model, calculate the difference between each face region in the target video face region sequence and each face region in the source video face region sequence, and take the face region in the source video face region sequence corresponding to the smallest difference as the closest face region; Step 3: Replace each face region in the target video with its closest corresponding face region; The method for detecting the face region of each frame in the source video and the target video in step 1 includes: using a pre-trained multi-scale multi-resolution model based on a regression tree to slide windows of different sizes in each frame of the source video and the target video, extracting weighted multi-scale multi-resolution values ​​in the window, and if the weighted multi-scale multi-resolution value is greater than a set threshold, judging that it belongs to the face region, and the face region of each frame of the image includes all parts covered by the window belonging to the face region; The pre-trained multi-scale and multi-resolution model based on regression tree is trained through the following process: for each pixel in each sample image in the training set, a label is corresponding to each pixel, and the label marks whether the pixel belongs to the face area; for each pixel in each sample image in the training set, a radius is set with itself as the center pixel and number of sampling points to define the neighborhood range, and within the neighborhood range, calculate the difference between the weighted grayscale value at different scales and different resolutions and the grayscale value of the central pixel to obtain a weighted multi-scale multi-resolution value; construct a weighted multi-scale multi-resolution regression tree, specifically including: taking each pixel as a node in the regression tree, and at each node, selecting a split feature and a split point so that the weighted mean square error of the weighted multi-scale multi-resolution value inside the child node after the split is minimized, thereby selecting the best split point to recursively construct the tree structure until the preset tree depth is reached.

2. The automatic target locking algorithm for video face swapping based on target tracking as claimed in claim 1, characterized in that: Step 1: Before detecting the face area of ​​each frame in the source video and the target video, first extract the video frames of the source video and the target video to obtain a source video frame set and a target video frame set respectively; then perform image denoising on each frame image in the source video frame set and the target video frame set respectively, and convert them into grayscale images.

3. The automatic target locking algorithm for video face swapping based on target tracking as claimed in claim 2, characterized in that: Weighted multi-scale multi-resolution values Use the following formula to calculate: ; in, is the gray value of the center pixel; It is Sampling points at scale and radius Gray value under ; is the number of scales; The combined weight of distance, scale and radius is calculated using the following formula: ; in, It is Sampling points at scale and radius The distance from the center pixel to the and It is a scale and radius The standard deviation of is the order of the preset polynomial.

4. The automatic target locking algorithm for video face swapping based on target tracking as claimed in claim 3 is characterized in that: Weighted mean square error of weighted multi-scale and multi-resolution values Calculated using the following formula: ; in, is the first Sample images at scale and radius The weight on is the sample size, It is The labels of sample images, is the predicted label.

5. The automatic target locking algorithm for video face swapping based on target tracking as claimed in claim 4 is characterized in that: In step 2: let the face region sequence in the target video be The face area is ; The first face region in the sequence of the source video The face area is , use the following formula to calculate the difference : ; in, Indicates the calculation of the histogram mean; is the number of buckets of the histogram; Indicates the calculation The histogram mean of the buckets.

6. The automatic target locking algorithm for video face swapping based on target tracking as claimed in claim 5, characterized in that: According to the difference calculation results of step 2, determine each face area in the target video The corresponding closest face area ; Using facial keypoint detection and Alignment; key point collection and Respectively and The key point position of arrive The TPS nonlinear transformation matrix is ​​the closest to the face area Based on the nonlinear transformation matrix, affine transformation to the face area The position of; is the face area after affine transformation Create a multi-layer mask using the multi-resolution fusion method ; The face area after affine transformation With face area Perform multi-resolution fusion; the fused face area Replace to the corresponding position of the target video.

7. The automatic target locking algorithm for video face swapping based on target tracking as claimed in claim 6, characterized in that: Use the Laplace pyramid fusion method to transform the face area after affine transformation With face area Perform multi-resolution fusion.

Citation Information

Patent Citations

  • Intelligent design system for face replacement

    CN118052723A

  • Face recognition and tracking method

    WO2018170864A1