Training method and training system of video special effect processing large model and electronic equipment
Through the training methods of freezing motion modeling and image processing modules, combined with diffusion model and unet network, the consistency of facial expressions and movements in video face change is solved, high-quality video special effects processing is achieved, and the fidelity and fluency of video face change is improved.
Patent Information
- Application Number
- CN202510398260.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing video face-changing technology has difficulties in maintaining the consistency and authenticity of faces in each frame of the video, especially in the dynamic matching of facial expressions and movements.
By freezing the video special effect processing large model trained by the motion modeling module and the image processing module, combining the diffusion model and the unet network, the pre-trained motion modeling module is used to insert it into the image processing module, superimposing parameters layer by layer, learning the priorities of facial movements and expressions, and realizing video special effect processing.
It achieves consistency of facial expressions and movements, reduces the shaking of video face change, increases the details and texture of the image, and improves the fidelity and fluency of video face change.
Smart Images

Figure CN120339044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and specifically provides a training method, a training system and an electronic device for a large model of video special effect processing. Background Art
[0002] Face swapping refers to replacing the face of A with the face of B through technical means. Face swapping technology is generally divided into two types:
[0003] Single-frame face swapping: Single-frame face swapping refers to replacing the face of A in a single picture with the face of B.
[0004] Video face swapping: Video face swapping refers to replacing the faces in a target video with the faces in another given picture. The application scenarios of face swapping technology are very extensive, for example:
[0005] Film production: In film production, face swapping technology can be used to replace the facial features of actors to better meet the needs of the roles.
[0006] Virtual reality and games: In virtual reality and games, face swapping technology can be used to create more realistic virtual characters and improve the realism and immersion of the games.
[0007] The difficulties of video face swapping are mainly reflected in the following aspects:
[0008] 1. Consistency: In a video, the faces in each frame image need to be replaced, and the replaced faces need to maintain temporal consistency with other parts of the original video.
[0009] Video face swapping requires real-time tracking of facial features, including dynamic changes in facial shapes, expressions, eyes, mouths and other parts. This requires high-precision face recognition and tracking algorithms to ensure that the replaced faces can match the facial dynamics in the original video.
[0010] 2. Authenticity: The goal of video face swapping is to create a realistic effect so that viewers cannot perceive the difference between the original face and the replaced face. This requires the algorithm to not only accurately identify and track facial features, but also perform delicate image processing and synthesis work to ensure that the result looks natural.
[0011] There is a need to develop a technology that can efficiently perform video feature processing such as face swapping, with natural face swapping effects and the ability to match the facial dynamics in the original video. Summary of the Invention
[0012] In order to overcome the above defects, the present invention provides a training method, a training system and an electronic device for a large model of video special effect processing, which can achieve smooth video processing.
[0013] In a first aspect, the present invention provides a training method for a large video special effect processing model, including:
[0014] Obtain a video to be face-swapped and a first image of the face-swapping target;
[0015] Based on the video and the first image, obtain the second facial key feature points of the face to be swapped and the first facial key feature points of the face-swapping target;
[0016] Based on the diffusion model, insert a motion modeling module into the image processing module to obtain a face-swapping model to be trained;
[0017] Based on the extracted first facial key feature points and second facial key feature points, freeze the image processing module and train the motion modeling module, and freeze the motion modeling module and train the image processing module to obtain a trained large video special effect processing model.
[0018] Further, before the step of obtaining the first facial key feature points of the face-swapping target and the second facial key feature points of the face to be swapped based on the video and the first image, the method includes:
[0019] Continuously extract frames from the video to form a plurality of video sequences, each video sequence including: a second image and a frame number.
[0020] Further, the obtaining the first facial key feature points of the face-swapping target and the second facial key feature points of the face to be swapped based on the video and the first image includes:
[0021] Perform face detection on the first image, and obtain the first facial key feature points according to the detected image;
[0022] Perform face detection on each second image respectively, and obtain the second facial key points according to the detected image.
[0023] Further, the inserting the motion modeling module into the image processing module to obtain a face-swapping model to be trained includes:
[0024] Insert the pre-trained motion modeling module into the unet network in a layer-by-layer stacking manner to obtain a face-swapping model to be trained.
[0025] Further, the freezing the image processing module and training the motion modeling module includes:
[0026] Input the continuous eigenvalue obtained according to the second image into the face-swapping model to be trained, and obtain a first random noise;
[0027] Using the continuous eigenvalue of the second image as the true value, calculate the first loss for the first random noise through a preset first loss function;
[0028] Backpropagation is performed based on the first loss to update the parameters of the face-swapping model, completing the current round of iterative training.
[0029] Further, freezing the motion modeling module and training the image processing module includes:
[0030] According to the continuous eigenvalue of the second image, the id feature to be face-swapped is obtained;
[0031] The id feature and the first facial key feature points of the first image are input into the face-swapping model to obtain a predicted ID value;
[0032] Taking the face-swapping target id feature of the extracted first image as the ground truth, the second loss is calculated for the predicted ID value through a preset second loss function;
[0033] Backpropagation is performed based on the second loss to update the parameters of the face-swapping model, completing the current round of iterative training.
[0034] Further, freezing the motion modeling module and training the image processing module includes:
[0035] Extract the id of the first image corresponding to the current frame of the second image;
[0036] Obtain the first facial key feature points of the first image corresponding to the current frame and the next frame of the second image;
[0037] The id of the first image and the obtained first facial key feature points are input into the face-swapping model to be trained to obtain a second random noise;
[0038] Taking the third random noise of the first image corresponding to the current frame of the second image as the ground truth, the third loss is calculated for the second random noise output by the face-swapping model through a preset third loss function;
[0039] Backpropagation is performed based on the third loss to update the parameters of the face-swapping model, completing the current round of iterative training.
[0040] Further, after obtaining the trained large video special effect processing model, it further includes:
[0041] Input the video to be face-swapped into the trained large video special effect processing model;
[0042] Based on the large model, a predicted face-swapping result is obtained.
[0043] In a second aspect, the present invention provides a training system for a large video special effect processing model, including:
[0044] An acquisition unit for acquiring a video to be face-swapped and a first image of a face-swapping target;
[0045] A feature point extraction unit, configured to obtain second facial key feature points to be face-swapped and first facial key feature points of a face-swapping target based on the video and the first image;
[0046] A processing unit, configured to insert a motion modeling module into an image processing module based on a diffusion model to obtain a face-swapping model to be trained;
[0047] A training unit, configured to freeze the image processing module and train the motion modeling module, and freeze the motion modeling module and train the image processing module respectively based on the extracted first facial key feature points and second facial key feature points, to obtain a trained large video special effect processing model.
[0048] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the training method of the large video special effect processing model as described in the first aspect is implemented.
[0049] One or more of the above technical solutions of the present invention have at least one or more of the following beneficial effects:
[0050] In implementing the technical solution of the present invention, by separately training the frozen motion modeling module and the image processing module in the large video special effect processing model, the present invention can not only maintain the consistency of facial expressions and movements, but also increase the details and textures of the image and reduce jitter, so as to achieve smooth video special effect processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Referring to the accompanying drawings, the disclosure of the present invention will become more understandable. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present invention. In addition, similar numbers in the figures are used to represent similar components, where:
[0052] Figure 1 is a schematic main step flow diagram of a training method of a large video special effect processing model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] Some embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention.
[0054] In the description of the present invention, a "module" and a "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various appropriate sensors, communication ports, memory, and may also include a software part, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other appropriate processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. A non-transitory computer-readable storage medium includes any appropriate medium that can store program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and so on.
[0055] The present invention provides a training method for a large video special effect processing model, referring to Figure 1 , including:
[0056] S1, obtaining a video to be face-swapped and a first image of the face-swapping target;
[0057] S2, based on the video and the first image, obtaining a second set of facial key feature points of the face to be swapped and a first set of facial key feature points of the face-swapping target;
[0058] S3, based on diffusion, inserting a motion modeling module into an image processing module to obtain a face-swapping model to be trained;
[0059] S4, based on the extracted first set of facial key feature points and second set of facial key feature points, freezing the image processing module and training the motion modeling module, and freezing the motion modeling module and training the image processing module, to obtain a trained large video special effect processing model.
[0060] The present invention first collects a video to be face-swapped for training and a first image of the face-swapping target.
[0061] For example, the video is a video of person A. The face-swapping target is person B, and the purpose is to replace person A in the video with person B, and finally form a video of person B generated according to the trajectory of person A.
[0062] Step S1 is to collect the video of person A and the image of person B (i.e., the first image).
[0063] Step S2 is to obtain a first set of facial key feature points of the face-swapping target B and a second set of facial key feature points of the person A to be face-swapped through the video of person A and the image of person B.
[0064] Step S3 is to combine a motion modeling module and an image processing module based on the framework structure of a diffusion model to form a face-swapping model to be trained.
[0065] Diffusion Model: A machine learning model used to generate data similar to the training data. It works by continuously adding Gaussian noise to corrupt the training data and then learning the inverse denoising process to recover the data. Diffusion models can be used for various tasks, such as image generation, text generation, etc. Through training, it can generate samples similar to the training data and has high application value in the field of content generation.
[0066] The image processing module can achieve the effect of single-frame face swapping.
[0067] The motion modeling module can extract the facial motion prior.
[0068] In step S4, the motion modeling module and the image processing module are trained separately. By training the motion modeling module alone, a large number of videos of various qualities can be used to learn the motion prior. And by training the image processing module alone, the details and textures of the human face can be retained, making the face swapping result more realistic.
[0069] In one embodiment, before the step of obtaining the second facial key feature points to be face-swapped and the first facial key feature points of the face-swapping target based on the video and the first image in step S2, the method includes:
[0070] Continuously extract frames from the video to form a plurality of video sequences, each video sequence including: a second image and a frame number.
[0071] Extract frames from the video of person A to form several video sequences, including consecutive picture frames of person A, i.e., the second image, and frame numbers that can reflect the frame rate. In this way, the pictures and frame numbers of person A are obtained, and the frame numbers are used to reflect the time sequence.
[0072] Next, the collected pictures of person A can be further preprocessed, including: face detection (locating the face area through a face detection box to obtain the relative position of the face in the picture), extraction of facial key feature points landmarks, alignment, and cropping, etc., in order to facilitate subsequent feature extraction and model training.
[0073] Face detection is used to detect the position of the face and output a detection box.
[0074] Extract facial key feature points through an existing model, such as dense or ordinary, etc. The more points, the finer but the greater the computational cost.
[0075] The alignment can be achieved by rectification.
[0076] Cropping is to cut off the part outside the face detection box.
[0077] Finally, the position of the face in the picture and the spatio-temporal relationship of the face change are obtained.
[0078] The data preprocessing stage can be implemented through a preprocessing module. First, the video is extracted into a frame sequence, and then the preparatory work of face detection, clipping, scaling, alignment, and face key point detection is completed. At the same time, the face image dataset is processed to complete the preparatory work of face detection, clipping, scaling, alignment, and face key point detection.
[0079] In one embodiment, in step S2, the obtaining of the second facial key feature points to be face-swapped and the first facial key feature points of the face-swapping target based on the video and the first image includes:
[0080] S21, perform face detection on the first image, and obtain the first facial key feature points according to the detected image.
[0081] S22, perform face detection on each second image respectively, and obtain the second facial key points according to the detected image.
[0082] Perform face detection on the image of person B to obtain the first facial key feature points. Perform face detection on the image of person A to obtain the second facial key feature points.
[0083] In one embodiment, in step S3, the inserting of the motion modeling module into the image processing module to obtain the face-swapping model to be trained includes:
[0084] S31, in a way of stacking layer by layer, insert the pre-trained motion modeling module into the unet network to obtain the face-swapping model to be trained.
[0085] The image processing module can select the unet network.
[0086] Based on the diffusion model framework, select the backbone network of the unet network (image processing module). Insert the newly added motion modeling module into the backbone network and combine it with the unet network. The motion modeling module and the unet backbone network perform parameter integration in a way of stacking layer by layer. After the two modules are combined, they have the ability to process videos.
[0087] The obtaining process of the pre-trained motion modeling module can be: copy the unet network and retrain it, or select the existing vit or TRAFOMER network.
[0088] The pre-training method of the motion modeling module can be: add a Facial Expression Coding System, construct a correlation coefficient matrix of different facial motion units, and by observing and recording the details of facial movements, the emotional state, the intensity of emotional expression, and other related features can be inferred, and further accurately learn the dynamic changes of the face.
[0089] In one embodiment, in step S4, the process of freezing the image processing module and training the motion modeling module includes:
[0090] S411: Input the continuous eigenvalue obtained from the second image into the face-swapping model to be trained. After being encoded as inputlatent by the vae encoder in the diffusion model of the face-swapping model, and then through the forward diffusion step, Gaussian noise is gradually added to the inputlatent until the data becomes the first random noise.
[0091] S412: Using the continuous eigenvalue of the second image as the ground truth, calculate the first loss for the first random noise through a preset first loss function.
[0092] S413: Based on the first loss, perform backpropagation to update the parameters of the face-swapping model and complete the training of the current round of iteration.
[0093] Model the features of the face image of person A using an existing clip encoder application to obtain continuous eigenvalues.
[0094] Input the continuous eigenvalues into the combined model in step S3 to obtain the diffused data with motion parameters added. This realizes inserting the initialized motion modeling module into the frozen unet model.
[0095] The ground truth is the continuous eigenvalue of the face image of person A, the predicted value is the first random noise, calculate the loss, and adjust the hyperparameters. Finally, the training of the motion modeling module is completed.
[0096] In this embodiment, the images of person A are used in the training process, so that the facial motion law and trajectory changes of person A in the video can be obtained.
[0097] In one embodiment, in step S4, the process of freezing the motion modeling module and training the image processing module includes:
[0098] S421: Obtain the id feature to be face-swapped according to the continuous eigenvalue of the second image. Obtain the id feature of person A according to the continuous eigenvalue of the image of person A.
[0099] S422: Input the id feature of person A and the first facial key feature points of person B in the first image into the face-swapping model to obtain the predicted ID value.
[0100] Combine the id feature of person A and the first facial key feature points of person B, and use it as the conditional input of the face-swapping model into the face-swapping model. Then, through the vae decoder to obtain the model and get the predicted ID value.
[0101] S423. Using the face-swapping target ID feature of the first image person B extracted as the ground truth, calculate the second loss, i.e., the ID loss, for the predicted ID value through a preset second loss function.
[0102] S424. Based on the second loss, perform backpropagation to update the parameters of the face-swapping model and complete the current round of iterative training. And so on, calculate the loss of each frame of the image and adjust the hyperparameters.
[0103] Adjust the model hyperparameters through the second loss to make the similarity between the actually replaced person image and the target replaced person B in each frame of the image.
[0104] The ID feature extraction can also train the VAE model separately for appearance feature extraction.
[0105] In one embodiment, in step S4, the freezing of the motion modeling module and the training of the image processing module include:
[0106] S425. Extract the ID of the first image corresponding to the current frame of the second image.
[0107] The first image person B extracts its ID feature through the ID encoder.
[0108] S426. Obtain the first facial key feature points of the first image corresponding to the current frame and the next frame of the second image.
[0109] S427. Input the ID of the first image corresponding to the current frame and the first facial key feature points of the first image corresponding to the current frame and the next frame of the second image into the face-swapping model to be trained to obtain the second random noise.
[0110] The ID and the first facial key feature points are input into the model trained in step S424, and the predicted z_0 noise is taken as the predicted value, denoted as the z_0 noise at the initial moment, and predicted frame by frame.
[0111] S428. Using the third random noise of the first image corresponding to the current frame of the second image as the ground truth, calculate the third loss for the second random noise feature output by the face-swapping model through a preset third loss function, which is the MSE loss.
[0112] S429. Based on the third loss, perform backpropagation to update the parameters of the face-swapping model and complete the current round of iterative training.
[0113] Exemplary illustration. In an application scenario, for example, before face swapping, in the video of person A, there are 4 frames of pictures, which are the 1st frame, the 2nd frame, the 3rd frame, and the 4th frame in chronological order. After face swapping, the 1st frame becomes the 1'st frame, and the 2'nd frame, 3'rd frame, and 4'rd frame. The 1'st frame, 2'nd frame, 3'rd frame, and 4'rd frame are all images of person B.
[0114] S427 is to input the first facial key feature points of person B in the 1'st frame and 2'nd frame and the id feature of person B in the 1'st frame into the model trained in S424. The predicted z_0 noise obtained is the predicted value, and the true value is the third random noise of the 1'st frame. The acquisition method of the third random noise is as follows: Encode the image of the 1'st frame into inputlatent through the vaeencoder in the diffusion model of the face swapping model. Then, through the forward diffusion step, gradually add Gaussian noise to the inputlatent until the data becomes the third random noise of the image of the 1'st frame.
[0115] And so on, input the first facial key feature points of person B in the 2'nd frame and 3'rd frame and the id feature of person B in the 2'nd frame into the model trained in S424. The predicted z_1 noise obtained is the predicted value, and the true value is the third random noise of the 2'nd frame. The acquisition method of the third random noise is to encode the image of the 2'nd frame into input latent through the vae encoder in the diffusion model of the face swapping model. Then, through the forward diffusion step, gradually add Gaussian noise to the inputlatent until the data becomes the third random noise of the image of the 2'nd frame.
[0116] The present invention judges the approximation degree between two consecutive frames through the third loss, making the change of person B in the video more natural and smooth.
[0117] The above-trained large model can be used for video special effect processing. In the application stage, after obtaining the trained large model for video special effect processing, the present invention further includes:
[0118] Input the video to be face-swapped into the trained large model for video special effect processing;
[0119] Based on the large model, obtain the predicted face-swapping result.
[0120] The video to be face-swapped here is the video to be face-swapped in the actual application process.
[0121] The video to be face-swapped in step S1 is the video pre-collected for training the model during the training process of the large model.
[0122] The present invention provides a training system for a large model for video special effect processing, including:
[0123] An acquisition unit, used for acquiring the video to be face-swapped and the first image of the face-swapped target;
[0124] A feature point extraction unit, used for obtaining the second key feature points of the face to be replaced and the first key feature points of the face replacement target based on the video and the first image;
[0125] A processing unit, used for inserting a motion modeling module into an image processing module based on a diffusion model to obtain a face-changing model to be trained;
[0126] The training unit is used to freeze the image processing module and train the motion modeling module and freeze the motion modeling module and train the image processing module based on the extracted first facial key feature points and the second facial key feature points, so as to obtain a trained video special effects processing large model.
[0127] Advantages of the present invention:
[0128] 1. Reduce jitter: Reduce the jitter caused by face swapping in videos.
[0129] 2. Increase details and textures: The clarity of the face images in the commonly used face swapping algorithms is low, and additional face super-resolution algorithms are usually required for post-processing after face swapping. However, the current face super-resolution often smooths out the facial details, and the processing results are unnatural. The present invention can ignore the image quality in the training motion modeling module stage, so a large amount of video data can be used to learn facial movement priors. However, high-definition images and videos are used in the image processing module training stage to directly generate high-definition face images and retain the authenticity of facial textures.
[0130] 3. Maintaining the consistency of facial expressions and movements: Facial expressions and movements are one of the key factors affecting the realism of video face swapping. In the process of video face swapping, the present invention learns the subtle movements and expression changes of the face through facial motion representation, tries to maintain the consistency of these details, and improves the realism of video face swapping.
[0131] In the actual application of the present invention, corresponding identification settings can be made according to different needs, such as adding a watermark, or clearly marking in the video that the video has been processed with a face swap. In order to ensure that the overall viewing experience of the video is not affected and at the same time achieve the purpose of identification, the added watermark is designed to be invisible to the naked eye, which not only ensures the aesthetics of the video, but also effectively informs the audience of the particularity of the video content.
[0132] The large model for video special effect processing provided by the present invention can not only perform face swapping operations on faces in videos, but also achieve various transformation processes on the faces of specific individuals. These transformations include but are not limited to "facial aging" effects, such as transforming a person's facial features into those of a middle-aged or elderly person; "facial rejuvenation" effects, for example, making facial features younger or juvenile; "transgender" effects, achieving the transformation of facial features from one gender to another; and "facial style transfer", including but not limited to transforming facial features into styles such as mature, cute, elegant, masculine, feminine or neutral. In addition, the video special effect processing of the model of the present invention also supports the application of "facial special effects", such as overall face slimming, overall beauty enhancement, enlarging eyes, raising the nose bridge, enlarging lips, etc. The implementation principles of all these technologies are based on step S3 of the present invention. Based on the diffusion model, the motion modeling module is inserted into the image processing module to obtain the face swapping model to be trained, that is, the "image processing module" and the "motion modeling module" are added. Through the collaborative work of these two modules, the conversion of the above various facial effects becomes possible.
[0133] When the present invention is in use, it strictly complies with the requirements of laws and regulations to ensure data security, follows the principles of legality, legitimacy and necessity, and processes the personal information actively provided by users during the use of products / services or generated due to the use of products / services, as well as the personal information obtained with user authorization, based on reasonable purposes of business scenarios.
[0134] During the process of providing video special effect processing, the present invention makes prominent markings according to the scenarios of services with functions of generating or significantly changing information content, to avoid public confusion or misidentification, and requires that no organization or individual shall use technical means to delete, tamper with or conceal relevant markings. It should be noted that although the above embodiments describe each step in a specific order, those skilled in the art can understand that in order to achieve the effects of the present invention, different steps do not necessarily have to be executed in such an order, and they can be executed simultaneously (in parallel) or in other orders, and these variations are all within the protection scope of the present invention.
[0135] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the training method of the large model for video special effect processing as described above. Further, it should be understood that since the setting of each module is only to illustrate the functional units of the device of the present invention, the physical devices corresponding to these modules can be the processor itself, or a part of the software in the processor, a part of the hardware, or a part combined by software and hardware. Therefore, the number of each module in the figure is only illustrative.
[0136] Those skilled in the art can understand that the various modules in the device can be adaptively split or combined. Such splitting or combining of specific modules will not cause the technical solution to deviate from the principle of the present invention. Therefore, the technical solutions after splitting or combining will all fall within the protection scope of the present invention.
[0137] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, those skilled in the art can easily understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A training method for a large model of video special effect processing, characterized in that, Including: Obtain the video to be face-swapped and the first image of the face-swapping target; Based on the video and the first image, obtain the second facial key feature points of the face to be swapped and the first facial key feature points of the face-swapping target; Based on the diffusion model, insert the motion modeling module into the image processing module to obtain the face-swapping model to be trained; Based on the extracted first facial key feature points and second facial key feature points, freeze the image processing module and train the motion modeling module, and freeze the motion modeling module and train the image processing module to obtain the trained large video special effect processing model.
2. The training method of the large model for video special effect processing according to claim 1, wherein Before the step of obtaining the first facial key feature points of the face-swapping target and the second facial key feature points of the face to be swapped based on the video and the first image, the method includes: Continuously extract frames from the video to form a plurality of video sequences, each video sequence including: a second image and a frame number.
3. The training method of the large model for video special effect processing according to claim 2, wherein, The obtaining the first facial key feature points of the face-swapping target and the second facial key feature points of the face to be swapped based on the video and the first image includes: Perform face detection on the first image, and obtain the first facial key feature points according to the detected image; Perform face detection on each second image respectively, and obtain the second facial key points according to the detected image.
4. The training method of the large model for video special effect processing according to claim 2, wherein The inserting the motion modeling module into the image processing module to obtain the face-swapping model to be trained includes: Insert the pre-trained motion modeling module into the unet network in a layer-by-layer stacking manner to obtain the face-swapping model to be trained.
5. The training method of the large model for video special effect processing according to claim 2, wherein, The freezing the image processing module and training the motion modeling module includes: Input the continuous eigenvalues obtained according to the second image into the face-swapping model to be trained, and obtain the first random noise; Using the continuous eigenvalues of the second image as the true value, calculate the first loss for the first random noise through a preset first loss function; Perform backpropagation based on the first loss to update the parameters of the face-swapping model, and complete the current round of iterative training.
6. The training method of the large model for video special effect processing according to claim 5, wherein The freezing the motion modeling module and training the image processing module includes: Obtain the id feature of the face to be swapped according to the continuous eigenvalues of the second image; Input the id feature and the first facial key feature points of the first image into the face-swapping model to obtain the predicted ID value; Using the id feature of the face-swapping target of the extracted first image as the true value, calculate the second loss for the predicted ID value through a preset second loss function; Perform backpropagation based on the second loss to update the parameters of the face-swapping model, and complete the current round of iterative training.
7. The training method of the large model for video special effect processing according to claim 5, wherein, The freezing the motion modeling module and training the image processing module includes: Extract the id of the first image corresponding to the current frame of the second image; Obtain the first facial key feature points of the first image corresponding to the current frame and the next frame of the second image; Input the id of the first image and the obtained first facial key feature points into the face-swapping model to be trained to obtain the second random noise; Using the third random noise of the first image corresponding to the current frame of the second image as the true value, calculate the third loss for the second random noise output by the face-swapping model through a preset third loss function; Perform backpropagation based on the third loss to update the parameters of the face-swapping model, and complete the current round of iterative training.
8. The training method of the large model for video special effect processing according to any one of claims 1 to 7, characterized in that, After obtaining the trained large video special effect processing model, it further includes: Input the video to be face-swapped into the trained large video special effect processing model; Based on the large model, obtain the predicted face-swapping result.
9. A training system for a large model of video special effect processing, characterized in that, It includes: An acquisition unit for acquiring the video to be face-swapped and the first image of the face-swapping target; A feature point extraction unit for obtaining the second facial key feature points of the face to be swapped and the first facial key feature points of the face-swapping target based on the video and the first image; A processing unit for inserting a motion modeling module into an image processing module based on a diffusion model to obtain a face-swapping model to be trained; A training unit for respectively freezing the image processing module and training the motion modeling module and freezing the motion modeling module and training the image processing module based on the extracted first facial key feature points and second facial key feature points to obtain a trained large video special effect processing model.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the training method of the large video special effect processing model according to any one of claims 1 to 8.