Video face changing dynamic optimization method based on multi-modal fusion

Through multimodal fusion and dynamic optimization technology, the problem of insufficient utilization of face information in video face swap technology is solved, and a more natural and stable face swap effect is achieved, adapting to complex scenes and avoiding details distortion.

CN120219145AInactive Publication Date: 2025-06-27BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256170.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing video face-changing technology fails to fully utilize the information of the face in the video, resulting in poor effect of the replaced video, inconsistent face with the surrounding environment, unnatural expressions and movements, inconsistent face positions, expressions, and lighting between frames, flickering or jumping, and unnatural movement.

Method used

A dynamic optimization method for video face swap based on multimodal fusion is adopted to generate multiple data sets through data acquisition and preprocessing, and the multimodal fusion and model fusion determination coefficients are analyzed, dynamic smoothing and detail optimization are optimized, and the lighting optimization coefficients, background parameter optimization coefficients and facial expression optimization coefficients are generated for analysis and feedback.

Benefits of technology

It achieves a more natural and stable face change effect, can adapt to complex and changeable scenes, avoid unnatural or distorted face change caused by details, and ensures that the transition between the face and the background is more natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219145A_ABST
    Figure CN120219145A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of videos, and discloses a video face changing dynamic optimization method based on multi-modal fusion, which not only considers independent optimization of illumination, background and facial expression, but also provides a more comprehensive optimization scheme by fusing a plurality of data sources. Through dynamic smoothing and detail optimization, each detail can be accurately adjusted, and unnatural face changing or distortion caused by the detail problem is avoided. Finally, the face changing effect based on the method is more natural and stable, the method can adapt to complex and changeable scenes, the facial expression optimization coefficient MBX is calculated through a formula, multiple factors such as the facial key point detection precision G, the expression change amplitude H and the facial posture change angle I are comprehensively considered, and facial expression recognition and matching can be accurately optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video technology, and in particular to a video face-changing dynamic optimization method based on multimodal fusion. Background Art

[0002] Video Face Swapping is a technology that replaces the face in a video with the target face. Video Face Swapping has a wide range of applications, such as film and television production, role swapping, entertainment and social media entertainment, advertising videos, education and training, virtual reality, etc., but it also faces many technical difficulties.

[0003] The main technical difficulties of video face swapping include identity consistency. The replaced face needs to be consistent with the identity in the target video, including skin color, lighting, expression and other details. Time consistency. The time consistency between frames needs to be maintained after the face swap. Detail loss. Current videos often make faces look unreal, with overly smooth details and a loss of realism. Lighting and shadow problems. The replaced face needs to be consistent with the lighting and shadows in the target video. Movement and posture problems. The replaced face needs to be consistent with the head movement and posture in the target video.

[0004] In recent years, the field of video face swapping has developed rapidly. Many achievements have been made, such as 3DMM, Deepface, FSGAN, and diffusion models. However, the current technology does not fully utilize the information of the face in the video, which leads to poor video effects after replacement. The replaced face is not coordinated with the surrounding environment, the expression and movement are unnatural, the face position, expression, and lighting between frames are inconsistent, flickering or jumping occurs, and the movement is unnatural.

[0005] Therefore, we propose a video face-changing dynamic optimization method based on multimodal fusion to solve the above problems. Summary of the invention

[0006] The purpose of the present invention is to provide a dynamic optimization method for video face swapping based on multimodal fusion to solve the above-mentioned background technology. In recent years, the field of video face swapping has developed rapidly. There are many achievements such as 3DMM, Deepface, FSGAN, diffusion model, etc., but the current technology does not make full use of the information of the face in the video, which leads to poor effect of the replaced video, the replaced face is not coordinated with the surrounding environment, the expression and movement are unnatural, the face position, expression and lighting between frames are inconsistent, flickering or jumping phenomenon occurs, and the movement is unnatural.

[0007] To achieve the above purpose, the present invention provides the following technical solution: a video face-changing dynamic optimization method based on multimodal fusion, the specific steps are as follows:

[0008] S1. Data collection and preprocessing: comprehensively extract various information of human faces in the video and perform preprocessing to generate the first dataset, the second dataset, and the third dataset;

[0009] S2. Multimodal fusion: integrate the collected first dataset, second dataset, and third dataset, construct a model fusion determination coefficient, and analyze the model fusion determination coefficient;

[0010] S3. Dynamic smoothing and detail optimization: independently analyze the first dataset, second dataset, and third dataset to generate a lighting optimization coefficient, a background parameter optimization coefficient, and a facial expression optimization coefficient, and analyze the lighting optimization coefficient, the background parameter optimization coefficient, and the facial expression optimization coefficient;

[0011] S4. Feedback: display the analysis results of the model fusion determination coefficient, the lighting optimization coefficient, the background parameter optimization coefficient, and the facial expression optimization coefficient.

[0012] Preferably, in the step S1, the first dataset includes the lighting intensity A, the lighting uniformity B, and the lighting contrast C;

[0013] The second dataset includes the background complexity D, the background movement speed E, and the background blur F;

[0014] The third dataset includes the facial key point detection accuracy G, the facial expression change amplitude H, and the facial pose change angle I.

[0015] Preferably, in the step S2, the specific calculation method for constructing the model fusion determination coefficient is as follows:

[0016]

[0017] In the formula: a1, a2, a3, and a4 are weight values, and the values of a1, a2, a3, and a4 are adjusted and set by the user, A is the lighting intensity, B is the lighting uniformity, C is the lighting contrast, D is the background complexity, E is the background movement speed, F is the background blur, G is the facial key point detection accuracy, H is the facial expression change amplitude, and I is the facial pose change angle.

[0018] Preferably, in the step S2, the specific analysis method for the model fusion determination coefficient is as follows:

[0019] When MXR ≤ 0.7, it means that the current model can be directly used;

[0020] When MXR > 0.7, it means that the current model cannot be directly used and problem tracing is required.

[0021] Preferably, in the step S3, the light optimization coefficient is calculated by the following formula:

[0022]

[0023] In the formula: b1, b2, b3, and b4 are weight values, and the values of b1, b2, b3, and b4 are adjusted and set by the user. A is the light intensity, B is the light uniformity, and C is the light contrast.

[0024] Preferably, in the step S3, the specific analysis method of the light optimization coefficient is as follows:

[0025] When GZX ≤ 0.77, it means that there is no light problem in the current model;

[0026] When GZX > 0.77, it means that there is a light problem in the current model.

[0027] Preferably, in the step S3, the background parameter optimization coefficient is calculated by the following formula

[0028]

[0029] In the formula: b1, b2, b3, and b4 are weight values, and the values of b1, b2, b3, and b4 are adjusted and set by the user. D is the background complexity, E is the background movement speed, and F is the background blur degree.

[0030] Preferably, in the step S3, the specific analysis method of the background parameter optimization coefficient is as follows:

[0031] When BJX ≤ 0.4, it means that there is no problem with the background parameters in the current model;

[0032] When BJX > 0.4, it means that there is a problem with the background parameters in the current model.

[0033] Preferably, in the step S3, the facial expression optimization coefficient is calculated by the following formula;

[0034]

[0035] In the formula: b1, b2, b3, and b4 are weight values, and the values of b1, b2, b3, and b4 are adjusted and set by the user. G is the facial key point detection accuracy, H is the change range of the facial expression, and I is the change angle of the facial pose.

[0036] Preferably, in the step S3, the specific analysis method of the facial expression optimization coefficient is as follows:

[0037] When MBX ≤ 0.35, it means that there is no facial expression recognition problem in the current model;

[0038] When MBX > 0.35, it indicates that there are problems with facial expression recognition in the current model.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] 1. This method not only considers the independent optimization of illumination, background, and facial expressions, but also provides a more comprehensive optimization solution by fusing multiple data sources. Through dynamic smoothing and detail optimization, every detail can be precisely adjusted, avoiding unnatural or distorted face swapping caused by detail problems. Ultimately, the face swapping effect based on this method is more natural and stable and can adapt to complex and changing scenarios.

[0041] 2. By calculating the facial expression optimization coefficient MBX through a formula and comprehensively considering multiple factors such as the facial key point detection accuracy G, the amplitude of expression change H, and the angle of facial pose change I, the facial expression recognition and matching can be precisely optimized. Through this optimization, the face swapping model can better adapt to the dynamic changes of facial expressions, ensure a more natural transition of facial expressions, and avoid face swapping distortion or disharmony caused by expression changes. The optimization of facial expressions is adjusted through the weight values b1, b2, b3, and b4, enabling users to flexibly control the optimization focus of different facial features according to specific needs. For example, in some videos where the facial expression changes greatly, the amplitude of expression change H may be more critical, while in other scenarios, the facial pose change I may require more attention. Such flexibility enables the face swapping effect to be more accurate and natural under different facial expressions and poses. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flowchart of the method steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Embodiment 1: Please refer to Figure 1 , a dynamic optimization method for video face swapping based on multimodal fusion, and the specific steps are as follows:

[0045] S1. Data collection and preprocessing, comprehensively extract various information of the human face in the video and perform preprocessing to generate the first data set, the second data set, and the third data set;

[0046] S2. Multimodal fusion: Integrate the collected first dataset, second dataset, and third dataset, construct a model fusion determination coefficient, and analyze the model fusion determination coefficient;

[0047] S3. Dynamic smoothing and detail optimization: Independently analyze the first dataset, second dataset, and third dataset to generate a lighting optimization coefficient, a background parameter optimization coefficient, and a facial expression optimization coefficient, and analyze the lighting optimization coefficient, the background parameter optimization coefficient, and the facial expression optimization coefficient;

[0048] S4. Feedback: Display the analysis results of the model fusion determination coefficient, the lighting optimization coefficient, the background parameter optimization coefficient, and the facial expression optimization coefficient.

[0049] In this embodiment: By comprehensively extracting various information of the human face in the video, such as lighting, background, facial features, etc., and performing fine preprocessing on them, it is possible to ensure a reliable data basis for subsequent multimodal fusion and optimization analysis. The generated first dataset, second dataset, and third dataset respectively represent different information categories, such as lighting, background, and facial features, and these datasets provide accurate and detailed input data for the subsequent optimization process. The main purpose of this step is to ensure the quality and diversity of the data, provide high-quality raw data for subsequent analysis and optimization, and thus improve the accuracy and naturalness of the face swapping effect.

[0050] By integrating the first, second, and third datasets, a model fusion determination coefficient is constructed and analyzed, which can comprehensively consider the influences of multiple factors such as lighting, background, and facial expressions in the video. This step can fuse different data sources into a unified determination coefficient, enabling the model to adjust various parameters more precisely according to the fusion result. Through this process, the system can identify the correlations between the data and adjust the applicability and reliability of the face swapping model according to the comprehensive evaluation result, thereby improving the naturalness and accuracy of the face swapping effect.

[0051] Through the independent analysis of the datasets, a lighting optimization coefficient, a background parameter optimization coefficient, and a facial expression optimization coefficient are generated. The key to this process lies in the independent analysis of each parameter to ensure that each item can be precisely optimized. The lighting optimization coefficient can effectively solve the problems of uneven or excessive lighting, the background parameter optimization coefficient can cope with the influence of complex backgrounds, and the facial expression optimization coefficient helps to solve the problem of face swapping distortion caused by facial dynamic changes. Through the meticulous optimization of these parameters, the details of the face swapping effect are further improved, making the transition between the face and the background more natural and avoiding the situation where the overall effect is affected by detail problems.

[0052] By presenting the analysis results of the model fusion determination coefficient, the lighting optimization coefficient, the background parameter optimization coefficient, and the facial expression optimization coefficient, users can intuitively see the feedback of various optimization effects and make corresponding adjustments. This process helps developers or users understand the advantages and disadvantages of the current model and make adjustments and corrections for possible problems. Through timely feedback and presentation, the face-swapping model can be improved more effectively to ensure its adaptation to various different video scenarios and ultimately achieve a higher-quality face-swapping effect.

[0053] Existing face-swapping technologies usually focus on single-dimensional optimization, such as only paying attention to factors like facial features or lighting, while ignoring the relationships between them. This multi-modal fusion method not only considers the independent optimization of lighting, background, and facial expressions but also provides a more comprehensive optimization solution by fusing multiple data sources. Through dynamic smoothing and detail optimization, every detail can be precisely adjusted, avoiding unnatural or distorted face-swapping caused by detail problems. Ultimately, the face-swapping effect based on this method is more natural, stable, and can adapt to complex and changing scenarios.

[0054] Embodiment 2: Please refer to Figure 1 , in step S1, the first data set includes the light intensity A, the light uniformity B, and the light contrast C;

[0055] The second data set includes the background complexity D, the background movement speed E, and the background blur degree F;

[0056] The third data set includes the facial key point detection accuracy G, the facial expression change amplitude H, and the facial pose change angle I.

[0057] In this embodiment: By accurately collecting data such as the light intensity A, the uniformity B, and the contrast C, it is possible to better analyze and optimize the lighting changes in different scenarios in the video. This plays an important role in avoiding unnatural or uneven lighting during face-swapping, thus making the face-swapping effect more natural and realistic. During the multi-modal fusion process, different features of lighting, such as intensity, uniformity, and contrast, are independently extracted and optimized to ensure the lighting match between the face and the background during face-swapping and avoid visual incoordination caused by lighting differences. Lighting is a key factor affecting the face-swapping effect. Being able to precisely control the intensity, uniformity, and contrast of lighting helps ensure a natural transition of lighting between the face and the background in various different environments, reducing artifacts and unrealistic visual effects.

[0058] By collecting information such as the complexity D, motion speed E, and blur F of the background, the characteristics of the background can be dynamically evaluated and optimized to better meet the needs of face swapping in complex environments. Optimizing the complexity and motion of the background can effectively avoid background interference with the face swapping effect. The blurriness and complexity of the background often affect the realism of the face swapping. By optimizing these background parameters, the background distortion or confusion that occurs during the face swapping process can be reduced, ensuring that the fusion of the character's face and the background is more natural. For different background situations such as complex backgrounds or fast-moving backgrounds, the optimization of background parameters can ensure that the face swapping runs stably in dynamic scenes, avoiding the unsmooth connection of the face swapping effect due to excessive background changes or too fast movement.

[0059] By collecting the facial key point detection accuracy G, the recognition accuracy of facial details during the face-changing process can be improved, so that facial features can be accurately matched during face-changing, avoiding deformities or dislocations. The data analysis of the facial expression change amplitude H and the facial posture change angle I can better cope with changes in facial expressions and postures during the face-changing process, ensuring the natural and smooth expression of the characters during face-changing. The accurate capture of dynamic facial expressions and the optimization of posture changes during face-changing can greatly improve the naturalness of the face-changing effect and avoid the inharmonious visual effects of facial expressions during face-changing, especially in scenes with rapid expression or posture changes.

[0060] Example 3: Please refer to Figure 1 In step S2, the specific calculation method of constructing the model fusion determination coefficient is as follows:

[0061]

[0062] Wherein: a1, a2, a3 and a4 are weight values, and the values ​​of a1, a2, a3 and a4 are adjusted by the user, A is the light intensity, B is the light uniformity, C is the light contrast, D is the background complexity, E is the background motion speed, F is the background blur, G is the facial key point detection accuracy, H is the facial expression change amplitude, and I is the facial posture change angle.

[0063] In step S2, the specific analysis method of the model fusion determination coefficient is as follows:

[0064] When MXR≤0.7, it means that the current model can be used directly;

[0065] When MXR>0.7, it means that the current model cannot be used directly and problem tracing is required.

[0066] In this embodiment, by integrating multiple factors such as illumination, background, and facial features into a determination coefficient MXR, a comprehensive evaluation of the video face-swapping model can be achieved. Each factor, such as illumination, background complexity, and facial features, is weighted in the model by different weights a1, a2, a3, a4 to ensure that the influence of each factor on the final face-swapping effect can be appropriately reflected. This enables the model to be optimized based on the comprehensive analysis results of multi-dimensional data, avoiding unnatural face-swapping effects caused by the influence of a single factor.

[0067] By allowing the user to adjust the weight values a1, a2, a3, a4, the user can flexibly adjust the degree of importance the model attaches to different data sources according to specific application scenarios or requirements. This adjustability enables the system to adaptively adjust parameters according to different video scenarios, such as scenes with strong illumination or complex backgrounds, improving the accuracy and naturalness of the face-swapping effect.

[0068] Through the analysis of MXR, when MXR ≤ 0.7, it indicates that the current model is reliable enough to be directly used. This analysis method can quickly screen out models with excellent performance, avoiding unnecessary additional adjustments and debugging, and improving efficiency.

[0069] When MXR > 0.7, it means that the current model does not meet the expectations and requires problem tracing and further optimization. This mechanism can help developers quickly identify the reasons for poor face-swapping effects, such as illumination problems, background interference, or mismatched facial expressions, and make targeted corrections to the model to ensure the accuracy and naturalness of the final result.

[0070] The construction of MXR takes into account the interactive effects of multiple datasets, including factors such as illumination, background, and facial expressions, which makes the model perform more stably in a changing environment. For example, in situations with uneven illumination, complex backgrounds, or drastic changes in facial expressions, the system can optimize the model by adjusting the weights to ensure the stability of the face-swapping effect. With the real-time analysis of the model fusion determination coefficient, the system can continuously monitor various parameters during the face-swapping process and make dynamic adjustments according to the analysis results. Through this feedback mechanism, the performance degradation of the model in complex scenarios can be effectively avoided, and the applicability and robustness of the face-swapping technology in multiple scenarios can be improved.

[0071] The user can automatically adjust the parameter settings of the model according to the feedback results of the MXR determination coefficient without manual intervention or complex debugging processes. This automated adjustment mechanism can greatly improve the application efficiency of the face-swapping technology, reduce manual intervention, and can quickly adjust and optimize in various complex environments. By setting the weights for each data dimension, such as illumination, background, and facial features, the system can more precisely optimize the face-swapping effect, avoid common distortion or unnatural phenomena, and improve the quality of video face-swapping.

[0072] Example 4: Please refer to Figure 1 In step S3, the illumination optimization coefficient is calculated by the following formula:

[0073]

[0074] Wherein: b1, b2, b3 and b4 are weight values, and the values ​​of b1, b2, b3 and b4 are adjusted and set by the user, A is the light intensity, B is the light uniformity, and C is the light contrast.

[0075] In step S3, the specific analysis method of the illumination optimization coefficient is as follows:

[0076] When GZX≤0.77, it means that the current model has no lighting problems;

[0077] When GZX>0.77, it means that there is a lighting problem with the current model.

[0078] In this embodiment, the illumination optimization coefficient GZX is calculated by a formula, and key parameters such as illumination intensity A, uniformity B, and contrast C are weighted by different weights, so as to accurately optimize the illumination. This process can effectively avoid unnatural effects caused by illumination problems during the face-changing process, such as mismatch between the face and the background illumination, or too strong or too weak illumination.

[0079] The weight values ​​b1, b2, b3, and b4 can be adjusted by the user to flexibly adjust the focus of lighting optimization according to the needs of different scenes. For example, in some scenes, light intensity is more important, while in other cases, optimization of lighting uniformity or contrast may be more critical. In this way, users can make customized adjustments based on the actual lighting environment of the video to ensure a more natural face-changing effect.

[0080] By analyzing GZX, the lighting performance of the current face-changing model can be automatically evaluated. If GZX ≤ 0.77, it means that the current model has no lighting problems and the face-changing effect is good. This automated lighting problem detection can reduce manual intervention and make the system more intelligent and efficient.

[0081] When GZX>0.77, it means that there is a lighting problem in the model. This analysis mechanism can provide timely feedback to users, indicating that the lighting optimization needs to be further adjusted to avoid distortion or unnatural face-swapping effects due to lighting differences. For example, if the lighting is too strong or too weak, the model may appear inconsistent with the background, affecting the overall effect. Through this mechanism, this situation can be avoided and the face-swapping effect can be adjusted and optimized in time.

[0082] Through the calculation and analysis of the light optimization coefficient GZX, the system can be made more intelligent. Without manual intervention by the user, the system can automatically detect lighting problems and make corresponding adjustments, greatly improving the automation level of the face-swapping technology. The analysis of the light optimization coefficient enables the user to avoid cumbersome manual debugging. The system automatically judges and feedbacks lighting problems based on the calculation results, reducing the workload of manual adjustment and improving the convenience of system use.

[0083] Example 5: Please refer to Figure 1 , in step S3, the background parameter optimization coefficient is calculated and obtained through the following formula

[0084]

[0085] In the formula: b1, b2, b3, and b4 are weight values, and the values of b1, b2, b3, and b4 are adjusted and set by the user. D is the background complexity, E is the background motion speed, and F is the background blur degree.

[0086] In step S3, the specific analysis method of the background parameter optimization coefficient is as follows:

[0087] When BJX ≤ 0.4, it means that there is no problem with the background parameters in the current model;

[0088] When BJX > 0.4, it means that there is a problem with the background parameters in the current model.

[0089] In this embodiment: By calculating the background parameter optimization coefficient BJX through the formula and comprehensively analyzing factors such as the background complexity D, the background motion speed E, and the background blur degree F, the fine optimization of the background features can be achieved. Through this optimization, the face-swapping technology can better adapt to complex background environments, avoid the negative impact of the background on the face-swapping effect, ensure a more natural transition between the background and the face. The optimization of the background complexity, motion speed, and blur degree is adjusted through the weight values b1, b2, b3, and b4, so that the user can flexibly adjust the influence of different background features according to the specific requirements of the video scene. For example, for a dynamic background, the optimization of the background motion speed and blur degree is more important, while for a static background, the background complexity may be the focus. This flexible control enables the system to adaptively adjust according to different background conditions, thus providing a more natural and realistic face-swapping effect.

[0090] Through the analysis of BJX, potential problems in the background can be detected in a timely manner. When BJX ≤ 0.4, it indicates that there are no problems in the background of the current model, and the face-swapping effect is natural and stable. When BJX > 0.4, it means that there are problems in the background, such as too high background complexity, too fast movement speed, or too large blur degree, which will affect the face-swapping effect. Through this real-time feedback mechanism, background problems can be quickly identified and corresponding measures can be taken to avoid the impact of background problems on the face-swapping effect. When problems occur in background parameters such as background movement or blur degree, the system can guide the user to adjust the corresponding parameters or optimize the model, thereby reducing the face-swapping distortion caused by background problems. This mechanism not only improves the naturalness of the face-swapping effect but also enhances the stability of the model in dynamic and complex backgrounds.

[0091] By calculating the background parameter optimization coefficient BJX, the system can intelligently detect and analyze background problems without manual intervention by the user. When the BJX value exceeds the preset threshold of 0.4, the system will automatically prompt and provide corresponding background optimization suggestions, reducing the complexity of manual operations. By automatically detecting background problems and feeding them back to the user, the system can achieve automatic adjustment and optimization of background features without complex manual intervention. This automated process makes the application of face-swapping technology more intelligent and convenient.

[0092] Embodiment Six: Please refer to Figure 1 , in step S3, the facial expression optimization coefficient is calculated and obtained through the following formula;

[0093]

[0094] In the formula: b1, b2, b3, and b4 are weight values, and the values of b1, b2, b3, and b4 are adjusted and set by the user, G is the facial key point detection accuracy, H is the facial expression change range, and I is the facial pose change angle.

[0095] In step S3, the specific analysis method of the facial expression optimization coefficient is as follows:

[0096] When MBX ≤ 0.35, it means that there are no facial expression recognition problems in the current model;

[0097] When MBX > 0.35, it means that there are facial expression recognition problems in the current model.

[0098] In this embodiment: By calculating the facial expression optimization coefficient MBX through the formula and comprehensively considering multiple factors such as the facial key point detection accuracy G, the expression change range H, and the facial pose change angle I, the facial expression recognition and matching can be accurately optimized. Through this optimization, the face-swapping model can better adapt to the dynamic changes of facial expressions, ensure that the transition of facial expressions is more natural, and avoid face-swapping distortion or disharmony caused by expression changes.

[0099] Facial expression optimization is adjusted by weight values ​​b1, b2, b3, and b4, allowing users to flexibly control the optimization focus of different facial features according to specific needs. For example, in some videos, when facial expressions change greatly, the expression change amplitude H may be more critical, while in other scenes, facial posture changes I may require more attention. This flexibility allows the face-changing effect to be more accurate and natural under different facial expressions and postures.

[0100] Through MBX analysis, the facial expression matching effect of the current face-changing model can be automatically determined. If MBX≤0.35, it means that the current model has no problem in facial expression recognition, and the face-changing effect is natural and stable. When MBX>0.35, it indicates that there is a problem with facial expression matching, which may be caused by inaccurate facial expression recognition or excessive facial posture changes. This mechanism can effectively identify and feedback expression problems to avoid affecting the face-changing effect due to unnatural or uncoordinated expressions.

[0101] When there is a problem with facial expression recognition, the system can promptly remind the user and provide adjustment suggestions. The user can adjust the relevant parameters of facial expression optimization based on the feedback, thereby optimizing facial expression recognition and face-changing effects, ensuring that the facial expressions after face-changing are natural and real. It should be noted that all actions to obtain signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located, and with the authorization given by the owner of the corresponding device.

[0102] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.

[0103] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A video face-changing dynamic optimization method based on multimodal fusion, characterized by: The specific steps are as follows: S1, data collection and preprocessing, comprehensively extracting various information of the face in the video and preprocessing it, so as to generate a first data set, a second data set and a third data set; S2, multimodal fusion, integrating the collected first data set, second data set and third data set, constructing a model fusion determination coefficient, and analyzing the model fusion determination coefficient; S3, dynamic smoothing and detail optimization, independently analyzing the first data set, the second data set, and the third data set to generate an illumination optimization coefficient, a background parameter optimization coefficient, and a facial expression optimization coefficient, and analyzing the illumination optimization coefficient, the background parameter optimization coefficient, and the facial expression optimization coefficient; S4, feedback, displays the analysis results of the model fusion determination coefficient, illumination optimization coefficient, background parameter optimization coefficient and facial expression optimization coefficient.

2. According to claim 1, a video face-changing dynamic optimization method based on multimodal fusion is characterized in that: In step S1, the first data set includes light intensity A, light uniformity B and light contrast C; The second data set includes background complexity D, background motion speed E, and background blur F; The third data set includes the facial key point detection accuracy G, the facial expression change amplitude H, and the facial posture change angle I.

3. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 2, characterized in that: In step S2, the specific calculation method of constructing the model fusion determination coefficient is as follows: Wherein: a1, a2, a3 and a4 are weight values, and the values ​​of a1, a2, a3 and a4 are adjusted by the user, A is the light intensity, B is the light uniformity, C is the light contrast, D is the background complexity, E is the background motion speed, F is the background blur, G is the facial key point detection accuracy, H is the facial expression change amplitude, and I is the facial posture change angle.

4. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 3, characterized in that: In step S2, the specific analysis method of the model fusion determination coefficient is as follows: When MXR≤0.7, it means that the current model can be used directly; When MXR>0.7, it means that the current model cannot be used directly and problem tracing is required.

5. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 4, characterized in that: In step S3, the illumination optimization coefficient is calculated by the following formula: Wherein: b1, b2, b3 and b4 are weight values, and the values ​​of b1, b2, b3 and b4 are adjusted and set by the user, A is the light intensity, B is the light uniformity, and C is the light contrast.

6. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 5, characterized in that: In step S3, the specific analysis method of the illumination optimization coefficient is as follows: When GZX≤0.77, it means that the current model has no lighting problems; When GZX>0.77, it means that there is a lighting problem with the current model.

7. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 6, characterized in that: In step S3, the background parameter optimization coefficient is calculated by the following formula: Wherein: b1, b2, b3 and b4 are weight values, and the values ​​of b1, b2, b3 and b4 are adjusted and set by the user, D is the background complexity, E is the background motion speed, and F is the background blur.

8. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 7, characterized in that: In step S3, the specific analysis method of the background parameter optimization coefficient is as follows: When BJX≤0.4, it means that the current model has no background parameter problem; When BJX>0.4, it means that there is a problem with the background parameters of the current model.

9. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 8, characterized in that: In step S3, the facial expression optimization coefficient is calculated by the following formula: Wherein: b1, b2, b3 and b4 are weight values, and the values ​​of b1, b2, b3 and b4 are adjusted and set by the user, G is the detection accuracy of facial key points, H is the amplitude of facial expression change, and I is the angle of facial posture change.

10. The method for dynamic optimization of video face swapping based on multimodal fusion according to claim 9, characterized in that: In step S3, the specific analysis method of the facial expression optimization coefficient is as follows: When MBX≤0.35, it means that the current model has no facial expression recognition problem; When MBX>0.35, it means that the current model has problems with facial expression recognition.