Super-realistic digital human video detection method and system
By extracting the personalized motion characteristics of the designated characters and combining them with the general character motion model, the problem of ignoring the physical motion characteristics and relying on a large number of real videos in the prior art is solved, and efficient and accurate fake video detection is achieved.
Patent Information
- Application Number
- CN202510297680.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art ignores the physical movement characteristics of the designated person in the forged detection, resulting in weakening of detection capabilities and requiring a large number of real videos as references, affecting the detection interaction efficiency.
Through the action sequence extraction based on the human body's three-dimensional model and the motion decoupling of the visual model, the personalized motion characteristics of the designated characters are extracted, and combined with the pre-trained general character motion model, the differential characteristics in digital human videos are accurately identified to achieve efficient detection of fake videos.
It improves the accuracy and robustness of fake video detection, reduces dependence on real video reference sets, optimizes the detection interaction process, and can efficiently and accurately identify the digital human videos of designated characters in various application scenarios.
Smart Images

Figure CN120236227A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of forgery detection, and particularly to a method and system for detecting forgery of hyper-realistic digital humans. Background Art
[0002] In recent years, with the rapid development of digital human generation technology, the generation quality of digital humans has been greatly improved, but at the same time, it has also brought potential social security risks. This current situation urgently requires the development of innovative forgery detection technologies to effectively address the possible threats posed by forged videos to society and individuals.
[0003] The core goal of forgery detection is to ensure the authenticity and credibility of video content through multi-disciplinary technical means such as artificial intelligence, computer vision, and graphics, and to curb the adverse effects of forged videos on personal privacy, social order, and institutional credibility.
[0004] Currently, the forgery detection methods for specific persons mainly rely on extracting temporal semantic features to detect the high-level semantic features that are difficult to reproduce in forged videos. These methods usually focus on the matching of speech and lip movement consistency, while ignoring the uniqueness of the specific person's body movements. For example, certain habitual head movements or specified body behaviors can often more accurately distinguish between real and forged content.
[0005] To improve the detection effect, the key lies in efficiently extracting the high-level features of human body movements in videos. In existing deep learning models, the Vision Transformer network (ViT), as a model based on the Transformer architecture, has obvious advantages in capturing high-level semantic and detailed features compared to traditional convolutional neural networks (CNNs) due to its ability to process image data using the self-attention mechanism. However, existing ViT-based forgery detection methods still focus on identifying low-frequency forgery traces. With the widespread application of diffusion models in the generation field, these forged videos have gradually gotten rid of the feature limitations of low-frequency forgery information, resulting in a significant weakening of the detection ability of traditional methods. Therefore, exploring detection methods focusing on the high-level semantic features of videos has become the core direction of hyper-realistic digital human forgery detection technology.
[0006] In summary, the following disadvantages exist in the prior art:
[0007] Existing methods for detecting forgery of specific persons mainly detect forgery of specific persons through the movement patterns of the face, the inconsistency between the mouth and spoken language factors, and the inconsistency between the inner and outer faces. These methods mainly rely on visual information, audio information, and some even require supplementary classification information, while ignoring the most unique body movement information of the specific person. Moreover, such methods often require several hours of real videos as a reference, which causes great inconvenience to the detection interaction.
[0008] Based on the above analysis, it is urgent to research and develop a set of hyper-realistic digital human forgery detection methods for specific individuals. The method of the present invention focuses on mining and analyzing the unique high-level semantic features of specific individuals, thereby effectively improving the accuracy and robustness of forged video detection, and providing solid technical support for ensuring the authenticity of videos. Summary of the Invention
[0009] The main technical problems to be solved by the present invention include: how to accurately extract the personalized motion features of specific individuals; how to distinguish the authenticity of the detected video by learning a small-scale motion reference set of specific individuals. A hyper-realistic digital human forgery detection method is proposed. Through a decoupling module, the personalized motion features of specific individuals can be learned, so as to achieve the effect of detecting hyper-realistic digital humans of specific individuals, and has strong generalization.
[0010] To solve the above problems, the present invention proposes a hyper-realistic digital human forgery detection for specific individuals, and proposes a method for accurately identifying the differences between the motion features of digital human videos and real human motion features by learning the unique motion patterns of specific individuals and combining a pre-trained general human motion model, so as to achieve efficient detection of forged videos.
[0011] In a first aspect, an embodiment of the present application provides a hyper-realistic digital human video detection method, and the method includes:
[0012] For a specific individual, based on a human body three-dimensional model, an action sequence is extracted, and based on a visual model, motion decoupling is performed to extract the personalized motion features of the specific individual;
[0013] By extracting the fusion of the personalized motion features and general motion features of the person in the video to be detected, personalized motion perception of the specific individual is performed;
[0014] A binary classification model is pre-trained to enable the binary classification model to obtain the ability to perceive the personalized motion of the specific individual, and the video to be detected for the specific individual is input to complete forgery detection.
[0015] In a specific embodiment of the present invention, the above-mentioned motion decoupling based on the visual model to extract the personalized motion features of the specific individual includes the steps of:
[0016] A self-supervised learning framework is used to train a proprietary motion encoder,
[0017] Among them, the general motion encoder identifies general motion features through the general human motion representation ability established in the pre-training stage; the general motion encoder adopts a Transformer architecture and learns the basic laws of human motion through self-supervised tasks such as human motion sequence reconstruction and motion prediction; the proprietary motion encoder identifies personalized motion features through a feature decoupling mechanism.
[0018] In a specific embodiment of the present invention, for the above-mentioned motion decoupling based on a visual model to extract the personalized motion features of a specified person, the steps further include:
[0019] Input multiple motion sequences of the specified person, and extract the corresponding proprietary motion features and general motion features through a proprietary motion encoder and a general motion encoder respectively;
[0020] Cross-exchange the proprietary motion features and general motion features of the multiple extracted motion sequences to generate multiple reconstructed motion features;
[0021] Input the multiple reconstructed motion features into the general motion encoder to remove the interference of general features, and extract the individualized motion features through the proprietary motion encoder.
[0022] In a specific embodiment of the present invention, for the above-mentioned personalized motion perception of a specified person by fusing the personalized motion features and general motion features of the person in the video to be detected, the steps include:
[0023] Convert the video to be detected into motion sequence features, and extract the proprietary features and general features of the motion sequence features of the video to be detected through a proprietary motion encoder and a general motion encoder respectively;
[0024] Concatenate the proprietary features and general features of the motion sequence features of the video to be detected into an overall motion representation;
[0025] Take the extracted personalized motion features of the specified person as reference motion features. If it is necessary to detect whether a video of a specified person is generated by a hyper-realistic digital human generation technology, first, a real video of the specified person with a certain duration is required. The real video of the specified person with a certain duration is defined as a reference set, reference set in academia. Therefore, a personalized motion decoupling model for the specified person is trained through the reference set, and the personalized motion features of the specified person are learned, that is, the reference motion features mentioned here. As for the general motion features and specific motion features that appear in the detection stage, they refer to the corresponding motion features of the video to be determined extracted through a general motion decoder and a specific motion decoder, and then feature interaction is performed and input into a classifier for detection. The overall motion representation of the video to be identified and the reference motion features are input together for feature interaction;
[0026] Through feature interaction, the difference between the input overall motion features and the reference motion features is amplified to highlight the motion signals of non-specified persons.
[0027] In a specific embodiment of the present invention, for the above-mentioned pre-trained binary classification model, the binary classification model is enabled to obtain the ability to perceive the personalized movement of a specified person. By inputting the video to be detected for the specified person, forgery detection is completed. The steps include:
[0028] Input the overall motion features after feature interaction into the binary classification model to complete the true / false binary classification task and realize the personalized perception of the hyper-realistic digital human video.
[0029] In a second aspect, an embodiment of the present application provides a hyper-realistic digital human video detection system, which adopts the above-mentioned hyper-realistic digital human video detection method. The system includes:
[0030] Motion decoupling module: For a specified person, based on the human body three-dimensional model, action sequences are extracted, and motion decoupling is performed based on the visual model to extract the personalized motion features of the specified person;
[0031] Personalized motion perception module: By fusing the personalized motion features and general motion features of the person in the video to be detected, personalized motion perception of the specified person is performed;
[0032] Video detection module: Pre-train a binary classification model to enable the binary classification model to obtain the ability to perceive the personalized motion of a specified person. Input the video to be detected for the specified person to complete forgery detection.
[0033] In a specific embodiment of the present invention, the above-mentioned motion decoupling module includes:
[0034] Use a self-supervised learning framework to train a proprietary motion encoder and complete the pre-training of the general motion encoder;
[0035] Input multiple motion sequences of a specified person, and respectively extract the corresponding proprietary motion features and general motion features through the proprietary motion encoder and the general motion encoder;
[0036] Cross-exchange the proprietary motion features and general motion features of the multiple extracted motion sequences to generate multiple reconstructed motion features;
[0037] Input the multiple reconstructed motion features into the general motion encoder to remove the interference of general features, and extract individualized motion features through the proprietary motion encoder.
[0038] In a specific embodiment of the present invention, the above-mentioned personalized motion perception module includes:
[0039] Convert the video to be detected into motion sequence features, and respectively extract the proprietary features and general features of the motion sequence features of the video to be detected through the proprietary motion encoder and the general motion encoder;
[0040] The proprietary features and general features of the motion sequence features of the video to be detected are spliced into an overall motion representation;
[0041] The extracted personalized motion features of the specified person are used as reference motion features, and the overall motion representation of the video to be authenticated and the reference motion features are input into the feature interaction module; wherein, the feature interaction module realizes feature interaction through the cross-attention mechanism;
[0042] Through the feature interaction module, the difference between the input overall motion features and the reference motion features is amplified to highlight the motion signals of non-specified persons.
[0043] Thirdly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned hyper-realistic digital human video detection method are realized.
[0044] Fourthly, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the hyper-realistic digital human video detection method as described above are realized.
[0045] Compared with the related prior art, it has the following outstanding beneficial effects:
[0046] The method of the present invention can successfully extract the unique action features of the specified person through the action decoupling and feature fusion operations on the real data set of the specified person, and combine them with the pre-trained general action features. In this way, the model can accurately perceive the personalized motion pattern of the specified person and is highly sensitive to abnormal motion signals, so as to effectively identify the authenticity of the hyper-realistic digital human video of the specified person. Compared with the traditional forgery detection methods, the present invention is no longer limited to the detection of single features such as faces and voices, but fully explores the uniqueness of the limb movements of the specified person, greatly improving the accuracy and reliability of detection. At the same time, the present invention reduces the dependence on a large-scale real video reference set, optimizes the detection interaction process, and can realize the efficient and accurate identification and detection of the digital human video of the specified person in various application scenarios, has broad application prospects and important practical significance, and provides a strong guarantee for maintaining the authenticity and security of video content. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0048] Figure 1 It is a schematic diagram of the hyper-realistic digital human video detection method of the present invention;
[0049] Figure 2 Schematic diagram of the efficient decoupling extraction method in the embodiment of the present invention;
[0050] Figure 3 Schematic diagram of the specified human body personalized motion perception method in the embodiment of the present invention;
[0051] Figure 4 Schematic diagram of the hyper-realistic digital human video detection system of the present invention;
[0052] Figure 5 Schematic diagram of the computer hardware of the present invention. Detailed implementation manners
[0053] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or a similar expression thereof means any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0054] It should also be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0055] It should also be understood that in various embodiments of the present invention, the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0056] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0057] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0058] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist physically separately for each unit, or two or more units may be integrated in one unit.
[0059] If the above function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0060] To make the above features and effects of the present invention more clearly understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0061] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0062] The method of the present invention aims at the detection of hyper-realistic digital human forgery for a specified person, and proposes a method for accurately identifying the differences in the motion characteristics of a real person in a digital human video by learning the unique motion pattern of the specified person and combining a pre-trained general human motion model, so as to achieve efficient detection of forged videos.
[0063] The present invention proposes a method for detecting the forgery of hyper-realistic digital humans. Its purpose is to extract the unique motion features of a specified person from the provided real dataset of the specified person, and then combine and infer with the pre-trained motion features to achieve the effect of detecting hyper-realistic digital human videos.
[0064] In the formation process of the present invention, the following main technical points emerged:
[0065] 1. Decoupling the motion of the specified person. The present invention extracts the action sequence based on SMPL-X (Skinned Multi-Person Linear Model - eXpressive) and decouples the motion based on the Vision Transformer model (ViT) to extract the personalized motion features of the specified person.
[0066] 2. Perceiving the personalized motion of the specified human body. The present invention is developed based on feature fusion. By fusing the personalized motion features of the specified person extracted with the general motion features and training a binary classification model, the model is enabled to have the ability to perceive the personalized motion of the specified person, achieving the purpose of detecting digital human forgery.
[0067] The following elaborates on the method of the embodiments of the present application in detail in combination with specific embodiments:
[0068] Embodiment 1
[0069] As Figure 1 shown, the embodiments of the present application provide a method for detecting hyper-realistic digital human videos, and the method includes:
[0070] Step 101, for a specified person, extract the action sequence based on the human body three-dimensional model, decouple the motion based on the vision model, and extract the personalized motion features of the specified person;
[0071] Among them, in the specific embodiments of the present invention, for the decoupling of the motion of the specified human body, the 3D mesh data of the human body is extracted from the video, and the mouth shape, full body posture, and hand actions of the person are saved with SMPL-X parameters, and the corresponding action sequence is extracted as the input for model training. The present invention separates the general human motion features and personalized proprietary motion features through decoupling operations, and proposes a scheme for extracting the motion features of the specified person based on self-supervised learning. The general motion encoder is used to extract general features and fix their weights, while the proprietary motion encoder is trained to extract personalized features. By cross-exchanging the general and proprietary features of the two motion sequences of the specified person and reconstructing the original motion sequence, the model learns personalized motion features during the optimization process of the reconstruction error while maintaining the stability of the general features. The present invention realizes the efficient decoupling of feature representation, accurately captures the personalized motion features of the specified person, and provides strong technical support for the research and forgery detection of hyper-realistic digital humans.
[0072] Step 102: Perform personalized motion perception of a specified person by extracting and fusing the personalized motion features and general motion features of the video to be detected.
[0073] In a specific embodiment of the present invention, the personalized motion perception of a specified human body is specifically realized by decoupling the personalized motion features and general motion features and combining the reference proprietary features of the specified person to achieve personalized perception of a hyper-realistic digital human video. The video to be identified is first extracted as motion sequence features. After extracting the proprietary and general features through a proprietary motion encoder and a general motion encoder and splicing them into an overall representation, they are input into a feature interaction module together with the reference proprietary features of the specified person. The feature interaction module refers to a multi-modal feature fusion architecture based on difference perception, and its core lies in establishing a dynamic association mapping between the features of the video to be identified and the reference features. This module realizes feature interaction through a cross-attention mechanism: First, the reference proprietary features are used as Key-Value pairs, and the spliced features (proprietary + general) of the video to be identified are used as Query to input into the multi-head attention layer to calculate the cross-feature similarity weights. Then, a gated residual connection mechanism is adopted to enhance the difference between the original features to be identified and the attention-weighted reference features. Specifically, a learnable gated network is used to generate a difference mask, and the dimensions that deviate greatly from the reference features in the attention output are non-linearly amplified using the Sigmoid activation, while the consistent dimensions are suppressed. The finally output interaction features not only retain the general motion laws but also strengthen the personalized feature difference signals.
[0074] The interaction module highlights the motion signals of non-specified persons by amplifying the differences between the input features and the reference features, and sends the interaction features into a classifier. The classifier is a binary classification decision network based on a multi-layer perceptron, which includes three fully connected layers: the first layer reduces the interaction features to 256 dimensions and applies the ReLU activation; the middle layer is further compressed to 64 dimensions, and LeakyReLU is used to prevent gradient disappearance; the last layer outputs the forgery probability in the range of 0-1 through the Sigmoid function. After receiving the 512-dimensional feature vector output by the interaction module, the classifier captures the discriminative patterns of the difference signals through layer-by-layer non-linear transformation to complete the true-false binary classification task. This solution effectively fuses multi-level motion features, accurately identifies forged content through a differential amplification mechanism, and provides an efficient and reliable ability to perceive abnormal motions of specified persons.
[0075] Step 103: Pre-train a binary classification model to enable the binary classification model to obtain the ability to perceive the personalized motion of a specified person, input the video to be detected for the specified person, and complete forgery detection.
[0076] In a specific embodiment of the present invention, in the above step 101, motion decoupling is performed based on a vision model to extract the personalized motion features of the specified person, including:
[0077] The proprietary motion encoder is trained using a self-supervised learning framework (the general motion encoder is pre-trained); among them, in the self-supervised learning framework, the general motion encoder identifies general motion features through the general human motion representation ability established in the pre-training stage. This encoder is an encoder pre-trained based on the large-scale human motion dataset Human3.6M, adopting a Transformer architecture, and learning the basic laws of human motion (such as joint kinematic constraints, lip movement laws, motion periodicity, etc.) through self-supervised tasks such as human motion sequence reconstruction and motion prediction. Its parameters remain frozen during subsequent training processes, and through forward propagation, it directly extracts general motion patterns independent of individuals in the input motion sequence, including common features such as basic gait laws and standard limb movement ranges.
[0078] The proprietary motion encoder identifies personalized motion features through a feature decoupling mechanism. In the self-supervised training stage, this encoder receives two motion sequence inputs from the same specified person, adopting a dynamic feature separation strategy: first, it extracts high-frequency detailed features in the motion sequence through a learnable multi-layer temporal attention network, and then constrains its orthogonality with the general motion features through a contrast loss function. During the feature cross-exchange reconstruction process, the proprietary encoder is forced to focus on the motion residual information that cannot be explained by the general features (such as specific head shaking frequencies, small habitual hand movements, etc.). By iteratively optimizing the motion sequence reconstruction error, this encoder gradually converges to the ability to extract personalized motion features directionally while avoiding representation overlap with the general motion features.
[0079] Input multiple motion sequences of a specified person, and respectively extract the corresponding proprietary motion features and general motion features through the proprietary motion encoder and the general motion encoder;
[0080] Cross-exchange the proprietary motion features and general motion features of the extracted multiple motion sequences to generate multiple reconstructed motion features;
[0081] Input the multiple reconstructed motion features into the general motion encoder to remove the interference of general features, and extract individualized motion features through the proprietary motion encoder.
[0082] In the specific embodiments of the present invention, the decoupling of the specified person's movement specifically aims to deeply explore the personalized movement preferences of the specified person. The present invention proposes a technical solution to separate the general human movement characteristics and personalized proprietary movement characteristics through decoupling operations. Specifically, the present invention designs two encoders respectively for extracting movement characteristics at different levels: the general movement encoder is used to extract general movement characteristics. The general movement encoder is an encoder pre-trained based on the human movement dataset Human3.6M, adopting the Transformer architecture, and learning the basic laws of human movement (such as joint kinematic constraints, lip movement laws, movement periodicity, etc.) through self-supervised tasks such as human movement sequence reconstruction and movement prediction. Its parameters are kept frozen during the subsequent training process, and the general movement patterns independent of individuals in the input movement sequence are directly extracted through forward propagation, including common features such as basic gait laws and standard limb movement ranges.
[0083] The proprietary movement encoder is used to extract the proprietary movement characteristics of the specified person. The proprietary movement encoder identifies personalized movement characteristics through a feature decoupling mechanism. In the self-supervised training stage, this encoder receives two movement sequence inputs from the same specified person, and adopts a dynamic feature separation strategy: first, the high-frequency detailed features in the movement sequence are extracted through a learnable multi-layer temporal attention network, and then the orthogonality between it and the general movement characteristics is constrained through a contrast loss function. In the process of feature cross-exchange reconstruction, the proprietary encoder is forced to focus on the movement residual information that cannot be explained by the general features (such as specific head shaking frequencies, small habitual hand movements, etc.). Among them, the model weights of the general movement encoder remain fixed, and the main task of the present invention is to train the proprietary movement encoder.
[0084] To achieve the efficient decoupling and extraction of general and proprietary movement characteristics, as Figure 2 shown, the present invention uses a self-supervised learning framework to train the proprietary movement encoder. The specific steps are as follows: First, two movement sequences of the specified person, namely movement sequence 1 and movement sequence 2, are input, and their corresponding proprietary movement characteristics and general movement characteristics are extracted through the proprietary movement encoder and the general movement encoder respectively. Subsequently, the proprietary and general movement characteristics of the two movement sequences are cross-exchanged. Next, the decoder D is used to decode the exchanged and spliced characteristics to reconstruct the original movement sequence, that is, the reconstructed sequence 1 and the reconstructed sequence 2 are obtained.
[0085] Through the optimization process of the reconstruction error, this self-supervised framework can effectively guide the model to learn the personalized movement characteristics of the specified person while ensuring the stability of the general movement characteristics. Combined with the pre-trained general movement encoder, the proprietary movement encoder can accurately extract individualized movement characteristics while removing the interference of general characteristics, thus supporting the in-depth analysis and modeling of the movement patterns of the specified person.
[0086] The design of the present invention not only strengthens the ability to extract the unique motion characteristics of a specified person, but also realizes the efficient decoupling of feature representation by separating general and proprietary features, providing strong technical support for the research of hyper-realistic digital humans and the field of forgery detection.
[0087] In a specific embodiment of the present invention, in step 102 above, by extracting the personalized motion features of the person in the video to be detected and fusing them with the general motion features, the personalized motion perception of the specified person is performed. The steps include:
[0088] Convert the video to be detected into motion sequence features, and respectively extract the proprietary features and general features of the motion sequence features of the video to be detected through a proprietary motion encoder and a general motion encoder;
[0089] Concatenate the proprietary features and general features of the motion sequence features of the video to be detected into an overall motion representation;
[0090] Use the extracted personalized motion features of the specified person as reference motion features, and input the overall motion representation of the video to be authenticated and the reference motion features together for feature interaction;
[0091] Through feature interaction, amplify the difference between the input overall motion features and the reference motion features, highlighting the motion signals of non-specified persons.
[0092] In a specific embodiment of the present invention, the personalized motion perception of the specified human body is specifically based on the personalized motion features decoupled by the specified person motion decoupling module, fusing them with the general motion features, and referring to the proprietary motion features of the specified person, so as to realize the personalized perception of the hyper-realistic digital human of the specified person. For the video to be authenticated, first convert it into motion sequence features, and respectively extract the proprietary features and general features of the sequence through a proprietary motion encoder and a general motion encoder. Subsequently, concatenate the two features as the overall representation of the motion sequence. At the same time, introduce the reference proprietary motion features of the specified person, which are the personalized motion features of the specified person extracted in step 1, and input the features of the video to be authenticated and the reference features into the feature interaction module together. This module amplifies the difference between the input motion features and the reference motion features through information interaction, highlighting the motion signals of non-specified persons.
[0093] In a specific embodiment of the present invention, in step 103 above, pre-train a binary classification model to enable the binary classification model to obtain the ability to perceive the personalized motion of the specified person, and input the video to be detected for the specified person to complete forgery detection. The steps include:
[0094] Input the overall motion features after feature interaction into the binary classification model to complete the true / false binary classification task and realize the personalized perception of the hyper-realistic digital human video.
[0095] Finally, the interacted features are fed into a classifier to complete the true / false binary classification task, thus realizing the personalized perception of hyper-realistic digital human videos. This technical solution can not only effectively extract and fuse multi-level motion features, but also accurately identify forged content through the differential amplification mechanism between features, providing an efficient and reliable abnormal motion perception ability for the specified person. The specific steps are as Figure 3 shown.
[0096] Embodiment 2
[0097] As Figure 4 shown, an embodiment of the present application provides a hyper-realistic digital human video detection system, which adopts the above hyper-realistic digital human video detection method. The system includes:
[0098] Motion decoupling module 201: For a specified person, extract the action sequence based on the human body three-dimensional model, perform motion decoupling based on the visual model, and extract the personalized motion features of the specified person;
[0099] Personalized motion perception module 202: Perform personalized motion perception of the specified person by fusing the personalized motion features and general motion features of the person in the video to be detected;
[0100] Video detection module 203: Pre-train a binary classification model to enable the binary classification model to obtain the ability to perceive the personalized motion of the specified person, and input the video to be detected for the specified person to complete forgery detection.
[0101] In a specific embodiment of the present invention, the above motion decoupling module 201 includes:
[0102] Train the proprietary motion encoder using a self-supervised learning framework and complete the pre-training of the general motion encoder;
[0103] Input multiple motion sequences of the specified person, and extract the corresponding proprietary motion features and general motion features through the proprietary motion encoder and the general motion encoder respectively;
[0104] Cross-exchange the proprietary motion features and general motion features of the multiple extracted motion sequences to generate multiple reconstructed motion features;
[0105] Input the multiple reconstructed motion features into the general motion encoder to remove the interference of general features, and extract the individualized motion features through the proprietary motion encoder.
[0106] In a specific embodiment of the present invention, the above personalized motion perception module 202 includes:
[0107] Convert the video to be detected into motion sequence features, and extract the proprietary features and general features of the motion sequence features of the video to be detected through the proprietary motion encoder and the general motion encoder respectively;
[0108] The proprietary features and general features of the motion sequence features of the video to be detected are concatenated into an overall motion representation;
[0109] The personalized motion features of the specified person extracted are used as reference motion features, and the overall motion representation of the video to be authenticated and the reference motion features are input into the feature interaction module together; among them, the feature interaction module refers to a multi-modal feature fusion architecture based on difference perception, and its core lies in establishing a dynamic association mapping between the features of the video to be authenticated and the reference features. This module realizes feature interaction through the cross-attention mechanism: First, the reference proprietary features are used as Key-Value pairs, and the concatenated features (proprietary + general) of the video to be authenticated are used as Query to input into the multi-head attention layer to calculate the cross-feature similarity weights. Then, the gated residual connection mechanism is adopted to enhance the difference between the original features to be authenticated and the reference features weighted by attention. Specifically, when implemented, a learnable gated network is used to generate a difference mask, and the dimensions that deviate greatly from the reference features in the attention output are non-linearly amplified using the Sigmoid activation, while the consistent dimensions are suppressed. The finally output interaction features not only retain the general motion rules but also strengthen the personalized feature difference signals.
[0110] Through the feature interaction module, the difference between the input overall motion features and the reference motion features is amplified to highlight the motion signals of non-specified persons.
[0111] Embodiment III
[0112] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned hyper-realistic digital human video detection method are realized.
[0113] Embodiment IV
[0114] The embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the hyper-realistic digital human video detection method as described above are realized.
[0115] In addition, combined with Figure 1 the hyper-realistic digital human video detection method described in the embodiment of the present application can be realized by an electronic device, such as a computer device. Figure 5 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.
[0116] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 5 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.
[0117] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.
[0118] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.
[0119] By reading and executing the computer program instructions stored in the memory 82, the processor 81 implements any one of the above-mentioned hyper-realistic digital human video detection methods in the embodiments.
[0120] The technical features of the above-mentioned embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0121] The above-mentioned embodiments only represent several implementation manners of the present application, and the description thereof is relatively specific and detailed, but it cannot be understood as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A hyper-realistic digital human video detection method, characterized in that: The method comprises: For a specified person, action sequence extraction is performed based on a three-dimensional human body model, motion decoupling is performed based on a visual model, and personalized motion features of the specified person are extracted; The personalized motion perception of the designated person is performed by extracting the personalized motion features of the person in the video to be detected and fusing them with the general motion features; A binary classification model is pre-trained to enable the binary classification model to acquire the ability to perceive the personalized movements of a designated person, and a video to be detected of the designated person is input to complete forgery detection.
2. The method for detecting hyper-realistic digital human video according to claim 1, characterized in that: The step of performing motion decoupling based on the visual model to extract the personalized motion features of the designated person includes: A self-supervised learning framework is used to train the proprietary motion encoder; Among them, the universal motion encoder recognizes universal motion features through the universal human motion representation capabilities established in the pre-training stage; the universal motion encoder adopts the Transformer architecture to learn the basic laws of human motion through self-supervised tasks such as human motion sequence reconstruction and motion prediction; the proprietary motion encoder recognizes personalized motion features through a feature decoupling mechanism.
3. The method for detecting hyper-realistic digital human video according to claim 2, characterized in that: The step of performing motion decoupling based on the visual model to extract the personalized motion features of the designated person also includes: Input multiple motion sequences of the designated person, and extract corresponding dedicated motion features and universal motion features through the dedicated motion encoder and the universal motion encoder respectively; the universal motion encoder receives two motion sequence inputs from the same designated person, adopts a dynamic feature separation strategy to extract high-frequency detail features in the motion sequence through a learnable multi-layer temporal attention network, and then constrains its orthogonality with the universal motion features through a contrast loss function; The extracted proprietary motion features and universal motion features of the plurality of motion sequences are cross-exchanged to generate a plurality of reconstructed motion features; the proprietary encoder receives motion residual information, and by iteratively optimizing the motion sequence reconstruction error, the proprietary encoder gradually converges to the directional extraction capability of the personalized motion features while avoiding the overlap of representation with the universal motion features; The plurality of reconstructed motion features are input into the universal motion encoder to remove universal feature interference, and the individualized motion features are extracted through the proprietary motion encoder.
4. The method for detecting hyper-realistic digital human video according to claim 2, characterized in that: The steps of extracting the personalized motion features of the person in the video to be detected and fusing them with the universal motion features to perform the personalized motion perception of the designated person include: Converting the video to be detected into motion sequence features, and extracting the proprietary features and universal features of the motion sequence features of the video to be detected by the proprietary motion encoder and the universal motion encoder respectively; Splicing the extracted specific features and common features of the motion sequence features of the video to be detected into an overall motion representation; The extracted personalized motion features of the designated person are used as reference motion features, and the overall motion representation of the video to be identified and the reference motion features are input together for feature interaction; Through feature interaction, the difference between the input overall motion feature and the reference motion feature is amplified to highlight the motion signal of the non-specified person.
5. The method for detecting hyper-realistic digital human video according to claim 4, characterized in that: The pre-trained binary classification model enables the binary classification model to acquire the ability to perceive the personalized movement of a designated person, inputs a video to be detected for the designated person, and completes the forgery detection, the steps comprising: The overall motion feature after feature interaction is input into the binary classification model to complete the true and false binary classification task, thereby realizing personalized perception of the hyper-realistic digital human video.
6. A hyper-realistic digital human video detection system, using the hyper-realistic digital human video detection method as claimed in any one of claims 1 to 5, characterized in that: The system comprises: Motion decoupling module: for a specified person, the action sequence is extracted based on the human body 3D model, and the motion decoupling is performed based on the visual model to extract the personalized motion features of the specified person; Personalized motion perception module: extracts the personalized motion features of the person in the video to be detected and fuses them with the general motion features to perform personalized motion perception of the specified person; Video detection module: pre-train a binary classification model to enable the binary classification model to acquire the ability to perceive the personalized movements of a specified person, input a video to be detected for the specified person, and complete forgery detection.
7. The hyper-realistic digital human video detection system according to claim 6, characterized in that: The motion decoupling module comprises: A self-supervised learning framework is used to train the proprietary motion encoder and pre-train the universal motion encoder. Input a plurality of motion sequences of the designated person, and extract corresponding dedicated motion features and universal motion features through the dedicated motion encoder and the universal motion encoder respectively; Cross-exchanging the extracted specific motion features and general motion features of the plurality of motion sequences to generate a plurality of reconstructed motion features; The plurality of reconstructed motion features are input into the universal motion encoder to remove universal feature interference, and the individualized motion features are extracted through the proprietary motion encoder.
8. The hyper-realistic digital human video detection system according to claim 6, characterized in that: The personalized motion perception module includes: Converting the video to be detected into motion sequence features, and extracting the proprietary features and universal features of the motion sequence features of the video to be detected by the proprietary motion encoder and the universal motion encoder respectively; Splicing the extracted specific features and common features of the motion sequence features of the video to be detected into an overall motion representation; The extracted personalized motion features of the designated person are used as reference motion features, and the overall motion representation of the video to be identified and the reference motion features are input into a feature interaction module; wherein the feature interaction module realizes feature interaction through a cross-attention mechanism; Through the feature interaction module, the difference between the input overall motion feature and the reference motion feature is amplified to highlight the motion signal of the non-specified person.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the hyper-realistic digital human video detection method described in any one of claims 1 to 5 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the hyper-realistic digital human video detection method according to any one of claims 1 to 5 are implemented.