Hyperrealistic digital human detection method and system based on motion modeling

Through the method based on action modeling, the action features in hyperrealistic digital human videos are extracted and analyzed using SMPL-X and visual Transformer models, and the shortcomings of the existing technology in detecting the forged traces of hyperrealistic digital human videos are solved, and efficient and generalized forged detection effects are achieved.

CN120236326APending Publication Date: 2025-07-01INST OF COMPUTING TECH CHINESE ACAD OF SCI +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510297683.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect forgery traces in hyperrealistic digital human videos, especially in unseen artifact recognition, and methods focusing on advanced semantics are difficult to meet the needs of detecting hyperrealistic digital humans under the limitations of feature extraction efficiency and training set scale.

Method used

A hyperrealistic digital man forged detection method based on action modeling is used to extract action sequences through the SMPL-X model, and motion features are extracted in combination with the visual Transformer model, a mixed model of action feature distribution is constructed, and anomaly scores are calculated to complete forged detection.

Benefits of technology

It realizes efficient identification and detection of hyperrealistic digital human videos, has strong generalization, can cross data sets and cross forgery methods, and significantly improves the ability to identify digital human videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236326A_ABST
    Figure CN120236326A_ABST
Patent Text Reader

Abstract

The invention discloses a super-realistic digital human detection method based on motion modeling, and the method comprises the steps: carrying out the motion sequence extraction of a designated person based on a human body three-dimensional model; extracting motion features of the specified character based on the motion sequence and a visual model; and for the extracted motion features, mapping the motion features into a feature space, constructing an action feature distribution hybrid model, calculating an abnormal score of the digital human under normal feature distribution based on the model, and completing digital human counterfeiting detection according to the abnormal score. According to the method and the system, a super-realistic digital human forgery detection method based on action modeling is matched with an existing action capture technology, more universal action characteristics can be learned, so that the effect of detecting the super-realistic digital human is achieved, and the method and the system have relatively high generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of forgery detection, and particularly to a method and system for detecting forgery of hyper-realistic digital humans based on action modeling. Background Art

[0002] In recent years, with the rapid development of video generation technology, the generation quality of digital humans has been rapidly improved, which also prompts us to need new forgery detection methods to handle potential social security problems. The hyper-realistic digital human forgery detection technology aims to characterize the inter-frame motion law through the change of the inter-frame motion flow by means of action modeling, use the sequence model to learn the human motion law, and learn to perceive the false motion behavior of digital humans on the digital human forgery dataset to achieve the purpose of detecting digital humans.

[0003] The goal of forgery detection technology is to ensure the authenticity and integrity of video content and prevent forged videos from having an adverse impact on individuals, society, and institutions by comprehensively applying various technical means such as artificial intelligence, computer vision, and computer graphics. The main technical routes of existing forgery detection technologies are divided into two types. One is to detect forgery traces by designing network architectures, data augmentation methods, and loss functions to achieve the purpose of identifying the authenticity of videos, that is, detecting video artifacts, model fingerprints, etc. Although this type of technology performs outstandingly when detecting artifacts that the detection model has seen, it performs mediocrely in identifying unseen artifacts and has poor generalization. The other is to detect the high-level semantics that are difficult to replicate in the deep forgery process by learning the temporal semantics extracted by the feature network. This type of technology no longer focuses on artifacts but on some high-level semantic information, so it has better generalization. However, some of these methods need to add a reference set for testing, which affects the user interaction experience.

[0004] To achieve an ideal hyper-realistic digital human forgery detection effect, it is necessary to extract efficient video human body features. In existing network models, the Vision Transformer network (ViT) is a deep learning model based on the Transformer architecture, which is specifically used for computer vision tasks. It uses the self-attention mechanism to process image data. Compared with traditional convolutional neural networks (CNNs), ViT can capture higher-level semantics and model features that the model wants to learn more effectively. Existing forgery detection technologies based on the ViT network often still focus on detecting forgery traces. However, with the wide application of diffusion models in the generation field, forged videos based on diffusion models often no longer have low-frequency forgery information, and such methods no longer have an advantage in detecting forged videos. Therefore, forgery detection methods that focus on high-level semantics will be the focus of hyper-realistic digital human forgery detection technology.

[0005] In summary, the following disadvantages exist in the prior art:

[0006] On the one hand, the existing forgery detection technologies that focus on forgery traces mainly focus on forgery traces by means of data augmentation, designing network architectures, etc., and do not have the ability to detect the forgery traces of hyper-realistic digital humans. On the other hand, the existing forgery detection technologies that focus on high-level semantics often still mainly rely on features such as audio-visual inconsistency and facial expressions to identify whether a video is forged. This technology is limited by the scale of the training set and the efficiency of feature extraction, and it is difficult to meet the application requirements for detecting hyper-realistic digital humans.

[0007] Based on the above analysis, it is urgent to research and develop a set of forgery detection methods for hyper-realistic digital humans based on action modeling. Based on the powerful feature extraction ability of the ViT model, on the premise of ensuring the extraction of human motion sequences, the model is trained to perceive the general action features of hyper-realistic digital humans, so as to realize an available forgery detection technology for hyper-realistic digital humans. Summary of the Invention

[0008] The main technical problems to be solved by the present invention include: how to accurately extract the action features in the dataset; how to learn a more general action model. A forgery detection method for hyper-realistic digital humans based on action modeling is proposed. By cooperating with the existing motion capture technology, more general action features can be learned, so as to achieve the effect of detecting hyper-realistic digital humans, and it has strong generalization.

[0009] To solve the above problems, the present invention proposes a forgery detection method for hyper-realistic digital humans based on action modeling. Its purpose is to extract the general forgery action features of digital humans on the provided hyper-realistic digital human dataset, so as to achieve the effect of detecting hyper-realistic digital human videos.

[0010] In the first aspect, the embodiments of the present application provide a detection method for hyper-realistic digital humans based on action modeling. The method includes:

[0011] For a specified person, extract the action sequence based on the human body three-dimensional model;

[0012] Based on the action sequence, extract the motion features of the specified person based on the visual model;

[0013] For the extracted motion features, map the motion features into the feature space, construct a mixture model of action feature distributions, and calculate the anomaly score of the digital human under the normal feature distribution based on the model, and complete the forgery detection of the digital human according to the anomaly score.

[0014] In the specific embodiments of the present invention, the above-mentioned step of extracting the action sequence based on the human body three-dimensional model for the specified person includes:

[0015] Extract the action sequence of the hyper-realistic digital human video based on the SMPL-X model; extract the 3D mesh data of the human body from the video, save the local feature information of the body of the person through the SMPL-X parameters, calculate the action flow of the front and back frames in the video using the SMPL-X model, and further extract and describe the human motion characteristics.

[0016] In a specific embodiment of the present invention, the above-mentioned extraction of the action sequence of the hyper-realistic digital human video based on the SMPL-X model includes the following steps:

[0017] Adopt a first-order motion estimation method based on motion anchors. By assuming the motion anchors in two adjacent frames of images, approximate the motion flow of the surrounding pixel points using the local affine transformation of the anchors, so as to accurately describe human motion;

[0018] By combining the parameterization method of the SMPL-X model with the action flow estimation based on motion anchors, finally generate an action sequence.

[0019] In a specific embodiment of the present invention, the above-mentioned extraction of the motion characteristics of a specified person based on the visual model from the action sequence further includes the following steps:

[0020] Based on the feature extraction network of the visual model, through the extracted action sequence, use the feature extraction network for learning, randomly cover a part of the action sequence, and then input the covered sequence into the visual model to let the visual model predict the motion flow at the covered place; wherein, the feature extraction network includes multiple encoder layers;

[0021] The encoder layer encodes the input action sequence into a series of feature vectors, captures the global and local information in the sequence, and generates an output sequence of motion characteristics.

[0022] In a specific embodiment of the present invention, for the above-mentioned extracted motion characteristics, map the motion characteristics into the feature space and construct a mixture model of action feature distributions, including the following steps:

[0023] Use a pre-trained feature extraction network to extract features from the input action sequence. The multi-head self-attention mechanism of the feature extraction network captures the long-range dependence relationship and high-dimensional features of spatio-temporal dynamic features in the action sequence;

[0024] After feature extraction, input the high-dimensional features into a Gaussian mixture model to construct a distribution model of normal motion patterns, and characterize the feature distribution of normal digital human motion through the pre-trained model.

[0025] In a specific embodiment of the present invention, the above-mentioned calculation of the anomaly score of the digital human under the normal feature distribution based on the model and the completion of digital human forgery detection according to the anomaly score include the following steps:

[0026] Input the video to be detected, extract its feature vector through the feature extraction network, and use the Gaussian mixture model to calculate the anomaly score of the video to be detected relative to the normal distribution. A high anomaly score indicates that the motion pattern of the video deviates from the normal distribution, and it is judged that the video may be forged.

[0027] In a specific embodiment of the present invention, the above steps of using the Gaussian mixture model to calculate the anomaly score of the video to be detected relative to the normal distribution, where a high anomaly score indicates that the motion pattern of the video deviates from the normal distribution, include:

[0028] By setting an anomaly score threshold, when the anomaly score is greater than the threshold, it is identified as a forged video.

[0029] In a second aspect, an embodiment of the present application provides a hyper-realistic digital human detection system based on action modeling, which adopts the hyper-realistic digital human detection method based on action modeling as described above. The system includes:

[0030] Action sequence extraction module: For a specified person, extract the action sequence based on the three-dimensional human model;

[0031] Motion feature extraction module: Based on the action sequence, extract the motion features of the specified person based on the visual model;

[0032] Video detection module: For the extracted motion features, map the motion features into the feature space, construct a mixture model of action feature distributions, and calculate the anomaly score of the digital human under the normal feature distribution based on the model, and complete the forgery detection of the digital human according to the anomaly score.

[0033] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the hyper-realistic digital human detection method based on action modeling are implemented.

[0034] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the hyper-realistic digital human detection method based on action modeling as described above are implemented.

[0035] Compared with the related prior art, it has the following outstanding beneficial effects:

[0036] The present invention can perform motion modeling on the collected videos of hyper-realistic digital humans and learn general and robust motion features through training. Through the sensitivity to abnormal motion patterns, the present invention can effectively identify hyper-realistic digital humans. By capturing and analyzing the motion features of hyper-realistic digital humans, the model can accurately identify forged videos that do not conform to the motion of real human bodies, thereby achieving efficient identification and detection of digital human videos in various application scenarios. This technology greatly improves the identification ability of hyper-realistic digital human videos, has broad application prospects and important practical significance. Compared with the methods for detecting forged videos, the present invention has stronger generalization ability for cross-datasets and cross-forgery methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0038] Figure 1 It is a schematic diagram of the method for detecting hyper-realistic digital human videos of the present invention;

[0039] Figure 2 It is a schematic diagram of the human motion estimation process of the embodiment of the present invention;

[0040] Figure 3 It is a schematic diagram of the hyper-realistic digital human video detection system of the present invention;

[0041] Figure 4 It is a schematic diagram of the computer hardware of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0043] It should also be understood that the term "and / or" in this article is only a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0044] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0045] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0046] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0047] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0048] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0049] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0050] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0051] The method of the present invention aims to propose a method for detecting hyper-realistic digital human forgery based on action modeling. Its purpose is to extract the general forgery action features of digital humans on the provided hyper-realistic digital human dataset, so as to achieve the effect of detecting hyper-realistic digital human videos.

[0052] In the formation process of the present invention, the following technical key points mainly emerged:

[0053] 1. Accurate action sequence extraction. The present invention extracts action sequences based on SMPL-X (Skinned Multi-Person Linear Model-eXpressive). SMPL-X combines geometric, pose, and expression information, and is naturally more adaptable to learning the action sequences of hyper-realistic digital humans than technologies such as skeletal key points.

[0054] 2. General motion feature modeling. The present invention is developed based on the Vision Transformer model (ViT). By using positional encoding as the information extraction method, it can accurately learn body pose actions and smooth lip shapes of the mouth, enabling the model to learn more robust action features.

[0055] 3. Action anomaly detection. Feature extraction is performed on the action sequences extracted from the training data, which are mapped into the feature space. A Gaussian Mixture Model (GMM) of the action feature distribution is constructed, and its anomaly score under the normal feature distribution is calculated. The higher the anomaly score, the more likely the action sequence is forged.

[0056] The following describes the method of the embodiment of the present application in detail with specific embodiments:

[0057] Embodiment 1

[0058] As Figure 1 shown, the embodiment of the present application provides a method for detecting hyper-realistic digital human videos, and the method includes:

[0059] A hyper-realistic digital human detection method based on action modeling, the method comprising:

[0060] Step 101, for a specified person, extract an action sequence based on a human body three-dimensional model;

[0061] In a specific embodiment of the present invention, the extraction of general action features of the digital human is specifically to collect a variety of hyper-realistic digital human videos based on neural radiance fields and structural squares, extract human 3D mesh data from the videos, save the mouth shapes, full-body postures and hand movements of the person with SMPL-X parameters, and extract the corresponding action sequences as the input for model training.

[0062] Step 102, extract the motion features of the specified person based on the action sequence based on a vision model;

[0063] In a specific embodiment of the present invention, the perception of digital human action features based on ViT is specifically as follows: In order to be able to learn more general action features, the present invention proposes a feature extraction network based on ViT. Through the extracted action sequence, use the ViT network for learning. Given an input sequence {T1, T2, … Tn}, randomly cover a part of it, and then input the covered sequence into the model to let the model predict the motion flow at the covered part. Adopt the encoder-decoder structure of Transformer, where the encoder encodes the input sequence into a fixed-length vector, and the decoder uses this vector as the input to generate the output sequence.

[0064] Step 103, for the extracted motion features, map the motion features into a feature space, construct a mixture model of action feature distributions, and calculate the anomaly score of the digital human under the normal feature distribution based on the model, and complete the forgery detection of the digital human according to the anomaly score.

[0065] In a specific embodiment of the present invention, the forgery detection of the digital human based on anomaly detection is specifically as follows: The pre-trained model already has the ability to model the real human body motion distribution, and will further perform true-false binary classification training on the digital human motion video. Adopt binary cross-entropy as the loss function to calculate the classification loss to ensure that the model can quickly learn and adapt to the motion pattern of the digital human. Through the digital human body motion distribution learned from a large amount of digital human data, construct a robust and well-generalized classification surface so that the model can accurately perceive and identify false motion patterns. This method uses the sensitivity of the model to abnormal motion patterns through the anomaly detection framework to achieve efficient and accurate detection of forged videos.

[0066] In a specific embodiment of the present invention, in step 101, the above-mentioned extraction of the action sequence based on the human body three-dimensional model for the specified person comprises the steps of:

[0067] Extract the action sequence of the hyper-realistic digital human video based on the SMPL-X model; extract the 3D mesh data of the human body from the video, and save the local feature information of the human body through the SMPL-X parameters. Use the SMPL-X model to calculate the action flow of the front and back frames in the video, and further extract and describe the human motion characteristics.

[0068] In a specific embodiment of the present invention, the above-mentioned extraction of the action sequence of the hyper-realistic digital human video based on the SMPL-X model includes the steps of:

[0069] Adopt a first-order motion estimation method based on motion anchors. By assuming the motion anchors in two adjacent frames of images, use the local affine transformation of the anchors to approximate the motion flow of the surrounding pixel points, so as to accurately describe human motion;

[0070] By combining the parameterization method of the SMPL-X model with the action flow estimation based on motion anchors, finally generate an action sequence.

[0071] In a specific embodiment of the present invention, the extraction of the digital human action sequence is specifically as follows:

[0072] First, collect a variety of hyper-realistic digital human videos based on neural radiance fields and structured squares. Extract the 3D mesh data of the human body from the videos, and save the mouth shape, full-body posture and hand movements of the characters through the SMPL-X parameters. Use the SMPL-X model to calculate the action flow of the front and back frames in the videos, and further extract and describe the human motion characteristics. To improve the calculation efficiency and robustness, the present invention proposes a first-order motion estimation method based on motion anchors. By assuming the motion anchors in two adjacent frames of images, use the local affine transformation of these anchors to approximate the motion flow of the surrounding pixel points, so as to accurately describe human motion. By combining the parameterization method of the SMPL-X model with the action flow estimation based on motion anchors, the present invention finally generates a complete and accurate action sequence. This sequence not only retains the full-body posture and hand movements of the human body, but also reflects the dynamic change characteristics through the affine transformation of the motion anchors. Combining these two methods, the present invention can efficiently extract and model the human action characteristics in hyper-realistic digital human videos, and use these characteristics to achieve accurate detection of forged videos. The specific calculation formula is as follows:

[0073]

[0074] Among them, and are the action anchors in the k-th frame of the front and back two frames, T k represents the action flow estimation obtained from the k-th pair of motion anchors, and θ k represents the local affine transformation corresponding to the k-th anchor. The specific process is as Figure 2 shown.

[0075] In a specific embodiment of the present invention, in step 102, for extracting the motion features of a specified person based on the action sequence and the visual model, the steps further include:

[0076] Based on the feature extraction network of the visual model, through the extracted action sequence, use the feature extraction network for learning, randomly cover a part of the action sequence, and then input the covered sequence into the visual model to let the visual model predict the motion flow at the covered part; wherein, the feature extraction network includes multiple encoder layers;

[0077] The encoder layer encodes the input action sequence into a series of feature vectors, captures the global and local information in the sequence, and generates an output sequence of motion features.

[0078] In a specific embodiment of the present invention, the digital human action feature perception module based on ViT is specifically:

[0079] In order to model long-distance motion features and learn more general action features, the present invention proposes a feature extraction network based on Vision Transformer (ViT). The ViT network was initially used for image classification tasks, but its self-attention mechanism is also very suitable for capturing long-term dependencies in human motion sequences. Through the extracted action sequence, use the ViT network for learning. Given an input sequence {T1, T2, …, Tn}, randomly cover a part of it, and then input the covered sequence into the model to let the model predict the motion flow at the covered part. The architecture of ViT includes multiple Transformer encoder layers, and these encoder layers encode the input sequence into a series of feature vectors, capturing the global and local information in the sequence. Using this encoding method, ViT can generate an output sequence containing rich motion features, thereby more accurately describing the human motion pattern and realizing the effective detection of forged videos.

[0080] In a specific embodiment of the present invention, in step 103, for the extracted motion features, map the motion features into the feature space and construct a mixture model of action feature distributions. The steps include:

[0081] Use a pre-trained feature extraction network to extract features from the input action sequence. The multi-head self-attention mechanism of the feature extraction network captures the long-range dependencies and high-dimensional features of spatio-temporal dynamic features in the action sequence;

[0082] After feature extraction, input the high-dimensional features into a Gaussian mixture model to construct a distribution model of normal motion patterns, and characterize the feature distribution of normal digital human motion through the pre-trained model.

[0083] In a specific embodiment of the present invention, the above-mentioned method for calculating the anomaly score of a digital human under the normal feature distribution based on a model and completing the forgery detection of the digital human includes the following steps:

[0084] Input the video to be detected, extract its feature vector through a feature extraction network, and use a Gaussian mixture model to calculate the anomaly score of the video to be detected relative to the normal distribution. A high anomaly score indicates that the motion pattern of the video deviates from the normal distribution, and it is judged that the video may be a forged video.

[0085] In a specific embodiment of the present invention, the above-mentioned method for calculating the anomaly score of the video to be detected relative to the normal distribution using a Gaussian mixture model, where a high anomaly score indicates that the motion pattern of the video deviates from the normal distribution, includes the following steps:

[0086] By setting an anomaly score threshold, when the anomaly score is greater than the threshold, it is identified as a forged video.

[0087] In a specific embodiment of the present invention, the digital human forgery detection module based on anomaly detection is specifically:

[0088] The present invention already has the ability to model the motion distribution of digital humans and will further perform true / false binary classification training on digital human motion videos. Binary cross-entropy is used as the loss function to calculate the classification loss, ensuring that the model can quickly learn and adapt to the motion patterns of digital humans. Through the motion distribution of digital humans learned from a large amount of digital human data, a robust and well-generalized classification surface is constructed, enabling the model to accurately perceive and identify false motion patterns. To improve the anomaly detection ability of the model, the present invention introduces a multi-level feature extraction and anomaly detection framework. First, a pre-trained ViT network is used to extract features from the input action sequence. The multi-head self-attention mechanism of ViT can capture long-range dependencies and complex spatio-temporal dynamic features in the action sequence. The extracted high-dimensional feature vector not only reflects the local details of the motion but also retains the global motion pattern. After feature extraction, these high-dimensional features are input into a Gaussian mixture model (GMM) to construct a distribution model of normal motion patterns. By training a large amount of normal data, the model can accurately characterize the feature distribution of normal digital human motion. In the detection stage, for a new input video, first, its feature vector is extracted through the ViT network, and then the GMM model is used to calculate its anomaly score relative to the normal distribution. A high anomaly score indicates that the motion pattern of the video deviates from the normal distribution and is very likely to be a forged video. By setting a reasonable threshold, the model can efficiently and accurately identify forged videos. The present invention utilizes the feature extraction advantage of ViT and the anomaly detection accuracy of GMM to ensure the robustness and generalization ability of the model, and finally realizes the efficient and accurate detection of forged videos, which has broad application prospects and important practical significance.

[0089] Embodiment 2

[0090] AsFigure 3 As shown in the figure, an embodiment of the present application provides a hyper-realistic digital human detection system based on action modeling, which adopts the above-mentioned hyper-realistic digital human detection method based on action modeling. The system includes:

[0091] Action sequence extraction module 201: For a specified person, extract the action sequence based on the human body three-dimensional model;

[0092] Motion feature extraction module 202: Extract the motion features of the specified person based on the action sequence and the visual model;

[0093] Video detection module 203: For the extracted motion features, map the motion features into the feature space, construct a mixture model of action feature distributions, and calculate the anomaly score of the digital human under the normal feature distribution based on the model, and complete the digital human forgery detection according to the anomaly score.

[0094] Embodiment III

[0095] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above-mentioned hyper-realistic digital human detection method based on action modeling are implemented.

[0096] Embodiment IV

[0097] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned hyper-realistic digital human detection method based on action modeling are implemented.

[0098] In addition, the hyper-realistic digital human detection method based on action modeling described in the embodiments of the present application can be implemented by an electronic device, such as a computer device. Figure 1 FIG. is a schematic hardware structure diagram of a computer device according to an embodiment of the present application. Figure 4 FIG.

[0099] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 4 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.

[0100] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0101] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0102] By reading and executing the computer program instructions stored in the memory 82, the processor 81 implements any one of the ultra-realistic digital human video detection methods in the above embodiments.

[0103] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0104] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A hyper-realistic digital human detection method based on action modeling, characterized in that: The method comprises: For a specified person, action sequence extraction is performed based on the 3D human body model; Extracting motion features of the designated person based on a visual model based on the action sequence; The extracted motion features are mapped into feature space, a motion feature distribution hybrid model is constructed, and an abnormal score of the digital human under normal feature distribution is calculated based on the model, and digital human forgery detection is completed according to the abnormal score.

2. The method for detecting hyper-realistic digital human video according to claim 1, characterized in that: The steps of extracting action sequences for a designated person based on a three-dimensional human body model include: The action sequence of hyper-realistic digital human video is extracted based on the SMPL-X model. The 3D mesh data of the human body is extracted from the video, and the local feature information of the human body is saved through SMPL-X parameters. The action flow of the previous and next frames in the video is calculated using the SMPL-X model to further extract and describe the human motion characteristics.

3. The method for detecting hyper-realistic digital human video according to claim 2, characterized in that: The steps of extracting action sequences from hyper-realistic digital human videos based on the SMPL-X model include: A first-order motion estimation method based on motion anchor points is adopted. By assuming motion anchor points in two adjacent frames of images, the motion flow of surrounding pixels is approximated by using the local affine transformation of the anchor points, thereby accurately describing human body motion. By combining the parameterization method of the SMPL-X model with the motion flow estimation based on motion anchors, an action sequence is finally generated.

4. The method for detecting hyper-realistic digital human video according to claim 1, characterized in that: The step of extracting the motion features of the designated person based on the action sequence and the visual model further comprises: A feature extraction network based on a visual model is used to learn the extracted action sequence using the feature extraction network, randomly masking a portion of the action sequence, and then inputting the masked sequence into the visual model to allow the visual model to predict the motion flow at the masked portion; wherein the feature extraction network includes a plurality of encoder layers; The encoder layer encodes the input action sequence into a series of feature vectors, capturing global and local information in the sequence, and generates an output sequence of motion features.

5. The method for detecting hyper-realistic digital human video according to claim 1, characterized in that: The steps of mapping the extracted motion features into a feature space and constructing a motion feature distribution mixed model include: Using the pre-trained feature extraction network to extract features from the input action sequence, the multi-head self-attention mechanism of the feature extraction network captures the high-dimensional features of the long-range dependencies and spatiotemporal dynamic features in the action sequence; After feature extraction, the high-dimensional features are input into a Gaussian mixture model to construct a distribution model of normal motion patterns, and the feature distribution of normal digital human motion is characterized by a pre-trained model.

6. The method for detecting hyper-realistic digital human video according to claim 1, characterized in that: The steps of calculating the abnormal score of the digital human under normal feature distribution based on the model and completing the digital human forgery detection according to the abnormal score include: The video to be detected is input, its feature vector is extracted through the feature extraction network, and the Gaussian mixture model is used to calculate the anomaly score of the video to be detected relative to the normal distribution. A high anomaly score indicates that the motion pattern of the video deviates from the normal distribution, and it is judged that it may be a forged video.

7. The method for detecting hyper-realistic digital human video according to claim 6, characterized in that: The step of calculating the abnormality score of the video to be detected relative to the normal distribution by using the Gaussian mixture model, wherein a high abnormality score indicates that the motion pattern of the video deviates from the normal distribution, comprises: By setting the abnormality score threshold, when the abnormality score is greater than the threshold, it is identified as a forged video.

8. A hyper-realistic digital human detection system based on action modeling, using the hyper-realistic digital human detection method based on action modeling as claimed in any one of claims 1 to 7, characterized in that: The system comprises: Action sequence extraction module: for a specified person, action sequence extraction is performed based on the human body 3D model; Motion feature extraction module: extracting the motion features of the designated person based on the action sequence and the visual model; Video detection module: for the extracted motion features, the motion features are mapped into the feature space, a motion feature distribution hybrid model is constructed, and the abnormal score of the digital human under the normal feature distribution is calculated based on the model, and the digital human forgery detection is completed according to the abnormal score.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the hyper-realistic digital human detection method based on action modeling described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the hyper-realistic digital human detection method based on action modeling are implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Three-dimensional digital human detection method and device based on large model, electronic equipment, medium, program product and three-dimensional digital human

    CN121169888A

  • Methods, devices, electronic equipment, media, and software products for 3D digital human detection based on large models, and 3D digital humans.

    CN121169888B