Deep forgery detection method and device based on visual detection, equipment and medium

By dynamically enhancing and mapping high-dimensional features of the original video frames, combined with enhanced domain label analysis and source obfuscation processing, the problem of insufficient feature robustness of existing forgery detection methods is solved, and more accurate forgery probability detection is achieved.

CN121962871APending Publication Date: 2026-05-01PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video frame forgery detection methods are limited by the single feature dimension of the original video frame and are easily interfered with by simple operations, resulting in insufficient robustness of the extracted features. They cannot effectively distinguish between real and forged frames and fail to fully eliminate the domain differences caused by enhancement operations, making it difficult to achieve accurate and stable forgery probability prediction in complex and ever-changing user-generated content scenarios.

Method used

By dynamically enhancing the original video frames, an enhanced domain view is generated and converted into a high-dimensional feature vector for normalization mapping. Enhanced domain label analysis and source obfuscation are then performed to extract stable feature vectors to detect the probability of forgery.

Benefits of technology

It improves the accuracy and stability of forgery detection, avoids misjudging normal enhancement mutations as forgery traces, enhances the model's anti-interference ability, and improves the ability to capture the core features of forged videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962871A_ABST
    Figure CN121962871A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of security detection, can be applied to business system platforms of medical health, financial science and technology and the like, and discloses a depth forgery detection method, device and equipment based on visual inspection and a medium, and the method comprises the steps: obtaining an original video frame issued by a user, carrying out the dynamic enhancement of the original video frame, and generating an enhanced domain view; converting the enhanced domain view into a view high-dimensional feature vector, and performing normalized mapping on the view high-dimensional feature vector to obtain a view normalized feature vector; performing enhanced domain label analysis on the view normalized feature vector to obtain a vector enhanced domain label; performing enhancement domain source confusion processing on the view normalized feature vector according to a vector enhancement domain label to obtain a stable feature vector; and detecting the forgery probability of the original video frame according to the stable feature vector. Through enhanced domain label analysis and enhanced domain source confusion processing on the normalized feature vector of the view, normal enhanced variation is prevented from being misjudged as a counterfeit trace, and the detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of security detection technology, and in particular to a method, apparatus, equipment and medium for deepfake detection based on visual inspection. Background Technology

[0002] With the rapid development of deepfake technology, the number of fake media on user-generated content (UGC) platforms has surged, posing a serious challenge to social trust and cybersecurity. Against this backdrop, video frame forgery detection methods have emerged and have deeply penetrated into many key areas. For example, in the healthcare field, they can identify forged lesion image interpretation videos to prevent unauthorized personnel from misleading patients' diagnosis and treatment decisions by tampering with medical images. In the fintech field, they can identify forged facial videos in remote account opening and identity verification processes to prevent others from impersonating the user to conduct financial transactions.

[0003] Currently, traditional video frame forgery detection methods mostly extract features directly from the original video frames and classify them. However, these methods are often limited by the single feature dimension of the original video frames and are easily interfered with by simple operations, resulting in insufficient robustness of the extracted features and an inability to effectively distinguish between real frames and forged frames that have undergone enhancement. At the same time, traditional methods fail to fully eliminate the domain differences brought about by enhancement operations, making it difficult to achieve accurate and stable forgery probability prediction in complex and ever-changing user-generated content scenarios.

[0004] Therefore, in the face of the growing demand for depth forgery detection in visual inspection, current depth forgery detection methods in visual inspection urgently need to be improved to solve the problem of insufficient accuracy in predicting forgery probability in existing methods. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and medium for depth forgery detection based on visual inspection, which addresses how to improve the accuracy of probability detection of view forgery.

[0006] Firstly, a visual detection-based deepfake detection method is provided, including: Obtain the original video frames published by the user, and dynamically enhance the original video frames to generate an enhanced domain view; The enhanced domain view is converted into a high-dimensional feature vector of the view, and the high-dimensional feature vector of the view is normalized and mapped to obtain a normalized feature vector of the view. Enhanced domain label analysis is performed on the normalized feature vector of the view to obtain vector enhanced domain labels; Based on the vector enhancement domain label, the view normalized feature vector is subjected to enhancement domain source obfuscation processing to obtain a stable feature vector; The probability of forgery of the original video frame is detected based on the stable feature vector.

[0007] Secondly, a vision-based deepfake detection device is provided, comprising: The view dynamic enhancement module is used to acquire the original video frames published by the user and dynamically enhance the original video frames to generate an enhanced domain view. The feature vector conversion module is used to convert the enhanced domain view into a high-dimensional feature vector of the view; The feature vector mapping module is used to normalize and map the high-dimensional feature vectors of the view to obtain normalized feature vectors of the view. The enhanced domain label analysis module is used to perform enhanced domain label analysis on the normalized feature vector of the view to obtain vector enhanced domain labels. The feature vector enhancement domain source obfuscation processing module is used to perform enhancement domain source obfuscation processing on the view normalized feature vector according to the vector enhancement domain label to obtain a stable feature vector. The forgery probability detection module is used to detect the forgery probability of the original video frame based on the stable feature vector.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned vision-based deepfake detection method.

[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned deepfake detection method based on visual inspection.

[0010] The aforementioned scheme for deepfake detection based on visual inspection, including the method, apparatus, computer equipment, and storage medium, dynamically enhances the original video frames to simulate video variations under different devices, environments, and transmission conditions, expanding the diversity of training data and avoiding insufficient generalization ability due to single data. It extracts the core discriminative features of the video frames by converting the enhanced domain view into high-dimensional features and normalizing them. Simultaneously, normalization eliminates feature scale / distribution differences caused by different enhancement operations, ensuring consistency of features for the same target under different enhanced views. By analyzing the normalized feature vectors of the enhanced domain labels, it accurately identifies the enhancement type corresponding to the features, avoiding misjudging normal enhancement variations as forgery traces and reducing the false alarm rate. Obfuscation processing is performed based on the enhanced domain labels of the vectors, stripping away domain differences unrelated to forgery detection, thus allowing features to focus more on the essential traces of forgery and improving detection accuracy. By detecting the forgery probability of the original video frames using stable feature vectors, it can more accurately capture the core features of forged videos, improving detection accuracy and stability. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of an application environment for a visual detection-based deepfake detection method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a deepfake detection method based on visual detection in one embodiment of the present invention; Figure 3 This is a schematic diagram of a vision-based deepfake detection device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to one embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] This invention provides a vision-based deepfake detection method that can be applied to applications such as... Figure 1In this application environment, the client communicates with the server via a network. The server can acquire the original video frames published by the user and dynamically enhance them to generate an enhanced domain view. The enhanced domain view is converted into a high-dimensional feature vector, and this high-dimensional feature vector is normalized to obtain a normalized feature vector. Enhancement domain label analysis is performed on the normalized feature vector to obtain vector enhancement domain labels. Enhancement domain source obfuscation is applied to the normalized feature vector based on the vector enhancement domain labels to obtain a stable feature vector. The forgery probability of the original video frame is detected based on the stable feature vector, and the forgery probability is fed back to the client. This invention provides a deep forgery detection device based on visual detection. For forgery probability detection, by analyzing the enhancement domain labels and obfuscating the enhancement domain source of the normalized feature vector, it avoids misjudging normal enhancement variations as forgery traces, thus improving detection accuracy. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention is described in detail below through specific embodiments.

[0015] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a vision-based deepfake detection method provided in this embodiment of the invention includes the following steps: S1. Obtain the original video frames published by the user, and dynamically enhance the original video frames to generate an enhanced domain view.

[0016] In this embodiment of the invention, the original video frame refers to an image or video directly posted by the user on social media, e-commerce, or other platforms.

[0017] In the medical and health field, real-time video frames from surgical endoscopy, dynamic ultrasound images, and pathological slide scanning video frames from remote diagnosis are acquired. The brightness, contrast, and noise levels of the video frames are adjusted to match the output effects of endoscopes and ultrasound machines from different manufacturers. At the same time, JPEG compression distortion and Gaussian blur are added to prevent some people from tampering with endoscopic surgical videos and forging ultrasound reports to defraud medical insurance.

[0018] In fintech scenarios, the system acquires dynamic video frames of a user's face during remote account opening, adjusts the exposure and white balance of the video frames, reproduces the shooting effect of the user's mobile phone or computer camera in strong light / low light, adds video compression block effects and stuttering frame skipping, matches the transmission loss of real-time video calls, and prevents the use of deepfake technology to generate fake facial videos to impersonate others for remote account opening, loan face-to-face signing and other operations.

[0019] In this embodiment of the invention, the step of dynamically enhancing the original video frame to generate an enhanced domain view includes: The original video frame is subjected to affine transformation and elastic warping to form a deformed video frame; The photometric adjustment is performed on the deformed video frame to obtain the photometric adjusted video frame; The photometric adjustment video frame is subjected to signal interference to generate an interfering intermediate frame; An independent domain transformation is performed on the interfering intermediate frame to obtain an enhanced domain view.

[0020] In this embodiment of the invention, during the process of performing affine transformation and elastic distortion on the original video frame, the original video frame is first rotated, scaled, and translated using an affine transformation function to output a pre-deformed video frame. Then, random stretching / compression forces are applied to the local pixel regions of the pre-deformed video frame, such as slightly distorting the left region of the video frame, thereby generating a deformed video frame.

[0021] Furthermore, during the process of adjusting the brightness of the intermediate video frames, color jitter is first applied to the deformed video frames. For example, the brightness of the deformed video frames is randomly increased by 15% and the contrast is decreased by 20%. Then, motion blur effect is superimposed to simulate the ghosting effect when the video is playing dynamically, so that the color and light and shadow texture of the frames change, and finally the brightness of the adjusted video frames are obtained.

[0022] Furthermore, Gaussian noise and JPEG compression are applied to the photometric adjustment video frames to interfere with their signal. For example, Gaussian noise is first added to the photometric adjustment video frames to generate random pixel noise with a mean of 0 and a variance of 0.02. The generated pixel noise is then added to the pixel values ​​of the photometric adjustment video frames. Next, JPEG compression is performed on the noisy photometric adjustment video frames, with a compression quality of 50 selected.

[0023] Furthermore, during the independent domain transformation of the interfering intermediate frames, the affine transformation, elastic distortion, photometric adjustment, and signal interference operations are combined into an independent enhancement function. For example, "30° rotation + 20% contrast adjustment + JPEG50 compression" is integrated into a transformation process. This combined operation is performed on the interfering video frames, and the corresponding enhanced domain view is finally output.

[0024] In this embodiment of the invention, affine transformation and elastic distortion of the original video frames are used to simulate the spatial deformation of the video frames during shooting or transmission, such as rotation, stretching, and translation, increasing the spatial diversity of the samples and enabling subsequent learning of more robust spatial features. By adjusting the photometric value of the video frames, the effects of video frames under different lighting, contrast, and color environments are simulated, compensating for the lack of photometric uniformity in the original video frames. By interfering with the video frames through signal interference, the enhanced video frames are made closer to the degradation situation in actual applications, enhancing the model's anti-interference ability. By interfering with intermediate frames through independent domain transformation, adversarial training can be used to make the model ignore the differences in the enhancement domain and focus on the true and false features of the video frames.

[0025] S2. Convert the enhanced domain view into a high-dimensional feature vector of the view, and normalize the high-dimensional feature vector of the view to obtain a normalized feature vector of the view.

[0026] In the medical and health field, a feature extraction network specifically designed for medical imaging is used to extract high-dimensional features such as "lesion edges, tissue texture, and blood flow signals" from surgical endoscopy enhanced frames and dynamic ultrasound enhanced images. These high-dimensional features are then normalized to eliminate feature scale differences caused by different enhancement operations (such as brightness adjustment and noise addition), making the features of the same lesion more consistent under different enhanced views, thereby stably identifying fake lesions.

[0027] In the fintech field, a dedicated facial / ID feature extraction network is used to extract high-dimensional features such as facial key points, document text outlines, and liveness trajectories from enhanced video frames, ID cards, and facial linkage enhanced frames used in remote account opening. Cosine normalization is then applied to these high-dimensional features to eliminate feature distribution shifts caused by different enhancement operations (such as exposure adjustment and background interference). This makes the "identity features of the same user under different enhanced views" more consistent, preventing deepfake faces from "disguising" themselves as real identities due to enhancement operations, and improving the security of remote identity verification.

[0028] In this embodiment of the invention, converting the enhanced domain view into a high-dimensional feature vector of the view includes: Perform convolution operations on the enhanced domain view to generate initial low-dimensional features; The initial low-dimensional features are convolved and stacked to generate intermediate high-dimensional features; The intermediate high-dimensional features are fully connected and mapped to obtain the view high-dimensional feature vector.

[0029] In this embodiment of the invention, during the convolution operation on the enhanced domain view, a preset convolution kernel (such as a 3×3 kernel) is invoked. The weighted sum of local pixels in the enhanced domain view is calculated by sliding the convolution kernel on the enhanced domain view, thereby forming a numerical matrix. Each value in the numerical matrix is ​​substituted into the ReLU activation function formula for processing, and the processed matrix is ​​used as the initial low-dimensional feature.

[0030] Furthermore, in the process of convolutional stacking of the initial low-dimensional features, the initial low-dimensional features are input into multiple concatenated convolutional layers, and multiple rounds of convolution and pooling stacking operations are performed. Each round of convolution extracts more complex combined features, while the pooling operation reduces the feature dimension and retains key information. After multiple rounds of stacking, the output is a fused intermediate high-dimensional feature.

[0031] Furthermore, the intermediate high-dimensional features are flattened into a one-dimensional feature matrix and input into a fully connected layer. The fully connected layer maps the flattened one-dimensional feature matrix to a feature space of a preset dimension through a preset weight matrix, and finally outputs a high-dimensional feature vector of the view.

[0032] In this embodiment of the invention, the step of normalizing the high-dimensional feature vector of the view to obtain the normalized feature vector of the view includes: The high-dimensional feature vector of the view is matched according to the preset input layer dimension to obtain the matching feature vector; Perform a fully connected layer operation on the matched feature vector to generate an initial feature vector; The initial feature vector is normalized by its magnitude to obtain the view-normalized feature vector.

[0033] In this embodiment of the invention, during the process of dimension matching of the high-dimensional feature vector of the view, the target dimension (such as 768 dimensions) is first preset. A dimension mapping matrix is ​​constructed with the current dimension of the high-dimensional feature vector of the view as the number of rows and the target dimension as the number of columns. The high-dimensional feature vector of the view is multiplied by the dimension mapping matrix to finally obtain the dimension-adapted matching feature vector.

[0034] Furthermore, during the fully connected layer operation on the matching feature vector, the matching feature vector is input into the first fully connected layer. The first fully connected layer performs a linear transformation on the matching feature vector. The transformed matching feature vector is then substituted into the activation function to filter out negative values ​​and retain nonlinear features. The nonlinearly transformed matching feature vector is then passed into the second fully connected layer, and the linear transformation operation is repeated to output the initial feature vector.

[0035] Furthermore, in the process of normalizing the initial feature vector, the magnitude of the initial feature vector is first calculated, and each element in the vector is divided by this magnitude to make the magnitude of the processed vector 1, so that the feature distribution is more stable and the view-normalized feature vector is obtained.

[0036] In this embodiment of the invention, initial low-dimensional features are obtained through convolution operations to extract basic visual features such as image edges and textures, while reducing the data dimensionality and subsequent computational load. Intermediate high-dimensional features are obtained through convolution stacking, and multiple convolutions gradually fuse basic features to generate more complex semantic features (such as object contours and local structures), thereby improving the expressive power of the features. A high-dimensional feature vector of the view is obtained through fully connected mapping, which maps the high-dimensional semantic features to a fixed-dimensional vector space, unifies the feature format, and facilitates subsequent tasks such as comparison and classification.

[0037] In this embodiment of the invention, the model input error caused by dimension mismatch is eliminated by adjusting the high-dimensional feature vector of the view to the same dimension as the preset input layer; the matching feature vector is reorganized and enhanced by linear transformation and nonlinear activation of multiple fully connected layers to extract more discriminative deep semantic features and improve the expressive power of the features; the magnitude of the initial feature vector is standardized to 1 to eliminate the scale difference between different feature vectors, making the feature distribution more stable, and improving the accuracy of subsequent adversarial training and forgery probability detection.

[0038] S3. Perform enhanced domain label analysis on the normalized feature vector of the view to obtain the vector enhanced domain label.

[0039] In the medical and health field, if label 1 is "enhanced by differences in endoscope equipment", label 2 is "enhanced by transmission distortion of PACS system", and label 3 is "enhanced by interference in surgical scene", the normalized feature vector of the view is used to identify the corresponding enhancement operation traces (such as "equipment noise distribution" in the feature corresponding to label 1), and the vector enhancement domain label is output to help determine the source of enhancement of medical images and avoid misjudging "normal variations caused by equipment differences" as forgery.

[0040] In fintech scenarios, if label 1 is "enhanced lighting of mobile phone camera", label 2 is "enhanced compression of video call", and label 3 is "enhanced background interference for liveness detection" (corresponding to real data collection scenarios for financial identity verification), the feature vector is identified by its "light distortion features" and "compression block effect features", and the corresponding enhancement domain labels are matched to distinguish between "enhanced real data collection scenarios" and "tampering traces in fake videos", thereby improving the accuracy of identity verification.

[0041] In this embodiment of the invention, the step of performing augmented domain label analysis on the view normalized feature vector to obtain vector augmented domain labels includes: Calculate the original score of the normalized feature vector of the view; The augmentation domain probability value of the view normalized feature vector is calculated based on the original score; Select the highest probability value among the augmentation domain probability values, and use the selected probability value as the initial augmentation domain analysis label; The initial augmented domain analysis label is matched with the preset augmented domain numbering strategy to obtain the vector augmented domain label.

[0042] In this embodiment of the invention, the original score refers to the K values ​​directly output by the view normalized feature vector after it has been processed by the domain discriminator (3 fully connected layers). Each value corresponds to an enhancement domain, representing the degree of matching between the feature vector and the corresponding enhancement domain. The larger the value, the more likely the view normalized feature vector is to belong to the enhancement domain.

[0043] In this embodiment of the invention, during the calculation of the original score of the view normalized feature vector, the view normalized feature vector is first multiplied by a preset first-layer weight matrix (dimension 512×256), and then the bias vector of this layer is added to obtain an intermediate vector. This intermediate vector is then substituted into an activation function to filter out negative values ​​and retain non-linear features. The intermediate vector retaining non-linear features is then multiplied by a preset second-layer weight matrix (e.g., dimension 256×K, where K is the total number of enhancement domains) and a bias is added to obtain a K-dimensional intermediate vector. After being processed by an activation function, it is passed to the third layer, where a final linear transformation (without an activation function) is performed on the K-dimensional intermediate vector output from the second layer, directly outputting K values. These K values ​​are used as the original scores of the K enhancement domains corresponding to the view normalized feature vector, with each value representing the degree of matching between the feature vector and the corresponding enhancement domain.

[0044] An enhancement domain refers to the set of features of all generated video frames after applying a fixed combination of enhancement operations to the original video frame. For example, "30° affine rotation + 15% brightness adjustment + 0.02 Gaussian noise with variance + JPEG50 compression" defines an enhancement domain. Video frames within the same enhancement domain will have similar enhancement features. Video frames in different enhancement domains have significant feature differences, which makes it easier for the model to distinguish different enhancement scenarios.

[0045] For example, if the total number of augmentation domains K=3 (corresponding to 3 combinations of augmentation operations), the view normalized feature vector is 512-dimensional. Multiplying the 512-dimensional view normalized feature vector by a 512×256 weight matrix and adding a bias, we get a 256-dimensional vector. Then, we use the ReLU activation function to turn all the negative values ​​into 0, resulting in a processed 256-dimensional intermediate vector. Subsequently, we multiply the 256-dimensional intermediate vector by a "256×3" weight matrix and add a bias, resulting in a 3-dimensional intermediate vector. After ReLU processing, the effective features are retained. Finally, we perform a final linear transformation on this 3-dimensional intermediate vector, outputting 3 values, such as [2.1, 0.8, 1.5]. These 3 values ​​are the original scores.

[0046] Furthermore, in the process of calculating the augmentation domain probability value of the view normalized feature vector based on the original scores, the Softmax function is performed on the K original scores to convert each score into a probability value between 0 and 1, and the sum of all probability values ​​is 1. The probability value represents the likelihood that the view normalized feature vector belongs to each augmentation domain. Then, the augmentation domain corresponding to the largest probability value is selected from the K augmentation domain probability values, and this probability value is used as the initial augmentation domain analysis label.

[0047] Furthermore, during the process of matching the initial enhanced domain analysis label domain with the preset enhanced domain numbering strategy, the preset enhanced domain numbering strategy is invoked ( This represents the augmentation field number corresponding to the i-th sample. This indicates the range of values ​​for the number. The numbering rule (representing the total number of preset augmentation domains) stipulates that "augmentation domain 1 corresponds to number 1, augmentation domain 2 corresponds to number 2, and augmentation domain 3 corresponds to number 3". The augmentation domains corresponding to the initial augmentation domain analysis labels are mapped to the corresponding numbers, which are the vector augmentation domain labels.

[0048] For example, first substitute the original score [2.1, 0.8, 1.5] into the Softmax function to calculate the probability values ​​of the augmentation domains: [0.6, 0.1, 0.3]. These probability values ​​correspond to augmentation domain 1, augmentation domain 2, and augmentation domain 3, respectively. The largest probability value is 0.6, and the augmentation domain corresponding to the largest probability value of 0.6 is augmentation domain 1. Use augmentation domain 1 as the initial augmentation domain analysis label, match the initial label "augmentation domain 1" with the numbering strategy, and the final vector augmentation domain label is "1".

[0049] S4. Perform source obfuscation processing on the view normalized feature vector according to the vector enhancement domain label to obtain a stable feature vector.

[0050] In this embodiment of the invention, the step of performing augmentation domain source obfuscation processing on the view normalized feature vector based on the vector augmentation domain label to obtain a stable feature vector includes: Based on the vector enhancement domain labels, the view normalized feature vectors are grouped by enhancement domain number to obtain an enhancement domain feature vector grouping table; Extract positive sample feature pairs from the enhanced domain feature vector grouping table and calculate the similarity of the positive sample feature pairs; Based on the similarity, negative sample features are extracted from the enhanced domain feature vector grouping table; Weights are assigned to the positive sample feature pairs and the negative sample features using preset hyperparameters, and the positive sample feature pairs and the negative sample features are fused into preliminary confusion features according to the assigned weights; The distribution compactness of the initial confusion features is adjusted to obtain intermediate confusion features; Obtain the true augmented domain label corresponding to the original score, and perform domain obfuscation verification on the intermediate obfuscated features based on the true augmented domain label to obtain a stable feature vector that passes the verification.

[0051] In this embodiment of the invention, during the process of grouping the view normalized feature vector enhancement domain numbers according to the vector enhancement domain labels, the view normalized feature vectors are classified according to the vector enhancement domain label numbers. For example, label 1 corresponds to the enhancement domain feature of "30° affine rotation + brightness enhancement", and label 2 corresponds to the feature of "Gaussian blur + contrast enhancement", thus forming an enhancement domain feature vector grouping table.

[0052] Furthermore, in the process of extracting positive sample feature pairs from the grouping table of enhancement domain feature vectors, based on the positive sample pair rules of cross-domain contrastive learning, two types of positive sample feature pairs are extracted from the grouping table, including different enhanced views of the same sample (e.g., the features of the original video frame A after processing by enhancement domains 1 and 2 to form a positive sample feature pair) or adjacent frame features of the same video frame (e.g., the enhanced view features of the 5th and 6th frames of the video). The cosine similarity algorithm is used to calculate the similarity of each positive sample feature pair.

[0053] Furthermore, according to the preset dynamic adaptive negative sample strategy, in the enhanced domain feature vector grouping table, all features except the original samples to which the positive sample feature pairs belong are traversed, the cosine similarity between the feature and the positive sample feature pairs is calculated, features with similarity higher than a set threshold are selected as negative sample features, and features with high similarity to positive sample features in historical batches are extracted and added to the selected negative sample features, and the negative sample features are output.

[0054] Furthermore, utilizing the hyperparameters of gradient reversal Weights are assigned based on the similarity between positive and negative samples, with positive samples having a higher weight. Negative sample feature weights Based on the weighted fusion features, the specific formula is as follows:

[0055]

[0056] in, This refers to hyperparameters. This refers to the training progress. This refers to controlling and adjusting the rate. This refers to positive sample pairs of features. This refers to negative sample features. This refers to the initial confusing characteristics.

[0057] Furthermore, in the process of adjusting the distribution compactness of the initial confusion features, the temperature coefficient can first be calculated using the following formula:

[0058] in, This refers to the training progress. It refers to the temperature coefficient.

[0059] Furthermore, for each feature vector in the initial confusion feature set, its cosine similarity value with other features in the set is calculated, and the cosine similarity value is substituted into the temperature coefficient. The scale is applied, and the distribution of the initial confusing features is normalized and calibrated based on the scaled similarity to obtain intermediate confusing features.

[0060] In this embodiment of the invention, the true enhancement domain label is directly labeled by the enhancement operation combination corresponding to the original video frame, and corresponds one-to-one with the view normalized feature vector. For example, three sets of enhancement operation combinations (K=3) are set and assigned numbers. Enhancement domain 1: operation combination = "30° affine rotation + 15% brightness enhancement" is numbered = 1; Enhancement domain 2: operation combination = "Gaussian blur (σ=0.5) + 20% contrast enhancement" is numbered = 2; Enhancement domain 3: operation combination = "JPEG compression (quality 50) + color shift (R channel + 10)" is numbered = 3. The operation combination of "enhancement domain 2" is randomly selected to process the original video frame A, resulting in enhanced video frame A'. Since this sample uses the operation of "enhancement domain 2", the number "2" is directly labeled as the true enhancement domain label of this sample. "2" is the true enhancement domain label corresponding to the original score (the result output by the subsequent domain discriminator).

[0061] Furthermore, in the process of performing domain obfuscation verification on the intermediate obfuscation features based on the real augmented domain labels, the intermediate obfuscation features are first input into the domain discriminator D. The domain discriminator D outputs a probability vector of length K, where K is the total number of augmented domains. The probability vector containing the intermediate obfuscation features that match the real augmented domain labels is then extracted. The value at the corresponding position will be used as the predicted probability. .

[0062] For example, if an original video frame A uses an operation combination of augmentation domain 2, then its true augmentation domain label... =2, input the intermediate obfuscation feature into the domain discriminator D. The domain discriminator outputs a probability vector of length 3: [0.1, 0.8, 0.1]; the three values ​​of this vector correspond to the predicted probabilities of the feature belonging to augmentation domain 1, augmentation domain 2, and augmentation domain 3, respectively. The true augmentation domain label of this intermediate obfuscation feature is... The value 0.8 is extracted from the second position of the probability vector and used as the prediction probability. That is, the prediction probability that the intermediate confusion feature belongs to the true enhancement domain 2 is 80%.

[0063] Subsequently, the cross-entropy loss formula was used to calculate the difference between the true augmentation domain label and the domain discriminator's predicted probability that the intermediate confusing feature belongs to the true augmentation domain. The specific formula is as follows:

[0064] in, Indicates batch size, The domain discriminator represents the predicted probability that the intermediate confusing features belong to the true augmentation domain.

[0065] like If the difference threshold is less than or equal to the preset threshold, then the intermediate confusion vector is normalized twice to obtain a stable feature vector. If the preset difference threshold is reached, the weights of the positive sample feature pairs and the negative sample features will be redistributed and then re-fused until the verification is successful.

[0066] S5. Detect the forgery probability of the original video frame based on the stable feature vector.

[0067] In this embodiment of the invention, detecting the forgery probability of the original video frame based on the stable feature vector includes: The stable feature vector is dimension-mapped to obtain a dimension-adapted feature vector; Activation function operations are performed on the dimension-adapted feature vectors to generate intermediate feature vectors; The intermediate feature vector is converted into an initial forgery probability value, and the initial forgery probability value is corrected for deviation using a preset probability calibration parameter to obtain the forgery probability of the original video frame.

[0068] In this embodiment of the invention, during the process of dimensional mapping of stable feature vectors, the target dimension for dimensional adaptation is first determined according to the requirements of the preset classification task (e.g., binary classification "fake / real"). The target dimension is 2. The stable feature vector is input into the fully connected layer according to the target dimension. The dimensional transformation is completed by matrix multiplication (stable feature vector × fully connected layer weight + bias), and the dimension-adapted feature vector is output.

[0069] Furthermore, during the activation function operation on the dimension-adapted feature vector, the dimension-adapted feature vector is input into a preset activation function, and a non-linear transformation is performed on each element of the vector (for example, Softmax will map the vector elements to the interval of 0~1, and the sum of the elements is 1), and an intermediate feature vector is output (for example, a 2-dimensional vector is obtained as [0.2,0.8] after Softmax).

[0070] Furthermore, in the process of converting the intermediate feature vector into the initial forgery probability value, if the intermediate feature vector corresponds to the binary classification of "forgery / real", the element value corresponding to the "forgery" category is directly extracted as the initial forgery probability value. Using the preset probability calibration parameters (e.g., calibration parameters α=0.95, β=0.02) learned through the validation set data, the initial forgery probability value is corrected using the calibration parameters (e.g., the correction formula is: final forgery probability = initial forgery probability × α + β), and the forgery probability of the original video frame is output.

[0071] For example, the intermediate feature vector is a binary classification vector [0.15, 0.85] after Softmax activation (corresponding to index 0 of the "real" category and index 1 of the "fake" category). The element value 0.85 corresponding to the "fake" category is directly extracted as the initial forgery probability value. Then, the preset probability calibration parameters are α=0.95 and β=0.02, which are substituted into the correction formula: the final forgery probability is 0.8275.

[0072] In this embodiment of the invention, dimension-adapted feature vectors are obtained through dimension mapping, adjusting the dimensions of stable features to those required for subsequent tasks, adapting to the input requirements of downstream classification modules, and avoiding dimension mismatch problems. By performing activation function operations on the dimension-adapted feature vectors, the features are mapped to a reasonable range, enhancing the discriminativeness of the features and making subsequent probability calculations more intuitive. By first extracting initial probabilities and then correcting deviations with calibration parameters, both the semantic rationality of the probabilities and the accuracy of the probabilities are ensured.

[0073] As can be seen, in the above scheme, for the forgery probability detection, the original video frames published by the user are acquired, and the original video frames are dynamically enhanced to generate an enhanced domain view; the enhanced domain view is converted into a high-dimensional feature vector of the view, and the high-dimensional feature vector of the view is normalized and mapped to obtain a normalized feature vector of the view; the normalized feature vector of the view is subjected to enhanced domain label analysis to obtain vector enhanced domain labels; the normalized feature vector of the view is subjected to enhanced domain source obfuscation processing based on the vector enhanced domain labels to obtain a stable feature vector; the forgery probability of the original video frame is detected based on the stable feature vector. By analyzing the enhanced domain labels and obfuscating the enhanced domain sources of the view normalized feature vector, normal enhancement variations are avoided from being misjudged as forgery traces, thus improving the detection accuracy.

[0074] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0075] In one embodiment, a vision-based deepfake detection device is provided, which corresponds one-to-one with the vision-based deepfake detection method described in the above embodiments. For example... Figure 3 As shown, this vision-based deepfake detection device includes a view dynamic enhancement module 101, a feature vector conversion module 102, a feature vector mapping module 103, an enhancement domain label analysis module 104, and a feature vector enhancement domain source obfuscation processing module 105. Detailed descriptions of each functional module are as follows: The view dynamic enhancement module 101 is used to acquire the original video frames published by the user and dynamically enhance the original video frames to generate an enhanced domain view. Feature vector conversion module 102 is used to convert the enhanced domain view into a high-dimensional feature vector of the view; Feature vector mapping module 103 is used to normalize and map the high-dimensional feature vector of the view to obtain the normalized feature vector of the view. The enhanced domain label analysis module 104 is used to extract product features from the terms and conditions to obtain product features. The feature vector enhancement domain source obfuscation processing module 105 is used to perform enhancement domain source obfuscation processing on the view normalized feature vector according to the vector enhancement domain label to obtain a stable feature vector. The forgery probability detection module 106 is used to detect the forgery probability of the original video frame based on the stable feature vector.

[0076] In one embodiment, the view dynamic enhancement module 101, when dynamically enhancing the original video frame to generate an enhanced domain view, is used to: The original video frame is subjected to affine transformation and elastic warping to form a deformed video frame; The photometric adjustment is performed on the deformed video frame to obtain the photometric adjusted video frame; The photometric adjustment video frame is subjected to signal interference to generate an interfering intermediate frame; An independent domain transformation is performed on the interfering intermediate frame to obtain an enhanced domain view.

[0077] In one embodiment, the feature vector conversion module 102, when converting the enhanced domain view into a high-dimensional feature vector of the view, is used to: Perform convolution operations on the enhanced domain view to generate initial low-dimensional features; The initial low-dimensional features are convolved and stacked to generate intermediate high-dimensional features; The intermediate high-dimensional features are fully connected and mapped to obtain the view high-dimensional feature vector.

[0078] In one embodiment, when the feature vector mapping module 103 performs normalization mapping on the high-dimensional feature vector of the view to obtain the normalized feature vector of the view, it is used to: The high-dimensional feature vector of the view is matched according to the preset input layer dimension to obtain the matching feature vector; Perform a fully connected layer operation on the matched feature vector to generate an initial feature vector; The initial feature vector is normalized by its magnitude to obtain the view-normalized feature vector.

[0079] In one embodiment, the enhanced domain label analysis module 104, when performing enhanced domain label analysis on the view normalized feature vector to obtain vector enhanced domain labels, is used to: Calculate the original score of the normalized feature vector of the view; The augmentation domain probability value of the view normalized feature vector is calculated based on the original score; Select the highest probability value among the augmentation domain probability values, and use the selected probability value as the initial augmentation domain analysis label; The initial augmented domain analysis label is matched with the preset augmented domain numbering strategy to obtain the vector augmented domain label.

[0080] In one embodiment, when the feature vector enhancement domain source obfuscation processing module 105 performs enhancement domain source obfuscation processing on the view normalized feature vector according to the vector enhancement domain label to obtain a stable feature vector, it is used to: Based on the vector enhancement domain labels, the view normalized feature vectors are grouped by enhancement domain number to obtain an enhancement domain feature vector grouping table; Extract positive sample feature pairs from the enhanced domain feature vector grouping table and calculate the similarity of the positive sample feature pairs; Based on the similarity, negative sample features are extracted from the enhanced domain feature vector grouping table; Weights are assigned to the positive sample feature pairs and the negative sample features using preset hyperparameters, and the positive sample feature pairs and the negative sample features are fused into preliminary confusion features according to the assigned weights; The distribution compactness of the initial confusion features is adjusted to obtain intermediate confusion features; Obtain the true augmented domain label corresponding to the original score, and perform domain obfuscation verification on the intermediate obfuscated features based on the true augmented domain label to obtain a stable feature vector that passes the verification.

[0081] In one embodiment, when the forgery probability detection module 106 detects the forgery probability of the original video frame based on the stable feature vector, it is used to: The stable feature vector is dimension-mapped to obtain a dimension-adapted feature vector; Activation function operations are performed on the dimension-adapted feature vectors to generate intermediate feature vectors; The intermediate feature vector is converted into an initial forgery probability value, and the initial forgery probability value is corrected for deviation using a preset probability calibration parameter to obtain the forgery probability of the original video frame.

[0082] This invention provides a deepfake detection device based on visual inspection. For forgery probability detection, it acquires user-submitted original video frames and dynamically enhances them to generate an enhanced domain view. The enhanced domain view is then converted into a high-dimensional feature vector, which is normalized to obtain a normalized feature vector. Enhancement domain label analysis is performed on the normalized feature vector to obtain vector enhancement domain labels. Based on these labels, the normalized feature vector is subjected to enhancement domain source obfuscation to obtain a stable feature vector. The forgery probability of the original video frame is detected based on the stable feature vector. By analyzing the enhancement domain labels and obfuscating the enhancement domain source of the normalized feature vector, the device avoids misidentifying normal enhancement variations as forgery traces, thus improving detection accuracy.

[0083] For specific limitations regarding a vision-based deepfake detection device, please refer to the limitations of a vision-based deepfake detection method described above, which will not be repeated here. Each module in the aforementioned vision-based deepfake detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0084] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a visual detection-based deepfake detection method on the server side.

[0085] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the client-side functions or steps of a vision-based deepfake detection method.

[0086] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain the original video frames published by the user, and dynamically enhance the original video frames to generate an enhanced domain view; The enhanced domain view is converted into a high-dimensional feature vector of the view, and the high-dimensional feature vector of the view is normalized and mapped to obtain a normalized feature vector of the view. Enhanced domain label analysis is performed on the normalized feature vector of the view to obtain vector enhanced domain labels; Based on the vector enhancement domain label, the view normalized feature vector is subjected to enhancement domain source obfuscation processing to obtain a stable feature vector; The probability of forgery of the original video frame is detected based on the stable feature vector.

[0087] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain the original video frames published by the user, and dynamically enhance the original video frames to generate an enhanced domain view; The enhanced domain view is converted into a high-dimensional feature vector of the view, and the high-dimensional feature vector of the view is normalized and mapped to obtain a normalized feature vector of the view. Enhanced domain label analysis is performed on the normalized feature vector of the view to obtain vector enhanced domain labels; Based on the vector enhancement domain label, the view normalized feature vector is subjected to enhancement domain source obfuscation processing to obtain a stable feature vector; The probability of forgery of the original video frame is detected based on the stable feature vector.

[0088] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0089] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0091] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. If any software tools or components other than those of our company appear in the embodiments, they are merely illustrative examples and do not represent actual use. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A deepfake detection method based on visual inspection, characterized in that, include: Obtain the original video frames published by the user, and dynamically enhance the original video frames to generate an enhanced domain view; The enhanced domain view is converted into a high-dimensional feature vector of the view, and the high-dimensional feature vector of the view is normalized and mapped to obtain a normalized feature vector of the view. Enhanced domain label analysis is performed on the normalized feature vector of the view to obtain vector enhanced domain labels; Based on the vector enhancement domain label, the view normalized feature vector is subjected to enhancement domain source obfuscation processing to obtain a stable feature vector; The probability of forgery of the original video frame is detected based on the stable feature vector.

2. The deepfake detection method based on visual detection as described in claim 1, characterized in that, The step of dynamically enhancing the original video frame to generate an enhanced domain view includes: The original video frame is subjected to affine transformation and elastic warping to form a deformed video frame; The photometric adjustment is performed on the deformed video frame to obtain the photometric adjusted video frame; The photometric adjustment video frame is subjected to signal interference to generate an interfering intermediate frame; An independent domain transformation is performed on the interfering intermediate frame to obtain an enhanced domain view.

3. The deepfake detection method based on visual inspection as described in claim 1, characterized in that, The step of performing enhanced domain label analysis on the normalized feature vector of the view to obtain vector enhanced domain labels includes: Calculate the original score of the normalized feature vector of the view; The augmentation domain probability value of the view normalized feature vector is calculated based on the original score; Select the highest probability value among the augmentation domain probability values, and use the selected probability value as the initial augmentation domain analysis label; The initial augmented domain analysis label is matched with the preset augmented domain numbering strategy to obtain the vector augmented domain label.

4. The deepfake detection method based on visual detection as described in claim 1, characterized in that, The step of performing augmentation domain source obfuscation processing on the view normalized feature vector based on the vector augmentation domain label to obtain a stable feature vector includes: Based on the vector enhancement domain labels, the view normalized feature vectors are grouped by enhancement domain number to obtain an enhancement domain feature vector grouping table; Extract positive sample feature pairs from the enhanced domain feature vector grouping table and calculate the similarity of the positive sample feature pairs; Based on the similarity, negative sample features are extracted from the enhanced domain feature vector grouping table; Weights are assigned to the positive sample feature pairs and the negative sample features using preset hyperparameters, and the positive sample feature pairs and the negative sample features are fused into preliminary confusion features according to the assigned weights; The distribution compactness of the initial confusion features is adjusted to obtain intermediate confusion features; Obtain the true augmented domain label corresponding to the original score, and perform domain obfuscation verification on the intermediate obfuscated features based on the true augmented domain label to obtain a stable feature vector that passes the verification.

5. The deepfake detection method based on visual detection as described in claim 1, characterized in that, The step of detecting the forgery probability of the original video frame based on the stable feature vector includes: The stable feature vector is dimension-mapped to obtain a dimension-adapted feature vector; Activation function operations are performed on the dimension-adapted feature vectors to generate intermediate feature vectors; The intermediate feature vector is converted into an initial forgery probability value, and the initial forgery probability value is corrected for deviation using a preset probability calibration parameter to obtain the forgery probability of the original video frame.

6. The deepfake detection method based on visual detection as described in claim 1, characterized in that, The step of converting the enhanced domain view into a high-dimensional feature vector of the view includes: Perform convolution operations on the enhanced domain view to generate initial low-dimensional features; The initial low-dimensional features are convolved and stacked to generate intermediate high-dimensional features; The intermediate high-dimensional features are fully connected and mapped to obtain the view high-dimensional feature vector.

7. The deepfake detection method based on visual inspection as described in claim 1, characterized in that, The step of normalizing the high-dimensional feature vector of the view to obtain the normalized feature vector of the view includes: The high-dimensional feature vector of the view is matched according to the preset input layer dimension to obtain the matching feature vector; Perform a fully connected layer operation on the matched feature vector to generate an initial feature vector; The initial feature vector is normalized by its magnitude to obtain the view-normalized feature vector.

8. A deepfake detection device based on visual inspection, characterized in that, include: The view dynamic enhancement module is used to acquire the original video frames published by the user and dynamically enhance the original video frames to generate an enhanced domain view. The feature vector conversion module is used to convert the enhanced domain view into a high-dimensional feature vector of the view; The feature vector mapping module is used to normalize and map the high-dimensional feature vectors of the view to obtain normalized feature vectors of the view. The enhanced domain label analysis module is used to perform enhanced domain label analysis on the normalized feature vector of the view to obtain vector enhanced domain labels. The feature vector enhancement domain source obfuscation processing module is used to perform enhancement domain source obfuscation processing on the view normalized feature vector according to the vector enhancement domain label to obtain a stable feature vector. The forgery probability detection module is used to detect the forgery probability of the original video frame based on the stable feature vector.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the deepfake detection method based on visual detection as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the deepfake detection method based on visual detection as described in any one of claims 1 to 7.