A Remote Interview Anti-Fraud System and Method Based on Multimodal Deepfake Detection
By employing multimodal deepfake detection, hashing job seeker identities and biometrics, and combining blockchain storage and dynamic decision-making, the system addresses identity verification vulnerabilities and deepfake issues in remote recruitment. This achieves efficient identity verification and fraud detection, ensuring the security and privacy of the recruitment system.
Patent Information
- Application Number
- CN202511271883.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing remote recruitment systems suffer from vulnerabilities in identity verification, insufficient defense against deepfakes, and rigid risk decision-making, leading to frequent fraudulent incidents involving virtual job seekers. Furthermore, existing solutions are prone to privacy leaks and harming genuine users.
A multimodal deepfake detection method is adopted. By hashing job seeker identity and biometric information and combining it with blockchain storage, the method performs live physiological signal detection, deepfake spatiotemporal analysis, voiceprint frequency domain diagnosis and cross-modal synchronous verification, dynamically adjusts confidence weights, constructs an RNN neural network for behavior analysis, and finally executes a hierarchical response.
It effectively identifies identity theft and deepfakes while ensuring privacy and security, reduces false positives, improves recruitment security and efficiency, and provides an end-to-end trusted solution.
Smart Images

Figure CN120805071B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data analysis technology, and specifically relates to a remote interview anti-fraud system and method based on multimodal deepfake detection. Background Technology
[0002] With the widespread adoption of remote recruitment, the misuse of AI deepfake technologies (such as Deepfake and voice cloning) has led to a surge in "virtual job seeker" fraud incidents. Existing solutions suffer from several problems: centralized storage of raw biometric data can easily lead to privacy breaches; modules such as liveness detection, forgery analysis, and lip-syncing operate independently, lacking collaborative decision-making; manual review is inefficient, while automated interception can easily mislead genuine users; the industry still needs a fraud prevention system that integrates zero-trust architecture, dynamic multimodal analysis, and adaptive decision-making.
[0003] Existing technologies suffer from vulnerabilities in identity verification, insufficient defense against deepfakes, and rigid risk decision-making. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] In response to the problems in related technologies, this invention provides a remote interview anti-fraud method based on multimodal deepfake detection to overcome the aforementioned technical problems existing in the existing related technologies.
[0006] (II) Technical Solution
[0007] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:
[0008] S1. Collect job seekers' identity information and biometrics; hash the job seekers' identity information and biometrics and then perform identity matching; after the matching is completed, upload it to the blockchain for storage.
[0009] S2. Using multimodal dynamic detection combined with data stored in the blockchain, live physiological signal detection, deepfake spatiotemporal analysis, voiceprint frequency domain diagnosis, and cross-modal synchronous verification are performed on job seekers to obtain physiological feature pairs, abnormality index, synthetic feature value, and maximum offset.
[0010] S3. Based on the physiological feature pairs, abnormality index, synthetic feature value, and maximum offset, calculate the liveness confidence, forgery detection confidence, and cross-modal confidence.
[0011] The weights of the liveness confidence, forgery detection confidence, and cross-modal confidence are dynamically adjusted to obtain a fusion score;
[0012] S4. Construct an RNN neural network, and use historical remote interview behavior feature data combined with optimization algorithms to optimize the RNN neural network to obtain an RNN behavior analysis network.
[0013] Use a behavioral analysis network to obtain a real-time index of abnormal behavior in job seekers;
[0014] S5. Based on the fusion score and the real-time abnormal behavior index, obtain the final risk score; execute the tiered response plan based on the final risk score.
[0015] This invention, with zero-knowledge verification, multimodal collaborative perception, and dynamic decision-making as its core, overcomes the technical challenges of identity theft, deepfakes, and behavioral fraud in remote interviews. It effectively solves the problems of identity theft, deepfakes, and behavioral fraud while ensuring privacy and security, and provides an end-to-end trusted solution for recruitment security.
[0016] Preferably, step S1 includes the following steps:
[0017] S11. Collect high-resolution images of job seekers' ID cards, extract names and ID numbers using OCR to obtain job seeker ID information; hash the job seeker ID information to obtain job seeker information hash;
[0018] S12. Collect photoplethysmography signal data, 3D structured light face point cloud data, and voiceprint sample data of job seekers through the encrypted SDK to obtain the job seekers' biometric features.
[0019] S13. Perform anti-collision fuzzy hashing on the job seeker's biometric features to obtain job seeker biometric hash pairs; the biometric hash pairs include face hash and voiceprint hash, and the anti-collision fuzzy hashing process includes face hash and voiceprint hash.
[0020] S14. Use zkSNARKs to verify that the biometric hash of the job seeker is matched with the hash of the job seeker's identification information without revealing the original biometric features.
[0021] S15. After a successful match, the job seeker's biometric hash and the job seeker's document information hash are used to construct a triplet to obtain the job seeker information triplet; the job seeker information triplet is written into the Chia blockchain, which includes a BLS signature + spacetime proof mechanism.
[0022] This invention extracts and hashes document information using OCR, employs an encrypted SDK to collect photoplethysmograms, 3D face point clouds, and voiceprint samples, and fuses curvature gradients with physiological frequency band energy to generate collision-resistant face hashes. These hashes are then combined with voiceprint mutation feature hashes to form biometric hash pairs. ZK-SNARKs zero-knowledge proofs are used to verify the match between the document and the biometric hash, and the hash triples are ultimately written into the Chia blockchain. The original biometric data is not stored throughout the process; hashing and zero-knowledge proofs completely eliminate the risk of leakage. The fusion of 3D point cloud curvature gradients and live pulse energy effectively resists photo / mask attacks. The Chia blockchain's spatiotemporal proof ensures the hash triples are immutable, providing an auditable and traceable chain.
[0023] Preferably, step S2 includes the following steps:
[0024] S21. By analyzing the dual-band video stream of the job seeker's face captured by the camera, signals related to physiological activities are extracted to obtain live physiological signals; by combining the live physiological signals with a live physiological detection algorithm, physiological feature pairs are obtained; the physiological feature pairs include signal-to-noise ratio and main pulse frequency.
[0025] S22. Collect consecutive frames of the job seeker's interview video to obtain a video frame sequence; the video frame sequence includes all consecutive frames of the job seeker's real-time interview video;
[0026] Extract the abnormal features of each frame in the video frame sequence to obtain the video frame abnormal feature sequence; the video frame abnormal feature sequence includes the abnormal features of each frame in the video frame sequence.
[0027] An anomaly index is obtained by combining the deepfake anomaly index algorithm with the anomaly feature sequence of video frames;
[0028] S23. Extract the speech phase spectrum of the job seeker's voice signal during the interview, calculate the synthesized speech features, and obtain the synthesized feature values;
[0029] S24. Use the DTW algorithm to verify the lip-sound synchronization of job seekers and obtain the maximum offset.
[0030] This invention extracts subcutaneous blood pulsation signals from dual-band video streams, calculates the signal-to-noise ratio and main pulse frequency to verify physiological liveness; analyzes abnormal features such as micro-expressions and muscle movements in video frame sequences, and calculates an anomaly index based on statistics from real datasets; extracts second-order mixed partial derivatives of the speech phase spectrum to quantize and synthesize feature values; uses the DTW algorithm to align lip shapes with sound event stamps and outputs the maximum offset; dual-band physiological signals accurately lock human blood pulsation, resisting 3D mask attacks; the anomaly index exposes spatiotemporal violations in Deepfake videos; the second-order partial derivative of the phase spectrum captures defects in the synthesized speech audio domain; and the maximum offset quantifies the asynchronous flaws of AI lip sound forgery.
[0031] Preferably, step S3 includes the following steps:
[0032] S31. Based on the physiological feature pairs, abnormality index, synthetic feature value, and maximum offset, calculate the liveness confidence, forgery detection confidence, and cross-modal confidence.
[0033] S32. Collect historical attack rates; combine the liveness confidence, forgery detection confidence, and cross-modal confidence with the historical attack rates, and obtain the liveness confidence weight, forgery detection confidence weight, and cross-modal confidence weight through a dynamic weight adjustment algorithm;
[0034] The liveness confidence, forgery detection confidence, and cross-modal confidence are vectorized to obtain a confidence vector; the liveness confidence weight, forgery detection confidence weight, and cross-modal confidence weight are vectorized to obtain a weight vector.
[0035] The fusion score is calculated by combining the confidence vector with the weight vector.
[0036] This invention achieves intelligent risk decision-making through a dynamic weight fusion mechanism; it quantifies four-dimensional features—physiological liveness, video forgery, voiceprint synthesis, and lip-sync—into confidence levels; it dynamically adjusts the weights of each confidence level based on historical attack rates to construct an adaptive fusion scoring model; it breaks through the limitations of static weights, reduces the false positive rate in novel deepfake attack scenarios, improves real-time defense response speed, and significantly enhances the system's ability to resist evolving attacks.
[0037] Preferably, step S31 includes the following steps:
[0038] S311. Set parameters for the steepness of the control logic function and threshold parameters for the logic function; calculate the liveness confidence level based on the parameters for the steepness of the control logic function and threshold parameters for the logic function, combined with physiological characteristics.
[0039] S312. Set the weight parameters for video anomalies and voice anomalies; calculate the forgery detection confidence level based on the anomaly index, the synthesized feature value, and the weight parameters for video anomalies and voice anomalies.
[0040] S313. Calculate the cross-modal confidence level based on the maximum offset;
[0041] This invention achieves dynamic risk assessment through a three-level confidence quantification system; improves the accuracy of liveness verification by using physiological feature parameterized logic functions for liveness detection; enhances the robustness of forgery identification by fusing video anomaly index and voiceprint synthesis feature values through weight allocation; directly maps cross-modal consistency through lip sound offset; and establishes a multi-dimensional adjustable threshold mechanism to reduce the false alarm rate while increasing the attack detection rate, significantly improving the adaptive capability of the anti-fraud system.
[0042] Preferably, step S4 includes the following steps:
[0043] S41. Collect behavioral data and corresponding behavioral anomaly indices from historical remote interviews of job seekers to obtain historical remote interview behavioral data; extract features from historical remote interview behavioral data to obtain historical remote interview behavioral feature data.
[0044] S42. Construct an RNN neural network, set the initial weight values of the RNN neural network; set the prediction accuracy and prediction accuracy threshold of the RNN neural network;
[0045] The RNN neural network is trained using historical remote interview behavior data. During the training process, the optimal weight values of the RNN neural network are found by combining the training accuracy and the training accuracy threshold through the ant colony algorithm, thus obtaining the optimal solution.
[0046] The optimal solution is used as the weight value of the RNN neural network to obtain the RNN behavior analysis network;
[0047] S43. Collect real-time job seeker interview behavior data and extract features to obtain real-time job seeker behavior feature data;
[0048] Real-time job seeker behavioral characteristic data are input into an RNN behavioral analysis network to obtain a real-time abnormal behavior index.
[0049] This invention achieves dynamic monitoring of abnormal behavior by optimizing the RNN model; it trains the network using historical behavior data and optimizes the weights using the ant colony algorithm to avoid local optima; it analyzes job seekers' behavior in real time to output an anomaly index; it improves the accuracy of anomaly detection, reduces the false alarm rate, improves model training efficiency, and effectively identifies novel deepfake behaviors.
[0050] Preferably, the historical remote interview behavior data includes behavioral data of both abnormal and non-abnormal job seekers.
[0051] Preferably, step S5 includes the following steps:
[0052] S51. Establish a tiered response plan; the tiered response plan is divided into a Level 1 execution plan, a Level 2 execution plan, and a Level 3 execution plan based on the final risk score;
[0053] S52. The final risk score is calculated based on the fusion score and the real-time abnormal behavior index.
[0054] S53. Implement the tiered response plan based on the final risk score;
[0055] This invention uses a quantitative and graded response mechanism to combine dynamic fusion scoring with behavioral anomaly index to calculate the final risk value, and automatically triggers differentiated handling processes based on thresholds; it achieves accurate interception of deepfake attacks while minimizing interference with real users, significantly improving the reliability and execution efficiency of the anti-fraud system.
[0056] The remote interview anti-fraud system based on multimodal deepfake detection is used to implement the aforementioned remote interview anti-fraud method based on multimodal deepfake detection. It includes an identity authentication and blockchain storage module, a multimodal dynamic detection module, a confidence fusion and dynamic scoring module, a behavior analysis network module, and a risk decision and response module.
[0057] The identity authentication and blockchain storage module is responsible for collecting job seekers' identity information and biometrics, generating job seeker information hashes and biometric hashes through a collision-resistant fuzzy hash algorithm; using zkSNARKs technology to implement zero-knowledge proofs for hash matching to ensure privacy and security; after a successful match, the hash triple (document information hash + face hash + voiceprint hash) is written into the Chia blockchain to achieve tamper-proof distributed storage.
[0058] The multimodal dynamic detection module extracts physiological signals from dual-band video streams, calculates the signal-to-noise ratio and main pulse frequency, and verifies whether physiological activities conform to human characteristics; it analyzes the abnormal features of video frame sequences and calculates the abnormality index by combining statistics from real datasets; it extracts the second-order mixed partial derivatives of the speech phase spectrum, calculates synthetic feature values to identify synthetic speech; and it uses the DTW algorithm to align lip shapes with the timestamps of sounds and outputs the maximum offset.
[0059] The confidence fusion and dynamic scoring module generates a liveness confidence score by combining a logical function with physiological features; obtains a forgery detection confidence score by weighted fusion of video anomaly index and speech synthesis feature value; generates a cross-modal confidence score based on the exponential decay function of lip sound offset; and updates the weights of each confidence score in real time by using a dynamic weight adjustment algorithm combined with historical attack rates, finally outputting a fusion score. The calculation process integrates the confidence vector, weight vector, and minimum value adjustment mechanism.
[0060] The behavior analysis network module utilizes historical interview behavior data, optimizes RNN weights using the ant colony algorithm, and outputs the optimal RNN behavior analysis network; it extracts job seeker interview behavior features and inputs them into the trained network to generate a real-time abnormal behavior index.
[0061] The risk decision-making and response module integrates and merges scoring and behavioral anomaly index, calculates the final risk score through weighted average, and executes a graded response based on the scoring threshold.
[0062] (III) Beneficial Effects
[0063] The present invention has the following beneficial effects:
[0064] This invention, with zero-knowledge verification, multimodal collaborative perception, and dynamic decision-making as its core, overcomes the technical challenges of identity theft, deepfakes, and behavioral fraud in remote interviews, providing an end-to-end trusted solution for recruitment security.
[0065] This invention constructs a zero-leakage identity credential chain, achieving a dual breakthrough in trusted evidence storage and privacy protection. Through anti-collision fuzzy hashing technology, it fuses the curvature gradient of 3D face point clouds with the physiological frequency band energy of photoplethysmography to generate irreversible face hashes. Simultaneously, it extracts the maximum value feature of the absolute derivative of voiceprints to generate voiceprint hashes, eliminating the risk of leakage of original biometric data. Combined with zk-SNARKs, it implements "hash matching zero-knowledge proof," concealing sensitive information throughout the process of verifying that document information and biometric features belong to the same person. Using the Chia blockchain to store hash triples, it constructs an immutable and auditable evidence storage chain, meeting the highest privacy compliance standards.
[0066] This invention employs multimodal collaborative detection of deepfakes to accurately identify cross-modal attacks; it innovatively integrates four-dimensional anti-spoofing verification; liveness detection extracts subcutaneous blood pulsation signals through dual-band video streams, using signal-to-noise ratio and main pulse frequency to lock onto physiological live bodies, effectively resisting 3D mask attacks; deepfake detection is based on statistics from real video datasets, quantifying unnatural micro-expressions and muscle movements through anomaly indices; voiceprint diagnosis utilizes the second-order mixed partial derivative of the speech phase spectrum to capture frequency domain defects in synthesized speech; cross-modal verification uses the DTW algorithm to calculate the maximum offset between lip movements and sound events, thoroughly exposing AI-generated fake synchronization attacks.
[0067] This invention employs a dynamic weighted fusion mechanism to achieve adaptive risk decision-making. It constructs a three-layer confidence model; the liveness confidence dynamically responds to physiological signal quality via a logistic function; forgery confidence is weighted and fused with video and speech anomaly indicators; cross-modal confidence is quantified with exponential decay to achieve lip-sync tolerance; a weight evolution algorithm driven by historical attack rates is introduced, enabling the system to automatically increase forgery detection weights during periods of high attack incidence; and finally, the fusion score combined with a minimum adjustment mechanism significantly reduces the false negative rate of complex attacks.
[0068] This invention improves risk control efficiency through a tiered response closed loop; it calculates the final risk value by weighting based on dynamic fusion scoring and behavioral indices; and it implements a multi-level response mechanism, which improves fraud interception rate and interview process efficiency while reducing false positive rate.
[0069] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0070] To more clearly illustrate the technical solutions of the embodiments of the invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the invention. For those skilled in the art, the drawings can be obtained from these drawings without creative effort.
[0071] Figure 1 This is a flowchart illustrating the remote interview anti-fraud method based on multimodal deepfake detection of the present invention.
[0072] Figure 2 This is a flowchart illustrating the multimodal dynamic detection process in the remote interview anti-fraud method based on multimodal deepfake detection of the present invention.
[0073] Figure 3 This is a schematic diagram of the modules of the remote interview anti-fraud system based on multimodal deepfake detection of the present invention. Detailed Implementation
[0074] The technical solutions of the embodiments of the invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the invention, and not all embodiments. Based on the embodiments of the invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the invention.
[0075] In the description of this invention, it should be understood that the terms "opening", "upper", "lower", "top", "middle", "inner", etc., which indicate orientation or positional relationship, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the components or elements referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the invention.
[0076] Example 1
[0077] Please see Figure 1 , Figure 2 This invention discloses a remote interview anti-fraud method based on multimodal deepfake detection, comprising the following steps:
[0078] S1. Collect job seekers' identity information and biometrics; hash the job seekers' identity information and biometrics and then perform identity matching; after the matching is completed, upload it to the blockchain for storage.
[0079] S1 includes the following steps:
[0080] S11. Collect high-resolution images of job seekers' ID cards, extract names and ID numbers using OCR to obtain job seeker ID information; hash the job seeker ID information to obtain job seeker information hash;
[0081] S12. Collect photoplethysmography signal data, 3D structured light face point cloud data (≥30,000 feature points), and voiceprint sample data (obtained by reading random encrypted text) from job seekers through the encryption SDK to obtain the job seeker's biometric features.
[0082] S13. Perform anti-collision fuzzy hashing on the job seeker's biometric features to obtain job seeker biometric hash pairs; the biometric hash pairs include face hash and voiceprint hash, and the anti-collision fuzzy hashing process includes face hash and voiceprint hash.
[0083] The face hash formula and voiceprint hash formula are as follows.
[0084] ;
[0085] in, H face Represents face hash, BLAKE 3() represents a hash function. d Represents the gradient operator. e This represents the curvature of the point cloud in 3D structured light face point cloud data. d ε (P 3D This represents the curvature gradient of a 3D structured light face point in cloud computing. β { S rPPG} represents the complex spectrum of the output photoplethysmogram signal data. df Represents the frequency differential variable. This represents the total energy of the photoplethysmogram signal data within the 0.82.0 Hz frequency band, used as a physiological biomarker.
[0086] in, H voice This represents the voiceprint hash, and Sphincs+ represents the hash function used for voiceprint hashing. V MFCC This represents voiceprint sample data. This represents the maximum value of the absolute derivative over the entire time series (used to extract the most significant variation features in voiceprints).
[0087] S14. Use zkSNARKs to verify that the biometric hash of the job seeker is matched with the hash of the job seeker's identification information without revealing the original biometric features.
[0088] S15. After a successful match, the job seeker's biometric hash and the job seeker's document information hash are used to construct a triplet to obtain the job seeker information triplet; the job seeker information triplet is written into the Chia blockchain, which includes a BLS signature + spacetime proof mechanism.
[0089] S2. Using multimodal dynamic detection combined with data stored in the blockchain, live physiological signal detection, deepfake spatiotemporal analysis, voiceprint frequency domain diagnosis, and cross-modal synchronous verification are performed on job seekers to obtain physiological feature pairs, abnormality index, synthetic feature value, and maximum offset.
[0090] S2 includes the following steps:
[0091] S21. By analyzing the dual-band video stream of the job seeker's face captured by the camera, signals related to physiological activities are extracted to obtain live physiological signals; a live physiological detection algorithm is used to combine the live physiological signals to obtain physiological feature pairs; the physiological feature pairs include signal-to-noise ratio and main pulse frequency, used to verify whether the blood pulsation frequency is within the human physiological range of 0.82Hz; the formula for the live physiological detection algorithm is as follows.
[0092] ;
[0093] in, SNR It represents the signal-to-noise ratio (i.e., the ratio of the intensity of signal fluctuations caused by blood pulsation to the intensity of ambient noise). I 530 , I 660 These represent the 530nm and 660nm wavelength signals in in vivo physiological signals, respectively. s Indicates standard deviation, f p Indicates the main pulse rate, This indicates that the frequency value corresponding to the peak value of the power spectrum is taken;
[0094] S22. Collect consecutive frames from the job seeker's interview video to obtain a video frame sequence. ,in v i This indicates the first video of the job applicant's interview. i frame, N This indicates the total number of frames in the video; the video frame sequence includes all consecutive frames of the job seeker's real-time interview video.
[0095] Extracting abnormal features from each frame in a video frame sequence yields a video frame abnormal feature sequence; the video frame abnormal feature sequence includes abnormal features from each frame in the video frame sequence; the abnormal features include micro-expression duration, facial muscle movement amplitude, and head rotation acceleration;
[0096] An anomaly index is obtained by combining the deepfake anomaly index algorithm with the anomaly feature sequence of video frames; the formula for the deepfake anomaly index algorithm is as follows.
[0097] ;
[0098] in, T abn Indicates an abnormality index. N This represents the total number of video frames in the video frame sequence. l i Indicates the first element in the video frame anomaly feature sequence. i One abnormal feature value, This represents the mean of the i-th feature in a real human video dataset. Indicates the first i The standard deviation of each feature in a real human video dataset;
[0099] S23. Extract the speech phase spectrum of the job seeker's voice signal during the interview, calculate the synthesized speech features, and obtain the synthesized feature values; the calculation formula for the synthesized speech features is as follows.
[0100] ;
[0101] Where Pfake represents a synthetic feature value. f Indicates frequency, t Indicates time, ηft Indicates frequency f and time t The partial derivative, P The represented speech phase spectrum. The second-order mixed partial derivative of the speech phase spectrum. d To represent the differential;
[0102] S24. The DTW algorithm is used to verify lip-sync of job seekers and obtain the maximum offset; the calculation formula is as follows.
[0103] ;
[0104] in, E max Indicates the maximum offset. k This indicates the sequence number of the event during the job seeker's interview process. t k lip Indicates the first k The timestamps of significant changes in the lower lip shape during the event. t k voice Indicates the first k Timestamps of significant changes in sound during each event;
[0105] S3. Based on the physiological feature pairs, abnormality index, synthetic feature value, and maximum offset, calculate the liveness confidence, forgery detection confidence, and cross-modal confidence.
[0106] The weights of the liveness confidence, forgery detection confidence, and cross-modal confidence are dynamically adjusted to obtain a fusion score;
[0107] S31 includes the following steps:
[0108] S31. Based on the physiological feature pairs, abnormality index, synthetic feature value, and maximum offset, calculate the liveness confidence, forgery detection confidence, and cross-modal confidence.
[0109] S31 includes the following steps:
[0110] S311. Set parameters for the steepness of the control logic function and the threshold parameters of the logic function; calculate the liveness confidence score based on the parameters for the steepness of the control logic function and the threshold parameters of the logic function, combined with physiological characteristics; the calculation formula is as follows.
[0111] ;
[0112] Among them, C live Indicates the confidence level in living individuals. oh A parameter indicating the steepness of the set control logic function. x This indicates the threshold parameter for setting the logic function;
[0113] S312. Set weight parameters for video anomalies and voice anomalies; calculate the forgery detection confidence level based on the anomaly index, synthesized feature value, and the weight parameters for video and voice anomalies; the calculation formula is as follows.
[0114] ;
[0115] in, C fake This indicates a falsified test confidence level. x 1 indicates the weight parameter for video anomalies. x 2 represents the weighting parameter for speech anomalies;
[0116] S313. Calculate the cross-modal confidence score based on the maximum offset; the calculation formula is as follows.
[0117] ;
[0118] in, C sync Indicates cross-modal confidence. g This parameter represents the rate of exponential decay.
[0119] S32. Collect historical attack rates; the historical attack rate represents the frequency or proportion of attacks (such as forgery, liveness detection, etc.) detected by the system over a past period. It reflects the threat level of the current environment; based on the liveness confidence, forgery detection confidence, and cross-modal confidence combined with the historical attack rate, and through a dynamic weight adjustment algorithm, obtain the liveness confidence weight, forgery detection confidence weight, and cross-modal confidence weight; the formula for the dynamic weight adjustment algorithm is as follows...
[0120] ;
[0121] in, w i t+1 Indicates time t +1 moment i Each confidence score ( C live , C fake , C sync The corresponding weights, w i t Indicates time t Time of the first i The weights corresponding to each confidence score. c This represents the set learning rate. r Indicates historical attack rate. C i t This indicates that at time t, the first... i Each confidence score;
[0122] The liveness confidence, forgery detection confidence, and cross-modal confidence are vectorized to obtain a confidence vector; the liveness confidence weight, forgery detection confidence weight, and cross-modal confidence weight are vectorized to obtain a weight vector.
[0123] The fusion score is calculated based on the confidence vector and the weight vector; the calculation formula is as follows.
[0124] ;
[0125] in, Score displayed Integration, w i Represents the weight vector of the first element. i A vector of weights, C i Represents the confidence vector of the first... i A vector of confidence levels. t This indicates the set adjustment coefficient (used to control the degree of influence of the minimum value item on the final score). CRepresents the confidence vector;
[0126] S4. Construct an RNN neural network, and use historical remote interview behavior feature data combined with optimization algorithms to optimize the RNN neural network to obtain an RNN behavior analysis network.
[0127] Use a behavioral analysis network to obtain a real-time index of abnormal behavior in job seekers;
[0128] S4 includes the following steps:
[0129] S41. Collect behavioral data and corresponding abnormal behavior indices from historical remote interviews of job seekers to obtain historical remote interview behavioral data; the historical remote interview behavioral data includes behavioral data of abnormal job seekers and non-abnormal job seekers.
[0130] Extract features from historical remote interview behavior data to obtain historical remote interview behavior feature data;
[0131] S42. Construct an RNN neural network, set the initial weight values of the RNN neural network; set the prediction accuracy and prediction accuracy threshold of the RNN neural network;
[0132] The RNN neural network is trained using historical remote interview behavior data. During the training process, the optimal weight values of the RNN neural network are found by combining the training accuracy and the training accuracy threshold through an optimization algorithm, thus obtaining the optimal solution.
[0133] The optimal solution is used as the weight value of the RNN neural network to obtain the RNN behavior analysis network;
[0134] The optimization algorithm of this invention can be the tern algorithm, genetic algorithm, fish swarm algorithm, and ant algorithm in the prior art;
[0135] This embodiment uses the ant colony algorithm as an example for introduction:
[0136] The optimal weight values for an RNN neural network are found using the ant colony algorithm, which combines training accuracy and a training accuracy threshold. The optimal solution is obtained through the following steps:
[0137] S421. Construct an ant colony, setting the colony size to be... m The ant colony is then represented as ,in, n i Indicates the first in the ant colony i Number of ants; Set the maximum number of iterations;
[0138] S422. Based on the weight values of the RNN neural network, randomly set the initial positions of the ant colony to obtain the initial position set of the ant colony. ,in u i Indicates the first in the ant colony i The initial position of an ant; the position of the ant can reflect the distance of the ant from the food;
[0139] S423. Based on the prediction accuracy and prediction accuracy threshold of the RNN neural network, define the fitness function for the initial position of ants in an ant colony. The fitness function formula is as follows.
[0140] ;
[0141] in ,o Let z represent the fitness function, z1 represent the prediction accuracy, and z2 represent the prediction accuracy threshold. b Indicates the bias amount;
[0142] S424. Iterate over the initial location set of the ant population. The higher the fitness value, the closer the ant is to the food. In each iteration, calculate the fitness value of each location in the initial location set of the ant population according to the fitness function. Update the location concentration of each ant in the initial location set of the ant population according to the fitness value from high to low. In each iteration, obtain the best individual ant location and the global best ant location in the ant population.
[0143] S425, Repeat 424. When the maximum number of iterations is reached, stop iterating and take the global best ant position as the optimal solution.
[0144] The intelligent behavior analysis of this invention optimizes the accuracy of anomaly recognition; it uses the ant colony algorithm to optimize the RNN neural network; with prediction accuracy as the objective function, it improves training efficiency by intelligently searching for the optimal weight combination through ant colony; the optimized RNN behavior analysis network can capture hidden fraud features such as micro-expression temporal contradictions and response delays in real time, and output an abnormal behavior index to make up for the blind spots of biological detection.
[0145] S43. Collect real-time job seeker interview behavior data and extract features to obtain real-time job seeker behavior feature data;
[0146] Real-time job seeker behavioral characteristic data are input into an RNN behavioral analysis network to obtain a real-time abnormal behavior index.
[0147] S5. Based on the fusion score and the real-time abnormal behavior index, obtain the final risk score; execute the tiered response plan based on the final risk score.
[0148] S5 includes the following steps:
[0149] S51. Set a tiered response scheme; the tiered response scheme is divided into a Level 1 execution scheme, a Level 2 execution scheme, and a Level 3 execution scheme based on the final risk score; for example, the Level 1 execution scheme is when the final risk score is ≤0.1, a verifiable digital resume is generated, the resume hash is stored in IPFS, and CID+blockchain notarization is returned; the Level 2 execution scheme is when 0.1 < final risk score ≤0.3, a dynamic challenge response is initiated, and the Level 1 execution scheme is executed after passing the challenge; the Level 3 execution scheme is when the final risk score is greater than 0.3, multi-party video arbitration is triggered.
[0150] S52. The final risk score is calculated based on the fusion score and the real-time abnormal behavior index; the calculation formula is as follows.
[0151] ;
[0152] RiskScore represents the final risk score. q 1 represents the weighting coefficient of the fusion score. q 2 represents the weighting coefficient of the real-time abnormal behavior index, indicating that... B SCORE Real-time abnormal behavior index;
[0153] S53. Implement the tiered response plan based on the final risk score.
[0154] Example 2:
[0155] Please see Figure 3 A remote interview anti-fraud system based on multimodal deepfake detection is used to implement the aforementioned remote interview anti-fraud method based on multimodal deepfake detection. It includes an identity authentication and blockchain storage module, a multimodal dynamic detection module, a confidence fusion and dynamic scoring module, a behavior analysis network module, and a risk decision and response module.
[0156] The identity authentication and blockchain storage module is responsible for collecting job seekers' identity information and biometrics, generating job seeker information hashes and biometric hashes through a collision-resistant fuzzy hash algorithm; using zkSNARKs technology to implement zero-knowledge proofs for hash matching to ensure privacy and security; after a successful match, the hash triple (document information hash + face hash + voiceprint hash) is written into the Chia blockchain to achieve tamper-proof distributed storage.
[0157] The multimodal dynamic detection module extracts physiological signals from dual-band video streams, calculates the signal-to-noise ratio and main pulse frequency, and verifies whether physiological activities conform to human characteristics; it analyzes the abnormal features of video frame sequences and calculates the abnormality index by combining statistics from real datasets; it extracts the second-order mixed partial derivatives of the speech phase spectrum, calculates synthetic feature values to identify synthetic speech; and it uses the DTW algorithm to align lip shapes with the timestamps of sounds and outputs the maximum offset.
[0158] The confidence fusion and dynamic scoring module generates a liveness confidence score by combining a logical function with physiological features; obtains a forgery detection confidence score by weighted fusion of video anomaly index and speech synthesis feature value; generates a cross-modal confidence score based on the exponential decay function of lip sound offset; and updates the weights of each confidence score in real time by using a dynamic weight adjustment algorithm combined with historical attack rates, finally outputting a fusion score. The calculation process integrates the confidence vector, weight vector, and minimum value adjustment mechanism.
[0159] The behavior analysis network module utilizes historical interview behavior data, optimizes RNN weights using the ant colony algorithm, and outputs the optimal RNN behavior analysis network; it extracts job seeker interview behavior features and inputs them into the trained network to generate a real-time abnormal behavior index.
[0160] The risk decision-making and response module integrates and merges scoring and behavioral anomaly index, calculates the final risk score through weighted average, and executes a graded response based on the scoring threshold.
[0161] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0162] The preferred embodiments of the invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A remote interview anti-fraud method based on multimodal deepfake detection, characterized in that, Includes the following steps: S1. Collect job seekers' identity information and biometric characteristics; After hashing the job seeker's identity information and biometrics, identity matching is performed, and the results are uploaded to the blockchain for storage. S2. Using multimodal dynamic detection combined with data stored in the blockchain, live physiological signal detection, deepfake spatiotemporal analysis, voiceprint frequency domain diagnosis, and cross-modal synchronous verification are performed on job seekers to obtain physiological feature pairs, abnormality index, synthetic feature value, and maximum offset. S3. Based on the physiological feature pairs, abnormality index, synthetic feature value, and maximum offset, calculate the liveness confidence, forgery detection confidence, and cross-modal confidence. The weights of the liveness confidence, forgery detection confidence, and cross-modal confidence are dynamically adjusted to obtain a fusion score; S4. Construct an RNN neural network, and use historical remote interview behavior feature data combined with optimization algorithms to optimize the RNN neural network to obtain an RNN behavior analysis network. Use a behavioral analysis network to obtain a real-time index of abnormal behavior in job seekers; S5. The final risk score is obtained by combining the fusion score with the real-time abnormal behavior index; A tiered response plan will be implemented based on the final risk score.
2. The remote interview anti-fraud method based on multimodal deepfake detection according to claim 1, characterized in that, S1 includes the following steps: S11. Collect high-resolution images of job seekers' ID cards, extract names and ID numbers using OCR to obtain job seeker ID information; hash the job seeker ID information to obtain job seeker information hash; S12. Collect photoplethysmography signal data, 3D structured light face point cloud data, and voiceprint sample data of job seekers through the encrypted SDK to obtain the job seekers' biometric features. S13. Perform anti-collision fuzzy hashing on the job seeker's biometric features to obtain job seeker biometric hash pairs; the biometric hash pairs include face hash and voiceprint hash, and the anti-collision fuzzy hashing process includes face hash and voiceprint hash. S14. Use zkSNARKs to verify that the biometric hash of the job seeker is matched with the hash of the job seeker's identification information without revealing the original biometric features. S15. After a successful match, the job seeker's biometric hash and the job seeker's document information hash are used to construct a triplet to obtain the job seeker information triplet. The job seeker information triplet is written into the Chia blockchain, which includes a BLS signature + spacetime proof mechanism.
3. The remote interview anti-fraud method based on multimodal deepfake detection according to claim 1, characterized in that, S2 includes the following steps: S21. By analyzing the dual-band video stream of the job seeker's face captured by the camera, signals related to physiological activities are extracted to obtain live physiological signals; by combining the live physiological signals with a live physiological detection algorithm, physiological feature pairs are obtained; the physiological feature pairs include signal-to-noise ratio and main pulse frequency. S22. Collect consecutive frames of the job seeker's interview video to obtain a video frame sequence; the video frame sequence includes all consecutive frames of the job seeker's real-time interview video; Extract the abnormal features of each frame in the video frame sequence to obtain the video frame abnormal feature sequence; the video frame abnormal feature sequence includes the abnormal features of each frame in the video frame sequence. An anomaly index is obtained by combining the deepfake anomaly index algorithm with the anomaly feature sequence of video frames; S23. Extract the speech phase spectrum of the job seeker's voice signal during the interview, calculate the synthesized speech features, and obtain the synthesized feature values; S24. Use the DTW algorithm to verify the lip-sync of job seekers and obtain the maximum offset.
4. The remote interview anti-fraud method based on multimodal deepfake detection according to claim 1, characterized in that, S3 includes the following steps: S31. Based on the physiological feature pairs, abnormality index, synthetic feature value, and maximum offset, calculate the liveness confidence, forgery detection confidence, and cross-modal confidence. S32. Collect historical attack rates; combine the liveness confidence, forgery detection confidence, and cross-modal confidence with the historical attack rates, and obtain the liveness confidence weight, forgery detection confidence weight, and cross-modal confidence weight through a dynamic weight adjustment algorithm; The liveness confidence, forgery detection confidence, and cross-modal confidence are vectorized to obtain a confidence vector; the liveness confidence weight, forgery detection confidence weight, and cross-modal confidence weight are vectorized to obtain a weight vector. The fusion score is calculated by combining the confidence vector with the weight vector.
5. The remote interview anti-fraud method based on multimodal deepfake detection according to claim 4, characterized in that, S31 includes the following steps: S311. Set parameters for the steepness of the control logic function and threshold parameters for the logic function; calculate the liveness confidence level based on the parameters for the steepness of the control logic function and threshold parameters for the logic function, combined with physiological characteristics. S312. Set the weight parameters for video anomalies and voice anomalies; calculate the forgery detection confidence level based on the anomaly index, the synthesized feature value, and the weight parameters for video anomalies and voice anomalies. S313. Calculate the cross-modal confidence level based on the maximum offset.
6. The remote interview anti-fraud method based on multimodal deepfake detection according to claim 1, characterized in that, S4 includes the following steps: S41. Collect behavioral data and corresponding behavioral anomaly indices from historical remote interviews of job seekers to obtain historical remote interview behavioral data; extract features from historical remote interview behavioral data to obtain historical remote interview behavioral feature data. S42. Construct an RNN neural network, set the initial weight values of the RNN neural network; set the prediction accuracy and prediction accuracy threshold of the RNN neural network; The RNN neural network is trained using historical remote interview behavior data. During the training process, the optimal weight values of the RNN neural network are found by combining the training accuracy and the training accuracy threshold through the ant colony algorithm, thus obtaining the optimal solution. The optimal solution is used as the weight value of the RNN neural network to obtain the RNN behavior analysis network; S43. Collect real-time job seeker interview behavior data and extract features to obtain real-time job seeker behavior feature data; Real-time job seeker behavioral characteristic data are input into an RNN behavioral analysis network to obtain a real-time abnormal behavior index.
7. The remote interview anti-fraud method based on multimodal deepfake detection according to claim 6, characterized in that, The historical remote interview behavior data includes behavioral data of both abnormal and non-abnormal job seekers.
8. The remote interview anti-fraud method based on multimodal deepfake detection according to claim 1, characterized in that, S5 includes the following steps: S51. Establish a tiered response plan; the tiered response plan is divided into a Level 1 execution plan, a Level 2 execution plan, and a Level 3 execution plan based on the final risk score; S52. The final risk score is calculated based on the fusion score and the real-time abnormal behavior index. S53. Implement the tiered response plan based on the final risk score.
9. A remote interview anti-fraud system based on multimodal deepfake detection, characterized in that, The remote interview anti-fraud method based on multimodal deepfake detection as described in any one of claims 1-8 is provided, wherein the system includes an identity authentication and blockchain storage module, a multimodal dynamic detection module, a confidence fusion and dynamic scoring module, a behavior analysis network module, and a risk decision and response module.
10. A storage medium, characterized in that, It stores a program that, when executed by a processor, implements the remote interview anti-fraud method based on multimodal deep forgery detection as described in any one of claims 1-8.
Citation Information
Patent Citations
ETC fraud detection method
CN120493249A
Authentication ID interview method and apparatus
US20070078668A1