Video call forgery detection method based on high-frequency sound signal

By transmitting high-frequency acoustic signals carrying random sequence codes during video calls, simultaneously acquiring audio and video signals, and extracting acoustic and visual motion features for consistency analysis, the limitations of existing video forgery detection methods are overcome, achieving imperceptible forgery detection and enhancing the defense against video replay and synthetic forgery.

CN121842373APending Publication Date: 2026-04-10ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-04
Publication Date
2026-04-10

Smart Images

  • Figure CN121842373A_ABST
    Figure CN121842373A_ABST
Patent Text Reader

Abstract

The invention discloses a video call forgery detection method based on a high-frequency sound signal, and the method comprises the steps: enabling a terminal device to actively transmit a high-frequency sound signal which is imperceptible to human ears in a video call process, and embedding a random sequence code in the high-frequency sound signal; the method comprises the following steps: synchronously acquiring an audio signal and a video signal of a video call, verifying a random sequence carried in the audio, extracting a human body motion feature of a target user reflected by high-frequency sound response, and extracting a corresponding visual motion feature from a video picture; and judging whether the video call is a real-time real call or not by analyzing the time sequence consistency between the acoustic motion feature and the visual motion feature. According to the method, a high-frequency sound active detection mechanism of random sequence coding and cross-modal motion consistency check are utilized, and playback and synthesis class forgery in a video call are defended from the physical level. The system can be deployed in terminal equipment such as a computer, a smart phone and the like with a loudspeaker, a microphone and a camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of information security and multimedia communication technology, specifically, a method for detecting video call spoofing based on high-frequency audio signals. Background Technology

[0002] With the popularization of video call technology, forgery attacks based on recorded or synthesized videos are increasing. Attackers can impersonate legitimate users to participate in video calls through video replay, deep synthesis, and other methods, causing identity fraud and security risks.

[0003] Current video spoofing detection methods are mostly passive, relying on video content features and are easily affected by image compression and algorithm evolution, and struggle to effectively distinguish between real-time and replay videos. Existing active detection methods often require displaying specific patterns in a large window or demanding specific user actions, impacting the video call experience and making them difficult to deploy in real-world systems. Therefore, there is an urgent need for a spoofing detection method that requires no user cooperation, operates seamlessly during real-time video calls, and possesses defenses against video replaying and fabrication. Summary of the Invention

[0004] This invention addresses the limitations of current video forgery detection methods by proposing a novel active detection method: a video call forgery detection method based on high-frequency acoustic signals. This method achieves reliable detection of the authenticity of video calls through active detection of high-frequency acoustic signals using random sequence encoding and consistency analysis of cross-modal motion features.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: during a video call, a high-frequency acoustic signal carrying a random sequence code is transmitted; audio and video signals are simultaneously acquired, and the random sequence in the high-frequency acoustic response is verified to distinguish between the real-time acoustic response and the replay signal; if the random code verification is successful, acoustic motion features and visual motion features are extracted, and consistency analysis is performed on the two to determine whether the video call is fake.

[0006] This invention is achieved through the following technical solution:

[0007] This invention discloses a video call spoofing detection method based on high-frequency acoustic signals. During a video call, a high-frequency acoustic signal carrying a random sequence code is emitted; audio and video signals are acquired simultaneously, and the random sequence in the high-frequency acoustic response is verified to distinguish between the real-time acoustic response and the replay signal; if the random code verification passes, acoustic motion features and visual motion features are extracted, and consistency analysis is performed on the two to determine whether the video call is spoofed.

[0008] As a further improvement, the present invention specifically includes:

[0009] S1. During a video call, the platform transmits high-frequency sound signals that are imperceptible to the human ear through the target user's terminal device. These high-frequency sound signals carry random sequence codes.

[0010] S2. Synchronously acquire audio and video signals of the video call, wherein the audio signal includes acoustic response signals formed by high-frequency sound signals reflected by the human body and the environment;

[0011] S3. Extract acoustic response features from the audio signal and verify them based on a random sequence to determine whether the acoustic response signal is generated in real time;

[0012] S4. If the random sequence verification passes, extract acoustic motion features reflecting changes in the target user's human motion from the acoustic response features;

[0013] S5. Extract visual motion features from the video signal that reflect changes in the target user's human motion;

[0014] S6. Perform a consistency check between acoustic motion characteristics and visual motion characteristics;

[0015] S7. Based on the random sequence verification results and the motion feature consistency test results, determine whether the target user used a fake video call.

[0016] As a further improvement, in S1 of the present invention, the high-frequency acoustic signal includes multiple continuous signal units, each signal unit consisting of a fixed number and length of linear continuous frequency-modulated chirp periods, and the random sequence is embedded by changing the frequency modulation parameters of different signal units.

[0017] As a further improvement, in S3 of this invention, the verification of the random sequence is accomplished by comparing the consistency of the received acoustic response with the known random sequence encoding in terms of time-frequency structure. Specifically, template-based correlation analysis is employed, including the following steps:

[0018] 1) Construct the corresponding high-frequency acoustic template waveform based on the known random sequence, perform cross-correlation operation with the acquired audio signal, and extract each signal unit with the highest peak position of the correlation curve as the starting point;

[0019] 2) Construct a single chirp template for each signal unit based on the known random sequence, and perform cross-correlation calculations with the extracted signal units. Use three metrics from the correlation curve—peak-to-noise ratio (PNR), peak-to-mean ratio (PMR), and peak spacing error (PIE)—to confirm whether each signal unit meets expectations. (PNR...) The significance of the correlation peak relative to background random fluctuations is measured, where Peak amplitude in the correlation curve The standard deviation of peak external noise; peak-to-mean ratio (PMR) This reflects the prominence of the correlation peak relative to the overall correlation baseline, where... The average value of the correlation curves; peak spacing error PIE = Examine the relative deviation between the time intervals between adjacent correlation peaks and the known chirp period, where The time difference between adjacent peaks The chirp period is known.

[0020] 3) When all three indicators within a signal unit meet the specific threshold requirements, the signal unit is considered to have successfully matched the template; when all signal units in the acquired audio signal are successfully matched, the random sequence verification is passed.

[0021] As a further improvement, in S4 of this invention, the acoustic motion feature is a time-varying feature quantity characterizing the change of acoustic response with the movement of the target user over a continuous time period. Specifically, the method for extracting the acoustic motion feature is as follows:

[0022] 1) High-pass filtering is applied to the audio signal to suppress interference in the audible frequency band;

[0023] 2) After mixing each chirp in the audio signal with its template, the channel impulse response (CIR) is calculated using fast Fourier transform;

[0024] 3) Differentiate the CIRs of adjacent chirps within each signal unit to obtain the DiffCIR, which reflects the motion intensity of each distance interval;

[0025] 4) For the DiffCIR within each time frame, select the distance range of interest and aggregate them to obtain the feature sequence of audio modalities reflecting the motion intensity of the target user. .

[0026] As a further improvement, in S5 of this invention, visual motion features are feature quantities characterizing the motion changes of the target user in consecutive video frames. Specifically, the method for extracting visual motion features is as follows:

[0027] 1) Calculate the optical flow field between consecutive video frames ,in Represents pixels From frame arrive Horizontal and vertical displacements; amplitude spectrum extracted from these. ;

[0028] 2) Use a semantic segmentation model on the video frames to generate a binary mask that extracts the region where the target user is located;

[0029] 3) Multiply the amplitude spectrum with a binarized mask to mask background interference, and then aggregate all pixels of the amplitude spectrum of each frame to obtain a feature sequence of video modalities reflecting the motion intensity of the target user. .

[0030] As a further improvement, in S6 of this invention, the consistency test is used to determine the synchronicity or correlation between acoustic motion features and visual motion features in the time dimension. Specifically, it employs a sliding window-based cross-correlation calculation, including the following steps:

[0031] 1) Motion feature sequences extracted from the two modes and Perform normalization and filtering;

[0032] 2) Using a length of Step size is The sliding window is used to segment the sequence, and the cross-correlation is calculated within each window. ,in Indicate the lag; extract the lag corresponding to the highest peak value of the correlation curve. Peak-to-noise ratio (PNR), if If both PNR and PNR exceed the specified threshold, the window is successfully matched;

[0033] 3) When the number of successfully matched windows exceeds the predefined ratio threshold, the acquired video passes the motion feature consistency test.

[0034] As a further improvement, in S7 of the present invention, a video call is determined to be real if and only if both the random sequence verification and the motion feature consistency test pass; otherwise, the target user is determined to have made a fake video call.

[0035] The beneficial effects of this invention are as follows:

[0036] This invention involves the terminal device actively transmitting a high-frequency acoustic signal carrying a random sequence code during a video call, and then verifying the consistency of the received acoustic response using the random sequence, thus physically distinguishing between real-time calls and replay spoofing. Because the random sequence is generated in real-time and unpredictably during the call, attackers find it difficult to construct a matching acoustic response through pre-recording or simple replay. Therefore, this invention effectively ensures the detection capability against video replay attacks.

[0037] This invention, based on verifying the real-time performance of high-frequency acoustic response, further extracts acoustic motion features reflecting changes in the target user's human body movement from the acoustic signal and extracts corresponding visual motion features from the video frame. Consistency analysis is used to determine the synchronicity of these two features in the temporal dimension, thereby identifying synthetic video forgeries. This consistency analysis is calculated based on physical signal features and temporal correlation, without relying on complex model training or large amounts of sample data, resulting in low computational overhead. Even if attackers can forge high-fidelity video frames, it is difficult to simultaneously generate an acoustic response that is strictly synchronized with real human body movement. Therefore, this invention effectively enhances the defense against advanced forgery techniques such as deep synthesis and virtual avatars.

[0038] This invention requires no specific user action or additional verification process. It can run in the background during normal video calls, and the high-frequency sound signal used is imperceptible to the human ear, avoiding the interference with user experience found in existing active detection methods. It does not affect the video call experience and is suitable for deployment in practical applications such as remote conferencing. This invention can be directly applied to general-purpose terminal devices such as computers and smartphones equipped with speakers, microphones, and cameras, requiring no additional hardware support, and boasts low deployment costs and good engineering scalability. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0040] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0041] This invention proposes a method for detecting spoofed video calls based on high-frequency acoustic signals. Figure 1 This is a schematic diagram of the method flow of the present invention.

[0042] The specific implementation method of the present invention is as follows:

[0043] Step 1: During the video call, the platform generates a segment of length [length missing] in real time each time. A random sequence is generated, and the random sequence is embedded into a high-frequency acoustic signal by changing the frequency modulation parameters; the generated high-frequency acoustic signal contains... Each signal unit contains [number] signal units. A series of consecutive chirp cycles are emitted by the speaker of the terminal device. In this embodiment, the random sequence is a ternary sequence. ), , Each chirp cycle lasts 0.05s; the starting frequency of each chirp is 18kHz, and the ending frequency is determined by the corresponding random sequence character. The termination frequencies of all chirps within each signal unit are .

[0044] Step 2: Simultaneously acquire audio and video signals of the video call using the microphone and camera of the terminal device. The audio signal includes the acoustic response signal formed by the high-frequency sound signal emitted in Step 1 after reflection from the human body and the environment. In this embodiment, the audio sampling rate is 48kHz and the video frame rate is 20FPS.

[0045] Step 3: Random sequence verification is achieved by comparing the received acoustic response with the known random sequence encoding in terms of time-frequency structure. Specifically, firstly, a corresponding high-frequency acoustic template waveform is constructed based on the known random sequence, and cross-correlation is performed with the acquired audio signal. Using the highest peak of the correlation curve as the starting point, each signal unit is extracted. Then, a single chirp template is constructed for each signal unit according to the known random sequence, and cross-correlation is performed with the extracted signal units. The peak-to-noise ratio (PNR), peak-to-mean ratio (PMR), and peak spacing error (PIE) of the correlation curve are calculated. Peak-to-noise ratio (PNR) The significance of the correlation peak relative to background random fluctuations is measured, where The peak amplitude in the correlation curve. The standard deviation of peak external noise; peak-to-mean ratio (PMR) This reflects the prominence of the correlation peak relative to the overall correlation baseline, where... The average value of the correlation curves; peak spacing error PIE = Examine the relative deviation between the time intervals between adjacent correlation peaks and the known chirp period, where The time difference between adjacent peaks The chirp period is known. When all three indicators within a signal unit meet specific threshold requirements, the signal unit is considered to have successfully matched the template. When all signal units in the acquired audio signal successfully match their templates, the random sequence verification passes, and the following steps continue. Otherwise, steps four through six are skipped, and the judgment in step seven is performed directly.

[0046] Step 4: If the random sequence verification in Step 3 passes, extract acoustic motion features reflecting the target user's human motion changes from the audio signal. Specific steps include: high-pass filtering the audio signal to suppress interference in the audible frequency band; mixing each chirp in the audio signal with its template and calculating the channel impulse response (CIR) using a fast Fourier transform; differentiating the CIRs of adjacent chirs within each signal unit to obtain the DiffCIR, reflecting the motion intensity of each distance interval; and then aggregating the DiffCIRs within each time frame for a selected distance range to obtain a feature sequence of audio modes reflecting the target user's motion intensity. In this embodiment, the distance range selected during DiffCIR polymerization is 0~50cm.

[0047] Step 5: Extract visual motion features reflecting changes in the target user's human motion from the video signal. Specific steps include: calculating the optical flow field between consecutive video frames. ,in Represents pixels From frame arrive The horizontal and vertical displacements were measured, and the amplitude spectrum was extracted from them. A semantic segmentation model is used on the video frames to generate a binary mask that extracts the region where the target user is located. Then, the amplitude spectrum is multiplied by the binary mask to mask background interference. Finally, all pixels of the amplitude spectrum of each frame are aggregated to obtain a feature sequence of video modalities that reflects the motion intensity of the target user. In this embodiment, the readily available RAFT model is used to calculate optical flow, and the MASK R-CNN model is used for semantic segmentation to extract human body regions.

[0048] Step Six: Examine the synchronicity or correlation between acoustic motion features and visual motion features in the temporal dimension. Specific steps include: extracting motion feature sequences from the two modalities. and Perform normalization and filtering; use a length of Step size is The sliding window is used to segment the sequence, and the cross-correlation is calculated within each window. ,in Indicate the lag; calculate the lag corresponding to the highest peak of the correlation curve. Peak-to-noise ratio (PNR), if If both the sliding window length and the point of repetition (PNR) exceed a specified threshold, the window is successfully matched. When the number of successfully matched windows exceeds a predefined proportion threshold, the acquired video passes the motion feature consistency check. In this embodiment, the sliding window length... Step length .

[0049] Step 7: Combining the random sequence verification results from Step 3 and the motion feature consistency test results from Step 6, determine whether the video call is forged. The video call is determined to be genuine if and only if both the random sequence verification and the motion feature consistency test pass; otherwise, the target user is determined to have forged the video call.

[0050] The above description is not intended to limit the present invention. It should be noted that, for those skilled in the art, various changes, modifications, additions or substitutions can be made without departing from the essential scope of the present invention, and these improvements and refinements should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting spoofing in video calls based on high-frequency acoustic signals, characterized in that, During a video call, a high-frequency acoustic signal carrying a random sequence code is transmitted; audio and video signals are acquired simultaneously, and the random sequence in the high-frequency acoustic response is verified to distinguish between the real-time acoustic response and the replay signal. If the random coding verification passes, acoustic motion features and visual motion features are extracted, and consistency analysis is performed on the two to determine whether the video call is fake.

2. The video call forgery detection method based on high-frequency acoustic signals according to claim 1, characterized in that, Specifically, it includes: S1. During a video call, the platform transmits a high-frequency sound signal imperceptible to the human ear through the target user's terminal device, the high-frequency sound signal carrying a random sequence code; S2. Synchronously acquire audio and video signals of the video call, wherein the audio signal includes the acoustic response signal formed by the high-frequency sound signal reflected by the human body and the environment; S3. Extract acoustic response features from the audio signal and verify them based on the random sequence to determine whether the acoustic response signal is generated in real time; S4. If the random sequence verification passes, extract acoustic motion features reflecting changes in the target user's human motion from the acoustic response features; S5. Extract visual motion features from the video signal that reflect changes in the target user's human body movement; S6. Perform a consistency check on the acoustic motion features and visual motion features; S7. Based on the random sequence verification results and motion feature consistency test results, determine whether the target user used a fake video call.

3. The video call forgery detection method based on high-frequency acoustic signals according to claim 1, characterized in that, In S1, the high-frequency acoustic signal includes multiple continuous signal units, each of which consists of a fixed number and length of linear continuous frequency-modulated chirp cycles. The random sequence is embedded by changing the frequency modulation parameters of different signal units.

4. The video call forgery detection method based on high-frequency acoustic signals according to claim 1, 2, or 3, characterized in that, In S3, the verification of the random sequence is accomplished by comparing the consistency of the received acoustic response with the known random sequence encoding in terms of time-frequency structure. Specifically, template-based correlation analysis is used, including the following steps: 1) Construct the corresponding high-frequency acoustic template waveform based on the known random sequence, perform cross-correlation operation with the acquired audio signal, and extract each signal unit with the highest peak position of the correlation curve as the starting point; 2) Construct a single chirp template for each signal unit based on the known random sequence, and perform cross-correlation calculations with the extracted signal units. Use three metrics from the correlation curve—peak-to-noise ratio (PNR), peak-to-mean ratio (PMR), and peak spacing error (PIE)—to confirm whether each signal unit meets expectations. (PNR...) The significance of the correlation peak relative to background random fluctuations is measured, where The peak amplitude in the correlation curve. The standard deviation of peak external noise; peak-to-mean ratio (PMR) This reflects the prominence of the correlation peak relative to the overall correlation baseline, where... The average value of the correlation curves; peak spacing error PIE = Examine the relative deviation between the time intervals between adjacent correlation peaks and the known chirp period, where The time difference between adjacent peaks The chirp period is known. 3) When all three indicators within a signal unit meet the specific threshold requirements, the signal unit is considered to have successfully matched the template; when all signal units in the acquired audio signal are successfully matched, the random sequence verification is passed.

5. The video call forgery detection method based on high-frequency acoustic signals according to claim 4, characterized in that, In S4, the acoustic motion feature is a time-varying feature quantity that characterizes the change of acoustic response with the target user's motion over a continuous time period. Specifically, the method for extracting the acoustic motion feature is as follows: 1) High-pass filtering is applied to the audio signal to suppress interference in the audible frequency band; 2) After mixing each chirp in the audio signal with its template, the channel impulse response (CIR) is calculated using fast Fourier transform; 3) Differentiate the CIRs of adjacent chirps within each signal unit to obtain the DiffCIR, which reflects the motion intensity of each distance interval; 4) For the DiffCIR within each time frame, select the distance range of interest and aggregate them to obtain the feature sequence of audio modalities reflecting the motion intensity of the target user. .

6. The video call forgery detection method based on high-frequency acoustic signals according to claim 5, characterized in that, In S5, visual motion features are feature quantities that characterize the motion changes of the target user in consecutive video frames. Specifically, the method for extracting visual motion features is as follows: 1) Calculate the optical flow field between consecutive video frames ,in Represents pixels From frame arrive Horizontal and vertical displacements; amplitude spectrum extracted from these. ; 2) Use a semantic segmentation model on the video frames to generate a binary mask that extracts the region where the target user is located; 3) Multiply the amplitude spectrum with a binarized mask to mask background interference, and then aggregate all pixels of the amplitude spectrum of each frame to obtain a feature sequence of video modalities reflecting the motion intensity of the target user. .

7. The video call forgery detection method based on high-frequency acoustic signals according to claim 1, 2, 3, 5, or 6, characterized in that, In S6, the consistency test is used to determine the synchronicity or correlation between acoustic motion features and visual motion features in the time dimension. Specifically, it employs a sliding window-based cross-correlation calculation, including the following steps: 1) Motion feature sequences extracted from the two modes and Perform normalization and filtering; 2) Using a length of Step size is The sliding window is used to segment the sequence, and the cross-correlation is calculated within each window. ,in Indicate the lag; extract the lag corresponding to the highest peak value of the correlation curve. Peak-to-noise ratio (PNR), if If both PNR and PNR exceed the specified threshold, the window is successfully matched; 3) When the number of successfully matched windows exceeds the predefined ratio threshold, the acquired video passes the motion feature consistency test.

8. The video call forgery detection method based on high-frequency acoustic signals according to claim 7, characterized in that, In S7, a video call is determined to be genuine if and only if both the random sequence verification and the motion feature consistency test pass; otherwise, the target user is determined to have made a fake video call.