Audio and video identity recognition system based on multimode clue driving

Through the audio and video identity recognition system driven by multi-mode clues, the stability and anti-interference problems of traditional audio and video multi-modal fusion methods in complex scenarios are solved, and efficient and accurate identity verification is achieved. It is suitable for a variety of hardware environments and meets the application needs of highly sensitive fields such as finance, security and government.

CN120260146APending Publication Date: 2025-07-04NANJING LONGYUAN INFORMATION TECH CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510357577.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The traditional audio and video multimodal fusion method has a lot of room for improvement in computing efficiency and recognition accuracy, especially in complex scenarios, the stability and anti-interference ability of the system cannot be guaranteed.

Method used

The audio and video identity recognition system based on multi-mode clue-driven is adopted, including audio feature extraction module, video feature extraction module, multi-mode clue fusion module, living body detection module and identity identification and verification module. Feature weighted fusion is performed through the attention mechanism, combining deep learning and time stamp synchronization, and intelligent fusion and living body detection of audio and video modes are realized.

Benefits of technology

It improves the accuracy and robustness of identity verification, can maintain the stability of the system in complex environments, effectively prevent forgery attacks, and is suitable for a variety of hardware environments to meet the application needs of highly sensitive fields such as finance, security and government.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260146A_ABST
    Figure CN120260146A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio and video identity recognition, in particular to an audio and video identity recognition system based on multimode clue driving. The audio and video identity recognition system comprises an audio feature extraction module, a video feature extraction module, a multimode clue fusion module, a living body detection module and an identity recognition and verification module. The audio feature extraction module is used for capturing a voice signal of a user and extracting key voiceprint features, the video feature extraction module is used for acquiring facial features or limb features of the user and extracting related features, and the multi-mode clue fusion module is used for performing intelligent weighted fusion on the features of audio and video modes. The method comprises the following steps: firstly, using a living body detection module to comprehensively analyze the dynamic characteristics of audio and video, judging whether a user is a real individual, and finally, completing the final verification of the identity of the user through an identity recognition and verification module. The problem of feature extraction and modal fusion in the prior art is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio - video identity recognition, and particularly to an audio - video identity recognition system based on multi - mode clue driving. Background Art

[0002] With the rapid development of information technology, the demand for identity authentication and security verification is increasing day by day, especially in the fields of finance, government affairs, medical care, etc. How to ensure the security and accuracy of identity verification has become a key technical challenge. Most traditional identity authentication technologies rely on single - modality biometric means, such as voiceprint recognition, face recognition, fingerprint recognition, etc. Although these technologies provide convenient verification methods to a certain extent, there are also problems such as forgery attacks, environmental interference, and low recognition accuracy. Therefore, how to effectively improve the security and accuracy of the identity authentication system has become a hot topic in current technical research.

[0003] Existing identity authentication technologies are mainly divided into two categories: single - modality and multi - modality. Among single - modality technologies, voiceprint recognition and face recognition are the two most common methods, which verify identities through the voice characteristics and face characteristics of users respectively. Multi - modality identity recognition technologies improve the accuracy and robustness of identity verification by integrating biometric information from different sources (such as audio, video, images, etc.) and combining the advantages of each modality. However, multi - modality identity recognition technologies still face the following challenges: Difficulty in inter - modality fusion: Audio and video information have different data structures and time characteristics. How to effectively fuse this heterogeneous data is a key difficulty in technical implementation; Real - time problem: When a multi - modality recognition system processes and fuses data from multiple modalities, it may be affected by computational complexity, resulting in a delay in response time; Poor environmental adaptability: Multi - modality recognition systems have poor adaptability in complex environments. Especially in environments with low light, high noise, etc., the recognition effect of a single modality will decrease significantly. How to ensure the stability of the multi - modality system in these environments is another technical problem.

[0004] In summary, there is still much room for improvement in the computational efficiency and recognition accuracy of traditional audio - video multi - modality fusion methods. Especially how to ensure the stability and anti - interference ability of the system in complex scenarios is still an urgent problem to be solved. Summary of the Invention

[0005] The purpose of the present invention is to provide an audio - video identity recognition system based on multi - mode clue driving, so as to solve the problems that there is still much room for improvement in the computational efficiency and recognition accuracy of traditional audio - video multi - modality fusion methods, and the system cannot ensure stability and anti - interference ability in complex scenarios.

[0006] To achieve the above object, the present invention provides an audio-visual identity recognition system based on multi-modal cue driving. The audio-visual identity recognition system based on multi-modal cue driving includes an audio feature extraction module, a video feature extraction module, a multi-modal cue fusion module, a live detection module, and an identity recognition and verification module. The audio feature extraction module and the video feature extraction module are both connected to the input end of the multi-modal cue fusion module. The input end of the live detection module is connected to the output end of the multi-modal cue fusion module. The output end of the live detection module is connected to the identity recognition and verification module;

[0007] The audio feature extraction module is used to capture the user's voice signal and extract key voiceprint features;

[0008] The video feature extraction module is used to obtain the user's facial features or limb features through a camera device and extract relevant features;

[0009] The multi-modal cue fusion module is used to intelligently weight and fuse the features of the audio and video modalities through a fusion method based on the attention mechanism;

[0010] The live detection module is used to comprehensively analyze the dynamic features of the audio and video to determine whether the user is a real individual;

[0011] The identity recognition and verification module completes the final verification of the user's identity based on the fusion features and the live detection results.

[0012] Among them, the specific content of the audio feature extraction module includes speech preprocessing, speech feature extraction, and deep feature learning;

[0013] Speech preprocessing: Preprocess the speech signal through endpoint detection and pre-emphasis filtering;

[0014] Speech feature extraction: Extract features from the speech signal through Mel-frequency cepstral coefficients and perceptual linear prediction;

[0015] Deep feature learning: Use the pre-trained ResNet221 model. After inputting the features extracted by the feature extraction unit, extract the high-dimensional feature vector of the audio through a deep convolutional neural network to improve the learning ability of complex patterns in the speech signal.

[0016] Among them, the specific content of the video feature extraction module includes face detection and key point extraction, facial feature extraction, and dynamic feature capture;

[0017] Face detection and key point extraction: Extract the face information in the video information through face detection and key point extraction;

[0018] Facial feature extraction: Based on the ArcFace pre-trained model, 128-dimensional or 512-dimensional highly recognizable feature vectors are extracted from the detected faces as the core data for identity authentication or matching;

[0019] Dynamic feature capture: The dynamic changes of the face are extracted through the optical flow method, and the temporal convolutional network is used to capture the movement trajectory of the face in the video stream to enhance the expressiveness of dynamic features.

[0020] The specific contents of the multi-mode clue fusion module include feature alignment, weighted fusion and fusion feature extraction;

[0021] Feature alignment: The timestamp synchronization mechanism is used to ensure the consistency of audio and video features in the same time window. The frame-level alignment strategy is used to fill the time misalignment caused by different sampling rates. Interpolation compensation is used to interpolate missing or misaligned feature data to ensure that the audio and video modes are consistent in timing.

[0022] Weighted fusion: The importance scores of audio features and video features to the task objectives are calculated through a multi-head attention mechanism, the weights are adjusted dynamically, and the audio and video features are weighted and combined through feature weighting to generate a joint feature vector.

[0023] Fusion feature extraction: Extract joint features through the residual network to enhance the nonlinear expression ability of features while avoiding the gradient vanishing problem, and map the fusion features to a low-dimensional space through a fully connected layer for final classification or verification.

[0024] The specific contents of the liveness detection module include audio liveness detection, video liveness detection and multi-modal liveness fusion;

[0025] Audio liveness detection: Liveness detection of audio information is completed through the combination of time domain analysis and frequency domain analysis;

[0026] Video liveness detection: Liveness detection of video information is completed through the combination of facial dynamic detection and 3D depth information analysis;

[0027] Multimodal living body fusion: Jointly model audio and video features through fusion model design, mine the dynamic correlation between modalities, and make judgment outputs.

[0028] The specific contents of the identity recognition and verification module include identity matching, decision threshold adjustment and result output;

[0029] Identity matching: Calculate the similarity between the fused features and the pre-stored feature library;

[0030] Decision threshold adjustment: Dynamically adjust the recognition threshold according to the application scenario to balance the false recognition rate and missed recognition rate;

[0031] Result output: Output the authentication result and record the verification log for auditing at the same time.

[0032] Among them, when the identity matching is completed by "calculating the similarity between the fused features and the pre-stored feature library", the cosine similarity or Euclidean distance is used for similarity calculation.

[0033] An audio-visual identity recognition system based on multi-modal clue driving of the present invention includes an audio feature extraction module, a video feature extraction module, a multi-modal clue fusion module, a liveness detection module, and an identity recognition and verification module. The audio feature extraction module is used to capture the user's voice signal and extract key voiceprint features, and the video feature extraction module is used to obtain the user's facial features or limb features and extract relevant features. The features of the audio and video modalities are intelligently weighted and fused by the multi-modal clue fusion module, and then the liveness detection module comprehensively analyzes the dynamic features of the audio and video to determine whether the user is a real individual. Finally, the identity recognition and verification module completes the final verification of the user's identity. By adopting this technical solution, through the deep fusion of audio features and video features, the problems in feature extraction and modality fusion in the prior art are effectively solved, providing an efficient, accurate, and robust technical solution for multi-modal identity verification and forgery detection. Brief Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 is a schematic block diagram of the audio-visual identity recognition system based on multi-modal clue driving provided by the present invention.

[0036] Figure 2 is a schematic diagram of the operation principle of the audio feature extraction module provided by the present invention. Detailed Embodiments

[0037] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation of the present invention.

[0038] Please refer to Figure 1 and Figure 2, the present invention provides an audio - video identity recognition system based on multi - modal cue driving. The audio - video identity recognition system based on multi - modal cue driving includes an audio feature extraction module, a video feature extraction module, a multi - modal cue fusion module, a live detection module, and an identity recognition and verification module. The audio feature extraction module and the video feature extraction module are both connected to the input end of the multi - modal cue fusion module. The input end of the live detection module is connected to the output end of the multi - modal cue fusion module, and the output end of the live detection module is connected to the identity recognition and verification module;

[0039] The audio feature extraction module is used to capture the user's voice signal and extract key voiceprint features;

[0040] The video feature extraction module is used to obtain the user's facial features or limb features through a camera device and extract relevant features;

[0041] The multi - modal cue fusion module is used to intelligently weight - fuse the features of the audio and video modalities through a fusion method based on the attention mechanism;

[0042] The live detection module is used to comprehensively analyze the dynamic features of the audio and video to determine whether the user is a real individual;

[0043] The identity recognition and verification module completes the final verification of the user's identity based on the fused features and the live detection result.

[0044] In this embodiment, the audio feature extraction module is used to capture the user's voice signal and extract key voiceprint features, the video feature extraction module is used to obtain the user's facial features or limb features and extract relevant features, the multi - modal cue fusion module is used to intelligently weight - fuse the features of the audio and video modalities, then the live detection module is used to comprehensively analyze the dynamic features of the audio and video to determine whether the user is a real individual, and finally the identity recognition and verification module completes the final verification of the user's identity. By adopting this technical solution, through the deep fusion of audio features and video features, the problems in feature extraction and modality fusion in the prior art are effectively solved, providing an efficient, accurate, and robust technical solution for multi - modal identity verification and forgery detection.

[0045] Further, the specific content of the audio feature extraction module includes speech pre - processing, speech feature extraction, and deep feature learning;

[0046] Speech pre - processing: The speech signal is pre - processed through endpoint detection and pre - emphasis filtering;

[0047] Speech feature extraction: The speech signal is feature - extracted through Mel - Frequency Cepstral Coefficients and Perceptual Linear Prediction;

[0048] Deep feature learning: After using the pre-trained ResNet221 model and inputting the features extracted by the input feature extraction unit, a high-dimensional feature vector of the audio is extracted through a deep convolutional neural network to enhance the learning ability of complex patterns in the speech signal.

[0049] In this embodiment, endpoint detection specifically refers to removing background noise and silent segments through an endpoint detection algorithm based on energy and zero-crossing rate, and only retaining the effective speech signal;

[0050] Pre-emphasis filtering specifically refers to applying a first-order high-pass filter H(z)=1 - \alpha z^{-1} (usually \alpha = 0.97) to improve the quality of high-frequency signals and enhance important features in the speech;

[0051] MFCC (Mel Frequency Cepstral Coefficients) specifically refers to converting the speech signal into frequency-domain features, analyzing the spectral energy distribution of the speech signal using a Mel filter bank, and generating low-dimensional acoustic features;

[0052] PLP (Perceptual Linear Prediction) specifically refers to further extracting linear prediction features related to perception by simulating the human auditory system to enhance the representation ability of timbre and pitch;

[0053] Deep feature learning specifically refers to using the pre-trained ResNet221 model, inputting MFCC and PLP features, and extracting a high-dimensional feature vector of the audio through a deep convolutional neural network to enhance the learning ability of complex patterns in the speech signal.

[0054] Furthermore, the audio feature extraction module further includes audio cue-driven feature extraction, and the audio cue-driven feature extraction needs to meet the following requirements:

[0055] Time-frequency domain spatial feature consistency: Analyze the spectral dynamics through the Short-Time Fourier Transform (STFT) to ensure the continuity and consistency of the speech signal in the frequency domain;

[0056] Time-domain context fluency: Adopt a time series analysis method to verify the front-back continuity and context fluency of the speech signal;

[0057] Blank silent segment background noise continuity: Detect the background noise characteristics of the silent segment and judge whether its continuity conforms to the natural scene;

[0058] High-frequency feature truncation continuity: Analyze the truncation characteristics of the high-frequency band signal and identify forged or tampered signals through consistency detection;

[0059] Speaker feature front-back clustering: Adopt a clustering algorithm based on K-means to verify the coherence of speaker features between audio segments;

[0060] Consistency of front and rear noise: Check whether the background noise of different speech segments remains consistent to determine the possibility of speech synthesis or splicing;

[0061] Consistency of timbre features: Classify and compare timbre features through a multi-layer perceptron (MLP) to ensure consistency;

[0062] Consistency of accent features: Learn the accent features of speech sequences based on the LSTM or Transformer model and evaluate their consistency.

[0063] Furthermore, the specific content of the video feature extraction module includes face detection and key point extraction, facial feature extraction, and dynamic feature capture;

[0064] Face detection and key point extraction: Extract face information in video information through face detection and key point extraction;

[0065] Facial feature extraction: Based on the ArcFace pre-trained model, extract 128-dimensional or 512-dimensional highly discriminative feature vectors from the detected face as the core data for identity verification or matching;

[0066] Dynamic feature capture: Extract the dynamic changes of the face through the optical flow method and use a temporal convolutional network to capture the movement trajectory of the face in the video stream to enhance the expressiveness of dynamic features.

[0067] In this embodiment, face detection specifically refers to using MTCNN (Multi-Task Cascaded Convolutional Networks) to accurately locate the face area to ensure clear facial images are captured;

[0068] Key point extraction specifically refers to annotating facial key points (such as eyes, nose tip, corners of the mouth, etc.) in the face area to lay a foundation for subsequent feature extraction and dynamic analysis;

[0069] Facial feature extraction specifically refers to based on the ArcFace pre-trained model, extracting 128-dimensional or 512-dimensional highly discriminative feature vectors from the detected face as the core data for identity verification or matching.

[0070] The optical flow method specifically refers to calculating the pixel motion vectors between video frames to extract the dynamic changes of the face, such as features like slight lip movement and blinking, for live detection;

[0071] The temporal convolutional network (TCN) specifically refers to performing temporal context modeling on the video sequence to capture the movement trajectory of the face in the video stream to enhance the expressiveness of dynamic features.

[0072] Furthermore, the audio feature extraction module further includes video clue-driven feature extraction, and the video clue-driven feature extraction needs to meet the following requirements:

[0073] Continuity of facial light and shadow changes: Based on the assumption of illumination invariance, analyze whether the light and shadow distribution of the face in the video conforms to natural laws, and detect the authenticity of light source changes (for example, evaluate the continuity of brightness gradients frame by frame);

[0074] Consistency of feature point distribution: By analyzing the spatial distribution of facial key points (such as the relative positions of the corners of the eyes and the tip of the nose, the angle of the corners of the mouth, etc.), verify whether their geometric shapes conform to natural physiological characteristics and identify potential forgeries;

[0075] Tracking of facial key points: Use a Kalman filter or a long short-term memory network (LSTM) to model the trajectories of facial key points to ensure that the key points are continuous in the video stream and have no abnormal jumps;

[0076] Analysis of the authenticity of facial texture: Based on LBP (Local Binary Pattern) or a CNN texture detector, evaluate the facial skin texture features to determine whether it is a screen playback or a forged image;

[0077] Analysis of blink frequency and subtle facial muscle movements: Through frequency statistics and muscle movement modeling, detect the natural movement characteristics of blinks and facial muscles to further enhance the ability of live detection.

[0078] Furthermore, the specific content of the multi-modal clue fusion module includes feature alignment, weighted fusion, and fused feature extraction;

[0079] Feature alignment: Ensure the consistency of audio and video features within the same time window through a timestamp synchronization mechanism, fill in the time misalignment that may be caused by different sampling rates through a frame-level alignment strategy, and perform interpolation processing on missing or misaligned feature data through interpolation compensation to make the audio and video modalities consistent in time series;

[0080] Weighted fusion: Calculate the importance scores of audio features and video features for the task objective respectively through a multi-head attention mechanism, dynamically adjust the weights, and combine the audio and video features by feature weighting to generate a joint feature vector;

[0081] Fused feature extraction: Extract joint features through a residual network to enhance the non-linear expression ability of the features, avoid the problem of gradient disappearance at the same time, and map the fused features to a low-dimensional space through a fully connected layer for final classification or verification.

[0082] In this embodiment, the timestamp synchronization mechanism specifically refers to using the timestamp information of audio and video to accurately align the data of the two modalities to ensure the consistency of audio and video features within the same time window;

[0083] The frame-level alignment strategy specifically refers to matching video features (such as the dynamic trajectory of key points) and audio features (such as MFCC sequences) in units of frames to fill in the time misalignment that may be caused by different sampling rates;

[0084] The interpolation compensation specifically refers to interpolating the missing or misaligned feature data to make the audio-visual modalities consistent in time series;

[0085] Multi-head attention mechanism: Adopt the multi-head attention model in Transformer to calculate the importance scores of audio features and video features for the task objective respectively, and dynamically adjust the weights:

[0086] \text{Attention}(Q,K,V)

[0087] =\text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V

[0088] Q, K, and V are the query, key, and value vectors respectively, and d_k is the feature dimension;

[0089] Feature weighting specifically refers to weighted combination of audio and video features based on the attention scores to generate a joint feature vector;

[0090] F_{\text{joint}}

[0091] =w_{\text{audio}}\cdot F_{\text{audio}}+w_{\text{video}}

[0092] \cdot F_{\text{video}}

[0093] Among them, w_{\text{audio}} and w_{\text{video}} are dynamic weights;

[0094] Residual network (ResNet) specifically refers to using residual modules to further extract joint features, enhance the non-linear expression ability of features, and avoid the problem of gradient disappearance:

[0095] \text{Output}=F(x)+x

[0096] F(x) specifically refers to the non-linear features extracted by the residual block;

[0097] The fully connected layer specifically refers to mapping the fused features to a low-dimensional space for final classification or verification.

[0098] Furthermore, the specific content of the liveness detection module includes audio liveness detection, video liveness detection, and multi-modal liveness fusion;

[0099] Audio liveness detection: The liveness detection of audio information is completed through the cooperation of time-domain analysis and frequency-domain analysis;

[0100] Video liveness detection: The liveness detection of video information is completed through the cooperation of facial dynamic detection and 3D depth information analysis;

[0101] Multi-modal liveness fusion: Through the fusion model design, joint modeling of audio-visual features is carried out, the dynamic correlation between modalities is mined, and the judgment is output.

[0102] In this embodiment, the time-domain analysis includes short-time energy change detection and zero-crossing rate (ZCR) detection;

[0103] The short-time energy change detection specifically refers to analyzing the naturalness through the temporal variation of the short-time energy of the speech signal. For example, forged speech usually has problems such as overly smooth energy or unnatural fluctuations. The formula is as follows:

[0104] E[n]=\sum_{m=0}^{M-1}|x[m+n]|^2

[0105] Where E[n] is the frame energy and x[m+n] is the signal amplitude.

[0106] The zero-crossing rate (ZCR) detection specifically refers to analyzing the speech continuity by using the frequency of zero-crossing points of the speech signal. Forged speech often shows mutations in the transition frequency band;

[0107] The frequency-domain analysis includes formant analysis and speech continuity detection;

[0108] The formant analysis specifically refers to extracting the formant frequencies of the speech. The formant distribution of forged speech often does not conform to the natural characteristics of human speech;

[0109] The speech continuity detection specifically refers to analyzing the high-frequency attenuation characteristics of the speech through spectral time series. Forged speech usually shows abnormalities during the attenuation process.

[0110] The facial dynamic detection includes blink detection and lip movement detection;

[0111] The blink detection specifically refers to analyzing the blink frequency and naturalness based on the key points of the human face (such as the contour points of the eyes) to detect 2D photo forgery:

[0112] EAR=\frac{|p_2-p_6|+|p_3-p_5|}{2\cdot|p_1-p_4|}

[0113] EAR is the eye aspect ratio. If it is continuously lower than the threshold, it is judged as a closed eye;

[0114] Lip motion detection specifically refers to detecting the opening and closing movement of the lips. It is difficult to simulate natural mouth dynamic changes by forging videos or static images.

[0115] 3D depth information analysis specifically refers to using a stereo camera or optical flow method to analyze 3D depth information to determine whether the face has a real depth distribution. Forgery attacks (such as 2D photos) usually cannot generate consistent depth maps.

[0116] Fusion model design includes weight allocation and feature combination:

[0117] Weight allocation specifically refers to using a weighted fusion mechanism to dynamically weight the results of audio and video liveness detection according to their reliability:

[0118] R_{\text{Living}}

[0119] =w_{\text{audio}}\cdot R_{\text{audio}}+w_{\text{video}}

[0120] \cdot R_{\text{Video}}

[0121] Among them, R_{\text{live}} is the final liveness determination result, and the weights w_{\text{audio}} and w_{\text{video}} can be dynamically adjusted by the attention mechanism;

[0122] Feature combination specifically refers to the joint modeling of audio and video features through a deep fusion model (such as Transformer or Bi-LSTM) to explore the dynamic correlation between modalities.

[0123] Furthermore, the specific contents of the identity recognition and verification module include identity matching, decision threshold adjustment and result output;

[0124] Identity matching: Calculate the similarity between the fused features and the pre-stored feature library;

[0125] Decision threshold adjustment: Dynamically adjust the recognition threshold according to the application scenario to balance the false recognition rate and missed recognition rate;

[0126] Result output: Output the authentication result and record the verification log for auditing.

[0127] In this embodiment, identity matching specifically refers to calculating the similarity between the fused feature and the pre-stored feature library (using cosine similarity or Euclidean distance);

[0128] Decision threshold adjustment specifically refers to dynamically adjusting the recognition threshold according to the application scenario to balance the false recognition rate and missed recognition rate.

[0129] The result output specifically refers to outputting the authentication result and recording the verification log for auditing at the same time.

[0130] In summary, in the process of audio signal processing, this technical solution fully combines traditional acoustic features (such as MFCC and PLP), deep learning features (such as high-dimensional features extracted by ResNet221), and clue-driven feature analysis (such as frequency-domain consistency and noise distribution continuity), ensuring that forged audio is difficult to evade detection. By comprehensively analyzing the time-domain dynamics (short-time energy and zero-crossing rate changes) and frequency-domain features (formant distribution and its changes before and after), this technical solution has extremely strong robustness for the dynamic consistency detection of forged audio and can cope with the forgery challenges in various complex scenarios. Thanks to the timestamp synchronization mechanism and the multi-head attention mechanism, this technical solution not only achieves the precise alignment of audio and video features, but also deeply explores the correlation information between modalities through fusion feature modeling (residual network), showing significant superiority in liveness detection and forged audio-visual recognition. Through modular design and lightweight feature extraction process, this technical solution can be adapted to a variety of hardware environments and can be efficiently deployed from mobile devices to servers while ensuring real-time performance. Combining multi-modal feature extraction and deep learning, this technical solution significantly improves the ability of the authentication system to prevent complex forgery attacks (such as deepfake voice and video), meeting the application requirements of high-sensitivity fields such as finance, security, and government.

[0131] The above-disclosed is only a preferred embodiment of the present invention, and of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. An audio-visual identity recognition system based on multi-modal clue driving, characterized in that it includes an audio feature extraction module, a video feature extraction module, a multi-modal clue fusion module, a live detection module, and an identity recognition and verification module. The audio feature extraction module and the video feature extraction module are both connected to the input end of the multi-modal clue fusion module. The input end of the live detection module is connected to the output end of the multi-modal clue fusion module, and the output end of the live detection module is connected to the identity recognition and verification module; The audio feature extraction module is used to capture the user's voice signal and extract key voiceprint features; The video feature extraction module is used to obtain the user's facial features or limb features through a camera device and extract relevant features; The multi-modal clue fusion module is used to intelligently weight and fuse the features of the audio and video modalities through a fusion method based on the attention mechanism; The live detection module is used to comprehensively analyze the dynamic features of the audio and video to determine whether the user is a real individual; The identity recognition and verification module completes the final verification of the user's identity based on the fusion features and the live detection results.

2. The audio-visual identity recognition system based on multi-modal clue driving according to claim 1, characterized in that The specific content of the audio feature extraction module includes speech preprocessing, speech feature extraction, and deep feature learning; Speech preprocessing: Preprocess the speech signal through endpoint detection and pre-emphasis filtering; Speech feature extraction: Extract features from the speech signal through Mel-frequency cepstral coefficients and perceptual linear prediction; Deep feature learning: Use the pre-trained ResNet221 model. After inputting the features extracted by the feature extraction unit, extract the high-dimensional feature vector of the audio through a deep convolutional neural network to improve the learning ability of complex patterns in the speech signal.

3. The audio-visual identity recognition system based on multi-modal clue driving according to claim 2, characterized in that The specific content of the video feature extraction module includes face detection and key point extraction, facial feature extraction, and dynamic feature capture; Face detection and key point extraction: Extract the face information in the video information through face detection and key point extraction; Facial feature extraction: Based on the pre-trained ArcFace model, extract a 128-dimensional or 512-dimensional highly recognizable feature vector from the detected face as the core data for identity verification or matching; Dynamic feature capture: Extract the dynamic changes of the face through the optical flow method, and use the temporal convolutional network to capture the movement trajectory of the face in the video stream to enhance the expressiveness of the dynamic features.

4. The audio-visual identity recognition system based on multi-modal clue driving according to claim 3, characterized in that The specific content of the multi-modal clue fusion module includes feature alignment, weighted fusion, and fusion feature extraction; Feature alignment: The timestamp synchronization mechanism is used to ensure the consistency of audio and video features in the same time window. The frame-level alignment strategy is used to fill the time misalignment caused by different sampling rates. Interpolation compensation is used to interpolate missing or misaligned feature data to ensure that the audio and video modes are consistent in timing. Weighted fusion: The importance scores of audio features and video features to the task objectives are calculated through a multi-head attention mechanism, the weights are adjusted dynamically, and the audio and video features are weighted and combined through feature weighting to generate a joint feature vector. Fusion feature extraction: Extract joint features through the residual network to enhance the nonlinear expression ability of features while avoiding the gradient vanishing problem, and map the fusion features to a low-dimensional space through a fully connected layer for final classification or verification.

5. The multi-mode clue-driven audio and video identity recognition system according to claim 4, characterized in that: The specific contents of the liveness detection module include audio liveness detection, video liveness detection and multi-modal liveness fusion; Audio liveness detection: Liveness detection of audio information is completed through the combination of time domain analysis and frequency domain analysis; Video liveness detection: Liveness detection of video information is completed through the combination of facial dynamic detection and 3D depth information analysis; Multimodal living body fusion: Jointly model audio and video features through fusion model design, mine the dynamic correlation between modalities, and make judgment outputs.

6. The multi-mode clue-driven audio and video identity recognition system according to claim 5, characterized in that: The specific contents of the identity recognition and verification module include identity matching, decision threshold adjustment and result output; Identity matching: Calculate the similarity between the fused features and the pre-stored feature library; Decision threshold adjustment: Dynamically adjust the recognition threshold according to the application scenario to balance the false recognition rate and missed recognition rate; Result output: Output the authentication result and record the verification log for auditing.

7. The multi-mode clue-driven audio and video identity recognition system according to claim 6, characterized in that: When "calculating similarity between fused features and pre-stored feature library" completes identity matching, cosine similarity or Euclidean distance is used for similarity calculation.

Citation Information

Cited By

  • Multi-mode audio authentic identification system and method based on time-space consistency characteristics

    CN120564759A

  • Voiceprint verification system based on CANN architecture

    CN120853584A

  • Voiceprint verification system based on cann architecture

    CN120853584B

  • Intelligent lock identity authentication method and system for multi-mode voiceprint verification

    CN121075020A

  • A multi-modal voiceprint verification intelligent lock identity authentication method and system

    CN121075020B