Multi-modal data fusion identity authentication and security monitoring system based on deep learning

Through multimodal data fusion technology based on deep learning, the existing identity authentication system has been solved, and the problem of insufficient recognition accuracy and susceptible to environmental interference is achieved, achieving higher recognition accuracy and better user experience.

CN120048010APending Publication Date: 2025-05-27NANJING LONGYUAN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510113075.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing identity authentication and security monitoring systems have insufficient identification accuracy during the identification process and are susceptible to interference from environmental factors, resulting in poor security and user experience of the overall system.

Method used

The multimodal data fusion identity authentication and security monitoring system based on deep learning is adopted. The multimodal data of facial images, sounds and behavior patterns are collected through the identity authentication module to be integrated. The real-time analysis module collects and fuses video and audio data, and the remote control module stores and transmits data.

Benefits of technology

It improves recognition accuracy, reduces interference from environmental factors, and enhances the security and user experience of the overall identity authentication system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048010A_ABST
    Figure CN120048010A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of identity recognition, in particular to a multi-modal data fusion identity authentication and security monitoring system based on deep learning. Comprising an identity authentication module, a real-time analysis module and a remote control module, and the identity authentication module is used for collecting face images, sound and behavior pattern multi-modal data of a user, performing multi-modal data fusion and outputting identity authentication data of the user; the real-time analysis module is used for respectively collecting real-time video data and audio data, fusing the picture detection data and the sound detection data and outputting real-time abnormal behavior data; the remote control module is used for storing the identity authentication data and the real-time abnormal behavior data and transmitting the data in a remote access mode; through the above mode, the identification precision is improved during analysis, and the interference of environmental factors is reduced, so that the safety of the whole identity authentication system and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of identity recognition, and in particular, to a multi-modal data fusion identity authentication and security monitoring system based on deep learning. Background Art

[0002] Currently, the development of identity authentication and security monitoring systems has experienced a transformation from single biometric recognition technology to multi-factor authentication technology, and then to the combination of artificial intelligence and multi-sensor fusion technology. The progress of these technologies has not only improved the security and reliability of the systems, but also enhanced the user experience to a certain extent. However, these technologies still face many challenges in practical applications.

[0003] In the prior art, when independently applying face recognition, voiceprint recognition, and behavior pattern analysis, there are problems of insufficient recognition accuracy and susceptibility to environmental factor interference, resulting in poor security and user experience of the overall identity authentication system. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-modal data fusion identity authentication and security monitoring system based on deep learning, which solves the problems in the prior art that during the recognition process, the recognition accuracy is insufficient and it is susceptible to environmental factor interference, resulting in poor security and user experience of the overall identity authentication system.

[0005] To achieve the above purpose, a multi-modal data fusion identity authentication and security monitoring system based on deep learning adopted by the present invention includes an identity authentication module, a real-time analysis module, and a remote control module. The identity authentication module is respectively connected to the real-time analysis module and the remote control module, and the remote control module is respectively connected to the identity authentication module and the real-time analysis module;

[0006] The identity authentication module is used to collect multi-modal data of a user's face image, voice, and behavior pattern, perform multi-modal data fusion, and output the user's identity authentication data;

[0007] The real-time analysis module is used to respectively collect real-time video data and audio data, and fuse the picture detection data and the sound detection data to output real-time abnormal behavior data;

[0008] The remote control module is used to store the identity authentication data and the real-time abnormal behavior data, and transmit the data in a remote access manner.

[0009] Among them, the identity authentication module includes a face recognition unit, a voiceprint recognition unit, a behavior analysis unit, and a multi-modal data fusion unit. The face recognition unit, the voiceprint recognition unit, and the behavior analysis unit are respectively connected to the multi-modal data fusion unit, and the multi-modal data fusion unit is respectively connected to the real-time analysis module and the remote control module.

[0010] Among them, the face recognition unit is used to recognize the user's face features through the training of face images by using a convolutional neural network based on deep learning, and output face matching data.

[0011] The voiceprint recognition unit is used to analyze the user's voice features by using a recurrent neural network and output voice matching data.

[0012] The behavior analysis unit is used to output behavior matching data by analyzing the user's operation habits and behavior patterns.

[0013] The multi-modal data fusion unit is used to obtain face matching data, voice matching data, and behavior matching data, fuse the data, comprehensively evaluate the matching results, and output authentication data.

[0014] Among them, a labeled face image dataset is used to train the face recognition model.

[0015] The user's face images are collected in real time, the image data is input into the face recognition model, and the face feature vectors output by the face recognition model are obtained.

[0016] A first threshold is set, the face feature vectors output by the face recognition model are compared with the stored user face feature vectors, and the first cosine similarity is calculated.

[0017] The first threshold and the first cosine similarity are compared, and face matching data is output.

[0018] Among them, the process executed by the voiceprint recognition unit is as follows:

[0019] Mel frequency cepstral coefficients are obtained through short-time Fourier transform, the Mel frequency cepstral coefficients are used as voice features, and a labeled voice dataset is used to train the voiceprint recognition model.

[0020] The user's voice data is collected in real time, the features are extracted and then input into the voiceprint recognition model, and the voiceprint feature vectors output by the voiceprint recognition model are obtained.

[0021] A second threshold is set, the voiceprint feature vectors output by the voiceprint recognition model are compared with the stored user voiceprint feature vectors, and the second cosine similarity is calculated.

[0022] Compare the second threshold value and the second cosine similarity, and output voiceprint matching data.

[0023] Among them, the process executed by the behavior analysis unit is as follows:

[0024] Train the behavior analysis model using the labeled operation data set.

[0025] Collect the operation data of the user in real time, extract key features, and input them into the behavior analysis model to obtain classification results; among them, the operation data includes key press frequency, sliding trajectory, input speed, and the key features include statistical features such as mean, standard deviation, maximum value, and minimum value in the time series and signal frequency domain features.

[0026] Set a third threshold value, compare the classification result output by the behavior analysis model with the stored user behavior pattern, and calculate the confidence level of the classification result.

[0027] Compare the third threshold value and the confidence level, and output behavior result data.

[0028] Among them, the real-time analysis module includes a video processing unit, an audio processing unit, and a time series analysis unit. The video processing unit and the audio processing unit are respectively connected to the identity authentication module, and the time series analysis unit is respectively connected to the video processing unit, the audio processing unit, and the remote control module.

[0029] Among them, the video processing unit is used to obtain real-time video data, detect abnormal behaviors in the real-time picture, and output picture detection data.

[0030] The audio processing unit is used to obtain real-time audio data, detect abnormal behaviors in the real-time sound, and output sound detection data.

[0031] The time series analysis unit is used to fuse the picture detection data and the sound detection data, analyze the abnormal situations in the environment, and output abnormal behavior data.

[0032] A multi-modal data fusion identity authentication and security monitoring system based on deep learning according to the present invention, wherein the identity authentication module is used to collect multi-modal data of the user's facial image, voice, and behavior pattern, perform multi-modal data fusion, and output user identity authentication data; the real-time analysis module is used to respectively collect real-time video data and audio data, and fuse the picture detection data and the sound detection data to output real-time abnormal behavior data; the remote control module is used to store the identity authentication data and the real-time abnormal behavior data, and transmit the data in a remote access manner; achieving improved recognition accuracy during analysis, reducing interference from environmental factors, and thus improving the security and user experience of the overall identity authentication system. Description of the Drawings

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is a partial structural schematic diagram of the multi-modal data fusion identity authentication and security monitoring system based on deep learning of the present invention.

[0035] Figure 2 It is a structural schematic diagram of the multi-modal data fusion identity authentication and security monitoring system based on deep learning of the present invention.

[0036] Figure 3 It is a schematic flowchart of the process executed by the face recognition unit of the present invention.

[0037] Figure 4 It is a schematic flowchart of the process executed by the voiceprint recognition unit of the present invention.

[0038] Figure 5 It is a schematic flowchart of the process executed by the behavior analysis unit of the present invention.

[0039] 100 - Identity authentication module, 101 - Face recognition unit, 102 - Voiceprint recognition unit, 103 - Behavior analysis unit, 104 - Multi-modal data fusion unit, 200 - Real-time analysis module, 201 - Video processing unit, 202 - Audio processing unit, 203 - Timing analysis unit, 300 - Remote control module. Detailed implementation manners

[0040] Please refer to Figures 1 to 5 , the present invention provides a multi-modal data fusion identity authentication and security monitoring system based on deep learning, including an identity authentication module 100, a real-time analysis module 200, and a remote control module 300. The identity authentication module 100 is respectively connected to the real-time analysis module 200 and the remote control module 300, and the remote control module 300 is respectively connected to the identity authentication module 100 and the real-time analysis module 200;

[0041] The identity authentication module 100 is used to collect multi-modal data of the user's face image, voice, and behavior pattern, perform multi-modal data fusion, and output the user's identity authentication data;

[0042] The real-time analysis module 200 is configured to collect real-time video data and audio data respectively, fuse the picture detection data and the sound detection data, and output real-time abnormal behavior data;

[0043] The remote control module 300 is configured to store the identity authentication data and the real-time abnormal behavior data, and transmit the data in a remote access manner.

[0044] In this embodiment, the identity authentication module 100 collects multi-modal data of the user's facial image, voice, and behavior pattern, performs multi-modal data fusion, and outputs the user's identity authentication data; the real-time analysis module 200 collects real-time video data and audio data respectively, fuses the picture detection data and the sound detection data, and outputs real-time abnormal behavior data; the remote control module 300 stores the identity authentication data and the real-time abnormal behavior data, and transmits the data in a remote access manner; through the above method, the recognition accuracy is improved during analysis, the interference of environmental factors is reduced, and thus the security and user experience of the overall identity authentication system are improved.

[0045] Further, the identity authentication module 100 includes a face recognition unit 101, a voiceprint recognition unit 102, a behavior analysis unit 103, and a multi-modal data fusion unit 104. The face recognition unit 101, the voiceprint recognition unit 102, and the behavior analysis unit 103 are respectively connected to the multi-modal data fusion unit 104, and the multi-modal data fusion unit 104 is respectively connected to the real-time analysis module 200 and the remote control module 300.

[0046] Further, the face recognition unit 101 is configured to use a convolutional neural network based on deep learning to identify the user's facial features through the training of facial images, and output facial matching data;

[0047] The voiceprint recognition unit 102 is configured to analyze the user's voice features using a recurrent neural network and output voice matching data;

[0048] The behavior analysis unit 103 is configured to output behavior matching data by analyzing the user's operation habits and behavior patterns;

[0049] The multi-modal data fusion unit 104 is configured to obtain the facial matching data, the voice matching data, and the behavior matching data, fuse the data, comprehensively evaluate the matching results, and output authentication data.

[0050] In this embodiment, the face recognition unit 101 adopts a convolutional neural network based on deep learning to identify user face features through the training of face images and outputs face matching data; the voiceprint recognition unit 102 uses a recurrent neural network to analyze user voice features and outputs voice matching data; the behavior analysis unit 103 outputs behavior matching data by analyzing user operation habits and behavior patterns; the multimodal data fusion unit 104 obtains face matching data, voice matching data and behavior matching data, fuses the data, comprehensively evaluates the matching results, and outputs authentication data.

[0051] Further, the process executed by the face recognition unit 101 is as follows:

[0052] S401: Train the face recognition model using the labeled face image dataset;

[0053] S402: Real-time collect the user's face image, input the image data into the face recognition model, and obtain the face feature vector output by the face recognition model;

[0054] S403: Set a first threshold, compare the face feature vector output by the face recognition model with the stored user face feature vector, and calculate the first cosine similarity;

[0055] S404: Compare the first threshold and the first cosine similarity, and output the face matching data.

[0056] In this embodiment, the face recognition unit 101 adopts a convolutional neural network based on deep learning to identify user face features through the training of face images and output face matching data. The process executed by the face recognition unit 101 is as follows: ResNet-50 is used as the basic network, leveraging its powerful feature extraction ability. ResNet-50 consists of multiple residual blocks, and each residual block contains multiple convolutional layers, which can effectively extract face features at different levels. The input to the network is a 224x224 pixel color image, and the output is a face feature vector with a dimension of 128. A variant of ResNet based on the attention mechanism, called CBAM-ResNet, is used to further improve the effect of face feature extraction. It is trained with a large number of labeled face image datasets, such as CelebA, LFW, etc., using the cross-entropy loss function and the Adam optimizer. The learning rate starts from 0.001 and decays to half of the original every 5 epochs for end-to-end learning to optimize the network parameters. Data augmentation techniques such as random rotation, scaling, and translation are also adopted to prevent overfitting and improve the generalization ability of the model. The user's face image is collected in real time, input into the trained CBAM-ResNet model, the face feature vector is output, and compared with the stored user face feature vector. The first cosine similarity is calculated. If the first threshold is set above 0.85, the identity is considered to match; if it is lower than the first threshold, further identity authentication measures are triggered.

[0057] Further, the process executed by the voiceprint recognition unit 102 is as follows:

[0058] S501: Obtain Mel-frequency cepstral coefficients through short-time Fourier transform, use the Mel-frequency cepstral coefficients as voice features, and train the voiceprint recognition model with a labeled voice dataset;

[0059] S502: Collect the user's voice data in real time, extract features and input them into the voiceprint recognition model, and obtain the voiceprint feature vector output by the voiceprint recognition model;

[0060] S503: Set a second threshold, compare the voiceprint feature vector output by the voiceprint recognition model with the stored user voiceprint feature vector, and calculate the second cosine similarity;

[0061] S504: Compare the second threshold and the second cosine similarity, and output the voiceprint matching data.

[0062] In this embodiment, the voiceprint recognition unit 102 analyzes the user's voice features using a recurrent neural network and outputs voice matching data. The process executed by the voiceprint recognition unit 102 is as follows: By using a long short-term memory network, it can process sequential data and capture long-term dependencies. The input of the LSTM network is a sequence of 128-dimensional MFCC feature vectors, and the output is a voiceprint feature vector with a dimension of 128. To improve the accuracy of voiceprint recognition, a BiLSTM (bidirectional LSTM) model is introduced, which combines forward and backward voiceprint features and improves the recognition ability of the model. Mel-frequency cepstral coefficients (MFCC) are used as voice features, and the MFCC features are obtained through short-time Fourier transform (STFT). Each voice segment is divided into frames of 20 ms with an overlap of 10 ms. To enhance the robustness of voice features, the PLP feature extraction technique is adopted. Training is performed using a large number of labeled voice datasets, such as VoxCeleb, etc. The mean squared error loss function and the Adam optimizer are used. The learning rate starts from 0.001 and is decayed to half of the original value every 10 epochs for end-to-end learning to optimize the network parameters. Data augmentation techniques such as time stretching and random noise addition are also adopted. The voice data of the user is collected in real time, the PLP features are extracted, input into the trained BiLSTM model, the voiceprint feature vector is output, and compared with the stored user voiceprint feature vector. The second cosine similarity is calculated. If the second threshold is set above 0.75, the identity is considered to match; if it is lower than the second threshold, further identity authentication measures are triggered.

[0063] Further, the process executed by the behavior analysis unit 103 is as follows:

[0064] S601: Train the behavior analysis model using a labeled operation dataset;

[0065] S602: Collect the operation data of the user in real time, extract key features, and input them into the behavior analysis model to obtain a classification result; where the operation data includes key press frequency, sliding trajectory, and input speed, and the key features include statistical features such as mean, standard deviation, maximum value, and minimum value in the time series and signal frequency domain features;

[0066] S603: Set a third threshold, compare the classification result output by the behavior analysis model with the stored user behavior pattern, and calculate the confidence of the classification result;

[0067] S604: Compare the third threshold and the confidence, and output behavior result data.

[0068] In this embodiment, the behavior analysis unit 103 outputs behavior matching data by analyzing the user's operation habits and behavior patterns. The process executed by the behavior analysis unit 103 is as follows: It collects the user's operation data in real time, including key press frequency, sliding trajectory, input speed, etc. The data collection frequency is 10 times per second, and the sliding window technology is used to process the continuous data stream. In order to better capture the changes in behavior characteristics, the dynamic time warping (DTW) technology is adopted for trajectory matching. The collected data is processed to extract key features, including statistical features such as mean, standard deviation, maximum value, minimum value, etc. in the time series, and the frequency domain features of the signal. In order to improve the expression ability of the features, an autoencoder is used for feature learning to extract high-dimensional feature vectors. The random forest algorithm is used to train the model, and the model parameters are optimized through a large number of labeled data sets. The training data includes the behavior pattern samples of different users, and each sample contains 10 seconds of continuous behavior data. In order to improve the generalization ability of the model, an ensemble learning method is adopted to perform weighted averaging on the prediction results of multiple random forest models. The user's behavior data is collected in real time and input into the trained random forest model, and the classification result is output. The classification result is compared with the stored user behavior pattern. If the confidence level of the classification result is higher than the third threshold of 0.8, it is considered that the identity matches; if it is lower than the third threshold, further identity authentication measures are triggered.

[0069] The multi-modal data fusion unit 104 comprehensively evaluates the results of face recognition, voiceprint recognition, and behavior pattern analysis to provide high-precision identity authentication services. The specific implementation is as follows:

[0070] The results of face recognition, voiceprint recognition, and behavior pattern analysis are normalized, and the range of the normalized results is [0,1]. Fusion is performed through a weighted average or voting mechanism. The weighted average weights are 0.4 for face recognition, 0.35 for voiceprint recognition, and 0.25 for behavior pattern analysis. In order to more effectively fuse data of different modalities, a mixture Gaussian model is adopted for feature fusion to improve the final recognition accuracy.

[0071] The Bayesian decision theory is adopted to calculate the probability of user identity authentication to ensure the accuracy and reliability of the authentication. By calculating the prior probability and conditional probability, combined with the Bayesian formula:

[0072] P(Authentication successful|Feature data) = \frac{P(Feature data|Authentication successful) \cdot P(Authentication successful)}{P(Feature data)}

[0073] Finally, the probability result of user identity authentication is obtained.

[0074] Through the above multi-modal data fusion technology, the system can provide more accurate and reliable identity authentication services. For users who fail the preliminary identity authentication, the system will trigger a further secondary authentication process and conduct a dynamic risk assessment of the real-time analysis module 200 to ensure the security of the system and the user experience.

[0075] Furthermore, the real-time analysis module 200 includes a video processing unit 201, an audio processing unit 202, and a timing analysis unit 203. The video processing unit 201 and the audio processing unit 202 are respectively connected to the identity authentication module 100, and the timing analysis unit 203 is respectively connected to the video processing unit 201, the audio processing unit 202, and the remote control module 300.

[0076] Furthermore, the video processing unit 201 is used to obtain real-time video data, detect abnormal behaviors in the real-time video, and output video detection data.

[0077] The audio processing unit 202 is used to obtain real-time audio data, detect abnormal behaviors in the real-time sound, and output sound detection data.

[0078] The timing analysis unit 203 is used to fuse the video detection data and the sound detection data, analyze abnormal situations in the environment, and output abnormal behavior data.

[0079] In this embodiment, the video processing unit 201 obtains real-time video data, detects abnormal behaviors in the real-time video, and outputs video detection data; the audio processing unit 202 obtains real-time audio data, detects abnormal behaviors in the real-time sound, and outputs sound detection data; the timing analysis unit 203 fuses the video detection data and the sound detection data, analyzes abnormal situations in the environment, and outputs abnormal behavior data.

[0080] The video processing unit 201 adopts an image recognition technology based on deep learning, and specifically uses the YOLOv5 model for abnormal behavior detection. YOLOv5 is a real-time object detection algorithm, and its network structure is based on CSPDarknet53, which can quickly and accurately identify intrusion behaviors, abnormal actions, etc. in the video. Its detection process can be expressed as:

[0081] Output = YOLOv5(Input\Image);

[0082] Among them, Output contains the location and type of the detected abnormal behavior. During the implementation process, the training data of the YOLOv5 model includes a large number of abnormal behavior samples, such as intrusion behaviors, abnormal actions, etc. Through the training of a large-scale dataset, the system can accurately identify various abnormal situations. The data input of the real-time analysis module 200 is transmitted in real time through a camera. The video stream data is preprocessed and then input into the YOLOv5 model for analysis. The results output by the video processing unit 201 will be transmitted to the identity authentication module 100 and the remote control module 300 through the API interface.

[0083] The audio processing unit 202 adopts real-time analysis technology based on a deep neural network. Specifically, the ResNet18 model is used to classify environmental sounds to identify potential security threats, such as abnormal noises, alarm sounds, etc. ResNet18 is a residual network with 18 layers, which can effectively extract audio features and perform classification. Its processing process can be described as:

[0084] \text{Class}=\text{ResNet18}(Input\Audio);

[0085] Among them, Class represents the audio classification result. During the implementation process, the audio signal is collected in real time through a microphone, preprocessed and then input into the ResNet18 model for analysis. The training data of the model includes various environmental noises and abnormal sound samples to ensure that the system can accurately identify potential security threats. The output result of the audio processing unit 202 is also transmitted to the identity authentication module 100 and the remote control module 300 through the API interface.

[0086] The timing analysis unit 203 improves the accuracy and real-time performance of event detection by combining the time series analysis of video and audio data. Specifically, the timing analysis unit 203 uses a long short-term memory network for comprehensive analysis of timing data. The long short-term memory network can capture the time dependence of the data, so as to more accurately identify and predict the occurrence of events. Its mathematical expression is:

[0087] \text{Output}_{t}=\text{LSTM}(Input_{t},\text{Hidden}_{t - 1});

[0088] Among them, $\text{Output}_{t}$ is the output result at time $t$, $Input_{t}$ is the input video and audio data, and $\text{Hidden}_{t - 1}$ is the state at the previous moment. The processing of the long short-term memory network in the time series enhances the real-time performance and accuracy of the system. Through the fusion analysis of multi-modal data, the system can more comprehensively understand the abnormal situations in the environment and reduce false alarms and missed alarms. The output result of the time series analysis unit 203 is transmitted to the identity authentication module 100 and the remote control module 300 through the API interface to ensure that the system can perform further identity verification and remote operations according to the detection results.

[0089] The above-disclosed are only one or more preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A multimodal data fusion identity authentication and security monitoring system based on deep learning, characterized in that: It includes an identity authentication module, a real-time analysis module and a remote control module, wherein the identity authentication module is connected to the real-time analysis module and the remote control module respectively, and the remote control module is connected to the identity authentication module and the real-time analysis module respectively; The identity authentication module is used to collect multimodal data of the user's facial image, voice and behavior pattern, perform multimodal data fusion, and output the user's identity authentication data; The real-time analysis module is used to collect real-time video data and audio data respectively, and fuse the picture detection data and the sound detection data to output real-time abnormal behavior data; The remote control module is used to store identity authentication data and real-time abnormal behavior data, and transmit data in a remote access manner.

2. The multimodal data fusion identity authentication and security monitoring system based on deep learning as claimed in claim 1, characterized in that: The identity authentication module includes a facial recognition unit, a voiceprint recognition unit, a behavior analysis unit and a multimodal data fusion unit. The facial recognition unit, the voiceprint recognition unit and the behavior analysis unit are respectively connected to the multimodal data fusion unit, and the multimodal data fusion unit is respectively connected to the real-time analysis module and the remote control module.

3. The multimodal data fusion identity authentication and security monitoring system based on deep learning as claimed in claim 2, characterized in that: The facial recognition unit is used to use a convolutional neural network based on deep learning to recognize the user's facial features through training of facial images and output facial matching data; The voiceprint recognition unit is used to analyze the user's voice features using a recurrent neural network and output voice matching data; The behavior analysis unit is used to output behavior matching data by analyzing user operation habits and behavior patterns; The multimodal data fusion unit is used to obtain facial matching data, voice matching data and behavior matching data, fuse the data, conduct a comprehensive evaluation of the matching results, and output authentication data.

4. The multimodal data fusion identity authentication and security monitoring system based on deep learning as claimed in claim 3, characterized in that: The process performed by the facial recognition unit is: Use labeled facial image datasets to train facial recognition models; Collect user facial images in real time, input the image data into the facial recognition model, and obtain the facial feature vector output by the facial recognition model; Setting a first threshold, comparing the facial feature vector output by the facial recognition model with the stored facial feature vector of the user, and calculating a first cosine similarity; Compare the first threshold and the first cosine similarity, and output facial matching data.

5. The multimodal data fusion identity authentication and security monitoring system based on deep learning as claimed in claim 3, characterized in that: The process performed by the voiceprint recognition unit is: The Mel frequency cepstral coefficients are obtained through short-time Fourier transform, and the Mel frequency cepstral coefficients are used as speech features. The voiceprint recognition model is trained using the labeled speech data set. Collect user voice data in real time, extract features and input them into the voiceprint recognition model, and obtain the voiceprint feature vector output by the voiceprint recognition model; Setting a second threshold, comparing the voiceprint feature vector output by the voiceprint recognition model with the stored user voiceprint feature vector, and calculating a second cosine similarity; Compare the second threshold and the second cosine similarity, and output the voiceprint matching data.

6. The multimodal data fusion identity authentication and security monitoring system based on deep learning as claimed in claim 3, characterized in that: The process performed by the behavior analysis unit is: Use labeled action datasets to train behavior analysis models; Collect user operation data in real time, extract key features, and input them into the behavior analysis model to obtain classification results; wherein the operation data includes key frequency, sliding trajectory, and input speed, and the key features include statistical features of mean, standard deviation, maximum value, minimum value in time series and signal frequency domain features; Setting a third threshold, comparing the classification result output by the behavior analysis model with the stored user behavior pattern, and calculating the confidence of the classification result; The third threshold value and the confidence level are compared, and the behavior result data is output.

7. The multimodal data fusion identity authentication and security monitoring system based on deep learning as claimed in claim 1, characterized in that: The real-time analysis module includes a video processing unit, an audio processing unit and a timing analysis unit. The video processing unit and the audio processing unit are respectively connected to the identity authentication module, and the timing analysis unit is respectively connected to the video processing unit, the audio processing unit and the remote control module.

8. The multimodal data fusion identity authentication and security monitoring system based on deep learning as claimed in claim 7, characterized in that: The video processing unit is used to obtain real-time video data, perform abnormal behavior detection on the real-time picture, and output picture detection data; The audio processing unit is used to obtain real-time audio data, perform abnormal behavior detection on the real-time sound, and output sound detection data; The time series analysis unit is used to integrate the picture detection data and the sound detection data, analyze the abnormal situation in the environment, and output abnormal behavior data.

Citation Information

Patent Citations

  • Multi-modal data intelligent analysis system and method

    CN116881335A

  • Face recognition system based on machine learning

    CN117115881A

  • Multi-modal identity authentication method and system based on deep fusion of incomplete information

    CN117155583A

  • Multi-mode identity verification system and method

    CN118245994A

  • Multi-modal biological feature fusion personnel identity recognition method and platform based on deep learning

    CN118247818A

Cited By

  • Remote identity authentication method based on multiple video recognition

    CN120223444A

  • Television terminal education application multi-mode user authentication method and system

    CN120956974A

  • Multi-modal user authentication method and system for television-based education applications

    CN120956974B