Remote authentication method based on multiple video recognition

By extracting and fusing facial, behavioral, and speech features through multiple video streams and an improved self-supervised dynamic graph neural network, the adaptability problem of single-video stream facial recognition in complex environments is solved, achieving efficient and secure identity authentication.

CN120223444BActive Publication Date: 2025-11-14XIAOYUAN PERCEPTION (HULUDAO) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510695537.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-11-14
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing identity authentication methods are poorly adaptable to complex environments. Single-stream facial recognition technology is easily affected by changes in lighting, angle deviations, and occlusions, while multimodal recognition methods lack effective feature fusion and environmental adaptability.

Method used

We employ a multi-video stream combined with an improved self-supervised dynamic graph neural network. By constructing a dynamic graph structure, we extract facial, behavioral, and speech features. We utilize dynamic edge weights and a self-supervised learning mechanism to optimize the consistency and correlation of modal features and adjust the recognition strategy in real time to improve authentication accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and anti-counterfeiting capabilities of identity authentication in complex environments, enhances adaptability to changes in lighting and angular deviations, improves the fusion effect of multimodal features, and provides an efficient and secure remote identity authentication solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223444B_ABST
    Figure CN120223444B_ABST
Patent Text Reader

Abstract

This invention discloses a remote identity authentication method based on multi-video recognition, comprising the following steps: S1, acquiring real-time video stream data of the target user; S2, preprocessing the acquired video stream data to generate preprocessed data; S3, using an improved self-supervised dynamic graph neural network to extract features from the preprocessed data and generate corresponding feature vector representations; S4, fusing the feature vector representations; S5, comparing and verifying the fused feature vectors, assigning weights to the authentication results of each video stream based on the authentication results of multiple video streams, and generating a final authentication result; S6, monitoring environmental changes in the video streams in real time and updating network parameters in real time; S7, determining whether to authorize access or perform subsequent operations based on the final authentication result. This invention combines multiple video streams and an improved self-supervised dynamic graph neural network to extract and fuse multimodal features, effectively improving the accuracy, security, and robustness of remote identity authentication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of identity authentication technology, and in particular to a remote identity authentication method based on multiple video recognition. Background Technology

[0002] With the widespread adoption of the internet and mobile devices, remote identity authentication has gradually become an important means of ensuring information security and privacy protection. Traditional authentication methods, such as password-based, SMS verification code, and fingerprint recognition, while improving security to some extent, still have many limitations. Especially in scenarios where users face high security requirements or need contactless authentication, traditional authentication methods are vulnerable to attacks or forgery. Meanwhile, with the development of biometric recognition technology, facial recognition, fingerprint recognition, and iris recognition have gradually become important means in the field of identity authentication. Facial recognition technology, as a contactless identity verification method, has been widely adopted in practical applications due to its convenience and high security. However, existing facial recognition technology still faces some technical challenges, particularly in complex environments and diverse application scenarios.

[0003] Currently, facial recognition technologies based on a single video stream typically rely on image or video information from a single viewpoint for authentication. While this approach performs well in standard environments, in real-world applications, facial recognition technology is often affected by changes in lighting, angular deviations, occlusion, and video quality, leading to decreased recognition accuracy. For example, faces in a video may change due to the user's head rotation, or facial features may not be clearly recognizable in low-light conditions. In such cases, single-video-stream facial recognition methods may fail to effectively handle these complex situations, resulting in authentication failure. Therefore, existing single-video-stream-based facial recognition technologies have poor adaptability in complex scenarios and cannot provide sufficient robustness.

[0004] Besides facial recognition, behavioral recognition and voice recognition, as other forms of biometric identification technologies, are also widely used. Behavioral recognition technology authenticates users by analyzing their actions, postures, or gestures, while voice recognition authenticates users by analyzing their voice characteristics. While these technologies improve the security of identity authentication to some extent, due to their inherent limitations, they often cannot work independently of facial recognition technology in practical applications. Behavioral and voice recognition are often susceptible to environmental interference, such as background noise, the number of people in the video, and unclear speech, which may lead to unstable and inaccurate recognition results.

[0005] Existing identity authentication systems typically rely on single biometric features for authentication, thus their security and accuracy fall short of ideal levels. For example, facial recognition alone can be easily fooled by photos or videos, speech recognition can be affected by differences in audio quality or noise, and behavioral recognition can be affected by background interference or changes in user location. To overcome the shortcomings of single biometric features, multimodal biometric recognition methods have gradually become a research hotspot in recent years. These methods improve the security and robustness of identity authentication by combining multiple biometric features such as facial, speech, and behavioral features. However, the challenge of multimodal recognition methods lies in how to efficiently fuse information from different modalities to fully leverage the advantages of each modality.

[0006] Existing multimodal recognition methods typically employ separate models to extract each biometric feature, followed by feature fusion. This approach has several drawbacks. First, because each modality uses a different processing model, the feature extraction process for each modality is conducted independently, failing to fully utilize the inherent relationships between modalities. Second, most existing multimodal fusion methods rely on simple feature concatenation and weighted averaging, lacking in-depth analysis of the dynamic correlations between modalities. These methods often ignore the complex dependencies between modal features, resulting in limited fusion effectiveness and ultimately affecting the accuracy of authentication.

[0007] With the development of graph neural networks (GNNs), graph-based structured learning methods have been gradually introduced into multimodal recognition tasks. GNNs possess excellent spatiotemporal modeling capabilities, effectively propagating information within graph structures and capturing complex relationships between nodes. In multimodal biometric recognition, GNNs construct graph structures of multimodal features, enabling information propagation between different nodes and fusing information from different modalities. In this way, GNNs not only preserve spatiotemporal dependencies but also achieve deep fusion of multimodal features, further improving the accuracy and robustness of identity authentication.

[0008] However, existing graph neural networks still have some shortcomings in multimodal recognition applications. Especially when processing large-scale video data, existing graph neural network models often have high computational complexity, cumbersome training processes, and insufficient adaptability to environmental changes. Therefore, how to optimize the structure of graph neural networks to better adapt to the multimodal characteristics and spatiotemporal dependencies of video data remains an urgent problem to be solved.

[0009] In summary, existing identity authentication methods still face many challenges in multimodal recognition and video data processing. The shortcomings of current technologies mainly lie in the limitations of single-stream recognition, the inadequacy of multimodal feature fusion methods, and the deficiencies of traditional deep learning models in spatiotemporal modeling. Therefore, a new method is needed that can effectively fuse data from different video streams and adaptively adjust recognition strategies in different environments, thereby improving the accuracy, robustness, and security of identity authentication. Summary of the Invention

[0010] One objective of this invention is to propose a remote identity authentication method based on multiple video recognition. This invention combines multiple video streams and an improved self-supervised dynamic graph neural network. By constructing a dynamic graph structure and extracting multimodal features of face, behavior, and speech, it effectively handles spatiotemporal dependencies, improves feature fusion performance, and enhances consistency and correlation between different modalities by utilizing dynamic edge weights and a self-supervised learning mechanism, thereby further improving authentication accuracy.

[0011] A remote identity authentication method based on multiple video recognition according to an embodiment of the present invention includes the following steps:

[0012] S1. Collect real-time video stream data of the target user from different angles using a camera;

[0013] S2. Preprocess the acquired video stream data to remove noise, enhance the image, segment and calibrate the video frames, locate key feature regions in the video, and generate preprocessed data.

[0014] S3. Use an improved self-supervised dynamic graph neural network to extract facial features, behavioral features and speech features from the preprocessed data. The improved self-supervised dynamic graph neural network constructs a dynamic graph structure to represent the spatiotemporal dependencies in the preprocessed data. Combined with a self-supervised learning mechanism, it performs unsupervised training through a comparative learning task to maximize the consistency between modal features, optimize the correlation between modal features, and extract facial features, behavioral features and speech features to generate corresponding feature vector representations.

[0015] S4. The feature vector representations of facial features, behavioral features, and voice features are fused to generate a comprehensive identity authentication feature vector.

[0016] S5. Compare and verify the generated comprehensive identity authentication feature vector, authenticate the user features with the preset template stored in the database, assign weights to the authentication results of each video stream based on the authentication results of multiple video streams using a weighted method based on dynamic learning, adjust the weights in real time according to the video stream quality and the reliability score of the authentication features in the video stream, and calculate the final authentication result using a weighted average method.

[0017] S6. During the authentication process, monitor the environmental changes of the video stream in real time, adjust the recognition strategy, and update the parameters of the improved self-supervised dynamic graph neural network in real time.

[0018] S7. Output the final authentication result. Based on the final authentication result, determine whether to authorize access or perform subsequent operations. If the authentication fails, take corresponding security protection measures.

[0019] Optionally, S3 specifically includes:

[0020] S31. Using an improved method to extract multimodal features from preprocessed data, a dynamic graph structure is constructed to represent the spatiotemporal dependencies in the preprocessed data. The dynamic graph structure changes dynamically over time. The multimodal features include facial features, behavioral features, and speech features. The improved graph convolution operation introduces dynamic edge weights and an improved self-supervised learning contrastive loss function. The contrastive loss function is optimized using different loss functions at different training stages.

[0021] ;

[0022] Where G(t) represents the dynamic graph structure, V(t) represents the node set of the dynamic graph, which is the feature in the video frame, E(t) represents the edge set of the dynamic graph, which is the spatiotemporal dependency between video frames, and X(t) represents the node feature matrix.

[0023] S32. Introduce dynamic edge weights and update the node representations through graph convolution operations:

[0024] ;

[0025] in, Indicates that node v is at the th The feature vector representation of the layer, where v represents a node in the dynamic graph. Let N(v) represent the non-linear activation function, and let N(v) represent the set of neighboring nodes of node v. This represents the degree of node k. This represents the weight matrix of the graph convolutional layer. Indicates that node v is at the th The feature vector representation of the layer, where k represents a node in the dynamic graph. This represents the dynamic edge weight between node v and its neighbor node k at time t:

[0026] ;

[0027] Where exp represents the natural exponential function, and N(v) represents the set of neighboring nodes of node v. Represents nodes in a dynamic graph. This represents the feature vector representation of node v. This represents the feature vector representation of node k. Represents a node The feature vector representation, k represents a node in the dynamic graph, and Sim represents the cosine similarity between the feature vector representations. This represents the hyperparameter that is being adjusted. The cosine similarity between the feature vector representations of node k and node v is denoted as . Represents a node The cosine similarity between the feature vector representations of node v and node v;

[0028] S33. Extract facial features, use node representations in a self-supervised dynamic graph neural network to capture facial key point information, update facial features through graph convolution operations, and dynamically adjust based on the similarity between facial nodes to generate feature vector representations of facial features:

[0029] ;

[0030] in, This indicates that the facial feature node v is at the 1st... The feature vector representation of the layer, This represents a non-linear activation function, where N(face) represents the set of neighboring nodes of the facial feature node. This represents the degree of node u. This represents the dynamic edge weights between face nodes at time t. Let represent the feature vector representation of node u at layer l, where u represents a node in the dynamic graph. The graph convolution weight matrix representing facial features. Indicates the bias term;

[0031] S34. Extract behavioral features and use time node representations in a self-supervised dynamic graph neural network to process motion information in the video. Process the dynamic behavioral data at each time node through graph convolutional layers, combine spatiotemporal dependency modeling, update the behavioral features, and generate feature vector representations of the behavioral features:

[0032] ;

[0033] in, The behavioral feature time node t represents the time node at the tth time node. The feature vector representation of the layer, This represents a non-linear activation function, where N(action) represents the set of neighboring nodes of the behavioral feature node. Represents a node The degree, This represents the dynamic edge weights between behavioral feature time nodes t. Representing behavioral characteristics at specific time points In the The feature vector representation of the layer, The time points that represent the behavioral characteristics of a dynamic graph. The graph convolution weight matrix represents behavioral features. Indicates the bias term;

[0034] S35. Extract speech features through graph convolution operations to generate feature vector representations of the speech features:

[0035] ;

[0036] in, The speech feature node v represents the first... Layer feature vector representation, Let N(speech) represent a non-linear activation function, and let N(speech) represent the set of neighboring nodes of the speech feature node. Represents a node The degree, Represents the dynamic edge weights between speech features. Represents speech feature nodes In the The feature vector representation of the layer, Represents the speech feature nodes of a dynamic graph. The graph convolution weight matrix representing speech features. Indicates the bias term;

[0037] S36. The modal feature consistency of facial features, behavioral features and speech features is optimized through a contrastive loss function of self-supervised learning. The contrastive loss function minimizes the distance between features of the same class and maximizes the distance between features of different classes.

[0038] S37. The feature vector representation is further optimized through the reconstruction loss function obtained by self-supervised learning, and the parameters of the dynamic graph neural network are updated through the backpropagation algorithm. The reconstruction loss function is adjusted by comparing the difference between the input features and the reconstructed features to minimize the reconstruction error.

[0039] ;

[0040] in, Let X represent the reconstruction loss function, and let X represent the original input feature matrix. This represents the feature matrix obtained from the reconstruction by the dynamic graph neural network. Let Frobenius norm be the matrix X and ? The differences between them.

[0041] Optionally, S36 specifically includes:

[0042] S361. Define a contrastive learning task: by minimizing the distance between samples of the same class and maximizing the distance between samples of different classes, capture the similarity and difference between modal features of different video streams. The modal features include facial features, behavioral features, and speech features.

[0043] S362. Select positive and negative samples from the facial features, behavioral features and voice features extracted from the dynamic graph neural network. The positive samples represent samples belonging to the same category, and the negative samples represent samples with different features and belonging to different categories.

[0044] S363. For each pair of samples, calculate the cosine similarity between the feature vector representations:

[0045] ;

[0046] in, Eigenvector representation and Cosine similarity between them Indicates sample The eigenvector representation, Indicates sample The eigenvector representation, Eigenvector representation L2 norm, Eigenvector representation L2 norm, and Indicates a sample;

[0047] S364. Define a contrastive loss function. During training, the learned features are continuously optimized by minimizing the contrastive loss function. Different loss functions are used for optimization at different training stages.

[0048] ;

[0049] in, This represents the contrastive loss function, where N represents the total number of samples. Indicates a label, when and When the sample is from the same class, the label value is 1; otherwise, the label value is 0. Indicates a positive sample. Indicates a negative sample. Indicates sample eigenvector representation and positive samples eigenvector representation Cosine similarity between them Indicates sample eigenvector representation and negative samples eigenvector representation The cosine similarity between the two values, where max represents the maximum value between them. Indicates sample The eigenvector representation, Indicates positive samples The eigenvector representation, Indicates negative samples The feature vector representation of , where p represents the training phase;

[0050] S365, By minimizing the contrastive loss function The Adam optimizer optimizes the learned features by reducing the distance between positive samples and increasing the distance between negative samples.

[0051] Optionally, S5 specifically includes:

[0052] S51. The generated comprehensive identity authentication feature vector is compared with a preset template stored in the database for authentication. The cosine similarity between the feature vector of the user to be authenticated and the template is calculated using weighted cosine similarity. The preset template contains trained user features, and each user feature is stored in vector form in the preset template.

[0053] ;

[0054] in, This represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template. This represents the comprehensive identity authentication feature vector of the user to be authenticated. This represents a preset template feature vector stored in the database. and These represent the Euclidean norms of the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template, respectively.

[0055] S52. When calculating similarity, the matching weights of the feature vectors corresponding to each video stream are dynamically adjusted using a weighted average method based on the quality score of each video stream:

[0056] ;

[0057] in, This represents the cosine similarity between the weighted average feature vector of the user to be authenticated and the feature vector of the preset template, where M represents the total number of video streams. Let cosine similarity represent the feature vector of the i-th video stream and the feature vector of the preset template. Let represent the comprehensive identity authentication feature vector of the user to be authenticated in the i-th video stream, and represent . The preset template feature vector of the i-th video stream stored in the database. The weight of the i-th video stream is:

[0058] ;

[0059] in, This represents the quality score of the i-th video stream. This represents the feature reliability score of the i-th video stream;

[0060] S53. After calculating the weighted cosine similarity, compare the obtained weighted cosine similarity with the set authentication threshold. If the calculated weighted cosine similarity is greater than or equal to the preset authentication threshold, the user identity is considered to match and authentication is successful; otherwise, authentication fails.

[0061] S54. If authentication fails, security protection measures will be triggered in the system. The system will start abnormal behavior detection and record relevant logs.

[0062] The beneficial effects of this invention are:

[0063] First, by introducing multiple video streams, this invention effectively alleviates the limitations of a single perspective by utilizing video data from different viewpoints, ensuring the robustness of the authentication process in complex or undesirable environments.

[0064] Secondly, an improved self-supervised dynamic graph neural network is used to simultaneously extract facial, behavioral, and speech features from videos, and a dynamic graph structure is constructed to represent the spatiotemporal dependencies in the video data. The improvement lies in the introduction of dynamic edge weights and an optimized self-supervised learning mechanism, which enhances the consistency and correlation among multimodal features, thereby more accurately extracting and fusing features from various modalities. This enables the system to effectively adapt to complex authentication scenarios, improving not only authentication accuracy but also anti-counterfeiting capabilities.

[0065] Finally, during the authentication process, the system uses a dynamic weighting algorithm to adjust the weight of each video stream in real time based on the quality score of the video stream and the reliability of the authentication features. The final authentication result is calculated using weighted cosine similarity, which enhances the authentication accuracy under video streams of different quality.

[0066] In summary, this invention optimizes the extraction and fusion process of multimodal features, improves the accuracy and security of identity authentication through multiple video streams and a self-supervised learning mechanism, avoids the weaknesses of traditional methods that are susceptible to attacks and forgery, and provides an efficient, secure, and robust remote identity authentication solution that is suitable for various practical application scenarios, such as financial payment, smart access control, and security monitoring. Attached Figure Description

[0067] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0068] Figure 1 This is a flowchart of a remote identity authentication method based on multiple video recognition proposed in this invention;

[0069] Figure 2 This is a structural diagram of an improved self-supervised dynamic graph neural network for a remote identity authentication method based on multiple video recognition proposed in this invention. Detailed Implementation

[0070] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0071] refer to Figure 1 and Figure 2 A remote identity authentication method based on multi-video recognition includes the following steps:

[0072] S1. Collect real-time video stream data of the target user from different angles using a camera;

[0073] S2. Preprocess the acquired video stream data to remove noise, enhance the image, segment and calibrate the video frames, locate key feature regions in the video, and generate preprocessed data.

[0074] S3. Use an improved self-supervised dynamic graph neural network to extract facial features, behavioral features and speech features from the preprocessed data. The improved self-supervised dynamic graph neural network constructs a dynamic graph structure to represent the spatiotemporal dependencies in the preprocessed data. Combined with a self-supervised learning mechanism, it performs unsupervised training through a comparative learning task to maximize the consistency between modal features, optimize the correlation between modal features, and extract facial features, behavioral features and speech features to generate corresponding feature vector representations.

[0075] S4. The feature vector representations of facial features, behavioral features, and voice features are fused to generate a comprehensive identity authentication feature vector.

[0076] S5. Compare and verify the generated comprehensive identity authentication feature vector, authenticate the user features with the preset template stored in the database, assign weights to the authentication results of each video stream based on the authentication results of multiple video streams using a weighted method based on dynamic learning, adjust the weights in real time according to the video stream quality and the reliability score of the authentication features in the video stream, and calculate the final authentication result using a weighted average method.

[0077] S6. During the authentication process, monitor the environmental changes of the video stream in real time, adjust the recognition strategy, and update the parameters of the improved self-supervised dynamic graph neural network in real time.

[0078] S7. Output the final authentication result. Based on the final authentication result, determine whether to authorize access or perform subsequent operations. If the authentication fails, take corresponding security protection measures.

[0079] In this embodiment, S3 specifically includes:

[0080] S31. An improved self-supervised dynamic graph neural network is used to extract multimodal features from the preprocessed data, and a dynamic graph structure is constructed to represent the spatiotemporal dependencies in the preprocessed data. The dynamic graph structure changes dynamically over time. The multimodal features include facial features, behavioral features, and speech features. The improvements of the improved self-supervised dynamic graph neural network include introducing dynamic edge weights in the graph convolution operation and an improved self-supervised learning contrastive loss function. The contrastive loss function is optimized using different loss functions at different training stages.

[0081] ;

[0082] Where G(t) represents the dynamic graph structure, V(t) represents the node set of the dynamic graph, which is the feature in the video frame, E(t) represents the edge set of the dynamic graph, which is the spatiotemporal dependency between video frames, and X(t) represents the node feature matrix.

[0083] S32. Introduce dynamic edge weights and update the node representations through graph convolution operations:

[0084] ;

[0085] in, Indicates that node v is at the th The feature vector representation of the layer, where v represents a node in the dynamic graph. Let N(v) represent the non-linear activation function, and let N(v) represent the set of neighboring nodes of node v. This represents the degree of node k. This represents the weight matrix of the graph convolutional layer. Indicates that node v is at the th The feature vector representation of the layer, where k represents a node in the dynamic graph. This represents the dynamic edge weight between node v and its neighbor node k at time t:

[0086] ;

[0087] Where exp represents the natural exponential function, and N(v) represents the set of neighboring nodes of node v. Represents nodes in a dynamic graph. This represents the feature vector representation of node v. This represents the feature vector representation of node k. Represents a node The feature vector representation, k represents a node in the dynamic graph, and Sim represents the cosine similarity between the feature vector representations. This represents the hyperparameter that is being adjusted. The cosine similarity between the feature vector representations of node k and node v is denoted as . Represents a node The cosine similarity between the feature vector representations of node v and node v;

[0088] S33. Extract facial features, use node representations in a self-supervised dynamic graph neural network to capture facial key point information, update facial features through graph convolution operations, and dynamically adjust based on the similarity between facial nodes to generate feature vector representations of facial features:

[0089] ;

[0090] in, This indicates that the facial feature node v is at the 1st... The feature vector representation of the layer, This represents a non-linear activation function, where N(face) represents the set of neighboring nodes of the facial feature node. This represents the degree of node u. This represents the dynamic edge weights between face nodes at time t. Let represent the feature vector representation of node u at layer l, where u represents a node in the dynamic graph. The graph convolution weight matrix representing facial features. Indicates the bias term;

[0091] S34. Extract behavioral features and use time node representations in a self-supervised dynamic graph neural network to process motion information in the video. Process the dynamic behavioral data at each time node through graph convolutional layers, combine spatiotemporal dependency modeling, update the behavioral features, and generate feature vector representations of the behavioral features:

[0092] ;

[0093] in, The behavioral feature time node t represents the time node at the tth time node. The feature vector representation of the layer, This represents a non-linear activation function, where N(action) represents the set of neighboring nodes of the behavioral feature node. Represents a node The degree, This represents the dynamic edge weights between behavioral feature time nodes t. Representing behavioral characteristics at specific time points In the The feature vector representation of the layer, The time points that represent the behavioral characteristics of a dynamic graph. The graph convolution weight matrix represents behavioral features. Indicates the bias term;

[0094] S35. Extract speech features through graph convolution operations to generate feature vector representations of the speech features:

[0095] ;

[0096] in, The speech feature node v represents the first... Layer feature vector representation, Let N(speech) represent a non-linear activation function, and let N(speech) represent the set of neighboring nodes of the speech feature node. Represents a node The degree, Represents the dynamic edge weights between speech features. Represents speech feature nodes In the The feature vector representation of the layer, Represents the speech feature nodes of a dynamic graph. The graph convolution weight matrix representing speech features. Indicates the bias term;

[0097] S36. The modal feature consistency of facial features, behavioral features and speech features is optimized through a contrastive loss function of self-supervised learning. The contrastive loss function minimizes the distance between features of the same class and maximizes the distance between features of different classes.

[0098] S37. The feature vector representation is further optimized through the reconstruction loss function obtained by self-supervised learning, and the parameters of the dynamic graph neural network are updated through the backpropagation algorithm. The reconstruction loss function is adjusted by comparing the difference between the input features and the reconstructed features to minimize the reconstruction error.

[0099] ;

[0100] in, Let X represent the reconstruction loss function, and let X represent the original input feature matrix. This represents the feature matrix obtained from the reconstruction by the dynamic graph neural network. Let Frobenius norm be the matrix X and ? The differences between them.

[0101] In this embodiment, S36 specifically includes:

[0102] S361. Define a contrastive learning task: by minimizing the distance between samples of the same class and maximizing the distance between samples of different classes, capture the similarity and difference between modal features of different video streams. The modal features include facial features, behavioral features, and speech features.

[0103] S362. Select positive and negative samples from the facial features, behavioral features and voice features extracted from the dynamic graph neural network. The positive samples represent samples belonging to the same category, and the negative samples represent samples with different features and belonging to different categories.

[0104] S363. For each pair of samples, calculate the cosine similarity between the feature vector representations:

[0105] ;

[0106] in, Eigenvector representation and Cosine similarity between them Indicates sample The eigenvector representation, Indicates sample The eigenvector representation, Eigenvector representation L2 norm, Eigenvector representation L2 norm, and Indicates a sample;

[0107] S364. Define a contrastive loss function. During training, the learned features are continuously optimized by minimizing the contrastive loss function. Different loss functions are used for optimization at different training stages.

[0108] ;

[0109] in, This represents the contrastive loss function, where N represents the total number of samples. Indicates a label, when and When the sample is from the same class, the label value is 1; otherwise, the label value is 0. Indicates a positive sample. Indicates a negative sample. Indicates sample eigenvector representation and positive samples eigenvector representation Cosine similarity between them Indicates sample eigenvector representation and negative samples eigenvector representation The cosine similarity between the two values, where max represents the maximum value between them. Indicates sample The eigenvector representation, Indicates positive samples The eigenvector representation, Indicates negative samples The feature vector representation of , where p represents the training phase;

[0110] S365, By minimizing the contrastive loss function The Adam optimizer optimizes the learned features by reducing the distance between positive samples and increasing the distance between negative samples.

[0111] In this embodiment, S5 specifically includes:

[0112] S51. The generated comprehensive identity authentication feature vector is compared with a preset template stored in the database for authentication. The cosine similarity between the feature vector of the user to be authenticated and the template is calculated using weighted cosine similarity. The preset template contains trained user features, and each user feature is stored in vector form in the preset template.

[0113] ;

[0114] in, This represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template. This represents the comprehensive identity authentication feature vector of the user to be authenticated. This represents a preset template feature vector stored in the database. and These represent the Euclidean norms of the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template, respectively.

[0115] S52. When calculating similarity, the matching weights of the feature vectors corresponding to each video stream are dynamically adjusted using a weighted average method based on the quality score of each video stream:

[0116] ;

[0117] in, This represents the cosine similarity between the weighted average feature vector of the user to be authenticated and the feature vector of the preset template, where M represents the total number of video streams. Let cosine similarity represent the feature vector of the i-th video stream and the feature vector of the preset template. Let represent the comprehensive identity authentication feature vector of the user to be authenticated in the i-th video stream, and represent . The preset template feature vector of the i-th video stream stored in the database. The weight of the i-th video stream is:

[0118] ;

[0119] in, This represents the quality score of the i-th video stream. This represents the feature reliability score of the i-th video stream;

[0120] S53. After calculating the weighted cosine similarity, compare the obtained weighted cosine similarity with the set authentication threshold. If the calculated weighted cosine similarity is greater than or equal to the preset authentication threshold, the user identity is considered to match and authentication is successful; otherwise, authentication fails.

[0121] S54. If authentication fails, security protection measures will be triggered in the system. The system will start abnormal behavior detection and record relevant logs.

[0122] Example 1:

[0123] To verify the feasibility of this invention in practice, it was applied to the remote identity authentication system of a financial payment platform. This platform primarily provides online payments, transfers, and other financial services. Users need to authenticate their identities to ensure transaction security when making payments. Traditional identity authentication methods, such as passwords and SMS verification codes, are largely unable to meet users' demands for high security and convenience. Therefore, this invention proposes a remote identity authentication method based on multi-video recognition, aiming to improve the accuracy and anti-counterfeiting capabilities of identity authentication and solve the problems of traditional methods.

[0124] In this scenario, when a user makes a payment, facial recognition is performed using the camera of their mobile phone or computer. The system collects real-time video stream data from different angles. According to the method of this invention, the system uses video streams from different perspectives captured by multiple cameras to extract the user's facial features, and simultaneously combines behavioral features (e.g., the user's gestures or movements) and voice features for identity authentication. Through an improved self-supervised dynamic graph neural network, the system can automatically extract and fuse facial, behavioral, and voice features from the video stream to form a comprehensive identity authentication feature vector, which is then compared with a preset template stored in the database to confirm the user's identity.

[0125] During implementation, the system adjusts the weight of each video stream using a dynamic learning-based weighting method based on the authentication results of multiple video streams, thereby deriving the final authentication result. To verify the feasibility of this method in complex environments, multiple authentication experiments were conducted using test data under different lighting conditions, angles, and backgrounds. Test results show that this invention has significant improvements over traditional single-video-stream authentication methods, especially exhibiting higher accuracy and robustness in low-light environments, under angle deviations, and with complex backgrounds.

[0126] In the experiment, 1000 different users were selected for identity authentication testing. Each user's identity characteristics included facial features, behavioral features, and voice features. The system performed authentication using data collected from multiple video streams, verifying its accuracy under varying environmental conditions.

[0127] Table 1. Comparison of Experimental Data

[0128] ;

[0129] Based on the comparison table of the experimental data above, it can be seen that the multi-stream video authentication method adopted in this invention is superior to the traditional single-stream video authentication method under various environmental conditions, especially in low-light and large-angle deviation scenarios, where it shows obvious advantages.

[0130] First, under well-lit conditions, the authentication pass rate of multi-stream video authentication is generally higher than that of traditional methods. For example, under well-lit conditions, User 1's authentication pass rate is 99.8%, while the traditional method only achieves 96.5%. Other users also showed high authentication accuracy under the same conditions, with the overall authentication pass rate significantly better than that of traditional methods. This indicates that utilizing authentication data from multiple video streams not only enhances the robustness of recognition but also effectively reduces the impact of lighting changes on the recognition results.

[0131] Secondly, in poor lighting conditions, traditional single-stream video authentication methods exhibit low pass rates, especially in low-light environments where accuracy drops significantly. For example, in low-light conditions, the pass rate for user 5 using traditional methods is only 68.5%, while the pass rate using multi-stream video authentication is 92.4%, significantly improving accuracy. This demonstrates that in low-light environments, multi-stream video authentication can compensate for the shortcomings of single-stream authentication by fusing data and information from multiple angles, reducing the impact of uneven lighting.

[0132] Furthermore, the advantages of the multi-stream authentication method become even more pronounced in situations involving angular deviations and background complexity. Experimental data from Users 3 and 8 clearly demonstrate this. Under conditions of significant angular variation (User 3), the traditional method achieved an authentication pass rate of only 80.3%, while the multi-stream method increased this to 98.5%. However, under conditions of complex backgrounds and low lighting (User 8), the traditional method achieved a pass rate of only 60.2%, while the multi-stream method achieved 89.3%. This significant difference indicates that multi-stream authentication, through the fusion of multi-angle and multi-modal information, effectively mitigates issues such as occlusion, angular deviations, and background noise in the video stream, thereby ensuring the stability of the authentication process.

[0133] In summary, the experimental results show that the multi-video stream authentication method of this invention outperforms traditional methods under various environmental conditions, especially in complex environments such as low light, angular deviations, and complex backgrounds, where it demonstrates stronger robustness and accuracy. By introducing multiple video streams and a dynamic weighting mechanism, this invention not only effectively improves the accuracy and anti-counterfeiting capabilities of identity authentication but also optimizes the overall performance of existing identity authentication systems.

[0134] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A remote identity authentication method based on multiple video recognition, characterized in that, Includes the following steps: S1. Collect real-time video stream data of the target user from different angles using a camera; S2. Preprocess the acquired video stream data to remove noise, enhance the image, segment and calibrate the video frames, locate key feature regions in the video, and generate preprocessed data. S3. Use an improved self-supervised dynamic graph neural network to extract facial features, behavioral features and speech features from the preprocessed data. The improved self-supervised dynamic graph neural network constructs a dynamic graph structure to represent the spatiotemporal dependencies in the preprocessed data. Combined with a self-supervised learning mechanism, it performs unsupervised training through a comparative learning task to maximize the consistency between modal features, optimize the correlation between modal features, and extract facial features, behavioral features and speech features to generate corresponding feature vector representations. S4. The feature vector representations of facial features, behavioral features, and voice features are fused to generate a comprehensive identity authentication feature vector. S5. Compare and verify the generated comprehensive identity authentication feature vector, authenticate the user features with the preset template stored in the database, and use a weighted method based on dynamic learning to assign weights to the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template. Adjust the weights in real time according to the video stream quality and the reliability score of the authentication features in the video stream, and use the weighted average method to calculate the final authentication result. S6. During the authentication process, monitor the environmental changes of the video stream in real time, adjust the recognition strategy, and update the parameters of the improved self-supervised dynamic graph neural network in real time. S7. Output the final authentication result. Determine whether to authorize access or perform subsequent operations based on the final authentication result. If authentication fails, take corresponding security protection measures. S3 specifically includes: S31. An improved self-supervised dynamic graph neural network is used to extract multimodal features from the preprocessed data, and a dynamic graph structure is constructed to represent the spatiotemporal dependencies in the preprocessed data. The dynamic graph structure changes dynamically over time. The multimodal features include facial features, behavioral features, and speech features. The improvements of the improved self-supervised dynamic graph neural network include introducing dynamic edge weights in the graph convolution operation and an improved self-supervised learning contrastive loss function. The contrastive loss function is optimized using different loss functions at different training stages. ; in, Represents a dynamic graph structure. This represents the set of nodes in a dynamic graph, which are features within a video frame. The edge set represents the spatiotemporal dependencies between video frames in a dynamic graph. Represents the node feature matrix; S32. Introduce dynamic edge weights and update the node representations through graph convolution operations: ; in, Represents a node In the The feature vector representation of the layer, Represents nodes in a dynamic graph. Represents a non-linear activation function. Represents a node The set of neighboring nodes, Represents a node The degree, This represents the weight matrix of the graph convolutional layer. Represents a node In the The feature vector representation of the layer, Represents nodes in a dynamic graph. Indicates time Time node with neighboring nodes Dynamic edge weights between them: ; in, This represents the natural exponential function. Represents a node The set of neighboring nodes, Represents nodes in a dynamic graph. Represents a node The eigenvector representation, Represents a node The eigenvector representation, Represents a node The eigenvector representation, Represents nodes in a dynamic graph. The cosine similarity between feature vectors is represented. This represents the hyperparameter that is being adjusted. Represents a node and nodes The feature vectors represent the cosine similarity between them. Represents a node and nodes The feature vectors represent the cosine similarity between them; S33. Extract facial features, use node representations in a self-supervised dynamic graph neural network to capture facial key point information, update facial features through graph convolution operations, and dynamically adjust based on the similarity between facial nodes to generate feature vector representations of facial features: ; in, Represents facial feature nodes In the The feature vector representation of the layer, Represents a non-linear activation function. Represents the set of neighbor nodes of a facial feature node. Represents a node The degree, Indicates time Dynamic edge weights between facial nodes Represents a node In the The feature vector representation of the layer, Represents nodes in a dynamic graph. The graph convolution weight matrix representing facial features. Indicates the bias term; S34. Extract behavioral features and use time node representations in a self-supervised dynamic graph neural network to process motion information in the video. Process the dynamic behavioral data at each time node through graph convolutional layers, combine spatiotemporal dependency modeling, update the behavioral features, and generate feature vector representations of the behavioral features: ; in, Representing behavioral characteristics at specific time points In the The feature vector representation of the layer, Represents a non-linear activation function. This represents the set of neighboring nodes of a node exhibiting behavioral characteristics. Represents a node The degree, Representing behavioral characteristics at specific time points Dynamic edge weights between them Representing behavioral characteristics at specific time points In the The feature vector representation of the layer, The time points that represent the behavioral characteristics of a dynamic graph. The graph convolution weight matrix represents behavioral features. Indicates the bias term; S35. Extract speech features through graph convolution operations to generate feature vector representations of the speech features: ; in, Represents speech feature nodes The Layer feature vector representation, Represents a non-linear activation function. Represents the set of neighbor nodes of a speech feature node. Represents a node The degree, Represents the dynamic edge weights between speech features. Represents speech feature nodes In the The feature vector representation of the layer, Represents the speech feature nodes of a dynamic graph. The graph convolution weight matrix representing speech features. Indicates the bias term; S36. The modal feature consistency of facial features, behavioral features and speech features is optimized through a contrastive loss function learned by self-supervised learning. The contrastive loss function minimizes the distance between features of the same class and maximizes the distance between features of different classes. S37. The feature vector representation is further optimized through the reconstruction loss function obtained by self-supervised learning, and the parameters of the dynamic graph neural network are updated through the backpropagation algorithm. The reconstruction loss function is adjusted by comparing the difference between the input features and the reconstructed features to minimize the reconstruction error. ; in, Represents the reconstruction loss function. Represents the original input feature matrix. This represents the feature matrix obtained from the reconstruction by the dynamic graph neural network. Let Frobenius norm be a matrix. and The differences between them.

2. The remote identity authentication method based on multiple video recognition according to claim 1, characterized in that, Specifically, S36 includes: S361. Define a contrastive learning task: by minimizing the distance between samples of the same class and maximizing the distance between samples of different classes, capture the similarity and difference between modal features of different video streams. The modal features include facial features, behavioral features, and speech features. S362. Select positive and negative samples from the facial features, behavioral features and voice features extracted from the dynamic graph neural network. The positive samples represent samples belonging to the same category, and the negative samples represent samples with different features and belonging to different categories. S363. For each pair of samples, calculate the cosine similarity between the feature vector representations: ; in, Eigenvector representation and Cosine similarity between them Indicates sample The eigenvector representation, Indicates sample The eigenvector representation, Eigenvector representation L2 norm, Eigenvector representation L2 norm, and Indicates a sample; S364. Define a contrastive loss function. During training, the learned features are continuously optimized by minimizing the contrastive loss function. Different loss functions are used for optimization at different training stages. ; in, This represents the contrastive loss function. Represents the total number of samples. Indicates a label, when and When the sample is from the same class, the label value is 1; otherwise, the label value is 0. Indicates a positive sample. Indicates a negative sample. Indicates sample eigenvector representation and positive samples eigenvector representation Cosine similarity between them Indicates sample eigenvector representation and negative samples eigenvector representation Cosine similarity between them This means taking the maximum value between the two. Indicates sample The eigenvector representation, Indicates positive samples The eigenvector representation, Indicates negative samples The eigenvector representation, Indicates the training phase; S365, By minimizing the contrastive loss function The Adam optimizer optimizes the learned features by reducing the distance between positive samples and increasing the distance between negative samples.

3. The remote identity authentication method based on multiple video recognition according to claim 1, characterized in that, S5 specifically includes: S51. The generated comprehensive identity authentication feature vector is compared with a preset template stored in the database for authentication. The cosine similarity between the feature vector of the user to be authenticated and the template is calculated using weighted cosine similarity. The preset template contains trained user features, and each user feature is stored in vector form in the preset template. ; in, This represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template. This represents the comprehensive identity authentication feature vector of the user to be authenticated. This represents a preset template feature vector stored in the database. and These represent the Euclidean norms of the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template, respectively. S52. When calculating similarity, the matching weights of the feature vectors corresponding to each video stream are dynamically adjusted using a weighted average method based on the quality score of each video stream: ; in, This represents the cosine similarity between the weighted average feature vector of the user to be authenticated and the feature vector of the preset template. Indicates the total number of video streams. Indicates the first Cosine similarity between the feature vectors of each video stream and the feature vectors of a preset template. Indicates the first The comprehensive identity authentication feature vector of the user to be authenticated in each video stream represents No. Each video stream is stored in a database as a preset template feature vector. Indicates the first Weights of each video stream: ; in, Indicates the first The quality score of each video stream Indicates the first A feature reliability score for each video stream; S53. After calculating the weighted cosine similarity, compare the obtained weighted cosine similarity with the set authentication threshold. If the calculated weighted cosine similarity is greater than or equal to the preset authentication threshold, the user identity is considered to match and authentication is successful; otherwise, authentication fails. S54. If authentication fails, security protection measures will be triggered in the system. The system will start abnormal behavior detection and record relevant logs.

Citation Information

Patent Citations

  • Video network asset identity credibility identification method and system

    CN118821029A

  • Intelligent access control management method and system based on multi-mode identification and Internet of Things technology

    CN118968665A