Remote identity authentication method based on multiple video recognition

By adopting multiple video recognition and an improved self-supervised dynamic graph neural network in identity authentication, extracting and fusing facial, behavior and speech features, the problem of reduced recognition accuracy and insufficient robustness of a single video stream facial recognition technology in complex environments is solved, and a more efficient, secure and robust identity authentication effect is achieved.

CN120223444AActive Publication Date: 2025-06-27XIAOYUAN PERCEPTION (HULUDAO) TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510695537.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing identity authentication methods have problems such as degradation of recognition accuracy and insufficient robustness in complex environments and diverse application scenarios, especially the single video stream facial recognition technology is difficult to cope with the impact of lighting changes, angle deviations and video quality.

Method used

The remote identity authentication method based on multiple video recognition is adopted, and the improved self-supervised dynamic graph neural network combines multiple video streams to extract facial, behavior and speech multimodal features, construct dynamic graph structures to represent space-time dependencies, and use dynamic edge weights and self-supervised learning mechanisms to enhance the consistency and correlation of modal features.

Benefits of technology

Improve the accuracy and robustness of identity authentication, enhance adaptability in complex environments, reduce the impact of lighting changes, angle deviations and background noise on the identification results, and provide a more efficient, secure and robust identity authentication solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223444A_ABST
    Figure CN120223444A_ABST
Patent Text Reader

Abstract

The invention discloses a remote identity authentication method based on multiple video recognition. The method comprises the following steps: S1, collecting real-time video stream data of a target user; s2, preprocessing the collected video stream data to generate preprocessed data; s3, extracting features in the preprocessed data by using an improved self-supervised dynamic graph neural network and generating corresponding feature vector representation; s4, fusing the feature vector representations; s5, comparing and verifying the fused feature vectors, and weighting the authentication result of each video stream according to the authentication results of the plurality of video streams to generate a final authentication result; s6, monitoring the environment change of the video stream in real time, and updating network parameters in real time; and S7, judging whether to authorize access or carry out subsequent operation according to the final authentication result. According to the method, multiple video streams and the improved self-supervised dynamic graph neural network are combined, multi-modal features are extracted and fused, and the accuracy, safety and robustness of remote identity authentication are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of identity authentication, and particularly to a remote identity authentication method based on multi-video recognition. Background Art

[0002] With the popularization of the Internet and mobile devices, remote identity authentication has gradually become an important means to ensure information security and privacy protection. Traditional identity authentication methods, such as those based on passwords, SMS verification codes, fingerprint recognition, etc., although they have improved security to a certain extent, still have many limitations. Especially in scenarios where users face high security requirements or need non-contact authentication, traditional authentication methods are vulnerable to attacks or forgery threats. At the same time, with the development of biometric recognition technology, biometric recognition technologies such as face recognition, fingerprint recognition, iris recognition, etc. have gradually become important means in the field of identity authentication. As a non-contact identity verification method, face recognition technology has been widely promoted in practical applications due to its convenience and high security. However, existing face recognition technologies still face some technical challenges, especially in complex environments and diverse application scenarios.

[0003] Currently, face recognition technology based on a single video stream usually relies on image or video information from a single perspective for identity verification. Although this method performs well in standard environments, in practical applications, face recognition technology is often affected by factors such as lighting changes, angle deviations, occlusions, and video quality, resulting in a decrease in recognition accuracy. For example, the face in the video may change due to the user's head rotation, or in low-light conditions, facial features cannot be clearly recognized. At this time, the face recognition method based on a single video stream may not be able to effectively handle these complex situations, leading to authentication failures. Therefore, existing face recognition technology based on a single video stream has poor adaptability in complex scenarios and cannot provide sufficient robustness.

[0004] In addition to face recognition, behavior recognition and voice recognition, as other forms of biometric recognition technology, have also been widely used. Behavior recognition technology can authenticate a user by analyzing the user's actions, postures, or gestures, while voice recognition authenticates a user by analyzing the user's voice characteristics. Although these technologies have improved the security of identity authentication to a certain extent, due to their inherent limitations, they often cannot work independently with face recognition technology in practical applications. Behavior recognition and voice recognition are usually easily affected by environmental interference, such as background noise, the number of people in the video, the fuzziness of the voice, etc., which may lead to instability and inaccuracy of the recognition results.

[0005] Existing identity authentication systems usually rely on a single biometric for authentication, so their security and accuracy cannot reach an ideal level. For example, standalone face recognition may be easily deceived by photos or videos, voice recognition may be affected by differences in audio quality or noise, and behavior recognition may also be affected by background interference or changes in the user's location. To make up for the deficiencies of single biometrics, in recent years, multi-modal biometric recognition methods have gradually become a research hotspot. These methods improve the security and robustness of identity authentication by combining multiple biometrics such as face, voice, and behavior. However, the challenge faced by multi-modal recognition methods lies in how to efficiently fuse information from different modalities to give full play to the advantages of each modality.

[0006] Existing multi-modal recognition methods usually use separate models to extract each biometric and perform feature fusion subsequently. This approach has certain drawbacks. First, since different processing models are used for each modality, the feature extraction processes of each modality are carried out independently, and the inherent relationships between modalities cannot be fully utilized. Second, most existing multi-modal fusion methods rely on simple methods such as feature concatenation and weighted averaging, lacking in-depth exploration of the dynamic correlations between modalities. These methods often ignore the complex dependencies between modality features, resulting in limited fusion effects and ultimately affecting the accuracy of authentication.

[0007] With the development of graph neural networks, graph-based structured learning methods have gradually been introduced into multi-modal recognition tasks. Graph neural networks have good spatio-temporal modeling capabilities and can effectively propagate information in a graph structure to capture the complex relationships between nodes. In multi-modal biometric recognition, graph neural networks can build a graph structure of multi-modal features, propagate information between different nodes in the graph, and fuse information from different modalities. In this way, graph neural networks can not only retain spatio-temporal dependencies but also achieve deep fusion of multi-modal features, further improving the accuracy and robustness of identity authentication.

[0008] However, there are still some deficiencies in the application of existing graph neural networks in multi-modal recognition. Especially when dealing with large-scale video data, existing graph neural network models often have a high computational complexity, a cumbersome training process, and insufficient adaptability to environmental changes. Therefore, how to optimize the structure of graph neural networks to better adapt to the multi-modal features and spatio-temporal dependencies of video data remains an urgent problem to be solved.

[0009] In summary, there are still many challenges in the existing identity authentication methods in terms of multi-modal recognition and video data processing. The deficiencies of the existing technologies are mainly reflected in the limitations of single video stream recognition, the inadequacies of multi-modal feature fusion methods, and the lack of spatio-temporal modeling in traditional deep learning models. Therefore, a new method is needed that can effectively fuse data from different video streams and adaptively adjust the recognition strategy in different environments, thereby improving the accuracy, robustness, and security of identity authentication. Summary of the Invention

[0010] An object of the present invention is to propose a remote identity authentication method based on multi-video recognition. The present invention combines multi-video streams and an improved self-supervised dynamic graph neural network. By constructing a dynamic graph structure and extracting multi-modal features of face, behavior, and voice, it effectively processes spatio-temporal dependencies, improves the feature fusion effect, and uses dynamic edge weights and self-supervised learning mechanisms to enhance the consistency and relevance between different modalities, further improving the authentication accuracy.

[0011] A remote identity authentication method based on multi-video recognition according to an embodiment of the present invention includes the following steps: S1. Collect real-time video stream data of a target user from different angles through a camera; S2. Preprocess the collected video stream data, remove noise, perform image enhancement, segment and calibrate video frames, locate key feature regions in the video, and generate preprocessed data; S3. Use an improved self-supervised dynamic graph neural network to extract face features, behavior features, and voice features from the preprocessed data. The improved self-supervised dynamic graph neural network represents the spatio-temporal dependencies in the preprocessed data by constructing a dynamic graph structure, combines a self-supervised learning mechanism, conducts unsupervised training through a contrast learning task, maximizes the consistency between modal features, optimizes the relevance between modal features, and extracts and generates corresponding feature vector representations for face features, behavior features, and voice features; S4. Fuse the feature vector representations of face features, behavior features, and voice features to generate a comprehensive identity authentication feature vector; S5. Compare and verify the generated comprehensive identity authentication feature vector, authenticate the user features with a preset template stored in the database, and use a weighted method based on dynamic learning to weight the authentication results of each video stream according to the authentication results of multiple video streams, adjust the weights in real time according to the video stream quality and the reliability score of the authentication features in the video stream, and calculate the final authentication result using the weighted average method; S6. During the authentication process, monitor the environmental changes of the video stream in real time, adjust the recognition strategy, and update the parameters of the improved self-supervised dynamic graph neural network in real time; S7. Output the final authentication result, and determine whether to authorize access or perform subsequent operations based on the final authentication result. If the authentication fails, take corresponding security protection measures.

[0012] Optionally, the S3 specifically includes: S31. Use improved methods to extract multi-modal features from the preprocessed data, construct a dynamic graph structure to represent the spatio-temporal dependencies in the preprocessed data. The dynamic graph structure changes dynamically over time. The multi-modal features include facial features, behavioral features, and speech features. In the improved graph convolution operation, dynamic edge weights and a contrast loss function for improved self-supervised learning are introduced. The contrast loss function is optimized using different loss functions at different training stages: ; Among them, G(t) represents the dynamic graph structure, V(t) represents the node set of the dynamic graph, is the feature in the video frame, E(t) represents the edge set of the dynamic graph, which is the spatio-temporal dependency between video frames, and X(t) represents the node feature matrix; S32. Introduce dynamic edge weights and update the node representations through graph convolution operations: ; Among them, represents the feature vector representation of node v at the layer, v represents the node of the dynamic graph, represents the non-linear activation function, N(v) represents the neighbor node set of node v, represents the degree of node k, represents the weight matrix of the graph convolution layer, represents the feature vector representation of node v at the layer, k represents the node of the dynamic graph, represents the dynamic edge weight between node v and neighbor node k at time t: ; Among them, exp represents the natural exponential function, N(v) represents the neighbor node set of node v, represents the node of the dynamic graph, represents the feature vector representation of node v, represents the feature vector representation of node k, represents the feature vector representation of node , k represents the node of the dynamic graph, Sim represents the cosine similarity between feature vector representations, represents the adjusted hyperparameter, represents the cosine similarity between the feature vector representations of node k and node v, represents node The cosine similarity between the eigenvector representation of node v; S33. Extract facial features, capture facial key point information using the node representation in the self-supervised dynamic graph neural network, update the facial features through graph convolution operations, and dynamically adjust in combination with the similarity between facial nodes to generate the eigenvector representation of the facial features: ; Among them, represents the eigenvector representation of the facial feature node v at the layer, represents the non-linear activation function, N(face) represents the set of neighbor nodes of the facial feature node, represents the degree of node u, represents the dynamic edge weight between facial nodes at time t, represents the eigenvector representation of node u at layer l, u represents the node of the dynamic graph, represents the graph convolution weight matrix of the facial features, represents the bias term; S34. Extract behavioral features, process the motion information in the video using the time node representation in the self-supervised dynamic graph neural network, process the dynamic behavior data of each time node through the graph convolution layer, combine spatio-temporal dependency modeling, and update the behavioral features to generate the eigenvector representation of the behavioral features: ; Among them, represents the eigenvector representation of the behavioral feature time node t at the layer, represents the non-linear activation function, N(action) represents the set of neighbor nodes of the behavioral feature node, represents the degree of node , represents the dynamic edge weight between the behavioral feature time nodes t, represents the behavioral feature time node at the layer of eigenvector representation, represents the behavioral feature time node of the dynamic graph, represents the graph convolution weight matrix of the behavioral features, represents the bias term; S35. Extract speech features through graph convolution operations to generate the eigenvector representation of the speech features: ; Among them, represents the eigenvector representation of the speech feature node v at the layer, represents a non-linear activation function, and N(speech) represents the set of neighbor nodes of the speech feature nodes. represents a node degree of represents the dynamic edge weight between speech features represents the speech feature node at the layer's feature vector representation represents the speech feature nodes of the dynamic graph represents the graph convolution weight matrix of speech features represents the bias term; S36. Optimize the modal feature consistency of facial features, behavioral features, and speech features through a contrast loss function of self-supervised learning. The contrast loss function minimizes the distance between similar features and maximizes the distance between different features; S37. Further optimize the feature vector representation through a reconstruction loss function of self-supervised learning, and update the parameters of the dynamic graph neural network through the backpropagation algorithm. The reconstruction loss function is adjusted by comparing the difference between the input features and the reconstructed features to minimize the reconstruction error: ; where represents the reconstruction loss function, X represents the original input feature matrix, represents the feature matrix reconstructed by the dynamic graph neural network, represents the Frobenius norm, which is the difference between the matrix X and .

[0013] Optionally, the S36 specifically includes: S361. Define a contrast learning task to capture the similarity and difference between modal features of different video streams by minimizing the distance between similar samples and maximizing the distance between different samples. The modal features include facial features, behavioral features, and speech features; S362. Select positive samples and negative samples from the facial features, behavioral features, and speech features extracted by the dynamic graph neural network. The positive samples represent samples belonging to the same category, and the negative samples represent samples with different features and belonging to different categories; S363. For each pair of samples, calculate the cosine similarity between the feature vector representations: ; where represents the cosine similarity between the feature vector representations and and represents the sample 's feature vector representation. Represents a sample The eigenvector representation of Represents the eigenvector representation The L2 norm of Represents the eigenvector representation The L2 norm of And Represents a sample; S364. Define a contrast loss function, and continuously optimize the learned features by minimizing the contrast loss function during the training process. The contrast loss function is optimized using different loss functions at different training stages: ; Wherein, Represents the contrast loss function, N represents the total number of samples, Represents the label. When And Are samples of the same class, the value of the label is 1. If they are not samples of the same class, the value of the label is 0, Represents a positive sample, Represents a negative sample, Represents a sample The eigenvector representation of And the positive sample The eigenvector representation of The cosine similarity between them, Represents a sample The eigenvector representation of And the negative sample The eigenvector representation of The cosine similarity between them, max represents taking the maximum value between the two, Represents a sample The eigenvector representation of Represents a positive sample The eigenvector representation of Represents a negative sample The eigenvector representation of, p represents the training stage; S365. By minimizing the contrast loss function , reduce the distance between positive samples and increase the distance between negative samples, and optimize the learned features through the Adam optimizer.

[0014] Optionally, the S5 specifically includes: S51. Authenticate the generated comprehensive identity authentication eigenvector with a preset template stored in the database. Calculate the cosine similarity between the eigenvector of the user to be authenticated and the template using weighted cosine similarity. The preset template contains trained user features, and each user feature is stored in vector form in the preset template: ; in, Represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the preset template feature vector, Represents the comprehensive identity authentication feature vector of the user to be authenticated, represents the preset template feature vector stored in the database, and Respectively represent the Euclidean norm of the comprehensive identity authentication feature vector of the user to be authenticated and the feature vector of the preset template; S52. When performing similarity calculation, the matching weight of the feature vector corresponding to each video stream is dynamically adjusted by weighted average method according to the quality score of each video stream: ; in, represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the preset template feature vector after weighted averaging, M represents the total number of video streams, represents the cosine similarity between the feature vector of the i-th video stream and the feature vector of the preset template, Represents the comprehensive identity authentication feature vector of the user to be authenticated in the i-th video stream, The preset template feature vector stored in the database for the i-th video stream, Represents the weight of the i-th video stream: ; in, represents the quality score of the i-th video stream, represents the feature reliability score of the i-th video stream; S53, after calculating the weighted cosine similarity, the obtained weighted cosine similarity is compared with the set authentication threshold. If the calculated weighted cosine similarity is greater than or equal to the preset authentication threshold, it is considered that the user identity matches and the authentication is passed. Otherwise, the authentication fails. S54. If the authentication fails, security protection measures will be triggered in the system, the system will start abnormal behavior detection, and record relevant logs.

[0015] The beneficial effects of the present invention are: Firstly, the present invention introduces multiple video streams and utilizes video data from different perspectives, thereby effectively alleviating the limitation of a single perspective and ensuring the robustness of the authentication process in complex or undesirable environments.

[0016] Secondly, through an improved self-supervised dynamic graph neural network, facial, behavioral, and speech features in the video are extracted simultaneously, and the spatio-temporal dependencies in the video data are represented by constructing a dynamic graph structure. The improvement of this network lies in the introduction of dynamic edge weights and an optimized self-supervised learning mechanism, which can enhance the consistency and correlation between multi-modal features, thereby more accurately extracting and fusing features of each modality. This enables the system to effectively adapt to complex authentication scenarios, not only improving the accuracy of authentication but also increasing the anti-counterfeiting ability.

[0017] Finally, during the authentication process, the system adjusts the weight of each video stream in real time according to the quality score of the video stream and the reliability of the authentication features through a dynamic weighting algorithm, and calculates the final authentication result through weighted cosine similarity, enhancing the authentication accuracy under different quality video streams.

[0018] In summary, the present invention optimizes the extraction and fusion process of multi-modal features, improves the accuracy and security of identity authentication through multiple video streams and a self-supervised learning mechanism, avoids the weaknesses of traditional methods being vulnerable to attacks and forgeries, and provides an efficient, secure, and robust remote identity authentication solution applicable to various practical application scenarios such as financial payment, intelligent access control, and security monitoring. Brief Description of the Drawings

[0019] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of a remote identity authentication method based on multiple video identifications proposed by the present invention; Figure 2 is a structural diagram of an improved self-supervised dynamic graph neural network of a remote identity authentication method based on multiple video identifications proposed by the present invention. Detailed Embodiment

[0020] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0021] Refer to Figure 1 and Figure 2 , a remote identity authentication method based on multiple video identifications, includes the following steps: S1. Collect real-time video stream data of the target user from different angles through a camera; S2. Preprocess the collected video stream data, remove noise, perform image enhancement, segment and calibrate video frames, locate key feature regions in the video, and generate preprocessed data; S3. Use an improved self-supervised dynamic graph neural network to extract facial features, behavioral features, and speech features from the preprocessed data. The improved self-supervised dynamic graph neural network constructs a dynamic graph structure to represent the spatio-temporal dependencies in the preprocessed data, combines a self-supervised learning mechanism, and performs unsupervised training through a contrastive learning task to maximize the consistency between modal features, optimize the correlation between modal features, and extract and generate corresponding feature vector representations for facial features, behavioral features, and speech features; S4. Fuse the feature vector representations of facial features, behavioral features, and speech features to generate a comprehensive identity authentication feature vector; S5. Compare and verify the generated comprehensive identity authentication feature vector, authenticate the user features with the preset templates stored in the database, and according to the authentication results of multiple video streams, use a weighted method based on dynamic learning to weight the authentication results of each video stream, adjust the weights in real time according to the video stream quality and the reliability score of the authentication features in the video stream, and use the weighted average method to calculate the final authentication result; S6. During the authentication process, monitor the environmental changes of the video stream in real time, adjust the recognition strategy, and update the parameters of the improved self-supervised dynamic graph neural network in real time; S7. Output the final authentication result, judge whether to authorize access or perform subsequent operations according to the final authentication result, and take corresponding security protection measures if the authentication fails.

[0022] In this embodiment, the S3 specifically includes: S31. Use an improved self-supervised dynamic graph neural network to extract multi-modal features from the preprocessed data, construct a dynamic graph structure to represent the spatio-temporal dependencies in the preprocessed data. The dynamic graph structure changes dynamically over time. The multi-modal features include facial features, behavioral features, and speech features. The improvements of the improved self-supervised dynamic graph neural network include introducing dynamic edge weights and an improved contrast loss function for self-supervised learning in graph convolution operations. The contrast loss function is optimized using different loss functions at different training stages: ; Among them, G(t) represents the dynamic graph structure, V(t) represents the node set of the dynamic graph, is the feature in the video frame, E(t) represents the edge set of the dynamic graph, is the spatio-temporal dependency between video frames, and X(t) represents the node feature matrix; S32. Introduce dynamic edge weights and update the node representation through graph convolution operations: ; Among them, represents the feature vector representation of node v at the th layer, v represents the node of the dynamic graph, represents a non - linear activation function, \(N(v)\) represents the set of neighbor nodes of node \(v\), represents the degree of node \(k\), represents the weight matrix of the graph convolutional layer, represents the feature vector representation of node \(v\) at the - th layer, \(k\) represents the nodes of the dynamic graph, represents the dynamic edge weight between node \(v\) and neighbor node \(k\) at time \(t\): ; where, \(\exp\) represents the natural exponential function, \(N(v)\) represents the set of neighbor nodes of node \(v\), represents the nodes of the dynamic graph, represents the feature vector representation of node \(v\), represents the feature vector representation of node \(k\), represents the feature vector representation of node , \(k\) represents the nodes of the dynamic graph, \(Sim\) represents the cosine similarity between feature vector representations, represents the adjusted hyperparameter, represents the cosine similarity between the feature vector representations of node \(k\) and node \(v\), represents node and the cosine similarity between the feature vector representations of node \(v\); S33. Extract facial features. Use the node representation in the self - supervised dynamic graph neural network to capture the information of facial key points, update the facial features through graph convolution operations, and dynamically adjust in combination with the similarity between facial nodes to generate the feature vector representation of facial features: ; where, represents the feature vector representation of facial feature node \(v\) at the - th layer, represents a non - linear activation function, \(N(face)\) represents the set of neighbor nodes of facial feature nodes, represents the degree of node \(u\), represents the dynamic edge weight between facial nodes at time \(t\), represents the feature vector representation of node \(u\) at layer \(l\), \(u\) represents the nodes of the dynamic graph, represents the graph convolution weight matrix of facial features, represents the bias term; S34. Extract behavioral features. Use the time node representation in the self - supervised dynamic graph neural network to process the motion information in the video, process the dynamic behavior data of each time node through the graph convolutional layer, combine spatio - temporal dependency modeling, update the behavioral features, and generate the feature vector representation of behavioral features: ; Among them, represents the feature vector representation of the behavior feature time node t at the layer, represents the non-linear activation function, N(action) represents the set of neighbor nodes of the behavior feature node, represents the node degree, represents the dynamic edge weight between the behavior feature time nodes t, represents the behavior feature time node at the layer feature vector representation, represents the behavior feature time node of the dynamic graph, represents the graph convolution weight matrix of the behavior feature, represents the bias term; S35. Extract speech features through graph convolution operations to generate the feature vector representation of speech features: ; Among them, represents the layer feature vector representation of the speech feature node v, represents the non-linear activation function, N(speech) represents the set of neighbor nodes of the speech feature node, represents the node degree, represents the dynamic edge weight between speech features, represents the speech feature node at the layer feature vector representation, represents the speech feature node of the dynamic graph, represents the graph convolution weight matrix of the speech feature, represents the bias term; S36. Optimize the modal feature consistency of facial features, behavior features, and speech features through the contrast loss function of self-supervised learning. The contrast loss function minimizes the distance between similar features and maximizes the distance between different features; S37. Further optimize the feature vector representation through the reconstruction loss function of self-supervised learning, and update the parameters of the dynamic graph neural network through the backpropagation algorithm. The reconstruction loss function is adjusted by comparing the differences between the input features and the reconstructed features to minimize the reconstruction error: ; Among them, represents the reconstruction loss function, X represents the original input feature matrix, represents the feature matrix reconstructed by the dynamic graph neural network, denotes the Frobenius norm, which is the difference between matrix X and the difference between.

[0023] In this embodiment, S36 specifically includes: S361. Define a contrastive learning task. By minimizing the distance between similar samples and maximizing the distance between different samples, capture the similarities and differences between the modal features of different video streams. The modal features include facial features, behavioral features, and speech features; S362. Select positive and negative samples from the facial features, behavioral features, and speech features extracted by the dynamic graph neural network. The positive samples represent samples belonging to the same category, and the negative samples represent samples with different features and belonging to different categories; S363. For each pair of samples, calculate the cosine similarity between the feature vector representations: ; where, represents the cosine similarity between the feature vector representations and , represents the sample 's feature vector representation, represents the sample 's feature vector representation, represents the L2 norm of the feature vector representation , represents the L2 norm of the feature vector representation , and represent samples; S364. Define a contrastive loss function. During training, continuously optimize the learned features by minimizing the contrastive loss function. The contrastive loss function is optimized using different loss functions at different training stages: ; where, represents the contrastive loss function, N represents the total number of samples, represents the label. When and are samples of the same category, the value of the label is 1. If they are not samples of the same category, the value of the label is 0, represents the positive sample, represents the negative sample, represents the sample 's feature vector representation and the positive sample 's feature vector representation , represents the sample Eigenvector representation and negative samples Eigenvector representation The cosine similarity between, max represents taking the maximum value between the two, represents the sample Eigenvector representation, represents the positive sample Eigenvector representation, represents the negative sample Eigenvector representation, p represents the training phase; S365. By minimizing the contrast loss function , reduce the distance between positive samples and increase the distance between negative samples, and optimize the learned features through the Adam optimizer.

[0024] In this embodiment, the S5 specifically includes: S51. Authenticate the generated comprehensive identity authentication feature vector with a preset template stored in the database. Calculate the cosine similarity between the feature vector of the user to be authenticated and the template using weighted cosine similarity. The preset template contains trained user features, and each user feature is stored in the preset template in vector form: ; Among them, represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the preset template feature vector, represents the comprehensive identity authentication feature vector of the user to be authenticated, represents the preset template feature vector stored in the database, and respectively represent the Euclidean norms of the comprehensive identity authentication feature vector of the user to be authenticated and the preset template feature vector; S52. When calculating the similarity, dynamically adjust the matching weight of the feature vector corresponding to each video stream by weighted average according to the quality score of each video stream: ; Among them, represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the preset template feature vector after weighted average. M represents the total number of video streams, represents the cosine similarity between the feature vector of the i-th video stream and the preset template feature vector, represents the comprehensive identity authentication feature vector of the user to be authenticated of the i-th video stream, represents the preset template feature vector stored in the database of the i-th video stream, represents the weight of the i-th video stream: ; wherein, represents the quality score of the i-th video stream, represents the feature reliability score of the i-th video stream; S53. After calculating the weighted cosine similarity, compare the obtained weighted cosine similarity with the set authentication threshold. If the calculated weighted cosine similarity is greater than or equal to the preset authentication threshold, it is considered that the user identity matches and the authentication passes; otherwise, the authentication fails. S54. If the authentication fails, trigger security protection measures in the system. The system will start abnormal behavior detection and record relevant logs.

[0025] Example 1: To verify the feasibility of the present invention in implementation, the present invention is applied to the remote identity authentication system of a certain financial payment platform. This platform mainly provides online payment, transfer, and other financial services. When users make payments, they need to pass identity authentication to ensure the security of transactions. Traditional identity authentication methods, such as passwords and SMS verification codes, can no longer meet the users' requirements for high security and convenience to a large extent. Therefore, the remote identity authentication method based on multiple video identifications proposed by the present invention aims to improve the accuracy and anti-counterfeiting ability of identity authentication and solve the problems in traditional methods.

[0026] In this scenario, when users make payments, they perform face recognition through the cameras of mobile phones or computers. The system will collect real-time video stream data from different angles. According to the method of the present invention, the system uses video streams from different perspectives captured by multiple cameras to extract the facial features of users, and at the same time combines behavioral features (such as users' gestures or actions) and voice features for identity authentication. Through the improved self-supervised dynamic graph neural network, the system can automatically extract and fuse facial, behavioral, and voice features in the video stream to form a comprehensive identity authentication feature vector, and then compare it with the preset template stored in the database to confirm the identity of the user.

[0027] During the implementation process, the system will adopt a weighted method based on dynamic learning to adjust the weights of each video stream according to the authentication results of multiple video streams, so as to obtain the final authentication result. To verify the feasibility of this method in complex environments, the test data has been subjected to multiple authentication tests under different lighting conditions, angles, and backgrounds. The test results show that the present invention has a significant improvement compared with the traditional single video stream authentication method, especially in low-light environments, angle deviations, and complex backgrounds, showing higher accuracy and robustness.

[0028] In the experiment, 1000 different users were selected for identity authentication testing. The identity characteristics of each user included facial features, behavioral features, and voice features. The system authenticated through the data collected from multiple video streams, verifying the accuracy under variable environmental conditions.

[0029] Table 1 Comparison Table of Experimental Data ; According to the above comparison table of experimental data, it can be seen that the multi-video stream authentication method adopted by the present invention is superior to the traditional single-video stream authentication method under various environmental conditions. Especially in scenarios with low light and large angle deviations, it shows obvious advantages.

[0030] First of all, in an environment with good lighting, the authentication passing rate of multi-video stream authentication is generally higher than that of the traditional method. For example, under good lighting conditions, the authentication passing rate of User 1 is 99.8%, while the traditional method is only 96.5%. Other users also showed high authentication accuracy in the same environment, and the overall authentication passing rate is significantly better than the traditional method. This shows that using the authentication data of multiple video streams not only enhances the robustness of recognition but also effectively reduces the impact of lighting changes on the recognition results.

[0031] Secondly, in the case of poor lighting, the traditional single-video stream method shows a low authentication passing rate. Especially in low-light environments, the authentication accuracy drops significantly. For example, for User 5 in a low-light environment, the authentication passing rate of the traditional method is only 68.5%, while the authentication passing rate of the multi-video stream method is 92.4%, significantly improving the authentication accuracy. This shows that in low-light environments, multi-video stream authentication can make up for the deficiencies of single-video stream authentication through multi-angle data and information fusion, reducing the impact of uneven lighting.

[0032] In addition, in the case of angle deviation and background complexity, the advantages of the multi-video stream authentication method are more prominent. The experimental data of User 3 and User 8 clearly demonstrate this. In the case of large angle changes (User 3), the authentication passing rate of the traditional method is only 80.3%, while the passing rate of the multi-video stream method is increased to 98.5%. And in the case of a more complex background and weak lighting (User 8), the authentication passing rate of the traditional method is only 60.2%, while the passing rate of the multi-video stream method is 89.3%. This obvious gap shows that multi-video stream authentication can effectively alleviate the problems of occlusion, angle deviation, and background noise in the video stream through the fusion of multi-angle and multi-modal information, thus ensuring the stability of the authentication process.

[0033] Overall, the experimental results show that the multi-video stream authentication method of the present invention has shown improvement compared to traditional methods under various environmental conditions. Especially in complex environments, such as low light, angular deviation, and complex backgrounds, it has shown stronger robustness and accuracy. By introducing multi-video streams and a dynamic weighting mechanism, the present invention not only effectively improves the accuracy and anti-counterfeiting ability of identity authentication, but also optimizes the overall performance of the existing identity authentication system.

[0034] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A remote identity authentication method based on multiple video recognitions, characterized in that, It includes the following steps: S1. Collect real-time video stream data of the target user from different angles through a camera; S2. Preprocess the collected video stream data, remove noise, perform image enhancement, segment and calibrate video frames, locate key feature regions in the video, and generate preprocessed data; S3. Use an improved self-supervised dynamic graph neural network to extract facial features, behavioral features, and speech features from the preprocessed data. The improved self-supervised dynamic graph neural network constructs a dynamic graph structure to represent the spatio-temporal dependencies in the preprocessed data, combines a self-supervised learning mechanism, and performs unsupervised training through a contrastive learning task to maximize the consistency between modal features, optimize the correlation between modal features, and extract and generate corresponding feature vector representations for facial features, behavioral features, and speech features; S4. Fuse the feature vector representations of facial features, behavioral features, and speech features to generate a comprehensive identity authentication feature vector; S5. Compare and verify the generated comprehensive identity authentication feature vector, authenticate the user features with a preset template stored in the database, and weight the authentication results of each video stream using a weighted method based on dynamic learning according to the authentication results of multiple video streams. Adjust the weights in real time based on the video stream quality and the reliability score of the authentication features in the video stream, and calculate the final authentication result using the weighted average method; S6. During the authentication process, monitor the environmental changes of the video stream in real time, adjust the recognition strategy, and update the parameters of the improved self-supervised dynamic graph neural network in real time; S7. Output the final authentication result, and judge whether to authorize access or perform subsequent operations according to the final authentication result. If the authentication fails, take corresponding security protection measures.

2. The remote identity authentication method based on multiple video identifications according to claim 1, characterized in that, The specific content of S3 includes: S31. Use an improved self-supervised dynamic graph neural network to extract multi-modal features from the preprocessed data, construct a dynamic graph structure to represent the spatio-temporal dependencies in the preprocessed data. The dynamic graph structure changes dynamically over time. The multi-modal features include facial features, behavioral features, and speech features. The improvements of the improved self-supervised dynamic graph neural network include introducing dynamic edge weights and an improved contrast loss function for self-supervised learning in graph convolution operations. The contrast loss function is optimized using different loss functions at different training stages: ; Among them, G(t) represents the dynamic graph structure, V(t) represents the node set of the dynamic graph, is the feature in the video frame, E(t) represents the edge set of the dynamic graph, is the spatio-temporal dependency between video frames, and X(t) represents the node feature matrix; S32. Introduce dynamic edge weights and update the node representation through graph convolution operations; ; Among them, represents the feature vector representation of node v at the th layer, where v represents a node in the dynamic graph, represents a non-linear activation function, and N(v) represents the set of neighbor nodes of node v, represents the degree of node k, represents the weight matrix of the graph convolutional layer, represents the feature vector representation of node v at the th layer, where k represents a node in the dynamic graph, represents the dynamic edge weight between node v and neighbor node k at time t: ; where exp represents the natural exponential function, N(v) represents the set of neighbor nodes of node v, represents the nodes of the dynamic graph, represents the eigenvector representation of node v, represents the eigenvector representation of node k, represents node 's eigenvector representation, k represents the nodes of the dynamic graph, Sim represents the cosine similarity between eigenvector representations, represents the adjusted hyperparameter, represents the cosine similarity between the eigenvector representations of node k and node v, represents node and the cosine similarity between the eigenvector representations of node v; S33. Extract facial features, use the node representation in the self-supervised dynamic graph neural network to capture the information of facial key points, update the facial features through graph convolution operations, and dynamically adjust in combination with the similarity between facial nodes to generate a feature vector representation of facial features: ; Among them, represents the feature vector representation of the facial feature node v at the layer, represents the non-linear activation function, N(face) represents the set of neighbor nodes of the facial feature node, represents the degree of node u, represents the dynamic edge weight between facial nodes at time t, represents the feature vector representation of node u at the l-th layer, u represents the node of the dynamic graph, represents the graph convolution weight matrix of facial features, represents the bias term; S34. Extract behavioral features, use the temporal node representation in the self-supervised dynamic graph neural network to process the motion information in the video, process the dynamic behavior data of each temporal node through the graph convolutional layer, combine spatio-temporal dependency modeling, update the behavioral features, and generate the feature vector representation of the behavioral features: ; Among them, represents the feature vector representation of the behavior feature time node t at the layer, represents the non-linear activation function, and N(action) represents the set of neighbor nodes of the behavior feature node, represents the node degree, represents the dynamic edge weight between the behavior feature time nodes t, represents the behavior feature time node at the layer of the feature vector representation, represents the behavior feature time node of the dynamic graph, represents the graph convolution weight matrix of the behavior feature, represents the bias term; S35. Extract speech features through graph convolutional operations and generate the feature vector representation of the speech features: ; Among them, represents the -layer feature vector representation of the speech feature node v, represents the non-linear activation function, N(speech) represents the set of neighbor nodes of the speech feature node, represents the degree of the node, represents the dynamic edge weight between speech features, represents the speech feature node at the -layer feature vector representation, represents the speech feature node of the dynamic graph, represents the graph convolution weight matrix of the speech feature, represents the bias term; S36. Optimize the modality feature consistency of the facial features, behavioral features, and speech features through the contrast loss function of self-supervised learning. The contrast loss function minimizes the distance between the same-class features and maximizes the distance between different-class features; S37. Further optimize the feature vector representation through the reconstruction loss function of self-supervised learning, and update the parameters of the dynamic graph neural network through the backpropagation algorithm. The reconstruction loss function is adjusted by comparing the difference between the input features and the reconstructed features to minimize the reconstruction error: ; Among them, represents the reconstruction loss function, X represents the original input feature matrix, represents the feature matrix reconstructed by the dynamic graph neural network, represents the Frobenius norm, which is the difference between the matrix X and the difference between them.

3. The remote identity authentication method based on multiple video identifications according to claim 2, characterized in that, The specific content of S36 includes: S361. Define the contrast learning task, capture the similarity and difference between the modality features of different video streams by minimizing the distance between the same-class samples and maximizing the distance between different-class samples. The modality features include facial features, behavioral features, and speech features; S362. Select positive samples and negative samples from the facial features, behavioral features, and speech features extracted from the dynamic graph neural network. The positive samples represent samples belonging to the same class, and the negative samples represent samples with different features and belonging to different classes; S363. For each pair of samples, calculate the cosine similarity between the feature vector representations: ; Among them, represents the cosine similarity between and ; represents the feature vector representation of sample ; represents the feature vector representation of sample ; represents the L2 norm of the feature vector representation ; represents the L2 norm of the feature vector representation ; and represent samples; S364. Define the contrast loss function, and continuously optimize the learned features by minimizing the contrast loss function during the training process. The contrast loss function uses different loss functions for optimization at different training stages: ; Among them, represents the contrast loss function, N represents the total number of samples, represents the label. When and are samples of the same class, the value of the label is 1; if they are not samples of the same class, the value of the label is 0. represents the positive sample, represents the negative sample, represents the sample 's feature vector representation and the positive sample 's feature vector representation the cosine similarity between them, represents the sample 's feature vector representation and the negative sample 's feature vector representation the cosine similarity between them, max represents taking the maximum value between the two, represents the feature vector representation of the sample , represents the feature vector representation of the positive sample , represents the feature vector representation of the negative sample , p represents the training stage; S365. By minimizing the contrastive loss function , the distance between positive samples is reduced, and the distance between negative samples is increased, and the learned features are optimized by the Adam optimizer.

4. The remote identity authentication method based on multiple video identifications according to claim 1, wherein The specific content of S5 includes: S51. Authenticate the generated comprehensive identity authentication feature vector with the preset template stored in the database. Calculate the cosine similarity between the feature vector of the user to be authenticated and the template using the weighted cosine similarity. The preset template contains the trained user features, and each user feature is stored in the preset template in vector form: ; Among them, represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated and the preset template feature vector, represents the comprehensive identity authentication feature vector of the user to be authenticated, represents the preset template feature vector stored in the database, and respectively represent the Euclidean norms of the comprehensive identity authentication feature vector of the user to be authenticated and the preset template feature vector; S52. When calculating the similarity, dynamically adjust the matching weight of the feature vector corresponding to each video stream through the weighted average method according to the quality score of each video stream: ; Among them, represents the cosine similarity between the comprehensive identity authentication feature vector of the user to be authenticated after the weighted average method and the preset template feature vector, M represents the total number of video streams, represents the cosine similarity between the feature vector of the i-th video stream and the preset template feature vector, represents the comprehensive identity authentication feature vector of the user to be authenticated of the i-th video stream, represents the preset template feature vector stored in the database of the i-th video stream, represents the weight of the i-th video stream: ; Among them, represents the quality score of the i-th video stream, represents the feature reliability score of the i-th video stream; S53. After calculating the weighted cosine similarity, compare the obtained weighted cosine similarity with the set authentication threshold. If the calculated weighted cosine similarity is greater than or equal to the preset authentication threshold, the user identity is considered to match and the authentication is passed; otherwise, the authentication fails; S54. If the authentication fails, trigger the security protection measures in the system. The system will start the abnormal behavior detection and record the relevant logs.

Citation Information

Patent Citations

  • Block chain-fused physical exercise action analysis and correction method and system

    CN116844084A

  • Video network asset identity credibility identification method and system

    CN118821029A

  • Intelligent access control management method and system based on multi-mode identification and Internet of Things technology

    CN118968665A

  • Multi-modal data fusion identity authentication and security monitoring system based on deep learning

    CN120048010A

  • Face verification method and apparatus based on triplet loss, and computer device and storage medium

    WO2019128367A1