Video two-factor identity authentication method and system, electronic device and storage medium

CN117892280BActive Publication Date: 2026-09-29E SURFING IOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311864021.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2026-09-29
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

手势认不具备明显的身份特征,攻击者可以使用用户照片和其他用户的手势组合尝试进行身份认证,因此,这种身份认证方式安全性较低

Benefits of technology

[0046]本申请提出的视频双因子动态认证方法、系统、电子设备及存储介质,其通过先获取静态的人脸图像数据,对人脸图像数据进行人脸特征提取处理得到对象人脸特征,根据对象人脸特征进行人脸识别,当人脸识别通过,表明其为系统用户,进一步通过摄像头采集动态的视频数据,对视频数据进行第一处理得到口型特征,以及对视频数据进行第二处理得到手势语义特征,然后根据口型特征和手势语义特征双因子分析对应的密码字符串,根据密码字符串进行身份认证,得到身份认证结果。本申请通过利用视频数据中的口型和手势的组合获得复杂度更高的密码输入形式,提高攻击者欺骗验证的难度,从而提高身份认证的安全性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117892280B_ABST
    Figure CN117892280B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a video double-factor dynamic authentication method, system, electronic equipment and storage medium, and belongs to the technical field of artificial intelligence. The embodiment of the application obtains a static face image data first, carries out face feature extraction processing on the face image data to obtain an object face feature, carries out face recognition according to the object face feature, and when the face recognition passes, it indicates that it is a system user, further collects dynamic video data through a camera, carries out first processing on the video data to obtain a mouth shape feature, carries out second processing on the video data to obtain a gesture semantic feature, then carries out double-factor analysis on a corresponding password string according to the mouth shape feature and the gesture semantic feature, carries out identity authentication according to the password string, and obtains an identity authentication result. The embodiment of the application obtains a password input form with higher complexity by using the combination of the mouth shape and the gesture in the video data, improves the difficulty of the attacker to cheat the verification, and thus improves the security of the identity authentication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a video two-factor dynamic authentication method, system, electronic device and storage medium. Background Technology

[0002] Facial recognition is increasingly used due to its convenience, speed, and low barrier to entry. However, facial recognition access control systems also have security vulnerabilities. Some systems lack liveness detection, allowing for "spoofing" verification using photos or videos under certain conditions. To improve security, some technologies combine facial recognition with gesture recognition for password input, addressing the lack of liveness detection and enhancing authentication security. However, gesture recognition lacks clear identifying characteristics, allowing attackers to attempt authentication using user photos and combinations of gestures from other users. Therefore, this authentication method has relatively low security. Summary of the Invention

[0003] The main objective of this application is to propose a video two-factor dynamic authentication method, system, electronic device, and storage medium, which aims to improve the security of identity authentication.

[0004] To achieve the above objectives, one aspect of this application proposes a video two-factor dynamic authentication method, comprising the following steps:

[0005] Acquire facial image data through a camera;

[0006] The facial image data is processed to extract facial features to obtain the facial features of the object;

[0007] Facial recognition is performed based on the facial features of the object. Once facial recognition is successful, video data is captured through a camera.

[0008] The video data is first processed to obtain lip-sync features, and the video data is second processed to obtain gesture semantic features;

[0009] The password string is determined based on the lip shape features and the gesture semantic features;

[0010] The identity is authenticated based on the password string, and the authentication result is obtained.

[0011] In some embodiments, the first processing of the video data to obtain lip-sync features includes the following steps:

[0012] Facial feature recognition and segmentation are performed on each image frame in the video data to obtain a facial image sequence;

[0013] Each image in the facial image sequence is input into the mouth key point localization model to obtain the coordinates of multiple mouth key points;

[0014] The mouth shape features of each image in the facial image sequence are determined based on the coordinates of multiple mouth key points.

[0015] In some embodiments, the plurality of mouth key point coordinates include the coordinates of the left corner of the mouth, the right corner of the mouth, the upper lip, and the lower lip. Determining the mouth shape features of each image in the facial image sequence based on the plurality of mouth key point coordinates includes the following steps:

[0016] The mouth width is determined based on the coordinates of the left and right corners of the mouth.

[0017] The mouth shape length is determined based on the coordinates of the upper lip point and the lower lip point;

[0018] The ratio of mouth opening degree is determined based on the mouth width and the mouth length;

[0019] The mouth opening degree ratio is compared with the preset target mouth shape ratio. If the mouth opening degree ratio is greater than the target mouth shape ratio, the mouth shape feature is determined to be not open; otherwise, the mouth shape feature is open.

[0020] In some embodiments, the second processing of the video data to obtain gesture semantic features includes the following steps:

[0021] Determine the skin color likelihood map based on the skin color features in the facial features of the object;

[0022] Based on the skin color likelihood map, hand feature recognition and segmentation are performed on the image frames in the video data to obtain a hand binarized image;

[0023] The hand binarized image is input into a gesture recognition model based on a convolutional neural network to obtain gesture semantic features.

[0024] In some embodiments, determining the password string based on the lip shape features and the gesture semantic features includes the following steps:

[0025] Determine whether the lip shape feature of the current frame image is an open mouth;

[0026] When the mouth shape feature is an open mouth, the current frame image in the video data is subjected to a second processing to determine the corresponding gesture semantic features;

[0027] If the mouth shape feature indicates that the mouth is not open, then no second processing is performed on the current frame image in the video data;

[0028] The password string is determined based on multiple gesture semantic features obtained after the second processing.

[0029] In some embodiments, determining the password string based on multiple gesture semantic features obtained through the second processing includes the following steps:

[0030] The multiple gesture semantic features obtained after the second processing are combined into a gesture feature sequence;

[0031] Examine the gesture feature sequence to see if there are any abnormal sequence segments with consecutive identical gesture semantic features;

[0032] If there are abnormal sequence segments, the abnormal sequence segments are deduplicated to obtain a new gesture feature sequence;

[0033] The password string is determined based on the new sequence of gesture features.

[0034] In some embodiments, the authentication process based on the password string to obtain the authentication result includes the following steps:

[0035] Obtain the object's preset password based on the object's facial features;

[0036] The password string is compared with the object's preset password. If the password string is the same as the object's preset password, the authentication result is successful.

[0037] To achieve the above objectives, another aspect of this application proposes a video two-factor dynamic authentication system, comprising:

[0038] The first module is used to acquire facial image data through a camera;

[0039] The second module is used to perform facial feature extraction processing on the facial image data to obtain the facial features of the object;

[0040] The third module is used to perform facial recognition based on the facial features of the object. When the facial recognition is successful, video data is collected through the camera.

[0041] The fourth module is used to perform a first processing on the video data to obtain lip-sync features, and a second processing on the video data to obtain gesture semantic features;

[0042] The fifth module is used to determine the password string based on the lip shape features and the gesture semantic features;

[0043] The sixth module is used to perform identity authentication based on the password string and obtain the identity authentication result.

[0044] To achieve the above objectives, another aspect of the present application provides an electronic device, which includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory. When the program is executed by the processor, it implements the video two-factor dynamic authentication method described in the above embodiments.

[0045] To achieve the above objectives, another aspect of the embodiments of this application proposes a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the video two-factor dynamic authentication method described in the above embodiments.

[0046] This application proposes a video two-factor dynamic authentication method, system, electronic device, and storage medium. It first acquires static facial image data, performs facial feature extraction on the facial image data to obtain the target's facial features, and then performs facial recognition based on these features. If the facial recognition passes, it indicates the user is a system user. Next, it captures dynamic video data through a camera, performs a first processing to obtain lip-sync features, and a second processing to obtain gesture semantic features. Then, it performs two-factor analysis based on the lip-sync and gesture semantic features to obtain the corresponding password string, and performs authentication based on the password string to obtain the authentication result. This application utilizes the combination of lip-sync and gestures in video data to obtain a more complex password input format, increasing the difficulty for attackers to spoof the verification, thereby improving the security of authentication. Attached Figure Description

[0047] Figure 1 This is a flowchart of the video two-factor dynamic authentication method provided in the embodiments of this application;

[0048] Figure 2 yes Figure 1 The flowchart of step S104, which involves performing the first processing on the video data to obtain the lip-shape feature sequence, is shown below.

[0049] Figure 3 yes Figure 2 The flowchart of step S203 in the process;

[0050] Figure 4 yes Figure 1 The flowchart of the second processing step in step S104 to obtain gesture semantic features from video data;

[0051] Figure 5 yes Figure 1 The flowchart of step S105 in the process;

[0052] Figure 6 yes Figure 5 The flowchart of step S504 in the process;

[0053] Figure 7 yes Figure 1 The flowchart of step S106 in the process;

[0054] Figure 8 This is a schematic diagram of a video two-factor dynamic authentication system provided in an embodiment of this application;

[0055] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application;

[0056] Figure 10 This is a schematic diagram illustrating an application scenario of the video two-factor dynamic authentication method provided in this application embodiment;

[0057] Figure 11 This is a schematic diagram of the password string recognition process based on video input provided in the embodiments of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0059] It should be noted that although functional modules are divided in the system diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0061] First, let's analyze some of the terms used in this application:

[0062] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0063] Image captioning generates natural language descriptions for images, helping applications understand the semantics expressed in the visual scene. For example, image captioning can convert image retrieval into text retrieval, classify images, and improve retrieval results. While people can often describe the details of a visual scene with a quick glance, automatically adding descriptions to images is a comprehensive and challenging computer vision task, requiring the conversion of complex information contained in the image into natural language descriptions. Compared to ordinary computer vision tasks, image captioning not only requires identifying objects in an image but also associating the identified objects with natural semantics and describing them in natural language. Therefore, image captioning requires extracting deep features from the image, associating them with semantic features, and converting them to generate descriptions.

[0064] Convolutional Neural Networks (CNNs) are a type of feedforward neural network that includes convolutional computations and has a deep structure. They can perform supervised learning using labeled training data to accomplish tasks such as visual image recognition and object detection.

[0065] Deep learning: Deep learning learns the inherent patterns and representational layers of sample data (such as images, speech, and text), enabling machines to have analytical and learning capabilities like humans. It can recognize data such as text, images, and sounds and is widely used in the field of artificial intelligence. Convolutional neural networks are a commonly used structure in deep learning.

[0066] A skin likelihood map is an image used in skin color detection that represents the probability that a skin color pixel belongs to the foreground (i.e., a face). If the color of a pixel is very close to skin color, its skin likelihood value is close to 1, indicating that this pixel is very likely to belong to a face. If the color of a pixel differs significantly from skin color, its skin likelihood value will be close to 0, indicating that this pixel is unlikely to belong to a face. In skin color detection, the skin likelihood map is a very important intermediate result. Using the skin likelihood map, the original image can be transformed into a binary image (with only foreground and background colors), facilitating subsequent face detection and processing.

[0067] Otsu's algorithm is an image binarization algorithm, also known as the maximum inter-class variance method. Its purpose is to determine a threshold that divides an image into black and white parts. Specifically, Otsu's algorithm performs thresholding binarization on the grayscale values ​​of an image. If the input is a color image, it needs to be converted to grayscale before calculation. The goal of Otsu's algorithm is to find a grayscale threshold that maximizes the sum of the variances of pixels above and below that threshold. Variance represents the dispersion of pixels; the larger the variance, the more dispersed the pixel distribution, i.e., the lower the correlation and the clearer the distinction between black and white.

[0068] Based on this, embodiments of this application provide a video two-factor dynamic authentication method, system, electronic device, and storage medium, aiming to improve the security of identity authentication.

[0069] The video two-factor dynamic authentication method, system, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the video two-factor dynamic authentication method in this application is described.

[0070] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0071] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0072] The video two-factor dynamic authentication method provided in this application relates to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the video two-factor dynamic authentication method, but is not limited to the above forms.

[0073] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0074] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0075] Figure 1 This is an optional flowchart of the video two-factor dynamic authentication method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.

[0076] Step S101: Acquire facial image data through the camera;

[0077] Step S102: Perform facial feature extraction processing on the facial image data to obtain the facial features of the object;

[0078] Step S103: Perform facial recognition based on the facial features of the object. If the facial recognition is successful, video data is collected through the camera.

[0079] Step S104: Perform a first processing on the video data to obtain lip-sync features, and perform a second processing on the video data to obtain gesture semantic features;

[0080] Step S105: Determine the password string based on lip shape features and gesture semantic features;

[0081] Step S106: Perform identity authentication based on the password string to obtain the authentication result.

[0082] Steps S101 to S106 of this embodiment involve first acquiring static facial image data, then performing facial feature extraction processing on the facial image data to obtain the object's facial features, and performing facial recognition based on the object's facial features. If the facial recognition passes, it indicates that the user is a system user. Next, dynamic video data is captured via a camera, and the video data undergoes a first processing step to obtain lip-sync features, and a second processing step to obtain gesture semantic features. Then, a two-factor analysis of the lip-sync features and gesture semantic features is performed on the corresponding password string, and identity authentication is performed based on the password string to obtain the authentication result. This embodiment utilizes the combination of lip-sync and gestures in video data to obtain a more complex password input format, increasing the difficulty for attackers to spoof the verification, thereby improving the security of identity authentication.

[0083] According to some embodiments of this application, in conjunction with Figure 10 This application describes the application scenario of the video two-factor dynamic authentication method. The video two-factor dynamic authentication method of this application is applied in a backend server cluster, which includes an algorithm server and a password management server. The backend server cluster is connected to the camera terminal, and the specific authentication process is as follows:

[0084] When a user appears in front of the camera terminal, the camera terminal initiates face capture to obtain static face image data.

[0085] The camera terminal sends facial image data to the algorithm server. The algorithm server performs facial recognition feature analysis based on the facial image data and inputs the facial recognition result as a user account (user ID) into the password management server for matching. If the password management server has the same user account, it means that the user's facial recognition has passed; otherwise, the facial recognition has failed.

[0086] After facial recognition is successful, the algorithm server triggers the camera terminal to collect image data a second time.

[0087] The camera terminal transmits video data containing complete image frames to the algorithm server;

[0088] The algorithm server performs first and second processing on the image frames in the video data to obtain lip-sync features and gesture semantic features. Based on the processing results, it determines the characters represented in the image frames, concatenates the characters to obtain a complete password string, and transmits the password string as the user password to the password management server.

[0089] The password management server authenticates the user based on the pre-stored password corresponding to the user ID, and then sends the verification result back to the user.

[0090] In step S101 of some embodiments, the face image data refers to the data captured by the camera terminal facing the user's face, and the face is represented by the depth values ​​of each pixel.

[0091] In step S102 of some embodiments, commonly used face feature extraction algorithms can be employed to extract the user's facial feature information (i.e., the object's facial features) contained in the face image data. Face feature extraction algorithms can be based on geometric features, algebraic features, or deep learning. Geometric feature-based methods describe facial features by extracting geometric information such as the position, distance, and angle of facial feature points. Algebraic feature-based methods treat the face image as a matrix and extract features through algebraic operations; for example, methods such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA) can extract the main facial features through dimensionality reduction. Deep learning-based methods train deep neural networks using deep learning techniques, which can learn complex feature representations of face images; for example, convolutional neural networks (CNNs) are used in face recognition tasks.

[0092] In step S103 of some embodiments, the facial features of the object are matched with the facial features pre-stored in the server. When there are features in the server that are the same as the facial features of the object, it indicates that the current object is the target user of the system. For example, when the photo obtained by the access control system passes the facial recognition of the dormitory access control system, it indicates that the user in the photo is a resident of the dormitory.

[0093] In step S104 of some embodiments, a lip recognition algorithm can be used to perform a first processing on the video data to obtain lip features in the video data. The lip recognition algorithm can be a target recognition model for recognizing mouth features built on a convolutional neural network, such as the YOLO model. The lip recognition algorithm can first recognize the face part of the image frame in the video data based on the target recognition model for recognizing faces to obtain a face detection box. Then, the target recognition model for recognizing lip shapes can be used to detect part of the image in the face detection box to obtain a lip detection box. At the same time, the output lip detection box corresponds to lip features. Then, a classifier can be used to classify the lip features to obtain a classification result. The classification result is the lip feature. The classification result can be open mouth or not open mouth. If it is determined that the mouth is open, the number or letter corresponding to the lip feature can be further determined. A gesture recognition algorithm can be used to perform a second processing on video data to obtain gesture features in the video data. The gesture recognition algorithm uses image description technology and can be a target recognition model built on convolutional neural networks to identify hand features. The gesture recognition algorithm can use the target recognition model to identify hand images in video data to obtain hand detection boxes. Then, a classifier is used to classify the features of the detection boxes to obtain classification results. The classification results are the gesture semantic features, which can be the numbers or letters corresponding to the gesture.

[0094] Please see Figure 2 In some embodiments, step S104, which involves performing a first process on the video data to obtain a lip-sync feature sequence, may include, but is not limited to, steps S201 to S203:

[0095] Step S201: Perform facial feature recognition and segmentation on each image frame in the video data to obtain a facial image sequence;

[0096] Step S202: Input each image in the face image sequence into the mouth key point localization model to obtain the coordinates of multiple mouth key points;

[0097] Step S203: Determine the mouth shape features of each image in the face image sequence based on the coordinates of multiple mouth key points.

[0098] In this embodiment, each image frame in the video data is input into a detection box for the face portion obtained based on a face recognition model. This detection box is then cropped to obtain the face image in the image frame. The face images from all image frames in the video data constitute a face image sequence. Then, the face images in the face image sequence are sequentially input into a pre-trained mouth keypoint localization model to perform face keypoint localization, obtaining the coordinates of multiple mouth keypoints in the image. Geometric morphology analysis algorithms can be used to analyze the coordinates of these multiple mouth keypoints, and then a classifier can be used for classification to obtain the mouth shape features of each image in the face image sequence.

[0099] Understandably, mouth keypoints can be defined during model training and may include the left corner of the mouth, right corner of the mouth, upper lip, and lower lip. The training process for the mouth keypoint localization model is as follows: The input face image is preprocessed, including grayscale conversion, scaling, and rotation, to better train the neural network; a neural network structure similar to face recognition or face detection, such as VGG or ResNet, can be used, or a custom neural network structure can be defined, such as convolutional layers, pooling layers, and fully connected layers, to determine the model's network structure; the neural network is trained using the face image and labeled mouth keypoint location data; optimization algorithms such as gradient descent are used to update the neural network's weights and biases to minimize the error between the predicted keypoint locations and the actual locations. The network parameters are continuously updated using training data to obtain the mouth keypoint localization model.

[0100] Please see Figure 3 In some embodiments, the coordinates of multiple mouth key points include the coordinates of the left corner of the mouth, the right corner of the mouth, the upper lip, and the lower lip. In step S203, the step of determining the mouth shape features of each image in the facial image sequence based on the coordinates of multiple mouth key points may include, but is not limited to, steps S301 to S304:

[0101] Step S301: Determine the mouth width based on the coordinates of the left and right corners of the mouth;

[0102] Step S302: Determine the mouth shape length based on the coordinates of the upper lip and lower lip points;

[0103] Step S303: Determine the ratio of mouth opening degree based on mouth width and mouth length;

[0104] Step S304: Compare the mouth opening degree ratio with the preset target mouth shape ratio. If the mouth opening degree ratio is greater than the target mouth shape ratio, the mouth shape feature is determined to be not open; otherwise, the mouth shape feature is open.

[0105] For example, the coordinates of the left corner of the mouth are a1, the coordinates of the right corner of the mouth are a2, the coordinates of the upper lip are b1, and the coordinates of the lower lip are b2. The formula for calculating the ratio of mouth opening degree is as follows:

[0106]

[0107] The mouth opening degree ratio is input into a classifier for classification, resulting in a mouth shape feature indicating whether the mouth is open or closed. The classifier compares the mouth opening degree ratio with a preset object mouth shape ratio. If the mouth opening degree ratio is greater than the object mouth shape ratio, the mouth shape feature is determined to be closed; otherwise, it is determined to be open.

[0108] It should be noted that the preset object mouth shape ratio can be obtained by multiplying the ratio of the mouth width to the mouth length when the user's mouth is completely closed by the threshold coefficient. The threshold coefficient can be obtained by using an adaptive mechanism and adjusting the classifier based on the feedback of accurate recognition results. By setting the object mouth shape ratio in a personalized way, the accuracy of user mouth shape recognition can be improved.

[0109] Please see Figure 4 In some embodiments, step S104, which involves performing a second process on the video data to obtain gesture semantic features, may include, but is not limited to, steps S401 to S403:

[0110] Step S401: Determine the skin color likelihood map based on the skin color features in the facial features of the object;

[0111] Step S402: Based on the skin color likelihood map, perform hand feature recognition and segmentation on the image frames in the video data to obtain a hand binarized image.

[0112] Step S403: Input the hand binarized image into the gesture recognition model based on the convolutional neural network to obtain gesture semantic features.

[0113] In this embodiment, the skin color likelihood map generated during the facial feature extraction algorithm is obtained. Then, the Otsu algorithm is used to identify the hand image in the image, resulting in a binary image of the hand. After noise removal, gesture segmentation is completed. Then, the feature extraction function of a convolutional neural network is used to extract feature vectors. Finally, a random forest classifier is used to classify the extracted feature vectors, thereby obtaining the gesture semantic features. Convolutional neural networks have the ability to learn in layers, enabling them to collect more representative information from images. Random forests offer randomness in sample and feature selection and average the results of each decision tree, making them less prone to overfitting.

[0114] Skin color likelihood map generation algorithms include single Gaussian, Gaussian mixture, Bayesian, and elliptic models. Taking the elliptic model as an example, based on extensive skin statistics, if skin information is mapped to the YCrCb space, these skin pixels in the CrCb two-dimensional space approximate an elliptical distribution. Therefore, if we obtain a CrCb ellipse, for the next coordinate (Cr, Cb), we only need to determine whether it is inside the ellipse (including the boundary). If it is, it can be determined as a skin pixel; otherwise, it is a non-skin pixel. The elliptic model is as follows:

[0115]

[0116]

[0117] In step S105 of some embodiments, the password string can be determined by combining lip-shape features and gesture semantic features. The lip-shape features are mainly used to determine whether the gesture semantic features corresponding to the time sequence are accurate password characters, and the gesture semantic features are mainly used to determine each character that forms the password string.

[0118] Please see Figure 5 In some embodiments, step S105, which involves determining the password string based on lip-reading features and gesture semantic features, may include, but is not limited to, steps S501 to S502:

[0119] Step S501: Determine whether the mouth shape feature of the current frame image is an open mouth;

[0120] Step S502: When the lip shape feature is an open mouth, perform a second processing on the current frame image in the video data to determine the corresponding gesture semantic features;

[0121] Step S503: If the lip feature indicates that the mouth is not open, then no second processing is performed on the current frame image in the video data;

[0122] Step S504: Determine the password string based on the multiple gesture semantic features obtained after the second processing.

[0123] In this embodiment, after acquiring video data, each image frame in the video data can be processed sequentially to determine whether the lip shape feature in the current processing frame is an open mouth. If the mouth is open, the image frame is processed again to obtain the gesture semantic features for the current time. If the mouth is not open, the image frame is marked as an invalid frame and no further processing is performed. By performing the above processing logic on each image frame in the video data, valid gesture semantic features can be obtained, thereby determining the corresponding password string. In this embodiment, by first performing lip shape recognition on the image frames in the video data, and then performing the second processing on the image frame if the mouth is open in the current frame, and not performing the second processing on the image frame if the mouth is not open, the computational load can be reduced and the gesture recognition efficiency can be improved without affecting gesture recognition.

[0124] Please see Figure 6 In some embodiments, step 504 may include, but is not limited to, steps S601 to S602:

[0125] Step S601: Combine the multiple gesture semantic features obtained after the second processing into a gesture feature sequence;

[0126] Step S602: Check whether there are abnormal sequence segments with consecutive identical gesture semantic features in the gesture feature sequence;

[0127] Step S603: If there is an abnormal sequence segment, perform a deduplication operation on the abnormal sequence segment to obtain a new gesture feature sequence;

[0128] Step S604: Determine the password string based on the new gesture feature sequence.

[0129] In this embodiment, considering that the duration of each gesture is different when the user makes a gesture, in order to avoid repeatedly recognizing the same image frame of the same gesture and thus improve the accuracy of gesture recognition, it is possible to check whether there are consecutive identical gesture semantic features in the gesture feature sequence. If there are consecutive identical gesture semantic features, the consecutive identical gesture semantic features are deleted into one gesture semantic feature, and then the password string is determined based on the new gesture feature sequence.

[0130] Specifically, in combination Figure 11 The user-input-based password string recognition process in this application embodiment is as follows:

[0131] S010. Create an empty password string and a temporary recognition result string;

[0132] S020, Send the complete input video stream to the server;

[0133] S030. Obtain an image sequence containing lip movements and gestures based on the video stream;

[0134] S040. Input the next frame image of the image sequence into the algorithm server;

[0135] S050. Perform the first processing on the current frame image in the image sequence to recognize the user's lip movements, and obtain the lip movement recognition result;

[0136] S060. Determine whether the user has opened their mouth based on the lip shape recognition result;

[0137] S070. If the mouth is not open, discard the image frame and return to step S040.

[0138] S080. If the user opens their mouth, the second processing of the frame image is performed to recognize the user's static gesture, and the gesture recognition result is obtained.

[0139] S090. Determine whether the user's gesture is empty based on the gesture recognition result;

[0140] S110. If the gesture recognition result is not empty, determine whether it is the same as the recognition result of the previous frame in the temporary recognition string.

[0141] S111. If the image is identical to the previous recognition result in the temporary recognition result string, it is determined to be the same input. The current image frame is discarded, and the process returns to step S040.

[0142] S112. If the result is different from the previous recognition result, update the temporary recognition result string data and append it to the password string, then return to step S040.

[0143] S120. If the gesture recognition result is empty, then confirm that all recognition is complete and output the password string.

[0144] In another example, the password string recognition process based on user video input can also be as follows: First and second processing are performed on each frame of the video data to obtain a lip-sync feature sequence and a gesture feature sequence. The lip-sync feature sequence includes the lip-sync features of each captured frame, and the gesture feature sequence includes the gesture semantic features of each captured frame. The image number of the lip-sync feature in the lip-sync feature sequence that is not open-mouthed is identified. The corresponding gesture semantic features in the gesture feature sequence are deleted according to the image number, resulting in an updated gesture feature sequence. The gesture feature sequence is checked for any abnormal sequence segments with consecutive identical gesture semantic features. If abnormal sequence segments exist, deduplication is performed on these segments, resulting in a second updated gesture feature sequence. The corresponding characters are determined based on the gesture semantic features contained in the gesture feature sequence, thereby determining the password string.

[0145] In step S106 of some embodiments, after obtaining the password string input by the user, authentication can be performed based on the password preset by the user in the server, thereby obtaining the authentication result.

[0146] Please see Figure 7 In some embodiments, step S106 may include, but is not limited to, steps S701 to S702:

[0147] Step S701: Obtain the object's preset password based on the object's facial features;

[0148] Step S702: Compare the password string with the object's preset password. If the password string is the same as the object's preset password, the authentication result is successful.

[0149] In this embodiment, after obtaining the facial features of the object, the facial features of the object are used as the user ID to look up the pre-stored object preset password in the server. The password string is compared with the object preset password. If the password string is the same as the object preset password, the user passes the identity authentication.

[0150] This application uses a combination of intelligent image algorithms to obtain more complex user passwords and allows users to actively add interference when entering password features, which increases the difficulty and cost of spoofing verification and replay attacks, thus improving the security of identity authentication compared to the original single face recognition method.

[0151] Please see Figure 8 This application also provides a video two-factor dynamic authentication system, including:

[0152] The first module is used to acquire facial image data through a camera;

[0153] The second module is used to extract facial features from facial image data to obtain the facial features of the object.

[0154] The third module is used to perform facial recognition based on the facial features of the object. When the facial recognition is successful, video data is collected through the camera.

[0155] The fourth module is used to perform a first processing on the video data to obtain lip-sync features, and a second processing on the video data to obtain gesture semantic features;

[0156] The fifth module is used to determine the password string based on lip-reading features and gesture semantic features;

[0157] The sixth module is used to perform identity authentication based on the password string and obtain the authentication result.

[0158] It is understood that the content of the above-described video two-factor dynamic authentication method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above-described video two-factor dynamic authentication method embodiments, and the beneficial effects achieved are also the same as those achieved in the above-described video two-factor dynamic authentication method embodiments.

[0159] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned video two-factor dynamic authentication method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0160] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0161] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0162] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the video two-factor dynamic authentication method of the embodiments of this application.

[0163] The input / output interface 903 is used to implement information input and output;

[0164] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0165] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0166] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0167] This application embodiment also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described video two-factor dynamic authentication method.

[0168] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0169] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0170] It will be understood by those skilled in the art that Figure 1-7 The technical solutions shown do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0171] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0172] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0173] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0174] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0175] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0176] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0177] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0178] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0179] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A video two-factor dynamic authentication method, characterized in that, Includes the following steps: Acquire facial image data through a camera; The facial image data is processed to extract facial features to obtain the facial features of the object; Facial recognition is performed based on the facial features of the object. Once facial recognition is successful, video data is captured through a camera. The video data is first processed to obtain lip-sync features, and the video data is second processed to obtain gesture semantic features; The password string is determined based on the lip shape features and the gesture semantic features; The identity is authenticated based on the password string to obtain the authentication result. Determining the password string based on the lip shape features and the gesture semantic features includes the following steps: Determine whether the lip shape feature of the current frame image is an open mouth; When the mouth shape feature is an open mouth, the current frame image in the video data is subjected to a second processing to determine the corresponding gesture semantic features; If the mouth shape feature indicates that the mouth is not open, then no second processing is performed on the current frame image in the video data; The password string is determined based on multiple gesture semantic features obtained after the second processing.

2. The video two-factor dynamic authentication method according to claim 1, characterized in that, The first processing of the video data to obtain lip-sync features includes the following steps: Facial feature recognition and segmentation are performed on each image frame in the video data to obtain a facial image sequence; Each image in the facial image sequence is input into the mouth key point localization model to obtain the coordinates of multiple mouth key points; The mouth shape features of each image in the facial image sequence are determined based on the coordinates of multiple mouth key points.

3. The video two-factor dynamic authentication method according to claim 2, characterized in that, The multiple mouth key point coordinates include the coordinates of the left corner of the mouth, the right corner of the mouth, the upper lip, and the lower lip. Determining the mouth shape features of each image in the facial image sequence based on these multiple mouth key point coordinates includes the following steps: The mouth width is determined based on the coordinates of the left and right corners of the mouth. The mouth shape length is determined based on the coordinates of the upper lip point and the lower lip point; The ratio of mouth opening degree is determined based on the mouth width and the mouth length; The mouth opening degree ratio is compared with the preset target mouth shape ratio. If the mouth opening degree ratio is greater than the target mouth shape ratio, the mouth shape feature is determined to be not open; otherwise, the mouth shape feature is open.

4. The video two-factor dynamic authentication method according to claim 1, characterized in that, The second processing of the video data to obtain gesture semantic features includes the following steps: Determine the skin color likelihood map based on the skin color features in the facial features of the object; Based on the skin color likelihood map, hand feature recognition and segmentation are performed on the image frames in the video data to obtain a hand binarized image; The hand binarized image is input into a gesture recognition model based on a convolutional neural network to obtain gesture semantic features.

5. The video two-factor dynamic authentication method according to claim 4, characterized in that, Determining the password string based on multiple gesture semantic features obtained after the second processing includes the following steps: The multiple gesture semantic features obtained after the second processing are combined into a gesture feature sequence; Examine the gesture feature sequence to see if there are any abnormal sequence segments with consecutive identical gesture semantic features; If there are abnormal sequence segments, the abnormal sequence segments are deduplicated to obtain a new gesture feature sequence; The password string is determined based on the new sequence of gesture features.

6. The video two-factor dynamic authentication method according to any one of claims 1 to 4, characterized in that, The process of performing identity authentication based on the password string to obtain the authentication result includes the following steps: Obtain the object's preset password based on the object's facial features; The password string is compared with the object's preset password. If the password string is the same as the object's preset password, the authentication result is successful.

7. A video two-factor dynamic authentication system, characterized in that, include: The first module is used to acquire facial image data through a camera; The second module is used to perform facial feature extraction processing on the facial image data to obtain the facial features of the object; The third module is used to perform facial recognition based on the facial features of the object. When the facial recognition is successful, video data is collected through the camera. The fourth module is used to perform a first processing on the video data to obtain lip-sync features, and a second processing on the video data to obtain gesture semantic features; The fifth module is used to determine the password string based on the lip shape features and the gesture semantic features; The sixth module is used to perform identity authentication based on the password string and obtain the authentication result; The fifth module is specifically used for: Determine whether the lip shape feature of the current frame image is an open mouth; When the mouth shape feature is an open mouth, the current frame image in the video data is subjected to a second processing to determine the corresponding gesture semantic features; If the mouth shape feature indicates that the mouth is not open, then no second processing is performed on the current frame image in the video data; The password string is determined based on multiple gesture semantic features obtained after the second processing.

8. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory. When the program is executed by the processor, it implements the steps of the video two-factor dynamic authentication method as described in any one of claims 1 to 6.

9. A storage medium, said storage medium being a computer-readable storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the video two-factor dynamic authentication method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-dimensional identity authentication system and method based on face recognition

    CN111159676A

  • Identity authentication method and apparatus, terminal and server

    WO2016034069A1