Identity verification method and system based on image recognition and storage medium

Through methods such as multi-angle document image upload, encrypted database, live detection and secondary identity authentication, the accuracy and efficiency of image recognition authentication in the prior art are solved, and higher authentication accuracy and security are achieved.

CN120260143APending Publication Date: 2025-07-04IND & COMMERCIAL BANK OF CHINA LTD

Patent Information

Application Number
CN202510330486.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing image recognition authentication technology has low recognition accuracy under complex lighting conditions or face occlusion, and may fail when initiating verification based on trusted channels, resulting in insufficient verification efficiency and accuracy.

Method used

By uploading multiple ID images from different angles by users, storing basic information using encrypted databases, combining OCR text recognition and image analysis to determine authenticity, conducting preliminary identity authentication, guiding live life detection, extracting high-quality video information for secondary identity authentication, and using black and white lists and problem verification to ensure that the user identity is consistent.

Benefits of technology

Improve the accuracy and efficiency of identity verification, prevent forgery documents from passing, distinguish between real users and non-organisms, and ensure the authenticity and security of user identities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260143A_ABST
    Figure CN120260143A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and discloses an identity verification method and system based on image recognition and a storage medium. The method comprises the steps that a user uploads an identity certificate to an authentication server through a mobile terminal, the identity certificate comprises a plurality of image sequences at different angles, the authentication server comprises an encryption database, and basic information of the user is stored in the encryption database; the authentication server carries out authenticity judgment on the identity certificate, and if the identity certificate is true, preliminary identity authentication is carried out on the user based on the basic information; if the preliminary authentication is passed, the mobile terminal guides a user to carry out living body detection based on a preset rule, whether the user is a living body is judged, and if yes, video information of the user in a living body detection time period is acquired, and a target image is extracted from the video information; and performing secondary identity authentication on the user based on the target image, and if the secondary identity authentication is passed, determining that the user identity authentication is successful, thereby improving the accuracy of identity authentication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to an identity authentication method, system and storage medium based on image recognition. Background Art

[0002] With the rapid development of information technology, the importance of identity authentication has become increasingly prominent in many fields such as finance, security, and intelligent devices. Image recognition technology has gradually become an important research direction in the field of identity authentication. By analyzing and processing image data, image recognition technology can quickly and accurately identify an individual's identity. Among them, face recognition is one of the most widely used image recognition technologies at present. It uses the unique features of the human face for identity authentication, and has the advantages of non-contact, convenience, and high efficiency. However, face recognition technology also faces some challenges. For example, in the case of complex lighting conditions, background interference, or face occlusion, the recognition accuracy may be affected.

[0003] For example, the existing Chinese patent application with the publication number CN112330331A proposes an identity authentication method, device, computer device and storage medium based on face recognition. The method includes: obtaining an identity card image to be verified and a first face image to be verified; extracting key information from the identity card image to be verified, where the key information includes at least one of user name, identity card number or face feature, and generating an identity card file according to the key information according to a preset template; retrieving a preset face reference image library according to the identity card file to obtain the face reference image of the identity card image to be verified; calculating the similarity between the first face image to be verified and the face reference image, and when the similarity is greater than or equal to a preset first threshold, passing the identity authentication. This application has low requirements for the identity card image and improves the efficiency of identity authentication. Although this existing technology reduces the requirements for image quality, it still requires key information in the identity card image. If the identity card image is blurred or the key information is missing, the verification may fail.

[0004] Another example is the existing Chinese patent application with the publication number CN113486317A, which proposes an identity authentication method. The method includes: responding to a user's identity authentication operation, obtaining the user's contact information and the face image of a business person; performing face recognition on the business person based on the face image to obtain an identity authentication result; and sending the identity authentication result to the user according to the contact information. This existing technology relies on trusted channels (such as official public accounts, applications, etc.) to initiate verification. If the user does not operate through these channels, the verification may not be possible.

[0005] Therefore, there is a need for an identity authentication method based on image recognition that can improve the accuracy of identity authentication while ensuring the image quality. Summary of the Invention

[0006] The present application provides an identity authentication method, system and storage medium based on image recognition, which is used to improve the accuracy in the identity authentication process.

[0007] In a first aspect, the present application provides an identity authentication method based on image recognition, and the method includes:

[0008] Step S1: The user uploads an identity certificate to the authentication server through a mobile terminal. The identity certificate includes a sequence of images at multiple different angles. The authentication server includes an encrypted database, and the encrypted database stores the basic information of the user.

[0009] Step S2: The authentication server determines the authenticity of the identity certificate. If it is true, a preliminary identity authentication is performed on the user based on the basic information.

[0010] Step S3: If the preliminary authentication is passed, the mobile terminal guides the user to perform a live detection based on a preset rule to determine whether the user is a living organism. If so, video information of the user during the live detection time period is obtained, and a target image is extracted from the video information. The target image is an image that meets the authentication requirements.

[0011] Step S4: A secondary identity authentication is performed on the user based on the target image. If the secondary identity authentication is passed, the user's identity authentication is successful.

[0012] Combined with the first aspect, in the first implementation manner of the first aspect of the present application, the authentication server determines the authenticity of the identity certificate, including:

[0013] Obtain a historical image sequence at different light source angles including both genuine and fake types, and perform type annotation on each group of historical image sequences. Perform Fourier transform on each image in the historical image sequence to obtain the spectral feature vector of each image. Analyze the displacement of pixels in the historical image sequence based on the optical flow algorithm to obtain the optical flow feature vector of each image. Establish a classification model based on a convolutional neural network, input the historical image sequence and the corresponding spectral feature vector and optical flow feature vector into the classification model for training, and input the identity certificate into the trained classification model. The classification model outputs the authenticity of the identity certificate.

[0014] Combined with the first aspect, in the second implementation manner of the first aspect of the present application, performing a preliminary identity authentication on the user based on the basic information includes:

[0015] Extract text information and image information from the image sequence, locate the face region in the image information, extract target features from the face region. The basic information includes the user's standard image and personal information. Extract face features from the standard image, and calculate the similarity between the target features and the face features. If the similarity is greater than a first threshold, then compare the text information with the personal information in the basic information to determine whether they are consistent. If they are consistent, it is determined that the user passes the preliminary identity authentication.

[0016] Combined with the first aspect, in the third implementation manner of the first aspect of the present application, guiding the user to perform a liveness detection based on a preset rule includes:

[0017] The liveness detection includes a first detection and a second detection. The mobile terminal randomly sets a first target pattern, and the first target pattern is a gaze guidance pattern for a person's gaze from a start point to an end point. Define the gaze movement trajectory from the start point to the end point as a first trajectory. Each first target pattern has multiple first trajectories corresponding to it. The user selects an arbitrary initial point from the first target pattern and moves their gaze to the end point based on the first target pattern to obtain a second trajectory. The mobile terminal records the time interval from the start point to the end point and calculates the similarity between the second trajectory and a preset first trajectory. If the similarity is greater than a second threshold and the time interval is less than a third threshold, it indicates that the first detection passes.

[0018] Combined with the first aspect, in the fourth implementation manner of the first aspect of the present application, the mobile terminal guiding the user to perform a liveness detection based on a preset rule includes:

[0019] After passing the first detection, the mobile terminal sets multiple second target patterns, and the second target patterns are facial expression guidance patterns. The user makes corresponding real-time expressions based on the second target patterns within a preset time interval. The mobile terminal calculates the matching degree between the real-time expressions and the second target patterns. If the matching degree is greater than a fourth threshold, then verify the next second target pattern until all the second target patterns are verified, indicating that the second detection passes. If both the first detection and the second detection pass, it is determined that the user is a living organism.

[0020] Combined with the first aspect, in the fifth implementation manner of the first aspect of the present application, extracting a target image from the video information includes:

[0021] The authentication server obtains multiple consecutive image frames from the video information, identifies the face regions in each image frame, selects multiple feature points in the face regions, tracks the positions of the face regions in consecutive frames based on the feature points, sets multiple quality metrics for each frame, where the quality metrics include face size, face confidence score, brightness variance of the face region, direction of the face, and the position of the face region in the image frame, sets a threshold range for each quality metric, determines whether the image frame meets the threshold ranges of all the quality metrics, and if so, defines the image frame as an excellent frame. If the number of excellent frames is greater than a fifth threshold, any one of the excellent frames is selected as the current user image with the highest quality. If the number of excellent frames is less than a sixth threshold, a second selection method is used to select the target image.

[0022] Combined with the first aspect, in the sixth implementation manner of the first aspect of the present application, using the second selection method to select the target image includes:

[0023] Calculate the similarity between the face features in the image frame and the features of the standard face image in the encrypted database, set a first score for the image frame based on the size of the similarity, analyze the movement vectors of each feature point in the image frame in consecutive image frames within a preset time period, set a second score for the image frame based on the size of the movement vector, set corresponding weight coefficients for the first score and the second score respectively, calculate the third score of each image frame based on the weight coefficients, and select the image frame with the highest third score as the target image.

[0024] Combined with the first aspect, in the seventh implementation manner of the first aspect of the present application, based on the target image, performing secondary identity authentication on the user includes:

[0025] Extract the face features from the target image and the image information of the identity document respectively, calculate the similarity. If the similarity is greater than a ninth threshold, obtain the blacklist and whitelist information from the encrypted database, query whether the user exists in the blacklist based on the text information in the identity document. If so, directly notify the user that the authentication fails, and record the authentication time of the user in the blacklist. If the user does not exist in the blacklist, query whether the user exists in the whitelist. If so, update the weight value of the user, where the weight value increases based on the number of times the user logs in, and users with a weight value higher than a tenth threshold can be exempted from verification in the whitelist and the blacklist. If the user is neither in the whitelist nor in the blacklist, conduct a question verification on the user. If the question verification passes, write the user into the whitelist.

[0026] In a second aspect, the present application provides an identity authentication system based on image recognition, the system comprising:

[0027] An upload module, configured to enable a user to upload an identity certificate to an authentication server through a mobile terminal, the identity certificate including a plurality of image sequences at different angles, and the authentication server including an encrypted database in which basic information of the user is stored;

[0028] A preliminary authentication module, configured to enable the authentication server to determine the authenticity of the identity certificate, and if it is genuine, perform a preliminary identity authentication on the user based on the basic information;

[0029] A live detection module, configured to, if the preliminary authentication is passed, enable the mobile terminal to guide the user to perform a live detection based on a preset rule, determine whether the user is a living organism, and if so, obtain video information of the user during the live detection period and extract a target image from the video information, the target image being an image that meets the authentication requirements;

[0030] A secondary authentication module, configured to perform a secondary identity authentication on the user based on the target image, and if the secondary identity authentication is passed, the user identity authentication is successful.

[0031] In a third aspect of the present application, there is provided a computer-readable storage medium storing instructions which, when run on a computer, cause the computer to execute the above-mentioned identity authentication method based on image recognition.

[0032] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0033] By enabling the user to upload a plurality of certificate images at different angles, more comprehensive certificate information can be provided. The authentication server stores the basic information of the user through the encrypted database, and can quickly compare the uploaded certificate information to preliminarily screen whether the user meets the authentication conditions; by technical means (such as OCR character recognition, image analysis, etc.) to determine the authenticity of the certificate, effectively identify forged or tampered certificates, and prevent lawbreakers from using false certificates to pass the authentication; if the certificate is determined to be forged, the authentication process can be directly terminated to avoid wasting resources in subsequent steps and improve the authentication efficiency;

[0034] The present invention performs live detection through line-of-sight guidance and expression guidance, effectively differentiating real users from non-living objects such as photos and videos, and preventing lawbreakers from impersonating users through static images or recorded videos. By extracting target images from the video information during the live detection process, biometric data of higher quality and more in line with the authentication requirements (such as face images) can be obtained, providing a more reliable basis for subsequent identity authentication. Based on the target image, secondary identity authentication is performed, and the blacklist and whitelist are used to further verify the true identity of the user, ensuring that the user is consistent with the registered information and completing the final identity authentication. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 Schematic diagram of an embodiment of the identity authentication method based on image recognition in an embodiment of the present application;

[0037] Figure 2 Schematic diagram of the first trajectory of the first target pattern in an embodiment of the present application;

[0038] Figure 3 Threshold range diagram for each quality index set in an embodiment of the present application;

[0039] Figure 4 Schematic diagram of an embodiment of the identity authentication system based on image recognition in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The embodiments of the present application provide an identity authentication method, system and storage medium based on image recognition. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0041] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer toFigure 1 , one embodiment of the identity authentication method based on image recognition in the embodiments of the present application includes:

[0042] Step S1: The user uploads an identity document to the authentication server through a mobile terminal. The identity document includes a sequence of images at multiple different angles. The authentication server includes an encrypted database, and the basic information of the user is stored in the encrypted database.

[0043] Specifically, the user uses the camera of the mobile terminal (mobile phone) to capture a sequence of images of the identity document (such as ID card, passport, etc.) at different light source angles. For example: front light source - capture the front image, left light - capture the left image, right light source - capture the right image, etc. Encrypt the captured sequence of images, and upload the encrypted data to the encrypted database in the authentication server through a secure communication protocol (such as HTTPS, SSL / TLS). The basic information of the user is also stored in the encrypted database, and the basic information includes personal information, face feature information, and other information. Personal information such as name, ID number, contact information, etc., and other information such as account information, transaction records, login logs, etc.

[0044] Step S2: The authentication server determines the authenticity of the identity document. If it is true, a preliminary identity authentication of the user is performed based on the basic information.

[0045] Specifically, the authentication server determines the authenticity of the sequence of images of the identity document uploaded by the user, and detects whether there are signs of tampering in the images, such as splicing, blurring, pixelation, etc. If it is determined that the identity document is forged, the authentication server directly sends the discrimination result to the mobile terminal, and the mobile terminal directly notifies the user that the authentication fails and there is no need to continue the identity authentication. If the discrimination result is true, the preliminary identity authentication of the user is continued. The preliminary identity authentication is to query in the encrypted database whether there is a user record that matches the face feature information and personal information to be authenticated. If so, the user can pass the preliminary authentication.

[0046] Step S3: If the preliminary authentication is passed, the mobile terminal guides the user to perform a live detection based on a preset rule to determine whether the user is a living organism. If so, the video information of the user during the live detection period is obtained, and a target image is extracted from the video information. The target image is an image that meets the authentication requirements.

[0047] Specifically, after the preliminary identity authentication is passed, in order to prevent non-living attacks such as photo, video playback, and mask, the system will guide the user to perform a liveness detection to confirm whether the user is a real living organism. The mobile terminal will switch to the liveness detection interface and provide clear instructions, informing the user that they need to complete the liveness detection to continue the authentication process. The liveness detection in the present invention mainly includes a first detection and a second detection. The first detection is gaze detection, and the second detection is expression detection. Only by passing both detections can it be determined as a real living organism.

[0048] After passing the liveness detection, the mobile terminal analyzes the video information during the liveness detection, extracts the target image that meets the authentication requirements. Compared with the directly captured real-time image, the quality of the target image obtained from the video information is better, has more facial features, and also saves more shooting time.

[0049] Step S4: Perform secondary identity authentication on the user based on the target image. If the secondary identity authentication is passed, the user identity authentication is successful.

[0050] Specifically, when the liveness test is passed, the authentication server needs to match the obtained target image with the personal image in the basic information, and also needs to further determine whether the user exists in the blacklist or whitelist of the encrypted database to perform a more accurate secondary authentication on the user.

[0051] In a specific embodiment, the steps for the authentication server to determine the authenticity of the identity certificate are as follows:

[0052] Obtain a historical image sequence at different light source angles including both genuine and fake types, and perform type annotation on each group of historical image sequences. Perform Fourier transform on each image in the historical image sequence to obtain the spectral feature vector of each image. Analyze the displacement of pixels in the historical image sequence based on the optical flow algorithm to obtain the optical flow feature vector of each image. Establish a classification model based on a convolutional neural network. Input the historical image sequence and the corresponding spectral feature vector and optical flow feature vector into the classification model for training, and input the identity certificate into the trained classification model. The classification model outputs the authenticity of the identity certificate.

[0053] Specifically, when conducting online banking or financial business processing, users may use fake identity documents for verification. Therefore, it is necessary to determine the authenticity of identity documents, collect image sequences of identity documents including both genuine and fake types, obtain the image sequence of real documents from a legal identity document database, and collect image sequences of forged documents made by professionals, fake documents generated using digital synthesis technology, and highly realistic forged documents generated using deep learning technologies such as generative adversarial networks (GANs). Label each group of historical image sequences, that is, mark them as "genuine" or "fake", adjust all images to the same resolution and format, perform a two-dimensional discrete Fourier transform (DFT) on each image in the historical image sequence, and convert the image from the spatial domain to the frequency domain. The conversion formula is expressed as:

[0054] where f(x, y) is the pixel value of the image at coordinates (x, y), F(u, v) is the spectrum value of the image at frequency coordinates (u, v), M and N are the total number of pixels of the image in the horizontal and vertical directions respectively, and e -j2π(ux / M+vy / N) represents the complex exponential function.

[0055] Analyze the pixel displacement between adjacent frames in the historical image sequence, calculate the optical flow field. The optical flow represents the velocity vector of pixels in the image. For example, use the Lucas-Kanade method or the Gunnar Farneback method for calculation, extract features from the optical flow field, and combine the extracted features into an optical flow feature vector. Establish a classification model based on a convolutional neural network, input the image sequence, spectral features, and optical flow features into different convolutional layers for processing respectively. After passing through the pooling layer and the fully connected layer, use the softmax function to output the classification result, such as the probability distribution of "genuine" or "fake".

[0056] Input the identity document uploaded by the user into the trained classification model. The classification model outputs the probability distribution of the identity document being "genuine" or "fake". If the probability of being "genuine" is greater than the preset threshold, it is determined that the document is a real document. If the probability of being "fake" is greater than the preset threshold, it is determined that the document is a forged document.

[0057] In a specific embodiment, the preliminary identity authentication of the user based on basic information includes the following steps:

[0058] Extract text information and image information from the image sequence, locate the face region in the image information, extract target features from the face region. The basic information includes the user's standard image and personal information. Extract face features from the standard image, and calculate the similarity between the target features and the face features. If the similarity is greater than the first threshold, compare the text information with the personal information in the basic information to determine whether they are consistent. If they are consistent, it is determined that the user passes the preliminary identity authentication.

[0059] Specifically, the authentication server receives the image sequence of the user's identity certificate uploaded by the user, uses OCR technology (such as Tesseract, Baidu OCR) to identify the key text information in the image, such as name, ID number, etc., uses a face detection algorithm (such as Haar feature cascade classifier, deep learning model) to locate the face region in the image, crops the face image from the face region, and extracts the target features, and compares them with the standard face features stored in the database, and calculates the similarity of the two features based on Formula 1

[0060] α1, Formula 1 is: Where A and B are the target feature vector and the standard face feature vector respectively, and S is the covariance matrix, which is used to measure the correlation between features, and S -1 is the inverse matrix of the covariance matrix. If the similarity ranges from 0 to 1, the first threshold can be set to 0.9. If the similarity is greater than the preset first threshold, it means that the face features in the image sequence are similar to the face features in the encrypted database, and then the consistency of the basic information is judged.

[0061] Compare the extracted text information item by item with the personal information in the database. During the comparison process, set a certain error tolerance rate, for example, allow OCR recognition errors or user input errors, but the consistency must be as high as 95% to authenticate that the text information is consistent. If both the text information and the image information pass the verification, it means that the user passes the preliminary verification.

[0062] In a specific embodiment, guiding the user to perform a live detection based on a preset rule includes the following steps:

[0063] Liveness detection includes a first detection and a second detection. The mobile terminal randomly sets a first target pattern, where the first target pattern is a line-of-sight guiding pattern for a person's line of sight from a starting point to an ending point. The line-of-sight movement trajectory from the starting point to the ending point is defined as the first trajectory. There are multiple first trajectories corresponding to each first target pattern. The user selects any initial point from the first target pattern and moves their own line of sight to the ending point based on the first target pattern to obtain a second trajectory. The mobile terminal records the time interval from the starting point to the ending point and calculates the similarity between the second trajectory and a preset first trajectory. If the similarity is greater than a second threshold and the time interval is less than a third threshold, it indicates that the first detection passes.

[0064] Specifically, in order to improve the accuracy of identity verification, it is necessary to perform liveness detection on the user. This liveness detection method includes a first detection and a second detection.

[0065] The first detection is to detect the user's line of sight. The process is as follows: The mobile terminal randomly generates a first target pattern. As Figure 2 shown, it is a schematic diagram of the first trajectory of the first target pattern. This schematic diagram is a pattern in the shape of the letter W. It can also be other letters, animals, Chinese characters, etc. This pattern includes a starting point A: the starting point of the user's line-of-sight movement, an ending point B: the ending point of the user's line-of-sight movement, and a line-of-sight guiding pattern: a guiding line connecting the starting point A and the ending point B, such as a straight line, a curve, or a broken line, etc. The user selects the starting point according to the guidance of the first target pattern, clicks to start, and the mobile terminal starts timing. The user moves their line of sight from the starting point to the ending point. The mobile terminal uses devices such as a front camera or an infrared sensor to capture the user's eye movement trajectory in real time, converts the captured eye movement data into a line-of-sight trajectory (the second trajectory), compares the second trajectory generated by the user with the preset first trajectory, calculates the similarity, compares the calculated similarity with the preset second threshold, and if the time interval from the starting point to the ending point of the line of sight is less than the third threshold, it is determined that the first detection passes. The present invention combines the line-of-sight movement trajectory and time interval information for liveness detection, increasing the difficulty of attackers' forgery and improving the security of identity authentication.

[0066] In a specific embodiment, the mobile terminal guiding the user to perform liveness detection based on a preset rule specifically further includes the following steps:

[0067] After passing the first detection, the mobile terminal sets multiple second target patterns. The second target patterns are facial expression guiding patterns. The user makes corresponding real-time expressions based on the second target patterns within a preset time interval. The mobile terminal calculates the matching degree between the real-time expression and the second target pattern. If the matching degree is greater than a fourth threshold, the verification of the next second target pattern is carried out until all the second target patterns are verified, indicating that the second detection passes. If both the first detection and the second detection pass, it is determined that the user is a living organism.

[0068] Specifically, after the first detection (detection based on the line-of-sight guidance pattern), the system will guide the user to perform the second detection, that is, the detection based on the facial expression guidance pattern. The process is as follows: The mobile terminal randomly generates multiple second target patterns, each pattern corresponding to a facial expression guidance pattern, such as smiling, blinking, frowning, being surprised, etc. Text information can be set below for prompt, such as text prompts "smile", "blink", etc., to help the user understand the expression to be made. According to the guidance of each second target pattern, the user makes corresponding real-time expressions within a preset time interval. The mobile terminal captures the user's expression changes in real time through the front camera, compares the captured user expression with the currently presented second target pattern, can use a pre-trained expression recognition model to classify the user's expression, and output the probability distribution of each expression category. Compare the probability of the target expression category with a preset threshold. If it is greater than the preset threshold, it is considered that the matching is successful. The system also sets a preset time interval, such as 3 - 5 seconds, to limit the time for the user to make an expression. If the user fails to make the correct expression within the specified time, the second detection fails. Only when the corresponding action is made within the preset time interval and the matching degree is greater than the fourth threshold, does it indicate that the second detection is successful.

[0069] By combining the line-of-sight guidance pattern and the facial expression guidance pattern for liveness detection and using different types of biometric features for verification, the accuracy of the detection is improved, and the security of identity authentication is enhanced.

[0070] In a specific embodiment, extracting the target image from the video information specifically includes the following steps:

[0071] The authentication server obtains multiple consecutive image frames from the video information, identifies the face region in each image frame, selects multiple feature points in the face region, and tracks the position of the face region in the consecutive frames based on the feature points. The authentication server sets multiple quality metrics for each frame. The quality metrics include face size, face confidence score, brightness variance of the face region, face orientation, and the position of the face region in the image frame. Threshold ranges are set for each quality metric, and it is determined whether the image frame meets the threshold ranges of all quality metrics. If so, the image frame is defined as an excellent frame. If the number of excellent frames is greater than the fifth threshold, any one of the excellent frames is selected as the current user image with the highest quality. If the number of excellent frames is less than the sixth threshold, a second selection method is used to select the target image.

[0072] Specifically, the authentication server receives the video information generated by the mobile terminal. The video contains the process of the user completing actions (such as eye movement, facial expression change) according to the requirements of the live detection. Multiple consecutive image frames are extracted from the video information. For example, 10 frames are extracted per second (according to the video frame rate). The face detection algorithm is applied to each image frame to identify the face region in the image. Multiple feature points are selected in the detected face region, such as the positions of key parts like eyes, nose, mouth, eyebrows, etc. Based on the tracking algorithm of deep learning (such as Siamese network), the movement trajectories of the feature points are tracked, and the position information of the face region in each frame is generated, such as the coordinates of the face bounding box.

[0073] Multiple quality metrics are set for each image frame to evaluate the quality of the image frame. The quality metrics include: face size (expressed in terms of the number of pixels or area, for example), face confidence score (the confidence score given by the face detection algorithm, indicating the credibility of the detected face), brightness variance (used to measure the uniformity of the brightness distribution), face orientation (such as front, side, tilt angle, etc.), and face position (the position of the face region in the image frame, such as the distance from the center, used to evaluate the image quality). As Figure 3 shown in the threshold range graph set for each quality metric, if the number of excellent frames is greater than a preset fifth threshold (such as 3 frames), then the image with the highest quality is selected from the excellent frames as the target image. If the number of excellent frames is less than a preset sixth threshold (such as 1 frame), then the second selection method is used for the selection of the target image, which will be elaborated later. By extracting the target image from the video information, the system can obtain richer facial information, improving the accuracy and reliability of the authentication. At the same time, this step realizes automated processing without manual intervention, improving the efficiency.

[0074] In a specific embodiment, the second selection method for selecting the target image specifically includes the following steps:

[0075] Calculate the similarity between the face features in the image frame and the features of the standard face image in the encrypted database, set the first score for the image frame based on the size of the similarity, analyze the movement vectors of each feature point in the continuous image frames within a preset time period, set the second score for the image frame based on the size of the movement vector, set the corresponding weight coefficients for the first score and the second score respectively, calculate the third score of each image frame based on the weight coefficients, and select the image frame with the highest third score as the target image.

[0076] Specifically, face detection is performed on each image frame in the video, and face features are extracted, such as: facial feature points: the positions and shapes of key parts such as eyes, nose, mouth, etc., and features related to facial expressions are extracted, such as using an expression recognition model to extract features. The feature information of the standard face image corresponding to the user is obtained from the encrypted database. The standard face image is an image taken of the user in a normal state (with natural expression and natural body). The similarity between the face features of each image frame and the features of the standard face image is calculated using a feature matching algorithm (such as cosine similarity, Euclidean distance). The first score mainly evaluates the naturalness of the expression. The higher the naturalness of the expression, the higher the similarity and the higher the first score. For example, a real facial expression will be more similar to the features of the standard face image, while a deliberately made or forged expression will have a lower similarity.

[0077] The change in the position of feature points between consecutive image frames is analyzed, and the motion vector of each feature point is calculated. Each feature point has a motion vector representing its displacement between consecutive frames. The second score mainly evaluates the stability of the image. The smaller the motion vector, the more stable the image and the higher the second score.

[0078] Weight coefficients are set for the first score and the second score respectively. For example: the weight coefficient of the first score (w1): 0.6, the weight coefficient of the second score (w2): 0.4. The weighted sum formula is used to calculate the third score of each image frame. The formula is: Third score = (w1 * First score) + (w2 * Second score). The image frame with the highest third score is selected as the target image. By combining face feature similarity and feature point motion vector analysis, this step realizes the multi-dimensional evaluation of image frames and can effectively select the image frames with natural expressions and stable images as the target images, thereby improving the accuracy and reliability of identity authentication.

[0079] In a specific embodiment, the secondary identity authentication of the user based on the target image includes the following steps:

[0080] Face features are extracted from the target image and the image information of the identity certificate respectively, and the similarity is calculated. If the similarity is greater than the ninth threshold, the blacklist and whitelist information is obtained from the encrypted database. Based on the text information in the identity certificate, it is queried whether the user exists in the blacklist. If so, the user is directly notified that the authentication fails, and the authentication time of the user is recorded in the blacklist. If the user does not exist in the blacklist, it is queried whether the user exists in the whitelist. If so, the weight value of the user is updated. The weight value increases based on the number of times the user logs in, and users with a weight value higher than the tenth threshold can be exempted from the verification of the whitelist and the blacklist. If the user is neither in the whitelist nor in the blacklist, the user is subjected to a question verification. If the question verification passes, the user is written into the whitelist.

[0081] Specifically, using the same method as the preliminary identity authentication, facial features are extracted from the target image and compared with the facial features extracted from the image information in the identity document to calculate the similarity. Based on the text information in the identity document (such as name, ID number), it is queried whether the user exists in the blacklist, which records the information of users prohibited from access or with security risks. If the user exists, the user is directly notified that the authentication fails, and the authentication time of the user is recorded in the blacklist for subsequent analysis and tracking. If the user is not in the blacklist, it is queried whether the user exists in the white list, which records the information of authorized or trusted users. If the user exists in the white list, the weight value of the user is updated based on the number of times the user logs in. For example, each time the user successfully logs in, the weight value increases by a certain value. When the weight value of the user is higher than the tenth threshold, the user can be exempt from the verification of the white list and the blacklist and directly pass the authentication, saving the authentication time.

[0082] If the user is not in the white list, the question verification step is entered. For example, the user is asked preset personal information questions, such as "Where is your place of birth?" and "Where was the location of your last login?", and a time limit can be set for each question. If the user correctly answers all questions within the specified time, the user is written into the white list and given an initial weight value. By combining facial feature similarity, black and white list verification, weight mechanism, and question verification, the comprehensive verification of the user's identity is achieved, improving the security and accuracy of authentication.

[0083] The above describes a method for identity authentication based on image recognition in an embodiment of the present application. Next, the identity authentication system based on image recognition in the embodiment of the present application will be described. Please refer to Figure 4 , an embodiment of an identity authentication system based on image recognition in the embodiment of the present application includes:

[0084] An upload module for the user to upload an identity document to the authentication server through a mobile terminal. The identity document includes an image sequence at multiple different angles. The authentication server includes an encrypted database that stores the basic information of the user;

[0085] A preliminary authentication module for the authentication server to determine the authenticity of the identity document. If it is true, a preliminary identity authentication is performed on the user based on the basic information;

[0086] A live detection module. If the preliminary authentication passes, the mobile terminal guides the user to perform a live detection based on a preset rule to determine whether the user is a living organism. If so, the video information of the user during the live detection period is obtained, and a target image is extracted from the video information. The target image is an image that meets the authentication requirements;

[0087] The secondary authentication module is used to perform secondary identity authentication on the user based on the target image. If the secondary identity authentication is passed, the user identity authentication is successful.

[0088] This application also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the identity authentication method based on image recognition.

[0089] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0090] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0091] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application.

Claims

1. An image recognition-based authentication method, characterized in that, The method includes the following: Step S1: The user uploads an identity certificate to an authentication server through a mobile terminal. The identity certificate includes an image sequence at multiple different angles. The authentication server includes an encrypted database, and basic information of the user is stored in the encrypted database. Step S2: The authentication server determines the authenticity of the identity certificate. If it is genuine, a preliminary identity authentication of the user is performed based on the basic information. Step S3: If the preliminary authentication passes, the mobile terminal guides the user to perform a live detection based on a preset rule to determine whether the user is a living organism. If so, video information of the user during the live detection period is obtained, and a target image is extracted from the video information. The target image is an image that meets the authentication requirements. Step S4: A secondary identity authentication of the user is performed based on the target image. If the secondary identity authentication passes, the user's identity authentication is successful.

2. The method according to claim 1, wherein The authentication server's determination of the authenticity of the identity certificate includes the following steps: Obtain historical image sequences at different light source angles including both genuine and fake types, perform type annotation on each group of historical image sequences, perform Fourier transform on each image in the historical image sequences to obtain the spectral feature vector of each image, analyze the displacement of pixels in the historical image sequences based on the optical flow algorithm to obtain the optical flow feature vector of each image, establish a classification model based on a convolutional neural network, input the historical image sequences and the corresponding spectral feature vectors and optical flow feature vectors into the classification model for training, and input the identity certificate into the trained classification model. The classification model outputs the authenticity of the identity certificate.

3. The method according to claim 2, wherein Performing a preliminary identity authentication of the user based on the basic information includes the following steps: Extract text information and image information from the image sequence, locate the face region in the image information, extract target features from the face region. The basic information includes the standard image and personal information of the user. Extract face features from the standard image, and calculate the similarity between the target features and the face features. If the similarity is greater than a first threshold, compare the text information with the personal information in the basic information to determine whether they are consistent. If they are consistent, it is determined that the user passes the preliminary identity authentication.

4. The method according to claim 1, wherein Guiding the user to perform a live detection based on a preset rule includes the following steps: The living body detection includes a first detection and a second detection. The mobile terminal randomly sets a first target pattern, which is a line-of-sight guiding pattern for a person's line of sight from a starting point to an ending point. The line-of-sight movement trajectory from the starting point to the ending point is defined as a first trajectory. There are multiple first trajectories corresponding to each first target pattern. The user selects an arbitrary initial point from the first target pattern and moves their line of sight to the ending point based on the first target pattern to obtain a second trajectory. The mobile terminal records the time interval from the starting point to the ending point and calculates the similarity between the second trajectory and a preset first trajectory. If the similarity is greater than a second threshold and the time interval is less than a third threshold, it indicates that the first detection is passed.

5. The method according to claim 4, characterized in that, The steps for the mobile terminal to guide the user to perform living body detection based on a preset rule further include the following: After passing the first detection, the mobile terminal sets multiple second target patterns, which are facial expression guiding patterns. The user makes corresponding real-time expressions based on the second target patterns within a preset time interval. The mobile terminal calculates the matching degree between the real-time expression and the second target pattern. If the matching degree is greater than a fourth threshold, the verification of the next second target pattern is carried out until all the second target patterns are verified, indicating that the second detection is passed. If both the first detection and the second detection are passed, it is determined that the user is a living body.

6. The method according to claim 1, characterized in that, The steps for extracting a target image from the video information include the following: The authentication server obtains multiple consecutive image frames from the video information and identifies the face regions in each image frame. Multiple feature points are selected in the face regions, and the positions of the face regions in consecutive frames are tracked based on the feature points. The authentication server sets multiple quality metrics for each frame. The quality metrics include face size, face credibility score, brightness variance of the face region, face direction, and the position of the face region in the image frame. Threshold ranges are set for each quality metric, and it is determined whether the image frame meets the threshold ranges of all the quality metrics. If so, the image frame is defined as an excellent frame. If the number of excellent frames is greater than a fifth threshold, any one of the excellent frames is selected as the current user image with the highest quality. If the number of excellent frames is less than a sixth threshold, a second selection method is used to select the target image.

7. The method according to claim 6, wherein The steps for using the second selection method to select the target image include the following: Calculate the similarity between the face features in the image frame and the features of the standard face image in the encrypted database, set a first score for the image frame based on the size of the similarity, analyze the movement vectors of each feature point in the image frame in consecutive image frames within a preset time period, set a second score for the image frame based on the size of the movement vector, set corresponding weight coefficients for the first score and the second score respectively, calculate the third score of each image frame based on the weight coefficients, and select the image frame with the highest third score as the target image.

8. The method according to claim 7, wherein Performing secondary identity authentication on the user based on the target image includes the following steps: Extract face features from the target image and the image information of the identity document respectively, and calculate the similarity. If the similarity is greater than the ninth threshold, obtain the blacklist and whitelist information from the encrypted database, and query whether the user exists in the blacklist based on the text information in the identity document. If so, directly notify the user that the authentication fails, and record the authentication time of the user in the blacklist. If the user does not exist in the blacklist, query whether the user exists in the whitelist. If so, update the weight value for the user. The weight value is increased based on the number of times the user logs in, and users with a weight value higher than the tenth threshold can be exempt from verification in the whitelist and the blacklist. If the user is neither in the whitelist nor in the blacklist, perform question verification on the user. If the question verification passes, write the user into the whitelist.

9. An image recognition-based authentication system for implementing the method according to any one of claims 1-8, characterized in that, The system includes: An upload module for the user to upload an identity document to the authentication server through a mobile terminal. The identity document includes an image sequence at multiple different angles. The authentication server includes an encrypted database that stores basic information of users; A preliminary authentication module for the authentication server to determine the authenticity of the identity document. If it is true, perform preliminary identity authentication on the user based on the basic information; A live detection module for, if the preliminary authentication passes, the mobile terminal guiding the user to perform live detection based on a preset rule, determining whether the user is a living organism. If so, obtain the video information of the user during the live detection period, and extract a target image from the video information. The target image is an image that meets the authentication requirements; A secondary authentication module for performing secondary identity authentication on the user based on the target image. If the secondary identity authentication passes, the user's identity authentication is successful.

10. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instruction is executed by a processor, it implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Identity verification method and device based on face recognition, equipment and storage medium

    CN112330331A

  • Identity verification method, identity verification device, electronic equipment and storage medium

    CN113486317A

Cited By

  • Risk control identity authentication system and method based on facial expression recognition

    CN120823637A