An identity authentication method
By setting open trigger conditions and verification pass conditions for secondary verification, and combining user audio and verification photo binding, the problems of reduced security and poor user experience in existing technologies are solved, achieving higher personalization and security, and enhancing the applicability of the authentication method.
Patent Information
- Application Number
- CN202411371477.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2044-09-29
AI Technical Summary
Existing identity authentication methods have become less secure and less personalized due to the development of artificial intelligence, and frequent verification has led to a poor user experience.
By setting open trigger conditions and verification pass conditions for secondary verification, combined with user audio and verification photo binding, personalized settings can be made, and security and versatility can be improved through risk control detection and risk standard updates.
It achieves higher user personalization verification, improves security, reduces the negative impact of frequent verification on user experience, and enhances the applicability of the authentication method.
Smart Images

Figure CN119337348B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of authentication technology, and in particular to an identity authentication method. Background Technology
[0002] Authentication, or identity verification, is the process of verifying whether a user has the authority to access or manipulate data from a server. When sending a request, it must typically include appropriate verification parameters to ensure the request has the necessary access permissions and returns the required data. For example, opening a website requires a username and password to log in, and some services require login before operation. To mitigate the risks of malicious attacks on servers and the leakage of sensitive personal information, major platforms have implemented two-factor authentication mechanisms. Common two-factor authentication technologies include: SMS verification codes, email verification codes, identity verification applications, hardware security keys, biometrics, push notifications, security questions, and device identification. To prevent web crawlers and other bots from accessing servers, platforms commonly employ methods such as text verification codes, image verification codes, audio verification codes, sliding verification codes, behavioral analysis, time analysis, invisible verification codes, question-and-answer verification codes, and interactive verification codes.
[0003] However, with the development of artificial intelligence, the above technologies are becoming increasingly easy to break, leading to reduced security. In addition, existing verification methods are basically automatically generated by the system, resulting in low personalization for users. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide an identity authentication method that offers greater personalization for user authentication by allowing users to set open secondary verification trigger conditions and verification pass conditions; it also provides a novel verification method by collecting user audio and verification photos and setting verification pairs to improve the security and specificity of verification; and it reduces the defects of poor user experience caused by frequent verification triggers by risk control detection and risk standard updates, thereby improving the versatility of the authentication method.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] An authentication method includes:
[0007] Register a platform account to obtain a personal account;
[0008] Set risk standards;
[0009] Based on the aforementioned personal account, set the conditions for triggering secondary verification and the conditions for passing verification;
[0010] Several sets of one-to-one corresponding user audio and verification photos are entered, and the corresponding user audio and verification photos are bound together to obtain several sets of verification pairs;
[0011] When the secondary verification trigger condition is triggered, several segments of the user's audio, a preset number of the verification photos, and several randomly generated interference images are randomly sent to the client to be detected.
[0012] Collect the matching results of the client to be detected, analyze the matching results and the verification conditions to obtain the verification result;
[0013] When the verification result is successful, the client to be tested is allowed to access the server;
[0014] Once the client to be tested logs into the server, the client's behavior is monitored for risk control to obtain risk control data.
[0015] When the risk control data exceeds the risk standard, the client to be tested is repeatedly verified by forcibly triggering the secondary verification trigger condition to obtain the risk control identification result.
[0016] When the risk control identification result is a risk-free issue, the risk standard is increased.
[0017] Preferably, the formula for calculating the risk standard is:
[0018]
[0019] Where R is the risk standard; R0 is the basic risk value; and n is the number of times the repeated verification is successful.
[0020] Preferably, the secondary verification triggering conditions include: logging in with an uncommon address, changing the login device, frequent login attempts, accessing private files, and frequently deleting files.
[0021] Preferably, the verification pass conditions include: matching accuracy, number of matching errors, and matching time.
[0022] Preferably, several sets of one-to-one corresponding user audio and verification photos are entered, and the corresponding user audio and verification photos are bound together to obtain several verification pairs, including:
[0023] Collect uploaded audio and photos from target users;
[0024] The uploaded audio is denoised and compressed to obtain the user's audio;
[0025] The uploaded photo is subjected to noise reduction and resolution normalization to obtain the verification photo;
[0026] According to the matching scheme for the target user, the corresponding user audio and the verification photo are packaged and stored to obtain the verification pair.
[0027] Preferably, when the secondary verification trigger condition is triggered, several segments of the user's audio, a preset number of verification photos, and several randomly generated interference images are randomly sent to the client to be detected, including:
[0028] When the secondary verification trigger condition is triggered, several sets of verification pairs are randomly selected;
[0029] Randomly select several segments of the user audio and a preset number of the verification photos from the selected verification pairs;
[0030] Based on the difference between the selected user audio and the verification photo, several interference images are randomly generated;
[0031] The selected user audio, the verification photo, and the interference image are sent to the client to be detected.
[0032] Preferably, the matching results of the client to be detected are collected, and the matching results and the verification conditions are analyzed to obtain the verification results, including:
[0033] Receive feedback information sent by the client to be tested;
[0034] The feedback information is parsed to obtain the matching result;
[0035] The matching results are judged using the verification conditions to obtain a judgment result;
[0036] When all the judgment results are in compliance, the verification result is determined to be successful.
[0037] If any of the judgment results is inconsistent, the verification result is determined to be a verification failure.
[0038] Preferably, after the client to be tested logs into the server, risk control detection is performed on the behavior of the client to be tested to obtain risk control data, including:
[0039] Once the client to be tested logs into the server, its behavior information is recorded.
[0040] Based on a preset set of risk behavior scores, the total value of the behavior information is calculated to obtain the risk control data.
[0041] The present invention discloses the following technical effects:
[0042] This invention provides an identity authentication method that addresses the shortcomings of existing methods by allowing users to customize their authentication settings, such as fixed trigger conditions and pass conditions. It also overcomes the vulnerability of traditional authentication methods to AI by collecting user audio and verification photos and setting verification pairs, thus achieving a more secure authentication function. Furthermore, it addresses the issue of frequent triggering of authentication mechanisms in existing methods by implementing risk control detection and risk standard updates, allowing for updates to risk standards based on the number of successful authentications. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a schematic diagram of the identity authentication process provided in an embodiment of the present invention;
[0045] Figure 2 An identity authentication flowchart provided for an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of the verification pair acquisition process provided in an embodiment of the present invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] The purpose of this invention is to provide an identity authentication method that offers greater personalization for user authentication by allowing users to set open secondary verification trigger conditions and verification pass conditions; it also provides a novel verification method by collecting user audio and verification photos and setting verification pairs to improve the security and specificity of verification; and it reduces the shortcomings of frequent verification triggers that lead to poor user experience by updating risk control detection and risk standards, thereby improving the versatility of the authentication method.
[0049] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0050] Figure 1This is a schematic diagram of the identity authentication process provided in an embodiment of the present invention. Figure 2 The identity authentication flowchart provided in the embodiments of the present invention is as follows: Figure 1 and Figure 2 As shown, the present invention provides an identity authentication method, including:
[0051] Step 100: Register a platform account and obtain your personal account;
[0052] Step 200: Set risk criteria;
[0053] Step 300: Based on the personal account, set the secondary verification trigger conditions and verification pass conditions;
[0054] Step 400: Input several sets of one-to-one corresponding user audio and verification photos, and bind the corresponding user audio and verification photos to obtain several sets of verification pairs;
[0055] Step 500: When the secondary verification trigger condition is triggered, randomly send several segments of the user audio, a preset number of the verification photos, and several randomly generated interference images to the client to be detected;
[0056] Step 600: Collect the matching results of the client to be detected, analyze the matching results and the verification conditions to obtain the verification results;
[0057] Step 700: When the verification result is successful, allow the client to be tested to access the server;
[0058] Step 800: After the client to be tested logs into the server, risk control detection is performed on the behavior of the client to be tested to obtain risk control data;
[0059] Step 900: When the risk control data exceeds the risk standard, the client to be tested is re-verified by forcibly triggering the secondary verification trigger condition to obtain the risk control identification result;
[0060] Step 1000: When the risk control identification result is a risk-free problem, increase the risk standard.
[0061] Specifically, the formula for calculating the risk standard is as follows:
[0062]
[0063] Where R is the risk standard; R0 is the basic risk value; and n is the number of times the repeated verification is successful.
[0064] Optionally, the secondary verification triggering conditions include: logging in with an uncommon address, changing the login device, frequent login attempts, accessing private files, and frequently deleting files.
[0065] Preferably, the verification pass conditions include: matching accuracy, number of matching errors, and matching time.
[0066] refer to Figure 3 Several sets of one-to-one corresponding user audio and verification photos are entered, and the corresponding user audio and verification photos are bound together to obtain several verification pairs, including:
[0067] Step 401: Collect the target user's uploaded audio and uploaded photos;
[0068] Step 402: Perform noise reduction and compression on the uploaded audio to obtain the user audio;
[0069] Step 403: Perform noise reduction and resolution normalization on the uploaded photo to obtain the verification photo;
[0070] Step 404: Based on the matching scheme for the target user, package and store the corresponding user audio and the verification photo to obtain the verification pair.
[0071] Preferably, to ensure the normal operation and speed of the system authentication function, the bitrate of the uploaded audio must not be lower than 192kbps, and the resolution of the uploaded photo must not be lower than 720p; when the bitrate of the uploaded audio exceeds 192kbps, the system will convert it to 192kbps format; when the uploaded photo exceeds 1080p, the system will convert it to 1080p format.
[0072] Furthermore, common denoising techniques include: DnCNN, FFDNet, CBDNet, RIDNet, PMRID, and SID. The main characteristics of the DnCNN network are: the network is divided into three parts; the first part is Conv+ReLU (one layer); the second part is Conv+BN+ReLU (several layers); and the third part is Conv (one layer), with 17 or 20 layers. The network learns the image residual, that is, the difference between the noisy image and the noise-free image, and the loss function used is MSE. The main characteristics of the FFDNet network are: its structure is consistent with DnCNN, but the input and output are different. The network input consists of four sub-images obtained by downsampling the original noisy image and a noise level image generated by the user-input parameter σ. The output consists of four denoised sub-images, which are then upsampled to obtain the final denoised image. The loss function used is still MSE; because the input of this network includes a user-controlled parameter, this algorithm is better at adapting to different noise levels than DnCNN. The main features of the CBDNet network are as follows: The network structure consists of two parts: the first part is a five-layer fully convolutional network used for noise level estimation; the second part, unlike FFDNet, is an UnNet with residuals used for noise reduction. An asymmetric loss function is designed primarily to eliminate asymmetric sensitivity. Asymmetric sensitivity refers to the fact that noise reduction algorithms like BM3D and FFDNet perform poorly when the input noise reduction parameter σ is small, while achieving better noise reduction results even with some texture loss when the input noise reduction parameter σ is large. Based on the above description, this embodiment selects CBDNet technology as the method for noise reduction of the uploaded photos.
[0073] Preferably, the method for denoising the uploaded audio in this embodiment is DPCRN_DNS3. DPCRN_DNS3 uses the TensorFlow framework and incorporates libraries such as numpy, matplotlib, librosa, and soxfile to build a powerful and flexible model. It primarily relies on DPCRN, an innovative neural network architecture that combines the advantages of convolutional and recurrent neural networks to effectively suppress environmental noise while maintaining the naturalness and integrity of the speech. Its advantages include: the DPCRN model achieves a good balance between denoising effectiveness and computational efficiency; in addition to offline testing, the project includes real-time inference scripts, enabling real-time speech enhancement on devices; it supports quantization and pruning in TensorFlow Lite, allowing the model to be used on resource-constrained platforms such as mobile devices; it provides a concise command-line interface for convenient model training, testing, and real-time processing; and it is open-source with clear reference guidelines.
[0074] Specifically, when the secondary verification trigger condition is triggered, several segments of the user audio, a preset number of verification photos, and several randomly generated interference images are randomly sent to the client to be detected, including:
[0075] When the secondary verification trigger condition is triggered, several sets of verification pairs are randomly selected;
[0076] Randomly select several segments of the user audio and a preset number of the verification photos from the selected verification pairs;
[0077] Based on the difference between the selected user audio and the verification photo, several interference images are randomly generated;
[0078] The selected user audio, the verification photo, and the interference image are sent to the client to be detected.
[0079] Further, the matching results of the client to be detected are collected, and the matching results and the verification are analyzed according to the conditions to obtain the verification results, including:
[0080] Receive feedback information sent by the client to be tested;
[0081] The feedback information is parsed to obtain the matching result;
[0082] The matching results are judged using the verification conditions to obtain a judgment result;
[0083] When all the judgment results are in compliance, the verification result is determined to be successful.
[0084] If any of the judgment results is inconsistent, the verification result is determined to be a verification failure.
[0085] Preferably, after the client to be tested logs into the server, risk control detection is performed on the behavior of the client to be tested to obtain risk control data, including:
[0086] Once the client to be tested logs into the server, its behavior information is recorded.
[0087] The total value of the behavioral information is calculated based on a preset risk behavior score set to obtain the risk control data; the risk behavior score set can be set according to user needs.
[0088] Preferably, this embodiment provides a KNN-based similar image filtering method for providing the interference images to the client to be detected. Its design concept is as follows:
[0089] Step S1: Collect the verification photos and user audio from all users across the entire platform;
[0090] Step S2: Filter the verification photo and the user audio; this step is used to filter out verification photos with poor display quality and user audio with too many interference frequencies;
[0091] Step S3: Based on the matching information of the verification pair, combine the filtered verification photo and the user audio, and discard any redundant unpaired verification photos and user audio.
[0092] Step S4: Extract the text information of the user's audio, extract keywords based on the text information, and combine the keywords with the corresponding verification photos to obtain training samples;
[0093] Step S5: Randomly divide the training samples into a training set and a test set (7:3);
[0094] Step S6: Construct the k-nearest neighbor network;
[0095] Step S7: Construct the loss function and optimization function;
[0096] Step S8: Train the k-nearest neighbor network to obtain a classification model.
[0097] Specifically, kNN (k-nearest neighbor) is a basic and simple classification algorithm. As a supervised learning method, the KNN model requires labeled training data. The class of a new sample is determined by the k nearest training data points according to a classification decision rule. k-nearest neighbor (kNN) is a basic classification and regression method; it is a model based on labeled training data; and it is a supervised learning algorithm. The three key points of its basic approach are: first, determining the distance metric; second, choosing the value of k (finding the k closest instances in the training set to the estimated point); and third, the classification decision rule. In classification tasks, a "voting method" can be used, that is, selecting the most frequently occurring label class among these k instances as the prediction result; in regression tasks, an "averaging method" can be used, that is, the average of the real values of the k instances' labels is used as the prediction result; weighted averaging or weighted voting can also be performed based on distance, with closer instances having greater weight. kNN does not have an explicit learning process. It is a well-known example of lazy learning. This type of learning technique simply saves the samples during the training phase, with zero training time overhead, and then processes them after receiving the test samples.
[0098] Optionally, the main method of k-nearest neighbor clustering analysis is as follows: Initially, K (hyperparameter) cluster centers are randomly given, and these K cluster centers are randomly selected from the sample points. The sample points to be classified are assigned to each cluster according to the nearest neighbor principle. After the division, the centroids of each cluster are recalculated using the averaging method, thus determining the new cluster centers. This process is iterated until the movement distance of the cluster centers is less than a given value. The evaluation criterion for the algorithm is the sum of squared errors criterion; when the sum of squared errors reaches its optimum (minimum), the clusters are as compact as possible within each cluster and as far apart as possible between clusters. The loss function is a non-convex function, so it has many minima and may not always reach the global optimum, possibly reaching a local optimum. Global optimization is performed using the loss function during iteration. The remaining data are clustered according to the silhouette coefficient; the formula for calculating the silhouette coefficient is:
[0099]
[0100] Where s is the silhouette coefficient; b is the sum of distances between the data to be assigned and all points in the next nearest cluster; a is the sum of distances between the data to be assigned and other points in the same cluster; max() is the maximum value function. The silhouette coefficient ranges from (-1, 1), where a value closer to 1 indicates that the sample is very similar to samples in its own cluster and not similar to samples in other clusters. When a sample point is more similar to sample points outside the cluster, the silhouette coefficient is negative. When the silhouette coefficient is 0, it means that the two clusters have the same similarity, and the two clusters should belong to the same cluster.
[0101] Furthermore, the authentication method provided in this embodiment, during runtime, mandates that all user audio sent to the client under test for authentication originates from the characteristic account to be logged in, and that the number of user audio clips exceeds the number of verification photos. Ultimately, it requires that the number of images received by the client under test is twice the number of audio clips.
[0102] Specifically, the authentication method provided in this embodiment calculates the difference between the number of user audio and verification photos sent to the client to be detected based on the verification pass conditions set by the user in advance, and generates a corresponding number of interference images using the classification model.
[0103] Furthermore, the tool used to extract the text information from the user's audio is FFmpeg. FFmpeg is an open-source computer program that can record, convert, and stream digital audio and video. It is licensed under the LGPL or GPL. It provides a complete solution for recording, converting, and streaming audio and video. It includes the highly advanced audio / video codec library libavcodec, much of the code in libavcodec was developed from scratch to ensure high portability and encoding / decoding quality. FFmpeg was developed on the Linux platform, but it can also be compiled and run on other operating system environments, including Windows and Mac OS X. This project was originally initiated by Fabrice Bellard and was primarily maintained by Michael Niedermayer from 2004 to 2015. Many FFmpeg developers come from the MPlayer project, and FFmpeg is currently hosted on the MPlayer project team's servers. The project name comes from the MPEG video coding standard, with "FF" standing for "FastForward". The FFmpeg encoding library can be accelerated using GPUs. FFmpeg is an open-source computer program that can record, convert, and stream digital audio and video. It includes leading audio / video encoding libraries such as libavcodec. libavformat: used for generating and parsing various audio and video container formats, including obtaining information needed for decoding to generate the decoding context structure and reading audio and video frames; libavcodec: used for encoding and decoding various types of sound / images; libavutil: contains some common utility functions; libswscale: used for video scene scaling and color mapping conversion; libpostproc: used for post-processing effects; ffmpeg: a tool provided by the project that can be used for format conversion, decoding, or real-time encoding for TV cards; ffsever: an HTTP multimedia real-time broadcast streaming server; ffplay: a simple player that uses the ffmpeg library for parsing and decoding and displays via SDL. This example provides basic pseudocode for extracting video and audio files based on ffmpeg:
[0104] ffmpeg-i source_video.avi-vn-ar44100-ac 2-ab 192-fmp3 sound.mp3;
[0105] Here, source_video.avi is the original video file, but it can also be other types of video files, such as mp4; 192-f is the audio bitrate, that is, the audio bitrate of the audio file is set to 192kb / s; mp3 is the target output format; and sound.mp3 is the output audio file name.
[0106] The beneficial effects of this invention are as follows:
[0107] This invention provides greater personalization for user verification by allowing users to set secondary verification trigger and pass conditions; it offers a novel verification method by collecting user audio and verification photos and setting verification pairs, thus improving verification security and specificity; and it reduces the poor user experience caused by frequent verification triggers through risk control detection and risk standard updates, thereby improving the versatility of the authentication method.
[0108] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0109] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An identity authentication method, characterized in that, include: Register a platform account to obtain a personal account; Set risk standards; Based on the aforementioned personal account, set the conditions for triggering secondary verification and the conditions for passing verification; Several sets of one-to-one corresponding user audio and verification photos are entered, and the corresponding user audio and verification photos are bound together to obtain several sets of verification pairs; When the secondary verification trigger condition is triggered, several segments of the user's audio, a preset number of the verification photos, and several randomly generated interference images are randomly sent to the client to be detected. Collect the matching results of the client to be detected, analyze the matching results and the verification conditions to obtain the verification result; When the verification result is successful, the client to be tested is allowed to access the server; Once the client to be tested logs into the server, the client's behavior is monitored for risk control to obtain risk control data. When the risk control data exceeds the risk standard, the client to be tested is repeatedly verified by forcibly triggering the secondary verification trigger condition to obtain the risk control identification result. When the risk control identification result is a risk-free issue, the risk standard is increased.
2. The identity authentication method according to claim 1, characterized in that, The formula for calculating the risk standard is: Where R is the risk standard; R0 is the basic risk value; and n is the number of times the repeated verification is successful.
3. The identity authentication method according to claim 1, characterized in that, The secondary verification trigger conditions include: logging in with an unused address, changing the login device, frequent login attempts, accessing private files, and frequently deleting files.
4. The identity authentication method according to claim 1, characterized in that, The verification pass conditions include: matching accuracy, number of matching errors, and matching time.
5. The identity authentication method according to claim 1, characterized in that, Several sets of one-to-one corresponding user audio and verification photos are entered, and the corresponding user audio and verification photos are bound together to obtain several verification pairs, including: Collect uploaded audio and photos from target users; The uploaded audio is denoised and compressed to obtain the user's audio; The uploaded photo is subjected to noise reduction and resolution normalization to obtain the verification photo; According to the matching scheme for the target user, the corresponding user audio and the verification photo are packaged and stored to obtain the verification pair.
6. The identity authentication method according to claim 1, characterized in that, When the secondary verification trigger condition is triggered, several segments of the user's audio, a preset number of the verification photos, and several randomly generated interference images are randomly sent to the client to be detected, including: When the secondary verification trigger condition is triggered, several sets of verification pairs are randomly selected; Randomly select several segments of the user audio and a preset number of the verification photos from the selected verification pairs; Based on the difference between the selected user audio and the verification photo, several interference images are randomly generated; The selected user audio, the verification photo, and the interference image are sent to the client to be detected.
7. The identity authentication method according to claim 1, characterized in that, Collect the matching results of the client to be detected, analyze the matching results and the verification according to the conditions, and obtain the verification results, including: Receive feedback information sent by the client to be tested; The feedback information is parsed to obtain the matching result; The matching results are judged using the verification conditions to obtain a judgment result; When all the judgment results are in compliance, the verification result is determined to be successful. If any of the judgment results is inconsistent, the verification result is determined to be a verification failure.
8. The identity authentication method according to claim 1, characterized in that, After the client to be tested logs into the server, its behavior is monitored for risk control to obtain risk control data, including: Once the client to be tested logs into the server, its behavior information is recorded. Based on a preset set of risk behavior scores, the total value of the behavior information is calculated to obtain the risk control data.
Citation Information
Patent Citations
Method and system for verifying user identity, client, and server
CN105991590A
Identity verification method and system
CN114186209A