Privacy protection continuous authentication method based on user behavior characteristics
Patent Information
- Application Number
- CN202310563821.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-05-18
AI Technical Summary
[0005]本发明的目的在于针对上述已有技术的不足,提出一种基于用户行为特征的隐私保护持续认证方法,用于解决未使用特征向量训练机器学习分类器、直接使用特征向量作为身份认证模型导致认证准确度低的问题,以及在线认证阶段将行为特征数据以明文的形式传输到服务器带来的隐私泄露问题
[0026]第一,本发明为用户训练随机森林分类器作为身份认证模型,使用训练得到的身份认证模型对用户身份进行验证,克服了现有技术未使用特征向量训练机器学习分类器、直接使用特征向量作为身份认证模型导致认证准确度低的不足,使得本发明具有认证准确度高、实用性强的优点。
Smart Images

Figure CN116502269B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electronic digital data processing technology, and further relates to a privacy-preserving continuous authentication method based on user behavior characteristics within the field of behavioral data processing technology. This invention enables privacy-preserving continuous authentication of user identity. Background Technology
[0002] Continuous authentication methods based on user behavior features can provide continuous and transparent identity verification for users, offering stronger security guarantees through continuous and implicit authentication. In some continuous authentication schemes based on user behavior features, the training and authentication tasks of the user identity authentication model are completed by the server. Specifically, the server guides the user to complete a series of specific operations on the terminal interface, collects behavioral data, extracts behavioral features, trains the identity authentication model for the user, and continuously judges the legitimacy of the current user's identity during the authentication phase.
[0003] Abbas Acar et al. proposed a lightweight privacy-aware continuous authentication method in their paper "A Lightweight Privacy-Aware Continuous Authentication Protocol-PACA" (2021 ACM Transactions on Privacy and Security (TOPS), 2021, 24(4): 1-28.). This method employs a password-based key exchange mechanism to perform continuous privacy-aware authentication based on user behavioral characteristics. It uses a template matching method, calculating the distance (e.g., Euclidean distance, Manhattan distance) between the probe template and the stored template and comparing it with a specific threshold to obtain the final authentication result. A drawback of this method is that, in the identity authentication model generation stage, it does not use feature vectors to train a machine learning classifier but directly uses feature vectors as the identity authentication model, resulting in low authentication accuracy and thus reducing the practicality of the continuous authentication method.
[0004] Nanjing University of Aeronautics and Astronautics proposed a system-level continuous user authentication method for smartphones in its patent application, "A System-Level Continuous User Identity Authentication Method for Smartphones" (Patent Application No. 202011136858.2, Publication No. CN112261222A). This method constructs a user behavior feature model using data from sensors on the smartphone and then uses this model to continuously authenticate the user. In the online authentication phase, the current user's feature vector set needs to be sent to the server-side identity authenticator to obtain the authentication result. A drawback of this method is that it requires transmitting behavioral feature data to the server in plaintext during the online authentication phase. If the server is compromised, it could lead to user privacy leaks. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the existing technologies by proposing a privacy-preserving continuous authentication method based on user behavior features. This method solves the problems of low authentication accuracy caused by not using feature vectors to train machine learning classifiers and directly using feature vectors as identity authentication models, as well as the privacy leakage problem caused by transmitting behavioral feature data to the server in plaintext during the online authentication stage.
[0006] The idea behind this invention is to input a set of user behavior feature vectors into a random forest classifier for training. The trained random forest classifier is then used as the user's authentication model. Because random forest classifiers have advantages such as high classification accuracy, strong generalization ability, and fast training speed, they better meet the requirements for accuracy and practicality of the authentication model. This avoids the problem of low authentication accuracy caused by using feature vectors directly as the authentication model without training the machine learning classifier using feature vectors. This invention stores the encrypted authentication model on the user's end. During the authentication phase, the user does not need to send behavioral feature data to the server. Instead, they calculate the ciphertext set locally based on the behavioral features and the encrypted model, and then send the ciphertext set to the server for decryption. The entire authentication process is completed in the ciphertext domain, avoiding the privacy leakage problem caused by transmitting behavioral feature data to the server in plaintext during the online authentication phase of existing technologies, thus protecting the privacy of the user's behavioral data.
[0007] The implementation steps of this invention include the following:
[0008] Step 1: The server generates a training set of users to be authenticated.
[0009] The N1 behavioral feature vectors of the user to be authenticated are used to form a positive sample set. N2 feature vectors are randomly selected from the behavioral feature vector sets of other users to form a negative sample set. The label of each positive sample is set to 1, and the label of each negative sample is set to -1. The positive and negative sample sets and their corresponding labels are used to form a training set, where N1≥500 and N1=N2.
[0010] Step 2: The server trains a random forest authentication model for the users to be authenticated.
[0011] The training set is fed into a random forest classifier for training until the number of decision trees in the random forest classifier is greater than 50. Training is then stopped, and the trained random forest classifier is obtained. This classifier is used as the identity authentication model for the user to be authenticated and is saved in the local database.
[0012] Step 3: The server processes the random forest authentication model.
[0013] Extract the path of each decision tree in the identity authentication model and form a path set from all paths; add a pseudo node to the end of each path, and the length of the path after adding the pseudo node is equal to the maximum length of the path in the path set.
[0014] Step 4: The server generates a public / private key pair.
[0015] The key generation algorithm KeyGen from Paillier homomorphic encryption is used to generate seven parameters p, q, N, λ. g,μ, store N and g as public key pk and λ and μ as private key sk;
[0016] Step 5, Server-side encrypted random forest authentication model:
[0017] The encryption algorithm Enc from Paillier homomorphic encryption technology is used to encrypt each decision node in the path set with added pseudo nodes, resulting in an encrypted random forest identity authentication model.
[0018] Step 6: The server sends the encrypted random forest authentication model and public key pk to the user to be authenticated;
[0019] Step 7, the user to be authenticated calculates the ciphertext set:
[0020] The user to be authenticated uses a feature extraction method to collect a behavioral feature vector x from the client that is the same type and size as described in step 1. Based on x, the ciphertext corresponding to each decision node is retrieved from the encrypted random forest identity authentication model, and the ciphertext set is output.
[0021] Step 8: The user to be authenticated sends a set of encrypted data to the server.
[0022] The user to be authenticated scrambles the order of the ciphertext in the ciphertext set and sends the scrambled ciphertext set to the server.
[0023] Step 9: The server decrypts the data to obtain the final authentication result.
[0024] The server uses the decryption algorithm Dec from Paillier homomorphic encryption technology to decrypt the scrambled ciphertext set, obtain the plaintext set, count the number r of plaintexts with a value of 0 in the plaintext set, and determine whether the authentication is successful based on the authentication threshold τ. If r > τ, the authentication is successful; otherwise, the authentication fails.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] First, this invention trains a random forest classifier for users as an identity authentication model, and uses the trained identity authentication model to verify the user's identity. This overcomes the shortcomings of existing technologies that do not use feature vectors to train machine learning classifiers and directly use feature vectors as identity authentication models, resulting in low authentication accuracy. This invention has the advantages of high authentication accuracy and strong practicality.
[0027] Secondly, this invention requires the server to encrypt the identity authentication model and send it to the user's storage. During the authentication phase, the user calculates the ciphertext set based on behavioral characteristics and the encryption model, and sends the ciphertext set to the server for decryption. The entire authentication process is completed in the ciphertext domain, avoiding the privacy leakage problem caused by transmitting behavioral characteristic data to the server in plaintext during the online authentication phase of the prior art. This invention protects the privacy of user behavioral data during the continuous authentication of user identity, and has the advantage of strong privacy. Attached Figure Description
[0028] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0029] The following is in conjunction with the appendix Figure 1 The present invention will be further described in conjunction with the embodiments.
[0030] Step 1: The server generates a training set of users to be authenticated.
[0031] A positive sample set is formed by N1 behavioral feature vectors of the user to be authenticated, and a negative sample set is formed by randomly selecting N2 feature vectors from the behavioral feature vector sets of other users. The label of each positive sample is set to 1, and the label of each negative sample is set to -1. The positive and negative sample sets and their corresponding labels are combined to form a training set, where N1≥500 and N1=N2.
[0032] The feature vector is a vector extracted from any one of the behavioral features selected from touch screen, gait, and keystroke.
[0033] In this embodiment of the invention, N1 = 500, and the behavioral feature used is the touch screen behavioral feature. The feature extraction method is as follows: When a user interacts with the client device via touch screen, each touch stroke consists of a series of touch points located between the start position and the stop position. For each swipe, by collecting five original features—the horizontal coordinate, vertical coordinate, timestamp, pressure, and finger coverage area of the touch point—and performing statistical calculations, behavioral feature information that can accurately identify the user can be extracted, resulting in a feature vector containing 15 feature points. These 15 feature points are: the duration of a touch stroke, the horizontal coordinate of the touch start point, the vertical coordinate of the touch start point, the horizontal coordinate of the touch end point, the vertical coordinate of the touch end point, the straight-line distance between the touch start point and the end point, the straightness of the stroke, the stroke direction, the slope of the line segment defined by the two endpoints, the maximum distance between the touch point and the line segment defined by the two endpoints, the average slope of the stroke trajectory line segment, the length of the stroke, the average speed of the stroke, the pressure at the midpoint of the stroke, and the area of the stroke covered by the finger.
[0034] Step 2: The server trains a random forest identity authentication model for the users to be authenticated.
[0035] The training set is fed into a random forest classifier for training until the number of decision trees in the random forest classifier is greater than 50. Training is then stopped, and the trained random forest classifier is obtained. This classifier is used as the authentication model for the user to be authenticated and is saved in the local database.
[0036] The random forest classifier is a binary classifier, where the classification label for positive samples is 1 and the classification label for negative samples is -1.
[0037] Step 3: The server processes the random forest identity authentication model.
[0038] Extract the path from each decision tree in the identity authentication model, and group all paths into a path set {P1,...,P}. i ,...,P α}, where P i Let represent the i-th path, and α represent the total number of paths.
[0039] Add a pseudo-node to the end of each path. The length of the path after adding the pseudo-node is equal to the maximum length β of the paths in the path set.
[0040] The extraction of paths for each decision tree in the identity authentication model involves extracting all paths leading to the leaf nodes with a corresponding classification label value of 1 for each decision tree in the identity authentication model. Each path consists of several decision nodes.
[0041] The pseudo-node is added, with a feature index of 0, a comparison symbol of >, and a feature value of -1.
[0042] Step 4: The server generates a public / private key pair.
[0043] The key generation algorithm KeyGen from Paillier homomorphic encryption is used to generate seven parameters p, q, N, λ. g,μ, store N and g as public key pk, and store λ and μ as private key sk.
[0044] The Paillier homomorphic encryption technology is derived from Paillier's paper "Public-Key Cryptosystems Based on Composite Degree Residuosity Classes" published at the EUR / USD conference in 1999. It supports additive homomorphism and features high computational efficiency and complete security proof.
[0045] The seven parameters p, q, N, λ g,μ refers to p and q representing prime numbers of length ξ bits, N representing the plaintext modulus and N = p·q, and λ representing the least common multiple of p-1 and q-1. Let g denote the cyclic group of multiplication, and g denote the multiplication group from... The generators selected in the denominator have an order equal to (p-1)(q-1) / 2, and μ represents the auxiliary computation parameter, the value of which is equal to ((g λ modN 2 -1) / N) -1 modN.
[0046] In the embodiments of the present invention, ξ = 512, and its value satisfies the 80-bit security level of the public key encryption algorithm.
[0047] Step 5, Server-side encrypted random forest authentication model:
[0048] By employing the Enc encryption algorithm from Paillier homomorphic encryption technology, each decision node in the path set with added pseudo-nodes is encrypted to obtain an encrypted random forest identity authentication model.
[0049] The specific steps for encrypting each decision node in the path set with added pseudo-nodes are as follows: for the j-th decision node d in the i-th path... i,j Encryption is performed to obtain a set of ciphertext. Where σ represents d i,j The bit length corresponding to the eigenvalue, d i,jThis corresponds to the k-th Paillier ciphertext in the ciphertext set.
[0050] The The calculation formula is:
[0051]
[0052] Among them, v i,j d i,j The corresponding comparison sign value, t i,j d i,j The corresponding eigenvalues.
[0053] The v i,j The calculation formula is:
[0054]
[0055] The encrypted random forest authentication model contains a set of ciphertexts corresponding to each decision node. and feature index set {id 1,1 ,...,id 1,β ,id 2,1 ,...,id 2,β ,...,id α,1 ,...,id α,β}, where id α,β Represents decision node d α,β The corresponding feature index.
[0056] Step 6: The server sends the encrypted random forest authentication model and public key pk to the user to be authenticated.
[0057] Step 7, the user to be authenticated calculates the ciphertext set:
[0058] The user to be authenticated uses a feature extraction method to collect a behavioral feature vector x from the client that is the same type and size as described in step 1. Based on x, the ciphertext corresponding to each decision node is retrieved from the encrypted random forest identity authentication model, and the ciphertext set is output.
[0059] The ciphertext corresponding to each decision node is retrieved from the encrypted random forest identity authentication model based on x, and based on each decision node d... i,j Corresponding feature index value id i,j Select the value in x that matches id. i,j Corresponding eigenvalues Search for the ciphertext corresponding to the feature value in the ciphertext set corresponding to the decision node.
[0060] The output ciphertext set is denoted as {S1,...,S...}. j,...,S α}, where S j This represents the j-th Paillier ciphertext in the ciphertext set.
[0061] The S j The calculation formula is:
[0062]
[0063] in, This represents the homomorphic addition operation in Paillier homomorphic encryption.
[0064] In the embodiments of the present invention, the behavioral features used are touch screen behavioral features, and the feature extraction method used is the same as the feature extraction method described in step 1.
[0065] Step 8: The user to be authenticated sends a set of encrypted data to the server.
[0066] The user to be authenticated scrambles the order of the ciphertext in the ciphertext set and sends the scrambled ciphertext set to the server.
[0067] The purpose of scrambling the ciphertext order in the ciphertext set is to hide S from the server. j The plaintext value.
[0068] Step 9: The server decrypts the data to obtain the final authentication result.
[0069] The server uses the decryption algorithm Dec from Paillier homomorphic encryption technology to decrypt the scrambled ciphertext set, obtain the plaintext set, count the number r of plaintexts with a value of 0 in the plaintext set, and determine whether the authentication is successful based on the authentication threshold τ. If r > τ, the authentication is successful; otherwise, the authentication fails.
[0070] The authentication threshold τ is set to half the number of decision trees in the random forest classifier.
Claims
1. A privacy-preserving continuous authentication method based on user behavior characteristics, characterized in that, The server trains a random forest classifier for the user as an authentication model, encrypts the authentication model, and sends it to the user's storage. During the authentication phase, the user calculates a ciphertext set based on behavioral characteristics and the encryption model, and sends the ciphertext set to the server for decryption. The entire authentication process is completed within the ciphertext domain. The steps of this method include the following: Step 1: The server generates a training set of users to be authenticated. Users to be authenticated A set of positive samples is composed of behavioral feature vectors, which are randomly drawn from the set of behavioral feature vectors of other users. The negative sample set is composed of feature vectors. Each positive sample is labeled 1, and each negative sample is labeled -1. The positive and negative sample sets and their corresponding labels form the training set. ; Step 2: The server trains a random forest authentication model for the users to be authenticated. The training set is fed into a random forest classifier for training until the number of decision trees in the random forest classifier is greater than 50. Training is then stopped, and the trained random forest classifier is obtained. This classifier is used as the authentication model for the user to be authenticated and is saved in the local database. Step 3: The server processes the random forest authentication model. Extract the path of each decision tree in the identity authentication model and form a path set from all paths; add a pseudo node to the end of each path, and the length of the path after adding the pseudo node is equal to the maximum length of the path in the path set. Step 4: The server generates a public / private key pair. The key generation algorithm KeyGen from Paillier homomorphic encryption is used to generate seven parameters. ,Will and As a public key Storage, will and As a private key storage; Step 5, Server-side encrypted random forest authentication model: The encryption algorithm Enc from Paillier homomorphic encryption technology is used to encrypt each decision node in the path set with added pseudo nodes, resulting in an encrypted random forest identity authentication model. Step 6: The server will encrypt the random forest authentication model and public key. Send to the user to be authenticated; Step 7, the user to be authenticated calculates the ciphertext set: The user to be authenticated is selected using a feature extraction method, which collects a behavioral feature vector from the client that is the same type and size as in step 1. ,according to Retrieve the ciphertext corresponding to each decision node from the encrypted random forest identity authentication model and output the ciphertext set; Step 8: The user to be authenticated sends a set of encrypted data to the server. The user to be authenticated scrambles the order of the ciphertext in the ciphertext set and sends the scrambled ciphertext set to the server. Step 9: The server decrypts the data to obtain the final authentication result. The server uses the Dec decryption algorithm from Paillier homomorphic encryption to decrypt the scrambled ciphertext set, obtaining the plaintext set, and then counts the number of plaintext characters with a value of 0 in the plaintext set. And based on the authentication threshold To determine whether the authentication was successful, if... If the authentication passes, the authentication is successful; otherwise, the authentication fails.
2. The privacy-preserving continuous authentication method based on user behavior characteristics according to claim 1, characterized in that, The feature vector mentioned in step 1 is a vector extracted from any one type of behavioral feature selected from touch screen, gait, and keystroke.
3. The privacy-preserving continuous authentication method based on user behavior characteristics according to claim 1, characterized in that, The seven parameters mentioned in step 4 It refers to, and Indicates length is prime numbers with 1 digit, Represent the plaintext modulus and , express and The least common multiple of, Represents a multiplication cyclic group. Indicates from The generators selected in the set have an order equal to 1. , This represents an auxiliary calculation parameter, whose value is equal to... .
4. The privacy-preserving continuous authentication method based on user behavior characteristics according to claim 1, characterized in that, The encrypted random forest authentication model described in step 5 includes a set of ciphertexts and a set of feature indexes for each decision node.
5. The privacy-preserving continuous authentication method based on user behavior characteristics according to claim 4, characterized in that, According to step 7 Retrieve the ciphertext corresponding to each decision node from the encrypted random forest identity authentication model, and based on the feature index value corresponding to each decision node, in... Select the feature value corresponding to the feature index value, and search for the ciphertext at the corresponding position of the feature value in the ciphertext set corresponding to the decision node.
6. The privacy-preserving continuous authentication method based on user behavior characteristics according to claim 1, characterized in that, The authentication threshold mentioned in step 9 The value is half the number of decision trees in the random forest classifier.
Citation Information
Patent Citations
System-level user identity continuous authentication method on smart phone
CN112261222A
Model training method and device and electronic equipment
CN111144576A
Privacy protection biological characteristic authentication method based on decision tree
CN113239336A