Context-based deep mail password strength measurement method

By constructing a context-based deep email password strength measurement method, and using a multi-head self-attention mechanism and a character-level LSTM decoder to simulate the password guessing behavior of attackers with known user information, this method solves the problem that existing technologies fail to fully consider user context information, and achieves more accurate password security assessment and risk warning.

CN120934756APending Publication Date: 2025-11-11SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511352466.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing password strength assessment methods fail to fully consider user context information, resulting in inaccurate security assessments in actual attack scenarios, especially in their insufficient ability to identify personalized weak passwords.

Method used

We construct a context-based deep email password strength measurement method. By extracting and encoding user context information, such as email address, name and birthday, we use a multi-head self-attention mechanism and a character-level LSTM decoder to simulate the password guessing behavior of attackers with known user information. We then combine a base model and a context-sensitive model to evaluate password strength.

Benefits of technology

It significantly improves the accuracy of identifying personalized weak passwords, making the assessment results closer to real security risks, providing more accurate security feedback, and enhancing user experience and account security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120934756A_ABST
    Figure CN120934756A_ABST
Patent Text Reader

Abstract

The invention provides a context-based deep mail password strength measurement method, which comprises the following steps of: introducing user context information such as an e-mail, a name and a birthday on a basic password measurement model framework, and realizing information fusion through a multi-head self-attention mechanism; and a password sequence probability calculation model is constructed in combination with a character-level long-short-term memory network (LSTM). An evaluation system taking guess times as a core index is established to quantify the performance difference of different models under the condition that whether user context information exists or not. Experiments show that the method can more truly reflect the cracking capability of an attacker when the attacker masters the context information of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of cryptographic security technology, specifically relating to a context-based method for measuring the strength of deep email ciphers. Background Technology

[0002] As the most common authentication mechanism, passwords are directly related to the protection of user accounts and sensitive information. Traditional password strength assessment methods are mainly based on information entropy, rule checking, or static probability models. These methods often assume that passwords are randomly generated or rely solely on the complexity of character combinations for evaluation, failing to fully consider the correlation between user personal information and passwords.

[0003] In reality, users tend to use passwords that are related to their personal information, such as their name, birthday, and email username. This allows attackers to significantly increase their password guessing efficiency once they obtain user context information. While existing deep learning cryptographic models can learn password distributions from large-scale data, most do not explicitly incorporate user context information and cannot accurately simulate password guessing behavior in real-world attack scenarios.

[0004] Therefore, it is necessary to propose a password strength measurement method that can integrate user context information to more realistically reflect the vulnerability of passwords under actual attack conditions and provide users and systems with more accurate security assessments. Summary of the Invention

[0005] Therefore, the purpose of this invention is to provide a context-based deep email password strength measurement method to more realistically reflect the vulnerability of passwords under actual attack conditions, and to provide users and systems with more accurate security assessments.

[0006] The technical solution of this invention is: a context-based deep email password strength measurement method, comprising:

[0007] S1: Build a basic password generation model to simulate the behavior of an attacker guessing passwords without having user context information;

[0008] S2: Construct a context-sensitive cryptographic strength measurement model; the model includes the following:

[0009] Extract and encode user context information: The email address is split into username, subdomain, and top-level domain, and encoded using Word2Ve word embedding and character-level LSTM respectively; one-hot encoding is used for high-frequency domains, and hash encoding is used for low-frequency domains; the name and date of birth are encoded into dense vectors through word embedding and LSTM; all context encoding vectors are concatenated and positional encoding is added to form a unified context representation;

[0010] Fusing contextual information: The context representation is input into a multi-head self-attention mechanism to extract the contextual features most relevant to password generation and output the fused context vector;

[0011] Context-aware password decoder: It adopts a character-level LSTM decoder, whose initial state is obtained by mapping the context vector output from the third step through a fully connected layer; the decoder generates a password sequence character by character, and at each step predicts the probability distribution of the next character based on the current state and the previous character;

[0012] S3: The user enters their password, email, birthday, and name information. The number of password guesses is calculated. The probability of password generation is calculated based on the basic password generation model and the context-sensitive password strength measurement model. The number of guesses is calculated with and without context information.

[0013] S4: Output password strength assessment results. Based on the comparison between the number of guesses and the preset threshold, determine the vulnerability of the password in actual attack scenarios and provide users with strength feedback and risk warnings.

[0014] Preferably, the basic password generation model described in S1 uses a character-level Long Short-Term Memory (LSTM) network to capture the contextual dependencies in the password character sequence, thereby learning the probability distribution of password generation.

[0015] The architecture of the basic cryptographic generation model consists of a two-layer stacked LSTM network and a fully connected network for modeling cryptographic character sequences;

[0016] The training objective of the basic password generation model is to maximize the likelihood function of the real password sequence under the model's generation probability distribution. At each time step, the model receives the embedding vector of the current character as input, calculates its hidden state through the LSTM unit, and outputs the conditional probability distribution of the next character using the Softmax function.

[0017] After training, the model is used to evaluate the probability of generating any password sequence. The higher the probability, the easier the password is to guess and the lower its strength.

[0018] Preferably, the encoding method for user context information in S2 includes:

[0019] The username is used to generate a semantic embedding vector through the Word2Vec model, and the context features are extracted through a character-level two-layer LSTM to form an embedding representation of the username. For subdomains and top-level domains, if they belong to the top 60% of high-frequency domains in the dataset, one-hot encoding is used; otherwise, a hash function is used to map them to a fixed-length index, and a fixed-dimensional embedding vector is generated through direct encoding.

[0020] (1)

[0021] (2)

[0022] in, This represents the semantic embedding vector of the username generated by the Word2Vec model. This represents a character-level long short-term memory network. This represents a fully connected neural network layer. Indicates one-hot encoding. Represents a hash function;

[0023] Name and birthdate information are encoded using a similar method to email usernames, and processed using the Word2Vec model and character-level LSTM to generate corresponding embedding vectors;

[0024] (3)

[0025] (4)

[0026] in, , , , These represent the semantic embedding vectors for year, month, day, and user name generated by the Word2Vec model, respectively.

[0027] The above information is concatenated to form a unified user context representation:

[0028] (5)

[0029] Positional encoding is introduced on top of the embedding vector. Positional encoding uses sine and cosine functions in different dimensions to ensure that each position in the sequence has a unique representation. The calculation method is as follows:

[0030] (6)

[0031] (7)

[0032] in, This indicates that the input is in the embedding vector. The position in the middle, Indicates a dimension index. It is the dimension of the embedding vector; each embedding vector With the corresponding Adding elements together yields a representation that includes location information.

[0033] Finally, the fused sequence representation is fed into the multi-head self-attention mechanism to obtain a representation of the context information:

[0034] (8)

[0035] in, This represents a multi-head self-attention mechanism.

[0036] Preferably, the context-aware cryptographic decoder includes: the decoder's initial state. The password is calculated by encoding user context information, and the decoding process uses a character-by-character generation method. Let the password sequence be represented as... ,in, Indicates the first Each character; the decoding process for each step is as follows:

[0037] (1) Input embedding: embed the characters selected in the previous step The decoder state is taken as input, and the initial decoder state is... The start character is The subsequent decoder state is derived from the hidden state passed from the previous time step. express:

[0038] (10)

[0039] (2) Probability prediction: using The function transforms the output of the fully connected layer into a probability distribution for each possible character:

[0040] (11)

[0041] in, and These are the output layer parameters;

[0042] (3) Character selection: Select the character with the highest probability in the current probability distribution as the output:

[0043] (12)

[0044] in, This indicates that the character with the highest probability is selected;

[0045] The above process is repeated until a stop marker is generated or the maximum password length is reached. .

[0046] Preferably, in S3, based on the basic model With context-sensitive models The generation probability difference was calculated, and the number of guesses for the two types were calculated separately to measure the impact of the user's public context information on password predictability.

[0047] (1) Calculate the probability: for the password Calculate the generation probabilities of the two types of models separately:

[0048] Basic model:

[0049] (13)

[0050] in, Password The One character;

[0051] Context-sensitive models:

[0052] (14)

[0053] in, Password The One character, Indicates user context information;

[0054] (2) Number of basic model guesses: Password guessing attempts without user context information. The number of guesses is:

[0055] (15)

[0056] in, This represents the set of all possible passwords (sample space). The total number of samples, Based on the model for the first The probability of generating a password;

[0057] (3) Number of guesses in the context-sensitive model: Given the user's context information ( When, password The number of guesses is:

[0058] (16)

[0059] in, For context-sensitive models, under user information conditions, the first The probability of generating a password.

[0060] This invention provides a context-based deep email password strength measurement method. By incorporating user context information and constructing a context-aware deep generative model, this method can simulate the password cracking capabilities of attackers with partial user background knowledge, significantly improving the accuracy of identifying "personalized weak passwords" (such as zhangwei1985, john_doe@mail), making the strength assessment results closer to real security risks. Simultaneously, by constructing a dual-model comparison mechanism of a basic model and a context-sensitive model, it simulates password guessing behavior under both context-free and context-sensitive conditions, quantifying the vulnerability of passwords under real attacks using the number of guesses, significantly improving the accuracy of identifying personalized weak passwords. Finally, it can provide feedback based on the evaluation results, prompting users to optimize their password structure, forming a "evaluation-prompt-optimization" security closed loop, improving user experience and account security. The overall solution has clear technical implementation, high engineering feasibility, and can be widely integrated into practical application scenarios such as email systems and identity authentication platforms, possessing good practicality, scalability, and promotional value. Attached Figure Description

[0061] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 A schematic diagram of the structure of the context-sensitive cryptographic strength measurement model provided by the present invention;

[0064] Figure 2 A schematic diagram of location encoding provided by the present invention;

[0065] Figure 3 This invention provides a correspondence between the number of guesses for the two types of models.

[0066] Figure 4 The present invention provides the ratio of the number of guesses for the two types of models under different conditions. Detailed Implementation

[0067] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of systems consistent with some aspects of the invention as detailed in the appended claims.

[0068] Existing password strength assessment methods rely solely on character complexity (such as length, case sensitivity, and symbols) or static rules, failing to reflect attackers' targeted guessing behavior using publicly available user information (such as email username, name, and date of birth). This invention provides a context-based deep email password strength measurement method. This method introduces user context information such as email address, name, and date of birth into a basic password measurement model framework, achieving information fusion through a multi-head self-attention mechanism, and constructing a password sequence probability calculation model using a character-level long short-term memory (LSTM) network. An evaluation system with the number of guesses as the core indicator quantifies the performance differences of different models with and without user context information. Experiments show that this method can more realistically reflect the cracking ability of attackers when they possess user context information.

[0069] Specifically, this includes: S1: Constructing a basic password generation model. A character-level long short-term memory (LSTM) network is used to train a large-scale plaintext password dataset to learn the probability distribution of password character sequences, which is used to simulate password guessing behavior under conditions without context information.

[0070] To effectively assess password strength, this invention first constructs a basic password generation probability model to simulate the behavior of an attacker guessing passwords without user context information. This model employs a character-level Long Short-Term Memory (LSTM) network, capable of capturing contextual dependencies within password character sequences, thereby learning the probability distribution of password generation.

[0071] The core architecture of the base model consists of a two-layer stacked LSTM network and a fully connected network for modeling password character sequences. The model is trained on approximately 1.4 billion plaintext credentials collected from public password leaks such as Exploit.in and Anti Public. The training objective is to maximize the likelihood function of real password sequences under the model's generation probability distribution. At each time step, the model receives the embedding vector of the current character as input, computes its hidden state through LSTM units, and outputs the conditional probability distribution of the next character using the Softmax function. After training, the model can be used to evaluate the generation probability of any password sequence; a higher probability indicates that the password is easier to guess and has lower strength. This base model does not contain any user-specific contextual information, thus providing a user-independent baseline for password strength, which helps in comparing the performance of subsequent models that consider user information.

[0072] S2: Construct a context-sensitive password strength measurement model: extract and encode user context information.

[0073] To further improve the accuracy of password strength assessment, this invention introduces user context information (email, name, date of birth) into the basic model, constructing a context-sensitive password strength measurement model. This model comprehensively models the correlation between publicly available user information and passwords, thus more realistically reflecting an attacker's password guessing ability when partial user information is known.

[0074] The overall architecture of the model is as follows Figure 1 As shown, it mainly consists of three parts: a user information encoder, a multi-head self-attention mechanism, and a context-based decoder. Its core idea is to encode user context information into vector form and influence the password generation process through an attention mechanism, thereby forming a user-specific password distribution modeling capability.

[0075] First, in the user information encoder, the user context information includes email address, user name, and date of birth. The email address is split into username, subdomain, and top-level domain, which are encoded using Word2Ve word embedding and character-level LSTM respectively. One-hot encoding is used for high-frequency domains, and hash encoding is used for low-frequency domains. The name and date of birth are encoded into dense vectors through word embedding and LSTM. All context encoding vectors are concatenated and positional encoding is added to form a unified context representation.

[0076] Specifically, the encoding methods for the aforementioned user context information include:

[0077] The username part is generated by Word2Vec word embedding, and then the sequence features are extracted by character-level LSTM;

[0078] For the domain name portion, one-hot encoding is used for the top 60% of high-frequency categories, and hash encoding is used for the rest.

[0079] Names and birthdates are encoded into fixed-dimensional vectors using word embeddings and LSTM, respectively.

[0080] All encoded vectors are concatenated and then sine and cosine positional encoding is added.

[0081] Preferably, the Word2Vec model is pre-trained on a corpus of usernames and names containing no less than 1 billion plaintext passwords to improve the quality of semantic embedding.

[0082] The hash encoding uses the MD5 hash function to map low-frequency domain names to fixed-length strings, then converts them into indexes through a lookup table, and finally maps them to an embedding vector.

[0083] The wavelength range for position encoding is set from 1 to 10000, the embedding dimension d_model=512, and it conforms to the standard configuration of sine and cosine functions.

[0084] The encoding of user context information covers multiple dimensions, including email address (split into username, subdomain, and top-level domain), name, and date of birth, employing differentiated encoding strategies for different feature types. For the username, a Word2Vec model is used to generate word embeddings, combined with a two-layer character-level LSTM network to model the character sequence, effectively capturing semantic features and spelling patterns in the username. Regarding domain name processing, the top 60% of the most frequent subdomains and top-level domains are encoded using one-hot encoding (OHE), while the remaining low-frequency categories are converted into fixed-dimensional vector representations through hash mapping, thus maintaining high information density while controlling feature dimensions. For other context information such as name and date of birth, the encoding method is similar to that of the username, converting them into dense vector representations through an embedding model.

[0085] The encoded contextual information is enhanced with positional encoding to improve its sequence information before being fed into a multi-head self-attention layer. This module automatically learns the parts of the password sequence's probability calculation process that are most relevant to the user's contextual information, thus providing contextual guidance for each character generation. The probability calculation process of the password sequence uses a character-level LSTM structure similar to the base model, but its initial state is determined by the aforementioned attention output. The model predicts the probability of the next character character by character and dynamically adjusts the prediction distribution based on contextual information. Compared to the base model, this context-sensitive model can generate password probabilities that are closer to the true distribution, even when an attacker knows some user information, thus providing a more targeted basis for password strength assessment.

[0086] To more accurately assess password security, this invention designs a context-sensitive password sequence probability calculation model by combining publicly available user information (such as email address, name, and birthday). User context information includes the encoding of email address, name, and birthday, and is fused using a multi-head self-attention mechanism. Preferably, the multi-head self-attention mechanism contains eight attention heads, each with a dimension of 64, resulting in a total context vector dimension of 512, consistent with the hidden state dimension of the LSTM. The hidden layer dimension of the character-level long short-term memory network (LSTM) is set to 512, the word embedding dimension is 128, the Adam optimizer is used during training, the learning rate is set to 0.001, and an early stopping mechanism is used to prevent overfitting.

[0087] Email addresses are broken down into username, subdomain, and top-level domain (TLD). The username is used to generate a semantic embedding vector using a Word2Vec model, and contextual features are extracted using a character-level two-layer LSTM to form an embedding representation of the username. For subdomains and top-level domains, if they belong to the top 60% of high-frequency domains in the dataset, one-hot encoding is used; otherwise, a hash function is used to map them to a fixed-length index, and a fixed-dimensional embedding vector is generated using direct encoding.

[0088] (1)

[0089] (2)

[0090] in, This represents the semantic embedding vector of the username generated by the Word2Vec model. This represents a character-level long short-term memory network. This represents a fully connected neural network layer. Indicates one-hot encoding. This represents a hash function. (It will appear later in the text.) , Both refer to the same type of structure.

[0091] Name and birthdate information are encoded using a similar method to email usernames, and processed using the Word2Vec model and character-level LSTM to generate corresponding embedding vectors.

[0092] (3)

[0093] (4)

[0094] in, , , , These represent the semantic embedding vectors for year, month, day, and user name generated by the Word2Vec model, respectively.

[0095] This information is pieced together to form a unified user context representation:

[0096] (5)

[0097] To enable the model to capture the relative structure between different user contexts, this invention introduces positional encoding based on the embedding vector. Positional encoding introduces sine and cosine functions in different dimensions, ensuring that each position in the sequence has a unique representation, thereby helping the model identify the sequential relationships of contextual information. Its calculation method is as follows:

[0098] (6)

[0099] (7)

[0100] in, This indicates that the input is in the embedding vector. The position in the middle (e.g., 0 corresponds to the email username, 1 corresponds to the subdomain, and so on). Indicates a dimension index. This is the dimension of the embedding vector. Each embedding vector... With the corresponding Adding elements together yields a representation that includes location information.

[0101] Finally, the fused sequence representation is fed into the multi-head self-attention mechanism to obtain a representation of the context information:

[0102] (8)

[0103] in, This represents a multi-head self-attention mechanism.

[0104] In order to use the fused context vector as the decoder's initial state, the model further reshapes it into the state dimension required by the decoder through a dense layer:

[0105] (9)

[0106] Context-Aware Password Decoder: In the probability calculation stage of the password sequence, the model uses a character-level LSTM decoder to predict the password character by character. The character-level LSTM decoder's initial state is obtained by mapping the context vector output from the third step through a fully connected layer. The decoder generates the password sequence character by character, predicting the probability distribution of the next character based on the current state and the previous character at each step. The fully connected layer maps the context vector output from the multi-head self-attention layer to the decoder's initial hidden state, using tanh as the activation function to ensure the initial state value is within the range [-1, 1].

[0107] The multi-head self-attention mechanism contains 8 attention heads, each with a dimension of 64, and a total context vector dimension of 512, which is consistent with the hidden state dimension of LSTM.

[0108] Unlike the base model, the decoder's initial state The password is derived from the aforementioned user context information through encoding and calculation. The decoding process uses a character-by-character generation method, assuming the password sequence is represented as... ,in, Indicates the first One character.

[0109] The decoding process for each step is as follows:

[0110] (1) Input embedding: embed the characters selected in the previous step The decoder state is taken as input, and the initial decoder state is... The start character is The subsequent decoder state is derived from the hidden state passed down from the previous time step. express:

[0111] (10)

[0112] (2) Probability prediction: using The function transforms the output of the fully connected layer into a probability distribution for each possible character:

[0113] (11)

[0114] in, and These are the output layer parameters.

[0115] (3) Character selection: Select the character with the highest probability in the current probability distribution as the output:

[0116] (12)

[0117] in, This indicates that the character with the highest probability is selected.

[0118] The above process is repeated until a stop marker is generated or the maximum password length is reached. This context-sensitive decoder utilizes global semantics extracted from user context information to influence the distribution of the entire generated sequence by initializing the state, thereby simulating an attacker's password guessing behavior under conditions of known user background information. Compared to the basic model, this structure more realistically reflects the distribution patterns of personalized passwords and has higher targeting and accuracy in prediction.

[0119] S3: The user enters their password, email address, birthday, and name. The number of password guesses is calculated. The probability of password generation is calculated based on both the base model and the context-sensitive model. Based on this, the number of guesses is calculated with and without context information, thereby quantifying the password strength.

[0120] Password strength is quantified by the Guessing Number (GN), reflecting the average number of attempts an attacker needs to make to crack a password under a specific model. Based on the base model ( ) and context-sensitive models ( The generation probability difference of the two types of guesses is calculated separately to measure the impact of the user's public context information on the predictability of the password.

[0121] (1) Probability calculation: for the cipher Calculate the generation probability of the two types of models separately:

[0122] Basic model:

[0123] (13)

[0124] in, Password The One character.

[0125] Context-sensitive models:

[0126] (14)

[0127] in, Password The One character, This indicates user context information.

[0128] (2) Number of basic model guesses: Password guessing attempts without user context information. The number of guesses is:

[0129] (15)

[0130] in, This represents the set of all possible passwords (sample space). The total number of samples, Based on the model for the first The probability of generating a password.

[0131] (3) Number of guesses in the context-sensitive model: Given the user's context information ( When, password The number of guesses is:

[0132] (16)

[0133] in, For context-sensitive models, under user information conditions, the first The probability of generating a password.

[0134] Calculation steps and examples:

[0135] Assuming sample space ,password The generation probability is as follows:

[0136] Table 1 Calculation Example

[0137]

[0138] Number of guesses for the basic model:

[0139]

[0140] Number of guesses in the context-sensitive model:

[0141]

[0142] Ultimately, password security is measured by the number of guesses. A threshold q is set. If the number of guesses for a user's input password is less than q, the user is advised to improve their password security until it exceeds q.

[0143] S4: Output password strength assessment results. Based on the comparison between the number of guesses and a preset threshold, determine the password's vulnerability in actual attack scenarios and provide the user with strength feedback and risk warnings.

[0144] To verify the differences in the number of guesses among different methods, this invention compares data collected from publicly disclosed password leaks such as Exploit.in and AntiPublic. The obtained data samples were normalized and denoised to ensure data integrity and validity. Based on this, information such as username, domain name, name, and birthday were extracted from the email field, and auxiliary features were formed that could be used by the context-sensitive model. In the experiment, the basic password sequence probability calculation model and the context-sensitive model incorporating user information completed parameter learning on the same training set, while the test set was used to test the differences in the distribution of guess counts between the two models in the password generation task, thereby quantifying the performance differences of different models with and without user context information.

[0145] Table 2 shows the number of guesses for some real email-password pairs under the two models. It can be seen that, compared to the basic model, the probability of the correct password sequence is generally higher after introducing user context information, leading to a significant reduction in the number of guesses. This indicates that when attackers have background information such as the user's email address, the actual security of the password is greatly weakened.

[16] .

[0146] Table 2 compares the probability and number of guesses for some samples under different models.

[0147]

[0148] The results in the table show that some seemingly complex or long passwords (such as #ESR%T6y7u8i(O)P containing special characters or combinations of uppercase and lowercase letters) still see a dramatic drop in the number of guesses after considering user context information. This reflects the advantage of the context-sensitive model in simulating real-world attack scenarios, and also reminds users to avoid using content highly related to personal information in their passwords.

[0149] To further reveal the relationship between the number of guesses for a single password under the two models, Figure 3 A scatter plot comparing the two was created. The horizontal axis represents the base model. The vertical axis represents the context-sensitive model. The results show that most data points are distributed in The dense distribution below the straight line, especially in the low guess count range, indicates that the context-sensitive model has higher cracking efficiency in weak password scenarios.

[0150] Figure 4A heatmap showing the difference in guess counts between the two models under different combinations of user context information types and password lengths is presented. The horizontal axis represents the user information type, the vertical axis represents the password length range, and the color depth represents the difference in guess counts. The results show that the difference is highest when the password is short (6–8 characters) and the user information is complete (email + name + birthday); the difference gradually decreases as the password length increases or the user context information decreases. This result reflects that the context-sensitive model has the most significant advantage in scenarios with rich information and short passwords, while the performance difference between the two models narrows when the password is long or the information is limited.

[0151] In summary, the context-sensitive model incorporating user information significantly improves the ability to identify and predict weak passwords, which is of great significance for improving the accuracy of password security risk assessment. It also reminds users to avoid using character combinations that are highly related to personal information when setting passwords.

[0152] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these changes and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A context-based deep email password strength measurement method, characterized in that, include: S1: Build a basic password generation model to simulate the behavior of an attacker guessing passwords without having user context information; S2: Construct a context-sensitive cryptographic strength measurement model; the model includes the following: Extract and encode user context information: Split email addresses into username, subdomain, and top-level domain, and encode them using Word2Ve word embedding and character-level LSTM, respectively; High-frequency domain names are encoded using one-hot encoding, while low-frequency domain names are encoded using hash encoding. Names and dates of birth are encoded into dense vectors using word embeddings and LSTM; all context encoding vectors are concatenated and positional encoding is added to form a unified context representation; Fusing contextual information: The context representation is input into a multi-head self-attention mechanism to extract the contextual features most relevant to password generation and output the fused context vector; Context-aware cryptographic decoder: It adopts a character-level LSTM decoder, whose initial state is obtained by mapping the context vector output in the third step through a fully connected layer; The decoder generates a password sequence character by character, and at each step predicts the probability distribution of the next character based on the current state and the previous character; S3: The user enters their password, email, birthday, and name information. The number of password guesses is calculated. The probability of password generation is calculated based on the basic password generation model and the context-sensitive password strength measurement model. The number of guesses is calculated with and without context information. S4: Output password strength assessment results. Based on the comparison between the number of guesses and the preset threshold, determine the vulnerability of the password in actual attack scenarios and provide users with strength feedback and risk warnings.

2. The context-based deep email password strength measurement method according to claim 1, characterized in that, The basic password generation model described in S1 uses a character-level Long Short-Term Memory (LSTM) network to capture the contextual dependencies in the password character sequence, thereby learning the probability distribution of password generation. The architecture of the basic cryptographic generation model consists of a two-layer stacked LSTM network and a fully connected network for modeling cryptographic character sequences; The training objective of the basic password generation model is to maximize the likelihood function of the real password sequence under the model's generation probability distribution. At each time step, the model receives the embedding vector of the current character as input, calculates its hidden state through the LSTM unit, and outputs the conditional probability distribution of the next character using the Softmax function. After training, the model is used to evaluate the probability of generating any password sequence. The higher the probability, the easier the password is to guess and the lower its strength.

3. The context-based deep email password strength measurement method according to claim 1, characterized in that, The encoding methods for user context information in S2 include: The username is used to generate a semantic embedding vector through the Word2Vec model, and the context features are extracted through a character-level two-layer LSTM to form an embedding representation of the username. For subdomains and top-level domains, if they belong to the top 60% of high-frequency domains in the dataset, one-hot encoding is used; otherwise, a hash function is used to map them to a fixed-length index, and a fixed-dimensional embedding vector is generated through direct encoding. (1) (2) in, This represents the semantic embedding vector of the username generated by the Word2Vec model. This represents a character-level long short-term memory network. This represents a fully connected neural network layer. Indicates one-hot encoding. Represents a hash function; Name and birthdate information are encoded using a similar method to email usernames, and processed using the Word2Vec model and character-level LSTM to generate corresponding embedding vectors; (3) (4) in, , , , These represent the semantic embedding vectors for year, month, day, and user name generated by the Word2Vec model, respectively. The above information is concatenated to form a unified user context representation: (5) Positional encoding is introduced on top of the embedding vector. Positional encoding uses sine and cosine functions in different dimensions to ensure that each position in the sequence has a unique representation. The calculation method is as follows: (6) (7) in, This indicates that the input is in the embedding vector. The position in the middle, Indicates a dimension index. It is the dimension of the embedding vector; each embedding vector With the corresponding Adding elements together yields a representation that includes location information. Finally, the fused sequence representation is fed into the multi-head self-attention mechanism to obtain a representation of the context information: (8) in, This represents a multi-head self-attention mechanism.

4. The context-based deep email password strength measurement method according to claim 1, characterized in that, The context-aware cryptographic decoder includes: the decoder's initial state. The password is calculated by encoding user context information, and the decoding process uses a character-by-character generation method. Let the password sequence be represented as... ,in, Indicates the first Each character; the decoding process for each step is as follows: (1) Input embedding: embed the characters selected in the previous step The decoder state is taken as input, and the initial decoder state is... The start character is The subsequent decoder state is derived from the hidden state passed from the previous time step. express: (10) (2) Probability prediction: using The function transforms the output of the fully connected layer into a probability distribution for each possible character: (11) in, and These are the output layer parameters; (3) Character selection: Select the character with the highest probability in the current probability distribution as the output: (12) in, This indicates that the character with the highest probability is selected; The above process is repeated until a stop marker is generated or the maximum password length is reached. .

5. The context-based deep email password strength measurement method according to claim 1, characterized in that, In S3, based on the base model With context-sensitive models The generation probability difference was calculated, and the number of guesses for the two types were calculated separately to measure the impact of the user's public context information on password predictability. (1) Calculate the probability: for the password Calculate the generation probabilities of the two types of models separately: Basic model: (13) in, Password The One character; Context-sensitive models: (14) in, Password The One character, Indicates user context information; (2) Number of basic model guesses: Password guessing attempts without user context information. The number of guesses is: (15) in, This represents the set of all possible passwords (sample space). The total number of samples, Based on the model for the first The probability of generating a password; (3) Number of guesses in the context-sensitive model: Given the user's context information ( When, password The number of guesses is: (16) in, For context-sensitive models, under user information conditions, the first The probability of generating a password.

Citation Information

Cited By

  • Dynamic password strength evaluation method for Internet of Things equipment

    CN122053048A

  • A dynamic password strength evaluation method for internet of things devices

    CN122053048B