Method and device for evaluating password strength in real time
By using polynomial naive Bayes model and feature extraction technology, the problem of difficult to identify complex password patterns and potential security risks in the prior art is solved, and a more accurate and comprehensive evaluation of cardiac password strength is achieved, improving password security.
Patent Information
- Application Number
- CN202510091481.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
Existing password strength evaluation methods are difficult to effectively identify complex password patterns and potential security risks, especially when facing deformed weak passwords, the evaluation results are quite different from the actual security.
By using a polynomial naive Bayes model, combining cracked weak password samples and validated strong password samples from the cipher dataset, a model of assessment that can identify complex patterns of passwords and security risks is trained. This model extracts character-level features and semantic features of the user input password, generates feature vectors, and uses the polynomial naive Bayes model to score to obtain a password strength score.
It realizes effective identification of complex password patterns and potential security risks, improves the accuracy and comprehensiveness of password strength evaluation, and can promptly discover potential security risks in passwords, help users improve password security.
Smart Images

Figure CN120046139A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer security technology, and particularly to a method and device for real-time evaluation of password strength. Background Art
[0002] With the development of the information age, the number of various network accounts of users has increased sharply, and password security has become an important barrier to protecting users' digital assets and privacy. To ensure account security, it is necessary to perform real-time evaluation of password strength when users set passwords to help users create more secure passwords.
[0003] Currently, common password strength evaluation methods are mainly based on rule matching and scoring mechanisms. For example, the password strength is evaluated by checking rules such as password length, whether it contains uppercase and lowercase letters, numbers, and special characters, or the complexity of the password is calculated using information entropy methods such as Shannon entropy. These methods are simple to implement and have high computational efficiency, and are widely used in practical applications.
[0004] Another type of method uses machine learning algorithms, such as decision trees or support vector machines, etc., to train through known password data sets to establish a password strength evaluation model. This type of method can better identify potential patterns and security risks in passwords by learning password features in historical data.
[0005] However, existing password strength evaluation methods often rely too much on predefined rules or a single feature dimension, and it is difficult to effectively identify complex password patterns and potential security risks. Especially when facing various deformed weak passwords, such as passwords constructed by using similar-shaped character replacements, keyboard combination patterns, etc., there is a large deviation between the evaluation results of existing methods and the actual password security. Summary of the Invention
[0006] In view of this, this application provides a method and device for real-time evaluation of password strength, which solves the problem in the prior art that it is difficult to effectively identify complex password patterns and potential security risks.
[0007] An embodiment of this application provides a method for real-time evaluation of password strength, including:
[0008] Training a multinomial naive Bayes model through a password data set, where the password data set includes cracked weak password samples and verified strong password samples;
[0009] Responding to a trigger event of a user input password, and extracting character-level features and semantic features of the user input password;
[0010] Generating sub-feature vectors according to the character-level features and the semantic features respectively, and weighting the sub-feature vectors to obtain a feature vector;
[0011] Score the feature vector using the trained multinomial Naive Bayes model to obtain a password strength score;
[0012] Determine the password strength level according to the password strength score.
[0013] Optionally, training the multinomial Naive Bayes model with a password dataset includes:
[0014] Learn the importance of different features in the multinomial Naive Bayes model according to the password dataset, and determine a feature subset with a high contribution degree to the multinomial Naive Bayes model;
[0015] Use the feature subset as the input features of the multinomial Naive Bayes model, and train the multinomial Naive Bayes model based on the input features;
[0016] Adjust the parameters of the multinomial Naive Bayes model through cross-validation to balance accuracy and recall.
[0017] Optionally, the learning the importance of different features in the multinomial Naive Bayes model according to the password dataset, and determining a feature subset with a high contribution degree to the multinomial Naive Bayes model includes:
[0018] Clean the password dataset;
[0019] Use the principal component analysis method to reduce the feature dimension of the password dataset to obtain a password dataset with reduced dimension;
[0020] Calculate the information gain ratio of each feature in the password dataset with reduced dimension;
[0021] Sort the features in the password dataset with reduced dimension based on the information gain ratio;
[0022] Use the mutual information method to evaluate the feature correlation degree of each feature in the password dataset with reduced dimension;
[0023] From the sorted features, initially screen out the features with an information gain ratio greater than a preset threshold, and perform secondary screening according to the feature correlation degree to obtain a feature subset with a high contribution degree to the model.
[0024] Optionally, the scoring the feature vector using the trained multinomial Naive Bayes model to obtain a password strength score includes:
[0025] Calculate the basic security score of the feature vector;
[0026] Based on preset security rules, perform weighted adjustment on the basic security score to obtain an adjusted security score;
[0027] Calculate the probability that the password is cracked based on the adjusted security score;
[0028] Calibrate the probability that the password is cracked based on the feature relevance of the feature subset to obtain a calibrated password strength score.
[0029] Optionally, the real-time password strength evaluation method further includes:
[0030] Dynamically adjust the scoring threshold of the calibrated password strength score according to the historical evaluation accuracy;
[0031] Use smoothing processing to avoid drastic fluctuations in the calibrated password strength score.
[0032] Optionally, the triggering event in response to the user inputting a password includes:
[0033] Set a password input listener to capture each character input by the user in real time;
[0034] When it is detected that the user pauses input for more than a preset time threshold or presses the Enter key, trigger a password evaluation event;
[0035] Store the password input by the user in a temporary cache for feature extraction.
[0036] Optionally, the step of weighting the sub-feature vectors to obtain the feature vector includes:
[0037] Statistically analyze the contribution of each feature to password strength prediction in historical evaluation data;
[0038] Dynamically adjust the character-level feature weight and semantic feature weight according to the size of the contribution;
[0039] Normalize the character-level feature weight and semantic feature weight so that the sum of all weights is 1;
[0040] Weight the sub-feature vectors according to the character-level feature weight and semantic feature weight to obtain the feature vector.
[0041] Optionally, the step of dynamically adjusting the character-level feature weight and semantic feature weight according to the size of the contribution includes:
[0042] Set a weight adjustment period to recalculate the feature weights at the end of each period;
[0043] Record the occurrence frequency of the character-level feature and semantic feature in correct predictions within each period;
[0044] Update the character-level feature weight and semantic feature weight using a weighted average algorithm according to the importance score and occurrence frequency of the feature.
[0045] Optionally, extract the character-level features of the user-entered password, including:
[0046] Count the number of uppercase letters, lowercase letters, digits, and special characters in the password, and calculate the distribution ratio of each type of character in the password; and / or,
[0047] Identify the continuously increasing or decreasing character sequences in the password, and record the length of the character sequences; and / or,
[0048] Detect the repeatedly occurring characters or character combinations in the password, and count the number of repetitions; and / or,
[0049] Calculate the total length of the password, and calculate the proportion of each type of character in the total length;
[0050] Extract the semantic features of the user-entered password, including:
[0051] Set the N value range of the N-gram model to 2 to 5, and perform word segmentation analysis on the password; and / or,
[0052] Match the password with a predefined common word dictionary, and record the number and positions of the matched words; and / or,
[0053] Detect the leetspeak encoding method in the password, including detecting the situation where letters are replaced by similar digits or special characters; and / or,
[0054] Identify common digital combination patterns such as date formats and phone number formats in the password.
[0055] An embodiment of the present application further provides a real-time password strength evaluation device, including:
[0056] A model training unit, configured to train a multinomial naive Bayes model through a password data set, where the password data set includes cracked weak password samples and verified strong password samples;
[0057] A feature extraction unit, configured to respond to a trigger event of the user-entered password, extract the character-level features and semantic features of the password, and generate a feature vector;
[0058] A vector generation unit, configured to generate sub-feature vectors according to the character-level features and the semantic features respectively, and weight the sub-feature vectors to obtain a feature vector;
[0059] A scoring execution unit, configured to score the feature vector using the multinomial naive Bayes model to obtain a password strength score;
[0060] A level determination unit, configured to determine a password strength level according to the password strength score.
[0061] An embodiment of the present application further provides a computer device, where the computer device includes:
[0062] At least one processor; and,
[0063] A memory communicatively connected to the at least one processor; wherein,
[0064] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned real-time password strength evaluation method.
[0065] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the above-mentioned real-time password strength evaluation method.
[0066] An embodiment of the present application further provides a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the steps of the above-mentioned real-time password strength evaluation method are implemented.
[0067] In an embodiment of the present application, a polynomial naive Bayes model is trained by using a data set including cracked weak password samples and verified strong password samples, and the recognition ability and evaluation accuracy of the model are improved by using real data. In an embodiment of the present application, character-level features and semantic features can be extracted in real time when a user inputs a password, and a feature vector is formed through a dual feature extraction and weighted fusion mechanism, so that the system can comprehensively analyze the composition characteristics and potential rules of the password. In the scoring stage, the non-linear characteristics of the polynomial naive Bayes model are used to process the correlation relationship between features, and finally the scoring result is converted into an easily understandable password strength level. The combination of these technical features not only improves the accuracy and comprehensiveness of the evaluation, but also ensures the real-time nature of the evaluation process and the practicality of the result, and can help users timely discover potential security risks in the password. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the disclosed embodiments in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0069] Figure 1 It is a schematic flowchart of the real-time password strength evaluation method provided by the embodiment of the present application;
[0070] Figure 2Schematic flowchart of the polynomial naive Bayes model training method provided by an embodiment of the present application;
[0071] Figure 3 Schematic flowchart of the feature extraction process provided by an embodiment of the present application;
[0072] Figure 4 Schematic flowchart of the password strength scoring process provided by an embodiment of the present application;
[0073] Figure 5 Schematic diagram of the structure of the password strength real-time evaluation device provided by an embodiment of the present application. Detailed implementation manners
[0074] To make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments.
[0075] As Figure 1 shown, the password strength real-time evaluation method provided by an embodiment of the present application includes:
[0076] S1: Train a polynomial naive Bayes model through a password data set;
[0077] Exemplarily, as Figure 2 shown, training a polynomial naive Bayes model through a password data set specifically includes:
[0078] S1.1: Learn the importance of different features in the polynomial naive Bayes model according to the password data set, and determine a feature subset with a high contribution degree to the model;
[0079] In specific implementation, it is first necessary to train a polynomial naive Bayes model to build the basis for password strength evaluation. During the training process, the password data set used contains two types of samples: weak password samples that have been cracked and strong password samples that have been verified. Weak password samples can be obtained from publicly available password leakage databases, and these passwords usually have low security and high predictability; strong password samples come from a set of high-strength passwords certified by a security assessment agency, and these passwords usually have high complexity and low predictability.
[0080] Specifically, S1.1 includes:
[0081] S1.1.1: Clean the password data set;
[0082] Cleaning the password data set includes operations such as removing duplicate samples, handling outliers, and unifying character encoding to ensure the quality and consistency of the data.
[0083] S1.1.2: Use the principal component analysis method to reduce the feature dimension of the password data set to obtain a password data set with reduced dimensions;
[0084] The principal component analysis (PCA) method is used to reduce the dimensionality of the feature space, which can not only reduce the computational complexity but also reduce the redundancy between features.
[0085] S1.1.3: Calculate the information gain ratio of each feature in the password dataset after dimensionality reduction;
[0086] The information gain ratio can measure the contribution degree of features to password strength judgment. Specifically, for feature X and password strength category Y, calculate the conditional entropy H(Y|X) and the information gain ratio G(Y,X) to evaluate the importance of feature X for judging password strength Y.
[0087] S1.1.4: Sort the features in the password dataset after dimensionality reduction based on the information gain ratio;
[0088] Based on the calculated information gain ratio, sort all features in descending order.
[0089] S1.1.5: Use the mutual information method to evaluate the feature correlation degree of each feature in the password dataset after dimensionality reduction;
[0090] Mutual information can measure the mutual dependence degree between two variables, which can not only capture linear correlation but also identify non - linear correlation. By calculating the mutual information value between features, redundant features with high correlation (feature correlation degree greater than a certain threshold) can be avoided.
[0091] S1.1.6: From the sorted features, initially screen out the features whose information gain ratio is greater than the preset threshold, and conduct a secondary screening according to the feature correlation degree to obtain a feature subset with a high contribution degree to the model.
[0092] From the sorted features, initially screen out the features whose information gain ratio is greater than the preset threshold (adjusted according to the actual application scenario and security requirements). The initially screened features may contain redundant features with high correlation. Then, through secondary screening according to the feature correlation degree, only one of the multiple highly correlated features can be retained, and thus the final feature subset can be obtained.
[0093] S1.2: Use the feature subset as the input features of the multinomial naive Bayes model and train the model based on the input features;
[0094] After determining the feature subset, use these features to train the multinomial naive Bayes model.
[0095] S1.3: Adjust the parameters of the multinomial naive Bayes model through cross - validation to balance the accuracy rate and the recall rate.
[0096] In the model parameter adjustment stage, the embodiment of the present invention adopts the k-fold cross-validation method to optimize the parameters of the polynomial naive Bayes model. Specifically, first, the password dataset is randomly divided into k subsets with similar sizes. Each time, k-1 of these subsets are selected as the training set, and the remaining 1 subset is used as the validation set. The model performance is evaluated through multiple trainings and validations. Usually, the value of k is 5 or 10, which can achieve a balance between ensuring the sufficiency of validation and computational efficiency.
[0097] In each round of cross-validation, the embodiment of the present invention focuses on the smoothing parameter alpha and the feature weight coefficient of the model. The smoothing parameter alpha is used to handle the zero-probability problem, and its value range is usually between 0.1 and 10. A smaller alpha value makes the model more dependent on the distribution characteristics of the training data, while a larger alpha value increases the generalization ability of the model. The feature weight coefficient is used to adjust the importance of different features in the prediction process, and the initial value can be set to a uniform distribution.
[0098] To balance accuracy and recall, the embodiment of the present invention adopts a grid search strategy to find the optimal parameter combination in the preset parameter space. The evaluation criterion uses the F1 score, which takes both accuracy and recall into account. The calculation formula is 2×(accuracy×recall) / (accuracy + recall). By comparing the F1 scores under different parameter combinations, the parameter combination that can maximize the F1 score is selected as the final model parameter.
[0099] Considering the particularity of password strength evaluation, the embodiment of the present invention pays special attention to avoiding false negatives of strong passwords during the parameter adjustment process. By introducing a weighted F1 score and assigning higher weights to strong password samples, the model can more accurately identify high-strength passwords while maintaining overall balance. In addition, learning curve analysis is also used to monitor the training process of the model to timely detect and solve overfitting or underfitting problems.
[0100] The model parameters optimized through cross-validation can not only ensure that the model has a high accuracy in identifying weak passwords but also ensure good recognition ability for strong passwords. This parameter optimization method enables the model to effectively warn of potential security risks in practical applications and avoid overly interfering with users' password setting behaviors.
[0101] S2: Respond to the trigger event of the user inputting a password;
[0102] S2 can specifically include:
[0103] S2.1: Set a password input listener to capture each character input by the user in real time;
[0104] In the embodiments of the present application, a password input listener is used to capture the user's input behavior in real time. For example, a character listener is set in the password input box. When the user inputs any character during the input process, the listener can immediately capture this input event. The listener can not only capture the input of ordinary letters, numbers and other characters, but also capture the input operations of special characters, spaces and backspace keys, etc. This real-time capture mechanism provides basic data support for subsequent password strength evaluation.
[0105] S2.2: When it is detected that the user pauses input for more than a preset time threshold or presses the enter key, trigger a password evaluation event;
[0106] Exemplarily, a preset time threshold is set, such as 1 - 2 seconds. When the time that the user pauses input during the input process exceeds this threshold, it is considered that the user has completed the input of the current password segment, and at this time, a password evaluation will be triggered. At the same time, if the user directly presses the enter key, the evaluation event will also be triggered immediately. This mechanism can not only achieve real-time evaluation, but also will not affect the performance of the embodiments of the present application due to overly frequent evaluation.
[0107] S2.3: Store the password input by the user in a temporary cache for feature extraction.
[0108] The password input by the user is temporarily stored in the memory cache instead of being directly written to the persistent storage. This temporary storage not only ensures the security of the password information, but also provides convenience for subsequent feature extraction. The feature extraction module can directly read the password string from the cache and analyze various features contained therein, such as the distribution of character types, repeated patterns, etc. After the feature extraction is completed, the embodiments of the present application will promptly clear the cache to ensure that the password information will not be leaked.
[0109] In addition, in order to protect user privacy and data security, this cache adopts encrypted storage and is immediately cleared after the feature extraction is completed. This mechanism ensures that the password information will not be retained in the system for a long time, effectively reducing the security risk.
[0110] S3: Extract password features and generate a feature vector;
[0111] Specifically, as Figure 3 shown, extracting password features and generating a feature vector includes:
[0112] S3.1: Extract character-level features;
[0113] S3.1.1: Count the number of uppercase letters, lowercase letters, numbers and special characters in the password, and calculate the distribution ratio of each type of character in the password;
[0114] In character-level feature extraction, first, the distribution of various types of characters in the password is counted, including the number of uppercase letters, lowercase letters, digits, and special characters, as well as their distribution ratios in the password. These basic statistical features can intuitively reflect the compositional complexity of the password. It should be noted that the range of special characters includes, but is not limited to, punctuation marks, mathematical symbols, and other printable characters in the ASCII code.
[0115] Suppose the password entered by the user is "ABC123!@#". Then, it is counted that there are 3 uppercase letters (A, B, C), 3 digits (1, 2, 3), 3 special characters (!, @, #), and 0 lowercase letters. The uppercase letters are concentrated at the beginning of the password, the digits are in the middle, and the special characters are at the end, showing obvious block characteristics.
[0116] S3.1.2: Identify the continuously increasing or decreasing character sequences in the password and record the lengths of the character sequences;
[0117] Suppose the password entered by the user is "ABC123!@#". The total length of this password is 9 characters. S3.1.3: Detect the repeatedly occurring characters or character combinations in the password and count the number of repetitions;
[0118] Specifically, continuously increasing or decreasing character sequences in the password, such as "abc", "123", etc., are identified by means of a sliding window, and the lengths of these sequences are recorded. Such sequences usually represent consecutive inputs on the keyboard layout or simple alphanumeric sequences, and these patterns are easily recognized by cracking tools. At the same time, repeatedly occurring characters or character combinations in the password, such as repeated patterns like "aaa", "123123", etc., are also detected. These repeated patterns will reduce the entropy value of the password and increase the risk of being cracked.
[0119] S3.1.4: Calculate the total length of the password and calculate the proportion of each type of character in the total length;
[0120] The calculation of the total length of the password is relatively straightforward, which is to count the total number of characters in the password string. For example, the total length of the password "P@ssw0rd#2023" is 12 characters. When calculating the length, all visible characters are considered, including letters, digits, special characters, spaces, etc., but hidden control characters are not included. The total length is a basic indicator for evaluating password strength. Generally, the longer the length, the higher the strength of the password.
[0121] When calculating the proportion of each type of character, first classify and count the characters in the password. Taking the above password as an example: there is 1 capital letter (P), accounting for 8.33%; there are 5 lowercase letters (s, s, w, r, d), accounting for 41.67%; there are 5 digits (0, 2, 0, 2, 3), accounting for 41.67%; there are 2 special characters (@, #), accounting for 16.67%. These proportion data can reflect the character diversity of the password, and a reasonable distribution of character types is very important for improving password strength.
[0122] The significance of calculating these proportions lies in: First, it can reflect whether the distribution of character types in the password is balanced. Over-reliance on a certain type of character will reduce password strength; Second, the proportions of each type of character can be important dimensions of the feature vector for subsequent password strength evaluation; Finally, these proportion data can also be used to provide improvement suggestions for users, such as prompting users to increase the usage proportion of a certain type of character.
[0123] By combining the total length and the analysis of the proportion of character types, not only can the complexity of the password be evaluated, but also potential weaknesses can be identified. For example, even if the password is very long, if the vast majority of characters are of the same type (such as all lowercase letters), the actual strength of the password is still not high. On the contrary, a shorter password with a balanced distribution of character types may get a better score.
[0124] In an example, the password entered by the user is "ABC123!@#". First, extract the character type features. Count that there are 3 capital letters (A, B, C), 3 digits (1, 2, 3), 3 special characters (!, @, #), and 0 lowercase letters. Second, extract the position features. The capital letters are concentrated at the beginning of the password, the digits are in the middle, and the special characters are at the end, showing an obvious block feature. Third, extract the length feature. The total length of this password is 9 characters.
[0125] In addition, deeper features can also be extracted. For example, the keyboard pattern feature, analyze that "ABC" is a continuous letter sequence, "123" is a continuous digit sequence, and "!@#" is a continuous special character sequence on the keyboard. Another example is the character repetition feature. There are no repeated characters in this password. And the character conversion feature, such as the number of conversions from letters to digits, from digits to special characters, etc.
[0126] Finally, convert these features into a feature vector. A possible form is: [3, 0, 3, 3, 9, 2, 0, 2], where each dimension represents: the number of capital letters, the number of lowercase letters, the number of digits, the number of special characters, the total length, the number of continuous sequences, the number of repeated characters, and the number of character type conversions. Such a feature vector is convenient for subsequent password strength evaluation algorithms to process.
[0127] For another password example "Pa$$w0rd", its characteristics will be very different. It contains 1 capital letter, 4 lowercase letters, 1 digit, and 2 repeated special characters. This password has no obvious chunking characteristics, and the character types are interspersed. It does not contain consecutive sequences, but has repeated characters ($). Converting these characteristics into a vector might be: [1, 4, 1, 2, 8, 0, 1, 4], which is significantly different from the feature vector of the previous password.
[0128] S3.2: Extract semantic features;
[0129] S3.2.1: Set the N value range of the N-gram model from 2 to 5, and perform word segmentation analysis on the password;
[0130] By setting the N value range from 2 to 5, character combination patterns of different lengths can be identified. For example, when N = 2, bigrams like "pa", "ss" can be captured; when N = 3, trigrams like "pas", "123" can be identified. This multi-scale N-gram analysis helps to discover potential semantic units in the password.
[0131] S3.2.2: Match the password with a predefined common word dictionary, and record the number and positions of the matched words;
[0132] The predefined common word dictionary contains common English words, various types of nouns (such as personal names, place names), and common abbreviations, etc. Fuzzy matching the password with the entries in the dictionary not only records the number of matched words, but also records the position information of these words in the password. This analysis helps to discover weak passwords constructed with simple word combinations.
[0133] S3.2.3: Detect the leetspeak encoding method in the password, including detecting the situation where letters are replaced by similar-looking digits or special characters;
[0134] Leetspeak is a common character replacement encoding method where users often replace letters with similar-looking digits or special characters, such as replacing "A" with "4", "E" with "3", etc. By establishing a character replacement mapping table, such deformed words can be identified, thus more accurately evaluating the actual strength of the password.
[0135] S3.2.4: Identify common digital combination patterns in the password, such as date formats, phone number formats, etc.
[0136] Use regular expressions to match various common date formats (such as "YYYYMMDD", "DDMMYYYY", etc.) and number combinations (such as phone numbers, postal codes, etc.). Although these patterns increase the apparent complexity of the password, they actually have obvious semantic features and are easily recognized by specialized cracking tools.
[0137] S3.3: Feature vector generation;
[0138] During the generation of feature vectors, a dynamic weight adjustment mechanism is adopted. Specifically, the contribution of each feature to password strength prediction is statistically analyzed in historical evaluation data, and the occurrence frequency of each feature in correct predictions is calculated. This statistical process continues within a preset weight adjustment period to ensure that the weights can be continuously optimized as new data accumulates. Using the weighted average algorithm, the weights of character-level features and semantic features are dynamically updated based on the importance scores and historical performance of the features. To ensure the rationality of the weights, the updated weights are normalized so that the sum of all feature weights remains 1.
[0139] S3.3.1: Statistically analyze the contribution of each feature to password strength prediction in historical evaluation data;
[0140] Collect a large amount of historical password evaluation data and analyze the correlation between each feature and the final password strength through data mining. For example, it is found that features such as password length and character diversity are strongly correlated with password strength, while the impact of some features is relatively small. This analysis based on historical data can help more accurately evaluate the actual contributions of each feature.
[0141] S3.3.2: Dynamically adjust the weights of character-level features and semantic features according to the magnitude of the contribution;
[0142] According to the contribution analysis results obtained in the previous step, different weights are assigned to character-level features (such as character type, length, etc.) and semantic features (such as common phrases, keyboard patterns, etc.). The weight assignment is dynamic and will be continuously updated with new evaluation data. For example, if it is found that semantic features contribute more to the accuracy of predicting password strength, the weights of semantic features are correspondingly increased.
[0143] S3.3.3: Normalize the weights of character-level features and semantic features so that the sum of all weights is 1;
[0144] Sum up the weight values of all features, and then divide the original weight of each feature by the sum, so that the sum of all weights is equal to 1. This normalization process not only makes the weights easier to understand and compare, but also avoids numerical calculation problems caused by overly large or small weights.
[0145] S3.3.4: According to the character-level feature weights and semantic feature weights, weight the sub-feature vectors to obtain the feature vectors.
[0146] Multiply the normalized weights with the corresponding character-level features and semantic features respectively to obtain the weighted feature values. These weighted feature values together constitute the final feature vector. This feature vector comprehensively considers the importance of each feature and can more accurately represent the overall features of the password, providing a reliable basis for subsequent strength evaluation.
[0147] Take the password "P@ssw0rd2023#" as an example to illustrate the generation process of the feature vector.
[0148] First, extract the character-level features of this password: The total length of the password is 12 characters, including 1 uppercase letter (P), 5 lowercase letters (sswrd), 4 digits (2023), and 2 special characters (@#); the character type conversion ratio is 0.33 (4 type conversions occur among 12 characters); the repeated character ratio is 0.25 (the letter s is repeated 1 time, accounting for the ratio of the total length); the number of consecutive sequences is 2 (2023 is a consecutive digit sequence). Accordingly, generate the character-level feature vector [12, 1, 5, 4, 2, 0.33, 0.25, 2].
[0149] Next, extract the semantic features of this password: Through N-gram analysis (N = 2 to 5), it is found that this password contains a variant "P@ssw0rd" of the common word "password"; identify the year pattern "2023" and the special character ending pattern "#"; detect the leetspeak encoding method, that is, replace "a" with "@" and "o" with "0"; find 1 keyboard sequence pattern. Thus, generate the semantic feature vector [1, 1, 2, 1], and each dimension represents: the number of common word variants, the number of digit patterns, the number of leetspeak replacements, and the number of keyboard patterns respectively.
[0150] According to the statistical results of historical evaluation data, the weight of the character-level features is set to 0.6, and the weight of the semantic features is set to 0.4. After normalizing the two sub-feature vectors respectively, multiply them with the corresponding weights and combine them. Finally, obtain a 9-dimensional feature vector [12, 1, 5, 4, 2, 0.33, 0.25, 2, 1]. The first 8 dimensions of this vector come from the character-level features (multiplied by the weight of 0.6), and the last 1 dimension comes from the semantic features (rounded after multiplying by the weight of 0.4), where the values of each dimension of the semantic features are weighted and combined to obtain the comprehensive index of the number of keyboard patterns. This final feature vector contains both the character composition information of the password and reflects the semantic patterns contained therein, and can be used as the input for subsequent strength evaluation.
[0151] S4: Score the feature vector using the trained multinomial Naive Bayes model;
[0152] Suppose a user enters a password "P@ssw0rd2023#", which is converted into a feature vector. This vector may contain data in the following dimensions: [12, 1, 5, 4, 2, 0.33, 0.25, 2, 1], where each dimension represents: total length (12 characters), number of uppercase letters (1 'P'), number of lowercase letters (5'ssw rd'), number of digits (4 '2023'), number of special characters (2 '@#'), character type conversion ratio (0.33), repeated character ratio (0.25, repeated's'), number of consecutive sequences (2, '2023' is consecutive digits), number of keyboard patterns (1).
[0153] Then, the multinomial Naive Bayes model calculates the probabilities that the password belongs to different strength levels. Suppose the password strength is divided into four levels: weak, medium, strong, very strong. The model calculates the probabilities that the password belongs to each level, for example: weak (0.05), medium (0.15), strong (0.55), very strong (0.25). These probabilities are calculated based on the feature distribution rules learned by the model from a large amount of training data.
[0154] The model considers the correlations between features during calculation. For example, although the password contains bonus items such as uppercase letters, lowercase letters, digits, and special characters, since it contains the variant "P@ssw0rd" of the common word "password" which is a deduction item, the final score will be reduced accordingly. At the same time, the combination of the common year "2023" and special characters appended to the password, "2023#", will also be recognized by the model and the score will be appropriately reduced.
[0155] Based on the above probability distribution, the password will finally be rated as the "strong" level because the probability of this level (0.55) is the highest. At the same time, improvement suggestions can also be provided to the user based on this probability distribution, such as suggesting avoiding using common word variants and increasing the randomness of the password. This probability-based scoring method can comprehensively consider various features of the password and give a relatively objective strength assessment.
[0156] Specifically, as Figure 4 shown, scoring the feature vector using the trained multinomial Naive Bayes model includes:
[0157] S4.1: Calculate the basic security score of the feature vector;
[0158] Based on Bayes' theorem, the model calculates the posterior probabilities of the password belonging to each strength category and selects the category with the highest probability as the preliminary judgment result. It should be noted that the model takes into account the non-linear relationships between features during the calculation process, and this characteristic enables it to better handle complex feature combinations in passwords.
[0159] S4.2: Based on preset security rules, perform weighted adjustment on the basic security score to obtain the adjusted security score;
[0160] To improve the accuracy of scoring, a security rule weighted adjustment mechanism is introduced on the basis of the basic security score. These security rules are formulated based on professional knowledge and practical experience in the field of password security, including password minimum length requirements, required character type combinations, common weak password pattern checks, etc. Each rule has a corresponding weight coefficient, and these coefficients reflect the degree of influence of the rule on password security. Through the weighted adjustment of these rules, the actual security of the password can be evaluated more comprehensively.
[0161] S4.3: Calculate the probability of the password being cracked based on the adjusted security score;
[0162] When calculating the probability of the password being cracked, a statistical model based on historical data is adopted. Specifically, maintain a password cracking difficulty database that records the resistance performance of different types of passwords under various cracking methods. By matching the characteristics of the current password with the records in the database, the probability of the password being cracked in the face of different attack methods can be estimated. This method takes into account the actual attack scenarios and makes the evaluation results more practical.
[0163] S4.4: Calibrate the probability of the password being cracked based on the feature correlation of the feature subset to obtain the calibrated password strength score;
[0164] To improve the reliability of scoring, the initially calculated cracking probability can be calibrated based on the feature correlation of the feature subset. This process mainly considers the mutual influence between features and avoids scoring bias caused by feature overlap. For example, when the password contains multiple highly correlated features, appropriately reduce the combined influence of these features to prevent the score from being overly magnified.
[0165] S4.5: Dynamically adjust the scoring threshold of the calibrated password strength score according to the historical evaluation accuracy;
[0166] In practical applications, the scoring threshold is dynamically adjusted according to the historical evaluation accuracy. Specifically, regularly count the degree of compliance between the evaluation results and the actual password security. When it is found that there is a systematic deviation in the scoring, the scoring threshold will be automatically adjusted to improve the accuracy. This adaptive mechanism can continuously optimize the scoring criteria and adapt to the changing password security situation.
[0167] S4.6: Apply smoothing to avoid drastic fluctuations in the password strength score after calibration;
[0168] To avoid drastic fluctuations in the scoring results, a smoothing mechanism is adopted. In the specific implementation, smoothing algorithms such as exponential moving average are used to perform weighted averaging on the results of consecutive multiple evaluations to obtain the final password strength score. This smoothing not only retains the sensitivity of the scoring but also avoids scoring instability caused by temporary factors.
[0169] S5: Determine the password strength level based on the password strength score.
[0170] In the specific implementation, the password strength is divided into multiple levels, such as extremely weak, weak, medium, strong, extremely strong, etc. Each level corresponds to a score interval, and the division of these intervals is based on a large amount of experimental data and professional security standards. The final password strength score is mapped to the corresponding level, and the evaluation results are presented to the user in an intuitive way (such as color, level description, etc.).
[0171] It should be noted that the present application also provides a corresponding device implementation. As Figure 5 shown, the device includes a model training unit, a feature extraction unit, a vector generation unit, a scoring execution unit, and a level determination unit. These functional units can be implemented in software, in hardware, or in a combination of software and hardware. The various functional units are communicatively connected via a data bus to jointly implement each step of the above password strength real-time evaluation method.
[0172] In the specific implementation process, the password strength real-time evaluation device provided by the present application can exist in the form of an independent device or can be integrated into an existing computer system, mobile terminal, or server. To better illustrate the working principle of the device, the following will be combined with Figure 5 to introduce in detail the specific implementation methods of each functional unit.
[0173] The model training unit 10 is used to train a polynomial naive Bayes model through a password data set, and the password data set includes cracked weak password samples and verified strong password samples;
[0174] The feature extraction unit 20 is used to respond to the trigger event of the user input password, extract the character-level features and semantic features of the password, and generate a feature vector;
[0175] The vector generation unit 30 is used to generate sub-feature vectors according to the character-level features and the semantic features respectively, and weight the sub-feature vectors to obtain a feature vector;
[0176] A scoring execution unit 40 is configured to score the feature vector using the polynomial naive Bayes model to obtain a password strength score;
[0177] A level determination unit 50 is configured to determine a password strength level according to the password strength score.
[0178] Specifically, the model training unit may adopt a distributed computing architecture to distribute the processing tasks of a large-scale password data set to multiple computing nodes for parallel execution, thereby improving the training efficiency. In addition, the unit may also adopt an incremental learning mechanism, which can continuously learn feature patterns from new password data without losing the existing knowledge, enabling the model to adapt to the evolving password construction methods.
[0179] The feature extraction unit adopts a multi-threaded design and can simultaneously execute the extraction tasks of character-level features and semantic features. The unit maintains a feature extraction task queue and, when receiving a password input event, immediately adds the feature extraction task to the queue. Through a task scheduling algorithm, it can reasonably utilize computing resources while ensuring real-time response. In addition, the unit may also adopt a feature caching mechanism. For frequently occurring feature patterns, the results can be directly obtained from the cache, further improving the processing efficiency.
[0180] The vector generation unit implements a dynamic feature fusion framework. Through a configurable feature combination strategy, the unit can flexibly adjust the generation method of the feature vector according to the requirements of different application scenarios. For example, in scenarios with high security requirements, the weight of semantic features can be increased; while in scenarios focusing on user experience, the weight of character-level features can be appropriately increased. In addition, the unit also implements a feature compression mechanism to reduce the dimension of the feature vector through a dimensionality reduction algorithm, reducing the computational complexity of subsequent processing while retaining key information.
[0181] The scoring execution unit adopts a pipeline architecture and divides the scoring process into multiple independent processing stages. This design can process multiple password evaluation requests in parallel, significantly improving the throughput. Each processing stage is equipped with an independent error handling mechanism. When an exception occurs in a certain stage, it can quickly perform fault recovery to ensure stability. At the same time, the unit can also implement a caching mechanism for the scoring results. For the same or similar password inputs, the cached evaluation results can be directly returned, reducing unnecessary repeated calculations.
[0182] The level determination unit is not only responsible for determining the final level, but also implements a complete evaluation feedback mechanism. When giving the evaluation result, targeted improvement suggestions will be generated at the same time, such as prompting the user to add specific types of characters and avoid using easily predictable patterns. These suggestions will be dynamically generated according to the user's specific input, taking into account practicality and operability. In addition, this unit also maintains an evaluation effect tracking system, recording the changes in password strength after the user adopts the suggestions, and these data can be used to optimize the suggestion generation strategy.
[0183] In terms of the integration of the embodiments of the present application, the device provided by the present application supports multiple deployment methods. It can be deployed as an independent service in the cloud, providing password strength evaluation services for multiple clients through the API interface; it can also be integrated into local applications and run as a component of the password management system. In the cloud deployment mode, through the load balancing and service degradation mechanisms, the processing strategy can be dynamically adjusted according to the server load situation to ensure the availability and responsiveness of the service.
[0184] To protect user privacy, the device adopts a full-link encryption mechanism during the processing. All password inputs and intermediate calculation results are encrypted and stored in memory, and are immediately cleared after the evaluation is completed. In addition, it can also have access control and audit log functions, capable of recording the execution of key operations, which is convenient for subsequent security audits and problem tracking.
[0185] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A real-time password strength assessment method, characterized in that: include: A multinomial naive Bayes model is trained using a password dataset, wherein the password dataset includes cracked weak password samples and verified strong password samples; In response to a trigger event of a user inputting a password, character-level features and semantic features of the password input by the user are extracted; Generate sub-feature vectors according to the character-level features and the semantic features respectively, and weight the sub-feature vectors to obtain a feature vector; Using the trained multinomial naive Bayes model to score the feature vector to obtain a password strength score; A password strength level is determined based on the password strength score.
2. The method according to claim 1, characterized in that Train a multinomial naive Bayes model on the password dataset, including: Learning the importance of different features in the polynomial naive Bayes model according to the password data set, and determining a feature subset with a high contribution to the polynomial naive Bayes model; Using the feature subset as input features of the polynomial naive Bayes model, and training the polynomial naive Bayes model based on the input features; The parameters of the multinomial naive Bayes model were adjusted through cross-validation to balance precision and recall.
3. The method according to claim 2, characterized in that The learning of the importance of different features in the polynomial naive Bayes model according to the password data set to determine a feature subset with a high contribution to the polynomial naive Bayes model includes: Performing data cleaning on the password data set; Using principal component analysis to reduce the feature dimension of the password data set to obtain a password data set after dimension reduction; Calculate the information gain ratio of each feature in the password data set after dimensionality reduction; Sorting the features in the password data set after dimensionality reduction based on the information gain ratio; Using mutual information method to evaluate the feature relevance of each feature in the password data set after dimensionality reduction; From the sorted features, the features with information gain ratio greater than the preset threshold are initially screened out, and secondary screening is performed based on the feature relevance to obtain a feature subset with high contribution to the model.
4. The method according to claim 3, characterized in that The feature vector is scored using the trained multinomial naive Bayes model to obtain a password strength score, including: Calculating a basic safety score of the feature vector; Performing weighted adjustment on the basic safety score based on preset safety rules to obtain an adjusted safety score; Calculating the probability of the password being cracked according to the adjusted security score; The probability of the password being cracked is calibrated based on the feature correlation of the feature subset to obtain a calibrated password strength score.
5. The method according to claim 4, characterized in that Also includes: Dynamically adjust the scoring threshold of the calibrated password strength score according to the historical evaluation accuracy; Smoothing is used to avoid drastic fluctuations in the calibrated password strength score.
6. The method according to claim 1, characterized in that The triggering event in response to the user inputting a password includes: Set up a password input listener to capture every character entered by the user in real time; When it is detected that the user pauses input for more than the preset time threshold or enters the enter key, a password evaluation event is triggered; The password entered by the user is stored in a temporary cache for feature extraction.
7. The method according to claim 1, characterized in that The step of weighting the sub-feature vectors to obtain the feature vector comprises: Statistically calculate the contribution of each feature in the historical evaluation data to the password strength prediction; Dynamically adjust the character-level feature weights and semantic feature weights according to the contribution; Normalize the character-level feature weights and semantic feature weights so that the sum of all weights is 1; The sub-feature vectors are weighted according to the character-level feature weights and the semantic feature weights to obtain the feature vector.
8. The method according to claim 7, characterized in that The step of dynamically adjusting the character-level feature weight and the semantic feature weight according to the contribution degree comprises: Set the weight adjustment period to recalculate feature weights at the end of each period; Record the frequency of character-level features and semantic features in correct predictions in each cycle; According to the importance score and occurrence frequency of the features, the weighted average algorithm is used to update the character-level feature weights and semantic feature weights.
9. The method according to claim 1, characterized in that: Extract character-level features of the password entered by the user, including: Counting the number of uppercase letters, lowercase letters, numbers and special characters in the password, and calculating the distribution ratio of each type of character in the password; and / or, identifying a sequence of characters that continuously increases or decreases in the password and recording the length of the sequence of characters; and / or, Detecting repeated characters or character combinations in the password and counting the number of repetitions; and / or, Calculate the total length of the password, and calculate the proportion of each type of character in the total length; Extract the semantic features of the password entered by the user, including: Setting the N value of the N-gram model to a range of 2 to 5, and performing word segmentation analysis on the password; and / or, Matching the password with a predefined dictionary of common words, and recording the number and position of the matched words; and / or, detecting leetspeak encoding in the password, including detecting replacement of letters with similar-looking numbers or special characters; and / or, Identify common number combination patterns such as date format, phone number format, etc. in the password.
10. A real-time password strength assessment device, characterized in that: include: A model training unit, used for training a multinomial naive Bayes model through a password data set, wherein the password data set includes cracked weak password samples and verified strong password samples; A feature extraction unit, configured to respond to a triggering event of a user inputting a password, extract character-level features and semantic features of the password, and generate a feature vector; A vector generating unit, used to generate sub-feature vectors according to the character-level features and the semantic features, and weight the sub-feature vectors to obtain a feature vector; A scoring execution unit, configured to score the feature vector using the polynomial naive Bayes model to obtain a password strength score; A level determination unit is used to determine the password strength level according to the password strength score.
Citation Information
Cited By
Weak password detection method based on multi-modal feature fusion and dynamic behavior analysis
CN120893032A
Password strength intelligent detection system and method based on exhaustion algorithm optimization
CN121441481A
Abnormal password detection method and device and electronic equipment
CN121603405A
Dynamic password strength evaluation method for Internet of Things equipment
CN122053048A
A dynamic password strength evaluation method for internet of things devices
CN122053048B