Password strength evaluation method based on self-labeling and integrated text classification

Through the method based on self-notation and integrated text classification, the problems of insufficient accuracy of password strength evaluation and large calculation overhead in the prior art are solved, and higher evaluation accuracy and better spatio-temporal performance are achieved.

CN120105402APending Publication Date: 2025-06-06NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510188393.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy, large dependence on prior knowledge, and high computational time and space overhead in password strength evaluation.

Method used

Password strength evaluation method based on self-notation and integrated text classification is adopted, and password guessing model training, password strength self-notation and stacked integrated learning are achieved intelligent evaluation of password strength.

Benefits of technology

It improves the accuracy of password strength evaluation, reduces the workload of manually labeling data, and achieves a trade-off between accuracy and space-time overhead to a certain extent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105402A_ABST
    Figure CN120105402A_ABST
Patent Text Reader

Abstract

The invention discloses a password strength evaluation method based on self-labeling and integrated text classification. The method comprises the following steps: training a password guessing model; performing password intensity self-labeling, performing simulation attack based on the password guessing model obtained through S1 training, and completing password intensity value labeling according to the number of times of simulation attack; and carrying out password strength evaluation based on integrated classification, carrying out stacked integrated text classification learning on the data set marked with the password strength, and finally carrying out automatic evaluation on the password strength by utilizing a text classification model to obtain a password strength value. According to the method, the password intensity of the to-be-learned password data set is automatically labeled by using the password guessing attack method, so that the problems of large workload and strong subjective difference of manual data labeling are solved. According to the password strength evaluation method based on stacked ensemble learning, different password strength evaluators based on classification are integrated to mutually take the advantages, so that higher password strength evaluation accuracy is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of information security technology, and in particular relates to a password strength assessment method based on self-annotation and integrated text classification. Background Art

[0002] Passwords are short keys that humans can remember and are the first line of defense for information systems and data security. They were first introduced into the computer field by Professor Fernando Corbato and are used for local file access control on computers. With the rapid development and popularization of the Internet, computers and mobile terminals, various information services are becoming increasingly networked. In today's highly information-based and networked era, password technology provides security functions such as data encryption, digital signatures, and identity authentication, and has become an important means of protecting users' digital assets.

[0003] Currently, the biggest security threat facing passwords is password guessing attacks, that is, attackers use leaked password libraries and combine password rules, statistical methods or deep learning technology to launch guessing attacks. The most direct and effective way to resist password guessing attacks is that when users set passwords, the target system will feedback the strength value of the pre-set password in real time to guide or recommend high-strength user passwords. Therefore, password strength assessment research has received continuous attention in the field of information security.

[0004] Password strength assessment is used to quantify the strength of a given plain text password. Its main purpose is to timely assess the password strength of network users and remind users to set high-security passwords to avoid security risks caused by weak passwords. Password strength factors include password length, the size of the character or symbol set used, and the order of character arrangement. If the password strength is too low, it means that the network security risks faced are greater.

[0005] Current research on password strength assessment can be roughly divided into four directions: rule-based, pattern detection, classification, and guessing attacks. The rule-based password strength assessment is represented by the NIST password strength assessment, which aims to measure the password strength based on the length and character types of the password. The advantages of this method are easy implementation and low latency, but the accuracy of its password strength is unsatisfactory. The password strength assessment method based on pattern detection detects and assigns the construction mode (such as keyboard mode, sequential character mode, palindrome mode, etc.) to each sub-segment of the password to obtain the password strength value. Its representative work is zxcvbn. However, the password strength assessment method based on classification uses classification technology to learn some well-annotated passwords to fit the underlying features of password strength, and finally uses a classifier to distinguish the strength of the password to be evaluated. The above two methods have a strong dependence on prior knowledge (construction mode and well-annotated password data sets). Finally, the password strength assessment method based on password guessing attacks measures the password strength by the number of guesses of the password guessing model. This type of method is more accurate, but it has a large space and time cost during calculation. Summary of the invention

[0006] The purpose of this application is to overcome the problems of the prior art and disclose a password strength assessment method based on self-annotation and integrated text classification for intelligent password strength assessment. The assessment result can provide password strength guidance for users or systems.

[0007] The purpose of this application is achieved through the following technical solutions:

[0008] A password strength assessment method based on self-annotation and integrated text classification, the password strength assessment method comprising:

[0009] S1: password guessing model training;

[0010] S2: Password strength self-labeling: Based on the password guessing model trained in S1, simulated attacks are carried out, and password strength values ​​are labeled according to the number of simulated attacks to obtain a labeled password strength dataset;

[0011] S3: Password strength assessment based on ensemble classification. Stacked ensemble text classification learning is carried out on a data set with labeled password strengths. Finally, the text classification model is used to automatically assess password strength and obtain the password strength value.

[0012] According to a preferred embodiment, in step S1, the guessing model includes: a statistical guessing method based on PCFG and Markov and a guessing method based on artificial intelligence.

[0013] According to a preferred embodiment, step S2 comprises:

[0014] S21: Perform a simulated attack on the password text to be marked, and count the number of guesses required to guess the password to be marked;

[0015] S22: Automatically mark according to the number of guesses in S21, if the number of guesses is less than 10 8 If the number of guesses is greater than 10, the strength value of the corresponding password is marked as weak. 8 Less than 10 12 It is marked as medium password strength, with more than 10 guesses. 12 The password strength is marked as strong;

[0016] S23: Execute the self-annotation process of S21 and S22 on all the annotated password texts to obtain an annotated password strength dataset.

[0017] According to a preferred embodiment, in step S3, stacking-based ensemble learning refers to training a model for combining other models.

[0018] According to a preferred embodiment, step S3 includes:

[0019] S31: Cross-validation: For the password strength dataset obtained in step S2, K-fold cross-validation is performed on K classification models to obtain K intermediate classification results, which are merged into a new training dataset NewTrainData and a new test dataset NewTestData respectively;

[0020] S32: Logistic regression, using the new training data set NewTrainData and the new test data set NewTestData obtained in S31, train a logistic regression model to obtain the final password strength value.

[0021] According to a preferred embodiment, step S31 includes:

[0022] Based on the three algorithms of TextCNN, Bilstm and DPCNN, the training set is divided into 3 folds, namely TrainData_i, where i = 1, 2, 3; 2 folds are used as training sets and 1 fold is used as validation set;

[0023] Use the training set to train the model Model_i, where i = 1, 2, 3; use Model_i to predict and get a 1-fold one-dimensional prediction sequence ValPredict_i, where i = 1, 2, 3; at the same time, use Model_i to predict the test set TestData, and also get a prediction sequence TestPredict_i;

[0024] The loop is repeated until all combinations are traversed, and finally 3 ValPredict_i and 3 TestPredict_i are obtained. The 3 ValPredict_i columns are merged into a prediction sequence NewTrainData_j, and the 3 TestPredict_i are averaged to obtain a prediction sequence NewTestData_j based on the test set.

[0025] Different algorithms are used respectively, and finally 3 NewTrainData_j and 3 NewTestData_j are obtained. The 3 NewTrainData_j and 3 NewTestData_j are merged row-wise to obtain a new training data set NewTrainData and a new test data set NewTestData.

[0026] The aforementioned main scheme of the present application and its further options can be freely combined to form multiple schemes, all of which are schemes that can be adopted and claimed for protection in the present application. After understanding the scheme of the present application, those skilled in the art can understand that there are multiple combinations based on the prior art and common knowledge, all of which are technical schemes to be protected by the present application, and they are not exhaustively listed here.

[0027] Beneficial effects of this application:

[0028] (1) This application uses a password guessing attack method to automatically label the password strength of the learning password dataset, thereby alleviating the problem of large workload and strong subjective differences in manual data labeling.

[0029] (2) This application proposes a password strength assessment method based on stacked ensemble learning, which integrates different classification-based password strength assessors to take advantage of each other, thereby achieving a higher password strength assessment accuracy.

[0030] (3) The password strength assessment method proposed in this application achieves a trade-off between accuracy and time and space overhead to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a flowchart of the password strength assessment method of this application;

[0032] Figure 2 It is a flowchart of password strength self-labeling in the password strength assessment method of this application;

[0033] Figure 3 It is a flowchart of password strength assessment based on integrated classification in the password strength assessment method of the present application;

[0034] Figure 4 This is a diagram of 3-fold cross validation. DETAILED DESCRIPTION

[0035] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0036] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0037] In the description of this application, it should also be noted that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0038] In addition, the present application would like to point out that, in the present application, unless the specific structure, connection relationship, positional relationship, power source relationship, etc. are specifically written out, the structure, connection relationship, positional relationship, power source relationship, etc. involved in the present application are all known by those skilled in the art on the basis of the prior art without creative work.

[0039] Example 1

[0040] refer to Figures 1 to 4 As shown, the present application discloses a password strength assessment method based on self-annotation and integrated text classification, and the password strength assessment method includes the following steps.

[0041] Step S1: Password guessing model training. The guessing model is trained to provide the self-annotation and self-attack methods and models for the next step. The guessing model can be a statistical guessing method based on PCFG and Markov, a guessing method based on artificial intelligence, etc.

[0042] Step S2: Password strength self-labeling. A simulated attack is carried out based on the password guessing model trained in S1, and the password strength value is labeled according to the number of simulated attacks to obtain a labeled password strength data set.

[0043] Specifically, step S2 includes:

[0044] S21: Perform a simulated attack on the password text to be marked, and count the number of guesses required to guess the password to be marked;

[0045] S22: Automatically mark according to the number of guesses in S21, if the number of guesses is less than 10 8 If the number of guesses is greater than 10, the strength value of the corresponding password is marked as weak. 8 Less than 10 12 It is marked as medium password strength, with more than 10 guesses. 12 The password strength is marked as strong;

[0046] S23: Execute the self-annotation process of S21 and S22 on all the annotated password texts to obtain an annotated password strength dataset.

[0047] Step S3: Password strength assessment based on integrated classification. Stacked integrated text classification learning is carried out on the data set with labeled password strength. Finally, the text classification model is used to automatically assess the password strength and obtain the password strength value.

[0048] Stacking-based ensemble learning refers to training a model to combine other models, which is essentially a stacking generalization method. The overall learning process of this method can be roughly divided into two steps, such as Figure 3 As shown in the figure: The first step is to carry out cross-validation on the training corpus for different classification models, and use the prediction results of the classification models as the input of the second-stage combined model.

[0049] Specifically, step S3 includes:

[0050] S31: Cross-validation: For the password strength dataset obtained in step S2, K-fold cross-validation is performed on K classification models to obtain K intermediate classification results, which are merged into a new training dataset NewTrainData and a new test dataset NewTestData respectively;

[0051] Further, step S31 includes:

[0052] refer to Figure 4 As shown, this embodiment is based on the three algorithms of TextCNN, Bilstm and DPCNN, and divides the training set into three folds, namely TrainData_i, where i=1,2,3; 2 folds are used as training sets and 1 fold is used as validation set; the specific classification model and cross-validation number can be adjusted according to the actual situation.

[0053] Use the training set to train the model Model_i, where i = 1, 2, 3; use Model_i to predict and get a 1-fold one-dimensional prediction sequence ValPredict_i, where i = 1, 2, 3; at the same time, use Model_i to predict the test set TestData, and also get a prediction sequence TestPredict_i;

[0054] The loop is repeated until all combinations are traversed, and finally 3 ValPredict_i and 3 TestPredict_i are obtained. The 3 ValPredict_i columns are merged into a prediction sequence NewTrainData_j, and the 3 TestPredict_i are averaged to obtain a prediction sequence NewTestData_j based on the test set.

[0055] Different algorithms are used respectively, and finally 3 NewTrainData_j and 3 NewTestData_j are obtained. The 3 NewTrainData_j and 3 NewTestData_j are merged row-wise to obtain a new training data set NewTrainData and a new test data set NewTestData.

[0056] S32: Logistic regression, using the new training data set NewTrainData and the new test data set NewTestData obtained in S31, train a logistic regression model to obtain the final password strength value. This method can avoid overfitting, learn the information of the combination of features, and improve the accuracy of prediction.

[0057] This application uses a password guessing attack method to automatically annotate the password strength of a learning password dataset, thereby alleviating the problem of large workload and strong subjective differences in manually annotated data. This application proposes a password strength assessment method based on stacked ensemble learning, which integrates different classification-based password strength evaluators to take advantage of each other, thereby achieving a higher password strength assessment accuracy. The password strength assessment method proposed in this application achieves a trade-off between accuracy and time and space overhead to a certain extent.

[0058] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A password strength assessment method based on self-annotation and integrated text classification, characterized in that: The password strength assessment method comprises: S1: password guessing model training; S2: Password strength self-labeling: Based on the password guessing model trained in S1, simulated attacks are carried out, and password strength values ​​are labeled according to the number of simulated attacks to obtain a labeled password strength dataset; S3: Password strength assessment based on ensemble classification. Stacked ensemble text classification learning is carried out on a data set with labeled password strengths. Finally, the text classification model is used to automatically assess password strength and obtain the password strength value.

2. The password strength evaluation method according to claim 1, characterized in that: In step S1, the guessing model includes: a statistical guessing method based on PCFG and Markov and a guessing method based on artificial intelligence.

3. The password strength evaluation method according to claim 1, characterized in that: Step S2 includes: S21: Perform a simulated attack on the password text to be marked, and count the number of guesses required to guess the password to be marked; S22: Automatically mark according to the number of guesses in S21, if the number of guesses is less than 10 8 If the number of guesses is greater than 10, the strength value of the corresponding password is marked as weak. 8 Less than 10 12 It is marked as medium password strength, with more than 10 guesses. 12 The password strength is marked as strong; S23: Execute the self-annotation process of S21 and S22 on all the annotated password texts to obtain an annotated password strength dataset.

4. The password strength assessment method according to claim 1, characterized in that: In step S3, stacking-based ensemble learning refers to training a model for combining other models.

5. The password strength assessment method according to claim 4, characterized in that: Step S3 includes: S31: Cross-validation: For the password strength dataset obtained in step S2, K-fold cross-validation is performed on K classification models to obtain K intermediate classification results, which are merged into a new training dataset NewTrainData and a new test dataset NewTestData respectively; S32: Logistic regression, using the new training data set NewTrainData and the new test data set NewTestData obtained in S31, train a logistic regression model to obtain the final password strength value.

6. The password strength assessment method according to claim 5, characterized in that: Step S31 includes: Based on the three algorithms of TextCNN, Bilstm and DPCNN, the training set is divided into 3 folds, namely TrainData_i, where i = 1, 2, 3; 2 folds are used as training sets and 1 fold is used as validation set; Use the training set to train the model Model_i, where i = 1, 2, 3; use Model_i to predict and get a 1-fold one-dimensional prediction sequence ValPredict_i, where i = 1, 2, 3; at the same time, use Model_i to predict the test set TestData, and also get a prediction sequence TestPredict_i; The loop is repeated until all combinations are traversed, and finally 3 ValPredict_i and 3 TestPredict_i are obtained. The 3 ValPredict_i columns are merged into a prediction sequence NewTrainData_j, and the 3 TestPredict_i are averaged to obtain a prediction sequence NewTestData_j based on the test set. Different algorithms are used respectively, and finally 3 NewTrainData_j and 3 NewTestData_j are obtained. The 3 NewTrainData_j and 3 NewTestData_j are merged row-wise to obtain a new training data set NewTrainData and a new test data set NewTestData.