PCFG-based password set similarity quantification method

By analyzing the password set based on PCFG, generating a password structure probability table and calculating the similarity, the problem of insufficient quantification of password set similarity in the prior art is solved, and the password guessing effect and traceability accuracy are improved.

CN120124045APending Publication Date: 2025-06-10UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332285.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology has defects in the quantification of password set similarity, and cannot effectively capture structural distribution characteristics, have high computational complexity, and the evaluation latitude of a single password strength evaluation model, resulting in poor cross-domain password guessing, inconsistent differences in crowd passwords obtained by password analysis, and inaccurate traceability of leaked password sets.

Method used

Using a PCFG-based method, the password in the password set is analyzed, the password structure probability table is generated, the password structure vector is obtained through vector space mapping, and the similarity is calculated using the cosine similarity formula.

Benefits of technology

By retaining the macro structure distribution information of the password set, the problem that traditional methods cannot quantify the similarity of password sets is solved, and the cross-domain password guessing effect is improved, the password difference obtained is more typical, and the leaked password set traceability is more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124045A_ABST
    Figure CN120124045A_ABST
Patent Text Reader

Abstract

The invention discloses a PCFG-based password set similarity quantification method, which comprises the following steps of: analyzing passwords in a password set based on a PCFG probability context-independent grammar model to obtain a password structure probability table; carrying out vector space mapping on the obtained password structure vector to obtain a password structure vector; and based on the obtained password structure vector, calculating the similarity by adopting a cosine similarity formula. According to the scheme, the PCFG method is utilized to obtain the password structure probability table, then the password structure probability table is vectorized, the similarity between the password structure vectors is calculated, the problem of password set similarity quantification is solved, and a simple, convenient and effective method is provided for password set similarity quantitative analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of password sets, and in particular, to a method for quantifying the similarity of password sets based on PCFG. Background Art

[0002] Different password sets are generated by different groups of people, so the password sets contain the patterns and preferences of password setting by specific groups. There are differences or similarities among different password sets. Studying the similarity of password sets helps us understand passwords and utilize passwords. For example: in password guessing, using similar cross-domain password sets can achieve better guessing results; when studying the differences among people in passwords, using password sets with lower similarity can obtain more obvious difference features; when tracing the source of leaked password sets, similarity measurement can be used to narrow down the source of the leaked password sets.

[0003] Currently, the research on password set similarity or difference mostly stays in the qualitative stage, and there is little quantitative research. However, these password set similarity quantification methods also have obvious defects: 1. Traditional password similarity comparison relies on character-level overlap (such as Jaccard similarity) and cannot capture structural distribution features.

[0004] 2. The clustering algorithm directly acts on the original password data, with high computational complexity and ignoring the grammar generation pattern.

[0005] 3. A single password strength evaluation model (such as zxcvbn) has a low evaluation dimension and cannot reflect password feature information.

[0006] Due to the above problems in traditional password similarity quantification methods, problems such as poor cross-domain password guessing effect, atypical differences in passwords of different people obtained from password analysis, and inaccurate tracing of the source of leaked password sets occur. Summary of the Invention

[0007] In view of the above technical problems, the present invention provides a method for quantifying the similarity of password sets based on PCFG.

[0008] The present invention is implemented by the following technical solutions: A method for quantifying the similarity of password sets based on PCFG, comprising the following steps: Step S1: Analyze the passwords in the password set based on the PCFG (Probabilistic Context-Free Grammar) model to obtain a password structure probability table; Step S2: Perform vector space mapping on the result obtained in Step S1 to obtain a password structure vector; Step S3: Based on the password structure vector obtained in Step S2, calculate the similarity using the cosine similarity formula.

[0009] Specifically, the Step S1 includes the following sub-steps: Step S11: Preprocess two original leaked password data a and b that need to calculate similarity. Only retain the correct passwords and store them line by line, denoted as password set A and password set B respectively; Step S12: Read the passwords in password set A in sequence, split the passwords into basic syntax units, and record the length distribution; Step S13: Record the password structure and the number of times the password structure appears in the password structure frequency table; Step S14: Loop and repeat Step S12 - Step S13 until all the passwords in the password set are traversed; Step S15: Calculate the occurrence probability of the password structure in the password set according to the password structure frequency in the password structure frequency table, and generate a password structure probability table; Step S16: Perform the operations of Step S12~Step S15 on password set B in sequence, and finally obtain the password structure probability tables Ag and Bg of password set A and password set B.

[0010] Specifically, the password data preprocessing in Step S11 specifically includes: Remove the redundant information in the original leaked password data and only retain the password field; Perform a legality check on the passwords retained in the previous step, and only retain passwords that are ASCII printable strings with a length of 1 to 128 bits.

[0011] Specifically, the password splitting rule in Step S12 is: Split according to the LDS letter - number - symbol pattern.

[0012] Specifically, the calculation of the occurrence probability of the password structure in the password set in Step S15 is specifically expressed as: Password structure occurrence frequency = Password structure occurrence frequency / Number of passwords in password set A.

[0013] Specifically, Step S2 includes the following sub - steps: Step S21: Align and supplement the password structures in the password structure probability tables Ag and Bg to obtain the aligned password structure probability tables AG and BG; Step S22: Extract the probability values of AG and BG to obtain the password structure vectors AV of password set A and the password structure vector BV of password set B.

[0014] Specifically, the calculation formula for the similarity in Step S3 is expressed as: ; where, is the dot product of the two vectors AV and BV, ; is the modulus length of the vector AV, that is, the Euclidean distance, is the modulus of the vector BV, , .

[0015] Specifically, the output value range of the similarity calculation formula is [0, 1]. The larger the value, the higher the structural similarity, that is, the higher the similarity between the password set A and the password set B.

[0016] The beneficial effects of the present invention are as follows: Based on the characteristic that passwords created by the same kind of people have similarities, the present invention uses the PCFG (Probabilistic Context-Free Grammar) model to perform probabilistic analysis on the password set to obtain the password structure probability table, and quantifies the similarity between password sets based on the password structure vector, solving the deficiencies existing in the traditional password set similarity quantification method, providing a feasible solution for the quantification analysis of password set similarity, and having the following advantages: (1) The password structure probability table generated based on the PCFG model retains the macrostructure distribution information of the password set, and its features are more profound compared to the macrostructure distribution information such as characters and lengths; (2) The method of calculating the cosine similarity is adopted to solve the problem that the similarity of password sets cannot be quantified; (3) The password set similarity quantification method proposed based on the present invention can make the cross-domain password guessing effect better, can more easily obtain the typical differences of passwords of different groups of people, and can make the tracing of the leaked password set more accurate. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0018] Figure 1 is the principle framework diagram of the password set similarity quantification method based on PCFG in the embodiments of the present invention; Figure 2 is the working flow diagram of the password set similarity quantification method based on PCFG in the embodiments of the present invention. Detailed Embodiments

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown here can be arranged and designed in various different configurations.

[0020] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0021] The following is combined with Figure 1-2 , some embodiments of the present invention are described in detail. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0022] The present invention proposes a password set similarity quantification method based on PCFG. Based on the fact that passwords created by similar people have similarities, PCFG is used to perform probability analysis on the password set to obtain a password structure probability table, and the similarity between the password sets is quantified based on the password structure vector. In a preferred embodiment, the method comprises the following steps: Step S1: Based on the PCFG probabilistic context-free grammar model, the passwords in the password set are analyzed to obtain a password structure probability table; Step S2: Perform vector space mapping on the vector obtained in step S1 to obtain a password structure vector; Step S3: Based on the password structure vector obtained in step S2, the similarity is calculated using the cosine similarity formula.

[0023] In a specific embodiment, the architecture and process of the password set similarity quantification method based on PCFG are respectively as follows: Figure 1 , 2 As shown, the detailed technical composition of each step is as follows: Step 1, PCFG password analysis: Step 1.1: Preprocess the two original leaked password data a and b that need to be calculated for similarity, retain only the correct passwords, and store them by row (recorded as password set A and password set B).

[0024] In this embodiment, the specific operations of preprocessing are: (1) Remove the redundant information in the original leaked password data and only keep the password field; (2) Determine the legitimacy of the passwords retained in the previous step and only retain passwords that are ASCII printable strings with a length of 1 to 128 bits.

[0025] Step 1.2: Read the passwords in password set A in order, split the passwords into basic syntax units according to the LDS (letter-number-symbol) pattern, and record the length distribution. Example: The structure of the password "Pass123!" is L4D3S1.

[0026] Step 1.3: Record the password structure and the number of times the password structure appears in the password structure frequency table.

[0027] Step 1.4: Loop through Step 1.2 and Step 1.3 until all the passwords in the password set have been traversed.

[0028] Step 1.5: Calculate the occurrence probability of the password structure in the password set according to the password structure frequency in the password structure frequency table (password structure occurrence frequency = password structure occurrence frequency / number of passwords in password set A), and generate a password structure probability table; Step 1.6: Perform the operations in Step 1.2 to 1.5 on password set B in sequence, and finally obtain the password structure probability tables Ag and Bg of password set A and password set B.

[0029] Step 2, Vector space mapping: Step 2.1: Align and supplement Ag and Bg according to the password structure to obtain the aligned password structure probability tables AG and BG. For example: Ag is {[D6, 0.1], [D8, 0.05], [L6, 0.01],...}; Bg is {[D8, 0.2], [L6, 0.05], [L3D3, 0.01],...}.

[0030] Then, after alignment and supplementation, we get: AG is {[D6, 0.1], [D8, 0.05], [L6, 0.01], [L3D3, 0],...}; BG is {[D6, 0], [D8, 0.2], [L6, 0.05], [L3D3, 0.01],...}.

[0031] Step 2.2: Extract the probability values of AG and BG to obtain the password structure vector AV of password set A and the password structure vector BV of password set B; for example: In the previous step, AV is [0.1, 0.05, 0.01, 0,...]; BV is [0, 0.2, 0.05, 0.01,...].

[0032] Step 3: Similarity calculation: Use the cosine similarity formula to calculate the similarity between password set A and password set B. The calculation formula is: ; where is the dot product of the two vectors AV and BV, ; is the modulus length of vector AV, that is, the Euclidean distance, is the modulus length of vector BV, ,, , and the output value range is [0, 1]. The larger the value, the higher the structural similarity.

[0033] In one embodiment, the method proposed by the present invention is used to calculate the similarity of leaked password sets such as 126, 136, linkedin, and rockyou. The similarity between the 126 and 163 password sets, both of which are of the email type, is as high as 0.99, and the similarity between the linkedin and rockyou password sets, both of which are social networking sites, is also as high as 0.94. This indicates that the password creation patterns of users of the same type of websites in the same region are relatively similar.

[0034] In this solution, a traditional PCFG is used to represent the password structure, and a PCFG-like method with added password fragment types can also be used to represent the password structure more precisely.

[0035] For the foregoing embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to this application.

[0036] In the above embodiments, the basic principles, main features, and advantages of the present invention are described. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any changes and modifications made by those skilled in the art should fall within the protection scope of the appended claims of the present invention.

Claims

1. A password set similarity quantification method based on PCFG, characterized in that: The following steps are involved: Step S1: Based on the PCFG probabilistic context-free grammar model, the passwords in the password set are analyzed to obtain a password structure probability table; Step S2: Perform vector space mapping on the vector obtained in step S1 to obtain a password structure vector; Step S3: Based on the password structure vector obtained in step S2, the similarity is calculated using the cosine similarity formula.

2. A password set similarity quantification method based on PCFG as claimed in claim 1, characterized in that: The step S1 comprises the following sub-steps: Step S11: pre-processing the two original leaked password data a and b that need to be calculated for similarity, retaining only the correct passwords, and storing them by row, respectively recorded as password set A and password set B; Step S12: Read the passwords in the password set A in order, divide the passwords into basic grammatical units, and record the length distribution; Step S13: recording the password structure and the number of times the password structure appears in a password structure frequency table; Step S14: Repeat steps S12 to S13 in a loop until all the passwords in the password set are traversed; Step S15: Calculate the occurrence probability of the password structure in the password set according to the password structure frequency table, and generate a password structure probability table; Step S16: Execute the operations of step S12 to step S15 on password set B in sequence, and finally obtain password structure probability tables Ag and Bg of password set A and password set B.

3. A password set similarity quantification method based on PCFG as claimed in claim 2, characterized in that: The step S11 password data preprocessing specifically includes: Remove redundant information from the original leaked password data and only keep the password field; The passwords retained in the previous step are judged for legitimacy, and only passwords with a length of 1 to 128 bits and ASCII printable strings are retained.

4. A password set similarity quantification method based on PCFG as claimed in claim 2, characterized in that: The password segmentation rule in step S12 is: segmentation according to the LDS letter-number-symbol pattern.

5. A password set similarity quantification method based on PCFG as claimed in claim 2, characterized in that: The calculation of the occurrence probability of the password structure in the password set in step S15 is specifically expressed as: password structure occurrence frequency = password structure occurrence frequency / number of passwords in password set A.

6. A password set similarity quantification method based on PCFG as claimed in claim 2, characterized in that: The step S2 comprises the following sub-steps: Step S21: aligning and supplementing the password structures of the password structure probability tables Ag and Bg to obtain aligned password structure probability tables AG and BG; Step S22: Extract the probability values ​​of AG and BG to obtain the password structure vector AV of password set A and the password structure vector BV of password set B.

7. A password set similarity quantification method based on PCFG as claimed in claim 6, characterized in that: The calculation formula of the similarity in step S3 is expressed as: ; in, is the dot product of the two vectors AV and BV, ; is the modulus of vector AV, i.e., the Euclidean distance, is the modulus of vector BV, , .

8. A password set similarity quantification method based on PCFG as claimed in claim 7, characterized in that: The output value range of the similarity calculation formula is [0, 1]. The larger the value, the higher the structural similarity, that is, the password set A and the password set B are highly similar.