Refined password structure generation method
Through the refined password structure generation method based on WordNet dictionary system, the problem of semantic information missing in the traditional PCFGs method is solved, and the model accuracy and interpretability are achieved, which is suitable for password analysis and intensity evaluation.
Patent Information
- Application Number
- CN202510332283.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-10
AI Technical Summary
The traditional PCFGs password structure analysis method has problems such as missing semantic information, inability to distinguish meaningful vocabulary from random strings, and analytical vocabulary, resulting in classification deviation, resulting in misjudgment of password strength evaluation system and low penetration testing efficiency.
A refined password structure generation method based on WordNet dictionary system is adopted, and passwords are entered and segmented, and a refined structure is finally generated.
It solves the problem of missing semantic information, improves model accuracy, can mine richer password structures and semantics, reduces computing power requirements, and enhances interpretability.
Smart Images

Figure CN120124044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of password analysis, and in particular to a method for generating a refined password structure. Background Art
[0002] At present, the main password analysis methods include those based on probability models (Markov, PCFG), machine learning, deep learning, and large language models. The advantage of the PCFG (Probabilistic Context-Free Grammar) method is that it converts password construction rules into quantifiable probability models. It has outstanding performance in efficiency, explainability, resource consumption, etc., and is widely used in the industry.
[0003] The traditional PCFG method decomposes the password into three categories: letter segment (L), number segment (D) and symbol segment (S), but there are the following problems: 1) The letter segments are divided only by length, ignoring semantic information (e.g., the structures of "apple123" and "asdfg123" are treated as L5D3); 2) It is unable to distinguish meaningful words from random strings, resulting in insufficient accuracy of the structural model; 3) Polysemous words (e.g. "bank" can refer to a bank or a riverbank) lead to classification bias; In recent years, some improved PCFGs methods have refined the password structure, but they are mostly supplemented from certain aspects and do not fully cover all the letter segments and words. Due to the above problems of the PCFGs password analysis method, the password strength assessment system makes misjudgments due to the rough structure analysis, and the penetration test or electronic forensics also has poor guessing efficiency due to inaccurate models. Summary of the invention
[0004] In view of the above technical problems, the present invention provides a method for generating a refined password structure.
[0005] The present invention is implemented by adopting the following technical scheme: a method for generating a refined password structure, implemented based on the WordNet dictionary system, comprising the following steps: Step S1: Enter a password and identify the special character string in the password; Step S2: segment and mark the passwords not identified in step S1 according to their types, and obtain letter segments L, number segments D and symbol segments S; Step S3: sub-classify the letter segments L separated in step S2; Step S4: Obtain a refined structure of the password based on the detailed classification results.
[0006] Specifically, step S1 includes the following sub-steps: Step S11: Determine whether the input password is an ASCII printable string with a length of 1 to 128 bits, if so, proceed to the next step, otherwise terminate the process; Step S12: lowercase the password; Step S13: Obtain a character string from the sensitive vocabulary table, and check whether the password contains the character string. If so, mark it; Step S14: traverse the sensitive vocabulary list, and execute step S13 in a loop until the sensitive vocabulary list is completely traversed; Step S15: Get a character string from the keyboard travel table, check whether the password contains the character string, and mark it if it exists; Step S16: traverse the keyboard travel table, and execute step S15 in a loop until the keyboard travel table is completely traversed; Step S17: The password matches the regular expression of the email address and the website.
[0007] Specifically, the sensitive vocabulary list and keyboard navigation list are created by a user or generated using a tool.
[0008] Specifically, the segment marking in step S2 further includes: marking the lengths of the letter segment L, the number segment D and the symbol segment S.
[0009] Specifically, the step S2 further includes: performing year recognition on the password segment classified as the digital segment D, and if the match is successful, replacing the digital segment D mark and re-marking it with the year mark Y of the corresponding length.
[0010] Specifically, step S3 includes the following sub-steps: Step S31: using the directed acyclic graph-based minimum path method to perform password segmentation, balancing the semantic and structural complexity of the letter segment L; Step S32: calling the WordNet dictionary system to classify the password segmentation obtained in step S31. If there are segmentations that do not conform to the refined structure segmentation type, then go to step S33, otherwise directly go to step S4; Step S33: performing name recognition and classification on unclassified words; Step S34: The words that have not been classified are classified into other letter segment class word.other.
[0011] Specifically, the classification of password segmentation in step S32 specifically includes the following steps: Call the synset method to get all synonyms; Call the synset.lexname method for each synonym to obtain its classification; Call the lemma.count method for each synonym to get the word frequency; Count the word frequency of each category and sort them; The classification with the highest frequency is used as the classification result of the word.
[0012] Specifically, the step S33 includes the following sub-steps: Get the Chinese and English name databases, record the English name database as the name.eng table, and perform the next operation on the Chinese name database; Use the Python version of the Chinese character pinyin conversion library pypinyin to convert all Chinese names in the Chinese name database into various forms of Latin letters, and store the conversion results in the name.xm name table, name.mx name and surname table, name.x surname table, name.xms abbreviation name table, and name.s abbreviation name table respectively; The word is looked up in a table, and the word is classified into the category in which the table is hit.
[0013] The beneficial effects of the present invention are as follows: the present invention fully considers the characteristics of human password settings, uses WordNet to perform full-coverage semantic recognition and classification of English strings, solves the problem of insufficient model accuracy caused by the lack of semantic information in the traditional PCFGs password structure analysis method, and provides a highly practical refined structure representation and generation method for password analysis, which has the following advantages: (1) It can mine the richer structure and semantics of passwords, which helps us further understand passwords; (2) By analyzing the detailed password structure, we can discover the hidden characteristics of the password set, which helps us understand the password setting characteristics and preferences of different groups of people; (3) Compared with the AI-based password structure analysis method, the method adopted by the present invention has lower computing power requirements, is more interpretable, and is easier to be put into practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0015] Figure 1 It is a principle block diagram of a method for generating a refined password structure in an embodiment of the present invention; Figure 2 A schematic diagram of word classification of the WordNet dictionary system in an embodiment of the present invention; Figure 3It is a workflow diagram of a method for generating a refined password structure in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0017] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0018] The following is combined with Figures 1 to 3 , some embodiments of the present invention are described in detail. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0019] The present invention proposes a method for generating a refined password structure, which is implemented based on the WordNet dictionary system. WordNet is an English dictionary database system based on a lexical semantic network developed by Princeton University. It organizes English nouns, verbs, adjectives and adverbs into synsets, each of which represents a basic lexical concept, and establishes multiple lexical semantic relationships between these lexical concepts, including synonymy, antonymy, hypernymy, hyponymy, partial relationship and complete relationship. In a preferred embodiment, the method for generating a refined password structure includes the following steps: Step S1: Enter a password and identify the special character string in the password; Step S2: segment and mark the passwords not identified in step S1 according to their types, and obtain letter segments L, number segments D and symbol segments S; Step S3: sub-classify the letter segments L separated in step S2; Step S4: Obtain a refined structure of the password based on the detailed classification results.
[0020] The following is combined with Figures 1 to 3 Each step is explained in detail.
[0021] Step 1: Special string identifier: Step 1.1: Determine whether the input password is an ASCII printable string with a length of 1 to 128 bits. If so, proceed to the next step, otherwise terminate the process; Step 1.2: Lowercase the password; Step 1.3: Get a string from the sensitive word list (created by the user), check whether the password contains the string, and mark it if it exists (for example, if satan in satan1qaz@ is a sensitive word, the satan part is marked as a sensitive word word.sensitive5 with a length of 5, see Table 1); Step 1.4: Traverse the sensitive vocabulary list and execute step 1.3 repeatedly until the sensitive vocabulary list is completely traversed; Step 1.5: Get a string from the keyboard wandering table (created by the user, which can be generated using the kwprocessor tool), check whether the password contains the string, and mark it if it exists (for example, if 1qaz in good1qaz@ belongs to the keyboard wandering string, the 1qaz part is marked as a keyboard wandering string K4 with a length of 4, see Table 1); Step 1.6: Traverse the keyboard travel table and repeat step 1.5 until the keyboard travel table is completely traversed; Step 1.7: The password matches the regular expression of the email address and the website. In this embodiment, the regular expression of the email address is: "^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$"; The regular expression for the URL is: "^(https?:\ / \ / |ftp:\ / \ / |www\.)[a-zA-Z0-9-]+(\.[a-zA-Z]{2,})+(:\d+)?(\ / [^\s]*)?$"; If the match is successful, the password fragment is marked (for example, if woaini@163.com in woaini@163.com! is an email address, the woaini@163.com part is marked as an email address string E14 with a length of 14, see Table 1).
[0022] Step 2: Segment by type Step 2.1: Mark the unidentified password segments in step 1 by category, that is, divide them into three categories: letter segment (L), number segment (D) and symbol segment (S). For example, good in good1qaz@ is marked as a letter string with a length of 4, and @ is marked as a special character S1 with a length of 1); Step 2.2: Identify the year for the numeric segment classified in step 2.1. If the match is successful, replace the numeric segment tag marked in step 2.1 (e.g. 1986 in woaini@1986 is marked as a year number Y4 with a length of 4, see Table 1); Step 3: Sub-classification by letter segment Step 3.1: Segment the string marked as letter segments in step 2.1; Since there is no obvious word separator in the letter segment of the password, and there may be multiple word segmentation methods for the same letter segment, in order to balance the semantics and structural complexity of the letter segment, the minimum path method based on the directed acyclic graph is used to segment the password. For example, loverain is divided into two sub-segments: love+rain.
[0023] Step 3.2: Call the method provided by WordNet to classify the word segmentation obtained in step 3.1. The specific classification method is as follows: Call the synset method to get all synonyms; Call the synset.lexname method for each synonym to obtain its classification; Call the lemma.count method for each synonym to get the word frequency; Count the word frequency of each category and sort them; Use the most frequent classification as the classification result of the word; WordNet's synset can classify words into 45 categories, including 26 nouns (such as noun.time), 16 verbs (such as verb.body), 2 adjectives (such as adj.pert) and 1 adverb (such as adv.all). The specific classification types are shown in the first 45 categories in Table 1.
[0024] The following are the synonyms, classifications, and word frequencies obtained by calling the WordNet method for the word segmentation "dog":
[0025] Among them, the word frequency of noun.animal category is the highest, so the participle "dog" belongs to the noun.animal category.
[0026] The following are the synonyms, classifications, and word frequencies obtained by calling the WordNet method for the word segmentation "feeling":
[0027] Similarly, the participle "feeling" belongs to the verb.emotion class.
[0028] In the same way, the participle "rain" belongs to the noun.phenomenon class, and the participle "apple" belongs to the noun.food class.
[0029] If there are words that do not belong to the type in Table 1, proceed to the next step, otherwise go to step 4.
[0030] Step 3.3: Perform name recognition and classification on unclassified words. The specific steps are as follows: (1) Obtain the Chinese and English name databases (https: / / gitee.com / lemtasev / Chinese-Names-Corpus), where the Chinese and English name databases are denoted as the name.eng table, and perform the next operation on the Chinese name database; (2) Use the Python version of the Chinese character pinyin conversion library (pypinyin) to convert all Chinese names in the Chinese name database into various forms of Latin letters, such as "张三" is converted to "zhangsan" and stored in the name.xm table; such as "张三" is converted to "sanzhang" and stored in the name.mx table; such as "张三" only retains the surname, converted to "zhang", stored in the name.x table; such as "张三" is converted to "zhangs", stored in the name.mxs table; such as "张三" is converted to "zs", stored in the name.s table. English name database; (3) Look up the word in the table in the following order: name.xm, name.mx, name.x, name.xms, name.s, name.eng. The word is classified into the name category according to the table it hits. For example, if it hits the name.xm table, it is classified into the name.xm name category (see Table 1 for specific types). Step 3.4: Classify the unclassified words into other letter segment categories word.other (see Table 1); Step 4: Generate a refined structure of the password.
[0031] The refined structure obtained by this method contains up to 59 classification types. The complete classification types are shown in Table 1.
[0032] Table 1 59 segment types involved in the refined structure
[0033] In one embodiment, the structure of liuclovedog! obtained by using the Chuangtong PCFG method is L11S1, and the refined structure obtained by using the refined password structure identification method of this scheme is name.xms4verb.emotion4none.animal3S1, that is, the password consists of a name with a length of 4 bytes (name is an abbreviation), an emotion verb with a length of 4 bytes, an animal name with a length of 3 bytes, and a special character with a length of 1 byte.
[0034] This paper uses the WordNet dictionary system to perform full semantic recognition and classification of English strings, solving the problem of insufficient model accuracy caused by the lack of semantic information in the traditional PCFGs password structure analysis method, and provides a highly practical refined structure representation and generation method for password analysis. It will bring the following advantages: (1) It can mine the richer structure and semantics of passwords, which helps us further understand passwords; (2) By analyzing the detailed password structure, we can discover the hidden characteristics of the password set, which helps us understand the password setting characteristics and preferences of different groups of people; (3) Compared with the AI-based password structure analysis method, the method adopted by this patent requires lower computing power, is more interpretable, and is easier to be put into practice.
[0035] Experimental verification shows that the method adopted by the present invention has a significant advantage in reducing the number of meaningless password fragments compared with the traditional PCFG method. Taking the 12306 leaked password set as an example, the number of meaningless password fragments is reduced by 63%.
[0036] For the aforementioned embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily required by the present application.
[0037] The above embodiments describe the basic principles and main features of the present invention and the advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the changes and modifications made by those skilled in the art shall be within the scope of protection of the appended claims of the present invention without departing from the spirit and scope of the present invention.
Claims
1. A method for generating a refined password structure, characterized in that: The implementation based on the WordNet dictionary system includes the following steps: Step S1: Enter a password and identify the special character string in the password; Step S2: segment and mark the passwords not identified in step S1 according to their types, and obtain letter segments L, number segments D and symbol segments S; Step S3: sub-classify the letter segments L separated in step S2; Step S4: Obtain a refined structure of the password based on the detailed classification results.
2. A method for generating a refined password structure as claimed in claim 1, characterized in that: The step S1 comprises the following sub-steps: Step S11: Determine whether the input password is an ASCII printable string with a length of 1 to 128 bits, if so, proceed to the next step, otherwise terminate the process; Step S12: lowercase the password; Step S13: Obtain a character string from the sensitive vocabulary table, and check whether the password contains the character string. If so, mark it; Step S14: traverse the sensitive vocabulary list, and execute step S13 in a loop until the sensitive vocabulary list is completely traversed; Step S15: Get a character string from the keyboard travel table, check whether the password contains the character string, and mark it if it exists; Step S16: traverse the keyboard travel table, and execute step S15 in a loop until the keyboard travel table is completely traversed; Step S17: The password matches the regular expression of the email address and the website.
3. A method for generating a refined password structure as claimed in claim 2, characterized in that: The sensitive vocabulary list and keyboard navigation table are created by a user or generated using a tool.
4. A method for generating a refined password structure as claimed in claim 1, characterized in that: The segment marking in step S2 also includes: marking the lengths of the letter segment L, the number segment D and the symbol segment S.
5. A method for generating a refined password structure as claimed in claim 4, characterized in that: The step S2 also includes: performing year recognition on the password segment classified as the digital segment D, and if the match is successful, replacing the digital segment D mark and re-marking it with the year mark Y of the corresponding length.
6. A method for generating a refined password structure as claimed in claim 1, characterized in that: The step S3 comprises the following sub-steps: Step S31: using the directed acyclic graph-based minimum path method to perform password segmentation, balancing the semantic and structural complexity of the letter segment L; Step S32: calling the WordNet dictionary system to classify the password segmentation obtained in step S31. If there are segmentations that do not conform to the refined structure segmentation type, then go to step S33, otherwise directly go to step S4; Step S33: performing name recognition and classification on unclassified words; Step S34: The words that have not been classified are classified into other letter segment class word.other.
7. A method for generating a refined password structure as claimed in claim 6, characterized in that: The classification of password segmentation in step S32 specifically includes the following steps: Call the synset method to get all synonyms; Call the synset.lexname method for each synonym to obtain its classification; Call the lemma.count method for each synonym to get the word frequency; Count the word frequency of each category and sort them; The classification with the highest frequency is used as the classification result of the word.
8. A method for generating a refined password structure as claimed in claim 7, characterized in that: The step S33 includes the following sub-steps: Get the Chinese and English name databases, record the English name database as the name.eng table, and perform the next operation on the Chinese name database; Use the Python version of the Chinese character pinyin conversion library pypinyin to convert all Chinese names in the Chinese name database into various forms of Latin letters, and store the conversion results in the name.xm name table, name.mx name and surname table, name.x surname table, name.xms abbreviation name table, and name.s abbreviation name table respectively; The word is looked up in a table, and the word is classified into the category in which the table is hit.