Directional password guessing method and system based on user personal information

Through the password generation algorithm that meticulously classifies personal information and optimizes memory, the structure with the greatest probability of generating passwords is solved, and the problem of rough classification and repeated statistics of personal information in the traditional password cracking model is improved, and the success rate of password guessing is improved.

CN119989332APending Publication Date: 2025-05-13HAINAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510081802.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The traditional targeted password cracking model is too rough in the classification of personal information and has a lot of repeated statistics, resulting in a low success rate of password guessing.

Method used

Personal information is divided into 6 major categories and 36 subcategories, and N password structures with the greatest probability of password generation through memory-optimized password generation algorithm are generated, and guess passwords are generated based on the personal information entered by the user.

Benefits of technology

It improves the success rate of password guessing, solves the problems of rough classification and repeated statistics of personal information in traditional methods, and avoids the problem of memory overflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989332A_ABST
    Figure CN119989332A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric data processing, in particular to a directional password guessing method and system based on user personal information, and the method comprises the following steps: S10, carrying out the preprocessing of an obtained data set, so as to remove invalid data; s20, dividing the personal information into six large classes, namely 36 small classes, and performing statistical analysis on a data set obtained after preprocessing based on the classification mode to obtain a composition structure and probability of a password; and S30, generating a guess password based on the personal information input by the user and the composition structure and probability of the password. According to the method, six categories of personal information are set, and 36 subcategories are further divided, so that the problems that a traditional directional password cracking model is too rough in division of the types of the personal information and has a large amount of repeated statistics are solved, meanwhile, the limitation that the personal information is represented by digits is eliminated, and semantics are directly captured and matched. Therefore, the distribution rule of the personal information in the password can be reflected more truly, and a better cracking effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic data processing, and in particular to a method and system for directional password guessing based on user personal information. Background Art

[0002] In order to protect the privacy of users, users are usually required to set passwords. Password guessing refers to guessing the user's password through technical means. Password guessing has positive significance. For example, it can detect the strength of the user's password to prompt the user to set a more secure password. Therefore, password guessing has always been a topic that scholars have worked hard to study.

[0003] For example, the Chinese patent with application number 201611079933.X discloses "a method for generating a password guessing set based on username information and a password cracking method". Although it can segment and annotate the semantic structure of usernames and passwords in the leaked data set, the classification of personal information types is too rough and has a large number of repeated statistical problems. The success rate of passwords guessed based on this method is low. Summary of the invention

[0004] The purpose of the present invention is to provide a method and system for directional password guessing based on user personal information, so as to solve the problem of over-coarse classification and repeated statistics of personal information in traditional directional password cracking models.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a method for directional password guessing based on user personal information, comprising the following steps:

[0007] S10, preprocessing the obtained data set to remove invalid data;

[0008] S20, dividing the personal information into 6 major categories with a total of 36 subcategories, and performing statistical analysis on the data set obtained after preprocessing based on the classification method to obtain the composition structure and probability of the password;

[0009] S30, generating a guessed password based on the personal information input by the user and the composition structure and probability of the password.

[0010] The step S30 comprises the following steps:

[0011] S301, generating N password structures with the highest password probabilities according to a memory-optimized password generation algorithm;

[0012] S302, escaping the personal information tag in the password structure by the corresponding personal information;

[0013] S303, for the remaining structures in the password structure, escape the letter string, special character string or digital string with a higher probability in the data set to obtain a guessed password.

[0014] The step S301 includes the following steps:

[0015] S3011, when the child process generates the initialization of the pre-terminal, it finds the index of the START node in the grammar, and reads the first replacement structure of the START node into the maximum heap;

[0016] S3012, pop out the element with the highest probability in the maximum heap, perform increment and expand operations, and generate a substructure of the element with the highest probability;

[0017] S3013, putting the substructure into a maximum heap, and popping out the element with the highest probability in the maximum heap;

[0018] And so on, each time the element with the highest probability is popped out from the maximum heap, and then its substructure is generated. If the substructure is completely filled, it is put into the pre-terminal list as a pre-terminal, waiting for the accumulated pre-terminals in the pre-terminal list to reach the set number, and then sent to the main process to generate a specific guess password.

[0019] The algorithm used in traditional password generation is prone to memory overflow problems under non-ideal conditions. The above scheme designs a password generation algorithm based on memory optimization. The algorithm uses variable dynamic nodes to delay the rhythm of low-probability nodes entering the priority queue, reduce the number of nodes added to the maximum heap, reduce memory usage, and improve the speed of password generation.

[0020] In a second aspect, an embodiment of the present invention provides a directional password guessing system based on user personal information, including:

[0021] A data preprocessing module is used to preprocess the obtained data set to remove invalid data;

[0022] The statistical analysis module is used to classify personal information into 6 major categories with a total of 36 subcategories, and to perform statistical analysis on the data set obtained after preprocessing based on the classification method to obtain the composition structure and probability of the password;

[0023] The password generation module is used to generate a guessed password based on the personal information input by the user and the composition structure and probability of the password.

[0024] In a third aspect, the present invention provides a computer program product, comprising computer-readable instructions, characterized in that the computer-readable instructions, when executed by a processor, implement the steps of the method for directional password guessing based on user personal information of the present invention.

[0025] In a fourth aspect, the present invention provides a computer-readable storage medium comprising computer-readable instructions, characterized in that the computer-readable instructions, when executed by a processor, implement the steps of the method for directional password guessing based on user personal information of the present invention.

[0026] In a fifth aspect, the present invention provides an electronic device, comprising: a memory storing program instructions; a processor connected to the memory, executing the program instructions in the memory, and implementing the steps of the directional password guessing method based on user personal information of the present invention.

[0027] Compared with the prior art, the present invention has the following technical advantages:

[0028] (1) The present invention divides personal information into 6 categories and further divides it into 36 subcategories, solving the problem that the traditional directional password cracking model is too rough in the classification of personal information and has a large number of repeated statistics. At the same time, it breaks away from the limitation of using digits to represent personal information and directly captures and matches the semantics. Therefore, the present invention can more realistically reflect the distribution law of personal information in passwords and achieve a better cracking effect.

[0029] (2) The present invention designs a password generation algorithm based on memory optimization. The algorithm uses variable dynamic nodes to delay the rhythm of low-probability nodes entering the priority queue, reduce the number of nodes added to the maximum heap, reduce memory occupancy, and solve the problem that traditional password generation algorithms may have memory overflow under non-ideal conditions.

[0030] For other advantages of the present invention, please refer to the relevant description in the embodiment section. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1 The flowchart is a method for directional password guessing based on user personal information given as an example in an embodiment of the present invention.

[0033] Figure 2 Flow chart of performing statistical analysis based on a data set in an embodiment.

[0034] Figure 3 This is a data structure diagram of the personal information used as an example in the embodiment.

[0035] Figure 4 Flow chart of the process of guessing password generation in the embodiment.

[0036] Figure 5 A schematic diagram of a sub-node list used as an example in the embodiment.

[0037] Figure 6 The figure is a block diagram of a directional password guessing system based on user personal information in an embodiment of the present invention.

[0038] Figure 7 A block diagram of the components of an electronic device. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0040] See also Figure 1 In this embodiment, a method for guessing a directional password based on user personal information is provided, comprising the following steps:

[0041] S10, preprocessing the obtained data set to remove invalid data.

[0042] In this step, according to application requirements, the preprocessing operation may specifically include the following steps:

[0043] S101, read the data set line by line, remove duplicate data, and store the retained data in the data set.

[0044] Especially when there are multiple data sets, the possibility of duplicate data is high. This step can effectively remove duplicate data and reduce the amount of useless data.

[0045] S102, further screening the data in the data set to remove all other characters that are not legal characters of the PCFGs algorithm.

[0046] Legal characters for the PCFGs algorithm include letters az and AZ, numbers 0-9, and special characters (specifically including !@#$%^&*()-_=+[]{};':”,. / <>?).

[0047] S103, remove incomplete data due to missing "personal information", such as ID card numbers with insufficient digits or not meeting the standards.

[0048] S104, counting the number of digits of the names in the data set, and eliminating a very small number of names with 4 or more digits, to facilitate subsequent algorithm processing.

[0049] S20, divide the personal information into 6 major categories with a total of 36 subcategories, and perform statistical analysis on the data set obtained after preprocessing based on the classification method to obtain the composition structure and probability of the password.

[0050] See also Figure 2 Specifically, this step may include the following steps:

[0051] S201, reading data in a data set line by line to obtain user information of the same user.

[0052] After preprocessing, information belonging to the same user is stored in the same row. Reading data row by row means processing information of different users in sequence.

[0053] S202, segment the read character string to obtain the plain text password, name, ID number, user name, mobile phone number and email address, and escape and save the segmented personal information according to the set data structure.

[0054] Take the 12306 data set as an example (to avoid privacy leakage, the real information is processed accordingly), assuming that the data in a row is: 123456788@qq.com----6837605----Zhang San----123456199503047890----z6837605----15012345678----123456788@qq.com. When splitting, remove the "----" in the string to get each data item.

[0055] It can be seen that 6837605 is the plain text of the password, and the remaining data items are all personal information, so the remaining data items are escaped and saved.

[0056] See also Figure 3 In this embodiment, personal information is divided into six categories, including name, ID number, birthday (which can be obtained from the ID number), user name, mobile phone number and email prefix. These six categories are further divided into 36 subcategories, each of which is represented by an identifier, which is the personal information tag.

[0057] For example, for the name category, it is divided into 11 subcategories, N1 to N 11 For example, the full name is represented by N1, and the initials of the name are represented by N2. Before escaping, use the "PyPin" library in Python to obtain the pinyin of each Chinese character from the real name, set the function parameters to output the pinyin without tone, and then follow Table 1 or Figure 3The data structure shown generates 11 kinds of strings, which are escaped and saved. For example, in the above example, for the data item "Zhang San", first obtain "zhangsan", and then generate 11 structures such as "zhangsan", "zs", "zhang", etc., and correspondingly escape them as N1, N2, N3, etc., and save them.

[0058] Another example is for the large category of user names, which is divided into 3 small categories, represented by U1 to U3 respectively. For example, the full user name composed of letters and numbers is represented by U1, the letter string of the user name is represented by U2, and the number string of the user name is represented by U3. For example, in the above example, for the data item "z6837605", first separate the character and number strings through regular expressions, and then generate 3 strings according to the data structure shown in Table 1 or Figure 3 shown, which are "z6837605", "z", "6837605" respectively, and then correspondingly escape them as U1, U2, U3, and save them.

[0059] Another example is for the large category of email prefixes, which is divided into 3 small categories, represented by E1 to E3 respectively. For example, the complete email prefix composed of letters and numbers is represented by E1, the continuous letter string of the email prefix is represented by E2, and the continuous number string is represented by E3. For example, in the above example, for the data item "123456788@qq.com", first separate the character and number strings through regular expressions, and then generate 1 (because there is no letter string in this email prefix) string according to the data structure shown in Table 1 or Figure 3 shown, which is "123456788", and then correspondingly escape it as E1, and save it.

[0060] Another example is for the large category of mobile phone numbers, which is divided into 4 small categories, represented by T1 to T4 respectively. For example, the complete mobile phone number is represented by T1, the first three digits are represented by T2, the middle three digits are represented by T3, and the last three digits are represented by T4. For example, in the above example, for the data item "15012345678", according to the data structure shown in Table 1 or Figure 3 shown, generate 4 strings, which are "15012345678", "150", "1234", "5678" respectively, and then correspondingly escape them as T1, T2, T3, T4, and save them.

[0061] Another example is for the large category of ID numbers, which is divided into 3 small categories, represented by I1 to I3 respectively. For example, the first six digits are represented by I1, the last six digits of the ID number are represented by I2, and the last four digits are represented by I3. For example, in the above example, for the data item "123456199503047890", according to the data structure shown in Table 1 or Figure 3The data structure shown generates three character strings, namely "123456", "19950304", and "7890", which are then escaped to I1, I2, and I3 respectively and saved.

[0062] For example, for the birthday category, it is divided into 12 subcategories, B1 to B 12 For example, birthdays of yyyy-mm-dd are represented by B1, birthdays of mm-dd-yyyy are represented by B2, birthdays of mm-dd are represented by B4, and birthdays of m-dd are represented by B 12 For example, in the above example, for the birthday "19950304", according to Table 1 or Figure 3 The data structure shown generates 12 structures such as "19950304", "03041995", "1995304", etc., and converts them into B1, B2, B3, etc., and saves them. Since personal information often does not have birthday information alone, but birthday information can usually be obtained from the ID card, the user's birthday is first intercepted from the 7th to 14th digits of the ID card number, and then according to Table 1 or Figure 3 The personal information structure shown is escaped and saved.

[0063] More perfectly, personal information is escaped and saved according to the data structure shown in Table 1.

[0064] Table 1

[0065]

[0066]

[0067]

[0068] S203, matching the plain text password with the data structure obtained by escaping in step S202, and determining the personal information tag composition structure of the plain text password.

[0069] For example, assuming that the plain text of a password is "zs0304", it is detected that "zs" matches a data structure in the name, and the personal information tag corresponding to the data structure is N2; it is also detected that "0304" matches a data structure in the birthday, and the personal information tag corresponding to the data structure is B4. Therefore, the personal information tag composition structure of the plain text of the password is N2B4.

[0070] For another example, in the above example, for the data item of the plain text password "6837605", since it is composed of pure numbers, it can be matched directly with the data structure of other items instead of the data structure of the name, thereby reducing the amount of data processing. In the process of matching the data structure of other items, it is detected that "6837605" matches a data structure of the user name, and the personal information tag corresponding to the data structure is U3, so the personal information tag structure of the plain text password is U3.

[0071] For "user name" and "email prefix", these information usually include character and numeric strings and may be subsets of other information types. In this case, the character and numeric strings are separated by regular expressions, and the username is matched in order from left to right to see whether it is U1, U2, and U3, and the email prefix is ​​matched in order from left to right to see whether it is E1, E2, and E3.

[0072] During the matching process, a character string in the plain text of the password may successfully match multiple personal information tags, for example, E3 and U3 are matched at the same time. At this time, all the matched personal information tags are retained, that is, the tag composition structure containing E3 and the tag composition structure containing U3 are retained at the same time.

[0073] S204, using a matching algorithm based on the traditional PCFGs model to perform secondary matching on the remaining characters in the password plain text to obtain the final composition structure of the password plain text.

[0074] In this embodiment, the remaining characters are defined as a filling character string.

[0075] In the PCFGs model, password characters are divided into three categories according to their types: letter sequence L, number sequence D, and special sequence S. The PCFGs model classifies and marks the consecutive letters, numbers, and special symbols in the password into different types of sequence structures, and also indicates the length of each sequence.

[0076] For example, after segmenting a user's information, the data items obtained are:

[0077] Login email: zw1985a@163.com; Password: !W19850607zw; Name: Zhao Wu; ID number: 123456198506071234; User name: zhaowu1985; Mobile phone number: 17712345678; Binding email: zw1985a@163.com.

[0078] For the plain text password "!W19850607zw", in step S203, the special character ! does not belong to personal information, so it cannot be matched successfully; the escaped data structure does not contain the structure of the single letter "w", so the single letter w cannot be matched successfully, but 19850607 will be recognized as B1, and zw will be recognized as N2. Therefore, after executing step S203, the recognition result is: !WB1N2. Then execute this step S204, use the PCFGs model to recognize "!" as S1 and "W" as L1. The subscript 1 in S1 indicates that the length of the special string is 1, and the subscript 1 in L1 indicates that the length of the letter string is 1. Therefore, the final recognition result is: S1L1B1N2.

[0079] It is easy to understand here that if all characters in the password plain text are matched successfully after executing step S203, step S204 will not be executed, that is, step S204 will be executed only if there are remaining characters that have not been matched successfully after one match.

[0080] S205, determine whether the data in the data set has been read, that is, whether each row of data in the data set has been processed by steps S202-204. If not, return to step S201, that is, read the next row of data, and execute S202-S204 until each row of data in the data set has been processed; if yes, calculate the structural probability of each personal information tag and the structural probability of each filling string. The password probability can also be calculated based on the structural probability of all personal information tags that constitute the password plain text and the structural probability of all filling strings. The lower the password probability, the less likely it is to generate a password in this structural manner, and the lower the priority in the password generation link.

[0081] Structural probability refers to the proportion of a small category label to the corresponding large category label during the password plain text matching process. For example, the frequency of mm-dd-yyyy corresponding to the personal information label B2 is 40 times in the data set (that is, B2 is matched when 40 password plain texts are matched), while B1-B 12 There are a total of 17,300 (that is, 17,300 plaintext passwords that match B1-B 12 Any one of them), so the structural probability of the personal information tag B2 is calculated to be 40 / 17300=0.23%.

[0082] The probability of a padding string refers to the proportion of a padding string in the corresponding category during the password plain text matching process. For example, the padding string "W" appears 7 times in the data set, and there are a total of 1500 letter strings in the data set, then the probability of the padding string "W" is 7 / 1500 = 0.47%.

[0083] The password probability is obtained by multiplying the probabilities of all structures that make up the plain text of the password. For example, the probability of the password "!W19850607zw" is P(S1L1B1N2) = P(S1)*P(L1)*P(B1)*P(N2).

[0084] After executing step S205, not only the various structural probabilities can be obtained, but also the password composition structure dictionary and the filling character string dictionary can be obtained.

[0085] S30, generating a guessed password based on the personal information input by the user and the composition structure and probability of the password.

[0086] More specifically, N password structures with the largest password probabilities are generated according to the structural probability of each personal information tag and the structural probability of each filling character string, and the password structures are converted into guessed passwords according to the personal information input by the user.

[0087] As an implementation method, the structural probabilities of personal information tags and filling character strings are sorted, m personal information tags and n filling character strings with larger structural probabilities are selected, and M password structures are obtained by combination, and the probabilities of obtaining the M password structures are calculated, and N password structures with larger probabilities are selected, and N guessed passwords are obtained according to the personal information escaped by the user. M, n, m, and N are all integers greater than 1.

[0088] N≤M, that is, only some guessed passwords with a higher probability are selected for display. The lower the probability, the lower the possibility that the user sets the password. Therefore, only some guessed passwords with a higher probability can avoid generating too many useless passwords. m is less than 36 (that is, the number of all personal information tags), and n is less than the number of all strings in the data set. By selecting only some personal information tags with a higher structural probability and filler strings to combine and generate the password structure, the amount of calculation can be reduced, while the success rate of guessing passwords can be guaranteed.

[0089] Or as another implementation, N password structures with higher password probabilities are selected from the data set, and N guessed passwords are obtained based on the personal information input by the user.

[0090] The higher probability mentioned in this article refers to sorting the probabilities, for example, sorting them from large to small, and selecting several probabilities with the highest sorting according to the number of requirements.

[0091] This step uses a multi-process and multi-threaded architecture when it is executed. The child thread reads the data containing personal information entered by the user and hands it over to the main thread to print the current guessing status. "Sub-process 1" and "Sub-process 2" communicate through queues and handle the storage of low-probability structures to reduce the queue size. "Sub-process 2" generates a pre-terminal and a main thread. A guessed password is generated based on the pre-terminal generated by "Sub-process 2". The pre-terminal is a password structure that is generated according to the algorithm and has not yet been fully filled. These password structures are intermediate steps in generating guessed passwords. They represent a node in the password generation process, but have not yet reached the final password form.

[0092] See also Figure 4 More specifically, this step may include the following steps:

[0093] S301, generating N password structures with the highest password probabilities according to a memory-optimized password generation algorithm.

[0094] In this embodiment, the step of generating a password structure may include the following steps:

[0095] S3011, when the subprocess 2 generates the initialization of the pre-terminal, it will find the index of the START node in the grammar, for example, 92 as shown in the table below, and read the first replacement structure M of START into the queue. Depending on the type of node represented by the number, it can be a subscript in the Grammar list or a subscript in the replacements list.

[0096] The example pre-terminal attributes are shown in Table 2.

[0097] Table 2

[0098] Self.grammar

[92] ["replacements"][0] <dict0x19d86a13630;len=5> Function "Transparent" is_terminal False POS <list0x19d86a11c48;len=1> prob 0.4 values <list0x19d86a11c48;len=1>

[0099] At this time, there is a structure "[92,0,[]]" in the queue, recorded as node, 92 is the subscript of the START node, 0 represents a certain state or attribute of the node, and the empty list [] represents the child node list of the node. It is represented as a tree like Figure 5 shown.

[0100] S3012, pop up the structure node, perform "increment" and "expand" operations, and generate two structures "[92,1,[]]" and "[92,0,[[91,0,[]]]]", namely "node[1]+=1" and the expanded "node[2]" list. These two structures are the substructures of node. Because the node structure has not been fully filled and is not a pre-terminal, it will not be used as a pre-terminal structure.

[0101] S3013, put the two substructures of node into the max heap, and pop out the element with the highest probability "[92,1,[]]" in the max heap. Then generate the substructure of the element with the highest probability, and get "[92,2,[]]" and "[92,1,[[37,0,[]]]]" and put them into the max heap. Pop out the element with the highest probability "[92,2,[]]", generate its substructures "[92,3,[]]" and "[92,2,[[41,0,[]]]]" and put them into the max heap.

[0102] And so on. Each time the element with the highest probability is popped from the maximum heap, its substructure is generated according to the password generation algorithm based on parallel optimization. If the substructure is completely filled, that is, all the child nodes of the substructure have been correctly filled and no further filling operations need to be performed, then it is put into the pre-terminal list as a pre-terminal, and waits for the accumulated elements in the pre-terminal list to reach the set number, for example, 10, and then sent to the main process to generate a specific password guess. Setting the number of accumulated elements to 10 can make full use of the computing power of multi-core processors and improve processing efficiency.

[0103] S302: For the personal information tag in the password structure, the corresponding personal information is directly escaped.

[0104] For example, the password structure obtained by combination is S1L1B1N2, and the user information entered by the user is "Name: Zhao Wu; ID number: 123456198506071234; User name: zhaowu1985", then the user information corresponding to B1 is "19850607", and the user information corresponding to N2 is "zw", so the result after escaping is S1L119850607zw.

[0105] S303: For the remaining structure in the password structure, that is, the padding string label, the padding string with a higher probability counted in the data set is escaped.

[0106] For example, in the above example, for S1L119850607zw obtained after S302 processing, for the remaining S1, a special character with the highest structural probability is selected from the filling string counted in the data set for escape. For L1, a letter with the highest structural probability is selected from the filling string counted in the data set for escape. Assuming that the escaped special character is @ and the escaped letter is S, the guessed password is: @S19850607zw.

[0107] It is easy to understand here that step S303 is executed only when the password structure contains a padding string, and for a password structure consisting only of personal information tags, only steps S301-302 need to be executed.

[0108] See also Figure 6 , this embodiment also provides a directional password guessing system based on user personal information, including:

[0109] For more detailed processing operations of the above modules, please refer to the relevant description in the above method, which will not be repeated here.

[0110] like Figure 7 As shown, this embodiment also provides an electronic device, which may include a processor 41 and a memory 42, wherein the memory 42 is coupled to the processor 41. It is worth noting that this figure is exemplary, and other types of structures may be used to supplement or replace this structure to achieve data extraction, report generation, communication or other functions.

[0111] like Figure 7 As shown, the electronic device may further include: an input unit 43, a display unit 44 and a power supply 45. It is worth noting that the electronic device does not necessarily have to include Figure 7 In addition, electronic devices may also include Figure 7 For components not shown, reference may be made to the prior art.

[0112] The processor 41 is sometimes also called a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The processor 41 receives inputs and controls the operations of various components of the electronic device.

[0113] The memory 42 may be, for example, one or more of a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices, and may store information such as configuration information of the processor 41 and instructions executed by the processor 41. The processor 41 may execute the program stored in the memory 42 to implement information storage or processing. In one embodiment, the memory 42 also includes a buffer memory, i.e., a buffer, to store intermediate information.

[0114] An embodiment of the present invention further provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed in an electronic device, the program product enables the electronic device to execute the operation steps included in the method of the present invention.

[0115] An embodiment of the present invention further provides a storage medium storing computer-readable instructions, wherein the computer-readable instructions enable an electronic device to execute the operation steps included in the method of the present invention.

[0116] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0117] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-On l yMemory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0118] The above-described embodiments are only specific implementations of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications, replacements and improvements within the technical scope disclosed by the present invention, and these modifications, replacements and improvements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A directional password guessing method based on user personal information, characterized in that: The following steps are involved: S10, preprocessing the obtained data set to remove invalid data; S20, dividing the personal information into 6 major categories with a total of 36 subcategories, and performing statistical analysis on the data set obtained after preprocessing based on the classification method to obtain the composition structure and probability of the password; S30, generating a guessed password based on the personal information input by the user and the composition structure and probability of the password.

2. The method for directional password guessing based on user personal information according to claim 1, characterized in that: The pre-processing comprises the following steps: S101, reading the data set line by line, removing duplicate data, and storing the retained data in the data set; S102, further screening the data in the data set to remove all other characters that are not legal characters of the PCFGs algorithm; S103, remove incomplete data due to missing personal information; S104, counting the number of digits of names in the data set, and eliminating a very small number of data with 4 or more digits of names.

3. The method for directional password guessing based on user personal information according to claim 1, characterized in that: The step S20 comprises the following steps: S201, reading data in a data set line by line to obtain user information of the same user; S202, segmenting the read character string to obtain the plain text password, name, ID number, user name, mobile phone number and email address, and escaping and saving the segmented personal information according to the data structure set in the classification method; S203, matching the plain text of the password with the data structure obtained by escaping in step S202, and determining the personal information tag composition structure of the plain text of the password; S204, using a matching algorithm based on a traditional PCFGs model to perform secondary matching on the padding string in the password plain text to obtain a composition structure of the password plain text; S205, determine whether the data in the data set has been read, that is, whether each row of data in the data set has been processed by steps S202-204. If not, return to step S201, that is, read the next row of data, and execute S202-S204 until each row of data in the data set has been processed; if so, calculate the structural probability of each personal information tag, the structural probability of each filling character string and the password probability.

4. The method for directional password guessing based on user personal information according to claim 3, characterized in that: In step S30, N password structures with the largest password probabilities are generated according to the structural probability of each personal information tag and the structural probability of each filling character string, and the password structures are converted into guessed passwords according to the personal information input by the user.

5. The method for directional password guessing based on user personal information according to claim 4, characterized in that: The step S30 comprises the following steps: S301, generating N password structures with the highest password probabilities according to a memory-optimized password generation algorithm; S302, escaping the personal information tag in the password structure by the corresponding personal information; S303, for the remaining structures in the password structure, escape the letter string, special character string or digital string with a higher probability in the data set to obtain a guessed password.

6. The method for directional password guessing based on user personal information according to claim 5, characterized in that: The step S301 includes the following steps: S3011, when the child process generates the initialization of the pre-terminal, it finds the index of the START node in the grammar, and reads the first replacement structure of the START node into the maximum heap; S3012, pop out the element with the highest probability in the maximum heap, perform increment and expand operations, and generate a substructure of the element with the highest probability; S3013, putting the substructure into a maximum heap, and popping out the element with the highest probability in the maximum heap; And so on, each time the element with the highest probability is popped out from the maximum heap, and then its substructure is generated. If the substructure is completely filled, it is put into the pre-terminal list as a pre-terminal, waiting for the accumulated pre-terminals in the pre-terminal list to reach the set number, and then sent to the main process to generate a specific guess password.

7. A directional password guessing system based on user personal information, characterized in that: include: A data preprocessing module is used to preprocess the obtained data set to remove invalid data; The statistical analysis module is used to classify personal information into 6 major categories with a total of 36 subcategories, and to perform statistical analysis on the data set obtained after preprocessing based on the classification method to obtain the composition structure and probability of the password; The password generation module is used to generate a guessed password based on the personal information input by the user and the composition structure and probability of the password.

8. A computer program product comprising computer readable instructions, characterized in that: The computer-readable instructions, when executed by a processor, implement the steps in the method for directional password guessing based on user personal information as described in any one of claims 1 to 6.

9. A computer-readable storage medium comprising computer-readable instructions, characterized in that: The computer-readable instructions, when executed by a processor, implement the steps in the method for directional password guessing based on user personal information as described in any one of claims 1 to 6.

10. An electronic device, characterized in that: include: Memory, which stores program instructions; A processor is connected to the memory and executes program instructions in the memory to implement the steps of the directional password guessing method based on user personal information described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Password guessing set generating method and password cracking method based on user name information

    CN106803035A