A Structured Data Hierarchical Method, Device, and Storage Device
Through reservoir sampling and pre-order traversal optimization data identification process, the high cost and low accuracy of structured data hierarchical classification in the prior art are solved, and efficient and accurate hierarchical classification are achieved.
Patent Information
- Application Number
- CN202310522176.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-05-05
AI Technical Summary
The existing structured data hierarchical classification methods consume a lot of human resources and time costs, and are not very accurate, which has problems such as human factors and low data identification efficiency.
Data samples are obtained by tank sampling method, combined with preorder traversal method and hierarchical classification rules to optimize the data identification process, and efficient hierarchical classification is achieved through multi-threaded processing.
It reduces labor and time costs, improves the accuracy and efficiency of grading and classification, and reduces the influence of human factors.
Smart Images

Figure CN116701546B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data classification, and particularly to a method, device, and storage device for hierarchical classification of structured data. Background Art
[0002] Structured data is a type of data with a standardized format, which has a highly organized and neat structure.
[0003] Generally, this type of data is data logically expressed and implemented by a two-dimensional table structure, strictly following data format and length specifications, and mainly stored and managed through a relational database.
[0004] In relevant regulations, it is necessary to classify and grade this type of structured data stored and managed through a relational database. Currently, the general classification and grading methods often consume a large amount of human resources and time costs, and the results have large errors.
[0005] Currently, the classification and grading processing methods for structured data mainly include two types:
[0006] 1. Arrange database management specialists to access various database services within the company, and perform statistical classification and grading on the data field by field, table by table, and database by database.
[0007] The advantage of this solution is that it can ensure that each piece of data classified and graded manually actually implements the classification and grading standard specifications, and performs accurate classification and grading according to the promulgated data protection measures. However, the problems are that relying on human resources for classification and grading will consume a large amount of human resources and time costs, and there will also be human factors affecting, resulting in the lack or absence of classification and grading of some unstructured data, and it will also be affected by the personal experience of database management specialists, resulting in misjudgment and missed judgment.
[0008] 2. Rely on a specific program system, use methods such as JDBC interface, SSL protocol, SSH authentication, or HTTP channel to remotely connect to the database. After connection, collect all or part of the data field by field, table by table, and database by database for the target database content, and then traverse the structured data according to specific classification and grading rule sets for the collected data, so as to classify and grade the structured data.
[0009] The advantage of this solution is that the specific program system can enable multi-threading to simultaneously collect data from different databases, different tables, and different fields, and then classify and grade the target structured data, reducing human resources and time costs. However, the disadvantages are also obvious:
[0010] ①In the data collection stage: If all data is selected for collection, then the amount of data obtained is large, imposing a relatively high memory pressure on the server; if only partial data is selected for collection, then the collected data may lack characteristics, universality, and effectiveness, being a relatively one-sided set of data. ②In the data recognition stage, double traversal is required. Not only do we need to traverse the set of hierarchical classification rules, but we also need to traverse the data set in the target database. When both are fully traversed, the resulting hierarchical classification will be a Cartesian product. This resulting set will also be a large amount of data, and the efficiency is not high enough. Summary of the Invention
[0011] To solve the problems of the existing hierarchical classification methods that consume a large amount of human resources, time costs, and have insufficient accuracy, the present invention is based on the existing program system:
[0012] ①Optimize the data collection module by using a set of model algorithms to enhance the characteristics, universality, and effectiveness of the collected data, thereby improving the accuracy of hierarchical classification.
[0013] ②Optimize the problem of the large time cost caused by double traversal in the data recognition stage by using a set of priority algorithms and a supporting data recognition scheme, thereby reducing the cost consumption in terms of manpower and time.
[0014] Specifically, the present invention provides a method, device, and storage device for hierarchical classification of structured data. The method specifically includes the following steps:
[0015] S1. Configure the target structured database.
[0016] S2. Remotely link to the target structured database.
[0017] S3. Use the reservoir sampling method to obtain data samples from the target structured database.
[0018] S4. Obtain the preliminary range of hierarchical classification rules based on the attributes of the data samples, and use the preorder traversal method to sort the hierarchical classification rules in the preliminary range to obtain a set of hierarchical classification rules.
[0019] S5. According to the set of hierarchical classification rules, use a supporting data recognition method to obtain the hierarchical classification results.
[0020] S6. Archive and store the hierarchical classification results.
[0021] A storage device stores instructions and data for implementing a method for hierarchical classification of structured data.
[0022] A structured data grading device, comprising: a processor and a storage device; the processor loads and executes instructions and data in the storage device to implement a structured data grading method.
[0023] The beneficial effects provided by the present invention are: reducing the cost consumption of manpower and time during data grading, and improving the accuracy of grading and classification at the same time. Description of the Drawings
[0024] Figure 1 is a flowchart of the method of the present invention;
[0025] Figure 2 is a working diagram of the hardware device of the present invention. Detailed Embodiments
[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described below in conjunction with the accompanying drawings.
[0027] Please refer to Figure 1 , Figure 1 which is a flowchart of the method of the present invention.
[0028] The present invention provides a structured data grading method, device and storage device, wherein the method specifically comprises the following steps:
[0029] S1. Configure a target structured database;
[0030] In step S1, configuring the target structured database specifically includes: configuring the ip address, port number, linked database name, database type and linking method of the target structured database.
[0031] For example:
[0032] Linking method selection: jdbc link;
[0033] Database type selection: mysql database (other types such as: pgsql, oracle, SqlServer, etc.);
[0034] Linked database name: the instance name your_database of the scanned target database;
[0035] Database ip address: 192.168.182.153;
[0036] Port number: 3306;
[0037] S2. Remotely link the target structured database;
[0038] It should be noted that in step S2, when remotely linking to the target structured database, according to the link method, the connection to the target structured database is completed through the account password, protocol, authentication certificate, channel account, and proxy server for accessing the specified table permissions.
[0039] For example, various methods such as using the JDBC interface or SSL protocol or SSH authentication or HTTP channel are used to remotely link to the database.
[0040] S3. Obtain a data sample of the target structured database by using the reservoir sampling method;
[0041] It should be noted that in the present invention, the reservoir sampling algorithm is used to optimize the data collection process, only collecting partial feature data, avoiding obtaining full-scale data or one-sided data, and improving the sample acquisition speed and sample characteristics, universality, and effectiveness.
[0042] Step S3 is specifically as follows:
[0043] S31. Obtain the total number of rows m of the specified data table in the target structured database;
[0044] S32. Determine whether the total number of rows m is less than or equal to the preset sampling quantity n. If so, obtain all the data of the specified data table as the data sample; otherwise, go to step S33;
[0045] S33. Set the capacity of the data pool to m according to the total number of rows m; set the capacity of the reservoir to n according to the sampling quantity n; initialize the n data in the reservoir as the first n data in the data pool;
[0046] S34. Initialize j to 1;
[0047] S35. Determine whether j is less than the data pool capacity m. If so, go to step S36; otherwise, go to step S38;
[0048] S36. Obtain a random integer d within [0, j], where the initial value of j is 1. If d falls within the range [0, n - 1], write the data at the j-th position in the data pool to the (d + 1)-th position in the reservoir;
[0049] S37. Increment j by 1 and return to step S35;
[0050] S38. Obtain a reservoir with an integer array whose data content value range is [1, m] and is obtained with equal probability, and the length of the array is n;
[0051] S39. Obtain the row number data corresponding to the integers in the reservoir from the target database as the data sample.
[0052] In the above steps S31 to S39, m is the size of the data pool; n is the size of the reservoir; j is the number of loops, and it also represents the sampling position when looping m times. d is a random integer within the range of [0, j]. If d ∈ [0, n - 1], it means a hit and replacement is required; otherwise, no replacement is done.
[0053] For example, if the data pool is [1, 2, 3, 4, 5] and the size of the data pool m = 5; the reservoir is temporarily unknown, [X1, X2, X3], and the size of the reservoir n = 3.
[0054] When the number of loops j is 3, the random integer d = 3. At this time, d is not within the range of [0, n - 1], indicating that the data for this time is not hit and no operation is performed.
[0055] When the number of loops j is 5 and the random integer d = 2, at this time d is within the range of [0, n - 1]. Write the number at the j-th position in the data pool, which is 5, to the position d + 1 in the reservoir, that is, the third cell in the reservoir, that is, X3 is written with 5.
[0056] Generally speaking, step S3 of the present invention first implements a reservoir algorithm function to obtain an array set with "a data content within the set of [1, the total data volume of the table], with equal probability and a length of the sample data volume", and then obtains the data corresponding to the row number of the target table according to the values in the array set as the sample set.
[0057] This set of reservoir sampling algorithm solutions can logically ensure that each sampled sample is obtained with equal probability among all the data, and the probability of obtaining each sample is "sample quantity / total data volume of the table * 100%".
[0058] S4. Obtain the preliminary range of the classification rules according to the data sample attributes, and use the pre-order traversal method to sort the classification rules in the preliminary range to obtain a set of classification rules;
[0059] It should be noted that step S4 specifically obtains a set of classification rules that are screened according to conditions such as the target field name, field comment, table name, table comment, library name, library comment, etc., and sorted by pre-order traversal according to the priority and creation time.
[0060] The specific process of step S4 is as follows:
[0061] S41. Determine whether the data sample has attributes. If so, match the meaning of each rule attribute in the classification library according to the data sample attribute information to obtain the preliminary range of classification; otherwise, obtain all the rule sets in the classification library and enter step S43;
[0062] It should be noted that the rules of the hierarchical classification library are as follows: Based on the national hierarchical protection measures, the characteristics of each type of data are manually summarized and compiled into a set of rules that are convenient for computers to identify. For example, taking the mobile phone number of a person as the target, assuming that it belongs to the classification of personal information and the second-level data according to the hierarchical protection measures. Then its corresponding rule is expressed as:
[0063] When the data uses the regular expression: "^(13[0-9]|14[01456879]|15[0-35-9]|16
[2567] |17[0-8]|18[0-9]|19[0-35-9])\\d{8}$", the data that can be matched can be recognized as his mobile phone number. Therefore, this group of data will be marked as the classification of personal information and the second-level data. This is a simple example here, and the specific rules will be more complex, and support multi-condition re-judgment of AND, OR, and NOT, and can more accurately identify the hierarchical classification of the target data.
[0064] It should be noted that the attributes of the data sample include: field name, field comment, table name, table comment, library name, and library comment;
[0065] For example, for the mysql database, if you need to extract the cell data in the mysql database, you need to obtain which library, which table, and which field the target data is in; then the cell data of the table at this time, that is, the attributes of the data source sample, are the library name, library comment, table name, table comment, field name, field comment, table name, and table comment where it is located.
[0066] The hierarchical classification library described in step S4 is completed through pre-configuration. The specific process is as follows: Use the attribute or content matching method to perform fuzzy or full matching or regular matching mode, and configure the recognition and judgment rules for each hierarchical classification.
[0067] It should be noted that, for example, a complete hierarchical classification library is as follows:
[0068] First-level classification
[0069] Second-level classification 01
[0070] Third-level classification 01 (hooked to the first-level classification)
[0071] The rules below it
[0072] Rule 1: The content satisfies the regular expression "xxxx", and satisfies that the content contains the field "123456";
[0073] Rule 2: The content satisfies the regular expression "yyyy", and satisfies that the content contains the field "654321";
[0074] Second-level classification 02
[0075] Tertiary Classification 02 (Hook Secondary Classification)
[0076] The rules below it
[0077] Rule 1: The content satisfies the regular expression "xxxx" and contains the field "123456".
[0078] Rule 2: The content satisfies the regular expression "yyyy" and contains the field "654321".
[0079] S42. According to the preliminary scope of hierarchical classification, obtain the rule list in pre-order traversal order according to the priority and creation order marked for hierarchical classification when creating the hierarchical classification library; the priority is the preset first level, second level, and third level; the creation order refers to the sequence of creation times.
[0080] The description of the rule list for pre-order traversal is as follows:
[0081] For the rule list of pre-order traversal, the creation process is to guess the possible hierarchical classification range of the target using the attribute field names, field comments, table names, table comments, library names, and library comments of the sample data.
[0082] For example, for the sample data 15926436889, its attributes are: field name phone, field comment, mobile phone number, table name user, and table comment user table. Based on this, it can be inferred that this sample belongs to the data of personnel classification.
[0083] Then the obtained rule of pre-order traversal is the hierarchical classification of the range inferred from the sample data attributes in the total hierarchical classification library, and then the rule list is obtained by sorting according to pre-order traversal. The purpose of pre-order traversal is to sort the rules.
[0084] S43. For all the rule sets in step S41 and the rule list of pre-order traversal in step S42, use the pre-order traversal method to obtain the hierarchical classification rule set.
[0085] It should be noted that step S43 is specifically as follows:
[0086] S431. Call the recursive function to obtain the input data of the recursive function; the input data includes: the set A to be sorted, the sorted result set B, and the parent id in the recursion; the set A is all the rule sets to be sorted or the rule list of pre-order traversal.
[0087] It should be noted that in the input data, the input sorted result set B is empty.
[0088] S432. Traverse the set A, group the data in the set A whose parent id is equal to the parent id in the input data of the recursion, and store them in the temporary set C.
[0089] S433. Determine whether the length of the temporary set C is less than 1. If so, jump out of the recursion and output the sorted result set B as the hierarchical classification rule set; otherwise, proceed to step S434;
[0090] S434. Traverse set A and remove the part of set A that intersects with C;
[0091] S435. Sort the content of set C according to the priority and creation time;
[0092] S436. Traverse set C, put the content of a single loop of set C into the result set B, and recursively call to determine whether the preset number of recursions has been reached. If so, output the result set B as the hierarchical classification rule set; otherwise, return to step S431 and input the parent id in sets A, B, and C.
[0093] For the above recursive process, when entering for the first time
[0094] 1. The set to be sorted is named A
[0095] 2. The sorted result set is named B
[0096] 3. The parent id, and the initial parent id is a default value such as 0
[0097] If the length of C is less than 1 at the beginning, it indicates that there is no required root node in set A.
[0098] The purpose of the recursive algorithm is to obtain set B with the specified parent id as the root node from set A.
[0099] S5. According to the hierarchical classification rule set, adopt the supporting data recognition method to obtain the hierarchical classification result;
[0100] Step S5 is specifically as follows:
[0101] S51. Obtain each row of data in the data sample in step S3;
[0102] S52. Input each row of data into the hierarchical classification rule set in step S4 for loop matching, and match whether each data in each row of data conforms to the data rules in the hierarchical classification rules. If so, mark it; otherwise, skip it;
[0103] S53. When the rule hit rate in the hierarchical classification rule set is greater than the preset value and the number of marked times is greater than the preset number of times, jump out of the loop matching;
[0104] S54. Obtain the hit situation of each data sample in the hierarchical classification rule set, and set the hierarchical classification with the highest hit rate and the highest priority as the default hierarchical classification result of this structured database.
[0105] S6. Archive and store the hierarchical classification result.
[0106] As an embodiment, the task flow of automatic hierarchical classification of structured data is as follows:
[0107] (1) According to the configured database information, use the corresponding link method to remotely link the target database; execute the statements of the target database to obtain the structure of the target database.
[0108] (2) According to the template database table structure, perform multi-threaded processing by database and by table; in multi-threaded processing, set the number of threads according to the capabilities of the server cores.
[0109] (3) Thread 1 executes the thread for table a in database A; Thread 2 executes the thread for table b in database B...
[0110] Each thread adopts the same processing method, which is specifically as follows:
[0111] (4) Data sampling: Invoke the reservoir sampling method to obtain the data samples of the target database (for the specific process, see steps S31 - S39).
[0112] (5) Hierarchical classification set: Invoke the pre-order traversal sorting method to obtain the hierarchical classification set (for the specific process, see steps S41 - S43, S431 - S436).
[0113] (6) Data recognition: Obtain the hierarchical classification result according to the supporting data recognition method.
[0114] (7) Archive and storage.
[0115] Please refer to Figure 2 , Figure 2 is the schematic diagram of the working of the hardware device in the embodiment of the present invention. The hardware device specifically includes: a structured data grading device 401, a processor 402, and a storage device 403.
[0116] A structured data grading device 401: The structured data grading device 401 implements the structured data grading method.
[0117] Processor 402: The processor 402 loads and executes the instructions and data in the storage device 403 to implement the structured data grading method.
[0118] Storage device 403: The storage device 403 stores instructions and data; the storage device 403 is used to implement the structured data grading method described above.
[0119] The beneficial effects of the present invention are as follows:
[0120] Compared with the manual identification by a specialist, this solution solves the problems of high human resource consumption, large time cost consumption, and significant differences in results caused by human factors in the specialist's processing. By using an efficient system multi-threading mechanism, the labor cost and time cost are greatly reduced; through the grading and classification rules set by the system, the misjudgment and omission caused by personal experience in manual identification are reduced.
[0121] Compared with the identification by an ordinary system, this solution also proposes a systematic data collection scheme, a grading and classification rule acquisition scheme, and a grading and classification identification scheme. Through the reservoir sampling algorithm in the data collection scheme, the sample data obtained is more characteristic, universal, and effective, eliminating the acquisition of all data or one-sided data in the conventional collection, and improving the expressiveness of the sample; through the grading and classification rule scheme and the supporting grading and classification identification scheme, the scope of the grading and classification rules to be matched is narrowed and accurately positioned, reducing the overall number of identifications and improving the grading and classification identification efficiency. It can optimize the time required for the overall identification task, improve the identification efficiency, and reduce the time cost while ensuring the identification quality.
[0122] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for classifying structured data, characterized in that: It includes the following steps: S1. Configure the target structured database; S2. Remotely link to the target structured database; S3. Use the reservoir sampling method to obtain a data sample of the target structured database; S4. Obtain the preliminary range of the hierarchical classification rules based on the data sample attributes, and use the preorder traversal method to sort the hierarchical classification rules in the preliminary range to obtain a set of hierarchical classification rules; The specific process of step S4 is as follows: S41. Determine whether the data sample has attributes. If so, match the meaning of each rule attribute in the hierarchical classification library according to the data sample attribute information to obtain the preliminary range of hierarchical classification; otherwise, obtain all the rule sets in the hierarchical classification library and go to step S43; S42. According to the preliminary range of hierarchical classification, obtain the rule list for preorder traversal according to the priority and creation order marked for hierarchical classification when creating the hierarchical classification library; the priority is the preset first level, second level, and third level; the creation order refers to the sequence of creation times; S43. For all the rule sets in step S41 and the rule list for preorder traversal in step S42, use the preorder traversal method to obtain a set of hierarchical classification rules; S5. According to the set of hierarchical classification rules, use the supporting data recognition method to obtain the hierarchical classification result; The specific details of step S5 are as follows: S51. Obtain each row of data in the data sample in step S3; S52. Input each row of data into the set of hierarchical classification rules in step S4 for loop matching, and check whether each data in each row of data conforms to the data rules in the hierarchical classification rules. If so, mark it; otherwise, skip it; S53. When the rule hit rate in the set of hierarchical classification rules is greater than the preset value and the number of marked times is greater than the preset number of times, jump out of the loop matching; S54. Obtain the hit situation of each data sample in the set of hierarchical classification rules, and set the hierarchical classification with the highest hit rate and the highest priority as the default hierarchical classification result of the structured database; S6. Archive and store the hierarchical classification results.
2. The structured data grading method according to claim 1, characterized in that: In step S1, configuring the target structured database specifically includes: configuring the ip address, port number, linked database name, database type, and link method of the target structured database.
3. The structured data grading method according to claim 2, wherein: In step S2, when remotely linking to the target structured database, according to the link method, complete the connection with the target structured database through the account password, protocol, authentication certificate, channel account, and proxy server for accessing the specified table permissions.
4. The structured data grading method according to claim 1, characterized in that: The specific details of step S3 are as follows: S31. Obtain the total number of rows m of the specified data table in the target structured database; S32. Determine whether the total number of rows m is less than or equal to the preset sampling quantity n. If so, obtain all the data in the specified data table as the data sample; otherwise, go to step S33; S33. Set the capacity of the data pool to m according to the total number of rows m; set the capacity of the reservoir to n according to the sampling quantity n; initialize the n data in the reservoir as the first n data in the data pool; S34. Initialize j to 1; S35. Determine whether j is less than the capacity m of the data pool. If so, go to step S36; otherwise, go to step S38; S36. Obtain a random integer d within [0, j], where the initial value of j is 1. If d falls within the range [0, n - 1], write the data at the j-th position in the data pool to the (d + 1)-th position in the reservoir. S37. Increment j by 1 and return to step S35. S38. Obtain a reservoir with an integer array whose value range of the data content is [1, m] and is obtained with equal probability, and the length of the array is n. S39. Obtain the row number data corresponding to the integers in the reservoir from the target database as the data sample.
5. A method for classifying structured data as claimed in claim 1, characterized in that: The hierarchical classification library described in step S4 is completed through pre-configuration. The specific process is as follows: Use the attribute or content matching method to perform fuzzy or full matching or regular matching mode to configure the recognition and judgment rules for each hierarchical classification.
6. The structured data grading method according to claim 1, wherein: Step S43 is specifically as follows: S431. Call the recursive function to obtain the input data of the recursive function; the input data includes: the set A to be sorted, the sorted result set B, and the parent id in the recursion; the set A is the entire set of rules to be sorted or the pre-order traversal rule list. S432. Traverse the set A, group the data in the set A whose parent id is equal to the parent id in the input data recursion, and store them in the temporary set C. S433. Judge whether the length of the temporary set C is less than 1. If so, jump out of the recursion and output the sorted result set B as the hierarchical classification rule set; otherwise, go to step S434. S434. Traverse the set A and remove the part where A intersects with C from the set A. S435. Sort the content of the set C according to the priority and creation time. S436. Traverse the set C, put the content of the single loop of the set C into the result set B, and recursively call to judge whether the preset recursion times are reached. If so, output the result set B as the hierarchical classification rule set; otherwise, return to step S431, input the set A, the set B, and the parent id in the set C.
7. A storage device, characterized in that: The storage device stores instructions and data for implementing a structured data hierarchical method according to any one of claims 1 to 6.
8. A structured data grading device, characterized in that: Including: A processor and a storage device; the processor loads and executes the instructions and data in the storage device for implementing a structured data hierarchical method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data asset classification modeling and grading protection method based on artificial intelligence technology
CN113590698A