A data classification grading method, device, equipment and storage medium

By combining dictionary matching and manual methods to generate new data classification and grading rules, the problems of low efficiency and high cost in existing data classification and grading technologies are solved, achieving efficient and accurate data classification and grading. Furthermore, historical data dictionaries can be reused in new projects, reducing manpower and time costs.

CN116127372BActive Publication Date: 2026-07-03HANGZHOU DBAPPSECURITY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211595528.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2026-07-03
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

Existing data classification and grading methods are inefficient, costly, unstable, inaccurate, and have low coverage, resulting in large errors in data classification and grading results, missed sensitive data, and a large amount of manpower required for repeated data collection and processing, which leads to a waste of human resources.

Method used

By combining dictionary matching and manual methods to process the data to be classified and graded, an initial classification and grading result is generated. New classification and grading rules are generated using preset classification and grading rules, and the historical data dictionary is updated using data dictionary maintenance methods. A conflict data dictionary is constructed to obtain the final classification and grading result.

Benefits of technology

It improves the efficiency of data classification and grading, reduces manpower and time costs, enhances the accuracy of classification and grading, and enables the reuse of knowledge accumulated in historical data dictionaries in new projects, reducing the difficulty of building and maintaining data dictionaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116127372B_ABST
    Figure CN116127372B_ABST
Patent Text Reader

Abstract

This application discloses a data classification and grading method, apparatus, device, and storage medium based on knowledge accumulation, relating to the field of data security governance. The method includes: acquiring data to be classified and graded; classifying and grading the data using a dictionary matching method based on a historical data dictionary and a manual method to obtain a first classification and grading result and a second classification and grading result, and merging them to obtain a third classification and grading result; processing the second classification and grading result and the data to be classified and graded based on a preset classification and grading rule generation specification to obtain a first classification and grading rule; updating the historical data dictionary based on preset data dictionary maintenance rules and the first classification and grading rule to obtain the current output data dictionary, and constructing a conflict data dictionary; updating the third classification and grading result based on the conflict data dictionary to obtain the final classification and grading result. In this way, this application can reuse and update the data dictionary manually, avoiding the waste of manpower.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data security governance, and in particular to a data classification and grading method, apparatus, device, and storage medium based on knowledge accumulation. Background Technology

[0002] The rapid development of modern information technology and the efficient and convenient information services have brought about a complete revolution. The widespread adoption of the internet has improved people's quality of life and work; however, it has also made people more vulnerable to data security and privacy breaches. Data classification and grading are prerequisites for data security governance. The data classification and grading process is the most important step in determining whether data is compliant. Only by effectively classifying and grading data can we avoid a one-size-fits-all approach to control and adopt more refined measures in data security management, achieving a balance between shared and secure data use.

[0003] Many common data classification and grading methods, such as manual methods and dictionary matching, suffer from drawbacks such as low efficiency, high cost, poor stability, inaccuracy, and low coverage. These often result in large errors in data classification and grading results and the omission of sensitive data. Furthermore, these methods involve repeated data collection and processing at significant human cost, failing to translate into tangible human output and thus causing substantial waste of human resources. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a data classification and grading method, apparatus, device, and storage medium based on knowledge accumulation, which can reuse the data dictionary obtained from previous data classification and grading tasks, avoiding the waste of the work results of manual data classification and grading. The specific solution is as follows:

[0005] Firstly, this application provides a data classification and grading method based on knowledge accumulation, including:

[0006] The data to be classified and graded is obtained, and the data is classified and graded using dictionary matching based on the historical data dictionary of the previous output and manual methods, respectively. The first classification and grading results are then merged with the second classification and grading results to obtain the third classification and grading result.

[0007] The first classification and grading rule is obtained by processing the second classification and grading result and the corresponding data to be classified and graded based on the preset classification and grading rules.

[0008] The historical data dictionary is updated based on the preset data dictionary maintenance rules and the first classification and grading rules to obtain the data dictionary for this output, and a conflict data dictionary is constructed based on the data conflicts between the first classification and grading rules and the historical data dictionary.

[0009] The third classification and grading results are updated based on the conflict data dictionary to obtain the final classification and grading results corresponding to the data to be classified and graded.

[0010] Optionally, the classification and grading of the data to be classified and graded using dictionary matching based on the historical data dictionary of the previous output and manual methods respectively includes:

[0011] The first classification and grading result is obtained by processing the data covered by the historical data dictionary in the data to be classified and graded using the dictionary matching method.

[0012] The second classification result is obtained by processing the data in the data to be classified and graded that is not covered by the historical data dictionary using the manual method.

[0013] Optionally, the step of processing the second classification and grading result and the corresponding data to be classified and graded based on the preset classification and grading rules to obtain the first classification and grading rules includes:

[0014] Preprocessing is performed on the data to be classified and graded corresponding to the second classification and grading result to obtain preprocessed data;

[0015] Extract keywords from the preprocessed data, and select target keywords from all extracted keywords that correspond to each single classification and grading result in the second classification and grading result to obtain the corresponding keyword set;

[0016] The second classification and grading results and the keyword set are transformed into an expression of data classification and grading rules to obtain the first classification and grading rules.

[0017] Optionally, the step of preprocessing the data to be classified and graded corresponding to the second classification and grading result to obtain preprocessed data includes:

[0018] Identify the text information of the data to be classified and graded corresponding to the second classification and grading result, and filter the text information based on a preset thesaurus to obtain filtered text information;

[0019] The target text information that meets the preset conversion conditions in the filtered text information is converted to obtain the preprocessed data.

[0020] Optionally, the step of updating the historical data dictionary based on preset data dictionary maintenance rules and the first classification and grading rules to obtain the current output data dictionary, and constructing a conflict data dictionary based on data conflicts between the first classification and grading rules and the historical data dictionary, includes:

[0021] Copy all contents of the historical data dictionary and generate a new data dictionary;

[0022] Based on a preset rule conflict judgment method, it is determined whether there is a conflict between the target classification and grading rule in the first classification and grading rule and the historical classification and grading rule in the historical data dictionary; the target classification and grading rule is any one of the classification and grading rules in the first classification and grading rule.

[0023] If there is no conflict, the target classification and grading rules are stored in the new data dictionary;

[0024] If a conflict exists, the conflict situation is verified manually. If a conflict is found through manual verification, it is determined manually whether the new data dictionary is allowed to support the target classification and grading rules or the historical classification and grading rules.

[0025] If the new data dictionary is allowed to support the target classification and grading rules, then the historical classification and grading rules are replaced with the target classification and grading rules in the new data dictionary;

[0026] If the new data dictionary is not allowed to support the target classification and grading rule or the historical classification and grading rule, then the historical classification and grading rule in the new data dictionary is deleted, and the historical classification and grading rule and the target classification and grading rule are stored in a temporary data dictionary.

[0027] The newly obtained data dictionary is determined as the data dictionary for this output, and the temporary data dictionary is determined as the conflict data dictionary.

[0028] Optionally, updating the third classification and grading result based on the conflict data dictionary to obtain the final classification and grading result corresponding to the data to be classified and graded includes:

[0029] The classification and grading results are obtained by manually verifying the data to be classified and graded covered by the conflict data dictionary.

[0030] The third classification result is updated using the processed classification result to obtain the final classification result.

[0031] Optionally, the method further includes:

[0032] If this is the first time the data to be classified and graded is being processed, the classification and grading results are obtained by manually processing the data.

[0033] The second classification and grading rule is obtained by processing the current classification and grading result and the data to be classified and graded based on the preset classification and grading rule generation specification.

[0034] Create an empty data dictionary and store the second classification and grading rules into the empty data dictionary to obtain the data dictionary for this output.

[0035] Secondly, this application provides a classification and grading device based on knowledge accumulation, comprising:

[0036] The initial result determination module is used to obtain the data to be classified and graded, and to classify and grade the data to be classified and graded using dictionary matching based on the historical data dictionary of the previous output and manual methods, respectively. The obtained first classification and grading results are then combined with the second classification and grading results to obtain the third classification and grading result.

[0037] The classification and grading rule determination module is used to process the second classification and grading result and the corresponding data to be classified and graded based on the preset classification and grading rules to obtain the first classification and grading rule;

[0038] The data dictionary determination module is used to update the historical data dictionary based on the preset data dictionary maintenance rules and the first classification and grading rules to obtain the data dictionary output this time, and to construct a conflict data dictionary based on the data conflicts between the first classification and grading rules and the historical data dictionary.

[0039] The final result determination module is used to update the third classification and grading result based on the conflict data dictionary to obtain the final classification and grading result corresponding to the data to be classified and graded.

[0040] Thirdly, this application provides an electronic device, comprising:

[0041] Memory, used to store computer programs;

[0042] A processor is used to execute the computer program to implement the aforementioned knowledge-based data classification and grading method.

[0043] Fourthly, this application provides a computer-readable storage medium for storing a computer program that, when executed by a processor, implements the aforementioned data classification and grading method based on knowledge accumulation.

[0044] Therefore, this application can obtain data to be classified and graded, and classify and grade the data using dictionary matching based on the historical data dictionary of the previous output and manual methods respectively, and merge the corresponding first classification and grading results with the second classification and grading results to obtain the third classification and grading result; then, based on the preset classification and grading rule generation specification, process the second classification and grading result and the corresponding data to be classified and graded to obtain the first classification and grading rule; then, based on the preset data dictionary maintenance rule and the first classification and grading rule, update the historical data dictionary to obtain the data dictionary of the current output, and construct a conflict data dictionary based on the data conflicts between the first classification and grading rule and the historical data dictionary; finally, the third classification and grading result can be updated based on the conflict data dictionary to obtain the final classification and grading result corresponding to the data to be classified and graded. In this way, this application can achieve classification and grading through dictionary matching and manual methods, which improves processing efficiency, reduces processing time, and reduces labor costs. Furthermore, this application can generate new classification and grading rules based on preset classification and grading rules, thereby generating new data dictionaries and conflict data dictionaries, reducing the difficulty of building and maintaining data dictionaries, improving the accuracy of classification and grading, and generating reusable data dictionaries through manual processing. Moreover, new data dictionaries obtained from previous data classification and grading processes, i.e., historical data dictionaries, can be reused in new classification and grading projects. When the classification and grading rules in the historical data dictionary accumulate to a certain extent, they can be applied to many real-time data analysis scenarios. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 This is a flowchart of a data classification and grading method based on knowledge accumulation disclosed in this application;

[0047] Figure 2 This is a data preprocessing flowchart disclosed in this application;

[0048] Figure 3 This is a flowchart of a data classification and grading rule generation method disclosed in this application;

[0049] Figure 4 This is a flowchart of a data dictionary maintenance method disclosed in this application;

[0050] Figure 5This is a flowchart of a specific data dictionary maintenance method disclosed in this application;

[0051] Figure 6 This application discloses a specific data classification and grading method based on knowledge accumulation as a flowchart.

[0052] Figure 7 This is a flowchart of a data dictionary construction method disclosed in this application;

[0053] Figure 8 This application discloses a specific data classification and grading method based on knowledge accumulation as a flowchart.

[0054] Figure 9 This is a schematic diagram of a data classification and grading device based on knowledge accumulation disclosed in this application;

[0055] Figure 10 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Commonly used data classification and grading methods include manual methods and dictionary matching. Manual data classification and grading is inefficient, requires high manpower and time costs, and necessitates specialized knowledge from staff, making it highly dependent on human expertise and thus unstable. Dictionary matching generally has lower accuracy and precision. Incomplete or incorrectly constructed data dictionaries can easily lead to incorrect data searches, and manual maintenance of the data dictionary is also costly in terms of time and manpower. Furthermore, dictionary matching carries the risk of low data classification and grading coverage. This application proposes a method to first automatically generate data classification and grading results using dictionary matching, and then generate data classification and grading results for data not covered by the first method using manual methods. The two sets of results are then combined. Specifically, this application can first generate new data classification and grading rules based on the manually generated data classification and grading results and the corresponding data to be classified and graded using a pre-defined data classification and grading rule generation specification. Then, based on the new rules, a newly proposed data dictionary construction / maintenance method is used to build / maintain the data dictionary for reuse in new projects.

[0058] See Figure 1As shown, this embodiment of the invention discloses a data classification and grading method based on knowledge accumulation, including:

[0059] Step S11: Obtain the data to be classified and graded, and classify and grade the data to be classified and graded using dictionary matching based on the historical data dictionary output from the previous time and manual methods respectively, and merge the corresponding first classification and grading results with the second classification and grading results to obtain the third classification and grading result.

[0060] In this embodiment, the data to be classified and graded is first obtained. Then, the data can be classified and graded using dictionary matching and manual methods. Dictionary matching is based on existing data classification rules in a data dictionary. It uses pattern matching to automatically generate classification results by analyzing specific fields in the data. Common methods include regular expression pattern matching, which is highly efficient for processing the data and suitable for large datasets even in today's rapidly developing big data environment, with low labor and time costs. Manual methods, on the other hand, involve manually analyzing specific fields in the data to generate classification results. It is important to emphasize that manual methods do not generate the corresponding classification criteria / rules used in the manual analysis process; they only generate the final classification results. Manual classification can cover more data and ensure accuracy.

[0061] This embodiment may include: processing data covered by the historical data dictionary in the data to be classified and graded using the dictionary matching method to obtain the first classification and grading result; and processing data not covered by the historical data dictionary in the data to be classified and graded using the manual method to obtain the second classification and grading result. Specifically, based on the data dictionary obtained from the previous data classification and grading task, i.e., the historical data dictionary, the dictionary matching method can be used to classify and grade the data covered by the historical data dictionary in the data to be classified and graded, thus obtaining the first classification and grading result; correspondingly, for data not covered by the historical data dictionary in the data to be classified and graded, manual classification and grading can be used, thus obtaining the second classification and grading result; then, the first classification and grading result and the second classification and grading result can be merged to obtain the third classification and grading result.

[0062] Step S12: Based on the preset classification and grading rules, the second classification and grading results and the corresponding data to be classified and graded are processed to obtain the first classification and grading rules.

[0063] In this embodiment, after obtaining the second classification and grading result by manually processing the data to be classified and graded in step S11, the second classification and grading result and the corresponding data to be classified and graded can be processed based on the preset classification and grading rules to obtain the first classification and grading rules. It is understood that data classification and grading usually requires a lot of manpower to process the data. At this time, it is necessary to accurately analyze the knowledge existing in the data and transform it into reusable data classification and grading basis / rules. In this way, the results of human work can be accumulated, and it can avoid consuming a lot of manpower to process the data again.

[0064] In a specific embodiment, the first classification rule can be obtained by processing the second classification result and the corresponding data to be classified based on a preset classification and grading rule. This process may include: preprocessing the data to be classified corresponding to the second classification result to obtain preprocessed data; extracting keywords from the preprocessed data; and selecting target keywords corresponding to each individual classification result in the second classification result from all extracted keywords to obtain a corresponding keyword set; and converting the second classification result and the keyword set into an expression form of data classification and grading rules to obtain the first classification rule. Further, the data preprocessing may include: identifying the text information of the data to be classified corresponding to the second classification result, and filtering the text information based on a preset thesaurus to obtain filtered text information; and converting the target text information in the filtered text information that meets preset conversion conditions to obtain the preprocessed data. Specifically, as shown... Figure 2 As shown, the data to be classified and graded corresponding to the second classification and grading results first needs to be preprocessed to obtain preprocessed data. It is understood that the data classification and grading rule generation method is automatically implemented by a machine. Therefore, before using the keyword extraction method, the data needs to be converted from a machine-incomprehensible text form into a machine-comprehensible word data form through word segmentation, and then the word data form is used to calculate keywords. Furthermore, this application can achieve data cleaning by performing data preprocessing. Further, the preprocessing can include: identifying text information with classification and grading data, and then filtering the text information according to a preset thesaurus. The preset thesaurus can be a thesaurus, a stop word thesaurus, a useless symbol library, etc., which can achieve the effect of data cleaning; then, the information in the filtered text information that meets the preset conversion conditions can be converted accordingly, including simplified / traditional character conversion, special symbol conversion, etc., thus obtaining the preprocessed data, i.e., the word list in the figure.

[0065] Correspondingly, based on the preprocessed data, i.e., the word list, corresponding data classification and grading rules can be further generated; specifically, such as... Figure 3 As shown, the process begins by obtaining the data classification and grading results, along with the corresponding original data (i.e., the data to be classified and graded). Then, the original data corresponding to the classification and grading results is set as the corresponding corpus. It's understood that each data classification and grading result has a corresponding corpus. It should be noted that preprocessing the corpus yields a corresponding word list, i.e., the preprocessed data described above. Then, keywords in all corpora can be calculated using keyword extraction methods. In one specific embodiment, keyword extraction can utilize the TF-IDF (Term Frequency–Inverse Document Frequency) algorithm. TF-IDF is a commonly used weighting technique in information retrieval and data mining. TF stands for Term Frequency, and IDF is the Inverse Document Frequency Index. TF-IDF is a statistical method used to evaluate the importance of a word / character to a document within a document set or corpus. The importance of a word / character, i.e., its TF-IDF value, increases proportionally to the number of times it appears in the document, but decreases inversely proportionally to its frequency in the corpus. Furthermore, the keyword extraction algorithms that can be used include, but are not limited to, TF-IDF and Latent Dirichlet Allocation (LDA). It should be noted that after calculating the keywords for all corpora using keyword extraction methods, deduplication can be performed. Specifically, if a keyword corresponds to more than one data classification result, that keyword can be deleted. Then, the deduplicated keywords and the original data corresponding to the data classification results are transformed into an expression of data classification rules. For example, if the keywords in the corpus corresponding to data classification result 'a' still have N1, N2, ..., Nn after deduplication, the following data classification rules can be generated: (.*N1.*)->a, (.*N2.*)->a, ..., (.*Nn.*)->a.

[0066] Step S13: Update the historical data dictionary based on the preset data dictionary maintenance rules and the first classification and grading rules to obtain the data dictionary for this output, and construct a conflict data dictionary based on the data conflicts between the first classification and grading rules and the historical data dictionary.

[0067] In this embodiment, after obtaining the first classification and grading rule, the historical data dictionary can be updated using the first classification and grading rule according to the preset data dictionary maintenance rule to obtain the data dictionary output for this data classification and grading task; at the same time, a conflict data dictionary can be output. It is understood that the first classification and grading rule is generated based on data classification and grading results obtained manually. Here, the first classification and grading rule can be used to update the historical data dictionary, and knowledge obtained manually can be accumulated in the data dictionary. This allows the knowledge mined manually accumulated in the data dictionary to be reused in subsequent data classification and grading tasks. It should be noted that the knowledge rules accumulated in the data dictionary can not only be used in future data classification and grading tasks, but also in other scenarios, such as data leakage prevention (DLP) systems. In these application scenarios where dictionary matching can be used, the data dictionary accumulated and maintained in several data classification and grading tasks in this application can be used, thus demonstrating the applicability of this application and the practicality of the data dictionary obtained by this application.

[0068] Furthermore, if the first data classification and grading rule conflicts with the classification and grading rules in the historical data dictionary, the conflicting classification and grading rule can be stored in the conflict data dictionary for subsequent manual processing.

[0069] Step S14: Update the third classification and grading result based on the conflict data dictionary to obtain the final classification and grading result corresponding to the data to be classified and graded.

[0070] In this embodiment, after obtaining the conflict data dictionary according to step S13, the third classification and grading results can be updated manually using the conflict data dictionary. It is understood that the manual method can cover a wider range of knowledge and can reasonably process the conflict data dictionary, thus obtaining a more accurate classification and grading result by processing the third classification and grading results.

[0071] Therefore, this application can achieve data classification and grading through dictionary matching and manual methods, which improves processing efficiency, reduces processing time, and minimizes labor costs. Furthermore, this application can generate new classification and grading rules based on preset classification and grading rules, and then generate a reusable data dictionary through manual processing based on the obtained classification and grading rules, which can avoid the waste of classification and grading results obtained through manual methods.

[0072] See Figure 4As shown, this embodiment will detail the process of updating the historical data dictionary using classification and grading rules, including:

[0073] Step S21: Copy all contents of the historical data dictionary and generate a new data dictionary.

[0074] In this embodiment, the contents of the historical data dictionary can be copied to obtain the new data dictionary. It can be understood that the new data dictionary can be updated based on the contents of the historical data dictionary. This is equivalent to using the new data dictionary to inherit from the historical data dictionary and then updating the data classification and grading rules in the new data dictionary.

[0075] Step S22: Determine whether there is a conflict between the target classification and grading rule in the first classification and grading rule and the historical classification and grading rule in the historical data dictionary based on the preset rule conflict judgment method.

[0076] In this embodiment, the preset rule conflict judgment method can be used to determine whether there is a conflict between the target classification and grading rule and the historical classification and grading rule in the historical data dictionary; specifically, the preset rule conflict judgment method can be used to determine whether there is a conflict between each new data classification and grading rule in the first classification and grading rule and each historical data classification and grading rule in the historical data dictionary; it can be understood that the target classification and grading rule is any one of the classification and grading rules in the first classification and grading rule.

[0077] Furthermore, in a specific embodiment, the preset rule conflict judgment method may include: if the new data classification and grading rule is condition A -> result a, the historical data classification and grading rule is condition A -> result b, and results a and b are different. In this case, because the results a and b corresponding to the same condition A are different, a conflict exists. It should be noted that the above judgment situation applies to the historical data dictionary being the data dictionary obtained in the data classification and grading task prior to this application. It is understood that this application may also use historical data dictionaries provided externally. Accordingly, in a specific embodiment, the preset rule conflict judgment method may include: the case of "OR" relationship, the case of "AND" relationship 1, the case of "AND" relationship 2, the case of "NOT" relationship, and the case where two / more relationships exist simultaneously. Specifically, the case of "OR" relationship may be that the new data classification and grading rule is condition A -> result a, and the historical data classification and grading rule is condition A|condition B…|condition Z -> result b. At this point, the historical data classification and grading rules can be broken down into condition A -> result b, condition B -> result b, ... and condition Z -> result b. Since for the part "condition A -> result b", the results a and b corresponding to the same condition A are different, there is also a conflict. The first scenario of the "AND" relationship can be that the new data classification and grading rule is condition A -> result a, and the historical data classification and grading rule is condition A & condition B ... & condition Z -> result b. Since the AND relationship cannot be directly decomposed, there is no conflict in this case. The second scenario of the "AND" relationship can be that the new data classification and grading rule is condition A -> result a, and the historical data classification and grading rule is condition A & condition B ... & condition Z -> result a. Since different conditions A and conditions A & B ... & Z correspond to the same result a, there is a conflict in this case. The third scenario of the "NOT" relationship can be that the new data classification and grading rule is condition A -> result a, and the historical data classification and grading rule is no condition A -> result a. Since different conditions A and no A correspond to the same result a, there is a conflict. The fourth scenario of two or more relationships existing simultaneously can be if the historical data classification and grading rules contain two or more AND, OR, and NO relationships that are not covered by the above special cases. In this application, the corresponding rules can be decoupled, split into multiple data classification and grading rules, and re-stored in the data dictionary.

[0078] Step S23: If there is no conflict, the target classification and grading rules are stored in the new data dictionary.

[0079] In this embodiment, if the target classification and grading rule in the first classification and grading rule does not conflict with the historical classification and grading rule in the historical data dictionary, the target classification and grading rule can be stored in the new data dictionary. It can be understood that this allows the new data classification and grading rule to be stored in the data dictionary to preserve the data classification and grading rules obtained by manual methods, which can be reused in subsequent data classification and grading tasks, thereby avoiding the waste of the data classification and grading results obtained by manual methods.

[0080] Step S24: If a conflict exists, the conflict situation is verified manually. If a conflict is found through manual verification, it is determined manually whether the new data dictionary is allowed to support the target classification and grading rules or the historical classification and grading rules.

[0081] In this embodiment, if there is a conflict between the target classification and grading rule in the first classification and grading rule and the historical classification and grading rule in the historical data dictionary, the conflicting classification and grading rule can be verified manually. By manually judging whether the target classification and grading rule or the historical classification and grading rule is supported, it can be understood that the dictionary matching method is maintained manually, and the corresponding knowledge system is also established and maintained manually. In this way, the manual method supported by professional knowledge background can more accurately judge the corresponding data classification and grading rule.

[0082] Step S25: If the new data dictionary is allowed to support the target classification and grading rule, then the historical classification and grading rule is replaced with the target classification and grading rule in the new data dictionary.

[0083] In this embodiment, if the target classification and grading rule is supported after manual verification, the historical classification and grading rule can be replaced with the target classification and grading rule in the new data dictionary. It is understood that the knowledge system of manual methods supported by professional knowledge is also constantly updated. The historical data classification and grading rule was correct and can be retained for the previous data classification and grading task, but the entire knowledge system may change in subsequent data classification and grading tasks, and the corresponding data classification and grading rule is also maintained and updated manually. Thus, if the target classification and grading rule is determined to be supported by manual methods, the target classification and grading rule can be used to replace the corresponding historical classification and grading rule in the new data dictionary to achieve the effect of updating the data dictionary.

[0084] Step S26: If the new data dictionary is not allowed to support the target classification and grading rule or the historical classification and grading rule, then delete the historical classification and grading rule in the new data dictionary, and store the historical classification and grading rule and the target classification and grading rule in a temporary data dictionary.

[0085] In this embodiment, if any of the above rules are not supported after manual verification, the corresponding historical classification and grading rules in the new data dictionary can be deleted, and the historical classification and grading rules and the target classification and grading rules can be stored in the temporary data dictionary.

[0086] It should be noted that, in one specific embodiment, if manual judgment determines that there is no real conflict in this rule combination, the next rule combination can be ignored and processed. It is understood that manual methods can determine whether there is a real conflict between the data classification and grading rules in the first classification and grading rule and the historical classification and grading rules; for example, if manual judgment determines that two data classification and grading rules have the same effect, then this rule combination can be ignored and the next rule combination processed.

[0087] Step S27: Determine the newly obtained data dictionary as the data dictionary for this output and determine the temporary data dictionary as the conflict data dictionary.

[0088] In this embodiment, it should be noted that, as Figure 5 As shown, if no conflicting rule combinations are found after manually judging all combinations, a new data dictionary and a temporary data dictionary can be output. Accordingly, the final new data dictionary can be designated as the final data dictionary for this data classification and grading task, and can be used for other data classification and grading tasks. Correspondingly, the temporary data dictionary can be designated as the conflict data dictionary. It is understood that the conflict data dictionary stores classification and grading rules that have been processed manually and determined to be unsupported.

[0089] It is understood that the classification and grading rules in the conflict data dictionary are part of the aforementioned first classification and grading rules. After obtaining the conflict data dictionary, the third classification and grading result obtained above can be updated based on the conflict data dictionary. Specifically, this can include: manually verifying the data to be classified and graded covered by the conflict data dictionary to obtain a processed classification and grading result; and using the processed classification and grading result to update the third classification and grading result to obtain the final classification and grading result. It is understood that the data to be classified and graded covered by the conflict data dictionary needs to be manually verified to obtain a more accurate processed classification and grading result. Then, the processed classification and grading result can be used to update the third classification and grading result obtained above, thus obtaining the final classification and grading result for this data classification and grading task.

[0090] Therefore, this application can generate a new data dictionary and a conflict data dictionary based on the historical data dictionary and the newly generated first classification and grading rules through a data dictionary maintenance method. The results of the current data classification and grading can then be updated based on the conflict data dictionary. In this way, this application can better utilize data obtained manually, generate reusable knowledge from it, accumulate it, and the resulting data dictionary has stronger data coverage capabilities.

[0091] The following examples will describe the relevant steps when processing the data to be classified and graded using this application for the first time. See [link to example]. Figure 6 As shown, this embodiment of the invention discloses a data classification and grading method based on knowledge accumulation, including:

[0092] Step S31: If this is the first time the data to be classified and graded is being processed, the classification and grading results are obtained by manually processing the data.

[0093] It is understood that in this embodiment, since the historical data dictionary obtained through this application does not allow for dictionary matching, the data to be classified and graded can be processed directly by manual methods to obtain the corresponding classification and grading results.

[0094] Step S32: Based on the preset classification and grading rules, the current classification and grading results and the data to be classified and graded are processed to obtain the second classification and grading rules.

[0095] Step S33: Create an empty data dictionary and store the second classification and grading rules into the empty data dictionary to obtain the data dictionary for this output.

[0096] In this embodiment, after obtaining the second classification and grading rule, the data dictionary output this time can be obtained based on the preset data dictionary construction method and the second classification and grading rule. In a specific embodiment, such as Figure 7 As shown, after obtaining the new data classification and grading rules, namely the second classification and grading rules, a new empty data dictionary can be created. Then, all the new data classification and grading rules can be stored in this new data dictionary to obtain the final new data dictionary. It is understood that this new data dictionary can be used as a historical data dictionary in subsequent data classification and grading tasks, or it can be used for data processing tasks in other scenarios.

[0097] For a more detailed description of the process of step S32, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0098] Therefore, when processing data to be classified and graded for the first time, this application can obtain the corresponding classification and grading results manually. Then, it can generate the corresponding data classification and grading rules based on the preset classification and grading rules. The data classification and grading rules can be stored in the data dictionary through the preset data dictionary construction method for use in subsequent data classification and grading tasks. This can accurately analyze the knowledge existing in the data and transform it into reusable data classification and grading rules. Then, the data classification and grading rules can be accumulated and reused by building the data dictionary.

[0099] The following embodiments will be based on, as follows Figure 8 The flowchart shown below provides a detailed description of the technical solution of this application, including:

[0100] In this embodiment, the data to be classified and graded is first obtained, and then it is determined whether a historical data dictionary exists. It is understood that if this method is not used for the first time to perform a data classification and grading task, a corresponding historical data dictionary will be obtained in previous data classification and grading tasks. It should be noted that, in a specific embodiment, regardless of whether this method is used for the first time to perform a data classification and grading task, a data dictionary unrelated to this method can be used to perform the data classification and grading task of this method.

[0101] In this embodiment, if the historical data dictionary exists, the data to be classified and graded covered by the historical data dictionary can be processed using dictionary matching to obtain a first data classification and grading result. Then, the data to be classified and graded not covered by the historical data dictionary can be processed manually to obtain a second data classification and grading result. Next, the first data classification and grading result and the second data classification and grading result can be merged to obtain a third data classification and grading result. Furthermore, this application can use the second data classification and grading result and the corresponding data to be classified and graded to generate a first data classification and grading rule through a preset data classification and grading rule generation specification. Then, based on the preset data dictionary maintenance rule, the historical data dictionary can be updated according to the first data classification and grading rule to obtain a first data dictionary and a conflict data dictionary. Next, it can be determined whether the conflict data dictionary is empty. If it is, it means that all the first data classification and grading rules have been stored in the first data dictionary. In this case, the first data dictionary and the third data classification and grading result can be output, and the task ends. Correspondingly, if the conflict data dictionary is not empty, the data to be classified and graded covered by the conflict data dictionary needs to be processed manually to generate the fourth data classification and grading result. Then, the fourth data classification and grading result can be used to update the corresponding data in the third data classification and grading result, so as to obtain the final fifth data classification and grading result. Then, the first data dictionary and the fifth data classification and grading result are output, and the task ends.

[0102] In this embodiment, if it is determined that the historical data dictionary does not exist, the classified and graded data can be processed manually to generate a sixth data classification and grading result. Then, the sixth data classification and grading result can be processed based on the preset data classification and grading rules to obtain a second data classification and grading rule. Furthermore, the second data classification and grading rule can be processed based on the preset data dictionary construction rules to obtain a second data dictionary. Finally, the second data dictionary and the sixth data classification and grading result are output.

[0103] Therefore, it is evident that the dictionary matching method used in this application for data classification and grading is applicable to scenarios with large data volumes, requiring very low manpower and time costs. Furthermore, it can process data classification and grading rules obtained manually to build or maintain a new data dictionary. This preserves the results of manual data processing while reducing the difficulty of building and maintaining the data dictionary, and lowering the time and manpower costs. The results of manual methods are also transformed into a data dictionary, covering more data and providing good noise resistance. Moreover, this application allows the data dictionary to be reused in subsequent data classification and grading tasks, reducing manpower consumption. After accumulating sufficient knowledge in the data dictionary, it can also be used in other relevant real-time analysis scenarios.

[0104] like Figure 9 As shown, this application discloses a classification and grading device based on knowledge accumulation, comprising:

[0105] The initial result determination module 11 is used to obtain the data to be classified and graded, and to classify and grade the data to be classified and graded using dictionary matching based on the historical data dictionary of the previous output and manual methods, respectively, and to merge the corresponding first classification and grading results with the second classification and grading results to obtain the third classification and grading result.

[0106] The classification and grading rule determination module 12 is used to process the second classification and grading result and the corresponding data to be classified and graded based on the preset classification and grading rule generation specification to obtain the first classification and grading rule;

[0107] The data dictionary determination module 13 is used to update the historical data dictionary based on the preset data dictionary maintenance rules and the first classification and grading rules to obtain the data dictionary output this time, and to construct a conflict data dictionary based on the data conflicts between the first classification and grading rules and the historical data dictionary.

[0108] The final result determination module 14 is used to update the third classification and grading result based on the conflict data dictionary to obtain the final classification and grading result corresponding to the data to be classified and graded.

[0109] Therefore, this application can obtain the corresponding classification and grading results by processing the data to be classified and graded using dictionary matching and manual methods. Then, it can generate the corresponding first classification and grading rules for the manual method based on the preset classification and grading rules. Then, it can update the historical data dictionary based on the first classification and grading rules. This can retain the results of the manual method in processing the data to be classified and graded, so that the data dictionary can be reused in subsequent data classification and grading tasks, and the data coverage capability of the data dictionary can be improved.

[0110] In one specific embodiment, the initial result determination module 11 may include:

[0111] The first result determination unit is used to process the data covered by the historical data dictionary in the data to be classified and graded by the dictionary matching method to obtain the first classification and grading result;

[0112] The second result determination unit is used to process the data in the data to be classified and graded that is not covered by the historical data dictionary through the manual method to obtain the second classification and grading result.

[0113] In one specific embodiment, the classification and grading rule determination module 12 may include:

[0114] The data preprocessing submodule is used to preprocess the data to be classified and graded corresponding to the second classification and grading result to obtain preprocessed data.

[0115] The keyword determination unit is used to extract keywords from the preprocessed data and filter out target keywords corresponding to each single classification and grading result in the second classification and grading result from all the extracted keywords to obtain the corresponding keyword set.

[0116] The first rule-determining unit is used to transform the second classification and grading result and the keyword set into an expression of data classification and grading rules to obtain the first classification and grading rules.

[0117] Accordingly, in one specific embodiment, the data preprocessing submodule may include:

[0118] The text filtering unit is used to identify the text information of the data to be classified and graded corresponding to the second classification and grading result, and to filter the text information based on a preset thesaurus to obtain the filtered text information.

[0119] The text conversion unit is used to convert the target text information that meets the preset conversion conditions in the filtered text information to obtain the preprocessed data.

[0120] In one specific embodiment, the data dictionary determination module 13 may include:

[0121] The new dictionary generation unit is used to copy all the contents of the historical data dictionary and generate a new data dictionary;

[0122] The conflict determination unit is used to determine whether there is a conflict between the target classification and grading rule in the first classification and grading rule and the historical classification and grading rule in the historical data dictionary, based on a preset rule conflict determination method; the target classification and grading rule is any one of the classification and grading rules in the first classification and grading rule.

[0123] The first rule storage unit is used to store the target classification and grading rules into the new data dictionary when there is no conflict.

[0124] The conflict determination unit is used to manually verify the conflict situation when a conflict exists. If a conflict is verified by manual verification, the unit will manually determine whether the new data dictionary is allowed to support the target classification and grading rule or the historical classification and grading rule.

[0125] A rule replacement unit is used to replace the historical classification and grading rule with the target classification and grading rule in the new data dictionary when the new data dictionary is allowed to support the target classification and grading rule.

[0126] The second rule storage unit is used to delete the historical classification and grading rules in the new data dictionary and store the historical classification and grading rules and the target classification and grading rules in a temporary data dictionary when the new data dictionary is not allowed to support the target classification and grading rules or the historical classification and grading rules.

[0127] The first dictionary determination unit is used to determine the newly obtained data dictionary as the data dictionary for this output and to determine the temporary data dictionary as the conflict data dictionary.

[0128] In one specific embodiment, the final result determination module 14 may include:

[0129] The conflict dictionary processing unit is used to manually verify the data to be classified and graded covered by the conflict data dictionary to obtain the processed classification and grading results.

[0130] The final result determination unit is used to update the third classification and grading result using the processed classification and grading result to obtain the final classification and grading result.

[0131] Furthermore, in one specific embodiment, the method may further include:

[0132] The manual processing unit is used to process the data to be classified and graded manually to obtain the classification and grading result when it is the first time the data to be classified and graded is processed.

[0133] The second rule determination unit is used to process the current classification and grading results and the data to be classified and graded based on the preset classification and grading rules to obtain the second classification and grading rules;

[0134] The second dictionary determination unit is used to create an empty data dictionary and store the second classification and grading rules into the empty data dictionary to obtain the data dictionary for this output.

[0135] Furthermore, embodiments of this application also disclose an electronic device, Figure 10 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0136] Figure 10 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the knowledge accumulation-based data classification and grading method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0137] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0138] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0139] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including computer programs capable of performing the knowledge-accumulation-based data classification and grading method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0140] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned data classification and grading method based on knowledge accumulation. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0141] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0142] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0143] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0144] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0145] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A data classification and grading method based on knowledge accumulation, characterized in that, include: The data to be classified and graded is obtained, and the data is classified and graded using dictionary matching based on the historical data dictionary of the previous output and manual methods, respectively. The first classification and grading results are then merged with the second classification and grading results to obtain the third classification and grading result. The first classification and grading rule is obtained by processing the second classification and grading result and the corresponding data to be classified and graded based on the preset classification and grading rules. The historical data dictionary is updated based on the preset data dictionary maintenance rules and the first classification and grading rules to obtain the data dictionary for this output, and a conflict data dictionary is constructed based on the data conflicts between the first classification and grading rules and the historical data dictionary. The third classification and grading results are updated based on the conflict data dictionary to obtain the final classification and grading results corresponding to the data to be classified and graded. The process of classifying and grading the data to be classified and graded using dictionary matching based on the historical data dictionary of the previous output and manual methods includes: The first classification and grading result is obtained by processing the data covered by the historical data dictionary in the data to be classified and graded using the dictionary matching method. The second classification result is obtained by processing the data in the data to be classified and graded that is not covered by the historical data dictionary using the manual method.

2. The data classification and grading method based on knowledge accumulation according to claim 1, characterized in that, The first classification and grading rule is obtained by processing the second classification and grading result and the corresponding data to be classified and graded based on the preset classification and grading rules, including: Preprocessing is performed on the data to be classified and graded corresponding to the second classification and grading result to obtain preprocessed data; Extract keywords from the preprocessed data, and select target keywords from all extracted keywords that correspond to each single classification and grading result in the second classification and grading result to obtain the corresponding keyword set; The second classification and grading results and the keyword set are transformed into an expression of data classification and grading rules to obtain the first classification and grading rules.

3. The data classification and grading method based on knowledge accumulation according to claim 2, characterized in that, The preprocessing of the data to be classified and graded corresponding to the second classification and grading result to obtain preprocessed data includes: Identify the text information of the data to be classified and graded corresponding to the second classification and grading result, and filter the text information based on a preset thesaurus to obtain filtered text information; The target text information that meets the preset conversion conditions in the filtered text information is converted to obtain the preprocessed data.

4. The data classification and grading method based on knowledge accumulation according to claim 1, characterized in that, The process of updating the historical data dictionary based on preset data dictionary maintenance rules and the first classification and grading rules to obtain the current output data dictionary, and constructing a conflict data dictionary based on data conflicts between the first classification and grading rules and the historical data dictionary, includes: Copy all contents of the historical data dictionary and generate a new data dictionary; Based on a preset rule conflict judgment method, it is determined whether there is a conflict between the target classification and grading rule in the first classification and grading rule and the historical classification and grading rule in the historical data dictionary; the target classification and grading rule is any one of the classification and grading rules in the first classification and grading rule. If there is no conflict, the target classification and grading rules are stored in the new data dictionary; If a conflict exists, the conflict situation is verified manually. If a conflict is found through manual verification, it is determined manually whether the new data dictionary is allowed to support the target classification and grading rules or the historical classification and grading rules. If the new data dictionary is allowed to support the target classification and grading rules, then the historical classification and grading rules are replaced with the target classification and grading rules in the new data dictionary; If the new data dictionary is not allowed to support the target classification and grading rule or the historical classification and grading rule, then the historical classification and grading rule in the new data dictionary is deleted, and the historical classification and grading rule and the target classification and grading rule are stored in a temporary data dictionary. The newly obtained data dictionary is determined as the data dictionary for this output, and the temporary data dictionary is determined as the conflict data dictionary.

5. The data classification and grading method based on knowledge accumulation according to claim 1, characterized in that, The step of updating the third classification and grading result based on the conflict data dictionary to obtain the final classification and grading result corresponding to the data to be classified and graded includes: The classification and grading results are obtained by manually verifying the data to be classified and graded covered by the conflict data dictionary. The third classification result is updated using the processed classification result to obtain the final classification result.

6. The data classification and grading method based on knowledge accumulation according to any one of claims 1 to 5, characterized in that, Also includes: If this is the first time the data to be classified and graded is being processed, the classification and grading results are obtained by manually processing the data. The second classification and grading rule is obtained by processing the current classification and grading result and the data to be classified and graded based on the preset classification and grading rule generation specification. Create an empty data dictionary and store the second classification and grading rules into the empty data dictionary to obtain the data dictionary for this output.

7. A classification and grading device based on knowledge accumulation, characterized in that, include: The initial result determination module is used to obtain the data to be classified and graded, and to classify and grade the data to be classified and graded using dictionary matching based on the historical data dictionary of the previous output and manual methods, respectively. The obtained first classification and grading results are then combined with the second classification and grading results to obtain the third classification and grading result. The classification and grading rule determination module is used to process the second classification and grading result and the corresponding data to be classified and graded based on the preset classification and grading rules to obtain the first classification and grading rule; The data dictionary determination module is used to update the historical data dictionary based on the preset data dictionary maintenance rules and the first classification and grading rules to obtain the data dictionary output this time, and to construct a conflict data dictionary based on the data conflicts between the first classification and grading rules and the historical data dictionary. The final result determination module is used to update the third classification and grading result based on the conflict data dictionary to obtain the final classification and grading result corresponding to the data to be classified and graded; The initial result determination module includes: The first result determination unit is used to process the data covered by the historical data dictionary in the data to be classified and graded by the dictionary matching method to obtain the first classification and grading result; The second result determination unit is used to process the data in the data to be classified and graded that is not covered by the historical data dictionary through the manual method to obtain the second classification and grading result.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the knowledge accumulation-based data classification and grading method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Used to store computer programs, which, when executed by a processor, implement the data classification and grading method based on knowledge accumulation as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Industry characteristics analyzer with artificial behavior learning capability

    CN105512191A

  • Data standardization method and device and electronic equipment

    CN111078639A