A method for constructing domain knowledge base of data classification and grading based on information extraction

By applying information extraction technology in the field of data security, data classification and grading information can be automatically extracted from policies and regulations, and a knowledge base is built, the problem of low information extraction efficiency in the existing technology has been solved, the work efficiency of practitioners has been improved, and a comprehensive data classification and grading ontology has been formed.

CN115292450BActive Publication Date: 2025-05-02SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210896400.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-05-02
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

The prior art is difficult to extract useful data classification and hierarchical information from a large number of unstructured policies and regulations in an automated manner, resulting in data security practitioners spending a lot of time and effort on learning and querying.

Method used

Using information extraction methods, machine learning and natural language processing technology are used to extract useful information from unstructured texts and tables to build a knowledge base for data classification and grading. Specific steps include document acquisition, preprocessing, text and tabular data extraction, tuple extraction, and knowledge base construction.

Benefits of technology

It has achieved rapid and automated extraction of data classification and grading information from policies and regulations, improved the work efficiency of data security practitioners, saved query time, and formed an information type dictionary and data classification and grading ontology covering multiple industries and fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115292450B_ABST
    Figure CN115292450B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a data classification and grading domain knowledge base based on information extraction, and relates to the field of natural language processing technology. The present invention includes a document acquisition step, a document preprocessing step, a text data extraction step, a table information extraction step, a data classification and grading tuple extraction step, and a data classification and grading domain knowledge base construction step. The present invention proposes an automatic policy and regulation parsing framework, and constructs the classification and grading information into a domain knowledge base, giving full play to the guiding role of policies and regulations on data classification and grading, which can effectively bridge the gap between national-level data protection concerns and specific countermeasures of organizations, and the framework can carry out more research in the future.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and more specifically to a method for constructing a classified and graded domain knowledge base based on information extraction. Background Art

[0002] The rapid development of big data technology has brought together a massive amount of data assets in organizations, which are stored in various places in structured and unstructured forms. The huge and scattered amount of data makes it difficult for organizations to effectively implement data security management, resulting in frequent data security incidents.

[0003] Data classification and grading serves as the basic support and prerequisite for solving this problem. Classifying data according to different attributes and characteristics and distinguishing protection measures at different levels have become the focus of attention of national, industrial and local governments in recent years. Laws, regulations and policy standards related to data security and classification and grading have been successively introduced (for the sake of simplicity, this invention refers to laws, regulations, policies and standards collectively as "policies and regulations").

[0004] These policies and regulations contain a lot of valuable information that can guide data security practitioners in the implementation of data classification and grading, such as data classification dimensions, recommended data security levels, etc. However, this information usually exists in an unstructured form, which means that data security personnel often need to spend a lot of time and energy learning the valuable experience provided by the country and other industries. Therefore, automatically extracting useful information from a large number of policies and regulations and presenting it in a structured way has become one of the effective ways to improve the level of automation of classification and grading. Summary of the invention

[0005] In order to overcome the defects and shortcomings existing in the above-mentioned prior art, the present invention provides a method for constructing a knowledge base in the field of data classification and grading based on information extraction. The purpose of the present invention is to design a classification and grading information extraction framework, and to extract useful information from unstructured texts and tables using techniques such as machine learning and natural language processing. Data classification practices, such as information types and security levels of information types, can be automatically extracted from policies and regulations related to data classification and grading, thereby completing the construction of a knowledge base in the field of data classification and grading. The knowledge base constructed by the present invention can summarize useful information such as extracted data categories and security levels, so that data security practitioners can quickly find the desired data, improve the work efficiency of data security practitioners, and save data query time.

[0006] In order to solve the above problems existing in the prior art, the present invention is implemented through the following technical solutions.

[0007] The present invention provides a method for constructing a data classification and grading domain knowledge base based on information extraction, the method comprising the following steps:

[0008] S1, the document acquisition step, finds the target document in the target website or target database by keyword search, and aggregates it to form a corpus;

[0009] S2, document preprocessing step, separating the target document obtained in step S1 into two categories: plain text and table;

[0010] S3, text data extraction step, constructing a semantically embedded naive Bayes classifier, classifying the plain text separated in step S2 through the constructed naive Bayes classifier, and generating data classification and grading sentence labels;

[0011] S4, table information extraction step, according to the table features and interesting information in the table in the corpus separated in step S2, split the merged cells, fill the empty cells according to the cell text before splitting, and then extract information based on pattern matching;

[0012] S5, data classification and grading tuple extraction step, using a combination of pattern matching and natural language processing technology, based on the identified classification and grading sentence labels, extracting classification and grading elements to achieve joint extraction of information types and their relationships;

[0013] S6, data classification and grading domain knowledge base construction steps, after plain text and table extraction, data classification and grading tuples are obtained, and after semantic similarity calculation formula, duplicate removal and fusion are performed according to high and low levels to complete the construction of classification and grading domain knowledge base.

[0014] Furthermore, the step S2 specifically includes the following sub-steps:

[0015] S201, deleting irrelevant content in the target document;

[0016] S202, format conversion, if the target document is in PDF format, convert the PDF format files into word format files, and then use the python-docx library to separate text and tables;

[0017] S203, using the language processing platform LTP to segment the plain text into sentences; wherein the hierarchical relationship of the table in the text is saved as a tuple, that is, the line number of the text and its direct parent node, to ensure that the sentence has a complete semantics and a simple structure;

[0018] S204, using a Chinese word segmentation tool to segment all sentences in the text to form a list of segmentation phrases; then applying the stop word list cn_stopwords to remove function words, auxiliary words and general words.

[0019] Furthermore, the S3 step specifically includes the following sub-steps:

[0020] S301, calling the Python Sklearn library interface to generate a TF-IDF vector for each word in the word segmentation phrase list; at the same time, the chi-square statistic is used to further filter out some redundant features; the output feature engineering vector is represented as fe(s i ), s i Indicates a sentence;

[0021] S302, use weighted Word2vec to generate a corpus-specific word embedding vector, assign a weight to each word embedding vector, and introduce additional semantic features. The weighted word embedding vector is Among them, w2v() means using the Word2vec word embedding method, t j represents a word in a sentence, M represents the total number of words in the sentence, w(t j ) is represented by the word t j The term frequency-inverse document frequency (TF-IDF) weight of

[0022] S303: concatenate the feature engineering vector output from step S301 and step S302 with the weighted word embedding vector vec(s i ) represents the concatenated vector; semantic information is introduced, and the concatenated vector is input into the naive Bayes classifier to complete the recognition of data classification and grading sentences.

[0023] Furthermore, given a sentence set D = {(s1, y1), (s2, y2), ..., (s N ,y N )},s i ={t1, t2, ..., t M},y i ∈{0, 1} represents the category label; the goal of the naive Bayes classifier is to i Find a correct label y i , its formal expression is as follows:

[0024]

[0025] For a given sentence s i , the naive Bayes classifier calculates the posterior probability p(y i |s i ), the class with the maximum probability value is called the maximum a posteriori estimate, which is expressed as:

[0026]

[0027] Furthermore, the S4 step specifically includes the following sub-steps:

[0028] S401, table structure analysis; the table structure analysis mainly includes processing merged cells, complex header processing, filtering irregular tables and cross-page table determination;

[0029] S402, table field extraction; first, formulate roles and matching rules for the fields to be extracted, specifically, by observing the classification and grading table in the target document, formulate a rule set according to the expression of the fields to be extracted in the table header and table content; then, traverse each column in order, match the content of each cell with the rule set, if a column is matched successfully, record the position index and role of the matched column in the table; finally, traverse the index column and extract the cell, which is the required field information.

[0030] Furthermore, the step S5 specifically includes the following sub-steps:

[0031] S501, semantic dependency analysis: use the semantic dependency analysis tool provided by LTP to parse the classified and graded sentences;

[0032] S502, simplifying the semantic dependency tree: after parsing the classified and graded sentences in step S501, the original semantic dependency tree in the classified and graded sentences is obtained, and the original semantic dependency tree is simplified;

[0033] S503, Tregex pattern generation: by observing the simplified semantic dependency tree, find the shortest path covering the classification and grading tuples, and express it as a Tregex pattern.

[0034] Furthermore, in the step S501, the semantic dependency analysis tool provided by LTP is used to parse the classified and graded sentences, the semantic dependency relationship and part of speech in the classified and graded sentences are parsed, and the semantic dependency relationship tags and part of speech tags are marked in the classified and graded sentences.

[0035] The semantic dependency relationship includes the subject EXP, the guest CONT, the point mark mPUNC, the dependency mark mDEPD, the parallel eCOO, the modification FEAT and the link LINK.

[0036] Furthermore, in step S502, the method of simplifying the original semantic dependency tree includes adding, merging and deleting, wherein adding refers to adding a node with a part-of-speech tag u, c or mPUNC relationship tag to the previous node; merging refers to merging nodes marked as FEAT or rFEAT or dFEAT or MEAS or eCOO, or object nodes whose parent nodes are marked as n; and deleting refers to deleting nodes marked as FEAT or mRELA or mDEPD or situational roles.

[0037] Furthermore, the classification and grading domain knowledge base constructed in step S6 consists of two parts: domain = <L domain , O domain >, one is the information type dictionary L domain ={c1, c2, ..., c n |c i ∈info_type, i∈N +},N + represents a set of positive integers; the other is a data classification and grading ontology consisting of classification and grading tuples O domain =<C,A,R> ; Among them, info_type indicates the information type, c i Represents the concepts in the ontology, C = c1, c2, ..., c n , A represents the security level attribute, and R represents the relationship between concepts.

[0038] Compared with the prior art, the beneficial technical effects brought by the present invention are as follows:

[0039] 1. The present invention can take data classification and grading of concern at the national and industry levels as the research object, collect promulgated policies and regulations as a corpus, use machine learning, natural language processing and other technologies to extract classification and grading-related knowledge from policies and regulations, build an information type dictionary and a comprehensive data classification and grading ontology, and display important expert experience hidden in unstructured files in a structured manner, helping data practitioners quickly understand existing classification and grading practices.

[0040] 2. The information type dictionary that can be formed by the present invention covers multiple industries and fields, which can help data practitioners formulate corresponding identification rules according to the sensitive personal information types in the dictionary, so as to quickly locate sensitive personal information and help relevant organizations achieve compliance.

[0041] 3. The data classification and grading ontology formed by the present invention can be used for data security level recommendation. With the powerful reasoning function of the ontology, data security level can be recommended to a certain extent. For example, "communication information" and its child nodes: "email address" and "telephone number", etc., if the security level of "communication information" is given, it can be inferred that the security levels of the two child nodes are not lower than it. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A schematic diagram of the architecture for constructing a domain knowledge base for data classification and grading in the present invention.

[0043] Figure 2 It is a schematic diagram of the preprocessing of the nested list in the present invention.

[0044] Figure 3Extract a flow chart for the table.

[0045] Figure 4 Schematic diagram of the semantic dependency tree before and after pruning. DETAILED DESCRIPTION

[0046] The technical solution of the present invention is further described in detail below in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] Example 1

[0048] As a preferred embodiment of the present invention, as shown in the attached Figure 1 As shown, a schematic diagram of the data classification and grading domain knowledge base construction architecture in the present invention is given. This embodiment discloses a data classification and grading domain knowledge base construction method based on information extraction, and the method includes the following steps:

[0049] S1. The step of acquiring documents is to find the target documents in the target website or target database by keyword search and summarize them into a corpus. In this embodiment, the target documents can be retrieved in the target website or database by keyword search through a crawler tool.

[0050] As an example, using "data", "security", and "classification and grading" as keywords to search from standard libraries, official websites, etc., we eventually formed a corpus containing 28 PDF documents and 10 Word documents.

[0051] S2, document preprocessing step. Given that both text and tables contain valuable information, the target document obtained in step S1 is first separated into two categories: plain text and tables, so as to facilitate subsequent targeted extraction.

[0052] S3, text data extraction step, data classification and grading sentences are often obscured by lengthy regulations. In this embodiment, a semantically embedded naive Bayes classifier (SE-NBC) is constructed. Given a sufficient set of training sentences, SE-NBC can automatically generate data classification and grading sentence labels for the sentences in the test set; then the plain text separated in step S2 is classified by the constructed naive Bayes classifier to generate data classification and grading sentences.

[0053] S4, table information extraction step. Tables usually involve complex situations, such as row (column) merging, page spanning, table titles and table contents appearing alternately, etc. Direct parsing will result in failure to extract content or inaccurate extraction. Therefore, according to the table features and information of interest in the table in the corpus separated in step S2, the merged cells are split, the empty cells are supplemented according to the cell text before splitting, and then the information is extracted based on pattern matching;

[0054] S5, data classification and grading tuple extraction step, using a combination of pattern matching and natural language processing technology, based on the identified classification and grading sentence labels, extracting classification and grading elements to achieve joint extraction of information types and their relationships;

[0055] S6, data classification and grading domain knowledge base construction steps, after plain text and table extraction, the data classification and grading tuples are obtained, and after the semantic similarity calculation formula, the data classification and grading tuples are removed and integrated according to the high and low levels to complete the construction of the classification and grading domain knowledge base.

[0056] Example 2

[0057] As another preferred embodiment of the present invention, refer to the attached specification Figure 1-4 As shown, this embodiment discloses a method for constructing a data classification and grading domain knowledge base based on information extraction, and the method includes the following steps:

[0058] S1. Document acquisition step: find the target document in the target website or target database by keyword search and summarize it to form a corpus.

[0059] S2, document preprocessing step, separates the target document obtained in step S1 into two categories: plain text and table; the details are as follows:

[0060] S201, deleting irrelevant content in the target document;

[0061] S202, format conversion, if the target document is in PDF format, convert the PDF format files into word format files, and then use the python-docx library to separate text and tables;

[0062] S203, using the language processing platform LTP to segment the plain text into sentences; wherein the hierarchical relationship of the table in the text is saved as a tuple, that is, the line number of the text and its direct parent node, to ensure that the sentence has a complete semantics and a simple structure;

[0063] S204, using a Chinese word segmentation tool to segment all sentences in the text to form a list of segmentation phrases; then applying the stop word list cn_stopwords to remove function words, auxiliary words and general words.

[0064] As an example, it is observed that policies and regulations usually have a fixed writing structure. For example, the catalog, scope, references, etc. contained in the standards and the annexes contained in the regulations are irrelevant to data classification and grading. By deleting these irrelevant chapters, the efficiency of subsequent processing can be improved.

[0065] S3, text data extraction step, construct a semantically embedded naive Bayes classifier, classify the plain text separated in step S2 through the constructed naive Bayes classifier, and generate data classification and grading sentence labels; the details are as follows:

[0066] S301, calling the Python Sklearn library interface to generate a TF-IDF vector for each word in the word segmentation phrase list; at the same time, the chi-square statistic is used to further filter out some redundant features; the output feature engineering vector is fe(s i ), s i Indicates a sentence;

[0067] S302, use weighted Word2vec to generate word embedding vectors that are specific to the corpus, assign weights to each word embedding vector, and introduce additional semantic features. The weighted word embedding vector is Among them, w2v() means using the Word2vec word embedding method, M represents the total number of words contained in the sentence, and t j represents a word in a sentence, w(t j ) is represented by the word t j The term frequency-inverse document frequency (TF-IDF) weight of

[0068] S303: concatenate the feature engineering vector output from step S301 and step S302 with the weighted word embedding vector vec(s i ) represents the concatenated vector; semantic information is introduced, and the concatenated vector is input into the naive Bayes classifier to complete the recognition of data classification and grading sentences.

[0069] Furthermore, given a sentence set D = {(s1, y1), (s2, y2), ..., (s N ,y N )},s i ={t1, t2, ..., t M},y i ∈{0, 1} represents the category label; the goal of the naive Bayes classifier is to i Find a correct label y i , its formal expression is as follows:

[0070]

[0071] For a given sentence s i , the naive Bayes classifier calculates the posterior probability p(y i |s i ), the class with the maximum probability value is called the maximum a posteriori estimate, which is expressed as:

[0072]

[0073] S4, table information extraction step, according to the table features and interesting information in the table in the corpus separated in step S2, the merged cells are split, the empty cells are supplemented according to the cell text before splitting, and then the information is extracted based on pattern matching; the details are as follows:

[0074] S401, table structure analysis; the table structure analysis mainly includes processing merged cells, complex header processing, filtering irregular tables and cross-page table determination;

[0075] S402, table field extraction; first, formulate roles and matching rules for the fields to be extracted, specifically, by observing the classification and grading table in the target document, formulate a rule set according to the expression of the fields to be extracted in the table header and table content; then, traverse each column in order, match the content of each cell with the rule set, if a column is matched successfully, record the position index and role of the matched column in the table; finally, traverse the index column and extract the cell, which is the required field information.

[0076] S5, data classification and grading tuple extraction step, using the combination of pattern matching and natural language processing technology, based on the identified classification and grading sentence labels, extract classification and grading elements to achieve joint extraction of information types and their relationships; the details are as follows:

[0077] S501, semantic dependency analysis: use the semantic dependency analysis tool provided by LTP to parse the classified and graded sentences;

[0078] S502, simplifying the semantic dependency tree: after parsing the classified and graded sentences in step S501, the original semantic dependency tree in the classified and graded sentences is obtained, and the original semantic dependency tree is simplified;

[0079] S503, Tregex pattern generation: by observing the simplified semantic dependency tree, find the shortest path covering the classification and grading tuples, and express it as a Tregex pattern.

[0080] As an example, the semantic dependency analysis tool provided by LTP is used to parse the classified and graded sentences, the semantic dependency relationship and part of speech in the classified and graded sentences are parsed, and the semantic dependency relationship tags and part of speech tags are marked in the classified and graded sentences.

[0081] As an example, the semantic dependency relationship includes the party EXP, the guest CONT, the point mark mPUNC, the dependency mark mDEPD, the parallel eCOO, the modification FEAT and the link LINK.

[0082] As an example, in step S502, the methods of simplifying the original semantic dependency tree include adding, merging and deleting. Adding refers to adding a node with a part-of-speech tag u, c or mPUNC relationship tag to the previous node; merging refers to merging nodes marked as FEAT or rFEAT or dFEAT or MEAS or eCOO, or object nodes whose parent nodes are marked as n; and deleting refers to deleting a node marked as FEAT or mRELA or mDEPD or situational role.

[0083] Example 3

[0084] As another preferred embodiment of the present invention, refer to the attached specification Figures 1 to 4 As shown, attached Figure 1 A schematic diagram of the data classification and grading domain knowledge base construction architecture in the present invention is given, including the following steps:

[0085] Step 1: Document acquisition steps

[0086] The target documents are found in the target website or target database by keyword search and summarized into a corpus. In this embodiment, the target documents can be retrieved in the target website or database by keyword search through a crawler tool. As an example, the keywords "data", "security" and "classification and grading" are used to search from standard libraries, official websites, etc., and finally a corpus containing 28 PDF documents and 10 word documents is formed.

[0087] Step 2: Policy and Regulation Preprocessing

[0088] (1) Delete irrelevant content

[0089] Through observation, it is found that policies and regulations usually have a fixed text structure. For example, the catalog, scope, references, etc. contained in the standards and the annexes contained in the regulations are irrelevant to data classification and grading. By deleting these irrelevant chapters, the efficiency of subsequent processing can be improved.

[0090] (2) PDF format conversion

[0091] Most policies and regulations are given in PDF format, and a few are in Word format. For ease of processing, all these policy documents are converted to Word format, and the python-docx library is used to separate text and tables for subsequent targeted information extraction;

[0092] (3) List processing

[0093] The plain text is segmented into sentences using the language processing platform LTP of Harbin Institute of Technology. Among them, the lists in the text (ordered or unordered) need special processing for two reasons: 1) LTP will mistake the items in the list as separate sentences, resulting in semantically incomplete sentences. 2) If we manually integrate the list into a sentence, multiple subjects, predicates and objects will appear, resulting in a very complex sentence structure. Both of these situations will hinder the semantic dependency analysis of the sentence. Figure 2 The following is an example of the most complex structure of nested lists and their processing results. A list is considered to consist of a list title and several list items. The list title description is usually represented by some indicative words ("below", "as follows" (as shown below)) and symbols (:), etc. List items usually start with bullets, Arabic numerals and English letters. Based on this, a list preprocessing scheme is designed to save the hierarchical relationship of the list as a pair of tuples, namely the line number of the text and its direct parent node, to ensure the integrity of the sentence semantics while having a simple structure;

[0094] (4) Word segmentation and stop word removal

[0095] Jieba, a Chinese word segmentation tool, is used to segment all sentences and form a list of segmentation phrases. Then, the stop word list cn_stopwords is used to remove function words, auxiliary words, general words, and other words that are not helpful for further analysis.

[0096] Step 3: Classification and grading sentence recognition;

[0097] Data classification and grading sentence recognition includes feature engineering, semantic embedding, and naive Bayes classifier;

[0098] (1) Feature Engineering

[0099] Call the Python Sklearn library interface to generate TF-IDF vectors for each word in the word segmentation phrase list. At the same time, the chi-square statistic is used to further filter out redundant features, and the output vector is represented as fe(s_i), where s_i represents a sentence;

[0100] (2) Semantic Embedding

[0101] Considering that the vocabulary involved in data classification and grading is highly domain-specific, Word2vec is used to generate corpus-specific word embedding vectors, expressed as Among them, w2v() means using the Word2vec word embedding method, t jRepresents a word in a sentence. In the word embedding model, the word itself lacks the ability to distinguish categories. In order to further distinguish the difference in the contribution of different words to category distinction, each word vector is given a weight, which not only brings additional semantic features but also distinguishes the degree of contribution of the word. The weighted word embedding vector is Among them, w(t j ) is represented by the word t j The term frequency-inverse document frequency (TF-IDF) weight of M represents the total number of words contained in the sentence.

[0102] (3) Naive Bayes Classifier

[0103] Given a sentence set D = {(s1, y1), (s2, y2), ..., (s N ,y N )},s i ={t1, t2, ..., t M},y i ∈{0, 1} represents the category label; the goal of the naive Bayes classifier is to i Find a correct label y i , its formal expression is as follows:

[0104]

[0105] For a given sentence s i , the naive Bayes classifier calculates the posterior probability p(y i |s i ), the class with the maximum probability value is called the maximum a posteriori estimate, which is expressed as:

[0106]

[0107] Concatenate the feature engineering vectors from step (1) and step (2) with the weighted word embedding vector Semantic information is introduced and the concatenated vector is input into the naive Bayes classifier to complete the recognition of data classification and grading sentences.

[0108] Step 4: Extract table information;

[0109] Attached Figure 3 A schematic diagram of the table information extraction process in the present invention is given. The specific extraction process is as follows:

[0110] 1. Table structure analysis

[0111] like Figure 2 As shown, table structure parsing mainly includes processing merged cells, complex header processing, filtering irregular tables, and cross-page table determination.

[0112] 2. Table field extraction

[0113] a) Create roles and matching rules for the fields to be extracted. By observing the table in the policy file, create a rule set based on the description of the fields to be extracted in the table header and table content;

[0114] b) Assign roles by column. Traverse each column in order and match the content of each cell with the rule set. If a column is matched successfully, record the position index and role of the matched column in the table;

[0115] c) Form a tree structure. Traverse the index column and extract the cells. The cell text is the required field information.

[0116] Step 5: Classification and grading tuple extraction;

[0117] Classification and grading tuple extraction includes sentence semantic dependency analysis, semantic dependency tree simplification, and Tregex pattern generation;

[0118] 1. Semantic Dependency Analysis

[0119] Use LTP semantic dependency analysis tools to parse classified and graded sentences. The advantage of semantic dependency analysis is that it can obtain deep semantic information, not just the structural information of the sentence. Commonly used semantic dependency relations include: "Party EXP", "Guest CONT", "Punctuation mark mPUNC", "Dependency mark mDEPD", "Parallel relationship eCOO", "Modification FEAT", "Link" and so on. In addition, in order to further facilitate the extraction of information, in addition to the relationship tag, each word is also annotated with a part-of-speech tag.

[0120] 2. Simplification of semantic dependency tree

[0121] like Figure 4 As shown in the figure, the structure of the originally generated semantic dependency tree is relatively complex. Based on the observation, the following pruning rules are formulated to simplify the original dependency tree:

[0122] a) Add: add a node with part-of-speech tag u, c or mPUNC relation tag to the previous node;

[0123] b) Merge: merge nodes marked as FEAT or rFEAT or dFEAT or MEAS or eCOO, or object nodes whose parent node is marked as n;

[0124] c) Delete: Delete the nodes marked as FEAT or mRELA or mDEPD or situational roles.

[0125] Attached Figure 4The following diagram shows the semantic dependency diagram before and after pruning for the sentence "Enterprise data refers to data held by basic telecommunications companies that is not related to users, including network and system data, enterprise management data, partner data, etc."

[0126] 3. Tregex pattern generation

[0127] By observing the simplified dependency tree, we find the shortest path covering the classification and hierarchical tuples and represent it as a Tregex pattern.

[0128] Step 6: Construction of domain knowledge base for data classification and grading;

[0129] The domain knowledge base consists of two parts: domain = <L domain , O domain >, one is the information type dictionary

[0130] L domain ={c1, c2, ..., c n |c i ∈info_type, i∈N +}; the other is a data classification and grading ontology consisting of classification and grading tuples O domain =<C,A,R> ; Among them, info_type indicates the information type, c i Represents the concepts in the ontology, C = c1, c2, ..., c n , A represents the security level attribute, R represents the relationship between concepts, and the present invention specifically refers to the hierarchical relationship.

[0131] Table 1 shows the performance of SE-NBC in identifying categorical and graded sentences. The results show that SE-NBC can identify categorical and graded sentences with high accuracy and outperforms the machine learning and deep learning baselines to a large extent, verifying the effectiveness of the naive Bayes classifier based on semantic embedding.

[0132] Table 1 Performance comparison of SE-NBC and baseline models

[0133] Classifier Accuracy Accuracy Recall F1 value Random Forest 88.05% 90.19% 80.15% 83.38% AdaBoost 85.26% 82.40% 80.34% 81.26% LR 88.05% 86.54% 83.13% 84.58% DT 82.07% 78.69% 75.13% 76.53% KNN 81.67% 82.55% 70.16% 72.87% SVM 88.44% 86.59% 84.27% 85.31% CNN 84.46% 78.93% 81.55% 80.05% LSTM 91.63% 88.12% 90.11% 89.05% SE-NBC 94.42% 93.24% 93.24% 93.24%

[0134] Table 2 shows the performance of categorized and graded tuple extraction. The three components of information type and relationship are evaluated separately. The accuracy of information type extraction is 87.13%, the recall is 81.34%, and the F1 score is 83.98%. The F1 score of relationship extraction is lower than that of information type because the accuracy calculation method of relationship extraction is more stringent.

[0135] Table 2 Classification and grading tuple extraction performance

[0136] Evaluation object Accuracy Recall F1 value Information Type 0.8713 0.8134 0.8398 relation 0.8872 0.6211 0.7307

[0137] The method of the present invention can be compiled into a program code, which can be stored in a computer-readable storage medium, and the program code can be transmitted to a processor, which can then execute the method of the present invention.

[0138] This invention proposes an automatic parsing framework for policies and regulations, and constructs classification and grading information into a domain knowledge base, giving full play to the guiding role of policies and regulations on data classification and grading, which can effectively bridge the gap between national-level data protection concerns and specific countermeasures of organizations. More research can be carried out on this framework in the future.

Claims

1. A method for constructing a domain knowledge base for data classification and grading based on information extraction, characterized in that: The method comprises the following steps: S1, the document acquisition step, finds the target document in the target website or target database by keyword search, and summarizes it to form a corpus; S2, document preprocessing step, separating the target document obtained in step S1 into two categories: plain text and table; S3, text data extraction step, constructing a semantically embedded naive Bayes classifier, classifying the plain text separated in step S2 through the constructed naive Bayes classifier, and generating data classification and grading sentence labels; S4, table information extraction step, according to the table features and interesting information in the table in the corpus separated in step S2, split the merged cells, fill the empty cells according to the cell text before splitting, and then extract information based on pattern matching; S5, data classification and grading tuple extraction step, using a combination of pattern matching and natural language processing technology, based on the identified classification and grading sentence labels, extracting classification and grading elements to achieve joint extraction of information types and their relationships; S6, data classification and grading domain knowledge base construction step, after plain text and table extraction, data classification and grading tuples are obtained, and after semantic similarity calculation formula, duplicate removal and fusion are performed according to high and low levels to complete the classification and grading domain knowledge base construction; The S5 step specifically includes the following sub-steps: S501, semantic dependency analysis: use the semantic dependency analysis tool provided by LTP to parse the classified and graded sentences; S502, simplifying the semantic dependency tree: after parsing the classified and graded sentences in step S501, the original semantic dependency tree in the classified and graded sentences is obtained, and the original semantic dependency tree is simplified; S503, Tregex pattern generation: by observing the simplified semantic dependency tree, find the shortest path covering the classification and grading tuples, and express it as a Tregex pattern; In the step S501, the semantic dependency analysis tool provided by LTP is used to parse the classified and graded sentences, the semantic dependency relationship and part of speech in the classified and graded sentences are parsed, and the semantic dependency relationship tags and part of speech tags are marked in the classified and graded sentences; The semantic dependency relationship includes the subject EXP, the guest CONT, the point mark mPUNC, the dependency mark mDEPD, the parallel eCOO, the modification FEAT and the link LINK; In step S502, the method of simplifying the original semantic dependency tree includes adding, merging and deleting, wherein adding refers to adding a node with a part-of-speech tag u, c or mPUNC relationship tag to the previous node; merging refers to merging nodes marked as FEAT or rFEAT or dFEAT or MEAS or eCOO, or object nodes whose parent nodes are marked as n; and deleting refers to deleting nodes marked as FEAT or mRELA or mDEPD or situational roles; The classification and grading domain knowledge base constructed in step S6 consists of two parts: , one is the information type dictionary ; The other is a data classification and grading ontology consisting of classification and grading tuples ; in, Indicates the type of information. Represents the concepts in the ontology, , A Represents the security level attribute, R Represents the relationship between concepts.

2. A method for constructing a domain knowledge base for data classification and grading based on information extraction as claimed in claim 1, characterized in that: The S2 step specifically includes the following sub-steps: S201, deleting irrelevant content in the target document; S202, format conversion, if the target document is in PDF format, convert the PDF format files into word format files, and then use the python-docx library to separate text and tables; S203, using the language processing platform LTP to segment the plain text into sentences; wherein the hierarchical relationship of the table in the text is saved as a tuple, that is, the line number of the text and its direct parent node, to ensure that the sentence has a complete semantics and a simple structure; S204, using a Chinese word segmentation tool to segment all sentences in the text to form a list of segmentation phrases; then applying the stop word list cn_stopwords to remove function words, auxiliary words and general words.

3. A method for constructing a domain knowledge base for data classification and grading based on information extraction as claimed in claim 2, characterized in that: The S3 step specifically includes the following sub-steps: S301, calling the interface of the Python Sklearn library to generate a TF-IDF vector for each word in the word segmentation phrase list; at the same time, using the chi-square statistic to further filter out some redundant features; The output feature engineering vector is , For a sentence; S302, using weighted Word2vec to generate word embedding vectors that are specific to the corpus, assigning weights to each word embedding vector, and introducing additional semantic features. The weighted word embedding vector is: ; in, To use the Word2vec word embedding method, M is the total number of words contained in the sentence, For a word in a sentence, Words The term frequency-inverse document frequency weight; S303: concatenate the feature engineering vector output from step S301 and step S302 with the weighted word embedding vector ; in, is the concatenated vector; semantic information is introduced, and the concatenated vector is input into the naive Bayes classifier to complete the recognition of data classification and grading sentences.

4. A method for constructing a data classification and grading domain knowledge base based on information extraction as claimed in any one of claims 1 to 3, characterized in that: A set of sentences given , , represents the category label; the goal of the naive Bayes classifier is to Find the right tag , its formal expression is as follows: ; For the given sentence , the naive Bayes classifier calculates the posterior probability of the sentence under various variables , the class with the largest probability value is called the maximum a posteriori estimate, which is expressed as: 。 5. A method for constructing a data classification and grading domain knowledge base based on information extraction as claimed in any one of claims 1 to 3, characterized in that: The S4 step specifically includes the following sub-steps: S401, table structure analysis; the table structure analysis includes processing merged cells, complex header processing, filtering irregular tables and cross-page table determination; S402, table field extraction; first, formulate roles and matching rules for the fields to be extracted, specifically, by observing the classification and grading table in the target document, formulate a rule set according to the expression of the fields to be extracted in the table header and table content; then, traverse each column in order, match the content of each cell with the rule set, if a column is matched successfully, record the position index and role of the matched column in the table; finally, traverse the index column and extract the cell, which is the required field information.

Citation Information

Patent Citations

  • Multi-source network encyclopedia-oriented knowledge base construction method

    CN107239481A

  • Security analysis and automatic evaluation method based on index threshold value and semantic analysis

    CN114528848A