A data security grading method and device, electronic equipment and storage medium

By using a data classification model based on a knowledge graph architecture, the classification of data security levels is automated, solving the problems of low quality control and efficiency in traditional data classification methods, and achieving fast and accurate data security level classification and protection.

CN118797706BActive Publication Date: 2025-10-24CHINA MOBILE GROUP ZHEJIANG +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311631842.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-10-24
Estimated Expiration
2043-11-29

AI Technical Summary

Technical Problem

Traditional data classification methods suffer from problems such as uncontrollable data classification quality and low efficiency, making it difficult to quickly classify data security levels.

Method used

We adopt a data hierarchical model based on a knowledge graph architecture. Through knowledge extraction, representation, fusion, and security level classification, we use hierarchical knowledge graphs to automatically classify data security levels. This includes knowledge extraction, representation, fusion, and classification modules, thus constructing an efficient data security classification method.

Benefits of technology

It enables rapid classification of data security levels, avoids data tampering, leakage and illegal use, improves the accuracy and efficiency of data classification, and reduces model maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118797706B_ABST
    Figure CN118797706B_ABST
Patent Text Reader

Abstract

One or more embodiments of the specification disclose a data security grading method and device, electronic equipment and storage medium. The method comprises: inputting a target data set into a knowledge extraction module of a data grading model for knowledge extraction, obtaining first knowledge corresponding to the target data set, which contains entity data, relationship data and attribute data; using a knowledge representation module of the data grading model, performing entity relationship prediction on the first knowledge to obtain second knowledge; inputting the second knowledge into a knowledge fusion module of the data grading model for knowledge fusion processing, merging the second knowledge with a similarity greater than a threshold to obtain third knowledge corresponding to the target data set; inputting the third knowledge into a knowledge graph module of the data grading model, and using a hierarchical knowledge graph previously constructed in the knowledge graph module to divide the data security level of the third knowledge to obtain a hierarchical result of the security level corresponding to the target data set.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of data processing, and in particular, to a data security grading method and device, electronic equipment and a storage medium. BACKGROUND

[0002] With data security rising to the national security level and the national strategy level, data classification and grading has become a must for enterprise data security governance. Data grading is defined according to the sensitivity of data and the impact on victims after data is tampered with, destroyed, leaked or illegally used, according to certain principles and methods. Data grading is essentially a data classification of data sensitivity dimensions.

[0003] The traditional data grading method is a manual grading method: experts in data security governance analyze structured data, semi-structured data and unstructured data within the managed range according to data grading rules extracted from relevant national laws and regulations, analyze the fields and case data of these data, and match data grading rules to add corresponding data tags to data that meets the conditions. However, this manual data grading mode has problems such as uncontrollable data grading quality and low data grading efficiency. Therefore, there is an urgent need to provide a more optimal data security grading scheme. SUMMARY

[0004] The embodiments of the present specification provide a data security grading method and device, electronic equipment and a storage medium to solve the problems of uncontrollable data grading quality and low data grading efficiency.

[0005] In a first aspect, one or more embodiments of the present specification provide a data security grading method, comprising: inputting a target data set into a knowledge extraction module of a data grading model for knowledge extraction to obtain first knowledge corresponding to the target data set, the first knowledge comprising entity data, relationship data and attribute data;

[0006] Using a knowledge representation module of the data grading model, performing entity relationship prediction on the first knowledge to obtain second knowledge in the form of a triple, the triple comprising a first entity, a second entity, and a relationship between the first entity and the second entity, or the triple comprising an entity, an attribute of the entity, and an attribute value;

[0007] Inputting the second knowledge into a knowledge fusion module of the data grading model for knowledge fusion processing to merge the second knowledge with a similarity greater than a threshold to obtain third knowledge corresponding to the target data set;

[0008] The third knowledge is input into a knowledge graph module of the data classification model, the third knowledge is classified according to a data security level by using a hierarchical knowledge graph pre-constructed in the knowledge graph module, and a hierarchical result of a security level corresponding to the target data set is obtained, the hierarchical result of the security level corresponding to the target data set including the third knowledge and a hierarchical result of a security level corresponding to each piece of the third knowledge.

[0009] Preferably, the knowledge extraction module is used to extract the first knowledge corresponding to the target data set, including:

[0010] A knowledge dictionary is constructed according to the entity data, the relationship data and the attribute data included in the hierarchical knowledge graph.

[0011] The target data set is input into the knowledge extraction module, and the first knowledge is extracted from the target data set based on the knowledge dictionary.

[0012] Preferably, the knowledge representation module of the data classification model is used to predict an entity relationship of the first knowledge, and second knowledge in a triple form is obtained, including:

[0013] The knowledge representation module is constructed based on a distance model, a single-layer neural network model, a bilinear model, a neural tensor model, a matrix decomposition model or a translation model.

[0014] The second knowledge is determined by using the knowledge representation module based on semantic correlation between different pieces of the first knowledge in the semantic space.

[0015] Preferably, the second knowledge is input into the knowledge fusion module of the data classification model for knowledge fusion processing, and the second knowledge with a similarity greater than a threshold value is merged to obtain third knowledge corresponding to the target data set, including:

[0016] The second knowledge is input into the knowledge fusion module for entity alignment processing to obtain aligned second knowledge.

[0017] The aligned second knowledge is processed by using the knowledge fusion module based on an ontology matched with the target data set to obtain the third knowledge, and the knowledge processing is used to determine a dependency relationship between the second knowledge.

[0018] Preferably, the second knowledge is input into the knowledge fusion module for entity alignment processing to obtain aligned second knowledge, including:

[0019] determining, from the knowledge fusion module, a target entity alignment method matched with the second knowledge based on a preset target of entity alignment;

[0020] inputting the second knowledge into the knowledge fusion module, and performing entity alignment on the second knowledge using the target entity alignment method to obtain the second knowledge after entity alignment.

[0021] Preferably, the target data set is obtained from a preset address, and the method further comprises:

[0022] obtaining a target data set every interval of a predetermined time period from the preset address;

[0023] comparing third knowledge corresponding to a target data set of the N time periods before a time period corresponding to the latest target data set with the third knowledge corresponding to the latest target data set to obtain update knowledge for updating the hierarchical knowledge graph, N being a positive integer;

[0024] updating the hierarchical knowledge graph based on the update knowledge and a security level corresponding to the update knowledge.

[0025] Preferably, the updating the hierarchical knowledge graph based on the update knowledge and a security level corresponding to the update knowledge comprises:

[0026] quantifying the credibility of the update knowledge using the knowledge graph module;

[0027] sorting the credibility of the quantified update knowledge from large to small, and taking the update knowledge corresponding to a credibility less than a preset value as target update knowledge;

[0028] updating the hierarchical knowledge graph based on the target update knowledge and a security level corresponding to the target update knowledge.

[0029] In a second aspect, an apparatus for data security grading is provided, comprising: an extraction module configured to input a target data set into a knowledge extraction module of a data grading model to perform knowledge extraction, and obtain first knowledge corresponding to the target data set, the first knowledge comprising entity data, relationship data and attribute data;

[0030] a representation module configured to use a knowledge representation module of the data grading model to perform entity relationship prediction on the first knowledge, and obtain second knowledge in the form of a triple, the triple comprising a first entity, a second entity, and a relationship between the first entity and the second entity, or the triple comprising an entity, an attribute of the entity, and an attribute value;

[0031] a fusion module, configured to input the second knowledge into the knowledge fusion module of the data hierarchical model for knowledge fusion processing, so as to merge the second knowledge having a similarity greater than a threshold to obtain third knowledge corresponding to the target data set;

[0032] A grading module is used to input the third knowledge into the knowledge graph module of the data grading model, and use the hierarchical knowledge graph pre-constructed in the knowledge graph module to divide the data security level of the third knowledge, to obtain the grading result of the security level corresponding to the target data set, and the grading result of the security level corresponding to the target data set includes the third knowledge and the grading result of the security level corresponding to each piece of the third knowledge.

[0033] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0034] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0035] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0036] In the embodiments of this specification, expert experience is converted into a hierarchical knowledge graph. Through a data grading model based on a hierarchical knowledge graph architecture, knowledge extraction, knowledge representation, knowledge fusion and security level division of the data in the target data set are performed, and a grading result of the security level corresponding to the target data set is obtained. This security data grading method based on the knowledge graph architecture can quickly complete the data security level setting of the data middle station, avoid the tampering, destruction, leakage or illegal acquisition and illegal use of security-related data, and thus harm the legitimate rights and interests of the data owner, etc., and realize the hierarchical protection of security data. At the same time, the data grading model can be trained end-to-end to improve the generalization performance of the data security grading method. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the one or more embodiments of the present specification or the prior art, the accompanying drawings needed to be used in the embodiment or prior art description will be briefly introduced as follows. Obviously, the accompanying drawings in the following description are only some embodiments of the present specification. For those skilled in the art, other drawings can also be obtained from these accompanying drawings without any creative effort.

[0038] Figure 1 is a schematic flow chart of a method for data security grading according to an embodiment of the present specification.

[0039] Figure 2 is a schematic flow chart of a method for data security grading according to an embodiment of the present specification.

[0040] Figure 3 is a schematic flow chart of a method for knowledge extraction according to an embodiment of the present specification.

[0041] Figure 4 is a schematic flow chart of a method for constructing a hierarchical knowledge graph according to an embodiment of the present specification.

[0042] Figure 5 is a schematic diagram of an application scenario of a method for data security grading according to an embodiment of the present specification.

[0043] Figure 6 is a structural schematic diagram of an apparatus for data security grading according to an embodiment of the present specification.

[0044] Figure 7 is a structural schematic diagram of an electronic device according to an embodiment of the present specification. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present specification will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present specification. Obviously, the described embodiments are some of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without any creative effort belong to the scope of protection of the present application.

[0046] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this specification can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of a class, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0047] The data security classification method and device, electronic device, and storage medium provided in the embodiments of this specification are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0048] Figure 1 Show the existing data classification scheme, such as Figure 1 As shown in the figure, the existing data classification scheme has the following disadvantages:

[0049] First, the quality of data classification cannot be controlled: the quality of data classification depends on the individual capabilities of data security experts, and expert experience cannot be solidified or improved;

[0050] Second, the efficiency of data classification is relatively low: manual data classification analysis cannot quickly complete data security classification, and the processing efficiency is insufficient.

[0051] The above shortcomings limit the rapid classification of data security levels. Therefore, in the field of data security classification, how to achieve rapid classification of data security levels is one of the difficulties in this field.

[0052] Figure 2 A flowchart illustrating a data security classification method provided by an embodiment of the present invention is provided. The method can be executed by an electronic device, which may include a server and / or a terminal device, such as an in-vehicle terminal or a mobile phone terminal. In other words, the method can be executed by software or hardware installed on the aforementioned electronic device, and includes the following steps:

[0053] S202: Input the target data set into the knowledge extraction module of the data classification model to perform knowledge extraction, and obtain first knowledge corresponding to the target data set including entity data, relationship data and attribute data.

[0054] The first knowledge is knowledge directly extracted from each data of the target data set. The core of the knowledge extraction module is to configure the judgment rule for classifying the security data, so as to ensure that different data to be processed can be quickly marked after being processed by the subsequent knowledge graph module. The knowledge extraction module includes three functional units of entity extraction, attribute extraction and relationship extraction.

[0055] The knowledge extraction module can complete the identification and analysis of data, Figure 3 The specific processing flow of the knowledge extraction module is shown in FIG. 1. Figure 3 As shown in FIG. 1, the data extraction module needs to understand the data format, data length, data interval range, rationality, integrity, consistency and the like. After comprehensive analysis of the extracted data, the data security level label is obtained.

[0056] Natural person information will be represented by certificate type and certificate number. The most commonly used certificate types at present are ID card, passport and the like. Among these valid certificates, the ID card accounts for more than 97% of all valid certificate number information under normal circumstances. Therefore, the subsequent case is described by taking the ID card and its number as the representation of natural person information. The following takes the knowledge related to the ID card number as an example to illustrate the knowledge extracted from the ID card number. The knowledge related to the ID card number that can be sorted out at present includes:

[0057] 1) Data format: the ID card number is a character field;

[0058] 2) Data length: the data length is eighteen characters long;

[0059] 3) Interval range:

[0060] ① Address code: represents the administrative code of the county where the coding object resides, which is executed according to GB / T 2260;

[0061] ② Birth date code: represents the year, month and day of birth of the coding object, which is executed according to GB / T 7408. The year, month and day are represented by 4 digits, 2 digits and 2 digits respectively, without separator;

[0062] ③ Sequence code: represents the sequence number of people born on the same day in the same address code identified area, with odd sequence code assigned to male and even sequence code assigned to female;

[0063] ④ Police station code: the 15th and 16th digits represent the code of the local police station;

[0064] ⑤ Gender code: the 17th digit represents gender: odd for male and even for female;

[0065] 6) Check code: the 18th digit is a check code: it is calculated by the number preparation unit according to a unified formula. If the tail number of a person is 0-9, X will not appear, but if the tail number is 10, X will be replaced by X, which is the Roman numeral 10. Replacing 10 with X can ensure that the citizen's ID card meets the national standard;

[0066] 4) Rationality: the birthday should be within 120 years of the current time under normal circumstances, and cannot exceed the current date. The address code is within the coding range of GB / T2260;

[0067] 5) Integrity: the ID number is 18 digits long.

[0068] S204: using the knowledge representation module of the data hierarchical model, performing entity relationship prediction on the first knowledge to obtain second knowledge in the form of triples.

[0069] Among them, the extracted knowledge needs to be represented in a reasonable way. The knowledge obtained after performing entity relationship prediction (i.e., knowledge representation) on the first knowledge is the second knowledge. The triple form of the second knowledge can include the first entity, the second entity, and the relationship between the first entity and the second entity, or the triple form can include the entity, the attribute of the entity, and the attribute value.

[0070] The traditional resource relationship expression method is to symbolically describe the relationship between entities in the form of triples SPO (subject, property, object) of Resource Description Framework (RDF), but it faces many problems in terms of computational efficiency and data sparsity.

[0071] Representative representation learning techniques, represented by deep learning, have made important progress, which can represent the semantic information of entities as dense low-dimensional real-valued vectors, and then efficiently compute entities, relationships, and their complex semantic associations in a low-dimensional space, which is of great significance to the construction, reasoning, fusion, and application of knowledge bases.

[0072] The representative models of knowledge representation learning include distance models, single-layer neural network models, bilinear models, neural tensor models, matrix factorization models, and translation models. The specific knowledge expression method uses which model to exhibit, and selects the appropriate model in a specific application scenario. In an example, a bilinear model scheme can be used.

[0073] S206, input the second knowledge into the knowledge fusion module of the data hierarchical model for knowledge fusion processing, to merge the second knowledge with a similarity greater than a threshold value, to obtain third knowledge corresponding to the target data set.

[0074] Through the knowledge extraction operation of step S202, the goal of obtaining entity, relationship and entity attribute information from unstructured and semi-structured data is achieved. However, due to the wide range of knowledge sources, there are problems such as uneven quality of knowledge, repetition of knowledge from different data sources, and lack of hierarchical structure, so knowledge fusion must be performed. Knowledge fusion is a high-level knowledge organization, which integrates heterogeneous data, disambiguates, processes, reasons, verifies, updates and other steps under the same framework specification, and achieves the fusion of data, information, methods, experiences and human thoughts, forming a high-quality knowledge base.

[0075] The descriptions of the same entity in different sub-datasets in the target dataset are often different. Normalizing these data is an important step to improve the accuracy of subsequent security classification. In an example, the knowledge fusion processing performed by the knowledge fusion module can be performed according to the similarity between different second knowledge. Specifically, the knowledge fusion processing can include unifying different labels of the same entity or unifying the same entity in different information sources. The threshold used for merging can be determined according to actual conditions, which is not specifically limited in the specification.

[0076] S208, input the third knowledge into the knowledge graph module of the data classification model, use the pre-constructed classification knowledge graph in the knowledge graph module to divide the third knowledge into data security levels, and obtain the classification result of the security level corresponding to the target dataset.

[0077] The classification knowledge graph can be a knowledge graph constructed according to the experience of experts in data security management. The classification knowledge graph can be constructed manually or automatically according to expert experience. The construction process of the classification knowledge graph is not specifically limited in the specification and can be determined according to actual conditions. Figure 4 A process diagram for constructing a classification knowledge graph is shown. Specifically, the classification knowledge graph can be constructed according to the data security classification standards in the expert experience shown in Table 1.

[0078] The classification result of the security level corresponding to the target dataset can include the third knowledge and the classification result of the security level corresponding to each third knowledge. In an example, the classification result of the security level of the first knowledge related to the third knowledge can be determined according to the classification result of the security level corresponding to the third knowledge, and the classification result of the security level of the first knowledge is labeled in the target database for relevant personnel to review.

[0079] Table 1 Data Classification Standard

[0080]

[0081] In the embodiments of the present specification, the expert experience is converted into a hierarchical knowledge graph. Through a data hierarchical model based on the hierarchical knowledge graph architecture, knowledge extraction, knowledge representation, knowledge fusion and security level division of the data in the target data set are performed, and the hierarchical result of the security level corresponding to the target data set is obtained. The security data hierarchical method based on the knowledge graph architecture can quickly complete the setting of the data security level of the data center, avoid the security-related data from being tampered with, damaged, leaked or illegally obtained and utilized, and thus cause harm to the legitimate rights and interests of the data owner, etc., and realize hierarchical protection of the security data. At the same time, the data hierarchical model can be trained end-to-end to improve the generalization performance of the data security hierarchical method.

[0082] Figure 5 An application scenario diagram of the data security hierarchical method is provided, as shown in Figure 5 The user sends a service request for data security hierarchical of the target data set to the server through a terminal such as a mobile phone, the server obtains the hierarchical result by using the data security hierarchical method in the present specification after receiving the service request, and sends the hierarchical result to the terminal for the user to refer.

[0083] In an implementation manner, the target data set is obtained from a preset address, and the method further comprises:

[0084] The target data set is obtained from the preset address every predetermined time period;

[0085] The third knowledge corresponding to the target data set of the previous N time periods of the time period corresponding to the latest target data set is compared with the third knowledge corresponding to the latest target data set, to obtain update knowledge for updating the hierarchical knowledge graph, and N is a positive integer;

[0086] The hierarchical knowledge graph is updated based on the update knowledge and the security level corresponding to the update knowledge.

[0087] Specifically, the hierarchical knowledge graph can be updated according to the predetermined time period. The length of the predetermined time period can be set according to actual needs. In an example, after a special event occurs, the length of the predetermined time period can be shortened to improve the update frequency of the hierarchical knowledge graph, and thus improve the accuracy of the data security hierarchical. The special event can be a hot topic or news.

[0088] In an example, the newly emerging third knowledge can be taken as updated knowledge by comparing the third knowledge corresponding to the latest time period with the previous time period, and the hierarchical knowledge graph can be updated according to the updated knowledge and the security level of the updated knowledge. Further, a preset update value can be set for the updated knowledge, and when the number of the same updated knowledge exceeds the preset update value, the hierarchical knowledge graph is updated again, so as to shield the occasional updated knowledge in the hierarchical knowledge graph and improve the accuracy of data security classification.

[0089] In an implementation, the hierarchical knowledge graph is updated based on the updated knowledge and the security level corresponding to the updated knowledge, including:

[0090] The credibility of the updated knowledge is quantified using the knowledge graph module;

[0091] The credibility of the quantified updated knowledge is sorted from large to small, and the updated knowledge corresponding to the credibility less than a preset value is taken as target updated knowledge;

[0092] The hierarchical knowledge graph is updated based on the target updated knowledge and the security level corresponding to the target updated knowledge.

[0093] Specifically, in the process of updating the hierarchical knowledge graph, the credibility of the updated knowledge can be quantified, and the credibility is higher, and the credibility is lower, and then the quality of the target updated knowledge used to update the hierarchical knowledge graph can be effectively ensured, and the updating efficiency and accuracy of the hierarchical knowledge graph are improved.

[0094] When new knowledge appears, if the data classification model is retrained, the training time of the data classification model is longer, and the data classification model cannot be quickly updated with the frequency of special events, so the effect of the data classification model on data security classification is affected. In contrast, in the method in the embodiments of the present application, the accuracy of data security classification can be quickly improved by updating the hierarchical knowledge graph, and the model does not need to be frequently retrained, thereby reducing the maintenance cost of the data classification model.

[0095] In an implementation, the target data set is input into the knowledge extraction module of the data classification model to extract knowledge, and the first knowledge corresponding to the target data set containing entity data, relationship data and attribute data is obtained, including:

[0096] A knowledge dictionary is constructed according to the entity data, relationship data and attribute data contained in the hierarchical knowledge graph;

[0097] The target data set is input into the knowledge extraction module, and the first knowledge is extracted from the target data set based on the knowledge dictionary.

[0098] The knowledge dictionary can be constructed based on various knowledge contained in the hierarchical knowledge graph. Specifically, the knowledge dictionary can be constructed at the same time as the hierarchical knowledge graph is constructed.

[0099] After obtaining the knowledge dictionary, the first knowledge can be extracted from the target data set based on various knowledge for security classification contained in the knowledge dictionary.

[0100] Since the hierarchical knowledge graph already contains knowledge of various data that can be used for security classification, in the embodiments of the present specification, the knowledge dictionary is constructed according to the hierarchical knowledge graph, and the extraction of the first knowledge is based on the knowledge dictionary. The first knowledge more relevant to security classification can be obtained, thereby improving the efficiency and accuracy of security classification of data in the target database.

[0101] In one implementation, the knowledge representation module of the data classification model is used to perform entity relationship prediction on the first knowledge to obtain second knowledge in the form of a triple, including:

[0102] The knowledge representation module is used to map the first knowledge to a low-dimensional semantic space to obtain a semantic vector corresponding to the first knowledge;

[0103] Based on the semantic correlation between different first knowledge in the semantic space, the knowledge representation module is used to determine the second knowledge.

[0104] As mentioned earlier, the traditional resource relationship expression method faces many problems in terms of computational efficiency, data sparsity, etc. In an example, the knowledge representation module can be used to map the first knowledge to a low-dimensional semantic space, and determine the second knowledge according to the semantic correlation between different first knowledge in the semantic space, thereby improving the efficiency of knowledge expression and reducing the complexity of knowledge expression. Specifically, the second knowledge in the form of a triple can be constructed according to the first knowledge with high semantic correlation in the semantic space.

[0105] The knowledge representation module can be constructed based on a distance model, a single-layer neural network model, a bilinear model, a neural tensor model, a matrix decomposition model, or a translation model. The following describes these models.

[0106] 1) Distance model

[0107] The distance model proposes a structured embedding (SE) of entities and relations in a knowledge base. The basic idea is to first represent entities as vectors, then project entities into the vector space of entity-relation pairs through a relation matrix, and finally calculate the distance between the projected vectors to determine the confidence of the existing relationship between entities. The relation matrix in the distance model is two different matrices, which makes the synergy poor.

[0108] 2) Single layer neural network model

[0109] To address the above-mentioned defects in the distance model, the industry proposes to use a single layer neural network nonlinear model (SLM). The model defines the following evaluation function for each triple (h, r, t) in the knowledge base:

[0110]

[0111] where, is the vector representation of relation r; g() is the tanh function; M r,1 , M r,2 ∈R k are two matrices defined by relation r. Although the nonlinear operation of the single layer neural network model can further depict the semantic correlation of entities under the relationship, it greatly increases the computational overhead.

[0112] 3) Bilinear model

[0113] The bilinear model is also called a latent factor model (LFM). The evaluation function defined by the model for each triple in the knowledge base has the following form:

[0114]

[0115] In this formula, M r ∈R d×d is a bilinear transformation matrix defined by relation r; is the vector representation of the head entity h and the tail entity t in the triple. The bilinear model mainly depicts the semantic correlation of entities under the relationship through a bilinear transformation based on the relationship between entities. The model not only has a simple form and is easy to calculate, but also can effectively depict the synergy between entities.

[0116] 4) Neural tensor model

[0117] The basic idea of the neural tensor model is to link entities in different dimensions to represent the complex semantic relationship between entities. The model defines the following evaluation function for each triple (h, r, t) in the knowledge base:

[0118]

[0119] In the formula, is a vectorized representation of the relation r; g() is a tanh function; M r ∈R d×k×k is a third-order tensor; Mr,1, Mr,2∈R d×k are two matrices defined by the relation r.

[0120] In an example, a model used by the knowledge representation module to perform entity relation prediction (i.e., knowledge representation) on the target data set can be determined according to the field of the target data set, so as to improve the matching degree of the knowledge representation module and the target data set, and thus improve the accuracy of the second knowledge.

[0121] In the embodiments of the present specification, the first knowledge is mapped to a low-dimensional semantic space, and the second knowledge is determined according to the semantic correlation between different first knowledge in the semantic space, which improves the efficiency of knowledge expression and reduces the complexity of knowledge expression.

[0122] In an implementation manner, the second knowledge is input into a knowledge fusion module of a data hierarchical model to perform knowledge fusion, so as to merge the second knowledge with a similarity greater than a threshold to obtain third knowledge corresponding to the target data set, including:

[0123] The second knowledge is input into the knowledge fusion module to perform entity alignment processing to obtain the second knowledge after entity alignment;

[0124] Based on the ontology matched with the target data set, the knowledge fusion module is used to perform knowledge processing on the second knowledge after alignment to obtain third knowledge, and the knowledge processing is used to determine the dependency relationship between the second knowledge.

[0125] The entity alignment is also called entity matching, entity resolution or entity linking, and is mainly used to eliminate entity conflicts, pointing to unknown inconsistencies and other inconsistencies in heterogeneous data. Specifically, a large-scale unified knowledge base can be created from the top layer, thereby helping machines to understand multi-source heterogeneous data and form high-quality knowledge.

[0126] In the environment of big data, affected by the size of the knowledge base, the following challenges will be faced when performing knowledge base entity alignment:

[0127] 1) Computational complexity. The computational complexity of the matching algorithm will increase quadratically with the size of the knowledge base, which is difficult to accept;

[0128] 2) Data quality. Due to the different purposes and ways of constructing different knowledge bases, there may be problems such as uneven quality of knowledge, similar and repetitive data, isolated data, and inconsistent data time granularity;

[0129] 3) Prior training data. It is very difficult to obtain such prior data in large-scale knowledge bases. Usually, researchers need to manually construct prior training data.

[0130] Based on the above, the main process of knowledge base entity alignment can include:

[0131] 1) Partition index of the data to be aligned to reduce the complexity of the calculation;

[0132] 2) Use similarity function or similarity algorithm to find matching instances;

[0133] 3) Use entity alignment algorithm for instance fusion;

[0134] 4) Combine the results of steps 2) and 3) to form the final alignment result.

[0135] The alignment algorithm can be divided into pairwise entity alignment and collective entity alignment, and the collective entity alignment can be divided into local collective entity alignment and global collective entity alignment. Specifically:

[0136] 1) Pairwise entity alignment method

[0137] ① Entity alignment method based on traditional probability model

[0138] The entity alignment method based on traditional probability model mainly considers the similarity of the attributes of two entities, and does not consider the relationship between entities. The problem of judging whether the entities match based on attribute similarity score is converted into a classification problem, and a probability model of the problem is established, the disadvantage is that it does not reflect the influence of important attributes on entity similarity. Based on the probability entity linking model, different weights are assigned to each matched attribute pair, and the matching accuracy is improved. Also combined with Bayesian network to model the correlation of attributes, and use maximum likelihood estimation method to estimate the parameters in the model.

[0139] ② Entity alignment method based on machine learning

[0140] The entity alignment method based on machine learning mainly converts the entity alignment problem into a binary classification problem. According to whether the labeled data is used, it can be divided into supervised learning and unsupervised learning, the entity alignment method based on supervised learning can be mainly divided into pairwise entity alignment, clustering-based alignment, and active learning.

[0141] The method of judging the match of entity pairs by comparing attribute vectors is called pair-wise entity alignment. Typical representatives of this method include decision tree, support vector machine, and ensemble learning. The entity resolution is completed by using classification regression tree and linear analysis discrimination. Based on the two-stage entity linking analysis model, a new SVM classification method is proposed, which has a much higher matching accuracy than the hybrid algorithm in TAILOR.

[0142] The main idea of the clustering-based entity alignment algorithm is to gather similar entities together and then perform entity alignment. An adaptive entity name matching and clustering algorithm with strong scalability is proposed, which can generate an adaptive distance function through training samples. Using a similar method, a distance function is trained in the conditional random field entity alignment model using supervised learning, and then the weights are adjusted to maximize the product of the feature function and the learning parameters.

[0143] In active learning, the problem of insufficient training data can be solved by continuous interaction with personnel. The ALIAS system can complete the tasks of entity linking and deduplication through human-computer interaction. The ActiveAtlas system is constructed using a similar method.

[0144] 2) Local collective entity alignment method

[0145] The local collective entity alignment method sets different weights for the attributes of the entity itself and the attributes of the entities associated with it, and calculates the overall similarity by weighted summation. It also uses the vector space model and cosine similarity to determine the similarity of entities in large-scale knowledge bases. The algorithm establishes a name vector and a virtual document vector for each entity. The name vector is used to identify the attributes of the entity, and the virtual document vector is used to represent the weighted sum of the attribute values of the entity and its neighbor nodes. To evaluate the importance of each component in the vector, the algorithm mainly uses TF-IDF to set the weight of each component, and establishes an inverted index for the component vector. Finally, the cosine similarity function is used to calculate their similarity. This algorithm has high recall rate and fast execution speed, but its accuracy is insufficient. The fundamental reason is that it does not truly consider semantics.

[0146] 3) Global collective entity alignment method

[0147] ① Similarity propagation-based collective entity alignment method

[0148] The similarity propagation-based method is a typical collective entity alignment method. The two entities being matched and their directly associated other entities also have high similarity, which in turn affects the associated other entities.

[0149] The similarity propagation collective entity alignment method is originally derived from the set relation clustering algorithm proposed in 2003. The algorithm mainly generates matching objects through an improved hierarchical agglomerative algorithm. On the basis of the above algorithm, the algorithm SiGMa suitable for large-scale knowledge base entity alignment is proposed. The algorithm regards the entity alignment problem as an optimization problem of a global matching score objective function, which belongs to the quadratic assignment problem, and its approximate solution can be obtained by a greedy optimization algorithm. The SiGMa method can comprehensively consider the attributes and relations of the entity pairs, and iteratively discover all matching pairs through the domain of collective entities.

[0150] The collective entity alignment method based on a probability model mainly uses statistical relation learning for calculation and reasoning. Commonly used methods include the LDA model, the CRF model, and the Markov logic network.

[0151] The LDA model is applied to the entity resolution process to obtain the relationship between entities through the hidden variables. However, the effect is generally poor on large-scale data sets. A CRF entity resolution model based on graph partitioning technology is proposed. The model produces entity discrimination decisions based on observations, which is conducive to processing data with dependent attributes. Based on the CRF entity resolution model, a multi-relation entity linking algorithm based on the conditional random field model is proposed. The canopy-based index is introduced to improve the efficiency of collective entity alignment in large-scale knowledge base environments. A Markov logic network-based entity resolution method is proposed. Through the Markov logic network, a Markov network can be constructed to convert the maximum likelihood calculation problem in the probabilistic graphical model into a typical maximization weighted satisfiability problem. However, when performing entity resolution based on the Markov network, a series of equivalent predicate axioms need to be defined to complete the collective entity alignment of the knowledge base.

[0152] In an example, the second knowledge is input into the knowledge fusion module for entity alignment processing to obtain entity-aligned second knowledge, including:

[0153] Based on the preset entity alignment target, a target entity alignment method matching the second knowledge is determined from the multiple entity alignment methods preset in the knowledge fusion module.

[0154] The second knowledge is input into the knowledge fusion module, and the target entity alignment method is used for entity alignment of the second knowledge to obtain entity-aligned second knowledge.

[0155] The target of entity alignment can be to reflect the influence of important attributes on entity similarity, high matching accuracy, fast matching speed, etc. The plurality of preset entity alignment methods can be two or more of the aforementioned entity alignment methods. The advantages of the preset entity alignment methods can correspond to the entity alignment targets one by one. Specifically, when the target of the preset entity alignment is fast matching speed, the local collective entity alignment method can be used as the target entity alignment method. Since this method does not consider semantics, the execution speed can be greatly improved, but the accuracy of this entity alignment method is insufficient.

[0156] In the embodiments of the present specification, by presetting the target of entity alignment, the target entity alignment method is selected from the plurality of preset entity alignment methods, which can improve the matching degree of the entity alignment method and the target of entity alignment, and further achieve the effect of improving user satisfaction.

[0157] Through entity alignment, a series of basic fact expressions or preliminary ontology prototypes can be obtained. However, facts are not equal to knowledge, but are basic units of knowledge. To form high-quality knowledge, knowledge processing is needed to determine the dependency relationship between different knowledge and form a large-scale knowledge system from the level to manage knowledge uniformly. Knowledge processing mainly includes ontology construction and quality evaluation.

[0158] Ontology is the semantic basis for communication and connection between different subjects in the same field. It mainly presents a tree structure, and the adjacent hierarchical nodes or concepts have a strict "IsA" relationship, which is conducive to constraint and reasoning, but not conducive to expressing the diversity of concepts. The position of ontology in the knowledge graph is equivalent to the mold of the knowledge base. The knowledge base formed by the ontology library not only has a strong hierarchical structure, but also has a small degree of redundancy. In an example, the ontology can be preset, and the ontology matching the target data set is selected from the preset ontology, that is, the ontology matching the scene of the target data set is selected, so that the knowledge processing process is simpler and more flexible, and the efficiency of obtaining the third knowledge is improved.

[0159] In the embodiments of the present specification, the second knowledge is subjected to knowledge fusion including entity alignment and knowledge processing to obtain the third knowledge, which can eliminate the repeated entities, relationships, attributes, etc. in the second knowledge, form high-quality third knowledge, and further improve the efficiency of data security classification.

[0160] It should be noted that the method for data security classification provided in the embodiments of the present specification can be executed by a data security classification device or a control module in the data security classification device for executing the method for data security classification. In the embodiments of the present specification, the data security classification device executes the method for data security classification as an example to illustrate the data security classification device provided in the embodiments of the present specification.

[0161] Figure 6 is a structural schematic diagram of the data security grading apparatus according to an embodiment of the present application. As shown in the figure, the data security grading apparatus 600 comprises: Figure 6

[0162] The extraction module 610 is configured to input the target data set into a knowledge extraction module of the data grading model to perform knowledge extraction, and obtain first knowledge corresponding to the target data set, the first knowledge comprising entity data, relationship data and attribute data.

[0163] The representation module 620 is configured to use a knowledge representation module of the data grading model to perform entity relationship prediction on the first knowledge, and obtain second knowledge in the form of a triple, the triple comprising a first entity, a second entity and a relationship between the first entity and the second entity, or the triple comprising an entity, an attribute of the entity and an attribute value.

[0164] The fusion module 630 is configured to input the second knowledge into a knowledge fusion module of the data grading model to perform knowledge fusion processing, so as to merge second knowledge with a similarity greater than a threshold, and obtain third knowledge corresponding to the target data set.

[0165] The grading module 640 is configured to input the third knowledge into a knowledge graph module of the data grading model, and use a pre-constructed grading knowledge graph in the knowledge graph module to perform data security level division on the third knowledge, so as to obtain a grading result of a security level corresponding to the target data set, the grading result of the security level corresponding to the target data set comprising the third knowledge and a grading result of a security level corresponding to each piece of third knowledge.

[0166] In an implementation manner, the extraction module 610 comprises:

[0167] The construction unit is configured to construct a knowledge dictionary according to knowledge of entity data, relationship data and attribute data contained in the grading knowledge graph.

[0168] The extraction unit is configured to input the target data set into the knowledge extraction module, and extract the first knowledge from the target data set based on the knowledge dictionary.

[0169] In an implementation manner, the representation module 620 comprises:

[0170] The mapping unit is configured to use the knowledge representation module to map the first knowledge to a low-dimensional semantic space, so as to obtain a semantic vector corresponding to the first knowledge, the knowledge representation module being constructed based on a distance model, a single-layer neural network model, a bilinear model, a neural tensor model, a matrix decomposition model or a translation model.

[0171] ​The indicating unit is configured to determine the second knowledge by using the knowledge representation module based on semantic correlations between different first knowledge in the semantic space.

[0172] In an implementation manner, the fusing module 630 comprises:

[0173] The aligning unit is configured to input the second knowledge into the knowledge fusing module for entity alignment processing to obtain the second knowledge after entity alignment.

[0174] The processing unit is configured to perform knowledge processing on the second knowledge after alignment by using the knowledge fusing module based on the ontology matched with the target data set to obtain third knowledge, and the knowledge processing is used to determine the dependency relationship between the second knowledge.

[0175] In an implementation manner, the aligning unit is configured to:

[0176] Based on the preset target of entity alignment, determine the target entity alignment method matched with the second knowledge from the plurality of preset entity alignment methods in the knowledge fusing module.

[0177] Input the second knowledge into the knowledge fusing module, and perform entity alignment on the second knowledge by using the target entity alignment method to obtain the second knowledge after entity alignment.

[0178] In an implementation manner, the target data set is obtained from a preset address, and the data security grading apparatus 600 further comprises:

[0179] The collecting module is configured to obtain the target data set from the preset address every interval of a predetermined time period.

[0180] The comparing module is configured to compare the third knowledge corresponding to the target data set of the first N time periods of the time period corresponding to the latest target data set with the third knowledge corresponding to the latest target data set to obtain update knowledge used to update the graded knowledge graph, and N is a positive integer.

[0181] The updating module is configured to update the graded knowledge graph based on the update knowledge and the security level corresponding to the update knowledge.

[0182] In an implementation manner, the updating module comprises:

[0183] The quantifying unit is configured to quantify the credibility of the update knowledge by using the knowledge graph module.

[0184] The ranking unit is configured to sort the credibility of the update knowledge after quantification from large to small, and take the update knowledge corresponding to the credibility less than a preset value as target update knowledge.

[0185] The updating unit is configured to update the hierarchical knowledge graph based on the target updated knowledge and a security level corresponding to the target updated knowledge.

[0186] The data security grading apparatus in the embodiments of the present specification can be an apparatus, or a component in a terminal, an integrated circuit, or a chip. The apparatus can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., which are not specifically limited in the embodiments of the present specification.

[0187] The data security grading apparatus in the embodiments of the present specification can be an apparatus with an operating system. The operating system can be an Android operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present specification.

[0188] The data security grading apparatus provided in the embodiments of the present specification can implement the processes implemented in the method embodiments, which are not repeated here to avoid repetition. Figure 1

[0189] Based on the same idea, one or more embodiments of the present specification also provide an electronic device, as shown in Figure 7 The electronic device can have a large difference due to different configurations or performances, and can include one or more processors 701 and memories 702, and the memories 702 can store one or more storage applications or data. The memory 702 can be temporary storage or persistent storage. The applications stored in the memory 702 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the electronic device. Further, the processor 701 can be configured to communicate with the memory 702 and execute a series of computer executable instructions in the memory 702 on the electronic device. The electronic device can also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input and output interfaces 705, and one or more keyboards 706.

[0190] ​In particular embodiments, an electronic device includes memory, and one or more programs, wherein one or more programs are stored in the memory and accessible by one or more processors for one or more programs include one or more modules, and each module can include a set of computer-executable instructions that, when executed by the one or more processors, cause the electronic device to perform a computer process. The one or more programs include instructions for:

[0191] extracting knowledge from the target dataset into a knowledge extraction module of the data classification model to obtain first knowledge corresponding to the target dataset, the first knowledge including entity data, relationship data, and attribute data;

[0192] performing entity relationship prediction on the first knowledge using a knowledge representation module of the data classification model to obtain second knowledge in the form of triples, the triples including a first entity, a second entity, and a relationship between the first entity and the second entity, or the triples including an entity, an attribute of the entity, and an attribute value;

[0193] performing knowledge fusion processing on the second knowledge in a knowledge fusion module of the data classification model to merge second knowledge with a similarity greater than a threshold to obtain third knowledge corresponding to the target dataset;

[0194] performing data security level classification on the third knowledge using a pre-constructed hierarchical knowledge graph in a knowledge graph module of the data classification model to obtain a hierarchical result of a security level corresponding to the target dataset, the hierarchical result of the security level corresponding to the target dataset including the third knowledge and a hierarchical result of a security level corresponding to each piece of third knowledge.

[0195] One or more embodiments of the present specification also propose a storage medium storing one or more computer programs, the one or more computer programs including instructions capable of causing an electronic device including a plurality of applications to perform each process of the above-mentioned data security classification method embodiment when the instructions are executed by the electronic device, and specifically for performing:

[0196] extracting knowledge from the target dataset into a knowledge extraction module of the data classification model to obtain first knowledge corresponding to the target dataset, the first knowledge including entity data, relationship data, and attribute data;

[0197] performing entity relationship prediction on the first knowledge using a knowledge representation module of the data classification model to obtain second knowledge in the form of triples, the triples including a first entity, a second entity, and a relationship between the first entity and the second entity, or the triples including an entity, an attribute of the entity, and an attribute value;

[0198] Inputting the second knowledge into the knowledge fusion module of the data classification model for knowledge fusion processing, so as to merge the second knowledge with similarity greater than a threshold to obtain the third knowledge corresponding to the target data set;

[0199] The third knowledge is input into the knowledge graph module of the data classification model, and the data security level of the third knowledge is divided into levels using the hierarchical knowledge graph pre-built in the knowledge graph module to obtain the classification results of the security level corresponding to the target data set. The classification results of the security level corresponding to the target data set include the third knowledge and the classification results of the security level corresponding to each piece of third knowledge.

[0200] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the aforementioned storage medium embodiment is generally similar to the method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the method embodiment.

[0201] The methods, devices, modules, or units described in the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0202] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0203] It will be understood by those skilled in the art that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0204] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or a combination thereof. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or a combination thereof. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks.

[0205] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or a combination thereof. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or a combination thereof. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks.

[0206] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or a combination thereof. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or a combination thereof. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks.

[0207] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0208] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) for example, for storage of information and instructions to be executed by the processor. The memory can also include non-volatile memory, such as read only memory (ROM) and / or flash memory for example, for storage of static information and instructions.

[0209] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0210] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0211] One or more embodiments of the specification can be described in the general context of computer-executable instructions being executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The specification can also be practiced in a distributed computing environment in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0212] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment focuses on the difference from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0213] The above merely provides one or more embodiments of the present specification, and is not intended to limit the present application. One of ordinary skill in the art can make various modifications and changes to the one or more embodiments of the present specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the one or more embodiments of the present specification shall be included in the scope of the claims of the one or more embodiments of the present specification.

Claims

1. A method for data security grading, comprising: inputting a target data set into a knowledge extraction module of a data grading model for knowledge extraction, to obtain first knowledge corresponding to the target data set, the first knowledge comprising entity data, relationship data, and attribute data; using a knowledge representation module of the data grading model to perform entity relationship prediction on the first knowledge, to obtain second knowledge in the form of triples, the triples comprising a first entity, a second entity, and a relationship between the first entity and the second entity, or the triples comprising an entity, an attribute of the entity, and an attribute value; inputting the second knowledge into a knowledge fusion module of the data grading model for knowledge fusion processing, to merge the second knowledge with a similarity greater than a threshold, to obtain third knowledge corresponding to the target data set; inputting the third knowledge into a knowledge graph module of the data grading model, and using a pre-constructed hierarchical knowledge graph in the knowledge graph module to perform data security level division on the third knowledge, to obtain a hierarchical result of a security level corresponding to the target data set, the hierarchical result of the security level corresponding to the target data set comprising the third knowledge and a hierarchical result of a security level corresponding to each piece of the third knowledge.

2. The method of claim 1, wherein the inputting a target data set into a knowledge extraction module of a data grading model for knowledge extraction, to obtain first knowledge corresponding to the target data set, the first knowledge comprising entity data, relationship data, and attribute data, comprises: constructing a knowledge dictionary according to entity data, relationship data, and attribute data contained in the hierarchical knowledge graph; inputting the target data set into the knowledge extraction module, and extracting the first knowledge from the target data set based on the knowledge dictionary.

3. The method of claim 1, wherein the using a knowledge representation module of the data grading model to perform entity relationship prediction on the first knowledge, to obtain second knowledge in the form of triples, comprises: using the knowledge representation module to map the first knowledge to a low-dimensional semantic space, to obtain a semantic vector corresponding to the first knowledge, the knowledge representation module being constructed based on a distance model, a single-layer neural network model, a bilinear model, a neural tensor model, a matrix factorization model, or a translation model; determining the second knowledge based on semantic correlations between different pieces of the first knowledge in the semantic space using the knowledge representation module.

4. The method of claim 1, wherein the inputting the second knowledge into a knowledge fusion module of the data grading model for knowledge fusion processing, to merge the second knowledge with a similarity greater than a threshold, to obtain third knowledge corresponding to the target data set, comprises: inputting the second knowledge into the knowledge fusion module for entity alignment processing, to obtain second knowledge after entity alignment; performing knowledge processing on the aligned second knowledge using the knowledge fusion module based on an ontology matching the target data set, to obtain the third knowledge, the knowledge processing being used to determine a dependency relationship between the second knowledge.

5. The method of claim 4, wherein the inputting the second knowledge into the knowledge fusion module for entity alignment processing to obtain entity-aligned second knowledge comprises: determining, based on a preset entity alignment target, a target entity alignment method that matches the second knowledge from a plurality of preset entity alignment methods in the knowledge fusion module; inputting the second knowledge into the knowledge fusion module, and using the target entity alignment method to perform entity alignment on the second knowledge to obtain the entity-aligned second knowledge.

6. The method of claim 1, wherein the target data set is obtained from a preset address, and the method further comprises: obtaining a target data set every interval of a predetermined time period from the preset address; comparing third knowledge corresponding to target data sets of the last N time periods with the third knowledge corresponding to the latest target data set to obtain update knowledge for updating the hierarchical knowledge graph, N being a positive integer; updating the hierarchical knowledge graph based on the update knowledge and a security level corresponding to the update knowledge.

7. The method of claim 6, wherein the updating the hierarchical knowledge graph based on the update knowledge and a security level corresponding to the update knowledge comprises: quantifying a credibility of the update knowledge using the knowledge graph module; ranking the credibility of the quantified update knowledge from large to small, and regarding the update knowledge corresponding to a credibility less than a preset value as target update knowledge; updating the hierarchical knowledge graph based on the target update knowledge and a security level corresponding to the target update knowledge.

8. An apparatus for data security grading, comprising: an extraction module configured to input a target data set into a knowledge extraction module of a data grading model to perform knowledge extraction, and obtain first knowledge corresponding to the target data set, the first knowledge including entity data, relationship data, and attribute data; a representation module configured to use a knowledge representation module of the data grading model to perform entity relationship prediction on the first knowledge, and obtain second knowledge in the form of a triple, the triple including a first entity, a second entity, and a relationship between the first entity and the second entity, or the triple including an entity, an attribute of the entity, and an attribute value; a fusion module configured to input the second knowledge into a knowledge fusion module of the data grading model to perform knowledge fusion processing, and merge the second knowledge with a similarity greater than a threshold to obtain third knowledge corresponding to the target data set; a grading module configured to input the third knowledge into a knowledge graph module of the data grading model, and use a pre-constructed hierarchical knowledge graph in the knowledge graph module to divide the third knowledge into data security levels to obtain a hierarchical result of a security level corresponding to the target data set, the hierarchical result of the security level corresponding to the target data set including the third knowledge and a hierarchical result of a security level corresponding to each piece of the third knowledge.

9. An electronic device, comprising: a processor; and ​ a memory arranged to store computer-executable instructions configured to be executed by the processor, the computer-executable instructions comprising instructions for performing the steps in the method of any of claims 1 to 7.

10. A storage medium, characterized by the storage medium is for storing computer-executable instructions that cause a computer to perform the method of any of claims 1 to 7.

Citation Information

Patent Citations

  • Security monitoring method and device for network data, equipment and storage medium

    CN113726784A

  • Knowledge graph biased classification for data

    IN201747000240A