Data Sensitivity Identification Method and Device Based on Sensitivity Identification Model
By processing metadata based on the feature extraction and recognition layer of the sensitivity identification model, the problem of time-consuming and laborious manual identification and high missed-reception rate is solved, efficient and accurate automatic data sensitivity recognition is achieved, and the risk of data leakage is reduced.
Patent Information
- Application Number
- CN202110139667.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-02-01
AI Technical Summary
In the prior art, it is time-consuming and labor-intensive to identify data sensitivity through manual means and the probability of missing sensitive data is high, resulting in an increase in the risk of data leakage.
Using a method based on the sensitivity recognition model, the metadata is processed through the feature extraction layer and the sensitivity recognition layer, and the data sensitivity is automatically identified, including obtaining metadata, feature extraction and sensitivity recognition, and using machine learning technology to train the model to improve recognition efficiency.
It improves the efficiency of data sensitivity identification, reduces the probability of missing sensitive data, and ensures data security.
Smart Images

Figure CN114840869B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence and Internet technologies, and in particular, to a data sensitivity recognition method based on a sensitivity recognition model, a training method for the sensitivity recognition model, an apparatus, a device, and a computer-readable storage medium. Background Art
[0002] In the data asset management of Internet enterprises, with the development of business and the improvement of user activity, a large amount of valuable data will be deposited in database tables or texts. As a part of metadata, data sensitivity classifies data according to leakage risks, which is convenient for developers to use and keep confidential. However, if some valuable data lacks specific data sensitivity or risk levels and is not managed and maintained by developers, this part of the data may be leaked during use, which will have a great impact on the business.
[0003] In related technologies, data sensitivity is identified manually, that is, a database administrator identifies and determines the data sensitivity of the data to be identified according to personal experience. However, this method is time-consuming and laborious, and the probability of missing sensitive data is relatively high. Summary of the Invention
[0004] Embodiments of the present application provide a data sensitivity recognition method based on a sensitivity recognition model, a training method for the sensitivity recognition model, an apparatus, a device, and a computer-readable storage medium, which can improve the recognition efficiency of data sensitivity and reduce the probability of missing sensitive data.
[0005] The technical solution of the embodiments of the present application is implemented as follows:
[0006] Embodiments of the present application provide a data sensitivity recognition method based on a sensitivity recognition model. The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer, and the method includes:
[0007] Obtain metadata of the data to be recognized, where the metadata is used to describe the data to be recognized;
[0008] Through the feature extraction layer, perform feature extraction on the metadata of the data to be recognized to obtain data features of the metadata;
[0009] Through the sensitivity recognition layer, based on the data features of the metadata, perform sensitivity recognition on the data to be recognized to obtain a sensitivity recognition result;
[0010] Wherein, the sensitivity recognition result is used to indicate the data sensitivity corresponding to the data to be recognized.
[0011] An embodiment of the present application provides a method for training a sensitivity recognition model. The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer. The method includes:
[0012] Obtain the metadata of the data sample. The data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample;
[0013] Through the feature extraction layer, perform feature extraction on the metadata of the data sample to obtain the sample data features of the metadata of the data sample;
[0014] Through the sensitivity recognition layer, based on the sample data features, perform sensitivity recognition on the data sample to obtain a sample sensitivity recognition result;
[0015] Obtain the difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample, and based on the difference, update the model parameters of the sensitivity recognition model;
[0016] Among them, the sensitivity recognition model is used to output a sensitivity recognition result indicating the data sensitivity corresponding to the data to be recognized after inputting the metadata of the data to be recognized into the sensitivity recognition model.
[0017] An embodiment of the present application provides a data sensitivity recognition device based on a sensitivity recognition model. The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer. The device includes:
[0018] A first acquisition module, configured to acquire the metadata of the data to be recognized, where the metadata is used to describe the data to be recognized;
[0019] A first extraction module, configured to perform feature extraction on the metadata of the data to be recognized through the feature extraction layer to obtain the data features of the metadata;
[0020] A first recognition module, configured to perform sensitivity recognition on the data to be recognized through the sensitivity recognition layer based on the data features of the metadata to obtain a sensitivity recognition result;
[0021] Among them, the sensitivity recognition result is used to indicate the data sensitivity corresponding to the data to be recognized.
[0022] In the above solution, the first acquisition module is further configured to, when the storage form of the data to be recognized is a data table, acquire at least one of the following table elements from the data table: the data table name, the table description corresponding to the data to be recognized in the data table, and the attribute fields corresponding to the data to be recognized in the data table;
[0023] Determine the acquired table elements as the metadata of the data to be recognized.
[0024] In the above solution, the first acquisition module is further configured to, when the storage form of the data to be recognized is a document, obtain at least one of the following document contents from the document: document title, document abstract, and document keywords;
[0025] Determine the obtained document content as the metadata of the data to be recognized.
[0026] In the above solution, the first extraction module is further configured to perform word segmentation on the metadata of the data to be recognized to obtain a plurality of words corresponding to the metadata;
[0027] Perform feature encoding on each of the words to obtain word features corresponding to each of the words;
[0028] Perform feature splicing on the word features corresponding to each of the words to obtain data features corresponding to the metadata.
[0029] In the above solution, the first extraction module is further configured to perform bidirectional encoding processing on the word features of each word to obtain an upstream encoding feature and a downstream encoding feature corresponding to each word;
[0030] Perform feature splicing on the upstream encoding feature and the downstream encoding feature of each word to obtain a corresponding spliced encoding feature;
[0031] Perform feature splicing on the spliced encoding features corresponding to each of the words to obtain data features corresponding to the metadata.
[0032] In the above solution, the first recognition module is further configured to, through the sensitivity recognition layer, perform classification prediction on the data features of the metadata corresponding to at least two sensitivity levels to obtain probabilities corresponding to each of the sensitivity levels of the metadata;
[0033] Select the sensitivity level with the highest probability as the sensitivity recognition result of the data to be recognized.
[0034] In the above solution, the first extraction module is further configured to, when the metadata includes at least two keywords, perform feature extraction on each of the keywords through the feature extraction layer to obtain features corresponding to each of the keywords as the data features of the metadata;
[0035] Correspondingly, the first extraction module is further configured to, through the sensitivity recognition layer, match the features corresponding to each of the keywords with the features corresponding to at least two sensitive words to obtain corresponding matching degrees;
[0036] Select the data sensitivity corresponding to the sensitive word with the highest matching degree as the sensitivity recognition result of the data to be recognized.
[0037] In the above solution, the device further includes:
[0038] A processing module, configured to establish an association relationship between the sensitivity recognition result and the data to be recognized, and store the association relationship;
[0039] Wherein, the association relationship is used to find the data sensitivity corresponding to the data to be recognized based on the data to be recognized.
[0040] In the above solution, the processing module is further configured to store the sensitivity recognition result in a target area associated with the data to be recognized, and the target area is an area corresponding to the data sensitivity in the storage area corresponding to the metadata.
[0041] In the above solution, the device further includes:
[0042] A return module, configured to obtain the data sensitivity corresponding to the data to be recognized in response to a data display request for the data to be recognized;
[0043] When the data sensitivity corresponding to the data to be recognized reaches a sensitivity threshold, return a shielding indication message corresponding to the data to be recognized;
[0044] The shielding indication message is used to indicate to perform a shielding display on the data to be recognized.
[0045] In the above solution, the device further includes:
[0046] An output module, configured to output an encryption prompt message corresponding to the data to be recognized when the sensitivity recognition result indicates that the data sensitivity of the data to be recognized reaches a target data sensitivity;
[0047] Wherein, the encryption prompt message is used to prompt to perform an encryption process on the data to be recognized.
[0048] An embodiment of the present application provides a training device for a sensitivity recognition model. The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer. The device includes:
[0049] A second acquisition module, configured to acquire metadata of a data sample. The data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample;
[0050] A second extraction module, configured to perform feature extraction on the metadata of the data sample through the feature extraction layer to obtain sample data features of the metadata of the data sample;
[0051] A second recognition module, configured to perform sensitivity recognition on the data sample based on the sample data features through the sensitivity recognition layer, so as to obtain a sample sensitivity recognition result;
[0052] An update module, configured to obtain a difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample, and update model parameters of the sensitivity recognition model based on the difference;
[0053] Wherein, the sensitivity recognition model is configured to output a sensitivity recognition result indicating the data sensitivity corresponding to the data to be recognized after inputting metadata of the data to be recognized into the sensitivity recognition model.
[0054] An embodiment of the present application provides an electronic device, including:
[0055] A memory, configured to store executable instructions;
[0056] A processor, configured to implement the data sensitivity recognition method provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0057] An embodiment of the present application provides an electronic device, including:
[0058] A memory, configured to store executable instructions;
[0059] A processor, configured to implement the training method of the sensitivity recognition model provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0060] An embodiment of the present application provides a computer-readable storage medium, storing executable instructions, configured to cause a processor to implement the data sensitivity recognition method provided by the embodiment of the present application when executed.
[0061] The embodiment of the present application further provides a computer-readable storage medium, storing executable instructions, configured to cause a processor to implement the training method of the sensitivity recognition model provided by the embodiment of the present application when executed.
[0062] The embodiment of the present application has the following beneficial effects:
[0063] The server performs sensitivity recognition on the metadata of the data to be recognized through a sensitivity recognition model. Specifically, it obtains the metadata used to describe the data to be recognized, extracts features from the metadata of the data to be recognized through the feature extraction layer of the sensitivity recognition model to obtain the data features of the metadata; through the sensitivity recognition layer of the sensitivity recognition model, based on the data features of the metadata, it performs sensitivity recognition on the data to be recognized to obtain a sensitivity recognition result; in this way, by inputting the metadata to be recognized into the sensitivity recognition model, the sensitivity recognition result indicating the data sensitivity corresponding to the data to be recognized can be automatically recognized. Compared with the manual recognition method, it can greatly improve the recognition efficiency of data sensitivity and reduce the probability of missing sensitive data inspection. Description of the Drawings
[0064] Figure 1 FIG. 1 is an optional schematic architecture diagram of a data sensitivity recognition system 100 based on a sensitivity recognition model provided by an embodiment of the present application;
[0065] Figure 2 FIG. 2 is an optional structural schematic diagram of an electronic device 500 provided by an embodiment of the present application;
[0066] Figure 3 FIG. 3 is a schematic flowchart of a data sensitivity recognition method based on a sensitivity recognition model provided by an embodiment of the present application;
[0067] Figure 4 FIG. 4 is a schematic structural diagram of a sensitivity recognition model provided by an embodiment of the present application;
[0068] Figure 5 FIG. 5 is a schematic structural diagram of a sensitivity recognition model provided by an embodiment of the present application;
[0069] Figure 6 FIG. 6 is a schematic structural diagram of a sensitivity recognition model provided by an embodiment of the present application;
[0070] Figure 7 FIG. 7 is a schematic flowchart of a training method of a sensitivity recognition model provided by an embodiment of the present application;
[0071] Figure 8 FIG. 8 is a schematic flowchart of a data sensitivity recognition method based on a sensitivity recognition model provided by an embodiment of the present application;
[0072] Figure 9 FIG. 9 is a schematic flowchart of a data sensitivity recognition method based on a sensitivity recognition model provided by an embodiment of the present application;
[0073] Figure 10 FIG. 10 is a schematic structural diagram of a sensitivity recognition model provided by an embodiment of the present application;
[0074] Figure 11Schematic structural diagram of a data sensitivity recognition device based on a sensitivity recognition model provided by an embodiment of the present application;
[0075] Figure 12 Schematic structural diagram of a training device for a sensitivity recognition model provided by an embodiment of the present application. Detailed implementation manners
[0076] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0077] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0078] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0079] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0080] 1) Metadata is data that describes data, or structural data used to provide information about a certain resource (i.e., the data to be recognized). It is mainly used to describe the data attribute information of the data to be recognized and support functions such as indicating the storage location, historical data, resource search, and file recording. Metadata can be called an electronic catalog. To achieve the purpose of compiling a catalog, it is necessary to describe and collect the content or characteristics of the data, and then achieve the purpose of assisting data retrieval.
[0081] 2) In response to is used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or have a set delay. Without special instructions, there is no limit on the execution order of the multiple executed operations.
[0082] Based on the above explanations of the nouns and terms involved in the embodiments of the present application, the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application will be described next. See Figure 1 , Figure 1FIG. 0 is an alternative schematic architecture diagram of a data sensitivity recognition system 100 based on a sensitivity recognition model provided by an embodiment of the present application. To support an exemplary application, terminals (exemplarily shown as terminal 400-1 and terminal 400-2) are connected to a server 200 through a network 300. The network 300 can be a wide area network, a local area network, or a combination of both, and uses a wireless link to implement data transmission.
[0083] In practical applications, a client is set on the terminal, such as Weibo, Zhihu, enterprise applications, etc., which is used to provide data to be recognized related to the business or data to be recognized related to the user's behavior, and send the data to be recognized to the server 200. The server 200 can be either a separately configured server that supports various services, can also be configured as a server cluster, or can be a cloud server, etc. For example, it can be the background server of the client or an information flow platform.
[0084] In actual implementation, the server 200 is used to obtain the metadata of the data to be recognized, where the metadata is used to describe the data to be recognized; through the feature extraction layer of the sensitivity recognition model, extract the features of the metadata of the data to be recognized to obtain the data features of the metadata; through the sensitivity recognition layer of the sensitivity recognition model, based on the data features of the metadata, perform sensitivity recognition on the data to be recognized to obtain a sensitivity recognition result; where the sensitivity recognition result is used to indicate the data sensitivity corresponding to the data to be recognized.
[0085] Next, an electronic device for implementing the data sensitivity recognition method based on the sensitivity recognition model according to the embodiment of the present application will be described. Refer to Figure 2 , Figure 2 FIG. 13 is an alternative structural schematic diagram of an electronic device 500 provided by an embodiment of the present application. In practical applications, the electronic device 500 can be Figure 1 the terminals in Figure 1 such as terminal 400-1 and terminal 400-2) or the server 200. Taking the server 200 shown in Figure 2 as an example, the electronic device 500 shown in Figure 2 includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each component in the electronic device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. The bus system 540 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 540.
[0086] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0087] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0088] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 510.
[0089] The memory 550 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0090] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0091] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0092] A network communication module 552, for reaching other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB);
[0093] A presentation module 553 for enabling presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 531 associated with the user interface 530 (e.g., a display screen, a speaker, etc.);
[0094] An input processing module 554 for detecting and translating one or more user inputs or interactions from one of one or more input devices 532.
[0095] In some embodiments, the data sensitivity recognition device based on the sensitivity recognition model provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is a data sensitivity recognition device 555 based on a sensitivity recognition model stored in the memory 550, which can be software in the form of a program, a plug-in, etc., including the following software modules: a first acquisition module 5551, a first extraction module 5552, and a first recognition module 5553. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.
[0096] In other embodiments, the data sensitivity recognition device based on the sensitivity recognition model provided by the embodiments of the present application can be implemented in hardware. As an example, the data sensitivity recognition device based on the sensitivity recognition model provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0097] Based on the above description of the data sensitivity recognition system and electronic device based on the sensitivity recognition model in the embodiments of the present application, the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application will be described next. In some embodiments, this method can be implemented independently by a terminal or a server, such as by Figure 1 the terminal 400-1, the terminal 400-2 or the server 200 in Figure 1The terminal 400-1 and the server 200 in Figure 1 and Figure 3 , Figure 3 is a schematic flowchart of the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application. Taking the server 200 in Figure 1 as an example to implement the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application for illustration.
[0098] Step 101: The server obtains the metadata of the data to be recognized, where the metadata is used to describe the data to be recognized.
[0099] In practical applications, the data to be recognized can be enterprise business-related data, or personal user-related data. It can be data obtained from a database, or real-time obtained data. The storage form of the data to be recognized can be a data table, or a text form such as text or log. The metadata is mainly used to describe the attributes of the data to be recognized. For example, if the data to be recognized is shopping business-related data, the metadata can be data such as shopping accounts, order numbers, names, mobile phone numbers, delivery addresses, etc.; if the data to be recognized is personal user-related data, the metadata can be data such as names, ID numbers, mobile phone numbers, email addresses, bank card numbers, home addresses, work units, etc.
[0100] In some embodiments, the server can obtain the metadata of the data to be recognized in the following way: when the storage form of the data to be recognized is a data table, obtain at least one of the following table elements from the data table: data table name, table description corresponding to the data to be recognized in the data table, and attribute fields corresponding to the data to be recognized in the data table; determine the obtained table elements as the metadata of the data to be recognized.
[0101] Here, in practical applications, when the storage form of the data to be recognized is a data table, the data table name, table description, or attribute fields in the data table are used as the metadata of the data to be recognized. For example, attribute fields such as the data table name, Chinese table name, table responsible person, field name, or field type in the data table corresponding to the data to be recognized are used as metadata.
[0102] In some embodiments, the server can also obtain the metadata of the data to be recognized in the following way: when the storage form of the data to be recognized is a document, obtain at least one of the following document contents from the document: document title, document abstract, document keywords; determine the obtained document contents as the metadata of the data to be recognized.
[0103] Here, the document can be a Word document or a text document (such as a TXT document). When the storage form of the data to be recognized is a document, the document title, document abstract, and document keywords of the data to be recognized are used as metadata. When the data to be recognized includes a document title and a document body, in addition to using the document title as metadata, key abstract content (i.e., the document abstract) can also be extracted from the document body as metadata. This is because in practice, the document body of the data to be recognized may be quite long. If all of the document body is recognized, it will inevitably bring a great computational pressure and lead to low recognition efficiency. Usually, the core theme of the data to be recognized can be summarized by one or a few sentences. Therefore, in order to effectively extract the core theme of the data to be recognized and improve the recognition efficiency at the same time, the abstract content used to represent the core theme of the data to be recognized can be extracted from the document body as the metadata of the data to be recognized.
[0104] In some embodiments, the corresponding document keywords can be obtained by extracting keywords from the document body, and the document abstract of the data to be recognized can be obtained in the following way: sentence extraction is performed on the document body of the data to be recognized to obtain multiple target sentences corresponding to the data to be recognized; according to the word weights of multiple keywords in each target sentence, the sentence weight of the corresponding target sentence is determined; based on each sentence weight, the target sentences are sorted in descending order to obtain the corresponding sentence sequence; starting from the first target sentence in the sentence sequence, a target number of target sentences are selected, and the target number of target sentences are used as the document abstract of the corresponding data to be recognized.
[0105] Among them, the server can perform the following operations on each target sentence respectively to implement determining the sentence weight of the corresponding target sentence according to the word weights of multiple keywords in each target sentence: keyword extraction is performed on the target sentence to obtain the corresponding multiple keywords; the word frequency of each keyword corresponding in the document body and the inverse document frequency of each keyword are respectively obtained; based on the word frequency and the inverse document frequency, the word weight of the corresponding keyword is determined; the word weights of each keyword are summed to obtain the sentence weight of the corresponding target sentence.
[0106] Here, the term frequency represents the ratio of the frequency of occurrence of the keyword in the data to be recognized to the total number of words in the data to be recognized, and the inverse document frequency represents the rarity of the keyword, which is represented by the logarithm of the ratio of the total number of data in the dataset to which the data to be recognized belongs to the number of data containing the corresponding keyword in the dataset to which the data to be recognized belongs. In addition, in addition to considering the term frequency of the keyword, the rarity of the keyword is also comprehensively considered. In actual implementation, the importance of a keyword is not only proportional to its frequency in the data to be recognized, but also inversely proportional to how many data in the dataset to which the data to be recognized belongs contain it. Generally speaking, the more data containing the keyword, the more general it is and the less it can reflect the characteristics of the data. Finally, the sum of the word weights of the keywords in the target sentence is determined as the sentence weight of the target sentence. In this way, the sentence weight of each target sentence is obtained. The larger the sentence weight, the more the corresponding target sentence can represent the core theme of the data to be recognized.
[0107] Through the above method, subsequent data sensitivity recognition is performed based on the metadata of the data to be recognized. Since the metadata can not only represent the attribute characteristics of the data to be recognized, but also the data volume is greatly reduced compared to the data to be recognized, not only the accuracy of recognition can be guaranteed, but also the efficiency of recognition can be improved.
[0108] Step 102: Through the feature extraction layer, extract features from the metadata of the data to be recognized to obtain the data features of the metadata.
[0109] In some embodiments, refer to Figure 4 , Figure 4 which is the structural schematic diagram of the sensitivity recognition model provided by the embodiment of the present application. As Figure 4 shown, the sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer. The metadata of the data to be recognized is input into the sensitivity recognition model. Through the feature extraction layer, features are extracted from the metadata to obtain the corresponding data features. Through the sensitivity recognition layer, sensitivity recognition is performed on the data features to obtain the sensitivity recognition result.
[0110] In some embodiments, refer to Figure 5 , Figure 5 which is the structural schematic diagram of the sensitivity recognition model provided by the embodiment of the present application. As Figure 5 shown, the server can extract features from the metadata of the data to be recognized in the following way to obtain the data features of the metadata:
[0111] Perform word segmentation on the metadata of the data to be recognized to obtain multiple words corresponding to the metadata; perform feature encoding on each word to obtain the word features corresponding to each word; perform feature splicing on the word features corresponding to each word to obtain the data features corresponding to the metadata.
[0112] Here, in actual implementation, word segmentation is performed on the metadata of the data to be recognized, such as the data table name or table description, to obtain multiple words or characters corresponding to the metadata. In practical applications, in order to facilitate the retrieval of corresponding words or characters, a unique index value can also be set for each word or character, that is, the corresponding word or character is obtained based on the index value of each word or character, and then each word or character is feature-encoded, such as word vector conversion, to obtain the corresponding word feature, that is, the word vector; then the word features corresponding to each word or character are feature-stitched to obtain the corresponding data feature, that is, the sentence vector.
[0113] In some embodiments, the server can perform feature stitching on the word features corresponding to each word in the following manner to obtain the data feature corresponding to the metadata:
[0114] Perform bidirectional encoding processing on the word features of each word respectively to obtain the upstream encoding feature and downstream encoding feature corresponding to each word; perform feature stitching on the upstream encoding feature and downstream encoding feature of each word respectively to obtain the corresponding stitched encoding feature; perform feature stitching on the stitched encoding features corresponding to each word to obtain the data feature corresponding to the metadata.
[0115] Here, considering the word context feature, after obtaining the word vector of each word, the word vector of each word is input into a bidirectional encoding layer, such as a bidirectional long short-term memory network (Bi-LSTM) layer. Among them, the Bi-LSTM layer includes two LSTMs: one is a forward input sequence and the other is a backward input sequence. The upstream encoding feature corresponding to each word is extracted through the forward process (such as from left to right), and the downstream encoding feature vector corresponding to each word is extracted through the backward process (such as from right to left). Finally, the upstream encoding feature and the downstream encoding feature are stitched to obtain the stitched encoding feature of the corresponding word.
[0116] Step 103: Through the sensitivity recognition layer, based on the data feature of the metadata, perform sensitivity recognition on the data to be recognized to obtain a sensitivity recognition result.
[0117] In some embodiments, the server can perform sensitivity recognition on the data to be recognized through the sensitivity recognition layer based on the data feature of the metadata in the following manner to obtain a sensitivity recognition result:
[0118] Through the sensitivity recognition layer, perform classification prediction on the data feature of the metadata corresponding to at least two sensitivity levels to obtain the probability of each sensitivity level corresponding to the metadata; select the sensitivity level with the highest probability as the sensitivity recognition result of the data to be recognized.
[0119] Among them, the sensitivity recognition result is used to indicate the data sensitivity corresponding to the data to be recognized. There are various forms of data sensitivity, such as it can be characterized by a sensitivity level or a sensitivity value, etc. When the data sensitivity of the data to be recognized is characterized by a sensitivity value, the larger the sensitivity value, the more sensitive the data to be recognized is; when the data sensitivity of the data to be recognized is characterized by a sensitivity level, the custom-defined sensitivity levels for data sensitivity are: publicly disclosed, internally disclosed, generally sensitive, particularly sensitive, and highly confidential, and they correspond to the five natural numbers from 1 to 5 in sequence. The enterprise's own data sensitivity level standard can be defined with reference to industry standards and relevant regulations of the national legislative department in terms of data security.
[0120] It should be noted that the determination of the number of sensitivity levels should not only be conducive to the reasonable distinction of data sensitivity but also consider the feasibility of implementing security control measures based on different sensitivity levels. Generally, 4 to 5 levels are more reasonable. When 5 levels are selected, from high to low, they are: 5 (highly confidential), 4 (particularly sensitive), 3 (generally sensitive), 2 (internally disclosed), and 1 (publicly disclosed); for the definition of sensitivity levels here, for a data table, it should be accurate to the sensitivity level of the field. For example, for fields such as ID number and mobile phone number, the level is 5, and for name, email address, delivery address, etc., it is 4. In addition, the data sensitivity of the data to be recognized can also be characterized only by the custom-defined sensitivity levels. For example, the sensitivity levels can be divided into five types: top secret, confidential, highly sensitive, medium sensitive, and low sensitive.
[0121] Here, assume that the sensitivity levels corresponding to the data to be recognized are the following five types: top secret, confidential, highly sensitive, medium sensitive, and low sensitive. If through the sensitivity recognition layer, the data characteristics of the metadata are classified and predicted, and the probabilities corresponding to the above sensitivity levels are: top secret (90%), confidential (40%), highly sensitive (30%), medium sensitive (15%), and low sensitive (10%) in sequence, then it can be known that the sensitivity level with the largest probability (90%) is selected as top secret, and top secret is used as the sensitivity recognition result of the data to be recognized.
[0122] In some embodiments, refer to Figure 6 , Figure 6 which is the structural schematic diagram of the sensitivity recognition model provided by the embodiment of the present application. As Figure 6As shown, the server can also extract features from the metadata of the data to be recognized through the feature extraction layer to obtain the data features of the metadata, including: when the metadata includes at least two keywords, through the feature extraction layer, respectively extract features from each keyword to obtain the features corresponding to each keyword as the data features of the metadata; correspondingly, the server can identify the sensitivity of the data to be recognized based on the data features of the metadata through the sensitivity recognition layer in the following manner to obtain the sensitivity recognition result: through the sensitivity recognition layer, respectively match the features corresponding to each keyword with the features corresponding to at least two sensitive words to obtain the corresponding matching degrees; select the data sensitivity corresponding to the sensitive word with the highest matching degree as the sensitivity recognition result of the data to be recognized.
[0123] Here, the server prestores the correspondence between sensitive words and the corresponding data sensitivities. For example, the data sensitivity corresponding to sensitive word 1 is top secret, the data sensitivity corresponding to sensitive word 2 is confidential, the data sensitivity corresponding to sensitive word 3 is highly sensitive, the data sensitivity corresponding to sensitive word 4 is low sensitive, and the data sensitivity corresponding to sensitive word 5 is low sensitive. Assume that the keywords corresponding to the metadata of the data to be recognized include: keyword 1 and keyword 2. Then, through the feature extraction layer, respectively extract features from keyword 1 and keyword 2 to obtain the feature corresponding to keyword 1 and the feature corresponding to keyword 2; through the sensitivity recognition layer, respectively match the feature of keyword 1 with the features of the above-mentioned sensitive words (such as sensitive words 1 to 5) to obtain the corresponding matching degrees in turn: 10%, 20%, 30%, 40%, 80%. Respectively match the feature of keyword 2 with the features of the above-mentioned sensitive words (such as sensitive words 1 to 5) to obtain the corresponding matching degrees in turn: 20%, 10%, 30%, 40%, 60%. Then select the low sensitivity corresponding to sensitive word 5 with the highest matching degree of 80% as the sensitivity recognition result of the data to be recognized.
[0124] In some embodiments, after obtaining the sensitivity recognition result, the server can also establish an association relationship between the sensitivity recognition result and the data to be recognized and store the association relationship; wherein, the association relationship is used to search for the data sensitivity corresponding to the data to be recognized based on the data to be recognized.
[0125] In some embodiments, the server can establish an association relationship between the sensitivity recognition result and the data to be recognized in the following manner:
[0126] Store the sensitivity recognition result in the target area associated with the data to be recognized, and the target area is the area corresponding to the data sensitivity in the storage area corresponding to the metadata.
[0127] Here, after the server determines the data sensitivity of the data to be recognized, the data sensitivity can also be added to the area associated with the data to be recognized for indicating data sensitivity. For example, the data sensitivity of the data to be recognized can be filled in the "sensitive level" column in the data label, and the data sensitivity is used as part of the metadata for users to use and maintain.
[0128] In some embodiments, after obtaining the sensitivity recognition result, the server can also obtain the data sensitivity corresponding to the data to be recognized in response to a data display request for the data to be recognized; when the data sensitivity corresponding to the data to be recognized reaches the sensitivity threshold, a shielding indication message for the corresponding data to be recognized is returned; wherein, the shielding indication message is used to indicate to shield and display the data to be recognized.
[0129] Here, when the data sensitivity of the data to be recognized reaches the sensitivity threshold, it indicates that the data to be recognized is relatively sensitive, such as confidential or top-secret data. At this time, the server returns a shielding indication message for the corresponding data to be recognized to the terminal so that the user can perform security maintenance on the data to be recognized at the terminal, such as shielding confidential or top-secret data to avoid leakage; in addition, some data in the data to be recognized can be selectively displayed. For example, some sensitive information in the table, such as the user's ID number, does not want to be shown to others, and this field can be shielded with a view.
[0130] In some embodiments, after obtaining the sensitivity recognition result, when the sensitivity recognition result indicates that the data sensitivity of the data to be recognized reaches the target data sensitivity, the server can also output an encryption prompt message for the corresponding data to be recognized; wherein, the encryption prompt message is used to prompt to perform an encryption process on the data to be recognized.
[0131] By the above method, when the data sensitivity of the data to be recognized reaches a certain level, the user is prompted to encrypt the data to be recognized when maintaining or using the data to be recognized to avoid leakage. For example, when the data to be recognized needs to be transferred from one database to another database, the method provided by the embodiments of the present application automatically recognizes the data sensitivity of the data to be recognized, and further performs an encryption process or a blurring process on the data to be recognized that meets a certain sensitivity, so as to further improve the security of the data to be transferred.
[0132] Next, the training of the sensitivity recognition model will be described. Refer to Figure 7 , Figure 7 which is a schematic flowchart of the training method of the sensitivity recognition model provided by the embodiments of the present application. In some embodiments, the sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer. The method includes:
[0133] Step 201: The server obtains the metadata of the data sample, where the data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample.
[0134] Step 202: Through the feature extraction layer, feature extraction is performed on the metadata of the data sample to obtain the sample data features of the metadata of the data sample.
[0135] Step 203: Through the sensitivity recognition layer, based on the sample data features, sensitivity recognition is performed on the data sample to obtain the sample sensitivity recognition result.
[0136] Step 204: Obtain the difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample, and based on the obtained difference, update the model parameters of the sensitivity recognition model.
[0137] In actual implementation, the value of the loss function of the sensitivity recognition model can be determined according to the difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample; when the value of the loss function reaches the preset threshold, the corresponding error signal is determined based on the value of the loss function of the sensitivity recognition model; the error signal is backpropagated in the sensitivity recognition model, and the model parameters of each layer of the sensitivity recognition model are updated during the propagation process.
[0138] Here, an explanation of backpropagation is given. The training data sample is input into the input layer of the neural network model, passes through the hidden layer, and finally reaches the output layer and outputs the result. This is the forward propagation process of the neural network model. Since there is an error between the output result of the neural network model and the actual result, the error between the output result and the actual value is calculated, and the error is backpropagated from the output layer to the hidden layer until it reaches the input layer. During the backpropagation process, the values of the model parameters are adjusted according to the error; the above process is continuously iterated until convergence.
[0139] In the above manner, the server inputs the metadata to be recognized into the sensitivity recognition model, and can automatically recognize the sensitivity recognition result used to indicate the data sensitivity corresponding to the data to be recognized. Compared with the manual recognition method, it can greatly improve the recognition efficiency of data sensitivity and reduce the probability of missing sensitive data.
[0140] Next, continue to describe the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application. In some embodiments, in combination with Figure 1 and Figure 8 , Figure 8 is the flow schematic diagram of the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application. Taking Figure 1Taking the cooperation between the terminal in [ID] and the server 200 to implement the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application as an example, the sensitivity recognition model provided by the embodiments of the present application includes a feature extraction layer and a sensitivity recognition layer. The method includes:
[0141] Step 301: The server obtains the metadata of the data sample. Among them, the data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample.
[0142] Step 302: The server extracts features from the metadata of the data sample through the feature extraction layer to obtain the sample data features of the metadata of the data sample.
[0143] Step 303: The server performs sensitivity recognition on the data sample based on the sample data features through the sensitivity recognition layer to obtain the sample sensitivity recognition result.
[0144] Step 304: The server obtains the difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample, and updates the model parameters of the sensitivity recognition model based on the obtained difference.
[0145] Through the above method, a sensitivity recognition model is trained.
[0146] Step 305: The terminal transmits the data to be recognized by the user to the server.
[0147] Step 306: If the storage form of the data to be recognized is a data table, the server obtains the data table name, table description, or attribute field in the data table as metadata.
[0148] Step 307: The server extracts features from the metadata of the data to be recognized through the feature extraction layer to obtain the data features of the metadata.
[0149] Step 308: The server performs sensitivity recognition on the data to be recognized based on the data features of the metadata through the sensitivity recognition layer to obtain the sensitivity recognition result.
[0150] Step 309: The server stores the sensitivity recognition result in the target area associated with the data to be recognized.
[0151] Among them, the target area is the area corresponding to the data sensitivity in the storage area corresponding to the metadata of the data to be recognized.
[0152] In the above manner, the data sensitivity of the data to be recognized is recognized by the trained sensitivity recognition model, and the corresponding sensitivity recognition result is stored in the target area associated with the data to be recognized, so that the data sensitivity of the data to be recognized becomes a part of the metadata, greatly improving the recognition efficiency of data sensitivity and avoiding missed checks of the data sensitivity of the data to be recognized.
[0153] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described. The data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application mainly uses machine learning to recognize the data sensitivity of the data to be recognized. Refer to Figure 9 , Figure 9 which is a schematic flowchart of the data sensitivity recognition method based on the sensitivity recognition model provided by the embodiments of the present application. As Figure 9 shown, the recognition of data sensitivity provided by the embodiments of the present application includes: training of the sensitivity recognition model (i.e., the training stage) and recognition of the data sensitivity of the data to be recognized based on the trained sensitivity recognition model (i.e., the recognition stage), which will be described one by one below.
[0154] In the training stage, the metadata of the data sample is obtained. Among them, the data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample. That is, the metadata of the data sample input to the sensitivity recognition model includes: the data table name, the table description corresponding to the data sample in the data table, and the data sensitivity (i.e., the sensitivity label) to which the data table (i.e., the data sample) belongs. Among them, the data sensitivity can be characterized by a sensitivity level, and the sensitivity level is divided into five types: top secret, confidential, highly sensitive, medium sensitive, and low sensitive.
[0155] Generally speaking, the sensitivity level of data related to user account security is top secret, the sensitivity level of user personal information and financial-related data is confidential, the sensitivity level of user behavior data is highly sensitive, the sensitivity level of the up-rolled large-grained data of data with a sensitivity level of confidential is medium sensitive, and the sensitivity level of ordinary statistical data is low sensitive.
[0156] During training, the data table name and table description of the data sample are used as sample points, and the sensitivity level is used as the sensitivity label. By optimizing and training the sensitivity recognition model, the relationship between the data table name and table description and the sensitivity level is learned. After training, the model parameters of the sensitivity recognition model are saved.
[0157] In the recognition stage, during the recognition process, first load the model parameters of the sensitivity recognition model saved in the training stage. Then, input the metadata of the data to be recognized, that is, the data table name and table description of the data to be recognized, into the trained sensitivity recognition model to recognize the data sensitivity of the data to be recognized, and obtain the sensitivity level indicating the corresponding data sensitivity of the data to be recognized.
[0158] Next, the structure of the sensitivity recognition model will be described. Refer to Figure 10 , Figure 10 which is the structural schematic diagram of the sensitivity recognition model provided by the embodiment of the present application. As Figure 10 shown, the sensitivity recognition model includes an input layer, a feature extraction layer, and a sensitivity recognition layer. Among them, the feature extraction layer includes: an embedding layer, a bidirectional encoding layer, and a pooling layer. Next, taking the application of recognizing the data sensitivity of the data to be recognized as an example, the sensitivity recognition model will be described.
[0159] 1. Input layer
[0160] In the input layer, first perform word segmentation on the metadata of the data to be recognized, such as the data table name or table description, to obtain multiple words or characters corresponding to the metadata. Then, set a unique index value for each word or character. For example, the i-th word input is w i , and after indexing, a unique integer number I i = I(w i ); Finally, transmit each word or character obtained by word segmentation, and the corresponding index value to the feature extraction layer.
[0161] 2. Feature extraction layer
[0162] 1) Embedding layer
[0163] Here, in the embedding layer, first obtain the corresponding word or character based on the index value of each word or character, and then perform word vector conversion (i.e., feature encoding) on each word or character to obtain the corresponding word vector (i.e., word feature).
[0164] Assume that the matrix of the embedding layer is E ∈ R V*D , where V is the total number of all words, and D is the dimension of each word vector. To obtain the word vector of the i-th word, first convert its index value into a One-Hot encoding vector with a vector length of V. There is only an element 1 at the position of I i , and the elements at the remaining positions are all 0. Multiply the One-Hot encoding vector by the matrix E to obtain the word vector e i corresponding to the word segmentation. The specific expression is:
[0165] O i ∈ 0 V
[0166]
[0167] 2) Bidirectional Encoding Layer
[0168] After obtaining the word vectors of each word, the word vectors of each word are input into the bidirectional encoding layer, such as the bidirectional long short-term memory (Bi-LSTM) layer. The Bi-LSTM layer includes two LSTMs: one for the forward input sequence and one for the backward input sequence, which can consider the context features simultaneously and play a role in fully integrating and understanding the context semantics.
[0169] In actual implementation, the word vectors of each word can be processed by bidirectional encoding respectively to obtain the upper-context encoding features and lower-context encoding features corresponding to each word; the upper-context encoding features and lower-context encoding features of each word are respectively concatenated to obtain the corresponding concatenated encoding features. The specific expression is:
[0170]
[0171] where l represents from left to right, r represents from right to left, represents the hidden state of the previous word, represents the current input word, represents the cell state of the previous word; represents the upper-context encoding features extracted through the forward process (such as from left to right), represents the lower-context encoding feature vector extracted through the backward process (such as from right to left), C t ,h t represents the concatenated encoding features of the corresponding word obtained by concatenating the upper-context encoding features and the lower-context encoding features .
[0172] 3) Pooling Layer
[0173] Through the above bidirectional encoding layer, the concatenated encoding features corresponding to each word are obtained. Through the pooling layer, the concatenated encoding features corresponding to each word are concatenated to obtain the corresponding sentence vector (i.e., the data feature of the metadata). The specific expression is:
[0174]
[0175] where z represents the sentence vector corresponding to the metadata of the data to be recognized, C t ,h t is the concatenated encoding feature corresponding to the current input word, and L represents the total number of input words.
[0176] 3. Sensitivity recognition layer
[0177] Here, the sensitivity recognition layer is also called the multi-layer perceptron (MLP) layer. The multi-layer perceptron consists of multiple fully-connected neural networks. The data features corresponding to the metadata of the data to be recognized pass through the sensitivity recognition layer, and the probability of the data to be recognized corresponding to each sensitive level is output. Taking a 3-layer fully-connected neural network as an example, the probability of belonging to each sensitive level can be referred to the following expression:
[0178] a i = f(W3f(W2f(W1z + b1)+b2)+b3)
[0179]
[0180] where f is a non-linear activation function, z is the sentence vector corresponding to the metadata of the data to be recognized obtained from the above pooling layer, W1 is the weight of the first-layer fully-connected neural network, W2 is the weight of the second-layer fully-connected neural network, W3 is the weight of the third-layer fully-connected neural network, the weights are trainable, b1, b2, and b3 are the corresponding trainable bias parameters, a i represents the i-th sensitive level, A represents the number of sensitive levels, and p i represents the probability that the sensitive level of the data to be recognized belongs to a i .
[0181] Then, select the sensitive level with the maximum probability as the sensitive level corresponding to the data to be recognized, output the finally determined sensitive level, and supplement the output sensitive level to the metadata of the data to be recognized.
[0182] It should be noted that the structure of the above sensitivity recognition model can be set according to the actual situation. For example, the metadata of the data to be recognized can be input to the input layer, and through the input layer, the metadata of the data to be recognized is transmitted to the feature extraction layer to perform word segmentation or indexing operations on the metadata in the feature extraction layer, etc. The present application does not specifically limit the structure of sensitivity recognition.
[0183] After the structural layout of the sensitivity recognition model is completed, the sensitivity recognition model can be trained using the method of stochastic gradient descent to make the model parameters optimal or locally optimal. For example, the metadata of the acquired data sample is transmitted through the input layer to the feature extraction layer. Through the feature extraction layer, the metadata of the data sample is subjected to feature extraction to obtain the sample data features of the metadata of the data sample. Through the sensitivity recognition layer, based on the sample data features, the sensitivity of the data sample is recognized to obtain the sample sensitivity recognition result. The difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample is obtained, and based on the obtained difference, the model parameters of the sensitivity recognition model are updated.
[0184] In addition, the sensitivity recognition model provided in the embodiments of the present application can also be trained based on traditional machine learning methods, such as FastText; or in a deep learning manner, such as a general model based on the deformed Bidirectional Encoder Representations from Transformers (BERT), TextCNN model, Chinese pre-trained RoBERTa model, Chinese-trained ELECTRA model, etc. The fully connected neural network provided in the embodiments of the present application can also adopt an attention network, a recurrent neural network, and a convolutional neural network, etc.
[0185] Through the above method, the metadata to be recognized is input into the sensitivity recognition model, and the corresponding sensitivity level is automatically recognized by using machine learning, and the recognized sensitivity level is supplemented into the metadata of the data to be recognized. Compared with the manual recognition method, it can greatly improve the recognition efficiency of data sensitivity and reduce the probability of missing sensitive data.
[0186] Next, the implementation of the data sensitivity recognition device 555 based on the sensitivity recognition model provided in the embodiments of the present application as a software module will be continued. In some embodiments, as Figure 11 shown Figure 11 is a schematic structural diagram of the data sensitivity recognition device based on the sensitivity recognition model provided in the embodiments of the present application. Among them, the sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer. The device includes:
[0187] A first acquisition module 5551, configured to acquire metadata of data to be recognized, where the metadata is used to describe the data to be recognized;
[0188] A first extraction module 5552, configured to perform feature extraction on the metadata of the data to be recognized through the feature extraction layer to obtain the data features of the metadata;
[0189] The first recognition module 5553 is configured to perform sensitivity recognition on the data to be recognized through the sensitivity recognition layer based on the data characteristics of the metadata, so as to obtain a sensitivity recognition result;
[0190] Wherein, the sensitivity recognition result is used to indicate the data sensitivity corresponding to the data to be recognized.
[0191] In some embodiments, the first acquisition module is further configured to, when the storage form of the data to be recognized is a data table, acquire at least one of the following table elements from the data table: the data table name, the table description corresponding to the data to be recognized in the data table, and the attribute fields corresponding to the data to be recognized in the data table;
[0192] Determine the acquired table elements as the metadata of the data to be recognized.
[0193] In some embodiments, the first acquisition module is further configured to, when the storage form of the data to be recognized is a document, acquire at least one of the following document contents from the document: the document title, the document abstract, and the document keywords;
[0194] Determine the acquired document contents as the metadata of the data to be recognized.
[0195] In some embodiments, the first extraction module is further configured to perform word segmentation on the metadata of the data to be recognized to obtain a plurality of words corresponding to the metadata;
[0196] Perform feature encoding on each of the words respectively to obtain word features corresponding to each of the words;
[0197] Perform feature splicing on the word features corresponding to each of the words to obtain the data features corresponding to the metadata.
[0198] In some embodiments, the first extraction module is further configured to perform bidirectional encoding processing on the word features of each word respectively to obtain an upstream encoding feature and a downstream encoding feature corresponding to each word;
[0199] Perform feature splicing on the upstream encoding feature and the downstream encoding feature of each word respectively to obtain a corresponding spliced encoding feature;
[0200] Perform feature splicing on the spliced encoding features corresponding to each word to obtain the data features corresponding to the metadata.
[0201] In some embodiments, the first recognition module is further configured to perform classification prediction on the data features of the metadata corresponding to at least two sensitive levels through the sensitivity recognition layer to obtain the probabilities corresponding to each of the sensitive levels of the metadata;
[0202] Select the sensitivity level with the highest probability as the sensitivity recognition result of the data to be recognized.
[0203] In some embodiments, the first extraction module is further configured to, when the metadata includes at least two keywords, respectively extract features of each keyword through the feature extraction layer to obtain features corresponding to each keyword as data features of the metadata.
[0204] Correspondingly, the first extraction module is further configured to respectively match the features corresponding to each keyword with the features corresponding to at least two sensitive words through the sensitivity recognition layer to obtain corresponding matching degrees.
[0205] Select the data sensitivity corresponding to the sensitive word with the highest matching degree as the sensitivity recognition result of the data to be recognized.
[0206] In some embodiments, the apparatus further includes:
[0207] A processing module, configured to establish an association relationship between the sensitivity recognition result and the data to be recognized, and store the association relationship.
[0208] Wherein, the association relationship is used to search for the data sensitivity corresponding to the data to be recognized based on the data to be recognized.
[0209] In some embodiments, the processing module is further configured to store the sensitivity recognition result in a target area associated with the data to be recognized, and the target area is an area corresponding to the data sensitivity in the storage area corresponding to the metadata.
[0210] In some embodiments, the apparatus further includes:
[0211] A return module, configured to obtain the data sensitivity corresponding to the data to be recognized in response to a data display request for the data to be recognized.
[0212] When the data sensitivity corresponding to the data to be recognized reaches a sensitivity threshold, return shielding indication information corresponding to the data to be recognized.
[0213] The shielding indication information is used to indicate shielding display of the data to be recognized.
[0214] In some embodiments, the apparatus further includes:
[0215] An output module, configured to output encryption prompt information corresponding to the data to be recognized when the sensitivity recognition result indicates that the data sensitivity of the data to be recognized reaches a target data sensitivity.
[0216] Among them, the encrypted prompt information is used to prompt to perform encryption processing on the data to be recognized.
[0217] Next, continue to describe the training device of the sensitivity recognition model provided by the embodiments of the present application. Refer to Figure 12 , Figure 12 which is a schematic structural diagram of the training device of the sensitivity recognition model provided by the embodiments of the present application. The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer. The training device 120 of the sensitivity recognition model includes:
[0218] A second acquisition module 121, configured to acquire metadata of a data sample. The data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample;
[0219] A second extraction module 122, configured to perform feature extraction on the metadata of the data sample through the feature extraction layer to obtain sample data features of the metadata of the data sample;
[0220] A second recognition module 123, configured to perform sensitivity recognition on the data sample based on the sample data features through the sensitivity recognition layer to obtain a sample sensitivity recognition result;
[0221] An update module 124, configured to obtain the difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample, and update the model parameters of the sensitivity recognition model based on the difference.
[0222] The embodiments of the present application provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described above in the embodiments of the present application.
[0223] The embodiments of the present application provide a computer-readable storage medium storing executable instructions, where the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the method provided by the embodiments of the present application.
[0224] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0225] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0226] As an example, the executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).
[0227] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or, on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0228] As described above, the above are only embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. A data sensitivity recognition method based on a sensitivity recognition model, characterized in that, The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer, and the method includes: Obtain the metadata of the data to be recognized, where the metadata is used to describe the data to be recognized; Through the feature extraction layer, perform feature extraction on the metadata of the data to be recognized to obtain the data features of the metadata; when the metadata includes at least two keywords, the data features include the features of each keyword; Through the sensitivity recognition layer, based on the data features of the metadata, perform sensitivity recognition on the data to be recognized to obtain a sensitivity recognition result; wherein, the sensitivity recognition result is the data sensitivity corresponding to the sensitive word with the highest matching degree after respectively matching the features corresponding to each keyword with the features corresponding to at least two sensitive words through the sensitivity recognition layer; Among them, the sensitivity recognition result is used to indicate the data sensitivity corresponding to the data to be recognized.
2. The method according to claim 1, wherein The obtaining of the metadata of the data to be recognized includes: When the storage form of the data to be recognized is a data table, obtain at least one of the following table elements from the data table: the data table name, the table description corresponding to the data to be recognized in the data table, and the attribute field corresponding to the data to be recognized in the data table; Determine the obtained table element as the metadata of the data to be recognized.
3. The method according to claim 1, characterized in that, The obtaining of the metadata of the data to be recognized includes: When the storage form of the data to be recognized is a document, obtain at least one of the following document contents from the document: the document title, the document abstract, and the document keywords; Determine the obtained document content as the metadata of the data to be recognized.
4. The method according to claim 1, characterized in that The performing of feature extraction on the metadata of the data to be recognized to obtain the data features of the metadata includes: Perform word segmentation processing on the metadata of the data to be recognized to obtain multiple words corresponding to the metadata; Perform feature encoding on each of the words respectively to obtain the word features corresponding to each of the words; Perform feature splicing on the word features corresponding to each of the words to obtain the data features corresponding to the metadata.
5. The method according to claim 4, characterized in that The performing of feature splicing on the word features corresponding to each of the words to obtain the data features corresponding to the metadata includes: Perform bidirectional encoding processing on the word features of each word respectively to obtain the upstream encoding feature and the downstream encoding feature corresponding to each word; Perform feature splicing on the upstream encoding feature and the downstream encoding feature of each word respectively to obtain the corresponding spliced encoding feature; Perform feature splicing on the spliced encoding features corresponding to each of the words to obtain the data features corresponding to the metadata.
6. The method according to claim 1, wherein The performing of sensitivity recognition on the data to be recognized through the sensitivity recognition layer based on the data features of the metadata to obtain a sensitivity recognition result includes: Through the sensitivity recognition layer, perform classification prediction on the data features of the metadata corresponding to at least two sensitivity levels to obtain the probabilities of the metadata corresponding to each of the sensitivity levels; Select the sensitivity level with the highest probability as the sensitivity recognition result of the data to be recognized.
7. The method according to claim 1, wherein The method further includes: Establish an association relationship between the sensitivity recognition result and the data to be recognized, and store the association relationship; Among them, the association relationship is used to find the data sensitivity corresponding to the data to be recognized based on the data to be recognized.
8. The method according to claim 7, wherein Establishing the association relationship between the sensitivity recognition result and the data to be recognized includes: Storing the sensitivity recognition result in a target area associated with the data to be recognized, where the target area is the area corresponding to the data sensitivity in the storage area corresponding to the metadata.
9. The method according to claim 1, characterized in that, The method further includes: In response to a data display request for the data to be recognized, obtaining the data sensitivity corresponding to the data to be recognized; When the data sensitivity corresponding to the data to be recognized reaches a sensitivity threshold, returning shielding indication information corresponding to the data to be recognized; The shielding indication information is used to indicate shielding display of the data to be recognized.
10. The method according to claim 1, characterized in that, The method further includes: When the sensitivity recognition result indicates that the data sensitivity of the data to be recognized reaches a target data sensitivity, outputting encryption prompt information corresponding to the data to be recognized; Among them, the encryption prompt information is used to prompt encryption processing of the data to be recognized.
11. A training method for a sensitivity recognition model, characterized in that, The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer, and the method includes: Obtaining the metadata of a data sample, where the data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample; Through the feature extraction layer, performing feature extraction on the metadata of the data sample to obtain sample data features of the metadata of the data sample; when the metadata includes at least two keywords, the sample data features include the features of each keyword; Through the sensitivity recognition layer, based on the sample data features, performing sensitivity recognition on the data sample to obtain a sample sensitivity recognition result; among them, the sample sensitivity recognition result is the data sensitivity corresponding to the sensitive word with the highest matching degree after respectively matching the features corresponding to each keyword with the features corresponding to at least two sensitive words through the sensitivity recognition layer; Obtaining the difference between the sample sensitivity recognition result and the sensitivity label, and updating the model parameters of the sensitivity recognition model based on the difference; Among them, the sensitivity recognition model is used to output a sensitivity recognition result indicating the data sensitivity corresponding to the data to be recognized after inputting the metadata of the data to be recognized into the sensitivity recognition model.
12. A data sensitivity recognition device based on a sensitivity recognition model, characterized in that, The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer, and the device includes: A first acquisition module, configured to acquire the metadata of the data to be recognized, where the metadata is used to describe the data to be recognized; A first extraction module, configured to perform feature extraction on the metadata of the data to be recognized through the feature extraction layer to obtain data features of the metadata; when the metadata includes at least two keywords, the data features include the features of each keyword; A first recognition module, configured to perform sensitivity recognition on the data to be recognized based on the data characteristics of the metadata through the sensitivity recognition layer, so as to obtain a sensitivity recognition result; wherein, the sensitivity recognition result is the data sensitivity corresponding to the sensitive word with the highest matching degree after respectively matching the characteristics corresponding to each keyword with the characteristics corresponding to at least two sensitive words through the sensitivity recognition layer; Wherein, the sensitivity recognition result is used to indicate the data sensitivity corresponding to the data to be recognized.
13. A training device for a sensitivity recognition model, characterized in that, The sensitivity recognition model includes a feature extraction layer and a sensitivity recognition layer, and the device includes: A second acquisition module, configured to acquire the metadata of a data sample, where the data sample carries a sensitivity label, and the sensitivity label is used to indicate the data sensitivity corresponding to the data sample; A second extraction module, configured to perform feature extraction on the metadata of the data sample through the feature extraction layer to obtain the sample data characteristics of the metadata of the data sample; when the metadata includes at least two keywords, the sample data characteristics include the characteristics of each keyword; A second recognition module, configured to perform sensitivity recognition on the data sample based on the sample data characteristics through the sensitivity recognition layer, so as to obtain a sample sensitivity recognition result; wherein, the sample sensitivity recognition result is the data sensitivity corresponding to the sensitive word with the highest matching degree after respectively matching the characteristics corresponding to each keyword with the characteristics corresponding to at least two sensitive words through the sensitivity recognition layer; An update module, configured to obtain the difference between the sample sensitivity recognition result and the sensitivity label carried by the data sample, and update the model parameters of the sensitivity recognition model based on the difference; Wherein, the sensitivity recognition model is configured to output a sensitivity recognition result indicating the data sensitivity corresponding to the data to be recognized after inputting the metadata of the data to be recognized into the sensitivity recognition model.
14. A computer-readable storage medium, characterized in that, Stores executable instructions, which are used to implement the method according to any one of claims 1 to 11 when executed by a processor.
15. An electronic device, characterized in that, Including: A memory, configured to store executable instructions; A processor, configured to implement the method according to claims 1 to 11 when executing the executable instructions stored in the memory.
16. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Identification method and device for sensitive content
CN107818077A