A method, device, medium, and program product for marking sensitive data
By sampling and collecting database log data and identifying classification and grading rules, sensitive data are automatically marked, which solves the problem of sensitive data labeling lacking intelligent process in the existing technology, and realizes intelligent sensitive data identification and storage.
Patent Information
- Application Number
- CN202211295439.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-10-21
AI Technical Summary
In the prior art, the labeling of sensitive data lacks intelligent process and cannot effectively identify and label sensitive data in the database.
By sampling and collecting database log data, classification and grading rules are used to identify sensitive classification and level information, and automatically label sensitive marks to achieve intelligent classification and grading.
It realizes intelligent classification and grading and automatic labeling of database log data, improves the efficiency and accuracy of sensitive data identification, and supports flexible selection of classification and grading rules.
Smart Images

Figure CN115659396B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and in particular, to a technology for marking sensitive data. Background Art
[0002] Sensitive data refers to data that may cause serious harm to society or individuals after leakage. It includes personal privacy data such as name, ID number, address, phone number, bank account number, email, password, medical information, educational background, etc. In the prior art, the industry's solution to sensitive data is in the form of manual marking, and there is no intelligent and process-based form. Summary of the Invention
[0003] An object of this application is to provide a method, device, medium, and program product for marking sensitive data.
[0004] According to one aspect of this application, a method for marking sensitive data is provided. The method includes:
[0005] Sampling and collecting the log data of the database to obtain one or more sampling data;
[0006] Identifying the sampling data according to the classification and grading rules, determining the sensitive classification corresponding to the sampling data, and obtaining the sensitive level information corresponding to the sampling data under the sensitive classification;
[0007] According to the sensitive level information, sensitive marks are added to at least one sampling data, and the at least one sampling data is stored, where the sensitive marks include the sensitive classification and sensitive level information corresponding to each sampling data.
[0008] According to one aspect of this application, a computer device for marking sensitive data is provided. The device includes:
[0009] A first module for sampling and collecting the log data of the database to obtain one or more sampling data;
[0010] A second module for identifying the sampling data according to the classification and grading rules, determining the sensitive classification corresponding to the sampling data, and obtaining the sensitive level information corresponding to the sampling data under the sensitive classification;
[0011] A third module for adding sensitive marks to at least one sampling data according to the sensitive level information and storing the at least one sampling data, where the sensitive marks include the sensitive classification and sensitive level information corresponding to each sampling data.
[0012] According to one aspect of the present application, there is provided a computer device for marking sensitive data, including a memory, a processor, and a computer program stored on the memory. Wherein, the processor executes the computer program to implement the operations of any of the above-mentioned methods.
[0013] According to one aspect of the present application, there is provided a computer-readable storage medium with a computer program stored thereon, characterized in that when the computer program is executed by a processor, it implements the operations of any of the above-mentioned methods.
[0014] According to one aspect of the present application, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, it implements the steps of any of the above-mentioned methods.
[0015] Compared with the prior art, the present application obtains one or more sampled data by sampling and collecting the log data of the database; identifies the sampled data according to the classification and grading rules, determines the sensitive classification corresponding to the sampled data, and obtains the sensitive level information corresponding to the sampled data under the sensitive classification; according to the sensitive level information, sensitive marks are added to at least one sampled data, and the at least one sampled data is stored, wherein the sensitive marks include the sensitive classification and sensitive level information corresponding to each sampled data, so that the log data of the database can be intelligently classified and graded through the preset classification and grading rules, and the corresponding sensitive marks are automatically added, and the classification and grading rules used can be flexibly selected, thus realizing an intelligent and process-based form. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:
[0017] Figure 1 A flowchart of a method for marking sensitive data according to an embodiment of the present application is shown;
[0018] Figure 2 A system diagram of an intelligent data classification and grading system for data security according to an embodiment of the present application is shown;
[0019] Figure 3 A structural diagram of a computer device for marking sensitive data according to an embodiment of the present application is shown;
[0020] Figure 4 An exemplary system that can be used to implement the various embodiments described in the present application is shown.
[0021] Identical or similar reference numerals in the drawings represent identical or similar components. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present application will be further described in detail below with reference to the accompanying drawings.
[0023] In a typical configuration of the present application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a Central Processing Unit (CPU)), an input / output interface, a network interface, and a memory.
[0024] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory. The memory is an example of a computer-readable medium.
[0025] The computer-readable medium includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0026] The devices referred to in this application include, but are not limited to, terminals, network devices, or devices formed by integrating terminals and network devices through a network. The terminal includes, but is not limited to, any mobile electronic product that can perform human-computer interaction with users (such as through a touchpad), such as a smart phone, a tablet computer, etc. The mobile electronic product can adopt any operating system, such as the Android operating system, the iOS operating system, etc. Among them, the network device includes an electronic device that can automatically perform numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an Application Specific Integrated Circuit (ASIC), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Digital Signal Processor (DSP), an embedded device, etc. The network device includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud composed of multiple servers; here, the cloud is composed of a large number of computers or network servers based on Cloud Computing. Among them, Cloud Computing is a type of distributed computing, which consists of a virtual supercomputer composed of a group of loosely coupled computers. The network includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, a wireless ad hoc network (Ad Hoc network), etc. Preferably, the device can also be a program running on the terminal, the network device, or a device formed by integrating the terminal and the network device, the network device, the touch terminal, or the network device and the touch terminal through a network.
[0027] Of course, those skilled in the art should understand that the above devices are only examples, and other existing or future devices that are applicable to this application should also be included within the protection scope of this application and are hereby incorporated by reference.
[0028] In the description of this application, "a plurality of" means two or more, unless otherwise specifically defined.
[0029] Figure 1A flowchart of a method for marking sensitive data according to an embodiment of the present application is shown. The method includes step S11, step S12, and step S13. In step S11, the computer device samples and collects the log data of the database to obtain one or more sampled data. In step S12, the computer device identifies the sampled data according to the classification and grading rules, determines the sensitive classification corresponding to the sampled data, and obtains the sensitive level information corresponding to the sampled data under the sensitive classification. In step S13, the computer device marks at least one sampled data as sensitive according to the sensitive level information and stores the at least one sampled data, where the sensitive mark includes the sensitive classification and sensitive level information corresponding to each sampled data.
[0030] In step S11, the computer device samples and collects the log data of the database to obtain one or more sampled data. In some embodiments, the log data refers to the operation behavior log data of the database, and the log data includes, but is not limited to, the behavior time, behavior content (for example, the stored data read, inserted, modified, or deleted in the database), behavior result (for example, whether it is successful), behavior object (that is, at least one stored data in the database), etc. of a certain operation behavior on the database. In some embodiments, the log data of the database can be sampled and collected according to a predetermined sampling rate, and the sampling method can be interval sampling in the order of the behavior time. For example, if the sampling rate is 10%, then for every 10 operation behaviors that have occurred in the database, the log data of 1 operation behavior is collected. That is, if the log data of operation behavior 1 has been collected, the next collected log data is that of operation behavior 2, which occurs tenth in the order of behavior time after operation behavior 1. In some embodiments, the sampling method can also be unordered random sampling. For example, first collect multiple log data of all operation behaviors on the database. If the sampling rate is 10%, then according to the number Num of the multiple log data, randomly select Num * 10% of the log data as the sampled data.
[0031] In step S12, the computer device identifies the sampled data according to the classification and grading rules, determines the sensitive classification corresponding to the sampled data, and obtains the sensitive level information corresponding to the sampled data under the sensitive classification. In some embodiments, the classification and grading rules include multiple predefined sensitive classifications (for example, personal identity information category, personal property information category) and the identification strategies for each sensitive classification. The identification strategy is used to identify whether the data includes the sensitive classification corresponding to the sensitive classification. The specific identification methods include, but are not limited to, regular expression identification, keyword identification, model feature identification, etc. The classification and grading rules also include a grading strategy. The grading strategy is used to determine the sensitive level corresponding to the data under the sensitive classification when it is identified that the data includes the sensitive classification corresponding to the identification strategy. The sensitive level can be represented in numerical form. For example, the larger the value, the more sensitive or less secure the corresponding data is. Or, the sensitive level can also be represented in text form, such as "slightly sensitive", "moderately sensitive", "highly sensitive", etc. If it is identified according to the classification and grading rules that the data includes the sensitive classification corresponding to a certain identification strategy, then it is used as the sensitive classification corresponding to the data, and then the sensitive level corresponding to the data under the sensitive classification is determined according to the classification and grading rules. In some embodiments, one sensitive classification may only correspond to one sensitive level. If it is identified that the data includes the sensitive classification corresponding to a certain classification strategy, then the sensitive level corresponding to the sensitive classification is directly used as the sensitive level corresponding to the data under the sensitive classification. Or, one sensitive classification may also correspond to multiple different sensitive levels. In this case, the grading strategy needs to be used to determine the sensitive level corresponding to the data under the sensitive classification. The specific determination methods include, but are not limited to, semantic analysis method, keyword extraction method, model feature method, etc. For example, the sensitive level corresponding to the data under the sensitive classification is determined by semantic analysis. Or, a mapping relationship between a number of predefined keywords and sensitive levels is set for the sensitive classification. The sensitive level corresponding to the data under the sensitive classification is determined according to the sensitive level mapped by the keywords included in the data. Or, the data is input into the sensitive level model corresponding to the sensitive classification that has been trained, and the sensitive level corresponding to the data under the sensitive classification output by the sensitive level model is obtained.
[0032] In step S13, the computer device marks at least one sampled data with a sensitivity label according to the sensitivity level information, and stores the at least one sampled data, where the sensitivity label includes the sensitivity classification and sensitivity level information corresponding to each sampled data. In some embodiments, the sampled data whose sensitivity level meets a predetermined level threshold may be determined as sensitive data according to the sensitivity level corresponding to the sampled data. For example, the sampled data whose corresponding sensitivity level is greater than or equal to a predetermined level numerical threshold may be determined as sensitive data. Alternatively, the sampled data that all identify at least one sensitivity classification in the classification and grading rules may also be determined as sensitive data. In some embodiments, a corresponding sensitivity label is marked for the sensitive data (i.e., at least one sampled data), and the sensitivity label includes the sensitivity classification and sensitivity level corresponding to the sensitive data, and the sensitive data is stored. In some embodiments, the user may retrieve the required data from the stored sensitive data, or may also perform asset visualization analysis on the stored sensitive data and export it through a report. In some embodiments, external calling of the stored sensitive data is also supported and / or sending the stored sensitive data out is supported. This application can intelligently classify and grade the log data of the database through preset classification and grading rules, automatically mark the corresponding sensitivity label, and can flexibly select the classification and grading rules used, so as to achieve an intelligent and process-based form.
[0033] In some embodiments, step S11 includes: the computer device samples and collects the log data of the database through a plug-in deployed on the database to obtain one or more sampled data. In some embodiments, the log data of the database can be automatically sampled and collected by deploying (or installing) a plug-in on the database. This plug-in belongs to a semi-intrusive software application, with low cost, good convenience, and strong scalability. By directly obtaining the database log data through this plug-in, more accurate results can be achieved.
[0034] In some embodiments, the method further includes: the computer device determines the classification and grading rules corresponding to the database. In some embodiments, it is necessary to first determine the classification and grading rules used for the database. The specific determination method may be to obtain the classification and grading rules corresponding to the database selected by the user from multiple default classification and grading rules, or may also be to perform semantic analysis on the stored data in the database to determine the characteristics of the stored data corresponding to the database, and automatically determine the classification and grading rules that match the characteristics of the stored data according to the characteristics of the stored data.
[0035] In some embodiments, determining the classification and grading rules corresponding to the database includes: obtaining the classification and grading rules corresponding to the database selected by the user from multiple default classification and grading rules. In some embodiments, multiple default classification and grading rules are preset, and the user can freely and flexibly select one from the multiple default classification and grading rules as the classification and grading rules corresponding to the database. For example, the classification and grading rules may be rules classified according to personal information protection, which include multiple sensitive classifications such as personal identity information and personal property information. Another example is that the classification and grading rules may also be rules classified according to the telecommunications operator industry, which include multiple sensitive classifications such as user basic information, location data, and consumption information.
[0036] In some embodiments, determining the classification and grading rules corresponding to the database includes: determining the characteristics of the stored data corresponding to the database by performing semantic analysis on the stored data in the database; and determining the classification and grading rules that match the characteristics of the stored data according to the characteristics of the stored data. In some embodiments, the characteristics of the stored data in the database can be obtained by first performing semantic analysis on the stored data in the database. The characteristics of the stored data are used to characterize the types and characteristics of the data mainly stored in the database. Then, according to the characteristics of the stored data, the classification and grading rules suitable for this type and characteristic of stored data are automatically determined from multiple default classification and grading rules and used as the classification and grading rules corresponding to the database. For example, if the characteristics of the stored data indicate that the database mainly stores message text data of the conversation type, the classification and grading rules suitable for this type of message text data of the conversation type may be rules classified according to social information inclusion, which include multiple sensitive classifications such as personal chat information, personal space post information, and personal profile information.
[0037] In some embodiments, determining the classification and grading rules that match the characteristics of the stored data according to the characteristics of the stored data includes: determining the sensitive scenario information corresponding to the database according to the characteristics of the stored data; and determining the classification and grading rules that match the sensitive information according to the sensitive scenario information. In some embodiments, the sensitive scenario information involved in the database can be determined first according to the characteristics of the stored data, and then, according to the sensitive scenario information, the classification and grading rules suitable for the sensitive scenario information are automatically determined from multiple default classification and grading rules. For example, if the characteristics of the stored data indicate that the database mainly stores data such as product links, product prices, and delivery addresses, it can be determined that the sensitive scenario involved in the database is the shopping scenario, and the classification and grading rules suitable for this shopping scenario are automatically determined from multiple default classification and grading rules, which include multiple sensitive classifications such as personal consumption information, personal payment information, personal contact information, and personal address information.
[0038] In some embodiments, the classification and grading rules include a classification strategy and a sensitivity level strategy; wherein, step S12 includes: the computer device identifies the sampled data according to the classification strategy to determine the sensitive classification corresponding to the sampled data; according to the sensitivity level strategy, determines the sensitivity level information corresponding to the sampled data under the sensitive classification. In some embodiments, the classification strategy usually includes a plurality of preset sensitive classifications. The classification strategy is used to identify whether the sampled data includes one or several of the sensitive classifications in the classification strategy. The specific identification methods include, but are not limited to, regular expression identification, keyword identification, model feature identification, etc. In some embodiments, a sampled data may correspond to multiple sensitive classifications, that is, if it is identified according to the classification strategy that the sampled data includes a certain sensitive classification, it can be directly used as one of the sensitive classifications corresponding to the sampled data. In some embodiments, a sampled data can only correspond to one sensitive classification. If it is identified according to the classification and grading rules that the sampled data includes multiple different sensitive classifications, one of the multiple different sensitive classifications needs to be determined as the sensitive classification corresponding to the sampled data. For example, each sensitive classification corresponds to a different recognition credibility or matching degree, and the sensitive classification with the highest recognition credibility or matching degree among the multiple different sensitive classifications can be used as the sensitive classification corresponding to the sampled data. In some embodiments, after determining the sensitive classification corresponding to the sampled data, further according to the sensitivity level strategy, determines the sensitivity level corresponding to the sampled data under the sensitive classification. The sensitivity level can be represented in numerical form. For example, the larger the value, the more sensitive or less secure the corresponding data is. Or, the sensitivity level can also be represented in text form, such as "slightly sensitive", "moderately sensitive", "highly sensitive", etc. In some embodiments, the sensitivity level strategy may include the sensitivity levels corresponding to each sensitive classification respectively, and the sensitivity level corresponding to the sensitive classification to which the sampled data corresponds can be directly used as the sensitivity level corresponding to the sampled data under the sensitive classification. In some embodiments, the sensitivity level strategy may include the specific determination methods for the sensitivity levels corresponding to each sensitive classification. The specific determination methods include, but are not limited to, semantic analysis methods, keyword extraction methods, model feature methods, etc. After determining the sensitive classification corresponding to the sampled data, it is necessary to specifically determine the sensitivity level corresponding to the sampled data under the sensitive classification according to the sensitivity level determination method corresponding to the sensitive classification. For example, determine the sensitivity level corresponding to the data under the sensitive classification through semantic analysis, or the sensitive classification presets a mapping relationship between several keywords and the sensitivity level, and determine the sensitivity level corresponding to the data under the sensitive classification according to the sensitivity level mapped by the keywords included in the data, or input the data into the sensitivity level model corresponding to the trained sensitive classification to obtain the sensitivity level corresponding to the data under the sensitive classification output by the sensitivity level model.In some embodiments, if the sampled data corresponds to multiple sensitive classifications, it is necessary to first determine the sensitive level corresponding to the sampled data under each sensitive classification respectively, and then determine the sensitive level corresponding to the sampled data under the multiple sensitive classifications according to the multiple sensitive levels. For example, if the sensitive level is in numerical form, the average value corresponding to the multiple sensitive level values can be used as the sensitive level corresponding to the sampled data under the multiple sensitive classifications.
[0039] In some embodiments, the sensitive level policy includes first sensitive level information corresponding to each sensitive classification under the classification policy; wherein, determining the sensitive level information corresponding to the sampled data under the sensitive classification according to the sensitive level policy includes: obtaining the first sensitive level information corresponding to the sensitive classification according to the sensitive level policy; and determining the sensitive level information corresponding to the sampled data under the sensitive classification according to the first sensitive level information. In some embodiments, each sensitive classification corresponds to a sensitive level, and the sensitive level policy includes the first sensitive level corresponding to each sensitive classification under the classification policy. In some embodiments, after determining the sensitive classification corresponding to the sampled data, the first sensitive level corresponding to the sensitive classification can be directly used as the sensitive level corresponding to the sampled data under the sensitive classification, or the first sensitive level corresponding to the sensitive classification can be input into a predetermined functional relationship, and the output of the functional relationship can be used as the sensitive level corresponding to the sampled data under the sensitive classification.
[0040] In some embodiments, the sensitivity level policy further includes a plurality of sub-sensitivity classifications corresponding to each sensitivity classification under the classification policy and second sensitivity level information corresponding to each sub-sensitivity classification; wherein, based on the first sensitivity level information, determining the sensitivity level information corresponding to the sampled data under the sensitivity classification includes: determining the sub-sensitivity classification corresponding to the sampled data under the sensitivity classification according to the sensitivity level policy; and determining the sensitivity level information corresponding to the sampled data under the sensitivity classification according to the first sensitivity level information and the second sensitivity level information corresponding to the sub-sensitivity classification. In some embodiments, each sensitivity classification further corresponds to a plurality of sub-sensitivity classifications, each sub-sensitivity classification corresponds to a sensitivity level, and the sensitivity level policy further includes the plurality of sub-sensitivity classifications corresponding to each sensitivity classification under the classification policy and the second sensitivity level corresponding to each sub-sensitivity classification. In some embodiments, after determining the sensitivity classification corresponding to the sampled data, it is necessary to determine the sub-sensitivity classification corresponding to the sampled data under the sensitivity classification according to the sensitivity level policy. The specific determination methods include, but are not limited to, regular expression recognition, keyword recognition, semantic analysis, model feature recognition, etc. Through the sensitivity level policy, it can be identified which specific sub-sensitivity classification the sampled data belongs to under the sensitivity classification. Then, based on the first sensitivity level corresponding to the sensitivity classification and the second sensitivity level corresponding to the sub-sensitivity classification to which the sampled data belongs, the sensitivity level corresponding to the sampled data under the sensitivity classification is determined. For example, if the sensitivity level is in numerical form, the average value, maximum value, or minimum value of the first sensitivity level and the second sensitivity level can be used as the sensitivity level corresponding to the sampled data under the sensitivity classification. Alternatively, the first sensitivity level and the second sensitivity level can also be input into a predetermined functional relationship, and then the output of the functional relationship is used as the sensitivity level corresponding to the sampled data under the sensitivity classification.
[0041] In some embodiments, the sensitivity level policy includes multiple sub-sensitivity classifications corresponding to each sensitive classification under the classification policy and third sensitivity level information corresponding to each sub-sensitivity classification; wherein, determining the sensitivity level information corresponding to the sampled data under the sensitive classification according to the sensitivity level policy includes: determining the sub-sensitivity classification corresponding to the sampled data under the sensitive classification according to the sensitivity level policy; and determining the sensitivity level information corresponding to the sampled data under the sensitive classification according to the third sensitivity level information corresponding to the sub-sensitivity classification. In some embodiments, the third sensitivity level is the same as or similar to the second sensitivity level described above, and will not be elaborated here. In some embodiments, after determining the sensitive classification corresponding to the sampled data, the sub-sensitivity classification corresponding to the sampled data under the sensitive classification is determined according to the sensitivity level policy, and then the sensitivity level corresponding to the sampled data under the sensitive classification can be determined only according to the third sensitivity level corresponding to the sub-sensitivity classification. For example, the third sensitivity level corresponding to the sub-sensitivity classification is used as the sensitivity level corresponding to the sampled data under the sensitive classification.
[0042] In some embodiments, the method further includes: obtaining a review result of the user on the stored target sampled data; if at least one of the sensitive classification and the sensitivity level information corresponding to the target sampled data is inconsistent with the review result, adjusting the sensitive mark of the target sampled data according to the review result. In some embodiments, all sensitive data (i.e., at least one sampled data) can be presented to the user for review, or only sensitive data with an identification credibility or matching degree lower than or equal to a predetermined threshold can be presented to the user for review, where the review means that the user manually checks whether the sensitive mark assigned to the sensitive data is accurate. In some embodiments, if the review result of the user on the target sensitive data (i.e., the target sampled data) indicates that the sensitive mark assigned to the target sensitive data is inaccurate, that is, if at least one of the sensitive classification and the sensitivity level corresponding to the target sensitive data is inconsistent with the review result, then the sensitive mark of the target sensitive data needs to be adjusted according to the review result. The specific adjustment methods include, but are not limited to, modifying the sensitive classification in the sensitive mark, modifying the sensitivity level information in the sensitive mark, removing the sensitive mark for the target sensitive data, and canceling the storage of the target sensitive data.
[0043] In some embodiments, adjusting the sensitive label of the target sampling data includes at least one of the following: modifying the sensitive classification in the sensitive label; modifying the sensitive level information in the sensitive label; removing the sensitive label from the target sampling data and canceling the storage of the target sampling data. In some embodiments, if the re-verified sensitive level of the target sensitive data does not meet a predetermined level threshold, for example, the re-verified sensitive level is less than a predetermined level numerical threshold, then the sensitive label attached to the target sensitive data is deleted, and the storage of the target sensitive data is canceled. In some embodiments, if the verification result indicates that the target sensitive data does not include any of the predetermined sensitive classifications in the classification and grading rule, then the sensitive label attached to the target sensitive data is deleted, and the storage of the target sensitive data is canceled.
[0044] In some embodiments, the method further includes: the computer device adjusts the sampling rate of the database according to the sensitive level information, so that the adjusted sampling rate is subsequently used to sample and collect the new log data of the database. In some embodiments, the ratio of the number of target sampling data (for example, the target sampling data may be the at least one sampled data with a sensitive label, i.e., sensitive data) whose sensitive level meets a predetermined level threshold to the total number of one or more sampled data sampled and collected can be calculated according to the sensitive level of the sampled data. For example, the ratio of the number of sampled data whose sensitive level is greater than or equal to a predetermined level numerical threshold to the total number of one or more sampled data sampled and collected is calculated. If the ratio value is greater than or equal to a predetermined first ratio threshold, the original sampling rate of the database can be increased, the original sampling rate can be increased by a default amplitude, or the increase amplitude of the original sampling rate can also be dynamically determined according to the size of the ratio value. If the ratio value is less than or equal to a predetermined second ratio threshold, the original sampling rate of the database can be decreased, the original sampling rate can be decreased by a default amplitude, or the decrease amplitude of the original sampling rate can also be dynamically determined according to the size of the ratio value. Then, when the database needs to be sampled and collected again subsequently, the adjusted sampling rate is used to sample and collect the new log data newly generated by the database again.
[0045] Figure 2 A flowchart of a method for marking sensitive data according to an embodiment of the present application is shown.
[0046] As Figure 2As shown, the intelligent data classification and grading system includes an audit log and database asset collection module, a data classification and grading engine, an intelligent policy configuration module, a manual review module, a storage module, and an asset retrieval and analysis module. The audit log and database asset collection module is used to sample and collect the log data of the database in a plug-in form. The intelligent policy configuration module is used to intelligently configure the classification and grading rules used by the data classification and grading engine. The data classification and grading engine is used to identify sensitive data from the log data using the classification and grading rules. The manual review module is used to manually review the identification results of the log data. The storage module is used to store the identified data. The asset retrieval and analysis module is used to retrieve the identified data by calling the stored assets and perform visual analysis on the identified data.
[0047] Figure 3 Shows a computer device structure diagram for marking sensitive data according to an embodiment of the present application. The device includes a module 11, a module 12, and a module 13. The module 11 is used to sample and collect the log data of the database to obtain one or more sampled data. The module 12 is used to identify the sampled data according to the classification and grading rules, determine the sensitive classification corresponding to the sampled data, and obtain the sensitive level information corresponding to the sampled data under the sensitive classification. The module 13 is used to mark at least one sampled data with a sensitive mark according to the sensitive level information and store the at least one sampled data, where the sensitive mark includes the sensitive classification and sensitive level information corresponding to each sampled data.
[0048] A module 11 for sampling and collecting log data of a database to obtain one or more sampled data. In some embodiments, the log data refers to the operation behavior log data of the database, and the log data includes but is not limited to the behavior time, behavior content (for example, the stored data read, inserted, modified, or deleted in the database), behavior result (for example, whether it is successful), behavior object (i.e., at least one stored data in the database), etc. of a certain operation behavior on the database. In some embodiments, the log data of the database can be sampled and collected according to a predetermined sampling rate, and the sampling method can be interval sampling in the order of the behavior time. For example, if the sampling rate is 10%, then for every 10 operation behaviors that have occurred in the database, the log data of 1 operation behavior is collected. That is, if the log data of operation behavior 1 has been collected, the next log data to be collected is the log data of operation behavior 2 that occurs tenth after operation behavior 1 in the order of the behavior time. In some embodiments, the sampling method can also be unordered random sampling. For example, first collect multiple log data of all operation behaviors on the database. If the sampling rate is 10%, then according to the number Num of the multiple log data, randomly select Num * 10% of the log data as the sampled data.
[0049] Module 12 for the first and second parts is used to identify the sampled data according to the classification and grading rules, determine the sensitive classification corresponding to the sampled data, and obtain the sensitive level information corresponding to the sampled data under the sensitive classification. In some embodiments, the classification and grading rules include multiple predefined sensitive classifications (e.g., personal identity information category, personal property information category) and the identification strategies for each sensitive classification. The identification strategy is used to identify whether the data includes the sensitive classification corresponding to the sensitive classification. The specific identification methods include, but are not limited to, regular expression identification, keyword identification, model feature identification, etc. The classification and grading rules also include a grading strategy, which is used to determine the sensitive level corresponding to the data under the sensitive classification when it is identified that the data includes the sensitive classification corresponding to the identification strategy. The sensitive level can be represented in numerical form. For example, the larger the value, the more sensitive or less secure the corresponding data is. Or, the sensitive level can also be represented in text form, such as "slightly sensitive", "moderately sensitive", "highly sensitive", etc. If it is identified according to the classification and grading rules that the data includes the sensitive classification corresponding to a certain identification strategy, then it is used as the sensitive classification corresponding to the data, and then the sensitive level corresponding to the data under the sensitive classification is determined according to the classification and grading rules. In some embodiments, one sensitive classification may only correspond to one sensitive level. If it is identified that the data includes the sensitive classification corresponding to a certain classification strategy, then the sensitive level corresponding to the sensitive classification is directly used as the sensitive level corresponding to the data under the sensitive classification. Or, one sensitive classification may also correspond to multiple different sensitive levels. In this case, the grading strategy is needed to determine the sensitive level corresponding to the data under the sensitive classification. The specific determination methods include, but are not limited to, semantic analysis method, keyword extraction method, model feature method, etc. For example, the sensitive level corresponding to the data under the sensitive classification is determined by semantic analysis. Or, a mapping relationship between several predefined keywords and the sensitive level is set for the sensitive classification. The sensitive level corresponding to the data under the sensitive classification is determined according to the sensitive level mapped by the keywords included in the data. Or, the data is input into the sensitive level model corresponding to the sensitive classification that has been trained, and the sensitive level corresponding to the data under the sensitive classification output by the sensitive level model is obtained.
[0050] A 13-module 13 is used to mark at least one sampled data with a sensitivity label according to the sensitivity level information and store the at least one sampled data, where the sensitivity label includes the sensitivity classification and sensitivity level information corresponding to each sampled data. In some embodiments, the sampled data with a sensitivity level meeting a predetermined level threshold may be determined as sensitive data according to the sensitivity level corresponding to the sampled data. For example, the sampled data with a corresponding sensitivity level greater than or equal to a predetermined level value threshold may be determined as sensitive data, or alternatively, all sampled data identified as including at least one sensitivity classification in the classification and grading rules may also be determined as sensitive data. In some embodiments, corresponding sensitivity labels are marked on the sensitive data (i.e., at least one sampled data), and the sensitivity label includes the sensitivity classification and sensitivity level corresponding to the sensitive data, and the sensitive data is stored. In some embodiments, a user may retrieve required data from the stored sensitive data, or alternatively, perform asset visualization analysis on the stored sensitive data and export it through a report. In some embodiments, external calls to the stored sensitive data are also supported and / or sending the stored sensitive data out is supported. This application can intelligently classify and grade the log data of the database through preset classification and grading rules, automatically mark corresponding sensitivity labels, and can flexibly select the classification and grading rules used, thereby realizing an intelligent and process-based form.
[0051] In some embodiments, the 11-module 11 is used to: sample and collect the log data of the database through a plug-in deployed on the database to obtain one or more sampled data. Here, the related operations are the same as or similar to those in Figure 1 the embodiments shown, so they will not be elaborated here and are included herein by reference.
[0052] In some embodiments, the device is further used to: determine the classification and grading rules corresponding to the database. Here, the related operations are the same as or similar to those in Figure 1 the embodiments shown, so they will not be elaborated here and are included herein by reference.
[0053] In some embodiments, the determining the classification and grading rules corresponding to the database includes: obtaining the classification and grading rules corresponding to the database selected by the user from multiple default classification and grading rules. Here, the related operations are the same as or similar to those in Figure 1 the embodiments shown, so they will not be elaborated here and are included herein by reference.
[0054] In some embodiments, the determining the classification and grading rules corresponding to the database includes: performing semantic analysis on the stored data in the database to determine the characteristics of the stored data corresponding to the database; and determining the classification and grading rules matching the characteristics of the stored data according to the characteristics of the stored data. Here, the related operations are the same asFigure 1 is the same as or similar to the illustrated embodiment, and thus will not be described in detail again, and is hereby incorporated by reference herein.
[0055] In some embodiments, determining a classification and grading rule that matches the stored data characteristics according to the stored data characteristics includes: determining, according to the stored data characteristics, sensitive scenario information corresponding to the database; and determining, according to the sensitive scenario information, a classification and grading rule that matches the sensitive information. Here, the related operations are Figure 1 is the same as or similar to the illustrated embodiment, and thus will not be described in detail again, and is hereby incorporated by reference herein.
[0056] In some embodiments, the classification and grading rule includes a classification policy and a sensitive level policy; wherein, the first and second modules 12 are configured to: identify the sampled data according to the classification policy to determine a sensitive classification corresponding to the sampled data; and determine, according to the sensitive level policy, sensitive level information corresponding to the sampled data under the sensitive classification. Here, the related operations are Figure 1 is the same as or similar to the illustrated embodiment, and thus will not be described in detail again, and is hereby incorporated by reference herein.
[0057] In some embodiments, the sensitive level policy includes first sensitive level information corresponding to each sensitive classification under the classification policy; wherein, determining, according to the sensitive level policy, sensitive level information corresponding to the sampled data under the sensitive classification includes: obtaining, according to the sensitive level policy, the first sensitive level information corresponding to the sensitive classification; and determining, according to the first sensitive level information, sensitive level information corresponding to the sampled data under the sensitive classification. Here, the related operations are Figure 1 is the same as or similar to the illustrated embodiment, and thus will not be described in detail again, and is hereby incorporated by reference herein.
[0058] In some embodiments, the sensitive level policy further includes a plurality of sub-sensitive classifications corresponding to each sensitive classification under the classification policy and second sensitive level information corresponding to each sub-sensitive classification; wherein, determining, according to the first sensitive level information, sensitive level information corresponding to the sampled data under the sensitive classification includes: determining, according to the sensitive level policy, a sub-sensitive classification corresponding to the sampled data under the sensitive classification; and determining, according to the first sensitive level information and the second sensitive level information corresponding to the sub-sensitive classification, sensitive level information corresponding to the sampled data under the sensitive classification. Here, the related operations are Figure 1 is the same as or similar to the illustrated embodiment, and thus will not be described in detail again, and is hereby incorporated by reference herein.
[0059] In some embodiments, the sensitivity level policy includes multiple sub-sensitivity classifications corresponding to each sensitive classification under the classification policy and third sensitivity level information corresponding to each sub-sensitivity classification; wherein, determining the sensitivity level information corresponding to the sampled data under the sensitive classification according to the sensitivity level policy includes: determining the sub-sensitivity classification corresponding to the sampled data under the sensitive classification according to the sensitivity level policy; and determining the sensitivity level information corresponding to the sampled data under the sensitive classification according to the third sensitivity level information corresponding to the sub-sensitivity classification. Here, the related operations are the same as or similar to those in Figure 1 the embodiments shown, so they will not be elaborated here and are incorporated herein by reference.
[0060] In some embodiments, the device is further configured to: obtain a review result of the user on the stored target sampled data; if at least one of the sensitive classification and the sensitivity level information corresponding to the target sampled data is inconsistent with the review result, adjust the sensitive mark of the target sampled data according to the review result. Here, the related operations are the same as or similar to those in Figure 1 the embodiments shown, so they will not be elaborated here and are incorporated herein by reference.
[0061] In some embodiments, adjusting the sensitive mark of the target sampled data includes at least one of the following: modifying the sensitive classification in the sensitive mark; modifying the sensitivity level information in the sensitive mark; removing the sensitive mark for the target sampled data and canceling the storage of the target sampled data. Here, the related operations are the same as or similar to those in Figure 1 the embodiments shown, so they will not be elaborated here and are incorporated herein by reference.
[0062] In some embodiments, the device is further configured to: adjust the sampling rate of the database according to the sensitivity level information, so that the new log data of the database is sampled and collected using the adjusted sampling rate subsequently. Here, the related operations are the same as or similar to those in Figure 1 the embodiments shown, so they will not be elaborated here and are incorporated herein by reference.
[0063] In addition to the methods and devices described in the above embodiments, the present application also provides a computer-readable storage medium storing computer code, and when the computer code is executed, the method as described in any one of the previous items is executed.
[0064] The present application also provides a computer program product, and when the computer program product is executed by a computer device, the method as described in any one of the previous items is executed.
[0065] The present application also provides a computer device, which includes:
[0066] One or more processors;
[0067] A memory for storing one or more computer programs;
[0068] When the one or more computer programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of the preceding items.
[0069] Figure 4 An exemplary system that can be used to implement the various embodiments described in the present application is shown;
[0070] As Figure 4 As shown, in some embodiments, the system 300 can act as any one of the devices in the various embodiments. In some embodiments, the system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., (one or more) processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement modules to perform the actions described in the present application.
[0071] For one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the (one or more) processors 305 and / or any suitable device or component communicating with the system control module 310.
[0072] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.
[0073] The system memory 315 may be used, for example, to load and store data and / or instructions for the system 300. For one embodiment, the system memory 315 may include any suitable volatile memory, e.g., suitable DRAM. In some embodiments, the system memory 315 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0074] For one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide an interface to the NVM / storage device 320 and the (one or more) communication interfaces 325.
[0075] For example, the NVM / storage device 320 can be used to store data and / or instructions. The NVM / storage device 320 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more compact discs (CDs) drives, and / or one or more digital versatile discs (DVDs) drives).
[0076] The NVM / storage device 320 can include storage resources that are physically part of a device on which the system 300 is installed, or it can be accessed by the device without being part of the device. For example, the NVM / storage device 320 can be accessed via a network through the (one or more) communication interfaces 325.
[0077] The (one or more) communication interfaces 325 can provide an interface for the system 300 to communicate through one or more networks and / or with any other suitable device. The system 300 can wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols.
[0078] For one embodiment, at least one of the (one or more) processors 305 can be logically encapsulated with one or more controllers of the system control module 310 (e.g., the memory controller module 330). For one embodiment, at least one of the (one or more) processors 305 can be logically encapsulated with one or more controllers of the system control module 310 to form a system-in-package (SiP). For one embodiment, at least one of the (one or more) processors 305 can be logically integrated with one or more controllers of the system control module 310 on the same die. For one embodiment, at least one of the (one or more) processors 305 can be logically integrated with one or more controllers of the system control module 310 on the same die to form a system-on-chip (SoC).
[0079] In various embodiments, the system 300 can be, but is not limited to: a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet, a netbook, etc.). In various embodiments, the system 300 can have more or fewer components and / or a different architecture. For example, in some embodiments, the system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and speakers.
[0080] It should be noted that the present application can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application specific integrated circuit (ASIC), a general purpose computer, or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and the like. Additionally, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with the processor to execute each step or function.
[0081] In addition, a part of the present application can be applied as a computer program product, such as computer program instructions, which when executed by a computer, through the operation of the computer, can invoke or provide the methods and / or technical solutions according to the present application. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0082] The communication medium includes a medium through which a communication signal containing, for example, computer-readable instructions, data structures, program modules, or other data is transmitted from one system to another system. The communication medium can include a guided transmission medium (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and a wireless (unguided) medium that can propagate energy waves, such as sound, electromagnetic, RF, microwave, and infrared. The computer-readable instructions, data structures, program modules, or other data can be embodied as, for example, a modulated data signal in a wireless medium (such as a carrier wave or a similar mechanism that is part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are changed or set in a manner that encodes information in the signal. The modulation can be analog, digital, or a hybrid modulation technique.
[0083] By way of example and not limitation, a computer-readable storage medium can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media includes, but is not limited to, volatile memory such as random access memory (RAM, DRAM, SRAM); and non-volatile memory such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, tapes, CDs, DVDs); or other media now known or later developed that can store computer-readable information / data for use by a computer system.
[0084] Here, an embodiment according to the present application includes a device that includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to run based on the methods and / or technical solutions according to the foregoing multiple embodiments of the present application.
[0085] For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims can also be implemented by one unit or device through software or hardware. First, second, etc. are used to denote names and do not denote any particular order.
Claims
1. A method for marking sensitive data, wherein, The method includes: Sampling and collecting the log data of the database to obtain one or more sampling data; Identifying the sampling data according to the classification and grading rules, determining the sensitive classification corresponding to the sampling data, and obtaining the sensitive level information corresponding to the sampling data under the sensitive classification; According to the sensitive level information, marking at least one sampling data with a sensitive mark and storing the at least one sampling data, where the sensitive mark includes the sensitive classification and sensitive level information corresponding to each sampling data; Wherein, the classification and grading rules include a classification strategy and a sensitive level strategy; the identifying the sampling data according to the classification and grading rules, determining the sensitive classification corresponding to the sampling data, and obtaining the sensitive level information corresponding to the sampling data under the sensitive classification includes: Identifying the sampling data according to the classification strategy to determine the sensitive classification corresponding to the sampling data; According to the sensitive level strategy, determining the sensitive level information corresponding to the sampling data under the sensitive classification.
2. The method according to claim 1, wherein, The sampling and collecting the log data of the database to obtain one or more sampling data includes: Sampling and collecting the log data of the database through a plugin deployed on the database to obtain one or more sampling data.
3. The method according to claim 1, wherein The method further includes: Determining the classification and grading rules corresponding to the database.
4. The method according to claim 3, wherein The determining the classification and grading rules corresponding to the database includes: Obtaining the classification and grading rules corresponding to the database selected by the user from multiple default classification and grading rules.
5. The method according to claim 3, wherein The determining the classification and grading rules corresponding to the database includes: Determining the storage data characteristics corresponding to the database by performing semantic analysis on the stored data in the database; According to the storage data characteristics, determining the classification and grading rules that match the storage data characteristics.
6. The method according to claim 5, wherein, The according to the storage data characteristics, determining the classification and grading rules that match the storage data characteristics includes: According to the storage data characteristics, determining the sensitive scenario information corresponding to the database; According to the sensitive scenario information, determining the classification and grading rules that match the sensitive scenario information.
7. The method according to claim 1, wherein, The sensitive level strategy includes the first sensitive level information corresponding to each sensitive classification under the classification strategy; Wherein, the according to the sensitive level strategy, determining the sensitive level information corresponding to the sampling data under the sensitive classification includes: According to the sensitive level strategy, obtaining the first sensitive level information corresponding to the sensitive classification; According to the first sensitive level information, determining the sensitive level information corresponding to the sampling data under the sensitive classification.
8. The method according to claim 7, wherein The sensitive level strategy further includes multiple sub-sensitive classifications corresponding to each sensitive classification under the classification strategy and the second sensitive level information corresponding to each sub-sensitive classification; Wherein, the according to the first sensitive level information, determining the sensitive level information corresponding to the sampling data under the sensitive classification includes: According to the sensitive level strategy, determining the sub-sensitive classification corresponding to the sampling data under the sensitive classification; Determine the sensitive level information corresponding to the sampled data under the sensitive classification according to the first sensitive level information and the second sensitive level information corresponding to the sub-sensitive classification.
9. The method according to claim 1, wherein The sensitive level policy includes multiple sub-sensitive classifications corresponding to each sensitive classification under the classification policy and the third sensitive level information corresponding to each sub-sensitive classification; Among them, determining the sensitive level information corresponding to the sampled data under the sensitive classification according to the sensitive level policy includes: Determine the sub-sensitive classification corresponding to the sampled data under the sensitive classification according to the sensitive level policy; Determine the sensitive level information corresponding to the sampled data under the sensitive classification according to the third sensitive level information corresponding to the sub-sensitive classification.
10. The method according to claim 1, wherein, The method further includes: Obtain the review result of the user on the stored target sampled data; If at least one of the sensitive classification and the sensitive level information corresponding to the target sampled data is inconsistent with the review result, adjust the sensitive mark of the target sampled data according to the review result.
11. The method according to claim 10, wherein, Adjusting the sensitive mark of the target sampled data includes at least one of the following: Modify the sensitive classification in the sensitive mark; Modify the sensitive level information in the sensitive mark; Remove the sensitive mark for the target sampled data and cancel storing the target sampled data.
12. The method according to claim 1, wherein, The method further includes: Adjust the sampling rate of the database according to the sensitive level information, so that the new log data of the database is sampled and collected using the adjusted sampling rate subsequently.
13. A computer device for marking sensitive data, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 12.
14. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Method and device for identifying sensitive data of database
CN113919352A
Data security classification sampling and labeling
US20200380160A1