A data governance method and system based on data security identification level
Through the coordinated work of the data identification end, the screening end and the desensitization decision-making end, the problems of difficulty in identifying duplicate data, inefficiency and improper processing of sensitive vocabulary in data governance are solved, and the accuracy and security of data governance are improved.
Patent Information
- Application Number
- CN202510044200.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-11
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-01-11
AI Technical Summary
In the prior art, data governance at the data security identification level has problems such as difficulty in identifying duplicate data, low recognition accuracy, low efficiency and improper processing of sensitive vocabulary, resulting in confusion in data governance and insufficient security.
The data identification end is used to perform multi-level data comparison and repeated data identification, the data screening end screens and sequence numbers of valid and invalid data, and the desensitization decision end performs sensitive vocabulary search and desensitization strategy generation to ensure the accuracy, efficiency and security of data governance.
It realizes accurate identification in the data governance process, prevents duplicate data chaos, improves data governance efficiency, and ensures accurate processing of sensitive words, improving the safety and reliability of data governance.
Smart Images

Figure CN119961972B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance technology, and in particular to a data governance method and system based on data security identification levels. Background Art
[0002] Data governance at the data security identification level is a complex process involving multiple aspects and levels. It requires regular risk assessments, updating data security policies based on new threats and regulatory requirements, distinguishing data of different importance and sensitivity, and implementing corresponding levels of protection measures to ensure that data is properly managed throughout its life cycle.
[0003] Publication No. CN116541382B discloses a data governance method based on data security identification level, including: receiving a data governance request sent by a client in response to a data governance process triggered by a user; creating a data index for the data to be governed, querying the data index that matches the preset sensitive identification rule in the data index, and then generating multiple sensitive data based on the successfully matched data index and obtaining the label attributes of each of the sensitive data; obtaining a pre-set data governance configuration interface corresponding to the user according to the user identifier in the data governance request, and finding a data governance policy that matches the highest sensitivity level in the label level attributes from the data governance policy group corresponding to the most label category attributes; the data governance policy includes a desensitizing policy, and the label attribute corresponding to the data governance policy is the optimal governance label attribute.
[0004] After searching the above patents, it was found that there are still some deficiencies in data governance at the data security identification level: 1. During data security identification, multi-level data indexes are prone to large quantities of duplicate data. Large quantities of duplicate data make data security identification more difficult, and it is impossible to effectively identify different levels in real time. In particular, multi-level identification is prone to data identification confusion, which affects the accuracy of data identification; 2. Due to the presence of multiple data governance requests in the data governance identification process, valid data and invalid data cannot be screened in real time, and valid data and invalid data cannot be identified and divided in order, resulting in inefficient data governance; 3. There are sensitive words in the data governance process. These sensitive words are likely to cause data governance difficulties during data identification. It is impossible to set data governance strategies for sensitive words, which affects the security and reliability of the data governance process and is prone to data governance anomalies.
[0005] Therefore, a data governance method and system based on data security identification level is proposed to solve the above problems. Summary of the Invention
[0006] The main purpose of the present invention is to provide a data governance method and system based on data security identification level to solve the problems raised in the above background.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is: a data governance method and system based on data security identification level, including a data identification end, a data screening end and a desensitization decision end;
[0008] The data identification terminal is used to receive data management request data in real time, receive and sort multiple data at the same time, perform multi-level data comparison in real time according to the sequence number, identify whether duplicate data appears in real time, and issue a voice alarm in time when duplicate data appears and store the duplicate data in real time for sequence number arrangement;
[0009] The data screening end is used to receive data management requests in real time, identify duplicate data, and screen the duplicate data into valid data and invalid data in real time, and identify and divide the valid data and invalid data into the order of precedence and arrange the serial numbers;
[0010] The desensitization decision-making end is used to receive valid data and invalid data in real time, identify and divide them in order, and arrange them in sequence, set a data governance model, and retrieve sensitive words in the data in real time, and then automatically generate a desensitization strategy execution plan based on the sensitive words.
[0011] The data identification terminal includes a request receiving module, a level identification module and an identification tracking module;
[0012] The request receiving module includes a governance request unit and a multi-data receiving unit;
[0013] The governance request unit is used to receive data governance request data corresponding to the data governance process in real time through a data receiver;
[0014] The multi-data receiving unit is used to receive data management request data corresponding to multiple channels in real time through multiple channels, and arrange the data in sequence according to Arabic numerals from small to large.
[0015] The level identification module includes a repeat comparison unit and a hierarchical sorting unit;
[0016] The duplication comparison unit is used to compare data management request data corresponding to multiple channels in real time to see if they are duplicated.
[0017] The hierarchical sorting unit is used to arrange the duplicate data of the data governance request data by serial number through time-series grading, and arrange the serial numbers according to Arabic numerals from small to large.
[0018] The identification and tracking module includes a synchronous identification unit, a multi-level tracking unit and a tracking warning unit;
[0019] The synchronous identification unit is used to track the data duplication check results of the data management request received by multiple channels at the current moment in real time through the data tracker;
[0020] The multi-level tracking unit is used to perform multi-level tracking on the duplicate checking results of the data governance request data through time-series grading, and the multi-level tracking is divided into multiple levels according to the time sequence;
[0021] The tracking and warning unit is used to report to the system and issue a voice alarm when data duplication is tracked.
[0022] The data screening end includes an identification and receiving module, a data screening module and a sequence arrangement module;
[0023] The identification and receiving module includes an identification and receiving unit and a data division unit;
[0024] The identification receiving unit is used to receive the duplicate checking results of the data governance request received through multiple channels in real time through the data receiver;
[0025] The data partitioning unit is used to perform real-time screening of valid and invalid data on the data governance request data using the OS CFAR algorithm. The screening method is as follows:
[0026] If there are at least min P points (including p itself) in the ε-neighborhood of a point p, it means that p is a valid data.
[0027] The data screening module includes a valid data unit and an invalid data unit;
[0028] The valid data unit is used to receive the valid data after data screening in real time through the data receiver and record it in real time through the data recorder;
[0029] The invalid data unit is used to receive the invalid data after data screening in real time through a data receiver, and record it in real time through a data recorder.
[0030] The sequence arrangement module includes a sequence setting unit, a valid sequence unit, an invalid sequence unit and a sequence adjustment unit;
[0031] The sequence setting unit is used to set the data management sequence number of the valid data and extract the corresponding keyword text according to different valid data and invalid data;
[0032] The effective sorting unit is used to arrange the keyword texts corresponding to the effective data in order by Arabic numerals from small to large;
[0033] The invalid sorting unit is used to arrange the keyword texts corresponding to the invalid data in order by Arabic numerals from small to large;
[0034] The sequence adjustment unit is used to adjust the data management sequence number of the valid data and invalid data of the corresponding sequence number in real time according to actual needs, and compare the keyword text in real time to determine whether the adjusted valid data is correct. The formula is the same as the calculation formula of the repeated comparison unit. If the D calculated twice is sm If D is equal to 0, it means that the order adjustment is normal. sm If it is not equal to 0, it means that the order adjustment is abnormal.
[0035] The desensitization decision-making end includes a data receiving module, a sequential progressive module, a sensitive retrieval module and a governance strategy module;
[0036] The data receiving module includes a data receiving unit and a data management module;
[0037] The data receiving unit is used to receive data governance request data after data duplication checking and data screening in real time through a data receiver;
[0038] The data governance model is used to set the data governance model, which is formulated according to the data security identification standard. The data governance model includes sensitive keywords, sensitive data and sensitive numbers, as well as the desensitization strategy execution plan corresponding to sensitive keywords, sensitive data and sensitive numbers;
[0039] The sequential progressive module is used to perform progressive data management on the data management request data according to time and sequence number arrangement.
[0040] The sensitive search module is used to perform sensitive search on the data management request data of the current moment of progressive data management in real time by repeating the calculation formula of the comparison unit. sm If D is equal to 0, it means that sensitive words appear. sm If it is not equal to 0, it means that no sensitive words appear;
[0041] The governance strategy module includes a desensitization strategy execution governance plan and an execution plan tracking unit;
[0042] The desensitization strategy execution scheme is used to automatically generate a corresponding desensitization strategy execution scheme through a sensitive retrieval module and immediately execute the corresponding desensitization strategy execution scheme;
[0043] The execution plan tracking unit is used to track the execution results of the desensitization strategy execution plan in real time through a data tracker.
[0044] A data governance method based on data security identification level includes the following steps:
[0045] Step 1: Configure the IP address information of the data governance remote control area server;
[0046] Step 2: Enter the data identification terminal, receive data management request data in real time, and use multi-level real-time identification to determine whether duplicate data exists. The accuracy of data management identification is guaranteed by the multi-level synchronous identification method. If duplicate data exists, the duplicate data is stored in real time and serialized.
[0047] Step 3: Enter the data screening terminal, receive data governance requests in real time, identify duplicate data, and screen valid and invalid data in real time. The valid and invalid data are identified and sorted in order of priority.
[0048] Step 4: Enter the desensitization decision-making end, and identify and divide the valid and invalid data in order and arrange them in sequence through real-time reception. By searching the sensitive words in the data in real time, the desensitization strategy execution plan is automatically generated based on the sensitive words.
[0049] The present invention has the following beneficial effects:
[0050] 1. In the present invention, by setting up a data identification terminal, during data governance based on the data security identification level, multi-level data comparison is performed in real time, so that the accuracy of data governance identification can be guaranteed according to the multi-level synchronous identification method during the data governance process, avoiding the difficulty of data security identification caused by large quantities of duplicate data, and being able to effectively identify different levels in real time. In particular, multi-level identification can prevent data identification confusion through multi-level progressive identification, thereby increasing the accuracy of data identification.
[0051] 2. In the present invention, by setting up a data screening end, when managing data based on the data security identification level, the valid data and invalid data are screened in real time on the data after repeated identification, valid data and invalid data are defined, and the valid data and invalid data are identified and divided in order and arranged in sequence. By extracting keywords from the valid data and invalid data and adjusting the order of data management according to the type of keywords, the data management efficiency is increased and the occupation of the data management channel by invalid data is reduced.
[0052] 3. In the present invention, by setting a desensitization decision-making end, during data governance based on the data security identification level, by real-time retrieval of sensitive words in the data and setting a data governance model, a desensitization strategy execution plan can be automatically generated according to the retrieved sensitive words during the data governance process, ensuring that sensitive words can be accurately identified during data identification, so that the system can set corresponding desensitization strategies for sensitive words, further avoiding data governance anomalies and ensuring the security and reliability of data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1This is a flowchart of the steps of a data governance method based on data security identification level of the present invention;
[0054] Figure 2 This is a schematic diagram of the architecture of a data identification terminal of a data governance system based on data security identification levels according to the present invention;
[0055] Figure 3 This is a schematic diagram of the architecture of a data screening end of a data management system based on data security identification levels according to the present invention;
[0056] Figure 4 This is a schematic diagram of the architecture of a desensitizing decision-making end of a data governance system based on data security identification levels in the present invention. DETAILED DESCRIPTION
[0057] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.
[0058] Example 1
[0059] Please refer to Figure 1-Figure 2 Shown: A data governance method and system based on data security identification level, including a data identification end, a data screening end, and a desensitization decision end;
[0060] The data identification terminal is used to receive data management request data in real time, receive and sort multiple data at the same time, perform multi-level data comparison in real time according to the sequence number, and identify whether duplicate data appears in real time. When duplicate data appears, it will promptly issue a voice alarm and store the duplicate data in real time for sequence number arrangement;
[0061] The data screening end is used to receive data governance requests in real time, identify duplicate data, and screen the duplicate data into valid and invalid data in real time, and identify and divide the valid and invalid data into the order of priority and arrange the serial numbers.
[0062] The desensitization decision-making end is used to identify and divide valid and invalid data in real time, arrange them in sequence, set up data governance models, and retrieve sensitive words in the data in real time, and then automatically generate a desensitization strategy execution plan based on the sensitive words.
[0063] The data identification end includes a request receiving module, a level identification module and an identification tracking module;
[0064] The request receiving module includes a management request unit and a multi-data receiving unit;
[0065] The governance request unit is used to receive data governance request data corresponding to the data governance process in real time through a data receiver;
[0066] The multi-data receiving unit is used to receive data management request data corresponding to multiple channels in real time through multiple channels, and arrange the serial numbers in ascending order according to Arabic numerals.
[0067] The level identification module includes a repeat comparison unit and a hierarchical sorting unit;
[0068] The duplicate comparison unit is used to compare data governance request data corresponding to multiple channels in real time to see if they are duplicated. The calculation formula is as follows:
[0069]
[0070] Among them, D sm The repetition rate of the corresponding data in the request data for duplicate data management, sm j is the repetition rate of the j-th query data, L j is the number of characters in the jth query data, if D sm If it is equal to 0, it means that the data is not repeated. sm If it is equal to 0, it means the data is repeated.
[0071] The hierarchical sorting unit is used to arrange the duplicate data of the data governance request data in sequence by time-series grading, and arrange the sequence in ascending order according to Arabic numerals.
[0072] The identification and tracking module includes a synchronous identification unit, a multi-level tracking unit, and a tracking warning unit;
[0073] The synchronous identification unit is used to track the data duplication checking results of the data management request received by multiple channels at the current moment in real time through the data tracker;
[0074] The multi-level tracking unit is used to track the duplicate checking results of data governance request data in multiple levels through time-series grading. The multi-level tracking is divided into multiple levels according to the time sequence;
[0075] The tracking warning unit is used to send a voice alarm to the reporting system when duplicate data is tracked.
[0076] Example 2
[0077] Please refer to Figure 3 As shown: Based on the first embodiment, the data screening end includes an identification and receiving module, a data screening module and a sequence arrangement module;
[0078] The identification and receiving module includes an identification and receiving unit and a data division unit;
[0079] The identification receiving unit is used to receive the duplicate checking results of the data governance request received through multiple channels in real time through the data receiver;
[0080] The data partitioning unit is used to perform real-time screening of valid and invalid data in data governance request data using the OS CFAR algorithm. The screening method is as follows:
[0081] If there are at least min P points (including p itself) in the ε-neighborhood of a point p, then p is a valid data, defined as:
[0082]
[0083] Among them, |N E P| represents the number of valid data in the ε-neighborhood of p, Core(p) represents the data set of the ε-neighborhood, min P represents the valid data in the ε-neighborhood, and the rest represents the invalid data in the data governance request data.
[0084] The data screening module includes valid data units and invalid data units;
[0085] The valid data unit is used to receive the valid data after data screening in real time through the data receiver and record it in real time through the data recorder;
[0086] The invalid data unit is used to receive invalid data after data screening in real time through a data receiver, and record it in real time through a data recorder, screen valid data and invalid data in real time after repeated identification, define valid data and invalid data, and identify and divide valid data and invalid data in sequence and arrange them in sequence, so as to manage multiple data governance requests in the data governance identification process.
[0087] The sequence arrangement module includes a sequence setting unit, a valid sequence unit, an invalid sequence unit and a sequence adjustment unit;
[0088] The sequence setting unit is used to set the data management sequence number of valid data and extract the corresponding keyword text according to different valid data and invalid data;
[0089] The effective sorting unit is used to arrange the keyword texts corresponding to the effective data in order from small to large using Arabic numerals;
[0090] The invalid sorting unit is used to arrange the keyword texts corresponding to the invalid data in order from small to large using Arabic numerals;
[0091] The sequence adjustment unit is used to adjust the data management sequence number of the valid data and invalid data of the corresponding sequence number in real time according to actual needs, and compare the keyword text in real time to determine whether the adjusted valid data is correct. The formula is the same as the calculation formula of the repeated comparison unit. If the D calculated twice is sm If D is equal to 0, it means that the order adjustment is normal. smIf it is not equal to 0, it means that the order adjustment is abnormal. It can timely screen out valid data and invalid data in data governance. By extracting keywords from valid data and invalid data, the valid data and invalid data are identified and divided in order, and the order of data governance is adjusted according to the type of keywords, thereby increasing data governance efficiency and reducing the occupation of data governance channels by invalid data.
[0092] Example 3
[0093] Please refer to Figure 4 As shown: Based on the first embodiment, the desensitization decision-making end includes a data receiving module, a sequential progressive module, a sensitive retrieval module and a governance strategy module;
[0094] The data receiving module includes a data receiving unit and a data management module;
[0095] The data receiving unit is used to receive data governance request data after data duplication checking and data screening in real time through a data receiver;
[0096] The data governance model is used to set the data governance model. The data governance model is formulated according to the data security identification standard. The data governance model includes sensitive keywords, sensitive data and sensitive numbers, as well as the desensitization strategy execution plan corresponding to sensitive keywords, sensitive data and sensitive numbers;
[0097] The sequential progressive module is used to perform progressive data governance on data governance request data according to time and serial number arrangement. During the data governance process, it can automatically generate a desensitization strategy execution plan based on the retrieved sensitive words, ensuring that sensitive words can be accurately identified during data recognition, so that the system can set corresponding desensitization strategies for sensitive words, further avoiding data governance anomalies and ensuring the security and reliability of data governance.
[0098] The sensitive retrieval module is used to perform sensitive retrieval on the data management request data of the current moment of progressive data management in real time by repeating the calculation formula of the comparison unit. sm If D is equal to 0, it means that sensitive words appear. sm If it is not equal to 0, it means that no sensitive words appear;
[0099] The governance strategy module includes the desensitization strategy execution governance plan and the execution plan tracking unit;
[0100] The desensitization strategy execution plan is used to automatically generate the corresponding desensitization strategy execution plan through the sensitive retrieval module and immediately execute the corresponding desensitization strategy execution plan;
[0101] The execution plan tracking unit is used to track the execution results of the desensitization strategy execution plan in real time through a data tracker. During the data governance process, it can automatically generate a desensitization strategy execution plan based on the retrieved sensitive words, ensuring that sensitive words can be accurately identified during data recognition, so that the system can set corresponding desensitization strategies for sensitive words, further avoiding data governance anomalies and ensuring the security and reliability of data governance.
[0102] In the present invention, a data management method and system based on data security identification level is provided. When the system is in operation, the IP address information of the data management remote control area server is first configured; the data identification terminal is entered, and data management request data is received in real time, and whether duplicate data appears is identified in real time through multiple levels. The accuracy of data management identification is guaranteed according to the multi-level synchronous identification method. If duplicate data appears, the duplicate data is stored in real time and the duplicate data is arranged in sequence. By receiving data management request data in real time, multiple data are received and sorted at the same time, and multi-level data comparison is performed in real time according to the sequence number to identify whether duplicate data appears in real time, so that the data management process can be guaranteed according to the multi-level synchronous identification method. The accuracy of data governance recognition, when duplicate data appears, voice alarms are issued in time and duplicate data are stored in real time for serial number arrangement, so that the same data can be arranged and divided in time during data governance, avoiding the difficulty of data security identification caused by large quantities of duplicate data, and can effectively identify different levels in real time, especially multi-level recognition can prevent data recognition confusion through multi-level progressive recognition, and increase the accuracy of data recognition; enter the data screening end, through real-time reception of data governance requests for duplicate identification of data, and real-time screening of valid data and invalid data, and identification and division of valid data and invalid data in order of sequence, through real-time reception of data governance requests and It identifies the data after repeated recognition, and screens the valid data and invalid data in real time after repeated recognition, defines valid data and invalid data, and identifies and divides the valid data and invalid data in order and arranges the serial numbers, so that multiple data governance requests in the data governance recognition process can timely screen out the valid data and invalid data in the data governance, and by extracting keywords from the valid data and invalid data, identify and divide the valid data and invalid data in order, and adjust the order of data governance according to the keyword type, thereby increasing the efficiency of data governance and reducing the occupation of the data governance channel by invalid data; entering the desensitization decision-making end, it identifies and divides the order through real-time reception. The valid and invalid data are divided and numbered, and sensitive words in the data are retrieved in real time. A desensitization strategy execution plan is automatically generated according to the sensitive words. The valid and invalid data are identified and divided in sequence and numbered, and sensitive words in the data are retrieved in real time. The data governance model is set, and a desensitization strategy execution plan is automatically generated according to the sensitive words. In the data governance process, a desensitization strategy execution plan can be automatically generated according to the retrieved sensitive words, ensuring that sensitive words can be accurately identified during data identification, so that the system can set corresponding desensitization strategies for sensitive words, further avoiding data governance anomalies and ensuring the security and reliability of data governance.
[0103] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A data governance system based on data security identification level, characterized by: The system includes a data identification end, a data screening end and a desensitization decision end; The data identification terminal is used to receive data management request data in real time, receive and sort multiple data at the same time, perform multi-level data comparison in real time according to the sequence number, identify whether duplicate data appears in real time, and issue a voice alarm in time when duplicate data appears and store the duplicate data in real time for sequence number arrangement; The data screening end is used to receive data management requests in real time, identify duplicate data, and screen the duplicate data into valid data and invalid data in real time, and identify and divide the valid data and invalid data into the order of precedence and arrange the serial numbers; The desensitization decision-making end is used to receive valid and invalid data in real time, identify and divide them in order, and arrange them in sequence, set a data governance model, and search for sensitive words in the data in real time, and then automatically generate a desensitization strategy execution plan based on the sensitive words; The data identification terminal includes a request receiving module, a level identification module and an identification tracking module; The request receiving module includes a governance request unit and a multi-data receiving unit; The governance request unit is used to receive data governance request data corresponding to the data governance process in real time through a data receiver; The multi-data receiving unit is used to receive data management request data corresponding to multiple channels in real time through multiple channels, and arrange the data in sequence according to Arabic numerals from small to large; The level identification module includes a repeat comparison unit and a hierarchical sorting unit; The duplication comparison unit is used to compare data corresponding to multiple channels in real time to see if the data management request data is duplicated through data comparison; The hierarchical sorting unit is used to arrange the duplicate data of the data governance request data by serial number according to time sequence classification, and arrange the serial numbers in ascending order according to Arabic numerals; The identification and tracking module includes a synchronous identification unit, a multi-level tracking unit and a tracking warning unit; The synchronous identification unit is used to track the data duplication check results of the data management request received by multiple channels at the current moment in real time through the data tracker; The multi-level tracking unit is used to perform multi-level tracking on the duplicate checking results of the data governance request data through time-series grading, and the multi-level tracking is divided into multiple levels according to the time sequence; The tracking and warning unit is used to report the system to issue a voice alarm when the data duplication is tracked; The data screening end includes an identification and receiving module, a data screening module and a sequence arrangement module; The identification and receiving module includes an identification and receiving unit and a data division unit; The identification receiving unit is used to receive the duplicate checking results of the data governance request received through multiple channels in real time through the data receiver; The data partitioning unit is used to perform real-time screening of valid and invalid data on the data governance request data using the OS CFAR algorithm. The screening method is as follows: If there are at least min P points (including p itself) in the ε-neighborhood of a point p, it means that p is a valid data; The data screening module includes a valid data unit and an invalid data unit; The valid data unit is used to receive the valid data after data screening in real time through the data receiver and record it in real time through the data recorder; The invalid data unit is used to receive the invalid data after data screening in real time through a data receiver, and record it in real time through a data recorder; The sequence arrangement module includes a sequence setting unit, a valid sequence unit, an invalid sequence unit and a sequence adjustment unit; The sequence setting unit is used to set the data management sequence number of the valid data and extract the corresponding keyword text according to different valid data and invalid data; The effective sorting unit is used to arrange the keyword texts corresponding to the effective data in order by Arabic numerals from small to large; The invalid sorting unit is used to arrange the keyword texts corresponding to the invalid data in order by Arabic numerals from small to large; The sequence adjustment unit is used to adjust the data management sequence number of the valid data and invalid data of the corresponding sequence number in real time according to actual needs, and compare the keyword text in real time to determine whether the adjusted valid data is correct. The formula is the same as the calculation formula of the repeated comparison unit. If the D calculated twice is sm If D is equal to 0, it means that the order adjustment is normal. sm If it is not equal to 0, it means that the order adjustment is abnormal; The desensitization decision-making end includes a data receiving module, a sequential progressive module, a sensitive retrieval module and a governance strategy module; The data receiving module includes a data receiving unit and a data management module; The data receiving unit is used to receive data governance request data after data duplication checking and data screening in real time through a data receiver; The data governance model is used to set the data governance model, which is formulated according to the data security identification standard. The data governance model includes sensitive keywords, sensitive data and sensitive numbers, as well as the desensitization strategy execution plan corresponding to sensitive keywords, sensitive data and sensitive numbers; The sequential progressive module is used to perform progressive data management on the data management request data according to time and sequence number arrangement.
2. The system according to claim 1, wherein: The sensitive search module is used to perform sensitive search on the data management request data of the current moment of progressive data management in real time by repeating the calculation formula of the comparison unit. sm If D is equal to 0, it means that sensitive words appear. sm If it is not equal to 0, it means that no sensitive words appear; The governance strategy module includes a desensitization strategy execution governance plan and an execution plan tracking unit; The desensitization strategy execution scheme is used to automatically generate a corresponding desensitization strategy execution scheme through a sensitive retrieval module and immediately execute the corresponding desensitization strategy execution scheme; The execution plan tracking unit is used to track the execution results of the desensitization strategy execution plan in real time through a data tracker.
3. A data governance method based on data security identification level according to any one of claims 1-2, characterized in that: The following steps are involved: Step 1: Configure the IP address information of the data governance remote control area server; Step 2: Enter the data identification terminal, receive data management request data in real time, and use multi-level real-time identification to determine whether duplicate data exists. The accuracy of data management identification is guaranteed by the multi-level synchronous identification method. If duplicate data exists, the duplicate data is stored in real time and serialized. Step 3: Enter the data screening terminal, receive data governance requests in real time, identify duplicate data, and screen valid and invalid data in real time. The valid and invalid data are identified and sorted in order of priority. Step 4: Enter the desensitization decision-making end, and identify and divide the valid and invalid data in order and arrange them in sequence through real-time reception. By searching the sensitive words in the data in real time, the desensitization strategy execution plan is automatically generated based on the sensitive words.
Citation Information
Patent Citations
Data governance methods and systems based on data security identification levels
CN116541382B
Data desensitization method and device, electronic device and storage medium
CN109558746A
Security monitoring management system and management method based on big data
CN113625603A