Data personalized desensitization method and system and storage medium

By identifying and dividing the levels of sensitive data and selecting appropriate desensitization algorithms based on application scenarios, the problems of low desensitization efficiency and single processing methods in the prior art are solved, efficient and accurate data desensitization is achieved, and data security and usability are ensured.

CN120030590APending Publication Date: 2025-05-23ANHUI YACHUANG ELECTRONICS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510086373.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing data desensitization technology has low desensitization efficiency and single processing methods, which leads to poor data availability and business operations.

Method used

A personalized data desensitization method is proposed. By extracting the data to be analyzed in the database, identifying and dividing the levels of sensitive data, and selecting a suitable desensitization algorithm for desensitization according to the application scenario. This method uses semantic analysis to classify data and subdivide sensitive features, combines artificial intelligence models to build a desensitization model, and intelligently selects a desensitization algorithm.

Benefits of technology

It realizes accurate identification and sensitivity level classification of sensitive data, improves desensitization efficiency and accuracy, and ensures that data meets business needs while also complying with the standards of data security and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030590A_ABST
    Figure CN120030590A_ABST
Patent Text Reader

Abstract

The invention discloses a data personalized desensitization method and system and a storage medium, relates to the technical field of data security, and solves the technical problems of low desensitization efficiency and single processing mode of an existing data desensitization technology. The method is used for extracting to-be-analyzed data in a database, identifying sensitive data in the to-be-analyzed data, performing sensitive grade division on the sensitive data, and selecting an application scene of the sensitive data; selecting a desensitization algorithm for desensitization according to the sensitivity level of the sensitive data and the corresponding application scene; according to the method, the sensitive data is subjected to sensitive grade division, so that the highly sensitive data can be identified and protected more accurately, the method is not limited to a traditional single desensitization mode any more, and the desensitization algorithm is flexibly selected according to the sensitive grade of the sensitive data and the corresponding application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data security, and specifically relates to a method, system and storage medium for personalized data desensitization. Background Art

[0002] In recent years, the key digital economy has continued to develop rapidly, and its output value ratio has increased rapidly year by year, and it has become an important engine to promote economic growth; data security risk patterns have also become diversified and complex, posing serious threats and damages to individuals, organizations, social public interests and even national security; data desensitization, as one of the important measures for data security, plays an important role in data security, but currently no unified standard has been formed globally, and it is still in an imperfect and unreliable development stage. The specific problems are: 1. Inconsistent standards for sensitive data; 2. Poor data availability after data desensitization and the irreversibility of desensitization technology have an impact on business operations; 3. Low data desensitization efficiency and a single processing method; Therefore, the present invention provides a personalized data desensitization method, system and storage medium. Summary of the invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a data personalized desensitization method, system and storage medium to solve the technical problems of low desensitization efficiency and single processing method of the existing data desensitization technology.

[0004] To achieve the above object, the first aspect of the present invention provides a data personalized desensitization method, comprising the following steps:

[0005] Step 1: Extract the data to be analyzed from the database and identify whether the data to be analyzed is sensitive data; if yes, classify the sensitive data into sensitive levels; if no, do not process it;

[0006] Step 2: Select application scenarios for sensitive data; the application scenarios are several operational requirements of the server;

[0007] Step 3: Select a desensitizing algorithm to desensitize sensitive data based on the sensitivity level of the sensitive data and the corresponding application scenario.

[0008] Preferably, the identifying whether the data to be analyzed is sensitive data includes:

[0009] The data to be analyzed are classified according to preset data types; the classified data to be analyzed are divided into sensitive data and non-sensitive data according to preset sensitive data features; wherein the preset sensitive data features are composed of a number of preset sensitive sub-features.

[0010] Preferably, the classifying the data to be analyzed according to preset data types includes:

[0011] Extract the key information of the data to be classified, and semantically match the key information of the data to be classified with the key information of each data type through semantic analysis to obtain a matching value; classify the data to be classified into the data type corresponding to the maximum matching value.

[0012] The present invention classifies the data to be analyzed by calculating the matching values ​​between the data to be classified and each data type through the semantic analysis method; the semantic analysis method relies on clear algorithms and statistical models, so that the key information extraction and matching process of the data type is objective and consistent, which helps to reduce the impact of human factors on the classification results and improve the accuracy of classification; in addition, as the data types continue to increase and change, the semantic analysis method can adapt to new classification needs by updating the key information library and algorithm.

[0013] Preferably, the step of classifying the sensitive data into different sensitivity levels includes:

[0014] Set sensitivity levels and preset sensitive data features corresponding to each sensitivity level, and mark the sensitive data features corresponding to each sensitivity level as sensitive feature one; extract sensitive data features of sensitive data and mark them as sensitive feature two; calculate the ratio between sensitive feature one and sensitive feature two, and divide the sensitive data into the sensitivity level corresponding to the largest ratio; wherein, sensitive feature one is composed of several sensitive sub-features one, and sensitive feature two is composed of several sensitive sub-features two.

[0015] The present invention can more accurately describe and identify sensitive data by subdividing sensitive feature 1 and sensitive feature 2 into several sensitive sub-features 1 and several sensitive sub-features 2, thus avoiding overgeneralization or neglect of data. This meticulous processing method makes the identification of sensitive data more accurate and comprehensive; and by calculating the ratio of sensitive feature 1 of sensitive data to sensitive feature 2 of each sensitivity level, the sensitivity of sensitive data can be confirmed.

[0016] Preferably, the sensitivity levels include primary sensitivity, secondary sensitivity and tertiary sensitivity; the sensitivity levels are ordered from large to small as primary sensitivity < secondary sensitivity < tertiary sensitivity.

[0017] Preferably, the calculating the ratio between the first sensitive feature and the second sensitive feature includes:

[0018] Perform semantic similarity analysis between each sensitive sub-feature 1 and each sensitive sub-feature 2 to obtain the semantic similarity value between each sensitive sub-feature 1 and each sensitive sub-feature 2; mark the sensitive sub-feature 1 and the sensitive sub-feature 2 with the largest semantic similarity value as a match; count the number of matches in sensitive sub-feature 2, calculate the ratio of the number of matches to the total number of sensitive sub-features 2, and obtain the ratio between sensitive feature 1 and sensitive feature 2.

[0019] The semantic similarity analysis of the present invention can accurately capture the intrinsic connection between sensitive sub-features, and quantify the degree of similarity between sensitive sub-feature one and sensitive sub-feature two by calculating the semantic similarity value between them, so as to more accurately judge which sub-features are matched, overcome the limitations of traditional keyword-based matching, and improve the accuracy and comprehensiveness of matching; and by counting the number of matches in sensitive sub-feature two and performing ratio processing with the total number, the ratio between sensitive feature one and sensitive feature two is obtained, and this ratio reflects the degree of fit between sensitive data features and preset features, and provides an objective basis for the classification and grading of sensitive data.

[0020] Preferably, the application scenarios for selecting sensitive data include:

[0021] Extract the demand data of the application scenario, and determine the application scenario of the sensitive data based on the preset correspondence between the demand data and the sensitive data; wherein the demand data of the application scenario is the data required when the application scenario is running.

[0022] Preferably, the step of selecting a desensitizing algorithm for desensitization according to the sensitivity level of the sensitive data and the corresponding application scenario includes:

[0023] The sensitivity level of sensitive data and the corresponding application scenario are input into the desensitization model, and the desensitization model outputs a desensitization label. The desensitization label is matched with a preset algorithm to obtain a desensitization algorithm. The sensitive data is desensitized by the desensitization algorithm.

[0024] The desensitization model is constructed by an artificial intelligence model, and the construction process is as follows:

[0025] Obtain the sensitivity level, application scenarios and desensitization algorithms of sensitive data from historical data;

[0026] Integrate the sensitivity level, application scenarios and desensitizing labels of sensitive data into several groups of training data and test data; use the training data to train the artificial intelligence model; use the test data to test the trained artificial intelligence model, and adjust the artificial intelligence model according to the test results; finally obtain a desensitizing model whose input is the sensitive data level and the corresponding application scenario, and whose output is the desensitizing algorithm; wherein the artificial intelligence model is a BP neural network model or a RBF neural network model.

[0027] The present invention integrates the sensitivity level, application scenario and desensitization algorithm in historical data to construct training data and test data, and then trains and adjusts the artificial intelligence model. The resulting desensitization model can intelligently select the most appropriate desensitization algorithm according to the input sensitive data level and application scenario. This intelligent selection method greatly improves the efficiency and accuracy of the desensitization process.

[0028] Preferably, the second aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the steps of the method.

[0029] Preferably, the third aspect of the present invention provides a data personalized desensitization system, including a scene selection module, and a scene selection module and a desensitization module connected thereto;

[0030] Sensitive data analysis module: used to extract the data to be analyzed from the database and identify whether the data to be analyzed is sensitive data; if yes, the sensitive data is classified into sensitivity levels; if not, no processing is performed;

[0031] Scenario selection module: used to select application scenarios of sensitive data; application scenarios are several operational requirements of the server;

[0032] Desensitization module: select a desensitization algorithm for desensitization based on the sensitivity level of sensitive data and the corresponding application scenario.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] In order to ensure the security of data, the present invention firstly utilizes a preset sensitive data feature recognition mechanism to accurately identify the sensitive parts in the data to be analyzed, and further classifies these data in detail according to different sensitivity levels, assigns them corresponding sensitivity levels, thereby realizing accurate identification of sensitive data and sensitivity level classification, and ensuring that sensitive information is effectively protected; sensitive data is selected according to the operational needs of the application scenario, and a desensitizing algorithm that can effectively protect sensitive information is intelligently selected according to the level of sensitive data and the specific application scenario, thereby realizing scenario-based desensitization scheme formulation, thereby ensuring that the data meets the standards of data security and privacy protection while meeting the business needs; the present invention realizes scenario-based desensitization scheme formulation, and can tailor corresponding desensitization strategies for different types of application scenarios, thereby not only improving the security of data, but also ensuring the availability and integrity of data, and reducing the risk of abuse or leakage. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 It is a schematic diagram of the process of the present invention;

[0037] Figure 2A schematic diagram of a method for classifying sensitivity levels of sensitive data according to the present invention;

[0038] Figure 3 This is a system structure block diagram of the present invention. DETAILED DESCRIPTION

[0039] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] See also Figure 1 The first aspect of the present invention provides a method for personalized data desensitization, comprising the following steps:

[0041] Step 1: Extract the data to be analyzed from the database and identify whether the data to be analyzed is sensitive data; if yes, classify the sensitive data into sensitive levels; if no, do not process it;

[0042] Specifically, the data to be analyzed is classified according to preset data types; the classified data to be analyzed is divided into sensitive data and non-sensitive data according to preset sensitive data features; wherein the preset sensitive data features are composed of a number of preset sensitive sub-features.

[0043] Extract the key information of the data to be classified, and semantically match the key information of the data to be classified with the key information of each data type through semantic analysis to obtain a matching value; classify the data to be classified into the data type corresponding to the maximum matching value.

[0044] The classification of data to be analyzed can be divided into personal information and non-personal information according to the individual dimension of citizens; divided into public data and social data according to the public management dimension; divided into public dissemination information and non-public dissemination information according to the information dissemination dimension; determine whether there are industry classification rules for the data, and if so, classify them according to the industry classification rules, such as industrial data, telecommunications data, health data, education data, etc.; if there are no industry classification rules, classify the data according to the organizational operation dimension into user data, business data, business management data, system operation and security data.

[0045] See also Figure 2 , setting sensitivity levels, and presetting sensitive data features corresponding to each sensitivity level, marking the sensitive data features corresponding to each sensitivity level as sensitive feature one; extracting sensitive data features of sensitive data, and marking them as sensitive feature two;

[0046] Among them, the sensitivity levels include level one sensitivity, level two sensitivity and level three sensitivity; the sensitivity levels are ordered from large to small as level one sensitivity < level two sensitivity < level three sensitivity.

[0047] Calculate the ratio between sensitive feature 1 and sensitive feature 2 as follows:

[0048] Perform semantic similarity analysis between each sensitive sub-feature 1 and each sensitive sub-feature 2 to obtain the semantic similarity value between each sensitive sub-feature 1 and each sensitive sub-feature 2; mark the sensitive sub-feature 1 and the sensitive sub-feature 2 with the largest semantic similarity value as a match; count the number of matches in sensitive sub-feature 2, calculate the ratio of the number of matches to the total number of sensitive sub-features 2, and obtain the ratio between sensitive feature 1 and sensitive feature 2.

[0049] Sensitive data is divided into sensitivity levels corresponding to the largest proportion; wherein, sensitive feature one is composed of several sensitive sub-feature ones, and sensitive feature two is composed of several sensitive sub-feature twos.

[0050] For example: assuming that the sensitive feature 1 of some sensitive data is [A, B, C], A, B, and C are the sensitive sub-feature 1 of the sensitive data respectively; the sensitive data feature 2 corresponding to the first-level sensitivity is [D0, E0, F0], the sensitive data feature 2 corresponding to the second-level sensitivity is [D1, E1, F1, G1], and the sensitive data feature 2 corresponding to the third-level sensitivity is [D2, E2, F1, G2]; it should be noted that the higher the sensitivity level, the more sensitive the sensitive data is, and once it is damaged, the greater the degree of harm caused.

[0051] If the semantic similarity analysis between the sensitive feature 1 of the sensitive data and each sensitive feature 2 is performed, assuming that the semantic similarity value between A and D0 is 85%, and the semantic similarity value between A and D1 is 60%, then A matches D0, and A and D0 are marked as matches respectively, and the analysis method of B, C and A is the same; if the number of matches between the sensitive feature 1 of the sensitive data and the first-level sensitive data feature 2 is 2, the number of matches with the second-level sensitive data feature 2 is 1, and the number of matches with the third-level sensitive data feature 2 is 1, then the ratio of the sensitive feature 1 of the sensitive data to the first-level sensitive data feature 2 is the largest, so the sensitivity level of the sensitive data is level 1, and the sensitivity level is relatively low.

[0052] Step 2: Select application scenarios for sensitive data; the application scenarios are several operational requirements of the server;

[0053] Specifically, extract the demand data of the application scenario, and determine the application scenario of the sensitive data according to the preset corresponding relationship between the demand data and the sensitive data;

[0054] Among them, the demand data of the application scenario is the data required when the application scenario is running; the application scenarios include development and testing scenarios, teaching and training scenarios, analysis and mining scenarios, data reporting scenarios, sharing and exchange scenarios, business query scenarios, operation and maintenance management scenarios, etc.

[0055] Step 3: Select a desensitizing algorithm to desensitize sensitive data based on the sensitivity level of the sensitive data and the corresponding application scenario.

[0056] The sensitivity level of sensitive data and the corresponding application scenario are input into the desensitization model, and the desensitization model outputs a desensitization label. The desensitization label is matched with a preset algorithm to obtain a desensitization algorithm. The sensitive data is desensitized by the desensitization algorithm.

[0057] The desensitization model is constructed by an artificial intelligence model, and the construction process is as follows:

[0058] Obtain the sensitivity level, application scenarios and desensitization algorithms of sensitive data from historical data;

[0059] Integrate the sensitivity level, application scenarios and desensitizing labels of sensitive data into several groups of training data and test data; use the training data to train the artificial intelligence model; use the test data to test the trained artificial intelligence model, and adjust the artificial intelligence model according to the test results; finally obtain a desensitizing model whose input is the sensitive data level and the corresponding application scenario, and whose output is the desensitizing algorithm; wherein the artificial intelligence model is a BP neural network model or a RBF neural network model.

[0060] Among them, desensitizing algorithms such as jacquard, shuffling, spatiotemporal variation, numerical variation, cancellation or deletion, random selection, encryption, expression desensitization, key-value desensitization, etc. support combined desensitization and custom grouping desensitization.

[0061] A second aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method steps of the present invention.

[0062] See also Figure 3 , the third aspect of the present invention provides a data personalized desensitization system, including a scene selection module, and a scene selection module and a desensitization module connected thereto;

[0063] Sensitive data analysis module: used to extract the data to be analyzed from the database and identify whether the data to be analyzed is sensitive data; if yes, the sensitive data is classified into sensitivity levels; if not, no processing is performed;

[0064] Scenario selection module: used to select application scenarios of sensitive data; application scenarios are several operational requirements of the server;

[0065] Desensitization module: select a desensitization algorithm for desensitization based on the sensitivity level of sensitive data and the corresponding application scenario.

[0066] Part of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is a formula closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.

[0067] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A method for personalized data desensitization, characterized in that: The following steps are involved: Step 1: Extract the data to be analyzed from the database and identify whether the data to be analyzed is sensitive data; If yes, the sensitive data will be classified into different sensitivity levels; if no, no processing will be done; Step 2: Select application scenarios for sensitive data; the application scenarios are several operational requirements of the server; Step 3: Select a desensitizing algorithm to desensitize sensitive data based on the sensitivity level of the sensitive data and the corresponding application scenario.

2. A data personalized desensitization method according to claim 1, characterized in that: The identifying whether the data to be analyzed is sensitive data includes: The data to be analyzed are classified according to preset data types; the classified data to be analyzed are divided into sensitive data and non-sensitive data according to preset sensitive data features; wherein the preset sensitive data features are composed of a number of preset sensitive sub-features.

3. A data personalized desensitization method according to claim 2, characterized in that: The classifying the data to be analyzed according to preset data types includes: Extract the key information of the data to be classified, and semantically match the key information of the data to be classified with the key information of each data type through semantic analysis to obtain a matching value; classify the data to be classified into the data type corresponding to the maximum matching value.

4. A method for personalized data desensitization according to claim 3, characterized in that: The sensitivity classification of sensitive data includes: Set sensitivity levels and preset sensitive data features corresponding to each sensitivity level, and mark the sensitive data features corresponding to each sensitivity level as sensitive feature one; extract sensitive data features of sensitive data and mark them as sensitive feature two; calculate the ratio between sensitive feature one and sensitive feature two, and divide the sensitive data into the sensitivity level corresponding to the largest ratio; wherein, sensitive feature one is composed of several sensitive sub-features one, and sensitive feature two is composed of several sensitive sub-features two.

5. A method for personalized data desensitization according to claim 4, characterized in that: The sensitivity levels include primary sensitivity, secondary sensitivity and tertiary sensitivity; the sensitivity levels are ordered from large to small as primary sensitivity < secondary sensitivity < tertiary sensitivity.

6. A data personalized desensitization method according to claim 4, characterized in that: The calculating the ratio between the first sensitive feature and the second sensitive feature includes: Perform semantic similarity analysis between each sensitive sub-feature 1 and each sensitive sub-feature 2 to obtain the semantic similarity value between each sensitive sub-feature 1 and each sensitive sub-feature 2; mark the sensitive sub-feature 1 and the sensitive sub-feature 2 with the largest semantic similarity value as a match; count the number of matches in sensitive sub-feature 2, calculate the ratio of the number of matches to the total number of sensitive sub-features 2, and obtain the ratio between sensitive feature 1 and sensitive feature 2.

7. A method for personalized data desensitization according to claim 4, characterized in that: The application scenarios of selecting sensitive data include: Extract the demand data of the application scenario, and determine the application scenario of the sensitive data based on the preset correspondence between the demand data and the sensitive data; wherein the demand data of the application scenario is the data required when the application scenario is running.

8. A method for personalized data desensitization according to claim 7, characterized in that: The method of selecting a desensitizing algorithm for desensitization according to the sensitivity level of the sensitive data and the corresponding application scenario includes: The sensitivity level of sensitive data and the corresponding application scenario are input into the desensitization model, and the desensitization model outputs a desensitization label. The desensitization label is matched with a preset algorithm to obtain a desensitization algorithm. The sensitive data is desensitized by the desensitization algorithm. The desensitization model is constructed by an artificial intelligence model, and the construction process is as follows: Obtain the sensitivity level, application scenarios and desensitization algorithms of sensitive data from historical data; Integrate the sensitivity level, application scenarios and desensitizing labels of sensitive data into several groups of training data and test data; use the training data to train the artificial intelligence model; use the test data to test the trained artificial intelligence model, and adjust the artificial intelligence model according to the test results; finally obtain a desensitizing model whose input is the sensitive data level and the corresponding application scenario, and whose output is the desensitizing algorithm; wherein the artificial intelligence model is a BP neural network model or a RBF neural network model.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the steps of the method according to any one of claims 1 to 8.

10. A data personalized desensitization system, operating based on a data personalized desensitization method according to any one of claims 1 to 8, characterized in that: It includes a scene selection module, and a scene selection module and a desensitization module connected thereto; Sensitive data analysis module: used to extract the data to be analyzed from the database and identify whether the data to be analyzed is sensitive data; If yes, the sensitive data will be classified into different sensitivity levels; if no, no processing will be done; Scenario selection module: used to select application scenarios of sensitive data; application scenarios are several operational requirements of the server; Desensitization module: select a desensitization algorithm for desensitization based on the sensitivity level of sensitive data and the corresponding application scenario.

Citation Information

Cited By

  • Data security management device based on medical data operation mode

    CN120321027A