Sensitive data classification and grading protection method and device and electronic equipment

By automatically grading and intelligently desensitizing railway data, the problem of insufficient grading accuracy of railway data in the existing technology is solved, and flexible and efficient sensitive data protection is achieved to ensure that the data can still be used for analysis and decision-making after processing.

CN120296789APending Publication Date: 2025-07-11INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510407488.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing railway data protection methods rely on manual judgment and simple rule matching, resulting in insufficient accuracy and reliability of sensitive information grading, and lack of automation and intelligence in the selection of desensitization algorithms, making it difficult to adapt to diversified and dynamic data environments.

Method used

The preset sensitive data grading algorithm is used to classify the data to be processed, and the desensitization algorithm is used to select a suitable desensitization algorithm based on the sensitivity level of the data. Combined with natural language processing and machine learning technology, data sensitivity is identified through word segmentation, subject lexicon and hierarchical clustering, and suitable desensitization strategies are applied.

Benefits of technology

It improves the objectivity and consistency of data classification, reduces manual intervention errors, enhances the flexibility and adaptability of desensitization processing, ensures that the data can still be used for analysis and decision-making support after processing, and prevents the risk of data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296789A_ABST
    Figure CN120296789A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of information security, and discloses a sensitive data classification and classification protection method and device and electronic equipment, and the method comprises the steps: carrying out the classification processing of to-be-processed data through employing a preset sensitive data classification algorithm; inputting the to-be-processed data and the sensitivity level into a preset desensitization algorithm selection model to obtain a selection result; and performing desensitization processing on the to-be-processed data by using a desensitization algorithm corresponding to the selection result. Through the preset sensitive data grading algorithm, the sensitivity of the to-be-processed data can be accurately identified and graded, and the objectivity and consistency of data classification are ensured. The model is selected by using the desensitization algorithm, and the most suitable desensitization algorithm is automatically selected according to the sensitivity level of the data, so that the error of manual intervention is reduced, and the flexibility and adaptability of desensitization processing are also improved. Finally, by applying the selected desensitization algorithm, sensitive data are effectively protected, and a solid foundation is provided for safety management of railway data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security, and more particularly, to a method, apparatus, electronic device, and computer-readable storage medium for classifying and protecting sensitive data at different levels. Background Art

[0002] With the rapid development of the railway transportation industry, the data scale of the railway network has expanded rapidly, and the data sources have become more diverse. These data include, but are not limited to, passenger information, train scheduling, freight details, maintenance records, etc., which are crucial for the optimization of railway operations and safety management. However, a large amount of sensitive information, such as personal identity information and financial data, is contained in these data, and the leakage or improper use of this information may pose a serious threat to personal privacy and railway operation safety.

[0003] Currently, the protection of railway data mainly relies on the identification and desensitization of sensitive information. However, the sensitivity classification of railway data mainly depends on manual judgment or simple rule matching, which are inefficient and difficult to adapt to the changing data environment. In addition, the existing classification methods often lack a deep understanding of data semantics, resulting in insufficient accuracy and reliability of the classification results.

[0004] In the desensitization process of railway data, selecting a suitable desensitization algorithm is crucial for protecting data privacy and ensuring data availability. However, the existing desensitization algorithm selection process often relies on expert experience and lacks automated and intelligent decision support. This leads to an inaccurate selection of desensitization algorithms and an inability to effectively meet the requirements of different data types and sensitivity levels.

[0005] Therefore, how to accurately identify the sensitivity of railway data, automatically select the most suitable desensitization algorithm, and perform effective desensitization processing is a technical problem to be solved by those skilled in the art. Summary of the Invention

[0006] To solve the problems of low accuracy in data classification and the dependence of desensitization algorithm selection on expert experience in the prior art, the present invention provides a method, apparatus, electronic device, and computer-readable storage medium for classifying and protecting sensitive data at different levels.

[0007] A method for classifying and protecting sensitive data at different levels includes:

[0008] Obtaining data to be processed;

[0009] Performing classification processing on the data to be processed by using a preset sensitive data classification algorithm to obtain the sensitivity level of the data to be processed;

[0010] Input the data to be processed and the sensitivity level into a pre-set desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model;

[0011] Perform desensitization processing on the data to be processed using the desensitization algorithm corresponding to the selection result.

[0012] Optionally, the step of performing grading processing on the data to be processed using a pre-set sensitive data grading algorithm to obtain the sensitivity level of the data to be processed includes:

[0013] Perform word segmentation processing on the data to be processed based on a natural language processing tool to obtain at least one word unit to be processed;

[0014] For each word unit to be processed, calculate the distance between the word unit to be processed and each subject word in the subject word library; the corresponding relationship between the subject word and the sensitivity level is stored in the subject word library;

[0015] Determine the sensitivity level corresponding to the subject word with the closest distance to the word unit to be processed as the sensitivity level of the word unit to be processed;

[0016] Determine the probability of the sensitivity level of the data to be processed based on the sensitivity levels of all the word units to be processed;

[0017] Judge whether the difference between the highest value among all the probabilities and other values is within the error tolerance threshold range;

[0018] If so, determine the sensitivity level corresponding to the highest value among the probabilities as the sensitivity level of the data to be processed;

[0019] If not, perform hierarchical clustering on each word unit to be processed in the subject word library using a hierarchical clustering algorithm, and determine the number of subject words with different sensitivity levels contained in the cluster where the data to be processed is located according to the hierarchical clustering result;

[0020] Determine the sensitivity level with the largest number of subject words as the sensitivity level of the data to be processed.

[0021] Optionally, the step of determining the probability of the sensitivity level of the data to be processed based on the sensitivity levels of all the word units to be processed includes:

[0022] Generate a sensitivity level prediction probability vector of the data to be processed based on the sensitivity levels of all the word units to be processed;

[0023] Obtain the weight of each sensitivity level, and perform weighting and normalization processing on the sensitivity level prediction probability vector of the data to be processed according to the weight to obtain the probability of the sensitivity level of the data to be processed.

[0024] Optionally, the process of establishing the thesaurus includes:

[0025] Processing the input original corpus based on a natural language processing tool to obtain and output the word vector space of the original corpus;

[0026] Responding to the input first selection instruction, selecting initial topic words in the word vector space;

[0027] Annotating the sensitivity levels of the initial topic words through the analytic hierarchy process to generate an initial thesaurus;

[0028] Using the K-means clustering algorithm with the initial topic words as the initial centroids of clustering, calculating the distance between each word unit in the word vector space and the initial topic words;

[0029] Determining the word units in the word vector space that are closest to the initial topic words as the words to be added, and determining the sensitivity level of the initial topic words as the sensitivity level corresponding to the words to be added;

[0030] Adding the words to be added and the corresponding sensitivity levels of the words to be added to the initial thesaurus to obtain the thesaurus.

[0031] Optionally, calculating the distance between each word unit in the word vector space and the initial topic words includes:

[0032] Calculating the Euclidean distance between the word unit and the initial topic word according to the formula:

[0033]

[0034] where x i is the i-th coordinate of word unit a in the n-dimensional space, y i is the i-th coordinate of initial word unit b in the n-dimensional space, and S Euclidean (a, b) is the Euclidean distance between word unit a and the initial topic word b.

[0035] Optionally, the training process of the desensitization algorithm selection model includes:

[0036] Receiving the input training set; the training set includes training data and the corresponding desensitization algorithms and parameters of the training data;

[0037] Extracting the features of the training data according to the preset desensitization rule library and the sensitivity level of the training data to obtain the feature attributes of the training data; the feature attributes include at least one of the data name, data sensitivity level, and data attribute type;

[0038] Determine the desensitization algorithm and parameters corresponding to the training data as the target label, and perform random forest classification based on the feature attributes of the training data and the target label to train the desensitization algorithm selection model.

[0039] Optionally, the step of inputting the data to be processed and the sensitivity level into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model includes:

[0040] Calculate the class probability of each desensitization algorithm using the desensitization algorithm selection model

[0041]

[0042] where P(y = c|X) represents the probability that the data to be processed X belongs to the desensitization algorithm c, N represents the total number of trees in the random forest, and h i (X) represents the predicted class of the data to be processed X by the i-th tree, and I(·) is the indicator function.

[0043] Optionally, after using the preset sensitive data grading algorithm to grade the data to be processed to obtain the sensitivity level of the data to be processed, the method further includes:

[0044] Output the sensitivity level of the data to be processed;

[0045] In response to the input of the second selection instruction, select the training data from the data to be processed;

[0046] In response to the input of the annotation instruction, annotate the corresponding desensitization algorithm and parameters for the training data to obtain the training set.

[0047] A device for classifying and grading the protection of sensitive data, including:

[0048] An acquisition module for acquiring the data to be processed;

[0049] A grading processing module for grading the data to be processed using a preset sensitive data grading algorithm to obtain the sensitivity level of the data to be processed;

[0050] A desensitization algorithm selection module for inputting the data to be processed and the sensitivity level into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model;

[0051] A desensitization processing module for desensitizing the data to be processed using the desensitization algorithm corresponding to the selection result.

[0052] An electronic device, including:

[0053] A processor and a memory, where the memory is used to store at least one instruction, and when the instruction is loaded and executed by the processor, it is used to implement the method for classifying and grading the protection of sensitive data as described in any one of the above.

[0054] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for classifying and grading the protection of sensitive data as described in any one of the above.

[0055] In the method for classifying and grading the protection of sensitive data provided by the embodiments of the present invention, by using a preset sensitive data grading algorithm to perform grading processing on the data to be processed, the sensitive level of the data to be processed is obtained, and the data to be processed and the sensitive level are input into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model. Finally, the data to be processed is desensitized by using the desensitization algorithm corresponding to the selection result. Through the preset sensitive data grading algorithm of the present invention, the sensitivity of the data to be processed can be accurately identified and graded, ensuring the objectivity and consistency of data classification. By using the desensitization algorithm selection model, the most suitable desensitization algorithm is automatically selected according to the sensitive level of the data. This process not only reduces the error of manual intervention, but also improves the flexibility and adaptability of desensitization processing. Finally, by applying the selected desensitization algorithm, sensitive data is effectively protected, the risk of data leakage is prevented, and the usability of the data after processing is also ensured, providing a solid foundation for the safe sharing and analysis of railway data. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0057] Figure 1 It is a flowchart of a method for classifying and grading the protection of sensitive data provided by an embodiment of the present invention;

[0058] Figure 2 For Figure 1 It is a flowchart of an actual manifestation of S02 in a method for classifying and grading the protection of sensitive data provided;

[0059] Figure 3 It is a flowchart of the establishment process of a thesaurus provided by an embodiment of the present invention;

[0060] Figure 4 It is a flowchart of the training process of a desensitization algorithm selection model provided by an embodiment of the present invention;

[0061] Figure 5 This is a structural schematic diagram of a sensitive data classification and hierarchical protection device provided by an embodiment of the present invention. Detailed implementation manners

[0062] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0063] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0065] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0066] With the continuous development of intelligent railways, the data scale of railway networks has expanded rapidly, and the data sources have become more diverse. The transmission and sharing of massive data not only stimulate its potential value, but also pose more severe challenges to data security. Effectively identifying, classifying and grading railway sensitive data, and giving the best data desensitization strategies according to different sensitive levels and different data types are important topics for promoting railway digital upgrading and ensuring the security of railway network data.

[0067] Currently, the classification and grading of sensitive data have received extensive attention. The earliest method proposed was manual grading, where annotators classify sensitive data according to classification and grading rules and corresponding annotation processes. However, manual grading has a large degree of subjectivity. Based on this, a method was proposed to identify sensitive fields through K-means clustering and then combine association rule analysis for grading. K-means is an unsupervised learning algorithm that can divide data points into K clusters, maximizing the similarity within the clusters and the difference between the clusters. It is widely used in data classification and pattern recognition. However, the calculation efficiency of association rules is low, making it difficult to apply to scenarios with large-scale data. To improve the algorithm efficiency, a convolutional-based approach was proposed to enhance the performance of sensitive data recognition and achieve the effect of sensitive level division by mining non-linear features in the data. However, current methods are all difficult to handle problems such as short field lengths and weak semantic information commonly found in real-world business scenarios.

[0068] Data desensitization ensures the security of sensitive data during storage, transmission, and analysis by masking, encrypting, generalizing, or anonymizing the original data, preventing unauthorized access and leakage. Although these desensitization technologies can effectively protect sensitive data, they pose some significant challenges. A single desensitization method cannot adapt to diverse scenarios. At the same time, a trade-off needs to be made between desensitization and data availability. High-intensity desensitization processing may lead to a significant decrease in data availability, affecting subsequent data analysis or use. Most existing systems rely on manually selecting desensitization algorithms, and users need to configure appropriate desensitization strategies according to the data type, scenario, and security requirements. This manual adjustment is not only time-consuming and laborious but also prone to errors or security vulnerabilities. In a dynamic data environment, the sensitivity of data may change over time or with the data type, and static desensitization strategies cannot meet the needs of dynamic data.

[0069] Therefore, the present invention provides a method for classifying, grading, and protecting sensitive data to solve the above problems.

[0070] Please refer to Figure 1 , which is a flowchart of a method for classifying, grading, and protecting sensitive data provided by an embodiment of the present invention, including the following steps:

[0071] Step S01, obtain the data to be processed.

[0072] In this embodiment, obtaining the data to be processed is the primary step in the sensitive data classification and grading protection process, which involves collecting and preparing the original data that needs to be further analyzed or processed. In the context of the railway industry, the data to be processed may include key data such as passenger information, train scheduling data, freight records, maintenance logs, etc. The data to be processed may come from different systems and databases, so preprocessing operations such as cleaning, integrating, and formatting can also be performed on it to ensure that the data to be processed can be effectively used for subsequent grading and de-identification processing.

[0073] In some embodiments, due to the diversity and complexity of railway data, the original data may contain inconsistent or incomplete information, so it can be corrected through the data acquisition and preprocessing steps. By performing the step of obtaining the data to be processed, a solid foundation can be laid for the sensitivity analysis and de-identification processing of the data, thereby maximizing the value of the data while ensuring data security.

[0074] Step S02: Use a preset sensitive data grading algorithm to grade the data to be processed to obtain the sensitive level of the data to be processed.

[0075] In this embodiment, by applying a pre-defined algorithm to analyze and evaluate the data to be processed to determine the sensitivity level of each piece of data. The reason for performing sensitive data grading is to implement appropriate security measures during data use, storage, and transmission. Through grading, it can be ensured that sensitive information is protected at a higher level, while allowing non-sensitive information to be more widely shared and analyzed without violating privacy regulations. This step provides a basis for the de-identification processing of the data, enabling different de-identification strategies to be applied according to the sensitive levels of different data, thereby maximizing the availability and value of the data while protecting personal privacy and data security.

[0076] In some embodiments, natural language processing techniques, regular expression matching, and machine learning model development can be combined to obtain a sensitive data grading algorithm suitable for the characteristics of railway data to identify and evaluate sensitive information in the data. For example, the algorithm can scan sensitive fields in the data such as personal identity information, credit card numbers, health records, etc., and identify these patterns according to preset rules or training data.

[0077] Once the algorithm is trained and optimized, it will be applied to the data to be processed. During the application process, the algorithm will analyze each field in the data to be processed one by one, determine whether it contains sensitive information, and assign a level according to the degree of sensitivity. These levels can include four levels, namely S1, S2, S3, and S4, which can specifically depend on the requirements of the railway industry. The classified data will be marked so that corresponding measures can be taken according to its sensitivity level during subsequent desensitization processing. This process not only improves the automation degree of classified and graded protection of sensitive data, but also ensures that sensitive information is properly managed and protected.

[0078] In some embodiments, the classification requirements for the sensitivity levels of the data to be processed can be as shown in the following table:

[0079] Security level Hazard level S4-level data Severe S3-level data General S2-level data Minor S1-level data None

[0080] Step S03: Input the data to be processed and the sensitivity level into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model.

[0081] In this embodiment, after the sensitivity classification of the data is completed, the data to be processed and its corresponding sensitivity level are used as inputs and output to the desensitization algorithm selection model. The role of the desensitization algorithm selection model is to select the most suitable desensitization algorithm for each data item's sensitivity level. The desensitization algorithm selection model can be based on machine learning and can make an optimal choice from multiple desensitization techniques according to the characteristics and sensitivity level of the data.

[0082] The purpose of this embodiment is to ensure that the data desensitization process is both safe and effective. By considering the specific sensitivity of the data, the desensitization algorithm selection model can customize desensitization strategies for different data items, thereby maximizing the availability of the data while protecting privacy and data security. This method is more flexible and accurate than the traditional "one-size-fits-all" desensitization strategy, which helps to avoid over-desensitization and ensures that the data can still be used for analysis and decision support after processing.

[0083] This embodiment helps to automate and simplify the desensitization process, reduce manual intervention, and reduce the risk of errors and omissions. Through automated desensitization algorithm selection, the system can manage large amounts of data more efficiently, ensure the consistency and compliance of classified and graded protection of sensitive data, and at the same time improve the speed and quality of classified and graded protection of sensitive data.

[0084] In some embodiments, machine learning techniques can be used to learn the associations between various data types, sensitivity levels, and desensitization algorithms by analyzing historical data and desensitization results, and obtain a desensitization algorithm selection model that can understand data with different sensitivity levels and recommend appropriate desensitization strategies for them.

[0085] During specific operations, the data to be processed and its sensitivity level are input into the desensitization algorithm selection model as feature vectors. The desensitization algorithm selection model will evaluate the applicability of each desensitization algorithm to the current data according to the learned patterns and output a selection result, which can be a recommended list of desensitization algorithms, with each algorithm attached with a confidence score indicating the degree to which the model believes the algorithm is applicable to the current data.

[0086] For example, for highly sensitive data containing personal identity information, the model may recommend using strong encryption algorithms; while for low-sensitivity data, it may recommend using lighter desensitization methods such as data generalization or masking. In this way, the system can apply the most appropriate desensitization strategy to data of different sensitivity levels according to the model's recommendations, ensuring the security of data during sharing and use.

[0087] Step S04: Desensitize the data to be processed using the desensitization algorithm corresponding to the selection result.

[0088] In this embodiment, after the desensitization algorithm selection model determines the desensitization algorithm most suitable for the sensitivity level of the data to be processed, the system will automatically apply this desensitization algorithm to process the data to ensure that sensitive information is properly protected. This process can include data masking, encryption, generalization, or anonymization, etc., with the aim of allowing data to be safely used and analyzed to a certain extent while protecting personal privacy.

[0089] In this embodiment, the system will apply the corresponding desensitization rules to each data field according to the model's recommendations, such as encrypting credit card numbers, generalizing address information, or anonymizing email addresses. For example, if the model determines that a field containing a personal name needs to be masked, then this field will be replaced with asterisks or other placeholders to hide the actual information.

[0090] By applying desensitization algorithms that match the data sensitivity, this embodiment can maximize the retention of data availability while protecting sensitive information from being leaked. Since railway data not only contains a large amount of sensitive information but is also an important resource for operation optimization, safety monitoring, and customer service. Therefore, through precise desensitization processing, this embodiment can continue to extract value from the data, promote business development and innovation while ensuring data security and privacy. At the same time, the automated desensitization process can also improve the efficiency of classified and hierarchical protection of sensitive data, reduce errors and delays in manual operations, and ensure the consistency of classified and hierarchical protection of sensitive data.

[0091] Based on the above technical solution, the method for classified and graded protection of sensitive data provided by the embodiments of the present invention performs grading processing on the data to be processed by using a preset sensitive data grading algorithm to obtain the sensitive level of the data to be processed, and inputs the data to be processed and the sensitive level into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model. Finally, the data to be processed is desensitized by using the desensitization algorithm corresponding to the selection result. Through the preset sensitive data grading algorithm, the present invention can accurately identify and grade the sensitivity of the data to be processed, ensuring objectivity and consistency in data classification. By using the desensitization algorithm selection model, the most suitable desensitization algorithm is automatically selected according to the sensitive level of the data. This process not only reduces the error of manual intervention but also improves the flexibility and adaptability of desensitization processing. Finally, by applying the selected desensitization algorithm, sensitive data is effectively protected, the risk of data leakage is prevented, and the usability of the data after processing is ensured, providing a solid foundation for the safe sharing and analysis of railway data.

[0092] Please refer to Figure 2 , for Figure 1 a flowchart showing an actual manifestation of S02 in a method for classified and graded protection of sensitive data provided. In some embodiments, as mentioned in step S02, grading processing is performed on the data to be processed by using a preset sensitive data grading algorithm to obtain the sensitive level of the data to be processed, which may specifically include the following steps:

[0093] Step S11, perform word segmentation processing on the data to be processed based on a natural language processing tool to obtain at least one word unit to be processed.

[0094] In this embodiment, since the data to be processed usually exists in the form of a continuous character sequence, this form is difficult for a computer to directly understand and analyze. Through word segmentation, the data to be processed can be converted into structured data that can be processed by a computer, providing a basis for subsequent sensitive information identification, data grading, and desensitization processing.

[0095] In the context of the railway industry, word segmentation processing helps to identify key information from a large amount of text data, such as train numbers, times, locations, etc. These information are crucial for railway operation, scheduling, and management. Through word segmentation, data can be classified and analyzed more accurately, thereby improving the efficiency and accuracy of classified and graded protection of sensitive data.

[0096] In some embodiments, as mentioned in step S11, performing word segmentation processing on the data to be processed based on a natural language processing tool may specifically be:

[0097] Take the tabular data exported from the railway data system as the original corpus input. Then, apply natural language processing techniques, such as Jieba segmentation, to segment the text fields in the data. Jieba segmentation is an efficient Chinese word segmentation tool that can identify the word boundaries in the text and segment them into individual word units.

[0098] Based on the word segmentation, a skip-gram model can be further used to generate the word vector space of the corpus. The skip-gram model generates a vector representation for each word unit by learning to predict its context words given a center word, and these vectors can capture the semantic relationships between word units. In this process, the word segmentation results can be preprocessed to remove unnecessary information such as numbers, special symbols, English letters, and punctuation marks to reduce data noise and improve the efficiency of model training.

[0099] Furthermore, to improve the word segmentation effect, this embodiment provides a filtering strategy. For example, low-frequency words can be removed to reduce data noise, and words that are difficult to obtain word vectors in the skip-gram model should be avoided being replaced by zero vectors. In addition, for fields with a length of 2, if they lose their actual meaning after word segmentation, no word segmentation is performed. This strategy helps to retain the original meaning of the fields while ensuring the accuracy and practicality of the word segmentation results.

[0100] Through this method, not only can text data be more accurately identified and processed, but also high-quality word units and word vectors can be provided for subsequent sensitive data grading and desensitization processing, thereby improving the efficiency and effect of the entire sensitive data classification and grading protection process.

[0101] Step S12, for each word unit to be processed, calculate the distance between the word unit to be processed and each topic word in the topic word library.

[0102] In this embodiment, by comparing the distance between the word unit to be processed and the topic words in the topic word library, the potential sensitivity of the word unit to be processed can be inferred. The topic words in the topic word library are associated with specific sensitive levels, so the distance can be used as a basis for judging the sensitive level of the word unit to be processed.

[0103] In some embodiments, the Euclidean distance calculation formula can be used to calculate the Euclidean distance between the word unit to be processed and the topic words in the topic word library.

[0104] This embodiment compares the word unit to be processed with the topic words in the existing topic word library, which can quickly assign a sensitive level to the word unit to be processed, thereby providing guidance for subsequent desensitization processing. This not only improves the efficiency of sensitive data classification and grading protection but also helps to ensure the consistency of data protection measures.

[0105] Step S13: Determine that the sensitivity level corresponding to the subject word closest to the word unit to be processed is the sensitivity level of the word unit to be processed.

[0106] In this embodiment, each word unit to be processed is compared with all the subject words in the subject word library to calculate the distance between them, which is determined by the word vector space model. Once the subject word closest to the word unit to be processed is found, the sensitivity level of this subject word is assigned to the word unit to be processed.

[0107] This embodiment determines that words semantically similar have similar sensitivities. Therefore, by comparing the word unit to be processed with the subject words in the subject word library, the sensitivity of the word unit to be processed can be effectively inferred. The advantage of this is that it can automatically assign sensitivity levels to a large amount of data, reducing the need for manual judgment, while also improving the processing speed and consistency.

[0108] Step S14: Determine the probability of the sensitivity level of the data to be processed based on the sensitivity levels of all the word units to be processed.

[0109] Since the sensitivity of a single word unit does not fully represent the sensitivity of the entire data set, in order to more comprehensively evaluate the sensitivity of the data to be processed, this embodiment calculates the probability of the sensitivity level of the data to be processed through the sensitivity level of each word unit. Specifically, for each word unit to be processed, the system determines its most likely sensitivity level based on its distance from the subject words in the subject word library. Then, the system counts the distribution of each sensitivity level in the data to be processed, thereby calculating the attribution probability of the data to be processed for each sensitivity level.

[0110] In some embodiments, the method of calculating the probability of the sensitivity level of the data to be processed may include weighting the sensitivity level of each word unit, or adjusting its contribution to the overall sensitivity according to the frequency of occurrence of the word unit in the data set. For example, if a word unit with a high sensitivity level appears frequently in the data set, its impact on the overall sensitivity of the data to be processed will also be greater.

[0111] In some embodiments, the step S14 mentioned, determining the probability of the sensitivity level of the data to be processed based on the sensitivity levels of all the word units to be processed, may specifically include the following steps:

[0112] Step S21: Generate a predicted probability vector of the sensitivity level of the data to be processed based on the sensitivity levels of all the word units to be processed.

[0113] In this embodiment, the likelihood of the data to be processed at each sensitivity level is evaluated by generating a sensitivity level probability vector. This process involves counting the sensitivity levels of each word unit to be processed in the dataset and then calculating the relative frequency of each sensitivity level in the data to be processed, thereby forming a probability vector. Each element of this vector represents the probability that the dataset belongs to the corresponding sensitivity level.

[0114] The reason for implementing this step is that it can provide a clear perspective for understanding the sensitivity distribution of the entire dataset. By predicting the probability vector, the most likely sensitivity level in the dataset can be quickly identified, enabling the system to automatically evaluate and respond to different levels of sensitive information when processing a large amount of data, thereby improving the efficiency and security of data management.

[0115] Step S22: Obtain the weight of each sensitivity level, and perform weighting and normalization processing on the sensitivity level prediction probability vector of the data to be processed according to the weight to obtain the probability of the sensitivity level of the data to be processed.

[0116] In this embodiment, weights are first assigned to different sensitivity levels, and these weights reflect the importance or risk level of each level in actual applications. Then, these weights are used to adjust the corresponding probability values in the predicted probability vector to ensure that high-risk or important sensitivity levels are appropriately emphasized in the evaluation. Finally, through normalization processing, it is ensured that the sum of all adjusted probability values is 1, thereby obtaining a final probability distribution that accurately reflects the likelihood of each sensitivity level.

[0117] The reason for implementing this step is that it can make the sensitivity evaluation more in line with actual business requirements and risk management strategies. By assigning different weights to sensitivity levels, it can be ensured that the evaluation results more accurately reflect the true sensitivity of the data. Normalization processing ensures the consistency and comparability of probability values, making the final sensitivity level probability more reliable and useful.

[0118] For example, there is a set of data to be processed, which contains five word units to be processed. After sensitivity analysis, these word units are assigned different sensitivity levels: two word units are rated as level 1 sensitive, two word units are rated as level 2 sensitive, and one word unit is rated as level 3 sensitive.

[0119] To more accurately reflect the contribution of these word units to the sensitivity of the entire dataset, this embodiment assigns different weights to each sensitivity level to account for the imbalance of data at each level in the thesaurus. Specifically, the weight of level 1 sensitivity is set to 2, the weight of level 2 sensitivity is set to 3, and the weight of level 3 sensitivity is set to 5. Such a weight assignment reflects that level 3 sensitivity is relatively rare in the dataset, so its impact on the overall sensitivity of the dataset is greater.

[0120] Next, in this embodiment, the predicted probability vector is weighted. For the two word units that are sensitive at level 1, the weighted probability is calculated as 2×2 = 4; for the two word units that are sensitive at level 2, the weighted probability is 3×2 = 6; for the one word unit that is sensitive at level 3, the weighted probability is 5×1 = 5.

[0121] Finally, in this embodiment, these weighted probabilities are normalized to obtain the final probabilities of the data to be processed at each sensitivity level. The normalization process is done by dividing each weighted probability by the sum of all weighted probabilities. In this example, the sum is 4 + 6 + 5 = 15. Therefore, the final probability of the data to be processed being sensitive at level 1 is 4 / 15, at level 2 is 6 / 15, and at level 3 is 5 / 15.

[0122] Based on step S21 and step S22, this embodiment can comprehensively consider the sensitivity of each word unit in the dataset and generate a probability vector that comprehensively reflects the sensitivity distribution of the entire dataset. By assigning weights to different sensitivity levels, the importance and influence of each level in the dataset can be more accurately reflected. In addition, normalizing the predicted probability vector ensures the rationality and consistency of the probability values, making the finally obtained probabilities truly reflect the sensitivity of the dataset. This method not only improves the accuracy of sensitivity analysis but also provides a more reliable basis for subsequent desensitization processing.

[0123] Step S15: Determine whether the difference between the highest value among all probabilities and other values is within the error tolerance threshold range.

[0124] If yes, execute step S16; if not, execute step S17.

[0125] In this embodiment, step S15 is used to determine whether the sensitivity level of the data to be processed is clear enough for desensitization processing. In this step, the system compares the gap between the highest probability value and other probability values in the sensitivity level probability vector. If this gap is greater than or equal to the preset error tolerance threshold, it indicates that the sensitivity level of the data is relatively clear and desensitization processing can be carried out accordingly.

[0126] The purpose of this embodiment is to ensure that the desensitization strategy is implemented based on the most likely sensitivity level while avoiding over - or under - desensitization when the sensitivity level is unclear. This method helps to improve the accuracy and efficiency of desensitization processing, ensuring that while protecting privacy, the value and usability of the data are maximized. In addition, by setting the error tolerance threshold, this method also allows a certain degree of flexibility to adapt to the precision requirements for sensitivity judgment in different scenarios.

[0127] Step S16, determine the sensitivity level corresponding to the highest value in the probabilities as the sensitivity level of the data to be processed.

[0128] In this embodiment, if the probability gap is within an acceptable error range, the system will execute Step S16 and directly adopt the sensitivity level with the highest probability value as the final sensitivity level of the data to be processed.

[0129] Step S17, use the hierarchical clustering algorithm to perform hierarchical clustering on each word unit to be processed in the thesaurus, and determine the number of theme words with different sensitivity levels contained in the cluster where the data to be processed is located according to the hierarchical clustering result.

[0130] If the difference between the highest probability value and other probability values is not within the preset error tolerance threshold range, this indicates that the sensitivity level of the data is not clear enough, and there may be multiple sensitivity levels with similar possibilities, making it difficult to directly determine.

[0131] In this embodiment, if the direct probability comparison fails to clarify the sensitivity level of the data, the system will use the hierarchical clustering algorithm to perform a more detailed classification of the word units. Hierarchical clustering is a data analysis method that can group data points to form clusters at different levels. Through this method, it can be identified which theme words in the thesaurus are more similar to the word units in the data to be processed, thereby determining the cluster to which the data belongs.

[0132] The reason for implementing this step is that hierarchical clustering can provide a more detailed sensitivity analysis to help identify potential patterns and associations in the data. By determining the number of theme words with different sensitivity levels in the cluster where the data is located, the overall sensitivity of the data can be evaluated more accurately. This method helps to reveal the internal structure of the data through cluster analysis when the sensitivity level is not clear, thereby providing a more reliable basis for the desensitization process.

[0133] Step S18, determine the sensitivity level of the theme word with the largest quantity as the sensitivity level of the data to be processed.

[0134] In this embodiment, the system will count the number of theme words with each sensitivity level in the cluster where the data to be processed is located, and use the sensitivity level with the highest frequency as the sensitivity level of the data to be processed.

[0135] The reason for implementing this step is that through hierarchical clustering, the data is organized into different clusters, and the theme words within each cluster are semantically closer. Within the cluster, if the number of theme words with a certain sensitivity level is the largest, this indicates that the data to be processed is semantically most relevant to this sensitivity level. Therefore, determining this sensitivity level as the sensitivity level of the data to be processed can more accurately reflect the sensitivity characteristics of the data.

[0136] Based on the above technical solution, in this embodiment, by adopting advanced natural language processing tools and hierarchical clustering algorithms, the automatic and accurate identification of the sensitivity level of railway data to be processed is realized. This embodiment can meticulously analyze each word unit in the data, and by comparing with the thesaurus, determine the sensitivity level of each word unit, thereby more accurately reflecting the sensitivity of the entire data to be processed. This method not only improves the efficiency of sensitive data identification, but also enhances the accuracy and reliability of classification by considering the distance between the word unit and the thesaurus terms and the hierarchical clustering results.

[0137] In addition, by judging whether the difference between the highest value and other values in the probability is within the error tolerance threshold range, this method can intelligently decide whether to directly adopt the most likely sensitivity level or further hierarchical clustering analysis is required. This flexibility makes the classification process more robust and can adapt to data to be processed with different complexities. Finally, by determining the sensitivity level of the thesaurus terms with the largest number in the cluster where the data is located, this method provides a clear and reasonable sensitivity level for the data to be processed, which is crucial for subsequent desensitization processing and the formulation of data protection strategies.

[0138] Please refer to Figure 3 , which is a flowchart of the establishment process of a thesaurus provided by an embodiment of the present invention. In some embodiments, for the thesaurus mentioned in step S12, its establishment process may specifically include the following steps:

[0139] Step S31, process the input original corpus based on natural language processing tools to obtain and output the word vector space of the original corpus.

[0140] Since the original corpus usually exists in text form, and the computer cannot directly understand the context and semantics of natural language, it is difficult to directly analyze and process the text.

[0141] Therefore, in this embodiment, natural language processing tools (such as Jieba segmentation and skip-gram model) are used to convert the original text data into a numerical form for the computer to understand and process. By converting the text into a word vector space, the text data can be transformed into structured numerical data, so that machine learning algorithms can be applied for further analysis, such as pattern recognition, classification, and clustering. In addition, the word vector space model can capture the semantic relationships between words, and thus provide a quantitative way to measure the similarity and difference of words.

[0142] Step S32, in response to the input first selection instruction, select the initial thesaurus terms in the word vector space.

[0143] In this embodiment, after the user issues a first selection instruction, the system can identify and select the most representative initial topic words in the pre-constructed word vector space. These initial topic words are the basis for subsequent sensitivity analysis and clustering processes, and they can be related to specific topics or concepts in railway data, such as "train scheduling", "passenger information", or "maintenance records", etc.

[0144] The reason for implementing this step is that the selection of initial topic words is crucial for building an accurate and useful thesaurus. By carefully selecting words related to key aspects of railway data, it can be ensured that the thesaurus can comprehensively cover all aspects of railway data. In addition, these initial topic words will serve as the initial centroids of the K-means clustering algorithm, helping the algorithm to identify and classify other relevant word units in the word vector space, thus expanding the thesaurus and providing a basis for the annotation of sensitive levels.

[0145] Step S33, annotate the sensitive levels of the initial topic words through the analytic hierarchy process to generate an initial thesaurus.

[0146] In this embodiment, by using the analytic hierarchy process as a decision-making tool to determine the sensitivity levels of the initial topic words, and accordingly construct a thesaurus containing sensitive level information. The analytic hierarchy process decomposes the problem into different levels of criteria and alternative solutions, and quantifies the relative importance of each factor through pairwise comparisons, so as to assign a sensitive level to each initial topic word.

[0147] The analytic hierarchy process can systematically handle complex decision-making problems, making the annotation process of sensitive levels more objective and quantitative. Through this method, it can be ensured that each word in the thesaurus is assigned an appropriate level according to its sensitivity in railway data, such as "public", "internal", "confidential", etc. Such annotation of sensitive levels helps to take corresponding protection measures according to the sensitivity of the data in the subsequent classification, grading protection and de-sensitization processes of sensitive data.

[0148] In some embodiments, according to expert experience, 2% of the representative word units can be screened out from the railway industry dataset as the initial topic words. These word units are carefully selected to ensure that they can cover the key aspects of railway data, such as "train schedule", "ticket information", "customer name", etc. Then, using the analytic hierarchy process, these initial topic words are assigned sensitive levels according to their importance in data protection and privacy regulations. For example, "customer name" may be annotated as a high sensitive level, while "train schedule" may be annotated as a low sensitive level. After completing the annotation of the sensitive levels of the initial topic words, this embodiment constructs an initial thesaurus, which contains the word units and their corresponding sensitive levels.

[0149] This embodiment combines expert experience and unsupervised learning algorithms to construct a thesaurus in a semi-supervised manner. This method overcomes the limitations of completely unsupervised methods in sensitive level classification and reduces the dependence on a large amount of expert-labeled data. By relying on a small amount of expert-labeled data, the classification model in this embodiment is more authoritative and can more accurately reflect the sensitivity of railway data, providing an effective solution for data protection and privacy management in the railway industry.

[0150] Step S34: Using the K-means clustering algorithm, with the initial topic words as the initial centroids of clustering, calculate the distance between each word unit in the word vector space and the initial topic words.

[0151] In this embodiment, the K-means algorithm takes the initial topic words as the starting points of the clustering process, that is, the centroids, and then searches for the word units in the word vector space that are closest to these centroids. The K-means clustering algorithm can automatically identify patterns and groups in the word vector space, thus helping to expand the thesaurus. By taking the initial topic words as the centroids, the algorithm can identify the word units that are semantically similar to these topic words, and these word units have similar sensitivity characteristics to the topic words.

[0152] In some embodiments, for the calculation of the distance between each word unit in the word vector space and the initial topic words mentioned in step S34, specifically:

[0153] Calculate the Euclidean distance between the word unit and the initial topic word according to the formula:

[0154]

[0155] where x i is the i-th coordinate of word unit a in the n-dimensional space, y i is the i-th coordinate of the initial word unit b in the n-dimensional space, and S Euclidean (a, b) is the Euclidean distance between word unit a and the initial topic word b.

[0156] Step S35: Determine the word unit in the word vector space that is closest to the initial topic word as the word to be added, and determine the sensitive level of the initial topic word as the sensitive level corresponding to the word to be added.

[0157] In this embodiment, the system first identifies the word units in the word vector space that are closest to each initial topic word, and these word units are considered to be the most semantically similar to the initial topic words. Then, the system sets the sensitive levels of these words to be added as the sensitive levels of the initial topic words closest to them.

[0158] This embodiment utilizes semantic proximity in the word vector space to infer the sensitivity of word units. This method assumes that for words that are semantically similar, their sensitivities are likely to be similar. By matching the sensitivity level of the word to be added with that of its closest initial topic word, the consistency and accuracy of the sensitivity level can be ensured. In addition, this method simplifies the process of labeling sensitivity levels and does not require separate sensitivity assessment for each newly identified word unit.

[0159] Step S36: Add the word to be added and its corresponding sensitivity level to the initial topic word library to obtain the topic word library.

[0160] In this embodiment, by incorporating the word to be added into the library, the topic word library can more comprehensively reflect various sensitive information in railway data, thereby enriching the content of the topic word library and improving its coverage and accuracy. At the same time, since the sensitivity levels of these new word units are determined based on semantic similarity to the initial topic words, the consistency and reliability of sensitivity labeling for the entire topic word library can be ensured.

[0161] Based on the above technical solution, this embodiment can automatically expand the topic word library. By using the K-means clustering algorithm to identify word units similar to the initial topic words and assigning corresponding sensitivity levels, it not only improves the efficiency of constructing the topic word library, but also enhances the consistency and accuracy of sensitivity level labeling through quantitative clustering analysis.

[0162] Please refer to Figure 4 , which is a flowchart of the training process of a desensitization algorithm selection model provided by an embodiment of the present invention. In some embodiments, for the desensitization algorithm selection model mentioned in step S03, its training process may specifically include the following steps:

[0163] Step S41: Receive the input training set.

[0164] In this embodiment, the training set includes training data and the corresponding desensitization algorithm and parameters for the training data. The training data provides the basic instances for the model to learn, while the corresponding desensitization algorithm and parameters provide the correct learning targets for the model. In this way, the model can continuously adjust its internal parameters during the training process to better understand and predict which desensitization algorithm and parameters should be used for data with different characteristic attributes.

[0165] Based on the above embodiments, in some embodiments, after performing step S02 to perform hierarchical processing on the data to be processed using a preset sensitive data grading algorithm to obtain the sensitivity level of the data to be processed, the following steps can also be performed to establish the training set:

[0166] Step S51: Output the sensitivity level of the data to be processed.

[0167] Step S52: In response to the input second selection instruction, select training data from the data to be processed.

[0168] In this embodiment, to improve the training accuracy of the model, the user can divide the output results based on expert experience or industry requirements. Based on the feature distribution and sensitivity analysis of the data set, a small amount of data is used for the training of model construction. A small amount of representative training data can effectively improve the model training efficiency, avoid the cost and time consumption brought by large-scale data collection. These data can cover various types of sensitive information, including character type, numerical type, and date type fields, and at the same time reflect the characteristics of different sensitive levels in the data set.

[0169] This data set is screened by experts to ensure its general applicability and representativeness in the selection of desensitization methods. The remaining data is used for actual deployment testing, so as to achieve desensitization adaptation of multiple data with a small amount of data, and improve the actual work efficiency.

[0170] Step S53: In response to the input annotation instruction, annotate the corresponding desensitization algorithm and parameters for the training data to obtain a training set.

[0171] After obtaining the sensitive level of the data to be processed in this embodiment, by outputting the sensitive level, the user can clearly understand the sensitivity of its data. Based on this, select training data and annotate desensitization algorithms and parameters for these data, and a high-quality training set can be constructed.

[0172] Step S42: According to the preset desensitization rule library and the sensitive level of the training data, extract the feature attributes of the training data.

[0173] In this embodiment, the feature attributes include at least one of the data name, data sensitive level, and data attribute type. During the model training process, the system extracts key features from the original training data according to the predefined desensitization rules and the sensitivity evaluation results of the training data. Feature extraction is a key link in the training of machine learning models, which determines what information the model can learn. By carefully selecting features closely related to the desensitization algorithm selection, the prediction accuracy and efficiency of the model can be improved. The preset desensitization rule library provides a framework for guiding the feature selection and extraction process to ensure that the selected features are closely related to the desensitization strategy.

[0174] This embodiment can enable the model to focus more on the data features that have a substantial impact on the desensitization decision. In this way, the model can more effectively learn the association between different feature attributes and desensitization algorithms, so that when facing new data to be processed, it can quickly and accurately recommend the most suitable desensitization algorithm.

[0175] In some embodiments, the desensitization rule library can be a desensitization rule library constructed by selecting appropriate desensitization algorithms for different types and requirements of data based on expert experience.

[0176] In the current desensitization work, most of the desensitization methods for different data in different application scenarios are determined by experts according to specifications. Therefore, in this embodiment, a desensitization rule library is constructed relying on experts' experience and suggestions. This rule library includes a variety of desensitization algorithms (such as masking, generalization, encryption, etc.), and detailed usage standards and parameter configurations are formulated for different data types, sensitive levels, and application scenarios. Through dynamic adjustment, this desensitization rule library can be extended according to different requirements, and at the same time, new rule updates can be completed to adapt to new usage scenarios, ensuring wide applicability in complex data scenarios. For example, a desensitization rule library can be as shown in the following table:

[0177]

[0178] Step S43: Determine the desensitization algorithm and parameters corresponding to the training data as the target label, and perform random forest classification based on the characteristic attributes of the training data and the target label to train the desensitization algorithm selection model.

[0179] In this embodiment, the optimal desensitization algorithm and its parameters of each training data sample are set as the output target of the model, that is, the target label. At the same time, the characteristic attributes of the training data, such as data name, sensitive level, and data attribute type, are used as the input features of the model. Through the ensemble learning method of random forest, the model learns the mapping relationship between the input features and the target label, thereby training a desensitization algorithm selection model that can predict the best desensitization algorithm and parameters.

[0180] As a powerful classification algorithm, random forest can handle high-dimensional data and capture the complex relationships between features. By transforming the problem of desensitization algorithm selection into a classification problem, the model can automatically learn the association between different characteristic attributes and desensitization algorithms, thereby realizing automated desensitization algorithm recommendation for new data. This embodiment can improve the efficiency and accuracy of the desensitization process, reduce manual intervention, and at the same time ensure the consistency and compliance of the desensitization strategy.

[0181] Based on the above embodiments, in some embodiments, for what is mentioned in step S03, input the data to be processed and the sensitive level into the preset desensitization algorithm selection model, and obtain the selection result output by the desensitization algorithm selection model, which can specifically be:

[0182] Use the desensitization algorithm selection model to calculate the class probability of each desensitization algorithm

[0183]

[0184] Among them, P(y = c|X) represents the probability that the data X to be processed belongs to the desensitization algorithm c, N represents the total number of trees in the random forest, and h i (X) represents the predicted class of the i-th tree for the data X to be processed, and I(·) is the indicator function.

[0185] In some embodiments, after the training of the desensitization algorithm selection model is completed, the model can also be tested according to the input test set, where the test set includes test data and the corresponding desensitization algorithms and parameters. The test process may include:

[0186] 1) Preprocess the test data to ensure that the test data has the same feature attributes as the training data, such as data name, sensitivity level, and data attribute type, etc.

[0187] 2) Input the preprocessed test data into the trained desensitization adaptive model. The model predicts the desensitization algorithm and parameters most suitable for each data sample according to the feature attributes of the test data.

[0188] 3) Subsequently, compare the prediction results with the corresponding desensitization algorithms and parameters of the test data. By comparing the prediction results of the model with the actual results, the accuracy and effectiveness of the model can be evaluated.

[0189] 4) If there are deviations between the predictions of the model and the actual desensitization results, these deviations can be used as feedback for further optimizing the model. Through this iterative test and optimization process, the performance of the model can be continuously improved to make it more accurately adapt to different data and desensitization requirements.

[0190] Based on the above technical solutions, this embodiment enables the model to learn the correlation between different data features and the best desensitization algorithms. By extracting features according to the preset desensitization rule library and the sensitivity level of the training data, the model can identify the key attributes of the data. Setting the desensitization algorithms and parameters as the target labels, the model can learn how to predict the most suitable desensitization algorithms and parameters according to the feature attributes of the data through the random forest classification algorithm. This learning method based on data features and desensitization rules improves the prediction accuracy and generalization ability of the model, enabling it to adapt to different data types and sensitivity requirements.

[0191] Please refer to Figure 5 , which is a schematic structural diagram of a sensitive data classification and grading protection device provided by an embodiment of the present invention. The sensitive data classification and grading protection device may include:

[0192] An acquisition module 100, configured to acquire data to be processed;

[0193] The classification processing module 200 is used to perform classification processing on the data to be processed by using a preset sensitive data classification algorithm to obtain the sensitive level of the data to be processed;

[0194] The desensitization algorithm selection module 300 is used to input the data to be processed and the sensitive level into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model;

[0195] The desensitization processing module 400 is used to perform desensitization processing on the data to be processed by using the desensitization algorithm corresponding to the selection result.

[0196] Based on the above embodiments, in a specific embodiment, the classification processing module 200 may specifically be used for:

[0197] Performing word segmentation processing on the data to be processed based on a natural language processing tool to obtain at least one word unit to be processed;

[0198] For each word unit to be processed, calculate the distance between the word unit to be processed and each subject word in the subject word library; the corresponding relationship between the subject word and the sensitive level is stored in the subject word library;

[0199] Determine that the sensitive level corresponding to the subject word with the closest distance to the word unit to be processed is the sensitive level of the word unit to be processed;

[0200] Determine the probability of the sensitive level of the data to be processed according to the sensitive levels of all word units to be processed;

[0201] Judge whether the difference between the highest value among all probabilities and other values is within the error tolerance threshold range;

[0202] If so, determine the sensitive level corresponding to the highest value in the probability as the sensitive level of the data to be processed;

[0203] If not, perform hierarchical clustering on each word unit to be processed in the subject word library by using the hierarchical clustering algorithm, and determine the number of subject words with different sensitive levels contained in the cluster where the data to be processed is located according to the hierarchical clustering result;

[0204] Determine the sensitive level with the largest number of subject words as the sensitive level of the data to be processed.

[0205] Based on the above embodiments, in a specific embodiment, the classification processing module 200 may specifically be used for:

[0206] Generate a sensitive level prediction probability vector of the data to be processed according to the sensitive levels of all word units to be processed;

[0207] Obtain the weights for each sensitivity level, and perform weighting and normalization on the predicted probability vector of the sensitivity level of the data to be processed according to the weights to obtain the probability of the sensitivity level of the data to be processed.

[0208] Based on the above embodiments, in a specific embodiment, the classification processing module 200 may specifically be used for:

[0209] Process the input original corpus based on natural language processing tools to obtain and output the word vector space of the original corpus;

[0210] In response to the input first selection instruction, select the initial topic words in the word vector space;

[0211] Annotate the sensitivity levels of the initial topic words through the analytic hierarchy process to generate an initial topic word library;

[0212] Use the K-means clustering algorithm with the initial topic words as the initial centroids of the clustering, and calculate the distances between each word unit in the word vector space and the initial topic words;

[0213] Determine the word units closest to the initial topic words in the word vector space as the words to be added, and determine the sensitivity level of the initial topic words as the sensitivity level corresponding to the words to be added;

[0214] Add the words to be added and the corresponding sensitivity levels of the words to be added to the initial topic word library to obtain the topic word library.

[0215] Based on the above embodiments, in a specific embodiment, the desensitization algorithm selection module 300 may specifically be used for:

[0216] Calculate the Euclidean distance between the word unit and the initial topic word according to the formula:

[0217]

[0218] where x i is the i-th coordinate of the word unit a in the n-dimensional space, y i is the i-th coordinate of the initial word unit b in the n-dimensional space, and S Euclidean (a, b) is the Euclidean distance between the word unit a and the initial topic word b.

[0219] Based on the above embodiments, in a specific embodiment, the desensitization algorithm selection module 300 may specifically be used for:

[0220] Receive the input training set; the training set includes training data and the corresponding desensitization algorithms and parameters for the training data;

[0221] Extract features from the training data according to the preset desensitization rule library and the sensitivity level of the training data to obtain the feature attributes of the training data; the feature attributes include at least one of the data name, data sensitivity level, and data attribute type;

[0222] Determine the desensitization algorithm and parameters corresponding to the training data as the target label, and perform random forest classification based on the feature attributes of the training data and the target label to train the desensitization algorithm selection model.

[0223] Based on the above embodiments, in a specific embodiment, the desensitization algorithm selection module 300 can specifically be used for:

[0224] Use the desensitization algorithm selection model to calculate the class probability of each desensitization algorithm

[0225]

[0226] where P(y = c∣X) represents the probability that the data X to be processed belongs to the desensitization algorithm c, N represents the total number of trees in the random forest, and h i (X) represents the predicted class of the data X to be processed by the i-th tree, and I(·) is the indicator function.

[0227] Based on the above embodiments, in a specific embodiment, the desensitization processing module 400 can specifically be used for:

[0228] Output the sensitivity level of the data to be processed;

[0229] In response to the input of the second selection instruction, select training data from the data to be processed;

[0230] In response to the input of the annotation instruction, annotate the corresponding desensitization algorithm and parameters for the training data to obtain the training set.

[0231] This embodiment provides an electronic device, including a processor and a memory. The memory is used to store at least one instruction, and when the instruction is loaded and executed by the processor, it implements the above method for classifying and grading the protection of sensitive data. Its execution manner and beneficial effects are similar and will not be elaborated here.

[0232] The embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it implements the above method for classifying and grading the protection of sensitive data. Its execution manner and beneficial effects are similar and will not be elaborated here.

[0233] It should be noted that although the above steps are described in a specific order, it does not mean that the above steps must be executed in the specific order. In fact, some of these steps can be executed concurrently or even changed in order, as long as the required functions can be achieved.

[0234] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for classifying and hierarchically protecting sensitive data, characterized in that Including: Obtain the data to be processed; Use a preset sensitive data grading algorithm to grade the data to be processed, and obtain the sensitive level of the data to be processed; Input the data to be processed and the sensitive level into a preset desensitization algorithm selection model, and obtain the selection result output by the desensitization algorithm selection model; Use the desensitization algorithm corresponding to the selection result to desensitize the data to be processed.

2. The method according to claim 1, wherein The step of using a preset sensitive data grading algorithm to grade the data to be processed and obtain the sensitive level of the data to be processed includes: Perform word segmentation on the data to be processed based on a natural language processing tool to obtain at least one to-be-processed word unit; For each to-be-processed word unit, calculate the distance between the to-be-processed word unit and each subject word in the subject word library; the corresponding relationship between the subject word and the sensitive level is stored in the subject word library; Determine the sensitive level corresponding to the subject word closest to the to-be-processed word unit as the sensitive level of the to-be-processed word unit; Determine the probability of the sensitive level of the data to be processed according to the sensitive levels of all the to-be-processed word units; Judge whether the difference between the highest value among all the probabilities and other values is within the error tolerance threshold range; If so, determine the sensitive level corresponding to the highest value among the probabilities as the sensitive level of the data to be processed; If not, use the hierarchical clustering algorithm to perform hierarchical clustering on each to-be-processed word unit in the subject word library, and determine the number of subject words with different sensitive levels contained in the cluster where the data to be processed is located according to the hierarchical clustering result; Determine the sensitive level with the largest number of subject words as the sensitive level of the data to be processed.

3. The method according to claim 2, wherein The step of determining the probability of the sensitive level of the data to be processed according to the sensitive levels of all the to-be-processed word units includes: Generate a sensitive level prediction probability vector of the data to be processed according to the sensitive levels of all the to-be-processed word units; Obtain the weight of each sensitive level, and perform weighting and normalization processing on the sensitive level prediction probability vector of the data to be processed according to the weight to obtain the probability of the sensitive level of the data to be processed.

4. The method according to claim 2, characterized in that, The establishment process of the subject word library includes: Process the input original corpus based on a natural language processing tool, and obtain and output the word vector space of the original corpus; In response to the input first selection instruction, select an initial subject word in the word vector space; Mark the sensitive level of the initial subject word through the analytic hierarchy process to generate an initial subject word library; Use the K-means clustering algorithm with the initial subject word as the initial centroid of clustering, and calculate the distance between each word unit in the word vector space and the initial subject word; Determine the word unit closest to the initial subject word in the word vector space as the to-be-added word, and determine the sensitive level of the initial subject word as the sensitive level corresponding to the to-be-added word; Add the to-be-added word and the sensitive level corresponding to the to-be-added word to the initial subject word library to obtain the subject word library.

5. The method according to claim 4, characterized in that, Calculating the distance between each word unit in the word vector space and the initial topic word includes: Calculating the Euclidean distance between the word unit and the initial topic word according to the formula: Among them, x i is the i-th coordinate of the word unit a in the n-dimensional space, and y i is the i-th coordinate of the initial word unit b in the n-dimensional space. S Euclidean (a, b) is the Euclidean distance between the word unit a and the initial topic word b.

6. The method according to any one of claims 1-5, characterized in that, The training process of the desensitization algorithm selection model includes: Receiving the input training set; the training set includes training data and the desensitization algorithm and parameters corresponding to the training data; Extracting features from the training data according to a preset desensitization rule library and the sensitive level of the training data to obtain the feature attributes of the training data; the feature attributes include at least one of the data name, data sensitive level, and data attribute type; Determining the desensitization algorithm and parameters corresponding to the training data as the target label, and training the desensitization algorithm selection model by performing random forest classification based on the feature attributes of the training data and the target label.

7. The method according to claim 6, characterized in that, Inputting the data to be processed and the sensitive level into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model includes: Calculating the class probability of each desensitization algorithm using the desensitization algorithm selection model Among them, P(y = c|X) represents the probability that the data X to be processed belongs to the desensitization algorithm c, N represents the total number of trees in the random forest, and h i (X) represents the predicted category of the i-th tree for the data X to be processed, and I(·) is the indicator function.

8. The method according to claim 6, wherein After classifying the data to be processed using a preset sensitive data classification algorithm to obtain the sensitive level of the data to be processed, the method further includes: Outputting the sensitive level of the data to be processed; In response to the input of a second selection instruction, selecting the training data from the data to be processed; In response to the input of a labeling instruction, labeling the corresponding desensitization algorithm and parameters for the training data to obtain the training set.

9. A device for classifying and hierarchically protecting sensitive data, characterized in that, Includes: An acquisition module for acquiring data to be processed; A classification processing module for classifying the data to be processed using a preset sensitive data classification algorithm to obtain the sensitive level of the data to be processed; A desensitization algorithm selection module for inputting the data to be processed and the sensitive level into a preset desensitization algorithm selection model to obtain the selection result output by the desensitization algorithm selection model; A desensitization processing module for desensitizing the data to be processed using the desensitization algorithm corresponding to the selection result.

10. An electronic device, characterized in that, Includes: A processor and a memory, the memory is used to store at least one instruction, and when the instruction is loaded and executed by the processor, it realizes the method for classifying and grading the protection of sensitive data as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the method for classifying and grading the protection of sensitive data as described in any one of claims 1-8.