A data security supervision method for a public data operation platform
By using machine learning-based sensitive data automatic identification algorithm and dynamic desensitization technology on the public data operation platform, combined with the desensitization anomaly detection model, the problem of sensitive data identification and protection in the platform is solved, and efficient and accurate data privacy protection and task stability are achieved.
Patent Information
- Application Number
- CN202510355687.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In public data operation platforms, existing data security management technologies are difficult to effectively identify and process sensitive data, resulting in possible risks of sensitive data leakage, privacy violations and legal compliance.
The data set is scanned by machine learning-based automatic data recognition algorithm, accurately locate sensitive data, and dynamically adapt to desensitization technology based on data sensitivity and business attributes. A desensitization abnormality detection model is constructed, and the desensitization process is monitored in real time and the recovery mechanism is automatically triggered through extreme distribution information, desensitization intensity imbalance information and data transmission link interrupt information.
It realizes efficient identification and processing of sensitive data in public data operation platforms, ensures that the data meets privacy protection requirements during the sharing process, reduces the risks of data leakage and privacy violations, and improves the stability and security of data desensitization tasks.
Smart Images

Figure CN119885281B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security supervision, and more specifically, the present invention relates to a data security supervision method for a public data operation platform. Background Art
[0002] With the wide application of public data operation platforms in fields such as smart cities, financial services, and healthcare, a large amount of data plays a core role in the sharing and circulation across departments and institutions. However, this data often contains sensitive information (such as personal identity information, transaction records, health records, etc.), and if not effectively protected during data processing and sharing, it will face severe challenges such as sensitive data leakage, privacy infringement, and even legal compliance risks.
[0003] In data security management, data desensitization technology is widely adopted. By masking, forging, or differential privacy processing of sensitive information, data can meet business requirements during sharing while ensuring privacy protection. However, the desensitization process is not foolproof, and a series of technical problems may be encountered during its implementation. For example, data incompleteness or abnormal distribution may lead to incorrect application of desensitization rules. At the same time, deficiencies in algorithm design and implementation (such as uneven desensitization intensity or improper noise parameter settings) may result in insufficient protection intensity or failure of business logic. In addition, computational resource limitations and transmission interruptions at the system level may also lead to desensitization task failures or even direct exposure of sensitive data. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a data security supervision method for a public data operation platform to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A data security supervision method for a public data operation platform includes the following steps:
[0007] Step S1, using a machine learning-based sensitive data automatic recognition algorithm to scan the dataset of the public data operation platform to accurately locate sensitive data;
[0008] Step S2, classifying the located sensitive data according to data sensitivity and business attributes, and dynamically adapting data desensitization technology according to the classification results;
[0009] Step S3, obtaining extreme distribution information of desensitized data, uneven desensitization intensity information, and data transmission link interruption information;
[0010] Step S4: Build a desensitization anomaly detection model, output a desensitization anomaly detection index, identify potential anomalies during the data desensitization process, and automatically trigger a recovery mechanism for detected desensitization anomalies to reschedule desensitization tasks and optimize parameters.
[0011] In a preferred embodiment, extract a dataset from the public data operation platform, including structured data and unstructured data; conduct data exploration on the field distribution, field names, and types of the data, and process the dataset in chunks according to the data type, where the data types include tabular data, text data, and image data;
[0012] Perform explicit feature recognition and implicit feature mining on the chunk-processed dataset to extract sensitive data features;
[0013] For structured data: Select a traditional machine learning model to identify and locate sensitive data, train the traditional machine learning model according to the extracted sensitive data features, and use the trained traditional machine learning model to scan the entire dataset to accurately identify the fields and records of sensitive data and output the sensitivity score of the sensitive data;
[0014] For unstructured data, select a deep learning model to identify and locate sensitive data, train the deep learning model according to the extracted sensitive data features, and use the trained deep learning model to scan the entire dataset to accurately identify the fields and records of sensitive data and output the sensitivity score of the sensitive data.
[0015] In a preferred embodiment, divide sensitive data into high-sensitivity data, medium-sensitivity data, and low-sensitivity data according to the data sensitivity score, as follows:
[0016] Compare the data sensitivity score with the sensitivity stage thresholds, where the sensitivity stage thresholds include a first sensitivity threshold and a second sensitivity threshold, and the first sensitivity threshold is less than the second sensitivity threshold;
[0017] If the data sensitivity score is less than or equal to the first sensitivity threshold, mark the located sensitive data as low-sensitivity data; if the data sensitivity score is greater than the first sensitivity threshold and less than the second sensitivity threshold, mark the located sensitive data as medium-sensitivity data; if the data sensitivity score is greater than or equal to the second sensitivity threshold, mark the located sensitive data as high-sensitivity data;
[0018] Classify sensitive data into business types according to business attributes, where the business attributes include business scenarios, data usage, and processing requirements; the classification results of business types include but are not limited to financial business, medical business, and marketing business;
[0019] Dynamically adapt the data desensitization technology according to the classification results of data sensitivity scores and the classification results of business types.
[0020] In a preferred embodiment, the extreme distribution information of the desensitized data includes an extreme distribution coefficient , the desensitization intensity imbalance information includes a desensitization intensity imbalance coefficient , and the data transmission link interruption information includes a data transmission link interruption coefficient .
[0021] In a preferred embodiment, the acquisition logic of the extreme distribution coefficient is as follows:
[0022] Obtain the desensitized data set and calculate the mean value of the desensitized data set , and the expression is as follows , where represents the storage byte number of the i-th desensitized data in the desensitized data set, , is a positive integer; calculate the standard deviation of the desensitized data set , and the expression is as follows ;
[0023] Perform extreme value detection on the desensitized data in the desensitized data set according to the mean value and standard deviation of the desensitized data set, specifically as follows: If , then mark the i-th desensitized data as an extreme value;
[0024] Calculate the skewness of the desensitized data set , and the expression is as follows ;
[0025] Calculate the kurtosis of the desensitized data set , and the expression is as follows ;
[0026] Calculate the extreme distribution coefficient , and the expression is as follows , where represents the extreme value ratio, and the calculation expression is as follows , where represents the number of extreme values.
[0027] In a preferred embodiment, the acquisition logic of the desensitization intensity imbalance coefficient is as follows:
[0028] Obtain the change amount of each desensitization operation , and calculate the desensitization intensity , and the expression is as follows , where represents the original amount before desensitization;
[0029] Calculate the Gini coefficient of the desensitization intensity , the expression is as follows , where represents the desensitization intensity of the nth desensitized data, , is a positive integer;
[0030] Calculate the KL divergence of the desensitization intensity, and the expression is as follows , where represents the KL divergence of the desensitization intensity, represents the expected desensitization intensity of the nth desensitized data;
[0031] Calculate the desensitization intensity imbalance coefficient , the expression is as follows , where respectively represent the preset proportionality coefficients of the Gini coefficient and the KL divergence of the desensitization intensity, and are both greater than 0.
[0032] In a preferred embodiment, the acquisition logic of the data transmission link interruption coefficient is as follows:
[0033] Use a network monitoring tool to monitor each interruption event in the data transmission process, and record the number of link interruption events, the duration of link interruption, the amount of data loss, and the number of retransmitted data packets;
[0034] Calculate the link interruption frequency according to the number of link interruption events , the expression is as follows , where represents the number of link interruption events, represents the total data transmission time;
[0035] Calculate the proportion of the link interruption duration according to the link interruption duration , the expression is as follows , where represents the link interruption duration of the mth interruption event, , is a positive integer;
[0036] Calculate the data loss rate according to the amount of data loss , the expression is as follows , where represents the amount of data loss, represents the total number of transmitted data packets;
[0037] Calculate the retransmission rate according to the number of retransmitted data packets , the expression is as follows , where represents the number of retransmitted data packets;
[0038] Calculate the data transmission link interruption coefficient according to the link interruption frequency, the proportion of link interruption duration, the data loss rate, and the retransmission rate , and the expression is as follows , where respectively represent the preset proportionality coefficients of the link interruption frequency, the proportion of link interruption duration, the data loss rate, and the retransmission rate, and are all greater than 0.
[0039] In a preferred embodiment, construct a desensitization anomaly detection model according to the extreme distribution coefficient, the desensitization intensity imbalance coefficient, and the data transmission link interruption coefficient, and output the desensitization anomaly detection index , and the formula on which the model is based is as follows , in the formula, represents the extreme distribution coefficient, represents the desensitization intensity imbalance coefficient, represents the data transmission link interruption coefficient, respectively represent the preset proportionality coefficients of the extreme distribution coefficient, the desensitization intensity imbalance coefficient, and the data transmission link interruption coefficient, and are all greater than 0.
[0040] In a preferred embodiment, compare the desensitization anomaly detection index with the preset desensitization anomaly detection index threshold to identify potential anomalies in the data desensitization process, specifically as follows:
[0041] If the desensitization anomaly detection index is greater than the desensitization anomaly detection index threshold, it means that there is an anomaly in the desensitization process, automatically trigger the recovery mechanism, reschedule the desensitization task and optimize the parameters;
[0042] If the desensitization anomaly detection index is less than or equal to the desensitization anomaly detection index threshold, it means that the desensitization process is normal.
[0043] The technical effects and advantages of the present invention:
[0044] The present invention establishes a comprehensive data security supervision framework by combining machine learning, which can efficiently and accurately identify and process sensitive data in the public data operation platform, ensuring that the data always meets the privacy protection requirements during the sharing process. The automatic sensitive data recognition algorithm based on machine learning can accurately locate sensitive information in large-scale datasets. Classifying sensitive data according to sensitivity and business attributes can not only adopt the most suitable data desensitization technology for different types of sensitive data, but also automatically adjust the desensitization intensity and technology according to the sensitivity of the data, realizing flexible privacy protection. By obtaining the extreme distribution information of desensitized data, the imbalance information of desensitization intensity, and the data transmission link interruption information, the potential anomalies in the data desensitization process are deeply analyzed, and the problems in the execution of desensitization tasks are timely identified to avoid the risk of data leakage or privacy infringement in desensitization operations. The constructed desensitization anomaly detection model can not only monitor the desensitization process in real time, but also automatically trigger a recovery mechanism according to the detected abnormal situations, reschedule desensitization tasks, and adjust algorithm parameters, so as to ensure that the desensitization tasks can quickly recover when encountering problems and avoid task failure or unqualified desensitization effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings;
[0046] Figure 1 It is a flowchart of the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Embodiment: Figure 1 A data security supervision method for a public data operation platform of the present invention is given, including the following steps:
[0049] Step S1, scanning the dataset of the public data operation platform by using an automatic sensitive data recognition algorithm based on machine learning to accurately locate sensitive data;
[0050] Step S2, classifying the located sensitive data according to data sensitivity and business attributes, and dynamically adapting data desensitization technology according to the classification results;
[0051] Step S3, obtaining the extreme distribution information of desensitized data, the imbalance information of desensitization intensity, and the data transmission link interruption information;
[0052] Step S4, construct a desensitization anomaly detection model, output a desensitization anomaly detection index, identify potential anomalies during the data desensitization process, and automatically trigger a recovery mechanism for detected desensitization anomalies to reschedule desensitization tasks and optimize parameters;
[0053] Step S1, use a machine learning-based sensitive data automatic recognition algorithm to scan the dataset of the public data operation platform and accurately locate sensitive data;
[0054] Extract the dataset from the public data operation platform, including structured data (such as database tables, CSV files) and unstructured data (such as text documents, pictures); conduct data exploration on the field distribution, field names, and types of the data, and process the dataset in chunks according to the data type, where the data types include tabular data, text data, and image data;
[0055] Perform explicit feature recognition and implicit feature mining on the chunk-processed dataset to extract sensitive data features;
[0056] It should be noted that explicit feature recognition includes keyword matching and regular expression matching. Keyword matching: quickly match potential sensitive fields or content based on a keyword library of sensitive information (such as "ID number", "bank card number"); regular expression matching: define rules through common sensitive information patterns (such as date format, phone number format) to accurately extract sensitive information; implicit feature mining: analyze the sentence structure through a natural language processing model to identify personal information from the semantic context (such as the association between name and bank account);
[0057] For structured data: select a traditional machine learning model to identify and locate sensitive data, train the traditional machine learning model according to the extracted sensitive data features, and use the trained traditional machine learning model to scan the entire dataset to accurately identify the fields and records of sensitive data and output the sensitivity score of the sensitive data;
[0058] For unstructured data, select a deep learning model to identify and locate sensitive data, train the deep learning model according to the extracted sensitive data features, and use the trained deep learning model to scan the entire dataset to accurately identify the fields and records of sensitive data and output the sensitivity score of the sensitive data;
[0059] It should be noted that for traditional machine learning algorithms such as decision trees, random forests, support vector machines (SVM), etc.; for unstructured data (such as text and image data), deep learning models are used for sensitive data identification: Text data: Pre-trained models in deep learning (such as BERT, GPT, etc.) are used for context analysis and named entity recognition (NER) of sensitive information to identify sensitive information (such as names, ID numbers, addresses, etc.) in the text. Image data: Optical character recognition (OCR) technology is used to extract text information from pictures, and combined with image recognition algorithms (such as convolutional neural network CNN) to analyze potential sensitive information in the picture content;
[0060] In this embodiment, through the sensitive data automatic recognition algorithm based on machine learning, the public data operation platform can efficiently and accurately identify and locate sensitive data therein. First, by comprehensively extracting and preliminarily processing the data set, the platform can process structured data and unstructured data, and conduct exploratory analysis on the field distribution, field names, types, etc. of the data to ensure that each data type is reasonably divided and processed. In the sensitive data feature extraction stage, the platform combines explicit feature recognition and implicit feature mining to ensure that various types of sensitive information can be comprehensively identified. Explicit feature recognition can quickly and efficiently locate potential sensitive fields by constructing a sensitive information keyword library and common data format rules. Implicit feature mining analyzes the context relationship of the text through natural language processing technology to identify implicit personal information and other sensitive content, so that even unlabeled sensitive information can be effectively identified. For different types of data, corresponding machine learning and deep learning algorithms are selected to process structured data and unstructured data. For structured data, traditional machine learning models can accurately locate and mark sensitive data according to the extracted features during the training process. For unstructured data, deep learning models can process sensitive information in text and image data to ensure that sensitive information can be accurately identified even in a complex data environment. This not only improves the efficiency and accuracy of data privacy protection, but also provides solid data support for subsequent desensitization processing. By accurately locating sensitive data, the platform can achieve more efficient privacy protection and compliance management, reducing potential risks such as data leakage and privacy infringement. At the same time, the training and optimization process of the model can also continuously improve the accuracy and robustness of recognition to cope with the changing data environment and data types.
[0061] Step S2, classify the located sensitive data according to the data sensitivity score and business attributes, and dynamically adapt the data desensitization technology according to the classification results;
[0062] The sensitive data is divided into high-sensitivity data, medium-sensitivity data, and low-sensitivity data according to the data sensitivity score, as follows:
[0063] Compare the data sensitivity score with the sensitivity stage thresholds, where the sensitivity stage thresholds include a first sensitivity threshold and a second sensitivity threshold, and the first sensitivity threshold is less than the second sensitivity threshold;
[0064] If the data sensitivity score is less than or equal to the first sensitivity threshold, mark the located sensitive data as low-sensitivity data; if the data sensitivity score is greater than the first sensitivity threshold and less than the second sensitivity threshold, mark the located sensitive data as medium-sensitivity data; if the data sensitivity score is greater than or equal to the second sensitivity threshold, mark the located sensitive data as high-sensitivity data;
[0065] Classify the sensitive data according to the business attributes, where the business attributes include business scenarios, data usage, and processing requirements; the classification results of business types include but are not limited to financial business, medical business, marketing business;
[0066] Dynamically adapt the data desensitization technology according to the classification results of the data sensitivity score and the classification results of the business types;
[0067] It should be noted that for low-sensitivity data, the desensitization requirement is relatively low, and simple desensitization technologies such as data camouflage, data encryption, or no desensitization can be considered. Usually, this type of data does not require special protection; for medium-sensitivity data, some more strict desensitization technologies such as data masking, field nulling, generalization, etc. can be used. For some business data that does not involve personal privacy, weak encryption storage or partial encryption may be used. For high-sensitivity data, the strongest data protection measures need to be adopted, including but not limited to: Encryption technology: Encrypt data storage and transmission to ensure that data cannot be accessed by unauthorized parties during storage and transmission.
[0068] Data masking: For some particularly sensitive data (such as ID card numbers, bank account information), it can be completely masked or encrypted.
[0069] Hierarchical access control: For high-sensitivity data, apply role-based access control (RBAC) or attribute-based access control (ABAC) policies to ensure that only authorized personnel can access.
[0070] Adaptation based on business attributes:
[0071] Financial business data: High-strength encryption and audit log tracking are required to ensure that data is protected during transmission, storage, and processing.
[0072] Medical business data: According to the requirements of medical industry regulations (such as HIPAA), it may be necessary to comprehensively encrypt the data and apply strict access control.
[0073] Marketing business data: Data masking or anonymization can be appropriately used to ensure that personal identity information is not exposed while retaining sufficient data for analysis.
[0074] By classifying sensitive data according to the sensitivity score and business attributes of the data, the present invention can select the most suitable data desensitization technology for each data category, thereby ensuring data privacy protection while meeting business requirements. This dynamically adaptable data desensitization method can not only improve the accuracy and efficiency of data protection, but also ensure compliance with various business scenarios and regulatory requirements;
[0075] Step S3, obtain the extreme distribution information, desensitization intensity imbalance information, and data transmission link interruption information of the desensitized data;
[0076] The extreme distribution information of the desensitized data includes the extreme distribution coefficient , the desensitization intensity imbalance information includes the desensitization intensity imbalance coefficient , and the data transmission link interruption information includes the data transmission link interruption coefficient ;
[0077] The extreme distribution coefficient is used to measure the deviation degree of extreme data points in the desensitized data from the overall data distribution. By analyzing the extreme distribution coefficient of the desensitized data, it is possible to effectively identify whether there are extreme values in the data. These extreme values usually mean that the data distribution is uneven or abnormal. During the data desensitization process, the extreme distribution coefficient can reveal the deficiencies of the desensitization strategy, especially the situation of over-modification or over-desensitization when dealing with certain special data. For example, when the extreme distribution coefficient is high, it may indicate that during the desensitization process, some data fields are over-modified or over-desensitized, thus affecting the availability of the data or the accuracy of the business logic. On the contrary, if the extreme distribution coefficient is low, it may indicate that the desensitization process can effectively cover all sensitive data, avoid over-desensitization of data fields, and thus increase the availability of the data or the accuracy of the business logic. Therefore, by means of the extreme distribution coefficient, the effectiveness of the desensitization process can be comprehensively evaluated, anomalies in the data distribution can be discovered in a timely manner, ensuring that the data can fully protect privacy without affecting the accuracy and rationality of the business. This evaluation not only improves the accuracy of the desensitization strategy, but also helps optimize the decision-making in the desensitization process, prevent potential leakage of sensitive data, and enhance the security and compliance of the system.
[0078] The acquisition logic of the extreme distribution coefficient is as follows:
[0079] Obtain the desensitized data set and calculate the mean value of the desensitized data set , and the expression is as follows , where represents the number of storage bytes of the i-th de-sensitized data in the de-sensitized dataset, , is a positive integer; calculate the standard deviation of the de-sensitized dataset , and the expression is as follows ;
[0080] Perform extreme value detection on the de-sensitized data in the de-sensitized dataset according to the mean and standard deviation of the de-sensitized dataset, specifically as follows: If , then mark the i-th de-sensitized data as an extreme value;
[0081] Calculate the skewness of the de-sensitized dataset , and the expression is as follows ;
[0082] Calculate the kurtosis of the de-sensitized dataset , and the expression is as follows ;
[0083] Calculate the extreme distribution coefficient , and the expression is as follows , where represents the extreme value ratio, and the calculation expression is as follows , where represents the number of extreme values;
[0084] The de-sensitization intensity imbalance coefficient is used to measure whether the de-sensitization process of each field or dataset is consistent during the data de-sensitization process, whether there is over-processing or under-processing of some parts of the data, and reflects the degree of uneven distribution of the de-sensitization process among different parts of the data; analyzing the de-sensitization intensity imbalance coefficient of the de-sensitization process can provide a more accurate assessment for the entire de-sensitization process, especially in scenarios dealing with complex structures and diverse data. The de-sensitization intensity imbalance coefficient reflects the inconsistency of the de-sensitization efforts on different data records or fields. Its calculation and analysis help identify potential risks or anomalies in the de-sensitization process, especially data vulnerabilities that may not be detected during the application of de-sensitization techniques. When there is an imbalance in the intensity of data de-sensitization, it means that some sensitive data fields may be under-de-sensitized and unable to effectively hide the real data, resulting in a leakage risk; while some other data fields may be over-de-sensitized, causing a reduction in data usability or data distortion, thus affecting subsequent business operations or analysis. By calculating the de-sensitization intensity imbalance coefficient, the difference in this intensity distribution can be quantified, and it can be identified whether some data is not adequately protected or whether some data has been over-processed, thereby helping to adjust the de-sensitization strategy in a timely manner. Therefore, the de-sensitization intensity imbalance coefficient plays a dual warning role in the de-sensitization process. It can not only help identify defects in the de-sensitization strategy but also optimize the entire data protection process to ensure that the data can continue to be effective in different business scenarios without compromising privacy.
[0085] The acquisition logic of the desensitization intensity imbalance coefficient is as follows:
[0086] Obtain the change amount of each desensitization operation , calculate the desensitization intensity , and the expression is as follows , where represents the original amount before desensitization;
[0087] It should be noted that for the desensitization operation of character / text type data, it is mainly character replacement. Therefore, the change amount of the desensitization operation for character / text type data refers to the number of character replacements, and the original amount before desensitization refers to the number of characters before the desensitization operation; for the desensitization operation of numerical data, it is mainly the offset of the numerical range. Therefore, the change amount of the desensitization operation for numerical data refers to the offset amount of the numerical range, and the original amount before desensitization refers to the numerical value before the desensitization operation; the desensitization operations adopted for different types of data are also different, which will not be elaborated here;
[0088] Calculate the Gini coefficient of the desensitization intensity , and the expression is as follows , where represents the desensitization intensity of the nth desensitized data, , is a positive integer;
[0089] It should be noted that the Gini coefficient of the desensitization intensity is used to measure the imbalance of the desensitization intensity distribution;
[0090] Calculate the KL divergence of the desensitization intensity, and the expression is as follows , where represents the KL divergence of the desensitization intensity, represents the expected desensitization intensity of the nth desensitized data;
[0091] It should be noted that the KL divergence is used to measure the difference between the actual distribution and the ideal uniform distribution of the desensitization intensity;
[0092] Calculate the desensitization intensity imbalance coefficient , and the expression is as follows , where respectively represent the preset proportional coefficients of the Gini coefficient and the KL divergence of the desensitization intensity, and are both greater than 0;
[0093] It should be noted that before calculating the desensitization intensity imbalance coefficient, it is necessary to ensure that both the Gini coefficient of the desensitization intensity and the KL divergence of the desensitization intensity have been normalized; Set according to the actual situation. For example, the expert weighting method is adopted, that is, experts in relevant fields are invited to determine the preset proportional coefficients of various indicators through professional opinion surveys and comprehensive evaluations;
[0094] The data transmission link interruption coefficient is used to measure the stability of data during transmission, especially the impact of network interruption or instability on desensitized data. Link interruption may lead to the failure of the data desensitization process or the exposure of sensitive data, and it is to evaluate the link interruption frequency during data transmission and its impact on the execution of the desensitization task; by calculating the data transmission link interruption coefficient, it can be quantified whether there are serious interruptions during data transmission, thereby affecting the continuity and consistency of the data desensitization effect. When the link is unstable, it may occur that some data are not fully desensitized, resulting in the leakage or incomplete processing of sensitive information during transmission, increasing the risk of data leakage. In addition, data link interruption may lead to the asynchrony of desensitization, that is, the desensitization intensity of different data records is inconsistent, thus affecting the protection effect of the entire data set and even having a negative impact on business applications; being able to identify link anomalies in a timely manner and correct them makes the desensitization process more stable and effective. If the data transmission link interruption coefficient is high, it indicates that the interruption frequency of the data transmission link is high, and the data privacy protection measures may be interfered during transmission, and there may be a hidden danger of incomplete data desensitization. Therefore, it is necessary to further optimize the link stability to avoid data loss or leakage. On the contrary, a lower data transmission link interruption coefficient indicates that the link is more stable, the desensitization process is smooth, and data privacy is better protected; therefore, by monitoring and calculating the data transmission link interruption coefficient, not only can link interruption problems in the desensitization process be discovered and corrected, but also corresponding measures can be taken in a timely manner when anomalies occur, so as to ensure the security, integrity and effectiveness of the desensitization process. At the same time, the timely identification of link interruption helps to provide data support for subsequent optimization plans, further enhancing the collaborative effect of data desensitization and transmission, and improving the overall reliability of the system.
[0095] The acquisition logic of the data transmission link interruption coefficient is as follows:
[0096] Use network monitoring tools to monitor each interruption event in the data transmission process, and record the number of link interruption events, the duration of link interruption, the amount of data loss, and the number of retransmitted data packets;
[0097] Calculate the link interruption frequency according to the number of link interruption events , and the expression is as follows , where represents the number of link interruption events, represents the total data transmission time;
[0098] Calculate the proportion of the link interruption duration according to the link interruption duration , the expression is as follows , where represents the link interruption duration of the mth interruption event, , is a positive integer;
[0099] Calculate the data loss rate based on the amount of data loss , the expression is as follows , where represents the amount of data loss, represents the total number of transmitted data packets;
[0100] Calculate the retransmission rate based on the number of retransmitted data packets , the expression is as follows , where represents the number of retransmitted data packets;
[0101] It should be noted that during the transmission link interruption, data may be lost, resulting in incomplete processing of information during the desensitization process. At the same time, the link interruption may trigger the retransmission of data, increasing the transmission delay.
[0102] Calculate the data transmission link interruption coefficient based on the link interruption frequency, the proportion of link interruption duration, the data loss rate, and the retransmission rate , the expression is as follows , where respectively represent the preset proportionality coefficients of the link interruption frequency, the proportion of link interruption duration, the data loss rate, and the retransmission rate, and are all greater than 0;
[0103] It should be noted that before calculating the data transmission link interruption coefficient, it is necessary to ensure that the link interruption frequency, the proportion of link interruption duration, the data loss rate, and the retransmission rate have all been normalized; Set according to the actual situation. For example, use the expert weighting method, that is, invite experts in related fields to determine the preset proportionality coefficients of each index through professional opinion surveys and comprehensive evaluations;
[0104] Step S4, construct a desensitization anomaly detection model, output a desensitization anomaly detection index, identify potential anomalies during the data desensitization process, and automatically trigger a recovery mechanism for detected desensitization anomalies, reschedule desensitization tasks, and optimize parameters;
[0105] Construct a desensitization anomaly detection model based on the extreme distribution coefficient, the desensitization intensity imbalance coefficient, and the data transmission link interruption coefficient, and output a desensitization anomaly detection index , the formula on which the model is based is as follows , in the formula, represents the extreme distribution coefficient, represents the desensitization intensity imbalance coefficient, represents the data transmission link interruption coefficient, respectively represent the preset proportionality coefficients of the extreme distribution coefficient, the desensitization intensity imbalance coefficient, and the data transmission link interruption coefficient, and are all greater than 0;
[0106] It should be noted that before constructing the desensitization anomaly detection model, it is necessary to ensure that the extreme distribution coefficient, the desensitization intensity imbalance coefficient, and the data transmission link interruption coefficient have all been normalized; Set according to the actual situation. For example, the expert weighting method is adopted, that is, relevant experts are invited to determine the preset proportionality coefficients of each index through professional opinion surveys and comprehensive evaluations;
[0107] From the above calculation expressions, it can be seen that the larger the extreme distribution coefficient, the larger the desensitization intensity imbalance coefficient, and the larger the data transmission link interruption coefficient, the larger the desensitization anomaly detection index, which means that the desensitization process is more unstable and the potential security risks are greater. On the contrary, the smaller the extreme distribution coefficient, the smaller the desensitization intensity imbalance coefficient, and the smaller the data transmission link interruption coefficient, the smaller the desensitization anomaly detection index, which means that the degree of abnormality in the desensitization process is lower, and the data desensitization task is executed more stably and securely;
[0108] The desensitization anomaly detection model constructed based on the extreme distribution coefficient, the desensitization intensity imbalance coefficient, and the data transmission link interruption coefficient can comprehensively evaluate the potential risks in the data desensitization process from multiple dimensions and monitor the abnormal situations in data processing in real time. First of all, the extreme distribution coefficient reflects whether there are abnormal extreme values in the desensitized data, which is crucial for ensuring the balance of desensitized data in statistics and privacy protection. By monitoring the existence of extreme values, the model can timely detect possible serious data distortion or insufficient protection in the desensitization process, avoid over - or improper processing of sensitive data, and thus ensure the quality of data privacy protection.
[0109] Secondly, the desensitization intensity imbalance coefficient evaluates whether the intensity of the desensitization operation is uniform. Uneven intensity may lead to some data fields being over - desensitized while others may not be fully desensitized, which will affect the overall security and integrity of the data. By monitoring this coefficient, the model can detect desensitization intensity deviations and adjust the strategy in a timely manner to ensure that all sensitive data is properly protected, thus effectively avoiding the risks of data leakage or misuse.
[0110] Finally, the data transmission link interruption coefficient reflects whether the data is affected by network instability or interruption during the transmission process. The reliability of network transmission directly affects the execution efficiency of the desensitization task and the integrity of the data. Especially during large-scale data processing, link interruption may lead to data loss or desensitization failure. By monitoring this coefficient, the model can evaluate the stability of the transmission link and immediately take recovery measures in case of interruption to ensure the smooth completion of the desensitization task and avoid any information loss or leakage during the data transmission process.
[0111] In summary, the desensitization anomaly detection model constructed based on these three key coefficients can achieve comprehensive monitoring during the data desensitization process, evaluate the quality of desensitization operations in real time, detect potential risks and anomalies, and automatically trigger a recovery mechanism for optimization, thereby ensuring the stability, security, and compliance of the data desensitization process. This comprehensive evaluation method not only improves the execution efficiency of the data desensitization task but also effectively reduces the privacy leakage and compliance risks caused by improper desensitization operations, enhancing the overall reliability of the data privacy protection system.
[0112] Compare the desensitization anomaly detection index with the preset desensitization anomaly detection index threshold to identify potential anomalies during the data desensitization process, as follows:
[0113] If the desensitization anomaly detection index is greater than the desensitization anomaly detection index threshold, it indicates that there is an anomaly in the desensitization process, and the recovery mechanism will be automatically triggered to reschedule the desensitization task and optimize the parameters;
[0114] If the desensitization anomaly detection index is less than or equal to the desensitization anomaly detection index threshold, it indicates that the desensitization process is normal, and the system will continue to execute the task according to the predetermined desensitization strategy without taking any additional adjustment or recovery measures;
[0115] The present invention establishes a comprehensive data security supervision framework by combining machine learning, which can efficiently and accurately identify and process sensitive data in the public data operation platform, ensuring that the data always meets the privacy protection requirements during the sharing process. The automatic sensitive data identification algorithm based on machine learning can accurately locate sensitive information in large-scale datasets. Classifying sensitive data according to sensitivity and business attributes can not only adopt the most suitable desensitization technology for different types of sensitive data, but also automatically adjust the desensitization intensity and technology according to the sensitivity of the data, realizing flexible privacy protection. By obtaining extreme distribution information of desensitized data, unbalanced desensitization intensity information, and data transmission link interruption information, the potential anomalies in the data desensitization process are deeply analyzed, and the problems in the execution of desensitization tasks are timely identified to avoid the risk of data leakage or privacy infringement in the desensitization operation. The constructed desensitization anomaly detection model can not only monitor the desensitization process in real time, but also automatically trigger a recovery mechanism according to the detected abnormal situation, reschedule the desensitization task, and adjust the algorithm parameters, so as to ensure that the desensitization task can quickly recover when encountering problems and avoid task failure or unqualified desensitization effect.
[0116] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0117] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0118] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A data security supervision method for a public data operation platform, characterized in that: The steps include: Step S1, using a machine learning-based automatic sensitive data identification algorithm to scan the data set of the public data operation platform and accurately locate sensitive data; Step S2: classify the located sensitive data according to data sensitivity and business attributes, and dynamically adapt the data desensitization technology according to the classification results; Step S3, obtaining the extreme distribution information of desensitized data, the imbalance information of desensitization intensity, and the interruption information of data transmission link; Step S4, construct a desensitization anomaly detection model, output a desensitization anomaly detection index, identify potential anomalies in the data desensitization process, automatically trigger a recovery mechanism for detected desensitization anomalies, reschedule desensitization tasks, and optimize parameters; The extreme distribution information of the desensitized data includes the extreme distribution coefficient ; The logic for obtaining the extreme distribution coefficient is as follows: Get the desensitized data set and calculate the mean of the desensitized data set , the expression is as follows ,in Indicates the number of bytes stored in the i-th desensitized data in the desensitized data set. , is a positive integer; Calculate the standard deviation of the desensitized data set , the expression is as follows ; According to the mean and standard deviation of the desensitized data set, the extreme value detection of the desensitized data in the desensitized data set is performed as follows: , then mark the i-th desensitized data as an extreme value; Calculate the skewness of the masked dataset , the expression is as follows ; Calculate the kurtosis of the desensitized dataset , the expression is as follows ; Calculate the extreme distribution coefficient , the expression is as follows ,in Represents the extreme value ratio, and the calculation expression is as follows ,in Indicates the number of extreme values.
2. A data security supervision method for a public data operation platform according to claim 1, characterized in that: Extracting data sets from a public data operation platform, including structured data and unstructured data; performing data exploration on the field distribution, field names, and types of the data, and processing the data sets in blocks according to data types, including tabular data, text data, and image data; Perform explicit feature recognition and implicit feature mining on the block-processed data set to extract sensitive data features; For structured data: select traditional machine learning models to identify and locate sensitive data, train the traditional machine learning models based on the extracted sensitive data features, and use the trained traditional machine learning models to scan the entire data set, accurately identify the fields and records of sensitive data, and output the sensitivity scores of sensitive data; For unstructured data, a deep learning model is selected to identify and locate sensitive data. The deep learning model is trained based on the extracted sensitive data features, and the trained deep learning model is used to scan the entire data set to accurately identify the fields and records of sensitive data and output the sensitivity score of the sensitive data.
3. A data security supervision method for a public data operation platform according to claim 2, characterized in that: Sensitive data is divided into high sensitivity data, medium sensitivity data, and low sensitivity data according to the data sensitivity score, as follows: Comparing the data sensitivity score with a sensitivity stage threshold, wherein the sensitivity stage threshold includes a first sensitivity threshold and a second sensitivity threshold, and the first sensitivity threshold is less than the second sensitivity threshold; If the data sensitivity score is less than or equal to the first sensitivity threshold, the located sensitive data is marked as low sensitivity data; if the data sensitivity score is greater than the first sensitivity threshold and the data sensitivity score is less than the second sensitivity threshold, the located sensitive data is marked as medium sensitivity data; if the data sensitivity score is greater than or equal to the second sensitivity threshold, the located sensitive data is marked as high sensitivity data; Classify sensitive data into business types according to business attributes, including business scenarios, data usage, and processing requirements; The classification results of business types include but are not limited to financial business, medical business, and marketing business; Dynamically adapt data desensitization technology based on the data sensitivity score classification results and business type classification results.
4. According to claim 1, a data security supervision method for a public data operation platform is characterized in that: Desensitization intensity imbalance information includes desensitization intensity imbalance coefficient , the data transmission link interruption information includes the data transmission link interruption coefficient .
5. A data security supervision method for a public data operation platform according to claim 4, characterized in that: The logic for obtaining the desensitization intensity imbalance coefficient is as follows: Get the change amount of each desensitization operation , calculate the desensitization intensity , the expression is as follows ,in It represents the original amount before desensitization; Calculate the Gini coefficient of desensitization intensity , the expression is as follows ,in Indicates the desensitization intensity of the nth desensitized data. , is a positive integer; Calculate the KL divergence of desensitization strength, the expression is as follows ,in represents the KL divergence of desensitization strength, Indicates the expected desensitization strength of the nth desensitized data; Calculate the desensitization intensity imbalance coefficient , the expression is as follows ,in They represent the Gini coefficient of desensitization strength and the preset proportional coefficient of KL divergence, respectively, and Both are greater than 0.
6. A data security supervision method for a public data operation platform according to claim 4, characterized in that: The logic for obtaining the data transmission link interruption coefficient is as follows: Use network monitoring tools to monitor every interruption event in the data transmission process and record the number of link interruption events, link interruption duration, data loss, and number of retransmitted data packets; Calculate the link interruption frequency based on the number of link interruption events , the expression is as follows ,in Indicates the number of link interruption events. Indicates the total data transmission time; Calculate the link interruption duration ratio based on the link interruption duration , the expression is as follows ,in Indicates the link interruption duration of the mth interruption event, , is a positive integer; Calculate the data loss rate based on the amount of data loss , the expression is as follows ,in Indicates the amount of data loss. Indicates the total number of transmitted data packets; Calculate the retransmission rate based on the number of retransmitted packets , the expression is as follows ,in Indicates the number of retransmitted packets; Calculate the data transmission link interruption coefficient based on link interruption frequency, link interruption duration ratio, data loss rate, and retransmission rate , the expression is as follows ,in Respectively represent the preset proportional coefficients of link interruption frequency, link interruption duration ratio, data loss rate, and retransmission rate, and Both are greater than 0.
7. A data security supervision method for a public data operation platform according to claim 4, characterized in that: A desensitization anomaly detection model is constructed based on the extreme distribution coefficient, desensitization intensity imbalance coefficient, and data transmission link interruption coefficient, and a desensitization anomaly detection index is output. The model is based on the following formula , where represents the extreme distribution coefficient, Indicates the desensitization intensity imbalance coefficient, represents the data transmission link interruption coefficient, They represent the preset proportional coefficients of the extreme distribution coefficient, the desensitization intensity imbalance coefficient, and the data transmission link interruption coefficient, respectively, and Both are greater than 0.
8. A data security supervision method for a public data operation platform according to claim 7, characterized in that: The desensitization anomaly detection index is compared with the preset desensitization anomaly detection index threshold to identify potential anomalies in the data desensitization process, as follows: If the desensitization anomaly detection index is greater than the desensitization anomaly detection index threshold, it means that there is an abnormality in the desensitization process, and the recovery mechanism is automatically triggered to reschedule the desensitization task and optimize the parameters; If the desensitization anomaly detection index is less than or equal to the desensitization anomaly detection index threshold, it means that the desensitization process is normal.
Citation Information
Patent Citations
Archive management method and system based on big data
CN118551414A
E-commerce platform information security desensitization scheme analysis system based on artificial intelligence
CN119538316A