Power grid data sharing method based on intelligent data identification and differential privacy protection
Patent Information
- Application Number
- CN202510911658.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-21
Smart Images

Figure CN121000409A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of electric digital data processing, and particularly relates to a power grid data sharing method based on intelligent data recognition and differential privacy protection. BACKGROUND
[0002] With the rapid development and popularization of big data storage and computing, mobile Internet, Internet of Things and other big data related technologies, the total amount of global data presents an exponential growth trend. Along with the deepening of the digital transformation of society and economy, the construction requirements of digital enterprises are gradually increasing.
[0003] In the process of improving the digital level of enterprises, various information technologies will be widely used in enterprise management services and businesses to achieve public data opening and sharing, promote the joint construction and sharing of enterprise information, and improve the scientific nature of decision-making and service efficiency.
[0004] At present, clear requirements are put forward for the protection of enterprise big data platform data resources throughout their life cycle, and the data security of enterprise information resources in the process of development and utilization, sharing and exchange needs to be effectively protected. A unified, standardized, interconnected, safe and controllable enterprise big data opening and sharing platform needs to be established to promote enterprise data utilization and sharing and jointly create an open and safe digital new ecosystem.
[0005] In order to better realize organizational information security protection and meet the requirements of business development and regulatory compliance, units are actively developing information system security protection projects. The protection capability has basically covered the network, terminal, data and business system. However, with the development of security construction, the contradiction between security and ease of use is increasingly prominent, and there is still a lack of targeted measures to cooperate with protection for data security, mainly in the following aspects:
[0006] 1. Internal leakage is difficult to control:
[0007] In daily work, it is inevitable to send sensitive data through IM transmission, network sending, email and other ways. It is difficult to effectively manage and control these behaviors and prevent uncontrolled diffusion of important files.
[0008] 2. It is difficult to trace the leakage risk:
[0009] Important core data lacks monitoring and auditing of the entire circulation channel, making it difficult to trace to the person;
[0010] After a security incident occurs, if there is no perfect behavior auditing system, it is still difficult to respond in time, accurately locate the source of the incident, and bring great trouble and serious information security risks to enterprises and institutions.
[0011] 3. It is difficult to grasp the management effect:
[0012] Although a large number of security management work is carried out, and a large number of security management equipment is deployed, the effect is not clear, and the management effect cannot be quantified;
[0013] The security department cannot effectively grasp the distribution, flow direction, and external transmission of the core data of the enterprise in real time, and the audit work efficiency of the security administrator is low, and it is difficult to grasp the overall situation of security management;
[0014] 4、Security management is difficult to close loop,
[0015] It is difficult to sort out the current situation of enterprise security, difficult to regularly evaluate the current situation of enterprise security, and difficult to plan the future development of enterprise security.
[0016] At present, in the field of intelligent identification of data security levels, domestic and foreign research is less involved. Most of the existing research in this field focuses on encryption algorithms, information security interaction and access control identification under known security levels.
[0017] Differential privacy (Differential privacy) is a system for sharing information about data sets publicly, which protects individual information in the data set while describing the characteristics of the group in the data set. In the concept of differential privacy, if the impact of any single change in the database is small enough, the query result cannot be used to infer a large amount of information about any individual, so the privacy of the individual is guaranteed. Through this constraint condition for the algorithm of publishing statistical database information, the disclosure of personal information in the database is limited.
[0018] The research field of differential privacy in power grid mainly focuses on the defense of differential privacy attack of smart meter data, the improvement of sensor differential privacy utility, the desensitization release of power transaction big data, the load balancing of electric vehicles, and the energy management of heterogeneous power generation equipment. The research methods used are mainly combined with clustering and edge computing.
[0019] The invention patent with the authorization announcement date of January 3, 2023 and the authorization announcement number of CN 114614974 B discloses "a privacy set intersection method, system and device for cross-industry sharing of power grid data", which uses Bloom filter technology to map and process local data, greatly reducing the calculation cost of the set intersection process. By using random response technology instead of homomorphic encryption technology, the communication cost is effectively reduced, and the demand for processing large amounts of data in actual scenarios can be met, and the random response technology can realize local differential privacy by randomly flipping the Bloom filter. The protocol parameters for each run calculation are agreed by the user in advance and are not shared with the server, which can avoid malicious attacks by the server to some extent. And the disturbance Bloom filter shared by the user and the server is processed by random response, which further protects the local privacy data. Since it focuses on using filter technology to map and process local data to greatly reduce the calculation cost of the set intersection process, and uses random response technology to reduce communication cost, it does not consider the processing problem of power company business data, especially the regulation and control data and non-numerical information in differential privacy.
[0020] The invention patent application with the application publication date of April 25, 2025 and the application publication number of CN 119885258 A discloses a "data processing method and system of power grid and storage medium", which comprises: acquiring first power grid associated data collected by a plurality of embedded devices in the power grid, and performing initial differential privacy processing on the first power grid associated data for each group of the first power grid associated data to obtain second power grid associated data; performing aggregation processing on a plurality of groups of the second power grid associated data through an edge computing node to obtain target power grid associated data, and uploading the target power grid associated data to the cloud. This technical solution improves the security of data processing. However, it only considers the desensitization problem of power big data and does not involve the processing problem of regulation and control data and non-numerical information in hierarchical differential privacy.
[0021] Currently, there is no relevant report on the differential privacy application and management of power company business data, especially regulation and control data and non-numerical information. SUMMARY
[0022] The invention object of the present application is to provide a power grid data sharing method based on intelligent data recognition and differential privacy protection. First, an artificial intelligence method is used to identify the data security level of massive data, and then different differential disturbances are added to data of different security levels. Through specific identification and processing of data, differential data sharing is realized, and data security and sharing are realized without affecting the statistical situation.
[0023] The technical scheme of the present application is to provide a power grid data sharing method based on intelligent data recognition and differential privacy protection, characterized by:
[0024] 1) Using transformer model-based power scenario recognition, multi-level deep semantic mining of text information is performed to realize classification of power dispatch system data and form corresponding secret levels of different data;
[0025] 2) Recognizing data secret levels, and then adding different differential disturbances to data of different secret levels;
[0026] 3) Building a secure intelligent classification and secret protection service system, and using differential privacy to protect power dispatch system data;
[0027] 4) Through classification and secret protection recognition, processing and management of power dispatch system data, providing easy-to-use service interfaces including sensitive data classification and secret protection, file classification and secret protection, and text data stream classification and secret protection;
[0028] 5) External data accesses the business of the power dispatch system through a secure SSL interface-easy-to-use service interface to realize differentiated power dispatch system data sharing.
[0029] Specifically, the secure intelligent classification and secret protection service system comprises a classification and secret protection interface module, a classification and secret protection unified service module, a data cleaning and preprocessing module, a model training module, a classification and secret protection task scheduling module, a computing engine hardware adaptation module, a statistical machine learning algorithm module, a neural network algorithm module, and a configuration management module.
[0030] Further, the classification and secret protection interface module is also called a unified application and calling interface, which provides a unified service interface for the application layer through modular technology and interface encapsulation; the classification and secret protection interface module uses a domestic hardware platform including a domestic CPU and GPU, supports cluster deployment and elastic expansion, and provides classification and secret protection services on demand to ensure high reliability and scalability of the classification and secret protection service platform;
[0031] The classification and secret protection unified service module provides integrated, multi-level and scenario-based classification and secret protection service capabilities, and through application modes including general text file classification and secret protection services, typical document classification and secret protection services, and real-time text data stream classification and secret protection services, it adapts to different business scenarios and provides easy-to-use service interfaces including sensitive data classification and secret protection, file classification and secret protection, and text data stream classification and secret protection for easy-to-use and convenient classification and secret protection services for applications;
[0032] The data cleaning and preprocessing module cleans and converts the data of the files and data to be classified and coded into the required data of the model, and is linked with the scheduling module to cache the input data and output data;
[0033] The model training module adopts supervised and unsupervised training, including receiving batch-labeled sample files, training artificial intelligence models for classified and coded service use; the unsupervised training part uses the actual use data to realize unsupervised machine learning to correct the model parameters and improve the accuracy;
[0034] The classified and coded task scheduling module realizes asynchronous task inference prediction through Tensor-RT, and realizes high-performance processing capability for large-scale data;
[0035] The computing engine hardware adaptation module modifies the deep learning basic middleware platform including TensorFlow, and adapts to the hardware of the domestic GPU / DCU, and supports the national hardware;
[0036] The statistical machine learning algorithm module supports the statistical machine learning algorithm engine including Bayesian / CRF conditional random field, Gaussian mixture, K-means and LDA, and further improves the accuracy of the classified and coded model;
[0037] The neural network algorithm module is the main model of the file security intelligent classified and coded service system, specifically involving the CL-BERT model, which is an algorithm model based on attention mechanism;
[0038] The configuration management module provides system management functions including communication parameters, user permissions, model parameters and certificates.
[0039] Specifically, in step 1), a one-dimensional convolutional neural network kernel is used to convolve the time series, project each local window into an embedding vector, and each token carries the short-term pattern of the time series; add position embedding to the token, then pass it through the Transformer model to learn the long-term dependency between tokens; the Transformer model outputs the latent embedding vector of the time series, and then processes it through a multi-layer perception MLP with softmax activation to generate a flag classification output.
[0040] Specifically, in step 2), differential privacy is used to protect the data of the power dispatching system; based on the existing data of the dispatching, the data of different address sources is analyzed, a similarity variance matrix is established, a mechanism for adding verification noise to the matrix is verified, and the protection of key data is realized.
[0041] Further, in step 3), based on the SQL agent mechanism, combined with data classification, privacy calculation is performed; when data is queried, different level queries are realized according to the level of data.
[0042] Further, the different level queries are realized by recursively and JOINing operations to obtain data from the table according to the hierarchical structure and arranging the data according to the hierarchy to process a large amount of structured data and realize the level query of the data.
[0043] Specifically, in step 3), the secure intelligent classification and encryption service system responds to the external SQL query request according to the following steps:
[0044] 1) an external SQL query request is proposed;
[0045] 2) SQL analysis and rewriting are performed;
[0046] 3) sensitivity analysis is performed;
[0047] 4) a sensitivity query request is proposed to the metadata management server;
[0048] 5) the metadata management server returns the sensitivity query result;
[0049] 6) data source access is performed;
[0050] 7) a query step is performed to the data warehouse;
[0051] 8) the data warehouse returns the query result;
[0052] 9) privacy budget allocation is performed;
[0053] 10) a request is proposed to the privacy budget management server;
[0054] 11) the privacy budget management server provides the privacy budget;
[0055] 12) noise addition is performed;
[0056] 13) a query section is constituted;
[0057] 14) a query result is output.
[0058] Specifically, the secure intelligent classification and encryption service system adopts a hierarchical design, divides the application program into multiple levels, and each level has a specific responsibility and function:
[0059] 1) a service layer: responsible for processing the logic of the user interface and user input and output, and sending the user's request to the next layer of the application program;
[0060] 2) Business Logic Layer: responsible for handling the core business logic of the application, including functions such as data processing, calculation, and verification;
[0061] 3) Data Access Layer: responsible for interacting with data storage systems, performing read and write operations on data.
[0062] The power grid data sharing method based on intelligent data recognition and differential privacy protection has the characteristics that the power grid data sharing method based on intelligent data recognition and differential privacy protection uses data recognition and protection in the power dispatching system, realizes data privacy and sharing without affecting the statistical situation through specific identification and processing of data, and adds different differential disturbances to data of different levels of data security through artificial intelligence methods for massive data recognition, thereby realizing differentiated data sharing through specific identification and processing of data.
[0063] Compared with the prior art, the advantages of the present application are:
[0064] 1. The technical solution of the present application uses intelligent technology to use data recognition and protection in the power dispatching system, realizes data privacy and sharing without affecting the statistical situation through specific identification and processing of data;
[0065] 2. The technical solution of the present application uses artificial intelligence methods to identify data security levels for massive data, and adds different differential disturbances to data of different levels of data security for power grid company business data, especially differential privacy of control data and non-numeric information, thereby realizing differentiated data sharing through specific identification and processing of data.
[0066] 3. The technical solution of the present application introduces Tensor-RT real-time processing software architecture to complete network layer fusion and precision calibration, and performs reverse analysis and optimization on the model, reduces double precision to single precision, reduces GPU memory occupation by 50%, and makes the model more lightweight, so that the number of parallel models that a single card can load is doubled, which helps the system to obtain better domestic support.
[0067] 4. The technical solution of the application realizes differentiated data sharing with minimum redundancy by deploying a differentiated data sharing software platform based on intelligent data recognition and differential privacy protection in data sharing of various departments, various levels of dispatching and various companies, thereby providing strong support for the safe operation of the network of the State Grid Power Company and improving the overall security defense level of the power grid. BRIEF DESCRIPTION OF DRAWINGS
[0068] Figure 1 is a structural schematic diagram of a Transformer model encoder;
[0069] Figure 2 is a structural schematic diagram of a CLBERT scheme;
[0070] Figure 3 is an internal structural schematic diagram of the security intelligent grading and classification service system of the application;
[0071] Figure 4 is a software logic architecture schematic diagram of the application;
[0072] Figure 5 is a differential privacy architecture schematic diagram of the application;
[0073] Figure 6 is a differential privacy data processing flow schematic diagram of the application;
[0074] Figure 7 is a whole model schematic diagram of the adaptive model CL-BERT of the application;
[0075] Figure 8 is a whole flow schematic diagram of privacy computing of the application;
[0076] Figure 9 is a schematic diagram of the responsibilities and functions of each level of the whole flow of external data business access of the application;
[0077] Figure 10 is a whole flow schematic diagram of business access of external data of the application. DETAILED DESCRIPTION
[0078] The application will be further described below in combination with the drawings and embodiments.
[0079] Data, as a core production factor and asset, plays an important driving role in improving the company management level, labor productivity, technological innovation capability and the like of a power grid enterprise.
[0080] From the outside, with the implementation of new power systems, double carbon targets and other strategies, new digital infrastructure such as the Internet of Things is widely deployed and applied, and power grid enterprise data is undergoing tremendous changes in content and form; Government-enterprise data fusion, energy big data center construction, and power data value-added services are developing rapidly, and company data sharing and internal and external data fusion applications have become an inevitable trend.
[0081] From the inside, as digital transformation deepens, the demand for data at all levels of the company is increasing, and data applications are generally characterized by cross-professional and large-scale, requiring higher standards for data reliability; The data application needs of grassroots units are becoming more and more vigorous, and higher requirements for data availability, visibility, and traceability are put forward.
[0082] The technical solutions of the present application will be described in detail below.
[0083] 1. Machine learning algorithm for data recognition
[0084] The BERT model is based on the Transformer model, so it is necessary to first understand the structure of the Transformer model.
[0085] The Transformer model (transformer model, which realizes the sequence-to-sequence conversion task through an encoder-decoder architecture (Encoder-Decoder)) was proposed by Vaswani et al. in 2017, which adopts an attention mechanism (attention mechanism) to capture the dependency between input sequences.
[0086] The Transformer model consists of an encoder and a decoder, and the specific structure of the Transformer model encoder is as shown in Figure 1 .
[0087] In the encoder, the input sequence is first converted to a vector representation through an embedding layer. This vector representation serves as the initial representation of the input sequence, and after passing through multiple layers of self-attention mechanisms (self-attention) and feed-forward neural networks (feed-forward neural network), the representation of the output sequence is obtained.
[0088] Specifically, the self-attention mechanism refers to calculating the similarity between each position in the input sequence and all other positions, then assigning weights to each position according to the similarity, and finally weighting and summing the representations of all positions. This process can be seen as a special aggregation operation on the input sequence, resulting in a more comprehensive and richer representation.
[0089] A feedforward neural network is a network composed of two fully connected layers, with an activation function (such as ReLU) as a non-linear transformation between each layer. This network can map the representation of an input sequence once to obtain a higher-dimensional and more complex representation.
[0090] The Transformer model has the characteristic of parallel computing, so it has high efficiency in both training and inference. In natural language processing tasks, the Transformer model has been proven to have very good performance.
[0091] 2. Differential privacy:
[0092] Differential privacy proposes an important idea: adding or subtracting a record in a data set for a statistical query can obtain almost the same output.
[0093] That is, any record, whether it is in the data set or not, has a negligible impact on the result, so it is impossible to restore any original record from the result.
[0094] Suppose the original data set is D (which can be understood as a table), and adding or subtracting a record to it forms D'. At this time, D and D' are adjacent data sets; suppose a differential privacy algorithm is A(), and the result of operating on data set D and adding noise is A(D) = V; the result of operating on data set D' and adding noise is A(D') = V'; V and V' are the results of statistical operations, and differential privacy requires that the results of adjacent data sets be basically the same, i.e. V = V'.
[0095] Different inputs will produce different outputs, and P() represents the probability that A(D) = V. For all outputs V, it is required that:
[0096]
[0097] Mathematically, when ε is small, it is approximately:
[0098] e ε ≈1+ε
[0099] Therefore, when ε is small, the formula can also be written as:
[0100]
[0101] This formula can help understand the principle of differential privacy.
[0102] ε is called the privacy protection budget, which is used to control the probability ratio of the privacy protection algorithm A() obtaining the same output on adjacent data sets.
[0103] The smaller the epsilon, the smaller the risk of privacy leakage, but the larger the noise introduced, and the smaller the research value of the output data set.
[0104] If epsilon = 0, it means that the differential privacy algorithm obtains exactly the same output on all neighboring data sets, that is, it is impossible to leak any user's privacy, but it has no research value for research institutions.
[0105] Differential privacy is mathematically proven that even if an attacker has mastered all record information except a certain designated record (i.e. the maximum background knowledge assumption), it cannot determine the private data contained in the record.
[0106] The technical scheme of the present application applies the following key technologies:
[0107] I. Adopting a lightweight CL-BERT model to improve overall performance:
[0108] The original BERT model (Bidirectional Encoder Representations from Transformers, "Bidirectional Encoder Representations from Transformers" or simply "Bidirectional Transformer Model") was developed and trained by Google using TensorFlow.
[0109] BERT is released in two sizes, BERTBASE and BERTLARGE. The BASE model is used to measure the performance of an architecture comparable to another architecture, while the LARGE model produces the most advanced results reported in research papers. One of the main reasons why BERT performs well on different NLP tasks is the use of semi-supervised learning. This means that the model is trained for a specific task, allowing it to understand the patterns of language.
[0110] The trained model (BERT) has language processing capabilities and can be used to authorize other models built and trained using supervised learning.
[0111] The BERT model is basically an encoder stack of a Transformer architecture. The Transformer architecture is an encoder-decoder network that uses self-attention on the encoder side and attention on the decoder side. BERT-Base has 12 layers in the encoder stack, while BERT-Large has 24 layers in the encoder stack. These are not just the Transformer architecture described in the original paper (6 encoder layers). The BERT architecture (Base and Large) also has a larger feed-forward network (768 and 1024 hidden units, respectively) and more attention heads (12 and 16, respectively) than the Transformer architecture. It contains 512 hidden units and 8 attention heads.
[0112] BERTBASE contains 110M parameters, while BERTLARGE contains 340M parameters. So to summarize BERT-Base: 12 layer Encoder / Decoder, d = 768, 110M parameters BERT-Large: 24 layer Encoder / Decoder, d = 1024, 340M parameters where d is the dimension of the final hidden vector that BERT outputs. Both versions have Cased and Uncased versions (the Uncased version converts all words to lowercase).
[0113] The model first takes the CLS token as input, followed by a sequence of words as input. The CLS here is a classification token. It then passes the input through the layers above. Each layer applies self-attention, passing the result through a feed-forward network to the next encoder.
[0114] The model outputs a vector of the hidden size (768 for BERT BASE). If you want to output a classifier from this model, you can take the output corresponding to the CLS token. One of the biggest challenges in NLP is the lack of enough training data.
[0115] In total, there is a large amount of text data available, but if you want to create a task-specific dataset, you need to split this data into a very large number of different domains. When you do this, you end up with only a few thousand or a few hundred thousand human-labeled training samples. Deep learning-based NLP models require a lot of data - you see significant improvements when training on millions or billions of annotated training examples.
[0116] The pre-trained BERT model weights have encoded a lot of information about our language.
[0117] Thus, the time required to train the fine-tuned model is much less - as if the underpinnings of the network have already been widely trained, and only need to be lightly adjusted while using their output as features for the classification task.
[0118] This approach allows for fine-tuning of a task on a much smaller dataset than would be required for a model built from scratch, due to the pre-trained weights.
[0119] One major drawback of NLP models built from scratch is that a very large dataset is usually required to train the network to a reasonable accuracy, which means a lot of time and effort must be invested in creating the dataset. With fine-tuning of BERT, it is now possible to train a model on less training data to achieve good performance.
[0120] Generally, increasing the size of the pre-trained model will bring the improvement of the effect; however, when the model size reaches a certain degree, it is difficult to proceed any further because it is limited by GPU memory and training time.
[0121] In order to reduce the model parameters and the model training time, the CLBERT scheme is proposed, CLBERT (Contrastive Learning BERT) is also the Encoder structure of the Transformer as BERT, and the activation function is also GLUE.
[0122] Compared with BERT, the main improvements of CLBERT are as follows: Embedding factorization, inter-layer parameter sharing, and inter-sentence correlation loss.
[0123] The scheme mainly takes attention mechanism as the framework, through the way of multi-head attention joint learning, multi-level deep semantic mining is carried out on the input text information, so as to extract more fine-grained information about the security theme, so as to realize high-precision intelligent security grading, the specific structure of CLBERT scheme is shown in Figure 2 .
[0124] The model takes multi-head attention as the benchmark, and multiple attention mechanisms work together to extract multiple levels of text semantic information.
[0125] For each input text T, the text is input into multiple self-attention mechanisms, and the formula is:
[0126]
[0127] In the adaptive attention mechanism, Q = K = V and is the same as the input text feature. Then the semantic information of multiple attention mechanisms is fused, and the final semantic feature fused with multiple layers of context information is obtained through multi-layer feature extraction and integration. Then the information is aligned to identify and predict the security level of the intelligent privacy.
[0128] The model makes more accurate judgments on the security level of the text by capturing and understanding the context in a more fine-grained dimension. At the same time, the model is lightweight by sharing the model parameters, greatly reducing the training and testing dimension requirements of the model, and better meeting the needs of the industry for model complexity and testing time.
[0129] II. Use LDA topic model to improve business flexibility:
[0130] LDA (Latent Dirichlet Allocation) is a document topic generation model, also known as a three-layer Bayesian probability model, containing three layers of words, topics and documents.
[0131] A generative model is a model that assumes that each word in an article is obtained through a process of "the article selects a certain topic with a certain probability, and selects a certain word from the topic with a certain probability". The document to topic obeys a multinomial distribution, and the topic to word obeys a multinomial distribution.
[0132] LDA is a non-supervised machine learning technique that can be used to identify hidden topic information in large-scale document collections or corpora.
[0133] For each document in the corpus, LDA defines the following generation process:
[0134] For each file, a topic is drawn from the topic distribution.
[0135] A word is drawn from the word distribution corresponding to the topic drawn above.
[0136] Repeat the above process until each word in the document is obtained.
[0137] LDA believes that each article is a mixture of multiple topics, and each topic can be represented by the probability of multiple words
[0138] Based on the LDA topic model, the document clustering and the supervised LDA topic model can provide more reference keywords in the clustering results, and the business department can formulate more business-related document categories according to the reference keywords.
[0139] III. System Internal Structure:
[0140] like Figure 3 As shown in the figure, the security intelligent classification and confidentiality service system in this technical solution includes a classification and confidentiality interface module, a unified classification and confidentiality service module, a data cleaning and preprocessing module, a model training module, a classification and confidentiality task scheduling module, a computing engine hardware adaptation module, a statistical machine learning algorithm module, a neural network algorithm module, and a configuration management module.
[0141] (1) Classification and security interface module.
[0142] Also known as the Unified Application and Invocation Interface, the Classification and Confidentiality Determination Service Platform uses modular technology to encapsulate interfaces and provide a unified service interface for the application layer. The platform utilizes domestically produced hardware, including domestically produced CPUs and GPUs, supports cluster deployment and elastic scaling, and provides classification and confidentiality determination services on demand, ensuring high reliability and scalability.
[0143] (2) Unified service module for classifying and classifying confidentiality.
[0144] It provides integrated, multi-layered, and scenario-based classification and confidentiality services, including flexible application methods such as general text file classification and confidentiality services, typical official document classification and confidentiality services, and real-time text data stream classification and confidentiality services, adapting to the differentiated classification and confidentiality services under different business scenarios. Through scenario-based classification and confidentiality service encapsulation, it provides easy-to-use service interfaces for sensitive data classification and confidentiality, file classification and confidentiality, and text data stream classification and confidentiality, supporting development languages such as C and Java, providing applications with easy-to-use and convenient classification and confidentiality services.
[0145] (3) Data cleaning and preprocessing module.
[0146] The system cleans and transforms files and data that need to be classified and confidential into the data required by the model, and works in conjunction with the scheduling module to cache the input and output data.
[0147] (4) Model training module.
[0148] The system employs both supervised and unsupervised training. This includes receiving batches of labeled sample files to train an artificial intelligence model for use in the classification and confidentiality services. The unsupervised training component utilizes actual data to achieve unsupervised machine learning, thereby refining model parameters and improving accuracy.
[0149] (5) Classification and security level task scheduling module.
[0150] Asynchronous task inference and prediction are achieved through Tensor-RT, enabling high-performance processing of large-scale data.
[0151] (6) Computing engine hardware adaptation module.
[0152] The deep learning base middleware platform such as TensorFlow is reformed, hardware adaptation is performed with a domestic GPU / DCU, and support is provided for nationally produced hardware.
[0153] (7) Statistical machine learning algorithm module.
[0154] Support is provided for statistical machine learning algorithm engines such as Bayesian / CRF conditional random field, Gaussian mixture, K-means and LDA, and model precision of fixed grading and fixed classification is further improved.
[0155] (8) Neural network algorithm module.
[0156] The main model of the file security intelligent fixed grading and fixed classification service system is a CL-BERT model, which is an algorithm model based on an attention mechanism.
[0157] (9) Configuration management module.
[0158] System management functions such as communication parameters, user permissions, model parameters and certificates are provided.
[0159] Four, software logic architecture:
[0160] The software logic architecture of the technical scheme of the application is as shown in Figure 4
[0161] (1) CL-BERT model.
[0162] The classification of the document is mainly realized by supervised pre-training of the training data, the model principle is based on an attention mechanism, and the model is called CL-BERT. Compared with the general BERT model, the model is relatively small, and the performance is greatly improved. And because the power grid document has certain language characteristics, the prediction accuracy of the simplified model does not decrease.
[0163] (2) Training data imbalance solution.
[0164] The classification engine retains grouping classification algorithms based on a small amount of supervised training text such as Bayesian / CRF conditional random field.
[0165] (3) Grading classification precision completeness.
[0166] In order to avoid omission of the document classification category, unsupervised clustering algorithms such as Gaussian mixture and K-means are used, which can provide intuitive and effective reference for manual review.
[0167] (4) New category mining.
[0168] Document clustering based on LDA topic model, as well as supervised LDA topic model, can provide more reference keywords in the clustering results. Business departments can formulate more business-related document categories based on the reference keywords.
[0169] (5) Artificial intelligence + artificial intelligence.
[0170] It retains the traditional manual assistance methods such as regular expression rule base, thesaurus, and typical semantic statement annotation, and further meets the needs of more refined classification.
[0171] (6) Real-time issues.
[0172] This paper introduces the Tensor-RT real-time processing software architecture. Unlike the traditional TensorFlow architecture, this model requires importing the model into the Tensor-RT architecture, performing inter-layer fusion and accuracy calibration, and then performing reverse engineering and optimization to reduce double-precision to single-precision, thereby reducing memory usage by 50%, making the model more lightweight, and doubling the number of parallel models that a single GPU can handle. This work is related to specific GPUs or DCUs, and the system has excellent domestic support.
[0173] Furthermore, in terms of the specific handling of details, the technical solution of the present invention also adopts the following methods:
[0174] 1. Use transformer-based power scene recognition to classify power data:
[0175] A 1D CNN (1D Convolutional Neural Network) kernel is used to convolve the time series data, projecting each local window into an embedding vector. Each token carries a short-term pattern of the time series. Location embeddings are added to the tokens, which are then passed through a Transformer model to learn long-term dependencies between these tokens. The Transformer model outputs the latent embedding vectors of the time series, which are then processed by a Multilayer Perceptron (MLP) with softmax activation to generate a token classification output.
[0176] 2. Use differential privacy to protect data in the power dispatching system;
[0177] Based on existing scheduling data, we analyze data from different address sources, establish a similarity variance matrix, and verify the mechanism of adding verification noise to the matrix. This achieves the protection of critical data.
[0178] 3. Perform privacy-preserving computations based on the SQL proxy mechanism and data classification.
[0179] The SQL (Structured Query Language) agent is implemented, and different registration queries can be implemented according to the levels of data when data is queried.
[0180] Specifically, data is obtained from a table according to a hierarchical structure through recursion and JOIN operation, and the data is arranged according to the hierarchy, so that a large amount of structured data can be processed, and the hierarchical query of data can be efficiently implemented.
[0181] Five, differential privacy:
[0182] The differential privacy in the technical scheme of the application is implemented by using the architecture shown in the figure. Figure 5
[0183] The data processing flow is as shown in the figure. Figure 6
[0184] In summary, information system security is the prerequisite and necessary condition for realizing informatization, and security management is an important guarantee for the safe operation of the system. Under the current large-scale network security device distribution, heterogeneity and complexity environment, only by establishing a unified and dynamic security boundary management, can the information resources be fully utilized, and the efficiency of the system be brought into play. Security management refers to the management of all security problems and links of the system, and its purpose is:
[0185] (1) Combine the platform, hardware and security boundary strategy to form a unified layered defense system to block illegal users from entering to reduce the possibility of network system damage.
[0186] (2) Through log auditing and tracking of illegal activities, a method for quickly detecting illegal use of network resources and rapidly determining the location of illegal entry is provided.
[0187] (3) Network managers can quickly block damaged files or applications, quickly respond and update security policies to minimize losses and facilitate system recovery.
[0188] Embodiment:
[0189] An adaptive model CL-BERT for data security level recognition based on an attention mechanism is constructed, which is similar to the BERT and Baidu ERNIE deep natural language models, and the overall model is as shown in the figure. Figure 7
[0190] Privacy compute is a computing theory and method for protecting the whole life cycle of privacy information, which is a computable model and axiomatic system of privacy measurement, privacy leakage cost, privacy protection and privacy analysis complexity when the ownership, management right and use right of privacy information are separated, and the overall process is as shown in Figure 8
[0191] Specifically, the overall process of the computable model and axiomatic system is as follows:
[0192] First, a data warehouse of the power dispatching system data is constructed, and a corresponding association relationship is established with a metadata management server and a privacy management server; the corresponding secret levels of different data are established in the metadata management server to cope with the metadata extraction of the power dispatching system data; and a differential privacy module is established in the privacy management server to provide privacy budget management.
[0193] When external data of the power dispatching system needs to be extracted, a secure intelligent level and secret service system is used.
[0194] The secure intelligent level and secret service system receives a query request from the outside and performs the following steps:
[0195] 1) The outside proposes an SQL query request;
[0196] 2) SQL analysis and rewriting are performed;
[0197] 3) Sensitivity analysis is performed;
[0198] 4) A sensitivity query request is proposed to the metadata management server;
[0199] 5) The metadata management server returns the sensitivity query result;
[0200] 6) Data source access is performed;
[0201] 7) The query step is performed to the data warehouse;
[0202] 8) The data warehouse returns the query result;
[0203] 9) Privacy budget allocation is performed;
[0204] 10) A request is proposed to the privacy budget management server;
[0205] 11) The privacy budget management server provides the privacy budget;
[0206] 12) Noise addition is performed;
[0207] 13) A query section is constituted;
[0208] 14) The query result is output.
[0209] The order of the above execution steps, the query information flow direction, and the corresponding association relationship between the data warehouse, the metadata management server and the privacy management server can be seen in Figure 8 .
[0210] From the perspective of application function, the system of the technical solution of the application adopts hierarchical design, divides the application program into multiple levels, and each level has a specific responsibility and function as shown in Figure 9 .
[0211] The service layer is responsible for processing the logic of the user interface and user input and output, and sending the user's request to the next layer of the application program.
[0212] The business logic layer is responsible for processing the core business logic of the application program, including data processing, calculation and verification functions.
[0213] The data access layer is responsible for interacting with the data storage system (such as a database) and performing data reading and writing operations.
[0214] External data accesses business through a secure SSL interface to ensure the security of the business. The overall process is shown in Figure 10 .
[0215] The structure and functions include: file classification and encryption model, classification and encryption interface module, classification and encryption unified service module, data cleaning and preprocessing module, model training module, classification and encryption task scheduling module, computing engine hardware adaptation module, statistical machine learning algorithm module, neural network algorithm module and configuration management module.
[0216] Through the implementation of the technical solution of the application, the user electricity feature recognition accuracy reaches 93%, the privacy leakage risk is reduced to 0.3%, and the cross-provincial load prediction error rate can be controlled within 3%.
[0217] The technical solution of the application uses data recognition and protection in the power dispatching system, realizes data security and sharing without affecting the statistical situation by specific recognition and processing of data, adopts an artificial intelligence method to identify data security levels for massive data in the business data of the power grid company, especially the difference privacy of control data and non-numeric information, adds different differential disturbances to data of different security levels, realizes differentiated data sharing through specific recognition and processing of data, deploys a differentiated data sharing software platform based on intelligent data recognition and differential privacy protection in the data sharing of each department, each level of dispatching and each company, realizes minimum redundancy differentiated data sharing, provides strong support for the safe operation of the State Grid power company network, and improves the overall security defense level of the power grid.
[0218] The present application can be widely used in the field of security management of power grid business data and dispatch data.
Claims
1. A power grid data sharing method based on intelligent data identification and differential privacy protection, characterized by: 1) Using power scene recognition based on the transformer model, multi-level deep semantic mining is performed on the information in the text to classify the data of the power dispatching system and form corresponding security levels for different data. 2) Identify the data security level, and then add different differential perturbations for data with different security levels; 3) Construct a secure and intelligent classification and confidentiality service system, and use differential privacy to protect the data of the power dispatching system; 4) By identifying, processing, and managing the classification and confidentiality of data in the power dispatching system, it provides easy-to-use service interfaces, including classification and confidentiality of sensitive data, documents, and text data streams. 5) External data accesses the power dispatching system through a secure SSL interface and an easy-to-use service interface, enabling differentiated data sharing within the power dispatching system.
2. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 1, characterized in that: The aforementioned security intelligent classification and confidentiality service system includes a classification and confidentiality interface module, a unified classification and confidentiality service module, a data cleaning and preprocessing module, a model training module, a classification and confidentiality task scheduling module, a computing engine hardware adaptation module, a statistical machine learning algorithm module, a neural network algorithm module, and a configuration management module.
3. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 2, characterized in that: The aforementioned classification and confidentiality interface module, also known as the unified application and call interface, uses modular technology to encapsulate the interface and provide a unified service interface for the application layer. The classification and confidentiality interface module uses a domestic hardware platform, including domestic CPUs and GPUs, supports cluster deployment and elastic expansion, and provides classification and confidentiality services on demand, ensuring the high reliability and scalability of the classification and confidentiality service platform. The aforementioned unified classification and confidentiality service module provides integrated, multi-layered, and scenario-based classification and confidentiality service capabilities. Through application methods including general text file classification and confidentiality services, typical official document classification and confidentiality services, and real-time text data stream classification and confidentiality services, it adapts to differentiated low-estimate confidentiality services in different business scenarios. Through scenario-based classification and confidentiality service encapsulation, it provides easy-to-use service interfaces, including sensitive data classification and confidentiality, file classification and confidentiality, and text data stream classification and confidentiality, providing applications with easy-to-use and convenient classification and confidentiality services. The data cleaning and preprocessing module cleans and converts the files and data to be classified and confidential into the data required by the model, and works in conjunction with the scheduling module to cache the input and output data. The model training module employs supervised and unsupervised training, including receiving batch labeled sample files and training an artificial intelligence model for use in the classification and confidentiality services; the unsupervised training function uses actual data to achieve unsupervised machine learning, thereby correcting model parameters and improving accuracy. The aforementioned task scheduling module for classifying and classifying tasks uses Tensor-RT to achieve asynchronous task inference and prediction, thereby enabling high-performance processing of large-scale data. The aforementioned computing engine hardware adaptation module modifies deep learning foundational middleware platforms, including TensorFlow, to adapt to domestically produced GPUs / DCUs and supports all domestically produced hardware. The statistical machine learning algorithm module supports statistical machine learning algorithm engines including Bayesian / CRF conditional random fields, Gaussian mixtures, K-means, and LDA, further improving the accuracy of the classification and density model; The neural network algorithm module mentioned above is the main model of the file security intelligent classification and confidentiality service system, specifically involving the CL-BERT model, which is an algorithm model based on the attention mechanism; The configuration management module provides system management functions including communication parameters, user permissions, model parameters, and certificates.
4. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 1, characterized in that... In step 1), a one-dimensional convolutional neural network kernel is used to convolve the time series, projecting each local window into an embedding vector. Each token carries the short-term pattern of the time series. Positional embeddings are added to the tokens, and then the tokens are passed through a Transformer model to learn the long-term dependencies between these tokens. The Transformer model outputs the latent embedding vector of the time series, which is then processed by a multilayer perceptron (MLP) with softmax activation to generate a label classification output.
5. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 1, characterized in that... In step 2), differential privacy is used to protect the data of the power dispatching system. Based on the existing dispatching data, data from different address sources are analyzed to establish a similarity variance matrix. The mechanism of adding verification noise to the matrix is verified to achieve the protection of key data.
6. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 1, characterized in that: In step 3), privacy calculations are performed based on the SQL proxy mechanism and data classification; when querying data, different levels of queries are implemented according to the data classification.
7. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 6, characterized in that: The implementation of different level queries includes recursion and JOIN operations to retrieve data from the table according to the hierarchical structure and arrange it according to the hierarchy, in order to process a large amount of structured data and realize the level query of the data.
8. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 1, characterized in that: In step 3), the security intelligent classification and confidentiality service system responds to external SQL query requests according to the following steps: 1) An external SQL query request is submitted; 2) Perform SQL parsing and rewriting; 3) Conduct sensitivity analysis; 4) Submit a sensitivity query request to the metadata management server; 5) The metadata management server returns the sensitivity query results; 6) Perform data source access; 7) Perform a query on the data warehouse; 8) The data warehouse returns the query results; 9) Implement privacy budget allocation; 10) Make a request to the privacy budget management server; 11) The privacy budget management server provides a privacy budget; 12) Add noise; 13) Construct the query interface; 14) Output the query results.
9. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 1, characterized in that: The aforementioned security intelligent classification and confidentiality service system adopts a layered design, dividing the application into multiple layers, each with specific responsibilities and functions: 1) Service layer: Responsible for handling the logic of user interface and user input / output, and sending user requests to the next layer of the application; 2) Business Logic Layer: Responsible for handling the core business logic of the application, including functions such as data processing, calculation, and verification; 3) Data Access Layer: Responsible for interacting with the data storage system and performing data reading and writing operations.
10. The power grid data sharing method based on intelligent data identification and differential privacy protection according to claim 1, characterized in that: The power grid data sharing method based on intelligent data identification and differential privacy protection utilizes data identification and protection in the power dispatching system. Through specific identification and processing of data, it achieves data confidentiality and sharing without affecting statistical analysis. Regarding differential privacy for control data and non-numerical information, artificial intelligence methods are employed to identify data security levels in massive datasets. Different differential perturbations are then added to data of different security levels, achieving differentiated data sharing through specific identification and processing. By introducing a Tensor-RT real-time processing software architecture, inter-layer network fusion and accuracy calibration are completed. The model is reverse-analyzed and optimized, reducing double precision to single precision. By deploying a differentiated data sharing software platform based on intelligent data identification and differential privacy protection in data sharing across departments, dispatching levels, and companies, it achieves differentiated data sharing with minimal redundancy, providing strong support for the safe operation of the power company's network and improving the overall security defense level of the power grid.
Citation Information
Patent Citations
A method, system, and apparatus for finding the intersection of privacy sets for cross-industry sharing of power grid data.
CN114614974B
Data processing method and system of power grid and storage medium
CN119885258A