Power data security compliance management method and system based on intelligent grading
By constructing a power safety knowledge graph and a dynamic grading method, combined with federated learning and user behavior analysis, the problems of data format diversity and security risks in traditional power data management are solved, and efficient, real-time data management and security protection are achieved.
Patent Information
- Application Number
- CN202511178325.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-22
AI Technical Summary
The traditional power data management model faces the problems of wide data sources and diverse formats, low manual processing efficiency, and difficulty in meeting real-time and accuracy requirements. The static classification method cannot adapt to the dynamic changes in the power business. There are security risks in cross-regional data collaborative analysis. Traditional encrypted transmission and permission control cannot balance data sharing needs and privacy protection requirements.
A Transformer-based model is used to build a power safety knowledge graph, combined with federated learning and real-time data analysis, to dynamically adjust data classification. User behavior is analyzed through MPNN and LSTM models to generate dynamic access policies. SIEM systems and NLP technology are used to identify abnormal behavior, automatically isolate risky accounts, and ensure data security and compliance.
It achieves efficient fusion and semantic understanding of multi-source heterogeneous data, dynamically adjusts data classification, significantly improves the accuracy and adaptability of data management, reduces human intervention costs and security risks, and provides a sustainable intelligent protection system.
Smart Images

Figure CN120672510A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management, and in particular to a method and system for power data security and compliance management based on intelligent grading. Background Art
[0002] Driven by the development of new power systems, the power industry is accelerating its transformation toward digitalization and intelligence, with data becoming increasingly prominent as a core production factor. The power system encompasses multiple links, including power generation, transmission, distribution, and consumption. The resulting multi-source, heterogeneous data (such as equipment operating data, user electricity usage information, and grid topology data) is massive and highly valuable. Simultaneously, various regulations have been introduced, imposing strict requirements for the classified, graded, and protected use of power data. Data security has become a critical component of both energy and national security. However, traditional power data management models face numerous challenges. First, data comes from a wide range of sources and in diverse formats, with structured, unstructured, and semi-structured data coexisting. Manual processing is inefficient and error-prone, making it difficult to meet real-time and accuracy requirements. Second, static data classification methods cannot adapt to the dynamic changes in the power business. For example, the sensitivity of real-time data generated during equipment failures increases significantly, and traditional classification standards cannot respond in a timely manner. Furthermore, cross-regional data collaborative analysis poses security risks, and traditional encryption transmission and permission control methods cannot balance data sharing needs with privacy protection requirements. Summary of the Invention
[0003] In order to solve the above problems, the purpose of the present invention is to provide a power data security compliance management method and system based on intelligent classification, which significantly improves the accuracy and adaptability of data compliance management.
[0004] To achieve the above object, the present invention adopts the following technical solutions: A method for power data security compliance management based on intelligent grading includes the following steps: S1: Acquire multi-source power data, including structured data, unstructured data, and semi-structured data, and pre-process them; S2: Based on preprocessed multi-source power data, we construct a power safety knowledge graph. We use a Transformer-based model to automatically extract entities and their relationships from text data. We also regularly crawl the latest regulations and automatically update the regulatory nodes and relationships in the power safety knowledge graph through semantic matching. S3: Based on the real-time collection of equipment operating status and user behavior logs, the federated learning model analyzes cross-regional data, predicts changes in data sensitivity, and uses the rules in the power safety knowledge graph to perform logical reasoning combined with real-time data to generate dynamic classification results. S4: Based on the dynamic classification results, user behavior data, and permission rules in the knowledge graph, the MPNN-based user credit evaluation model is used to analyze user access behavior and evaluate user credit values. Dynamic access policies are generated by combining the permission rules and user credit values in the power safety knowledge graph. S5: Based on the generated dynamic access policy, logs are collected centrally through the SIEM system, and NLP technology is used to parse the log content and extract abnormal behaviors. Abnormal behaviors trigger emergency response rules in the power safety knowledge graph and automatically isolate related accounts.
[0005] Further preprocessing is as follows: Structured data preprocessing: Missing values are interpolated using time series, Z-Score standardization is used to eliminate dimensional differences, and wavelet transform is combined to filter out high-frequency noise. Peak and valley power consumption and load fluctuation characteristics are extracted from user electricity consumption records to construct a time series feature matrix. Unstructured data preprocessing, including text and image data: A pre-trained model is used to extract entities from fault reports, and a rule engine is used to extract time and location information. The equipment manual text is converted into word or sentence vectors. Image data uses the target detection model to identify device labels and abnormal areas and extract image features; Semi-structured data preprocessing: Map the protection device's XML configuration into key-value pairs, store them in a NoSQL database, parse the nodes and connection relationships in the CIM / E file, and construct a graph-structured adjacency matrix; And perform data desensitization, as follows: Use the BERT-based named entity recognition model to identify sensitive entities in text, including user names, address text, ID numbers, bank card numbers, and user electricity addresses; For user name and address text information, random replacement is used to desensitize it; For ID card number and bank card number information, AES-256 symmetric encryption algorithm is used for encryption to ensure the security of data during storage and transmission; For the user's electricity address, the detailed address is generalized to the province and city levels.
[0006] Furthermore, based on the pre-processed multi-source power data, we constructed a power safety knowledge graph and used a Transformer-based model to automatically extract entities and their relationships from the text data. The details are as follows: The entity types defined include five core entity types: device entity, user entity, regulatory entity, operation entity and risk entity; The relationships between entities are established through business logic and mathematical models, including equipment association relationships, regulatory constraint relationships, and risk impact relationships; The device association relationship describes the physical connection between devices, and the relationship weight is calculated using the Haversine formula to calculate the geographical distance between devices; Regulatory constraints, indicating compliance of data classification and access control with regulations; Risk impact relationship, establishing the causal relationship between equipment failure, operational abnormality and data security; Attribute definitions include device entity attributes and regulatory entity attributes. The device entity attributes include numeric attributes, timestamp attributes, and spatial coordinate attributes. Regulatory entity attributes are text-based and record the release time, scope of application, and revision history of regulations to ensure the accuracy of compliance checks. According to the defined entities, relations and attributes, entity relation extraction is performed based on BERT and GNN; The extracted entities, relationships, and attributes are converted into entity, relationship, and entity triples, and attribute information is added to form the basic unit of the knowledge graph. The triples are stored in the Neo4j graph database, and a network structure is constructed based on node labels and relationship types. When the device position changes, recalculate the attention weight of node k to neighbor node m ; When the new regulations new When joining, calculate k new The similarity sim(q,k new ), if sim(q,k new )>τ, τ is the preset threshold, then a constraint relationship is added; Adjust conditional probabilities based on real-time statistics and update risk relationships: ;
[0007] Where β is the smoothing factor and II is the indicator function; is the updated risk relationship; is the risk relationship before the update; Event is the risk triggering event; is the data sensitivity change.
[0008] Furthermore, according to the defined entities, relations, and attributes, entity relation extraction is performed based on BERT and GNN, as follows: Input text data into the BERT model to generate context-related word vector representations: ; Among them, x i is the i-th token in the text, n is the total number of tokens in the text, h iFor the corresponding hidden layer vector, semantic modeling of professional terms in the power field is realized; After the BERT output layer, the conditional random field CRF is connected to model the dependency between the labels y and output the optimal entity label sequence. ; Probability of CRF The calculation formula is: ;
[0009] Among them, ψ i is the transfer score of adjacent labels, Z(H) is the normalization factor; H is the output of the BERT output layer; y i-1 、y i are the i-1th and i-th labels in the sequence respectively; Combine the entities identified by NER in pairs to construct candidate relationship pairs to be classified; For each recognized entity k, the text span is [s k ,e k ], whose initial node embedding is : ; Among them, s k 、e k are the starting and ending positions of entity k in the text respectively; Through the graph attention network GAT feature aggregation, the text is converted into a dependency graph, with words as nodes and grammatical dependencies as edges. The graph attention network GAT is updated as follows: ; in, is the set of neighbor nodes of node k; W (l) is the trainable weight of the lth layer; is the attention weight of node k to neighbor node m; is the activation function; is the l+1th layer output of node k; is the output of the lth layer of node m; Calculate the attention weights between nodes through the graph attention network GAT , aggregate entity context information: ; in, is the activation function; W is the weight matrix; h k 、h m Represent the feature vectors of node k and node m respectively; a represents the trainable weight vector used to calculate the attention score; T represents transpose; The aggregated entity features are input into the multi-layer perceptron and the relationship type is predicted using the softmax function: ; Among them, W r is the weight of the multilayer perceptron; Represents a given entity , the probability that the relationship between them is r; and Entity Embedding output at layer L; It means concatenating two embedding vectors.
[0010] Furthermore, we regularly crawl the latest regulations and automatically update the regulatory nodes and relationships in the power safety knowledge graph through semantic matching, as follows: Regularly obtain the latest regulations through web crawlers, and convert them into structured text after sentence segmentation, word segmentation, and part-of-speech tagging using NLP to ensure the real-time nature of regulatory knowledge; Use the Sentence-BERT model to generate sentence vectors for regulatory texts, and use cosine similarity to determine the degree of correlation between new and old regulations: Sim(s1,s2): ; Among them, s1 and s2 represent the new regulations and the old regulations respectively; SBERT(s1) and SBERT(s2) represent the new regulations and the old regulations respectively. The sentence vectors of the regulatory text are generated using the Sentence-BERT model; represents the Euclidean norm; If Sim(s1,s2) exceeds the threshold, it is determined to be a relevant regulation and the update process is triggered; Through the entity linking technology DBpedia Spotlight, entities in new regulations are matched with existing entities in the graph, and the rule nodes associated with the entities in the new regulations are automatically updated to ensure consistency between regulatory requirements and security policies.
[0011] Furthermore, based on the real-time collection of device operating status and user behavior logs, a federated learning model is used to analyze cross-regional data and predict changes in data sensitivity; The federated learning model uses an LSTM neural network to process time series data and a fully connected layer to output sensitivity scores. The local training process is as follows: Equipment state feature sequence X=[x1,x2,…,x t ,…,x T’ ], x t represents the device state at time t, T' is the total time sequence; the user behavior feature is Y; LSTM layer extracts temporal dependency h t:h t =LSTM(h t-1 ,x t ),t=1,…,T' The fully connected layer outputs the sensitivity prediction value of the i-th token , represents the probability of data sensitivity; The central server collects the model parameters of each region and uses the FedAvg algorithm for weighted averaging to obtain the model parameters :
[0012] in, For the The model parameters of the region, For the The amount of data in each region, is the total amount of global data; is the number of regions; Output of federated learning Combined with historical sensitive events in the knowledge graph, a comprehensive sensitivity score is generated through a weighted fusion formula:
[0013] in, and is the weight coefficient, KG_RuleScore is the knowledge graph rule matching score; Using the Datalog language definition, the data classification standards are converted into executable rules for the knowledge graph. The association between real-time data features and knowledge graph nodes is queried through SPARQL, and mapped to levels 1-5 according to the comprehensive sensitivity score S: ;
[0014] Based on the predefined rules of the knowledge graph, sensitivity upgrade rules are triggered.
[0015] Furthermore, the user credit evaluation model based on MPNN is as follows: Hierarchical messaging: Message generation: Node a sends message m to its neighbor node b a,b , generate messages based on node attributes and edge attributes: ; Among them, v a represents the feature vector of node a, W m 、U m is the trainable weight matrix; ReLU is the activation function; Δt a,b Represents the edge features between node a and its neighbor node b; Message aggregation: Node a aggregates all neighbor messages {m a,b}, using mean pooling, we get the neighbor aggregation feature vector of node a : ;
[0016] Where N(a) is the neighbor set of node a; Node status is updated, and the updated feature vector of node a is obtained : ; Among them, W u is the weight coefficient; Global feature aggregation: Graph pooling is used to obtain the global representation g of the user behavior graph, which is then input into the fully connected layer to generate the credit value C∈[0,1]: ; Among them, W c and b c are the weights and biases of the fully connected layer.
[0017] Furthermore, by combining the permission rules and user credit values in the power safety knowledge graph, a dynamic access policy is generated, as follows: The knowledge graph predefined rules include hierarchical association rules and credit threshold rules, and the rules are formally represented using conditional expressions; Access control strength Control_Level calculation: Control_Level=W d ×DataRisk(d)+W c ×BehaviorRisk(u) Among them, DataRisk(d) is the risk value corresponding to the data classification, BehaviorRisk(u) is the user behavior risk value; W d 、W c The weights corresponding to data classification and user behavior respectively; Get the data d and data classification L requested by user u d ; Credit value query: Get the user's current credit value C from the real-time database u ; Rule matching: Check the mandatory rules of d in the knowledge graph, calculate the control strength Control_Level, match the policy mapping table, and generate an access policy that includes authentication method, operation permission, and time limit.
[0018] Furthermore, based on the generated dynamic access policy, logs are collected centrally through the SIEM system, and NLP technology is used to parse the log content and extract abnormal behaviors. Abnormal behaviors trigger emergency response rules in the power safety knowledge graph and automatically isolate related accounts. The details are as follows: The SIEM system detects an anomaly, generates an event ID, and calculates the risk level Risk_Level based on the event type, data classification, and user credit value: Risk_Level=data classification×L1+behavior risk value×L2; Among them, L1 and L2 are weight coefficients; If the risk level is ≥ P, account isolation measures are triggered. The account management system is called through the API to perform the following immediate isolation measures on the risky user account; Risk level = P-1, triggering access throttling, limiting the number of concurrent users to 1 and the amount of exported data to the preset value; After the isolation operation is completed, the current status of the user node in the knowledge graph is updated to isolated, and the disposal timestamp is recorded.
[0019] An electric power data security and compliance management system based on intelligent grading includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically performs the steps in the electric power data security and compliance management method based on intelligent grading as described above.
[0020] The present invention has the following beneficial effects: 1. This invention integrates structured, semi-structured, and unstructured data in the power system, automatically constructs and updates knowledge graphs through the Transformer model, and achieves efficient fusion and semantic understanding of multi-source heterogeneous data. It also combines federated learning and real-time data analysis to dynamically adjust data classification, significantly improving the accuracy and adaptability of data management and providing a reliable data foundation for subsequent security strategies. 2. This invention uses a message passing neural network (MPNN) and LSTM model to implement real-time analysis and credit assessment of user behavior, dynamically generating access policies (such as triggering multi-factor authentication or permission restrictions). Combined with rule-based reasoning within the knowledge graph, the system can automatically respond to changes in data sensitivity and abnormal behavior (such as batch export during off-hours), forming a closed-loop security mechanism from prediction to protection, significantly reducing the cost of human intervention and security risks. 3. Based on the SIEM system and NLP technology, the present invention can quickly parse logs and identify key events (such as unauthorized access), automatically isolate risky accounts through emergency rules in the knowledge graph, and at the same time, by crawling the latest regulations and automatically updating the knowledge graph, ensure that data classification and access policies always meet the requirements, taking into account the real-time and compliance of security management, and providing a sustainable intelligent protection system for the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0022] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: refer to Figure 1 In this embodiment, a method for power data security and compliance management based on intelligent grading is provided, comprising the following steps: S1: Acquire and preprocess multi-source power data, including structured data (including real-time equipment operation data (voltage / current), user electricity consumption records, and weather data), unstructured data (fault repair reports, equipment manuals, inspection photos, and laws and regulations), and semi-structured data (protection device parameter configuration and network topology files). S2: Based on preprocessed multi-source power data, we construct a power safety knowledge graph. We use a Transformer-based model to automatically extract entities and their relationships from text data. We also regularly crawl the latest regulations and automatically update the regulatory nodes and relationships in the power safety knowledge graph through semantic matching. S3: Based on real-time data collection of device operating status and user behavior logs, a federated learning model analyzes cross-regional data to predict changes in data sensitivity. For example, real-time operating data during equipment failures is automatically upgraded to core data. Leveraging rules in the power safety knowledge graph (such as "core data must be transmitted encrypted"), logical reasoning is combined with real-time data to generate dynamic classification results. For example, when the knowledge graph detects that a certain type of data is associated with a high-risk scenario, the classification is automatically adjusted to produce a data classification result (levels 1-5). S4: Based on dynamic classification results, user behavior data, and permission rules in the knowledge graph, a user credit assessment model based on MPNN (message passing neural network) is used to analyze user access behavior (such as login frequency and operation path) and assess user credit. For example, abnormally high access frequency triggers a credit score drop, automatically restricting permissions. Dynamic access policies are generated by combining permission rules in the power safety knowledge graph (such as "core data requires multi-factor authentication") with user credit. For example, when the credit score falls below a threshold, MFA is automatically triggered and access is restricted. S5: Based on the generated dynamic access policy, logs are collected centrally through the SIEM system, and NLP technology is used to parse the log content to extract abnormal behaviors (such as "unauthorized access" and "high-frequency data downloads" such as batch data exports during non-working hours). Abnormal behaviors trigger emergency response rules in the power safety knowledge graph and automatically isolate related accounts.
[0023] In this embodiment, the structured data includes real-time equipment operation data, user electricity consumption records, and weather data; Real-time equipment operation data: Using the SCADA system's Modbus / TCP and IEC 61850 protocols, edge computing gateways enable real-time data collection. Smart terminals are deployed in substations and distribution rooms to upload voltage, current, and other data to a time-series database (such as InfluxDB) at a frequency of seconds, ensuring real-time and continuous data. User electricity usage records: Utilize ETL tools (such as Apache NiFi) to periodically extract user electricity usage data, including electricity consumption and payment records, from the power marketing system's relational database (such as MySQL or Oracle). The extraction frequency can be set to hourly or daily based on business needs. Weather data: Obtain weather data, including temperature, humidity, and wind speed, through APIs (such as the China Meteorological Administration's data interface and OpenWeatherMap). Use Python's requests library to write scripts that call the APIs on a minute-by-minute basis and store the data in a relational database. Unstructured data includes troubleshooting reports, equipment manuals, inspection photos, and laws and regulations. Specifically: Troubleshooting reports and equipment manuals: Obtain PDF or Word files from a file management system (such as NAS storage) or corporate document repository. Use Python's PyPDF2 and python-docx libraries to read the file contents. Combined with OCR technology (such as Tesseract or Baidu AI Open Platform OCR), recognize the text in the scanned documents and convert them into processable text data. Inspection photos: Install an image capture app on mobile inspection terminals (such as smart helmets and handheld devices) and upload them to the data center in real time via the 5G network. Use a file naming convention (such as "device number_photographing time.jpg") to associate photos with devices.
[0024] Laws and Regulations: Use web crawlers to scrape regulatory documents from government websites and legal databases (such as Peking University Law Database). Use Python's Scrapy framework to write crawlers and schedule scheduled tasks (such as daily updates) to retrieve the latest regulatory text. Semi-structured data includes protection device parameters and network topology files, specifically: Protective device parameter configuration: Extract data from the protective device configuration file (XML, JSON format). Use Python's xml.etree.ElementTree library to parse the XML file and the json.loads() function to parse the JSON file. Then, convert the configuration parameters into structured data and store it in the database.
[0025] Network topology file: Obtain the network topology file (such as CIM format) from the power grid dispatching automation system, convert it into a format recognizable by a graph database (such as Neo4j) through a dedicated parsing tool, and construct the power grid topology relationship.
[0026] In this embodiment, the preprocessing is as follows: Structured data preprocessing: Missing values are interpolated using time series (e.g., linear interpolation, Z-Score normalization to eliminate dimensional differences, combined with wavelet transform (e.g., Daubechies basis function) to filter out high-frequency noise. Peak and valley power consumption and load fluctuation characteristics are extracted from user electricity consumption records to construct a time series feature matrix. Unstructured data preprocessing, including text and image data: For text data, pre-trained models (such as BERT) are used to extract entities (such as faulty device ID and fault type) from fault reports. Time and location information are then extracted using rule engines (such as regular expressions). The device manual text is converted into word embeddings (such as Word2Vec) or sentence embeddings (such as Sentence-BERT). Image data uses object detection models (such as YOLOv8) to identify device labels and abnormal areas (such as rust and cracks) and extract image features (such as ResNet-50). Semi-structured data preprocessing: Map the XML configuration of the protection device into key-value pairs and store them in a NoSQL database. Parse the nodes (such as substations and lines) and connection relationships in the CIM / E file to construct a graph adjacency matrix.
[0027] And perform data desensitization, as follows: Using a BERT-based named entity recognition (NER) model to identify sensitive entities in text, including user names, address text, ID numbers, bank card numbers, and user electricity addresses; For user name and address text information, random replacement is used for desensitization, such as replacing "Zhang San" with "random name"; For ID card number and bank card number information, AES-256 symmetric encryption algorithm is used for encryption to ensure the security of data during storage and transmission; For the user's electricity address, the detailed address is generalized to the province and city levels, such as "No. XX, XX Road, Haidian District, Beijing" is generalized to "Beijing City".
[0028] In this example, we construct a power safety knowledge graph based on preprocessed multi-source power data, and use a Transformer-based model to automatically extract entities and their relationships from text data, as follows: The entity types defined include five core entity types: device entity, user entity, regulatory entity, operation entity and risk entity; Equipment entities: Covers power equipment such as transformers, circuit breakers, and insulators. Data is derived from equipment records and inspection reports, and contains attributes such as equipment model, rated voltage, and geographic location.
[0029] User entities: include residential users and industrial users. The data comes from the power marketing system and is associated with user electricity usage records, sensitive information, and other attributes.
[0030] Regulatory entities: Includes the Data Security Law and industry standards. Data is crawled from government websites, recording attributes such as release date, revision version, and effective scope.
[0031] Operation entity: involves operations such as data access, encrypted transmission, and permission changes. The data comes from the log system and work order records, reflecting the trajectory of data processing behavior.
[0032] Risk entities: such as data leakage, unauthorized access, and equipment failure. The data comes from risk assessment and audit reports and is used to identify potential security threats.
[0033] The relationships between entities are established through business logic and mathematical models, including equipment association relationships, regulatory constraint relationships, and risk impact relationships; The device association relationship describes the physical connection between devices (such as the connection between a transformer and a busbar). The relationship weight is calculated using the Haversine formula to calculate the geographical distance between devices. Regulatory constraints, indicating compliance of data classification and access control with regulations (e.g., "data classification" complies with the "Electric Power Industry Data Classification and Grading Specification"); Risk impact relationship: establishing the causal relationship between equipment failure, operational anomalies, and data security (e.g., "equipment failure" leads to "increased data sensitivity"); Attribute definitions include device entity attributes and regulatory entity attributes. The device entity attributes include numerical attributes (rated voltage, operating temperature), timestamp attributes (commissioning time, maintenance time), and spatial coordinate attributes (geographic location). Regulatory entity attributes are text-based and record the release time, applicable scope, and revision history of the regulations to ensure the accuracy of compliance checks. According to the defined entities, relations and attributes, entity relation extraction is performed based on BERT and GNN; The extracted entities, relationships, and attributes are converted into entity, relationship, and entity triples (e.g., "Transformer A - Connection - Busbar B"), and attribute information (e.g., distance, voltage) is added to form the basic unit of the knowledge graph. A graph database such as Neo4j is used to store the triples, and a mesh structure is constructed using node labels and relationship types. When the device position changes, recalculate the attention weight of node k to neighbor node m ; When the new regulations new When joining, calculate k new The similarity sim(q,k new ), if sim(q,k new )>τ, τ is the preset threshold, then a constraint relationship is added; Adjust conditional probabilities based on real-time statistics and update risk relationships: ;
[0034] Where β is the smoothing factor and II is the indicator function; is the updated risk relationship; is the risk relationship before the update; Event is the risk triggering event; is the data sensitivity change.
[0035] In this embodiment, entity relationship extraction is performed based on the defined entities, relationships, and attributes using BERT and GNN, as follows: Input text data into the BERT model to generate context-related word vector representations: ; Among them, x i is the i-th token (word / character) in the text, n is the total number of tokens in the text, h i For the corresponding hidden layer vector, semantic modeling of professional terms in the power field is realized; After the BERT output layer, a conditional random field (CRF) is added to model the dependencies between labels y (e.g., “device model” followed by “parameters”) to output the optimal entity label sequence. ; Probability of CRF The calculation formula is: ;
[0036] Among them, ψ iis the transfer score of adjacent labels, Z(H) is the normalization factor to ensure the accuracy of entity recognition (such as distinguishing "transformer model" from "transformer location"), and H is the output of the BERT output layer; y i-1 、y i are the i-1th and i-th labels in the sequence respectively; Combine entities identified by NER in pairs (e.g., "transformer A" and "busbar B") to construct candidate relationship pairs to be classified; For each recognized entity k, the text span is [s k ,e k ], whose initial node embedding is : ; Among them, s k 、e k are the starting and ending positions of entity k in the text respectively; Through the graph attention network GAT feature aggregation, the text is converted into a dependency graph, with words as nodes and grammatical dependencies as edges. The graph attention network GAT is updated as follows: ; in, is the set of neighbor nodes of node k; W (l) is the trainable weight of the lth layer; is the attention weight of node k to neighbor node m; is the activation function; is the l+1th layer output of node k; is the output of the lth layer of node m; Calculate the attention weights between nodes through the graph attention network GAT , aggregate entity context information: ; in, is the activation function; W is the weight matrix; h k 、h m Represent the feature vectors of node k and node m respectively; a represents the trainable weight vector used to calculate the attention score; T represents transpose; The aggregated entity features are input into the multi-layer perceptron and the relationship type is predicted using the softmax function: ; Among them, W r is the weight of the multilayer perceptron; Represents a given entity , the probability that the relationship between them is r; and Entity Embedding output at layer L; Indicates concatenating two embedding vectors; In this embodiment, the latest regulations are crawled regularly, and the regulation nodes and associations in the power safety knowledge graph are automatically updated through semantic matching, as follows: Regularly obtain the latest regulations (such as the revised provisions of the Data Security Law) through web crawlers, and convert them into structured text after sentence segmentation, word segmentation, and part-of-speech tagging through NLP to ensure the real-time nature of regulatory knowledge; Use the Sentence-BERT model to generate sentence vectors for regulatory texts, and use cosine similarity to determine the degree of correlation between new and old regulations: Sim(s1,s2): ; Among them, s1 and s2 represent the new regulations and the old regulations respectively; SBERT(s1) and SBERT(s2) represent the new regulations and the old regulations respectively. The sentence vectors of the regulatory text are generated using the Sentence-BERT model; represents the Euclidean norm; If Sim(s1,s2) exceeds a threshold (e.g., 0.8), it is determined to be a relevant regulation and the update process is triggered; Through the entity linking technology DBpedia Spotlight, entities in new regulations (such as "data classification standards") are matched with existing entities in the graph, and their associated rule nodes are automatically updated (such as adjusting the judgment conditions for data classification) to ensure consistency between regulatory requirements and security policies.
[0037] In this embodiment, based on the real-time collection of device operating status and user behavior logs, a federated learning model is used to analyze cross-regional data and predict changes in data sensitivity. The federated learning model uses an LSTM neural network to process time series data and a fully connected layer to output sensitivity scores. The local training process is as follows: Equipment state feature sequence X=[x1,x2,…,x t ,…,x T’ ], x t represents the device state at time t, T' is the total time sequence; the user behavior feature is Y; LSTM layer extracts temporal dependency h t :h t =LSTM(h t-1 ,x t ),t=1,…,T'; The fully connected layer outputs the sensitivity prediction value of the i-th token , represents the probability of data sensitivity; The central server collects the model parameters of each region and uses the FedAvg algorithm for weighted averaging to obtain the model parameters : ; in, For the The model parameters of the region, For the The amount of data in each region, is the total amount of global data; is the number of regions; Output of federated learning Combined with historical sensitive events in the knowledge graph (e.g., "the probability of data sensitivity increases by 40% when equipment fails"), a comprehensive sensitivity score is generated through a weighted fusion formula: ;
[0038] in, and is the weight coefficient (the weight is adjusted according to the business scenario, and β=0.6 for the equipment failure scenario), and KG_RuleScore is the knowledge graph rule matching score; Using the Datalog language definition, the data classification standards are converted into executable rules for the knowledge graph. The association between real-time data features and knowledge graph nodes is queried through SPARQL, and mapped to levels 1-5 according to the comprehensive sensitivity score S: ; Based on the predefined rules of the knowledge graph, sensitivity upgrade rules are triggered; for example, the voltage fluctuation amplitude ΔV>20%×rated voltage, and the duration exceeds 10 minutes; the user access frequency exceeds 3 times the historical average, and the operation involves core data fields (such as "user ID" and "meter reading").
[0039] In this embodiment, the user credit evaluation model based on MPNN is as follows: Hierarchical messaging: Message generation: Node a sends message m to its neighbor node b a,b , generate messages based on node attributes and edge attributes: ; Among them, v a represents the feature vector of node a, W m 、U m is the trainable weight matrix; ReLU is the activation function; Δt a,b Represents the edge features between node a and its neighbor node b; Message aggregation: Node a aggregates all neighbor messages {m a,b}, using mean pooling, we get the neighbor aggregation feature vector of node a : ;
[0040] Among them, N(a) is the neighbor set of node a, and the pre-order and post-order operation nodes; Node status is updated, and the updated feature vector of node a is obtained : ; Among them, W u is the weight coefficient; Global feature aggregation: Graph pooling (such as SumPooling) is used to obtain the global representation g of the user behavior graph, which is then fed into a fully connected layer to generate a credit value C∈[0,1]: ; Among them, W c and b c are the weights and biases of the fully connected layer; a higher credit value indicates a more credible behavior, whereas a lower credit value indicates a higher risk (for example, abnormally high frequency of access to core data will reduce the credit value).
[0041] In this embodiment, the permission rules and user credit values in the power safety knowledge graph are combined to generate a dynamic access policy, as follows: Knowledge graph predefined rules include hierarchical association rules: such as "core data (level 5) → MFA authentication required" and "important data (level 4) → access requires approval"; credit threshold rules: such as "credit value C < 0.6 → trigger MFA" and "C < 0.4 → prohibit export operations"; The rules are formalized using conditional expressions, for example: Need_MFA(u,d)←DataLevel(d,5)∨(UserCredit(u,C)∧C<0.6) This means that when user u accesses data d, if the data is level 5 or the credit value is lower than 0.6, MFA authentication is required; Access control strength Control_Level calculation: Control_Level=W d ×DataRisk(d)+W c ×BehaviorRisk(u) Where DataRisk(d) is the risk value corresponding to the data classification (level 5 = 1, level 4 = 0.8, and so on), BehaviorRisk(u) is the user behavior risk value (1-C); W d 、W c The weights corresponding to data classification and user behavior respectively; Get the data d and its classification L requested by user u d ; Credit value query: Get the user's current credit value C from the real-time database u ; Rule matching: Check the mandatory rules of d in the knowledge graph (such as "Level 5 data must have MFA"), calculate the control strength Control_Level, match the policy mapping table, and generate an access policy that includes authentication method, operation permissions, and time limit.
[0042] In this embodiment, based on the generated dynamic access policy, logs are collected centrally through the SIEM system, and NLP technology is used to parse the log content and extract abnormal behaviors. Abnormal behaviors trigger emergency response rules in the power safety knowledge graph, and related accounts are automatically isolated. The details are as follows: SIEM system detects anomalies (such as LSTM-based time series anomaly detection, model architecture: Input layer: Encodes the log event sequence into a one-hot vector (e.g., “login failed” = 001, “data export” = 010).
[0043] LSTM layer: captures the temporal dependencies of event sequences. The hidden layer dimension is set to 128 and learns normal behavior patterns.
[0044] Reconstruction layer: reconstructs the input sequence through the autoencoder and calculates the reconstruction error.
[0045] If the LSTM reconstruction error exceeds the threshold or the rule matching succeeds, it is considered an exception and an event ID (such as E20250528_001) is generated. Calculate the risk level Risk_Level based on the event type, data classification, and user credit value: Risk_Level=data classification×L1+behavior risk value×L2; Among them, L1 and L2 are weight coefficients; Risk level ≥ 4 triggers account isolation measures, and the account management system is called through the API to immediately perform "isolation" measures on the risky user account (such as freezing, offline, revoking permissions, etc.); Risk level = Level 3, triggering access throttling, limiting the user's concurrent access to 1 and the amount of exported data to the preset value; After the isolation operation is completed, the current status of the user node in the knowledge graph is updated to isolated, and the disposal timestamp is recorded.
[0046] An electric power data security and compliance management system based on intelligent grading includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically performs the steps in the electric power data security and compliance management method based on intelligent grading as described above.
[0047] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0048] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0049] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0050] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0051] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.
Claims
1. A power data security compliance management method based on intelligent classification, characterized in that: The following steps are involved: S1: Acquire multi-source power data, including structured data, unstructured data, and semi-structured data, and pre-process them; S2: Based on pre-processed multi-source power data, we construct a power safety knowledge graph. We use a Transformer-based model to automatically extract entities and their relationships from text data. We also regularly crawl the latest regulations and automatically update the regulatory nodes and relationships in the power safety knowledge graph through semantic matching. S3: Based on the real-time collection of equipment operating status and user behavior logs, the federated learning model analyzes cross-regional data, predicts changes in data sensitivity, and uses the rules in the power safety knowledge graph to perform logical reasoning combined with real-time data to generate dynamic classification results. S4: Based on the dynamic classification results, user behavior data, and permission rules in the knowledge graph, the MPNN-based user credit evaluation model is used to analyze user access behavior and evaluate user credit values. Dynamic access policies are generated by combining the permission rules and user credit values in the power safety knowledge graph. S5: Based on the generated dynamic access policy, logs are collected centrally through the SIEM system, and NLP technology is used to parse the log content and extract abnormal behaviors. Abnormal behaviors trigger emergency response rules in the power safety knowledge graph and automatically isolate related accounts.
2. The power data security and compliance management method based on intelligent grading according to claim 1 is characterized in that: The preprocessing is as follows: Structured data preprocessing: Missing values are interpolated using time series, Z-Score standardization is used to eliminate dimensional differences, and wavelet transform is combined to filter out high-frequency noise. Peak and valley power consumption and load fluctuation characteristics are extracted from user electricity consumption records to construct a time series feature matrix. Unstructured data preprocessing, including text and image data: A pre-trained model is used to extract entities from fault reports, and a rule engine is used to extract time and location information. Convert the device manual text into word vectors or sentence vectors; Image data uses the target detection model to identify device labels and abnormal areas and extract image features; Semi-structured data preprocessing: Map the protection device's XML configuration into key-value pairs, store them in a NoSQL database, parse the nodes and connection relationships in the CIM / E file, and construct a graph-structured adjacency matrix; And perform data desensitization, as follows: Use the BERT-based named entity recognition model to identify sensitive entities in text, including user names, address text, ID numbers, bank card numbers, and user electricity addresses; For user name and address text information, random replacement is used to desensitize it; For ID card number and bank card number information, AES-256 symmetric encryption algorithm is used for encryption to ensure the security of data during storage and transmission; For the user's electricity address, the detailed address is generalized to the province and city levels.
3. The power data security compliance management method based on intelligent grading according to claim 1 is characterized in that: Based on the pre-processed multi-source power data, we construct a power safety knowledge graph and use a Transformer-based model to automatically extract entities and their relationships from text data. The details are as follows: The entity types defined include five core entity types: device entity, user entity, regulatory entity, operation entity and risk entity; The relationships between entities are established through business logic and mathematical models, including equipment association relationships, regulatory constraint relationships, and risk impact relationships; The device association relationship describes the physical connection between devices, and the relationship weight is calculated using the Haversine formula to calculate the geographical distance between devices; Regulatory constraints, indicating compliance of data classification and access control with regulations; Risk impact relationship, establishing the causal relationship between equipment failure, operational abnormality and data security; Attribute definitions include device entity attributes and regulatory entity attributes. The device entity attributes include numeric attributes, timestamp attributes, and spatial coordinate attributes. Regulatory entity attributes are text-based and record the release time, scope of application, and revision history of regulations to ensure the accuracy of compliance checks. According to the defined entities, relations and attributes, entity relation extraction is performed based on BERT and GNN; Convert the extracted entities, relationships, and attributes into entity, relationship, and entity triples, and add attribute information to form the basic unit of the knowledge graph; Neo4j graph database is used to store triples, and a network structure is constructed through node labels and relationship types; When the device position changes, recalculate the attention weight of node k to neighbor node m ; When the new regulations new When joining, calculate k new The similarity sim(q,k new ), if sim(q,k new )>τ, τ is the preset threshold, then a constraint relationship is added; Adjust conditional probabilities based on real-time statistics and update risk relationships: ; Where β is the smoothing factor and II is the indicator function; is the updated risk relationship; is the risk relationship before the update; Event is the risk triggering event; is the data sensitivity change.
4. The power data security compliance management method based on intelligent grading according to claim 3 is characterized in that: According to the defined entities, relations and attributes, entity relation extraction is performed based on BERT and GNN, as follows: Input text data into the BERT model to generate context-related word vector representations: ; Among them, x i is the i-th token in the text, n is the total number of tokens in the text, h i For the corresponding hidden layer vector, semantic modeling of professional terms in the power field is realized; After the BERT output layer, the conditional random field CRF is connected to model the dependency between the labels y and output the optimal entity label sequence. ; Probability of CRF The calculation formula is: ; Among them, ψ i is the transfer score of adjacent labels, Z(H) is the normalization factor; H is the output of the BERT output layer; y i-1 、y i are the i-1th and i-th labels in the sequence respectively; Combine the entities identified by NER in pairs to construct candidate relationship pairs to be classified; For each recognized entity k, the text span is [s k ,e k ], the initial node embedding of entity k is : ; Among them, s k 、e k are the starting and ending positions of entity k in the text respectively; Through the graph attention network GAT feature aggregation, the text is converted into a dependency graph, with words as nodes and grammatical dependencies as edges. The graph attention network GAT is updated as follows: ; in, is the set of neighbor nodes of node k; W (l) is the trainable weight of the lth layer; is the attention weight of node k to neighbor node m; is the activation function; is the l+1th layer output of node k; is the output of the lth layer of node m; Calculate the attention weights between nodes through the graph attention network GAT , aggregate entity context information: ; in, is the activation function; W is the weight matrix; h k 、h m Represent the feature vectors of node k and node m respectively; a represents the trainable weight vector used to calculate the attention score; T represents transpose; The aggregated entity features are input into the multi-layer perceptron and the relationship type is predicted using the softmax function: ; Among them, W r is the weight of the multilayer perceptron; Represents a given entity The probability that the relationship between them is r; and Entity Embedding output at layer L; It means concatenating two embedding vectors.
5. The power data security and compliance management method based on intelligent grading according to claim 4 is characterized in that: The latest regulations are crawled regularly, and the regulatory nodes and association relationships in the power safety knowledge graph are automatically updated through semantic matching, as follows: Regularly obtain the latest regulations through web crawlers, and convert them into structured text after sentence segmentation, word segmentation, and part-of-speech tagging using NLP to ensure the real-time nature of regulatory knowledge; Use the Sentence-BERT model to generate sentence vectors for regulatory texts, and use cosine similarity to determine the degree of correlation between new and old regulations: Sim(s1,s2): ; Among them, s1 and s2 represent the new regulations and the old regulations respectively; SBERT(s1) and SBERT(s2) represent the new regulations and the old regulations respectively. The sentence vectors of the regulatory text are generated using the Sentence-BERT model; represents the Euclidean norm; If Sim(s1,s2) exceeds the threshold, it is determined to be a relevant regulation and the update process is triggered; Through the entity linking technology DBpedia Spotlight, entities in new regulations are matched with existing entities in the graph, and the rule nodes associated with the entities in the new regulations are automatically updated to ensure consistency between regulatory requirements and security policies.
6. The power data security and compliance management method based on intelligent grading according to claim 1 is characterized in that: According to the real-time collection of device operation status and user behavior logs, cross-regional data is analyzed through a federated learning model to predict changes in data sensitivity; The federated learning model uses an LSTM neural network to process time series data and a fully connected layer to output sensitivity scores. The local training process is as follows: Equipment state feature sequence X=[x1,x2,…,x t ,…,x T’ ], x t represents the device state at time t, T' is the total time sequence; the user behavior feature is Y; LSTM layer extracts temporal dependency h t :h t =LSTM(h t-1 ,x t ),t=1,…,T' The fully connected layer outputs the sensitivity prediction value of the i-th token , represents the probability of data sensitivity; The central server collects the model parameters of each region and uses the FedAvg algorithm for weighted averaging to obtain the model parameters : ; in, For the The model parameters of the region, For the The amount of data in each region, is the total amount of global data; is the number of regions; Output of federated learning Combined with historical sensitive events in the knowledge graph, a comprehensive sensitivity score is generated through a weighted fusion formula: ; in, and is the weight coefficient, KG_RuleScore is the knowledge graph rule matching score; Using the Datalog language definition, the data classification standards are converted into executable rules for the knowledge graph. The association between real-time data features and knowledge graph nodes is queried through SPARQL, and mapped to levels 1-5 according to the comprehensive sensitivity score S: ; Based on the predefined rules of the knowledge graph, sensitivity upgrade rules are triggered.
7. The power data security and compliance management method based on intelligent grading according to claim 1 is characterized in that: The user credit evaluation model based on MPNN is as follows: Hierarchical messaging: Message generation: Node a sends message m to its neighbor node b a,b , generate messages based on node attributes and edge attributes: ; Among them, v a represents the feature vector of node a, W m 、U m is the trainable weight matrix; ReLU is the activation function; Δt a,b Represents the edge features between node a and its neighbor node b; Message aggregation: Node a aggregates all neighbor messages {m a,b }, using mean pooling, we get the neighbor aggregation feature vector of node a : ; Where N(a) is the neighbor set of node a; Node status is updated, and the updated feature vector of node a is obtained : ; Among them, W u is the weight coefficient; Global feature aggregation: Graph pooling is used to obtain the global representation g of the user behavior graph, which is then input into the fully connected layer to generate the credit value C∈[0,1]: ; Among them, W c and b c are the weights and biases of the fully connected layer.
8. The method for power data security and compliance management based on intelligent grading according to claim 7 is characterized in that: The above method combines the permission rules and user credit values in the power safety knowledge graph to generate a dynamic access policy, which is as follows: The knowledge graph predefined rules include hierarchical association rules and credit threshold rules, and the rules are formally represented using conditional expressions; Access control strength Control_Level calculation: Control_Level=W d ×DataRisk(d)+W c ×BehaviorRisk(u) Among them, DataRisk(d) is the risk value corresponding to the data classification, BehaviorRisk(u) is the user behavior risk value; W d 、W c The weights corresponding to data classification and user behavior respectively; Get the data d and data classification L requested by user u d ; Credit value query: Get the user's current credit value C from the real-time database u ; Rule matching: Check the mandatory rules of d in the knowledge graph, calculate the control strength Control_Level, match the policy mapping table, and generate an access policy that includes authentication method, operation permission, and time limit.
9. The power data security and compliance management method based on intelligent grading according to claim 8 is characterized in that: Based on the generated dynamic access policy, logs are collected centrally through the SIEM system, and NLP technology is used to parse the log content and extract abnormal behaviors. Abnormal behaviors trigger emergency response rules in the power safety knowledge graph and automatically isolate related accounts. The details are as follows: The SIEM system detects an anomaly, generates an event ID, and calculates the risk level Risk_Level based on the event type, data classification, and user credit value: Risk_Level=data classification×L1+behavior risk value×L2; Among them, L1 and L2 are weight coefficients; If the risk level is ≥ P, account isolation measures are triggered. The account management system is called through the API to immediately isolate the risk user account. Risk level = P-1, triggering access throttling, limiting the number of concurrent accesses to 1 and the amount of exported data to the preset value; After the isolation operation is completed, the current status of the user node in the knowledge graph is updated to isolated, and the disposal timestamp is recorded.
10. An intelligent hierarchical power data security and compliance management system, characterized by: It includes a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the power data security and compliance management method based on intelligent classification as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Transformer substation digital twinborn early warning decision-making method and system based on knowledge graph
CN118521433A
Power grid service intelligent auditing method and system based on AI enhancement
CN119599821A
Graph-model complementary driven power grid fault intelligent auxiliary analysis method and system
CN119940687A
Digital finance and tax auditing method based on artificial intelligence
CN120429359A
Power Grid Monitoring Using Network Devices in a Cable Network
US20210367448A1
Cited By
Construction method of trusted industrial data space
CN120934911A
Computer network security access control management method based on big data
CN121193499A
Computer network security access control management method based on big data
CN121193499B
Project ledger intelligent generation method and system based on AI
CN121235648A
Intelligent analysis method and system for industry power utilization transaction based on multi-modal large model
CN121765605A