Digital power grid network attack data numeralization method based on feature engineering
By employing feature engineering and data quantification methods, a dataset for digital power grid network attacks was constructed, resolving the issue of inconsistent data in existing technologies. This enabled rapid and reliable data support, providing an effective data foundation for the research and defense of digital power grid network attacks.
Patent Information
- Application Number
- CN202510868034.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies cannot provide rapid and unified network attack data, making it difficult to support research and defense against attacks on digital power grid networks.
By employing a feature engineering approach, threat intelligence data is collected, preprocessed, and mapped. The data is then quantified using the TRAM tool and TF-IDF method is used to quantify character-based data, thereby constructing an attack behavior dataset and providing fast and unified attack data.
It provides rapid and unified attack data for digital power grid network attacks, supporting subsequent research and defense analysis, and improving the integrity and reliability of the data foundation.
Smart Images

Figure BDA0005469127910000031 
Figure BDA0005469127910000032 
Figure BDA0005469127910000042
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network security technology, and in particular relates to a method for quantifying digital power grid network attack data based on feature engineering. Background Technology
[0002] Cyberattacks are generally defined as attempts to damage, expose, modify, disable, steal, or obtain unauthorized or illegal use of assets. They can be categorized into single-step attacks and multi-step attacks. Single-step attacks refer to cyberattacks with a single attack step, such as SQL injection attacks, port scanning, and Denial of Service (DoS) attacks. Complex multi-step attacks consist of at least two single-step attacks in a specific sequence, and may be launched simultaneously by one or more attackers against a specific target. For example, the Advanced Persistent Threat (APT) attacks currently affecting many higher education institutions, financial institutions, and government agencies fall under the category of multi-step attacks. APT attacks are highly targeted, covert, and destructive, with complex and varied attack methods and long durations. APT attacks have become a significant threat to the security of critical infrastructure.
[0003] With the widespread application of sensing components and communication equipment in power systems, modern power systems have become cyber-physical systems highly dependent on communication networks. Increased informatization not only enhances the system's real-time sensing and dynamic control capabilities but also exacerbates the risk of malicious cyberattacks. Current research on cyberattacks against power systems primarily focuses on improving system operating efficiency and the ability to cope with routine power failures, failing to provide rapid and unified attack data and thus lacking a data foundation for researching issues related to cyberattacks on digital power grids. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a numerical method for digital power grid network attack data based on feature engineering, so as to solve the technical problems that the current network attack analysis of the power system is mainly designed to improve the system's operating efficiency and the ability to deal with routine power failures, which cannot provide fast and unified attack data and cannot provide a data foundation for studying digital power grid network attack-related issues.
[0005] The technical solution of this invention is:
[0006] A method for quantifying digital power grid network attack data based on feature engineering, the analysis method comprising:
[0007] Step 1: Collect data on network attack behavior;
[0008] Step 2: Preprocess the data to construct an attack behavior dataset;
[0009] Step 3: Use the TRAM tool to map threat intelligence for the digital power grid;
[0010] Step 4: Data quantification based on feature engineering.
[0011] The collected network attack behavior data comes from the raw data of the network attack organization database provided by MITRE Corporation. Threat intelligence related to attacks on digital power grids is crawled through web crawlers, and non-research information is filtered out when obtaining the raw data.
[0012] The method for preprocessing data information includes: directly storing fields such as attack behavior IDs into a CSV file; for fields containing descriptive text data, performing unified field processing, first removing line breaks between paragraphs and converting them into single-line format; then performing sentence segmentation processing, dividing the single-line format text into single-sentence format text according to punctuation marks; and finally performing word segmentation processing before storing it into a CSV file.
[0013] The attack behavior dataset contains the attack tactics, techniques / sub-techniques, and corresponding textual descriptions of the attacking organization.
[0014] When using the TRAM tool to map threat intelligence for digital power grids, the attack behavior of attacking organizations is automatically extracted through TRAM, thereby mapping the threat intelligence into ATT&CK.
[0015] Methods for mapping threat intelligence to ATT&CK include: users submitting threat intelligence, submitting threat intelligence documents in txt or word format, or web addresses containing the threat intelligence to be analyzed.
[0016] The method of mapping threat intelligence to ATT&CK also includes: using the TRAM tool to retrieve submitted threat intelligence content, using the underlying logistic regression model to analyze and predict attack behavior, comparing it with information from the threat intelligence sharing platform, displaying the predicted attack behavior and its probability on the page, and automatically filtering the prediction results if the probability is lower than the user-defined threshold; finally, the user manually adjusts the feedback based on the prediction results and generates the corresponding threat intelligence in JSON format.
[0017] When quantifying data based on feature engineering, the term frequency-inverse document frequency method is used to calculate the contribution of each word; assuming the following situation, N texts contain M words, and each word is represented as t. i The text is represented as d j Therefore, a word set T = [t1, t2, ..., t] is formed. i ,...t mand text set D = d1, d2, ..., d j ,...d n ], Term Frequency (TF) represents word frequency (t). i In text d j The frequency of occurrence is calculated according to Formula 1. The larger the TF value, the more important the word is in the text.
[0018]
[0019] Where, n i,j The word t i In text d j Frequency of occurrence, ∑ k n k,j Represents text d j The number of words contained in the text is shown in Equation 2. Inverse text frequency (IDF) refers to the inverse relationship between the number of times a word appears and the importance of the word in the text. The larger the IDF value, the more important the word is in the text, and the more representative it is of the text features.
[0020]
[0021] Where N represents the total number of texts. Indicates that it contains the word t i The number of texts, where TF-IDF represents the importance of a word relative to the text set, is calculated according to Equation 3.
[0022]
[0023] Where w(t,d) represents the weight value of word t calculated using TE IDF.
[0024] The beneficial effects of this invention are:
[0025] This invention employs a feature-engineered digital power grid network attack data quantification technology. Targeting business scenarios with digital power grid characteristics, it analyzes the principles and conditions of applicable attack techniques, identifies effective technical links in the attack chain targeting the characteristics of the attacked points, and provides rapid and unified attack data. This provides a data foundation for research on digital power grid-related issues and offers effective data support for subsequent simulations and network attack defense.
[0026] This addresses the technical issue that current power system network attack analysis primarily focuses on improving system operating efficiency and the ability to cope with routine power failures, thus failing to provide a data foundation for research on network attacks related to digital power grids. Detailed Implementation
[0027] A method for quantifying digital power grid network attack data based on feature engineering, specifically including:
[0028] In the ATT&CK framework, attack tactics, techniques, and methods are presented in text form, which requires data processing before applying machine learning classification algorithms. Therefore, constructing a threat intelligence dataset and performing unified data processing is the first step in achieving classification.
[0029] This uses raw data from the Attack Organizations Repository provided by MITRE. The Attack Organizations Repository is an open-source knowledge base of attack behaviors, designed to organize threat intelligence and attack activities of attack organizations, thereby enabling security professionals to respond promptly and appropriately when facing system risks.
[0030] The fields included in the data of each attack organization are shown in the table below. This article selects ID, description, and technology / sub-technology to process the data of the digital power grid.
[0031] Table 1: Attack Organization Fields in the ATT&CK Framework
[0032]
[0033] The technical fields included in the organization are described in the table below. This article selects ID and description for data processing.
[0034] Table 2: Attack Technique Fields in the ATT&CK Framework
[0035]
[0036] To obtain the raw data from the ATT&CK cyberattack group database, a web crawler was used to extract threat intelligence related to attacks on digital power grids. This dataset contains over two thousand attack records from 130 attack groups. Therefore, this dataset fully reflects the threat intelligence information currently exposed by attack groups on the internet.
[0037] Use Python web scraping tools to crawl relevant information from HTML web pages. Since HTML web pages contain some non-research fields, such as references, preventative measures, and detection methods, it is necessary to filter out this non-research information when obtaining the raw data.
[0038] After the above filtering process, the initial raw data is obtained. However, since some text information is too lengthy and not conducive to model training, it is necessary to process the filtered raw data. On the one hand, fields such as attack behavior IDs are directly stored in a CSV file. On the other hand, fields containing descriptive text data need to be uniformly processed. First, the line breaks between paragraphs need to be removed, converting it to single-line format; then, sentence segmentation is required. Since single-line text is too long to be used as model input, this paper segments the single-line text into single-sentence text according to punctuation marks; finally, word segmentation is also required.
[0039] After the above processing, we can obtain the preprocessed data, namely the attack behavior dataset of the attacking organization, which mainly includes the attack tactics, techniques / sub-techniques of the attacking organization and corresponding text descriptions.
[0040] Due to the limited data in the ATT&CK attack group's database, much threat intelligence regarding digital power grids is not mapped to the ATT&CK framework, and mapping threat intelligence to the ATT&CK framework is difficult. The TRAM (ThreatReportATT&CKMapping) tool can effectively solve this problem. TRAM is an open-source network tool that can automatically extract the attack behavior of attack groups, thereby mapping threat intelligence to ATT&CK.
[0041] The process of mapping threat intelligence using the TRAM tool is mainly as follows: First, the user submits threat intelligence. The current submission specifications are threat intelligence documents in TXT or Word format, or web addresses containing the threat intelligence to be analyzed. Second, the TRAM tool retrieves the submitted threat intelligence content and uses an underlying logistic regression model to analyze and predict attack behaviors. This is mainly achieved by comparing the predicted attack behaviors with information from a threat intelligence sharing platform. The page displays the predicted attack behaviors and their probabilities. If the probability is lower than a user-defined threshold, the tool automatically filters out overly low prediction results to enhance the credibility and usability of the prediction results. Finally, the user can manually adjust the feedback based on the prediction results and generate corresponding threat intelligence in JSON format.
[0042] Through the above steps, we have gained a basic understanding of the fields in the acquired raw data and performed preliminary processing of the raw text data. However, when training the model, it is necessary to quantify the character-based data and convert it into computable numerical data, i.e., data quantification based on feature engineering.
[0043] Label encoding (LE) and one-hot encoding (OE) are commonly used for encoding character data, such as the attack tactics and techniques / sub-techniques studied in this invention. A comparison between label encoding and one-hot encoding in attack tactics is shown in the table below, where one-hot encoding is a vector composed of the first three horizontal columns of numbers, and label encoding is the last column of numbers.
[0044] Label encoding maps n data points to integers between 0 and n-1, meaning each data point is represented by a unique number. While this mapping method is simple and easy to understand, it does not consider the relationships between the data points, which may make the encoded data more difficult for the model to understand and learn.
[0045] One-hot encoding, also known as one-bit effective encoding, maps data to an n-dimensional vector, with the index containing the data mapped to 1 and the rest to 0. Therefore, the encoding result is a sparse vector representation, and the dimension of the vector is positively correlated with the amount of data n. If the amount of data is particularly large, it can easily cause the curse of dimensionality, increasing the difficulty of model training.
[0046] Table 3: Comparison of Tag Encoding and One-Hot Code
[0047]
[0048]
[0049] (2) For the processing of character text data such as related descriptions, this invention uses the Term Frequency-Inverse Document Frequency (TF-IDF) method to calculate the contribution of each word, which is the importance of the word in the text.
[0050] Suppose we have N texts containing M words, where each word is represented as t. i The text is represented as d j Therefore, a word set T = [t1, t2, ..., t] is formed. i ,...t m ] and text set D = [d1, d2, ..., d j ,...d n ] Term frequency (TF) represents word t i In text d j The frequency of occurrence is calculated according to Formula 1. The higher the TF value, the more important the word is in the text.
[0051]
[0052] Where, n i,j The word ti In text d j Frequency of occurrence, ∑ k n k,j Represents text d j The number of words contained within. As shown in Equation 2, Inverse Text Frequency (IDF) refers to the inverse relationship between the number of times a word appears and the importance of the text that the word represents. The larger the IDF value, the more important the word is in the text, and the more representative it is of the text features.
[0053]
[0054] Where N represents the total number of texts. Indicates that it contains the word t i The number of texts. TF-IDF represents the importance of a word relative to the text set, calculated according to Equation 3.
[0055]
[0056] Where w(t,d) represents the weight value of word t calculated using TE IDF.
[0057] This invention employs the aforementioned feature engineering-based digital power grid network attack data quantification technology. Targeting business scenarios with digital power grid characteristics, it analyzes the applicable attack technology principles and usage conditions, and identifies effective technical links in the attack chain targeting the characteristics of the attacked points. This provides rapid and unified attack data, offering effective data support for subsequent simulations and network attack defense.
Claims
1. A method for quantifying attack data in digital power grid networks based on feature engineering, characterized in that: The analytical method includes: Step 1: Collect data on network attack behavior; Step 2: Preprocess the data to construct an attack behavior dataset; Step 3: Use the TRAM tool to map threat intelligence for the digital power grid; Step 4: Data quantification based on feature engineering.
2. The method for quantifying digital power grid network attack data based on feature engineering according to claim 1, characterized in that: The collected network attack behavior data comes from the raw data of the network attack organization database provided by MITRE Corporation. Threat intelligence related to attacks on digital power grids is crawled through web crawlers, and non-research information is filtered out when obtaining the raw data.
3. The method for quantifying digital power grid network attack data based on feature engineering according to claim 1, characterized in that: The method for preprocessing data information includes: directly storing fields such as attack behavior IDs into a CSV file; for fields containing descriptive text data, performing unified field processing, first removing line breaks between paragraphs and converting them into single-line format; then performing sentence segmentation processing, dividing the single-line format text into single-sentence format text according to punctuation marks; and finally performing word segmentation processing before storing it into a CSV file.
4. The method for quantifying digital power grid network attack data based on feature engineering according to claim 1, characterized in that: The attack behavior dataset contains the attack tactics, techniques / sub-techniques, and corresponding textual descriptions of the attacking organization.
5. The method for quantifying digital power grid network attack data based on feature engineering according to claim 1, characterized in that: When using the TRAM tool to map threat intelligence for digital power grids, the attack behavior of attacking organizations is automatically extracted through TRAM, thereby mapping the threat intelligence into ATT&CK.
6. The method for quantifying digital power grid network attack data based on feature engineering according to claim 1, characterized in that: Methods for mapping threat intelligence to ATT&CK include: users submitting threat intelligence, submitting threat intelligence documents in txt or word format, or web addresses containing the threat intelligence to be analyzed.
7. A method for quantifying digital power grid network attack data based on feature engineering according to claim 6, characterized in that: The method of mapping threat intelligence to ATT&CK also includes: using the TRAM tool to retrieve submitted threat intelligence content, using the underlying logistic regression model to analyze and predict attack behavior, comparing it with information from the threat intelligence sharing platform, displaying the predicted attack behavior and its probability on the page, and automatically filtering the prediction results if the probability is lower than the user-defined threshold; finally, the user manually adjusts the feedback based on the prediction results and generates the corresponding threat intelligence in JSON format.
8. The method for quantifying digital power grid network attack data based on feature engineering according to claim 1, characterized in that: When quantifying data based on feature engineering, the term frequency-inverse document frequency method is used to calculate the contribution of each word; assuming the following situation, N texts contain M words, and each word is represented as t. i The text is represented as d j Therefore, a word set T = t1, t2, ..., t is formed. i ,...t m ] and text set D = [d1, d2, ..., d j ,...d n ], Term Frequency (TF) represents word frequency (t). i In text d j The frequency of occurrence is calculated according to Formula 1. The larger the TF value, the more important the word is in the text. Where, n i,j The word t i In text d j Frequency of occurrence, ∑ k n k,j Represents text d j The number of words contained in the text is shown in Equation 2. Inverse text frequency (IDF) refers to the inverse relationship between the number of times a word appears and the importance of the word in the text. The larger the IDF value, the more important the word is in the text, and the more representative it is of the text features. Where N represents the total number of texts. Indicates that it contains the word t i The number of texts, where TF-IDF represents the importance of a word relative to the text set, is calculated according to Equation 3. Where w(t,d) represents the weight value of word t calculated using TE IDF.