Multi-source unstructured threat intelligence identification and credibility evaluation method and system
By combining network crawling technology and grid-based Transformer model, the threat intelligence characteristic data related to the power system is extracted and evaluated, and the problem of identifying and evaluating the power grid when facing new types of cyber threats is solved, and efficient threat intelligence screening and security protection decision support is achieved.
Patent Information
- Application Number
- CN202411899483.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-02
AI Technical Summary
When the power grid faces new cyber threats, traditional defense technologies are difficult to identify new and unknown attack methods, and lack effective sources of threat intelligence information, resulting in a lack of forward-looking and targeted security protection.
By designing network crawler technology with the best limited search algorithm and depth-first algorithm, multi-source unstructured threat intelligence data is obtained; target threat intelligence feature data related to power systems is extracted using grid-based Transformer model; decision tree model is trained based on random forest algorithm to evaluate the confidence of target feature data.
Accurate identification and credibility assessment of multi-source unstructured threat intelligence are achieved, the accuracy of feature extraction and reliability of credibility assessment are improved, and the response capabilities of power grid security protection are enhanced.
Smart Images

Figure CN119917665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of threat assessment, and in particular to a method and system for identifying and evaluating the credibility of multi-source unstructured threat intelligence. Background Art
[0002] As the power grid actively participates in the digital transformation process, information technology is widely used in many business fields, and the level of informatization has been significantly improved. However, this has also led to a sharp increase in the complexity of its system architecture, with more devices, systems and interfaces exposed to the network environment, which has greatly increased the potential attack entry points. For example, key infrastructure such as smart meters, distributed energy management systems, and power dispatching automation systems are all connected to the network. Once attacked, it may cause serious consequences such as power grid failures and power outages, affecting the stability and reliability of power supply.
[0003] The power grid currently uses a variety of common network security defense technologies, such as network firewalls, intrusion detection systems, application firewalls, and log auditing. These technologies can resist known attack patterns and vulnerability exploits to a certain extent, and play a basic security protection role. However, in the face of new network threats, such as 0day vulnerability attacks and advanced threat attacks, traditional defense technologies mainly rely on known attack features and patterns for detection, and it is difficult to effectively identify new and unknown attack methods. In addition, these technical means are insufficient in the acquisition and utilization of threat intelligence information, and lack effective sources of threat intelligence information. The lack of threat intelligence makes security protection measures lack foresight and pertinence, and it is impossible to respond to new and complex network attacks in a timely and effective manner, resulting in the lack of overall security protection capabilities of the power grid.
[0004] In view of this, a method and system for identifying and evaluating the credibility of multi-source unstructured threat intelligence is needed. Summary of the invention
[0005] The embodiments of the present application provide a method and system for identifying and evaluating the credibility of multi-source unstructured threat intelligence, which are used to solve the problem of inaccurate identification of threat intelligence information.
[0006] A first aspect of an embodiment of the present application provides a method for identifying and evaluating the credibility of multi-source unstructured threat intelligence, including:
[0007] Design web crawler technology by combining the best finite search algorithm with the depth-first algorithm;
[0008] Use the designed web crawler technology to obtain multi-source unstructured data for threat intelligence information;
[0009] Extracting target threat intelligence feature data related to the power system from the multi-source unstructured data through a gridded Transformer model;
[0010] The decision tree model obtained by training based on the random forest algorithm performs credibility assessment on the target threat intelligence feature data to obtain corresponding credibility assessment results.
[0011] Furthermore, the multi-source unstructured data of threat intelligence information obtained by utilizing the designed web crawler technology includes:
[0012] Construct a priority queue and a set of visited URLs, and select the initial URL related to the power system as a seed;
[0013] Add the selected seed URL to the constructed priority queue according to the set initial priority, and determine the URL with the highest priority;
[0014] Use the network request library to send a request to obtain the web page content corresponding to the URL with the highest priority, and use the parsing library to parse the web page content to extract unstructured data.
[0015] Furthermore, the extracting target threat intelligence feature data related to the power system from the multi-source unstructured data through the gridded Transformer model includes:
[0016] Using a natural language processing tool to pre-process the text data in the multi-source unstructured data to generate a corresponding word sequence;
[0017] Determine the dimension and size of the text grid according to the length and structural characteristics of the text, fill the generated word sequence into the text grid, and obtain the gridded text data;
[0018] The preprocessed and gridded text data is input into the trained gridded Transformer model to output target threat intelligence feature data related to the power system.
[0019] Furthermore, the dimension and size of the text grid are determined according to the length and structural characteristics of the text, and the generated word sequence is filled into the text grid to obtain the gridded text data, including:
[0020]
[0021] Where: pos i,j is the coordinate position of the word in the grid, i is the row of the coordinate position, j is the column of the coordinate position, k is the dimension index of the position encoding vector, and d modelis the dimension size of the embedding representation such as word vector in the model. If k is even, it means k is an even number, and if k is odd, it means k is an odd number.
[0022] Furthermore, the preprocessed and gridded text data is input into the trained gridded Transformer model to output target threat intelligence feature data related to the power system, including:
[0023] Build a Transformer architecture with multiple encoder and decoder layers and set hyperparameters;
[0024] Map the words in the grid to the corresponding word vector representation, and organize the labeled power system threat intelligence data into a label form corresponding to the input text;
[0025] Select the cross entropy loss function and the configured optimizer to update the model parameters for gridded Transformer model training to obtain a trained gridded Transformer model.
[0026] Furthermore, the cross entropy loss function and the configured optimizer are selected to update the model parameters for gridded Transformer model training to obtain a trained gridded Transformer model, including:
[0027]
[0028] Where: L CE is the cross loss value, C is the number of categories, y true,c is the probability of category c in the power system threat intelligence data sample, y pred,c is the probability that the model predicts that it is class c.
[0029] Furthermore, the decision tree model trained based on the random forest algorithm performs credibility assessment on the target threat intelligence feature data to obtain corresponding credibility assessment results, including:
[0030] Determine the credibility label based on the threat intelligence feature data extracted by the gridded Transformer model, and divide the data set at the same time;
[0031] A random forest model is formed based on the constructed decision trees;
[0032] The test set is used to evaluate the performance of the random forest model, and the model is optimized based on the evaluation results to obtain the optimized random forest model.
[0033] Furthermore, the random forest model formed by the constructed decision tree includes:
[0034] The information gain calculation formula of the decision tree is:
[0035]
[0036] Where: D is the data set of the current node, that is, it contains the feature vectors of multiple samples and the corresponding credibility labels, A is the feature used for splitting, V is the number of values of feature A, D v is a subset of feature A whose value is v, that is, the set of samples in the data set whose feature A value is equal to v. H(D) is the entropy of data set D, K is the number of categories, and p k is the proportion of category k in the dataset D.
[0037] Furthermore, the use of the test set to evaluate the performance of the random forest model and optimizing the model according to the evaluation results to obtain an optimized random forest model includes:
[0038]
[0039] Among them: Accuracy is the accuracy, Precision is the precision, Recall is the recall, TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative.
[0040] A second aspect of an embodiment of the present application provides a system for identifying and evaluating the credibility of multi-source unstructured threat intelligence, including:
[0041] Web crawler technology design unit, used to design web crawler technology by combining the best limited search algorithm with the depth-first algorithm;
[0042] A multi-source unstructured data acquisition unit, used to acquire multi-source unstructured data of threat intelligence information using a designed web crawler technology;
[0043] A target threat intelligence feature data extraction unit, used to extract target threat intelligence feature data related to the power system from the multi-source unstructured data through a gridded Transformer model;
[0044] The credibility assessment result determination unit is used to perform credibility assessment on the target threat intelligence feature data based on the decision tree model trained by the random forest algorithm to obtain a corresponding credibility assessment result.
[0045] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0046] The present invention designs a network crawler technology by organically combining the best-first search algorithm with the depth-first algorithm, thereby improving the search efficiency and coverage of network information, and can accurately and comprehensively obtain multi-source unstructured data of threat intelligence information; the gridded Transformer model is used to extract features from these multi-source data, which can deeply mine the semantic and structural information in the text, accurately extract target threat intelligence feature data related to the power system, and improve the precision and accuracy of feature extraction; the decision tree model obtained by training based on the random forest algorithm performs credibility assessment on the target feature data, taking into account the influence of multiple factors on credibility, making the credibility assessment result more reliable and accurate, and being able to efficiently screen out highly credible threat intelligence, providing a solid and powerful decision-making basis for the security protection of the power system, and effectively enhancing the ability and level of the entire security protection system to cope with potential threats.
[0047] Other advantages, objectives, and features of the present invention will be set forth in part in the following description, and in part will be apparent to those skilled in the art based on an examination of the following or may be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 The present invention is a flowchart of an embodiment of a method for identifying and evaluating the credibility of multi-source unstructured threat intelligence. DETAILED DESCRIPTION
[0049] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0050] Embodiment 1
[0051] The implementation method in this embodiment can be implemented in the system, in the server, or in the terminal, and the specific implementation is not clearly limited. The following will introduce the identification and credibility assessment method of multi-source unstructured threat intelligence in this application from the perspective of system implementation. Figure 1 , the method provided in the embodiment of the present application comprises the following steps:
[0052] S11. Design web crawler technology by combining the best-first search algorithm with the depth-first algorithm;
[0053] The purpose of the best-first search algorithm is to first crawl those web pages that are most likely to contain valuable information according to certain priority rules. In web crawlers, by defining reasonable priority indicators, high-quality information sources can be quickly located. The depth-first algorithm is to deeply explore the internal link structure of web pages and obtain information hidden in deep pages. Some websites may place important details, related references, etc. in a deeper link hierarchy. The depth-first algorithm can ensure that the crawler will not only stay on the surface information, but can explore in depth and fully obtain all information related to the target topic. The website structure and power information distribution on the Internet are very complex and diverse. Some websites may have a clear hierarchical structure, while the information of some websites may be scattered in various cross-linked pages. Therefore, combining the best-first search algorithm with the depth-first algorithm to design web crawler technology can better adapt to this complexity. The best-first search algorithm can guide the crawler to move towards high-value information areas on a macro level, while the depth-first algorithm can explore deeply in local areas and flexibly respond to the structural characteristics of different websites, thereby effectively traversing various types of websites and improving the comprehensiveness of information acquisition.
[0054] S12. Use the designed web crawler technology to obtain multi-source unstructured data of threat intelligence information;
[0055] In this embodiment, step S12 also includes the following:
[0056] 1. Construct a priority queue and a set of visited URLs, and select the initial URL related to the power system as the seed;
[0057] Create a priority queue to store the URLs to be crawled and their priority information. Each element in the priority queue can be a tuple, containing the URL and its corresponding priority value. The smaller the value, the higher the priority. At the same time, establish a visited URL set to record the web page addresses that have been crawled to avoid repeated visits.
[0058] 2. Add the selected seed URL to the constructed priority queue according to the set initial priority, and determine the URL with the highest priority;
[0059] According to the goal of the crawler, here to capture threat intelligence information related to the power system, a set of initial seed URLs are selected. These seed URLs should be pages that are highly relevant to the target field and have high credibility, such as power safety report pages issued by authoritative organizations, official security announcement pages of well-known power companies, etc. The priority of each seed URL can be determined based on a variety of factors, such as the weight of the website where the URL is located, which can be measured by indicators such as the website's popularity, authority, and traffic ranking, the relevance of the page to the target topic, which can be evaluated through methods such as keyword matching and semantic analysis, and the timeliness of the information, such as recently released pages may have a higher priority. Add the seed URL and its priority information to the priority queue.
[0060] 3. Use the network request library to send a request to obtain the web page content corresponding to the URL with the highest priority, and use the parsing library to parse the web page content to extract unstructured data.
[0061] Take the URL with the highest priority from the priority queue, and use the network request library to send an HTTP request to obtain the webpage content corresponding to the URL. If a network anomaly is encountered during the request, an appropriate retry mechanism can be set, such as retrying 3-5 times and waiting for a certain time interval between each retry. Use the HTML parsing library to parse the webpage content and extract the text information, hyperlinks, and other possible metadata in the page, that is, the power unstructured data.
[0062] S13. Extract target threat intelligence feature data related to the power system from multi-source unstructured data through a gridded Transformer model;
[0063] In this embodiment, step S13 includes the following:
[0064] 1. Use natural language processing tools to preprocess text data in multi-source unstructured data to generate corresponding word sequences;
[0065] In the above steps, extract text data from multi-source unstructured data and preprocess the text data, where the preprocessing is word segmentation, that is, use appropriate natural language processing tools to segment the text and split the text into individual words or sequences of words. For example, for the Chinese sentence "There is a risk of cyber attacks on the power system", after word segmentation, the words "power system", "existence", "cyber attack", "risk" and so on are obtained. At the same time, tokenization operations are performed according to needs, such as adding corresponding part-of-speech tags to each word, which will help to better understand the semantic structure of the text later. However, part-of-speech tagging is not a necessary step and depends on the specific feature extraction requirements.
[0066] 2. Determine the dimension and size of the text grid according to the length and structural characteristics of the text, fill the generated word sequence into the text grid, and obtain the gridded text data;
[0067] The dimension and size of the grid are determined based on the length and structural characteristics of the text combined with experience. Here, a two-dimensional grid is constructed, for example, the number of sentences is used as the row dimension, and the average number of words in the sentence is used as the column dimension. Assuming that the text contains m sentences, and each sentence has n words after word segmentation on average, the size of the constructed two-dimensional grid can be expressed as M×N, where M represents the number of rows and N represents the number of columns. It can be adjusted according to actual conditions, such as M=m, and N is a suitable integer slightly larger than n to ensure that the words of each sentence can be accommodated.
[0068] The segmented words are filled into the grid in order. If the number of words in a sentence is less than N, they are filled with special filling symbols. For the position of each word in the grid, position encoding is required. The position encoding method here is similar to the absolute position encoding in Transformer, and the calculation formula is as follows:
[0069]
[0070] Where: pos i,j is the coordinate position of the word in the grid, i is the row of the coordinate position, j is the column of the coordinate position, k is the dimension index of the position encoding vector, and d model is the dimension size of the embedding representation such as word vector in the model. If k is even, it means k is an even number, and if k is odd, it means k is an odd number.
[0071] Through this position encoding, the model can be provided with the position information of words in the text structure, helping it capture the order and structural characteristics of the text.
[0072] 3. Input the preprocessed and gridded text data into the trained gridded Transformer model to output target threat intelligence feature data related to the power system.
[0073] Specifically, the step also includes the following:
[0074] Build a Transformer architecture with multiple encoder and decoder layers and set hyperparameters;
[0075] Map the words in the grid to the corresponding word vector representation, and organize the labeled power system threat intelligence data into a label form corresponding to the input text;
[0076] Select the cross entropy loss function and the configured optimizer to update the model parameters for gridded Transformer model training to obtain a trained gridded Transformer model.
[0077] Construct a Transformer architecture with multiple encoder and decoder layers. Each encoder layer is mainly composed of multi-head attention mechanism, feedforward neural network, residual connection and layer normalization modules. Hyperparameters include embedding dimension, number of heads of multi-head attention mechanism, dimension of intermediate layer of feedforward neural network, number of encoder layers, etc. The values of these parameters will affect the complexity and performance of the model and need to be adjusted and optimized through experiments.
[0078] Map the words in the grid to the corresponding word vector representation. Here, we use the pre-trained word vector model or randomly initialize the word vector during the training process and learn and update it through model training. Let the word w ij The corresponding word vector is v ij , whose dimension is d model , add the word vector to the corresponding position encoding vector to get the input representation x of the word in the model ij ,Right now After word vector embedding and position encoding, the entire text has a shape of M×N×d model The tensor is used as the input of the model.
[0079] If there is annotated power system threat intelligence data, organize the annotated information into labels corresponding to the input text for supervised training of the model. Suppose the training data set is {(X1,y1),(X2,y2),...,(X k ,y k )}, where X i represents the i-th preprocessed and gridded text input tensor, y i Indicates the corresponding feature annotation label. Select the cross entropy loss function according to the task type:
[0080]
[0081] Where: L CE is the cross loss value, C is the number of categories, y true,c is the probability of category c in the power system threat intelligence data sample, y pred,c is the probability that the model predicts that it is class c.
[0082] Configure a suitable optimizer to update the model parameters, set parameters such as the learning rate, and continuously adjust the model parameters through iterative training to gradually reduce the loss function value and improve the model's ability to extract threat intelligence features.
[0083] The preprocessed and gridded text data is input into the trained gridded Transformer model. Through the multi-layer multi-head attention mechanism and feedforward neural network operations, the model encodes and learns the semantic and structural information in the text.
[0084] At the output of the last encoder layer of the model, the feature representation vector corresponding to each word is obtained. For example, after passing through the L-layer encoder, the word w ij The characteristic is expressed as It integrates the semantic relationship, context information, and structural position information of the word in the entire text. The feature representation tensor corresponding to the entire text is: Its dimensions are M×N×d model .
[0085] According to the specific threat intelligence feature extraction requirements, the feature representation of the model output is further selected and integrated. For example, if you focus on the overall threat level description features in the text, you can average pool the feature representation vectors of all words to obtain a fixed-dimensional feature vector to represent the threat level features of the entire text. The average pooling calculation formula is as follows:
[0086]
[0087] Where: f avg is the overall feature vector of the text obtained by average pooling.
[0088] The above can also be done according to the specific feature type by designing a specific classifier or feature screening mechanism to extract the corresponding sub-feature vector from the feature representation output by the model for subsequent more detailed threat intelligence analysis.
[0089] S14. Perform credibility assessment on the target threat intelligence feature data based on the decision tree model trained by the random forest algorithm to obtain the corresponding credibility assessment results.
[0090] In this embodiment, step S14 includes the following:
[0091] 1. Determine the credibility label based on the threat intelligence feature data extracted by the gridded Transformer model and divide the data set at the same time;
[0092] Assume that n features are extracted from the threat intelligence text through the gridded Transformer model, forming a feature vector X = [F1, F2, ..., F n ], the meaning of each feature is as follows:
[0093] F1 is the credibility of the intelligence source, and its value range is [0,1]. For example, the value of intelligence released by an authoritative organization is close to 1, and the value of intelligence from anonymous and unreliable channels is close to 0; F2 is the integrity of key information in intelligence, which is measured by counting the proportion of key information to the total amount of key information that should be included, and its value range is [0,1]; F3 is the matching degree between intelligence and existing authoritative knowledge base, which calculates the text similarity between intelligence content and similar content in known authoritative power system safety knowledge base, and its value range is [0,1]; ...; F n For other related features defined according to business needs, there are also corresponding reasonable value ranges. Preprocess these feature vectors, and normalize them here. Determine the credibility label for each threat intelligence sample, such as manually annotating high credibility with 1, medium credibility with 0.5, low credibility with 0, etc., to form a data set {(X1, y1), (X2, y2), ..., (X m ,y m )}, where m is the number of samples.
[0094] 2. Form a random forest model based on the constructed decision tree;
[0095] Sub-training sets are obtained by random sampling with replacement from the training set to build each decision tree. At each node, a portion of features are randomly selected from all features, and then the best split point is calculated based on these features so that the purity of the child node after the split is the highest. Here, the purity change is measured by information gain, and the information gain calculation formula is as follows:
[0096]
[0097] Where: D is the data set of the current node, that is, it contains the feature vectors of multiple samples and the corresponding credibility labels, A is the feature used for splitting, V is the number of values of feature A, D v is a subset of feature A whose value is v, that is, the set of samples in the data set whose feature A value is equal to v. H(D) is the entropy of data set D, K is the number of categories, and p k is the proportion of category k in the dataset D.
[0098] Select the features and splitting points that maximize the information gain for node splitting, and recursively build a decision tree until the stopping condition is met. Repeat the above decision tree construction process T times, for example, T = 100, to obtain a random forest model consisting of T decision trees.
[0099] 3. Use the test set to evaluate the performance of the random forest model, and optimize the model based on the evaluation results to obtain the optimized random forest model.
[0100] The performance of the random forest model is evaluated using the test set. The evaluation includes:
[0101]
[0102] Where: Accuracy is the accuracy, Precision is the precision, Recall is the recall, TP is the true positive example, that is, the number of samples that are actually high-confidence and predicted by the model as high-confidence, TN is the true negative example, that is, the number of samples that are actually low-confidence and predicted by the model as low-confidence, FP is the false positive example, that is, the number of samples that are actually low-confidence but predicted by the model as high-confidence, and FN is the false negative example, that is, the number of samples that are actually high-confidence but predicted by the model as low-confidence.
[0103] The F1 value comprehensively considers the harmonic mean of precision and recall, and the calculation formula is:
[0104]
[0105] By calculating these indicators on the test set, the performance of the model can be evaluated to understand the prediction accuracy and recall rate of the model in different categories. According to the evaluation results, if the model performance is not ideal, the model can be optimized. For example, the number of decision trees T, the depth of the decision tree, the minimum number of samples required for node splitting and other parameters can be adjusted.
[0106] After model training and optimization are completed, for new threat intelligence feature data X new , and input it into the random forest model. Each decision tree predicts it and obtains its own credibility prediction results, such as the corresponding values of high, medium, and low credibility, and then obtains the final credibility evaluation result through voting. For example, when voting, the number of decision trees predicted as high credibility, medium credibility, and low credibility is counted. Whichever category receives the most votes is assigned to the new threat intelligence sample as the final credibility evaluation result, so as to sort and screen the threat intelligence according to credibility, give priority to high-credibility intelligence, and provide decision support for the security protection of the power system.
[0107] The above embodiment uses the gridded Transformer model to extract threat intelligence feature data related to the power system from multi-source unstructured data, providing a basis for further credibility assessment, intelligence analysis and other tasks. Based on the random forest algorithm, the threat intelligence feature data is used to effectively assess its credibility, helping to screen out more reliable and valuable threat intelligence information.
[0108] Embodiment 2
[0109] The embodiment of the system for identifying and evaluating the credibility of multi-source unstructured threat intelligence in the present invention comprises the following steps:
[0110] Web crawler technology design unit, used to design web crawler technology by combining the best limited search algorithm with the depth-first algorithm;
[0111] A multi-source unstructured data acquisition unit, used to acquire multi-source unstructured data of threat intelligence information using a designed web crawler technology;
[0112] A target threat intelligence feature data extraction unit is used to extract target threat intelligence feature data related to the power system from multi-source unstructured data through a gridded Transformer model;
[0113] The credibility assessment result determination unit is used to perform credibility assessment on the target threat intelligence feature data based on the decision tree model trained by the random forest algorithm to obtain the corresponding credibility assessment result.
[0114] For the specific definition of the system, please refer to the definition of the method above, which will not be repeated here. Each module in the above system can be implemented in whole or in part by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0115] Those of ordinary skill in the art will appreciate that the units of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition of each example has been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0116] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored. In addition, each functional unit in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units.
[0117] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-0nlyMemory), random access memory (RAM, RandomAccessMemory), mobile hard disk, magnetic disk or optical disk, etc., which can store program code.
[0118] It can be understood that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
Claims
1. A method for identifying and evaluating the credibility of multi-source unstructured threat intelligence, characterized in that: include: Design web crawler technology by combining the best finite search algorithm with the depth-first algorithm; Use the designed web crawler technology to obtain multi-source unstructured data for threat intelligence information; Extracting target threat intelligence feature data related to the power system from the multi-source unstructured data through a gridded Transformer model; The decision tree model obtained by training based on the random forest algorithm performs credibility assessment on the target threat intelligence feature data to obtain corresponding credibility assessment results.
2. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 1, characterized in that: The multi-source unstructured data of threat intelligence information obtained by utilizing the designed web crawler technology includes: Construct a priority queue and a set of visited URLs, and select the initial URL related to the power system as a seed; Add the selected seed URL to the constructed priority queue according to the set initial priority, and determine the URL with the highest priority; Use the network request library to send a request to obtain the web page content corresponding to the URL with the highest priority, and use the parsing library to parse the web page content to extract unstructured data.
3. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 1, characterized in that: The extracting target threat intelligence feature data related to the power system from the multi-source unstructured data by using the gridded Transformer model includes: Preprocessing the text data in the multi-source unstructured data using a natural language processing tool to generate a corresponding word sequence; Determine the dimension and size of the text grid according to the length and structural characteristics of the text, fill the generated word sequence into the text grid, and obtain the gridded text data; The preprocessed and gridded text data is input into the trained gridded Transformer model to output target threat intelligence feature data related to the power system.
4. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 3, characterized in that: The dimension and size of the text grid are determined according to the length and structural characteristics of the text, and the generated word sequence is filled into the text grid to obtain the gridded text data, including: Where: pos i,j is the coordinate position of the word in the grid, i is the row of the coordinate position, j is the column of the coordinate position, k is the dimension index of the position encoding vector, and d model is the dimension size of the embedding representation such as word vector in the model. If k is even, it means k is an even number, and if k is odd, it means k is an odd number.
5. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 4, characterized in that: The preprocessed and gridded text data is input into the trained gridded Transformer model, and target threat intelligence feature data related to the power system is output, including: Build a Transformer architecture with multiple encoder and decoder layers and set hyperparameters; Map the words in the grid to the corresponding word vector representation, and organize the labeled power system threat intelligence data into a label form corresponding to the input text; Select the cross entropy loss function and the configured optimizer to update the model parameters for gridded Transformer model training to obtain a trained gridded Transformer model.
6. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 5, characterized in that: The cross entropy loss function and the configured optimizer are selected to update the model parameters for gridded Transformer model training to obtain a trained gridded Transformer model, including: Where: L CE is the cross loss value, C is the number of categories, y true,c is the probability of category c in the power system threat intelligence data sample, y pred,c is the probability that the model predicts that it is class c.
7. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 1, characterized in that: The decision tree model obtained by training based on the random forest algorithm performs credibility assessment on the target threat intelligence feature data to obtain corresponding credibility assessment results, including: Determine the credibility label based on the threat intelligence feature data extracted by the gridded Transformer model, and divide the data set at the same time; A random forest model is formed based on the constructed decision trees; The test set is used to evaluate the performance of the random forest model, and the model is optimized based on the evaluation results to obtain the optimized random forest model.
8. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 7, characterized in that: The random forest model formed by the constructed decision tree includes: The information gain calculation formula of the decision tree is: Where: D is the data set of the current node, that is, it contains the feature vectors of multiple samples and the corresponding credibility labels, A is the feature used for splitting, V is the number of values of feature A, D v is a subset of feature A whose value is v, that is, the set of samples in the data set whose feature A value is equal to v. H(D) is the entropy of data set D, K is the number of categories, and p k is the proportion of category k in the dataset D.
9. The method for identifying and evaluating the credibility of multi-source unstructured threat intelligence according to claim 1, characterized in that: The test set is used to evaluate the performance of the random forest model, and the model is optimized according to the evaluation results to obtain an optimized random forest model, including: Among them: Accuracy is the accuracy, Precision is the precision, Recall is the recall, TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative.
10. A system for identifying and evaluating the credibility of multi-source unstructured threat intelligence, characterized in that: include: Web crawler technology design unit, used to design web crawler technology by combining the best limited search algorithm with the depth-first algorithm; A multi-source unstructured data acquisition unit, used to acquire multi-source unstructured data of threat intelligence information using a designed web crawler technology; A target threat intelligence feature data extraction unit, used to extract target threat intelligence feature data related to the power system from the multi-source unstructured data through a gridded Transformer model; The credibility assessment result determination unit is used to perform credibility assessment on the target threat intelligence feature data based on the decision tree model trained by the random forest algorithm to obtain a corresponding credibility assessment result.