Power system threat intelligence collection method and system based on topic correlation
By constructing a topic relevance evaluation model based on stacked BiGRU and combining semantic and keyword information, the accuracy and timeliness of power system threat intelligence collection are improved, the problem of insufficient accuracy in topic discrimination in existing technologies is solved, and efficient automated threat intelligence collection is achieved.
Patent Information
- Application Number
- CN202510749201.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
Existing topic crawler methods for threat intelligence lack accuracy in topic identification in power system scenarios and lack comprehensive consideration of semantics, word frequency and keyword information.
A topic relevance evaluation model based on stacked BiGRU is adopted, and semantics, word frequency and keywords are combined to extract web page text feature vectors. SBERT and TF-IDF algorithms are used to generate feature vectors, and the model parameters are optimized through the cross-entropy loss function. The model is integrated into the topic crawler for web page classification.
It improves the accuracy and timeliness of power system threat intelligence collection, reduces information overload, and realizes efficient and automated collection of power system-related information.
Smart Images

Figure CN120633668A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of power system network security, and specifically to a method and system for collecting power system threat intelligence based on topic relevance. Background Art
[0002] Threat intelligence refers to information related to network threats obtained through systematic collection and analysis. It is widely used in cyberspace security analysis and APT (Advanced Persistent Threat) defense.
[0003] Existing threat intelligence collection methods can be categorized as manual and automated. Manual collection relies on expert knowledge and human resources, making it difficult to collect and update threat intelligence from the public internet in a timely and comprehensive manner. Automated collection methods are gradually being replaced by automated methods. Automated collection uses crawler software built with automated tools or scripts to crawl, organize, and analyze relevant information from public resources scattered across the internet, and then save it as threat intelligence. Compared to manual collection, automated methods can collect threat intelligence more comprehensively and update it regularly or in real time, ensuring the timeliness and accuracy of threat intelligence.
[0004] Considering the inherent requirements of threat intelligence for web page topics, using a focused crawler to selectively crawl relevant web pages can improve the accuracy of threat intelligence collection, avoid interference from irrelevant information, and reduce the risk of information overload. Existing focused crawler methods for threat intelligence have gradually evolved into technologies that use deep learning models or word vector model fusion, which has great reference value for power system threat intelligence collection. However, there are currently no focused crawler methods for power system threat intelligence, and the effectiveness of existing methods in power system scenarios is difficult to guarantee. In addition, existing methods mainly use a single word vector model for topic identification, without integrating the semantic information and word frequency information of the text, and the effectiveness of topic relevance identification still needs to be improved. Summary of the Invention
[0005] The present application provides a method and system for collecting power system threat intelligence based on topic relevance, which can solve the technical problems existing in existing threat intelligence-oriented topic crawler technologies, such as insufficient accuracy in topic discrimination in power system scenarios and lack of comprehensive consideration of semantics, word frequency and keyword information.
[0006] In a first aspect, the present application provides a method for collecting power system threat intelligence based on topic relevance, comprising the following steps: Extract feature vectors of web page text from three perspectives: semantics, word frequency, and keywords; Construct a topic relevance evaluation model based on stacked BiGRU (Bi-directional Gated Recurrent Unit), set the feature vector of web page text as the model input, and use the constructed topic relevance evaluation model as the main body of the evaluation model; The feature vector of the extracted web page text and the topic relevance evaluation method are integrated into a topic crawler, and the topic crawler is used to crawl the classification results of the web page.
[0007] Furthermore, the feature vector of the webpage text is extracted from the three perspectives of semantics, word frequency and keywords, including the following steps: Encode webpage text based on the pre-trained SBERT model and generate SBERT vectors that reflect the semantic features of text topic evaluation; Based on the TF-IDF algorithm, the text is first preprocessed, a bag-of-words model is built, and the term frequency-inverse document frequency is calculated to generate a TF-IDF vector that reflects the importance of the keyword; Divide the SBERT vector and TF-IDF vector by label type and calculate the mean of each vector to obtain the mean vector divided by label type. Calculate the similarity between each label type vector and the corresponding mean vector to obtain the SBERT cosine similarity vector and the TF-IDF cosine similarity vector. Sort by TF-IDF value from high to low, select a preset number of top-ranked keywords, convert them into SBERT word vectors, and calculate the maximum cosine similarity with the preset label words to generate a keyword similarity vector.
[0008] Furthermore, the topic relevance evaluation method of constructing a stacked BiGRU-based topic relevance evaluation model, setting the feature vector of the webpage text as the model input, and using the constructed topic relevance evaluation model as the evaluation model body specifically includes the following steps: constructing a stacked BiGRU-based topic relevance evaluation model, the topic relevance evaluation model including a feature extraction layer, a fully connected layer, and a Softmax output layer, the feature extraction layer being two independent BiGRU models; Input the SBERT vector and TF-IDF vector into the BiGRU model respectively, use two independent BiGRU models to extract features from the SBERT vector and TF-IDF vector respectively, and output the extracted features of the SBERT vector and the TF-IDF vector; Combine the output SBERT vector extraction features, TF-IDF vector extraction features with the SBERT cosine similarity vector, TF-IDF cosine similarity vector and keyword similarity vector to obtain a combined feature vector, which is input into the fully connected layer for training, minimizing the loss value of the cross entropy loss function to obtain the final classification result of the web page.
[0009] Furthermore, the cross entropy loss function is shown as follows:
[0010] Where, is the sample size, is the number of categories, Represents a sample In category The true labels on , encoded in one-hot form, Represents a sample In category The model prediction probability is given by the Softmax function.
[0011] Furthermore, the classification result is calculated as follows:
[0012] Where, is the weight matrix, is the bias term, Predict probabilities for the categories.
[0013] Furthermore, the combined feature vector satisfies the following relationship:
[0014] Where, is the input of the fully connected layer, that is, the combined feature vector, They are the output features of SBERT vector and TF-IDF vector after BiGRU processing, is the cosine similarity vector, is the keyword similarity vector.
[0015] Furthermore, the framework of the subject crawler includes: Crawling module, used to download HTML documents based on URL seeds; Storage module, used to cache URL seeds and crawling results; Extraction module, used to parse HTML documents and extract body text; A classification module, for integrating the topic relevance evaluation model to perform real-time classification; Tracking module, used to recursively crawl new URLs and update the seed library.
[0016] Furthermore, when the tracking module performs recursive crawling, a maximum recursion depth threshold is set, and only URLs that do not exceed the threshold are prioritized, with the ranking based on the topic relevance score of the parent page.
[0017] In a second aspect, the present application provides a power system threat intelligence collection system based on topic relevance, including: Web page text feature vector acquisition module, used to extract feature vectors of web page text from three perspectives: semantics, word frequency, and keywords; The topic relevance evaluation strategy module is used to build a topic relevance evaluation model based on stacked BiGRU, and set the feature vector of the web page text as the model input, and use the constructed topic relevance evaluation model as the evaluation model body of the topic relevance evaluation method; The web page classification result acquisition module is in communication with the web page text feature vector acquisition module and the topic relevance evaluation strategy module, integrates the extracted feature vector of the web page text and the topic relevance evaluation method into the topic crawler, and uses the topic crawler to crawl the classification results of the web page.
[0018] Furthermore, the webpage text feature vector acquisition module includes: The SBERT vector acquisition module is used to encode web page text based on the pre-trained SBERT model and generate SBERT vectors that reflect the semantic features of text topic evaluation; The TF-IDF vector acquisition module is used to preprocess the text based on the TF-IDF algorithm, build a bag-of-words model, and calculate the term frequency-inverse document frequency to generate a TF-IDF vector that reflects the importance of the keyword. A cosine similarity vector acquisition module is communicatively connected to the SBERT vector acquisition module and the TF-IDF vector acquisition module, and is used to divide the SBERT vector and the TF-IDF vector according to the label type, and perform mean calculation on each of them to obtain the mean vector divided according to the label type, calculate the similarity between each label type vector and the corresponding mean vector, and obtain the SBERT cosine similarity vector and the TF-IDF cosine similarity vector; The keyword similarity vector acquisition module is used to sort the keywords from high to low according to the TF-IDF value, select a preset number of top-ranked keywords, convert them into SBERT word vectors, and calculate the maximum cosine similarity with the preset label words to generate a keyword similarity vector.
[0019] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least: By combining semantics, word frequency, and keywords, a more representative text feature vector is constructed. Stacked BiGRU is used to avoid interference between different types of vectors, improving the model's feature extraction effect on text sequences. The topic relevance evaluation is integrated with the topic crawler to enable the collection of threat intelligence for the open Internet. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A flowchart of a method for collecting power system threat intelligence based on topic relevance provided in an embodiment of the present application; Figure 2 A schematic diagram of a specific process of a method for collecting power system threat intelligence based on topic relevance provided in an embodiment of the present application; Figure 3 A structural diagram of the BiGRU model based on stacked learning provided in an embodiment of the present application; Figure 4 This is a diagram of the theme crawler architecture provided for an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0022] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit the "first", "second" and "third" to different types.
[0023] In the description of the embodiments of the present application, the words "exemplary," "for example," or "for example" are used as examples, illustrations, or explanations. Any embodiment or design described in the embodiments of the present application as "exemplary," "for example," or "for example" should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0024] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0025] In some processes described in the embodiments of the present application, multiple operations or steps are included that appear in a specific order. However, it should be understood that these operations or steps may not be performed in the order in which they appear in the embodiments of the present application or may be performed in parallel. The sequence numbers of the operations are only used to distinguish between different operations, and the sequence numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be performed in sequence or in parallel, and these operations or steps may be combined.
[0026] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0027] First, please refer to Figure 1-Figure 2 , this application provides a method for collecting power system threat intelligence based on topic relevance, including the following steps: Step S1: Obtain the web page to be crawled for threat intelligence information, and extract the feature vector of the web page text from the perspectives of semantics, word frequency, and keywords; Step S2: Construct a topic relevance evaluation model based on stacked BiGRU, set the feature vector of the webpage text as the model input, and use the constructed topic relevance evaluation model as the evaluation model body of the topic relevance evaluation method; Step S3: Integrate the feature vector of the extracted web page text and the topic relevance evaluation method into the topic crawler, and use the topic crawler to crawl the classification results of the web page. The classification results include those related to the specific topic of power system threat intelligence. If relevant, they are collected as power system threat intelligence; if not relevant, they are not collected.
[0028] This application constructs a more representative text feature vector by combining semantics, word frequency and keywords, uses stacked BiGRU to avoid interference between different types of vectors, improves the model's feature extraction effect on text sequences, and integrates topic relevance evaluation with topic crawlers to achieve threat intelligence collection for the public Internet.
[0029] In one embodiment, step S1: extracting feature vectors of webpage text from the perspectives of semantics, word frequency, and keywords. The feature vectors include an SBERT vector reflecting semantic features for text topic evaluation, a TF-IDF vector reflecting keyword importance, an SBERT cosine similarity vector and a TF-IDF cosine similarity vector, and a keyword similarity vector. The specific method for obtaining each vector is as follows: Step S1A: Encode the webpage text based on the pre-trained SBERT model to generate an SBERT vector reflecting the semantic features of the text topic evaluation as one of the feature vector components of the text topic evaluation; Step S1B: Based on the TF-IDF algorithm, a TF-IDF vector generation process is designed, which mainly includes three parts: text preprocessing, word bag construction and vector calculation. Specifically, first, all documents in the text document list are preprocessed to remove stop words, punctuation and other information, and all words are restored to lowercase stem form; then a word bag model is constructed to add non-repeated words in all documents to the vocabulary; finally, the TF-IDF vector values of all terms are calculated to form a TF-IDF vector. ; Step S1C: Divide the SBERT vector and TF-IDF vector according to the label type, and perform mean calculation on each of them to obtain the mean vector divided by the label type, calculate the similarity between each label type vector and the corresponding mean vector, and obtain the SBERT cosine similarity vector and the TF-IDF cosine similarity vector; wherein, the label types include SBERT and TF-IDF; specifically, based on the document text and stop words, the SBERT vector and TF-IDF vector can be obtained, and the two types of vectors are divided according to the label type, and the mean is calculated for each of them, and the mean vector divided according to the label type is obtained as the basic vector for cosine similarity calculation. Calculate vector and the mean vector for each label type The similarity of the cosine similarity can be calculated according to formula (1), and all the results are combined to form a length of the number of label categories. SBERT cosine similarity vector And TF-IDF cosine similarity vector : Formula (1) Where, is a vector, is a vector The mean vector of For the samples.
[0030] Step S1D: Sort by TF-IDF value from high to low, select a preset number of top-ranked keywords, convert them into SBERT word vectors, and calculate the maximum cosine similarity with the preset label words to generate a keyword similarity vector; further comprising the following steps: Using TF-IDF value as the evaluation standard, sort the calculated TF-IDF values of all terms in reverse order and take the top ranked terms. Item, restore it to the original word text as the keyword of the current web page; Evaluate the similarity between all keywords and tags to obtain the keyword similarity vector; specifically, The keyword is converted into SBERT word vector, and the cosine similarity is calculated with the SBERT word vector of the related words. The result with the highest similarity score among all the calculation results is taken as the similarity result between the current keyword and the corresponding input word. Assume that the number of related words input is , for each document, we can get the length Keyword similarity vector .
[0031] Based on step S1, feature vectors of webpage text are extracted using three dimensions: semantics (SBERT vectors), word frequency (TF-IDF vectors), and keyword similarity. This method comprehensively characterizes the relevance of webpage content to power system threat intelligence topics. Compared to traditional single-feature methods (such as keyword matching alone), this method effectively avoids misjudgments caused by differences in text representation (e.g., different wording in technical reports and vulnerability descriptions), thereby improving the robustness of feature representation. Furthermore, TF-IDF statistics on high-frequency terms, combined with keyword filtering, can quickly identify core threat terms (e.g., "power grid attack" and "SCADA vulnerability"), providing highly discriminative input data for subsequent model classification.
[0032] In one embodiment, if Figure 3 As shown, step S2: constructing a topic relevance evaluation model based on stacked BiGRU, setting the feature vector of the webpage text as the model input, and using the constructed topic relevance evaluation model as the evaluation model body of the topic relevance evaluation method, specifically includes the following steps: Step S21: Construct a topic relevance evaluation model based on stacked BiGRU, such as Figure 3 As shown, the topic relevance evaluation model includes a feature extraction layer, a fully connected layer and a Softmax output layer, and the feature extraction layer is two independent BiGRU models; Step S22: Input the SBERT vector and TF-IDF vector into the BiGRU model respectively, use two independent BiGRU models to extract features from the SBERT vector and TF-IDF vector respectively, and output the extracted features of the SBERT vector and the TF-IDF vector, as shown in formula (2) and formula (3): Formula (2) Formula (3) Step S23: Combine the output SBERT vector extraction features, TF-IDF vector extraction features, SBERT cosine similarity vector, TF-IDF cosine similarity vector and keyword similarity vector to form the input vector of the fully connected layer ,in, They are the output features of SBERT vector and TF-IDF vector after BiGRU processing, is the cosine similarity vector, is the keyword similarity vector, input The model continuously adjusts the weight matrix during training. and bias , minimize the loss value of the cross entropy loss function, thereby obtaining the optimal classification decision boundary, and thus obtaining the final classification result of the web page; wherein, the cross entropy loss function is shown as follows:
[0033] Where, is the loss value of the cross entropy loss function, is the sample size, Represents a sample In category The true labels on , encoded in one-hot form, Represents a sample In category The model prediction probability is given by the Softmax function.
[0034] Based on step S2, a topic relevance assessment method based on a stacked BiGRU model uses a bidirectional gated recurrent unit (BiGRU) to capture contextual dependencies within text sequences (such as the long-range semantic associations in "APT attacks targeting power industrial control systems"). This model, combined with a fully connected layer, fuses multi-dimensional features (SBERT, TF-IDF, and keyword similarity). This model adaptively learns the semantic patterns of power system threat intelligence, significantly improving the classification accuracy of complex technical text (such as vulnerability analysis reports). Furthermore, the model optimizes parameters by minimizing the cross-entropy loss function, ensuring filtering of low-relevance web pages (such as financial security news) and reducing noise interference.
[0035] In one embodiment, the classification result is calculated as follows:
[0036] Where, is the weight matrix, is the bias term, is the class prediction probability, is the input of the fully connected layer.
[0037] In one embodiment, the framework of the subject crawler includes: Crawling module, used to download HTML documents based on URL seeds; Storage module, used to cache URL seeds and the final web page crawling results; Extraction module, used to parse HTML documents and extract body text to obtain web page text; A classification module, for integrating the topic relevance evaluation model to classify web page text in real time; A tracking module is used to recursively crawl new URLs and update the seed library. When performing recursive crawling, the tracking module sets a maximum recursion depth threshold and only prioritizes URLs that do not exceed the threshold, with the ranking based on the topic relevance score of the parent page.
[0038] In a specific embodiment, based on the above-mentioned theme crawler framework, a typical threat intelligence crawling process is as follows: URL seed acquisition: read the currently available URL seeds from the storage module and send them to the crawling module.
[0039] HTML document acquisition: accept the URL seed provided by the storage module, access the corresponding web server according to the preset crawling strategy, and download the HTML document of the web page.
[0040] HTML text information extraction: After obtaining the HTML document of a web page, the extraction module parses the HTML structure, extracts the web page body information, and removes irrelevant elements.
[0041] Topic relevance classification: The extracted web page text content is input into the classification module, and the topic evaluation model is used to determine the relevance.
[0042] Data storage and recursive crawling: The output results of the classification model are returned to the extraction module, and the relevance of the web page is determined based on the relevance threshold. If relevant, the web page URL and its text data are returned to the storage module for storage, and the hyperlinks in the web page are further obtained, and the new URLs that meet the URL recursive layer number are added to the URL seed library.
[0043] Based on step S3, the feature extraction and topic discrimination models are integrated into the topic crawler system, enabling automated and accurate threat intelligence collection. The crawler dynamically adjusts its crawling strategy based on the model's classification output (relevant / irrelevant), retaining only highly relevant web pages (e.g., those related to topics like power industrial control vulnerabilities and malware attacks), thus avoiding the data redundancy caused by the "broad-net" approach of traditional crawlers. Furthermore, through recursive crawling and URL prioritization (based on parent page topic scores), the system can efficiently track deeper threat intelligence chains (e.g., forum discussions and vulnerability library links), significantly improving the depth and timeliness of intelligence coverage.
[0044] In a second aspect, the present application provides a power system threat intelligence collection system based on topic relevance, including: Web page text feature vector acquisition module, used to extract feature vectors of web page text from three perspectives: semantics, word frequency, and keywords; The topic relevance evaluation strategy module is used to build a topic relevance evaluation model based on stacked BiGRU, and set the feature vector of the web page text as the model input, and use the constructed topic relevance evaluation model as the evaluation model body of the topic relevance evaluation method; The web page classification result acquisition module is in communication with the web page text feature vector acquisition module and the topic relevance evaluation strategy module, integrates the extracted feature vector of the web page text and the topic relevance evaluation method into the topic crawler, and uses the topic crawler to crawl the classification results of the web page.
[0045] In one embodiment, the webpage text feature vector acquisition module includes: The SBERT vector acquisition module is used to encode web page text based on the pre-trained SBERT model and generate SBERT vectors that reflect the semantic features of text topic evaluation; The TF-IDF vector acquisition module is used to preprocess the text based on the TF-IDF algorithm, build a bag-of-words model, and calculate the term frequency-inverse document frequency to generate a TF-IDF vector that reflects the importance of the keyword. A cosine similarity vector acquisition module is communicatively connected to the SBERT vector acquisition module and the TF-IDF vector acquisition module, and is used to divide the SBERT vector and the TF-IDF vector according to the label type, and perform mean calculation on each of them to obtain the mean vector divided according to the label type, calculate the similarity between each label type vector and the corresponding mean vector, and obtain the SBERT cosine similarity vector and the TF-IDF cosine similarity vector; The keyword similarity vector acquisition module is used to sort the keywords from high to low according to the TF-IDF value, select a preset number of top-ranked keywords, convert them into SBERT word vectors, and calculate the maximum cosine similarity with the preset label words to generate a keyword similarity vector.
[0046] Among them, the functional implementation of each module in the above-mentioned power system threat intelligence collection system based on topic relevance corresponds to the various steps in the above-mentioned power system threat intelligence collection method embodiment based on topic relevance, and its functions and implementation processes will not be repeated here one by one.
[0047] On the third aspect, an embodiment of the present application provides a power system threat intelligence collection device based on topic relevance. The power system threat intelligence collection device based on topic relevance can be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.
[0048] Communication interfaces include input / output (I / O), physical, and logical interfaces, which interconnect components within the topic-related power system threat intelligence collection device, as well as other devices (such as other computing devices or user devices). Physical interfaces can include Ethernet, fiber, and ATM interfaces; user devices can include displays and keyboards.
[0049] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0050] The processor may be a general-purpose processor that can invoke a topic-related power system threat intelligence collection program stored in a memory and execute the topic-related power system threat intelligence collection method provided in the embodiments of the present application. For example, the general-purpose processor may be a central processing unit (CPU). The method executed when the topic-related power system threat intelligence collection program is invoked can be referenced in the various embodiments of the topic-related power system threat intelligence collection method of the present application and will not be further described here.
[0051] In a fourth aspect, an embodiment of the present application also provides a readable storage medium.
[0052] The readable storage medium of the present application stores a power system threat intelligence collection program based on topic relevance, wherein when the power system threat intelligence collection program based on topic relevance is executed by a processor, the steps of the power system threat intelligence collection method based on topic relevance as described above are implemented.
[0053] Among them, the method implemented when the power system threat intelligence collection program based on topic relevance is executed can refer to the various embodiments of the power system threat intelligence collection method based on topic relevance in this application, and will not be repeated here.
[0054] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0055] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes a number of instructions for enabling a terminal device to execute the methods described in each embodiment of this application.
[0056] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for collecting power system threat intelligence based on topic relevance, characterized in that: The following steps are involved: Extract feature vectors of web page text from three perspectives: semantics, word frequency, and keywords; Construct a topic relevance evaluation model based on stacked BiGRU, set the feature vector of web page text as the model input, and use the constructed topic relevance evaluation model as the main body of the evaluation model; The feature vector of the extracted web page text and the topic relevance evaluation method are integrated into a topic crawler, and the topic crawler is used to crawl the classification results of the web page.
2. The method for collecting power system threat intelligence based on topic relevance according to claim 1, characterized in that: The method of extracting feature vectors of web page text from the perspectives of semantics, word frequency, and keywords includes the following steps: Encode webpage text based on the pre-trained SBERT model and generate SBERT vectors that reflect the semantic features of text topic evaluation; Based on the TF-IDF algorithm, the text is first preprocessed, a bag-of-words model is built, and the term frequency-inverse document frequency is calculated to generate a TF-IDF vector that reflects the importance of the keyword; Divide the SBERT vector and TF-IDF vector by label type and calculate the mean of each vector to obtain the mean vector divided by label type. Calculate the similarity between each label type vector and the corresponding mean vector to obtain the SBERT cosine similarity vector and the TF-IDF cosine similarity vector. Sort by TF-IDF value from high to low, select a preset number of top-ranked keywords, convert them into SBERT word vectors, and calculate the maximum cosine similarity with the preset label words to generate a keyword similarity vector.
3. The method for collecting power system threat intelligence based on topic relevance according to claim 1, characterized in that: The topic relevance evaluation method of constructing a stacked BiGRU-based topic relevance evaluation model, setting the feature vector of the webpage text as the model input, and using the constructed topic relevance evaluation model as the evaluation model body specifically includes the following steps: Constructing a topic relevance evaluation model based on stacked BiGRU, wherein the topic relevance evaluation model includes a feature extraction layer, a fully connected layer, and a softmax output layer, wherein the feature extraction layer is two independent BiGRU models; Input the SBERT vector and TF-IDF vector into the BiGRU model respectively, use two independent BiGRU models to extract features from the SBERT vector and TF-IDF vector respectively, and output the extracted features of the SBERT vector and the TF-IDF vector; Combine the output SBERT vector extraction features, TF-IDF vector extraction features with the SBERT cosine similarity vector, TF-IDF cosine similarity vector and keyword similarity vector to obtain a combined feature vector, which is input into the fully connected layer for training, minimizing the loss value of the cross entropy loss function to obtain the final classification result of the web page.
4. The method for collecting power system threat intelligence based on topic relevance according to claim 3, wherein: The cross entropy loss function is shown below: Where, is the loss value of the cross entropy loss function, is the sample size, Represents a sample In category The true label on , encoded in one-hot form, Represents a sample In category The model prediction probability is given by the Softmax function.
5. The method for collecting power system threat intelligence based on topic relevance according to claim 3, characterized in that: The classification result is calculated as follows: Where, is the weight matrix, is the bias term, is the class prediction probability, is the input of the fully connected layer.
6. The method for collecting power system threat intelligence based on topic relevance according to claim 3, characterized in that: The combined eigenvector satisfies the following relationship: Where, is the input of the fully connected layer, that is, the combined feature vector, They are the output features of SBERT vector and TF-IDF vector after BiGRU processing, is the cosine similarity vector, is the keyword similarity vector.
7. The method for collecting power system threat intelligence based on topic relevance according to claim 1, characterized in that: The framework of the subject crawler includes: Crawling module, used to download HTML documents based on URL seeds; Storage module, used to cache URL seeds and crawling results; Extraction module, used to parse HTML documents and extract body text; A classification module, for integrating the topic relevance evaluation model to perform real-time classification; Tracking module, used to recursively crawl new URLs and update the seed library.
8. The method for collecting power system threat intelligence based on topic relevance according to claim 7, characterized in that: When the tracking module performs recursive crawling, it sets a maximum recursion depth threshold and prioritizes only URLs that do not exceed the threshold, with the ranking based on the topic relevance score of the parent page.
9. A power system threat intelligence collection system based on topic relevance, characterized in that: include: Web page text feature vector acquisition module, used to extract feature vectors of web page text from three perspectives: semantics, word frequency, and keywords; The topic relevance evaluation strategy module is used to build a topic relevance evaluation model based on stacked BiGRU, and set the feature vector of the web page text as the model input, and use the constructed topic relevance evaluation model as the evaluation model body of the topic relevance evaluation method; The web page classification result acquisition module is in communication with the web page text feature vector acquisition module and the topic relevance evaluation strategy module, integrates the extracted feature vector of the web page text and the topic relevance evaluation method into the topic crawler, and uses the topic crawler to crawl the classification results of the web page.
10. The power system threat intelligence collection system based on topic relevance according to claim 9, characterized in that: The webpage text feature vector acquisition module includes: The SBERT vector acquisition module is used to encode web page text based on the pre-trained SBERT model and generate SBERT vectors that reflect the semantic features of text topic evaluation; The TF-IDF vector acquisition module is used to preprocess the text based on the TF-IDF algorithm, build a bag-of-words model, and calculate the term frequency-inverse document frequency to generate a TF-IDF vector that reflects the importance of the keyword. A cosine similarity vector acquisition module is communicatively connected to the SBERT vector acquisition module and the TF-IDF vector acquisition module, and is used to divide the SBERT vector and the TF-IDF vector according to the label type, and perform mean calculation on each of them to obtain the mean vector divided according to the label type, calculate the similarity between each label type vector and the corresponding mean vector, and obtain the SBERT cosine similarity vector and the TF-IDF cosine similarity vector; The keyword similarity vector acquisition module is used to sort the keywords from high to low according to the TF-IDF value, select a preset number of top-ranked keywords, convert them into SBERT word vectors, and calculate the maximum cosine similarity with the preset label words to generate a keyword similarity vector.